Privacy notice
This site studies how automated clients discover and retrieve a controlled collection of public pages. It does not attempt to identify human visitors.
Last updated: 30 July 2026
Controller and contact
Controller: Benjamin Blatti
Privacy contact: privacy@mail.aiwiki.ch
Location: Dählenweg 5, 3603 Thun, Switzerland
Data categories and source
Request metadata can be personal data or pseudonymous personal data even when a request appears to come from a crawler. The service therefore treats user-agent strings, referrers, timestamps and short-lived session identifiers as potentially personal data. The data comes directly from the HTTP request and Cloudflare's delivery metadata.
For each relevant HTTP request, the service records the time, requested path without its query string or fragment, method, response status, resource type, protocol metadata, server-assigned experiment cohort and response profile, and a crawler classification. Declared crawler user-agents may be retained in a bounded form because they are a research signal, but remain unverified claims. Browser user-agents are reduced to a coarse browser family and major version.
For cache-validation experiments, the service records only whether a supported conditional request header was present, not its raw value. Experimental depth and cohort labels describe this site's fixed test design rather than the visitor.
Exact same-origin referrers are stored as origin and path without query strings or fragments. All other referrers are reduced to their origin. The referrer header is optional and client-controlled, so it is treated as a source claim rather than proof of a click or navigation.
A small paired Arena cohort publishes fixed edge labels in four link addresses. A label is identical for every visitor, accepted only for its manifest-bound target path, and stored as that fixed label while the request query is discarded. It identifies a site-defined link, not a person, browser, device, or durable session, and it does not prove causal traversal.
Session grouping and IP addresses
Cloudflare necessarily receives the source IP to deliver the HTTP request. The observatory itself does not write raw IP addresses or reusable IP hashes to its event store. It uses the source IP transiently inside a keyed HMAC together with a minimised user-agent and a 30-minute time bucket to create an approximate, pseudonymous session identifier. The source IP is discarded before the event is written; the identifier changes between time buckets and is not used to identify a person.
Purposes and legal basis
The data is used only to measure crawler discovery and retrieval behaviour, evaluate this controlled research corpus, protect and operate the service, troubleshoot faults and produce data-minimised aggregate statistics. It is not used for advertising, identity enrichment, sale of visitor profiles or decisions producing legal or similarly significant effects.
Where the EU GDPR applies, the intended legal basis is the controller's legitimate interest in operating, securing and studying this public research service, subject to necessity and balancing. Swiss data-protection principles, including proportionality, purpose limitation, transparency and security, are applied.
Some public Arena pages intentionally return robots directives or synthetic 401, 403 and 404 responses. Public capability puzzles, unlinked canaries and inert DOM-context simulations reveal harmless measurement markers only. They do not protect data, accept credentials, create privileged sessions or implement an intentionally exploitable vulnerability. The simulations accept only fixed, server-defined test classes, isolate rendered demonstrations, block script and network execution, and never execute, store or reflect arbitrary request input.
Recipients and processing abroad
The controller and authorised local operators can access the minimised research events. Raw event exports and session identifiers are not published. Cloudflare provides edge delivery, Worker execution and object storage as a processor and can engage group companies, data-centre operators and other subprocessors for those services. For customer-initiated runs there is one further recipient: the customer who created the run receives the redacted report for that run, as described below.
Cloudflare and its subprocessors may process network and event data in Switzerland, the EEA, the United States and other countries identified in Cloudflare's current subprocessor list. For restricted transfers, Cloudflare's Data Processing Addendum incorporates the applicable Standard Contractual Clauses, including the Swiss adaptations, and supplementary safeguards. Current details are available in Cloudflare's Data Processing Addendum and subprocessor list. Cloudflare's independent network-security logs, where applicable, follow Cloudflare's own contractual retention and privacy terms.
Customer-initiated crawl runs
Alongside the research corpus, this host serves crawl targets that a customer of the AIWiki Crawler Lab creates for their own test. A run issues separate single-use link addresses, one per crawler the customer wants to observe, and each address resolves to a page of this same bounded laboratory corpus. The purpose is to let the customer see which crawlers actually retrieve a page and along which path, without placing any measurement code on the customer's own website.
Retrieving such a link is recorded in a separate object store that is not part of the research corpus and is not used for research reporting. The recorded fields are the time, the requested path without query string or fragment, the method and response status, the resource type, the discovery cohort, a coarse crawler classification with an authenticity assessment, and the pseudonymous run, actor and link identifiers. No IP address and no full user-agent string is written to that store.
Recipient: the customer who created the run receives a redacted report containing exactly those fields. Capability tokens, referrers, full user-agent strings and session identifiers are not part of that report and are never disclosed to the customer.
Retention: events from customer-initiated runs are deleted after 24 hours. A scheduled job performs the deletion and an independent object-storage lifecycle rule provides a backstop. This period is deliberately much shorter than the research retention stated below and is not extended by any subscription.
The controller for these runs is the operator named above. Anyone who sends such a link to a third-party service causes that provider to retrieve the public laboratory pages under that provider's own terms and privacy notice.
Retention and deletion
Individual request events from the research corpus have a configured retention period of 30 days. A scheduled application job deletes expired events, and an object-storage lifecycle rule provides an independent deletion backstop. Reports may retain aggregated statistics that no longer contain request IDs, session identifiers, referrers, or full user-agents. Events from customer-initiated runs are held separately and deleted after 24 hours, as described above.
Cookies
The public field guide does not set analytics or advertising cookies. If an optional research journey explicitly offers a temporary session, it remains disabled until the participant activates it.
Your rights
Depending on the applicable law, individuals may request information about processing, access, correction or deletion of their data, restriction of processing, or object to processing by contacting the controller. A request may require enough technical information, such as an approximate request time and path, to locate the relevant event without collecting unnecessary additional identity data. Individuals may also contact the competent data-protection supervisory authority.
Further explanations of Swiss transparency and access rights are available from the Federal Data Protection and Information Commissioner (FDPIC).
Research limitations
A user-agent is a claim, not proof of provider identity. A fetch does not prove indexing, training use, citation, semantic reading, or that a human saw the page.