Article50.io
Comparison · Art. 50

Article50.io vs Manual Compliance Checklists: What Automated Scanning Catches That Forms Miss

· 9 min read

A self-assessment form records what the person filling it in believes is on your website. An automated scan loads the live page in a browser and records what is actually there, so it catches chat widgets and AI-generation markers that nobody told the form-filler about. What a scan can't do is the reverse: it can't know your role under the AI Act, see anything behind a login, inspect file metadata, or make legal judgements. You need both.

Two different kinds of evidence

A manual checklist, such as a self-assessment questionnaire, an internal audit spreadsheet or our own Article 50 compliance checklist, asks questions. "Do you use an AI chatbot?" "Do you publish AI-generated images?" The answers are only as good as the knowledge of whoever gives them, and they describe the site on the day the form was filled in.

An Article50.io scan works differently. It loads one public URL in a headless browser, waits for the page and its scripts to load, and then reads three things: the rendered HTML, the script sources the page pulled in, and the visible text, including text shown inside a chat widget as it first loads. It runs two checks against that evidence:

  • Undisclosed AI chat widget (flagged under Art. 50(1)). The page loads a known chat widget (Intercom, Drift, Zendesk or Crisp, detected by script source or JavaScript global) or custom chat markup (for example, class names such as chatbot or ai-assistant), and neither the visible page text nor the chat widget's own text contains an AI-interaction disclosure.
  • Undisclosed AI-generated content (flagged under Art. 50(4)). The page carries a concrete machine-readable marker, such as a generator meta tag naming a tool like Midjourney, DALL-E, Stable Diffusion, Runway, Synthesia, Sora or Firefly, a data-ai-generated="true" attribute or an ai-generated class, and the visible text contains no AI-content disclosure.

That's the full list of checks. It's deliberately narrow: the scanner only reports what it can point to in the markup, and it never infers from how an image looks that the image is AI-generated.

Key point: A manual checklist records what you believe your site does. A scan records what your public page actually renders. Neither one on its own tells you whether you comply with Article 50.

What a live scan catches that a form structurally can't

Widgets nobody mentioned

The person completing a compliance form is often in legal, product or engineering. Chat widgets are often added by marketing or support, sometimes through a tag manager rather than the codebase. If the form-filler doesn't know an Intercom or Crisp script is loading on the pricing page, the form will say "no chatbot" and be sincerely wrong. The scan doesn't depend on anyone's knowledge. If the script or global is on the page it loads, the scan sees it.

Disclosure on the wrong page

A common pattern is an AI disclosure in the terms of service or privacy policy, but nothing on the page where the widget actually appears. A checklist question like "Do you disclose AI use?" gets a truthful "yes". The scan reads only the visible text of the page it loaded, so a disclosure that lives on a different page doesn't count. That's closer to what the law asks for. Article 50(5) requires the information to be given in a clear and distinguishable manner, at the latest at the time of the first interaction or exposure, and to meet the applicable accessibility requirements.

Markers your tools added for you

Some generative tools and publishing pipelines write machine-readable signals into the page, such as a generator meta tag or a data-ai-generated attribute. The people who publish the page may never look at its source. If a marker says "AI-generated" and the visible page says nothing, the scan flags the gap.

Drift since the form was filled in

A form is a snapshot. Sites change: a new landing page, a new widget, a vendor's product update. A scan shows the live page as it is on the day you run it. It is also a snapshot, just a newer one. Article50.io doesn't monitor your site continuously. If you buy the full report, it includes one free recheck within 30 days of purchase so you can confirm your fixes, and that's all. For anything after that, run a new scan.

What a manual review catches that a scan can't

This part matters just as much, and it's where a careful human (ideally with legal input) is irreplaceable.

  • Provider or deployer? Art. 50(1) and 50(2) apply to providers, and Art. 50(4) applies to deployers. Which role you hold for a given system depends on facts the scan can't see, such as who developed the system and under whose name it is placed on the market. See who Article 50 applies to.
  • Is it "obvious from context"? Art. 50(1) doesn't require the interaction disclosure where it's obvious to a reasonably well-informed person that they are talking to an AI. The scan doesn't make that judgement. It treats a known widget as not obvious and asks for a disclosure.
  • Are AI features actually switched on? The scan treats any detected Intercom, Drift, Zendesk or Crisp widget as AI-capable. It can't see your vendor settings, so it can't tell whether the widget is staffed by humans or answered by AI. Only you can check that.
  • Does Art. 50(4) actually apply? The scan flags missing disclosures next to AI-generation markers under Art. 50(4). Whether the content really is a deep fake, or AI-generated text published to inform the public on matters of public interest, is a judgement for a person. So is whether the human review or editorial control exception applies, which also requires someone to hold editorial responsibility. See does Article 50 apply to AI-generated text? and deepfake disclosure.
  • Emotion recognition or biometric categorisation. Under Art. 50(3), deployers of these systems must inform the people exposed to them and process the personal data in line with the GDPR and related EU data protection law, with an exception for certain law-enforcement uses. The scan does not detect these systems.
  • Machine-readable marking inside files. Art. 50(2) requires providers of generative systems to mark outputs in a machine-readable format. That marking often lives inside the image, audio or video file itself (for example, C2PA or IPTC metadata). The scan reads page markup, not file metadata or watermarks. See labelling AI-generated content.
  • Anything behind a login. The scan doesn't sign in, so it can't see in-product assistants, customer portals or authenticated dashboards.
  • Emails, apps and APIs. AI-written support emails, mobile apps, voice agents and API outputs are all outside what a scan of a public URL can see.
  • Contracts. Whether your AI vendor's terms make the provider responsible for marking outputs, or leave the disclosure to you, is a question for your contracts.

There are also practical scan limits worth knowing. The scan checks one URL per run, not your whole site. It reads the page as loaded and doesn't click anything. A widget that only loads after a visitor accepts cookies may not appear. It reads text inside a chat widget as the widget first loads, but not text that only appears once the chat is opened or a conversation starts, such as a label on the bot's replies. The disclosure check also looks for common phrasings, so unusual wording could be missed. Treat a finding as a precise prompt to go and look, not a verdict.

Side-by-side comparison

Check Live scan (Article50.io) Manual checklist
Known chat widget actually loading on a public page Yes, from script sources and globals Only if the form-filler knows about it
Visible AI disclosure on that same page Yes, if worded in a common way Depends on the answer given
Page-level AI-generation markers (meta tag, data-ai-generated) Yes Rarely, because few people read page source
Change since the last review Shows the live page on the day you scan No, it's a snapshot of what someone believed
Provider vs deployer role No Yes
"Obvious from context" judgement No Yes
Whether AI features are enabled in a vendor widget No Yes, via vendor settings
Deep fake or public-interest text scope, editorial exception No Yes
Emotion recognition or biometric categorisation (Art. 50(3)) No Yes
C2PA or IPTC metadata inside media files (Art. 50(2)) No Yes, with file tooling
Logged-in product, emails, apps, APIs No Yes
Vendor contracts No Yes

Using them together

A sensible order: work through the manual checklist to settle your role, your systems and the legal judgement calls. Then run a free scan on your key public pages to check that what you decided is actually live on them. The free scan shows the single most severe finding, along with the Art. 99(4) penalty tier: up to €15,000,000 or 3% of total worldwide annual turnover, whichever is higher, or the lower of the two for SMEs, start-ups and small mid-cap enterprises. The full report costs €499, paid once. It lists every finding on the page scanned, with a fix snippet for each, and arrives by email. For disclosure wording you can adapt yourself, see the free AI chatbot disclaimer template and AI content label template.

Article 50 applies from 2 August 2026. Primary source: Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744. The Commission's AI Act Service Desk also has an Article 50 page.

Frequently asked questions

Does a clean scan mean I'm compliant with Article 50?

No. A clean Article50.io scan means that on the one public page it loaded, it found no known chat widget without a visible AI disclosure and no page-level AI-generation marker without a visible AI-content disclosure. It doesn't assess your provider or deployer role, logged-in areas, emails, apps, file metadata, emotion recognition or biometric categorisation, or whether exceptions apply. Compliance needs a manual review as well.

Can a scan give a false positive?

Yes. The scan treats any detected Intercom, Drift, Zendesk or Crisp widget as AI-capable, even if a human answers every chat. It reads the page and the chat widget as they first load, so a disclosure that only appears once the chat is opened or a conversation starts, or that uses unusual wording, may be missed. Treat each finding as a specific thing to check, then decide.

Why not just use a manual checklist?

A manual checklist is essential for legal judgements such as your role, the "obvious from context" test and the editorial-control exception. But it relies on what the person completing it knows. Widgets added through a tag manager, disclosures that sit on the wrong page and markers written by publishing tools are easy to miss on a form and easy to see on the live page.

Does Article50.io monitor my site over time?

No. Each scan checks one page at one point in time. The €499 full report includes one free recheck within 30 days of purchase to confirm your fixes. After that, run a new scan whenever the page changes.

Does the scan check C2PA or other metadata inside my images?

No. It reads the page's HTML, script sources and visible text. It doesn't open media files or read embedded metadata or watermarks such as C2PA or IPTC, so machine-readable marking under Art. 50(2) needs separate checking.

This article is general information, not legal advice.

Check your site automatically

Article50.io is an automated Article 50 transparency assessment platform that scans websites for potential EU AI Act transparency obligations and provides remediation guidance, implementation instructions, and compliance-ready disclosure language.

The free scan shows your single most severe finding in about 30 seconds — no signup, public pages only.

More from the blog

Automated technical guidance, not legal advice. Citations refer to Regulation (EU) 2024/1689.