A cardiologist records a patient education video and says "watch for signs of hypotension." The automatic caption writes "hypertension." Two letters, opposite condition, opposite patient action. Nobody catches it because caption review is not assigned to anyone. Six months later, an AI answer engine reads that caption file and repeats the error to someone searching for care information.

Metadata used to be an internal problem. It stopped being one when caption files, transcripts, and alt text started leaving the DAM and appearing on public websites where machines read them and speak for your organization.

What actually leaves your DAM

Start with inventory. Ask what lives in your library that ends up somewhere public. The list is longer than most teams expect.

Transcript and caption files ride along with video. Alt text follows images onto your website. Social copy sits beside the asset it describes, so whoever grabs the post grabs the words too. Product descriptions, campaign names, usage terms all travel. Every one of those fields is a place where your organization describes itself, and those words are increasingly read by machines that then describe you to other people.

When AI trains on something wrong, you have not made one mistake. You have started a game of telephone you no longer control.

The 2027 healthcare accessibility deadline

In May 2024, HHS finalized a rule requiring recipients of its federal financial assistance to make web content and mobile apps accessible under WCAG 2.1 Level AA. The rule reaches hospitals, doctors, dentists, clinics, emergency rooms, child welfare agencies, medical and nursing schools, and childcare and social service providers.

HHS extended the original compliance dates by one year in an interim final rule effective May 7, 2026. Recipients with fifteen or more employees must comply by May 11, 2027. Recipients with fewer than fifteen employees have until May 10, 2028.

WCAG 2.1 Level AA requires captions for prerecorded video. Not approximate captions. The standard expects them to carry the dialogue, identify who is speaking, and convey meaningful non-speech audio. Captions that actually match what happened.

Why automatic captions are not enough

Automatic transcription is a starting point, not a finish line. Medical vocabulary is where this shows fastest. A systematic review published in the Journal of the American Medical Informatics Association examined speech recognition in clinical settings from 1990 to 2018. Reported word error rates ranged from 7.4 percent to 38.7 percent. The share of documents containing at least one error ranged from 4.8 percent to 71 percent.

That research covers clinician dictation rather than captioning, but the problem underneath is the same. A pharmacist says "take 0.5 milligrams." The caption drops the decimal and publishes "take 5 milligrams." Ten times the dose, live on your website, with your logo on it. A clinician mentions a patient with a history of ileus, a bowel obstruction. The caption confidently introduces a man named Elias.

None of these are exotic. They are one letter, one decimal point, one homophone. Then an answer engine reads the file and passes it along.

What AI search reads today and what it will read tomorrow

Current AI search largely reads pixels and audio. It is not reading your controlled vocabulary, your product nicknames, or the fifteen years of institutional shorthand your team uses to find things. That is changing. Platforms are adding business-term features so the system can build a working vocabulary. MCP connectors rolling out across DAM platforms will link the AI agents organizations already built to the assets in the DAM.

The gap between what AI reads today and what it will read tomorrow is the argument for getting your metadata in order now rather than when the capability arrives. Stacks covers this shift in detail at their blog on metadata and the 2027 deadline, noting that organizations should start paying attention to metadata now so their data is ready when platform AI learns their terminology.

Build workflow and governance before you buy more AI

Before anyone buys another AI feature, two unglamorous things need to exist. The first is workflow. Someone has to be accountable for the alt text getting written, the copy getting attached, the transcript getting tied to the right video. AI does not remove that need. It relocates it. A human still has to review and validate that whatever the alt text or transcription the AI produces is accurate.

The second is governance, and it matters more the more sensitive your content is. Working with patients and HIPAA creates opportunity for serious issues when metadata is not reviewed. If AI is going to read your library and speak on your behalf, deciding who can publish and what gets reviewed is not optional.

Inventory what actually leaves your DAM. Decide where transcripts and caption files live and how they stay tied to the right video, language, and version. Build the review step into the workflow instead of hoping someone catches it, and write down who owns it. Then get your vocabulary in order, because when platform AI starts learning your terminology, it will learn from whatever you have.

Key takeaways

  • Metadata fields like captions, transcripts, and alt text now leave your DAM and appear on public websites where AI answer engines read and repeat them.
  • Healthcare organizations receiving HHS federal financial assistance with fifteen or more employees must meet WCAG 2.1 Level AA by May 11, 2027, which requires accurate captions for prerecorded video.
  • Automatic transcription produces word error rates ranging from 7.4 percent to 38.7 percent in medical contexts, making human review mandatory.
  • Current AI search reads pixels and audio, not your controlled vocabulary, but that is changing as platforms add business-term features and MCP connectors.
  • Workflow and governance must exist before AI features are added, with named owners responsible for caption review and decisions about who can publish.

Standards and sources