Who Gets a Say in Sign Language AI? Mapping the Hard Questions
- Marshall Hurst

- 5 days ago
- 3 min read
AI technologies are moving really fast for many spoken/written languages. They can transcribe speech, translate between languages, and answer questions. Sign languages are far behind, mostly because there isn't enough data to train the tools. That gap is starting to close, and as it does, we teamed up with a large, interdisciplinary group of researchers to stop and ask a set of questions before the field rushes ahead.
The paper is called "The FATE Landscape of Sign Language AI Datasets." FATE stands for Fairness, Accountability, Transparency, and Ethics. The team behind it is a mix of deaf and hearing computer scientists, accessibility researchers, cognitive scientists, and linguists.
Why sign language data is different
Sign language AI runs on video of people signing. And video of you signing is deeply personal. It shows your face, your body, and often where you are. You can't strip your identity out of it the way you can blank a name off a form.
That raises hard questions the paper lays out:
Whose data is it? A collection involves the people who signed, the people who gathered and store the videos, and the people who use them later. Each has rights and responsibilities, and they don't always line up.
Who is in the data, and who isn't? Sign language collections are tiny compared to text collections, and each one includes only a small number of signers. Tools trained on a narrow slice of signers won't work well for everyone else.
Why is trust fragile here? Sign languages are central to deaf culture and identity, and deaf people have faced a long history of discrimination and oppression, often over language itself. That history makes it especially important that technologists handle this work with care. When projects leave deaf communities out, signers often don't trust or adopt what gets built.
The paper's goal: better questions, not quick answers
The goal of the paper isn't to hand out answers, it is to map some of the big questions. We walk through the parameters that go into collecting and handling sign language data, and what each choice can mean for fairness, accountability, transparency, and ethics.
The through-line: deaf people have to be at the heart of development
One point runs through the whole paper. When sign language AI efforts don't involve deaf communities, and don't account for how sign languages actually work as languages, the results tend to fall short and could potentially do harm. This is a moment when AI-enabled sign language tools are no longer out of reach. That makes it the right time to be thoughtful, so these tools are useful to deaf people rather than harm them.
Sign languages are full languages, with their own grammar and vocabulary. Treating them that way, and treating deaf signers as integral to development rather than data sources, isn't a nice extra. It's what makes the technology worth building.
Key takeaways
Sign language AI learns from video of people signing, which is deeply personal and can't be made anonymous the way text can.
Big questions come with that data: who owns it, who is represented in it, and how to earn the trust of a community with a history of oppression.
The paper maps these fairness, accountability, transparency, and ethics questions rather than handing down quick answers.
Its core message: build these tools with deaf communities and respect sign languages as full languages, or they won't serve the people they're meant for.


