Who Owns Sign Language Data?
- Jan 1, 2021
- 2 min read
Data, recordings of people using sign language, is a key ingredient in building sign language technologies. But does that data come from? Did anyone agree to share their data? Will the technology be used to benefit deaf communities, or will it inadvertently cause harm? Our team helped write a paper that lays some of these questions out and starts to work through possible answers.
Why sign language data is personal
A signing video is not like a page of text. You cannot remove someone's face from it without losing the language itself, since the face carries grammar and meaning. So the data is, by nature, identifying. It shows a specific person.
Beyond that, sign language sits at the heart of Deaf cultural identity, in communities that have often had decisions made for them rather than with them. Put those two things together and sign language data becomes doubly sensitive.
Questions to ask
Who is included in the data, and who is left out? A dataset drawn from a narrow group will work well for some signers and poorly for others. Who owns the video once it is collected, the person who signed it, the group that gathered it, or the company that stores it? Who can access it, and what are they allowed to do with it? How do people give real consent, and are they fairly compensated? And finally, do the tools built on the data actually help the community, or could they cause harm?
What "good" looks like
The answers point in a consistent direction. Good sign language data is gathered with clear consent, led by deaf people, transparent about how it will be used, and built to include the full range of signers. Good tools support people rather than replace them, and they keep signers in control of their own communication. It is not just a technical problem, it is a question of trust, which has to be earned rather than assumed.
Key takeaways
Sign language AI depends on video of real people signing, which makes the data personal and impossible to anonymize.
Sign language is also central to Deaf cultural identity, so the data is sensitive in more ways than one.
There are important worries, about consent, replacement, ownership, and premature use
Responsible sign language data is consent-forward, community-led, transparent, and representative.
Based on: Bragg, D., Caselli, N., Hochgesang, J. A., Huenerfauth, M., Katz-Hernandez, L., Koller, O., Kushalnagar, R., Vogler, C., & Ladner, R. E. (2021). The FATE landscape of sign language AI datasets: An interdisciplinary perspective. ACM Transactions on Accessible Computing, 14(2).


