Responsible AI guidelines
What we ask of anyone who publishes data or a model here, and of anyone who uses one.
Before you publish data
- Say where it came from and how it was gathered. Provenance is not a formality: it is what lets somebody else judge whether your data fits their problem.
- Remove or mask personal details unless you have a lawful basis to publish them and, where consent is the basis, that consent.
- Say what is missing. Which districts, which dialects, which kinds of people are not in this data — every dataset is partial, and the ones that admit it are the useful ones.
- Say how it was labelled, by whom, and what they were told to do.
Before you publish a model
- Say what it is for, and what it is not for.
- Say what it was trained on, and under what licence.
- Report how it does — and where it does worst. A model that works on formal Bangla and fails on a dialect must say so.
- Say what it costs to run, so somebody with one machine knows whether it is for them.
Before you use one
- A model’s answer is a guess, however confident it sounds. Where a wrong answer would hurt somebody — health, money, justice, safety — a person decides, not the model.
- Test it on your own people’s data before you rely on it. Performance reported here was measured somewhere else.
- Tell the people affected that a machine is involved, and give them a way to reach a human.
Bangla in particular
- Bangla is written by people whose dialect, spelling and script habits differ. Treat variation as the language, not as noise to be cleaned away.
- Bengalis live in many countries. Data from one of them does not speak for the rest.
- Publish evaluation in Bangla too, so the people a model is about can read how well it works on them.
Where this is enforced
- Review asks for these things, and a curator may return work that lacks them.
- A dataset that carries personal data without a basis is taken down.
- Work that is used to target or harm a group is removed, and the account that published it is suspended.