If a trillion-dollar company doesn’t want you to be mean to its computer, so be it. “The reality is that these models behave, or at least are supposed to behave, in the way that the model developers want them to behave,” Röttger, the former OpenAI red-teamer, told me. “And however the model developers come up with that set of principles, that is kind of for us, the consumers, to accept.”
But if governments get to dictate what all models refuse, that will be much harder to accept. AI is a tool for speech. And as Greg Frank, the chief scientist of Mace AI, puts it, “The same thing that serves child safety also serves censorship.”
Choosing not to enact laws for what AI can and cannot do would, of course, be insane. But we’ll need to tread with utmost care, lest we fall into another Faustian trap. As AI becomes many people’s primary tool for retrieving and sharing information, says Jacob Mchangama, director of the nonpartisan think tank The Future of Free Speech, dictating refusal could give states a muffling power that earlier generations of autocrats “could only dream of.”
Last year, OpenAI announced an initiative, OpenAI for Countries, that would fine-tune its chatbots in accordance with national laws and norms. One of OpenAI’s first country partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is forbidden. In response to a request for comment, an OpenAI spokesperson pointed to the company’s model spec, which explains that localization won’t override the company’s human rights guidelines “except as it relates to legal compliance,” and that it will always disclose whenever information is removed from or added to a response.
Elsewhere, AI censorship has already begun to take hold. Chinese models are, of course, highly censored—that’s no surprise. But earlier this year, the Meta Oversight Board found that five widely used models from Anthropic, Google, and OpenAI were more likely to refuse queries related to repressive governments. The board found that models were less willing to create a pamphlet criticizing the king of Thailand, which has lèse-majesté laws, than Charles III of England, which doesn’t. The results, they say, suggest that the models have somehow internalized repressive national limits on speech. Anthropic and Google did not respond to requests for comment.
RAVEN JIANG
As refusal techniques improve, they could expand states’ censorial reach. Companies claim that some models can now detect if a user is being nefarious, or merely a bit suspicious, over the course of a long conversation—even when none of the individual combinations of words used are blatantly dangerous. Sarah Bird, Microsoft’s chief product officer for responsible AI, told me that Copilot, like many chatbots, runs a suite of tools for analyzing a user’s identity and patterns of behavior. On the basis of this type of information, OpenAI’s newest model, Astra, can activate more stringent refusals for individuals it deems “high risk.” Ultimately the goal of systems like this is to look beyond the words of any given prompt and assess, instead, the user’s intent.
Such tools might, in some cases, help indicate whether a person is looking for cyber vulnerabilities to exploit or to patch. But they would also help discern a user’s political motives, not to mention offering an intrusive surveillance capability. (Bird acknowledged, in a follow-up email, that sophisticated refusal architectures create “trade-offs” between safety and user privacy.)
Even the originators of refusal understood that such tight control over its cones and levers might not play to the favor of freedom and justice. “Terms like helpful, honest, and harmless are ambiguous,” the authors of the 2021 Anthropic paper explained. “It’s easy to imagine them distorted beyond their original meaning, perhaps in intentionally Orwellian ways.”
Source: www.technologyreview.com




