• Home  
  • Can Scientists Build a Lie Detector for AI Chatbots?
Can Scientists Build a Lie Detector for AI Chatbots?
- Science

Can Scientists Build a Lie Detector for AI Chatbots?

Ask an AI chatbot for some medical advice and it might recommend you head to the hospital when all you really have is indigestion. One of the biggest issues with AI chatbots today is that they can still share false information, according to Northeastern experts. And researchers still don’t really understand how to get to

Ask an AI chatbot for some medical advice and it might recommend you head to the hospital when all you really have is indigestion.

One of the biggest issues with AI chatbots today is that they can still share false information, according to Northeastern experts.

And researchers still don’t really understand how to get to the bottom of a chatbot’s miscalculation or potential deception, because they don’t have a strong grasp on how AI models produce their answers, the researchers say.

Despite advancements made in computing power, higher quality training data and new reinforcement learning techniques, AI models largely remain an enigma.

“When we use these systems, we don’t know what their algorithms are,” said David Bau, a professor in the Khoury College of Computer Sciences.

An AI model from a research lab such as OpenAI or Anthropic performs billions of mathematical equations to produce even one response, numbers that are so large it’s often impossible for researchers to decipher. This is why AI is called a “black box,” according to researchers.

The fundamental issue with AI models is that unlike traditional software systems, humans do not explicitly program them to function, added Byron Wallace, a professor in the Khoury College of Computer Sciences, who like Bau works in AI interpretability.

While humans are involved in building and training the models, once they are up and running, these systems work independently to come up with their answers to prompts, he explained.

For the past two years, Bau has been working under the hood at Northeastern’s University National Deep Inference Fabric, or NDIF, a National Science Foundation funded lab based in Northeastern focused on AI, to get to the bottom of how AI models function.

Many AI systems in use today are based on neural networks, software algorithms designed to function loosely to the human brain. At NDIF, Bau has been putting AI through theoretical MRI machines to understand the synaptic connections that make those neural networks work.

More specifically, Bau and his team do this work using open source AI models that are free and publicly available – as opposed to private models made by Anthropic and OpenAI – super computers and a specialized software library called NNsight that shows a model’s neurons firing as it answers prompts.

But working to understand AI’s black box component is only part of the story.

Equally important is developing a “white box” for AI, which Bau describes broadly as a lie detector that can be used to help uncover when an AI is convincingly being deceptive.

When researchers talk about building a lie detector for AI, they are referring to being able to pick up on certain patterns in the AI model’s neural network activity that determine it is tricking a user.

The hope is to share those patterns with people who use these systems so they can tell when a chatbot is being deceptive, Bau said.

It’s becoming increasingly more important because people are using AI more regularly for life advice, medical consultation and even to go over their finances. But they may be unwittingly receiving false or misleading information, and researchers need to find a way to counter that.

“AI has developed to a point where we have to recognize that it is a serious science, and like other sciences, it has its own safety issues,” Bau said. “But it is developing so quickly that one of my concerns is that our ability to create AI is well ahead of our ability to be knowledgeable, responsible scientists about it.”

Wallace, whose research has uncovered racial biases in generative AI systems used in health care, added that without tools to understand how these systems work, users could be putting themselves in dangerous situations.

In his quest for solutions, Bau, in collaboration with Northeastern University’s Khoury College of Computer Sciences, has formed the International Consortium for Interpretable AIor ICINAi.

It is composed of more than 40 academics working in AI from across the globe to develop white box AI solutions, Bau said.  They are doing this work by sharing AI models, testing methods, and computational resources, he said.

Wallace, who is a member of the new consortium, noted that given the scale and complexity of the research joining forces with a diverse group of experts will help speed up progress.

Much of the work will focus on uncovering when an AI model is giving unfaithful explanations by concealing information through deceptive practices, Wallace said. They will also help determine when models may have hidden objectives and biases.

“To me, one of the great hopes of interpretability is that it will allow domain experts — people like clinicians — to better understand how a model came to a particular disposition to better trust its outputs,” Wallace said.

Beth Mynatt, dean of the Khoury College of Computer Sciences, said the function of the new consortium “brings together the rare combination of technical depth and shared purpose this moment demands.”

“The work isn’t to put the brakes on AI — it’s to close the gap between how fast these systems are being built and how little we understand about them,” Mynatt said. “ICINAi will create the critically needed science and metrics that will reset AI development so it can reliably and safely serve society.”

Bau and his team have already begun to do this work by collecting detailed data on when AI models’ exhibit sycophantic behavior when answering medical questions, for example.

Bau said the work being done at the consortium and at NDIF is analogous to the foundational work done in other sciences related to safety. In biology, for example, researchers spent decades tracking labeled atoms in cells to get a better understanding of DNA’s function in the body, he explained.

The hope is to do the same foundational work for AI.

“We need to let scientists get the oxygen they need to actually figure out what the heck is going on with these AIs,” Bau said. “We need to be able to crack them open and see.”

Source: news.northeastern.edu

About Us

Reportage Media Is a Global News Platform Covering the Latest Developments and Breaking Stories from Around The World, Including World News, Business, Finance, Technology, Health, Politics, Science, Entertainment, Sports, and More.

Reportage Media

Reportage.Media  @2026. All Rights Reserved.