Academic research is quietly shaping the next wave of trustworthy, useful, and equitable AI, right here at U-M.
In just a few years, large language models (LLMs) such as ChatGPT, Gemini, and Claude have gone from research curiosity to household names. They have the power to chat, translate, and generate content, and are becoming increasingly embedded in domains spanning science, medicine, business, and education. Tech companies are pouring billions of dollars into expanding the size and capability of LLMs, rapidly transforming these models from niche tools into essential infrastructure that is reshaping how we work, learn, and connect.

Despite this revolutionary growth, today’s LLMs remain, in many respects, virtual black boxes. Beneath the surface of rapid innovation is a parallel race: U-M researchers, such as the affiliate members of the Michigan Institute for Data and AI in Society (MIDAS) featured in this story, are working to expose these models’ inner workings and hidden pitfalls. Through their explorations, these researchers are getting to the heart of what LLMs actually “know” and, in the process, ensuring safer, smarter, and more responsible AI tools for real
world use.
Language, vision, and what makes us human
Understanding LLMs means understanding the roots of language development itself. “Language didn’t emerge for humans until about 70,000 years ago, long after we’d already developed vision and physical navigation of the world,” said Joyce Chai, professor of Computer Science of Engineering.
This linguistic origin story continues to shape how people, and the machines we create, communicate and learn. “We learn to speak after we see, move, and act first,” Chai noted. “Language builds on our embodied, shared experiences in the world.”
Her research aims to close the gap between words and the world, building AIs that aren’t just text-only chatbots but agents that can both “see” and “say”, linking images, video, and physical context to language for a more grounded, robust understanding.
For instance, in a recent COLM paper, Chai’s team developed a novel framework to bootstrap the development of visual dialogue agents that can guide humans through complex tasks in physical environments. Other projects explore how robots learn tasks through trial, feedback, and spatial reasoning.
Across these projects, Chai’s team is devising new architectures and training paradigms to bring language, perception, and experience closer together. From examining how well AI reasons based on embodied context to exploring whether today’s models can represent spatial and physical relationships, her work is bringing us closer to AI tools that communicate, collaborate, and learn in ways that mirror human experience.
Chai points out that while LLMs are being widely deployed, they are trained based on people’s perspectives of the world through text; they are not grounded to the true physical world. Next-generation multimodal models – ones that can watch and act as well as chat – will require fresh architectures and a deeper understanding of both human and machine learning.
“Human learning is incredibly data efficient, incorporating feedback, grounding, and social interaction in a rapid process,” emphasized Chai. “If we could give models some of those advantages, the impact could be huge, across everything from healthcare to education.”
LLMs’ trust problem
Understanding is only half the story. As LLMs become more widespread, new dangers surface: misinformation, overconfidence, and hallucinations have become some of the field’s most urgent challenges. Lu Wang, associate professor of Computer Science and Engineering, is tackling these risks head on, designing techniques that make models smarter and more trustworthy.
“A lot of people have heard stories about AI chatbots making things up: false medical advice, legal citations that don’t exist, or flat-out dangerous recommendations,” said Wang. “This isn’t just annoying. These errors can have very real consequences in domains including healthcare, justice, and education.”
The root problem, Wang explains, is that LLMs absorb nearly the whole internet, which itself is riddled with errors, and learn to reproduce patterns rather than truths, sometimes with surprising confidence. Even newer, more powerful models aren’t necessarily better at sticking to the facts: “In some cases, hallucinations actually get worse with stronger models, because newer training methods are designed to boost general capabilities, but not factual grounding,” said Wang.
Wang’s group is addressing these shortcomings by pushing for models that actively retrieve and cite evidence from reliable sources. Such methods improve accuracy and build a transparent, attributable “factual trail” for users. She and her lab are also pioneering the approach of confidence calibration, encouraging LLMs to say “I’m not certain” or provide confidence scores when unsure, making it easier for users to judge and manage risk.

Figure from
Zhang et al.
(2025) – BASIS
presents a three
stage pipeline for
visual assistant
development
Beyond technical fixes, Wang believes that academic research plays a vital role in holding AI technology accountable and ultimately ensuring it serves society’s best interests: “Industry moves fast, but universities have the objectivity and independence to probe, critique, and improve these tools, identifying weaknesses and designing the guardrails that are sorely needed.”
Models for social insight
Trust and transparency matter even more when language touches on personal, cultural, or ethical dimensions. David Jurgens, associate professor of Computer Science and Engineering and Information, leads research exploring the complex social and cultural factors that shape language, and how LLMs can (or fail to) reflect them.
His group develops models and datasets that capture how people from different backgrounds and cultures communicate, and how language reflects underlying values and priorities. For instance, a recent project from his group, which won Best Resource Paper Award at ACL 2025, tackled the challenge of moral and cultural alignment in LLMs. Their findings revealed that models often struggle with classic moral dilemmas —such as whether to tell a difficult truth or protect someone’s feelings—especially when cultural context shapes possible answers.
“We found that today’s models are nowhere near capturing the rich diversity in human responses to these questions,” he said. “They’re not random, but they miss a lot of what drives moral decision-making.” This gap highlights both the remarkable potential and the existing shortcomings of LLMs in truly human-centered understanding.
Jurgens’ team also systematically tests LLM “psychology,” discovering that demographic differences drive disagreement in annotating subjective language, and highlighting ongoing gaps in model empathy, authenticity, and connection. As Jurgens put it, “The aim isn’t to replace human interaction but to meaningfully support it, so people can have deeper and more effective conversations.”
In another ongoing project, supported by a MIDAS pilot grant, Jurgens and his collaborators are working to improve patient-physician interactions by using LLMs to better understand how patients experience language used during clinical encounters. This work investigates how Black and White patients perceive physician comments and is part of a broader research effort supported by MIDAS.
Jurgens credits the collaborative, interdisciplinary culture at U-M for making this kind of research possible. “We’re constantly reaching across computing, psychology, political science, and more to work on questions that really matter for society,” he noted. “It’s a huge strength of Michigan.”
AI across culture
As LLMs shape communication and decision-making around the World, Janice M. Jenkins Collegiate Professor of Computer Science and Engineering and director of the Michigan AI Lab Rada Mihalcea works to confront the critical challenge of making these technologies truly inclusive. Much of her current research focuses on cross-cultural understanding, developing methods to identify people’s values, worldviews, and behaviors from language, and ensuring that AI systems serve not just English-speaking or Western users, but a diverse, global audience.
“If we want AI to be truly transformative, we need to address not only what models can do, but who they work for, how they handle diverse backgrounds, and whether they respect the values of the communities they support,” Mihalcea emphasized.

For example, her 2024 EMNLP paper analyzed how LLMs encode age-related values, demonstrating the importance of demographic-aware modeling for truly fair systems. In a 2025 NAACL study, Mihalcea’s team showed how prompting models with geographic and socioeconomic context can improve performance on data from low-income or marginalized communities—favoring perspectives too often overlooked by standard AI benchmarks.
Her work in multilingual sentiment and emotion analysis offers vital benchmarks to counteract bias and avoid narrow stereotypes, advancing language models that can understand and respond to emotional cues across different regions and languages.
From building cross-lingual semantic tools to investigating multimodal sensing of human behavior, Mihalcea’s work is shaping the next generation of AI to be not just smarter but also more empathetic and equitable. “AI technologies shouldn’t just be powerful,” she noted. “They should adapt, empathize, and reflect the richness of human diversity.”
Connecting the dots: Michigan’s unique contribution
Michigan’s AI researchers aren’t just tackling benchmarks or chasing the latest leaderboard. Their work is deeply collaborative and interdisciplinary, spanning computer science, psychology, information science, medicine, and robotics. At U-M, advancing the science of LLMs means thinking as much about meaning, fairness, and impact as about performance or scale.
By focusing on grounded, transparent, and user-centered models, researchers are charting a path beyond the current AI hype. They’re showing how LLMs can be made more reliable, more inclusive, and ultimately more useful for the complex world outside the lab.