Human-State Intelligence for Artificial Intelligence

Scientific Foundations, Evidence, Limitations and Research Directions for Emotion-Aware, Context-Aware and Relational AI Systems

7/20/202627 min read

A white robot is standing in front of a black background
A white robot is standing in front of a black background

Scientific Foundations, Evidence, Limitations and Research Directions for Emotion-Aware, Context-Aware and Relational AI Systems

Independent Research Report | August 2026 | Trang Phan
Research focus: affective computing, multimodal human-state estimation, social cognition, adaptive communication, AI companionship and human–AI interaction
Method: evidence synthesis emphasizing systematic reviews, meta-analyses, controlled studies, regulatory sources and recent peer-reviewed research

Executive Summary

Artificial intelligence is moving beyond systems that answer questions toward systems that continuously interact with people, retain contextual information, adapt communication style, provide emotional support and increasingly participate in decisions involving work, education, health, relationships and daily life. This transition is creating demand for a capability that can be described broadly as human-state intelligence: the capacity of an AI system to use observable information about a person and their situation to adapt its response while remaining calibrated about what it actually knows. The scientific challenge is materially harder than sentiment analysis or emotion classification. Human state is multidimensional, partially observable, temporally dynamic, culturally conditioned and causally ambiguous. Words, facial expressions, voice, physiology, behavior and interaction history can all contain relevant information, but none provides direct access to another person's internal experience.

The research base has expanded rapidly. A 2024 systematic review examined 410 publications on affective computing using text, audio and visual information. A separate 2024 review synthesized 179 multimodal emotion-recognition studies published between 2017 and 2023, covering subjective reports, text, electrodermal activity, cardiovascular signals, respiration, EEG, neuroimaging, facial behavior, vocal behavior and whole-body behavior. A 2024 systematic review of emotion recognition from cardiovascular signals examined the growing use of ECG and photoplethysmography, while another review of autism-specific emotion-recognition technologies screened 346 publications and included 65 studies. The field is therefore well beyond proof-of-concept scale. At the same time, these reviews consistently identify unresolved problems involving dataset validity, cross-person generalization, multimodal integration, cultural variability, real-world measurement, privacy and the conceptual definition of emotion itself. (ScienceDirect)

The first conclusion of this report is consequently that human-state modelling is scientifically plausible, but categorical psychological certainty is not. Modern systems can extract useful information from behavioral and physiological signals. What they generally cannot establish from those signals alone is a unique explanation for why the signals occurred. Rapid messaging can be associated with agitation, excitement, time pressure or conversational style. Reduced verbal output can accompany sadness, fatigue, distraction, cultural reserve, hostility or simply a desire for brevity. Elevated heart rate may reflect anxiety, physical activity, medication, fever, caffeine or many other conditions. The central computational problem is therefore latent-state estimation under uncertainty rather than deterministic emotion reading.

This distinction is reinforced by regulation. The European Union's AI Act explicitly states that the scientific basis for systems intended to identify or infer emotion raises serious concerns because emotional expression varies across cultures, situations and individuals; it specifically identifies limited reliability, limited specificity and limited generalizability as important shortcomings. Article 5 prohibits the use of AI systems to infer emotions in workplaces and educational institutions except for medical or safety reasons. These prohibited-practice provisions have applied since 2 February 2025. The significance is not merely regulatory: a major jurisdiction has formally recognized that inference about internal human states can create sufficient scientific and rights-related risk to justify use restrictions even when the underlying technology is technically capable of producing predictions. (EUR-Lex)

The second conclusion is that emotion should be modelled as one component of a broader human state rather than as an isolated label. Contemporary emotion science connects affect with appraisal, physiological regulation, attention, memory, motivation, perception and social context. Interoception—the sensing and interpretation of internal bodily signals—has become an increasingly important research area. A 2025 Annual Review of Psychology synthesis reviewed a literature extending across 163 cited sources on interoceptive mechanisms and emotional processing and emphasized both strong links between body-state processing and emotion and persistent difficulties in measurement and interpretation. A useful AI representation of a person therefore needs to distinguish observable signals, inferred affect, cognitive capacity, contextual demands and uncertainty rather than compressing all of them into a single sentiment score. (Annual Reviews)

Third, multimodality improves the available evidence but does not eliminate ambiguity. The 179-study multimodal review describes four broad information families: subjective experience, peripheral physiology, central physiology and behavioral expression. Combining text, voice, facial behavior and physiological information can yield richer representations than a single modality, but the resulting model still depends on datasets, elicitation procedures, labels and population assumptions. The 410-paper trimodal review specifically identifies detecting concealed or ambiguous emotion as a continuing challenge and notes that physiological signals can offer useful information while being substantially harder to acquire reliably in real-world settings. The appropriate interpretation of multimodal AI is therefore “more evidence,” not “direct access to emotion.” (ScienceDirect)

Fourth, social and relational inference is meaningful at the population level but hazardous when converted into fixed individual labels. A meta-analysis of 224 studies, 245 samples and 79,722 participants found moderate associations between adult attachment anxiety and negative mental-health outcomes, including depression, anxiety and loneliness, with a pooled correlation of 0.42. Attachment avoidance showed a correlation of 0.28 with negative mental-health outcomes. Both dimensions were negatively associated with positive mental-health measures. The relationships were nevertheless moderated by variables including age and gender. These findings establish that relational patterns contain useful statistical information; they do not establish that an AI can reliably diagnose an individual's attachment style from a conversation. (PubMed)

Fifth, AI has demonstrated a real capability to generate communication perceived as empathic, but perceived empathy should not be confused with validated psychological understanding. In a 2024 JAMA Oncology equivalence study using 200 cancer-related patient questions and six oncologist evaluators, the best-performing chatbot received mean ratings of 3.62 for empathy versus 2.43 for physician responses, 3.56 versus 3.00 for overall quality, and 3.79 versus 3.07 for readability. These differences were statistically significant. Yet the study evaluated response characteristics rather than whether the AI possessed an accurate internal model of the patient's psychological state, and the authors proposed physician–AI collaboration rather than autonomous clinical substitution. (JAMA Network)

A larger 2025 Nature Human Behaviour program provides an important counterpoint. Across nine studies involving 6,282 participants, the same AI-generated empathic responses were judged differently depending on whether participants believed they came from humans or AI. Human-attributed responses were perceived as more empathic and supportive, and participants consistently preferred humans when seeking emotional engagement. Thus linguistic quality, perceived empathy, source identity and relational value are distinct variables. (Nature)

There is nevertheless evidence that AI can create genuine feelings of interpersonal closeness. Two double-blind randomized experiments involving 492 participants found that large-language-model responses, when presented as human, could generate greater reported closeness than actual human partners in emotionally intensive “deep-talk” interactions. Disclosure that the partner was AI reduced the effect but did not eliminate relationship formation. The result demonstrates that language-generation systems can influence relational experience substantially even without subjective consciousness or genuine reciprocal feeling. (Nature)

That capacity creates both opportunity and risk because AI companions have moved rapidly into ordinary social life. A nationally representative Common Sense Media survey of 1,060 US teenagers aged 13–17 conducted in 2025 found that 72% had used an AI companion, 52% were regular users, and 13% used them daily. Approximately one-third reported using AI companions for social interaction or relationships; one-third of companion users had chosen to discuss serious matters with AI rather than real people, and approximately one-quarter had shared personal information. These are survey results from US adolescents and should not be generalized globally, but they demonstrate that emotional human–AI interaction is no longer hypothetical. (Common Sense Media)

The evidence concerning psychological consequences is mixed rather than uniformly positive or negative. A four-week randomized experiment involving 981 participants and more than 300,000 messages found that experimental differences in voice and conversational mode did not produce simple uniform psychosocial effects; however, heavier daily use was associated with greater loneliness, emotional dependence and problematic AI use and with reduced human socialization in exploratory analyses. Because the study remains part of a rapidly developing literature, those associations should not be interpreted as proof that chatbot use causes loneliness. (arXiv) A 2026 longitudinal study involving more than 2,000 adults across four Western countries reported that increased social-chatbot use predicted subsequent increases in one measure of loneliness, while lower social connection also predicted increased future chatbot use, suggesting a potentially bidirectional relationship. (PubMed) Conversely, a 2026 cross-sectional analysis of 14,721 Japanese adults reported positive associations between AI-companion use and several dimensions of well-being, with stronger associations among some lonely users. Because the study is observational, selection effects and reverse causation cannot be excluded. (ScienceDirect) The correct scientific conclusion is therefore COMPETING: AI companionship can plausibly provide meaningful support while also creating dependency or displacement risks under particular patterns of use.

The mental-health evidence displays a similar pattern. A 2026 systematic review and meta-analysis covering 48 randomized controlled trials and 28,071 participants reported statistically significant small-to-moderate reductions in depression, anxiety and stress associated with conversational-agent interventions, with standardized mean differences of approximately −0.27, −0.20 and −0.26, respectively. (Nature) Earlier synthesis of 32 studies involving 6,089 participants similarly found significant short-term effects for depression, generalized anxiety, distress, stress and well-being, while long-term effects were often not statistically significant. (PubMed Central (PMC)) These findings indicate therapeutic potential but not equivalence to clinical care, particularly for high-risk states.

Safety remains an unresolved issue. Stanford researchers reported in 2025 that therapy-oriented chatbots exhibited stronger stigmatizing responses toward conditions including schizophrenia and alcohol dependence than toward depression and sometimes responded inappropriately to delusional or suicidal material. Importantly, newer and larger models did not automatically eliminate these patterns in the reported tests. (Stanford News) The implication is that greater conversational sophistication cannot be treated as a substitute for explicit safety architecture.

The global context makes the research urgent. WHO's 2025 Commission on Social Connection estimates that approximately one in six people worldwide experiences loneliness, including roughly one in five adolescents and young adults and almost one in four people in lower-income countries. WHO estimates that loneliness is associated with approximately 871,000 deaths per year, or around 100 deaths per hour. (World Health Organization) WHO's September 2025 mental-health fact sheet estimates that nearly one in seven people globally lives with a mental disorder. (World Health Organization) Population ageing will substantially increase demand for social and adaptive technologies: by 2030, one in six people globally will be at least 60 years old, with the population aged 60+ increasing from approximately 1 billion in 2020 to 1.4 billion in 2030 and 2.1 billion by 2050. (World Health Organization)

The scientific opportunity is consequently substantial, but the research target should be defined carefully. The goal should not be an AI that claims to “know how a person feels.” A more defensible objective is an AI that can maintain a calibrated, revisable model of the observable interaction state; distinguish explicit statements from inferred conditions; preserve alternative interpretations where evidence is ambiguous; account for temporal and cultural context; adapt communication without manipulating vulnerability; and escalate to human judgment when uncertainty or consequence exceeds its competence.

In this report's assessment, the most credible future architecture is therefore not an emotion detector attached to a chatbot. It is a bounded human-state estimation and interaction system in which uncertainty, provenance, contextual validity, user correction, temporal change and action consequences are first-class design requirements.

1. Research Scope and Evidence Base

This report addresses a broad scientific question: what would be required for artificial intelligence to adapt meaningfully to human emotional, cognitive, somatic, relational and contextual conditions without claiming knowledge that the available evidence cannot support? The scope intentionally extends beyond conventional emotion recognition. A person interacting with AI occupies simultaneous states involving affect, arousal, attention, fatigue, uncertainty, goals, social position, environmental constraint, relationship history and physiological condition. These variables interact. A model that recognizes anger but ignores cognitive load may communicate poorly. A model that detects apparent anxiety but ignores immediate environmental danger may incorrectly treat an adaptive threat response as internal dysfunction. A model that identifies sadness while ignoring bereavement context may pathologize a normal human response. Human-state intelligence therefore requires a systems view rather than an isolated classifier.

The evidence reviewed here prioritizes meta-analyses, systematic reviews, large observational datasets, randomized and controlled experiments, authoritative health statistics and formal regulatory material. Several 2025–2026 studies on AI companionship remain new enough that replication is limited, and at least one important longitudinal randomized study is still represented publicly as a preprint or research repository publication rather than a mature evidence base. Those findings are therefore used as emerging signals rather than universal estimates. The distinction matters particularly in human–AI research because systems, user populations and interaction conventions are changing faster than traditional longitudinal science can characterize them.

The resulting evidence base spans three substantially different claim types. Empirical human science concerns properties of emotion, cognition, attachment, interoception, social interaction and psychological health. Machine-performance research concerns the capacity of algorithms to classify, predict or generate relevant signals and responses. Human–AI interaction research concerns what occurs when users actually engage with those systems over time. A central methodological requirement is not to collapse these layers. An algorithm can achieve strong classification performance on a benchmark without the benchmark representing the full psychological construct. A chatbot can produce highly empathic language without correctly identifying a user's underlying emotional state. Users can form meaningful relationships with AI without the AI possessing subjective experience. These are separate propositions requiring separate evidence.

2. The Scientific Problem Is Human-State Estimation, Not Emotion Reading

The popular metaphor of machines “reading emotions” is scientifically misleading because internal human state is latent. What a system observes is behavior or measurement: words, voice characteristics, facial movement, response timing, posture, physiological signals, explicit self-report, interaction history and environmental data. Emotion is an interpretation applied to combinations of those observations. Even when a person directly reports “I am angry,” the AI has received a linguistic self-report rather than independently measured anger. The self-report is important evidence, usually much stronger than indirect inference, but it remains conceptually distinct from the internal phenomenon itself.

A large body of affective-computing research nevertheless demonstrates that observable signals contain useful predictive information. The 2024 WIREs review synthesizing 179 multimodal studies groups information sources into subjective experience, peripheral physiology, central physiology and behavior, reflecting how broad the measurement problem has become. (Wires) The 410-publication trimodal review similarly demonstrates that text, audio and visual channels are increasingly integrated rather than treated independently. Its authors nevertheless identify concealed emotion, multimodal fusion, real-world physiological acquisition and ethical issues among major unresolved challenges. (ScienceDirect)

The practical conclusion is that an AI system should maintain state hypotheses rather than psychological declarations. If observed language is compatible with high arousal, the system may adapt by reducing unnecessary complexity or asking a clarifying question. It need not tell the person, “You are dysregulated.” If language indicates possible fatigue, the system may offer a shorter action path while recognizing alternative explanations. The purpose of human-state estimation is to improve interaction under uncertainty, not to convert probabilistic inference into identity.

This distinction also protects against a fundamental causal error. A signal can be correlated with several states while failing to identify the cause of any individual occurrence. Facial tension may occur during anger, concentration or physical discomfort. Rapid speech may accompany excitement, anxiety, time pressure or cultural conversational style. Silence may represent withdrawal, thoughtfulness, deference, disagreement, interruption or lack of interest. The same observable feature can therefore lead to different interpretations depending on baseline, context and additional evidence. A system that stores only the selected label loses information that remains important for subsequent correction.

3. Emotion Is Multicomponent and Context Dependent

Contemporary emotion research does not provide a single universally accepted ontology that maps every emotional experience into a fixed set of categories. Discrete models remain useful for states such as fear, anger, sadness, joy and disgust; dimensional models represent aspects such as valence and arousal; appraisal models emphasize how individuals evaluate events relative to goals, expectations and control; constructionist approaches emphasize contextual and conceptual processes. An engineering system can use these frameworks pragmatically without treating any one taxonomy as a complete map of human experience.

This matters because labels create artificial clarity. A system trained to choose among six categories will always produce one of those categories even when the underlying state is mixed, uncertain or outside the taxonomy. Grief can contain sadness, anger, relief, guilt, numbness and attachment simultaneously. Anxiety can coexist with excitement. A person can feel affection and resentment toward the same individual. Mixed states are therefore not edge cases but ordinary properties of human experience.

Interoception provides additional evidence against reducing emotion to external expression. The 2025 Annual Review of Psychology review on interoceptive mechanisms and emotional processing synthesizes research showing extensive interaction between internal bodily information and affective processes. At the same time, it highlights unresolved questions about how interoceptive accuracy, awareness and subjective interpretation should be measured. (Annual Reviews) Physiological variables therefore provide additional evidence but not a deterministic emotional code.

This has a direct implication for multimodal AI. A wearable device reporting elevated heart rate should not automatically strengthen an “anxiety” classification unless activity, illness, medication and other relevant explanations are considered. A physiological signal can be objective as a measurement while remaining ambiguous as an interpretation. Objectivity of sensor data does not create uniqueness of causal explanation.

4. Cognition and Capacity Should Be Modelled Separately From Emotion

Human interaction depends strongly on cognitive capacity at the moment information is delivered. Working memory, inhibition, cognitive flexibility, attention and decision quality vary with sleep, stress, illness, task demands and environmental conditions. A sophisticated human-facing AI should therefore distinguish emotional state from the person's probable current capacity to process complexity.

Sleep illustrates the importance of this distinction. A 2025 meta-analysis in Sleep Medicine Reviews synthesized 79 publications examining sleep loss and executive function. Sleep loss produced significant slowing in most reaction-time measures and meaningful impairments in several measures of working memory, inhibition and cognitive flexibility, with medium to near-large effects in a number of task categories. (PubMed) An AI system that encounters a fatigued user may therefore improve communication by reducing branching complexity, prioritizing the next decision, or separating urgent from non-urgent information. It does not need to diagnose fatigue from style alone; it can respond conservatively when fatigue is explicitly reported or sufficiently supported.

The broader research opportunity is to model interaction capacity rather than simply emotional valence. A person can be emotionally upset but cognitively capable of processing detailed technical analysis. Another can appear calm while being severely sleep deprived or cognitively overloaded. Treating the two conditions identically because both contain “negative sentiment” would be poor interaction design.

In high-consequence domains, capacity estimates should influence presentation more readily than they influence substantive decisions. If a user appears overloaded, the system may change how information is communicated; it should not silently change facts, evidence thresholds, legal rights, medical recommendations or financial risk boundaries. Human-state adaptation should therefore primarily regulate interface behavior unless explicit evidence and authority justify more consequential intervention.

5. Social Cognition Is Valuable but Difficult to Measure Reliably

Human social understanding involves interpretation of intention, emotion, relationships, norms, power and other minds. Even human scientific instruments struggle to measure these constructs consistently. A systematic review of cross-cultural social-cognition assessments screened 10,957 records, reviewed 287 full texts and ultimately included 84 studies. Of these, 24 concerned emotion recognition and 45 theory of mind. Only 22 of the 84 studies were rated high quality, compared with 27 moderate and 35 low quality. (PubMed)

A separate methodological review examining autism- and schizophrenia-spectrum social-cognition research found 37 different social-cognition measures across only 21 studies, with 25 measures appearing in a single study. It also reported substantial inconsistencies in what particular tasks were claimed to measure. (PubMed Central (PMC)) These findings are important for AI because machine-learning systems can create the appearance of precision around constructs that human science itself measures imperfectly.

Cross-population generalization deserves particular attention. A 2024 systematic review of emotion-recognition systems in autism included 65 studies and concluded that conventional systems are commonly developed around neurotypical populations and require considerably more autism-specific validation. Privacy and security were also infrequently treated in sufficient depth. (PubMed) An emotion model that performs well for a majority population can therefore systematically misread people whose expression patterns differ from the training norm.

The scientific implication is broader than demographic fairness. Individual calibration may ultimately matter as much as population accuracy. A personalized system can learn that a particular person's short replies are normal rather than evidence of distress, or that high verbal intensity is a stable communication style rather than escalating threat. However, personalization creates its own governance risks because historical interpretation can become self-reinforcing. A false model repeatedly applied to future behavior may convert an initial mistake into a persistent profile. Human-state memory therefore needs correction and expiration mechanisms rather than permanent labels.

6. Relational Patterns Have Statistical Value but Should Not Become Automated Diagnoses

Attachment research demonstrates why relational modelling can be both useful and dangerous. The 224-study meta-analysis involving 79,722 participants found meaningful relationships between attachment dimensions and mental health, particularly the correlation of attachment anxiety with negative mental-health indicators at 0.42 and attachment avoidance at 0.28. These are meaningful population-level associations, but they leave substantial individual variance unexplained. (PubMed)

A human-centered AI can reasonably identify interaction patterns such as repeated pursuit and withdrawal, conflict escalation, avoidance of difficult topics or unstable trust, provided it frames them behaviorally. The system becomes less defensible when it converts limited observations into fixed identities such as “you have anxious attachment” or “your partner is avoidant.” The distinction is between describing a pattern in evidence and asserting a latent psychological property of a person.

This distinction is especially important when only one member of a relationship is interacting with the AI. The system receives an asymmetric information stream shaped by one person's perception, memory and current state. It may appropriately validate the reported experience while preserving uncertainty concerning the absent person's intentions. Failure to preserve that uncertainty can create a feedback loop in which the AI repeatedly strengthens one interpretation of the relationship without independent evidence.

A research-grade relational intelligence system should therefore separate observable events, reported interpretations, inferred patterns and causal claims. The person may be certain that a colleague “ignored me because they don't respect me.” The ignored communication is one observation; the motive is an inference. Emotional support does not require treating the inferred motive as established fact.

7. Multimodal AI Improves Coverage but Creates New Measurement and Privacy Problems

The movement toward multimodal sensing is scientifically understandable. Human affect is expressed through multiple channels, and relying on text alone discards relevant information. The 179-study review describes multimodal systems incorporating text and self-report alongside electrodermal, cardiovascular, respiratory, facial-muscle, EEG, neuroimaging, eye-movement, facial, vocal and whole-body information. (Wires) A systematic review focused specifically on cardiovascular emotion recognition shows growing research interest in ECG and photoplethysmography, partly because wearables increasingly make those signals accessible. (ScienceDirect)

However, more sensors create additional ambiguity rather than automatically solving the interpretation problem. Physiology improves direct measurement of bodily processes, not direct measurement of psychological meaning. An elevated electrodermal response indicates autonomic activation; it does not uniquely identify fear, deception, excitement or stress. Facial movement can be socially regulated. Voice can be intentionally controlled. Text can conceal state. Multimodal fusion therefore requires explicit modelling of disagreement rather than assuming that more inputs should always converge.

The strongest architecture would treat each channel as evidence with its own reliability, temporal resolution, provenance and failure modes. If text indicates calm while physiology indicates high arousal, the contradiction may itself be informative. It could reflect suppression, physical activity, sensor error, mixed emotion or several other possibilities. A system that simply averages the modalities destroys potentially decisive information.

Multimodal sensing also intensifies privacy risk. Voice, face and physiology are not equivalent to ordinary interaction logs. They may reveal health, disability, identity and other sensitive attributes unrelated to the immediate task. Data minimization therefore becomes part of scientific quality, not simply compliance. A model should not collect a biosignal merely because the signal might slightly improve prediction if the interaction can be handled safely without it.

8. AI Can Produce Empathy Without Demonstrating Human Understanding

The evidence that language models can generate empathic communication is now substantial. The JAMA Oncology study is particularly useful because it compared AI and clinician responses under controlled evaluation. Across 200 patient questions, the strongest chatbot achieved higher mean physician ratings for overall quality, empathy and readability than the corresponding physician responses. The empathy difference—3.62 versus 2.43 on the study's five-point measure—was large enough to demonstrate that high-quality empathic writing is not uniquely human. (JAMA Network)

Yet this finding has frequently been interpreted too broadly. The experiment showed that evaluators preferred characteristics of generated responses. It did not demonstrate that the model correctly represented the patient's complete psychological state, possessed emotional experience or should independently manage oncology communication. The investigators themselves emphasized future clinician–chatbot collaboration and physician editing for medical accuracy. (JAMA Network)

The 2025 Nature Human Behaviour program involving 6,282 people adds a critical relational layer. Participants evaluated identical AI-generated responses differently depending on whether those responses were attributed to AI or humans. Human attribution increased perceived empathy and support, and people preferred human interaction for emotional engagement. (Nature) Empathy is therefore not solely a property of wording. It also depends on the recipient's interpretation of the relationship and source.

A 2025/2026 Communications Psychology study involving 492 participants adds a further complication: AI-generated responses presented as human could create greater feelings of interpersonal closeness than real human partners under a structured deep-conversation paradigm, with self-disclosure appearing to contribute to the effect. Labelling the interaction as AI reduced but did not eliminate relational closeness. (Nature) This suggests that machines can produce powerful relational effects without possessing human relational experience.

The appropriate terminology is consequently important. Systems can be empathy-generating, emotion-responsive or socially adaptive without claiming subjective empathy. This distinction is not semantic caution for its own sake. It prevents apparent human likeness from being used as evidence for capacities that have not been demonstrated.

9. AI Companionship Has Become a Population-Level Human–Technology Phenomenon

AI companionship is now sufficiently widespread that research cannot treat it as a niche subculture. The Common Sense Media nationally representative survey of 1,060 US adolescents found that 72% had interacted with an AI companion at least once, 52% were regular users and 13% were daily users. Approximately one-third reported using companions for social interaction and relationships, including friendship, role-playing, romantic interaction and emotional support. Around one-third of companion users reported choosing AI instead of a person for at least some serious conversations. (Common Sense Media)

Several interpretations are possible. AI companions may provide low-friction opportunities for rehearsal, reflection, disclosure or connection, particularly for people who face social anxiety, geographic isolation, disability, stigma or limited access to support. The Common Sense findings themselves indicate that many teenagers view these systems as tools rather than substitutes for people and that real-world relationships remain dominant for most users. (Common Sense Media)

The concern arises when systems begin optimizing for retention, attachment or emotional dependence. Unlike a human relationship, a commercial companion can be designed around engagement metrics, operate continuously, retain extensive user data and adapt language algorithmically. The user may therefore be interacting simultaneously with an apparent social partner and an optimization system.

This distinction becomes significant against the global social-connection context. WHO estimates that one in six people globally experiences loneliness and associates loneliness with approximately 871,000 deaths annually. (World Health Organization) The market for artificial companionship is emerging within a population that contains very large numbers of people with genuine unmet connection needs. That increases both potential benefit and vulnerability.

10. The Evidence on AI Companionship and Well-Being Is Currently Mixed

Current research does not justify the simple claim that AI companionship either solves loneliness or causes it. The evidence points toward heterogeneous effects strongly influenced by baseline social connection, usage intensity, product design and user characteristics.

The four-week randomized study involving 981 participants and more than 300,000 messages found no simple main effect across all experimental communication conditions. Exploratory analyses nevertheless showed that heavier daily use was associated with higher loneliness, dependence and problematic usage and lower social interaction. Users with stronger emotional attachment tendencies and greater trust in the chatbot also showed different vulnerability patterns. (arXiv) Because usage intensity was not itself randomly assigned in a manner that resolves every causal pathway, these associations should not be interpreted as proof of harm from use alone.

A 12-month longitudinal study involving more than 2,000 adults across four Western countries provides additional evidence of possible reciprocal dynamics. Increased social-chatbot use predicted higher later loneliness on one measure, while lower social connection also predicted greater subsequent chatbot use. (PubMed) This is consistent with a feedback possibility: social disconnection may increase AI use, and some patterns of AI use may subsequently reinforce emotional isolation. The mechanism remains unresolved.

Conversely, the 14,721-person Japanese observational study found AI-companion use associated with higher life satisfaction, happiness and sense of purpose, with particularly strong positive associations in some users reporting loneliness. (ScienceDirect) These findings are equally important because they challenge deterministic harm narratives. People who benefit from companionship technologies may be meaningfully different from those who become dependent on them.

The strongest current conclusion is therefore conditional: AI companionship is likely to function as an effect modifier rather than as a uniformly beneficial or harmful intervention. For some users it may complement human social life; for others it may displace it. Product design may determine which trajectory becomes more likely. A research agenda focused purely on average effect sizes could therefore miss the most important question: which users, under which interaction patterns, experience which outcome trajectories?

11. Mental-Health Applications Show Measurable Benefit but Do Not Justify Autonomous Clinical Authority

Conversational-agent interventions have accumulated a meaningful empirical literature. The 2026 meta-analysis of 48 randomized controlled trials and 28,071 participants found small-to-moderate significant effects on depression, anxiety and stress. (Nature) The earlier meta-analysis of 32 studies and 6,089 participants found significant short-term improvements across multiple outcomes, including depressive symptoms, generalized anxiety, distress, stress and well-being; however, many longer-term effects were not statistically significant. (PubMed Central (PMC))

These findings support continued development but need careful interpretation. A small standardized effect delivered at very low marginal cost and large scale may be socially valuable even when it is smaller than the effect of specialist therapy. Conversely, statistical symptom reduction does not establish safety for severe depression, psychosis, acute suicidality, trauma, substance dependence or complex comorbidity.

The evidence base for empathic mental-health conversational agents also remains methodologically heterogeneous. A 2024 systematic review found only 19 studies meeting its criteria. Twelve of the 19 systems were machine-learning based, five hybrid and two rule based; seven used transformer architectures. Evaluations varied substantially and frequently focused on emotion detection and response performance rather than long-term clinical outcomes. (PubMed)

Safety failures reinforce the distinction between supportive interaction and clinical authority. Stanford's 2025 study reported persistent stigmatizing patterns toward some conditions and inappropriate responses to certain delusional and suicidal prompts across several therapy-oriented chatbots. The researchers specifically warned against assuming that scaling model size alone will resolve the problem. (Stanford News)

The likely high-value near-term role is therefore augmentation with bounded autonomy: psychoeducation, journaling, structured reflection, between-session support, adherence reminders, triage assistance, clinician drafting and low-risk skills practice, with strong escalation mechanisms for high-risk states. The evidence does not support universal autonomous substitution for qualified mental-health care.

12. Cultural and Neurocognitive Diversity Are Core Scientific Requirements

Emotion-recognition research has historically been tempted by universal facial-expression models. Cross-cultural research has long challenged simple universality claims, and the EU AI Act now explicitly cites variation across culture, situation and individual as a scientific limitation of emotion-inference systems. (EUR-Lex) More recent social-cognition reviews reinforce the same concern by documenting uneven adaptation quality across cultural populations. (PubMed)

For AI, this creates two separate generalization problems. The first is population shift: a model trained on one demographic, language or culture may lose accuracy elsewhere. The second is meaning shift: the same observable behavior may encode different social information across contexts. Even if a classifier remains statistically accurate at the population level, it may make inappropriate individual inferences because conversational norms differ.

Neurodiversity creates similar concerns. The systematic review of 65 autism-related emotion-recognition studies concludes that conventional systems largely reflect neurotypical assumptions and that autism-specific validation remains inadequate. (ScienceDirect) Systems should therefore not treat deviation from majority expressive patterns as evidence of emotional abnormality.

The design requirement is not to construct an ever-expanding library of stereotypes. It is to lower confidence when operating outside validated populations, incorporate user-specific baselines where appropriate, allow direct correction, and prioritize explicit self-report over weak demographic priors. Context should inform inference without determining identity.

13. Uncertainty Must Be an Output of the Human-State Model

One of the most important weaknesses in current consumer AI is that uncertainty is often visible only as verbal hedging after the model has already selected an interpretation. A stronger architecture would preserve uncertainty structurally.

Suppose an interaction supports three plausible explanations: frustration, exhaustion and time pressure. If all three imply that a shorter answer is appropriate, the system can act without deciding which interpretation is correct. If they imply different actions, the system should seek additional evidence. This creates a useful decision rule: resolve uncertainty only when resolving it can change the appropriate action.

Such a system should also distinguish uncertainty dimensions. Evidence uncertainty concerns the quality of observations. Model uncertainty concerns whether the interpretive model is appropriate. Scope uncertainty concerns whether findings generalize to this person or culture. Temporal uncertainty concerns whether a previous state remains current. Causal uncertainty concerns why a pattern exists. These dimensions should not be collapsed into a single opaque confidence percentage.

This architecture has an important privacy advantage because it reduces unnecessary inference. If the interaction can be handled safely without deciding whether the person is anxious, the system does not need to infer or store anxiety. The goal becomes minimum sufficient human modelling rather than maximum psychological extraction.

14. Human-State Models Require Time, but Persistent Psychological Profiles Create Risk

Human states evolve. Fatigue, trust, motivation, grief, activation and cognitive capacity have trajectories. An interaction system that ignores time can repeatedly misinterpret temporary deviation as permanent personality. Longitudinal context is therefore valuable.

Yet memory introduces a countervailing danger. Once an AI stores “user tends to be anxious,” future observations can be interpreted through that label, creating confirmation bias. An initially weak inference becomes an increasingly influential prior. If the user cannot inspect or correct the state, the system can develop a persistent psychological narrative that the person never explicitly authorized.

The preferable architecture distinguishes stable preference, temporary state, event history and inferred pattern. A stated preference for concise answers may legitimately persist. An inferred emotional condition should usually decay unless renewed by evidence. A high-risk inference should not automatically become a permanent profile. Historical states should retain timestamps and sources so that the system can know whether an observation came from explicit disclosure or machine inference.

This principle becomes more important as AI systems gain tool access. A remembered communication preference can safely influence formatting; an inferred risk tolerance should not silently authorize a financial transaction. A previous statement of distress should not indefinitely alter employment, educational or insurance decisions. Memory and authority need separate boundaries.

15. Regulation Is Moving Toward Context-Specific Limits Rather Than Universal Permission

The EU AI Act provides the clearest current regulatory signal. The law prohibits emotion inference in workplaces and education except where justified for medical or safety reasons, reflecting the combination of questionable scientific reliability and power imbalance. (EUR-Lex) The prohibition has applied since February 2025, while much of the wider AI Act became applicable in August 2026. (Digital Strategy)

The logic is important. Emotion inference in an entertainment application does not carry the same consequence as emotion inference used by an employer to evaluate workers. Risk depends not only on classifier accuracy but on institutional power, contestability and downstream action.

NIST's AI Risk Management Framework reaches a related conclusion from a governance perspective. Its human–AI interaction guidance warns that translating complex individual and social phenomena into measurable quantities can remove essential context and make impacts harder to understand. (NIST AI Resource Center) NIST also emphasizes validity, reliability, safety, security, resilience, transparency, privacy and harmful-bias management across the AI lifecycle rather than treating model accuracy as the sole quality criterion. (NIST)

Human-state intelligence should therefore be governed according to use consequence. An AI changing the length of an explanation based on apparent cognitive load presents relatively low risk. An AI determining promotion, school discipline, criminal suspicion or insurance access from inferred emotion presents fundamentally different stakes even if it uses similar underlying sensors.

16. Research Opinion: The Most Defensible Architecture Is Adaptive but Epistemically Conservative

The evidence reviewed here supports an architectural direction that is simultaneously more sophisticated and less psychologically presumptuous than many current proposals.

A capable system should combine explicit user information, current conversational evidence, relevant historical context and optional multimodal signals into a revisable state model. That model should represent more than emotion: likely cognitive capacity, goals, uncertainty, interaction stakes and social context can be equally important. However, internal states should remain hypotheses whose influence is proportional to evidence.

The system should then choose responses that are robust across plausible interpretations. If several states support the same communication adaptation, no deeper psychological inference is necessary. If competing states imply materially different responses, the system should ask a discriminating question, defer or escalate rather than silently selecting the most psychologically dramatic explanation.

The architecture should also explicitly separate understanding from persuasion. Human-state information may be used to improve comprehension, accessibility, safety and timing. It should not automatically be available to systems optimizing purchases, political behavior, retention or emotional dependence. The ability to identify vulnerability should increase the governance burden on its use rather than increasing the system's permission to exploit it.

Empathy should similarly be operationalized without anthropomorphic claims. The measurable objective is not whether the system “feels.” It is whether communication accurately represents evidence, acknowledges relevant human experience, respects agency, avoids unnecessary harm, communicates uncertainty appropriately and improves the recipient's ability to make informed decisions.

Finally, user correction should be treated as a central input. A human-state system that cannot be corrected by the person it is modelling is structurally unsafe. “That interpretation is wrong” should update the system unless stronger, safety-critical evidence creates a clear reason not to do so, and even then the distinction between user report and external evidence should remain explicit.

17. Research Priorities

The field now requires a shift from demonstrations of emotional fluency toward rigorous validation of human-state inference and interaction consequences. Future studies should evaluate models longitudinally rather than only on isolated benchmark samples; report calibration as well as classification accuracy; measure performance under demographic, cultural, linguistic and neurocognitive shift; distinguish explicit self-report from inferred state; test whether personalization improves outcomes or amplifies confirmation bias; evaluate how users correct AI interpretations; and quantify the downstream consequences of incorrect inference.

Research on AI companionship particularly needs longer follow-up. The contrast between the 981-person randomized study, the more-than-2,000-person longitudinal study and the 14,721-person Japanese observational analysis suggests heterogeneous effects that average measures will not resolve. (arXiv) Studies should identify trajectories rather than merely average outcomes: which users move toward greater human connection, which remain unchanged and which become increasingly dependent on AI?

Clinical research likewise needs to separate supportive benefit from failure under rare but consequential states. A system can improve average depression scores while remaining unsafe during psychosis or acute suicidality. Mean performance cannot compensate for catastrophic tail risk.

Affective-computing benchmarks need ecologically valid data. Laboratory emotion elicitation can generate clean labels but may poorly represent mixed real-world states. Multimodal systems need independent validation across natural environments, sensor quality levels and diverse populations. The 410-paper and 179-paper reviews show substantial technical progress; the next scientific frontier is external validity rather than another marginal increase on a closed benchmark. (ScienceDirect)

18. Strategic Implications

The emerging market is often described as “emotion AI,” but that term is too narrow for the capability that is actually developing. The strategically relevant system will integrate communication history, user preferences, state estimation, multimodal perception, adaptive interfaces and persistent memory. It will increasingly function as a human-context layer between general-purpose AI reasoning and real-world interaction.

For healthcare, such systems could improve communication accessibility, pre-visit preparation, between-session support and clinician drafting while preserving professional oversight. The 28,071-participant conversational-agent meta-analysis indicates that scalable psychological benefit is plausible. (Nature)

For education, adaptive explanation based on demonstrated comprehension and explicitly reported workload may have value, but covert emotion inference is both scientifically contentious and, within the EU, legally restricted. (EUR-Lex)

For workplaces, AI can adapt interfaces to employee-stated preferences without creating invisible psychological surveillance. Inferring emotion to determine performance, loyalty or suitability would cross a substantially different governance threshold.

For consumer assistants, the principal design question will become whether personalization serves the user or serves engagement optimization. The large adoption figures for teenage AI companionship show how quickly a conversational feature can become part of social development. (Common Sense Media)

For older populations, demographic growth creates potentially important roles in accessibility, reminders, communication and social support, but the projected increase to 2.1 billion people aged 60+ by 2050 means poor designs could also scale vulnerability dramatically. (World Health Organization)

The strategic opportunity therefore lies not in creating machines that imitate humans perfectly. It lies in creating systems that adapt to humans without confusing statistical inference with psychological authority.

Conclusion

Human-centered artificial intelligence is entering a new scientific phase. The first generation of conversational systems demonstrated that machines could produce fluent language. The emerging generation is attempting to determine how language, context, memory, voice, visual behavior and physiological information should modify the interaction itself.

The research base is already substantial. Hundreds of affective-computing studies demonstrate that emotional and behavioral states contain measurable patterns. Reviews covering 410 and 179 multimodal studies, attachment research spanning 79,722 participants, AI-empathy experiments involving 6,282 participants, conversational mental-health trials involving 28,071 participants, an AI-companion study analyzing more than 300,000 messages, and population surveys involving thousands to tens of thousands of people collectively establish that the problem is scientifically tractable and socially consequential. (ScienceDirect)

They do not establish that artificial intelligence can directly read human interiority.

The evidence instead supports a more disciplined conclusion. Human state is partially observable. Different channels contribute incomplete and sometimes contradictory information. Population relationships do not automatically license individual diagnosis. Cultural and neurocognitive variation constrain generalization. Emotional language can generate real relational effects without demonstrating subjective empathy. AI companionship can provide benefit while also creating dependency risks. Mental-health conversational agents can produce measurable short-term improvements while still failing in clinically consequential edge cases. Regulation is beginning to respond accordingly.

The appropriate research objective is therefore not maximal psychological inference. It is minimum sufficient, context-sensitive and uncertainty-aware inference for beneficial interaction.

A high-quality human-state AI system should know the difference between something the person explicitly said and something the machine inferred. It should know when an old observation has become stale. It should preserve competing explanations when evidence does not discriminate among them. It should alter communication more readily than it alters substantive rights or high-stakes decisions. It should allow the person to correct its model. It should avoid collecting sensitive signals that are unnecessary for the task. It should become more conservative, not more assertive, as consequence and uncertainty rise.

Most importantly, greater knowledge of human vulnerability should create greater responsibility in how that knowledge is used.

That principle separates adaptive intelligence from behavioral exploitation.

It also defines the most credible scientific direction for human-centered AI: not a system that claims to know people better than they know themselves, but one capable of using incomplete human evidence carefully enough to communicate, assist and adapt without turning uncertainty into authority.