AI toy personalized interaction generation method and system based on multi-modal emotion recognition

By collecting children's voice and facial images in real time, calculating confidence scores and dynamically weighting and fusing them, and combining them with long-term memory profiles to generate interactive decisions, the system drives a large language model to generate personalized content. This solves the technical challenges of AI toys in emotion recognition and interaction continuity, and enhances the experience of emotional understanding and empathetic companionship.

CN122417023APending Publication Date: 2026-07-17SHENZHEN UASCENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN UASCENT TECH CO LTD
Filing Date
2026-06-08
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing AI toys are susceptible to interference in single-modal recognition of emotions, lack multimodal fusion decision-making, and lack personalized interaction continuity, resulting in inaccurate emotion judgment and a lack of continuity in content generation.

Method used

By collecting children's voice and facial images in real time, extracting emotional features and calculating confidence scores, and then dynamically weighting and fusing them with long-term memory profiles to generate interactive decision parameters, the large language model is driven to generate personalized interactive content and perform emotional speech synthesis.

Benefits of technology

It enables multimodal emotion-driven personalized interaction, improves the accuracy of emotion understanding, interaction coherence and empathic companionship experience, and solves the problems of single-modality susceptibility to interference and isolated content generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417023A_ABST
    Figure CN122417023A_ABST
Patent Text Reader

Abstract

本申请涉及AI玩具交互的技术领域,公开基于多模态情感识别的AI玩具个性化交互生成方法与系统,所述方法包括:S1、实时采集交互过程中儿童的语音信号与面部图像序列;S2、对语音信号与面部图像序列分别进行情感特征提取,得到语音情感特征向量与视觉情感特征向量;本发明通过引入多模态情感特征提取与置信度自适应的动态加权融合机制,有效解决了单一模态情感识别易受环境干扰、情感判断准确性与稳定性不足的问题;通过将实时融合情感状态与包含历史偏好及情绪记忆的儿童长期记忆画像进行联合分析,使每次交互均具备个性化连续性与历史感知能力,克服了现有技术中交互内容孤立、千人一面、缺乏情感延续性的缺陷。
Need to check novelty before this filing date? Find Prior Art