Emotion accompanying AI toy system based on multi-modal perception and fusion algorithm

The AI ​​toy system, which utilizes multimodal perception and fusion algorithms, solves the problems of limited interaction and insufficient emotion recognition in smart toys. It enables multi-dimensional feedback and personalized learning, thereby improving the accuracy and realism of emotional companionship.

CN121744196APending Publication Date: 2026-03-27GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-27

Smart Images

  • Figure CN121744196A_ABST
    Figure CN121744196A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion accompanying AI toy system based on a multi-modal perception and fusion algorithm, relates to the technical field of artificial intelligence and intelligent toys, and combines a deep learning model to perform fusion analysis on expressions, voices and tones and touch behaviors of children by using multi-modal perception technologies such as computer vision, voice recognition and touch sensing. Precise recognition of the emotional state of the user is realized through a multi-level weighted fusion algorithm, so that anthropomorphic feedback, including voice pacifying, light change, action response and the like, adaptive to the emotional state is output. The method can be widely applied to the fields of child emotion accompanying, intelligent toy interaction, auxiliary treatment of special children and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and smart toy technology, specifically to an AI toy system for emotional companionship based on multimodal perception and fusion algorithms. Background Technology

[0002] Toys, as important companions in children's growth, significantly impact their psychological development, social skills, and emotional management through interactive experiences and emotional connections. Traditional toys primarily rely on mechanical structures or simple electronic controls, offering limited interaction methods and failing to effectively perceive and respond to children's emotional states. With the development of artificial intelligence technology, smart toys have gradually become a research hotspot, but existing smart toys still have significant shortcomings.

[0003] Currently, most smart toys on the market rely on single-modal interaction, such as supporting only voice command recognition and achieving basic dialogue functions through a pre-set voice command library. However, these toys cannot perceive changes in a child's facial expressions, tone of voice, or physical contact, resulting in a rigid interactive experience lacking emotional warmth and making it difficult to establish a genuine emotional connection. This is especially true for young children or children with special needs whose expressive abilities are limited; relying solely on voice commands cannot meet their companionship needs.

[0004] Furthermore, while some toys equipped with cameras or touch sensors have emerged in the current technology, these products generally suffer from insufficient multimodal information fusion capabilities. Most products process only a single modality signal independently, failing to fully utilize the complementarity and correlation between different modalities. For example, when a child speaks in a sad tone, their facial expression often simultaneously displays sadness, and their tactile behavior may manifest as prolonged cuddling. Relying solely on speech recognition may lead to misjudgments due to speech quality issues; however, combining visual and tactile information for comprehensive judgment can significantly improve the accuracy and robustness of emotion recognition.

[0005] In recent years, multimodal fusion technology has demonstrated enormous potential in the field of human-computer interaction. By integrating information from multiple sensory channels such as speech, vision, and touch, the system can more comprehensively understand the user's state, thereby providing more natural and accurate feedback. However, existing research has largely focused on adult interaction scenarios, with fewer multimodal emotion recognition systems designed specifically for children. Children's emotional expressions are often more direct and diverse, accompanied by frequent physical movements and tactile behaviors, which places higher demands on the real-time performance, accuracy, and adaptability of multimodal fusion algorithms.

[0006] Furthermore, existing smart toys have limited emotional feedback mechanisms, mostly relying solely on voice playback of preset content, lacking multi-dimensional sensory stimulation such as light, movement, and temperature. Psychological research shows that children are more sensitive to emotional responses to multi-sensory stimuli, and the combined use of visual, auditory, and tactile feedback can significantly enhance the realism and intimacy of the companionship experience. At the same time, existing toys lack personalized learning capabilities, failing to adaptively adjust recognition thresholds and feedback strategies based on children's interaction history, leading to long-term homogenization and fatigue in the interactive experience. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides an AI toy system for emotional companionship based on multimodal perception and fusion algorithms, in order to solve the problems of limited interaction, insufficient emotion recognition capabilities, rigid feedback mechanisms, and lack of personalized learning capabilities in existing smart toys.

[0008] To achieve the above objectives, the present invention provides the following technical solution: an AI toy system for emotional companionship based on multimodal perception and fusion algorithms, comprising:

[0009] A multimodal perception module is used to simultaneously collect physiological and behavioral data of users through two different types of sensors, and to perform preliminary processing on the data to generate raw feature data containing at least two modalities.

[0010] The emotion computing module is used to receive the raw feature data and perform fusion processing to determine the user's current emotional state; the fusion processing includes: dynamically adjusting the fusion weight of each modality in the process of determining the final emotional state based on the real-time signal quality index and historical recognition accuracy index of at least one modality in the multimodal data;

[0011] The feedback control module is used to generate and execute multi-dimensional anthropomorphic feedback, which is a combination of two or more feedback forms, including voice, light and action, based on the emotional state determined by the emotion computing module.

[0012] The learning optimization module is used to adaptively optimize the feedback strategy of the feedback control module based on the user's subsequent emotional changes, so as to achieve personalized interaction.

[0013] Preferably, the multimodal sensing module specifically includes:

[0014] The speech recognition unit is used to collect the user's speech signal through a microphone array and extract acoustic features including pitch, energy and speech rate;

[0015] The visual recognition unit is used to capture facial images of users through a camera and extract facial features, including the angle of the corners of the mouth and the position of the eyebrows.

[0016] The tactile sensing unit is used to detect the user's touch behavior through a touch sensor array and extract tactile parameters including contact area, pressure intensity and duration.

[0017] Preferably, the emotion computing module employs a hierarchical weighted fusion mechanism, which includes:

[0018] The feature-level fusion unit is used to standardize the original feature data of each modality and then perform weighted combination by a first set of preset weights to generate a comprehensive feature vector.

[0019] The decision-level fusion unit is used to input the original feature data of each modality into an independent classifier to obtain the emotion probability distribution of each modality, and to perform a weighted summation of the emotion probability distribution based on the dynamically adjusted fusion weights to generate a comprehensive emotion probability distribution.

[0020] Preferably, the calculation logic for dynamically adjusting the fusion weights in the decision-level fusion unit includes:

[0021] The ratio of the number of times each modality was correctly recognized in historical interactions to the total number of recognitions is obtained to calculate the historical recognition accuracy index; the input signals of each modality are analyzed in real time, and the real-time signal quality index is obtained by calculating the signal-to-noise ratio of the speech signal, the sharpness of the image signal, or the stability of the tactile signal.

[0022] The historical recognition accuracy index and the real-time signal quality index are weighted and summed using a preset balance coefficient to calculate the updated fusion weight for the current interaction.

[0023] Preferably, the emotion computing module further includes an emotion confidence assessment unit, which is used to calculate an emotion confidence score after determining the emotion state;

[0024] When the emotional confidence level is lower than a preset confidence threshold, the feedback control module is triggered to execute a preset emotional exploration strategy. The emotional exploration strategy is a specific interactive action used to guide the user to generate clear emotional feedback. At the same time, the multimodal perception module is instructed to enter an enhanced perception mode to focus on collecting and analyzing the user's response data to the emotional exploration strategy, thereby verifying or correcting the user's emotional state.

[0025] Preferably, the feedback control module internally stores an emotional feedback mapping knowledge base, which defines the initial correspondence between different emotional states and multi-dimensional anthropomorphic feedback, wherein the multi-dimensional anthropomorphic feedback includes at least:

[0026] For sadness, output a combination of soothing voice, soft gradient lighting, and patting or hugging gestures; for happiness, output an encouraging voice, bright breathing lighting, and swaying or nodding gestures.

[0027] Preferably, the learning optimization module uses a reinforcement learning algorithm to adaptively optimize the feedback strategy, and its optimization logic includes:

[0028] The emotional state output by the emotion computing module is defined as a state in reinforcement learning; the multi-dimensional anthropomorphic feedback combination executed by the feedback control module is defined as an action in reinforcement learning; by comparing the degree of positive change in the user's emotional state before and after executing the action, it is quantified as a reward signal; through continuous iteration, a policy function for selecting actions is updated so that the action selected in a specific state can maximize the long-term accumulated reward signal, thereby achieving personalized adaptation of the feedback strategy.

[0029] Preferably, the calculation logic of the reward signal includes: after performing the feedback action, collecting the user's subsequent physiological and behavioral data through the multimodal perception module, and determining the user's subsequent emotional state; if the subsequent emotional state changes in a positive direction compared to the initial emotional state, a positive reward value is generated, and the greater the change, the greater the positive reward value; if the subsequent emotional state does not change or changes in a negative direction, a zero reward value or a negative reward value is generated.

[0030] Preferably, multimodal data of the user is collected using sensors of at least two different modalities; the multimodal data is processed by an emotion computing module to determine the user's emotional state, wherein the processing steps include:

[0031] Based on the real-time signal quality and historical recognition accuracy of at least one modality in the multimodal data, the fusion weights of each modality data in the process of determining the final emotional state are dynamically adjusted.

[0032] Based on the determined emotional state, a multi-dimensional anthropomorphic feedback is generated and executed; according to the change in the user's emotional state after receiving the feedback, the strategy for generating the feedback is iteratively optimized using a reinforcement learning algorithm.

[0033] Preferably, the confidence level of the current emotional state is calculated and it is determined whether the confidence level is lower than a preset threshold. If so, a preset emotion exploration action is executed to actively guide the user to generate a clear emotional signal, and the signal is enhanced, collected, and analyzed to verify the emotional state.

[0034] This invention provides an AI toy system for emotional companionship based on multimodal perception and fusion algorithms. It has the following beneficial effects:

[0035] This invention addresses the limitations and fragility of existing technologies that rely on a single modality in emotion recognition by constructing a comprehensive multimodal perception module. Specifically, the multimodal perception module includes a speech recognition unit that acquires speech signals via a microphone array and extracts acoustic features including pitch, energy, and speech rate; a visual recognition unit that acquires facial images via a camera and extracts facial expression features including the angle of the corners of the mouth and the position of the eyebrows; and a tactile perception unit that detects touch behavior via a touch sensor array and extracts tactile parameters including contact area, pressure intensity, and duration. Based on this, the emotion computing module employs a hierarchical weighted fusion mechanism. A feature-level fusion unit weights and combines the original feature data to generate a comprehensive feature vector, and a decision-level fusion unit fuses the emotion probability distributions of each modality. This multi-source information complementarity and cross-validation approach significantly improves the overall accuracy and robustness of emotion recognition, effectively avoiding misjudgments caused by poor quality signals from a single modality.

[0036] This invention addresses the technical shortcomings of existing technologies, such as static fusion strategies that cannot adapt to diverse interactive scenarios, by introducing a unique dynamic weight adjustment mechanism. Specifically, the decision-level fusion unit dynamically calculates the fusion weights when fusing the probability distributions of emotions across modalities. This calculation logic includes: first, obtaining the ratio of the number of times each modality was correctly recognized in historical interactions to the total number of recognitions to calculate the historical recognition accuracy index; simultaneously, analyzing the input signals of each modality in real time, and obtaining real-time signal quality indices by calculating the signal-to-noise ratio of the speech signal, the clarity of the image signal, or the stability of the tactile signal; finally, weighting and summing the historical recognition accuracy index and the real-time signal quality index using a preset balance coefficient to calculate the updated fusion weights for the current interaction. In addition, the system also includes an emotion confidence assessment unit. When the calculated emotion confidence is lower than a preset confidence threshold, the system actively executes an emotion exploration strategy and enters an enhanced perception mode. This dual adaptive mechanism ensures that the system can achieve high-precision emotion recognition under different environmental conditions, making it particularly suitable for children with diverse emotional expressions.

[0037] This invention addresses the pain points of existing technologies, such as single feedback mechanisms, homogenized experiences, and lack of evolution, by designing a closed-loop system with multi-dimensional feedback and personalized learning optimization. The feedback control module can call and execute multi-dimensional anthropomorphic feedback composed of voice, light, and action combinations based on emotional state from an emotional feedback mapping knowledge base. For example, it can output a combination of soothing voice, soft gradient lighting, and patting or hugging actions for sadness, significantly enhancing the realism of emotional companionship. More importantly, the learning optimization module uses a reinforcement learning algorithm, defining emotional state as a state and feedback combination as an action, and quantifying and generating a reward signal by comparing the positive change in the user's emotional state before and after executing the action. The system continuously iterates and updates a policy function for selecting actions, ensuring that the action selected in a specific state maximizes the long-term accumulated reward signal. This achieves adaptive evolution of the feedback strategy from standardization to deep personalization, effectively avoiding fatigue caused by long-term interaction and establishing a sustainable deep emotional connection. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of the framework structure of an AI toy system for emotional companionship based on multimodal perception and fusion algorithms according to the present invention.

[0039] Figure 2 This is a schematic diagram of the system workflow of an AI toy system for emotional companionship based on multimodal perception and fusion algorithms according to the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1

[0042] Please see Figure 1 This invention provides an AI toy system for emotional companionship based on multimodal perception and fusion algorithms, comprising:

[0043] A multimodal perception module is used to simultaneously collect physiological and behavioral data of users through two different types of sensors, and to perform preliminary processing on the data to generate raw feature data containing at least two modalities.

[0044] The emotion computing module is used to receive the raw feature data and perform fusion processing to determine the user's current emotional state; the fusion processing includes: dynamically adjusting the fusion weight of each modality in the process of determining the final emotional state based on the real-time signal quality index and historical recognition accuracy index of at least one modality in the multimodal data;

[0045] The feedback control module is used to generate and execute multi-dimensional anthropomorphic feedback, which is a combination of two or more feedback forms, including voice, light and action, based on the emotional state determined by the emotion computing module.

[0046] The learning optimization module is used to adaptively optimize the feedback strategy of the feedback control module based on the user's subsequent emotional changes, so as to achieve personalized interaction.

[0047] In this embodiment, within the sentiment computing module, the present invention employs a hierarchical weighted fusion algorithm to comprehensively analyze multimodal features; this algorithm comprises two levels: feature-level fusion and decision-level fusion. In the feature-level fusion stage, the original features of each modality are first standardized to eliminate dimensional differences.

[0048] Let the speech feature vector be f, the visual feature vector be fi, and the tactile feature vector be ft. After standardization, we obtain f, fi, and ft. The standardization formula is as follows:

[0049]

[0050] Where μk and σk are the feature mean and standard deviation of mode k, respectively. Feature-level fusion is then performed using a weighted linear combination.

[0051]

[0052] in, `at` represents the feature weight parameters for each modality, satisfying the constraint Q + ai + Qt = 1. This feature-level fusion vector F1 is then input into the emotion classification network to generate preliminary emotion judgment results.

[0053] In the decision-level fusion stage, each modality-independent emotion classifier outputs an emotion probability distribution. Let the total number of emotion categories be C (e.g., happiness, sadness, anger, fear, etc.), then the probability distribution of the speech modality output is p, = p1, p2, ..., pc, and the probability distribution of the visual modality output is...

[0054] pi = [pi1, pi2, ..., pic], where tactile modality output probability distribution is pt1, pt2, ..., ptc. The comprehensive emotion probability distribution is obtained by weighting with dynamic decision weights.

[0055]

[0056] in, βi and βt are decision-level weight parameters that satisfy... This two-layer fusion mechanism utilizes the complementarity of features while preserving the reliability of independent judgments of each modality, thereby significantly improving the accuracy and robustness of emotion recognition.

[0057] This embodiment further introduces a dynamic weight adaptive update mechanism, which automatically adjusts the fusion weights based on the signal quality of each modality and the historical recognition accuracy. The dynamic weight update formula is as follows:

[0058]

[0059] Among them, Acc (t) k Let Q be the recognition accuracy of modality k in the t-th interaction. (t) k γ is the signal quality index of mode k (including speech signal-to-noise ratio, image clarity, and touch signal stability), and δ is the balance coefficient, with typical values ​​of 0.7 and 0.3, respectively. This mechanism enables the system to automatically optimize the contribution of each mode under different environmental conditions and interaction scenarios, significantly enhancing the system's environmental adaptability and long-term stability.

[0060] To improve the stability and consistency of emotion recognition, this invention introduces an emotion confidence calculation model, comprehensively considering the consistency of multimodal judgments and historical emotion distribution. The final emotion output adopts a confidence-weighted maximization strategy: where Pfcusion is the probability value of the C-th emotion after fusion, and Confc is the confidence score of that emotion, calculated as follows:

[0061] confc=λ1×std(pc,pic,ptc)1+λ2×Historyc

[0062] Where std(PC, pic, ptc) is the standard deviation of the probability of the three-modal judgment of the C-th emotion. The smaller the standard deviation, the higher the consistency of the multimodal judgment and the higher the confidence. Historyc is the frequency weight of the emotion in the historical interactions. λ1 and λ2 are empirical weight coefficients. This confidence model can effectively suppress misjudgments caused by single-modal noise and improve the robustness of emotion recognition.

[0063] To avoid short-term fluctuations in emotion recognition results, this invention introduces a time smoothing mechanism, which uses a sliding window to perform a weighted average of emotional states over a continuous period of time.

[0064] Et=(1p)Et1+p×Emotint

[0065] Where Et represents the final output emotion state at time t, Emotiont represents the current recognition result, and Et1 represents the emotion state at the previous time step.

[0066] pe(0,1) is a smoothing coefficient that controls the weight of new recognition results. This mechanism makes emotion recognition results more consistent over time, avoiding feedback jumps caused by instantaneous noise, thereby improving the naturalness and fluency of the interactive experience.

[0067] The feedback control module generates multi-dimensional, human-like feedback instructions based on the identified emotional state Et, combined with the current interaction scenario and historical interaction data. The feedback decision function is expressed as:

[0068] Response = fct rl (Et, content, history)

[0069] Among them, fct rl The feedback mapping function is defined by `conte`, which represents the current scene parameters, including time period, ambient lighting, and interaction frequency. `History` records the user's historical emotional trends and preferences. The feedback output includes three dimensions:

[0070] First, the voice feedback unit selects appropriate soothing voices, encouraging words, or interactive dialogue content from a preset voice library or a generative voice synthesis module based on the emotion category.

[0071] Second, the LED light unit adjusts the color, brightness and breathing frequency of the light according to the emotional level. For example, happy emotions correspond to warm yellow breathing lights, sad emotions correspond to soft blue gradient lights, angry emotions correspond to red flashing lights, and fearful emotions correspond to soft pink lights.

[0072] Third, the motion-driven unit controls the toy's rocking, hugging, and nodding movements through servo motors or servo motors, enabling emotional expression at the physical level. This multi-dimensional feedback mechanism can simultaneously stimulate children through multiple sensory channels such as hearing, sight, and touch, significantly enhancing the realism and immersion of emotional companionship.

[0073] The learning optimization module employs reinforcement learning algorithms to adaptively optimize the emotion recognition threshold and feedback strategy. By recording the child's subsequent emotional changes after each interaction, the system constructs a reward function. When the child's emotion changes from negative to positive (e.g., from sadness to a smile), the reward value of the current feedback strategy is increased, and the weight of the corresponding modality is increased accordingly; conversely, the reward value is decreased, and the strategy parameters are adjusted. The reward function takes the following form:

[0074] Rt=w1×Epositive-W2×AEnegotive

[0075] Here, Epositive represents the increment of positive emotions, Enegative represents the increment of negative emotions, and W1 and w2 are reward weighting coefficients. Through iterative optimization, the system can gradually learn personalized recognition thresholds and feedback strategies for specific children, achieving continuous optimization and customized experiences in long-term companionship.

[0076] Example 2

[0077] Please see Figure 2 The multimodal sensing module specifically includes:

[0078] The speech recognition unit is used to collect the user's speech signal through a microphone array and extract acoustic features including pitch, energy and speech rate;

[0079] The visual recognition unit is used to capture facial images of users through a camera and extract facial features, including the angle of the corners of the mouth and the position of the eyebrows.

[0080] The tactile sensing unit is used to detect the user's touch behavior through a touch sensor array and extract tactile parameters including contact area, pressure intensity and duration.

[0081] The emotion computing module employs a hierarchical weighted fusion mechanism, which includes:

[0082] The feature-level fusion unit is used to standardize the original feature data of each modality and then perform weighted combination by a first set of preset weights to generate a comprehensive feature vector.

[0083] The decision-level fusion unit is used to input the original feature data of each modality into an independent classifier to obtain the emotion probability distribution of each modality, and to perform a weighted summation of the emotion probability distribution based on the dynamically adjusted fusion weights to generate a comprehensive emotion probability distribution.

[0084] The calculation logic for dynamically adjusting the fusion weights in the decision-level fusion unit includes:

[0085] The ratio of the number of times each modality was correctly recognized in historical interactions to the total number of recognitions is obtained to calculate the historical recognition accuracy index; the input signals of each modality are analyzed in real time, and the real-time signal quality index is obtained by calculating the signal-to-noise ratio of the speech signal, the sharpness of the image signal, or the stability of the tactile signal.

[0086] The historical recognition accuracy index and the real-time signal quality index are weighted and summed using a preset balance coefficient to calculate the updated fusion weight for the current interaction.

[0087] The emotion computing module also includes an emotion confidence assessment unit, which is used to calculate an emotion confidence score after determining the emotion state.

[0088] When the emotional confidence level is lower than a preset confidence threshold, the feedback control module is triggered to execute a preset emotional exploration strategy. The emotional exploration strategy is a specific interactive action used to guide the user to generate clear emotional feedback. At the same time, the multimodal perception module is instructed to enter an enhanced perception mode to focus on collecting and analyzing the user's response data to the emotional exploration strategy, thereby verifying or correcting the user's emotional state.

[0089] The feedback control module internally stores an emotional feedback mapping knowledge base. This knowledge base defines the initial correspondence between different emotional states and multi-dimensional anthropomorphic feedback. The multi-dimensional anthropomorphic feedback includes at least:

[0090] For sadness, output a combination of soothing voice, soft gradient lighting, and patting or hugging gestures; for happiness, output an encouraging voice, bright breathing lighting, and swaying or nodding gestures.

[0091] The learning optimization module uses a reinforcement learning algorithm to adaptively optimize the feedback strategy, and its optimization logic includes:

[0092] The emotional state output by the emotion computing module is defined as a state in reinforcement learning; the multi-dimensional anthropomorphic feedback combination executed by the feedback control module is defined as an action in reinforcement learning; by comparing the degree of positive change in the user's emotional state before and after executing the action, it is quantified as a reward signal; through continuous iteration, a policy function for selecting actions is updated so that the action selected in a specific state can maximize the long-term accumulated reward signal, thereby achieving personalized adaptation of the feedback strategy.

[0093] The calculation logic of the reward signal includes: after the feedback action is performed, the user's subsequent physiological and behavioral data are collected through the multimodal perception module, and the user's subsequent emotional state is determined; if the subsequent emotional state changes to a positive direction compared to the initial emotional state, a positive reward value is generated, and the greater the change, the greater the positive reward value; if the subsequent emotional state does not change or changes to a negative direction, a zero reward value or a negative reward value is generated.

[0094] Multimodal data of the user is collected using sensors with at least two different modalities; the multimodal data is processed by an emotion computing module to determine the user's emotional state, wherein the processing steps include:

[0095] Based on the real-time signal quality and historical recognition accuracy of at least one modality in the multimodal data, the fusion weights of each modality data in the process of determining the final emotional state are dynamically adjusted.

[0096] Based on the determined emotional state, a multi-dimensional anthropomorphic feedback is generated and executed; according to the change in the user's emotional state after receiving the feedback, the strategy for generating the feedback is iteratively optimized using a reinforcement learning algorithm.

[0097] Calculate the confidence level of the current emotional state and determine whether the confidence level is lower than a preset threshold. If so, execute a preset emotional exploration action to actively guide the user to generate a clear emotional signal, and enhance the collection and analysis of the signal to verify the emotional state.

[0098] In this embodiment, the voice recognition unit acquires the child's voice signal through a MEMS digital microphone array (typically configured with 2 to 4 microphone units) embedded in the toy's head or chest. The microphone array employs beamforming technology for spatial directional enhancement, suppressing environmental noise and reverberation interference, and improving the quality of the voice signal. The acquired raw audio signal is enhanced with a pre-emphasis filter to enhance high-frequency components, and then subjected to frame processing, with a typical frame length of 25 milliseconds and a frame shift of 10 milliseconds. For each frame of the speech signal, a Fast Fourier Transform (FFT) is performed to obtain the spectrum, and the following acoustic features are extracted: First, the fundamental frequency (Pitch) feature, extracted using autocorrelation or cepstral method, reflects the pitch of the speech and is used to identify the level of excitement or frustration in the tone of voice; Second, the energy feature, calculating the short-time energy of each frame of the signal, reflects the loudness and emotional intensity of the speech; Third, the speech rate feature, calculated by detecting the number of syllables per unit time, with fast speech rate usually corresponding to anxiety or excitement, and slow speech rate corresponding to sadness or fatigue; Fourth, Mel-frequency cepstral coefficients (MFCC), extracting 13 to 39-dimensional MFCC coefficients and their first and second-order difference features, as the core features for speech content recognition and emotion recognition. The above features are combined to form a speech feature vector. The data is then fed into a speech emotion classifier based on a recurrent neural network (RNN) or a long short-term memory network (LSTM), which outputs the emotion probability distribution P of the speech modality.

[0099] The visual recognition unit captures images of children's faces using a wide-angle camera embedded in the toy's face, typically with a resolution of 640×480 pixels or higher and a frame rate of 15 to 30 frames per second. First, face detection algorithms (such as those based on Haar cascade classifiers or deep learning-based MTCNN models) are used to locate the face region in the image. After face detection, face alignment is performed by detecting key points such as the eyes, nose tip, and corners of the mouth and performing affine transformations to normalize the face to a standard pose and size. Subsequently, a convolutional neural network (CNN) is used to extract facial expression features. In this embodiment, a pre-trained lightweight convolutional network (such as MobileNet or SqueezeNet) is used, fine-tuned on an expression recognition dataset (such as FER2013 or CK+) to adapt it to children's expression characteristics. The feature vector fi extracted by the network includes geometric features of key facial regions (such as eye curvature, mouth corner angle, and eyebrow height) and texture features (such as wrinkle distribution and muscle tension). The feature vector is input into the expression classifier, which outputs the emotional probability distribution pi of the visual modality, including categories such as happiness, sadness, anger, fear, surprise, disgust, and calmness.

[0100] The tactile sensing module units are distributed on the surface of the toy using an array of touch sensors to detect children's tactile behavior. In this embodiment, capacitive touch sensors are used, which sense the touch location and contact area by detecting changes in capacitance. A typical sensor array layout is as follows: 6 to 10 sensing units are arranged in the head area, 8 to 12 sensing units in the back area, and 2 to 4 sensing units in each of the limb areas. After the sensor output signal is filtered by a low-pass filter to remove high-frequency noise, the following tactile features are extracted: First, the contact area s, estimated by statistically counting the number of activated sensor units; a single-point touch corresponds to a small area, while hugging or stroking corresponds to a large area; Second, the pressure intensity pt, estimated by the amplitude of capacitance change; a light touch corresponds to low pressure, while forceful pressing corresponds to high pressure; Third, the pressure change rate Δp, calculated by the difference in pressure values ​​within adjacent time windows; rapid patting corresponds to a high change rate, while continuous stroking corresponds to a low change rate; Fourth, the contact duration T, the length of time from the start to the end of the touch; a brief touch lasts less than 1 second, while continuous cuddling can last for tens of seconds. The above feature combination forms a tactile feature vector ft = [s, pt, AP, T], which is then input into a tactile emotion classifier, outputting the emotion probability distribution pt of the tactile modality. The tactile emotion classification rules are designed based on the conclusions of psychological research, for example: continuous gentle stroking usually corresponds to a calm or comfort-seeking state, brief forceful slapping may correspond to anger or anxiety, and long-term hugging corresponds to attachment or fear;

[0101] The workflow of this invention is illustrated below through specific application scenarios:

[0102] Scene 1: Emotional comfort for children who are grieving.

[0103] A child cries because a toy is broken. The speech recognition module detects crying sound characteristics, including large fundamental frequency fluctuations, high energy, and slow speech rate. The visual recognition module captures sad facial expression features such as frowning, downturned corners of the mouth, and moist eyes. The tactile perception module detects that the child hugs the toy for a long time, with a large touching area, moderate pressure, and long duration. The emotion computing module fuses the three modal information. The speech modality outputs a sadness probability of 0.85, the visual modality outputs a sadness probability of 0.90, and the tactile modality outputs a seeking comfort probability of 0.80. Through decision-level fusion, the sadness category scores the highest in the overall emotion probability distribution. Confidence calculation shows high consistency among the three modal judgments, and the final output emotion state is "sad." The feedback control module executes the following combined feedback: the speech unit plays the comforting voice "Don't be sad, I'll always be with you," the light unit displays a soft blue gradient light effect, and the motion unit performs a patting motion, slowly raising and gently lowering the arm. After 30 seconds of feedback, the visual recognition module detected that the child's crying had lessened and the touching behavior had changed to a gentle stroking. The learning optimization module recorded that the feedback strategy was effective and increased the reward value of voice and action feedback, and prioritized the use of this strategy in subsequent interactions.

[0104] Scene 2: Interactive responses when children are happy.

[0105] A child successfully completes a jigsaw puzzle and excitedly tells the toy, "I did it!" The speech recognition module detects high-pitched, high-energy, and fast-paced speech; the visual recognition module captures happy facial expressions such as a smile, wide eyes, and upturned corners of the mouth; the tactile perception module detects the child briefly patting the toy's head, with a small contact area, high pressure, and short duration. The emotion computing module integrates these features and outputs the emotional state as "happy." The feedback control module executes the following combination of feedback: the speech unit plays the encouraging voice "You're great, keep it up!", the light unit displays a warm yellow breathing light effect, and the motion unit performs a swaying motion, moving the head from side to side. This interaction enhances the child's sense of accomplishment and positive emotions. The learning optimization module records that the feedback strategy is effective in this scenario, and subsequently prioritizes encouraging voices and swaying motions when the child is in a happy emotional state. Scenario 3: Robust recognition under environmental noise interference.

[0106] In noisy environments (such as when a television is playing sound), a child speaks to a toy, but the speech signal is interfered with by noise. This results in high uncertainty in the emotion probability distribution output by the speech recognition module, with the highest probability being only 0.50. Simultaneously, the visual recognition module works normally, capturing the child's calm expression and outputting a calm probability of 0.85; the tactile perception module detects gentle stroking and outputs a calm probability of 0.75. A dynamic weight adaptive mechanism detects low speech signal quality (signal-to-noise ratio less than 10dB) and automatically reduces the speech modality weights. The weights were increased to 0.15, raising the visual and tactile modal weights βi = 0.55 and βt = 0.30. The fused output emotional state was "calm," avoiding misjudgments caused by speech noise. This scenario demonstrates the robustness of the dynamic weight adaptation mechanism under environmental interference.

[0107] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An AI toy system for emotional companionship based on multimodal perception and fusion algorithms, characterized in that, include: A multimodal perception module is used to simultaneously collect physiological and behavioral data of users through two different types of sensors, and to perform preliminary processing on the data to generate raw feature data containing at least two modalities. The emotion computing module is used to receive the raw feature data and perform fusion processing to determine the user's current emotional state; The fusion process includes: dynamically adjusting the fusion weight of each modality in determining the final emotional state based on the real-time signal quality index and historical recognition accuracy index of at least one modality in the multimodal data; The feedback control module is used to generate and execute multi-dimensional anthropomorphic feedback, which is a combination of two or more feedback forms, including voice, light and action, based on the emotional state determined by the emotion computing module. The learning optimization module is used to adaptively optimize the feedback strategy of the feedback control module based on the user's subsequent emotional changes, so as to achieve personalized interaction.

2. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The multimodal sensing module specifically includes: The speech recognition unit is used to collect the user's speech signal through a microphone array and extract acoustic features including pitch, energy and speech rate; The visual recognition unit is used to capture facial images of users through a camera and extract facial features, including the angle of the corners of the mouth and the position of the eyebrows. The tactile sensing unit is used to detect the user's touch behavior through a touch sensor array and extract tactile parameters including contact area, pressure intensity and duration.

3. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The emotion computing module employs a hierarchical weighted fusion mechanism, which includes: The feature-level fusion unit is used to standardize the original feature data of each modality and then perform weighted combination by a first set of preset weights to generate a comprehensive feature vector. The decision-level fusion unit is used to input the original feature data of each modality into an independent classifier to obtain the emotion probability distribution of each modality, and to perform a weighted summation of the emotion probability distribution based on the dynamically adjusted fusion weights to generate a comprehensive emotion probability distribution.

4. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The calculation logic for dynamically adjusting the fusion weights in the decision-level fusion unit includes: The ratio of the number of times each modality was correctly recognized in historical interactions to the total number of recognitions is obtained to calculate the historical recognition accuracy index; the input signals of each modality are analyzed in real time, and the real-time signal quality index is obtained by calculating the signal-to-noise ratio of the speech signal, the sharpness of the image signal, or the stability of the tactile signal. The historical recognition accuracy index and the real-time signal quality index are weighted and summed using a preset balance coefficient to calculate the updated fusion weight for the current interaction.

5. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The emotion computing module also includes an emotion confidence assessment unit, which is used to calculate an emotion confidence score after determining the emotion state. When the emotional confidence level is lower than a preset confidence threshold, the feedback control module is triggered to execute a preset emotional exploration strategy. The emotional exploration strategy is a specific interactive action used to guide the user to generate clear emotional feedback. At the same time, the multimodal perception module is instructed to enter an enhanced perception mode to focus on collecting and analyzing the user's response data to the emotional exploration strategy, thereby verifying or correcting the user's emotional state.

6. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The feedback control module internally stores an emotional feedback mapping knowledge base. This knowledge base defines the initial correspondence between different emotional states and multi-dimensional anthropomorphic feedback. The multi-dimensional anthropomorphic feedback includes at least: For sadness, output a combination of soothing voice, soft gradient lighting, and patting or hugging gestures; for happiness, output an encouraging voice, bright breathing lighting, and swaying or nodding gestures.

7. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The learning optimization module uses a reinforcement learning algorithm to adaptively optimize the feedback strategy, and its optimization logic includes: The emotional state output by the emotion computing module is defined as a state in reinforcement learning; the multi-dimensional anthropomorphic feedback combination executed by the feedback control module is defined as an action in reinforcement learning; by comparing the degree of positive change in the user's emotional state before and after executing the action, it is quantified as a reward signal; through continuous iteration, a policy function for selecting actions is updated so that the action selected in a specific state can maximize the long-term accumulated reward signal, thereby achieving personalized adaptation of the feedback strategy.

8. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: The calculation logic of the reward signal includes: after the feedback action is performed, the user's subsequent physiological and behavioral data are collected through the multimodal perception module, and the user's subsequent emotional state is determined; if the subsequent emotional state changes to a positive direction compared to the initial emotional state, a positive reward value is generated, and the greater the change, the greater the positive reward value; if the subsequent emotional state does not change or changes to a negative direction, a zero reward value or a negative reward value is generated.

9. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: Multimodal data of the user is acquired through sensors with at least two different modalities; The multimodal data is processed by an emotion computing module to determine the user's emotional state, wherein the processing steps include: Based on the real-time signal quality and historical recognition accuracy of at least one modality in the multimodal data, the fusion weights of each modality data in the process of determining the final emotional state are dynamically adjusted. Based on the determined emotional state, a multi-dimensional anthropomorphic feedback is generated and executed; according to the change in the user's emotional state after receiving the feedback, the strategy for generating the feedback is iteratively optimized using a reinforcement learning algorithm.

10. The AI ​​toy system for emotional companionship based on multimodal perception and fusion algorithm according to claim 1, characterized in that: Calculate the confidence level of the current emotional state and determine whether the confidence level is lower than a preset threshold. If so, execute a preset emotional exploration action to actively guide the user to generate a clear emotional signal, and enhance the collection and analysis of the signal to verify the emotional state.