Children emotion pacifying intelligent doll

By collecting multi-dimensional data and recognizing emotions in real time, personalized execution command sequences are generated, which solves the problem of insufficient comforting of dolls when children experience sudden emotional changes, and achieves rapid response and synchronous comforting, thus improving the effect of children's emotional comforting.

CN121808466APending Publication Date: 2026-04-07XUZHOU UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing scenario-based emotional interactive companion dolls are unable to provide effective comfort when children experience sudden emotional shifts, lacking an instant intervention mechanism.

Method used

Employing a data acquisition module, an emotion recognition module, a personalized profiling module, and an adaptive decision-making module, the system generates personalized execution instruction sequences through multi-dimensional data acquisition and real-time emotion recognition. When emotions suddenly change, the current soothing action is interrupted, and a new strategy is quickly re-identified and executed.

Benefits of technology

It improves the speed of doll response to children's emotional shifts and the synchronicity of soothing behaviors, enhancing the speed and effectiveness of feedback to children's emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808466A_ABST
    Figure CN121808466A_ABST
Patent Text Reader

Abstract

The invention discloses a child emotion pacifying intelligent doll, and relates to the field of artificial intelligence. According to the child emotion pacifying intelligent doll, the emotion recognition module is used for processing multi-dimensional data, collected by the data collection module, of child users, the child emotion can be recognized with high accuracy, and the personalized portrait module defines different pacifying measures for the child users with different personalities; the self-adaptive decision module generates a highly targeted personalized execution instruction sequence according to the child emotion recognition result and the personalized portrait module, and the multi-modal execution module executes the personalized execution instruction sequence, so that the structure on the doll acts according to the personalized execution instruction sequence; the system is provided with a real-time response sub-module which is used for continuously monitoring emotional state vectors and defining emotional abrupt change criteria, and a trigger result is configured to interrupt a current pacifying action, start rapid re-recognition and execute a new strategy, so that the system can rapidly capture emotional turning of a child, and the feedback speed of the pacifying action of the doll is effectively increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to artificial intelligence technology, specifically to intelligent dolls for calming children's emotions. Background Technology

[0002] With increasing attention and deeper understanding of children's mental health in contemporary society, smart dolls are no longer merely entertainment toys, but are gradually evolving into important tools to assist children in recognizing, expressing, and regulating emotions. By integrating artificial intelligence and affective computing technologies, they provide children with a safe, private, and empathetic interactive partner, effectively compensating for the shortcomings of traditional educational methods in terms of immediacy, personalization, and stress-free intervention. These dolls can keenly capture children's emotional fluctuations and guide them to learn to manage emotions through personalized interaction strategies, such as soothing dialogue, gentle tactile feedback, and calming multi-sensory stimulation. This subtly cultivates their emotional resilience, making them indispensable digital emotional support partners in family and educational settings.

[0003] Chinese invention patent CN120179069B discloses a scene-based emotional interactive companion doll system and method, relating to the field of artificial intelligence technology. The system includes a data acquisition module, a feature analysis module, a profile construction module, an interaction decision module, an execution control module, and a database module. The data acquisition module collects interaction data and scene data; the feature analysis module uses a multimodal feature network to generate emotional feature vectors and scene feature vectors to determine the user's emotional state and scene type; the profile construction module constructs a user profile; the interaction decision module includes an emotional evolution unit, a scene interaction unit, and a physiological regulation unit, used to generate a sequence of execution instructions; and the execution control module controls the doll to complete emotional interaction behaviors. This invention, through scene perception and emotion analysis technology, enables the doll to accurately identify the user's emotional state and provide personalized interactive responses, offering intelligent emotional companionship services.

[0004] Existing scenario-based emotional interactive companion dolls require lengthy data processing and decision-making processes to execute soothing actions when in use. However, children's emotions are constantly changing and fluctuate greatly. For example, it is very common for children to experience sudden emotional shifts, such as from crying to laughing or from calm to irritability. Existing dolls lack effective instantaneous intervention mechanisms and are unable to provide effective soothing measures in the face of sudden emotional shifts in children. Summary of the Invention

[0005] The purpose of this invention is to provide a smart doll for calming children's emotions, in order to solve the problem that existing dolls are unable to effectively calm children's sudden emotional changes.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a child emotional comfort intelligent doll, the system comprising: The data acquisition module is used to collect interactive data and scene data to generate a dataset. It includes a core perception layer and an extended perception layer. The core perception layer consists of a tactile perception unit, a voice perception unit, and an environmental perception unit. The extended perception layer consists of a visual perception unit and a physiological perception unit. The emotion recognition module is used to identify children's emotions based on the dataset and output a structured emotion state vector. The personalized profile module includes a basic attribute layer, a reassurance feedback layer, an emotional pattern layer, and a physiological rhythm layer; The adaptive decision-making module includes a real-time response submodule and a personalized strategy generation submodule. The real-time response submodule is used to continuously monitor the emotional state vector and define the emotional change criterion. The trigger result is configured to interrupt the current soothing action, start rapid re-identification and execute a new strategy. The decision-making process of the personalized strategy generation submodule is configured to initially select candidate instructions from the strategy library based on the current emotional state vector and scene data, query the personalized profile module, select the instruction corresponding to the high-scoring strategy in the user soothing feedback layer, adjust the instruction parameters according to the data of the physiological rhythm layer, and finally output a personalized execution instruction sequence. The multimodal execution module includes a voice unit, a motion unit, and a light effect unit, which are used to control the actions of the doll's voice module, motion module, and light effect module. The database module includes a library of over 100 pre-stored soothing combination commands and a strategy library in the form of "audio, light, vibration".

[0007] Preferably, the database further includes: A children's emotional memory bank is used to store historical interaction data, emotion recognition results, and emotion evolution trends. Scene-Emotion Rule Base: Pre-stores children's life scenes and their typical emotion association rules; The children's profile database is used to store the personalized profile characteristics of each user; A physiological parameter library is used to store standard physiological parameters for children, for data standardization.

[0008] Preferably, in the data acquisition module, The tactile sensing unit consists of flexible pressure sensors distributed on the doll's torso and arms, used to collect pressure values, duration, pressure distribution, and waveform patterns. The voice perception unit consists of a microphone array and a noise reduction module integrated into the doll's head, used to collect voice streams and extract acoustic features; The environmental sensing unit consists of an environmental noise sensor and an ambient light sensor; The visual perception unit uses a camera integrated inside the doll to capture images of the user's facial micro-expressions. The physiological sensing unit uses a flexible photoplethysmography (PPG) sensor integrated inside the doll to collect heart rate data without being noticed.

[0009] Preferably, the emotion recognition module includes a feature extraction submodule, a dynamic weight fusion submodule, and an emotion output submodule. The feature extraction submodule is used to generate children's emotion representation features and children's scene representation features. The dynamic weight fusion submodule is used to adjust the weights of different data in the dataset of the emotion recognition model under different environments. The emotion output submodule processes the dataset through the emotion recognition model.

[0010] Preferably, the extraction method of the child's emotional representation features includes obtaining expression features by processing facial images through strong quantization CNN, obtaining fluctuation features by processing heart rate data through wavelet transform, extracting acoustic features by processing speech data through MFCC, and extracting temporal features by processing tactile data through LSTM. The multi-dimensional data is converted into a unified dimension, and then the unified dimension data is processed by the emotion fusion network to output the child's emotional representation features. The extraction method of the child scene representation features includes processing scene data through a scene encoder, encoding environmental parameters with a fully connected layer, encoding time features with an embedding layer, encoding scene type labels with one-hot encoding, and encoding guardian status with binary encoding, and then outputting the child scene representation features through a scene semantic understanding network.

[0011] Preferably, the operation steps of the dynamic weight fusion submodule include: The environmental noise decibel value and a binary vector used to identify the data of each extended perception layer sub-unit are input into the multi-factor weight decision function. The function outputs a weight vector containing all the perception modalities involved in the fusion, and the sum of all weight vectors is 1. The emotion output sub-module multiplies the original feature vector of each modality in the child's emotion representation features with its corresponding weight vector through the emotion recognition model to form a new feature vector. All the new feature vectors are concatenated to form a fused feature vector, and the emotion is classified through a classifier to output a structured emotion state vector.

[0012] Preferably, in the personalized profile module, the basic attribute layer includes age and gender, the soothing feedback layer records the historical effectiveness scores of different soothing strategies to form a preference list, the emotion pattern layer records historical emotion trigger points and the average duration of emotion calming, and the circadian rhythm layer learns the child's sensory sensitivity at different times of the day through long-term data.

[0013] Preferably, its usage includes: The data acquisition module collects interactive data and scene data in real time, generating a dataset that includes tactile data, voice data, environmental data, visual data, and physiological data, divided into a core perception layer and an extended perception layer. Feature extraction and emotion recognition: Extracting raw feature vectors from the dataset and generating emotion state vectors after processing; Personalized profile update and retrieval: The personalized profile module reads data from the database module, updates the basic attribute layer, comfort feedback layer, emotional pattern layer, and physiological rhythm layer of the child's personalized profile, and retrieves the current profile features; The system executes instructions adaptively. It continuously monitors the emotional state vector through the real-time response submodule. When the child's emotions do not change, the personalized strategy generation submodule selects candidate instructions from the strategy library based on the current emotional state vector and scene data. It queries the current profile features, selects the instruction corresponding to the high-scoring strategy in the user's soothing feedback layer, adjusts the instruction parameters based on the data of the physiological rhythm layer, and finally outputs a personalized execution instruction sequence. When a sudden change in emotion occurs, the current soothing action is interrupted, and a rapid re-identification and execution of a new strategy is initiated. Multimodal execution controls the doll's voice module, motion module, and light effect module to perform actions according to a personalized sequence of instructions; The data is updated by storing the emotional data, scene data, and execution effect of the commands in the emotional memory bank, and updating the child's profile.

[0014] Compared with existing technologies, the intelligent emotional soothing doll provided by this invention uses an emotion recognition module to process multi-dimensional data collected from child users by a data acquisition module. This allows for high accuracy in recognizing children's emotions. The personalized profiling module can define different emotional soothing measures for children with different personalities. The adaptive decision-making module generates a highly targeted personalized execution instruction sequence based on the child's emotion recognition results and the personalized profiling module. The multimodal execution module executes the personalized execution instruction sequence, causing the doll's structure to soothe the child according to the personalized execution instruction sequence. Furthermore, the system has a real-time response submodule to continuously monitor the emotional state vector and define emotional change criteria. The trigger result is configured to interrupt the current soothing action, initiate rapid re-recognition, and execute a new strategy. This allows the system to quickly capture the child's emotional shifts, transforming the soothing behavior from following to synchronizing, effectively enhancing the feedback speed of the doll's soothing actions. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0016] Figure 1 This is an overall structural block diagram provided for an embodiment of the present invention; Figure 2 The appearance diagram of the smart doll provided in the embodiment of the present invention; Figure descriptions: 1. Eyeball assembly; 2. Heart simulator vibrator. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0018] As attached Figure 1-2 As shown: Example:

[0019] This invention provides a smart emotional soothing doll for children. The system includes a data acquisition module, an emotion recognition module, a personalized profile module, an adaptive decision-making module, a multimodal execution module, and a database module. The emotion recognition module processes multi-dimensional data collected from child users by the data acquisition module, enabling high-accuracy identification of children's emotions. The personalized profile module defines different emotional soothing measures for children with different personalities. The adaptive decision-making module generates highly targeted personalized execution command sequences based on the child's emotion recognition results and the personalized profile module. The multimodal execution module executes these personalized execution command sequences, causing the doll's structures to move according to the personalized command sequences to soothe the child. Furthermore, the system includes a real-time response submodule to continuously monitor the emotional state vector and define emotional change criteria. Triggering the change involves interrupting the current soothing action, initiating rapid re-identification, and executing a new strategy. This allows the system to quickly capture the child's emotional shifts, transforming the soothing behavior from following to synchronizing, effectively enhancing the feedback speed of the doll's soothing actions.

[0020] The data acquisition module is used to collect interactive data and scene data to generate a dataset, which includes a core perception layer and an extended perception layer. The core perception layer consists of a tactile perception unit, a voice perception unit, and an environmental perception unit, while the extended perception layer consists of a visual perception unit and a physiological perception unit.

[0021] The tactile sensing unit consists of flexible pressure sensors distributed on the torso and arms of the doll, used to collect pressure values, duration, pressure distribution and waveform patterns. The sensor surface is covered with flexible foam to ensure a comfortable touch. It can collect the pressure value (range of 0-50N) of each sensing point in real time, calculate the average pressure and the maximum pressure, and distinguish different interactive intentions such as "hugging", "patting" and "grasping" by calculating the distribution center of gravity and contact area of ​​the pressure points. It can also analyze the temporal pattern of pressure changes, such as the different waveform characteristics of rapid patting and gentle stroking. The voice perception unit consists of a microphone array and a noise reduction module integrated into the doll's head. It is used to collect voice streams and extract acoustic features. The microphone array is integrated on both sides of the doll's head and uses beamforming technology for directional sound pickup, effectively suppressing environmental noise. The environmental perception unit consists of an environmental noise sensor and an ambient light sensor. The environmental noise sensor uses a digital microphone or a dedicated noise sensor to continuously monitor the A-weighted sound pressure level of the ambient background noise. The ambient light sensor senses the ambient light intensity to provide context for emotion recognition and light effect execution. The visual perception unit uses a camera integrated inside the doll to capture images of the user's facial micro-expressions. It employs a low-power, low-resolution CMOS camera and is equipped with a physical sliding cover. The system can only drive the micro-motor to open the sliding cover and initiate image acquisition after the guardian actively authorizes it via the APP. The physiological sensing unit uses a flexible photoplethysmography (PPG) sensor integrated inside the doll to collect heart rate data without being noticed. It employs a reflective PPG sensor, which is flexibly encapsulated in the doll's palm or the inside of its abdomen to ensure good contact with the child's skin.

[0022] The emotion recognition module is used to identify children's emotions based on the dataset and output a structured emotion state vector.

[0023] The emotion recognition module includes a feature extraction submodule, a dynamic weight fusion submodule, and an emotion output submodule. The feature extraction submodule is used to generate children's emotion representation features and children's scene representation features. The dynamic weight fusion submodule is used to adjust the weights of different data in the dataset of the emotion recognition model under different environments. The emotion output submodule processes the dataset through the emotion recognition model.

[0024] The extraction method of children's emotional representation features includes obtaining expression features by processing facial images through strong quantization CNN, obtaining fluctuation features by processing heart rate data through wavelet transform, extracting acoustic features by processing speech data through MFCC, and extracting temporal features by processing tactile data through LSTM. The multi-dimensional data is converted into a unified dimension, and then the unified dimension data is processed by the emotion fusion network to output the children's emotional representation features. The extraction method of the child scene representation features includes processing scene data through a scene encoder, encoding environmental parameters with a fully connected layer, encoding time features with an embedding layer, encoding scene type labels with one-hot encoding, and encoding guardian status with binary encoding, and then outputting the child scene representation features through a scene semantic understanding network.

[0025] The operation steps of the dynamic weight fusion submodule include: The environmental noise decibel value and a binary vector used to identify the data of each extended perception layer sub-unit are input into the multi-factor weight decision function. The function outputs a weight vector containing all the perception modalities involved in the fusion, and the sum of all weight vectors is 1. The emotion output sub-module multiplies the original feature vector of each modality in the child's emotion representation features with its corresponding weight vector through the emotion recognition model to form a new feature vector. All the new feature vectors are concatenated to form a fused feature vector, and the emotion is classified through a classifier to output a structured emotion state vector.

[0026] In this step, the multi-factor weight decision function can be divided into a basic mode and an enhanced mode. In the basic mode, when extended sensory data is unavailable, the tactile weight Wt is 0.5 + k * (NT), and the speech weight Wa is 1 - Wt, where k is the adjustment coefficient and T is the noise threshold. That is, in noisy environments, the proportion of tactile weight is automatically increased. In the enhanced mode, when physiological data is available, a higher fixed basic weight (e.g., 40%) is assigned to it, and the remaining 60% of the weight is distributed between tactile and speech according to the logic of the basic mode. When visual data is also available, a portion (e.g., 15% in total) is extracted from the tactile and speech weights and allocated to vision.

[0027] The personalized profile module includes a basic attribute layer, a soothing feedback layer, an emotion pattern layer, and a circadian rhythm layer. The basic attribute layer includes age and gender. The soothing feedback layer records the historical effectiveness scores of different soothing strategies, forming a preference list as a two-dimensional matrix. Rows represent different soothing strategy IDs, and columns represent different emotional scenarios. Each cell records the historical number of times the strategy was used in that scenario and its average effectiveness score. The effectiveness score is automatically calculated based on the degree of improvement in emotional state over a period of time after soothing (such as the rate of heart rate decrease and the magnitude of reduction in confidence in negative emotions). The emotion pattern layer records historical emotional trigger points and the average duration of emotional calming. Specifically, it records the frequency and average intensity of high-frequency scenario-emotion pairs, such as the frequency of guardian departure-sadness. It can even connect the presence of different guardians with emotional fluctuations and, through an emotion calming model based on historical data statistics, learn the average time required for the child to recover from different negative emotions to a calm state and the effective soothing methods. The circadian rhythm layer learns the child's sensory sensitivity at different times of the day through long-term data learning, constructing a sensory sensitivity curve for the child over a 24-hour period. For example, studies have shown that children’s tolerance to bright light and sudden movements is significantly reduced during the 8:00-9:00 PM (bedtime) period.

[0028] The adaptive decision-making module includes a real-time response submodule and a personalized strategy generation submodule. The real-time response submodule continuously monitors the emotion state vector and defines an emotion mutation criterion. The triggering result is configured to interrupt the current soothing action, initiate rapid re-identification, and execute a new strategy. This can be determined from two aspects: condition 1 is that the primary emotion label in the emotion state vector changes; condition 2 is that the confidence level of the new emotion jumps sharply from below 30% to above 60% within a very short time (e.g., Δt < 1.0s). At this time, a high-priority interruption mechanism is triggered. This mechanism runs on the highest priority thread of the system. Once the emotion mutation criterion is triggered, a hardware-level interrupt signal is immediately sent to the execution module, unconditionally stopping all current actions and calling the emotion recognition module for a one-time rapid re-identification. Then, control is transferred to the personalized strategy generation submodule. The decision-making process of the personalized strategy generation submodule is configured to initially select candidate instructions from the strategy library based on the current emotion state vector and scene data, query the personalized profile module, select the instruction corresponding to the high-scoring strategy in the user soothing feedback layer, adjust the instruction parameters according to the data of the physiological rhythm layer, and finally output a personalized execution instruction sequence.

[0029] The multimodal execution module includes a voice unit, a motion unit, and a light effect unit, which are used to control the actions of the doll's voice module, motion module, and light effect module.

[0030] The database module includes a strategy library with 100+ pre-stored soothing combination instructions in the form of "audio, light, vibration". The database also includes a children's emotion memory library for storing historical interaction data, emotion recognition results and emotion evolution trends, a scene-emotion rule library for pre-stored children's life scenarios and their typical emotion association rules, a children's profile library for storing personalized profile features of each user, a standard physiological parameter library for storing children, and a physiological parameter library for data standardization.

[0031] Its usage methods include: The data acquisition module collects interactive data and scene data in real time, generating a dataset that includes tactile data, voice data, environmental data, visual data, and physiological data, divided into a core perception layer and an extended perception layer. Feature extraction and emotion recognition: Extracting raw feature vectors from the dataset and generating emotion state vectors after processing; Personalized profile update and retrieval: The personalized profile module reads data from the database module, updates the basic attribute layer, comfort feedback layer, emotional pattern layer, and physiological rhythm layer of the child's personalized profile, and retrieves the current profile features; The system executes instructions adaptively. It continuously monitors the emotional state vector through the real-time response submodule. When the child's emotions do not change, the personalized strategy generation submodule selects candidate instructions from the strategy library based on the current emotional state vector and scene data. It queries the current profile features, selects the instruction corresponding to the high-scoring strategy in the user's soothing feedback layer, adjusts the instruction parameters based on the data of the physiological rhythm layer, and finally outputs a personalized execution instruction sequence. When a sudden change in emotion occurs, the current soothing action is interrupted, and a rapid re-identification and execution of a new strategy is initiated. Multimodal execution controls the doll's voice module, motion module, and light effect module to perform actions according to a personalized sequence of instructions; The data is updated by storing the emotional data, scene data, and execution effect of the commands in the emotional memory bank, and updating the child's profile.

[0032] physical reference Figure 2 As shown, the controllable LED light group and camera of the doll are integrated into the eye ball component 1 and set in the doll's eye area. The LED light group can be programmed to control the color and flashing frequency. For example, blue represents calm, red represents excitement, and yellow represents joy. Slow flashing, fast flashing, and breathing mode represent different functions. For example, using fast flashing represents data uploading.

[0033] A linear resonant actuator can be installed inside the chest cavity of the smart doll as a heart-simulating vibrator 2, which can accurately simulate the tactile sensation of heartbeats at different frequencies (such as 60-120 times / minute) and intensities (such as gentle and strong), enhancing the sense of comfort and immersion.

[0034] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A child's emotional comfort intelligent doll, characterized in that: Its system includes: The data acquisition module is used to collect interactive data and scene data to generate a dataset. It includes a core perception layer and an extended perception layer. The core perception layer consists of a tactile perception unit, a voice perception unit, and an environmental perception unit. The extended perception layer consists of a visual perception unit and a physiological perception unit. The emotion recognition module is used to identify children's emotions based on the dataset and output a structured emotion state vector. The personalized profile module includes a basic attribute layer, a reassurance feedback layer, an emotional pattern layer, and a physiological rhythm layer; The adaptive decision-making module includes a real-time response submodule and a personalized strategy generation submodule. The real-time response submodule is used to continuously monitor the emotional state vector and define the emotional change criterion. The trigger result is configured to interrupt the current soothing action, start rapid re-identification and execute a new strategy. The decision-making process of the personalized strategy generation submodule is configured to initially select candidate instructions from the strategy library based on the current emotional state vector and scene data, query the personalized profile module, select the instruction corresponding to the high-scoring strategy in the user soothing feedback layer, adjust the instruction parameters according to the data of the physiological rhythm layer, and finally output a personalized execution instruction sequence. The multimodal execution module includes a voice unit, a motion unit, and a light effect unit, which are used to control the actions of the doll's voice module, motion module, and light effect module. The database module includes a library of over 100 pre-stored soothing combination commands and a strategy library in the form of "audio, light, vibration".

2. The child emotional comfort intelligent doll according to claim 1, characterized in that, The database also includes: A children's emotional memory bank is used to store historical interaction data, emotion recognition results, and emotion evolution trends. Scene-Emotion Rule Base: Pre-stores children's life scenes and their typical emotion association rules; The children's profile database is used to store the personalized profile characteristics of each user; A physiological parameter library is used to store standard physiological parameters for children, for data standardization.

3. The child emotional comfort intelligent doll according to claim 1, characterized in that, In the data acquisition module The tactile sensing unit consists of flexible pressure sensors distributed on the doll's torso and arms, used to collect pressure values, duration, pressure distribution, and waveform patterns. The voice perception unit consists of a microphone array and a noise reduction module integrated into the doll's head, used to collect voice streams and extract acoustic features; The environmental sensing unit consists of an environmental noise sensor and an ambient light sensor; The visual perception unit uses a camera integrated inside the doll to capture images of the user's facial micro-expressions. The physiological sensing unit uses a flexible photoplethysmography (PPG) sensor integrated inside the doll to collect heart rate data without being noticed.

4. The child emotional comfort intelligent doll according to claim 1, characterized in that, The emotion recognition module includes a feature extraction submodule, a dynamic weight fusion submodule, and an emotion output submodule. The feature extraction submodule is used to generate children's emotion representation features and children's scene representation features. The dynamic weight fusion submodule is used to adjust the weights of different data in the dataset of the emotion recognition model under different environments. The emotion output submodule processes the dataset through the emotion recognition model.

5. The child emotional comfort intelligent doll according to claim 4, characterized in that, The extraction method of children's emotional representation features includes obtaining expression features by processing facial images through strong quantization CNN, obtaining fluctuation features by processing heart rate data through wavelet transform, extracting acoustic features by processing speech data through MFCC, and extracting temporal features by processing tactile data through LSTM. The multi-dimensional data is converted into a unified dimension, and then the unified dimension data is processed by the emotion fusion network to output the children's emotional representation features. The extraction method of the child scene representation features includes processing scene data through a scene encoder, encoding environmental parameters with a fully connected layer, encoding time features with an embedding layer, encoding scene type labels with one-hot encoding, and encoding guardian status with binary encoding, and then outputting the child scene representation features through a scene semantic understanding network.

6. The child emotional comfort intelligent doll according to claim 4, characterized in that, The operation steps of the dynamic weight fusion submodule include: The environmental noise decibel value and a binary vector used to identify the data of each extended perception layer sub-unit are input into the multi-factor weight decision function. The function outputs a weight vector containing all the perception modalities involved in the fusion, and the sum of all weight vectors is 1. The emotion output sub-module multiplies the original feature vector of each modality in the child's emotion representation features with its corresponding weight vector through the emotion recognition model to form a new feature vector. All the new feature vectors are concatenated to form a fused feature vector, and the emotion is classified through a classifier to output a structured emotion state vector.

7. The child emotional comfort intelligent doll according to claim 1, characterized in that, In the personalized profile module, the basic attribute layer includes age and gender, the soothing feedback layer records the historical effectiveness scores of different soothing strategies to form a preference list, the emotion pattern layer records historical emotion trigger points and the average duration of emotion calming, and the circadian rhythm layer learns the child's sensory sensitivity at different times of the day through long-term data.

8. The child emotional comfort intelligent doll according to claim 1, characterized in that, Its usage methods include: The data acquisition module collects interactive data and scene data in real time, generating a dataset that includes tactile data, voice data, environmental data, visual data, and physiological data, divided into a core perception layer and an extended perception layer. Feature extraction and emotion recognition: Extracting raw feature vectors from the dataset and generating emotion state vectors after processing; Personalized profile update and retrieval: The personalized profile module reads data from the database module, updates the basic attribute layer, comfort feedback layer, emotional pattern layer, and physiological rhythm layer of the child's personalized profile, and retrieves the current profile features; The system executes instructions adaptively. It continuously monitors the emotional state vector through the real-time response submodule. When the child's emotions do not change, the personalized strategy generation submodule selects candidate instructions from the strategy library based on the current emotional state vector and scene data. It queries the current profile features, selects the instruction corresponding to the high-scoring strategy in the user's soothing feedback layer, adjusts the instruction parameters based on the data of the physiological rhythm layer, and finally outputs a personalized execution instruction sequence. When a sudden change in emotion occurs, the current soothing action is interrupted, and a rapid re-identification and execution of a new strategy is initiated. Multimodal execution controls the doll's voice module, motion module, and light effect module to perform actions according to a personalized sequence of instructions; The data is updated by storing the emotional data, scene data, and execution effect of the commands in the emotional memory bank, and updating the child's profile.

Citation Information

Patent Citations

  • A scene-based emotional interactive companion doll system and method

    CN120179069B