Child emotion recognition and guidance interaction method based on multi-modal emotion computing
By combining multimodal perception and psychological schema mapping models, the problems of non-stationarity and context dependence in children's emotion recognition are solved, achieving high accuracy and personalized emotion guidance in complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SIMAI INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-16
AI Technical Summary
Existing technologies for children's emotion recognition ignore the non-stationarity and context-dependent nature of children's emotional expressions, resulting in decreased recognition accuracy in low signal-to-noise ratio scenarios, inability to accurately decode highly ambiguous deep psychological states, and inability of interactive systems to provide effective feedback.
By simultaneously acquiring facial video streams, head posture trajectories, and upper body behavior sequences through multimodal perception units, and combining optical flow field information and psychological schema mapping models, a physiological-psychological dual-layer mapping network is constructed to identify facial behavior units and generate contextual feature vectors, providing personalized emotion guidance strategies.
It can stably capture repressed micro-expressions under non-ideal lighting and complex scenes, accurately decode psychological motivations, improve the robustness and depth of emotion recognition, and realize personalized and precise emotion guidance.
Smart Images

Figure CN122208142A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and behavioral feature recognition technology, specifically involving a method for children's emotion recognition and guidance interaction based on multimodal emotion computing. Background Technology
[0002] With the deep integration of artificial intelligence and affective computing technologies, intelligent interactive systems are increasingly being applied in children's education, psychological assessment, and mental and physical health monitoring. Affective computing aims to construct an intelligent interactive environment with empathy by simulating the perception, recognition, and understanding processes related to human emotions through computers. During critical stages of children's growth and development, timely capture and understanding of their emotional fluctuations plays a crucial role in optimizing educational strategies and preventing psychological deviations. This requires systems capable of acquiring children's physiological characteristics and behavioral performance in real time and performing in-depth logical analysis.
[0003] Children's emotion recognition and interaction technology based on multimodal data has become a core research direction. This technology aims to integrate facial muscle movements, behavioral trajectories, and developmental psychology principles to establish a mapping model from underlying physical representations to higher-level psychological intentions. By introducing prior knowledge of children's emotional development, the system can more objectively assess children's cognitive state and emotional feedback in specific interactive situations, and achieve precise emotional guidance and behavioral direction by dynamically adjusting the interaction logic.
[0004] Current technologies for processing children's emotions often directly apply adult emotion models, ignoring the unique non-stationarity and context-dependent nature of children's emotional expression. Traditional visual feature extraction methods struggle to capture subtle transient facial muscle movements, leading to a significant drop in recognition accuracy in low signal-to-noise ratio scenarios such as when children are active or in uneven lighting. AI models that rely solely on data-driven approaches lack deep integration of psychological schemas, failing to distinguish between exaggerated and concealed elements in children's expressions and struggling to accurately decode highly ambiguous deep psychological states such as shame and repression. This results in interactive systems being unable to provide effective feedback for complex psychological motivations. Therefore, a multimodal affective computing-based interactive method for children's emotion recognition and guidance is desired. Summary of the Invention
[0005] The purpose of this invention is to provide a method for children's emotion recognition and guidance based on multimodal emotion computing, which can solve the problems in the background art mentioned above.
[0006] To achieve the above objectives, the technical solution adopted by this invention is: a children's emotion recognition and guidance interaction method based on multimodal emotion computing, comprising the following specific steps: Step 1: Simultaneously collect facial video streams, head posture trajectories, and upper body behavior sequences of children in natural interactive situations using a multimodal sensing unit; Step 2: Preprocess the facial video stream to eliminate image distortion caused by changes in lighting or motion blur, and extract optical flow field information between consecutive frames to characterize the transient micro-tremor features of facial muscles; Step 3: Identify facial behavior units based on the facial behavior coding system, and construct a dynamic expression time sequence map by combining the optical flow field information, which is used to characterize the speed, amplitude and direction characteristics of expression evolution; Step 4: Input the dynamic facial expression time sequence map into the psychological schema mapping model. This model has an embedded database of emotional development milestones divided according to the child's age stage, which is used to match the potential psychological state category corresponding to the current facial expression combination. Step 5: Integrate head posture trajectory and upper body behavior sequence to generate contextual feature vector, and use it as an auxiliary criterion in the final determination of psychological state to enhance the ability to identify camouflaged or exaggerated expressions. Step 6: Generate an appropriate set of psychological counseling strategy instructions based on the judgment results, and output interactive feedback with empathetic tone and anthropomorphic body language through the voice synthesis module and the virtual character animation module. Step 7: Monitor children's behavioral responses to interactive feedback in real time, dynamically adjust the intensity and pace of subsequent guidance strategies, and form a closed-loop emotion regulation mechanism.
[0007] Preferably, the multimodal perception unit in step 1 consists of a high frame rate visible light camera, a depth sensor, and an inertial measurement unit. The high frame rate visible light camera is used to capture a facial video stream of no less than 60 frames per second, the depth sensor is used to obtain the three-dimensional spatial position of the head, and the inertial measurement unit is used to record the changes in the angular velocity and acceleration of the upper body.
[0008] Preferably, the preprocessing of the facial video stream in step 2 includes adaptive histogram equalization, motion deblurring filtering, and face region cropping. The motion deblurring filtering uses a blind deconvolution algorithm based on gradient sparse prior, which preserves edge details while suppressing image motion blur caused by rapid head rotation.
[0009] Preferably, the facial action unit identification in step 3 is based on a joint architecture of local binary pattern and temporal convolutional neural network, which can detect at least 14 basic facial action units, and combined with the direction consistency and energy concentration index of optical flow field, to screen out micro-expression segments with emotional indication significance.
[0010] Preferably, in step 4, the psychological schema mapping model adopts a hierarchical Bayesian inference framework. Its prior distribution parameters are divided into multiple developmental stage intervals according to the child's age. Each interval corresponds to a set of typical emotional expression patterns and their common confusion situations, thereby achieving accurate classification of complex psychological states such as shame, repression, and anxiety at the high-level semantic level.
[0011] Preferably, in step 5, the head posture trajectory is obtained by fitting the head central axis with the point cloud data output by the depth sensor, and the upper body behavior sequence is generated by continuously tracking the key points of the shoulders and elbows. The contextual feature vector composed of the two includes three core dimensions: avoidance tendency index, body tension score, and interaction willingness level.
[0012] Preferably, in step 6, the psychological counseling strategy instruction set is automatically matched with a preset interactive script library based on the psychological state category. The interactive script library covers four basic strategies: soothing, guiding, diverting, and motivating. Each strategy is configured with corresponding voice tone parameters, vocabulary selection rules, and virtual character expression templates.
[0013] Preferably, in step 7, behavioral response monitoring is achieved by comparing the changes in the frequency of facial micro-expressions, head orientation stability, and upper body movement amplitude of the child before and after intervention. If the monitoring results show that the emotional fluctuation tends to be stable, the intensity of intervention is reduced; if the emotional reaction intensifies, a higher level of intervention strategy is switched.
[0014] Compared with the prior art, the present invention has the following beneficial effects: This invention overcomes the technical challenges of high ambiguity and low signal-to-noise ratio in children's emotional expression by constructing a physiological-psychological dual-layer mapping network that deeply integrates low-level visual computing with high-level psychological schemas. Even under non-ideal lighting conditions and in complex scenarios involving frequent child movement, it can stably capture suppressed micro-expressions and accurately decode their underlying psychological motivations, thus improving the robustness and semantic depth of emotion recognition.
[0015] By introducing contextual features and a closed-loop adjustment mechanism, the interactive system is made dynamically adaptable, enabling it to provide personalized and precise emotional guidance services for children at different developmental stages, thus achieving a leap from single facial expression recognition to deep psychological state understanding. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the psychological schema mapping model based on the hierarchical Bayesian inference framework and the division of children's age development stages in this invention. Figure 3 This is a logical flowchart of the synchronous acquisition of multimodal perception data and the construction of dynamic facial expression time sequence map in this invention; Figure 4 This is a flowchart illustrating the logical process of recognizing masquerade expressions and determining psychological states in this invention by combining contextual feature vectors. Figure 5This is a schematic diagram of the multi-level interaction relationship and data flow of the real-time monitoring of children's behavioral responses and the closed-loop emotion regulation mechanism in this invention. Detailed Implementation
[0017] Example 1: Please refer to the appendix Figure 1 To be continued Figure 5 To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.
[0018] In the interactive method for children's emotion recognition and guidance based on multimodal emotion computing provided by the present invention, the method realizes accurate recognition and closed-loop guidance of children's complex psychological states by constructing a physiological-psychological dual-layer mapping network.
[0019] In step 1, a multimodal sensing unit synchronously acquires facial video streams, head posture trajectories, and upper body behavior sequences of the child in a natural interactive context. In practice, the multimodal sensing unit is positioned in front of the interactive terminal or in its surrounding environment to ensure coverage of the child's effective activity area. The multimodal sensing unit consists of a high frame rate visible light camera, a depth sensor, and an inertial measurement unit.
[0020] A high-frame-rate visible light camera is configured to capture facial video streams at a rate of at least 60 frames per second. This frame rate is designed to ensure the complete recording of extremely short-lived micro-expression changes, such as facial muscle twitches lasting between 1 / 25 and 1 / 5 of a second. A depth sensor, utilizing structured light or time-of-flight principles, acquires the absolute position of the child's head in three-dimensional space in real time, achieving millimeter-level spatial resolution. An inertial measurement unit, through its built-in three-axis accelerometer and three-axis gyroscope, records the angular velocity and acceleration changes of the child's upper body at a sampling frequency of at least 100 Hz, used to sense minute movements such as body swaying, leaning forward, or leaning backward during interaction.
[0021] In step 1 above, the data acquisition process of the multimodal sensing unit involves a strict time synchronization mechanism. The system uses a global hardware trigger signal to ensure that the shutter opening time of the visible light camera, the infrared projection time of the depth sensor, and the data reading time of the inertial measurement unit are within a synchronization deviation range of microseconds. The acquired facial video stream data is encapsulated into a raw pixel matrix, the head pose data is processed into a three-dimensional coordinate sequence, and the upper body behavior data is recorded as a set of six-degree-of-freedom motion vectors.
[0022] In step 2, the facial video stream is preprocessed to eliminate image distortion caused by changes in lighting or motion blur, and optical flow field information between consecutive frames is extracted to characterize the transient micro-tremor features of facial muscles. The specific preprocessing process includes performing adaptive histogram equalization. This adaptive histogram equalization dynamically adjusts the contrast of each pixel by calculating the brightness distribution histogram of local image regions, ensuring that details of facial feature areas remain clearly discernible even in strong light overexposure or weak light shadow environments. To address motion blur caused by frequent head movements during interaction by children, the system employs a blind deconvolution algorithm based on gradient sparse priors for motion deblurring filtering.
[0023] In this blind deconvolution algorithm, it is assumed that the blurred image is generated by convolving a sharp image with an unknown point spread function. The system iteratively optimizes the process to simultaneously estimate the point spread function and recover the sharp image while maintaining the sparsity constraint of the image edge gradients. This process does not rely on pre-known blur kernel information and can effectively suppress image motion blur caused by rapid movement. A face detection operator is used to crop the face region of the processed image and scale it to a uniform pixel size to eliminate background noise interference with subsequent feature extraction.
[0024] In step 2, when extracting optical flow information, the system employs a dense optical flow estimation algorithm. This algorithm calculates the displacement components of each pixel in the horizontal and vertical directions by comparing the brightness constant assumptions of corresponding pixels in two consecutive frames. The extracted optical flow information reflects the motion vector distribution of the facial epidermis in a very short time. These vectors can characterize the transient micro-tremor features of facial muscles, such as eye twitching, nasal flaring, or lip tremors.
[0025] In step 3, facial action units are identified based on the facial behavior coding system, and a dynamic expression time-series map is constructed by combining the optical flow field information to characterize the speed, amplitude, and direction characteristics of expression evolution. The identification of facial action units is based on a joint architecture of local binary mode and temporal convolutional neural network. Local binary mode is used to extract subtle spatial features of facial skin texture. By comparing the size relationship between the central pixel and neighboring pixels, it encodes local regions into binary codes, exhibiting high robustness to illumination fluctuations. The temporal convolutional neural network specifically processes the temporal dimension information of the video stream, extracting the dynamic evolution patterns of facial movements by sliding multiple layers of one-dimensional convolutional kernels along the time axis. This architecture can detect at least 14 basic facial action units, including but not limited to corrugator supercilii contraction, levator labii superioris movement, and anguli oris movement.
[0026] Step 3 combines the directional consistency and energy concentration indices of the optical flow field to screen out micro-expression fragments with emotional indicative significance. Optical flow directional consistency is measured by calculating the mean cosine similarity of motion vectors of all pixels within a local region. If this mean is greater than a preset threshold of 0.85, the region is considered to have coordinated facial movements. Energy concentration is calculated by accumulating the sum of squares of motion vector amplitudes within a time window. When a pulse-like increase in energy within a very short time with highly consistent direction is detected in a region, the system identifies it as a micro-expression fragment. The resulting dynamic expression time-series map is a multi-dimensional data structure containing time, spatial location, and motion intensity, capable of accurately quantifying the speed, amplitude, and direction of motion of an expression from its emergence, peak, to its decline.
[0027] In step 4, the dynamic facial expression time series map is input into a psychological schema mapping model. This model embeds a database of emotional development milestones categorized according to children's age stages to match the potential psychological state category corresponding to the current facial expression combination. The psychological schema mapping model employs a hierarchical Bayesian inference framework. This framework decomposes psychological state determination into two levels: the bottom level is the observation layer, corresponding to the visual features in the dynamic facial expression time series map; the top level is the state layer, corresponding to the abstract psychological state category.
[0028] In step 4, the prior distribution parameters are the core of the model, which are divided into multiple developmental stage intervals based on the child's age. For example, for preschool children aged 3 to 5, the prior distribution parameters assign higher weights to exaggerated expressions because children at this stage tend to attract attention through facial movements; while for children aged 7 to 12, the prior distribution parameters increase the weights for repressive and concealing characteristics. The emotional development milestone database contains typical emotional expression patterns of children of different ages in specific situations.
[0029] Each interval corresponds to a set of typical emotional expression patterns and their common confounding scenarios. For example, when the system detects a combination of frowning and eye avoidance, the hierarchical Bayesian inference framework calculates the posterior probability of it belonging to anger, shame, or anxiety based on the prior distribution. If the child is in the early stages of schooling, the model will combine data from the milestone database to increase the prior probability of shame, thereby achieving accurate classification of complex psychological states such as shame, repression, and anxiety at a high-level semantic level.
[0030] In step 5, the head posture trajectory and upper body behavior sequence are fused to generate a contextual feature vector, which is used as an auxiliary criterion in the final determination of psychological state to enhance the ability to identify masquerading or exaggerated expressions. The head posture trajectory is obtained by fitting the head's central axis to point cloud data output by a depth sensor. The system performs spatial transformation on the point cloud of the face region, uses the least squares method to fit a central vector perpendicular to the facial plane, and obtains the temporal curves of yaw, pitch, and roll angles by tracking the rotation angle changes of this vector in three-dimensional space. The upper body behavior sequence is generated by continuously tracking key points of the shoulders and elbows. The system uses a human posture estimation algorithm to identify the two-dimensional or three-axis coordinates of the left shoulder, right shoulder, left elbow, and right elbow, and calculates the continuity of their motion trajectories.
[0031] In step 5 above, the generated contextual feature vector includes three core dimensions: avoidance tendency index, body tension score, and interaction willingness level. The avoidance tendency index is determined by the angle and frequency at which the head turns away from the interaction center; the body tension score is measured by calculating the elevation and shaking frequency of the shoulder muscles; and the interaction willingness level is assessed based on the rate of change in the physical distance between the child's upper body and the screen. When the facial expression identified in step 4 shows happy characteristics, but the avoidance tendency index calculated in step 5 is high and the body tension score exceeds a preset threshold, the system will correct the final psychological state judgment from pleasant to concealed anxiety. This multimodal data fusion judgment solves the problem of children deliberately performing or intentionally concealing their true emotions when facing adults or machines.
[0032] In step 6, a set of appropriate psychological counseling strategy instructions is generated based on the judgment result, and interactive feedback with empathetic tone and anthropomorphic body language is output in collaboration between the speech synthesis module and the virtual character animation module. The set of psychological counseling strategy instructions automatically matches a preset interactive script library according to the psychological state category. The interactive script library is organized into a tree-like logical structure, covering four basic strategies: reassurance, guidance, diversion, and motivation.
[0033] In step 6, when the assessment result is anxiety, the system automatically invokes a soothing strategy. Each strategy is configured with corresponding voice tone parameters, vocabulary selection rules, and virtual character expression templates. Upon receiving a soothing instruction, the speech synthesis module lowers the fundamental frequency by 15% to 20% and increases the smoothness of the speech rate, producing a gentle and patient listening experience. The vocabulary selection rules prioritize selecting words from a lexicon with positive and emotionally supportive meanings. The virtual character animation module controls the virtual image on the screen to perform specific expressions, such as slightly squinting and raising the corners of the mouth, along with listening-like body animations, such as slightly tilting the head or leaning forward. This collaborative feedback provides children with a safe and trustworthy interactive environment through dual visual and auditory empathy.
[0034] In step 7, the child's behavioral response to interactive feedback is monitored in real time, and the intensity and pace of subsequent guidance strategies are dynamically adjusted to form a closed-loop emotion regulation mechanism. This behavioral response monitoring is achieved by comparing the changes in the frequency of facial micro-expressions, head orientation stability, and upper body movement amplitude before and after guidance.
[0035] In the specific execution logic of step 7, the system establishes a time sliding window to continuously monitor changes in various physiological indicators of the child within 10 to 30 seconds after feedback output. If the monitoring results show that the frequency of facial micro-expressions decreases by more than 30%, and the stability of the head's orientation towards the center increases, and the upper body movement returns to the normal range, it is determined that the emotional fluctuation is stabilizing, and the system reduces the intensity of guidance and switches to a regular interaction mode. Conversely, if the system detects a higher frequency of frowning or mouth closure, an increased frequency of head turning away, or a large backward movement of the upper body, it is determined that the emotional reaction is intensifying. The closed-loop regulation mechanism will immediately intervene and switch to a higher-level intervention strategy, such as switching from verbal guidance to playing soothing music or displaying more engaging virtual animations, to forcibly interrupt the current negative emotional chain.
[0036] To further refine the technical implementation of this invention, at the hardware support level, the interactive terminal is equipped with a high-performance embedded processor specifically responsible for running the complex computer vision algorithms and Bayesian inference models described above. The embedded processor is connected to the multimodal perception unit via a dedicated bus to ensure that the data transmission bandwidth meets real-time requirements. At the software architecture level, the system adopts a hierarchical storage structure, storing the real-time acquired video stream in a high-speed random access memory for caching, while storing the emotion development milestone database and interaction script library in non-volatile storage media for rapid retrieval.
[0037] In the motion deblurring filtering of step 2, the iterative process of the blind deconvolution algorithm is limited to a fixed number of iterations, such as 20 iterations, or it stops when the change in mutual information of the image gradients generated by two adjacent iterations is less than a certain minimum value, in order to achieve a balance between computational efficiency and deblurring quality. For face region cropping, the system uses a cascaded classifier based on Haar features for initial localization, and then uses 68 facial key point detection to accurately locate the facial features, ensuring that the center of the cropped image is always aligned with the tip of the child's nose.
[0038] In the local binary pattern extraction of step 3, rotation-invariant equivalent pattern encoding is used to convert the eight-neighborhood relationship of each pixel into an eight-bit binary number. The temporal convolutional neural network employs a residual connection structure to prevent gradient vanishing during deep feature extraction. The 14 detected action units include: 1. Medial frontalis muscle elevation, 2. Lateral frontalis muscle elevation, 4. Corrugator supercilii muscle depression, 5. Levator palpebrae superioris muscle elevation, 7. Orbicularis oculi muscle contraction, 10. Levator labii superioris muscle elevation, 12. Zygomaticus major muscle pulling the corner of the mouth, 15. Depressor anguli oris muscle depressing the corner of the mouth, 17. Menti muscle elevation, 20. Smileis muscle pulling the corner of the mouth, 23. Orbicularis oris muscle tightening, 24. Orbicularis oris muscle tightening, 25. Lip separation, and 26. Mandibular descent. The intensity of each action unit is quantized into a continuous value from 0 to 5, which, together with the dynamic features of the optical flow field, forms a temporal feature vector.
[0039] In the hierarchical Bayesian inference of step 4, the model output is a probability distribution vector, representing the probability that the current state belongs to each psychological category. If the difference between the highest probability value and the second highest probability value is less than a preset threshold of 0.15, the system will initiate a secondary confirmation logic, that is, retrieve the historical states within the previous 5 seconds and perform a weighted average to improve the stability of the judgment. The age development stages are specifically divided into: the first stage is under 3 years old, focusing on emotions caused by basic physiological needs; the second stage is 3 to 6 years old, focusing on the emergence of self-awareness and social frustration emotions; the third stage is 7 to 12 years old, focusing on hidden emotions caused by academic pressure and complex interpersonal relationships.
[0040] In the feature fusion step 5, the contextual feature vector is not merely a static set of numerical values, but a dynamic trajectory with timestamps. The avoidance tendency index is calculated by combining the velocity and acceleration of head yaw angles; if a rapid head dodge is detected, the index spikes instantaneously. The body tension score incorporates spectral analysis, extracting micro-tremor components of shoulder key points in the 8-12 Hz frequency range using Fast Fourier Transform, which is typically highly correlated with internal anxiety or fear. The assessment of the willingness to interact level also references the body's midline tilt given by depth sensors; a backward tilt is considered a defensive posture, while a forward tilt is seen as active participation.
[0041] In the speech synthesis step 6, in addition to adjusting the fundamental frequency, prosodic control technology is also introduced. When outputting guided strategies, the system adds pauses before and after keywords and increases the rising slope of the sentence-end intonation to stimulate children's thinking and response. The virtual character animation module uses a skeletal skinning algorithm to achieve smooth limb movement switching by calculating vertex weights, ensuring that the virtual character's every move conforms to human empathic logic and avoiding resistance caused by stiff movements.
[0042] In the closed-loop adjustment of step 7, the strategy strength is adjusted following a stepwise algorithm of increasing or decreasing. Each adjustment records the current control parameters in the metadata and performs correlation analysis with subsequent behavioral feedback. In this way, the system can continuously optimize the mapping relationship between strategy and feedback, achieving personalized adaptation during long-term interaction.
[0043] Example 2: Based on the children's emotion recognition and guidance interaction method based on multimodal emotion computing described in Example 1 above, this example provides an alternative implementation scheme for a specific teaching assistance scenario.
[0044] In the above method, the data source acquired in step 1 is further expanded. In addition to facial video streams, head posture, and upper body behavior, the system also acquires the sequence of touch intensity during the child's operation via an external pressure-sensitive interactive whiteboard. A high frame rate visible light camera is used not only to capture facial information but is also configured to recognize the trajectory of hand movements.
[0045] In step 2, color constancy correction was added to the preprocessing to address the specific lighting conditions in the teaching environment. This operation adjusts the red, green, and blue channel balance of the entire image by calculating the gain coefficient of known white objects in the image, ensuring that facial skin color remains consistent under sunlight, fluorescent light, or mixed light sources. This is crucial for accurately extracting optical flow information, as color fluctuations may be misinterpreted by the algorithm as pixel displacement.
[0046] In step 3, the recognition of facial behavior units is not limited to the 14 types mentioned above, but also includes a combination of AU units specifically for attention detection, such as changes in the size of the eyelid slit. The system assesses a child's fatigue or excitement level when facing a teaching task by monitoring subtle contractions of the levator palpebrae superioris muscle. Frequency domain features have been added to the dynamic expression time sequence map; by performing discrete cosine transform on the sequence of muscle movement amplitudes, low-frequency components representing emotional stability and high-frequency components representing sudden disturbances are extracted.
[0047] In step 4, the emotional development milestone database of the psychological schema mapping model is supplemented with dimensions related to cognitive load. The hierarchical Bayesian inference framework introduces situational variables as latent variables, such as the difficulty level of the current teaching task. When children face high-difficulty tasks and exhibit frequent frowning, the model classifies their psychological state as "frustration due to cognitive overload," rather than simple anger as in general scenarios.
[0048] In step 5, the contextual feature vector incorporates the fusion of touch intensity. By analyzing changes in force when children write or click on the tablet, the system assesses their internal fluctuations. Research has found that touch intensity typically increases erratically when children experience anxiety or frustration. The avoidance tendency index here includes not only head deflection but also the proportion of time the hand is away from the operating area.
[0049] In step 6, the psychological guidance strategy instruction set is refined into a teaching guidance instruction set. Interactive feedback is no longer merely emotional reassurance, but also includes cognitive guidance. For example, when frustration is detected, the vocabulary rules output by the speech synthesis module trigger a "task breakdown" suggestion, guiding children to break down complex learning tasks into simpler sub-tasks. The virtual character animation module displays encouraging gestures, such as a thumbs-up, to enhance children's self-efficacy.
[0050] In step 7, the adjustment target of the closed-loop emotion regulation mechanism is set as the "optimal learning emotion range". If the system detects that a child is in a state of over-excitement or over-depression, it will adjust the difficulty of subsequent teaching content and the frequency of interactive feedback to bring the child's emotional state back to a stable range conducive to knowledge absorption.
[0051] This embodiment further enhances the breadth and depth of application of the method of the present invention by expanding and refining multimodal data in specific scenarios.
[0052] Example 3: Based on the methods described in the above examples, this example describes the system implementation details based on a collaborative architecture of edge computing and cloud computing.
[0053] In the above method, the processing steps 1 to 3 are configured to be executed in the local edge computing module of the interactive terminal. This configuration ensures that high-bandwidth facial video streams can be processed locally in real time, avoiding network latency issues caused by transmitting large amounts of raw image data to the cloud. The local edge computing module uses a dedicated neural network accelerator, which can complete the preprocessing and motion unit extraction of a single frame image within 10 milliseconds.
[0054] In step 4, the computational task of the psychological schema mapping model is undertaken by a cloud server cluster. The edge computing module compresses and packages the extracted facial behavior unit feature vectors, optical flow field statistical features, and the initially generated time series map, and transmits them to the cloud via an encrypted communication protocol. The cloud server utilizes its powerful computing capabilities to run a complex hierarchical Bayesian inference algorithm and calls upon a large-scale emotional development milestone dataset stored in a distributed database for matching in real time.
[0055] In step 5, the generation of contextual feature vectors employs a distributed fusion strategy. Head pose trajectories are calculated at the local edge, while the deep correlation analysis between upper body behavior sequences and contextual variables is performed in the cloud. The cloud server uses global historical interaction data to calibrate the current child's feature vector.
[0056] In step 6, the interactive script library is stored in the cloud and supports dynamic updates. Whenever psychological research yields new guidance strategies, the system can seamlessly push the new strategies to the interactive instruction set. The speech synthesis module employs deep learning-based parametric synthesis technology, which can generate empathic speech streams with specific timbres in real time in the cloud according to the child's personalized preferences, and download them to the terminal for playback in an audio compression format.
[0057] In step 7, the control parameters of the closed-loop regulation mechanism are recorded in the individual's growth profile in the cloud. Through long-term monitoring and feedback data accumulation, the cloud system can use reinforcement learning algorithms to continuously optimize the guidance parameter settings for specific children. For example, if it learns that a child is more sensitive to musical feedback than verbal feedback, it can increase the priority of musical intervention in subsequent regulation strategies.
[0058] The architecture described in this embodiment ensures real-time emotion recognition while making full use of cloud storage and computing power to achieve precise and personalized emotion guidance services for a large user group.
[0059] In any of the above embodiments, the characterization of transient micro-tremors of facial muscles is achieved by establishing a high-dimensional feature space. Each point in this space represents the set of displacement vectors of all valid sampling points on the face at a specific moment. During the extraction of optical flow field information, to address non-ideal lighting conditions, the system also introduces a brightness compensation operation based on the Lambertian volume assumption in step 2. This operation estimates the normal vector distribution of the facial surface and performs brightness-weighted compensation on the shadow areas caused by the offset of the light source position, ensuring the accuracy of the optical flow vector calculation.
[0060] In step 3, the process of constructing the dynamic facial expression temporal map involves the temporal alignment of facial action units. Since different expressions have varying durations, the system employs a dynamic temporal warping algorithm to map the captured action sequences onto a standard-length time axis. This allows the system to compare the statistical differences in speed and amplitude between expressions at different times and for different durations. The energy concentration index is calculated by integrating the square of the optical flow vector amplitude within a preset 200-millisecond time window and dividing by the average spatial area within that window.
[0061] In step 4, during the inference process of the psychological schema mapping model, the likelihood function in the hierarchical Bayesian framework is constructed as a Gaussian mixture model. This model is used to describe the probability density distribution of various combinations of facial action units under a given psychological state. The prior distribution parameters are updated using an online learning mechanism. Every fixed period (e.g., one month), the system fine-tunes the prior parameters using maximum a posteriori probability estimation based on the child's actual performance and feedback results, to adapt to changes in the child's emotional expression habits as they grow older.
[0062] In step 5, when generating the contextual feature vector, the calculation of the avoidance tendency index not only considers the physical head rotation angle but also incorporates eye-tracking data. If the system detects that the child's pupil center deviates from the center of the screen for more than 60% of the time, the weighted score of the avoidance tendency index is increased. In the body tension score, the tracking of key points on the shoulders and elbows uses a Kalman filter-based prediction algorithm to infer the continuous movement trend of the limbs even when light is obstructed or parts of the body are out of frame.
[0063] In step 6, the intonation parameter control of the speech synthesis module includes four dimensions: fundamental frequency mean, fundamental frequency dynamic range, speech intensity, and speech rate. A soothing strategy narrows the fundamental frequency dynamic range, making the intonation sound flat; while an encouraging strategy expands the fundamental frequency dynamic range and increases speech intensity, producing a high-spirited and invigorating effect. The facial expression templates of the virtual character animation module correspond one-to-one with the Facial Behavior Coding System (FACS), ensuring that the virtual character's expressions have biologically plausible rationality and can induce empathic responses from mirror neurons in children.
[0064] In the monitoring logic of step 7, the specific method for comparing the trend of change before and after intervention is to calculate the rate of change of the Euclidean distance between the feature vectors in two adjacent time periods. If the rate of change of this distance shows a trend of moving closer to the "neutral emotion template," the intervention is considered effective. If the rate of change is positive and the absolute value increases, the emotion is determined to have deviated from the normal range. The hierarchical switching of intervention strategies follows a logically rigorous decision tree: level one is fine-tuning the tone, level two is changing the script, level three is introducing multimedia materials, and level four is triggering manual intervention reminders.
[0065] The method provided by this invention, through precise capture of physiological signals and support from profound psychological logic, can not only identify surface expressions but also gain insight into the underlying psychological motivations. For example, when a child exhibits extreme excitement, the system will determine, in conjunction with the context, whether this is a "masked overreaction." By analyzing unnatural muscle tremors and high body tension scores in the child's photoflow field, it may ultimately determine that the child is experiencing excessive anxiety due to social pressure and provide timely stress-relief guidance strategies. This deep semantic understanding capability is the core competitiveness of this invention.
[0066] This invention also fully considers the robustness requirements under non-ideal environments. In the preprocessing of step 2 above, adaptive histogram equalization prevents excessive amplification of image background noise by limiting the magnitude of contrast enhancement. In feature extraction of step 3, pseudo-motion vectors caused by light flicker or camera thermal noise are filtered out by setting a local energy threshold. These engineered technical processes ensure that the system can operate stably in complex and ever-changing natural interaction scenarios such as homes and classrooms.
[0067] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A child emotion recognition and guidance interaction method based on multimodal affective computing, characterized in that, Includes the following steps: The multimodal perception unit synchronously collects children's facial video streams, head posture trajectories, and upper body behavior sequences in interactive situations. The facial video stream is preprocessed, and optical flow field information between consecutive frames is extracted to characterize the transient micro-tremor features of facial muscles. Facial behavior units are identified based on the facial behavior coding system, and dynamic expression time sequence maps are constructed by combining the optical flow field information to characterize the speed, amplitude and direction characteristics of expression evolution. The dynamic facial expression time sequence map is input into the psychological schema mapping model to match the potential psychological state category corresponding to the current facial expression combination; By fusing the head posture trajectory with the upper body behavior sequence, a contextual feature vector is generated, which serves as an auxiliary criterion in the final determination of the psychological state category. Based on the judgment results, a set of appropriate psychological counseling strategy instructions is generated, and interactive feedback is output through the speech synthesis module and the virtual character animation module. By monitoring children's behavioral responses to the interactive feedback, the intensity and pace of subsequent guidance strategies can be adjusted to form a closed loop of emotion regulation.
2. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 1, characterized in that: The process of synchronously acquiring children's facial video streams, head posture trajectories, and upper body behavior sequences in interactive scenarios through a multimodal perception unit specifically includes: Facial video streams are captured using a high frame rate visible light camera configured on the interactive terminal. The position information of a child's head in three-dimensional space is obtained using a depth sensor and fitted into a head posture trajectory; The inertial measurement unit is used to record the changes in angular velocity and acceleration of the child's upper body, forming a sequence of upper body behaviors; The shutter opening time of the visible light camera, the projection time of the depth sensor, and the data reading time of the inertial measurement unit are time-aligned by a global hardware trigger signal to ensure the synchronization of data acquisition.
3. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 2, characterized in that: The preprocessing of the facial video stream specifically includes: performing an adaptive histogram equalization operation, which calculates the brightness distribution histogram of the local region of the image and adjusts the pixel contrast to eliminate image distortion caused by changes in illumination. A blind deconvolution algorithm based on gradient sparse prior is used for motion deblurring filtering. Under the constraint of maintaining the sparsity of the gradient at the image edge, the point spread function is estimated alternately and the clear image is restored to suppress image ghosting. The face detection operator is used to identify and crop the face region of the processed image, and the cropped image is scaled to a uniform pixel size.
4. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 3, characterized in that: The method of identifying facial behavior units based on the facial behavior coding system and constructing a dynamic expression time series map by combining the optical flow field information specifically includes: extracting the spatial features of facial skin texture using a local binary mode, and extracting the dynamic evolution features of facial movements on the time axis using a temporal convolutional neural network. It can identify at least a variety of facial movement units, including: lifting the medial frontalis muscle, lifting the lateral frontalis muscle, depressing the corrugator supercilii muscle, lifting the levator palpebrae superioris muscle, contracting the orbicularis oculi muscle, lifting the levator labii superioris muscle, pulling the corner of the mouth with the zygomaticus major muscle, depressing the corner of the mouth with the anguli oris muscle, lifting the mentalis muscle, pulling the corner of the mouth with the risorius muscle, tightening the orbicularis oris muscle, pressing the orbicularis oris muscle, separating the lips, and lowering the jaw. The direction consistency index and energy concentration index of the optical flow vector are calculated. When the mean cosine similarity of all pixel motion vectors in a local area is greater than a preset direction threshold, and the cumulative value of the sum of squares of motion vector amplitudes within a preset time window is greater than a preset energy threshold, it is determined to be a micro-expression segment. Combining the intensity value of the facial action unit with the dynamic characteristics of the micro-expression segment, a multi-dimensional dynamic expression time sequence map including time dimension, spatial position dimension and motion intensity dimension is constructed.
5. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 4, characterized in that: The step of inputting the dynamic facial expression time sequence map into the psychological schema mapping model to match the potential psychological state category corresponding to the current facial expression combination specifically includes: constructing a psychological schema mapping model using a hierarchical Bayesian inference framework, with the bottom observation layer corresponding to the dynamic facial expression time sequence map and the high-level state layer corresponding to the psychological state category; The embedded emotional development milestone database is invoked, and the prior distribution parameters of the psychological schema mapping model are configured according to the child's age range. The age ranges include the first stage, which focuses on emotions triggered by basic physiological needs; the second stage, which focuses on emotions triggered by the emergence of self-awareness and social frustration; and the third stage, which focuses on emotions triggered by academic pressure and complex interpersonal relationships. Based on the prior distribution parameters, the posterior probability of the dynamic facial expression time series map belonging to different psychological state categories is calculated, and the preliminary psychological state category is determined according to the probability distribution vector. When the difference between the highest probability value and the second highest probability value is less than the preset stability threshold, the state records within the preset historical time period are retrieved and weighted averaged to output the final matching result.
6. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 5, characterized in that: The process of fusing the head posture trajectory with the upper body behavior sequence to generate a contextual feature vector specifically includes: obtaining the time-series curves of yaw angle, pitch angle and roll angle by tracking the rotation angle change of the head center vector in three-dimensional space; The coordinates of shoulder and elbow key points are tracked using a human posture estimation algorithm to generate upper body movement trajectories; the avoidance tendency index, characterized by the angle and frequency of the head turning away from the interaction center, is calculated. Calculate the body tension score, characterized by the elevation amplitude and microtremor component of the shoulder muscles; calculate the interaction willingness level, characterized by the rate of change of upper body displacement relative to the interactive interface. The avoidance tendency index, the physical tension score, and the interaction willingness level are combined into the situational context feature vector.
7. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 6, characterized in that: The step of generating an appropriate set of psychological counseling strategy instructions based on the judgment result and outputting interactive feedback through the speech synthesis module and the virtual character animation module specifically includes: matching the strategy type corresponding to the psychological state category from a preset interactive script library, wherein the strategy type includes soothing strategy, guiding strategy, transfer strategy and incentive strategy; Configure speech synthesis parameters according to the matched strategy type, the speech synthesis parameters including the fundamental frequency mean, fundamental frequency dynamic range, speech intensity and speech rate; The virtual character's facial expression template is invoked based on the matched strategy type, and a humanoid body language animation synchronized with the speech synthesis parameters is generated using a skeletal skinning algorithm. When the judgment result is a negative psychological state, a soothing feedback with a calm tone is output by reducing the mean of the fundamental frequency and narrowing the dynamic range of the fundamental frequency. When the judgment result is a positive psychological state or the completion of a specific task, the system outputs motivating feedback with a high-pitched tone by expanding the dynamic range of the fundamental frequency and increasing the intensity of the speech.
8. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 7, characterized in that: The monitoring of children's behavioral responses to the interactive feedback and the adjustment of the intensity and pace of subsequent guidance strategies specifically include: establishing a time sliding window and continuously statistically analyzing the changes in children's physiological indicators within a preset time after the interactive feedback is output. Compare the changes in the frequency of facial micro-expressions, the stability of head orientation, and the range of upper body movement of children before and after the feedback output; if the decrease in the frequency of facial micro-expressions is greater than the preset decrease threshold and the stability of the head orientation towards the center position is improved, it is determined that the emotions are stabilizing, the intensity of subsequent guidance strategies is reduced, and the interaction mode is switched to normal. If the facial movement unit detects frequent depressor anguli or corrugator supercilii movements (such as lowering the corners of the mouth or the lowering of the brow muscles), and the avoidance tendency index increases, then the emotional reaction is determined to be intensified, and the intervention strategy is switched to a higher level according to the preset decision tree logic.
9. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 8, characterized in that: The preprocessing of the facial video stream also includes brightness compensation and dynamic time warping operations. Specifically, the normal vector distribution of the facial surface is estimated based on the Lambertian body assumption, and brightness weighted compensation is performed on the shadow area caused by the offset of the light source position to ensure the accuracy of the optical flow field information extraction. A dynamic time warping algorithm is used to map captured facial motion sequences of varying lengths onto a standard-length time axis, thereby aligning the statistical features of expressions of different durations in terms of speed and amplitude.
10. The interactive method for children's emotion recognition and guidance based on multimodal emotion computing according to claim 9, characterized in that: The child emotion recognition and guidance interaction method based on multimodal emotion computing is also executed collaboratively by the cloud and the edge. Specifically, the preprocessing of the facial video stream, the extraction of the optical flow field information, and the recognition of the facial behavior unit are performed in the local edge computing module to obtain an intermediate feature vector. The intermediate feature vector is compressed and transmitted to the cloud server, where the reasoning of the psychological schema mapping model, the fusion and determination of the situational context feature vector, and the generation of the psychological counseling strategy instruction set are performed. Based on long-term monitoring data accumulation, the cloud server uses reinforcement learning algorithms to optimize the guidance strategy parameters for individuals, and then sends the optimized parameters to the local edge computing module.