A child psychological monitoring system based on multimodal spatiotemporal alignment
By integrating an RGB-D camera, microphone array, and infrared sensor, combined with a multi-branch lightweight network architecture and rule base, high-precision spatiotemporal alignment and analysis of multimodal data is achieved, solving the problem of insufficient multimodal data fusion in existing technologies and improving the recognition accuracy and reliability of the interaction process of the child psychological monitoring system.
Patent Information
- Application Number
- CN202511100674.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing child psychological monitoring systems lack spatiotemporal alignment when fusing multimodal data, resulting in low recognition accuracy and making it difficult to achieve accurate monitoring of children's daily psychological and behavioral states.
It employs an integrated RGB-D camera, microphone array, and infrared sensor to acquire multimodal data through temporal and spatial alignment. Combined with a multi-branch lightweight network architecture and feature fusion layer, it utilizes neural networks and rule bases to perform high-precision analysis and interactive decision-making of multimodal data.
It achieves high-precision, low-latency multimodal data acquisition and analysis, improves the accuracy of children's psychological state recognition and the timeliness of the interaction process, and solves the problems of semantic conflict and decision interpretability in complex scenarios.
Smart Images

Figure CN120636823B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a child psychological monitoring system based on multimodal spatiotemporal alignment. Background Technology
[0002] Children's mental health is a core element influencing their lifelong development and social adaptation. Early identification and intervention of abnormal mental states such as anxiety, depression, and social impairment have significant social value. However, traditional child mental health monitoring mainly relies on subjective observation by parents or teachers, periodic questionnaires, or invasive assessments by professional institutions. This approach has significant drawbacks, including poor real-time performance, narrow coverage, insufficient objectivity, and susceptibility to reporter bias. It struggles to capture subtle, transient, yet crucial psychological and behavioral signals in children's daily lives.
[0003] With the development of sensing technology and artificial intelligence, monitoring solutions based on multi-source data collection from wearable devices, video surveillance, and mobile terminals are gradually emerging. These technologies offer the possibility of overcoming the limitations of traditional methods. Existing technical solutions can be broadly categorized as follows:
[0004] Single-modal physiological monitoring systems primarily rely on devices such as smartwatches and wristbands to collect physiological indicators such as heart rate, skin conductance (GSR), and body temperature. Their limitation lies in the fact that physiological signals and psychological states do not have a simple one-to-one correspondence, and they severely lack behavioral and environmental context, easily leading to misjudgments.
[0005] Video-based behavior analysis systems utilize cameras to capture facial expressions, body postures, and movement trajectories. While these systems perform reasonably well in controlled laboratory environments, they face significant challenges in complex, dynamic scenes: changes in lighting, occlusion, limited field of view, and multi-person interaction greatly reduce recognition accuracy. More importantly, simple behavior analysis struggles to address underlying psychological states.
[0006] Voice emotion analysis systems use microphones to collect speech and analyze tone, speed, and keywords to infer emotions. However, children's language expression is often unclear or implicit, environmental noise interference is severe, and voice modalities can only reflect the state during communication, failing to reflect psychological activities when silent or alone.
[0007] Single-modal data such as physiological data, video, and voice can only reflect local features of psychological activities, lack the ability to perform multi-dimensional correlation analysis, and are prone to misjudgment due to environmental interference, resulting in low recognition accuracy. In addition, children's psychological states are dynamic and context-dependent, and the coordinated changes in their behavior, facial expressions, voice, and other signals across time and space have not yet been effectively explored.
[0008] Therefore, monitoring technology based on multimodal data has gradually become a research hotspot. Although existing technologies have made some progress in multimodal data acquisition and fusion, they still have significant shortcomings in spatiotemporal alignment, dynamic modeling, and interpretability, which limits the accuracy and practicality of monitoring systems.
[0009] For example, the invention patent titled "Children's ADHD Screening and Assessment System Based on Multimodal Deep Learning Technology" (CN111528859B) integrates multimodal data such as scale tests, eye tracking, and facial expression and posture analysis, and classifies abnormal behaviors through a temporal fusion model. Its core lies in using a hardware and software collaborative module to record eye movements, facial expressions, postures, and interactive actions during task tests, and combining this with a multimodal information fusion model to achieve screening and assessment. However, this system relies on the fusion of static multimodal data and does not consider the spatiotemporal asynchrony of different modal data. For example, the response delay between facial expression changes in videos and eye movement data is not accurately synchronized, leading to deviations in cross-modal correlation analysis.
[0010] In conclusion, the field of child psychological monitoring urgently needs a method that can effectively solve the problem of high-precision spatiotemporal alignment of multimodal data, achieve accurate monitoring of children's daily psychological and behavioral states, and provide truly reliable objective evidence for early warning and personalized intervention. Summary of the Invention
[0011] To address the issues of insufficient multimodal data fusion and low identification accuracy in existing child psychological monitoring processes, this invention provides a lightweight child psychological monitoring system based on multimodal spatiotemporal alignment.
[0012] The technical solution adopted by this invention to solve its technical problem is:
[0013] A child psychological monitoring system based on multimodal spatiotemporal alignment includes:
[0014] Data acquisition layer: integrates RGB-D camera, microphone array and infrared sensor, and performs spatiotemporal synchronous acquisition of multimodal data through time alignment and spatial alignment. The multimodal data includes image data, voice data and infrared data.
[0015] The image data includes color images and depth information of the user's facial expressions and three-dimensional pose;
[0016] The voice data includes the user's voice signal and voiceprint features;
[0017] The infrared data includes the user's pupil diameter, proximity status, and gesture commands;
[0018] Feature processing layer: includes a visual branch module, a speech branch module, and an infrared branch module, which respectively process the image data, speech data, and infrared data to obtain RGB-D fusion feature vector, age classification probability, expression classification probability, speech feature vector, emotion classification probability, intensity level, behavior feature vector, pupil diameter change rate, and gesture classification result;
[0019] The RGB-D fusion feature vector, speech feature vector, and behavior feature vector constitute a vector information set;
[0020] The age classification probability, expression classification probability, emotion classification probability, intensity level, pupil diameter change rate, and gesture classification results constitute a structured information set;
[0021] Feature fusion layer: By fusing the vector information set and the structured information set, fused features are obtained;
[0022] Knowledge fusion layer: Analyzes the fusion features based on neural networks to generate global sentiment state, user profile, and interaction commands;
[0023] The global sentiment state includes the global sentiment classification probability and the global intensity level;
[0024] The user profile includes age and ESI index;
[0025] Dynamic Interaction Decision Layer: By using PPO reinforcement learning on the global emotional state, user profile, and interaction commands, expression weights, voice weights, and action weights are obtained. The expression weights, voice weights, and action weights drive the terminal to adjust the voice, expression, and action interaction modes.
[0026] Furthermore, the time alignment is based on the PTP protocol to ensure that the device time bases of the RGB-D camera, microphone array and infrared sensor are consistent, and the time alignment of the multimodal data is achieved through a timestamp mechanism;
[0027] The spatial alignment unifies the infrared coordinate system and the RGB-D image coordinate system through an affine transformation. The affine transformation formula is as follows:
[0028]
[0029] In the formula: (x, y) are parameters in the infrared coordinate system; (x`, y`) are parameters in the RGB-D image coordinate system, with a range of 0 ≤ x` < 640, 0 ≤ y` < 480; a, b, c, d, t x t y These are affine transformation parameters, calibrated offline using a calibration board.
[0030] Furthermore, the visual branch module is based on the improved MobileNet-AGE network architecture, and introduces an inverse residual structure and an h-swish activation function to obtain the RGB-D fusion feature vector, age classification probability, and expression classification probability;
[0031] The age classifications include: 3-6 years old, 7-9 years old, 10-12 years old, and >12 years old;
[0032] The facial expression categories include: smiling, frowning, crying, staring, surprised, disgusted, and expressionless.
[0033] Furthermore, the speech branch module adopts a cascaded architecture of 1D CNN and bidirectional GRU to extract the speech feature vector and simultaneously output the emotion classification probability and scalar intensity value;
[0034] The emotion categories include: happiness, sadness, anger, fear, surprise, disgust, and neutrality;
[0035] The intensity level is obtained by classifying the scalar intensity value according to a threshold.
[0036]
[0037] In the formula: x t The scalar strength value, .
[0038] Furthermore, the infrared branch module analyzes the data through the infrared processing API to obtain the gesture classification results, behavioral feature vectors, and pupil diameter change rate.
[0039] The gesture classification results include: nodding, shaking head, waving, clenching fist, sliding, and stillness.
[0040] Furthermore, the feature fusion layer fuses the vector information set and the structured information set through gated cross-attention, and the steps for obtaining the fused features include:
[0041] Step 11: The RGB-D fusion feature vector, age classification probability, and expression classification probability constitute a visual feature set; the speech feature vector, emotion classification probability, and intensity level constitute a speech feature set; the behavior feature vector, pupil diameter change rate, and gesture classification result constitute a behavior feature set.
[0042] The visual feature set, speech feature set, and behavior feature set are mapped to a unified semantic space by a projection algorithm to obtain projected visual, speech, and behavior feature vectors, which correspond to visual, speech, and behavior modalities, respectively.
[0043] Step 12: Using the visual, speech, and behavioral modalities as target modalities respectively, calculate the attention weights of the target modalities relative to the other two modalities. The target modalities are used as query modalities, and the other two modalities are used as key-value modalities. After attention aggregation, each target modality yields a first-level fusion feature.
[0044] Step 13: Calculate the gating weights for each mode;
[0045] Step 14: Multiply the first-level fusion feature of each modality by the corresponding gating weight to obtain the weighted feature. Then sum the weighted features of all modalities and the projected visual feature vector to obtain the fusion feature.
[0046] Furthermore, the feature fusion layer achieves the fusion of multimodal feature vectors in the vector information set through a cross-modal spatiotemporal attention mechanism and a dynamic feature weighting strategy, thereby obtaining fused vector information. The age classification probability, expression classification probability, emotion classification probability, and intensity level in the structured information set are updated through the analysis of the fused vector information by a neural network. The updated structured information set is then fused through gated cross-attention fusion to obtain fused features.
[0047] Furthermore, the fusion of multimodal feature vectors in the vector information set includes the following steps:
[0048] Step 21: Map the RGB-D fused feature vector, speech feature vector, and behavior feature vector to query, key, and value vectors to obtain the query vector of the RGB-D fused feature vector, the key vector of the speech feature, and the key vector of the behavior feature;
[0049] Step 22: Based on the query vector of the RGB-D fusion feature vector, the key vector of the speech feature, and the key vector of the behavior feature, calculate the attention weights of the RGB-D fusion feature vector on the speech feature vector and the behavior feature vector, respectively.
[0050] Step 23: Obtain the multimodal fusion feature vector by applying attention weights to the speech feature vector and behavior feature vector respectively based on the RGB-D fusion feature vector;
[0051] Step 24: Compress the multimodal fusion feature vector using 1×1 convolution to obtain fusion vector information.
[0052] Furthermore, the knowledge fusion layer includes a rule base, which includes dynamic interaction strategy rules and user profile construction rules. The fusion features are analyzed using a neural network combined with the rule base, and a multi-task output head is used to generate global sentiment state, user profile, and interaction instructions in parallel.
[0053] Furthermore, the rule base also includes basic emotion recognition rules and multimodal conflict resolution rules. The basic emotion recognition rules are constructed based on Ekman's basic emotion theory and Russell's circular emotion model, and the multimodal conflict resolution rules are constructed based on Zuckerman's deception model. The global emotional state is modified by combining the Rete algorithm.
[0054] Furthermore, the step of modifying the global sentiment state using the Rete algorithm includes:
[0055] Step 31: Analyze the fused features using a neural network to generate a global sentiment classification prediction probability;
[0056] Step 32: Generate AU action units based on the RGB-D fusion feature vector, and use the Rete algorithm to match the AU action units, speech feature vector, pupil change rate with the basic emotion recognition rules to determine whether the basic emotion categories corresponding to different modalities are consistent. The basic emotion categories include: happiness, sadness, anger, fear, surprise, disgust, and neutrality.
[0057] When the basic emotion categories corresponding to different modalities are inconsistent, proceed to step 33;
[0058] When the basic emotion categories corresponding to different modalities are consistent, the global emotion classification probability is the global emotion classification prediction probability;
[0059] Step 33: Match the expression classification probability, emotion classification probability, pupil diameter change rate, and gesture classification result with the multimodal conflict resolution rule, update the emotion classification probability, and output the global emotion classification rule probability;
[0060] Step 34: Calculate the rule trigger weight by weighting the confidence level based on the expression classification probability and the sentiment classification probability;
[0061] Step 35: Based on the global sentiment classification rule probability, rule trigger weight, and global sentiment classification prediction probability, calculate the global sentiment classification correction probability, whereby the global sentiment classification probability is the global sentiment classification correction probability.
[0062] Furthermore, the multimodal conflict resolution rules include:
[0063] When the expression is classified as "smiling" and the emotion is classified as "sad," the forced smile correction rule is triggered:
[0064]
[0065] In the formula: Adjust the weights for the rules to adjust the probability distribution of emotion categories, setting the probability of the sadness category to... Output the global sentiment classification rule probability. Classify the probability of facial expressions. For sentiment classification probability, This represents the rate of change in pupil diameter.
[0066] Furthermore, the interactive decision-making method of the dynamic interactive decision-making layer includes the following steps:
[0067] Step 41: Encode the global emotional state, user profile, and interaction commands into a unified numerical state vector;
[0068] Step 42: Using the PPO algorithm, generate the expression weights, voice weights, and action weights based on the state vector;
[0069] Step 43: Based on the facial expression weight, voice weight, and action weight, drive the terminal to adjust the voice, facial expression, and action interaction modes.
[0070] Furthermore, step 42 also includes calculating a reward function to generate a reward value, which is used for parameter updates of the PPO algorithm.
[0071] Furthermore, the dynamic interaction decision layer also includes a weight correction module. Based on the expression classification probability, emotion classification probability, and pupil diameter change rate, the weight correction module determines whether there is a conflict scenario according to the Zuckerman deception model. When a conflict scenario exists, the module forcibly corrects the allocation of expression weight, voice weight, and action weight.
[0072] Furthermore, the voice interaction mode adjusts the pitch and speech rate of the synthesized speech based on speech weights;
[0073] The facial expression interaction mode adjusts the amplitude of the terminal's facial expressions according to the expression weight;
[0074] The action interaction mode adjusts the terminal's response sensitivity to user gestures based on action weights.
[0075] Furthermore, it also includes a parent platform that enables the tracking of overall emotional state, user profiles, interactive commands, facial expression weights, voice weights, and action weights; it can also push psychological counseling suggestions based on user profiles.
[0076] The beneficial effects of this invention are as follows:
[0077] This high-precision, high-synchronization-rate, and low-cost multimodal data acquisition system integrates an RGB-D camera, an infrared sensor, and a microphone array. Combined with hardware clock synchronization and data stream alignment mechanisms, it achieves high-precision acquisition and time consistency of multimodal data. The RGB-D camera provides 640×480 resolution color images and ±5mm depth information for capturing user facial expressions and body posture. The infrared sensor detects user proximity and supports user emotional state analysis. The microphone array clearly captures user voice signals through 2-channel beamforming and a signal-to-noise ratio of ≥60dB. The hardware clock synchronization module, based on the PTP protocol (synchronization error ≤1ms), ensures consistent time bases across all devices. The data stream alignment module uses a timestamp mechanism (alignment error ≤2ms) to achieve time alignment of the multimodal data streams. Time alignment unifies the infrared coordinate system and the RGB-D image coordinate system through affine transformation.
[0078] The feature processing layer adopts a multi-branch lightweight network architecture to process visual (RGB-D), speech (MFCC), and infrared (pupil diameter / gesture, manufacturer API interface) data respectively. Combined with the feature fusion layer, it achieves efficient feature extraction and fusion through cross-modal feature alignment and multi-task learning. It fully integrates the spatiotemporal correlation of multimodal data such as visual (RGB-D camera), speech (microphone array), and behavior (infrared sensor) data to improve the accuracy of children's facial expression and emotion recognition.
[0079] Different sensor data may generate semantic conflicts due to environmental noise, user behavior masquerading, or differences in physiological state (e.g., the coexistence of "smiling emojis" and "sad voices"). Traditional single-modal decision-making mechanisms struggle to effectively resolve such contradictions, leading to emotional misjudgments and interaction failures. To address this, this invention constructs a four-layer progressive rule base, including basic emotion recognition rules, multimodal conflict resolution rules, dynamic interaction strategy rules, and user profile construction rules. By fusing multimodal features and combining them with a knowledge system built upon the rule base, and through psychological theory-driven and computational model-assisted approaches, it achieves accurate detection and resolution of conflicts, solving the problems of semantic conflict and decision interpretability in complex scenarios.
[0080] The dynamic interaction decision layer, based on the high-precision global emotional state (7 types of emotions + 3 intensity levels), user profile (age, ESI index), and historical interaction data output by the knowledge fusion layer, dynamically optimizes multimodal weights (facial expressions, voice, and actions) through the PPO algorithm. This generates low-latency, personalized interaction commands and triggers terminal responses, improving the timeliness and accuracy of the interaction process in the child psychological care system. Furthermore, it uses the Zuckerman deception model to determine the existence of conflict scenarios. When conflict scenarios exist, it forcibly corrects the allocation of facial expression weights, voice weights, and action weights. Through a dynamic interaction decision-making method that collaborates with psychological rules, the accuracy of interaction decisions is enhanced. Attached Figure Description
[0081] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0082] Figure 1 A framework diagram of a child psychological monitoring system based on multimodal spatiotemporal alignment;
[0083] Figure 2 This is a schematic diagram illustrating the process of fusing multimodal feature vectors in the vector information set in the embodiment.
[0084] Figure 3 A flowchart illustrating the process of correcting the global sentiment state using the Rete algorithm;
[0085] Figure 4 A basic emotion recognition rule table;
[0086] Figure 5 For a multimodal conflict resolution rule table;
[0087] Figure 6 For state interaction strategy rule table;
[0088] Figure 7 Build rule tables for user profiles. Detailed Implementation
[0089] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0090] A child psychological monitoring system based on multimodal spatiotemporal alignment can be applied to child companion robots. It significantly improves the technical level of child psychological monitoring through high-precision multimodal fusion, low-latency real-time interaction and dynamic adaptation capabilities.
[0091] like Figure 1 As shown, the system includes:
[0092] Data Acquisition Layer: Integrates an RGB-D camera, microphone array, and infrared sensor. Through time and spatial alignment, it performs spatiotemporal synchronous acquisition of multimodal data, including image data, voice data, and infrared data. Image data includes color images and depth information of the user's facial expressions and three-dimensional pose. Voice data includes the user's voice signal and voiceprint features. Infrared data includes the user's pupil diameter, proximity status, and gesture commands.
[0093] Feature processing layer: Includes visual branch module, speech branch module, and infrared branch module, which process image data, speech data, and infrared data respectively to obtain RGB-D fusion feature vector, age classification probability, facial expression classification probability, speech feature vector, emotion classification probability, intensity level, behavior feature vector, pupil diameter change rate, and gesture classification result; RGB-D fusion feature vector, speech feature vector, and behavior feature vector constitute a vector information set; age classification probability, facial expression classification probability, emotion classification probability, intensity level, pupil diameter change rate, and gesture classification result constitute a structured information set;
[0094] Feature fusion layer: By fusing the vector information set and the structured information set, fused features are obtained;
[0095] Knowledge fusion layer: Based on neural networks, the fusion features are analyzed to generate global sentiment state, user profile, and interaction commands. The global sentiment state includes global sentiment classification probability and global intensity level, and the user profile includes age and ESI index.
[0096] Dynamic Interaction Decision Layer: By using PPO reinforcement learning on global emotional state, user profile, and interaction commands, we obtain facial expression weights, voice weights, and action weights. These weights then drive the terminal to adjust the voice, facial expression, and action interaction modes.
[0097] In this embodiment, multimodal information is collected through multiple modules. By setting up a visual branch module, a voice branch module, and an infrared branch module, visual, voice, and infrared data are processed respectively, realizing parallel analysis and processing of multimodal data. Then, a feature fusion module is used to fuse the multi-type output data of the feature processing layer. A neural network is used to learn the fused features to obtain global emotional state, user profile, and interaction commands. Based on the global emotional state, user profile, and interaction commands, PPO reinforcement learning is used to obtain expression weights, voice weights, and action weights. The expression weights, voice weights, and action weights drive the terminal to adjust the voice, expression, and action interaction modes, realizing multimodal data analysis and multimodal interaction.
[0098] In this embodiment, the data acquisition layer integrates an RGB-D camera, an infrared sensor, and a microphone array to achieve high-precision, high-synchronization-rate multimodal data acquisition. To achieve temporal and spatial synchronization of the multimodal data, another embodiment of the invention employs both temporal alignment and spatial alignment to maintain the spatiotemporal consistency of the multimodal data.
[0099] Time alignment is based on the PTP protocol to ensure that the device time bases of RGB-D cameras, microphone arrays and infrared sensors are consistent, and time alignment of multimodal data is achieved through a timestamp mechanism.
[0100] Hardware clock synchronization is based on the PTP protocol (IEEE 1588), with a synchronization error of ≤1ms. Through hardware clock synchronization technology, it is ensured that the data acquisition of the RGB-D camera, infrared sensor and microphone array are carried out under the same time reference.
[0101] Data stream alignment is based on timestamps, with an alignment error of ≤2ms. The timestamp mechanism ensures that the data streams of each modality are aligned on the time axis, providing an accurate foundation for subsequent multimodal fusion and analysis.
[0102] Spatial alignment unifies the infrared coordinate system and the RGB-D image coordinate system through affine transformation. Combined with affine transformation matrix calibration (spatial error ≤ 2 pixels), it ensures the spatial consistency between behavioral data and visual scenes.
[0103] The formula for affine transformation is:
[0104]
[0105] In the formula: (x, y) are parameters in the infrared coordinate system; (x`, y`) are parameters in the RGB-D image coordinate system, with a range of 0 ≤ x` < 640, 0 ≤ y` < 480; a, b, c, d, t x t y These are affine transformation parameters, calibrated offline using a calibration board.
[0106] Output spatiotemporally synchronized multimodal data streams through time and space alignment: including image, voice and infrared data, ensuring high temporal consistency among the various modalities.
[0107] Image data: Output color image (640×480) and depth information (±5mm) for visual feature extraction (age, expression).
[0108] Speech data: Output speech signal (16kHz, signal-to-noise ratio ≥60dB) and voiceprint features (MFCC) for speech emotion classification.
[0109] Infrared data: Outputs pupil diameter (accuracy ±0.1mm), proximity status (1-5 meters), and gesture commands (such as waving), used for emotion intensity analysis and action command generation.
[0110] This embodiment relates to a high-precision, high-synchronization-rate, and low-cost multimodal data acquisition system. By integrating an RGB-D camera, an infrared sensor, and a microphone array, combined with hardware clock synchronization and data stream alignment mechanisms, it achieves high-precision acquisition and time consistency of multimodal data. The RGB-D camera provides 640×480 resolution color images and ±5mm accuracy depth information for capturing user facial expressions and body posture. The infrared sensor, based on an 850nm wavelength, enables proximity detection within a 1-5 meter range, supporting user emotional state analysis. The microphone array, through 2-channel beamforming and a signal-to-noise ratio of ≥60dB, clearly captures user voice signals. The hardware clock synchronization module, based on the PTP protocol (synchronization error ≤1ms), ensures consistent time bases across all devices. The data stream alignment module, through a timestamp mechanism (alignment error ≤2ms), achieves time alignment of the multimodal data streams. Time alignment unifies the infrared coordinate system and the RGB-D image coordinate system through affine transformation.
[0111] The feature processing layer is designed with a multi-branch lightweight network architecture to process visual, speech and infrared data respectively.
[0112] The visual branch module is based on the improved MobileNet-AGE network architecture, and introduces the inverse residual structure and h-swish activation function to obtain RGB-D fusion feature vectors, age classification probability, and expression classification probability.
[0113] Age categories include: 3-6 years old, 7-9 years old, 10-12 years old, and >12 years old;
[0114] Facial expressions are categorized as follows: smiling, frowning, crying, staring, surprised, disgusted, and expressionless.
[0115] Traditional MobileNetV2 uses an inverse residual structure with an Expansion Ratio of 6, resulting in high channel redundancy (1152 channels after dimensionality upscaling), which is difficult to adapt to the characteristics of children's simple facial features and limited texture details. This embodiment improves the inverse residual structure by optimizing the expansion ratio to 4. By compressing the number of upscaling channels (reduced to 768), it significantly reduces the number of model parameters (2.1M → 1.46M) and computational cost (FLOPs reduced by 19.8%) while preserving key age-sensitive information (such as contour proportions). Multi-stage experiments validate that Expansion Ratio=4 achieves an age classification accuracy of 93.5% on the UTKFace dataset without introducing overfitting (validation set loss reduced by 15%).
[0116]
[0117] Output is the output of the inverse residual structure, and Input is the input. It is a 1×1 convolutional layer, with ReLU as the activation function. It uses a 3×3 convolutional layer. By first expanding the feature dimension (Expansion Ratio=4), then performing 3×3 convolution, and finally compressing the feature dimension, the feature representation capability is significantly improved. At the same time, the inverse residual structure reduces the number of parameters while maintaining high computational efficiency, making it suitable for deployment in embedded devices.
[0118] On the other hand, this embodiment uses the h-swish activation function instead of ReLU6 to improve the nonlinear expression capability of the network while maintaining computational efficiency.
[0119] h-swish performs more stably in low-precision quantization scenarios and can better preserve feature information. Compared to ReLU6, h-swish significantly improves the nonlinear expressive power of the network while maintaining computational efficiency.
[0120] The visual branch module employs a lightweight multi-task neural network based on the improved MobileNet-AGE network architecture, combined with a cross-modal feature alignment module, to achieve efficient multi-task processing of RGB-D data. By introducing an inverse residual structure (Expansion Ratio=4) and the h-swish activation function, feature representation capability and computational efficiency are significantly improved. Simultaneously, 1×1 convolutional compression of the depth channels, combined with the SE Block channel attention mechanism, achieves efficient fusion of RGB and depth features, solving the problem of poor cross-modal feature fusion performance in traditional networks.
[0121] The speech branch module adopts a cascaded architecture of 1D CNN and bidirectional GRU to extract speech feature vectors and simultaneously output sentiment classification probability and scalar intensity value.
[0122] Emotional categories include: happiness, sadness, anger, fear, surprise, disgust, and neutrality.
[0123] In this embodiment, the 1D CNN local feature extraction uses three layers of dilated convolution (Dilation=1, 2, 4) to gradually expand the receptive field. Each layer is followed by BatchNorm and ReLU activation, and finally outputs 256-dimensional high-order acoustic features, compressing the time step to 1 / 8 of the input.
[0124]
[0125] in: Indicates the first The output of layer at time step t, It is the weight index of the convolution kernel, and its value range is ( ),in It is the width of the convolution kernel (e.g.) =5 indicates that 5 time steps are covered. Indicates the first The weights of the convolutional kernels, For input features, The coefficient of thermal expansion is (1 / 2 / 4). This is a bias term.
[0126] After this step, the output is: (Batch, 256, T / 8), which is the 256-dimensional feature after downsampling the time dimension by 8 times.
[0127] The bidirectional GRU encodes temporal information from the forward and backward directions respectively, and finally splices the last hidden state at both ends to form a 256-dimensional global feature vector (128 dimensions from the forward direction + 128 dimensions from the backward direction).
[0128] After this step, the output will be (Batch, 256), which is the speech feature vector of the entire speech segment.
[0129] After extracting the speech feature vector from the speech signal, this embodiment achieves fine-grained emotion classification through a fully connected layer and Softmax normalization. The input is a 256-dimensional speech feature vector output by a bidirectional GRU, which is then processed by a learnable weight matrix. and bias After a linear transformation, the probability distribution of the seven sentiment categories is calculated using the Softmax function. Specifically, the module first performs feature mapping, projecting the speech feature vectors onto the sentiment category space; the second step is normalization, using exponential operations and normalization to ensure the validity of the output probability; finally, a decision is made, selecting the category with the highest probability as the final prediction result.
[0130] To quantify the intensity of emotional expression, this module employs a two-layer MLP+Sigmoid regression structure. The input speech feature vector is transformed nonlinearly to obtain a scalar intensity value. The values are categorized into low, medium, and high levels based on a threshold. This embodiment employs nonlinear modeling, using the ReLU activation function to capture the complex mapping relationship between intensity and features; the Sigmoid function compresses the output to the [0,1] interval; and a fixed threshold is used to convert continuous values to discrete levels.
[0131] Because after Sigmoid activation, the scalar intensity value naturally falls in the [0,1] interval, with sparse distribution at both ends (few values close to 0 or 1), and most samples concentrated in the middle region (0.3~0.7). Combined with the consistency between psychology and labeling and the nonlinearity of human perception, people's perception of emotional intensity is nonlinear, and the acceptance range of "medium intensity" is wider (in line with the Weber-Fechner law), thus obtaining the threshold division rule.
[0132] Strength levels are determined by classifying scalar strength values according to threshold values;
[0133]
[0134] In the formula: x t The scalar strength value, .
[0135] Through the above operations, the original output is a (Batch, 1) Sigmoid value, and after processing, each sample is output as an intensity level label.
[0136] In this embodiment, the speech branch module adopts a cascaded architecture of 1D CNN and bidirectional GRU to achieve speech feature vector extraction and emotional semantic modeling of speech signals. It captures the short-time spectral patterns of speech feature vectors through the local perceptual characteristics of 1D convolutional kernels (Kernel Size=5, covering 31.25ms of audio), and combines this with the long-range temporal dependency modeling capability of the bidirectional GRU to simultaneously output multi-task predictions of emotion categories (7 categories) and intensity levels (3 levels). After SoftMax normalization, the network output has a dimension of 7 for the emotion branch (corresponding to 7 basic emotions), and the intensity branch outputs scalar intensity values in the range of 0 to 1, divided into low / medium / high intensity levels.
[0137] The infrared branch module analyzes the data through the infrared processing API to obtain gesture classification results, behavioral feature vectors, and pupil diameter change rate. The gesture classification results include: nodding, shaking head, waving, clenching fist, sliding, and stillness.
[0138] Infrared data is input, digitally filtered (5th-order Butterworth) and motion detected, then fed into the Gesture_Detection_API for real-time gesture classification. This API, based on a lightweight CNN and dynamic time warping algorithm, outputs six basic gesture categories (nodding, shaking head, waving, clenching fist, swiping, and stillness), and supports OTA updates to expand the gesture library. Existing hardware sensors can be used directly; a ready-made sensor, such as the Sharp GP2Y0A02YK, is planned to acquire behavioral feature vectors and pupil diameter change rates. The processing is hardware-accelerated (DSP feature extraction + Cortex-M7 classification), completing the entire process from signal acquisition to action recognition within 50ms.
[0139] The feature processing layer mainly realizes the accurate extraction of multimodal features. Through parallel visual branch module, speech branch module and infrared branch module, it outputs RGB-D fusion feature vector, age classification probability, expression classification probability, speech feature vector, emotion classification probability, intensity level, behavior feature vector, pupil diameter change rate and gesture classification results.
[0140] The RGB-D fusion feature vector, speech feature vector, and behavior feature vector constitute a vector information set;
[0141] The structured information set consists of age classification probability, facial expression classification probability, emotion classification probability, intensity level, pupil diameter change rate, and gesture classification results.
[0142] The feature fusion layer enables the fusion analysis of multimodal features. In this embodiment, the feature fusion layer fuses vector information sets and structured information sets through gated cross-attention. The steps for obtaining fused features include:
[0143] Step 11: The RGB-D fusion feature vector, age classification probability, and expression classification probability constitute the visual feature set; the speech feature vector, emotion classification probability, and intensity level constitute the speech feature set; the infrared feature set includes: the behavior feature vector, pupil diameter change rate, and gesture classification results constitute the behavior feature set.
[0144] The visual feature set, speech feature set, and behavior feature set are mapped to a unified semantic space by a projection algorithm to obtain the projected visual, speech, and behavior feature vectors, which correspond to the visual, speech, and behavior modalities, respectively.
[0145]
[0146] In the formula: This is the projected visual feature vector. This is the projected speech feature vector. This is the projected behavioral feature vector;
[0147] For age-related classification probabilities, Classify the probability of facial expressions. For sentiment classification probability, For the gesture classification results, The rate of change of pupil diameter; This is an RGB-D fused feature vector. I is the speech feature vector, and I is the behavior feature vector; These are the learnable projection matrices corresponding to visual, speech, and behavioral features, respectively. These are the bias terms corresponding to visual, speech, and behavioral features, respectively.
[0148] Step 12: Using visual, speech, and behavioral modalities as target modalities respectively, calculate the attention weight of the target modality relative to the other two modalities. The target modality is used as the query modality, and the other two modalities are used as key-value modalities. After attention aggregation, each target modality is used to obtain the first-level fusion feature.
[0149] Specifically, when the visual modality is the target modality, the attention weights of the visual modality relative to the speech modality and the visual modality relative to the infrared modality are calculated respectively. When the speech modality is the target modality, the attention weights of the speech modality relative to the visual modality and the speech modality relative to the infrared modality are calculated respectively. When the infrared modality is the target modality, the attention weights of the infrared modality relative to the visual modality and the infrared modality relative to the speech modality are calculated respectively.
[0150] The formula for calculating the weights of bidirectional attention is:
[0151]
[0152] Where: Output This represents the attention weight of mode m to mode n, with a total of 6 groups (visual→speech, visual→infrared, speech→visual, speech→infrared, infrared→visual, infrared→speech). , Modal m and modality n The corresponding projected visual feature vector The projected speech feature vector The projected behavioral feature vector ; Representing modes m Query transformation matrix, modality n The key transformation matrix, represent Transpose of a matrix.
[0153] For each pair of modalities (e.g., vision to speech, vision to infrared, etc.), calculate the attention weight of one modality to another. For each target modality m (e.g., vision), treat the target modality as the query modality and the other modalities as the key modalities, and calculate the weighted features that the target modality m obtains from each of the other modalities n (excluding itself).
[0154] The formula for calculating the first-level fusion feature is:
[0155]
[0156] In the formula: This represents the first-order fusion feature of mode n; For modality n Corresponding visual feature vector after projection The projected speech feature vector The projected behavioral feature vector .
[0157] Step 13: Calculate the gating weights for each mode;
[0158] The formula for calculating the gating weight is:
[0159] In the formula: Let n be the gating weight for mode n; This represents the gate weight matrix for mode n; This is a gated bias term; By compressing the gating weights to the (0, 1) interval, feature-level filtering is achieved.
[0160] Step 14: Multiply the first-level fusion feature of each modality by the corresponding gating weight to obtain the weighted feature, and then sum the weighted features of all modalities to obtain the fusion feature.
[0161] A residual network is used to perform deep optimization on the fused features. By using skip connections to preserve the original information, it learns high-order nonlinear relationships to prevent gradient vanishing and improves the model's adaptability to complex scenarios.
[0162]
[0163] In the formula: As a feature of fusion, For normalization layer, This is the projected visual feature vector.
[0164] To achieve the fusion of multimodal feature vectors in the vector information set, in another embodiment of the present invention, the feature fusion layer uses a cross-modal spatiotemporal attention mechanism and a dynamic feature weighting strategy to achieve the fusion of multimodal feature vectors in the vector information set, thereby obtaining fused vector information. The age classification probability, expression classification probability, emotion classification probability, and intensity level in the structured information set are updated by analyzing the fused vector information through a neural network. The updated structured information set is then fused through gated cross-attention fusion to obtain fused features.
[0165] This embodiment first fuses and compresses the multimodal feature vectors in the vector information set to reduce the memory occupied by subsequent data processing and reduce computational latency. Specifically, for example... Figure 2 As shown, the fusion of multimodal feature vectors in a vector information set includes the following steps:
[0166] Step 21: Map the RGB-D fused feature vector, speech feature vector, and behavior feature vector to query, key, and value vectors to obtain the query vector of the RGB-D fused feature vector, the key vector of the speech feature, and the key vector of the behavior feature;
[0167]
[0168] In the formula: This is an RGB-D fused feature vector. I is the speech feature vector, and I is the behavior feature vector; RGB-D fusion features , For speech feature weight matrix, This is the behavioral feature weight matrix; The query vector is the RGB-D fused feature vector. The key vector for speech features. The key vector represents the behavioral characteristics.
[0169] Step 22: Based on the query vector of the RGB-D fusion feature vector, the key vector of the speech feature, and the key vector of the behavior feature, calculate the attention weights of the RGB-D fusion feature vector on the speech feature vector and the behavior feature vector, respectively.
[0170]
[0171]
[0172] In the formula: The attention weights of the RGB-D fusion feature vector on the speech feature vector. ; The attention weights of the RGB-D fused feature vector on the behavior feature vector. ; This is the transpose of the key vector of the speech features; This is the transpose of the key vector of the behavioral features; in this embodiment, the scaling factor is 64.
[0173] Step 23: Obtain the multimodal fusion feature vector by applying attention weights to the speech feature vector and behavior feature vector respectively based on the RGB-D fusion feature vector;
[0174]
[0175] In the formula: This is a multimodal fusion feature vector. This indicates a splicing operation.
[0176] Step 24: Compress the multimodal fusion feature vector using 1×1 convolution to obtain the fusion vector information;
[0177]
[0178] In the formula: This represents a 1×1 convolutional layer. To fuse vector information.
[0179] After obtaining the fused vector information, the fused vector information is learned through a neural network. Specifically, a multi-task prediction head is used to obtain the probability of age classification, the probability of expression classification, the probability of emotion classification, and the intensity level. The probability of age classification, the probability of expression classification, the probability of emotion classification, and the intensity level obtained based on the data processing layer are updated. Then, the updated structured information set is fused through gating cross-attention to obtain the fused features.
[0180] Specifically, the computational process for obtaining fusion features based on structured information is referenced in steps 11-14, with the difference being that the inputs in the visual, speech, and behavioral multimodal semantic alignment are different in step 11.
[0181] In this embodiment, the input is the updated structured information set, including the age classification probability, expression classification probability, and emotion classification probability obtained based on the fused vector information, as well as the gesture classification result and pupil diameter change rate obtained by the data processing layer. The specific calculation formula is as follows:
[0182]
[0183] In the formula: This is the projected visual feature vector. This is the projected speech feature vector. This is the projected behavioral feature vector;
[0184] For age-related classification probabilities, Classify the probability of facial expressions. For sentiment classification probability, For the gesture classification results, The rate of change of pupil diameter; These are the learnable projection matrices corresponding to visual, speech, and behavioral features, respectively. These are the bias terms corresponding to visual, speech, and behavioral features, respectively.
[0185] The feature processing layer obtains fused features from the fused vector information and the structured information set, which are then processed by the knowledge fusion layer.
[0186] To achieve deep learning of fusion features, a knowledge fusion layer was designed. The knowledge fusion layer includes a rule base, which includes dynamic interaction strategy rules and user profile construction rules. A neural network is used in conjunction with the rule base to analyze the fusion features. A multi-task output head is used to generate global sentiment state, user profile, and interaction instructions in parallel.
[0187] In multimodal perception scenarios of affective computing, semantic conflicts may arise from environmental noise, user behavior masquerading, or differences in physiological states (e.g., the coexistence of "smiling emojis" and "sad voice"). Traditional single-modal decision-making mechanisms struggle to effectively resolve such contradictions, leading to emotional misjudgments and interaction failures. To address this, this embodiment constructs a four-layer progressive rule base, including basic emotion recognition rules, multimodal conflict resolution rules, dynamic interaction strategy rules, and user profile construction rules. By fusing multimodal features and combining them with a knowledge system built upon the rule base, and through psychological theory-driven collaboration with computational models, accurate detection and resolution of conflicts are achieved, resolving semantic conflicts and decision interpretability issues in complex scenarios.
[0188] In another embodiment, the rule base also includes basic emotion recognition rules and multimodal conflict resolution rules. The basic emotion recognition rules are constructed based on Ekman's basic emotion theory and Russell's circular emotion model, while the multimodal conflict resolution rules are constructed based on Zuckerman's deception model. The Rete algorithm is used to correct the global emotional state. In this embodiment, the four-layer rule system achieves millisecond-level inference through the Rete algorithm, forming a complete processing chain from instantaneous perception to long-term cognition.
[0189] Basic emotion recognition rules such as Figure 4 As shown, the Facial Action Unit (AU) theory in the Basic Emotion Recognition Rule Table comes from the Ekman Facial Action Coding System (FACS). For example, the co-activation of AU6 (orbicularis oculi) and AU12 (zygomaticus major) is the gold standard for distinguishing between genuine and fake smiles.
[0190] In this embodiment, AU action units can be generated by learning RGB-D fused feature vectors based on neural networks for basic emotion category recognition. The basic emotion recognition rule table is based on Ekman's basic emotion theory and Russell's ring emotion model, and defines core emotion mapping rules by combining multimodal signals. It maps the original multimodal signals (visual, speech, infrared, etc.) to discrete emotion categories (such as anger, happiness, sadness), providing standardized input for higher-level decision-making. Micro-expressions are quantified by combining FACS action units (such as AU6+AU12), and a multi-evidence verification mechanism is constructed by combining speech MFCC bandwidth and pupil change rate, which significantly improves the robustness of cross-domain data (such as children's faces). For example, when AU4 (frowning) and a sudden increase in high-frequency speech energy occur simultaneously, anger can still be detected even if the expression intensity is low, avoiding missed detections.
[0191] The AU action unit, speech feature vector, pupil change rate and basic emotion recognition rules are matched to determine whether the basic emotion categories corresponding to different modalities are consistent. If the multimodal AU action unit, speech feature vector, pupil change rate and basic emotion recognition rules cannot match, a conflict is triggered.
[0192] To achieve conflict resolution, the expression classification probability, sentiment classification probability, intensity level, pupil diameter change rate, and gesture classification results are matched with multimodal conflict resolution rules to output a global sentiment classification rule probability. The multimodal conflict resolution rules are as follows: Figure 5 As shown.
[0193] Forced smile: Changing the emotional category through a weighting formula.
[0194] Suppressing anger: Intensity is calculated through linear combination (although the emotion category is not changed, the intensity is changed, which is actually an adjustment of the emotional state).
[0195] False surprise: Directly changes the emotion category (surprise becomes neutral).
[0196] Conflicting instructions: Instructions are selected for execution based on weights.
[0197] In this embodiment, to address the intermodal decision-making contradiction ("forced smile" scenario), a Zuckerman deception detection four-factor model is introduced to dynamically allocate modal weights and correct the global emotional state.
[0198] When the expression is classified as "smiling" and the emotion is classified as "sad," the forced smile correction rule is triggered:
[0199]
[0200] In the formula: Adjust the weights for the rules to adjust the probability distribution of emotion categories, setting the probability of the sadness category to... Output the global sentiment classification rule probability. Classify the probability of facial expressions. For sentiment classification probability, This represents the rate of change in pupil diameter.
[0201] For example, if the facial expression classifier outputs a probability of 0.9 for "happy," but the voice is "sad" and the pupils are dilated, the system will still classify it as "sad." This rule greatly improves accuracy in conflict scenarios.
[0202] Dynamic interaction strategy rules such as Figure 6 As shown, the emotion recognition results are transformed into personalized interactive behaviors, and a reinforcement mechanism is designed based on Skinner's operant conditioning (such as playing alpha wave music when anxious). The response intensity is adaptively achieved by parameterizing the threshold.
[0203] User profile building rules are as follows: Figure 7 As shown, the module quantifies users' long-term psychological characteristics, such as calculating extraversion based on the Big Five model and monitoring psychological fluctuations using the Emotional Stability Index (ESI). For example, a monthly average decrease of 0.15 in the ESI triggers a recommendation for psychological counseling. The module's prediction results are significantly correlated with the professional scale (…). (), reaching the clinical reference level.
[0204] The multi-task output head generates global emotional states (7 emotion categories + 3 intensity levels), interaction commands (9 command categories), and user profiles (ESI index, etc.) in parallel, achieving a comprehensive analysis from instantaneous emotions to long-term psychological trends. Its necessity lies in the need for children's interactions to simultaneously meet real-time response and personalized adaptation requirements. For example, the emotional intensity rating (threshold 0.3 / 0.7) directly determines the urgency of the soothing strategy, while the ESI index (… By quantifying mood fluctuations, this provides clinical evidence for psychological counseling recommendations (correlation with the scale). The multi-task design avoids the accumulation of delays in phased processing, ensuring that the system completes the "perception-decision-execution" closed loop within ≤28ms.
[0205] Global emotional state: Real-time identification and quantification of the user's current emotional state, outputting the probability distribution and intensity level (low / medium / high) of 7 basic emotions, providing core input for subsequent interaction decisions and user profiling, ensuring fine-grained and robust emotional understanding.
[0206] Output dimensions: 7 basic emotions + 3 intensity levels (high / medium / low).
[0207] Loss function: Improved Focal Loss alleviates class imbalance.
[0208]
[0209] in The predicted probability of sentiment type c. For real labels (one-hot encoded) The focusing parameter (γ) is used to alleviate the imbalance of children's emotion categories (e.g., the scarcity of "disgust" samples) by adjusting the focusing parameter (e.g., γ=2.0). The intensity grading is based on the Sigmoid output value (s∈[0,1]) divided by a threshold (0.3 / 0.7), which conforms to the nonlinear characteristics of human perception in Weber-Fechner's law.
[0210] Interactive command generation: Based on multimodal data (emotional state, voice intensity, gestures, etc.), dynamic interaction strategy rules are matched to dynamically trigger personalized interactive commands (such as playing music or issuing an alarm), achieving real-time conversion from emotion recognition to specific actions, meeting the immediate response needs of child companionship scenarios (such as emergency alarms and emotional soothing). This step utilizes a learnable weight matrix. and bias The fusion features are mapped to the triggering logic of 9 types of instructions.
[0211] Output content: 9 basic commands (such as "play music" and "emergency alarm").
[0212] Triggering conditions: Avoid false triggering by a single modality by using a multimodal feature joint threshold (e.g., voice intensity > 0.8 and gesture = "waving").
[0213] User profile building: This module tracks user emotional fluctuations over a long period and quantifies psychological stability (ESI index) by matching user profile building rules, providing a basis for mental health assessment. It supports long-term mental state monitoring, provides early warnings of potential risks (such as anxiety and depressive tendencies), and facilitates early intervention.
[0214] Output parameters: Age group, Emotional Stability Index (ESI), Social Preference Score. Calculation formula:
[0215]
[0216] in The time window length (600 seconds, or 10 minutes) is set according to the child psychological assessment standards. The number of sentiment categories (7 categories) is used to normalize the fluctuation range to [0,1]. The change in the probability distribution of emotion at time t is represented by the following formula: .
[0217] Specifically, this embodiment combines the Rete algorithm to correct the global sentiment state, and the correction process is as follows: Figure 3 As shown, the specific steps include:
[0218] Step 31: Analyze the fused features using a neural network to generate global sentiment classification prediction probabilities;
[0219] Step 32: Generate AU action units based on RGB-D fusion feature vectors. Use the Rete algorithm to match AU action units, speech feature vectors, pupil change rate and basic emotion recognition rules to determine whether the basic emotion categories corresponding to different modalities are consistent. The basic emotion categories include: happiness, sadness, anger, fear, surprise, disgust and neutrality.
[0220] When the basic emotion categories corresponding to different modalities are inconsistent, proceed to step 33;
[0221] When the basic emotion categories corresponding to different modalities are consistent, the global sentiment classification probability is the global sentiment classification prediction probability.
[0222] Step 33: Match the expression classification probability, emotion classification probability, pupil diameter change rate, and gesture classification result with the multimodal conflict resolution rules, update the emotion classification probability, and output the global emotion classification rule probability;
[0223] Step 34: Calculate the rule trigger weight based on the confidence weight of the expression classification probability and the sentiment classification probability. The calculation formula is as follows:
[0224]
[0225] Weights are triggered by rules. Classify the probability of facial expressions. The probability of sentiment classification is given by α, which is a decay factor determined through experimental optimization (grid search). Final choice Its function is to dynamically reduce the weight of high-confidence modes, thereby preventing a single mode from dominating the decision-making process.
[0226] A non-linear transformation is applied to the probabilities of facial expression classification and sentiment classification. Since α < 1, this transformation will relatively decrease the values with high confidence and relatively increase the values with low confidence (but the absolute values will still be low). For example:
[0227] when When it approaches 1, The value is still close to 1, but slightly less than 1. (For example, 0.9^0.7≈0.93, which is less than 0.9).
[0228] when At lower levels, It will be greater than (For example, 0.5^0.7≈0.62, which is greater than 0.5).
[0229] The purpose of this transformation is to avoid a single high-confidence mode dominating the decision. When multiplying directly, a high confidence level close to 1 will bias the entire product towards that mode. However, after α decay, the influence of high confidence is reduced, making the confidence levels of multiple modes more balanced.
[0230] Step 35: Based on the global sentiment classification rule probability, rule trigger weight, and global sentiment classification prediction probability, calculate the global sentiment classification correction probability. The global sentiment classification probability is the global sentiment classification correction probability.
[0231]
[0232] In the formula: Adjust the probability for global sentiment classification; Weights are assigned to trigger rules; Predict probabilities for global sentiment classification; Global sentiment classification rule probability.
[0233] In this embodiment, multimodal semantic alignment and knowledge constraints are employed to significantly improve the robustness of the companion robot's emotional understanding in complex scenarios, laying a solid technical foundation for personalized interaction. This provides the child companion robot with end-to-end capabilities encompassing "perception-decision-monitoring," balancing immediate interactive experiences with long-term psychological care.
[0234] The dynamic interaction decision layer adopts a dynamic interaction decision method that combines the PPO algorithm (based on the proximal policy optimization algorithm) with psychological rules. Based on the global emotional state (7 types of emotions + 3 levels of intensity), user profile (age, ESI index) and historical interaction data output by the knowledge fusion layer, this module dynamically optimizes multimodal weights (facial expressions, voice, and actions) through reinforcement learning algorithms to generate low-latency, personalized interaction commands and trigger the execution terminal response.
[0235] Specifically, the interactive decision-making method of the dynamic interactive decision-making layer includes the following steps:
[0236] Step 41: Encode the global sentiment state, user profile, and interaction commands into a unified numerical state vector;
[0237] The input global sentiment state, user profile, and interaction commands are encoded into a 128-dimensional vector Q. This step is achieved through... The concatenation operation encodes heterogeneous multimodal data into a unified numerical state vector, providing a comprehensive and structured environmental representation for reinforcement learning.
[0238]
[0239] in : For a unified numerical state vector; Representing the global sentiment probability distribution: the 7 sentiment probabilities (happiness, sadness, etc.) output by the knowledge fusion layer, satisfying... . It is a global strength scalar, output through the Sigmoid regression head, with a value range of [0, 1], and is divided into low (<0.3), medium (0.3-0.7), and high (>0.7) according to the threshold. `age` represents the one-hot encoding of age, derived from the age classification results of the feature processing layer (4 categories: 3-6 years, 7-9 years, 10-12 years, >12 years). For example, 7-9 years is encoded as [0, 1, 0, 0]. `ESI` represents the Emotional Stability Index, and `Hist_cmd` represents the 9-dimensional instruction frequency vector (normalized to [0, 1]). `Hist_weights` represents the 3-dimensional sliding window mean (the weight sequence of the past 5 times).
[0240] Step 42: Employ the PPO algorithm to generate facial expression weights, speech weights, and action weights based on the state vector. Based on the state vector, generate the weight assignments for facial expression, speech, and action to achieve dynamic multimodal interaction strategy optimization. The PPO algorithm includes a policy network (Actor) structure and a value network (Critic) structure.
[0241] Policy network (Actor) structure:
[0242] Input layer (128-dimensional): Receives the state vector.
[0243] Hidden layer (64-dimensional): Fully connected layer extracts high-order features, with ReLU as the activation function.
[0244] Output layer (3D): Outputs unnormalized weights The expression weights, voice weights, and action weights are generated by converting the expression weights into a probability distribution using Softmax.
[0245]
[0246] In the formula: N Used to iterate through three modalities (facial expression, voice, and action). Indicates the first N An unnormalized weight, For unnormalized expression weights, For unnormalized speech weights, These are the unnormalized action weights; For facial expression weight, For speech weights, For action weights.
[0247] Value network (Critic) structure:
[0248] LSTM layer (64 units): Captures temporal dependencies (such as the evolution of emotions in continuous interactions).
[0249] Fully connected layer (1D): Outputs state value V(s), evaluating the long-term benefits of the current policy.
[0250] In this embodiment, the PPO algorithm is used to limit the policy update magnitude (cutting range ϵ=0.2) by replacing the objective function with a cut-off algorithm to prevent training instability. At the same time, the generalized advantage estimation (GAE) balances the single-step TD error and multi-step reward, reduces variance and improves the convergence speed.
[0251] Step 43: Drive the terminal to adjust the voice, facial expression, and action interaction modes based on facial expression weight, voice weight, and action weight.
[0252] To improve the accuracy of the PPO algorithm, step 42 also includes calculating a reward function to generate a reward value, which is used for updating the parameters of the PPO algorithm.
[0253] The reward function quantifies the quality of the policy, guiding the reinforcement learning model towards higher user satisfaction, lower latency, and greater policy stability. Single-objective optimization (such as focusing solely on satisfaction) may lead to policy imbalance (e.g., excessively sacrificing real-time performance), necessitating multi-objective joint optimization. The detailed process is as follows:
[0254] 1) User Satisfaction (SUS Rating): A subjective evaluation that quantifies the interactive experience.
[0255] In the formula: The normalized satisfaction level. Original user rating.
[0256] The raw scores (1-5 points) are normalized, and the results directly reflect the subjective evaluation of children and parents on the interactive experience, driving the strategy to evolve towards higher satisfaction.
[0257] 2) Real-time punishment: Restrict end-to-end delays to prevent children from losing attention.
[0258] In the formula: For delayed penalty items, D The end-to-end latency (unit: ms) is used to constrain decision delay (≤150ms) to avoid loss of attention in children (experiments show that the loss rate increases by 35% when the delay is >150ms).
[0259] 3) Policy stability penalty (Var(a)): Suppresses abrupt changes in action weights to ensure interaction consistency.
[0260] In the formula: Let Variance be the action weights. For a historic moment ti Action weights, This is the average weight of the last 5 actions.
[0261] The variance of the action weight sequence is calculated using the above formula, and abrupt changes in the core inhibition weight (such as a sudden drop in speech weight from 0.7 to 0.3) are used to ensure the continuity of the interaction.
[0262] Output: Reward value R The reward value is a weighted sum of the normalized satisfaction, the delay penalty, and the variance of the action weights, and is used to update the policy network parameters.
[0263] To enhance the adaptability of the interaction process, the dynamic interaction decision layer also includes a weight correction module. The weight correction module determines whether there is a conflict scenario based on the probability of expression classification, the probability of emotion classification, and the rate of change of pupil diameter, according to the Zuckerman deception model. When a conflict scenario exists, the allocation of expression weight, voice weight, and action weight is forcibly corrected.
[0264] 1) Collision detection conditions:
[0265] Confidence difference threshold: (Experimental calibration).
[0266] Physiological signal verification: Pupil diameter change rate (Zuckerman deception model).
[0267] 2) Weighting adjustment formula:
[0268]
[0269] in, The corrected speech weights are represented by the numerator, which is the probability of emotion classification. ) and pupillary change rate ( The product of ) enhances the voice weight in deception scenarios, with the denominator being 1- This reduces the impact of unreliable visual modalities.
[0270] 3) Normalization process:
[0271] Through the Normalization is performed to adjust the weights of facial expressions, voice, and actions to ensure that the weights are 1 after adjustment.
[0272] Output: Corrected facial expression, speech, and motion weights, covering the original output of the PPO network.
[0273] Dynamic interactive decision-making is based on multimodal weight parameters to drive multimodal execution terminals, enabling voice interaction, facial expression interaction, and action interaction, thereby improving the user interaction experience.
[0274] The voice interaction mode adjusts the pitch and speech rate of the synthesized speech based on speech weights;
[0275] The facial expression interaction mode adjusts the amplitude of the terminal's facial expressions based on the weight of the expressions;
[0276] The motion interaction mode adjusts the terminal's response sensitivity to user gestures based on motion weights.
[0277] To effectively safeguard children's mental health, the system also includes a parent platform, which enables the tracking of overall emotional state, user profiles, interactive commands, facial expression weights, voice weights, and action weights; it can also push psychological counseling suggestions based on user profiles.
[0278] To enhance interaction between parents and children, the system also includes a parent-child collaboration platform, the core functions of which are described below:
[0279] Emotional Listening and Recording Module: Utilizing Natural Language Processing (NLP) technology, this module records children's daily conversations in real time and analyzes their language patterns and emotional tendencies. The aim is to provide children with a safe and private space to express themselves, creating an emotional journal to help them articulate their inner feelings.
[0280] Problem and Psychological Advice Report Generation Module: Based on emotional log data and combined with psychological theories, this module uses machine learning algorithms to generate personalized problem and psychological advice reports. The aim is to help parents understand potential emotional or behavioral problems their children may be facing and to provide scientific and feasible suggestions.
[0281] Data Acquisition and Privacy Protection Module: Adopting a low-threshold payment model, parents can obtain detailed data about their children through the API interface; AES-256 encryption technology and access control are used to ensure data security, providing parents with a convenient data access channel while ensuring the privacy of their children.
[0282] Real-time Emotional Status Sharing Module: Through real-time data analysis technology, this module presents children's emotional states to parents in a visual format (such as an emotion heatmap). This allows parents to understand their children's emotional changes in a timely manner and provide emotional support when needed.
[0283] The Growth Trend Analysis and Scientific Parenting Guidance module generates growth trend analysis reports based on long-term communication data and time series analysis algorithms. It combines educational and psychological theories to provide parents with scientific parenting guidance, helping them understand their children's developmental trajectory and offering targeted parenting advice.
[0284] In this embodiment, the parent-child collaboration platform, through functions such as emotional listening, data analysis, and scientific advice, builds a bridge of understanding and communication between parents and children. Its technical effects include:
[0285] 1) Provide a safe and private space for children to express their feelings.
[0286] 2) Through data-driven insights, help parents understand their children's emotional and behavioral problems.
[0287] 3) Ensure the convenience and security of data through a low-threshold payment model and a strict data protection mechanism.
[0288] 4) Provide parents with scientific parenting guidance through real-time sharing of emotional states and analysis of growth trends.
[0289] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0290] The above descriptions are merely preferred examples of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.
Claims
1. A child psychological monitoring system based on multimodal spatiotemporal alignment, characterized in that, The system includes: Data acquisition layer: integrates RGB-D camera, microphone array and infrared sensor, and performs spatiotemporal synchronous acquisition of multimodal data through time alignment and spatial alignment. The multimodal data includes image data, voice data and infrared data. The image data includes color images and depth information of the user's facial expressions and three-dimensional pose; The voice data includes the user's voice signal and voiceprint features; The infrared data includes the user's pupil diameter, proximity status, and gesture commands; Feature processing layer: includes a visual branch module, a speech branch module, and an infrared branch module, which respectively process the image data, speech data, and infrared data to obtain RGB-D fusion feature vector, age classification probability, expression classification probability, speech feature vector, emotion classification probability, intensity level, behavior feature vector, pupil diameter change rate, and gesture classification result; The RGB-D fusion feature vector, speech feature vector, and behavior feature vector constitute a vector information set; The age classification probability, expression classification probability, emotion classification probability, intensity level, pupil diameter change rate, and gesture classification results constitute a structured information set; Feature fusion layer: By fusing the vector information set and the structured information set, fused features are obtained; Knowledge fusion layer: Analyzes the fusion features based on neural networks to generate global sentiment state, user profile, and interaction commands; The global sentiment state includes the global sentiment classification probability and the global intensity level; The user profile includes age and Emotional Stability Index (ESI). Dynamic Interaction Decision Layer: By using PPO reinforcement learning on the global emotional state, user profile, and interaction commands, facial expression weights, voice weights, and action weights are obtained. The facial expression weights, voice weights, and action weights drive the terminal to adjust the voice, facial expression, and action interaction modes. The knowledge fusion layer includes a rule base, which includes dynamic interaction strategy rules and user profile construction rules. The fusion features are analyzed by combining a neural network with the rule base. A multi-task output head is used to generate global sentiment state, user profile, and interaction instructions in parallel. The rule base also includes basic emotion recognition rules and multimodal conflict resolution rules. The basic emotion recognition rules are constructed based on Ekman's basic emotion theory and Russell's circular emotion model. The multimodal conflict resolution rules are constructed based on Zuckerman's deception model and combined with the Rete algorithm to correct the global emotional state.
2. The child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The time alignment is based on the PTP protocol to ensure that the device time bases of the RGB-D camera, microphone array and infrared sensor are consistent, and the time alignment of the multimodal data is achieved through the timestamp mechanism; The spatial alignment unifies the infrared coordinate system and the RGB-D image coordinate system through an affine transformation. The affine transformation formula is as follows: In the formula: ( x , y () are parameters in the infrared coordinate system; ( x `, y `) is a parameter in the RGB-D image coordinate system, with a range of 0 ≤ x <640, 0≤ y <480; a , b , c , d , t x , t y These are affine transformation parameters, calibrated offline using a calibration board.
3. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The visual branch module is based on the improved MobileNet-AGE network architecture, and introduces an inverse residual structure and h-swish activation function to obtain the RGB-D fusion feature vector, age classification probability, and expression classification probability. The age classifications include: 3-6 years old, 7-9 years old, 10-12 years old, and >12 years old; The facial expression categories include: smiling, frowning, crying, staring, surprised, disgusted, and expressionless; The speech branch module adopts a cascaded architecture of 1D CNN and bidirectional GRU to extract the speech feature vector and simultaneously output the emotion classification probability and scalar intensity value. The emotion categories include: happiness, sadness, anger, fear, surprise, disgust, and neutrality; The intensity level is obtained by classifying the scalar intensity value according to a threshold. In the formula: x t The scalar strength value, ; The infrared branch module analyzes the data through the infrared processing API to obtain the gesture classification results, behavioral feature vectors, and pupil diameter change rate. The gesture classification results include: nodding, shaking head, waving, clenching fist, sliding, and stillness.
4. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The feature fusion layer fuses the vector information set and the structured information set through gated cross-attention. The steps for obtaining the fused features include: Step 11: The RGB-D fusion feature vector, age classification probability, and expression classification probability constitute a visual feature set; the speech feature vector, emotion classification probability, and intensity level constitute a speech feature set; the behavior feature vector, pupil diameter change rate, and gesture classification result constitute a behavior feature set. The visual feature set, speech feature set, and behavior feature set are mapped to a unified semantic space by a projection algorithm to obtain projected visual, speech, and behavior feature vectors, which correspond to visual, speech, and behavior modalities, respectively. Step 12: Using the visual, speech, and behavioral modalities as target modalities respectively, calculate the attention weights of the target modalities relative to the other two modalities. The target modalities are used as query modalities, and the other two modalities are used as key-value modalities. After attention aggregation, each target modality yields a first-level fusion feature. Step 13: Calculate the gating weights for each mode; Step 14: Multiply the first-level fusion feature of each modality by the corresponding gating weight to obtain the weighted feature. Then sum the weighted features of all modalities and the projected visual feature vector to obtain the fusion feature.
5. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The feature fusion layer achieves the fusion of multimodal feature vectors in the vector information set through a cross-modal spatiotemporal attention mechanism and a dynamic feature weighting strategy, thereby obtaining fused vector information. The age classification probability, expression classification probability, emotion classification probability, and intensity level in the structured information set are updated through the analysis of the fused vector information by a neural network. The updated structured information set is then fused through gated cross-attention fusion to obtain the fused features. The fusion of multimodal feature vectors in the vector information set includes the following steps: Step 21: Map the RGB-D fused feature vector, speech feature vector, and behavior feature vector to query, key, and value vectors to obtain the query vector of the RGB-D fused feature vector, the key vector of the speech feature, and the key vector of the behavior feature; Step 22: Based on the query vector of the RGB-D fusion feature vector, the key vector of the speech feature, and the key vector of the behavior feature, calculate the attention weights of the RGB-D fusion feature vector on the speech feature vector and the behavior feature vector, respectively. Step 23: Obtain the multimodal fusion feature vector by applying attention weights to the speech feature vector and behavior feature vector respectively based on the RGB-D fusion feature vector; Step 24: Compress the multimodal fusion feature vector using 1×1 convolution to obtain fusion vector information.
6. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The step of modifying the global sentiment state using the Rete algorithm includes: Step 31: Analyze the fused features using a neural network to generate a global sentiment classification prediction probability; Step 32: Generate AU action units based on the RGB-D fusion feature vector, and use the Rete algorithm to match the AU action units, speech feature vector, pupil change rate with the basic emotion recognition rules to determine whether the basic emotion categories corresponding to different modalities are consistent. The basic emotion categories include: happiness, sadness, anger, fear, surprise, disgust, and neutrality. When the basic emotion categories corresponding to different modalities are inconsistent, proceed to step 33; When the basic emotion categories corresponding to different modalities are consistent, the global emotion classification probability is the global emotion classification prediction probability; Step 33: Match the expression classification probability, emotion classification probability, pupil diameter change rate, and gesture classification result with the multimodal conflict resolution rule, update the emotion classification probability, and output the global emotion classification rule probability; Step 34: Calculate the rule trigger weight by weighting the confidence level based on the expression classification probability and the emotion classification probability; Step 35: Based on the global sentiment classification rule probability, rule trigger weight, and global sentiment classification prediction probability, calculate the global sentiment classification correction probability, whereby the global sentiment classification probability is the global sentiment classification correction probability.
7. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The multimodal conflict resolution rules include: When the expression is classified as "smiling" and the emotion is classified as "sad," the forced smile correction rule is triggered: In the formula: Adjust the weights for the rules to adjust the probability distribution of emotion categories, setting the probability of the sadness category to... Output the global sentiment classification rule probability. Classify the probability of facial expressions. For sentiment classification probability, This represents the rate of change in pupil diameter.
8. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The interactive decision-making method of the dynamic interactive decision-making layer includes the following steps: Step 41: Encode the global emotional state, user profile, and interaction commands into a unified numerical state vector; Step 42: Using the PPO algorithm, generate the expression weights, voice weights, and action weights based on the state vector; Step 43: Based on the facial expression weight, voice weight, and action weight, drive the terminal to adjust the voice, facial expression, and action interaction modes.
9. A child psychological monitoring system based on multimodal spatiotemporal alignment according to claim 1, characterized in that, The dynamic interaction decision layer also includes a weight correction module. Based on the expression classification probability, emotion classification probability, and pupil diameter change rate, the weight correction module determines whether there is a conflict scenario according to the Zuckerman deception model. When a conflict scenario exists, the module forcibly corrects the allocation of expression weight, voice weight, and action weight.
Citation Information
Patent Citations
A Childhood ADHD Screening and Assessment System Based on Multimodal Deep Learning Technology
CN111528859B
Child ADHD screening evaluation system based on a multi-modal deep learning technology
CN111528859A
Intelligent voice nursing system and nursing method based on physiological and emotional characteristics
CN116564561A