A comprehensive emotion recognition system integrating biosignals and behavioral features

By integrating biosignals and behavioral characteristics into a full-spectrum emotion recognition system, we have achieved accurate identification and personalized intervention for different levels of emotional problems. This solves the problem of insufficient full-spectrum coverage in existing technologies, improves the accuracy and applicability of emotion recognition, and supports continuous monitoring and management of emotions.

CN122123701APending Publication Date: 2026-06-02SUZHOU KUYUE NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU KUYUE NETWORK TECH CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current emotion recognition technologies cannot achieve full spectrum coverage, cannot accurately identify different levels of emotional problems, and lack a unified solution for ordinary children and special groups such as ASD.

Method used

By integrating biosignals and behavioral features, using multimodal data synchronization with unified timestamp calibration, and combining a CNN-Transformer hybrid model and a lightweight MobileNetV3-LSTM model, we can accurately identify the type and intensity of emotions, and then divide them into three groups based on the identification results for personalized intervention.

Benefits of technology

It achieves full-spectrum emotion recognition, ranging from ordinary emotional distress to severe emotional disorders, accurately distinguishes different degrees of emotional problems, provides personalized intervention measures, improves the accuracy and applicability of emotion recognition, reduces discomfort for children, and supports continuous monitoring and timely intervention of emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122123701A_ABST
    Figure CN122123701A_ABST
Patent Text Reader

Abstract

This invention discloses a full-spectrum emotion recognition system integrating biosignals and behavioral features, belonging to the field of emotion recognition technology. The system includes the following steps: S1, data acquisition; S2, data preprocessing; S3, emotion feature fusion; S4, emotion analysis and output; S5, full-spectrum adaptation. This full-spectrum emotion recognition system integrates biosignals and behavioral features to achieve full-spectrum emotion recognition, ranging from common emotional distress to severe emotional disorders. By integrating multi-dimensional data of biosignals and behavioral features, and based on scientific classification standards, it accurately distinguishes children with different degrees of emotional problems, allowing each group to receive appropriate intervention. This precise classification avoids the blindness of intervention measures, providing basic regulatory support for children with common emotional distress and securing timely intervention opportunities for children with severe emotional disorders. It comprehensively meets the emotional management needs of different children and promotes their healthy emotional development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of emotion recognition technology, specifically to a full-spectrum emotion recognition system that integrates biological signals and behavioral features. Background Technology

[0002] Insufficient emotional management skills have become a core issue that has a wide-ranging impact on children's development, encompassing the entire spectrum of needs from emotional regulation difficulties in ordinary children to emotional deficits in special groups such as autism spectrum disorder (ASD). Therefore, it is necessary to be able to identify children's emotional types and intervene and adjust emotions in a timely manner to prevent negative emotions from deepening.

[0003] Existing emotion recognition technologies are often designed for a single group, focusing on emotional distress in ordinary children or limited to special groups such as ASD, lacking the ability to cover the entire spectrum and failing to accurately classify different levels of emotional problems. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a full-spectrum emotion recognition system that integrates biological signals and behavioral features, thus solving the problems mentioned in the background section.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a full-spectrum emotion recognition system integrating biological signals and behavioral characteristics, comprising the following steps: S1. Data Acquisition: Collect children's biosignal data and behavioral characteristic data, and achieve multimodal data synchronization through unified timestamp calibration; S2, Data Preprocessing: Outlier processing, standardization, and time alignment are performed on the collected biosignal and behavioral feature data to form a complete emotion data sample; S3, Emotional Feature Fusion: Extract unimodal features, calculate adaptive weights for each modality feature based on the entropy weight method, and achieve cross-modal feature fusion through a CNN-Transformer hybrid model to output a unified feature vector; S4. Emotion Analysis and Output: A lightweight MobileNetV3-LSTM model is used to analyze the fused feature vectors and output the emotion type and emotion intensity. S5 full spectrum compatibility: The severity of emotional problems is assessed based on the final output of emotion type and intensity. Based on the identification results, the identified subjects are divided into three groups so that corresponding intervention measures can be taken for emotion regulation based on the specific group type.

[0006] Furthermore, in step S1, the biosignal data acquisition includes: Heart rate (HRV) acquisition: A flexible fabric wristband with a built-in PPG photoplethysmography sensor is worn on the inside of the child's wrist to convert heart rate variability data. Skin conductance EDA acquisition: A flexible conductive fabric chest patch is used to fit the child's chest skin to collect skin conductance values; EEG acquisition: A non-contact millimeter-wave radar device is deployed at the top of the intervention area, 1.5-2 meters away from the child's head, for data acquisition. Wave (8-13Hz), Wave (14-30Hz) signal.

[0007] Furthermore, the SDNN formula for calculating the central rate variability, a core indicator of biological signals, is as follows: in, For the first indivual Interval; For average Interval; for Total number of intervals; The peak frequency of the power spectrum is calculated by converting the EDA time-domain signal to the frequency domain using a Fast Fourier Transform (FFT) to extract the frequency corresponding to the peak power spectrum. The specific formula is as follows: Where, x The function representing the change of skin electrical activity signal over time reflects the change at any given moment. The state of skin electrical activity; It is a complex exponential function, used as the kernel function of the Fourier transform, through interaction with the original time-domain signal x. Perform integration operations to achieve the conversion from the time domain to the frequency domain; Indicates time An infinitesimal increment; EEG , The formula for calculating the wave power ratio is as follows: in, express frequency band and The frequency band EEG energy ratio is used to quantify the relative intensity of two frequency band EEG signals; EEG power spectral density describes the EEG signal at different frequencies. Power distribution under; Indicates the power spectral density of brain waves exist The integral of the frequency band is used to calculate the total energy of the EEG signal within that frequency band. Indicates the power spectral density of brain waves exist The integral of the frequency band is used to calculate the total energy of the EEG signal within that frequency band.

[0008] Furthermore, in step S1, the behavioral feature data collection utilizes a wearable, motion-sensing smart screen, specifically including: The motion-sensing smart screen has a built-in 3D depth camera to capture data from 342 skeletal points; the collected metrics include motion features, facial features, and voice features. The motion features include the three-dimensional coordinates of skeletal points, the standardization of the motion, and the duration; the facial expression features include the displacement and rate of change of four key points at the corners of the eyes and six key points at the corners of the mouth; and the speech features include the fundamental frequency of 80-500Hz, speech rate, and the frequency of emotional keywords. Furthermore, the aforementioned data is synchronized with the biosignal acquisition equipment via the NTP protocol to calibrate the timestamps.

[0009] Furthermore, in step S2, during outlier handling, the behavioral feature data uses 3D modeling. Outliers were identified using cubic spline interpolation; outliers in biosignals (HRV and EDA) were identified using box plots and corrected using median replacement; outliers in EEG were removed using independent component analysis (ICA), with the objective function being: in, Used to constrain the separation matrix to ensure that the transformation of the matrix does not cause unreasonable scaling of the signal energy; This is a separation matrix used to separate the skin electrodermal activity signal x. The goal of converting it into independent components is to optimize the objective function. This ensures that the converted components are independent of each other; This represents the total number of time points in the data sample, i.e., the number of EEG data samples collected. Used to measure the non-Gaussianity of the converted signal; During standardization, behavioral feature data is standardized using Min-Max. Biosignal data were mapped to the [0,1] interval and normalized using Z-score. ,in, Represents a single sample value in the original data; It is the maximum value in this set of behavioral feature data; It is the minimum value in this set of behavioral feature data; This represents the mean of the set of biological signal data; This represents the standard deviation of the set of biological signal data.

[0010] Furthermore, in step S2, time alignment employs the Dynamic Time Warping (DTW) algorithm to construct a distance matrix: The goal is to find the optimal path that minimizes the total path distance, where... Represents the first time series, the first... The first time point Dimensional feature values, which correspond to any specific dimension of biological signals or behavioral characteristics; Represents the second time series. The first time point dimensional eigenvalues, which are related to They are in the same feature dimension; This indicates the number of feature dimensions of a time series, that is, the total number of data features contained at each point in time.

[0011] Furthermore, in step S3, the emotion feature fusion specifically includes: Single-modal feature extraction: Action features are extracted into 128-dimensional vectors using the ResNet18 network; Facial expression features are extracted into 64-dimensional features using the HOG algorithm; Speech features are extracted into 39-dimensional features using MFCC; Biosignal features are extracted into 64-dimensional vectors using a CNN-LSTM hybrid model. Information entropy is used to measure the degree of disorder of features. The smaller the information entropy, the more effective information the feature carries, the higher its importance in emotion recognition, and the greater its corresponding weight. Therefore, feature weights are calculated based on the entropy weight method. The formula for calculating information entropy is as follows: in, For the first Information entropy of each feature; Indicates the number of samples; For the first In the nth sample The proportion of each eigenvalue must satisfy the following: The weight calculation formula is as follows: in, For the first The final weights of each feature; Information utility value is a quantitative representation of the amount of characteristic information; The total number of features.

[0012] Furthermore, cross-modal feature fusion is used to integrate multimodal data to achieve more accurate emotion recognition, which includes two stages; The first stage involves achieving same-dimensional weighted fusion using the following formula: in, This represents the features obtained after weighted fusion along the same dimension; Represents the number of features of the same dimension participating in the fusion; For the first Each feature vector has the same dimension; The second stage involves calculating cross-modal weights using an attention mechanism, with the specific formula as follows: in, Cross-modal attention weights; , These represent feature vectors of two different modalities; It is a learnable weight matrix used to adjust the mapping relationship between feature vectors; The function normalizes the dot product of the eigenvectors, converting it into a probability distribution form; Final output The 128-dimensional fused feature vector.

[0013] Furthermore, the specific process of step S4 is as follows: Fusion feature vectors As input to the MobileNetV3-LSTM model, the fused feature vector is used to perform deep feature extraction to extract representative features. After processing by MobileNetV3, a low-dimensional feature representation containing key information is output. ; Features output by MobileNetV3 The input is fed into the LSTM network; the computation process of the LSTM network is as follows: Forgotten Gate: in, yes The forget gate outputs at any given moment; Use the Sigmoid activation function; It is the weight matrix of the forget gate; The engraved hidden state; It is the input at the current moment; It is the bias of the forget gate; the forget gate determines which information from the previous moment will be retained in the current moment; Input Gate: in, yes Input gate output at any given time; It is the weight matrix of the input gate; It is the bias of the input gate; the input gate determines what new information will be added to the cell state at the current moment. Cell status update: in, It represents the current state of the candidate cells; It is the weight matrix for cell state updates; It is a bias; It is the hyperbolic tangent activation function; in, yes The updated cell state combines information retained from the previous time step with new information from the current time step; Output gate: in, yes The output gate outputs at any given time; It is the weight matrix of the output gate; It is the bias of the output gate; the output gate determines which information in the cell state will be output as the hidden state at the current moment; Hidden state: in, yes The hidden state at each time step; after processing by LSTM It contains temporal feature information of the input sequence; The hidden state at the last moment The input is fed into a fully connected layer, which maps the features to the output space of emotion type and emotion intensity. Let the emotion types be... The probability distribution of emotion types is calculated through a fully connected layer. ,in, It is the weight matrix of the fully connected layer; It is a bias; the emotion intensity output is a continuous value, calculated through a fully connected layer. ,in, It is a weight matrix; It is a bias; The final model outputs the emotion type. With emotional intensity .

[0014] Furthermore, in step S5, the three groups are children with ordinary emotional distress, children with mild emotional deficits, and children with severe emotional disorders such as ASD. Criteria for classifying children with common emotional distress: The frequency of negative emotions occurring ≤30% in the past 30 days; the average intensity of negative emotions... And single peak ; Criteria for classifying children with mild emotional deficits: The frequency of negative emotions occurring in 30% to 60% of the past 30 days; the average intensity of negative emotions. or single peak ; Criteria for classifying children with severe mood disorders such as ASD: Frequency of negative emotions occurring ≥60% in the past 30 days; mean intensity of negative emotions. or single peak ; Step S3 also includes the following sub-steps: Based on the physiological and behavioral response patterns of children's emotions, the time boundaries of the three-level characteristics are clearly defined: the immediate characteristic corresponds to 0-3 seconds after the emotional trigger, which is the instantaneous reaction stage; the short-term characteristic corresponds to 3-30 seconds, which is the emotional duration stage; and the cumulative characteristic corresponds to more than 30 seconds, which is the emotional stabilization or decay stage. Focusing on abrupt fluctuation indicators in biological signals and transient actions and vocal bursts in behavioral characteristics, abrupt fluctuation indicators include EDA instantaneous peak values, HRV abrupt changes, and EEG values. , Wave power ratio abrupt change value; instantaneous actions in behavioral characteristics include sudden limb swings and sudden facial muscle tension; vocal burst signals include abrupt changes in tone and instantaneous frequency of emotional keywords; a lightweight CNN network is used to quickly extract 32-dimensional instantaneous feature vectors to highlight key signals at the moment of emotional triggering; Using a 10-second sliding window and a 5-second step, time-series slices are made for each modality of data to extract trend features within the window, including HRV fluctuation slope, EDA rise or fall rate, speech intonation change trend, and cumulative displacement of facial expression key points. Temporal dependencies are captured through a single-layer LSTM network to output a 64-dimensional short-term temporal feature vector that reflects the changing patterns during the duration of emotion. It fully adopts the original single-modal feature extraction logic, extracts a 128-dimensional cumulative feature vector, ensures the stability and integrity of long-term emotional features, and seamlessly connects with the original process; A basic weighting framework for three layers of features is established, with short-term features having a fixed weight of 30%, and the sum of the weights for immediate and cumulative features being 70%, dynamically allocated based on the current emotion intensity; when the real-time output emotion intensity... When the score is ≥7, indicating a high level of negative emotion or a strong positive emotion, the weight of the immediate feature is increased to 30%, and the weight of the cumulative feature is reduced to 40%, thus strengthening the decision weight of the instantaneous explosive emotional signal. When emotional intensity When the emotion is determined to be of moderate intensity, the weight of the immediate feature is adjusted to 20%, and the weight of the cumulative feature is 50%, to balance the contribution of the instantaneous reaction and the long-term cumulative feature. When emotional intensity <4, when judged as a low-intensity emotion or calm state, the weight of the immediate feature drops to 10%, and the weight of the cumulative feature rises to 60%, highlighting the dominant role of long-term stable emotional features. Then, layered fusion is performed. The first step is to fuse immediate features and short-term features. An attention mechanism is used to calculate cross-layer weights. The formula is based on the cross-modal attention fusion logic, as follows: in, It is a 32-dimensional instantaneous feature vector; It is a 64-dimensional short-term eigenvector; The weight matrix is ​​a learnable matrix with dimensions 32×64. After obtaining the cross-layer attention weights using this formula, they are fused to generate a 96-dimensional dynamic feature vector. The second step is to fuse dynamic features and cumulative features. This is done by weighting the dynamic features based on the weight ratios determined in the first step, using the following formula: in, For dynamic feature weights; For cumulative feature weights; It is a 128-dimensional cumulative feature vector; during the fusion process, the dimension adaptation matrix is ​​used to... Map to 128 dimensions to ensure dimensional consistency; The third step is to standardize the fusion results. L2 normalization is performed to output a 128-dimensional final fused feature vector. This vector has the same dimension as the feature vector output from the cross-modal fusion above, and is directly used as the input for subsequent cross-modal fusion in step S3 without changing the processing logic of the original CNN-Transformer hybrid model.

[0015] This invention provides a full-spectrum emotion recognition system that integrates biological signals and behavioral features, and has the following beneficial effects: 1. This comprehensive emotion recognition system, integrating biosignals and behavioral characteristics, achieves full-spectrum emotion identification, ranging from common emotional distress to severe emotional disorders. By integrating multi-dimensional data on biosignals and behavioral characteristics, and based on scientific classification standards, it accurately distinguishes children with different degrees of emotional problems, ensuring that each group receives appropriate intervention. This precise classification avoids the blind application of intervention measures, providing basic regulatory support for children with common emotional distress and securing timely intervention opportunities for children with severe emotional disorders. It comprehensively meets the emotional management needs of different children, contributing to their healthy emotional development.

[0016] 2. This comprehensive emotion recognition system, integrating biosignals and behavioral characteristics, deeply mines effective information through multimodal feature fusion and optimization algorithms, significantly improving the accuracy of emotion type and intensity identification. The data collection device employs a flexible, wearable, and contactless design, conforming to children's physiological characteristics and behavioral habits, reducing discomfort and resistance. Simultaneously, the recognition results are intuitive and easy to understand, providing schools and families with clear decision-making support and facilitating rapid, targeted intervention. Its lightweight deployment characteristics also make the system easier to apply in everyday scenarios, enabling continuous monitoring and timely intervention of children's emotions, effectively helping to improve children's emotional problems.

[0017] 3. This full-spectrum emotion recognition system integrates biosignals and behavioral features. Children's emotions are characterized by both instantaneous bursts and phased persistence, and different time dimensions have different impacts on recognition. Therefore, a three-layer fusion architecture consisting of instantaneous features, short-term features, and cumulative features is adopted. First, the time dimension is divided according to the emotional response pattern and corresponding features are extracted. Then, a weight benchmark is set, and the weights of each layer are dynamically adjusted according to the emotional intensity, and the reasonableness is ensured by information entropy verification. Subsequently, the features are fused in three steps: first, instantaneous and short-term features are fused through the attention mechanism, then dynamic weights and cumulative features are combined for weighted fusion, and finally, the output is standardized. This process is highly compatible with the original cross-modal fusion, can completely capture the dynamic process of children's emotions, and significantly improve the accuracy of instantaneous emotion recognition and the accuracy of phased emotion trend judgment. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the steps of a full-spectrum emotion recognition system that integrates biological signals and behavioral features, as described in this invention. Detailed Implementation

[0019] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.

[0020] like Figure 1 As shown, the present invention provides a technical solution: a full-spectrum emotion accurate recognition system that integrates biological signals and behavioral features, comprising the following implementation steps: Wear a flexible fabric wristband with a built-in PPG photoplethysmography sensor on the inside of the child's wrist (fitting against the skin without pressure and avoiding hair obstruction), and a flexible conductive fabric chest patch fitted against the center of the child's chest (ensuring that the sensing area is not obstructed by clothing). Fix the non-contact millimeter-wave radar device at the top of the intervention scene, adjusting the height to be 1.5-2m away from the child's head, and ensuring that the radar detection direction is vertically aligned with the child's head area. Place the wearable motion-sensing smart screen 1.5-2m directly in front of the child, adjusting the screen height to be level with the child's line of sight. Ensure that the 3D depth camera can completely capture the child's entire skeletal structure (without limb obstruction) and the facial and voice capture range. Activate all acquisition devices and synchronize the biosignal acquisition devices with the motion-sensing smart screen through the NTP timestamp calibration interface of the data synchronization calibration module. Ensure that the multimodal data time synchronization error is ≤10ms, and simultaneously complete device self-checks (sensor signal stability, camera capture integrity). The flexible wristband collects children's heart rate variability (HRV) data at a sampling frequency of 250Hz and converts it into... Interphase sequence; conductive chest patch acquires skin conductance (EDA signal) at a sampling frequency of 100Hz; millimeter-wave radar acquires at a sampling frequency of 500Hz. Wave (8-13Hz), The system collects EEG signals (14-30Hz); the 3D depth camera of the motion-sensing smart screen captures the three-dimensional coordinates of 342 skeletal points at 30fps, and simultaneously extracts motion features (changes in skeletal point coordinates, motion standardization, and duration) and facial expression features (displacement and rate of change of 4 key points at the corners of the eyes and 6 key points at the corners of the mouth); and it collects speech signals through a built-in microphone at a sampling frequency of 16kHz, extracting the fundamental frequency (80-500Hz), speech rate, and frequency of emotional keywords (based on matching a database of commonly used emotional words for children). Behavioral characteristic data uses 3 Outliers were identified using criteria, and missing data were filled in using cubic spline interpolation. Outliers in HRV and EDA data were identified using box plots and corrected using median replacement. EEG data underwent independent component analysis (ICA) to remove artifacts related to electrooculography and electromyography, and the separation matrix was optimized according to the ICA objective function. To achieve artifact separation, the ICA objective function is as follows: in, Used to constrain the separation matrix to ensure that the transformation of the matrix does not cause unreasonable scaling of the signal energy; This is a separation matrix used to separate the skin electrodermal activity signal x. The goal of converting it into independent components is to optimize the objective function. This ensures that the converted components are independent of each other; This represents the total number of time points in the data sample, i.e., the number of EEG data samples collected. Used to measure the non-Gaussianity of the converted signal; Behavioral feature data were standardized using the Min-Max formula. Mapped to the [0,1] interval to eliminate dimensional differences between different feature dimensions; biosignal data were standardized using the Z-score formula. Transforming the data into a standard normal distribution with a mean of 0 and a standard deviation of 1 improves the model's convergence speed. Time alignment employs the Dynamic Time Warping (DTW) algorithm to construct the distance matrix: The goal is to find the optimal path that minimizes the total path distance, where... Represents the first time series, the first... The first time point Dimensional feature values, which correspond to any specific dimension of biological signals or behavioral characteristics; Represents the second time series. The first time point dimensional eigenvalues, which are related to They are in the same feature dimension; This indicates the number of feature dimensions in a time series, i.e., the total number of data features contained at each point in time. Action features were extracted into 128-dimensional feature vectors using a ResNet18 network (batch size 32, iterations 50); facial expression features were extracted into 64-dimensional feature vectors using the HOG algorithm (cell size 8×8, block size 16×16); speech features were extracted into 39-dimensional feature vectors using the MFCC algorithm (13-dimensional MFCC coefficients + first-order difference + second-order difference); and biosignal features were extracted into 64-dimensional feature vectors using a CNN-LSTM hybrid model (containing 2 convolutional layers and 1 128-neuron LSTM hidden layer). Information entropy is used to measure the degree of disorder of features. The smaller the information entropy, the more effective information the feature carries, the higher its importance in emotion recognition, and the greater its corresponding weight. Therefore, feature weights are calculated based on the entropy weight method. The formula for calculating information entropy is as follows: in, For the first Information entropy of each feature; Indicates the number of samples; For the first In the nth sample The proportion of each eigenvalue must satisfy the following: The weight calculation formula is as follows: in, For the first The final weights of each feature; Information utility value is a quantitative representation of the amount of characteristic information; The total number of features; Cross-modal feature fusion is used to integrate multimodal data to achieve more accurate emotion recognition, and it includes two stages; The first stage involves achieving same-dimensional weighted fusion using the following formula: in, This represents the features obtained after weighted fusion along the same dimension; Represents the number of features of the same dimension participating in the fusion; For the first Each feature vector has the same dimension; The second stage involves calculating cross-modal weights using an attention mechanism, with the specific formula as follows: in, Cross-modal attention weights; , These represent feature vectors of two different modalities; It is a learnable weight matrix used to adjust the mapping relationship between feature vectors; The function normalizes the dot product of the eigenvectors, converting it into a probability distribution form; Final output The 128-dimensional fused feature vector; The severity of emotional problems is assessed based on the final output of emotion type and intensity, and the identified subjects are divided into three groups based on the identification results, so that corresponding intervention measures can be taken for emotion regulation based on the specific group type; The three groups are children with common emotional distress, children with mild emotional deficits, and children with severe emotional disorders such as ASD; Criteria for classifying children with common emotional distress: The frequency of negative emotions occurring ≤30% in the past 30 days; the average intensity of negative emotions... And single peak ; Criteria for classifying children with mild emotional deficits: The frequency of negative emotions occurring in 30% to 60% of the past 30 days; the average intensity of negative emotions. or single peak ; Criteria for classifying children with severe mood disorders such as ASD: Frequency of negative emotions occurring ≥60% in the past 30 days; mean intensity of negative emotions. or single peak ; The above group classification criteria are consistent with the range of negative emotion intensity and scores on the Children's Emotional Ability Scale. Based on the physiological and behavioral response patterns of children's emotions, the time boundaries of the three-level characteristics are clearly defined: the immediate characteristic corresponds to 0-3 seconds after the emotional trigger, which is the instantaneous reaction stage; the short-term characteristic corresponds to 3-30 seconds, which is the emotional duration stage; and the cumulative characteristic corresponds to more than 30 seconds, which is the emotional stabilization or decay stage. Focusing on abrupt fluctuation indicators in biological signals and transient actions and vocal bursts in behavioral characteristics, abrupt fluctuation indicators include EDA instantaneous peak values, HRV abrupt changes, and EEG values. , Wave power ratio abrupt change value; instantaneous actions in behavioral characteristics include sudden limb swings and sudden facial muscle tension; vocal burst signals include abrupt changes in tone and instantaneous frequency of emotional keywords; a lightweight CNN network is used to quickly extract 32-dimensional instantaneous feature vectors to highlight key signals at the moment of emotional triggering; Using a 10-second sliding window and a 5-second step, time-series slices are made for each modality of data to extract trend features within the window, including HRV fluctuation slope, EDA rise or fall rate, speech intonation change trend, and cumulative displacement of facial expression key points. Temporal dependencies are captured through a single-layer LSTM network to output a 64-dimensional short-term temporal feature vector that reflects the changing patterns during the duration of emotion. It fully adopts the original single-modal feature extraction logic, extracts a 128-dimensional cumulative feature vector, ensures the stability and integrity of long-term emotional features, and seamlessly connects with the original process; A basic weighting framework for three layers of features is established, with short-term features having a fixed weight of 30%, and the sum of the weights for immediate and cumulative features being 70%, dynamically allocated based on the current emotion intensity; when the real-time output emotion intensity... When the score is ≥7, indicating a high level of negative emotion or a strong positive emotion, the weight of the immediate feature is increased to 30%, and the weight of the cumulative feature is reduced to 40%, thus strengthening the decision weight of the instantaneous explosive emotional signal. When emotional intensity When the emotion is determined to be of moderate intensity, the weight of the immediate feature is adjusted to 20%, and the weight of the cumulative feature is 50%, to balance the contribution of the instantaneous reaction and the long-term cumulative feature. When emotional intensity <4, when judged as a low-intensity emotion or calm state, the weight of the immediate feature drops to 10%, and the weight of the cumulative feature rises to 60%, highlighting the dominant role of long-term stable emotional features. Then, layered fusion is performed. The first step is to fuse immediate features and short-term features. An attention mechanism is used to calculate cross-layer weights. The formula is based on the cross-modal attention fusion logic, as follows: in, It is a 32-dimensional instantaneous feature vector; It is a 64-dimensional short-term eigenvector; The weight matrix is ​​a learnable matrix with dimensions 32×64. After obtaining the cross-layer attention weights using this formula, they are fused to generate a 96-dimensional dynamic feature vector. The second step is to fuse dynamic features and cumulative features. This is done by weighting the dynamic features based on the weight ratios determined in the first step, using the following formula: in, For dynamic feature weights; For cumulative feature weights; It is a 128-dimensional cumulative feature vector; during the fusion process, the dimension adaptation matrix is ​​used to... Map to 128 dimensions to ensure dimensional consistency; The third step is to standardize the fusion results. L2 normalization is performed to output a 128-dimensional final fused feature vector. This vector has the same dimension as the feature vector output by the cross-modal fusion above, and is directly used as the input for subsequent cross-modal fusion in step S3 without changing the processing logic of the original CNN-Transformer hybrid model. Among them, the cumulative features directly reuse the original single-modal feature extraction results in step S3, the weight allocation is based on the logical extension of the information utility value of the original entropy weight method, and the fusion method adopts the attention mechanism and weighted fusion framework, forming a two-level architecture consisting of hierarchical preprocessing and cross-modal integration with the original cross-modal fusion process. There is no technical conflict between the two. Children's emotions are characterized by both instantaneous bursts and phased persistence. Different time-dimensional features have varying impacts on recognition. To address this, a three-layer fusion architecture consisting of instantaneous features, short-term features, and cumulative features is employed. First, the time dimension is divided according to the emotional response pattern, and corresponding features are extracted. Then, a weight benchmark is set, and the weights of each layer are dynamically adjusted according to the intensity of the emotion, with information entropy verification to ensure reasonableness. Subsequently, features are fused in three steps: first, instantaneous and short-term features are fused through an attention mechanism; then, dynamic weights and cumulative features are combined for weighted fusion; finally, the output is standardized. This process is highly compatible with existing cross-modal fusion methods, can fully capture the dynamic process of children's emotions, and significantly improves the accuracy of instantaneous emotion recognition and the accuracy of phased emotion trend judgment.

[0021] Example: Raw biosignal data: HRV: Interphase sequence (unit: ms): 820, 815, 830, 805, 825, 818, 840, 832, 821, 810 (first 10 data points extracted); EDA: Skin conductivity (unit: μS): 1.2, 1.3, 1.5, 1.4, 1.6, 1.5, 1.7, 1.6, 1.8, 1.7; EEG: Wavelength (8-13Hz) power: 4.2μV² / Hz Wavelength (14-30Hz) power: 2.8μV² / Hz and The ratio = 1.5; Raw data on behavioral characteristics: Movement characteristics: 3D skeletal points (left shoulder: x=320, y=240, z=150; right shoulder: x=380, y=242, z=148), shoulder movement variation ≤5cm within 10 seconds; Facial features: The displacement of the key point at the corner of the eye (left outer canthus: x=290, y=210) was 0.3cm, and the displacement of the key point at the corner of the mouth (right corner of the mouth: x=350, y=260) was 0.5cm, with a rate of change of 0.05cm / s; Voice characteristics: base frequency 180Hz, speech rate 120 words / minute, emotional keywords "unhappy" appear twice and "don't want to play" appear once; After processing, the above data was used to construct a distance matrix using the DTW algorithm. The total distance of the optimal path was 12.3. After alignment, the data was divided into 1800 data samples with a step size of 100ms. Single-modal feature extraction results: Action features: ResNet18 outputs a 128-dimensional vector with dimensions of [0.12, 0.35, ..., 0.28] (dimensions 3 to 126 are omitted); Facial features: The HOG algorithm outputs a 64-dimensional vector, with the first 5 dimensions being: [0.08, 0.15, 0.22, 0.18, 0.09]; Speech features: MFCC outputs a 39-dimensional vector, with 13-dimensional coefficient mean values: [-2.1, -1.8, ..., 3.2]; Biological signal characteristics: CNN-LSTM outputs a 64-dimensional vector, with EEG-related dimension values ​​of [0.42, 0.38, ..., 0.51]; Feature weight calculation (entropy weight method): Information entropy calculation for each modality: EEG (0.32), HRV (0.41), EDA (0.45), action (0.58), expression (0.62), speech (0.65); Final weighting: EEG (0.28), HRV (0.22), EDA (0.18), motion (0.15), facial expression (0.10), voice (0.07); Cross-modal fusion: First-stage weighted summation: F1 = 0.28 × EEG features + 0.22 × HRV features + ... + 0.07 × speech features; Second-stage attention adjustment: diagonal values ​​of the cross-modal weight matrix [0.35, 0.28, ..., 0.12], ultimately outputting 128-dimensional fused features. Example dimensions: [0.25, 0.31, ..., 0.29]; MobileNetV3-LSTM model input Inference time: 120ms (edge ​​terminal deployment); Result calculation: Emotion type probability distribution: happy (0.08), angry (0.05), anxious (0.12), depressed (0.65), fear (0.03), calm (0.07); Emotion intensity value =6.8 (range 1-10); The smart screen displays the following emotion type: depressed; intensity: 6.8 / 10, progress bar filled to 68%, and icon showing "depressed expression"; Based on the child's data over the past 30 days, negative emotions occurred 42% of the time (30%-60% range), with an average intensity of 5.3 (4-7 range), classifying the child as having "mild emotional deficits." Intervention suggestions were provided. School-based: A counselor will provide one-on-one counseling once a week (with a focus on managing low mood). For families: We recommend reading the picture book "My Little Emotional Monster" for emotional regulation with your child for 15 minutes daily. Priority rating: Low mood regulation has a "high" priority and requires continuous monitoring for 2 weeks. Implementation Results: Through manual observation and professional scale assessment (Children's Depression Scale, CDI), the identification results were 92% consistent with the actual emotional state; the power consumption throughout the process was ≤10W (smart screen and wearable device), and the data transmission latency was ≤200ms, meeting the requirements for lightweight deployment in the home; after the intervention suggestions were adopted by parents and schools, the child's low mood intensity dropped to 4.2 and the frequency dropped to 28% after a retest two weeks later, and the child was classified as "general emotional distress".

[0022] Based on the above description, this invention achieves full-spectrum emotion recognition, ranging from common emotional distress to severe emotional disorders. By integrating multi-dimensional data of biosignals and behavioral characteristics, and according to scientific classification standards, it accurately distinguishes children with different degrees of emotional problems, allowing each group to receive appropriate intervention. This precise classification avoids the blindness of intervention measures, providing basic regulatory support for children with common emotional distress and securing timely intervention opportunities for children with severe emotional disorders. It comprehensively meets the emotional management needs of different children and promotes their healthy emotional development. By employing multimodal feature fusion and optimization algorithms, the system deeply mines effective information, significantly improving the accuracy of emotion type and intensity recognition. The data collection device utilizes a flexible, wearable, and contactless design, conforming to children's physiological characteristics and behavioral habits, reducing discomfort and resistance. Simultaneously, the recognition results are intuitive and easy to understand, providing schools and families with clear decision-making support and facilitating rapid, targeted intervention. Its lightweight deployment characteristics also make the system easier to apply in everyday scenarios, enabling continuous monitoring and timely intervention of children's emotions, effectively helping to improve children's emotional problems.

[0023] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A full-spectrum emotion recognition system integrating biological signals and behavioral features, characterized in that: Includes the following steps: S1. Data Acquisition: Collect children's biosignal data and behavioral characteristic data, and achieve multimodal data synchronization through unified timestamp calibration; S2, Data Preprocessing: Outlier processing, standardization, and time alignment are performed on the collected biosignal and behavioral feature data to form a complete emotion data sample; S3, Emotional Feature Fusion: Extract unimodal features, calculate adaptive weights for each modality feature based on the entropy weight method, and achieve cross-modal feature fusion through a CNN-Transformer hybrid model to output a unified feature vector; S4. Emotion Analysis and Output: A lightweight MobileNetV3-LSTM model is used to analyze the fused feature vectors and output the emotion type and emotion intensity. S5 full spectrum compatibility: The severity of emotional problems is assessed based on the final output of emotion type and intensity. Based on the identification results, the identified subjects are divided into three groups so that corresponding intervention measures can be taken for emotion regulation based on the specific group type.

2. The full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: In step S1, the acquisition of biosignal data includes: Heart rate (HRV) acquisition: A flexible fabric wristband with a built-in PPG photoplethysmography sensor is worn on the inside of the child's wrist to convert heart rate variability data. Skin conductance EDA acquisition: A flexible conductive fabric chest patch is used to fit the child's chest skin to collect skin conductance values; EEG acquisition: A non-contact millimeter-wave radar device is deployed at the top of the intervention area, 1.5-2 meters away from the child's head, for data acquisition. Wave (8-13Hz), Wave (14-30Hz) signal.

3. The full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 2, characterized in that: The formula for calculating the central rate variability, a core indicator of biological signals, using SDNN is as follows: in, For the first indivual Interval; For average Interval; for Total number of intervals; The peak frequency of the power spectrum is calculated by converting the EDA time-domain signal to the frequency domain using a Fast Fourier Transform (FFT) to extract the frequency corresponding to the peak power spectrum. The specific formula is as follows: Where, x The function representing the change of skin electrical activity signal over time reflects the change at any given moment. The state of skin electrical activity; It is a complex exponential function, used as the kernel function of the Fourier transform, through interaction with the original time-domain signal x. Perform integration operations to achieve the conversion from the time domain to the frequency domain; Indicates time An infinitesimal increment; EEG , The formula for calculating the wave power ratio is as follows: in, express frequency band and The frequency band EEG energy ratio is used to quantify the relative intensity of two frequency band EEG signals; EEG power spectral density describes the EEG signal at different frequencies. Power distribution under; Indicates the power spectral density of brain waves exist The integral of the frequency band is used to calculate the total energy of the EEG signal within that frequency band. Indicates the power spectral density of brain waves exist The integral of the frequency band is used to calculate the total energy of the EEG signal within that frequency band.

4. The full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: In step S1, the behavioral feature data collection utilizes a wearable, motion-sensing smart screen, specifically including: The motion-sensing smart screen has a built-in 3D depth camera to capture data from 342 skeletal points; the collected metrics include motion features, facial features, and voice features. The motion features include the three-dimensional coordinates of skeletal points, the standardization of the motion, and the duration; the facial expression features include the displacement and rate of change of four key points at the corners of the eyes and six key points at the corners of the mouth; and the speech features include the fundamental frequency of 80-500Hz, speech rate, and the frequency of emotional keywords. Furthermore, the aforementioned data is synchronized with the biosignal acquisition equipment via the NTP protocol to calibrate the timestamps.

5. The full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: In step S2, during outlier handling, the behavioral feature data uses 3D modeling. Outliers were identified using cubic spline interpolation; outliers in biosignals (HRV and EDA) were identified using box plots and corrected using median replacement; outliers in EEG were removed using independent component analysis (ICA), with the objective function being: in, Used to constrain the separation matrix to ensure that the transformation of the matrix does not cause unreasonable scaling of the signal energy; This is a separation matrix used to separate the skin electrodermal activity signal x. The goal of converting it into independent components is to optimize the objective function. This ensures that the converted components are independent of each other; This represents the total number of time points in the data sample, i.e., the number of EEG data samples collected. Used to measure the non-Gaussianity of the converted signal; During standardization, behavioral feature data is standardized using Min-Max. Biosignal data were mapped to the [0,1] interval and normalized using Z-score. ,in, Represents a single sample value in the original data; It is the maximum value in this set of behavioral feature data; It is the minimum value in this set of behavioral feature data; This represents the mean of the set of biological signal data; This represents the standard deviation of the set of biological signal data.

6. The full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: In step S2, time alignment employs the Dynamic Time Warping (DTW) algorithm to construct the distance matrix: The goal is to find the optimal path that minimizes the total path distance, where... Represents the first time series, the first... The first time point Dimensional feature values, which correspond to any specific dimension of biological signals or behavioral characteristics; Represents the second time series. The first time point dimensional eigenvalues, which are related to They are in the same feature dimension; This indicates the number of feature dimensions of a time series, that is, the total number of data features contained at each point in time.

7. The full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: In step S3, the fusion of emotional features specifically includes: Single-modal feature extraction: Action features are extracted into 128-dimensional vectors using the ResNet18 network; Facial expression features are extracted into 64-dimensional features using the HOG algorithm; Speech features are extracted into 39-dimensional features using MFCC; Biosignal features are extracted into 64-dimensional vectors using a CNN-LSTM hybrid model. Information entropy is used to measure the degree of disorder of features. The smaller the information entropy, the more effective information the feature carries, the higher its importance in emotion recognition, and the greater its corresponding weight. Therefore, feature weights are calculated based on the entropy weight method. The formula for calculating information entropy is as follows: in, For the first Information entropy of each feature; Indicates the number of samples; For the first In the nth sample The proportion of each eigenvalue must satisfy the following: The weight calculation formula is as follows: in, For the first The final weights of each feature; Information utility value is a quantitative representation of the amount of characteristic information; The total number of features.

8. A full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 7, characterized in that: Cross-modal feature fusion is used to integrate multimodal data to achieve more accurate emotion recognition, and it includes two stages; The first stage involves achieving same-dimensional weighted fusion using the following formula: in, This represents the features obtained after weighted fusion along the same dimension; Represents the number of features of the same dimension participating in the fusion; For the first Each feature vector has the same dimension; The second stage involves calculating cross-modal weights using an attention mechanism, with the specific formula as follows: in, Cross-modal attention weights; , These represent feature vectors of two different modalities; It is a learnable weight matrix used to adjust the mapping relationship between feature vectors; The function normalizes the dot product of the eigenvectors, converting it into a probability distribution form; Final output The 128-dimensional fused feature vector.

9. A full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: The specific process of step S4 is as follows: Fusion feature vectors As input to the MobileNetV3-LSTM model, the fused feature vector is used to perform deep feature extraction to extract representative features. After processing by MobileNetV3, a low-dimensional feature representation containing key information is output. ; Features output by MobileNetV3 The input is fed into the LSTM network; the computation process of the LSTM network is as follows: Forgotten Gate: in, yes The forget gate outputs at any given moment; Use the Sigmoid activation function; It is the weight matrix of the forget gate; The engraved hidden state; It is the input at the current moment; It is the bias of the forget gate; the forget gate determines which information from the previous moment will be retained in the current moment; Input Gate: in, yes Input gate output at any given time; It is the weight matrix of the input gate; It is the bias of the input gate; the input gate determines what new information will be added to the cell state at the current moment. Cell status update: in, It represents the current state of the candidate cells; It is the weight matrix for cell state updates; It is a bias; It is the hyperbolic tangent activation function; in, yes The updated cell state combines information retained from the previous time step with new information from the current time step; Output gate: in, yes The output gate outputs at any given time; It is the weight matrix of the output gate; It is the bias of the output gate; the output gate determines which information in the cell state will be output as the hidden state at the current moment; Hidden state: in, yes The hidden state at each time step; after processing by LSTM It contains temporal feature information of the input sequence; The hidden state at the last moment The input is fed into a fully connected layer, which maps the features to the output space of emotion type and emotion intensity. Let the emotion types be... The probability distribution of emotion types is calculated through a fully connected layer. ,in, It is the weight matrix of the fully connected layer; It is a bias; the emotion intensity output is a continuous value, calculated through a fully connected layer. ,in, It is a weight matrix; It is a bias; The final model outputs the emotion type. With emotional intensity .

10. A full-spectrum emotion recognition system integrating biological signals and behavioral features according to claim 1, characterized in that: In step S5, the three groups are children with common emotional distress, children with mild emotional deficits, and children with severe emotional disorders such as ASD. Criteria for classifying children with common emotional distress: The frequency of negative emotions occurring ≤30% in the past 30 days; the average intensity of negative emotions... And single peak ; Criteria for classifying children with mild emotional deficits: The frequency of negative emotions occurring in 30% to 60% of the past 30 days; the average intensity of negative emotions. or single peak ; Criteria for classifying children with severe mood disorders such as ASD: Frequency of negative emotions occurring ≥60% in the past 30 days; mean intensity of negative emotions. or single peak ; Step S3 also includes the following sub-steps: Based on the physiological and behavioral response patterns of children's emotions, the time boundaries of the three-level characteristics are clearly defined: the immediate characteristic corresponds to 0-3 seconds after the emotional trigger, which is the instantaneous reaction stage; the short-term characteristic corresponds to 3-30 seconds, which is the emotional duration stage; and the cumulative characteristic corresponds to more than 30 seconds, which is the emotional stabilization or decay stage. Focusing on abrupt fluctuation indicators in biological signals and transient actions and vocal bursts in behavioral characteristics, abrupt fluctuation indicators include EDA instantaneous peak values, HRV abrupt changes, and EEG values. , Wave power ratio abrupt change value; instantaneous actions in behavioral characteristics include sudden limb swings and sudden facial muscle tension; vocal burst signals include abrupt changes in tone and instantaneous frequency of emotional keywords; a lightweight CNN network is used to quickly extract 32-dimensional instantaneous feature vectors to highlight key signals at the moment of emotional triggering; Using a 10-second sliding window and a 5-second step, time-series slices are made for each modality of data to extract trend features within the window, including HRV fluctuation slope, EDA rise or fall rate, speech intonation change trend, and cumulative displacement of facial expression key points. Temporal dependencies are captured through a single-layer LSTM network to output a 64-dimensional short-term temporal feature vector that reflects the changing patterns during the duration of emotion. It fully adopts the original single-modal feature extraction logic, extracts a 128-dimensional cumulative feature vector, ensures the stability and integrity of long-term emotional features, and seamlessly connects with the original process; A basic weighting framework for three layers of features is established, with short-term features having a fixed weight of 30%, and the sum of the weights for immediate and cumulative features being 70%, dynamically allocated based on the current emotion intensity; when the real-time output emotion intensity... When the score is ≥7, indicating a high level of negative emotion or a strong positive emotion, the weight of the immediate feature is increased to 30%, and the weight of the cumulative feature is reduced to 40%, thus strengthening the decision weight of the instantaneous explosive emotional signal. When emotional intensity When the emotion is determined to be of moderate intensity, the weight of the immediate feature is adjusted to 20%, and the weight of the cumulative feature is 50%, to balance the contribution of the instantaneous reaction and the long-term cumulative feature. When emotional intensity <4, when judged as a low-intensity emotion or calm state, the weight of the immediate feature drops to 10%, and the weight of the cumulative feature rises to 60%, highlighting the dominant role of long-term stable emotional features. Then, layered fusion is performed. The first step is to fuse immediate features and short-term features. An attention mechanism is used to calculate cross-layer weights. The formula is based on the cross-modal attention fusion logic, as follows: in, It is a 32-dimensional instantaneous feature vector; It is a 64-dimensional short-term eigenvector; The weight matrix is ​​a learnable matrix with dimensions 32×64. After obtaining the cross-layer attention weights using this formula, they are fused to generate a 96-dimensional dynamic feature vector. The second step is to fuse dynamic features and cumulative features. This is done by weighting the dynamic features based on the weight ratios determined in the first step, using the following formula: in, For dynamic feature weights; For cumulative feature weights; It is a 128-dimensional cumulative feature vector; during the fusion process, the dimension adaptation matrix is ​​used to... Map to 128 dimensions to ensure dimensional consistency; The third step is to standardize the fusion results. L2 normalization is performed to output a 128-dimensional final fused feature vector. This vector has the same dimension as the feature vector output from the cross-modal fusion above, and is directly used as the input for subsequent cross-modal fusion in step S3 without changing the processing logic of the original CNN-Transformer hybrid model.