A Method for Establishing an Emotion Regulation Model Based on Emotion Monitoring

By combining multidimensional data acquisition and LSTM neural networks with the t-SNE algorithm, an emotion regulation model is constructed, which solves the problem of insufficient multidimensional feature capture in traditional emotion analysis techniques. This enables real-time monitoring of individual emotion baselines and targeted intervention for emotion regulation, thereby reducing mental health risks.

CN120413070BActive Publication Date: 2025-10-31HANGZHOU PIGEON NEST TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510916065.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-31
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional emotion analysis techniques rely on a single data source, making it difficult to fully capture the multidimensional characteristics of emotions. They cannot effectively model emotion latency, recovery speed, and expression diversity, and lack a systematic multimodal modeling framework, resulting in a lack of analysis on the correlation between emotion regulation and potential development.

Method used

By collecting multidimensional data, including heart rate variability, skin conductance response, facial micro-expressions, and voice analysis, and combining Ekman's basic emotion theory and LSTM neural network, an emotion regulation model is constructed to monitor emotional fluctuations in real time and predict the risk of emotional dysregulation. The t-SNE algorithm is used to visualize the trajectory of emotion development.

Benefits of technology

It enables real-time monitoring of individual emotional baselines, early warning and targeted intervention, reduces the incidence of mental health problems, adapts to individual differences, and improves the accuracy of emotional state identification and the targeted nature of emotion regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120413070B_ABST
    Figure CN120413070B_ABST
Patent Text Reader

Abstract

This invention discloses a method for establishing an emotion regulation model for emotion monitoring, relating to the fields of mental health and clinical intervention technology. The method includes: Step 1: Collecting multidimensional data, including static information samples, dynamic information samples, and self-assessment samples from children; Step 2: Based on Ekman's basic emotion theory, mapping the multidimensional data into emotion factors, including emotion latency, emotion recovery speed, and expression diversity; using principal component analysis to reduce dimensionality and screen for highly correlated features; using an LSTM neural network to predict the risk of emotional outbursts; Step 3: Based on an extended process model of emotion regulation, regulating emotions in a timely manner; combining the t-SNE algorithm to visualize the emotion development trajectory and generate customized regulation plans; collecting heart rate variability and skin conductance indicators through wearable devices to establish an individual emotion baseline; and using computer vision technology to extract facial micro-expressions and limb movement frequency behavioral features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mental health and clinical intervention technology, specifically to a method for establishing an emotion regulation model for emotion monitoring. Background Technology

[0002] By integrating Jungian psychological type theory and Daniel's emotional intelligence model through the SPA six-dimensional personality analysis system, a complete personality framework integrating Eastern and Western psychology is constructed. This theory, with binary multidimensional modeling at its core, transforms dimensions such as children's attention to environmental information (macro / detail), analytical focus (logic / ethics), and processing methods (decisiveness / cooperation) into 6 dimensions and 12 indicators, such as high / low desire, conceptual / detailed, artistic / outspoken. Through static data assignments, voice recordings, dynamic data situational behavior capture, and self-assessment, personality traits are visualized. Based on long-term personality profiles and dynamic talent profiles, Matching teaching resources, such as matching structured courses to ISTJ types and recommending self-directed inquiry learning for ENTJ types, and developing robotic tutors to provide real-time learning suggestions and emotional regulation interventions; Relationship and potential management: Optimizing communication strategies through relationship early warning modality analysis of parent-child / teacher-student interaction patterns; Identifying children's behavioral patterns in the comfort and stress zones of the four psychological zones using emotional personality modality analysis to assist parents in adjusting their educational methods; Interdisciplinary application: Combining MBTI educational research to verify the impact of teacher personality types, such as ENFP types, on innovative teaching on student learning outcomes, and improving knowledge transfer efficiency through teacher-mentor cognitive consistency.

[0003] Traditional emotion analysis often relies on a single data source, making it difficult to comprehensively capture the multidimensional characteristics of emotions. For example, physiological indicators can reflect the level of emotional arousal, but they lack dynamic correlation modeling with behavioral characteristics, which limits the accuracy of emotion state identification. At the same time, modeling of emotion changes often uses static threshold methods or simple statistical models, which cannot effectively capture the dynamic patterns of emotion latency, recovery speed, and expression diversity. Furthermore, traditional techniques often apply single-domain methods in isolation, without forming a systematic multimodal modeling framework, resulting in a lack of correlation analysis between emotion regulation and potential development. Summary of the Invention

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] Step 1: Collect multidimensional data, including children's static information samples, children's dynamic information samples, and self-assessment samples;

[0006] Step 2: Based on Ekman's basic emotion theory, multidimensional data is mapped into emotion factors, including emotion latency, emotion recovery speed, and expression diversity; dimensionality reduction is performed through principal component analysis to screen for highly correlated features; and an LSTM neural network is used to predict the risk of emotional outbursts.

[0007] Step 3: Based on the extended process model of emotion regulation, regulate emotions in a timely manner; combine the t-SNE algorithm to visualize the trajectory of emotion development and generate customized regulation plans.

[0008] Furthermore, the process of collecting multidimensional data is as follows:

[0009] By collecting heart rate variability and skin conductance indicators through wearable devices, an individual emotional baseline is established; computer vision technology is used to extract facial micro-expressions and limb movement frequency behavioral characteristics; and voice samples of children over a period of time are randomly captured, and changes in speech rate, tone, and loudness are analyzed through a speech analysis module.

[0010] Furthermore, the process of establishing an individual's emotional baseline is as follows:

[0011] Establishing an individual's emotional baseline includes both static and dynamic baselines;

[0012] By collecting HRV and EDA data from users in a resting state, calculating the mean and standard deviation, a personalized reference range is obtained, and a static baseline is acquired.

[0013] By analyzing the HRV / EDA variation patterns of users in different contexts using machine learning models, the current state is predicted to deviate from the baseline, and a dynamic baseline is obtained.

[0014] Furthermore, the process of mapping multidimensional data into emotion factors is as follows:

[0015] Based on multidimensional data, we identify emotion-triggered events and reaction times, and combine the Ekman emotion classification model to associate the dominant emotion type during the latency period. We quantify the peak and recovery speed of emotions through physiological indicators and behavioral characteristics. We use text emojis, speech rate and interactive operations to statistically analyze the distribution of emotion types, and calculate the diversity of emotion expression through Shannon entropy.

[0016] Furthermore, the process of quantifying the peak emotion and recovery speed is as follows:

[0017] Based on multidimensional data, emotional peaks are identified and their intensity is quantified; recovery speed is calculated based on baseline state; multimodal weighted fusion and half-life modeling are used, combined with Ekman emotion classification model, to dynamically monitor and evaluate emotional peaks and recovery processes.

[0018] Furthermore, the process of screening highly correlated features is as follows:

[0019] The method compresses n features from multidimensional data into a low-dimensional space and k principal components, retaining the main information. By analyzing the loading matrix of the principal components, the correlation between the original features and the principal components is analyzed, the features that contribute more to the principal components are identified, and the features that contribute less to the principal components are deleted. The effectiveness is verified by visualization or model performance evaluation.

[0020] Furthermore, the process of using an LSTM neural network to predict the risk of emotional dysregulation is as follows:

[0021] S201: Preprocess the collected multidimensional data, serialize and fill in the dynamic information samples, and unify the dimensions; fuse the static information samples with time series data to form a multimodal input;

[0022] S202: Input multidimensional sequence data, use stacked LSTM layers to capture long-term dependencies in the time series, add a fully connected layer after the LSTM to filter features that are highly correlated with the risk of emotional outbursts; through a gating mechanism, the input gate, forget gate, and output gate dynamically adjust the information flow; output the emotion prediction result at the current moment;

[0023] S203: Divide the multidimensional dataset into time series segments and monitor metrics by tracking accuracy indicators; evaluate model robustness using time series cross-validation; and optimize the number of layers, number of units, and learning rate parameters of the LSTM through grid search.

[0024] Furthermore, the process of regulating emotions is as follows:

[0025] Based on the risk of emotional outbursts, emotional states are categorized and the risk level is assessed. Key emotional factors are extracted through principal component analysis. If environmental triggers for negative emotions are identified, the intensity of emotions is reduced by adjusting the environment, modifying the situation, and guiding children to focus on positive stimuli. The intensity of negative emotions is reduced by reinterpreting the situation through language guidance. If the current strategy is ineffective, the strategy is adjusted in a timely manner.

[0026] Furthermore, the process of visualizing the emotion development trajectory using the t-SNE algorithm is as follows:

[0027] Continuous sentiment data is divided into fixed-length time windows to form time series segments. Key features of each time window are extracted to construct a high-dimensional feature matrix. The similarity of each sample point to other points is calculated, and the distance between each sample point and other points is converted into a similarity value. The closer the distance, the higher the similarity; the farther the distance, the lower the similarity. The similarity of each point to other points is symmetrically processed. The high-dimensional data is compressed into a low-dimensional space.

[0028] Furthermore, the process of compressing high-dimensional data into a low-dimensional space is as follows:

[0029] Based on the t-SNE algorithm, high-dimensional emotion data is mapped to a low-dimensional space. The similarity between points is modeled by the t-distribution, and the points are dynamically adjusted to make the low-dimensional distribution approximate the high-dimensional structure. Gradient descent optimization is used to bring high-dimensional similar points closer and push away dissimilar points, capturing the stable emotional state clustering state of children.

[0030] The present invention provides a method for establishing an emotion regulation model based on emotion monitoring, which has the following beneficial effects:

[0031] (1) This invention constructs an individual emotional baseline by using physiological indicators such as HRV and EDA and multimodal modeling of voice and behavioral data to monitor emotional fluctuations in real time; predicts the risk of emotional outbursts by using LSTM neural network and visualizes the trajectory of emotional development by combining t-SNE algorithm to achieve early warning and targeted intervention, thereby reducing the incidence of mental health problems.

[0032] (2) This invention uses the normal distribution method to establish a personalized static baseline and uses a machine learning model to dynamically model the situational baseline. It combines exponentially weighted moving average and incremental learning techniques to adapt to changes in children’s physiological or psychological state in real time. For example, it can automatically trigger baseline reconstruction after stress events, effectively solving the problem of insufficient adaptability of the traditional static threshold method to individual differences.

[0033] (3) This invention captures the long-term dependence of dynamic changes in emotions based on LSTM neural network, combines attention mechanism and Transformer structure to optimize model performance, outputs emotional out-of-control risk score and triggers early warning mechanism; through principal component analysis to screen highly correlated features, such as decreased HRV and increased skin conductance, combined with causal inference model to correlate behavioral data, to achieve forward-looking intervention in emotional crisis. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the method of the present invention;

[0035] Figure 2 This is an extended process model for emotion regulation in this invention;

[0036] Figure 3 The waveform diagram for calculating the decibel level of sound in this invention is shown. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Example

[0039] Please see Figures 1 to 3 This application provides a method for establishing an emotion regulation model based on emotion monitoring, the method comprising:

[0040] Step 1: Collect multidimensional data, including children's static information samples, children's dynamic information samples, and self-assessment samples;

[0041] Sample of children's dynamic information:

[0042] By collecting heart rate variability and skin conductance indicators through smart wearable devices, an individual emotional baseline is established; computer vision technology is used to extract facial micro-expressions and limb movement frequency behavioral characteristics; and two-minute voice samples of children are randomly captured, and changes in speech rate, tone and loudness are analyzed through a speech analysis module.

[0043] Smart wearable devices:

[0044] Heart rate sensor: A photoplethysmography sensor that uses a smart bracelet or dedicated medical device to optically detect changes in blood flow and calculate the heart rate interval (HRV).

[0045] Skin conductivity sensor: Employs an EDA sensor from an electronic skin or wearable device to reflect sympathetic nerve activity by measuring changes in the conductivity of the skin surface;

[0046] Wearing and Calibration:

[0047] Wear the device on a skin contact area such as the wrist, forearm, or chest; turn on the device and perform baseline calibration, recording the heart rate sensor and EDA values ​​at rest; continuously record the heart rate interval using the PPG sensor, calculating time-domain and frequency-domain metrics; measure skin conductance and transient response using the constant current method or voltage method, recording baseline fluctuations and peak values ​​triggered by events;

[0048] Establishing an individual's emotional baseline:

[0049] Establishing an individual's emotional baseline includes both static and dynamic baselines;

[0050] Static baselines are formed by collecting HRV and EDA data from users in a resting state, calculating the mean and standard deviation, and creating a personalized reference range.

[0051] Personalized reference range:

[0052] Calculate HRV and EDA indices separately; EDA indices include skin conductance level and transient response count, with skin conductance level reflecting baseline sympathetic tone; transient response count, such as the frequency of skin conductance response SCR; use the Shapiro-Wilk test to determine whether the data conforms to a normal distribution; if the data has influencing factors such as age and gender, modeling should be done separately for subgroups, such as by male / female gender and age;

[0053] Using the normal distribution method, the mean (μ) and standard deviation (σ) of the HRV and EDA indices are calculated, with μ±2σ as the reference range, covering approximately 95% of the resting state data. A dynamic adjustment strategy is employed to handle outliers and time decay weights. If extreme values ​​exist in the data, such as outliers after strenuous exercise, the Winsorization method can be used to replace these extreme values ​​with values ​​near the threshold before recalculation. More recent data is given higher weight, such as using an exponentially weighted moving average, to adapt to gradual changes in the user's physiological state.

[0054] Data is repeatedly collected during the child's subsequent resting state to check whether the new data falls within the reference range. If more than 90% of the data meets expectations, and more than 10% of the data is out of range, the data quality needs to be reassessed or the model parameters adjusted. At the same time, the data is refreshed periodically. If more than 10% of the data is out of range, the data quality needs to be reassessed or the model parameters adjusted. When the user experiences a major physiological change, such as a stressful event, the reference range is actively reconstructed.

[0055] Dynamic baseline: Analyzes the HRV / EDA change patterns of users in different contexts through machine learning models to predict the degree of deviation between the current state and the baseline;

[0056] For different contexts, such as "stress" and "relaxation," the distribution characteristics of HRV / EDA are modeled separately to build contextual baselines. Using labeled contextual data, including HRV / EDA features and contextual labels, classification models, such as random forests, XGBoost, or regression models, are trained to predict the HRV / EDA pattern corresponding to the current context. Clustering algorithms are used to identify the latent state categories of HRV / EDA, and the baseline is dynamically updated by combining sliding windows or exponentially weighted moving averages. Feature vectors are formed from the time and frequency domains of the HRV / EDA data.

[0057] Static baseline comparison: Calculate the deviation of the current feature value from the static baseline (mean ± 2σ), such as Z-score = (current value - static mean) / static standard deviation. If the Z-score exceeds the threshold, it is marked as a significant deviation.

[0058] Use a trained machine learning model to predict the "expected HRV / EDA pattern" of the current situation, and a dynamic baseline; calculate the Euclidean distance or cosine similarity between the current features and the predicted baseline to quantify the degree of deviation, such as the larger the distance, the more significant the deviation.

[0059] If the model predicts the current situation as "stress" or "anxiety", the status is further confirmed by combining the deviation of HRV / EDA; if the deviation exceeds the preset threshold, such as the 95% confidence interval of the dynamic baseline, an early warning or intervention recommendation is triggered.

[0060] Baseline adaptive updates are implemented by periodically retraining the model, incorporating newly acquired HRV / EDA data, and correcting the baseline; baseline reconstruction is triggered when a major physiological or psychological event is detected.

[0061] If a drift in the HRV / EDA distribution over time is detected, the model is adjusted through incremental learning or transfer learning; data augmentation is used to improve the model's adaptability to unknown situations.

[0062] Acquiring facial micro-expressions and body movement frequencies:

[0063] Using a computer vision system, high-definition cameras are used to capture facial and full-body movements; C3D networks are used in combination with optical flow or STSTNet to analyze micro-expression changes; the EMAGE framework is used to generate speech-synchronized body movements in combination with audio input, and the movement frequency is extracted by optical flow.

[0064] Facial feature points are extracted using the ASM or dlib library, and all video clips are aligned to the standard face coordinate system to reduce individual differences; RGB images are converted to grayscale images, and brightness and contrast are adjusted.

[0065] Feature extraction:

[0066] Micro-expression feature extraction utilizes optical flow to calculate pixel displacement between adjacent frames, i.e., horizontal / vertical optical flow, to capture subtle facial muscle movements; it learns spatial and facial region limb movements through a 3D CNN; it extracts the positions of joints throughout the body using the SMPLX model, calculates joint angles and movement speeds; and it counts the number of movements per unit time, such as the frequency of head nodding.

[0067] By fusing spatiotemporal features, the spatiotemporal features of facial micro-expressions and body movements are input into a classifier to identify emotion categories, such as happiness and anxiety. Through Fourier transform or sliding window statistics, the rhythm of body movements is quantified, such as rapid movements may be associated with tension.

[0068] Acquisition and analysis of children's voice samples:

[0069] Use a directional microphone or the built-in microphone of a smartwatch to ensure clear capture of children's voices in noisy environments; record 2 minutes of continuous voice by sampling at a 16kHz sampling rate, and process the voice in segments;

[0070] Voice decibel calculation: Everyone's voice is different. People with strong emotions naturally have a more varied and dramatic tone; people with calm emotions naturally have a more even and hesitant tone. A two-minute sample of a child's voice is randomly selected, and the difference between the peak and mean is calculated. If the result is greater than value A (e.g., 22 decibels), it indicates a high-desire type; if it is less than value A, it indicates a low-desire type. Figure 3 ;

[0071] Sentiment association analysis:

[0072] Speech rate and emotions: A fast speech rate may reflect excitement or anxiety, while a slower speech rate may be associated with sadness or fatigue;

[0073] Tone of voice and emotion: High-frequency fundamental frequencies may indicate surprise or anger, while low-frequency fundamental frequencies may be associated with calmness or depression;

[0074] Loudness fluctuations: A sudden increase in loudness may indicate excitement or tension, while a sustained low loudness may be associated with repression;

[0075] Children's static information samples: Combining social media logs, interaction records, etc., marking emotionally triggered events; including historical records such as children's workbooks, newspapers, calligraphy works, Douyin videos, WeChat records, singing recordings, and dance performance videos;

[0076] WebCrawler is used to obtain text, image, and video information from children's WeChat Moments, Weibo, QQ Space, Maimai, and audio / video content; text / multimedia content such as children's historical posts, comments, and private messages on social media platforms; behavioral data such as likes, reposts, follows, and private messages; as well as the time, frequency, and content of interactions with others; and behavioral records of children on learning platforms, game apps, and family environment data.

[0077] Data source:

[0078] Public content of children's accounts, user behavior logs on learning platforms and gaming platforms, chat logs and voice logs on smart speakers and tablets in the home;

[0079] Social media logs and interaction records can be obtained from the open APIs of social media platforms to access the content posted and interaction records of children's accounts; parents can also access the historical data of their children's accounts through the "data download" function provided by the platform, such as downloading a JSON data package from Instagram containing all posts, comments, and interaction records.

[0080] Metadata of posts, comments, private messages, images, and videos, including timestamps and frequency of likes, shares, and follows, and extracting key fields;

[0081] Key field extraction: Conflicts mentioned in the text (e.g., "My classmate said I'm ugly"), achievements (e.g., "I got first place"); emotional cues, including emoticons, interjections (e.g., "So annoying"), and repetitive behaviors, such as frequently deleting / modifying posts;

[0082] The learning platform's "Learning Report" function can be used to export children's learning records. At the same time, the app usage time and operation records on the device can be captured through the home router or parental control software. Voice conversations and video clips can be obtained through smart speakers or home monitoring systems, such as analyzing negative emotional keywords (e.g., "I don't want to go to school") in children's voice conversations with parents.

[0083] Parents can use a log app to record their children's emotional expressions and triggering events, such as "They cried today because the homework was too hard";

[0084] Clean text data, removing advertisements and meaningless characters; convert unstructured data, such as voice and video, into structured data, such as speech-to-text and video keyframe extraction; align data from different sources along a timeline to facilitate association with emotional events and trigger points; match social media post timestamps with timestamps from home surveillance videos;

[0085] Rule-based tagging uses an emotion dictionary to match emotional keywords in the text, such as "angry" and "happy." For example, it detects a child posting on social media, "I was scolded by my teacher today, I'm so sad!" and tags it as an "anger / sadness trigger event." It also identifies anomalous behaviors, such as suddenly deleting friends or reducing social interaction, as potential triggers. For example, if a child stops following any friends after a certain day, it may indicate a social frustration event.

[0086] The text sentiment is analyzed using a pre-trained model. The HuggingFace distilbert-base-uncased-finetuned-sst-2-english model is used to determine post sentiment. A causal inference model is trained to correlate behavioral data with emotional states; for example, combining the number of game failures with social media post sentiment to predict "frustration trigger events." The model also supports verification of machine-labeled trigger events by parents or psychologists to correct misjudgments. For instance, a parent points out that "the child cried because of failing an exam" is a real event, but the model misclassifies it as "boredom." A labeling platform is used to manually label the data, forming a training set where "classmates mocking" is labeled as a trigger event and "the child angrily retaliating" is labeled as an emotional behavior.

[0087] We identify cyberbullying trigger events by having children post on social media logs that "I was ridiculed by my classmates," or by having a large number of anonymous private messages or messages containing insulting language received on the same day. Through text analysis, we detect keywords such as "ridiculed" and "insults" and mark them as "cyberbullying trigger events."

[0088] It should be noted that the data is anonymized, and children's identity information (such as name and mobile phone number) is deleted or encrypted. Specific names in social media posts are replaced with "Classmate A". Written authorization from parents is obtained, the purpose and storage period of the data are clearly defined, and an agreement is signed stating that the data is only used for emotion research and not for commercial purposes. Children's data is stored on local servers and not uploaded to the cloud.

[0089] Self-assessment: The assessment is conducted using a scale for students in grades 3 and above. For students in grades 1 and 2 and kindergarten, the assessment is assisted by parents reading the questions aloud and children choosing the questions themselves.

[0090] A mental health assessment scale is used to evaluate scientific inquiry, technological cognition and other abilities through multiple-choice and fill-in-the-blank questions. The electronic scale is sent to students via mobile phone or computer, and students fill it out and submit it online. The total score and scores for each dimension are calculated according to the scale manual.

[0091] For example:

[0092] Students complete the "Primary School Students' Attention Assessment Scale" online. The system records the completion time and error rate of the Schulte Grid test. If the completion time exceeds the average of the same age group, it indicates that attention training is needed. By adapting assessment tools to different age groups and combining electronic and paper data collection methods, children's psychological, behavioral and developmental data can be systematically obtained.

[0093] Step 2: Based on Ekman's basic emotion theory, multidimensional data is mapped into emotion factors, including emotion latency, emotion recovery speed, and expression diversity; dimensionality reduction is performed through principal component analysis to screen for highly correlated features; and an LSTM neural network is used to predict the risk of emotional outbursts.

[0094] Mapping multidimensional data to sentiment factors:

[0095] From the time interval between emotional triggering events, such as conflict and achievement, and emotional reactions, such as posting or behavioral changes; during the emotional latency period, by identifying triggering events and calculating reaction time, and based on mapping logic, we can associate emotional types; the time required for emotions to recover from a peak, such as an outburst of anger, to a baseline state, such as calmness, we can obtain the emotional recovery speed; and the ways and ranges in which children express emotions in different situations, such as using emojis, tone of voice, and behavioral patterns, we can express diversity.

[0096] By identifying emotional triggers through social media logs, gaming behavior records, etc., we can identify triggered events; after the triggering event occurs, we can count the timestamps of the child's first emotional reaction, such as changes in posting or voice emotion, and calculate the reaction time; combined with Ekman's emotion classification model, such as anger and sadness, we can determine the dominant emotion during the latency period and associate emotion types.

[0097] Emotional peaks are identified using physiological data, such as heart rate and skin conductance, or behavioral data, such as voice volume and social media posting frequency. After statistical analysis of emotional peaks, relevant indicators, such as the time it takes for heart rate to return to normal and the time it takes for posting emotion scores to return to normal, are analyzed. Combined with Ekman's emotion classification, the differences in recovery speed for different emotions are analyzed, such as anger recovering faster than sadness.

[0098] Quantifying peak and recovery speed of emotions:

[0099] In emotion regulation models based on multidimensional data, quantifying emotional peaks and recovery speeds requires combining physiological indicators and behavioral characteristics. Emotional arousal states are collected through physiological signals such as skin conductance response (GSR) and heart rate variability (HRV); for example, a sudden increase in GSR reflects anger or fear, while a decrease in HRV indicates stress. Behavioral characteristics, such as abrupt changes in voice tone (increased pitch), a surge in social media posting frequency, or a high frequency of deletions, can help pinpoint emotional peaks. Noise removal and time alignment are performed during data preprocessing to ensure synchronization of multimodal data.

[0100] The quantification of emotional peaks is achieved by setting a threshold, such as recognizing GSR exceeding the 90th percentile, combining it with the Ekman emotion classification model to associate the dominant emotion type during the latency period, and calculating the comprehensive emotion intensity score through multimodal weighted fusion, such as GSR, HRV, and voice intonation weight allocation.

[0101] Recovery speed is modeled using regression analysis or half-life method, comparing it with baseline states, such as mean GSR and HRV range at rest, to assess the regression time of physiological and behavioral indicators, such as the time required for GSR to drop from peak to 50% or the time it takes for posting frequency to return to baseline.

[0102] Multimodal emotion recognition is performed by combining text data such as emojis and keywords, voice data such as speech rate and tone, and behavioral data such as like / delete actions. The frequency and intensity of each basic Ekman emotion are statistically analyzed, such as "happiness" accounting for 30% and "anger" accounting for 15%, to obtain the distribution of emotion types. The richness of emotion expression is quantified by Shannon entropy.

[0103] The uncertainty of emotion distribution is quantified using the Shannon entropy formula. When all emotions have equal probabilities, the Shannon entropy is the largest, indicating the richest expression of emotions.

[0104] When the probability of a certain type of emotion approaches 1, the Shannon entropy approaches 0, indicating that the emotion expression is singular; and the results are explained as follows:

[0105] High Shannon entropy (close to 2.58): Diverse emotional expression; users can flexibly experience and express multiple emotions, such as simultaneously expressing happiness, surprise, and disgust.

[0106] Low Shannon entropy (close to 0): Emotional expression is singular, and users mainly focus on a certain type of emotion, such as long-term expression of anger or sadness;

[0107] Principal component analysis is used for dimensionality reduction to screen for highly correlated features.

[0108] By calculating the correlation between all features, strongly correlated features are obtained, and the emotional state reflected by the combination of strongly correlated features is analyzed. For example, a decrease in HRV and an increase in skin conductance may jointly indicate stress. The core directions that can explain most of the emotional changes are extracted from the data, the main trends of emotional changes are determined, the top few directions that can explain the most emotional information are retained, and the secondary directions are discarded to reduce complexity.

[0109] The original data is mapped to a low-dimensional space composed of core directions, with each sample represented by fewer dimensions. The n features in the multidimensional data are compressed into a low-dimensional space with k principal components, preserving key information. The correlation between the original features and the principal components is analyzed using the principal component loading matrix, identifying features that contribute significantly to the principal components (typically the top 15% of influential features). Key emotional features are also retained. Furthermore, the original features that contribute most to the core directions are analyzed, and these highly correlated features are selected. For example, RMSSD in heart rate variability may significantly contribute to a certain principal component. Features that contribute less to the core directions, such as certain low-frequency behavioral patterns, are removed, retaining a concise and informative feature set. The data after dimensionality reduction is observed to see if the key distributions of emotional factors are preserved, such as the correlation between latency and recovery speed, and its effectiveness is verified through visualization or model performance evaluation. If some important features are found to be missing, the selection strategy can be readjusted, such as increasing the number of principal components or supplementing features, to ensure that the final results meet the needs of emotion analysis.

[0110] Using LSTM neural networks to predict the risk of emotional dysregulation:

[0111] S201: Clean the collected multidimensional data, such as physiological signals and behavioral data in children's dynamic information samples, to remove noise and outliers; normalize or standardize continuous data, such as heart rate and activity intensity, to make them fall within a uniform range, such as 0-1 or -1 to 1; perform one-hot encoding or embedding processing on discrete data, such as category labels in self-assessment samples.

[0112] Dynamic information samples are serialized and padded, such as time series data, which are divided into fixed-length sliding windows, for example, one window every 10 seconds, forming time steps; sequences of different lengths are padded to ensure that the tensor dimension of the input LSTM is consistent.

[0113] Static information samples, such as age, gender, and medical history, are fused with dynamic information samples, such as time series data, to form a multimodal input; for example, static features can be used as the initial state of an LSTM or concatenated with the LSTM output.

[0114] S202: For the input layer: input multidimensional sequence data;

[0115] LSTM layer: Use stacked LSTM layers or bidirectional LSTM to capture long-term dependencies in time series;

[0116] Through a gating mechanism, the information flow is dynamically adjusted by the input gate, forget gate, and output gate. For example:

[0117] Forget Gate: Determine which historical emotional states need to be ignored (such as information from low-risk periods).

[0118] Input gate: Updates the current emotional state (e.g., detection of high-risk signals);

[0119] Output gate: Outputs the current moment's sentiment prediction result;

[0120] Add a fully connected layer after LSTM or use principal component analysis to further reduce dimensionality and screen for features that are highly correlated with the risk of emotional dysregulation, such as emotional latency and recovery speed.

[0121] Output layer: Outputs the probability of the risk of emotional loss of control, such as binary classification to distinguish between high risk and low risk, or risk scoring through continuous values;

[0122] The model is optimized by using a Dropout layer for regularization to prevent overfitting; loss functions are selected based on the task type, such as binary cross-entropy for classification and mean squared error for regression; and the learning rate is adjusted using Adam or RMSprop optimizers to accelerate convergence.

[0123] S203: Divide the cube into time series, such as 80% training set, 10% validation set, and 10% test set, to avoid missing timestamps;

[0124] The batch size for training parameters is 32 or 64; the number of training epochs is 10-50, and early stopping is used to prevent overfitting; metrics are monitored by tracking accuracy, F1 score, AUC-ROC curve and other indicators; time series cross-validation is used to evaluate the robustness of the model, and the number of layers, number of units, learning rate and other parameters of LSTM are adjusted by grid search or Bayesian optimization.

[0125] S204: Input the real-time collected dynamic information samples into the trained LSTM model and output the current moment's emotional outburst risk score; if the risk score exceeds the threshold, trigger the early warning mechanism; use the hidden state or attention mechanism of LSTM to extract the features of the emotional development trajectory for long-term trend analysis; combine the t-SNE algorithm, as required by S203, to visualize the emotional change path and assist in formulating personalized adjustment plans; analyze which time steps or features have the greatest impact on the risk of emotional outburst risk through attention weight analysis, such as abnormal heart rate during a specific time period;

[0126] LSTM can capture the long-term dependencies of dynamic changes in emotions, and by combining static and dynamic information, the prediction accuracy can be improved. Furthermore, performance can be optimized by adding attention mechanisms or Transformer structures.

[0127] Step 3: Based on the extended process model of emotion regulation, regulate emotions in a timely manner; combine the t-SNE algorithm to visualize the trajectory of emotion development and generate customized regulation plans;

[0128] The extended process model of emotion regulation is a dynamic model, with the emotion assessment system composed of W, P, V, A, and V as the core component of the extended model. W refers to the internal or external world (situation selection and situation modification strategies), P is the perception of the assessment system (attention allocation strategy), V is the evaluation of feelings as good, neutral, or bad (cognitive alteration strategy), and A is the behavior caused by the evaluation (feedback adjustment strategy). This behavior becomes the next W (W1), which triggers the cycle of the next assessment system.

[0129] Timely regulation of emotions:

[0130] Children's emotional state data is collected in real time through sensors, behavioral observation, and subjective reports. Combined with static and dynamic information, a multidimensional emotional feature vector is formed. LSTM or random forest is used to classify emotional states, such as anger, anxiety, and happiness, and to assess the risk level of emotional outburst. Key emotional factors, such as emotional latency and recovery speed, are extracted through principal component analysis or feature engineering.

[0131] If negative emotions are identified as being triggered by the environment, such as anxiety caused by a noisy environment, the intensity of the emotion can be reduced by adjusting the environment, such as playing soothing music or diverting attention; and the situation can be modified by guiding the child to focus on positive stimuli, such as guiding their attention to games or drawing, thus reducing excessive focus on negative situations.

[0132] Reinterpret the situation through language guidance, such as "This failure is a learning opportunity," to reduce the intensity of negative emotions; if emotions have already been aroused, such as anger, use expressive inhibition or physiological regulation to control overt reactions, such as deep breathing or pausing action.

[0133] Interventions are implemented based on selected strategies, such as music therapy and mindfulness meditation, and physiological and behavioral data, such as decrease in heart rate and calming of facial expressions, are monitored in real time. The effectiveness of the strategies is evaluated by comparing data before and after the intervention, such as mood scores and physiological indicators.

[0134] If the current strategy is ineffective, such as if the child's anxiety is not relieved, switch strategies, such as switching from music therapy to sensory stimulation: holding ice cubes to cool down; based on feedback loops, continuously optimize the strategy combination;

[0135] Visualizing the trajectory of emotion development using the t-SNE algorithm:

[0136] Continuous emotional data is divided into fixed-length time windows to form time series segments, such as a window every 10 seconds; key features of each time window are extracted, such as the mean emotional intensity, fluctuation amplitude, and latency, to construct a feature matrix of high-dimensional data.

[0137] In high-dimensional data, the similarity of each sample point to other points is calculated to form a "similarity matrix". For each sample point, it is assumed that the points around it are closer to it, while the points further away are sparser. A distribution similar to a "bell curve" (such as a Gaussian distribution) is used to convert the distance of each sample point to other points into a similarity value; the closer the distance, the higher the similarity; the farther the distance, the lower the similarity. In order to balance the overall distribution, the similarity of each point to other points is symmetrically processed; for example, the similarity from point A to point B and the similarity from point B to point A are averaged.

[0138] t-SNE compresses high-dimensional data into a low-dimensional space, enabling complex emotional development trajectories to be displayed in the form of scatter plots, heat maps, etc., making it easier to observe the clustering, distribution and dynamic changes of emotional states.

[0139] High-dimensional data is compressed into a low-dimensional space, such as two or three dimensions, while preserving as much similarity as possible in the high-dimensional data. Initial point positions are randomly generated in the low-dimensional space. An alternative distribution, such as the t-distribution with 1 degree of freedom (whose tail is wider than that of a Gaussian distribution), is used to calculate the similarity between low-dimensional points. The greater the distance between points, the slower the similarity decreases, thus preventing overcrowding in the low-dimensional space. Finally, the positions of the low-dimensional points are adjusted to make the similarity in the low-dimensional space as close as possible to the similarity in the high-dimensional space.

[0140] By continuously adjusting the positions of low-dimensional points, the distribution of the low-dimensional space is made as consistent as possible with the distribution of the high-dimensional space. An "error index" is defined to measure the difference between high-dimensional similarity and low-dimensional similarity. The gradient descent method is used to gradually adjust the positions of the low-dimensional points.

[0141] If two points are very similar in a high-dimensional space but far apart in a low-dimensional space, then bring them closer together;

[0142] If two points are dissimilar in a high-dimensional space but are close in a low-dimensional space, then push them apart.

[0143] Repeat the above process until the error index no longer decreases significantly, at which point the point distribution in the low-dimensional space is basically stable.

[0144] Statistical methods are used to quantify the similarity between points in high-dimensional data. A wider-tailed distribution, such as the t-distribution, is used to simulate high-dimensional similarity in a low-dimensional space. By continuously bringing low-dimensional points closer or pushing them apart, the distribution in the low-dimensional space gradually approaches the distribution in the high-dimensional space.

[0145] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0147] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for establishing an emotion regulation model based on emotion monitoring, characterized in that, The method includes: Step 1: Collect multidimensional data, including children's static information samples, children's dynamic information samples, and self-assessment samples; Step 2: Based on Ekman's basic emotion theory, multidimensional data is mapped into emotion factors, including emotion latency, emotion recovery speed, and expression diversity; dimensionality reduction is performed through principal component analysis to screen for highly correlated features; and an LSTM neural network is used to predict the risk of emotional outbursts. Step 3: Based on the extended process model of emotion regulation, regulate emotions in a timely manner; combine the t-SNE algorithm to visualize the trajectory of emotion development and generate customized regulation plans; The process of mapping multidimensional data into emotion factors is as follows: Based on multidimensional data, we identify emotion-triggered events and reaction times, and combine the Ekman emotion classification model to associate the dominant emotion type during the latency period. We quantify the peak and recovery speed of emotions through physiological indicators and behavioral characteristics. We use text emojis, speech rate and interactive operations to statistically analyze the distribution of emotion types, and calculate the diversity of emotion expression through Shannon entropy. The process of quantifying emotional peaks and recovery speed is as follows: emotional peaks are identified and their intensity is quantified based on multidimensional data; recovery speed is calculated based on baseline state; multimodal weighted fusion and half-life method are used for modeling; and Ekman emotion classification model is combined to dynamically monitor and evaluate emotional peaks and recovery processes.

2. The method for establishing an emotion regulation model for emotion monitoring according to claim 1, characterized in that, The process of collecting multidimensional data is as follows: By collecting heart rate variability and skin conductance indicators through wearable devices, an individual emotional baseline is established; computer vision technology is used to extract facial micro-expressions and limb movement frequency behavioral characteristics; and voice samples of children over a period of time are randomly captured, and changes in speech rate, tone, and loudness are analyzed through a speech analysis module.

3. The method for establishing an emotion regulation model for emotion monitoring according to claim 2, characterized in that, The process of establishing an individual's emotional baseline is as follows: Establishing an individual's emotional baseline includes both static and dynamic baselines; By collecting HRV and EDA data from users in a resting state, calculating the mean and standard deviation, a personalized reference range is obtained, and a static baseline is acquired. By analyzing the HRV / EDA variation patterns of users in different contexts using machine learning models, the current state is predicted to deviate from the baseline, and a dynamic baseline is obtained.

4. The method for establishing an emotion regulation model for emotion monitoring according to claim 1, characterized in that, The process of screening highly correlated features is as follows: The method compresses n features from multidimensional data into a low-dimensional space and k principal components, retaining the main information. By analyzing the loading matrix of the principal components, the correlation between the original features and the principal components is analyzed, features that contribute significantly to the principal components are identified, features that contribute little to the principal components are deleted, and the effectiveness is verified through visualization or model performance.

5. The method for establishing an emotion regulation model for emotion monitoring according to claim 4, characterized in that, The process of using an LSTM neural network to predict the risk of emotional outburst is as follows: S201: Preprocess the collected multidimensional data, serialize and fill in the dynamic information samples, and unify the dimensions; fuse the static information samples with time series data to form a multimodal input; S202: Input multidimensional sequence data, use stacked LSTM layers to capture long-term dependencies in time series, add a fully connected layer after LSTM, and filter features that are highly correlated with the risk of emotional dysregulation. Through a gating mechanism, the information flow is dynamically adjusted by the input gate, forget gate, and output gate; the current moment's emotion prediction result is output. S203: Divide the multidimensional dataset into time series segments and monitor metrics by tracking accuracy indicators; evaluate model robustness using time series cross-validation; and optimize the number of layers, number of units, and learning rate parameters of the LSTM through grid search.

6. The method for establishing an emotion regulation model for emotion monitoring according to claim 1, characterized in that, The process of regulating emotions is as follows: Based on the risk of emotional outbursts, emotional states are categorized and the risk level is assessed. Key emotional factors are extracted through principal component analysis. If environmental triggers for negative emotions are identified, the intensity of emotions is reduced by adjusting the environment, modifying the situation, and guiding children to focus on positive stimuli. The intensity of negative emotions is reduced by reinterpreting the situation through language guidance. If the current strategy is ineffective, the strategy is adjusted in a timely manner.

7. The method for establishing an emotion regulation model for emotion monitoring according to claim 1, characterized in that, The process of visualizing the emotion development trajectory using the t-SNE algorithm is as follows: Continuous sentiment data is divided into fixed-length time windows to form time series segments. Key features of each time window are extracted to construct a high-dimensional feature matrix. Calculate the similarity between each sample point and other points, and convert the distance between each sample point and other points into a similarity value; the closer the distance, the higher the similarity; the farther the distance, the lower the similarity. The similarity of each point to other points is symmetrically processed, and the high-dimensional data is compressed into a low-dimensional space.

8. The method for establishing an emotion regulation model for emotion monitoring according to claim 7, characterized in that, The process of compressing high-dimensional data into a low-dimensional space is as follows: Based on the t-SNE algorithm, high-dimensional emotion data is mapped to a low-dimensional space. The similarity between points is modeled by the t-distribution, and the points are dynamically adjusted to make the low-dimensional distribution approximate the high-dimensional structure. Gradient descent optimization is used to bring high-dimensional similar points closer and push away dissimilar points, capturing the stable emotional state clustering state of children.

Citation Information

Patent Citations

  • Teenager mental health data analysis and early warning system

    CN117912710A

  • Emotional state monitoring method and system based on skin resistance change

    CN120036787A