Psychological emotion recognition and intervention system based on eye movement tracking

By using an eye-tracking-based psychological emotion recognition system, users' emotional changes can be captured in real time, generating personalized intervention strategies. This solves the problem of subjectivity in assessment results and mismatch between intervention strategies in traditional methods, and improves the timeliness and effectiveness of psychological intervention.

CN121400831APending Publication Date: 2026-01-27SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511594114.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Traditional mental health assessment and intervention methods struggle to capture rapid, dynamic fluctuations in emotions in real time. Assessment results are easily influenced by subjective will and memory bias, and intervention strategies lack real-time, objective data feedback, failing to accurately match an individual's true needs and long-term response patterns.

Method used

An eye-tracking-based psychological emotion recognition system is adopted. By acquiring high-definition eye-tracking data from users, features are extracted, and emotional state labels are generated using a knowledge rule engine and decision state machine. Combined with a personalized intervention strategy library, real-time intervention instructions are provided, and the intervention strategy is optimized through a feedback module.

Benefits of technology

It enables real-time capture of rapid and dynamic fluctuations in emotions, provides personalized intervention strategies, improves the timeliness and effectiveness of psychological intervention, and adapts to the individual's real needs and long-term reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121400831A_ABST
    Figure CN121400831A_ABST
Patent Text Reader

Abstract

The invention relates to a psychological emotion recognition and intervention system based on eye movement tracking. The system comprises a feature module used for acquiring original eye movement data of a user and extracting features of the original eye movement data to obtain eye movement features; the inference module is used for inputting the eye movement characteristics into a preset knowledge rule engine to obtain an inference conclusion of the physiological state of the user; the state module is used for generating an emotional state label through a decision state machine based on the inference conclusion and the environment context information; and the intervention module is used for querying a personalized intervention strategy library based on the emotional state tag to obtain an intervention instruction for the current emotion of the user. By adopting the system, the rapid dynamic fluctuation of the emotion of the user can be captured in real time, and the current real demand and long-term response of the user are accurately adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of behavior recognition and intervention, and in particular relates to a psychological emotion recognition and intervention system based on eye tracking. Background Technology

[0002] With the development of wearable physiological signal sensing and computing technologies, emotion recognition methods based on multimodal data-driven approaches have emerged. This technology is characterized by its non-invasive, continuous, and quantitatively objective nature, gradually changing the traditional mental health assessment and intervention model that primarily relies on subjective self-report scales or clinical observation. In traditional techniques, the assessment of an individual's emotional state is often indirectly inferred through self-reports or behavioral observations by professionals. Interventions typically rely on pre-set, standardized protocols or are entirely decided on-the-spot by therapists based on experience. Currently, traditional methods have significant limitations: assessment results are easily influenced by subjective will and recall bias, making it difficult to capture rapid, dynamic fluctuations in emotions in real time; the selection of intervention strategies lacks real-time, objective data feedback from the user's physiological state, failing to accurately match the individual's current real needs and long-term response patterns, and facing bottlenecks in timeliness, personalization, and effectiveness. Summary of the Invention

[0003] Therefore, it is necessary to provide an eye-tracking-based psychological emotion recognition and intervention system that can capture rapid dynamic fluctuations in emotions in real time and accurately adapt to an individual's current real needs and long-term reaction patterns, addressing the aforementioned technical problems.

[0004] Firstly, this application provides a psychological emotion recognition and intervention system based on eye-tracking, including:

[0005] The feature module is used to acquire the user's raw eye movement data and extract the features of the raw eye movement data to obtain eye movement features; the raw eye movement data is a series of high-definition videos of the surface of the eyeball;

[0006] The inference module is used to input eye-tracking features into a preset knowledge rule engine to obtain inferences about the user's physiological state.

[0007] The state module is used to generate emotional state labels based on inference conclusions and environmental context information through a decision state machine; the emotional state labels represent the user's psychological emotions.

[0008] The intervention module is used to query the personalized intervention strategy library based on the emotional state tag to obtain intervention instructions for the user's current emotion; the intervention instructions are used to instruct the user to perform psychological adjustment.

[0009] Furthermore, the system also includes a feedback module for:

[0010] The short-term average value of eye movement features was used as the baseline to obtain the assessment benchmark value; and based on the emotional state label, the physiological indicators that the intervention instructions need to regulate and the corresponding expected changes in physiological indicators were determined.

[0011] Collect raw eye-tracking data after the intervention command is issued; and calculate intervention eye-tracking features based on the raw eye-tracking data;

[0012] Based on trend analysis algorithms, the direction and magnitude of changes in interventional eye movement features relative to the assessment baseline are calculated to obtain quantitative conclusions.

[0013] By matching quantitative conclusions with expected changes in physiological indicators, the effectiveness of intervention instructions can be determined, and an evaluation conclusion can be obtained.

[0014] Based on the evaluation results, the personalized intervention instruction library for each user will be updated.

[0015] Furthermore, the feedback module is also used for:

[0016] For each physiological indicator in the intervention eye movement feature, a fixed-length time series data is extracted; and the corresponding evaluation benchmark value of the time series data is extracted to obtain the baseline value.

[0017] The average value of time series data can be calculated using the following formula:

[0018]

[0019] in, To evaluate the average value within the window, x j Let j be the j-th data point in the window, n be the total number of data points in the window, and j be the index of the data point.

[0020] Based on the average and baseline values, the direction and magnitude of changes in physiological indicators are calculated using the following formula:

[0021]

[0022] Among them, P i The value represents the magnitude of the change, and the sign indicates the direction of the change. B is the average value. i The baseline value is ΔX. i The change;

[0023] Quantitative conclusions are generated based on the direction and magnitude of the change.

[0024] Furthermore, the state module is also used for:

[0025] Based on a pre-defined rule mapping table, a decision state machine is used to map the inferred conclusions and environmental context information into specific emotions, thereby obtaining contextualized state hypotheses. The environmental context information includes at least one of the following: current task type, task stage, and custom label.

[0026] Confidence scores are calculated for contextualized hypotheses to obtain confidence ratings;

[0027] If the confidence rating is not lower than the preset threshold, the contextualized state hypothesis will be identified as the emotional state label.

[0028] Furthermore, the state module is also used for:

[0029] If the confidence rating is lower than the preset threshold, higher-order features are extracted from the original eye-tracking data to obtain higher-order feature vectors; and the eye-tracking features, environmental context information and higher-order feature vectors are integrated to obtain standardized feature vectors.

[0030] By inputting standardized feature vectors into the emotion distribution model, the probability distribution of the inferred conclusions belonging to each predefined emotion type is obtained.

[0031] Based on the probability distribution, calculate the contribution of each input feature in the standardized feature vector to the probability distribution;

[0032] Based on contribution and probability distribution, a model-assisted consultation report is generated; the consultation report is used as input to the decision state machine to assist in generating emotional state labels.

[0033] Furthermore, the intervention module is also used for:

[0034] Using emotional state tags as an index, the personalized intervention strategy library is traversed to extract the corresponding candidate strategies and obtain a list of candidate intervention strategies.

[0035] Calculate the expected utility score for each candidate strategy in the candidate intervention strategy list; and sort the candidate strategies based on the expected utility scores to obtain the candidate strategy score list.

[0036] Based on the selection rules, candidate strategies are selected from the candidate strategy scoring list to obtain basic candidate strategies;

[0037] Based on the user's personal preferences, the detailed parameters of the basic candidate strategy are modified to obtain the intervention instructions.

[0038] Furthermore, the feature module is also used for:

[0039] For each frame of video image in the raw eye-tracking data, the features of the eyeball are identified by computer vision algorithms to obtain the coordinates of two-dimensional feature points and the sequence of pupil diameter.

[0040] Based on the geometric mapping model, the coordinates of two-dimensional feature points are mapped to the coordinates of the user's gaze point on the screen to obtain the real-time gaze point coordinates.

[0041] Based on real-time fixation point coordinates and pupil diameter sequences, the raw eye movement data is segmented into eye movement event sequences;

[0042] Based on the analysis time window, quantitative statistical calculations are performed on the eye movement event sequence and pupil diameter sequence to obtain eye movement characteristics; the eye movement characteristics include at least one of pupil-related indicators, fixation-related indicators, saccade-related indicators and blink-related indicators.

[0043] Secondly, this application also provides a method for psychological emotion recognition and intervention based on eye-tracking, including:

[0044] The user's raw eye-tracking data is acquired, and features of the raw eye-tracking data are extracted to obtain eye-tracking features; the raw eye-tracking data is a series of high-definition videos of the surface of the eyeball;

[0045] Eye-tracking features are input into a pre-defined knowledge rule engine to obtain inferences about the user's physiological state.

[0046] Based on the inference conclusions and environmental context information, emotional state labels are generated through a decision state machine; the emotional state labels represent the user's psychological emotions.

[0047] Based on the emotional state tags, the system queries the personalized intervention strategy library to obtain intervention instructions for the user's current emotions; these instructions are used to guide the user to make psychological adjustments.

[0048] Thirdly, this application also provides a computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform any of the steps performed by the system provided in the first aspect of this application.

[0049] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs any of the steps performed by the system provided in the first aspect of this application.

[0050] The aforementioned eye-tracking-based psychological emotion recognition and intervention system comprises the following modules: a feature module for acquiring the user's raw eye-tracking data and extracting its features; the raw eye-tracking data consists of a series of high-definition videos of the eyeball surface; an inference module for inputting the eye-tracking features into a preset knowledge rule engine to obtain inferences about the user's physiological state; a state module for generating emotional state labels based on the inferences and environmental context information via a decision state machine; these emotional state labels represent the user's psychological emotions; and an intervention module for querying a personalized intervention strategy library based on the emotional state labels to obtain intervention instructions tailored to the user's current emotion; these intervention instructions instruct the user to perform psychological adjustment. This system can capture pupil changes in real time and, combined with environmental context information, accurately capture rapid dynamic fluctuations in emotions. The selection of intervention strategies is based on real-time user data feedback, enabling precise adaptation to individual needs and long-term responses, thus improving the timeliness, personalization, and effectiveness of psychological intervention. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a schematic diagram of the structure of a psychological emotion recognition and intervention system based on eye tracking, provided in an embodiment of the present invention.

[0053] Figure 2 This is a schematic diagram of the process of a psychological emotion recognition and intervention method based on eye tracking, provided in an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0055] In one embodiment, such as Figure 1 As shown, a psychological emotion recognition and intervention system 100 based on eye tracking is provided. This embodiment illustrates the application of the system to a terminal. It is understood that the system can also be applied to a server, or to a device including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the system includes the following structure:

[0056] The feature module 101 is used to acquire the user's raw eye movement data and extract the features of the raw eye movement data to obtain eye movement features; the raw eye movement data is a series of high-definition videos of the surface of the eyeball.

[0057] Raw eye-tracking data refers to a series of high-definition videos of the user's eye surface collected using a dedicated eye tracker. These videos record subtle eye movements, pupil changes, and corneal reflections at extremely high frame rates. Eye-tracking features refer to a series of quantitative indicators extracted from the raw videos. These indicators objectively reflect the patterns and physiological states of eye movements and contain information across multiple dimensions. For example, they may include pupil diameter changes reflecting neural excitability, fixation point coordinates and duration reflecting attention allocation, saccade speed and angle reflecting information search strategies, and blink frequency and duration reflecting cognitive load or fatigue levels. The module continuously and non-invasively records eye videos of the user while performing specific tasks via eye tracking. It processes each frame of video image and uses computer vision algorithms to accurately identify the outline of the pupil and key feature points on the eyeball. Through a geometric optics model, it maps the two-dimensional position of the eyeball in the image to the coordinates of the screen the user is viewing, thereby determining the location of the user on the screen. Based on the movement trajectory of the gaze point, the continuous eye movement data is segmented into different events, including fixation, saccades, and blinks. Within a set time window, these events and the pupil diameter are statistically calculated to output a series of numerical eye movement features.

[0058] The inference module 102 is used to input eye movement features into a preset knowledge rule engine to obtain inference conclusions about the user's physiological state.

[0059] Specifically, the knowledge rule engine is a logical reasoning system with built-in expert knowledge, containing a rule base that stores a large number of if-then conditional rules. These rules are predefined based on research findings in eye-tracking physiology and psychology. The inferred physiological state refers to a professional judgment of the user's current basic physiological state. For example, this might include high arousal level, heavy cognitive load, focused attention, or signs of fatigue; it is a purely physiological description and not yet linked to specific contexts or emotions. The terminal matches the calculated eye-tracking feature values ​​one by one with the rule conditions in the knowledge rule engine. When the input feature meets the conditions of one or more rules, the conclusion corresponding to that rule is activated.

[0060] The state module 103 is used to generate emotional state labels through a decision state machine based on inference conclusions and environmental context information; the emotional state labels represent the user's psychological emotions.

[0061] Specifically, environmental context information is auxiliary information used to understand the user's current scenario. This includes the type and stage of the task the user is currently performing, as well as any custom tags that may be added. It is crucial for accurately interpreting physiological signals. The decision state machine receives inferences and contextual information and, based on pre-defined mapping logic, determines the most likely emotional state. It can handle complex situations; for example, the same high-arousal physiological state might be interpreted as anxiety in an exam context, but as excitement while watching a game. The emotional state label is the result of the entire emotion recognition process. It is a specific label representing the user's current psychological emotion and can be categorized as anxiety, calmness, joy, frustration, or focus. The terminal combines the inferences about the physiological state with environmental context information through the decision state machine. Through internal decision logic, it evaluates various possible emotional hypotheses and ultimately selects the most reasonable emotional state label for the current situation.

[0062] The intervention module 104 is used to query the personalized intervention strategy library based on the emotional state tag to obtain intervention instructions for the user's current emotion; the intervention instructions are used to instruct the user to perform psychological adjustment.

[0063] The personalized intervention strategy library is a database storing various intervention methods, customized for different users. It contains strategies for coping with different emotions, each with specific execution steps and parameters. An intervention instruction is a specific, executable guide presented to the user in text, voice, animation, or other formats. The terminal uses emotion state tags as an index to search the personalized intervention strategy library for all candidate strategies applicable to that emotion. Based on the strategy's historical effectiveness, user preferences, and other factors, it selects one or more optimal strategies and translates them into a specific intervention instruction, which is then issued to the user.

[0064] This embodiment provides a psychological emotion recognition and intervention system based on eye-tracking. The system includes a feature module for acquiring the user's raw eye-tracking data and extracting its features to obtain eye-tracking characteristics. The raw eye-tracking data consists of a series of high-definition videos of the eye's surface. An inference module inputs the eye-tracking characteristics into a preset knowledge rule engine to obtain inferences about the user's physiological state. A state module generates emotional state labels based on the inferences and environmental context information using a decision state machine. These emotional state labels represent the user's psychological emotions. An intervention module queries a personalized intervention strategy library based on the emotional state labels to obtain intervention instructions tailored to the user's current emotion. These intervention instructions instruct the user to perform psychological adjustment. This structure enables real-time capture of precise pupil changes and, combined with environmental context information, effectively captures rapid emotional fluctuations. The selection of intervention strategies is based on real-time user data feedback, accurately adapting to individual needs and long-term responses, thus improving the timeliness, personalization, and effectiveness of psychological intervention.

[0065] In one embodiment, the system further includes a feedback module for:

[0066] Step 201: Use the short-term average value of eye movement features as the baseline to obtain the assessment benchmark value; and based on the emotional state label, determine the physiological indicators that the intervention instructions need to regulate and the corresponding expected changes in physiological indicators.

[0067] The assessment baseline, also known as the benchmark value, represents the user's physiological level before intervention, in a relatively stable or natural state. It is the short-term average of recent eye movement characteristics, serving as an objective reference point for subsequent measurement of the changes caused by the intervention. Physiological indicators are specific, quantifiable dimensions of eye movement characteristics that can be directly or indirectly affected by intervention methods; each indicator corresponds to a specific physiological or psychological state. The expected change in physiological indicators is a pre-set intervention goal based on emotional state labels, specifying the desired direction and approximate degree of change for a particular physiological indicator. The terminal analyzes user eye movement data within a short time window before issuing the intervention command, calculates the average value of various key eye movement characteristics, and establishes this average value as the assessment baseline. Based on the currently identified emotions and using a built-in psychological model, it determines which physiological indicators are the key targets for this intervention and sets the expected direction of change.

[0068] Step 202: Collect raw eye movement data after the intervention command is issued; and calculate the intervention eye movement features based on the raw eye movement data.

[0069] Specifically, intervention eye-tracking features are a new round of eye-tracking features calculated immediately after the system issues an intervention command and the user executes it. These features are derived from raw eye-tracking data collected and processed through a feature extraction process, representing the user's immediate physiological state during or after the intervention. After the intervention begins, the terminal continues to record video of the user's eyes using an eye tracker and employs the exact same processing flow as the feature module to extract new, analyzable intervention eye-tracking features from this new video data.

[0070] Step 203: Based on the trend analysis algorithm, calculate the direction and magnitude of the change in the intervention eye movement features relative to the assessment benchmark value, and obtain quantitative conclusions.

[0071] Specifically, trend analysis algorithms are a set of mathematical methods used to quantify the changes of a data sequence relative to another reference sequence, analyze trends over a period of time, and ensure the robustness of the results. The direction and magnitude of change are precise descriptions of the changes; direction refers to whether the physiological indicator has increased or decreased; magnitude refers to the size of the change, usually expressed as a percentage or absolute difference. The quantitative conclusion is the final quantitative description of the changes in each physiological indicator of interest. The terminal compares the values ​​of each indicator in the intervention eye-tracking features with the corresponding assessment benchmark values, and by calculating the percentage change, accurately calculates the actual extent of change in each indicator.

[0072] Step 204: Match the quantitative conclusions with the expected changes in physiological indicators to determine whether the intervention instructions are effective and obtain the evaluation conclusion.

[0073] The evaluation conclusion is a binary judgment regarding the effectiveness of the intervention instruction, directly derived from the matching result between the actual changes and the expected goals. The terminal performs logical matching and judgment, comparing the quantitative conclusion with the set expected changes in physiological indicators. If the direction of the actual change is consistent with the expectation, and the magnitude reaches or exceeds the expected threshold, it is judged as effective; otherwise, it is judged as ineffective.

[0074] Step 205: Based on the evaluation results, update the user's corresponding personalized intervention instruction library.

[0075] The terminal adjusts the personalized intervention strategy library based on the evaluation results. If an intervention instruction is proven to be effective for a user under a certain emotion multiple times, the priority or weight of the instruction is increased, making it more likely to be selected in the future. If an instruction is proven to be ineffective multiple times, its weight is reduced, or it is no longer recommended in similar situations, thereby achieving true personalized adaptation.

[0076] This embodiment continuously cycles through adaptation and self-learning to optimize intervention assessments, gaining a deeper understanding of each user's unique response patterns. This allows for increasingly precise and effective psychological adjustment suggestions, enhancing the personalization and accuracy of emotional interventions.

[0077] In one embodiment, the feedback module is further configured to:

[0078] Step 301: For each physiological indicator in the intervention eye movement features, extract time series data of a fixed length; and extract the evaluation benchmark value of the corresponding time series data to obtain the baseline value.

[0079] Physiological indicators refer to specific, quantifiable dimensions selected from the intervention's eye-tracking characteristics, and each indicator will be analyzed independently. Fixed-length time-series data consists of a series of observations arranged chronologically over a specific period after the intervention for a given physiological indicator; the fixed length ensures consistency within the analysis window. The baseline value is a reference value representing the stable state of the corresponding physiological indicator before the intervention, derived from the assessment benchmark. For the physiological indicator currently being analyzed, its corresponding benchmark value is directly used as the baseline value. The terminal performs data slicing, extracting a predefined data segment from the continuous data stream after the intervention begins for each physiological indicator to be evaluated. This segment contains dynamic information about the indicator's changes over time during the intervention. The average value of the indicator that is identical to the currently analyzed physiological indicator is found from the stored benchmark data and designated as the benchmark value for this comparison.

[0080] Step 302: Calculate the average value of the time series data using the following formula:

[0081]

[0082] in, To evaluate the average value within the window, x j Let j be the j-th data point in the window, n be the total number of data points in the window, and j be the index of the data point.

[0083] Specifically, the average value within the assessment window represents the arithmetic mean of all data points in the extracted fixed-length time series data. This smooths out short-term fluctuations and reflects the central trend or general level of the physiological indicator within the intervention window. The terminal performs data summation by adding all data points in the time series and dividing by the total number of data points, using a single, stable value to represent the typical performance over the entire post-intervention period.

[0084] Step 303: Based on the average and baseline values, calculate the direction and magnitude of changes in physiological indicators using the following formula:

[0085]

[0086] Among them, P i The value represents the magnitude of the change, and the sign indicates the direction of the change. B is the average value. i The baseline value is ΔX. i The variable is the amount of change.

[0087] Specifically, the change is the absolute difference between the post-intervention average level and the baseline level, directly reflecting the absolute magnitude of the change. The rate of change, also often called the percentage change, is the ratio of the change to the baseline value. Standardizing the change eliminates the influence of varying initial values ​​for different indicators, allowing for comparison of the magnitudes of change between different indicators. The terminal uses subtraction to obtain the absolute value of the post-intervention average deviation from the baseline; the sign directly indicates the direction of change. Division transforms the absolute change into a proportion relative to the initial level; the absolute value represents the magnitude of the change, and its sign is consistent with the change amount, further confirming the direction.

[0088] Step 304: Generate quantitative conclusions based on the direction and magnitude of change.

[0089] Specifically, the quantitative conclusion is the final output of this process. It is a structured description of the changes in each physiological indicator based on the calculation results. The terminal integrates the numerical values ​​and signs of the changes and rates of change, along with the corresponding names of the physiological indicators, to form a clear and quantitative description that objectively records the actual effects of the intervention on the physiological parameters.

[0090] This embodiment provides a complete mathematical description of the intervention effect by generating two key quantitative outputs, which enables the intervention effect to be quantified, thereby improving the accuracy of emotion recognition and intervention.

[0091] In one embodiment, the status module 103 is further configured to:

[0092] Step 401: Based on the preset rule mapping table, the inference conclusion and environmental context information are mapped to specific emotions through the decision state machine to obtain the contextualized state hypothesis; the environmental context information includes at least one of the current task type, task stage and custom label.

[0093] The rule mapping table is a lookup table or database storing expert knowledge, defining the mapping relationship between different combinations of physiological states and environmental context information and emotions. For example, a rule could be: if the inference conclusion is high arousal and the task type is coping with an urgent deadline, then the emotion is assumed to be anxiety. The decision state machine is a logical model that transitions from one state to another based on the current input conditions. In this embodiment, the decision state machine executes the logic in the rule mapping table, using the current input as a condition and matching rules to transition from an undetermined state to a specific emotion assumption state. The contextualized state assumption is a preliminary emotion judgment generated by the decision state machine based on rules; it is called an assumption and has not yet undergone credibility verification, representing only the most probable result based on the rules. The terminal inputs the inference conclusion from the inference module and the perceived environmental context information into the decision state machine, searching the rule mapping table for a rule that matches both conditions. If a matching rule is found, the emotion output by that rule is taken as the contextualized state assumption.

[0094] Step 402: Calculate the confidence level of the contextualized hypothesis to obtain the confidence level rating.

[0095] Specifically, a confidence rating is a quantified score or grade used to represent the system's degree of confidence in its generated contextualized state assumptions, reflecting the strength of currently available evidence supporting the sentiment hypothesis. The terminal evaluates a range of factors to calculate the confidence rating, including whether the input conditions perfectly match the rule conditions; whether the original eye-tracking data signals used to generate the inference are clear; the significance of key eye-tracking features reflecting the emotion; and the clarity of the current environmental context. These factors are then considered together to generate a confidence rating through rule-based scoring or a simple mathematical model.

[0096] Step 403: If the confidence rating is not lower than the preset threshold, then the contextualized state hypothesis is determined as the emotional state label.

[0097] Specifically, the terminal makes a threshold decision by comparing a pre-set confidence threshold with the calculated confidence rating. If the confidence rating is greater than or equal to the threshold, it means that there is sufficient confidence that the judgment is reliable, and the hypothesis is upgraded to a conclusion, formally determining the contextualized state hypothesis as the final emotional state label.

[0098] This embodiment establishes a rapid decision-making channel, prioritizing lightweight reasoning methods based on explicit rules. Its core advantage lies in efficiency. When the rules are clear and the evidence is conclusive, the system can immediately make a judgment without initiating more complex calculations, which ensures the real-time response of the system. If the conditions are not met, the system will enter a backup decision path. Through this dual guarantee mechanism, while ensuring efficiency, the robustness and accuracy of the system's judgment in complex and ambiguous scenarios are greatly improved.

[0099] In one embodiment, the status module 103 is further configured to:

[0100] Step 501: If the confidence rating is lower than the preset threshold, then higher-order features are extracted from the original eye-tracking data to obtain higher-order feature vectors; and the eye-tracking features, environmental context information and higher-order feature vectors are integrated to obtain standardized feature vectors.

[0101] Higher-order feature vectors (HEMs) are features extracted from raw eye-tracking data using more complex algorithms, building upon basic eye-tracking features. They describe more abstract patterns and, for example, can include dynamic characteristics such as the frequency of pupil diameter fluctuations and temporal patterns like the stability of saccade velocity; interactive characteristics such as synergistic or antagonistic relationships between different eye-tracking indicators; and morphological characteristics such as finer geometric features of pupil shape changes or eye movement trajectories, capable of capturing more subtle physiological signal patterns that may be related to complex emotions. Standardized feature vectors integrate all features from different sources with varying dimensions and ranges into a unified numerical vector that can be processed by machine learning models. This includes feature scaling and vectorization, ensuring that the model is not dominated by features with large numerical values. The terminal bypasses the pre-calculated basic features and returns to the original, most information-rich eye-tracking video data. It then uses more complex signal processing or temporal analysis algorithms to uncover deep patterns that have not been utilized in conventional feature extraction, forming high-order feature vectors. The basic eye-tracking features, environmental context information, and newly extracted high-order feature vectors are then merged. Through encoding and standardization techniques, these three types of information are combined into a unified, structured, and standardized feature vector.

[0102] Step 502: Input the standardized feature vector into the emotion distribution model to obtain the probability distribution of the inferred conclusion belonging to each predefined emotion type.

[0103] Specifically, the emotion distribution model is a trained machine learning classification model that outputs a probability distribution, rather than just a single label. The probability distribution is a set of probability values, each representing the likelihood that the input data belongs to each predefined emotion type, with the sum of all probabilities being 100%. The terminal inputs a standardized feature vector into the trained emotion distribution model, which, based on learned complex patterns, calculates and outputs a probability distribution for all possible emotions.

[0104] Step 503: Based on the probability distribution, calculate the contribution of each input feature in the standardized feature vector to the probability distribution.

[0105] Specifically, contribution, also known as feature importance, quantifies how much each input feature in the standardized feature vector influences the final generated probability distribution. Terminals use SHAP (SHapley Additive Explanation), LIME (Local Interpretable Model-Agnostic Explanations), or the model's built-in feature importance metrics to analyze how the model's output changes when the value of a particular feature is altered, thus deducing the contribution of each feature.

[0106] Step 504: Based on contribution and probability distribution, generate a model-assisted consultation report; the consultation report is used as input to the decision state machine to assist in generating emotional state labels.

[0107] The model-assisted consultation report is a structured summary report containing probability distributions (how likely the model perceives various emotions) and key evidence (which features contribute most to the current judgment). The terminal integrates the probability distribution and contribution analysis results into a concise report, clearly identifying the most likely emotions and their probabilities, and listing the key factors supporting these judgments. This bridges the gap between the data-driven model and the rule-based decision state machine. Upon receiving the consultation report, the decision state machine can use it as powerful supplementary information, combining it with existing rules to conduct a final comprehensive assessment, thereby generating more accurate and reliable emotion state labels.

[0108] This embodiment introduces contribution-based in-depth analysis by constructing a more expensive deep analysis process, which helps the decision state machine to perform the final emotion recognition when the original emotion recognition is at a low confidence level. This greatly improves the emotion recognition capability of marginal and complex cases.

[0109] In one embodiment, the intervention module 104 is further configured to:

[0110] Step 601: Using emotional state tags as indexes, traverse the personalized intervention strategy library, extract the corresponding candidate strategies, and obtain a list of candidate intervention strategies.

[0111] Among these, the emotion state tag is the judgment result of the user's current psychological emotion and is the core basis for this intervention. The personalized intervention strategy library is a database storing various intervention methods. Each strategy in the library indicates its targeted emotion, specific operation steps, required parameters, and historical usage effects, etc., with personalized descriptions recording the user's historical reactions to different strategies. The candidate intervention strategy list is a collection of all strategies applicable to the emotion retrieved from the personalized intervention strategy library using the emotion state tag as the keyword. The terminal performs database queries and searches, using the emotion state tag as the search condition, scanning the intervention strategy library for the applicable emotion tags of each strategy, collecting all matching strategies to form a candidate intervention strategy list.

[0112] Step 602: Calculate the expected utility score for each candidate strategy in the candidate intervention strategy list; and sort the candidate strategies based on the expected utility scores to obtain a candidate strategy score list.

[0113] Specifically, the expected utility score is a predictive value used to quantify the anticipated benefits a candidate strategy might bring to the current user if adopted. A higher score indicates a better expected effect. Calculating the score requires considering multiple factors, which may include: historical effectiveness (the success rate of the strategy on similar emotions experienced by the user in the past); user preference (whether the user explicitly expresses liking or disliking this type of strategy); and contextual suitability (whether the strategy's execution requirements match the current user's environment). The candidate strategy scoring list is a list obtained by attaching the calculated expected utility score to each strategy in the candidate intervention strategy list and sorting them from highest to lowest score. The terminal starts a scoring algorithm for each strategy in the list. The algorithm calls upon various data related to the strategy and the current user, calculates a comprehensive expected utility score using a weighted formula, and sorts all strategies in descending order based on this score to generate the candidate strategy scoring list.

[0114] Step 603: Based on the selection rules, select candidate strategies from the candidate strategy scoring list to obtain basic candidate strategies.

[0115] Specifically, selection rules are a set of strategies that determine how to make the final choice from the ranking list. For example, rules may include: optimal selection rules, which directly select the strategy with the highest score; rotational selection rules, which randomly select one of the top strategies to avoid user boredom; and exploratory selection rules, which occasionally select a new strategy with potential but not the highest score to collect data. The basic candidate strategy is the strategy ultimately selected from the candidate strategy scoring list according to the selection rules, representing the intervention plan that the system considers most suitable under the global rules. The terminal invokes the preset selection rules and applies them to the candidate strategy scoring list, generating a preliminary intervention plan. While theoretically reasonable, this is merely a standard template and has not yet incorporated the user's immediate personal preferences.

[0116] Step 604: Based on the user's personal preferences, modify the detailed parameters of the basic candidate strategy to obtain the intervention instruction.

[0117] The user's personal preferences are detailed settings stored in their user profile, reflecting their preferred intervention methods. For example, these might include a preference for male or female voice guidance; a preference for instrumental music or natural sounds; a preference for short 3-minute exercises or 10-minute in-depth exercises; or a preference for watching animated demonstrations or listening to only voice guidance. The intervention instruction is a specific, executable, and fully personalized guidance message sent to the user. The terminal performs parameter customization, obtains the selected basic candidate strategy, queries the user's personal preference profile, and adjusts the specific execution parameters of the strategy based on the preferences. Optionally, the basic strategy might be a deep breathing exercise, which would be specified based on preference as: "Please follow this female voice's guidance for a 5-minute deep breathing exercise accompanied by the sound of ocean waves."

[0118] This embodiment improves user compliance and the actual effectiveness of interventions by personalizing intervention instructions, ensuring that the system provides an intervention method that users are willing to accept and feel comfortable with.

[0119] In one embodiment, feature module 101 is further configured to:

[0120] Step 701: For each frame of video image in the original eye-tracking data, the features of the eyeball are identified using a computer vision algorithm to obtain the coordinates of two-dimensional feature points and the sequence of pupil diameters.

[0121] The two-dimensional feature point coordinates are the location information of key points identified on the eye image by computer vision algorithms. These can include the pupil center, corneal reflective points, etc. The coordinates are in pixels and describe the precise location of these points in the video image. The pupil diameter sequence is a sequence formed by arranging the pupil diameter values ​​calculated for each frame of the video image in chronological order, reflecting the dynamic changes in pupil size. The terminal performs image recognition and measurement. For each frame of the incoming high-definition video image, it executes a series of complex image processing algorithms. Through edge detection, threshold segmentation, and other methods, it accurately locates the pupil outline. Based on the pupil location, it finds the pupil center, identifies other feature points such as corneal reflective points, and calculates the pupil diameter for that frame based on the pupil outline. This process is continuously applied to each frame in the video stream, thereby outputting the two-dimensional feature point coordinates and pupil diameter sequence that change over time.

[0122] Step 702: Based on the geometric mapping model, the coordinates of the two-dimensional feature points are mapped to the coordinates of the user's gaze point on the screen to obtain the real-time gaze point coordinates.

[0123] Specifically, the geometric mapping model is a mathematical model that establishes the mathematical relationship between the positions of eye feature points in the camera image and the user's gaze point position on the physical screen. The model requires a specific calibration process to personalize it for each user. Real-time gaze point coordinates are the position of the user's gaze in the screen coordinate system, typically represented by screen pixel coordinates, clearly showing the trajectory of the user's visual attention. The terminal performs coordinate transformation, inputting the two-dimensional feature point coordinates representing eye posture into the pre-calibrated geometric mapping model. The model calculates and outputs the precise position on the screen corresponding to the current gaze direction, i.e., the real-time gaze point coordinates.

[0124] Step 703: Based on the real-time gaze point coordinates and pupil diameter sequence, the original eye movement data is segmented into an eye movement event sequence.

[0125] Specifically, an eye-tracking event sequence is a chronological sequence of different types of basic eye movements, obtained by classifying continuous fixation trajectory data. The main event types include fixation (where the gaze remains relatively stable on a target while the brain processes visual information), saccades (rapid movement of the gaze between two fixation points during which almost no visual information is acquired), and blinking (the process of the eyelids closing and reopening, which results in a data gap). The terminal uses a velocity-threshold algorithm to analyze the real-time fixation coordinate sequence, determining which time periods belong to fixation and which to saccades based on the fixation point's movement speed, acceleration, and positional stability. Sudden changes or data gaps in the pupil diameter sequence can be used to assist in detecting blinking events. The continuous trajectory is segmented into a series of discrete, meaningful eye-tracking events.

[0126] Step 704: Based on the analysis time window, perform quantitative statistical calculations on the eye movement event sequence and pupil diameter sequence to obtain eye movement features; the eye movement features include at least one of pupil-related indicators, fixation-related indicators, saccade-related indicators, and blink-related indicators.

[0127] The analysis time window is a time range used for statistical calculations. Sliding this window continuously calculates the latest features for real-time analysis. Eye movement features are generalized indicators obtained by statistically calculating eye movement event sequences and pupil diameter sequences within the analysis time window. Examples include: pupil-related indicators (mean pupil diameter, pupil diameter standard deviation); fixation-related indicators (number of fixations, average fixation duration); saccade-related indicators (number of saccades, average saccade amplitude and velocity); and blink-related indicators (blink frequency, average blink duration). Within the set analysis time window, the terminal extracts all relevant data, performs statistical calculations, and the calculated values ​​constitute the eye movement features.

[0128] This embodiment generates a fixed-dimensional numerical feature vector that represents the user's eye-tracking behavior pattern over a recent period. This vector serves as the direct data basis for the system's subsequent emotion inference. It extracts low-dimensional, meaningful, and computable behavioral biomarkers from high-dimensional, redundant raw video data, providing solid data support for upper-level emotion computing and improving the accuracy of emotion recognition and intervention.

[0129] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0130] Based on the same inventive concept, this application also provides an eye-tracking-based method for implementing the aforementioned eye-tracking-based psychological emotion recognition and intervention system. The solution provided by this method is similar to the implementation scheme described in the above system. Therefore, the specific limitations of one or more eye-tracking-based psychological emotion recognition and intervention method embodiments provided below can be found in the above-described limitations of an eye-tracking-based psychological emotion recognition and intervention system, and will not be repeated here.

[0131] In one exemplary embodiment, such as Figure 2 As shown, a psychological emotion recognition and intervention method 800 based on eye tracking is provided, including:

[0132] Step 801: Obtain the user's raw eye movement data and extract the features of the raw eye movement data to obtain eye movement features; the raw eye movement data is a series of high-definition videos of the surface of the eyeball;

[0133] Step 802: Input the eye-tracking features into the preset knowledge rule engine to obtain inferences about the user's physiological state;

[0134] Step 803: Based on the inference conclusion and environmental context information, generate emotional state labels through a decision state machine; the emotional state labels represent the user's psychological emotions.

[0135] Step 804: Based on the emotional state tag, query the personalized intervention strategy library to obtain intervention instructions for the user's current emotion; the intervention instructions are used to instruct the user to perform psychological adjustment.

[0136] Furthermore, after querying the personalized intervention strategy library based on the emotion state tag to obtain intervention instructions for the user's current emotion, it also includes:

[0137] The short-term average value of eye movement features was used as the baseline to obtain the assessment benchmark value; and based on the emotional state label, the physiological indicators that need to be regulated by the intervention instructions and the corresponding expected changes in the indicators were determined.

[0138] Collect raw eye-tracking data after the intervention command is issued; and calculate intervention eye-tracking features based on the raw eye-tracking data;

[0139] Based on trend analysis algorithms, the direction and magnitude of changes in interventional eye movement features relative to the assessment baseline are calculated to obtain quantitative conclusions.

[0140] By matching quantitative conclusions with expected changes in physiological indicators, the effectiveness of intervention instructions can be determined, and an evaluation conclusion can be obtained.

[0141] Based on the evaluation results, the personalized intervention instruction library for each user will be updated.

[0142] Furthermore, the quantitative conclusions are matched with the expected changes in physiological indicators to determine the effectiveness of the intervention instructions and obtain evaluation conclusions, including:

[0143] For each physiological indicator in the intervention eye movement feature, a fixed-length time series data is extracted; and the corresponding evaluation benchmark value of the time series data is extracted to obtain the baseline value.

[0144] The average value of time series data can be calculated using the following formula:

[0145]

[0146] in, To evaluate the average value within the window, x j Let j be the j-th data point in the window, n be the total number of data points in the window, and j be the index of the data point.

[0147] Based on the average and baseline values, the direction and magnitude of changes in physiological indicators are calculated using the following formula:

[0148]

[0149] Among them, P i The value represents the magnitude of the change, and the sign indicates the direction of the change. B is the average value. i The baseline value is ΔX. i The change;

[0150] Quantitative conclusions are generated based on the direction and magnitude of the change.

[0151] Furthermore, the eye-tracking features are input into a pre-defined knowledge rule engine to obtain inferences about the user's physiological state, including:

[0152] Based on a pre-defined rule mapping table, a decision state machine is used to map the inferred conclusions and environmental context information into specific emotions, thereby obtaining contextualized state hypotheses. The environmental context information includes at least one of the following: current task type, task stage, and custom label.

[0153] Confidence scores are calculated for contextualized hypotheses to obtain confidence ratings;

[0154] If the confidence rating is not lower than the preset threshold, the contextualized state hypothesis will be identified as the emotional state label.

[0155] Furthermore, after calculating the confidence level of the contextualized hypothesis and obtaining the confidence rating, the following steps are also included:

[0156] If the confidence rating is lower than the preset threshold, higher-order features are extracted from the original eye-tracking data to obtain higher-order feature vectors; and the eye-tracking features, environmental context information and higher-order feature vectors are integrated to obtain standardized feature vectors.

[0157] By inputting standardized feature vectors into the emotion distribution model, the probability distribution of the inferred conclusions belonging to each predefined emotion type is obtained.

[0158] Based on the probability distribution, calculate the contribution of each input feature in the standardized feature vector to the probability distribution;

[0159] Based on contribution and probability distribution, a model-assisted consultation report is generated; the consultation report is used as input to the decision state machine to assist in generating emotional state labels.

[0160] Furthermore, based on the emotion state tag, the personalized intervention strategy library is queried to obtain intervention instructions targeting the user's current emotion, including:

[0161] Using emotional state tags as an index, the personalized intervention strategy library is traversed to extract the corresponding candidate strategies and obtain a list of candidate intervention strategies.

[0162] Calculate the expected utility score for each candidate strategy in the candidate intervention strategy list; and sort the candidate strategies based on the expected utility scores to obtain the candidate strategy score list.

[0163] Based on the selection rules, candidate strategies are selected from the candidate strategy scoring list to obtain basic candidate strategies;

[0164] Based on the user's personal preferences, the detailed parameters of the basic candidate strategy are modified to obtain the intervention instructions.

[0165] Furthermore, features are extracted from the raw eye-tracking data to obtain eye-tracking features, including:

[0166] For each frame of video image in the raw eye-tracking data, the features of the eyeball are identified by computer vision algorithms to obtain the coordinates of two-dimensional feature points and the sequence of pupil diameter.

[0167] Based on the geometric mapping model, the coordinates of two-dimensional feature points are mapped to the coordinates of the user's gaze point on the screen to obtain the real-time gaze point coordinates.

[0168] Based on real-time fixation point coordinates and pupil diameter sequences, the raw eye movement data is segmented into eye movement event sequences;

[0169] Based on the analysis time window, quantitative statistical calculations are performed on the eye movement event sequence and pupil diameter sequence to obtain eye movement characteristics; the eye movement characteristics include at least one of pupil-related indicators, fixation-related indicators, saccade-related indicators and blink-related indicators.

[0170] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of an eye-tracking-based psychological emotion recognition and intervention system as described above.

[0171] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0172] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0173] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A psychological emotion recognition and intervention system based on eye-tracking, characterized in that, The system includes: The feature module is used to acquire the user's raw eye movement data and extract the features of the raw eye movement data to obtain eye movement features; the raw eye movement data is a series of high-definition videos of the surface of the eyeball; The inference module is used to input the eye movement features into a preset knowledge rule engine to obtain an inference conclusion about the user's physiological state. The state module is used to generate emotional state labels through a decision state machine based on the inference conclusion and environmental context information; the emotional state labels represent the user's psychological emotions. The intervention module is used to query a personalized intervention strategy library based on the emotional state tag to obtain intervention instructions for the user's current emotion; the intervention instructions are used to instruct the user to perform psychological adjustment.

2. The system according to claim 1, characterized in that, The system also includes a feedback module for: The short-term average of the eye movement features is used as a baseline to obtain the evaluation benchmark value; and based on the emotional state label, the physiological indicators that the intervention instruction needs to regulate and the corresponding expected changes in the physiological indicators are determined. Collect the raw eye movement data after the intervention command is issued; And based on the original eye movement data, the intervention eye movement features are calculated; Based on a trend analysis algorithm, the direction and magnitude of the change in the intervention eye movement features relative to the evaluation benchmark value are calculated to obtain quantitative conclusions. By matching the quantitative conclusions with the expected changes in the physiological indicators, the effectiveness of the intervention instructions is determined, and an evaluation conclusion is obtained. Based on the evaluation results, the personalized intervention instruction library corresponding to the user is updated.

3. The system according to claim 2, characterized in that, The feedback module is also used for: For each physiological indicator in the intervention eye movement feature, a fixed length of time series data is extracted; and the corresponding evaluation benchmark value of the time series data is extracted to obtain the baseline value. The average value of the time series data is calculated using the following formula: in, To evaluate the average value within the window, x j Let j be the j-th data point in the window, n be the total number of data points in the window, and j be the index of the data point. Based on the average value and the baseline value, the direction and magnitude of change of the physiological indicator are calculated using the following formula: Among them, P i The value represents the magnitude of the change, and the sign indicates the direction of the change. B is the average value. i The baseline value is ΔX. i The change; The quantitative conclusion is generated based on the direction and magnitude of the change.

4. The system according to claim 1, characterized in that, The status module is also used for: Based on a preset rule mapping table, the decision state machine maps the inference conclusion and the environmental context information to specific emotions to obtain a contextualized state hypothesis; the environmental context information includes at least one of the following: current task type, task stage, and custom label. The confidence level of the contextualized hypothesis is calculated to obtain a confidence rating; If the confidence rating is not lower than a preset threshold, then the contextualized state hypothesis is determined as the emotional state label.

5. The system according to claim 4, characterized in that, The status module is also used for: If the confidence rating is lower than a preset threshold, then higher-order features are extracted from the original eye-tracking data to obtain a higher-order feature vector; and the eye-tracking features, the environmental context information, and the higher-order feature vector are integrated to obtain a standardized feature vector. The standardized feature vector is input into the emotion distribution model to obtain the probability distribution of the inferred conclusion belonging to each predefined emotion type. Based on the probability distribution, calculate the contribution of each input feature in the standardized feature vector to the probability distribution; Based on the contribution and the probability distribution, a model-assisted consultation report is generated; the consultation report is used as input to the decision state machine to assist in generating the emotional state label.

6. The system according to claim 1, characterized in that, The intervention module is also used for: Using the emotional state tags as indexes, the personalized intervention strategy library is traversed to extract the corresponding candidate strategies and obtain a list of candidate intervention strategies. Calculate the expected utility score for each candidate strategy in the candidate intervention strategy list; and sort the candidate strategies based on the expected utility scores to obtain a candidate strategy score list; Based on the selection rules, the candidate strategy is selected from the candidate strategy scoring list to obtain the basic candidate strategy; Based on the user's personal preferences, the detailed parameters of the basic candidate strategy are modified to obtain the intervention instruction.

7. The system according to claim 1, characterized in that, The feature module is also used for: For each frame of video image in the original eye-tracking data, the features of the eyeball are identified by computer vision algorithms to obtain two-dimensional feature point coordinates and pupil diameter sequence; Based on the geometric mapping model, the coordinates of the two-dimensional feature points are mapped to the coordinates of the user's gaze point on the screen to obtain the real-time gaze point coordinates. Based on the real-time fixation point coordinates and the pupil diameter sequence, the raw eye movement data is segmented into an eye movement event sequence; Based on the analysis time window, quantitative statistical calculations are performed on the eye movement event sequence and the pupil diameter sequence to obtain the eye movement features; the eye movement features include at least one of pupil-related indicators, fixation-related indicators, saccade-related indicators and blink-related indicators.

8. A method for psychological emotion recognition and intervention based on eye-tracking, characterized in that, The method includes: The user's raw eye movement data is acquired, and the features of the raw eye movement data are extracted to obtain eye movement features; the raw eye movement data is a series of high-definition videos of the surface of the eyeball. The eye-tracking features are input into a preset knowledge rule engine to obtain an inference about the user's physiological state. Based on the inference conclusions and environmental context information, an emotional state label is generated through a decision state machine; the emotional state label represents the user's psychological emotions. Based on the emotional state tag, a personalized intervention strategy library is queried to obtain intervention instructions for the user's current emotion; the intervention instructions are used to instruct the user to perform psychological adjustment.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements any step of the system according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements any step of the system according to any one of claims 1 to 7.