Psychological state assessment method and system based on environmental perception, and readable storage medium
By using RGB cameras and microphones to analyze environmental and behavioral data, the method addresses the limitations of traditional psychological assessment by quantifying environmental influences, improving accuracy and robustness in real-time monitoring.
Patent Information
- Application Number
- CN202510780093.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing psychological state assessment method relies on user active participation and cannot be applied to real-time monitoring in natural scenarios. It ignores the impact of environmental stimulation on psychological state, resulting in distortion or inaccuracy of evaluation results.
Environmental scene and user behavior data were collected through RGB cameras and microphones, environmental features were extracted using semantic segmentation and sound recognition models, environmental-psychological causal maps were constructed, and psychological state evaluation was performed in combination with causal Transformer model, and causal weights were optimized through personalized calibration.
Real-time and accurate psychological state assessment in natural scenarios is achieved, which improves the accuracy and robustness of the assessment, provides a more comprehensive psychological state explanation, and adapts to diverse environments and individual differences.
Smart Images

Figure CN120304829A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mental state assessment, and particularly to a mental state assessment method, system, and readable storage medium based on environmental perception. Background Art
[0002] Mental state assessment has important application values in fields such as mental health monitoring and social interaction analysis. Traditional assessment methods mainly rely on users to actively fill out mental scales or wear contact physiological sensors, which have the following significant defects: First, users may cause the distortion of scale results due to privacy concerns or cognitive biases, and it requires users' active participation, which is not applicable to real-time monitoring in natural scenarios; second, existing non-contact technologies (such as text / image-based analysis, application document CN111477328B; multi-modal micro-expression analysis, application document CN117936032A) only focus on the user's own behavior data and ignore the direct impact of environmental stimuli, such as social scenarios, noise interference, etc. on the mental state. Common situations include that a closed space may cause anxiety, but the existing methods do not incorporate the spatial layout into the evaluation dimension. Summary of the Invention
[0003] The present invention aims to at least solve one of the technical problems in the related technologies to some extent. For this purpose, the object of the present invention is to propose a mental state assessment method, system, and readable storage medium based on environmental perception to achieve the quantitative impact analysis of the environment on the mental state.
[0004] To achieve the above object, the first aspect embodiment of the present invention proposes a mental state assessment method based on environmental perception, including the following steps: S1. Collect the environmental scene and the behavior image of the object to be evaluated through an RGB camera, and use a semantic segmentation model to extract the object category, spatial layout, and dynamic events; collect the environmental voiceprint through a microphone, and use a voice recognition model to identify the voice category; synchronously collect the facial micro-expression feature points, speech prosody features, and head-neck-shoulder vibration frequencies of the object to be evaluated; S2. Generate semantic vectors for the object / scene labels through a semantic embedding model, convert the spatial layout and voice labels into sparse vectors, and map them to the preset environmental emotion dimensions through a multi-layer perceptron to obtain the stress value, relaxation degree, and social intensity; S3. Construct an environment-psychology causal graph, define the causal relationship nodes and edge weights of the environmental emotion dimensions, behavior characteristics, and mental states; adjust the attention of the environmental characteristics to the behavior characteristics through a causal attention module, and fuse and generate a multi-modal feature vector; S4. Input the fused feature vector into a causal Transformer model including a dynamic causal weight update unit, and output the evaluation result of the mental state dimension; S5. Generate a psychological state interpretation report by combining environmental semantics, and dynamically calibrate the causal weights of the model based on the historical data of the evaluated object.
[0005] In some embodiments of the present invention, in step S1, the semantic segmentation model is Mask R-CNN, which is used to perform real-time semantic segmentation on the images collected by the RGB camera and output the set of object categories, spatial layout labels and dynamic event set in the environment; the sound recognition model is YAMNet, which is used to classify the environmental sound patterns into a predefined set of sound categories .
[0006] In some embodiments of the present invention, in step S2, the environmental emotion dimensions include stress value , relaxation level and social intensity . The quantitative mapping from object, space, and sound labels to emotion dimensions is realized through a predefined semantic-emotion mapping atlas, and the mapping atlas is iteratively optimized by an online knowledge graph update algorithm.
[0007] In some embodiments of the present invention, in step S3, the causal graph modeling step includes: Defining the node set as environmental emotion dimensions, behavioral characteristics, and psychological states; Calculating the causal effect through Do-Calculus as the edge weight, where = , representing the direct causal strength of node on node .
[0008] In some embodiments of the present invention, in step S3, the input of the causal attention module is the environmental feature query vector , the behavioral feature key vector and value vector , and the calculation formula is:
[0009] Among them, comes from the linear projection of the environmental emotion dimension , and come from the linear projection of the behavioral feature ; is the causal weight matrix, and the element represents the causal strength of the environmental feature on the behavioral feature ; is the key vector dimension, Indicates element-wise multiplication.
[0010] In some embodiments of the present invention, in step S4, the dynamic causal weight update unit realizes the adaptive adjustment of causal weights through a gating mechanism:
[0011] Wherein, is the fused feature vector, obtained by concatenating the environmental sentiment dimension and the behavioral features and then encoding through a Transformer; is the real-time environmental sentiment vector; is a learnable gating weight matrix, is the Sigmoid activation function; is the basic causal weight matrix, obtained by training with global data.
[0012] In some embodiments of the present invention, in step S4, during the process of outputting the evaluation result of the mental state dimension, the loss function considered for mental state evaluation is the causal mean square error:
[0013] Wherein, is the total number of samples, is the true value vector of the mental state of the th sample, is the model prediction value; is the set of directed edges in the causal graph, is the true causal strength, is the model-predicted causal strength; is the regularization parameter, used to balance the mental state prediction error and the causal weight fitting error.
[0014] In some embodiments of the present invention, in step S5, the dynamic calibration of the model causal weights based on the historical data of the object to be evaluated includes personalized calibration, and the personalized calibration is achieved in the following manner: For a new user, use the previous times ( is a natural number and ) evaluation results to fine-tune the causal weight matrix:
[0015] Wherein: is a learnable personalized bias matrix, with the same dimension as ; MLP is a multi-layer perceptron, used to map the historical evaluation results to a weight adjustment vector; the calibrated causal weight only acts on the evaluation process of the current user.
[0016] To achieve the above object, an embodiment of the second aspect of the present invention provides a psychological state evaluation system based on environmental perception, the system comprising: A multi-modal data acquisition module configured to collect environmental scene and user behavior images through an RGB camera, collect environmental voiceprints through a microphone, and synchronously collect user facial micro-expression feature points, speech prosody features, and head-neck-shoulder vibration frequencies; an environmental semantic parsing module configured to extract object categories, spatial layouts, and dynamic events using a semantic segmentation model, identify sound categories using a sound recognition model, and generate environmental emotion dimension values through a semantic-emotion mapping atlas; a causal fusion module configured to construct an environment-psychology causal graph, fuse environmental features and behavior features through a causal attention module to generate a multi-modal feature vector; a psychological state evaluation module configured to use a causal Transformer model including a dynamic causal weight update unit to output a psychological state dimension evaluation result; a feedback and optimization module configured to generate a psychological state explanation report and dynamically calibrate the model causal weight based on user historical data.
[0017] To achieve the above object, an embodiment of the third aspect of the present invention provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned psychological state evaluation method based on environmental perception is implemented.
[0018] Compared with the prior art, the beneficial effects of the present invention are: The psychological state evaluation method, system, and readable storage medium based on environmental perception according to the embodiments of the present invention break through the limitation of traditional psychological evaluation's neglect of environmental factors through three core innovations: quantitative modeling of environmental semantics, causal reasoning and dynamic attention mechanism, and personalized adaptive calibration, construct a "environment-behavior-psychology" trinity evaluation framework, and significantly improve the accuracy, robustness, and practicality of psychological state evaluation, providing a new technical paradigm for the field of intelligent mental health monitoring. Description of the Drawings
[0019] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them: Figure 1 is a flowchart of a psychological state evaluation method based on environmental perception in an embodiment of the present invention; Figure 2 is a heat map of environmental semantic-emotion mapping in an embodiment of the present invention; Figure 3 is a schematic diagram of the dynamic change curve of the environmental emotion dimension in an embodiment of the present invention; Figure 4 It is a comparison schematic diagram of personalized calibration in an embodiment of the present invention; Figure 5 It is a schematic structural diagram of a psychological state assessment system based on environmental perception in another embodiment of the present invention. Specific embodiments
[0020] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings below are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0021] The following refers to the accompanying drawings to describe a psychological state assessment method, system, and readable storage medium based on environmental perception according to an embodiment of the present invention.
[0022] Figure 1 It is a schematic flowchart of a psychological state assessment method based on environmental perception according to an embodiment of the present invention.
[0023] As Figure 1 shown, the psychological state assessment method based on environmental perception includes the following steps: S1. Collect the environmental scene and the behavior image of the object to be evaluated through an RGB camera, and use a semantic segmentation model (such as Mask R-CNN) to extract the object category, spatial layout, and dynamic events; Among them, the object category is such as "desk", "green plant"; the spatial layout is such as "open office", "closed meeting room"; the dynamic event is such as "multiple people talking", "object moving".
[0024] Collect the environmental soundprint through a microphone. The environmental soundprint is such as the sound of keyboard tapping, alarm sound, and use a sound recognition model (such as YAMNet) to identify the sound category; at the same time, synchronously collect the facial micro-expression feature points, speech prosody features, and head-neck-shoulder vibration frequencies of the object to be evaluated. The micro-expression feature points are 68 feature points such as the corners of the mouth rising, frowning, etc., the speech prosody features are such as speech rate, pitch, and the head-neck-shoulder vibration frequency is such as the slight tremor when nervous.
[0025] Since traditional psychological assessment only relies on user self-report or physiological indicators, this step first collects environmental stimulus data (such as scene layout, sound type) and user behavior data at the same time to construct a complete data set of "environment - person" interaction.
[0026] As an example, in a driving scenario, the stress level can be comprehensively evaluated by identifying "congested road conditions" (environment) and the driver's "frequent frowning" (behavior).
[0027] Moreover, this solution uses computer vision and speech recognition technologies to convert fuzzy environmental information into structured data, such as "enclosed space", "high-frequency noise", etc., laying a foundation for the subsequent quantitative analysis of psychological states.
[0028] S2. Generate semantic vectors for object / scene labels (such as "green plants", "conference table") through a semantic embedding model (such as BERTopic), and convert spatial layouts (such as "enclosed") and sound labels (such as "alarm sound") into sparse vectors (such as One-Hot encoding); Map the above vectors to a preset environmental emotion dimension through a multi-layer perceptron (such as MLP) to obtain stress values, relaxation levels, and social intensities. For example: Stress value like "conference table" → Stress +10, Relaxation level like "green plants" → Relaxation +15, Social intensity like "multiple people talking" → Social intensity +20.
[0029] Traditional methods cannot quantify the impact of the environment on psychology. This step uses a semantic-emotion mapping graph to convert abstract environmental features into computable emotion indicators; moreover, it iteratively optimizes the mapping relationship through an online knowledge graph algorithm. For example, when a user feedbacks that "green plants do not relieve stress", the weights are automatically adjusted to ensure that the model adapts to diverse scenarios.
[0030] As an example, "natural landscape" is automatically mapped to Relaxation +20, and "alarm sound" is mapped to Stress value +18, enabling environmental factors to directly participate in the calculation of psychological states.
[0031] S3. Construct an environment-psychology causal graph, defining the causal relationship nodes and edge weights of environmental emotion dimensions, behavioral characteristics, and psychological states. Traditional models only analyze data correlations, such as "noise and anxiety appear simultaneously". This step uses a causal graph to clarify the direct influence path, such as "noise → physiological arousal → anxiety", avoiding interference from spurious correlations. For example, if "music sound" and "relaxation" are only indirectly related through the mediation of a "pleasant scene", the model will weaken their direct weights to improve the reasoning accuracy.
[0032] Design a causal attention module to adjust the attention of environmental features to behavioral characteristics according to the causal attention module, that is, adjust their interaction weights, and fuse them to generate a multi-modal feature vector; for example: if the causal intensity between "high-frequency noise" and "voice tremor" is high, the model preferentially fuses these two types of features. The causal attention module automatically suppresses irrelevant environmental features, such as the impact of background bird chirping on the stress of the office scene, and focuses on feature combinations with high causal intensity, such as "enclosed space + rapid speech", improving the efficiency and accuracy of the model.
[0033] S4. Input the fused feature vectors into a causal Transformer model containing a dynamic causal weight update unit. The model has a built-in dynamic weight update unit that adjusts the causal weights in real time through a gating mechanism (Gate). For example, when "driving scenario + sudden braking sound" is detected, the weight of "stress value" on "micro-expression tension" is automatically enhanced, and the evaluation results of the psychological state dimension are output, such as an anxiety value of 75 points and a relaxation degree of 20 points.
[0034] Design a loss function to simultaneously constrain the prediction error and the causal weight fitting error to ensure that the model is both accurate and conforms to the true causal relationship.
[0035] S5. Generate a psychological state interpretation report by combining environmental semantics. For example, "The current psychological stress value is 68 points, because a closed space (stress + 8), 3 people talking (social intensity + 15), and voice frequency fluctuations (anxiety-related features) are detected. Through environmental semantic interpretation, such as 'the stress is increased due to the meeting scenario', the trust of the evaluated object is enhanced, which is convenient for the psychological counselor or the evaluated object to understand the evaluation basis.
[0036] In addition, the causal weights of the model need to be dynamically calibrated based on the historical data of the evaluated object to avoid misjudgment of individual differences by the general model.
[0037] As an example, it is applicable to the psychological monitoring of different drivers in the driving scenario.
[0038] Through the above steps, this solution constructs a complete technical chain of "environmental perception - causal modeling - dynamic evaluation - personalized feedback", significantly improving the accuracy, robustness, and interpretability of psychological state evaluation, and providing a new technical path for the field of intelligent mental health monitoring.
[0039] In some embodiments of the present invention, in step S1, the semantic segmentation model is Mask R-CNN, which is used for real-time semantic segmentation of the images collected by the RGB camera and outputs the set of object categories, spatial layout labels and the set of dynamic events .
[0040] Among them, the RGB camera uses a global shutter camera with ≥8 million pixels, supports a resolution of 1920×1080@30fps, and is equipped with automatic white balance and low-light enhancement functions to ensure image clarity in complex environments.
[0041] The processing process of the semantic segmentation model is as follows: Input: Input real-time RGB image frames into the semantic segmentation model Mask R-CNN (an instance segmentation model based on the Faster R-CNN architecture) (where H is the image height and W is the width); Then output: the set of object categories , where is a string, such as desk, green plant, and each element is attached with a confidence score , indicating the detection reliability; Spatial layout label Open, closed, semi-open, outdoor , output through the scene classification branch; Set of dynamic events , where is a description of dynamic behavior, such as "multiple people talking", "object moving", "door opening and closing", etc., attached with spatio-temporal coordinates (x, y are image coordinates and z is the timestamp).
[0042] It should be noted that the input image should be Resized to a fixed size of 640*480, normalized to the interval [-1, 1], and non-maximum suppression (NMS) is used to filter redundant detection boxes with a threshold set to 0.5.
[0043] During the acquisition of environmental voiceprints, the following configurations are used for acoustic data acquisition and sound recognition: Microphone array: Integrated 3-channel MEMS microphone, frequency response 20Hz - 20kHz, signal-to-noise ratio ≥65dB, supporting sound source localization; The sound recognition model is YAMNet (a CNN-based environmental sound classification model), which is used to classify environmental voiceprints into a predefined set of sound categories. Its processing flow is: Input: Real-time voiceprint signal (T is the number of sampling points and the sampling rate is 44.1kHz); Output: Predefined set of sound categories , where , attached with a class probability vector .
[0044] Preprocessing is also required during acoustic data acquisition. The preprocessing steps include: Voiceprint framing: Frame length 512ms, frame shift 256ms, adding Hamming window; Feature extraction: Calculating the Mel spectrogram and generating 40-dimensional MFCC features.
[0045] As an example, user behavior data also needs to be collected, including: Facial micro-expressions: Using the OpenFace algorithm to extract the coordinates of 68 facial landmark points (j = 1, …, 68), calculating the intensity of action units (AUs), such as AU1 (inner brow raise), AU12 (lip corner raise); Speech prosody features: Extract 23-dimensional features such as fundamental frequency (F0), short-time energy, speech rate, formant frequency, etc.; Head-neck-shoulder vibration: Collect three-axis vibration data through an IMU sensor (accelerometer + gyroscope). , with a sampling rate of 100 Hz, and extract the physiological tremor signal through a band-pass filter (2 - 20 Hz).
[0046] The following table shows the technical implementation and examples of some actions in the action unit (AU): AU Number Name Description of Facial Muscle Movements Association with Mental State Technical Function in This Solution AU1 Inner Brow Raising The medial part of the frontalis muscle contracts, and the eyebrows move upward and toward the middle Surprise, Attention, Confusion Extract the coordinates of the brow feature points through OpenFace, calculate the difference between the vertical displacement of the eyebrows and the baseline, and quantify it as an intensity value (0 - 5 levels) AU4 Frowning The corrugator supercilii muscle and the depressor supercilii muscle contract, and the eyebrows are pressed down and converge Stress, Anxiety, Thinking Calculate the change in the distance between the eyebrows and the angle of the brow corners, output a standardized intensity value, and associate it with the stress value dimension AU12 Lip Corners Pulled Upward The zygomatic major muscle contracts, and the lip corners are pulled upward to both sides Pleasure, Positive Emotion Extract the coordinates of the lip corner feature points, calculate the upward angle of the lip corners and the baseline angle, and input them into the causal attention module AU17 Chin Raising The masseter muscle contracts, and the chin is lifted upward Uncertainty, Thinking Monitor the movement amplitude of the chin for micro-expression analysis in social intensity assessment AU25 Lips Parted The orbicularis oris muscle relaxes, and the lips open Surprise, Expressing Emotions Detect the change in the distance between the upper and lower lips to assist in judging the impact of sudden environmental stimuli As an example, during the above environmental scene acquisition process, the traditional semantic segmentation model does not distinguish between the processing of dynamic objects, such as moving crowds and objects, and static scenes. However, dynamic objects often have a more significant immediate impact on the psychological state of the evaluated object. For example, a suddenly moving object is likely to cause a sudden change in attention and a psychological stress response. To accurately capture such dynamic environmental stimuli, a dynamic object priority mechanism is introduced to increase the weight of its corresponding semantic features through the dynamic object priority mechanism:
[0047] Among them, : The intersection over union of the dynamic object mask and the current frame, reflecting the coverage and detection accuracy of the dynamic object in the current frame; : The historical maximum intersection over union, serving as a reference benchmark to measure the saliency of the currently detected dynamic object; : The weight coefficient, determined through experiments, aiming to highlight the dynamic change characteristics of the dynamic object; When is close to , tends to , indicating that the current dynamic object detection is stable and significant, and the weight of its semantic features needs to be increased; conversely, if is small, tends to , to avoid over - attention caused by unstable detection.
[0048] Regarding the above dynamic object priority mechanism, two specific examples are presented here to verify the effect of the dynamic object priority mechanism.
[0049] Example 1: Evaluation of dynamic interference in an office scene Scene description: The evaluated object is working in an open - plan office, and a colleague walks quickly past their desk (dynamic object).
[0050] Traditional method processing: The semantic segmentation model treats the dynamic "colleague" and static scene elements such as the "desk" and "computer" equally, without highlighting the immediate impact of the dynamic object on the psychological state.
[0051] The final psychological state assessment may overlook this dynamic interference, resulting in inaccurate calculations of dimensions such as stress values.
[0052] The processing of this solution: 1. Data calculation: For dynamic objects (colleagues) in the current frame = 0.7, historical maximum = 0.8, substitute into the formula:
[0053] Significantly enhance the weight of the semantic features of this dynamic object.
[0054] 2. Evaluation effect: When generating the dynamic emotion vector of the environment, this dynamic interference is captured with emphasis. Combining the possible micro-expression changes (such as eye movement, frowning) of the object to be evaluated, the increase in stress value is calculated more accurately. Compared with traditional methods, this solution can reflect the impact of the dynamic environment on the psychological state in real time and accurately.
[0055] Example 2: Dynamic assessment of the crowd in a shopping mall scenario Scenario description: The object to be evaluated is in the promotional activity area of a shopping mall, surrounded by a surging crowd (dynamic objects).
[0056] Traditional method processing: The semantic segmentation model does not assign special weights to the dynamic element of the "surging crowd", making it difficult to accurately quantify its stimulation to the psychological state (such as excitement or irritability).
[0057] The processing of this solution: 1. Data calculation: Assume the current frame = 0.6, = 0.8, then
[0058] The semantic feature weight of the dynamic crowd is enhanced.
[0059] 2. Evaluation effect: Combining the emotional polarity of the environmental sound (such as the promotional broadcast sound, assumed to be positive), a more comprehensive dynamic emotion vector of the environment is generated. When evaluating the psychological state, it can more accurately reflect the emotional fluctuations (such as increased excitement) of the object to be evaluated due to the crowd dynamics and promotional atmosphere, while traditional methods are prone to evaluation deviations due to ignoring the weights of dynamic objects.
[0060] As can be seen from the above examples, the dynamic object priority mechanism quantifies the detection salience of dynamic objects ( and relationship), dynamically adjusts its semantic feature weight , making the environmental perception more focused on the dynamic elements that have a significant immediate impact on the psychological state. Compared with traditional methods, it can capture dynamic environmental stimuli more accurately, improve the accuracy and timeliness of psychological state assessment, and effectively verify the technical effect of this solution.
[0061] As an example, the ambient voiceprint is collected through a microphone above and the sound category is recognized, such as "alarm sound", "conversation sound", etc. However, only the category information is not sufficient to comprehensively reflect the impact of the environment on the mental state. For example, the "alarm sound" has a negative emotional polarity and is likely to trigger a tense emotion; the "bird song" may have a positive emotional polarity and bring a sense of relaxation.
[0062] Therefore, the output of the sound recognition model YAMNet can be extended to include the emotional polarity (positive / negative / neutral) of the ambient sound, which is fused with the semantic segmentation result to generate an ambient dynamic emotion vector, and more accurately mapped to the ambient emotion dimension. By supplementing the emotional attributes of the sound, the emotional modeling of the ambient semantics is deepened, so that the ambient dynamic emotion vector can more comprehensively characterize the potential impact of the environment on the mental state, providing richer and more accurate ambient emotion features for the causal fusion and mental state assessment in the above step S3.
[0063] Such as Figure 2 shows a heat map of the ambient semantics-emotion mapping. Among them, the horizontal axis represents: ambient labels, including objects (such as "green plants"), spatial layouts (such as "enclosed space"), sounds (such as "alarm sound"), and dynamic events (such as "multiple people talking"); the vertical axis represents: ambient emotion dimension, corresponding to the stress value, relaxation degree, and social intensity in this solution.
[0064] Figure 2 In it, a positive value represents a positive impact (such as "green plants → relaxation degree +15"), and a negative value represents a negative impact (such as "enclosed space → relaxation degree -3"); by visualizing the quantization logic of the ambient semantics-emotion mapping heat map, it is proved that the ambient features can be converted into computable emotion indicators.
[0065] In some embodiments of the present invention, to accurately achieve the quantization mapping from ambient semantics to the emotion dimension and solve the problems of insufficient consideration of the comprehensive effects of environmental factors and poor dynamic adaptability in traditional methods, the quantization process of the ambient emotion dimension is specifically refined here, as shown in the following technical means: In step S2, the ambient emotion dimension includes a stress value 、a relaxation degree and a social intensity ,and the quantization mapping from object, space, and sound labels to the emotion dimension can be achieved through a predefined semantics-emotion mapping atlas, and the mapping atlas is iteratively optimized through an online knowledge graph update algorithm.
[0066] However, traditional methods do not fully consider the combined effects of various environmental factors and their dynamic characteristics. For example, the stress value may be affected by the combined action of an enclosed space, high-frequency noise, and social density, but traditional methods do not systematically quantify and integrate these factors; the relaxation level also does not clarify the synergistic effects of natural elements and specific sounds. Therefore, this solution achieves precise mapping through specific formulas and echoes the "online knowledge graph update algorithm iterative optimization mapping graph" mentioned above to ensure dynamic adaptability.
[0067] As an example, the quantification formulas and explanations for each environmental emotion dimension are as follows: 1. Stress value:
[0068] Among them, : Proportion of the enclosed space, obtained by analyzing the static scene mask through a semantic segmentation model, reflecting the potential impact of the space enclosure degree on stress; : High-frequency noise intensity, obtained by combining an acoustic data collection with a sound recognition model to quantify the psychological stimulation of noise; : Social density, calculated based on detected dynamic objects (such as crowds) and semantic segmentation results, reflecting the intensity of social interaction; : Environmental factor weight, determined through experiments to balance the contributions of various factors to the stress value.
[0069] This formula is associated with the dynamic object priority mechanism in this solution. If a dynamic object (such as a fast-moving crowd) is detected, its corresponding social density will have its weight increased through the priority mechanism, and thus more significantly affect , ensuring that dynamic environmental stimuli are accurately captured.
[0070] 2. Relaxation level:
[0071] : Proportion of natural elements (such as green plants, water features), obtained by analyzing the static scene mask through a semantic segmentation model, reflecting the contribution of natural elements to the sense of relaxation; : Low-frequency white noise intensity, obtained through a sound recognition model and acoustic data processing. Such sounds usually have a soothing effect; : Environmental factor weight, highlighting the dominant role of natural elements in the relaxation level.
[0072] 3. Social intensity:
[0073] : The number of detected human faces, obtained through a vision module (such as an RGB camera combined with a face recognition algorithm), reflecting the number of social participants; : The frequency of voice interaction, obtained from a microphone array and a voice processing module, reflecting the frequency of social interaction; : The social factor coefficient, which adjusts the calculation of social intensity.
[0074] The above three formulas are the specific implementation of "realizing quantitative mapping through a predefined semantic-emotional mapping atlas", and the mapping atlas is iteratively optimized through an online knowledge graph update algorithm.
[0075] Such as Figure 3 shows a dynamic change curve graph of the environmental emotion dimension. Among them, the horizontal axis represents: time series, simulating environmental changes under different scenarios (such as "open office → green plants introduced → alarm sound → multiple people talking → enclosed space"). The vertical axis represents: emotion dimension value, reflecting the real-time impact of environmental stimuli.
[0076] Figure 3 In, the stress value (red curve) in the initial state (open space + keyboard sound) is 10 - 6 = 4; it suddenly rises by 18 when the alarm sound rings; it increases by another 12 when switching to an enclosed space (the direct impact of the space layout).
[0077] The relaxation level (green dashed line): it increases by 15 when green plants are introduced (the positive emotion of the object label); it decreases by 10 due to the alarm sound (negative sound emotion).
[0078] The social intensity (blue dotted line): it increases by 15 × 1.3 = 19.5 when multiple people are talking (in this solution, the dynamic object priority mechanism, with the weight increased by 30%).
[0079] Figure 3 Intuitively demonstrates the mechanism of "dynamic environmental stimuli → real-time update of emotion dimension", verifies the dynamic adaptability of the semantic-emotional mapping atlas, and also echoes the online update algorithm in step S2.
[0080] And through event annotation, it also highlights the immediate impact of dynamic objects (such as "multiple people talking") and sudden sounds (such as "alarm sound") on the mental state, proving the limitations of traditional methods in ignoring environmental factors.
[0081] As an example, if it is found through the dynamic object priority mechanism that a certain type of dynamic object (such as a suddenly emerging crowd) has a significant impact on the stress value it can be adjusted through the update algorithm the weight or The calculation method ensures that the mapping atlas adapts to diverse scenarios, which is consistent with the above-mentioned direction of "environmental semantic depth modeling and emotional association", and improves the accuracy and dynamic adaptability of mental state assessment.
[0082] In some embodiments of the present invention, in the above step S3, the core purpose of constructing the environment-psychology causal graph is to solve the limitations of correlation analysis in traditional psychological assessment techniques. In existing methods, such as application document CN111477328B and application document CN117936032A, only statistical correlations are used to associate the environment with the mental state, and it is impossible to distinguish causal relationships from spurious correlations. For example, "music sound" and "relaxation" may be indirectly related through the mediation of "pleasant scenario" rather than directly causally related, resulting in poor model interpretability and susceptibility to interference.
[0083] Therefore, it can be obtained that traditional methods cannot clarify the direct causal strength of environmental factors (such as "enclosed space") on mental states (such as "anxiety"), and can only judge the association through co-occurrence frequencies, lacking scientific basis; and when environmental factors (such as moving objects, sudden sounds) change, traditional models cannot adjust the feature interaction weights in real time, resulting in lagging or inaccurate assessments.
[0084] Here, a solution is proposed: By introducing the causal inference theory (Do-Calculus), explicitly model the direct causal relationships between the environmental emotional dimension (such as stress value ), behavioral characteristics (such as micro-expression frequency) and mental state (such as anxiety value), and quantify the causal effect as the edge weight to ensure that the model focuses on the real influence path and improves the interpretability and dynamic adaptability of the assessment.
[0085] And the specific steps of causal graph modeling include: 1. Definition of node set Environmental emotional dimension (U): includes stress value , relaxation degree , social intensity ; Behavioral characteristics (B): include the intensity of facial action units (such as the frequency of the corners of the mouth turning up), the fluctuation of the fundamental frequency of speech , the vibration frequency of the head, neck and shoulders ; Mental state (Y): the target assessment dimension, such as anxiety value, pleasure value, aggression index, which is output by the causal Transformer model in the above step S4.
[0086] 2. Edge weight calculation: Causal effect based on Do-Calculus Formula:
[0087] Among them, (environmental emotional dimension or behavioral feature node), (psychological state node); : Intervention variable When, And The covariance of, which represents the strength of causal dependence; : Variance of variable For normalizing the causal effect.
[0088] Physical meaning: Measure when Is actively intervened (such as artificially increasing the duration of the enclosed space), The degree of change of, excluding the interference of other variables, directly reflecting The causal strength of.
[0089] 3. Relevance with the dynamic object priority mechanism In step S1, the dynamic object priority mechanism enhances the semantic weight of dynamic environmental features (such as fast-moving crowds) through , and the causal graph modeling further incorporates this weight into the causal effect calculation. The causal graph modeling not only solves the defects of traditional correlation analysis, but also constructs a scientific and interpretable psychological state evaluation framework through quantifying causal strength, dynamic weight adjustment, and personalized adaptation.
[0090] For example: The social density of dynamic objects As A component of, its causal effect Will increase due to the increase of , enabling the model to prioritize the feature interactions with high causal strength in dynamic scenarios.
[0091] In some embodiments of the present invention, the design of the causal attention module aims at the limitations of the traditional attention mechanism in psychological state evaluation: Since the traditional multi-head attention mechanism only calculates the feature interaction weights through data correlation, it cannot distinguish the causal relationship between environmental features and behavioral features. For example, "office green plants" and "relaxation level" may be wrongly associated due to scene correlation, while the actual causal path may act indirectly through "cognition of natural elements → relaxation", and the traditional method cannot explicitly model this causal logic.
[0092] Moreover, when the environment changes dynamically (such as sudden noise, crowd movement), the traditional mechanism cannot adjust the feature weights in real time, resulting in a lag in the model's response to immediate stimuli. For example, the strong causal relationship between the sudden braking sound and the driver's micro-expression in a driving scenario needs to be captured in real time, and the traditional method may ignore such key interactions due to fixed weights.
[0093] Therefore, the core of solving the above problems lies in how to introduce causal constraints in multi-modal feature fusion, enabling the model to focus on the true causal path and enhancing the evaluation accuracy and interpretability in a dynamic environment.
[0094] In step S3, the input of the causal attention module is the environmental feature query vector , the behavioral feature key vector and the value vector , and the calculation formula is:
[0095] where, is the linear projection from the environmental emotion dimension , and are the linear projections from the behavioral feature ; is the causal weight matrix, and the element represents the causal strength of the environmental feature on the behavioral feature , which is obtained through Do-Calculus; is the key vector dimension, represents element-wise multiplication.
[0096] In some embodiments of the present invention, the design of the dynamic causal weight update unit stems from the defect that the causal weight in the traditional psychological assessment model is fixed.
[0097] The traditional method assumes that the causal relationship between environmental features and psychological states is static and cannot adapt to a dynamically changing environment. For example, in a driving scenario, the impact weight of a sudden alarm sound on anxiety should be significantly higher than that of the same sound in a quiet office, but the traditional model cannot adjust in real time. It is also manifested in that when the environmental semantics change (such as from "quiet" to "noisy"), the fixed causal weight cannot reflect the true impact strength between features in the current environment; and in different scenarios (such as home, workplace), the impact of the same environmental feature (such as "human voice") on the psychological state is different, and the traditional model lacks a weight adjustment mechanism for scene perception.
[0098] Therefore, the core of solving the problem lies in real-time perceiving environmental changes through a gating mechanism, dynamically adjusting the causal weight matrix, enabling the model to adapt to dynamic environments and individual scene differences, and enhancing the accuracy of psychological state assessment.
[0099] In step S4, the dynamic causal weight update unit realizes the adaptive adjustment of causal weights through a gating mechanism:
[0100]
[0101] Among them, : The fused feature vector is obtained by concatenating the environmental sentiment dimensions ( , , ) and behavioral features (such as micro-expressions, speech prosody), and then encoded by Transformer, containing the deep fusion information of the environment and behavior; : The real-time environmental sentiment vector is generated by semantic-sentiment mapping, reflecting the stress value, relaxation degree and social intensity of the current environment; : The learnable gating weight matrix learns the influence of the joint representation of the environment and the fused features on the gating signal through training; : The Sigmoid activation function compresses the gating signal into the interval [0, 1], and outputs Gate ∈ [0, 1], representing the update degree of the dynamic causal weight in the current environment; : The basic causal weight matrix is obtained by training with global data, reflecting the causal relationship between environmental features and mental states in general cases.
[0102] The mechanism of the above gating signal calculation and weight update is explained as follows: Gating signal calculation: Gate = ( ), After concatenating the fused feature H and the real-time environmental sentiment , input them into the gating network, and output a gating signal between 0 and 1. For example, when a sudden alarm sound (S_{stress} surges) is detected in a driving scenario, the gating signal Gate approaches 1, indicating that the causal weight needs to be updated significantly.
[0103] Weight update: , when Gate is close to 1, the dynamically adjusted weight plays a dominant role, and the model focuses on the causal relationship in the current environment; when Gate is close to 0, the basic weight is retained to ensure the evaluation consistency of the model in a stable environment.
[0104] Through the above steps, the dynamic causal weight update unit solves the problem that the traditional model has poor adaptability to dynamic environments, and adjusts the causal weight in real time through the gating mechanism, enabling the model to accurately capture the key causal relationships in the current environment.
[0105] In some embodiments of the present invention, the loss function of the traditional mental state evaluation model (such as mean square error ) only focuses on the predicted value and the true value The gap ignores the causal relationship constraints between environmental features and mental states. This causes the model to potentially learn superficial statistical correlations (such as the association between "office green plants" and "relaxation" due to co-occurrence in a scene), rather than the true causal path (such as "cognition of natural elements → neural relaxation → psychological pleasure"), resulting in poor model interpretability and insufficient generalization ability.
[0106] Therefore, the core of solving the problem lies in introducing the causal mean squared error loss function, while constraining the mental state prediction error and the causal weight fitting error, to ensure that the model not only accurately predicts the mental state but also learns the feature interaction relationships that conform to the causal logic of the real world.
[0107] In step S4, during the process of outputting the mental state dimension evaluation result, the loss function considered for mental state evaluation is the causal mean squared error:
[0108] where, : The total number of samples, which is the scale of the dataset covering different environmental scenarios and user states; : The true value vector of the mental state of the -th sample, for example , obtained through professional psychological assessment tools or prior knowledge; : The predicted value of the mental state of the -th sample by the model, output by the causal Transformer model; : The set of directed edges in the causal graph, for example , reflecting the causal relationship between the environmental emotion dimension and the mental state; : The true causal strength, determined by Do - Calculus or prior knowledge, for example , indicating the direct influence strength of the stress value on anxiety; : The causal strength predicted by the model, learned by the causal Transformer model during training; : The regularization parameter, used to balance the mental state prediction error and the causal weight fitting error. When is large, the model pays more attention to the accuracy of the causal relationship; when is small, it focuses on the accuracy of mental state prediction.
[0109] In the above steps, the mental state prediction error term: Ensures that the model output is close to the true mental state , is the retained term of the traditional loss function; And the causal weight fitting error term: Constrains the causal strength predicted by the model Approximates the true value . For example, if in the real scenario the causal strength of "high-frequency noise anxiety" , during model training, the parameters will be adjusted to make approach 0.8, to prevent the model from wrongly associating "high-frequency noise" with "relaxation".
[0110] It should be noted that the and in the loss function are directly related to the environment-psychology causal graph constructed in step S3, strengthening the learning of the causal relationships defined in the causal graph. Moreover, the dynamic causal weight update unit in step S4 adjusts through a gating mechanism, and the causal weight fitting error term in the loss function provides an optimization goal for this adjustment, making not only adapt to the dynamic environment (through Gate), but also conform to the real causal relationships, through .
[0111] Through the above steps, the causal mean square error loss function solves the problems of the lack of causal relationship modeling and insufficient robustness in traditional models, ensuring a balance between the prediction accuracy of the mental state and the consistency of causal logic of the model.
[0112] In some embodiments of the present invention, since the traditional mental state assessment model uses a unified causal weight matrix (such as ), it ignores the differences in individuals' responses to environmental stimuli. For example, in the environment-psychology causal graph constructed in step S3, the causal strength of "social density anxiety" is a general value trained based on global data, but in reality, introverted users are more sensitive to social density, while extroverted users are not.
[0113] This "one-size-fits-all" approach leads to a decrease in the assessment accuracy of the model when facing new users. Therefore, this solution proposes to dynamically fine-tune the causal weight matrix by using the first ( ) assessment results of new users, so that the model adapts to individual differences and improves the personalized assessment accuracy.
[0114] In step S5, dynamically calibrating the causal weights of the model based on the historical data of the object to be evaluated includes personalized calibration, and the personalized calibration is achieved in the following way: For new users, use the first times ( is a natural number and ) assessment results Fine-tune the causal weight matrix (the calibration effect is best when t = 3):
[0115] Where: : The causally weighted matrix after personalized calibration, which only acts on the evaluation process of the current user. For example, when calculating the causal strength of "enclosed space → anxiety", use to replace the global ; : The basic causal weight matrix, which is obtained by training with global data and reflects the causal relationship between environmental characteristics and mental states in general cases, such as the general causal strength calculated by Do-Calculus in step S3; : A learnable personalized bias matrix, with the same dimension as , used to capture the differences between individuals and the global. For example, if a user is more sensitive to "high-frequency noise", the corresponding item in will increase the weight of "high-frequency noise → anxiety"; MLP: Multilayer perceptron, which maps the historical evaluation results to a weight adjustment vector. For example, if the user's anxiety value is relatively high in social scenarios in the previous t evaluations, the output vector of the MLP will increase the weight of "social intensity →anxiety".
[0116] : The vector of the psychological state evaluation results of the new user in the previous t times, output by the causal Transformer model in step S4.
[0117] It should be noted that the historical evaluation result processing is to input the previous t evaluation results into the MLP, and the MLP learns the mapping relationship between the evaluation results and the weight adjustment through non-linear transformation (such as the ReLU activation function). For example, if the user's anxiety value is relatively high in the "multi-person conversation" scenario in the previous two times, the output vector of the MLP will indicate an increase in the weight of "social intensity →anxiety"; And the weight fine-tuning refers to calculating the personalized bias through and superimposing it on the basic weight to obtain . For example, the weight of "social intensity → anxiety" in is 0.3, and the personalized bias adjusts it to 0.5, which is more in line with the user's sensitivity to social interactions.
[0118] It should also be noted that the psychological state interpretation report generated by combining environmental semantics in step S5 provides feedback to the user, and this feedback can be indirectly used to adjust the evaluation results after times, forming a closed loop of "evaluation - calibration - re - evaluation" to continuously optimize the personalized experience, which is closely combined with "real - time feedback and model optimization" mentioned above. And through the personalized calibration scheme, the problems of ignoring individual differences and poor cold - start adaptation in traditional models are solved. The causal weights are dynamically fine - tuned based on historical evaluation results to make the model better adapt to new users.
[0119] After being processed by this step, the Under the constraint of the causal mean - squared error loss function in step S4, not only the prediction accuracy of the psychological state is optimized, but also the personalized causal weights are close to the user's real causal relationship, avoiding logical biases caused by excessive personalization.
[0120] For example, Figure 4 shows a comparison experimental graph of personalized calibration. Among them, the upper graph: the curve display under the same environmental stimulus (such as continuous noise); the middle graph: the display of the uncalibrated pressure value evaluation results (user A is sensitive and user B is tolerant); the lower graph: the display of the pressure value evaluation results after personalized calibration. Figure 4 It verifies the improvement of the personalized calibration mechanism on the evaluation accuracy.
[0121] For example, Figure 5 As shown, a psychological state evaluation system based on environmental perception includes: A multi - modal data acquisition module, specifically including: Environmental visual data acquisition: Use a high - resolution RGB camera (≥8 million pixels, global shutter, supporting 1920×1080@30fps), equipped with automatic white balance and low - light enhancement functions, to capture environmental scenes (such as desks, green plants) and user behavior images (such as facial expressions, body movements) in real time.
[0122] Environmental voiceprint acquisition: Integrate a 3 - channel MEMS microphone array (frequency response 20Hz - 20kHz, signal - to - noise ratio ≥65dB), support sound source localization and directional noise reduction, and capture environmental voiceprints (such as keyboard typing sounds, alarm sounds).
[0123] User behavior data synchronous acquisition: Facial micro - expressions: Extract the coordinates of 68 facial feature points through the OpenFace algorithm, and calculate the intensity of action units (AU) (such as AU12 the corners of the mouth turn up, AU4 frowns).
[0124] Voice prosody: Extract 23 - dimensional acoustic features such as fundamental frequency (F0), speech rate, and short - term energy, with a sampling rate of 44.1kHz and a frame length of 512ms.
[0125] Head, neck and shoulder vibration: Using a six-axis IMU sensor (accelerometer + gyroscope, sampling rate 100Hz), physiological tremor signals are extracted through band-pass filtering (2 - 20Hz) to quantify the state of tension or relaxation. Environmental semantic parsing module, specifically including: Semantic feature extraction: Visual semantic parsing: Mask R-CNN performs instance segmentation on images, outputs structured labels (object categories, spatial layouts, dynamic events), and enhances the semantic weights of dynamic elements (such as moving crowds) through a dynamic object priority mechanism (based on the intersection over union to calculate weights). ).
[0126] Voiceprint classification: YAMNet classifies voiceprint signals into predefined sound categories (such as "natural sounds") and extends the output to include emotional polarity (positive / negative / neutral). For example, "bird song" is mapped to a positive emotion.
[0127] Semantic-emotion mapping: Environmental labels are mapped to the emotional dimension (stress value , relaxation level , social intensity ) through a predefined atlas (such as "green plants relaxation level +15"), and dynamic optimization is achieved by combining an online knowledge graph update algorithm (such as automatically adjusting weights when the user feedbacks that "green plants do not relieve stress").
[0128] Causal fusion module, specifically including: Construction of environmental-psychological causal graph: Node definition: Environmental emotional dimension, behavioral characteristics, and psychological state.
[0129] Calculation of causal intensity: Based on Do Calculus, edge weights are calculated , clarifying the direct causal path of "noise → physiological arousal → anxiety" to avoid interference from spurious correlations.
[0130] Causal attention mechanism: Input the environmental feature query vector Q and the behavioral feature key-value pair K / V, and adjust the interaction weights through the formula , where is the causal weight matrix, which preferentially fuses features with high causal intensity (such as "high-frequency noise" and "voice tremor").
[0131] Psychological state assessment module, specifically including: Dynamic causal Transformer model: Encoding layer: Input the fused feature vector (environmental emotion + behavioral features) into the Transformer encoder to generate a context-aware hidden representation H.
[0132] Dynamic weight update unit: Through a gating mechanism Real-time adjustment of the causal weight matrix W_c’. For example, when a sudden braking sound is detected in a driving scenario, enhance the weight of "stress value → anxiety".
[0133] Loss function design: Adopt causal mean squared error to constrain the prediction error and the causal weight fitting error, ensuring that the model conforms to the true causal logic.
[0134] Feedback and optimization module, specifically including:[[]] Explanation report generation: Generate a readable report in combination with environmental semantics (such as "detecting a closed space + 3 people talking"), for example, "The current stress value is 68 points, mainly attributed to the closed environment (+8) and social density (+15)".
[0135] Personalized calibration: For new users, utilize the results of the previous t evaluations to fine-tune the causal weights through a formula. For example, the weight of "social density → anxiety" for introverted users is increased by 20%.
[0136] The interaction process between the above modules includes:[[]] Data collection: Multimodal sensors synchronously capture environmental and user behavior data; Semantic parsing: Extract structured environmental labels and map them to the emotional dimension; Causal fusion: Construct a causal graph and fuse features through the attention mechanism; State evaluation: The dynamic Transformer model outputs the mental state and adjusts the weights in real time; Feedback optimization: Generate an explanation report, calibrate the model in combination with historical data, and improve the personalized accuracy.
[0137] The system models through the trinity of "environment - behavior - psychology", breaking through the neglect of environmental factors in traditional methods. For example, in a driving scenario, it can accurately identify the stress correlation of "congested road conditions + frequent frowning"; and the dynamic weight update and personalized calibration mechanism (such as formula ) significantly improves the robustness and practicality of the model in diverse scenarios (home, office, cockpit).
[0138] Corresponding to the above embodiments, the present invention also proposes a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for evaluating mental state based on environmental perception is implemented.
[0139] The readable storage medium of the present invention can be any form of storage medium readable by a processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned method for evaluating mental state based on environmental perception can be implemented.
[0140] The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and practice. For example, according to legislation and practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0141] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or equipment (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, device or equipment and execute the instructions), or in combination with these instruction execution systems, devices or equipment. For the purposes of this specification, a "readable storage medium" can be any device that can contain, store, evaluate mental state, propagate or transmit a program for use by an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber device, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting or otherwise processing it as appropriate, and then storing it in a computer memory.
[0142] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0143] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0144] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0145] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
[0146] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for evaluating mental state based on environmental perception, characterized in that, It includes the following steps: S1. Collect the environmental scene and the behavior images of the object to be evaluated through an RGB camera, and use a semantic segmentation model to extract object categories, spatial layouts, and dynamic events; collect the environmental voiceprint through a microphone, and use a voice recognition model to identify the voice category; Synchronously collect the facial micro-expression feature points, speech prosody features, and head-neck-shoulder vibration frequencies of the object to be evaluated; S2. Generate semantic vectors for object / scene labels through a semantic embedding model, convert spatial layouts and voice labels into sparse vectors, and map them to preset environmental emotion dimensions through a multi-layer perceptron to obtain stress values, relaxation levels, and social intensities; S3. Construct an environment-psychology causal graph, define the causal relationship nodes and edge weights of environmental emotion dimensions, behavior characteristics, and mental states; adjust the attention of environmental features to behavior characteristics through a causal attention module, and fuse to generate a multi-modal feature vector; S4. Input the fused feature vector into a causal Transformer model containing a dynamic causal weight update unit, and output the evaluation result of the mental state dimension; S5. Combine the environmental semantics to generate a mental state explanation report, and dynamically calibrate the causal weights of the model based on the historical data of the object to be evaluated.
2. The evaluation method according to claim 1, wherein In step S1, the semantic segmentation model is MaskR-CNN, which is used to perform real-time semantic segmentation on the images collected by the RGB camera and output the set of object categories, spatial layout labels, and the set of dynamic events ; The voice recognition model is YAMNet, which is used to classify environmental voiceprints into a predefined set of voice categories .
3. The evaluation method according to claim 1, characterized in that In step S2, the environmental emotion dimension includes a stress value , relaxation level and social intensity . The quantitative mapping from object, space, and sound labels to the emotion dimension is realized through a predefined semantic-emotion mapping atlas, and the mapping atlas is iteratively optimized by an online knowledge graph update algorithm.
4. The evaluation method according to claim 1, wherein In step S3, the causal graph modeling step includes: Defining the node set as environmental emotion dimensions, behavior characteristics, and mental states; Calculating Causal Effects through Do-Calculus As an edge weight, where = represents the direct causal strength of node on node .
5. The evaluation method according to claim 1, characterized in that In step S3, the input of the causal attention module is the environmental feature query vector , the behavioral feature key vector and the value vector , and the calculation formula is: Among them, the linear projection from the environmental emotional dimension and the linear projection from the behavioral characteristics; is the causal weight matrix, and the element represents the causal intensity of the environmental feature on the behavioral characteristic ; is the key vector dimension, indicating element-wise multiplication. ; represents element-wise multiplication.
6. The evaluation method according to claim 3, characterized in that In step S4, the dynamic causal weight update unit realizes the adaptive adjustment of causal weights through a gating mechanism: Among them, is the fused feature vector, which is obtained by concatenating the environmental sentiment dimension and the behavior features and then encoding through Transformer; is the real-time environmental sentiment vector; is the learnable gating weight matrix, is the Sigmoid activation function; is the basic causal weight matrix, which is obtained by training with global data.
7. The evaluation method according to claim 1, characterized in that In step S4, during the process of outputting the evaluation result of the mental state dimension, the loss function considered for mental state evaluation is the causal mean square error: Among them, is the total number of samples, is the true value vector of the mental state of the th sample, is the set of directed edges in the causal graph, is the true causal strength, is the causal strength predicted by the model; is the regularization parameter, which is used to balance the prediction error of the mental state and the fitting error of the causal weight.
8. The evaluation method according to claim 1, wherein In step S5, the dynamic calibration of the causal weights of the model based on the historical data of the object to be evaluated includes personalized calibration, and the personalized calibration is achieved through the following methods: For new users, before utilization the results of the previous evaluation are used to finely tune the causal weight matrix, where is a natural number and ; Wherein: is a learnable personalized deviation matrix, with the dimension being the same as that of ; MLP is a multi-layer perceptron used to map historical evaluation results into a weight adjustment vector; the calibrated causal weight only acts on the evaluation process of the current user.
9. A psychological state assessment system based on environmental perception, characterized in that, It includes: A multi-modal data acquisition module configured to collect the environmental scene and user behavior images through an RGB camera, collect the environmental voiceprint through a microphone, and synchronously collect the facial micro-expression feature points, speech prosody features, and head-neck-shoulder vibration frequencies of the user; An environmental semantic parsing module configured to extract object categories, spatial layouts, and dynamic events using a semantic segmentation model, identify voice categories using a voice recognition model, and generate environmental emotion dimension values through a semantic-emotion mapping atlas; A causal fusion module configured to construct an environment-psychology causal graph, and fuse environmental features and behavior characteristics through a causal attention module to generate a multi-modal feature vector; A mental state evaluation module configured to use a causal Transformer model containing a dynamic causal weight update unit to output the evaluation result of the mental state dimension; A feedback and optimization module configured to generate a mental state explanation report and dynamically calibrate the causal weights of the model based on the user's historical data.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the environmental perception-based mental state evaluation method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-mode based emotion recognition method
CN108805089A
Psychological disorder evaluation method and system for psychiatric patient
CN117438048A
Environment simulation equipment based on VR and psychology
CN118629285A
Intelligent evaluation method for cognitive function detection
CN119150117A
AI-assisted limb rehabilitation system
CN119580933A
Cited By
Artificial intelligence psychological assessment method and device based on multiple modes
CN120600318A
Model training method and device, environment quality evaluation method and device and storage medium
CN121524602A