Intelligent dynamic assessment method for mental state of teenagers based on multi-modal data

By combining multimodal data acquisition with feature fusion and dynamic state extrapolation of machine learning models, the limitations of single-modal data assessment are overcome, enabling comprehensive and real-time assessment of psychological states and adapting to assessment needs in different scenarios.

CN121196543BActive Publication Date: 2026-08-25LUDONG UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511276737.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-08-25
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing methods for assessing psychological state mainly rely on single-modal data, which cannot fully reflect the complexity and dynamic changes of psychological state. This results in poor stability and reliability of assessment results, making it difficult to meet the needs of real-time and accurate assessment.

Method used

Multimodal data acquisition is employed, including voice emotion information and facial expression behavior relationships. Feature fusion and temporal correlation analysis are performed through machine learning models to generate feature fusion results containing emotional dimension features and cognitive pattern features. Dynamic state inference is then performed to output psychological state assessment elements.

Benefits of technology

It enables comprehensive and real-time assessment of psychological states, captures subtle changes, adapts to assessment needs in different scenarios, and provides more valuable assessment information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121196543B_ABST
    Figure CN121196543B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent evaluation of psychological state, and discloses a method for intelligent dynamic evaluation of psychological state of teenagers based on multi-modal data. The method acquires multi-modal psychological data of a target evaluation object, the multi-modal psychological data containing voice emotional information and expression behavior relationship, and extracts dynamic features therefrom; a trained machine learning model is used to perform feature fusion processing on the above psychological data, generating a multi-modal feature set containing a mapping relationship between modal type identification and time sequence coding; based on the multi-modal feature set, a state analysis network is used to perform time sequence correlation analysis on the dynamic features, obtaining a fusion result containing emotional dimension features and cognitive mode features; the fusion result is input into an evaluation model for dynamic state deduction processing, and psychological state evaluation elements are output. The present application can realize comprehensive and dynamic evaluation of psychological state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent assessment technology for psychological states, specifically to an intelligent dynamic assessment method for the psychological states of adolescents based on multimodal data. Background Technology

[0002] Psychological state assessment has urgent applications in many fields of modern society, such as clinical psychotherapy, campus mental health management, and psychological monitoring of specific occupational groups. For a long time, traditional psychological state assessment methods have relied primarily on the subjective observation of professionals and various psychological scales completed by the assessment subjects. This approach is overly dependent on the assessor's personal experience and subjective judgment; different assessors may have significantly different judgments of the same subject, leading to poor stability of the assessment results. Furthermore, when completing the scales, subjects may conceal their true situation due to social expectation effects or their own cognitive biases, further affecting the reliability of the assessment results.

[0003] With the development of digital technology, psychological state assessment methods based on single-modal data have emerged, such as inferring psychological state solely by analyzing voice or facial expression data. However, psychological state is a complex system formed by the interaction of multiple factors, including emotion, cognition, and behavior. Single-modal data can only reflect one aspect of psychological state and cannot depict its complete picture. For example, voice data can reflect the level of emotional excitement of the assessed individual, but it is difficult to reflect their deep-seated cognitive patterns; facial expression data can show immediate emotional reactions, but it cannot capture the underlying psychological activities contained in language.

[0004] Furthermore, most existing assessment methods are static, meaning they are based on a single assessment of data collected at a specific moment, ignoring the dynamic nature of psychological states. A person's psychological state constantly changes over time, due to environmental changes and experienced events. Static assessments only provide fragmented information at a particular moment, failing to reflect the trajectory and trend of these changes. In multimodal data processing, existing technologies often simply superimpose features from different sources without establishing the intrinsic connections between modalities. This prevents the full utilization of the collaborative information inherent in multimodal data, making it difficult to achieve the desired accuracy in assessment results.

[0005] In practical applications, these issues make it difficult for existing assessment methods to meet the needs for comprehensive, real-time, and accurate assessment of mental states. For example, in the daily monitoring of patients with depression, it is necessary to capture subtle changes in their mental state in a timely manner in order to adjust the treatment plan, which existing assessment methods often cannot do. In the monitoring of the mental state of workers at heights, it is necessary to grasp their emotional and cognitive states in real time, and the lag and one-sidedness of existing methods may pose safety hazards. Summary of the Invention

[0006] The purpose of this invention is to provide an intelligent dynamic assessment method for the psychological state of adolescents based on multimodal data, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, this invention provides an intelligent dynamic assessment method for the psychological state of adolescents based on multimodal data, the method comprising:

[0008] S1. Obtain multimodal psychological data of the target assessment object and extract dynamic features from the multimodal psychological data; wherein, the multimodal psychological data includes voice emotion information and facial expression behavior relationship;

[0009] S2. The trained machine learning model is used to perform feature fusion processing on the multimodal psychological data to generate a multimodal feature set; wherein, the multimodal feature set includes a mapping relationship between modality type identifiers and temporal codes;

[0010] S3. Based on the multimodal feature set, a state analysis network is used to perform temporal correlation analysis on the dynamic features to generate a feature fusion result that includes emotional dimension features and cognitive pattern features;

[0011] S4. Input the feature fusion result into the evaluation model for dynamic state deduction processing, and output psychological state evaluation elements; wherein, the psychological state evaluation elements include state labels and trend prediction parameters generated based on feature evolution sequences.

[0012] Optionally, step S2 uses a trained machine learning model to perform feature fusion processing on the multimodal psychological data to generate a multimodal feature set, including:

[0013] S21. Perform modality segmentation processing on the multimodal psychological data to generate multiple data domains and corresponding initial feature parameters;

[0014] S22. Extract the temporal feature sequence of each data domain, and perform attention weighting processing on the temporal feature sequence to generate a weighted multi-domain feature vector;

[0015] S23. Calculate the similarity between the multi-domain feature vector and the preset modality association threshold, and filter out candidate features that meet the confidence conditions.

[0016] S24. Redundancy elimination processing is performed on the candidate features to generate a multimodal feature set containing modality type identifiers. Each multimodal feature contains a normalized temporal code relative to the multimodal psychological data.

[0017] Specifically, based on the multimodal feature set, a state analysis network is used to perform temporal correlation analysis on the dynamic features, generating a feature fusion result that includes emotional dimension features and cognitive pattern features, including:

[0018] Based on the normalized temporal coding in the multimodal feature set, the multimodal psychological data is subjected to time window truncation processing to generate multiple feature time segments;

[0019] Each feature time segment is standardized and feature-enhanced to generate optimized state analysis input data.

[0020] The recurrent memory subnet in the state analysis network is used to extract sequence features from the state analysis input data to generate an emotion dimension feature vector.

[0021] In parallel, a pattern classification model is used to perform feature pattern matching on the feature time segments to generate a cognitive pattern probability distribution.

[0022] The emotion dimension feature vector and the cognitive pattern probability distribution are concatenated to generate a multimodal feature fusion result.

[0023] Optionally, step S4 inputs the feature fusion result into the evaluation model for dynamic state deduction processing, and outputs psychological state evaluation elements, including:

[0024] S41. Analyze the contextual relationships of the current evaluation scenario and generate state priority weights associated with the multimodal feature fusion results;

[0025] S42. Generate a state transition matrix based on the state priority weights, and traverse and sort the modality type identifiers in the multimodal feature set;

[0026] S43. Using a dynamic programming strategy, select the optimal evolution sequence from the state transition matrix to generate a basic evaluation set containing emotion labels, cognitive states, and trend predictions;

[0027] S44. Logically integrate the basic evaluation set and trend prediction parameters to generate a state description segment that conforms to the target evaluation framework.

[0028] Optionally, the method further includes:

[0029] During the evaluation process, status response data is captured in real time to generate feedback logs containing feature offsets and status change information.

[0030] Extract the abnormal state features from the feedback log, and perform similarity matching between the abnormal state features and the historical evaluation case library to generate adaptive adjustment instructions;

[0031] The modality association threshold parameters of the machine learning model are dynamically updated based on the adaptive adjustment instructions.

[0032] The updated modal association threshold parameters are injected into the evaluation model, and the state priority weights in the state transition matrix are recalculated.

[0033] Optionally, the step of extracting abnormal state features from the feedback log and performing similarity matching between the abnormal state features and the historical evaluation case library to generate adaptive adjustment instructions includes:

[0034] The abnormal state features are divided into time windows to generate multiple state segments and their corresponding multimodal data sequences;

[0035] The trained anomaly classification model is used to perform root cause analysis on each state segment, generating classification labels that include feature recognition bias, temporal drift, and response delay.

[0036] Retrieve solution templates that match the classification labels from the historical evaluation case library to generate a set of candidate adjustment strategies;

[0037] Based on the matching degree ranking between the multimodal data sequence and the candidate adjustment strategy set, the strategy with the highest confidence is selected to generate an adaptive adjustment instruction.

[0038] Optionally, the step of using a trained anomaly classification model to perform root cause analysis on each state segment and generating classification labels that include feature recognition bias, temporal drift, and response delay includes:

[0039] The state segments are processed by time series segmentation to generate multimodal data sequence slices containing start and end times;

[0040] The multimodal data sequence slices are subjected to keyframe sampling processing to generate a set of state change keyframes and corresponding timestamp indices;

[0041] Extract the feature offset between adjacent keyframes in the set of keyframes for state changes, and generate feature evolution trajectory vectors and time interval sequences;

[0042] The feature evolution trajectory vector is processed by trend calculation to generate an abnormal evolution pattern feature vector;

[0043] The feature vector of the abnormal evolution pattern is input into the spatiotemporal convolutional subnet of the trained abnormal classification model for local feature extraction, generating a spatial abnormal activation map.

[0044] The time interval sequence is subjected to sliding window mean filtering to generate a smoothed response delay time sequence;

[0045] The response delay time series is input into the gated recurrent subnet of the trained anomaly classification model for periodic pattern matching to generate a time dimension anomaly score; the spatial anomaly activation map is subjected to channel max pooling to generate a spatial dimension anomaly score; the spatial dimension anomaly score and the time dimension anomaly score are subjected to feature cross-fusion processing to generate a spatiotemporal joint anomaly probability distribution.

[0046] Peak detection results of feature recognition deviation probability, temporal drift probability, and response delay probability are extracted from the spatiotemporal joint anomaly probability distribution; dynamic threshold comparison processing is performed on the preset probability threshold based on the peak detection results to generate a candidate set of classification labels containing probability ranking;

[0047] The candidate classification labels are subjected to timestamp index alignment verification to generate target classification labels that are consistent with the phase of the abnormal state in the multimodal data sequence slice; the target classification labels and the feature evolution trajectory vector are subjected to anomaly type reverse verification to generate a final classification label set containing confidence weights.

[0048] Optionally, the method further includes:

[0049] Virtual feature perturbation parameters are injected before the evaluation model is executed. These virtual feature perturbation parameters are used to simulate the random offset scenario of multimodal data features.

[0050] Monitor the fault-tolerant processing results of the evaluation model on the perturbed features and generate stability evaluation indicators;

[0051] When the stability evaluation index is lower than a preset threshold, the online learning mode of the completed machine learning model is triggered.

[0052] The temporal feature parameters of the trained machine learning model are updated by gradient backpropagation based on the feature difference data before and after the perturbation.

[0053] The evaluation model monitors its fault-tolerant processing results for the perturbed features and generates stability evaluation indices, including:

[0054] The number of evaluation failures and the anomaly capture success rate caused by feature errors during the statistical evaluation process are recorded.

[0055] Calculate the ratio of the number of evaluation failures to the total number of evaluation steps to generate a first evaluation coefficient;

[0056] Extract the proportion of events that match the preset fault tolerance rules from the anomaly capture success rate, and generate a second evaluation coefficient;

[0057] The first evaluation coefficient and the second evaluation coefficient are weighted and summed to generate a comprehensive stability evaluation index.

[0058] Optionally, updating the temporal feature parameters of the trained machine learning model using gradient backpropagation based on the feature difference data before and after the perturbation includes:

[0059] The feature difference data before and after the perturbation are time-stamp aligned to generate a matching pair sequence containing the feature set before the perturbation and the feature set after the perturbation;

[0060] Extract the feature offsets from the matching pair sequence to generate the horizontal and vertical feature vectors for each multimodal data.

[0061] The distance between the horizontal feature vector and the vertical feature vector is calculated to generate a set of feature difference vectors for each multimodal data.

[0062] A feature regression loss function is constructed based on the set of feature difference vectors. The feature regression loss function includes the mean square error between the features predicted by the trained machine learning model and the actual features after perturbation.

[0063] The feature regression loss function is subjected to a differentiable transformation to generate the loss value tensor required for gradient backpropagation;

[0064] Iterate through the temporal feature parameters of the trained machine learning model, calculate the partial derivative of the loss tensor with respect to each temporal feature, and generate the temporal feature gradient matrix.

[0065] The temporal feature gradient matrix is ​​subjected to momentum smoothing to generate a decayed parameter update direction vector;

[0066] Based on the updated direction vector and the preset learning rate parameter, the temporal feature parameters of the trained machine learning model are iteratively and incrementally adjusted.

[0067] During the incremental adjustment process, the rate of change of the feature regression loss function is monitored in real time to generate a loss convergence status indicator.

[0068] When the loss convergence state indicator reaches a stable threshold, the gradient backpropagation update is terminated and the temporal feature parameters are frozen.

[0069] The updated temporal feature parameters are injected into the forward propagation path of the trained machine learning model, and the modality segmentation and feature fusion processes are re-executed to verify the improved feature extraction accuracy.

[0070] Optionally, the method further includes:

[0071] Construct a cross-scenario evaluation adapter, and use the cross-scenario evaluation adapter to parse the semantic differences of metrics in different evaluation frameworks;

[0072] The psychological state assessment elements are converted into underlying assessment instructions supported by the target scenario;

[0073] The temporal dependencies and exception handling context of the state labels are preserved during the transformation process;

[0074] Inject the evaluation optimization parameters that match the target scenario to generate an evaluation result descriptor that meets the cross-scenario execution conditions.

[0075] Optionally, the step of constructing a cross-scenario evaluation adapter and resolving the semantic differences of metrics across different evaluation frameworks using the cross-scenario evaluation adapter includes:

[0076] Establish a semantic rule base for the evaluation framework, and store the indicator mapping tables and parameter transmission paths for each scenario;

[0077] Abstract semantic tree parsing is performed on the psychological state assessment elements to generate intermediate representation layer data;

[0078] Based on the intermediate representation layer data, the indicator mapping table is traversed and queried to generate a scenario-compatible indicator replacement scheme;

[0079] Dependency injection is performed on conflicting parameter passing paths to generate unambiguous index transformation results;

[0080] The process of performing dependency injection on conflicting parameter passing paths to generate unambiguous index transformation results includes:

[0081] Identify conflicting nodes in the parameter passing path that have type ambiguity or overlapping scope, and generate a set of conflicting path identifiers;

[0082] Extract the context metadata of each node in the set of conflict path identifiers. The context metadata includes parameter type declarations, lifecycle markers, and call stack fingerprints.

[0083] Based on the context metadata, an interface proxy class name corresponding to the conflict node is generated. The interface proxy class name inherits the native interface of the target scenario and injects type casting constraints.

[0084] Parse the dependency chain of the interface proxy class name in the target scenario and generate a dynamic binding strategy that includes parameter scope isolation boundaries;

[0085] Based on the dynamic binding strategy, conflicting nodes in the parameter passing path are mapped to scenario-compatible interface method instances, generating a dynamic binding queue of parameter instances.

[0086] Traverse the interface method instances in the dynamic binding queue, verify the compatibility status of the interface method instances with the target scenario indicator semantics, and generate parameter binding validity status codes;

[0087] Based on the parameter binding validity status code, filter out interface method instances without type conflicts and generate a scenario-independent intermediate indicator sequence;

[0088] The intermediate indicator sequence is filled into the indicator template according to the semantic rules of the target scenario to generate an unambiguous indicator conversion result.

[0089] Compared with the prior art, the beneficial effects of the present invention are:

[0090] This intelligent dynamic assessment method for adolescent psychological states based on multimodal data can capture signs of psychological activity from different levels by collecting multimodal psychological data that includes vocal emotional information and facial expression-behavioral relationships. Vocal emotional information contains various emotion-related cues such as tone, speech rate, and intonation, while facial expression-behavioral relationships convey immediate psychological reactions through facial muscle movements and eye changes. The two complement each other, forming a rich source of information reflecting psychological states, thus avoiding the assessment limitations caused by the limited information in single-modal data.

[0091] In the feature extraction and fusion stage, dynamic features are extracted and processed using a well-trained machine learning model to generate a multimodal feature set containing a mapping relationship between modality type identifiers and temporal codes. This approach overcomes the limitations of simple concatenation in traditional feature fusion. By establishing a mapping between modality types and temporal codes, features from different modalities can be placed within a unified temporal framework, allowing the correlation between features of different modalities to be fully explored, and making the fused features better reflect the overall characteristics of psychological states.

[0092] State analysis networks perform temporal correlation analysis on dynamic features, generating a fusion result that includes emotional and cognitive pattern features. This process deeply considers the continuity and correlation of psychological states over time. Through temporal correlation analysis, the intrinsic connections between psychological features at different points in time can be discovered, clearly demonstrating the interaction between emotional changes and cognitive pattern evolution. This allows the understanding of psychological states to extend beyond isolated features to grasping their dynamic development patterns.

[0093] The feature fusion results are input into the evaluation model for dynamic state extrapolation, outputting psychological state evaluation elements. This dynamic extrapolation mechanism can adapt to the real-time changes in psychological states. By continuously receiving new multimodal data, constantly updating feature information and performing extrapolation, it can generate evaluation elements that reflect the current psychological state in real time, keeping the evaluation process synchronized with changes in psychological states, and promptly capturing subtle changes that might be overlooked by static evaluation.

[0094] This method establishes a complete workflow from multimodal data acquisition, dynamic feature extraction, multimodal feature fusion, temporal correlation analysis to dynamic state deduction. Each step works in concert to provide a more comprehensive, in-depth assessment of psychological states that accurately reflects their dynamic nature. This assessment approach can adapt to the needs of psychological state evaluation in various scenarios, providing more valuable assessment information for both long-term tracking of patients' psychological states in clinical diagnosis and real-time monitoring of students' learning psychology in education. Attached Figure Description

[0095] Figure 1 This is a schematic diagram illustrating the working principle of the intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in this invention.

[0096] Figure 2 This is a flowchart of feature fusion processing and temporal correlation analysis;

[0097] Figure 3 A flowchart for dynamic state deduction processing;

[0098] Figure 4 A flowchart for real-time feedback and dynamic updates;

[0099] Figure 5 A flowchart for generating abnormal state feature handling and adjustment instructions. Detailed Implementation

[0100] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0101] like Figure 1 As shown, this invention provides an intelligent dynamic assessment method for the psychological state of adolescents based on multimodal data. The following detailed description of the method is provided in conjunction with specific steps:

[0102] S1. Obtain multimodal psychological data of the target assessment object and extract dynamic features from the multimodal psychological data; wherein, the multimodal psychological data includes vocal emotion information and facial expression behavior relationships.

[0103] Dynamic features refer to the time-varying characteristic information extracted from the multimodal psychological data of the target assessment object. Combining the vocal emotion information and facial expression behavior relationships contained in the multimodal psychological data, specifically, the dynamic features in vocal emotion information include changes in fundamental frequency profiles, formant trajectories, and short-time energy spectra over time; the dynamic features in facial expression behavior relationships include the evolution of facial motor unit intensity, eye movement vectors, and the rate of change of muscle micro-expressions over time. These features can reflect the trajectory of the assessment object's psychological state at different times, providing a foundation for subsequent dynamic assessment.

[0104] One feasible implementation method is, for example Figure 2 The flowchart illustrates the feature fusion processing and temporal correlation analysis. Modal domain processing is performed on the collected speech emotion information and facial expression behavior relationship data. A predefined modal segmentation algorithm decomposes the multimodal psychological data into independent speech and facial expression data domains. Each data domain generates corresponding initial feature parameters; specifically, fundamental frequency contours, formant trajectories, and short-time energy spectra are extracted for the speech domain; facial action unit intensity, eye movement vectors, and muscle micro-expression change rates are extracted for the facial expression domain. Attention-weighted processing is applied to the temporal feature sequences of each data domain, and a multi-head attention mechanism is used to calculate feature importance scores. A learnable weight matrix is ​​constructed, which generates relevance scores through the dot product of feature vectors and context vectors, and is normalized using a softmax function. Features with relevance scores exceeding a dynamic threshold are assigned exponentially increasing weight coefficients, while low-relevance features are weighted attenuated. The weighted multi-domain feature vector is output, which retains the spatiotemporal structure information of the original data. Specifically, when performing attention-weighted processing on the temporal feature sequences of each data domain, a multi-head attention mechanism is first used to calculate feature importance scores. This mechanism can capture the correlation between features from multiple dimensions. By constructing a learnable weight matrix, the feature vectors of each data domain are multiplied by the context vector to measure the relevance of features to the context. The results are then normalized using a softmax function to obtain standardized relevance scores. A dynamic threshold is used to judge the relevance scores. Highly relevant features exceeding the threshold are assigned exponentially increasing weight coefficients to highlight their importance; low-relevance features below the threshold have their weights decayed to reduce their impact on subsequent processing. The final output weighted multi-domain feature vector fully preserves the temporal and spatial structural relationships of the original data, ensuring the temporal integrity and spatial correlation of the features.

[0105] S2. The trained machine learning model is used to perform feature fusion processing on the multimodal psychological data to generate a multimodal feature set; wherein, the multimodal feature set contains the mapping relationship between modality type identifier and temporal coding.

[0106] The process involves matching and filtering multi-domain feature vectors against preset modality association thresholds. The cosine similarity of weighted feature vectors in Hilbert space is calculated, and a dynamic similarity threshold adjustment mechanism is implemented. When feature stability exceeds a critical value within multiple consecutive time windows, the threshold strictness is automatically increased. Candidate feature sets that meet confidence criteria are selected; these sets include cross-modal feature intersections. Redundancy elimination is performed on the candidate feature sets, and statistical dependencies between features are analyzed using mutual information entropy calculation. A feature mutual information matrix is ​​constructed; when the values ​​of off-diagonal elements in the matrix exceed a preset entropy upper limit, feature dimensionality reduction is performed. Principal component analysis is used to retain orthogonalized feature subsets with mutual information entropy below the critical value. A multimodal feature set containing modality type identifiers is output, where each feature is appended with a normalized temporal code. Temporal codes are generated based on the ratio of the timestamp sequence to the feature sampling period, mapping the absolute time coordinates to a standard time axis of [0,1]. For example, the temporal code value increases by 0.02 for every 0.1 seconds of feature point.

[0107] In one feasible implementation, a fixed-duration sliding window is used to extract psychological data. The window length is dynamically configured according to the assessment scenario; a 5-second window is set by default for clinical diagnosis scenarios, and a 2-second window is set for behavioral observation scenarios. The extracted feature time segments are standardized by calculating the mean and standard deviation of each modality feature and applying a Z-score transform to convert the feature values ​​to a distribution space with zero mean and unit variance. Random noise conforming to a Gaussian distribution is injected into the standardized feature segments, with the noise intensity set to 5%-10% of the feature standard deviation to enhance the model's robustness to input perturbations. The optimized state analysis input data is output, which preserves the relative temporal relationships of each modality feature.

[0108] Optionally, the specific implementation process of S2 includes S21-S24:

[0109] S21. Perform modality segmentation on the multimodal psychological data to generate multiple data domains and corresponding initial feature parameters;

[0110] S22. Extract the temporal feature sequence of each data domain, and perform attention weighting on the temporal feature sequence to generate a weighted multi-domain feature vector.

[0111] S23. Calculate the similarity between the multi-domain feature vector and the preset modality association threshold, and select candidate features that meet the confidence conditions.

[0112] S24. Redundancy elimination is performed on the candidate features to generate a multimodal feature set containing modality type identifiers. Each multimodal feature contains a normalized temporal code relative to the multimodal psychological data.

[0113] Specifically, based on a multimodal feature set, a state analysis network is used to perform temporal correlation analysis on the dynamic features, generating a feature fusion result that includes emotional dimension features and cognitive pattern features, including:

[0114] Based on the normalized temporal coding in the multimodal feature set, the multimodal psychological data is processed by time window truncation to generate multiple feature time segments;

[0115] Each feature time segment is standardized and feature-enhanced to generate optimized state analysis input data.

[0116] The recurrent memory subnet in the state analysis network is used to extract sequence features from the state analysis input data to generate emotion dimension feature vectors.

[0117] Parallel pattern classification models are used to perform feature pattern matching on feature time segments to generate cognitive pattern probability distributions.

[0118] The feature vector of the emotion dimension is concatenated with the probability distribution of the cognitive pattern to generate a multimodal feature fusion result.

[0119] The recurrent memory subnet in the state analysis network is used to extract sequence features from the state analysis input data, and the specific content of the emotion dimension feature vector is as follows.

[0120] When using the recurrent memory subnet within a state analysis network for sequence feature extraction, this subnet employs a three-layer gated recurrent unit architecture. It processes the state analysis input data through the coordinated operation of reset and update gates. The reset gate controls the forgetting rate of historical state information, while the update gate adjusts the degree to which new feature information is integrated into the current state. The optimized state analysis input data is sequentially fed into the recurrent memory subnet step by step. The network processes temporal features time-series by time step, learning long-term dependency patterns in the data. After traversing all time-step features, the hidden layer state at the last time step is output as the emotion dimension feature vector. The dimension of this vector is consistent with the hidden layer dimension of the gated recurrent unit in the recurrent memory subnet, comprehensively reflecting the emotional changes of the evaluated object over time.

[0121] S3. Based on the multimodal feature set, a state analysis network is used to perform temporal correlation analysis on dynamic features, generating feature fusion results that include emotional dimension features and cognitive pattern features.

[0122] In one feasible implementation, the specific process of using a state analysis network to perform temporal correlation analysis on dynamic features includes:

[0123] A recurrent memory subnet of a state analysis network is used to process time-series data. The recurrent memory subnet employs a three-layer gated recurrent unit architecture with 256 neurons in the hidden layer. The unit state update process involves the coordinated operation of reset and update gates: the reset gate controls the forgetting rate of historical states, and the update gate adjusts the fusion degree of new feature information. Feature segments at each time step are sequentially input into the network, which learns long-term dependency patterns through backpropagation. The hidden layer state at the final time step is output as the emotion dimension feature vector, with the same dimension as the hidden layer of the recurrent unit. Simultaneously, a pattern classification model is used to process the feature time segments in parallel. This model uses a combination of convolutional neural networks and fully connected layers, with the input layer receiving multi-channel data from the feature segments. Pattern features from local receptive fields are extracted through three convolutional layers with a 3×3×3 kernel size. Max pooling is performed in the pooling layer to compress the feature dimension, and finally, a softmax classification layer generates a cognitive pattern probability distribution vector, with the vector dimension corresponding to the preset number of cognitive pattern classifications. The cognitive pattern probability distribution vector is generated by the pattern classification model after matching feature patterns to feature time segments. Its vector dimension corresponds to the preset number of cognitive pattern classifications. Each element in the vector represents the probability that the psychological characteristics of the assessment object conform to a certain preset cognitive pattern. These probability values ​​can intuitively show the distribution tendency of the assessment object at the cognitive level, providing a quantitative basis at the cognitive level for subsequent integration of emotional dimension features and generation of comprehensive assessment results.

[0124] The specific implementation process of fusing the emotion dimension feature vector with the cognitive pattern probability distribution vector includes:

[0125] Tensor concatenation is performed along the feature dimensions. The process of dimensional alignment of the vectors before concatenation includes: when the dimension of the emotion feature vector is higher than that of the cognitive pattern probability distribution vector, the cognitive pattern probability distribution vector is linearly projected to expand its dimension; when the cognitive distribution dimension is even higher, the emotion feature vector is increased in dimension by a fully connected layer. The resulting fusion matrix has a row dimension equal to the total number of feature segments and a column dimension equal to the sum of the dimensions of the two vectors. The resulting matrix retains the temporal encoding indexes of the original feature segments, forming a multimodal feature fusion result that includes the emotion state intensity gradient and the cognitive pattern distribution.

[0126] S4. Input the feature fusion results into the evaluation model for dynamic state deduction processing, and output psychological state evaluation elements; among them, psychological state evaluation elements include state labels and trend prediction parameters generated based on feature evolution sequences.

[0127] Among them, such as Figure 3 The diagram shown is a flowchart of the dynamic state deduction process; as follows: Figure 4 The flowchart shown is for real-time feedback and dynamic updates.

[0128] Optionally, the specific implementation process of S4 includes S41-S44:

[0129] S41. Analyze the contextual relationships of the current evaluation scenario and generate state priority weights associated with the multimodal feature fusion results;

[0130] S42. Generate a state transition matrix based on state priority weights, and traverse and sort the modality type identifiers in the multimodal feature set;

[0131] In one feasible implementation, a domain knowledge graph of the current evaluation scenario is loaded. This graph stores causal rules between emotional states, cognitive pattern constraints, and historical state transition records in a triplet structure. A graph neural network traverses the knowledge graph nodes, calculating the centrality index of each emotional state node. The in-degree value of the node is multiplied by the feature confidence to generate a state priority weight matrix. Based on this matrix, a state transition matrix is ​​constructed, where the dimension is determined by the number of emotional categories in the current multimodal feature set. The matrix element definition rule is: if a valid path exists in the knowledge graph for the transition from emotional state A to B, the element value is equal to the product of the priority weight and the time decay coefficient; otherwise, it is set to zero. Simultaneously, topological sorting is performed on the modality type identifiers in the multimodal feature set, and the Kahn algorithm is used to analyze the dependencies between modalities, forming an ordered feature processing queue. Specifically, the domain knowledge graph of the current evaluation scenario is loaded, which stores causal rules between emotional states, cognitive pattern constraints, and historical state transition records in triplet form. A graph neural network is used to traverse the nodes in the knowledge graph, calculating the centrality index of each emotion state node. The in-degree value of the node is multiplied by the feature confidence value to generate a state priority weight matrix. A state transition matrix is ​​constructed based on this weight matrix. The matrix dimension is determined by the number of emotion categories in the multimodal feature set. If a valid transition path exists between emotion states in the knowledge graph, the corresponding element value in the matrix is ​​the product of the priority weight and the time decay coefficient; otherwise, it is zero. The Kahn algorithm is used to analyze the dependencies between modality type identifiers in the multimodal feature set. Topological sorting is used to form an ordered feature processing queue, ensuring that the feature processing order conforms to the dependency logic between modalities.

[0132] S43. Use dynamic programming to select the optimal evolution sequence from the state transition matrix to generate a basic evaluation set containing emotion labels, cognitive states, and trend predictions.

[0133] In one feasible implementation, the present invention employs a dynamic programming strategy to search for the optimal state evolution path, specifically including:

[0134] (1) Define the state evolution cost function, which consists of three parts: the probability of the emotion label of the current state, which is the feature fusion result, the transition cost from the previous state to the current state in the state transition matrix, and the logical conflict value between the cognitive pattern features and the current state.

[0135] (2) Perform the initialization phase and assign an initial cost to each emotional state; where the cost is equal to the negative logarithm of its initial probability;

[0136] (3) The recursive phase is performed using the Viterbi algorithm, which includes: traversing the feature processing queue at each time step, calculating the minimum cumulative cost of all possible states at the current moment and recording the backtracking path; and selecting the path with the minimum cumulative cost as the optimal sequence output at the last time step. This sequence includes a time-sorted sequence of emotion labels, a cognitive state classification code corresponding to each label (taking the classification corresponding to the maximum probability), and trend prediction parameters (generated by extrapolating the costs of the next three time steps). The results obtained above form the basic evaluation set.

[0137] S44. Logically integrate the basic evaluation set and trend prediction parameters to generate a state description segment that conforms to the target evaluation framework.

[0138] The framework configuration file predefines syntax rules for description segments. For example, the clinical scenario template includes three required fields: "Subject's Emotional State," "Cognitive Pattern Performance," and "Short-Term Evolution Trend." The emotional label sequence is mapped along a timeline into a composite structure of time phrases and state descriptions (e.g., "Anxiety is evident for the first 2 seconds, then transitions to calmness for the next 3 seconds"). Cognitive state encoding is converted into natural language descriptions by looking up a predefined cognitive pattern dictionary. Trend prediction parameters generate trend quantification values ​​through linear interpolation, such as "Anxiety levels will decrease by 30% within 5 seconds." The integrated description segment retains the timestamp markers of the original sequence, forming the final psychological state assessment elements.

[0139] Optionally, the method further includes:

[0140] During the evaluation process, status response data is captured in real time to generate feedback logs containing feature offsets and status change information.

[0141] Extract abnormal state features from feedback logs and perform similarity matching between abnormal state features and historical evaluation case library to generate adaptive adjustment instructions;

[0142] In one feasible implementation, a real-time feedback mechanism is initiated synchronously during the evaluation process. A state response data stream is captured via an asynchronous thread. This data stream contains feature offsets of the intermediate layers of the evaluation model (i.e., the spatial distance between real-time features and baseline features) and state change events (emotion category switching or cognitive pattern mutation). The data stream is formatted to generate a structured feedback log, where log entries include timestamps, offset vectors, and event type codes.

[0143] Optionally, abnormal state features are extracted from the feedback log, and the abnormal state features are matched with the historical evaluation case library to generate adaptive adjustment instructions, including:

[0144] The abnormal state characteristics are divided into time windows to generate multiple state segments and their corresponding multimodal data sequences;

[0145] In a feasible implementation, the dynamic partitioning of the time window for anomaly feature extraction includes: setting an initial window length based on the typical anomaly duration in the historical evaluation case library, with a default value of 1.2 seconds; and triggering dynamic expansion (maximum expansion to 3 seconds) when the anomaly density within the window exceeds a critical value. Each window generates an independent state fragment and its corresponding multimodal data sequence fingerprint (MD5 digest value of the sequence features).

[0146] The trained anomaly classification model is used to perform root cause analysis on each state segment, generating classification labels that include feature recognition bias, temporal drift, and response delay.

[0147] Specifically, the process of the trained anomaly classification model for handling state segments includes: inputting the state segment into a three-dimensional convolutional layer to extract spatial features, capturing temporal patterns through a long short-term memory network, and finally outputting a multi-dimensional classification label containing three dimensions: feature recognition bias, temporal drift, and response delay.

[0148] The anomaly classification model is used for root cause analysis of anomalous state fragments. It can identify three types of anomalies: feature recognition bias, temporal drift, and response delay. Its structure combines a spatiotemporal convolutional subnet and a gated recurrent subnet. The spatiotemporal convolutional subnet is responsible for extracting spatial anomaly information from the anomalous state features and generating a spatial anomaly activation map; the gated recurrent subnet is used to capture anomalous patterns in the temporal dimension and generate a temporal anomaly score. The outputs of both are fused through feature cross-fertilization to obtain a spatiotemporal joint anomaly probability distribution, thereby achieving accurate classification of anomaly root causes. Building this model requires first collecting a labeled dataset containing features of various anomalous states. Using this dataset, the model is trained. By iteratively adjusting the feature extraction parameters of the spatiotemporal convolutional subnet and the temporal learning parameters of the gated recurrent subnet, the model's ability to identify and classify anomalous features is optimized to meet the accuracy requirements of root cause analysis.

[0149] Among them, such as Figure 5 The diagram shows the flowchart for abnormal state feature processing and adjustment instruction generation.

[0150] Optionally, a trained anomaly classification model is used to perform root cause analysis on each state segment, generating classification labels that include feature recognition bias, temporal drift, and response delay, including:

[0151] The state segments are processed by time series segmentation to generate multimodal data sequence slices containing start and end times;

[0152] Specifically, a dynamic temporal boundary detection mechanism is used to detect the zero-crossing points of the second derivative of the multimodal data stream to determine the start and end times, generating sequence slices with variable durations but satisfying minimum information content constraints. Each slice contains a speech spectral feature matrix and an facial expression geometric feature tensor under continuous timestamp indices.

[0153] Keyframe sampling is performed on slices of multimodal data sequences to generate a set of keyframes for state changes and their corresponding timestamp indices;

[0154] Among them, keyframe sampling processing is performed on the multimodal data sequence slices. The sampling algorithm is dynamically adjusted based on the feature evolution rate and uses the gradient change rate as the core criterion. The feature difference measurement formula (1) between adjacent time frames is defined. This formula describes the degree of drastic change of state in the feature space and is expressed by the following formula (1):

[0155]

[0156] Where Γ(t) represents the intensity of the characteristic change at time point t, and M is the total number of modal types. Let Γ(t) be the feature vector of the m-th mode at time t, where τ is the time base unit (default is 0.1 seconds). When Γ(t) exceeds the adaptive threshold (which is determined by multiplying the average intensity of the first ten time windows by a coefficient of 1.3), that time point is marked as a keyframe. The generated set of state change keyframes carries timestamp index offsets;

[0157] Extract the feature offset between adjacent keyframes in the set of keyframes for state changes, and generate feature evolution trajectory vectors and time interval sequences;

[0158] Extract the feature offsets of adjacent keyframes in the keyframe set; calculate t k With t k+1 The multidimensional feature differences between two keyframes are used to construct the feature evolution trajectory vector: D k It can be expressed by the following formula (2):

[0159]

[0160] Simultaneously record the keyframe time interval sequence

[0161] The feature evolution trajectory vector is processed by trend calculation to generate anomaly evolution pattern feature vector;

[0162] Among them, the linear regression model that fits the feature trajectory within the sliding time window extracts the slope parameter to form the pattern feature, which is expressed by the following formula (3):

[0163] S=[β1,β2,...,β M (3)

[0164] Where, β m The coefficient represents the evolution trend of the m-th modal feature; positive values ​​indicate an increasing trend, while negative values ​​indicate a decreasing trend. S represents the direction and intensity of the evolution of the anomalous state reflected by the mode feature vector.

[0165] The feature vector of the abnormal evolution pattern is input into the spatiotemporal convolutional subnet of the trained abnormal classification model to extract local features and generate a spatial abnormal activation map.

[0166] The spatiotemporal convolutional subnetwork receives the feature vector S of the anomalous evolution pattern and expands the receptive field step by step through three layers of dilated convolution operations: the first convolutional kernel size is 3×M (time dimension × modality dimension) with a dilation coefficient of 1; the second kernel size is 5×M with a dilation coefficient of 2; and the third kernel size is 7×M with a dilation coefficient of 4. After each convolutional layer, the ReLU activation function and batch normalization are applied, ultimately outputting a spatial anomalous activation map. Where T ′ This is the length after compression of the time dimension.

[0167] A sliding window mean filter is applied to the time interval sequence to generate a smoothed response delay time series.

[0168] In one feasible implementation, a time-dimensional processing path is used to apply a sliding window mean filter to the response delay time series ΔT. The window width is set to 20% of the series length, and a Hanning window is used to weight and smooth the data to obtain a smoothed series.

[0169] The response delay time series is input into the gated recurrent subnet of the trained anomaly classification model for periodic pattern matching to generate a time dimension anomaly score; channel max pooling is performed on the spatial anomaly activation map to generate a spatial dimension anomaly score; the spatial dimension anomaly score and the time dimension anomaly score are subjected to feature cross-fusion processing to generate a spatiotemporal joint anomaly probability distribution.

[0170] The smoothed sequence is input into a gated recurrent subnet, which employs a two-layer bidirectional GRU architecture with 64 hidden layer units. The forward and backward output states are concatenated and passed through a temporal attention layer to calculate the importance weight of each time step, generating a temporal dimension anomaly score vector. Where T ″ The number of time steps output by GRN.

[0171] Peak detection results of feature identification deviation probability, temporal drift probability, and response delay probability are extracted from the spatiotemporal joint anomaly probability distribution; dynamic threshold comparison processing is performed on preset probability thresholds based on the peak detection results to generate a candidate set of classification labels containing probability ranking;

[0172] Specifically, a feature fusion module receives dual-path outputs: a time-dimensional processing path and a spatial-dimensional anomaly output. Cross-channel max pooling is performed on the spatial anomaly activation map A to reduce its dimensionality and generate a spatial-dimensional anomaly score vector. Compare this vector with the time dimension anomaly score vector E t The input feature cross layer is used for calculation. The specific calculation process includes:

[0173] Calculate the outer product of vectors Generate T ′ ×T ″ Interaction matrix;

[0174] Add learnable location encoding to preserve spatiotemporal correlation;

[0175] Compressed into a three-dimensional probability space through a two-layer fully connected network;

[0176] Output spatiotemporal joint anomaly probability distribution Its three dimensions correspond to the feature recognition bias p. d Timing drift p s and response delay p r Probability estimates.

[0177] Peak detection and classification decisions are performed in the probability space, and a dynamic threshold generation rule is defined, wherein different anomaly probabilities θ are output, which are expressed by the following formula (4):

[0178] θ=μ+ασ(4)

[0179] Where μ is the historical mean probability of normal samples, σ is the standard deviation, and α is the sensitivity control factor (default 1.5). When p d >θ d or p s >θ s or p r >θ rCandidate label generation is triggered at certain times. The probabilities of the three types of anomalies are sorted from high to low, generating a candidate set of classification labels with a sorted index.

[0180] The candidate classification labels are subjected to timestamp index alignment verification to generate target classification labels that are consistent with the phase of abnormal states in the multimodal data sequence slices; the target classification labels and feature evolution trajectory vectors are subjected to anomaly type reverse verification to generate the final classification label set containing confidence weights.

[0181] In this embodiment of the invention, a dual verification mechanism is used to complete the timestamp index alignment verification process for the candidate set of classification labels, specifically including:

[0182] Align the timestamp index: align the timestamp t corresponding to the candidate label. k Map back to the original data slice and check if there are any underlying feature jumps that support the anomalous event within 1 second before and after that moment (such as a change in the speech fundamental frequency exceeding 50Hz or a jump in the intensity of facial expression units exceeding 40%).

[0183] Reverse validation of anomaly types: The anomaly type specified by the classification label (e.g., response delay) is reverse-mapped to the feature evolution trajectory vector to validate D. k Does the expected pattern exist (e.g., delayed categories should exhibit an exponential decay pattern)? The verified labels are assigned confidence weights, expressed by the following formula (5):

[0184]

[0185] The final generated form is {(type, w, t)} k The final set of category labels for t, where t kThe precise moment when the anomaly occurred is marked; w is the confidence weight, and Pi is the type value of the i-th label. Feature matching degree is a quantitative indicator measuring the degree of fit between the feature evolution trajectory vector of the current abnormal state (constructed from feature offsets such as the fundamental frequency contour and formant trajectory of speech emotion information between adjacent keyframes, and feature offsets such as the intensity of facial action units and eye movement vectors in relation to facial expressions and behaviors) and the corresponding classification labels (such as feature recognition deviation and temporal drift) feature patterns in a preset typical anomaly feature pattern library. It is obtained by calculating the cosine similarity between the current speech features and typical patterns, the normalized Euclidean distance between the facial expression features and typical patterns, and then summing them according to their weights. The value range is [0,1]. Temporal consistency is the average of the absolute values ​​of the Pearson correlation coefficients between the start-end time interval of the anomaly classification label and the actual time interval of feature jumps (such as sudden changes in speech fundamental frequency or displacement of facial expression key points) in the multimodal data sequence, and the time interval of the anomaly probability change trend corresponding to the label and the evolution trajectory trend of the multimodal features. The value range is [0,1]. It is used to ensure that the label is in phase with the actual occurrence of anomalies in the time dimension. Both serve as the basis for verifying the validity of the label and participate in the confidence weight calculation together with the anomaly probability (p_i) corresponding to the label.

[0186] Retrieve solution templates that match the category tags from the historical evaluation case library to generate a set of candidate adjustment strategies;

[0187] Based on the matching degree ranking of multimodal data sequences and candidate adjustment strategy sets, the strategy with the highest confidence is selected to generate adaptive adjustment instructions.

[0188] In one feasible implementation, an inverted index strategy is used to retrieve solutions from a historical evaluation case library. The specific process includes: constructing an inverted index table of anomaly type labels and case entries; retrieving matching candidate solution templates; filtering the candidate set based on similarity; calculating the Jaccard similarity coefficient between the current multimodal data sequence and the fingerprint of the case library sequence, which is calculated by matching the differences in the positions of key events in the data stream; selecting candidate solutions with a similarity coefficient greater than 0.85 to add to the adjustment strategy set; and finally selecting a strategy through a weighted voting mechanism, where each candidate strategy is assigned a basic weight based on its historical success record, while an additional weight is added based on the current scenario strategy matching degree, and the strategy with the highest total weight is selected to generate adaptive adjustment instructions.

[0189] The modal correlation threshold parameters of the machine learning model are dynamically updated based on adaptive adjustment instructions;

[0190] In one feasible implementation, an incremental gradient adjustment method is used to update the modal correlation threshold, specifically including:

[0191] Analyze the parameter update direction (boost / decrease) and magnitude (discrete gradient level) in the adaptive adjustment command;

[0192] Dynamically modifying modality association threshold parameters during the runtime of a machine learning model includes inserting a threshold correction operator into the similarity calculation layer. This operator performs a linear transformation on the original threshold based on the gradient level, such as gradient level 1 corresponding to a threshold of ±0.05.

[0193] The updated modal association threshold parameters take effect immediately, triggering a recalculation of the state transition matrix, clearing the existing matrix cache, and regenerating the state priority weights and transition matrix element values ​​using the new thresholds. The evaluation model automatically continues its analysis task in the next processing cycle using the updated logical parameters.

[0194] The updated modal association threshold parameters are injected into the evaluation model, and the state priority weights in the state transition matrix are recalculated.

[0195] Optionally, the method further includes:

[0196] Virtual feature perturbation parameters are injected before the evaluation model is executed. These virtual feature perturbation parameters are used to simulate the random offset scenario of multimodal data features.

[0197] For the voice emotion information stream, the amplitude modulation perturbation scheme includes: selecting 10% of the voice frames at a fixed sampling interval and applying a random fluctuation of ±15% to their amplitude spectrum. For the facial expression behavior data stream, a spatial coordinate offset perturbation is implemented: randomly selecting 20% ​​of the facial key points and generating a displacement perturbation within the range of [-0.2, 0.2] cm in three-dimensional space. The perturbation parameter configuration file is dynamically loaded through the system configuration interface. The configuration file uses a JSON structure to define the perturbation type and range of action. The key parameters are shown in Table 1 below.

[0198] Table 1

[0199]

[0200] After the perturbation is executed, the fault tolerance monitoring process of the evaluation model is initiated. A real-time diagnostic log system is established to record abnormal events in the evaluation process: when the emotion label changes abruptly within an adjacent time window (e.g., the label is "calm" for 0-2 seconds and changes to "angry" for 2-4 seconds), it is recorded as an evaluation failure event caused by feature error; when the feature offset exceeds the safety threshold but is successfully corrected by the model (e.g., a sudden change in voice amplitude is identified as device noise), it is recorded as an anomaly capture success event. A total of 42 evaluation steps are recorded within a 15-minute monitoring period. Example data records of the diagnostic log output are as follows:

[0201] At 3 minutes and 12 seconds: An alarm was triggered by a 0.18cm shift in facial expression coordinates. After correction, the emotion tag remained stable (anomaly capture successful).

[0202] At 5 minutes and 47 seconds: A sudden change in the fundamental frequency of the voice caused the label to jump from "calm" to "anxious" (assessment failed);

[0203] 11 minutes and 03 seconds: All three modal features drifted simultaneously, and the model automatically switched to alternative analysis paths (anomaly successfully captured).

[0204] The monitoring and evaluation model assesses the fault-tolerant processing results of the features after disturbance and generates stability evaluation indicators.

[0205] The process of generating stability indicators based on diagnostic logs includes: extracting the total number of evaluation steps (42) within a 15-minute monitoring period and counting 7 evaluation failure events. The first evaluation coefficient is calculated as: 7 / 42≈0.167. From the 28 anomaly capture events, 19 events conforming to the fault tolerance rules are identified (e.g., feature drift amplitude within the fault tolerance range, multimodal conflict within the threshold range), generating the second evaluation coefficient as: 19 / 28≈0.679. A weighted calculation is performed using a 0.6:0.4 weight allocation: 0.6×0.167+0.4×0.679≈0.365, resulting in a comprehensive stability evaluation index of 0.365 (range [0,1], lower values ​​indicate greater stability).

[0206] When the stability assessment metric falls below a preset threshold, the online learning mode of the completed machine learning model is triggered.

[0207] Specifically, when the stability index falls below a preset threshold of 0.4, the online learning mode of the machine learning model is triggered. Feature data from the sixth monitoring time window (8 minutes 30 seconds - 9 minutes 45 seconds) is selected for parameter updates. This period contains three valid perturbation events:

[0208] Speech feature perturbation: The fundamental frequency value in frames 52-55 changes from 198Hz to 215Hz to 226Hz;

[0209] Perturbation of facial features: coordinates of key points of the left eyebrow (3.12, 5.78) → (3.32, 5.62) → (3.15, 5.83);

[0210] Cross-modal joint perturbation: The speech energy spectrum is reduced by 3dB while the corner of the mouth coordinates shift by 0.12cm.

[0211] A multidimensional spatial distance metric was used to construct the difference matrix. The fundamental frequency features of the speech were used to generate a horizontal vector (198, 215, 226) and a vertical vector (198, 217.6, 219.2), the latter being the baseline value before perturbation. The corrected Euclidean distance between the two was calculated.

[0212] Among them, the coordinate displacement vector for calculating facial expression features is Δ=[(3.32-3.15),(5.62-5.83)]=(0.17,-0.21), with a modulus of 0.17. The difference vectors of all features within this time window are summarized to form a difference matrix with dimensions of 8 (number of features) × 15 (time steps).

[0213] The temporal feature parameters of the trained machine learning model are updated by gradient backpropagation based on the feature difference data before and after the perturbation.

[0214] The gradient backpropagation update is performed using a momentum acceleration mechanism. The initial learning rate is set to 0.002, and the momentum factor to 0.89. The temporal feature parameter update direction is calculated based on a weighted average of historical gradients. Specifically, the current iteration gradient vector is the column mean of the difference matrix, and the momentum vector is the exponentially weighted average of the gradients from the previous five iterations. Specifically, during the sixth iteration, the speech feature parameter update amount is:

[0215] New gradient = Current gradient (6.97,...) × 0.11 + Momentum vector (7.81,...) × 0.89;

[0216] Parameter update value = -learning rate 0.002 × new gradient;

[0217] After each full parameter update, a fast inference process is performed on the validation set, including: comparing the output features of the feature extraction layer with the baseline features using cosine similarity and recording the trajectory of the mean similarity change. The validation results after the sixth iteration are recorded as follows:

[0218] First update: Similarity 0.812 → Second update: 0.827 → Third update: 0.834

[0219] 4th time: 0.841 → 5th time: 0.843 → 6th time: 0.845

[0220] Specifically, when the similarity improvement for three consecutive iterations is less than 0.004 (improvement of +0.002, +0.002, and +0.002 for iterations 4-6), the loss convergence state is considered to have reached a stable condition. The current time-series feature parameters are frozen, a model parameter snapshot file (containing 32,119 floating-point parameter values) is generated, and the current update cycle is terminated.

[0221] The updated machine learning model was immediately put into online validation. Modal segmentation processing for the sixth time window was re-executed: after speech feature segmentation, the standard deviation of the fundamental frequency sequence decreased from 12.4Hz to 9.7Hz; after facial expression feature segmentation, the false alarm rate for coordinate offset detection decreased from 22% to 15%. In the feature fusion stage, the cross-modal feature alignment accuracy improved from 78.3% to 83.6%, demonstrating enhanced feature extraction accuracy. The entire update cycle took 137 seconds, during which the evaluation service remained uninterrupted.

[0222] The monitoring and evaluation model generates stability evaluation indicators based on the fault-tolerant processing results of the perturbation features, including:

[0223] The number of evaluation failures and the anomaly capture success rate caused by feature errors during the statistical evaluation process are recorded.

[0224] Calculate the ratio of the number of evaluation failures to the total number of evaluation steps to generate the first evaluation coefficient;

[0225] Extract the proportion of events that match the preset fault tolerance rules from the anomaly capture success rate, and generate a second evaluation coefficient;

[0226] The first evaluation coefficient and the second evaluation coefficient are weighted and summed to generate a comprehensive stability evaluation index.

[0227] Optionally, gradient backpropagation is performed to update the temporal feature parameters of the trained machine learning model based on the feature difference data before and after the perturbation, including:

[0228] The feature difference data before and after the perturbation are timestamped to generate a matching sequence containing the feature set before the perturbation and the feature set after the perturbation;

[0229] Extract the feature offsets from the matching pair sequences to generate the horizontal and vertical feature vectors for each multimodal data set;

[0230] The distance between the horizontal and vertical feature vectors is calculated to generate a set of feature difference vectors for each multimodal data.

[0231] A feature regression loss function is constructed based on the set of feature difference vectors. The feature regression loss function includes the mean square error between the features predicted by the trained machine learning model and the actual features after perturbation.

[0232] Perform a differentiable transformation on the feature regression loss function to generate the loss value tensor required for gradient backpropagation;

[0233] Iterate through the temporal feature parameters of the trained machine learning model, calculate the partial derivative of the loss tensor with respect to each temporal feature, and generate the temporal feature gradient matrix.

[0234] Momentum smoothing is applied to the temporal feature gradient matrix to generate a decayed parameter update direction vector;

[0235] Based on the parameter update direction vector and the preset learning rate parameter, the temporal feature parameters of the trained machine learning model are iteratively and incrementally adjusted.

[0236] During the incremental adjustment process, the rate of change of the feature regression loss function is monitored in real time to generate a loss convergence status indicator.

[0237] When the loss convergence state indicator reaches a stable threshold, the gradient backpropagation update is terminated and the temporal feature parameters are frozen.

[0238] The updated temporal feature parameters are injected into the forward propagation path of the trained machine learning model, and the modality segmentation and feature fusion processes are re-executed to verify the improved feature extraction accuracy.

[0239] Optionally, the method further includes:

[0240] Construct a cross-scenario evaluation adapter and use it to parse the semantic differences of metrics across different evaluation frameworks;

[0241] Optionally, a cross-scenario evaluation adapter is constructed, and the semantic differences of metrics across different evaluation frameworks are resolved using this adapter, including:

[0242] Establish a semantic rule base for the evaluation framework, and store the indicator mapping tables and parameter transmission paths for each scenario;

[0243] Among them, a triplet storage structure is used to organize data to establish an evaluation framework semantic rule base. The entry directory of the rule base includes the source scenario indicator name, the target scenario corresponding indicator name, the transformation constraint conditions, and the historical transformation power parameters.

[0244] Abstract semantic tree parsing is performed on the psychological state assessment elements to generate intermediate representation layer data;

[0245] In one feasible implementation, the input elements are structured as hierarchical objects in JSON format, with the top-level fields containing a sequence of state labels, trend prediction parameters, and anomaly context stacks. The parser decomposes the object layer by layer, breaking down complex state descriptions (such as "anxiety with distraction") into atomic emotion labels ("anxiety") and cognitive labels ("distraction"); and decomposing time-series predictions into three atomic parameters: base value, slope of change, and confidence interval. The parsing process preserves the nesting relationships of the original data, generating intermediate representation layer data—this data structure is a multi-branch tree, with leaf nodes storing atomic parameters and non-leaf nodes storing logical operations.

[0246] Based on the intermediate presentation layer data, the indicator mapping table is traversed and queried to generate a scenario-compatible indicator replacement scheme.

[0247] In one feasible implementation, the process of traversing and querying the indicator mapping table based on intermediate representation layer data includes: backtracking upwards from the leaf nodes and retrieving matching mapping entries from the semantic rule base. A dual matching strategy is employed for retrieval: first, precise matching of parameter names and data types is performed; if no results are found, a fuzzy matching algorithm is used to calculate the semantic similarity of the parameter description text. Successfully matched nodes are replaced with target scenario indicators, such as mapping "classroom participation" in an educational scenario to "social interaction willingness" in a clinical scenario. For cases with multiple candidate mappings, a priority ranking mechanism is used, prioritizing mapping schemes with high historical conversion success rates, while simultaneously verifying the compatibility of parameter value ranges.

[0248] Dependency injection is performed on conflicting parameter passing paths to generate unambiguous index transformation results;

[0249] Parameter passing path conflict resolution is performed after mapping. Nodes with type ambiguity in the parameter dependency graph are detected—for example, when two scenarios define the "response latency" parameter as milliseconds and a rating, respectively. Contextual metadata of conflicting nodes is extracted, including: data type declarations (float32 vs. int16), lifecycle markers (session-level vs. task-level), and call stack fingerprints (e.g., {func:analyze(), line:203}). An interface proxy class name is generated based on the metadata, using the naming convention [source scenario]_[target scenario]Proxy[number]. The proxy class inherits the target scenario's native interface and injects type casting logic: float32 values ​​are discretized into int16 ratings according to a preset scale, and session-level parameters are converted to task-level parameters through time slicing.

[0250] This includes dependency injection for conflicting parameter passing paths to generate unambiguous index transformation results, including:

[0251] Identify conflicting nodes in the parameter passing path that have type ambiguity or overlapping scope, and generate a set of conflicting path identifiers;

[0252] Extract the context metadata of each node in the conflict path identifier set. The context metadata includes parameter type declarations, lifecycle markers, and call stack fingerprints.

[0253] The interface proxy class name corresponding to the conflict node is generated based on the context metadata. The interface proxy class name inherits the native interface of the target scenario and injects type casting constraints.

[0254] The dependency chain of the interface proxy class name in the target scenario is resolved, and a dynamic binding strategy containing parameter scope isolation boundaries is generated.

[0255] Based on the dynamic binding strategy, conflicting nodes in the parameter passing path are mapped to scenario-compatible interface method instances, generating a dynamic binding queue of parameter instances.

[0256] In one feasible implementation, a dynamic binding strategy is used to resolve the dependency chain of the proxy class. The runtime environment of the target scenario is scanned, and a dependency graph containing the complete call chain is constructed. A scope isolation layer is inserted into the graph: independent memory pages are allocated for the source scenario parameters, and the address space offset is calculated based on the scenario ID; a lazy loading mechanism is implemented for dependency binding, triggering parameter conversion only when the target scenario actually calls the interface method. A breadth-first search algorithm is used to construct the binding queue, ensuring that parent and child nodes are initialized in dependency order.

[0257] Iterate through the interface method instances in the dynamic binding queue, verify the compatibility status of the interface method instances with the target scenario indicator semantics, and generate parameter binding validity status codes.

[0258] The binding validity verification is performed dynamically during queue execution. A snapshot of the parameter syntax tree is generated at compile time, and at runtime, the matching degree between the actual call tree and the predefined template is compared. Verification metrics include: parameter type matching status, scope boundary consistency, and lifetime constraint satisfaction. A parameter binding validity status code is output—a three-bit encoding representing the type / scope / lifetime verification result. For example, status code 0b110 indicates that type and scope verification passed, but lifetime verification was abnormal. Binding is considered successful only when the status code is 0b111.

[0259] Based on the parameter binding validity status code, filter out interface method instances without type conflicts and generate a scenario-independent intermediate indicator sequence;

[0260] The intermediate indicator sequence is filled into the indicator template according to the semantic rules of the target scenario to generate unambiguous indicator conversion results.

[0261] In one feasible implementation, template filling is performed based on target indicator transformation. The intermediate indicator sequence, validated for effectiveness, is recombined according to the template specifications of the target scenario; specifically, the clinical scenario uses an HL7 standard XML template, while the educational scenario uses the JSONSchema of the IMS global learning tool. Scenario-specific evaluation optimization parameters are injected during the filling process; the clinical scenario adds a sensitive patient label field, and the educational scenario supplements learning behavior analysis labels. A timestamp chain is embedded in the template header as an auxiliary field, recording key moments in state evolution. The exception context stack is compressed and stored in the template extension area, preserving a complete exception handling path fingerprint. The final output evaluation result description code is a binary encoded stream, containing a protocol header (scenario identifier + version number), payload (template content), and checksum (CRC32 code). The entire transformation process is transparently completed at the output layer of the evaluation model, with the average adapter processing latency controlled within 20 milliseconds.

[0262] Transform psychological state assessment elements into underlying assessment instructions supported by the target scenario;

[0263] The temporal dependencies of state labels and the exception handling context are preserved during the transformation process;

[0264] Inject the evaluation and optimization parameters that match the target scenario, and generate an evaluation result descriptor that meets the cross-scenario execution conditions.

[0265] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0266] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent dynamic assessment of adolescent psychological state based on multimodal data, characterized in that, include: S1. Obtain multimodal psychological data of the target assessment object and extract dynamic features from the multimodal psychological data; wherein, the multimodal psychological data includes voice emotion information and facial expression behavior relationship; S2. The trained machine learning model is used to perform feature fusion processing on the multimodal psychological data to generate a multimodal feature set; wherein, the multimodal feature set includes a mapping relationship between modality type identifiers and temporal codes; S3. Based on the multimodal feature set, a state analysis network is used to perform temporal correlation analysis on the dynamic features to generate a feature fusion result that includes emotional dimension features and cognitive pattern features; S4. Input the feature fusion result into the evaluation model for dynamic state deduction processing, and output psychological state evaluation elements, including: S41. Analyze the contextual relationships of the current evaluation scenario and generate state priority weights associated with the multimodal feature fusion results; S42. Generate a state transition matrix based on the state priority weights, and traverse and sort the modality type identifiers in the multimodal feature set; S43. Using a dynamic programming strategy, select the optimal evolution sequence from the state transition matrix to generate a basic evaluation set containing emotion labels, cognitive states, and trend predictions; S44. Logically integrate the basic evaluation set and trend prediction parameters to generate a state description segment that conforms to the target evaluation framework. The psychological state assessment elements include state labels and trend prediction parameters generated based on feature evolution sequences.

2. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 1, characterized in that, S2 employs a trained machine learning model to perform feature fusion processing on the multimodal psychological data, generating a multimodal feature set, including: S21. Perform modality segmentation processing on the multimodal psychological data to generate multiple data domains and corresponding initial feature parameters; S22. Extract the temporal feature sequence of each data domain, and perform attention weighting processing on the temporal feature sequence to generate a weighted multi-domain feature vector; S23. Calculate the similarity between the multi-domain feature vector and the preset modality association threshold, and filter out candidate features that meet the confidence conditions. S24. Redundancy elimination processing is performed on the candidate features to generate a multimodal feature set containing modality type identifiers. Each multimodal feature contains a normalized temporal code relative to the multimodal psychological data. Specifically, based on the multimodal feature set, a state analysis network is used to perform temporal correlation analysis on the dynamic features, generating a feature fusion result that includes emotional dimension features and cognitive pattern features, including: Based on the normalized temporal coding in the multimodal feature set, the multimodal psychological data is subjected to time window truncation processing to generate multiple feature time segments; Each feature time segment is standardized and feature-enhanced to generate optimized state analysis input data. The recurrent memory subnet in the state analysis network is used to extract sequence features from the state analysis input data to generate an emotion dimension feature vector. In parallel, a pattern classification model is used to perform feature pattern matching on the feature time segments to generate a cognitive pattern probability distribution. The emotion dimension feature vector and the cognitive pattern probability distribution are concatenated to generate a multimodal feature fusion result.

3. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 1, characterized in that, The method further includes: During the evaluation process, status response data is captured in real time to generate feedback logs containing feature offsets and status change information. Extract the abnormal state features from the feedback log, and perform similarity matching between the abnormal state features and the historical evaluation case library to generate adaptive adjustment instructions; The modality association threshold parameters of the machine learning model are dynamically updated based on the adaptive adjustment instructions. The updated modal association threshold parameters are injected into the evaluation model, and the state priority weights in the state transition matrix are recalculated.

4. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 3, characterized in that, The step of extracting abnormal state features from the feedback log and performing similarity matching between the abnormal state features and the historical evaluation case library to generate adaptive adjustment instructions includes: The abnormal state features are divided into time windows to generate multiple state segments and their corresponding multimodal data sequences; The trained anomaly classification model is used to perform root cause analysis on each state segment, generating classification labels that include feature recognition bias, temporal drift, and response delay. Retrieve solution templates that match the classification labels from the historical evaluation case library to generate a set of candidate adjustment strategies; Based on the matching degree ranking between the multimodal data sequence and the candidate adjustment strategy set, the strategy with the highest confidence is selected to generate an adaptive adjustment instruction.

5. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 4, characterized in that, The trained anomaly classification model is used to perform root cause analysis on each state segment, generating classification labels that include feature recognition bias, temporal drift, and response delay, including: The state segments are processed by time series segmentation to generate multimodal data sequence slices containing start and end times; The multimodal data sequence slices are subjected to keyframe sampling processing to generate a set of state change keyframes and corresponding timestamp indices; Extract the feature offset between adjacent keyframes in the set of keyframes for state changes, and generate feature evolution trajectory vectors and time interval sequences; The feature evolution trajectory vector is processed by trend calculation to generate an abnormal evolution pattern feature vector; The feature vector of the abnormal evolution pattern is input into the spatiotemporal convolutional subnet of the trained abnormal classification model for local feature extraction, generating a spatial abnormal activation map. The time interval sequence is subjected to sliding window mean filtering to generate a smoothed response delay time sequence; The response delay time series is input into the gated recurrent subnet of the trained anomaly classification model for periodic pattern matching to generate a time dimension anomaly score; the spatial anomaly activation map is subjected to channel max pooling to generate a spatial dimension anomaly score; the spatial dimension anomaly score and the time dimension anomaly score are subjected to feature cross-fusion processing to generate a spatiotemporal joint anomaly probability distribution. Peak detection results of feature recognition deviation probability, temporal drift probability, and response delay probability are extracted from the spatiotemporal joint anomaly probability distribution; dynamic threshold comparison processing is performed on the preset probability threshold based on the peak detection results to generate a candidate set of classification labels containing probability ranking; The candidate classification labels are subjected to timestamp index alignment verification to generate target classification labels that are consistent with the phase of the abnormal state in the multimodal data sequence slice; the target classification labels and the feature evolution trajectory vector are subjected to anomaly type reverse verification to generate a final classification label set containing confidence weights.

6. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 1, characterized in that, The method further includes: Virtual feature perturbation parameters are injected before the evaluation model is executed. These virtual feature perturbation parameters are used to simulate the random offset scenario of multimodal data features. Monitor the fault-tolerant processing results of the evaluation model on the perturbed features and generate stability evaluation indicators; When the stability evaluation index is lower than a preset threshold, the online learning mode of the completed machine learning model is triggered. The temporal feature parameters of the trained machine learning model are updated by gradient backpropagation based on the feature difference data before and after the perturbation. The evaluation model monitors its fault-tolerant processing results for the perturbed features and generates stability evaluation indices, including: The number of evaluation failures and the anomaly capture success rate caused by feature errors during the statistical evaluation process are recorded. Calculate the ratio of the number of evaluation failures to the total number of evaluation steps to generate a first evaluation coefficient; Extract the proportion of events that match the preset fault tolerance rules from the anomaly capture success rate, and generate a second evaluation coefficient; The first evaluation coefficient and the second evaluation coefficient are weighted and summed to generate a comprehensive stability evaluation index.

7. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 6, characterized in that, The step of updating the temporal feature parameters of the trained machine learning model through gradient backpropagation based on the feature difference data before and after the perturbation includes: The feature difference data before and after the perturbation are time-stamp aligned to generate a matching pair sequence containing the feature set before the perturbation and the feature set after the perturbation; Extract the feature offsets from the matching pair sequence to generate the horizontal and vertical feature vectors for each multimodal data. The distance between the horizontal feature vector and the vertical feature vector is calculated to generate a set of feature difference vectors for each multimodal data. A feature regression loss function is constructed based on the set of feature difference vectors. The feature regression loss function includes the mean square error between the features predicted by the trained machine learning model and the actual features after perturbation. The feature regression loss function is subjected to a differentiable transformation to generate the loss value tensor required for gradient backpropagation; Iterate through the temporal feature parameters of the trained machine learning model, calculate the partial derivative of the loss tensor with respect to each temporal feature, and generate the temporal feature gradient matrix. The temporal feature gradient matrix is ​​subjected to momentum smoothing to generate a decayed parameter update direction vector; Based on the updated direction vector and the preset learning rate parameter, the temporal feature parameters of the trained machine learning model are iteratively and incrementally adjusted. During the incremental adjustment process, the rate of change of the feature regression loss function is monitored in real time to generate a loss convergence status indicator. When the loss convergence state indicator reaches a stable threshold, the gradient backpropagation update is terminated and the temporal feature parameters are frozen. The updated temporal feature parameters are injected into the forward propagation path of the trained machine learning model, and the modality segmentation and feature fusion processes are re-executed to verify the improved feature extraction accuracy.

8. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 1, characterized in that, The method further includes: Construct a cross-scenario evaluation adapter, and use the cross-scenario evaluation adapter to parse the semantic differences of metrics in different evaluation frameworks; The psychological state assessment elements are converted into underlying assessment instructions supported by the target scenario; The temporal dependencies and exception handling context of the state labels are preserved during the transformation process; Inject the evaluation optimization parameters that match the target scenario to generate an evaluation result descriptor that meets the cross-scenario execution conditions.

9. The intelligent dynamic assessment method for adolescent psychological state based on multimodal data as described in claim 8, characterized in that, The construction of a cross-scenario evaluation adapter, and the parsing of semantic differences in metrics across different evaluation frameworks using the cross-scenario evaluation adapter, includes: Establish a semantic rule base for the evaluation framework, and store the indicator mapping tables and parameter transmission paths for each scenario; Abstract semantic tree parsing is performed on the psychological state assessment elements to generate intermediate representation layer data; Based on the intermediate representation layer data, the indicator mapping table is traversed and queried to generate a scenario-compatible indicator replacement scheme; Dependency injection is performed on conflicting parameter passing paths to generate unambiguous index transformation results; The process of performing dependency injection on conflicting parameter passing paths to generate unambiguous index transformation results includes: Identify conflicting nodes in the parameter passing path that have type ambiguity or overlapping scope, and generate a set of conflicting path identifiers; Extract the context metadata of each node in the set of conflict path identifiers. The context metadata includes parameter type declarations, lifecycle markers, and call stack fingerprints. Based on the context metadata, an interface proxy class name corresponding to the conflict node is generated. The interface proxy class name inherits the native interface of the target scenario and injects type casting constraints. Parse the dependency chain of the interface proxy class name in the target scenario and generate a dynamic binding strategy that includes parameter scope isolation boundaries; Based on the dynamic binding strategy, conflicting nodes in the parameter passing path are mapped to scenario-compatible interface method instances, generating a dynamic binding queue of parameter instances. Traverse the interface method instances in the dynamic binding queue, verify the compatibility status of the interface method instances with the target scenario indicator semantics, and generate parameter binding validity status codes; Based on the parameter binding validity status code, filter out interface method instances without type conflicts and generate a scenario-independent intermediate indicator sequence; The intermediate indicator sequence is filled into the indicator template according to the semantic rules of the target scenario to generate an unambiguous indicator conversion result.

Citation Information

Patent Citations

  • Dredging auxiliary platform based on psychological health education

    CN120299634A