Stress level evaluation and training guidance system fused with machine learning

By combining multimodal sensors and dynamic machine learning models, the problems of low accuracy and poor adaptability in traditional stress level assessment techniques have been solved. This enables high-precision quantitative assessment and personalized intervention in complex scenarios, improving assessment accuracy and intervention effectiveness.

CN122163967APending Publication Date: 2026-06-09ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ARMY ENG UNIV OF PLA
Filing Date
2026-01-29
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Traditional stress level assessment techniques suffer from low accuracy, weak anti-interference ability, and poor model adaptability, making it difficult to achieve accurate quantitative judgment and personalized intervention in complex scenarios.

Method used

Physiological signal data is collected in real time using multimodal sensors. Through cross-modal feature fusion and dynamic machine learning models, combined with transfer learning and incremental learning techniques, model parameters are dynamically optimized to generate personalized training guidance schemes.

Benefits of technology

It enables high-precision quantitative assessment and personalized intervention of stress levels in complex scenarios, improving the accuracy of assessment and the pertinence of intervention, and adapting to changes in different scenarios and user states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122163967A_ABST
    Figure CN122163967A_ABST
Patent Text Reader

Abstract

The application provides a stress level evaluation and training guidance system combined with machine learning, which comprises a data acquisition module, a feature extraction and fusion module, a machine learning model training module, a real-time stress level evaluation module, a model dynamic adjustment module and a personalized training guidance generation module; the data acquisition module acquires physiological signal data of a user in different scenes in real time, the feature extraction and fusion module integrates features of different signals through a cross-modal feature fusion algorithm to form a unified feature vector, the machine learning model training module trains a benchmark machine learning model and a dynamic machine learning model based on a historical scene data set and a current scene data set respectively; the real-time stress level evaluation module inputs the integrated feature vector into the dynamic machine learning model, outputs a stress level score in the current scene, and compares the score with a preset stress threshold to determine whether the user is in a stress state; the model dynamic adjustment module automatically adjusts feature weights of the dynamic model and standardization parameters of input data according to new data and environmental noise levels of the user in a real scene after each evaluation, thereby improving robustness of the model in a non-standard scene; and the personalized training guidance generation module generates a targeted training guidance scheme according to an evaluation result and historical behavior data of the user. The application realizes high-precision quantitative evaluation of stress levels in combination with a dynamically optimized machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biological signal processing technology, specifically relating to a stress level assessment and training guidance system that integrates machine learning. Background Technology

[0002] Stress level assessment and training guidance is a systematic approach that uses physiological indicators (such as electrocardiogram and skin conductance signals) combined with psychological questionnaires to quantitatively assess the intensity of an individual's physical and mental response to stress situations. It also provides personalized intervention programs based on methods such as cognitive behavioral therapy, mindfulness training, and progressive exposure.

[0003] Traditional stress level assessment and intervention techniques primarily employ two methods: First, methods based on single physiological signals, relying heavily on wearable devices to collect individual signals such as heart rate and skin conductance, assessing stress levels through simple threshold judgments or basic statistical analysis, lacking collaborative analysis of multimodal signals; second, methods based on subjective reports, where users subjectively fill out stress perception questionnaires, and professionals then develop fixed intervention plans based on the questionnaire results. At the model application level, traditional techniques often use machine learning models with fixed parameters, trained on historical data in controlled scenarios and then deployed immediately, without considering environmental noise and data drift in real-world scenarios, and lacking mechanisms for dynamic adjustment of model parameters.

[0004] Traditional stress level assessment and intervention techniques suffer from low accuracy and weak resistance to interference. They rely heavily on single physiological signals or user subjective reports, lacking collaborative analysis of multi-dimensional physiological signals and failing to establish signal optimization mechanisms for environmental noise and individual differences. As a result, signals are easily interfered with in complex scenarios such as changes in lighting and background noise, and they cannot eliminate assessment biases caused by differences in basic physiological indicators among different users, making it difficult to achieve accurate quantitative judgment of users' stress levels.

[0005] Secondly, traditional models have poor adaptability and struggle to cope with dynamic changes in real-world scenarios. Many of these models employ fixed parameter designs, being trained only on datasets from initial controlled scenarios before being deployed. They fail to consider changes in user physiological characteristics over time, data distribution shifts caused by environmental fluctuations, and lack dynamic optimization mechanisms for model parameters. When data features drift, the accuracy of model evaluation results significantly decreases, making it unable to stably adapt to different scenarios and user states. Summary of the Invention

[0006] The purpose of this invention is to provide a stress level assessment and training guidance system that integrates machine learning, which solves the problems of low assessment accuracy, weak anti-interference ability and poor model adaptability in traditional stress level assessment technology, and achieves high-precision quantitative assessment of stress level.

[0007] The technical solution to achieve the purpose of this invention is as follows:

[0008] A stress level assessment and training guidance system integrating machine learning achieves high-precision quantitative assessment of stress levels. It includes a data acquisition module, a feature extraction and fusion module, a machine learning model training module, a real-time stress level assessment module, a model dynamic adjustment module, and a personalized training guidance generation module. The data acquisition module acquires real-time physiological signal data of users in different scenarios through multimodal sensors. This physiological signal data includes heart rate, skin conductance response, facial micro-expressions, and voice tone. The feature extraction and fusion module preprocesses the acquired multimodal physiological signals, removing noise and extracting features. Then, it integrates the features of different signals using a cross-modal feature fusion algorithm to form a unified feature vector. The machine learning model training module trains a baseline machine learning model and a dynamic machine learning model based on historical scenario datasets and current scenario datasets, respectively. The baseline model is used to identify the core stress level under controlled scenarios. The dynamic model, using transfer learning and incremental learning techniques, optimizes the baseline model in real-world scenarios to adapt to data drift. The real-time stress level assessment module inputs the integrated feature vectors into the dynamic machine learning model, outputs a stress level score for the current scenario, and compares the score with a preset stress threshold to determine if the user is under stress. After each assessment, the model dynamic adjustment module automatically adjusts the feature weights and standardized parameters of the input data based on new user data and environmental noise levels in real-world scenarios to improve the model's robustness in non-standardized environments. The personalized training guidance generation module generates targeted training guidance plans based on assessment results and the user's historical behavioral data, including breathing regulation training, cognitive reconstruction training, and progressive exposure training. These plans dynamically adjust training intensity and content to ensure users receive the most effective interventions at different stress levels.

[0009] Furthermore, the data acquisition module includes a multimodal sensor group, which comprises a wearable heart rate monitoring device, a skin conductance sensor, a facial micro-expression camera, and a voice acquisition microphone. The wearable heart rate monitoring device acquires the user's heart rate data in real time through a photoelectric sensor or an electrocardiogram sensor. The skin conductance sensor measures the conductivity of the user's skin through skin surface electrodes to reflect sympathetic nerve activity. The facial micro-expression camera acquires the user's facial muscle movement data through high frame rate video and performs micro-expression recognition using optical flow and a facial action coding system (FACS). The voice acquisition microphone acquires the user's speech intonation data through audio signals and extracts emotional features of the speech using Mel-frequency cepstral coefficients (MFCC). The data acquisition process of the sensor group is performed synchronously to ensure time alignment of the multimodal data.

[0010] Furthermore, the feature extraction and fusion module includes a preprocessing unit and a cross-modal feature fusion unit. The preprocessing unit normalizes the collected raw data and fills in missing values ​​using a data preprocessing algorithm, thereby providing high-quality input for subsequent feature extraction. The cross-modal feature fusion unit adopts a cross-modal feature fusion algorithm based on attention mechanism and multi-scale feature extraction technology. First, the multi-scale feature extraction technology performs hierarchical processing on the features of different physiological signals. Then, the cross-modal feature fusion algorithm fuses the feature vectors of each modality in a unified high-dimensional feature space through adaptive attention weight allocation. The attention weight is dynamically adjusted according to the environmental noise level of the current scene to enhance the contribution of key modal features, thereby forming a more robust fused feature vector for subsequent model training and evaluation.

[0011] Furthermore, the multi-scale feature extraction technology performs hierarchical processing of features of different physiological signals, including: extracting time-domain features (such as heart rate variability) and frequency-domain features (such as heart rate spectrum) from heart rate data; extracting time-frequency features (such as fundamental frequency and energy distribution) and semantic features (such as emotional keywords) from speech tone data; and extracting spatial features (such as facial action unit combinations) and time-series features (such as expression duration) from facial micro-expression data.

[0012] Furthermore, the machine learning model training module includes a baseline model training unit and a dynamic model optimization unit. The baseline model training unit uses historical scene datasets collected in controlled scenarios to train a baseline machine learning model using deep neural networks (DNN) or support vector machines (SVM). The goal of the baseline model is to identify the core features of a user's stress level in a standard environment. The dynamic model optimization unit, in a real-world scenario, transfers the knowledge of the baseline model to the dynamic model using transfer learning techniques based on the current scene dataset and the initial parameters of the baseline model. It then continuously optimizes the dynamic model using an incremental learning algorithm. The incremental learning algorithm employs an online learning strategy, updating the model parameters only with newly added real-world scene data after each evaluation. This avoids overfitting due to repeated training with a large amount of historical data, thereby improving the adaptability and accuracy of the dynamic model in non-standard scenarios.

[0013] Furthermore, the formula for updating the model parameters using the incremental learning algorithm is as follows:

[0014] ;

[0015] in: Indicates after the first The model parameters of the dynamic model at the next time step (t+1) after the parameter update.

[0016] This represents the model parameters of the dynamic model at the current moment. It is a loss function Regarding the model parameters at the current moment gradient, This is the dynamic loss function.

[0017] Furthermore, the dynamic loss function Huber's losses are:

[0018] ;

[0019] This represents the dynamic loss function, with the actual stress level as the input. and the stress level predicted by the model The output is the loss between the model's prediction and the actual value. It is the true level of stress. It is the predicted stress level of the current scenario by a dynamic machine learning model. It is the absolute error between the actual value and the predicted value, measuring the degree of deviation in the prediction. This is the smoothing threshold constant.

[0020] Furthermore, the real-time stress level assessment module includes a dynamic score generation unit and a stress state determination unit. The dynamic score generation unit inputs the fused feature vector into a dynamic machine learning model and calculates the stress level score for the current scenario through the model's forward propagation. The score range is 0-100, where 0 represents no stress and 100 represents extreme stress. The stress state determination unit compares the score with preset stress thresholds, which include mild stress thresholds, moderate stress thresholds, and severe stress thresholds. When the score exceeds any threshold, the user is determined to be in the corresponding stress state, and the response process of the model dynamic adjustment module and the personalized training guidance generation module is triggered, thereby realizing immediate feedback of the assessment results and dynamic planning of subsequent intervention measures.

[0021] Furthermore, the model dynamic adjustment module includes a feature weight update unit and an input normalization parameter adjustment unit. After each evaluation by the real-time stress level assessment module, the feature weight update unit calculates the dynamic weight coefficients of each modality feature based on the signal-to-noise ratio and feature correlation of each modality data in the current scene, and feeds the weight coefficients back to the cross-modal feature fusion algorithm to adjust the contribution ratio of each modality feature in the fused feature vector. The input normalization parameter adjustment unit dynamically updates the normalization parameters of the input data based on the environmental noise level of the current scene and individual user differences.

[0022] Furthermore, based on the signal-to-noise ratio and feature correlation of each modality in the current scene, the dynamic weight coefficients of each modality feature are calculated as follows:

[0023] ;

[0024] in, , They represent the first species, first Attention weight coefficients corresponding to the feature vectors of each modality , They represent the first species, first The signal-to-noise ratio corresponding to the feature vectors of each modality; the attention weight coefficients The initial value is:

[0025] ;

[0026] in, , They are respectively with the first species, first Modal eigenvectors The relevant learnable weight vector, , The first species, first The single-modal feature vector is obtained after feature extraction of various physiological signals.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] (1) The system achieves high-precision quantitative assessment of stress levels through multimodal data acquisition and processing technology combined with a dynamically optimized machine learning model. During the data acquisition phase, the multimodal sensor array simultaneously acquires heart rate, electrical skin response (EDA), facial micro-expressions, and speech tone data. A time alignment formula ensures data temporal consistency, and a normalization formula eliminates individual differences and scale shifts. Missing data is then processed using… Completion provides high-quality input for feature extraction. In the feature fusion stage, weights are dynamically allocated based on an attention mechanism. To enhance the contribution of high signal-to-noise ratio modes, the dynamic model employs an incremental learning formula during model training. The parameters are updated, and the robustness and stability are balanced by combining the Huber loss function. Finally, a quantitative score of 0-100 is output. The score is accurately classified by three levels of thresholds (mild, moderate, and severe), which effectively reduces the impact of environmental noise and data drift on the evaluation results and improves the accuracy and reliability of the evaluation in different scenarios.

[0029] (2) The system achieves dynamic adaptation and precise implementation of stress intervention programs based on multi-dimensional parameter optimization and personalized data-driven approaches. The personalized training guidance generation module generates training programs such as breathing regulation and cognitive reconstruction based on real-time stress scores and historical behavioral data, and the program adjustment is deeply linked with the model's dynamic optimization parameters. The model dynamic adjustment module calculates dynamic weight coefficients based on the signal-to-noise ratio (SNR) of each modality through the feature weight update unit and feeds them back to the cross-modal fusion algorithm to adjust the feature contribution ratio; the input standardized parameter adjustment unit dynamically updates parameters such as the normalized range of heart rate and the fundamental frequency calibration coefficient of speech based on environmental noise and individual differences to ensure the consistency of model input. For example, when feature distribution drift is detected, the difference is measured by KL divergence; if it exceeds the threshold Then through Adjusting model parameters ensures stable evaluation results. This parameterized dynamic adjustment mechanism allows training guidance plans to be flexibly adapted to users' real-time stress levels and individual characteristics, ensuring that users with different stress levels and individuals can obtain the optimal intervention plan, significantly improving the pertinence and effectiveness of stress intervention. Attached Figure Description

[0030] Figure 1 A framework diagram for a stress level assessment and training guidance system that integrates machine learning.

[0031] Figure 2 Workflow view for the machine learning model training module.

[0032] Figure 3 This is a timeline diagram for a stress level assessment and training guidance system. Detailed Implementation

[0033] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0034] Combination Figure 1 and Figure 3This embodiment provides a stress level assessment and training guidance system integrating machine learning, including a data acquisition module, a feature extraction and fusion module, a machine learning model training module, a real-time stress level assessment module, a model dynamic adjustment module, and a personalized training guidance generation module. The data acquisition module acquires real-time physiological signal data of users in different scenarios through multimodal sensors. This physiological signal data includes heart rate, skin conductance, facial micro-expressions, and voice tone. The feature extraction and fusion module preprocesses the acquired multimodal physiological signals, removing noise and extracting features. Then, it integrates the features of different signals using a cross-modal feature fusion algorithm to form a unified feature vector. The machine learning model training module trains a baseline machine learning model and a dynamic machine learning model based on historical scenario datasets and current scenario datasets, respectively. The baseline model is used to identify the core features of stress level in controlled scenarios. The dynamic model is used to optimize the baseline model in real-world scenarios using transfer learning and incremental learning techniques to adapt to data drift. The real-time stress level assessment module inputs the integrated feature vector into the dynamic machine learning model, outputs a stress level score for the current scenario, and compares the score with a preset stress threshold to determine whether the user is in a stress state. After each assessment, the model dynamic adjustment module automatically adjusts the feature weights of the dynamic model and the standardized parameters of the input data based on new data from the user in the real-world scenario and the level of environmental noise, to improve the robustness of the model in non-standard scenarios. The personalized training guidance generation module generates targeted training guidance plans based on the assessment results and the user's historical behavioral data, including breathing regulation training, cognitive reconstruction training, and progressive exposure training. These plans dynamically adjust the training intensity and content to ensure that users receive the most effective interventions at different stress levels.

[0035] The data acquisition module further includes a multimodal sensor group, which comprises a wearable heart rate monitoring device, a skin conductance sensor, a facial micro-expression camera, and a voice acquisition microphone. The wearable heart rate monitoring device acquires the user's heart rate data in real time through a photoelectric sensor or an electrocardiogram sensor. The skin conductance sensor measures the conductivity of the user's skin through electrodes on the skin surface to reflect sympathetic nerve activity. The facial micro-expression camera acquires the user's facial muscle movement data through high frame rate video and performs micro-expression recognition using optical flow and a facial action coding system (FACS). The voice acquisition microphone acquires the user's speech intonation data through audio signals and extracts emotional features of the speech using Mel-frequency cepstral coefficients (MFCC). The data acquisition process of the sensor group is performed synchronously to ensure time alignment of the multimodal data. The acquired raw data is normalized and missing values ​​are filled in through data preprocessing algorithms, thereby providing high-quality input for subsequent feature extraction.

[0036] The feature extraction and fusion module includes a preprocessing unit and a cross-modal feature fusion unit. The preprocessing unit is used to clarify the time synchronization, normalization and noise suppression processes of each physiological signal to ensure the stability and consistency of the input data.

[0037] In time At any given time, the signal acquired from the multimodal sensor array is represented as:

[0038]

[0039] in:

[0040] This indicates the heart rate signal.

[0041] This indicates the electrodermal activity signal.

[0042] Indicates facial micro-expression features;

[0043] It represents the signal of speech intonation.

[0044] The acquired modal signals were normalized and time-aligned.

[0045]

[0046]

[0047] in and The first The mean and standard deviation of the modal signal within the time window are used to eliminate scale shifts caused by individual differences.

[0048] When data is missing, the system uses a time series interpolation algorithm to complete the signal:

[0049]

[0050]

[0051] in, Indicates time At that moment, the The signal values ​​obtained after interpolation of the physiological signals of a certain modality are used to fill in the missing data of the signal of that modality at that time point to ensure signal continuity; Indicates time The previous moment (i.e. (moment), the The actual acquired values ​​of various modal physiological signals are one of the core historical reference data for interpolation calculations; This indicates that at the two moments before time t (i.e., time t−2), the first... The actual acquired values ​​of the physiological signals of each modality, together with xt−1(i), constitute the historical data basis for interpolation calculation; These represent specific types of multimodal physiological signals, corresponding to heart rate (HR), electrical skin response (EDA), facial micro-expression feature (F), and voice tone (V). Represents the smoothing coefficient, used for balancing. Time signal value and The contribution weight of the signal value at time step in the interpolation calculation. The larger the value, the better. The greater the influence of the signal value at any given time on the completion result, the better it can match the recent trend of the signal.

[0052] Through the above parameterization process, the system achieves time alignment, scale uniformity, and noise suppression of multimodal signals, providing high-quality input for subsequent feature extraction and cross-modal feature fusion.

[0053] The cross-modal feature fusion unit adopts a cross-modal feature fusion algorithm based on attention mechanism and multi-scale feature extraction technology. The multi-scale feature extraction technology performs hierarchical processing on the features of different physiological signals. In order to make the feature extraction and fusion process quantifiable and feasible, the system has mathematically defined and parameterized the feature extraction and fusion of different modal physiological signals to ensure the effective expression and robust fusion of multimodal features in a unified high-dimensional space.

[0054] The specific details of heart rate signal feature extraction are as follows:

[0055] heart rate signal Calculate its time-domain and frequency-domain characteristics, where the time-domain characteristics are represented using heart rate variability (HRV):

[0056]

[0057] Frequency domain characteristics were obtained by power spectrum analysis, yielding the ratio of low-frequency to high-frequency components:

[0058]

[0059] in, The interval between adjacent heartbeats, and These represent the low-frequency and high-frequency power spectral densities, respectively.

[0060] The specific extraction of skin conductance response features is as follows:

[0061] Electrodermal signals It is broken down into slowly changing skin potential (SCL) and rapidly responding components (SCR):

[0062]

[0063] in, Indicates the first The magnitude of the reaction, It is the signal decay time constant, used to characterize the activity level of the sympathetic nervous system.

[0064] The specific steps for extracting facial micro-expression features are as follows:

[0065] Facial video signals are processed using optical flow and a Facial Action Coding System (FACS) to extract Action Units (AUs), which form facial muscle motion feature vectors.

[0066]

[0067] in, Indicates the first The activation amplitude of each action unit The number of action units detected.

[0068] The specific extraction of speech intonation features is as follows:

[0069] For speech signals Calculate Mel frequency cepstral coefficients (MFCC) to extract sentiment-related time-frequency features:

[0070]

[0071] in, The power spectrum of the speech signal at the 1st The output of the Mel filter This represents the number of cepstral coefficients.

[0072] The system employs an attention-weighted feature fusion approach to adaptively integrate different modalities within a unified feature space. Let the feature vectors of different modalities be... , indicating the first The single-modal feature vectors obtained after feature extraction from various modal physiological signals, such as the time-domain / frequency-domain feature vectors of heart rate signals and the slow potential / fast response feature vectors of skin conductance, are the basic input units for feature fusion. Therefore, the fused feature vector is:

[0073]

[0074] Attention weight The calculation formula is:

[0075]

[0076] Indicates the first The attention weight coefficients corresponding to the modal feature vectors are used to dynamically adjust the contribution of the modal feature in the cross-modal feature fusion process;

[0077] In order to be with the first Modal eigenvectors The relevant learnable weight vectors, obtained through model training, are used to measure... The importance of internal features and other related characteristics;

[0078] For the first The single-modal feature vector obtained after feature extraction of various modal physiological signals;

[0079] In order to be with the first Modal eigenvectors The relevant learnable weight vectors, their effects and Similarly, operations related to weight calculation for feature vectors of different modalities.

[0080] Represents the modality for all designs A summation operation is performed for normalization calculations, ensuring that the sum of the attention weights for all modalities is 1, thus maintaining the rationality of the weight allocation.

[0081] When the external noise level is high, the system dynamically adjusts the attention weights based on the signal-to-noise ratio (SNR):

[0082]

[0083] This enhances the contribution of higher-quality signal modes and improves fusion characteristics. The robustness and expressiveness of the final output fused feature vector As input to the machine learning model training and real-time evaluation module, it satisfies:

[0084]

[0085] in This represents the fusion mapping function, which is jointly modeled by a multi-scale feature extraction layer and an attention mechanism.

[0086] Combination Figure 2The machine learning model training module includes a baseline model training unit and a dynamic model optimization unit. The baseline model training unit uses historical scene datasets collected in controlled scenarios to train a baseline machine learning model using deep neural networks (DNN) or support vector machines (SVM). The goal of the baseline model is to identify the core features of a user's stress level in a standard environment. The dynamic model optimization unit, in a real-world scenario, uses transfer learning technology to transfer knowledge from the baseline model to the dynamic model based on the current scene dataset and the initial parameters of the baseline model. It then uses an incremental learning algorithm to continuously optimize the dynamic model. The incremental learning algorithm employs an online learning strategy, updating the model parameters only with newly added real-world scene data after each evaluation. This avoids overfitting caused by repeated training with a large amount of historical data, thereby improving the adaptability and accuracy of the dynamic model in non-standard scenarios.

[0087] In a controlled scenario, let the fusion feature input be... The corresponding stress level label is The training objective of the baseline model is:

[0088]

[0089] in For deep neural network (DNN) or support vector machine (SVM) models, These are the parameters for the baseline model.

[0090] The baseline model uses mean squared error (MSE) as the loss function:

[0091]

[0092] Minimize using the backpropagation algorithm To obtain the optimal model parameters under standard conditions:

[0093]

[0094] The parameters representing the baseline model are the set of parameters that are ultimately determined through subsequent optimization processes and minimize the loss of the baseline model.

[0095] It involves minimizing the parameters to find the loss function that makes the subsequent steps easier. The parameter corresponding to the minimum value.

[0096] These are the parameter variables to be adjusted during the optimization process. By adjusting them, we find the parameters that minimize the loss function Lbase, and ultimately, this optimal value is achieved. The parameters determined as the baseline model .

[0097] The loss function represents the baseline model and is used to measure the error between the baseline model's predictions and the actual results on the training data. The goal of optimization is to minimize the value of this loss function.

[0098] In real-world scenarios, the system introduces a dynamic model. Its parameters are derived from the baseline model parameters:

[0099]

[0100] Newly collected data in the current scenario An online incremental learning strategy is used for parameter updates:

[0101]

[0102] in: Indicates after the first After the parameter update, the model parameters of the dynamic model at the next time step (t+1 time); This represents the model parameters of the dynamic model at the current moment, and is the basis for this parameter update; The learning rate controls the step size of each parameter update. A larger learning rate can result in a larger range of parameter adjustments, while a smaller learning rate allows for more precise parameter adjustments. An appropriate value needs to be selected based on the training situation to ensure that the model can learn effectively without failing to converge due to step size issues. It is a loss function Regarding the model parameters at the current moment The gradient reflects the trend of the loss function at the current parameter point. Updating the parameters in the opposite direction of the gradient can minimize the loss function and bring the model closer to the optimal state. It is the fusion feature input of newly collected data in the current scenario. (These are the corresponding real tags) The dynamic model loss is calculated based on this set of inputs and outputs and is used to measure the difference between the model's predictions and the actual results. The dynamic loss function is the Huber loss, which balances stability and robustness.

[0103]

[0104] This represents the dynamic loss function, with the actual stress level as the input. and the stress level predicted by the model The output is the loss between the model's prediction and the actual value, i.e., a quantification of the degree of error. It is the true level of stress. It is the predicted stress level of the current scenario by a dynamic machine learning model. It is the absolute error between the actual value and the predicted value, measuring the degree of deviation in the prediction. The smoothing threshold constant is a core hyperparameter of Huber loss. It determines the critical point at which the loss function switches from "mean squared error-like" (small error scenario) to "absolute error-like" (large error scenario). When the error becomes absolute... At that time, the loss item is taken This part, similar to mean squared error, allows the loss under small errors to increase with the square of the error, making the model more refined for small biases. When the absolute error... At that time, the loss item is taken This part is similar to absolute error, which allows the loss under large error to grow linearly with the error, avoiding the loss from expanding sharply due to large error (caused by outliers), thereby preventing the model parameters from being biased by abnormal data and ensuring robustness.

[0105] Huber loss through It achieves the effect of penalizing small errors with squared penalties and large errors with linear penalties, balancing the model's pursuit of accuracy and its tolerance to anomalies.

[0106] The model dynamic adjustment module includes a feature weight update unit and an input normalization parameter adjustment unit. After each evaluation, the feature weight update unit calculates the dynamic weight coefficients of each modality feature based on the signal-to-noise ratio and feature correlation of each modality data in the current scene, and feeds the weight coefficients back to the cross-modal feature fusion algorithm to adjust the contribution ratio of each modality feature in the fused feature vector. The input normalization parameter adjustment unit dynamically updates the normalization parameters of the input data based on the environmental noise level of the current scene (such as light intensity and background volume) and individual differences of users (such as age, gender, and basic physiological indicators). For example, it adjusts the normalization range of heart rate data and the fundamental frequency calibration coefficient of speech pitch data, thereby ensuring the consistency and stability of the evaluation results of the dynamic machine learning model in different scenarios and user groups.

[0107] The updated prediction output of the dynamic model is:

[0108]

[0109] The rating result This indicates the user's stress level in a real-world scenario;

[0110] in, This represents the stress level prediction result output by the dynamic model at the current time (the t-th related process or time point t), such as a stress level score, used to reflect the user's current stress state. The forward computation function represents the core operational logic of a dynamic machine learning model, which generates predictions based on input data. This function transforms the fused features of the input into predicted values ​​of the stress level. This represents the fused feature vector input into the dynamic model at the current moment. It is a unified feature representation containing multimodal physiological signal information, obtained after cross-modal feature fusion, providing a data foundation for model prediction. This represents the set of parameters after the dynamic model update. The parameters are obtained by the model through online incremental learning, optimizing and adjusting using new data, compared to the set before the update. It can better adapt to the data characteristics of the current scenario, thereby improving prediction accuracy.

[0111] After completing the parameter update, the dynamic model uses the fused features at the current moment. As input, through the model function And combined with the updated parameters Finally, the predicted stress level is obtained. The process demonstrates the characteristics of dynamic models in continuous learning and real-time prediction.

[0112] When a change in feature distribution is detected, i.e. At that time, the system updates the transfer weights using a feature distribution distance metric (such as Kullback-Leibler divergence):

[0113]

[0114] like Exceeding the threshold Then, the weight decay term of the model parameters is dynamically adjusted:

[0115]

[0116] in This is the drift compensation coefficient, used to balance stability and adaptability between the dynamic model and the baseline model.

[0117] The convergence condition of the model is defined as:

[0118]

[0119] in This is the convergence threshold. Once this condition is met, the model enters a stable operating phase and the latest... As the main model parameter of the evaluation module.

[0120] The real-time stress level assessment module includes a dynamic score generation unit and a stress state determination unit. The dynamic score generation unit inputs the fused feature vector into a dynamic machine learning model and calculates the stress level score for the current scenario through the model's forward propagation. The score ranges from 0 to 100, where 0 represents no stress and 100 represents extreme stress. The stress state determination unit compares the score with preset stress thresholds, which include mild stress thresholds, moderate stress thresholds, and severe stress thresholds. When the score exceeds any threshold, the user is determined to be in the corresponding stress state, triggering the response process of the model dynamic adjustment module and the personalized training guidance generation module, thereby achieving immediate feedback of the assessment results and dynamic planning of subsequent intervention measures.

[0121] The model dynamic adjustment module includes a feature weight update unit and an input normalization parameter adjustment unit. After each evaluation, the feature weight update unit calculates the dynamic weight coefficients of each modality feature based on the signal-to-noise ratio and feature correlation of each modality data in the current scene, and feeds the weight coefficients back to the cross-modal feature fusion algorithm to adjust the contribution ratio of each modality feature in the fused feature vector. The input normalization parameter adjustment unit dynamically updates the normalization parameters of the input data based on the environmental noise level of the current scene and individual user differences. For example, it adjusts the normalization range of heart rate data and the fundamental frequency calibration coefficient of speech intonation data, thereby ensuring the consistency and stability of the evaluation results of the dynamic machine learning model in different scenarios and user groups.

[0122] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A stress level assessment and training guidance system integrating machine learning, characterized in that, It includes a data acquisition module, a feature extraction and fusion module, a machine learning model training module, a real-time stress level assessment module, a model dynamic adjustment module, and a personalized training guidance generation module; among which: The data acquisition module acquires users' physiological signal data in different scenarios in real time through multimodal sensors; the physiological signal data includes heart rate, skin conductance, facial micro-expressions, and voice tone. The feature extraction and fusion module preprocesses the collected physiological signal data, removes noise and extracts features, and then integrates the features of different signals through a cross-modal feature fusion algorithm to form a unified feature vector. The machine learning model training module trains a baseline machine learning model and a dynamic machine learning model based on historical scene datasets and current scene datasets, respectively. The baseline machine learning model is used to identify the core features of stress level in a controlled scenario, while the dynamic machine learning model is used to optimize the baseline machine learning model in a real scenario through transfer learning and incremental learning techniques to adapt to data drift. The real-time stress level assessment module inputs the integrated feature vector into the dynamic machine learning model, outputs the stress level score in the current scenario, and compares the score with the preset stress threshold to determine whether the user is in a stress state. After each evaluation, the model dynamic adjustment module automatically adjusts the feature weights of the dynamic model and the standardized parameters of the input data based on the user's new data and environmental noise level in the real scenario, thereby improving the robustness of the model in non-standard scenarios. The personalized training guidance generation module generates targeted training guidance plans based on the evaluation results and the user's historical behavior data, including breathing regulation training, cognitive reconstruction training, and progressive exposure training. The training guidance plans dynamically adjust the training intensity and content to ensure that users receive the most effective intervention measures under different stress levels.

2. The stress level assessment and training guidance system integrating machine learning according to claim 1, characterized in that, The data acquisition module includes a multimodal sensor group, which comprises a wearable heart rate monitoring device, a skin conductance sensor, a facial micro-expression camera, and a voice acquisition microphone. The wearable heart rate monitoring device acquires the user's heart rate data in real time through a photoelectric sensor or an electrocardiogram sensor. The skin conductance sensor measures the conductivity of the user's skin through electrodes on the skin surface to reflect sympathetic nerve activity. The facial micro-expression camera acquires the user's facial muscle movement data through high frame rate video and performs micro-expression recognition using optical flow and a facial motion coding system. The voice acquisition microphone acquires the user's voice intonation data through audio signals and extracts emotional features of the voice using Mel-frequency cepstral coefficients. The data acquisition process of the multimodal sensor group is performed synchronously to ensure time alignment of the acquired data.

3. The stress level assessment and training guidance system integrating machine learning according to claim 1, characterized in that, The feature extraction and fusion module includes a preprocessing unit and a cross-modal feature fusion unit. The preprocessing unit normalizes the collected raw data and fills in missing values ​​using a data preprocessing algorithm. The cross-modal feature fusion unit integrates the features of different signals using a cross-modal feature fusion algorithm based on attention mechanism and multi-scale feature extraction technology. First, it extracts features of different physiological signals hierarchically using multi-scale feature extraction technology. Then, it fuses the feature vectors of each modality in a unified high-dimensional feature space through adaptive attention weight allocation.

4. The stress level assessment and training guidance system integrating machine learning according to claim 3, characterized in that, The hierarchical extraction of features for different physiological signals using multi-scale feature extraction technology includes: extracting time-domain and frequency-domain features from heart rate data, extracting time-frequency and semantic features from speech tone data, and extracting spatial and time-series features from facial micro-expression data.

5. The stress level assessment and training guidance system integrating machine learning according to claim 1, characterized in that, The machine learning model training module includes a baseline model training unit and a dynamic model optimization unit. The baseline model training unit uses a historical scene dataset collected in a controlled environment to train a baseline machine learning model using a deep neural network or support vector machine algorithm. The goal of the baseline machine learning model is to identify the core features of the user's stress level in a standard environment. In a real-world scenario, the dynamic model optimization unit, based on the current scenario dataset and the initial parameters of the benchmark machine learning model, transfers the knowledge of the benchmark machine learning model to the dynamic machine learning model using transfer learning technology. It then continuously optimizes the dynamic machine learning model using an incremental learning algorithm. This incremental learning algorithm employs an online learning strategy, updating the model parameters only with newly added real-world scenario data after each evaluation.

6. The stress level assessment and training guidance system integrating machine learning according to claim 5, characterized in that, The formula for updating model parameters using the incremental learning algorithm is as follows: ; in: Indicates after the first After the parameter update, the model parameters of the dynamic model at the next time step (t+1) are... This represents the model parameters of the dynamic model at the current moment. It is a loss function Regarding the model parameters at the current moment gradient, For dynamic loss function, This represents the learning rate.

7. The stress level assessment and training guidance system integrating machine learning according to claim 6, characterized in that, Dynamic loss function Huber's losses are: ; This represents the dynamic loss function, with the actual stress level as the input. and the stress level predicted by the model The output is the loss between the model's prediction and the actual value. It is the true level of stress. It is the predicted stress level of the current scenario by a dynamic machine learning model. This is the absolute error between the actual value and the predicted value. This is the smoothing threshold constant.

8. The stress level assessment and training guidance system integrating machine learning according to claim 1, characterized in that, The real-time stress level assessment module includes a dynamic score generation unit and a stress state determination unit. The dynamic score generation unit inputs the fused feature vector into a dynamic machine learning model and calculates the stress level score for the current scenario through forward propagation of the model. The score range is 0-100, where 0 represents no stress and 100 represents extreme stress. The stress state determination unit compares the score with preset stress thresholds, which include mild stress threshold, moderate stress threshold, and severe stress threshold. When the score exceeds any threshold, the user is determined to be in the corresponding stress state, and the response process of the model dynamic adjustment module and the personalized training guidance generation module is triggered to realize the immediate feedback of the assessment results and the dynamic planning of subsequent intervention measures.

9. The stress level assessment and training guidance system integrating machine learning according to claim 1, characterized in that, The model dynamic adjustment module includes a feature weight update unit and an input standardized parameter adjustment unit. After each evaluation by the real-time stress level assessment module, the feature weight update unit calculates the dynamic weight coefficients of each modal feature based on the signal-to-noise ratio and feature correlation of each modal data in the current scene, and feeds the weight coefficients back to the cross-modal feature fusion algorithm to adjust the contribution ratio of each modal feature in the fused feature vector. The input standardized parameter adjustment unit dynamically updates the standardized parameters of the input data based on the environmental noise level of the current scene and individual user differences.

10. A stress level assessment and training guidance system integrating machine learning according to claim 9, characterized in that, Based on the signal-to-noise ratio and feature correlation of each modality in the current scenario, the dynamic weight coefficients of each modality feature are calculated as follows: ; in, , They represent the first species, first Attention weight coefficients corresponding to the feature vectors of each modality , They represent the first species, first The signal-to-noise ratio corresponding to the feature vectors of each modality; the attention weight coefficients The initial value is: ; in, , They are respectively with the first species, first Modal eigenvectors The relevant learnable weight vector, , The first species, first The single-modal feature vector is obtained after feature extraction of various physiological signals.