An ai mindfulness training analysis method and system based on multi-modal physiological perception
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING ANSHI DATA TECHNOLOGY CO LTD
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]针对上述方案可见,目前的其一,生理数据采集模态单一,仅依赖心率或呼吸等有限数据,难以全面反映用户在正念训练过程中的身心状态变化;其二,数据分析方法缺乏个性化适配,采用固定算法对所有用户进行统一分析,无法根据不同用户的生理特征、训练目标及实时状态动态调整训练方案;其三,反馈机制滞后且缺乏针对性,通常在训练结束后提供总结性反馈,难以实现训练效果的实时优化
(1)本发明通过采集并处理正念训练周期内用户多模态生理历史感知数据,处理完成后通过数据分析方式对处理后的多模态生理历史感知数据进行分析,分析完成后通过感知评估理论和AI正念分析方式对多模态生理历史感知数据进行量化,并构建心理评估模型;心理评估模型构建完成后实时采集用户多模态生理感知数据,同时通过心理评估模型实时评估用户心理健康状态,并生成正念流派个性化干预方案;生成完成后,基于生成正念流派个性化干预方案对用户进行正念引导,并通过数据预测方式实时预测正念引导后用户心理健康状态变化情况,最后基于实时预测正念引导后用户心理健康状态变化情况动态调整正念流派个性化干预方案,提高了用户心理健康状态分析的准确性。
Smart Images

Figure CN122531736A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, specifically to an AI mindfulness training and analysis method and system based on multimodal physiological perception. Background Technology
[0002] Mindfulness training, as a psychological intervention, has been widely used in scenarios such as stress management, emotion regulation, and cognitive function enhancement. With the development of artificial intelligence technology, existing technologies have attempted to combine physiological signal monitoring with mindfulness training, assisting users in mindfulness practice by collecting physiological data of single or limited modalities such as heart rate and respiration.
[0003] Existing technologies, such as the invention patent application with publication number CN121215191A, disclose an intelligent identification and intervention system for depressive mood. The method includes: a data acquisition module for real-time acquisition of the user's raw data; a data preprocessing module for hierarchical preprocessing of the raw data; a feature extraction module for extracting features from each modality; and a multimodal fusion module that achieves dynamic multimodal fusion based on a bidirectional long short-term memory network and a cross-attention mechanism, and supports depressive binary classification and PHQ through a fully connected network. The system offers multi-task outputs, including a 9-scale regression prediction system; a user interaction module provides users with task guidance, data collection control, and visualized intervention suggestions; and integrates a self-developed VR mindfulness game and an AI psychological model, offering intervention tools such as mindfulness training, dialogue intervention, and serious PC games for different levels of depression. The data management module stores user information, raw data, data on the depressive mood identification process, and assessment results, providing a feasible technical means for the early identification and auxiliary assessment of depression risk.
[0004] As can be seen from the above solutions, firstly, the physiological data collection modality is singular, relying only on limited data such as heart rate or respiration, which is difficult to comprehensively reflect the changes in the user's physical and mental state during mindfulness training; secondly, the data analysis methods lack personalized adaptation, using fixed algorithms to perform uniform analysis on all users, and cannot dynamically adjust the training plan according to the different physiological characteristics, training goals and real-time status of different users; thirdly, the feedback mechanism is lagging and lacks pertinence, usually providing summative feedback after training, which makes it difficult to achieve real-time optimization of training effects. Summary of the Invention
[0005] The purpose of this invention is to provide an AI mindfulness training and analysis method and system based on multimodal physiological perception, which solves the problems existing in the background technology.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides an AI mindfulness training and analysis method and system based on multimodal physiological perception, specifically including the following steps: S1. Collect multimodal physiological history perception data of users during the mindfulness training cycle based on visual perception devices, and process the collected multimodal physiological history perception data through data processing methods to obtain processed multimodal physiological history perception data. S2. The processed multimodal physiological history perception data is analyzed using data analysis methods to obtain the analyzed multimodal physiological history perception data. S3. Quantify the analyzed multimodal physiological history perception data using perception assessment theory and AI mindfulness analysis, and construct a psychological assessment model based on the quantified multimodal physiological history perception data. S4. Collect users' multimodal physiological perception data in real time, and assess users' mental health status in real time through a psychological assessment model. After the assessment is completed, generate a mindfulness-based personalized intervention plan based on the real-time assessment of users' mental health status. S5. Based on the generation of personalized intervention plans for mindfulness schools, users are guided to mindfulness, and changes in the user's mental health status after mindfulness guidance are predicted in real time through data prediction. S6. Based on real-time prediction of changes in users' mental health status after mindfulness guidance, dynamically adjust the personalized intervention plan of the mindfulness school and conduct mindfulness guidance; Preferably, the step of collecting multimodal physiological history perception data from the user's mindfulness training cycle using a visual perception device, and then processing the collected multimodal physiological history perception data to obtain processed multimodal physiological history perception data includes the following steps: S11. Set the acquisition frequency of the visual perception device, and acquire multimodal physiological history perception data based on the set acquisition frequency; Based on the user's mindfulness training cycle, the multimodal physiological history perception data collected within the corresponding cycle were divided into a control group and a standard group; The mindfulness training process includes four components: guided meditation, guided breathing regulation, guided mindful walking, and guided stretching. The control group consisted of multimodal physiological history perception data collected during mindfulness training, while the standard group consisted of multimodal physiological history perception data collected during non-mindfulness training. The collected multimodal physiological history perception data includes: mindfulness training movement data, facial behavior data, and physiological data; The physiological data includes: heart rate data and respiratory rate data, with the heart rate data sampling rate set at 1000Hz and the respiratory rate data sampling rate at 50Hz. Set the sampling rate for facial behavior data to 100Hz; S12. The collected multimodal physiological history perception data is processed through data processing methods to obtain processed multimodal physiological history perception data.
[0007] Preferably, the step of processing the collected multimodal physiological history perception data to obtain processed multimodal physiological history perception data includes the following steps: S121. Process the facial behavior data in the collected multimodal physiological history perception data to obtain processed facial behavior data. The facial behavior data in the multimodal physiological history perception data is collected by a visual perception device, and the facial behavior data includes: frowning, facial expressions, and blinking; Collect initial facial image data of the user, and uniformly set multiple reference points in the initial facial image data of the user; The system processes multiple frames of images captured by the visual perception device through a multi-frame comparison method. By comparing the position coordinates of reference points between adjacent frames, it records the facial behavior data generated by the user during the mindfulness training cycle and determines the data type of the user's current behavior based on the change in the position of the reference points. Arrange the types of user facial behavior data in order of frequency; The sorted facial behavior data is defined as the processed facial behavior data. S122. Process the physiological data in the collected multimodal physiological history perception data to obtain processed physiological data; The heart rate and respiratory rate data in the collected physiological data are filtered using a filtering algorithm. Based on the acquisition frequency of various types of data in physiological data, a corresponding filtering window is set, and various types of data in physiological data are filtered through the set filtering window to output the filtering results of various types of data in physiological data. The filtering results of various types of physiological data are summarized to obtain the processed physiological data. S123. Process the mindfulness training action data in the collected multimodal physiological history perception data to obtain the processed mindfulness training action data. The collected mindfulness training action data is decomposed into frames, and the contours of mindfulness training actions in each frame of the user's image are extracted by background separation. Set the segmentation threshold k for each frame of the user's image, and calculate the average grayscale value of the image based on the set segmentation threshold; The portion less than or equal to the threshold k is designated as the preset shape portion, and the portion greater than the threshold k is designated as the preset background portion; The variance of a preset shape and a preset background is calculated based on the image's grayscale mean, and the maximum calculated variance is set as the optimal threshold. ; The image is binarized based on a threshold to extract the mindfulness training action contours from each frame of the user's image. The extracted mindfulness training motion contours from each frame of the user's image are defined as the processed mindfulness training motion data. S124. The processed facial behavior data, mindfulness training action data, and physiological data are summarized to obtain the processed multimodal physiological history perception data.
[0008] Preferably, the step of analyzing the processed multimodal physiological history perception data through data analysis to obtain the analyzed multimodal physiological history perception data includes the following steps: S21. Extract the features of each type of data in the processed multimodal physiological history perception data by feature analysis. S22. Summarize the characteristics of each type of data in the processed multimodal physiological history perception data to obtain the analyzed multimodal physiological history perception data.
[0009] Preferably, the step of extracting features from each type of data in the processed multimodal physiological history perception data through feature analysis includes the following steps: S211. By employing sliding window technology, features of the processed physiological data are extracted; Set the sliding window size and move it in the time series of processed physiological data according to the set sliding window size. Determine the local features of the processed physiological data by calculating the mean, variance, standard deviation, mode, maximum value, minimum value, range and first difference of the data within the sliding window. The local features of the processed physiological data are summarized to obtain the characteristics of the processed physiological data. S212. Feature extraction is performed on the processed facial behavior data and mindfulness training action data using a convolutional neural network to obtain the features of the processed facial behavior data and mindfulness training action data. The convolutional neural network includes: convolutional layers, pooling layers, and fully connected layers; Set the kernel size, weights, and stride of the convolutional neural network; The processed facial behavior data and mindfulness training action data are input into the convolutional neural network. The convolutional layer determines the feature extraction window and movement direction based on the set convolutional kernel size, weights and stride. After the determination is completed, the feature data in the current window is extracted by convolution calculation. After extraction, the feature data within the current window is integrated and input into the pooling layer. The pooling layer performs dimensionality reduction on the received feature data, and then sends the dimensionality-reduced features to the fully connected layer. The fully connected layer performs feature fusion and outputs the features of the processed facial behavior data and mindfulness training action data.
[0010] Preferably, the step of quantifying the analyzed multimodal physiological history perception data using perception assessment theory and AI mindfulness analysis, and constructing a psychological assessment model based on the quantified multimodal physiological history perception data, includes the following steps: S31. The analyzed multimodal physiological history perception data is quantified using perception assessment theory and scoring quantification methods to obtain quantified multimodal physiological history perception data. We collected ratings from different experts on the analyzed multimodal physiological history perception data, and used a weighted average method to aggregate the ratings from different experts to obtain the aggregated multimodal physiological history perception data. The correlation between the scoring data of the aggregated multimodal physiological history perception data was calculated based on the Pearson correlation coefficient. The correlation between the score data of the aggregated multimodal physiological history perception data and the score data of the aggregated multimodal physiological history perception data is used to obtain the quantified multimodal physiological history perception data. S32. Construct a psychological assessment model based on quantified multimodal physiological history perception data and analyzed multimodal physiological history perception data.
[0011] Preferably, the construction of the psychological assessment model based on the quantified multimodal physiological history perception data and the analyzed multimodal physiological history perception data includes the following steps: An undirected psychological perception graph was constructed based on quantified and analyzed multimodal physiological history perception data. ; in, This represents an undirected mental perception map. This represents the set of psychological assessment state nodes, obtained from the scoring data of aggregated multimodal physiological history and perception data. This represents the set of edges between nodes in the psychological assessment state, obtained from the correlation between the scoring data of the aggregated multimodal physiological history perception data. The characteristics of the psychological assessment state nodes are obtained from the analyzed multimodal physiological history perception data; A graph convolutional neural network is used to train and infer a given undirected psychological perception graph, and a psychological assessment model is constructed based on the training and inference results. The psychological assessment model is designed to consist of a graph convolutional neural network and a multilayer perceptron. The graph convolutional neural network consists of 5 graph convolutional layers, 5 graph pooling layers and 1 fully connected layer, with the graph convolutional layers and graph pooling layers nested together. The undirected mental perception graph is input into the graph convolutional neural network, and the graph convolutional layer extracts the undirected mental perception graph through message passing and state update processes. Calculate the neural transmission information between each pair of nodes and their neighboring nodes, and aggregate the features of the neighboring nodes and the node itself through a weighted summation method via message passing; After aggregation is complete, the aggregation result is processed by setting an update function through the state update process to generate new features for the current node; The new features of the aggregated nodes are input into the graph pooling layer, which then transforms the undirected psychological perception graph extracted by the graph convolutional layer into a feature vector through mean calculation. The processing results of five nested graph convolutional and graph pooling layers are summarized, and the feature vectors of the five layers are concatenated through a fully connected layer to obtain the trained psychological evaluation features. The trained psychological assessment features are input into a multilayer perceptron. The multilayer perceptron then fuses and reduces the dimensionality of the trained psychological assessment features to obtain psychological assessment recognition features.
[0012] Preferably, the real-time acquisition of users' multimodal physiological perception data, and the real-time assessment of users' mental health status through a psychological assessment model, followed by the generation of a mindfulness-based personalized intervention plan based on the real-time assessment of users' mental health status, includes the following steps: Real-time collection of users' multimodal physiological perception data and input into the psychological assessment model. The psychological assessment model dynamically evaluates users' mental health status. When users' mental health status is lower than the standard mental health status, historical multimodal physiological perception data of users during the mindfulness training process is collected and fitted using data fitting methods to determine the impact curve of each mindfulness training process on users' mental health status. Calculate the slope of the curves showing the impact of each mindfulness training process on the user's mental health, and sort the calculated slopes in descending order; Frequency thresholds for each mindfulness training session are set based on the impact curves of each mindfulness training process on the user's mental health. Set the relevant weights for the ranking results and the frequency of each mindfulness training exercise, and generate personalized intervention plans for mindfulness schools based on the ranking results, the frequency thresholds of each mindfulness training exercise, and the set relevant weights.
[0013] Preferably, the step of guiding users to practice mindfulness based on a personalized intervention plan generated from a mindfulness school of thought, and predicting changes in users' mental health status after mindfulness guidance in real time through data prediction, includes the following steps: S51. Based on the generation of personalized intervention solutions for mindfulness schools, drive digital humans to conduct real-time mindfulness-guided dialogue interaction, and guide users to complete the corresponding personalized intervention solutions for mindfulness schools based on the real-time mindfulness-guided dialogue interaction method. Standard action videos for personalized intervention programs corresponding to mindfulness schools of thought are generated through digital human simulation. The system collects users’ mindfulness training action data in real time through visual perception devices, and guides users to complete the corresponding mindfulness-based personalized intervention program through multi-frame comparison and real-time mindfulness-guided dialogue interaction. Feature extraction is performed on real-time user mindfulness training action data using convolutional neural networks, and the mindfulness training action data is identified and feature points are saved based on the feature extraction results. Based on the saved feature points, the key coordinate positions of the user's mindfulness training actions are determined. At the same time, the key coordinate positions are compared with the corresponding frame images of the standard action video, and the action completion rate is judged based on the key coordinate positions. If the action completion rate is higher than 80%, it means that the user's mindfulness training action data is normal and no correction is needed. Otherwise, it means that the user's mindfulness training action data is abnormal and correction is guided through mindfulness-guided dialogue interaction. S52. Construct an LSTM network model to predict changes in users' mental health status in real time after mindfulness guidance. The LSTM network includes an input gate, an output gate, and a forget gate; S521. Input the user's mental health status set and the mindfulness-based personalized intervention program set as input gates into the LSTM network. S522. Process the current input and the user's mental health status from the previous moment through the forget gate, and update the user's mental health status at the current moment. S523. Output the user's mental health status after processing by the forget gate through the output gate to obtain the predicted user mental health status at the next moment. S524. Iteratively use the LSTM network to determine and obtain the trained LSTM network; Set an error threshold between the user's mental health status data predicted by the LSTM network and the user's multimodal physiological perception data collected in real time. When the error between the user's mental health status data predicted by the LSTM network and the user's multimodal physiological perception data collected in real time is within the set error threshold range, the iteration stops, and the trained LSTM network is obtained. S525. Set an LSTM network that meets the error threshold range as the state prediction model, and output the predicted running state data for the next time step.
[0014] This invention also provides a personalized mindfulness training system based on multimodal physiological perception and AI mindfulness analysis, which is used to implement a personalized mindfulness training method based on multimodal physiological perception and AI mindfulness analysis. The system includes: a data acquisition module, a data processing module, a data analysis module, an AI mindfulness analysis module, a psychological assessment module, a mindfulness training module, a dynamic adjustment module, and a display module. The data acquisition module is used to acquire multimodal physiological sensing data; The data processing module is used to process the collected multimodal physiological sensing data to obtain processed multimodal physiological sensing data. The data analysis module is used to analyze the processed multimodal physiological perception data to obtain the analyzed multimodal physiological perception data. The AI mindfulness analysis module is used to quantify the analyzed multimodal physiological perception data; The psychological assessment module is used to construct a psychological assessment model based on the quantified multimodal physiological perception data; The mindfulness training module is used to generate and execute personalized intervention plans based on real-time assessments of users' mental health status. The dynamic adjustment module is used to predict changes in the user's mental health status in real time after mindfulness guidance through data prediction, and to dynamically adjust the personalized intervention plan of the mindfulness school based on the prediction results.
[0015] The beneficial effects of this invention are as follows: (1) This invention collects and processes multimodal physiological history perception data of users during the mindfulness training period. After processing, the multimodal physiological history perception data is analyzed by data analysis. After analysis, the multimodal physiological history perception data is quantified by perception assessment theory and AI mindfulness analysis, and a psychological assessment model is constructed. After the psychological assessment model is constructed, the multimodal physiological perception data of users is collected in real time, and the psychological assessment model is used to assess the psychological health status of users in real time and generate a personalized intervention plan of the mindfulness school. After generation, the user is guided to mindfulness based on the generated personalized intervention plan of the mindfulness school, and the changes in the user's psychological health status after mindfulness guidance are predicted in real time by data prediction. Finally, the personalized intervention plan of the mindfulness school is dynamically adjusted based on the real-time prediction of the changes in the user's psychological health status after mindfulness guidance, thereby improving the accuracy of the analysis of the user's psychological health status. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the AI mindfulness training and analysis method for multimodal physiological perception according to the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] In a specific embodiment of the present invention, Reference Figure 1 As shown, this invention provides an AI mindfulness training and analysis method and system based on multimodal physiological perception, comprising: S1. Collect multimodal physiological history perception data of users during the mindfulness training cycle based on visual perception devices, and process the collected multimodal physiological history perception data through data processing methods to obtain processed multimodal physiological history perception data. S2. The processed multimodal physiological history perception data is analyzed using data analysis methods to obtain the analyzed multimodal physiological history perception data. S3. Quantify the analyzed multimodal physiological history perception data using perception assessment theory and AI mindfulness analysis, and construct a psychological assessment model based on the quantified multimodal physiological history perception data. S4. Collect users' multimodal physiological perception data in real time, and assess users' mental health status in real time through a psychological assessment model. After the assessment is completed, generate a mindfulness-based personalized intervention plan based on the real-time assessment of users' mental health status. S5. Based on the generation of personalized intervention plans for mindfulness schools, users are guided to mindfulness, and changes in the user's mental health status after mindfulness guidance are predicted in real time through data prediction. S6. Based on real-time prediction of changes in users' mental health status after mindfulness guidance, dynamically adjust the personalized intervention plan of the mindfulness school and conduct mindfulness guidance.
[0020] Furthermore, referring to Figure 1 As shown, multimodal physiological history perception data is collected from users during their mindfulness training period using visual perception devices. The collected multimodal physiological history perception data is then processed using data processing methods to obtain the processed multimodal physiological history perception data, including the following steps: S11. Set the acquisition frequency of the visual perception device, and acquire multimodal physiological history perception data based on the set acquisition frequency; Based on the user's mindfulness training cycle, the multimodal physiological history perception data collected within the corresponding cycle were divided into a control group and a standard group; The mindfulness training process includes four components: guided meditation, guided breathing regulation, guided mindful walking, and guided stretching. The control group consisted of multimodal physiological history perception data collected during mindfulness training, while the standard group consisted of multimodal physiological history perception data collected during non-mindfulness training. Furthermore, the collected multimodal physiological history perception data is defined to include: mindfulness training movement data, facial behavior data, and physiological data; The physiological data includes: heart rate data and respiratory rate data, with the heart rate data sampling rate set at 1000Hz and the respiratory rate data sampling rate at 50Hz. Set the sampling rate for facial behavior data to 100Hz; S12. The collected multimodal physiological history perception data is processed by data processing methods to obtain processed multimodal physiological history perception data. S121. Process the facial behavior data in the collected multimodal physiological history perception data to obtain processed facial behavior data. The facial behavior data in the multimodal physiological history perception data is collected by a visual perception device, and the facial behavior data includes: frowning, facial expressions, and blinking; Collect initial facial image data of the user, and uniformly set multiple reference points in the initial facial image data of the user; Furthermore, the multi-frame images captured by the visual perception device are processed by multi-frame comparison. By comparing the position coordinates of reference points between adjacent frames, the facial behavior data generated by the user during the mindfulness training cycle is recorded, and the data type of the user's current behavior is determined according to the change in the position of the reference points. Arrange the types of user facial behavior data in order of frequency; The sorted facial behavior data is defined as the processed facial behavior data. S122. Process the physiological data in the collected multimodal physiological history perception data to obtain processed physiological data; The heart rate and respiratory rate data in the collected physiological data are filtered using a filtering algorithm. Based on the acquisition frequency of various types of data in physiological data, a corresponding filtering window is set, and various types of data in physiological data are filtered through the set filtering window to output the filtering results of various types of data in physiological data. Furthermore, the filtering results of various types of data in the physiological data are summarized to obtain the processed physiological data; The filtering is performed based on the set filter window as follows: Set the filter window to , This represents the data value of the corresponding type of physiological data within the j-th window; ; in, Indicates window size, This represents the data value of the corresponding type of physiological data within the j-th window after filtering. This indicates the signal strength of the corresponding type of physiological data. This represents the signal strength collected per unit of time; S123. Process the mindfulness training action data in the collected multimodal physiological history perception data to obtain the processed mindfulness training action data. The collected mindfulness training action data is decomposed into frames, and the mindfulness training action data in each frame of the user's image is extracted by background separation. Set a segmentation threshold k for each frame of the user's image. Define the portion less than or equal to threshold k as the preset body portion and the portion greater than threshold k as the preset background portion. Set the number of pixels for the body portion and the background portion as follows: and The mean gray level of the image is shown in the formula: ; in, and These are the average grayscale values of the preset face portion and the average grayscale value of the preset background portion, respectively. This is a preset ratio of the number of pixels in the shape to the total number of pixels. The proportion of background pixels to the total number of pixels. This indicates the total number of pixels; Furthermore, the variance of the preset shape portion and the preset background portion is calculated based on the image grayscale mean, and the maximum calculated variance is set as the optimal threshold. ; Furthermore, the image is binarized based on a threshold to extract the mindfulness training action contours from each frame of the user's image; The extracted mindfulness training motion contours from each frame of the user's image are defined as the processed mindfulness training motion data. S124. The processed facial behavior data, mindfulness training action data and physiological data are summarized to obtain the processed multimodal physiological history perception data. Furthermore, referring to Figure 1 As shown, the processed multimodal physiological history perception data is analyzed using data analysis methods. The analyzed multimodal physiological history perception data includes the following steps: S21. Extract the features of each type of data in the processed multimodal physiological history perception data by feature analysis. S211. By employing sliding window technology, features of the processed physiological data are extracted; Set the sliding window size and move it in the time series of processed physiological data according to the set sliding window size. Determine the local features of the processed physiological data by calculating the mean, variance, standard deviation, mode, maximum value, minimum value, range and first difference of the data within the sliding window. Furthermore, the local features of the processed physiological data are summarized to obtain the features of the processed physiological data; S212. Feature extraction is performed on the processed facial behavior data and mindfulness training action data using a convolutional neural network to obtain the features of the processed facial behavior data and mindfulness training action data. The convolutional neural network includes: convolutional layers, pooling layers, and fully connected layers; Set the kernel size, weights, and stride of the convolutional neural network; The processed facial behavior data and mindfulness training action data are input into the convolutional neural network. The convolutional layer determines the feature extraction window and movement direction based on the set convolutional kernel size, weights and stride. After the determination is completed, the feature data in the current window is extracted by convolution calculation. Furthermore, after extraction, the feature data within the current window is integrated and input into the pooling layer. The pooling layer performs dimensionality reduction on the received feature data, and then the dimensionality-reduced features are sent to the fully connected layer. The fully connected layer performs feature fusion and outputs the features of the processed facial behavior data and mindfulness training action data. S22. Summarize the characteristics of each type of data in the processed multimodal physiological history perception data to obtain the analyzed multimodal physiological history perception data. Furthermore, referring to Figure 1 As shown, the analysis of multimodal physiological history perception data is quantified using perceptual assessment theory and AI mindfulness analysis. A psychological assessment model is then constructed based on this quantified data, including the following steps: S31. The analyzed multimodal physiological history perception data is quantified using perception assessment theory and scoring quantification methods to obtain quantified multimodal physiological history perception data. We collected ratings from different experts on the analyzed multimodal physiological history perception data, and used a weighted average method to aggregate the ratings from different experts to obtain the aggregated multimodal physiological history perception data. Furthermore, the correlation between the scoring data of the aggregated multimodal physiological history perception data was calculated based on the Pearson correlation coefficient; Furthermore, the correlation between the score data of the aggregated multimodal physiological history perception data and the score data of the aggregated multimodal physiological history perception data is used to obtain the quantified multimodal physiological history perception data. The formula for calculating correlation is as follows: ; in, Representing rating data and rating data Covariance between Representing rating data standard deviation Representing rating data standard deviation Representing rating data and rating data The correlation between them; S32. Construct a psychological assessment model based on quantified multimodal physiological history perception data and analyzed multimodal physiological history perception data; An undirected psychological perception graph was constructed based on quantified and analyzed multimodal physiological history perception data. ; in, This represents an undirected mental perception map. This represents the set of psychological assessment state nodes, obtained from the scoring data of aggregated multimodal physiological history and perception data. This represents the set of edges between nodes in the psychological assessment state, obtained from the correlation between the scoring data of the aggregated multimodal physiological history perception data. The characteristics of the psychological assessment state nodes are obtained from the analyzed multimodal physiological history perception data; Furthermore, a psychological assessment model is constructed based on the training and reasoning results of a graph convolutional neural network on a given undirected psychological perception graph. The psychological assessment model is designed to consist of a graph convolutional neural network and a multilayer perceptron. The graph convolutional neural network consists of 5 graph convolutional layers, 5 graph pooling layers and 1 fully connected layer, with the graph convolutional layers and graph pooling layers nested together. The undirected mental perception graph is input into the graph convolutional neural network, and the graph convolutional layer extracts the undirected mental perception graph through message passing and state update processes. Calculate the neural transmission information between each pair of nodes and their neighboring nodes, and aggregate the features of the neighboring nodes and the node itself through a weighted summation method via message passing; After aggregation is complete, the aggregation result is processed by setting an update function through the state update process to generate new features for the current node; Furthermore, the new features of the aggregated nodes are input into the graph pooling layer, which transforms the undirected psychological perception graph extracted by the graph convolutional layer into a feature vector through mean calculation. Furthermore, the processing results of five nested graph convolutional layers and graph pooling layers are summarized, and the feature vectors of the five layers are concatenated through a fully connected layer to obtain the trained psychological evaluation features. Furthermore, the trained psychological assessment features are input into a multilayer perceptron, which then fuses and reduces the dimensionality of the trained psychological assessment features to obtain psychological assessment recognition features. The multilayer perceptron includes: an input layer, a hidden layer, and an output layer; The output expression of the multilayer perceptron is shown below: ; in, This represents the activation function. This is the weight matrix of the hidden layer in layer 1. This is the bias vector for layer 1. This represents the output value of group T in layer 1. This represents the Tth group of input values, with 0 indicating the input layer; Furthermore, referring to Figure 1 As shown, the system collects multimodal physiological perception data from users in real time and assesses their mental health status in real time using a psychological assessment model. After the assessment, a personalized mindfulness-based intervention plan is generated based on the real-time assessment of the user's mental health status, including the following steps: Real-time collection of users' multimodal physiological perception data and input into the psychological assessment model. The psychological assessment model dynamically evaluates users' mental health status. When users' mental health status is lower than the standard mental health status, historical multimodal physiological perception data of users during the mindfulness training process is collected and fitted using data fitting methods to determine the impact curve of each mindfulness training process on users' mental health status. The normal fit formula is shown below: ; in, The function representing the normal fit of a user's mental health status over a period of time with respect to the user's mindfulness training process. It is the natural logarithm. Represents the logarithmic transformation variance Represents the logarithmic transformation standard deviation This represents the mean of a normal distribution. This represents the change in a user's mental health status over time during mindfulness training. Furthermore, the slope of the impact curve of each mindfulness training process on the user's mental health status was calculated, and the calculated slopes were sorted in descending order. Furthermore, frequency thresholds for each mindfulness training session are set based on the impact curves of each mindfulness training process on the user's mental health status; Furthermore, the relevant weights of the ranking results and the frequency of each mindfulness training exercise are set, and personalized intervention plans for mindfulness schools are generated based on the ranking results, the frequency thresholds of each mindfulness training exercise, and the set relevant weights. Furthermore, referring to Figure 1 As shown, the process of guiding users to practice mindfulness based on a personalized intervention plan generated from the mindfulness school, and predicting changes in users' mental health status in real time after mindfulness guidance through data prediction, includes the following steps: S51. Based on the generation of personalized intervention solutions for mindfulness schools, drive digital humans to conduct real-time mindfulness-guided dialogue interaction, and guide users to complete the corresponding personalized intervention solutions for mindfulness schools based on the real-time mindfulness-guided dialogue interaction method. Standard action videos for personalized intervention programs corresponding to mindfulness schools of thought are generated through digital human simulation. Furthermore, the system collects users' mindfulness training action data in real time through visual perception devices, and guides users to complete corresponding personalized intervention programs of mindfulness schools through multi-frame comparison and real-time mindfulness guided dialogue interaction. Feature extraction is performed on real-time user mindfulness training action data using convolutional neural networks, and the mindfulness training action data is identified and feature points are saved based on the feature extraction results. Furthermore, based on the saved feature points, the key coordinate positions of the user's mindfulness training actions are determined. At the same time, the key coordinate positions are compared with the corresponding frame images of the standard action video, and the action completion rate is judged according to the key coordinate positions. If the action completion rate is higher than 80%, it means that the user's mindfulness training action data is normal and no correction is needed. Otherwise, it means that the user's mindfulness training action data is abnormal and correction is guided through mindfulness-guided dialogue interaction. S52. Construct an LSTM network model to predict changes in users' mental health status in real time after mindfulness guidance. The LSTM network includes an input gate, an output gate, and a forget gate; Set user mental health status set ,in This represents the user's first state, where Represents the user's z-th state; a set of personalized intervention solutions from the mindfulness school of thought. ,in This refers to the first type of personalized intervention program from the mindfulness school, in which... This represents the personalized intervention plan for the nth mindfulness school of thought; S521. Input the user's mental health status set and the mindfulness-based personalized intervention program set as input gates into the LSTM network. S522. Process the current input and the user's mental health status from the previous moment through the forget gate, and update the user's mental health status at the current moment. S523. Output the user's mental health status after processing by the forget gate through the output gate to obtain the predicted user mental health status at the next moment. S524. Iteratively use the LSTM network to determine and obtain the trained LSTM network; Set an error threshold between the user's mental health status data predicted by the LSTM network and the user's multimodal physiological perception data collected in real time. When the error between the user's mental health status data predicted by the LSTM network and the user's multimodal physiological perception data collected in real time is within the set error threshold range, the iteration stops, and the trained LSTM network is obtained. S525. Set an LSTM network that meets the error threshold range as the state prediction model, and output the predicted running state data for the next time step. Furthermore, referring to Figure 1 As shown, the personalized intervention plan for mindfulness-based approaches is dynamically adjusted based on real-time predictions of changes in the user's mental health status after mindfulness guidance, and the mindfulness guidance includes the following steps: We collect and summarize data on changes in users’ mental health status after real-time prediction of mindfulness guidance, and determine the temporal distribution of the data on changes in users’ mental health status through data fitting. The data fitting formula is shown below: ; in, The probability density function representing the distribution of user mental health status changes over time after mindfulness guidance. For scale parameters, Represents the shape parameters of the data fit; Furthermore, a fluctuation threshold is set, and the training time and cycle of the personalized intervention program of the mindfulness school are dynamically adjusted based on the calculated time distribution. In one specific embodiment, the personalized mindfulness training system based on multimodal physiological perception and AI mindfulness analysis is used to implement a personalized mindfulness training method based on multimodal physiological perception and AI mindfulness analysis. The system includes: a data acquisition module, a data processing module, a data analysis module, an AI mindfulness analysis module, a psychological assessment module, a mindfulness training module, a dynamic adjustment module, and a display module. The data acquisition module is used to acquire multimodal physiological sensing data; The data processing module is used to process the collected multimodal physiological sensing data to obtain processed multimodal physiological sensing data. The data analysis module is used to analyze the processed multimodal physiological perception data to obtain the analyzed multimodal physiological perception data. The AI mindfulness analysis module is used to quantify the analyzed multimodal physiological perception data; The psychological assessment module is used to construct a psychological assessment model based on the quantified multimodal physiological perception data; The mindfulness training module is used to generate and execute personalized intervention plans based on real-time assessments of users' mental health status. The dynamic adjustment module is used to predict changes in the user's mental health status in real time after mindfulness guidance through data prediction, and to dynamically adjust the personalized intervention plan of the mindfulness school based on the prediction results.
[0021] It should be noted that, The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.
Claims
1. An AI mindfulness training and analysis method based on multimodal physiological perception, characterized in that, Includes the following steps: S1. Collect multimodal physiological history perception data of users during the mindfulness training cycle based on visual perception devices, and process the collected multimodal physiological history perception data through data processing methods to obtain processed multimodal physiological history perception data. S2. The processed multimodal physiological history perception data is analyzed using data analysis methods to obtain the analyzed multimodal physiological history perception data. S3. Quantify the analyzed multimodal physiological history perception data using perception assessment theory and AI mindfulness analysis, and construct a psychological assessment model based on the quantified multimodal physiological history perception data. S4. Collect users' multimodal physiological perception data in real time, and assess users' mental health status in real time through a psychological assessment model. After the assessment is completed, generate a mindfulness-based personalized intervention plan based on the real-time assessment of users' mental health status. S5. Based on the generation of personalized intervention plans for mindfulness schools, users are guided to mindfulness, and changes in the user's mental health status after mindfulness guidance are predicted in real time through data prediction. S6. Based on real-time prediction of changes in users' mental health status after mindfulness guidance, dynamically adjust the personalized intervention plan of the mindfulness school and conduct mindfulness guidance.
2. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 1, characterized in that, The process of collecting multimodal physiological history perception data from users during their mindfulness training cycle using visual perception devices, and then processing this data to obtain the processed multimodal physiological history perception data includes the following steps: S11. Set the acquisition frequency of the visual perception device, and acquire multimodal physiological history perception data based on the set acquisition frequency; Based on the user's mindfulness training cycle, the multimodal physiological history perception data collected within the corresponding cycle were divided into a control group and a standard group; The mindfulness training process includes four components: guided meditation, guided breathing regulation, guided mindful walking, and guided stretching. The control group consisted of multimodal physiological history perception data collected during mindfulness training, while the standard group consisted of multimodal physiological history perception data collected during non-mindfulness training. The collected multimodal physiological history perception data includes: mindfulness training movement data, facial behavior data, and physiological data; The physiological data includes: heart rate data and respiratory rate data, with the heart rate data sampling rate set at 1000Hz and the respiratory rate data sampling rate at 50Hz. Set the sampling rate for facial behavior data to 100Hz; S12. The collected multimodal physiological history perception data is processed through data processing methods to obtain processed multimodal physiological history perception data.
3. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 2, characterized in that, The process of processing the collected multimodal physiological history perception data to obtain processed multimodal physiological history perception data includes the following steps: S121. Process the facial behavior data in the collected multimodal physiological history perception data to obtain processed facial behavior data. The facial behavior data in the multimodal physiological history perception data is collected by a visual perception device, and the facial behavior data includes: frowning, facial expressions, and blinking; Collect initial facial image data of the user, and uniformly set multiple reference points in the initial facial image data of the user; The system processes multiple frames of images captured by the visual perception device through a multi-frame comparison method. By comparing the position coordinates of reference points between adjacent frames, it records the facial behavior data generated by the user during the mindfulness training cycle and determines the data type of the user's current behavior based on the change in the position of the reference points. Arrange the types of user facial behavior data in order of frequency; The sorted facial behavior data is defined as the processed facial behavior data. S122. Process the physiological data in the collected multimodal physiological history perception data to obtain processed physiological data; The heart rate and respiratory rate data in the collected physiological data are filtered using a filtering algorithm. Based on the acquisition frequency of various types of data in physiological data, a corresponding filtering window is set, and various types of data in physiological data are filtered through the set filtering window to output the filtering results of various types of data in physiological data. The filtering results of various types of physiological data are summarized to obtain the processed physiological data. S123. Process the mindfulness training action data in the collected multimodal physiological history perception data to obtain the processed mindfulness training action data. The collected mindfulness training action data is decomposed into frames, and the contours of mindfulness training actions in each frame of the user's image are extracted by background separation. Set the segmentation threshold k for each frame of the user's image, and calculate the average grayscale value of the image based on the set segmentation threshold; The portion less than or equal to the threshold k is designated as the preset shape portion, and the portion greater than the threshold k is designated as the preset background portion; The variance of a preset shape and a preset background is calculated based on the image's grayscale mean, and the maximum calculated variance is set as the optimal threshold. ; The image is binarized based on a threshold to extract the mindfulness training action contours from each frame of the user's image. The extracted mindfulness training motion contours from each frame of the user's image are defined as the processed mindfulness training motion data. S124. The processed facial behavior data, mindfulness training action data, and physiological data are summarized to obtain the processed multimodal physiological history perception data.
4. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 1, characterized in that, The process of analyzing the processed multimodal physiological history perception data to obtain the analyzed multimodal physiological history perception data includes the following steps: S21. Extract the features of each type of data in the processed multimodal physiological history perception data by feature analysis. S22. Summarize the characteristics of each type of data in the processed multimodal physiological history perception data to obtain the analyzed multimodal physiological history perception data.
5. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 4, characterized in that, The process of extracting features from each type of data in the processed multimodal physiological history perception data through feature analysis includes the following steps: S211. By employing the sliding window technique, features of the processed physiological data are extracted. Set the sliding window size and move it in the time series of processed physiological data according to the set sliding window size. Determine the local features of the processed physiological data by calculating the mean, variance, standard deviation, mode, maximum value, minimum value, range and first difference of the data within the sliding window. The local features of the processed physiological data are summarized to obtain the characteristics of the processed physiological data. S212. Extract features from the processed facial behavior data and mindfulness training action data using a convolutional neural network to obtain the features of the processed facial behavior data and mindfulness training action data. The convolutional neural network includes: convolutional layers, pooling layers, and fully connected layers; Set the kernel size, weights, and stride of the convolutional neural network; The processed facial behavior data and mindfulness training action data are input into the convolutional neural network. The convolutional layer determines the feature extraction window and movement direction based on the set convolutional kernel size, weights and stride. After the determination is completed, the feature data in the current window is extracted by convolution calculation. After extraction, the feature data within the current window is integrated and input into the pooling layer. The pooling layer performs dimensionality reduction on the received feature data, and then sends the dimensionality-reduced features to the fully connected layer. The fully connected layer performs feature fusion and outputs the features of the processed facial behavior data and mindfulness training action data.
6. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 1, characterized in that, The process of quantifying the analyzed multimodal physiological history perception data using perception assessment theory and AI mindfulness analysis, and then constructing a psychological assessment model based on the quantified multimodal physiological history perception data, includes the following steps: S31. The analyzed multimodal physiological history perception data is quantified using perception assessment theory and scoring quantification methods to obtain quantified multimodal physiological history perception data. We collected ratings from different experts on the analyzed multimodal physiological history perception data, and used a weighted average method to aggregate the ratings from different experts to obtain the aggregated multimodal physiological history perception data. The correlation between the scoring data of the aggregated multimodal physiological history perception data was calculated based on the Pearson correlation coefficient. The correlation between the score data of the aggregated multimodal physiological history perception data and the score data of the aggregated multimodal physiological history perception data is used to obtain the quantified multimodal physiological history perception data. S32. Construct a psychological assessment model based on quantified multimodal physiological history perception data and analyzed multimodal physiological history perception data.
7. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 6, characterized in that, The construction of the psychological assessment model based on quantified and analyzed multimodal physiological history perception data includes the following steps: An undirected psychological perception graph was constructed based on quantified and analyzed multimodal physiological history perception data. ; in, This represents an undirected mental perception map. This represents the set of psychological assessment state nodes, obtained from the scoring data of aggregated multimodal physiological history and perception data. This represents the set of edges between nodes in the psychological assessment state, obtained from the correlation between the scoring data of the aggregated multimodal physiological history perception data. The characteristics of the psychological assessment state nodes are obtained from the analyzed multimodal physiological history perception data; A graph convolutional neural network is used to train and infer a given undirected psychological perception graph, and a psychological assessment model is constructed based on the training and inference results. The psychological assessment model is designed to consist of a graph convolutional neural network and a multilayer perceptron. The graph convolutional neural network consists of 5 graph convolutional layers, 5 graph pooling layers and 1 fully connected layer, with the graph convolutional layers and graph pooling layers nested together. The undirected mental perception graph is input into the graph convolutional neural network, and the graph convolutional layer extracts the undirected mental perception graph through message passing and state update processes. Calculate the neural transmission information between each pair of nodes and their neighboring nodes, and aggregate the features of the neighboring nodes and the node itself through a weighted summation method via message passing; After aggregation is complete, the aggregation result is processed by setting an update function through the state update process to generate new features for the current node; The new features of the aggregated nodes are input into the graph pooling layer, which then transforms the undirected psychological perception graph extracted by the graph convolutional layer into a feature vector through mean calculation. The processing results of five nested graph convolutional and graph pooling layers are summarized, and the feature vectors of the five layers are concatenated through a fully connected layer to obtain the trained psychological evaluation features. The trained psychological assessment features are input into a multilayer perceptron. The multilayer perceptron then fuses and reduces the dimensionality of the trained psychological assessment features to obtain psychological assessment recognition features.
8. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 1, characterized in that, The process of collecting users' multimodal physiological perception data in real time and assessing their mental health status in real time using a psychological assessment model, followed by generating a personalized mindfulness-based intervention plan based on the real-time assessment of the users' mental health status, includes the following steps: Real-time collection of users' multimodal physiological perception data and input into the psychological assessment model. The psychological assessment model dynamically evaluates users' mental health status. When users' mental health status is lower than the standard mental health status, historical multimodal physiological perception data of users during the mindfulness training process is collected and fitted using data fitting methods to determine the impact curve of each mindfulness training process on users' mental health status. Calculate the slope of the curves showing the impact of each mindfulness training process on the user's mental health, and sort the calculated slopes in descending order; Frequency thresholds for each mindfulness training session are set based on the impact curves of each mindfulness training process on the user's mental health. Set the relevant weights for the ranking results and the frequency of each mindfulness training exercise, and generate personalized intervention plans for mindfulness schools based on the ranking results, the frequency thresholds of each mindfulness training exercise, and the set relevant weights.
9. The AI mindfulness training and analysis method based on multimodal physiological perception according to claim 1, characterized in that, The personalized intervention program based on the generative mindfulness school guides users in mindfulness, and uses data prediction to predict changes in users' mental health status after mindfulness guidance in real time, including the following steps: S51. Based on the generation of personalized intervention solutions for mindfulness schools, drive digital humans to conduct real-time mindfulness-guided dialogue interaction, and guide users to complete the corresponding personalized intervention solutions for mindfulness schools based on the real-time mindfulness-guided dialogue interaction method. Standard action videos for personalized intervention programs corresponding to mindfulness schools of thought are generated through digital human simulation. The system collects users’ mindfulness training action data in real time through visual perception devices, and guides users to complete the corresponding mindfulness-based personalized intervention program through multi-frame comparison and real-time mindfulness-guided dialogue interaction. Feature extraction is performed on real-time user mindfulness training action data using convolutional neural networks, and the mindfulness training action data is identified and feature points are saved based on the feature extraction results. Based on the saved feature points, the key coordinate positions of the user's mindfulness training actions are determined. At the same time, the key coordinate positions are compared with the corresponding frame images of the standard action video, and the action completion rate is judged based on the key coordinate positions. If the action completion rate is higher than 80%, it means that the user's mindfulness training action data is normal and no correction is needed. Otherwise, it means that the user's mindfulness training action data is abnormal and correction is guided through mindfulness-guided dialogue interaction. S52. Construct an LSTM network model to predict changes in users' mental health status in real time after mindfulness guidance. The LSTM network includes an input gate, an output gate, and a forget gate; S521. Input the user's mental health status set and the mindfulness-based personalized intervention program set as input gates into the LSTM network. S522. Process the current input and the user's mental health status from the previous moment through the forget gate, and update the user's mental health status at the current moment. S523. Output the user's mental health status after processing by the forget gate through the output gate to obtain the predicted user mental health status at the next moment. S524. Iteratively use the LSTM network to determine and obtain the trained LSTM network; Set an error threshold between the user's mental health status data predicted by the LSTM network and the user's multimodal physiological perception data collected in real time. When the error between the user's mental health status data predicted by the LSTM network and the user's multimodal physiological perception data collected in real time is within the set error threshold range, the iteration stops, and the trained LSTM network is obtained. S525. Set an LSTM network that meets the error threshold range as the state prediction model, and output the predicted running state data for the next time step.
10. A system for implementing the AI mindfulness training and analysis method based on multimodal physiological perception as described in claim 1, characterized in that, The system includes: a data acquisition module, a data processing module, a data analysis module, an AI mindfulness analysis module, a psychological assessment module, a mindfulness training module, a dynamic adjustment module, and a display module; The data acquisition module is used to acquire multimodal physiological sensing data; The data processing module is used to process the collected multimodal physiological sensing data to obtain processed multimodal physiological sensing data. The data analysis module is used to analyze the processed multimodal physiological perception data to obtain the analyzed multimodal physiological perception data. The AI mindfulness analysis module is used to quantify the analyzed multimodal physiological perception data; The psychological assessment module is used to construct a psychological assessment model based on the quantified multimodal physiological perception data. The mindfulness training module is used to generate and execute personalized intervention plans based on real-time assessments of users' mental health status. The dynamic adjustment module is used to predict changes in the user's mental health status in real time after mindfulness guidance through data prediction, and to dynamically adjust the personalized intervention plan of the mindfulness school based on the prediction results.
Citation Information
Patent Citations
Intelligent depression emotion recognition and intervention system
CN121215191A