AI-based Mental Health Assessment System
Through the AI-based mental health assessment system, the camera uses the camera to collect facial video streams for keyframe sampling and muscle unit partitioning, and extracts the characteristics of micromovement status, solving the problem of time-consuming and low accuracy of traditional mental health assessment, and achieving fast and accurate mental health detection.
Patent Information
- Application Number
- CN202510344181.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Traditional mental health assessment methods have problems such as time-consuming, relying on doctor experience, low accuracy, and inability to conduct rapid large-scale testing. Especially in scale analysis, individual impact assessment results are easily misjudged and concealed.
Using an AI-based mental health assessment system, facial video stream is collected through the camera, keyframe sampling and facial recognition, muscle units are partitioned, micromovement state features are extracted, and comprehensive representation is used to judge mental health types.
It realizes fast and accurate mental health testing, avoids subjective interference from patients, improves the efficiency and accuracy of the assessment, and provides a preliminary mental health type label.
Smart Images

Figure CN119867756B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent evaluation, and more specifically, to a mental health assessment system based on AI. Background Art
[0002] Currently, social mental health problems have become increasingly prominent. From teenagers to adults, they all face different degrees of troubles. Due to academic and social pressures, teenagers are prone to mood swings, depression and other problems, and extreme events occur from time to time; while adults are under work and life pressures, and mental illnesses such as anxiety and depression are gradually increasing. These problems not only affect the quality of personal life and work efficiency, but also pose potential threats to family harmony and social stability.
[0003] However, traditional mental health assessment methods have many drawbacks. Specifically, when seeing a doctor, patients need to go through long queuing, registration and waiting processes, and it is difficult to accurately express their feelings, resulting in the diagnosis results often relying on the personal experience of doctors and unable to meet the needs of large-scale rapid detection. Scale analysis also faces challenges: the question-solving cycle is long and not suitable for batch physical examinations; the question settings are complex and ambiguous, which is easy to lead to misjudgment; patients may randomly check answers due to impatience, further affecting the accuracy. In addition, some important position personnel may be worried about affecting their future and conceal their symptoms. Coupled with the fact that the scale cannot detect physiological index diseases, the accuracy and comprehensiveness of the assessment are significantly reduced.
[0004] Therefore, a mental health assessment solution based on AI is expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed.
[0006] According to one aspect of this application, a mental health assessment system based on AI is provided, which includes:
[0007] A face acquisition module, configured to acquire face data information of a target object to be measured through a camera, and the face data information is a face video stream;
[0008] An identification module, configured to perform face area recognition based on key frames on the face data information to obtain a time series set of face area images;
[0009] A displacement state characterization module is used to comprehensively analyze the micro-motion displacements of facial muscle points in the time series set of the facial region images to obtain a comprehensive characterization of the micro-motion displacement states of facial muscle points. Among them, the displacement state characterization module includes: a muscle point feature extraction unit, which is used to extract the micro-motion features of muscle points from the time series set of the facial region images to obtain a time series set of micro-motion state features of multiple facial muscle points; a displacement comprehensive transfer unit, which is used to comprehensively transfer the time series of the core facial muscle point displacements for the time series set of the micro-motion state features of multiple facial muscle points to obtain the comprehensive characterization of the micro-motion displacement states of facial muscle points.
[0010] An evaluation module is used to obtain a mental health evaluation result representing a mental health type label based on the comprehensive characterization of the micro-motion displacement states of facial muscle points.
[0011] Compared with the prior art, a mental health evaluation system based on AI provided by the present application collects facial data information (facial video stream) of a target object to be measured through a camera, uses an AI-based image recognition and analysis algorithm to perform key frame sampling and facial recognition on the facial data information. Then, it performs muscle unit partitioning and micro-motion state feature extraction on each recognized facial region image. Next, it performs core muscle displacement state time series transfer on the time series set of the micro-motion state features of each extracted facial muscle point, so as to intelligently judge the mental health type according to the comprehensive characterization between the transferred micro-motion displacement state information transfer features of each facial muscle point. The present application uses a camera to collect data and processes it through an AI algorithm, which can quickly complete operations such as key frame sampling, realizing efficient detection. At the same time, it accurately captures the subtle changes in facial muscles, avoids the subjective interference of patients, and improves the evaluation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0013] Figure 1 It is a block diagram of a mental health evaluation system based on AI according to an embodiment of the present application.
[0014] Figure 2 It is a block diagram of a displacement state characterization module in a mental health evaluation system based on AI according to an embodiment of the present application.
[0015] Figure 3 It is a block diagram of a muscle point feature extraction unit in a mental health evaluation system based on AI according to an embodiment of the present application.
[0016] Figure 4 It is a block diagram of a displacement comprehensive transmission unit in an AI-based mental health assessment system according to an embodiment of the present application. Detailed implementation manners
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0018] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0019] It is worth noting that in this application, all actions of obtaining signals, information or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.
[0020] In view of the above technical problems, the present application proposes an AI-based mental health assessment system, which collects facial data information (facial video stream) of a target object to be measured through a camera, uses AI-based image recognition and analysis algorithms to perform key frame sampling and facial recognition on the facial data information, then, performs muscle unit partitioning and micro motion state feature extraction on each recognized facial region image, and then performs core muscle displacement state time series transmission on the time series set of the micro motion state features of each facial muscle point after extraction, so as to intelligently judge the mental health type according to the comprehensive representation between the information transmission features of the micro displacement states of each transmitted facial muscle point. The present application uses a camera to collect data and processes it through an AI algorithm, which can quickly complete operations such as key frame sampling, achieving efficient detection. At the same time, it accurately captures the subtle changes in facial muscles, avoids subjective interference of patients, and improves the accuracy of evaluation.
[0021] Figure 1 It is a block diagram of an AI-based mental health assessment system according to an embodiment of the present application. Specifically, as Figure 1As shown, the AI-based mental health assessment system 100 according to an embodiment of the present application includes: a face acquisition module 110, configured to acquire face data information of a target subject to be measured through a camera, where the face data information is a face video stream; an identification module 120, configured to perform face region identification based on key frames on the face data information to obtain a temporal set of face region images; a displacement state characterization module 130, configured to perform comprehensive analysis of the micro-displacements of facial muscle points on the temporal set of face region images to obtain a comprehensive characterization of the micro-displacement state of facial muscle points; and an assessment module 140, configured to obtain a mental health assessment result representing a mental health type label based on the comprehensive characterization of the micro-displacement state of facial muscle points.
[0022] Specifically, the face acquisition module 110 is configured to acquire face data information of a target subject to be measured through a camera, where the face data information is a face video stream. It should be understood that the face contains a large amount of information related to the physiological and psychological states of the human body. When people are in different psychological and emotional states, the facial muscles will make subtle movements. For example, when anxious, they may unconsciously frown, and when happy, there will be corresponding movements at the corners of the mouth. This information exists in a dynamic form, and the face video stream can comprehensively and real-time capture these rich physiological and psychological signals, providing an adequate data basis for subsequent analysis.
[0023] First, in the selection of hardware devices, a suitable camera needs to be selected. Determine the parameters of the camera according to the actual usage scenario and requirements. For example, in scenarios with high requirements for image quality and precise capture of subtle facial muscle changes, select a high-resolution camera to ensure that the captured face video stream can clearly present every detail of the face, such as slight wrinkle changes at the corners of the eyes and subtle movements at the corners of the mouth. These details are crucial for subsequent mental health assessments. At the same time, the frame rate of the camera cannot be ignored. A higher frame rate can ensure the smoothness of the video stream, completely record the dynamic process of facial muscle movement, avoid frame freezing or action loss, and make the captured face data more continuous and accurate.
[0024] When installing and deploying the camera, fully consider the shooting angle and lighting conditions. Install the camera in a suitable position to ensure that the face of the subject to be measured can be completely captured without occlusion. At the same time, avoid direct strong light or too dark light. Direct strong light may cause facial reflection, making some details unable to be clearly presented; too dark light will affect the image quality and increase the difficulty of subsequent image analysis. A stable and suitable shooting environment can be created by adjusting the position of the camera, using a light shield or fill light and other devices.
[0025] When collecting data, it is necessary to communicate appropriately with the object to be measured, inform it of the collection process, and keep it in a natural state to avoid affecting the normal movement of facial muscles due to deliberate performance or nervousness. During the collection process, the camera works continuously, constantly capturing the facial images of the object to be measured and forming a continuous facial video stream. These video streams are stored in specific formats and encoding methods. Common formats include MP4, AVI, etc., and encoding methods include H.264, H.265, etc. These formats and encodings can effectively compress the data size while ensuring the video quality, facilitating subsequent data transmission and storage.
[0026] The collected facial video stream is not directly used for mental health assessment and analysis. A series of preprocessing work is still required. Using video processing technology, noise in the video stream is removed, such as the miscellaneous points and stripes generated by the camera device itself or environmental interference. At the same time, the video stream is stabilized to prevent the instability of the picture caused by the slight shaking of the object to be measured or the jitter of the camera, which may affect the accuracy of subsequent image analysis. Through these steps, it is ensured that the collected facial video stream has good quality, providing a reliable data basis for subsequent AI-based image recognition and analysis algorithms, thereby realizing the accurate assessment of mental health status.
[0027] Specifically, the recognition module 120 is configured to perform key-frame based facial region recognition on the facial data information to obtain a time series set of facial region images. More specifically, in the embodiments of the present application, the recognition module is configured to: perform key-frame sampling on the facial data information to obtain a time series set of facial key frames; perform facial region recognition on each facial key frame in the time series set of facial key frames respectively to obtain the time series set of facial region images. Correspondingly, considering the huge amount of facial video stream data, and that there is a lot of information in many frames of the video stream that is highly similar, especially in static or slowly changing scenarios, if all video frames are processed, it will greatly increase the computational amount, and at the same time consume a large amount of time and resources. However, it also contains the main features and significant change information of the face, and these key features play an important role in the evaluation of mental health. Based on this, in the present application, by performing key-frame sampling on the facial data information to retain the key information while effectively reducing the data amount, a time series set of facial key frames is obtained. It should be understood that key frames usually contain significant change information, such as expression changes, muscle micro-movements, etc., so that subsequent key information related to mental health assessment can be obtained more accurately, improving the accuracy of analysis. And the obtained time series set of facial key frames retains the sequential information of the facial state changing over time, and can reflect the dynamic change process of the facial features and expressions of the measured object, thereby helping to more comprehensively and deeply understand the change trend of the mental state of the measured object, and thus more accurately judging the type of mental health. Then, considering that different regions of the face have different characteristics and sensitivities when reflecting mental states. For example, the muscle movements around the eyes may be more related to the tension or relaxation of emotions, while the movements of the corners of the mouth may be more closely related to the positive or negative expression of emotions. Therefore, in order to be able to focus on the facial information of key regions, such as key feature points of the eyes, eyebrows, mouth, nose, etc., in the technical solution of the present application, facial region recognition is performed on each facial key frame in the time series set of facial key frames respectively to obtain the time series set of facial region images. In this way, by recognizing the facial regions, the system can extract more detailed expression features and muscle micro-movement information, thereby realizing high-precision analysis of individual emotions and mental states.
[0028] Specifically, facial region recognition on each facial key frame in the time series set of facial key frames respectively can be implemented in the following manner:
[0029] Before performing facial region recognition, the system first loads a trained facial region recognition model. These models are usually constructed based on deep learning algorithms, such as convolutional neural networks (CNNs). CNNs have powerful feature extraction capabilities and can automatically learn various feature patterns in facial images. To make the model more accurate in recognizing facial regions, a large amount of image data containing different facial expressions, poses, and lighting conditions is used for training. During the training process, the model continuously adjusts its own parameters to optimize the recognition ability of various facial region features.
[0030] When a facial key frame enters the recognition process, the image needs to be preprocessed first. This includes adjusting the size and resolution of the image to meet the input requirements of the recognition model. At the same time, the image is also normalized, mapping the pixel values of the image to a unified range to reduce the impact of factors such as lighting differences on the recognition results. For example, normalizing the pixel values of the image to the interval [0, 1] or [-1, 1].
[0031] The preprocessed facial key frame image is input into the facial region recognition model. The convolutional layer in the model performs convolution operations on the image, extracting low-level features such as edges and textures in the image through different convolutional kernels. As the convolutional layer deepens, these low-level features are gradually combined into higher-level and more abstract features, such as the features of facial regions like eyes, eyebrows, and mouth. The pooling layer downsamples the feature map output by the convolutional layer without losing key information, reducing the data volume and computational complexity.
[0032] During the process of the model processing the image, object detection algorithms are used to locate the facial region. Common object detection algorithms such as YOLO, Faster R-CNN, etc., can quickly and accurately detect the position and category of the target object in the image. In facial region recognition, these algorithms will detect key regions such as eyes, eyebrows, mouth, and nose in the facial key frame and mark their positions and bounding boxes.
[0033] After recognizing the facial region, the system will crop the images of each facial region from the facial key frame image according to the detected bounding box information. These cropped images form the facial region images. Arranging the corresponding cropped facial region images in the time sequence order of the facial key frames forms a time sequence set of facial region images. This set not only contains the image information of each facial region in each facial key frame but also retains the time order, providing an important data basis for subsequent analysis of the micro-motion characteristics of facial muscle points over time.
[0034] To improve the accuracy and stability of facial region recognition, the system also employs some post-processing techniques. For example, the Non-Maximum Suppression (NMS) algorithm is used to remove overlapping bounding boxes, ensuring that each facial region is accurately recognized only once. At the same time, some prior knowledge and rules are combined to verify and correct the recognition results. For instance, based on the relative positional relationship of facial features, it is judged whether the recognized facial region is reasonable. If an unreasonable situation is found, corresponding adjustments or re-recognition will be carried out. In actual application scenarios, various complex situations may also be faced, such as facial occlusion, overly exaggerated expressions, etc. To address these issues, researchers continuously improve and optimize facial region recognition technology. On the one hand, by increasing the diversity of training data, the model is enabled to learn more facial features in different situations. On the other hand, more advanced algorithms are developed to improve the model's adaptability to complex situations.
[0035] Specifically, the displacement state characterization module 130 is configured to perform a comprehensive analysis of the micro-movement displacements of facial muscle points on the temporal sequence set of the facial region images to obtain a comprehensive characterization of the micro-movement displacement state of facial muscle points. Figure 2 It is a block diagram of the displacement state characterization module in the AI-based mental health assessment system according to an embodiment of the present application. Specifically, as Figure 2 shown, the displacement state characterization module 130 includes: a muscle point feature extraction unit 131, configured to extract micro-movement features of muscle points from the temporal sequence set of the facial region images to obtain a temporal sequence set of micro-movement state features of multiple facial muscle points; and a displacement comprehensive transfer unit 132, configured to perform a comprehensive transfer of the temporal sequence of the core facial muscle point displacements on the temporal sequence set of the micro-movement state features of multiple facial muscle points to obtain the comprehensive characterization of the micro-movement displacement state of facial muscle points.
[0036] Specifically, in the embodiment of the present application, the muscle point feature extraction unit 131 is configured to extract micro-movement features of muscle points from the temporal sequence set of the facial region images to obtain a temporal sequence set of micro-movement state features of multiple facial muscle points. Figure 3 It is a block diagram of the muscle point feature extraction unit in the AI-based mental health assessment system according to an embodiment of the present application. Specifically, as Figure 3As shown, more specifically, in the embodiment of the present application, the muscle point feature extraction unit 131 includes: a muscle unit partitioning subunit 1311, configured to perform muscle unit partitioning on each facial area image in the time series set of the facial area images respectively to obtain a time series set of multiple facial muscle unit area images; a muscle point micro-motion state feature extraction subunit 1312, configured to respectively pass each facial muscle unit area image in the time series set of the multiple facial muscle unit area images through a muscle point micro-motion feature extractor based on a dilated convolutional neural network model to obtain a time series set of multiple facial muscle point micro-motion state feature vectors as the time series set of the multiple facial muscle point micro-motion states.
[0037] Specifically, the muscle unit partitioning subunit 1311 is configured to perform muscle unit partitioning on each facial area image in the time series set of the facial area images respectively to obtain a time series set of multiple facial muscle unit area images. It should be understood that considering that different facial expressions (such as smiling, frowning, staring, etc.) are controlled by specific facial muscle groups. And different facial muscles have different functions and response patterns when expressing emotions and psychological states. For example, the contraction of the corrugator supercilii may indicate anxiety or thinking, and the movement of the orbicularis oculi is related to the tension or relaxation of emotions. The facial area image contains a variety of muscle movement information, but it is difficult to accurately distinguish the specific activities of different muscles through overall analysis. Based on this, the present application can more clearly and accurately capture the movement details of each muscle by performing muscle unit partitioning on each facial area image in the time series set of the facial area images to subdivide the complex facial muscle movement into each small muscle unit, and obtain a time series set of multiple facial muscle unit area images. That is, the movement of each muscle unit may contain specific psychological information. After partitioning, the micro-motion state features can be extracted for each muscle unit area image to obtain more accurate data such as muscle vibration amplitude, frequency, displacement, etc., so as to more sensitively capture the subtle manifestations of psychological state changes in muscle movement.
[0038] Particularly, performing muscle unit partitioning on each facial area image in the time series set of the facial area images respectively can be implemented in the following manner:
[0039] After obtaining the facial region image, a facial muscle localization algorithm based on computer vision technology is used to determine the approximate positions of facial muscles. These algorithms are trained based on facial anatomy knowledge and a large amount of sample data, and can identify the positions and ranges of different facial muscles. For example, by learning a large number of facial images with different expressions and individuals, the algorithm can master the characteristics and position information of the main facial muscles such as the orbicularis oculi muscle and the orbicularis oris muscle in different states. Using these trained algorithms, the positions of each muscle are initially marked in the facial region image. These algorithms can quickly and accurately identify the orbicularis oculi muscle around the eyes, the corrugator supercilii muscle and frontalis muscle that control eyebrow movement, the orbicularis oris muscle and zygomatic major muscle around the mouth, etc. The algorithm extracts and analyzes the features in the image, judges which regions belong to specific muscles, and marks their positions and approximate ranges.
[0040] After initially locating the facial muscle regions, semantic segmentation technology is further used to achieve more refined muscle unit partitioning. Semantic segmentation aims to classify each pixel in the image into a specific category. For facial muscle partitioning, it means dividing each pixel into the corresponding muscle unit category. Commonly used semantic segmentation models, such as U-Net, DeepLab, etc., are trained on a large number of facial images with accurate muscle unit annotations, and learn the detailed features and boundary information of facial muscles. During the training process, the model continuously adjusts its own parameters to optimize the recognition ability of different muscle units. When processing a new facial region image, the model classifies each pixel in the image according to the learned features, thereby accurately dividing the facial muscle region into individual muscle units. For example, for the orbicularis oculi muscle, sub-units with different functions such as the orbital part and the palpebral part can be subdivided; for the orbicularis oris muscle, muscle tissues in different parts can also be distinguished. These subdivided muscle units are crucial for accurately capturing the subtle movement characteristics of the muscles.
[0041] In addition, considering the dynamic nature of facial muscle movement, when processing the time series set of facial region images, a time series analysis method is adopted. The movement changes of muscle units between adjacent frames are analyzed to track the movement trajectories of the muscles. For example, when a person smiles, the orbicularis oris muscle and zygomatic major muscle, etc. will move in coordination. By analyzing each frame image in the time series set, information such as the movement sequence and amplitude change of these muscles can be observed. This not only helps to more accurately divide the muscle unit regions, but also provides richer data for subsequent analysis of the relationship between muscle movement and mental state.
[0042] After completing the partitioning of muscle units, in the chronological order of facial region images, the muscle unit region images divided in each image are arranged in sequence to form a chronological set of multiple facial muscle unit region images. These sets record the state changes of facial muscles at different times, providing crucial data support for subsequent extraction of the micro-motion state features of facial muscle points, thus playing an important role in the AI-based mental health assessment system and helping to more accurately evaluate the mental health status of individuals.
[0043] Specifically, the micro-motion state feature extraction sub-unit 1312 is used to respectively pass each facial muscle unit region image in the chronological set of the multiple facial muscle unit region images through a muscle point micro-motion feature extractor based on a dilated convolutional neural network model to obtain a chronological set of multiple facial muscle point micro-motion state feature vectors as the chronological set of the multiple facial muscle point micro-motion states. Correspondingly, considering that the micro-motion information of muscle points contained in facial muscle unit region images is very subtle and complex, and these micro-motions often involve extremely small displacements and changes. Traditional feature extraction methods are difficult to accurately capture and effectively analyze. The dilated convolutional neural network model has unique advantages. It can expand the receptive field of the convolutional kernel without increasing the number of parameters and computational complexity, enabling it to better capture the micro-motion features of facial muscle points at different scales. Therefore, in the technical solution of this application, each facial muscle unit region image in the chronological set of the multiple facial muscle unit region images is respectively passed through a muscle point micro-motion feature extractor based on a dilated convolutional neural network model to extract the subtle change features in the facial muscle unit region image, obtaining a chronological set of multiple facial muscle point micro-motion state feature vectors. In particular, compared with ordinary convolution, dilated convolution can obtain richer context information. For the weak motion changes of muscle points in facial muscle unit region images, the dilated convolutional neural network can more sensitively perceive and extract them without missing key subtle features, so as to better achieve high-precision recognition of individual emotional states.
[0044] Specifically, in the embodiments of the present application, the displacement comprehensive transfer unit 132 is configured to perform core facial muscle point displacement time series comprehensive transfer on the time series set of the micro-motion state features of the multiple facial muscle points to obtain the comprehensive representation of the micro-motion displacement state of the facial muscle points. In particular, considering that the micro-motion of facial muscle points is a dynamic process, there is a close time series correlation between the micro-motion features of muscle points at different time points. That is, the micro-motion state of facial muscle points changes dynamically at different moments, and the motion features of each muscle point contain important information over time. Taking the corrugator supercilii muscle as an example, within a period of time, the degree of its contraction and relaxation changes over time. Therefore, in order to clarify the time series trend of muscle point movement, such as whether it is gradually tense or gradually relaxed, and then analyze the potential connection between this change and the mental state. For example, anxiety may start from mild tension and gradually develop into obvious frowning or teeth clenching. The present application performs time series information transfer of muscle displacement states based on core information focusing on each time series set of the micro-motion state feature vectors of the multiple facial muscle points respectively to obtain multiple micro-motion displacement state information transfer feature vectors of the facial muscle points. In this way, the information that is truly significant for judging mental health can be screened out, key features can be highlighted, interference can be excluded, so as to better utilize the time series correlation between the micro-motion features of muscle points at each time point and explore the law of mental state changes behind muscle movement.
[0045] Figure 4 It is a block diagram of the displacement comprehensive transfer unit in the AI-based mental health assessment system according to the embodiments of the present application. Specifically, as Figure 4 shown, the displacement comprehensive transfer unit 132 includes: a facial muscle point micro-motion displacement transfer subunit 1321, configured to perform time series information transfer of muscle displacement states based on core information focusing on each time series set of the micro-motion state feature vectors of the multiple facial muscle points respectively to obtain multiple micro-motion displacement state information transfer feature vectors of the facial muscle points; a facial muscle point micro-motion displacement global analysis subunit 1322, configured to perform comprehensive analysis of the micro-motion displacement of all facial muscle points on the multiple micro-motion displacement state information transfer feature vectors to obtain a comprehensive representation vector of the micro-motion displacement state of the facial muscle points as the comprehensive representation of the micro-motion displacement state of the facial muscle points.
[0046] Specifically, the facial muscle point micro-motion displacement transfer subunit 1321 is configured to perform muscle displacement state time-series information transfer based on core information focusing on each time-series set of facial muscle point micro-motion state feature vectors in the time-series set of the multiple facial muscle point micro-motion state feature vectors, so as to obtain multiple facial muscle point micro-motion displacement state information transfer feature vectors. In particular, considering that the micro-motion of facial muscle points is a dynamic process, there is a strong time-series correlation between the micro-motion features of muscle points at different time points. That is, the micro-motion state of facial muscle points changes dynamically at different moments, and the motion features of each muscle point contain important information over time. Taking the corrugator supercilii muscle as an example, within a certain period of time, the degree of its contraction and relaxation changes over time. Therefore, in order to clarify the time-series trend of muscle point movement, such as whether it is gradually tense or gradually relaxed, and then analyze the potential connection between this change and the mental state. For example, anxiety may start from mild tension and gradually develop into obvious frowning or teeth clenching. In this application, muscle displacement state time-series information transfer based on core information focusing is performed on each time-series set of facial muscle point micro-motion state feature vectors in the time-series set of the multiple facial muscle point micro-motion state feature vectors to obtain multiple facial muscle point micro-motion displacement state information transfer feature vectors. In this way, the information that is truly significant for judging mental health can be screened out, key features can be highlighted, interference can be excluded, and thus the time-series correlation between the micro-motion features of muscle points at each time point can be better utilized to explore the law of mental state changes behind muscle movement.
[0047] Specifically, in the embodiment of this application, the facial muscle point micro-motion displacement transfer subunit 1321 includes: a facial muscle point micro-motion state coarse-grained coding secondary subunit, configured to perform coarse-grained aggregation of micro-motion state information on the time-series set of facial muscle point micro-motion state feature vectors to obtain a facial muscle point micro-motion state feature coarse-grained aggregation coding vector; a facial muscle point micro-motion state feature fine-grained compensation secondary subunit, configured to perform fine-grained weight compensation of the muscle point micro-motion state on the time-series set of facial muscle point micro-motion state feature vectors based on the facial muscle point micro-motion state feature coarse-grained aggregation coding vector to obtain a facial muscle point micro-motion state feature fine-grained compensation aggregation coding vector; and a coarse-fine grain interaction analysis secondary subunit, configured to perform interaction analysis on the facial muscle point micro-motion state feature fine-grained compensation aggregation coding vector and the facial muscle point micro-motion state feature coarse-grained aggregation coding vector to obtain the facial muscle point micro-motion displacement state information transfer feature vector.
[0048] Specifically, the facial muscle point micro-motion state coarse-grained coding secondary subunit is configured to perform coarse-grained aggregation of micro-motion state information on the time-series set of facial muscle point micro-motion state feature vectors to obtain a facial muscle point micro-motion state feature coarse-grained aggregation coding vector, which can be expressed by the formula:
[0049]
[0050] Among them, is the time series set of the micro-motion state feature vectors of the facial muscle points, , , and are respectively the 1st, 2nd, th, and th micro-motion state feature vectors of the facial muscle points in the time series set of the micro-motion state feature vectors of the facial muscle points, and are respectively the maximum value and the minimum value of taking , is the th micro-motion state feature aggregation value of the facial muscle points in the time series set of the micro-motion state feature aggregation values of the facial muscle points, is the normalization function, is the th normalized micro-motion state feature aggregation value of the facial muscle points in the time series set of the normalized micro-motion state feature aggregation values of the facial muscle points, is the number of vectors in the , is the coarse-grained aggregation coding vector of the micro-motion state features of the facial muscle points.
[0051] It should be understood that the data volume of the time series set of the micro-motion state feature vectors of the facial muscle points is huge and complex, and it is difficult to grasp the key information directly. By performing micro-motion state information kernel coarse-grained aggregation on the time series set of the micro-motion state feature vectors of the facial muscle points, the concept of information kernel can be utilized. Just like constructing a macroscopic framework in complex data, a large amount of micro-motion information of the facial muscle points can be effectively integrated in a high-dimensional space, and a compressive model of the global features can be built, so as to quickly capture the overall feature trend, obtain the coarse-grained aggregation coding vector of the micro-motion state features of the facial muscle points, make the subsequent calculation and analysis focus on the key information, and provide a basis for the subsequent in-depth analysis.
[0052] More specifically, in the embodiments of the present application, the fine-grained compensation secondary subunit of the facial muscle point micro-motion state feature includes: a facial muscle point micro-motion state compensation factor calculation tertiary subunit, configured to calculate the kernel convergence compensation factors of each facial muscle point micro-motion state feature vector in the time series set of the facial muscle point micro-motion state feature vectors relative to the facial muscle point micro-motion state feature coarse-grained convergence coding vector to obtain a time series set of facial muscle point micro-motion state feature kernel convergence compensation factors; a weight factor calculation tertiary subunit, configured to perform micro-motion state gating function compensation dominance on the time series set of the facial muscle point micro-motion state feature kernel convergence compensation factors to obtain a time series set of facial muscle point micro-motion state feature kernel convergence compensation weight factors; a compensation convergence tertiary subunit, configured to perform muscle point micro-motion state fine-grained time series dynamic compensation convergence on the time series set of the facial muscle point micro-motion state feature kernel convergence compensation weight factors, the facial muscle point micro-motion state feature coarse-grained convergence coding vector, and the time series set of the facial muscle point micro-motion state feature vectors to obtain the facial muscle point micro-motion state feature fine-grained compensation convergence coding vector.
[0053] More specifically, the facial muscle point micro-motion state compensation factor calculation tertiary subunit is configured to: respectively perform deep extraction on the facial muscle point micro-motion state feature vector and the facial muscle point micro-motion state feature coarse-grained convergence coding vector to obtain a facial muscle point micro-motion state deep feature vector and a facial muscle point micro-motion state feature coarse-grained convergence deep coding vector; calculate a facial muscle point micro-motion state feature differential compensation coding vector between the facial muscle point micro-motion state deep feature vector and the facial muscle point micro-motion state feature coarse-grained convergence deep coding vector; multiply the facial muscle point micro-motion state feature differential compensation coding vector by a facial muscle point micro-motion state weight matrix and then perform element-wise addition with a facial muscle point micro-motion state bias value to obtain a facial muscle point micro-motion state feature differential compensation correction vector; multiply the facial muscle point micro-motion state feature differential compensation correction vector by a scoring weight vector to obtain the facial muscle point micro-motion state feature kernel convergence compensation factor corresponding to the facial muscle point micro-motion state feature vector. The above process is represented by the formula:
[0054]
[0055] Wherein, is the th facial muscle point micro-motion state feature vector in the time series set of the facial muscle point micro-motion state feature vectors, is the facial muscle point micro-motion state feature coarse-grained convergence coding vector, is point convolution coding, is an activation function, and are respectively and the corresponding weight matrix is the deep feature vector of the micro-motion state of the corresponding facial muscle points is the deep encoding vector of the coarse-grained aggregation of the micro-motion state features of the facial muscle points is subtraction by position points is the absolute value operation is and the differential compensation correction vector of the micro-motion state features of the facial muscle points between is the corresponding weight matrix of the micro-motion state of the facial muscle points is matrix multiplication is the corresponding bias value of the micro-motion state of the facial muscle points is the corresponding scoring weight vector is the th facial muscle point micro-motion state feature kernel aggregation compensation factor in the time series set of the facial muscle point micro-motion state feature kernel aggregation compensation factors
[0056] Correspondingly, considering that although the coarse-grained aggregation encoding vector can reflect the overall features, in this process, the unique subtle change information of each muscle point may be weakened or lost. By calculating the kernel aggregation compensation factors of each facial muscle point micro-motion state feature vector in the time series set of the facial muscle point micro-motion state feature vectors relative to the coarse-grained aggregation encoding vector of the facial muscle point micro-motion state features, it is possible to measure the difference between the micro-motion state feature vector of a single muscle point at each moment and the overall summary vector, and retrieve the personalized information ignored in the global summary through this difference, providing local supplementation for the subsequent accurate characterization of the micro-motion features of the muscle points
[0057] Specifically, in response to the two-norm of the deep feature vector of the micro-motion state of the facial muscle points being less than the two-norm of the coarse-grained aggregation deep encoding vector of the micro-motion state features of the facial muscle points, the logarithmic function value obtained by taking the logarithm to the base 2 after adding a constant one to the ratio between the two-norm of the deep feature vector of the micro-motion state of the facial muscle points and the two-norm of the coarse-grained aggregation deep encoding vector of the micro-motion state features of the facial muscle points is used as the bias value of the micro-motion state of the facial muscle points; in response to the two-norm of the deep feature vector of the micro-motion state of the facial muscle points being greater than or equal to the two-norm of the coarse-grained aggregation deep encoding vector of the micro-motion state features of the facial muscle points, the ratio obtained by dividing the two-norm of the deep feature vector of the micro-motion state of the facial muscle points by the two-norm of the coarse-grained aggregation deep encoding vector of the micro-motion state features of the facial muscle points is used as the bias value of the micro-motion state of the facial muscle points. This process is expressed by the formula as follows
[0058]
[0059] Among them, is the deep feature vector of the micro motion state of the corresponding facial muscle point, is the deep coding vector of the coarse-grained aggregation of the micro motion state features of the facial muscle point, is to calculate the two-norm of the vector, is the logarithmic function value with base 2, is the offset value of the micro motion state of the corresponding facial muscle point.
[0060] That is to say, here, for the deviation compensation between the micro motion state feature vector of the facial muscle point and the coarse-grained aggregation coding vector of the micro motion state feature of the facial muscle point, it can be measured by quantifying the regret metric based on the information kernel compression hypothesis in the kernel aggregation decision process, that is, the game-theoretic counterfactual regret value, to measure the performance deviation of the kernel aggregation strategy as a scenario strategy. Specifically, through the norm representation of the vector, a normalized decision point loss description based on the policy action, that is, the vector norm representation of the micro motion state feature vector of the facial muscle point and the coarse-grained aggregation coding vector of the micro motion state feature of the facial muscle point, is provided for the counterfactual regret value. Then, for the possible differences in the vector distribution action game scenarios, the compensation rule correction of the node personalized information is carried out respectively with the information distribution degree of the regret value and the relative distribution amplitude of the regret value, so as to consider the personalized information of the micro motion state of the facial muscle point as the unselected action in the decision-making, and perform the offset compensation in the way of assuming its potential benefit based on the information kernel aggregation hypothesis.
[0061] Specifically, the weight factor calculation three-level sub-unit is used to perform explicit compensation of the micro motion state gating function on the time series set of the micro motion state feature kernel aggregation compensation factor of the facial muscle point to obtain the time series set of the micro motion state feature kernel aggregation compensation weight factor of the facial muscle point. This process is expressed by the formula:
[0062]
[0063] Among them, is the th facial muscle point micro motion state feature kernel aggregation compensation factor in the time series set of the facial muscle point micro motion state feature kernel aggregation compensation factor, is to perform gating compensation on , is the preset threshold, is the th facial muscle point micro motion state feature kernel aggregation compensation weight factor in the time series set of the facial muscle point micro motion state feature kernel aggregation compensation weight factor.
[0064] It should be understood that not all the information contained in the nuclear convergence compensation factor is equally important, and there may be some redundant information caused by measurement errors or irrelevant factors. By performing micro-motion state gating function compensation on the time-series set of the micro-motion state feature nuclear convergence compensation factor of the facial muscle points, it is possible to screen and weight the nuclear convergence compensation factor based on the ability of the gating function to dynamically select information under non-linear constraints, highlighting those key differential information that can truly reflect the relationship between the micro-motion of the muscle points and the mental health state, and suppressing useless or interfering information, so that the model can more accurately grasp the effective information and obtain the time-series set of the micro-motion state feature nuclear convergence compensation weight factors of the facial muscle points.
[0065] Specifically, the compensation convergence three-level subunit is used to perform fine-grained time-series dynamic compensation convergence on the time-series set of the micro-motion state feature nuclear convergence compensation weight factors of the facial muscle points, the coarse-grained convergence coding vector of the micro-motion state feature of the facial muscle points, and the time-series set of the micro-motion state feature vectors of the facial muscle points to obtain the fine-grained compensation convergence coding vector of the micro-motion state feature of the facial muscle points. This process is expressed by the formula:
[0066]
[0067] Among them, is the th micro-motion state feature vector of the facial muscle points in the time-series set of the micro-motion state feature vectors of the facial muscle points, is the th micro-motion state feature nuclear convergence compensation weight factor in the time-series set of the micro-motion state feature nuclear convergence compensation weight factors of the facial muscle points, is the coarse-grained convergence coding vector of the micro-motion state feature of the facial muscle points, is the fine-grained compensation convergence coding vector of the micro-motion state feature of the facial muscle points.
[0068] Correspondingly, considering that in order to comprehensively and accurately describe the micro-motion state of the facial muscle points, it is necessary to comprehensively consider the overall features (coarse-grained convergence coding vector), the original detailed features (time-series set of the micro-motion state feature vectors of the facial muscle points), and the screened and weighted local differential features (time-series set of the nuclear convergence compensation weight factors). Through the fine-grained time-series dynamic compensation convergence of the micro-motion state of the muscle points, it is possible to dynamically adjust the expression of the micro-motion features of the muscle points according to the weights of this information at different times, so as to generate a fine-grained feature vector that better fits the actual muscle movement state.
[0069] Specifically, the coarse-grained and fine-grained interactive analysis secondary subunit is used to interactively analyze the fine-grained compensation convergence coding vector of the facial muscle point micro-motion state feature and the coarse-grained convergence coding vector of the facial muscle point micro-motion state feature to obtain the facial muscle point micro-motion displacement state information transmission feature vector. The process is expressed by the formula:
[0070]
[0071] in, is the coarse-grained aggregation encoding vector of facial muscle point micro-motion state features, is the fine-grained compensation aggregation encoding vector of facial muscle point micro-motion state features, and is the weighted hyperparameter, It is the characteristic vector for transmitting the micro-displacement state information of the facial muscle points.
[0072] Finally, since the fine-grained compensation convergence encoding vector focuses on local details and fine-tuning, and the coarse-grained convergence encoding vector focuses on overall feature grasping, the interactive analysis of the two can comprehensively examine the micro-motion state of facial muscle points from both global and local dimensions, dig out the deep connection and coordinated change information between the whole and the local, and further improve the understanding of the micro-motion displacement state of muscle points, so as to obtain a more comprehensive and accurate feature vector reflecting the relationship between muscle movement and mental health.
[0073] Specifically, in the embodiments of the present application, the global analysis sub-unit 1322 of the micro-displacement of facial muscle points is configured to perform a comprehensive analysis of the micro-displacement of facial muscle points on the global scale on the feature vectors of the transmitted states of the micro-displacements of the multiple facial muscle points to obtain a comprehensive representation vector of the micro-displacement state of the facial muscle points as the comprehensive representation of the micro-displacement state of the facial muscle points. More specifically, in the embodiments of the present application, the global analysis sub-unit of the micro-displacement of facial muscle points is configured to: pass the feature vectors of the transmitted states of the micro-displacements of the multiple facial muscle points through a comprehensive analysis module of the micro-displacement of facial muscle points based on LSTM to obtain the comprehensive representation vector of the micro-displacement state of the facial muscle points. Correspondingly, considering the actual situation of facial muscle movement, the movements of different muscle points do not occur in isolation, but cooperate with each other and act synergistically to express various expressions and psychological states. For example, when a person is surprised, not only will the eyebrows raise, the eyes widen, but the mouth may also open slightly. The movements of these muscle points are correlated in time and space and jointly constitute the expression of "surprise". This correlation is not only reflected in the coordinated actions of different muscle points at the same moment, but also in the sequence and trend of changes in the movements of muscle points at different times. Based on this, the present application passes the feature vectors of the transmitted states of the micro-displacements of the multiple facial muscle points through a comprehensive analysis module of the micro-displacement of facial muscle points based on LSTM to obtain a comprehensive representation vector of the micro-displacement state of the facial muscle points. It can be understood that LSTM (Long Short-Term Memory network) has a unique structure and gating mechanism, enabling it to effectively capture this context correlation effect. Its memory unit can store information for a long time, and the forget gate, input gate, and output gate can control the inflow, outflow, and retention of information. When processing the feature vectors of the transmitted states of the micro-displacements of multiple facial muscle points, LSTM can remember the movement state information of each muscle point at the previous moment and update and adjust according to the currently input information. In this way, it can comprehensively consider the movement conditions of different muscle points in the time dimension and capture the complex context correlation between them. For example, LSTM can learn how the movement of the eye muscle points affects the subsequent movement of the mouth muscle points during a certain emotional change process and the change law of this mutual influence in the time series. By capturing these context correlation effects, LSTM can effectively integrate the information of multiple facial muscle points and finally obtain a comprehensive representation vector that can comprehensively reflect the micro-displacement state of facial muscle points, providing strong support for subsequent accurate analysis and judgment of psychological states.
[0074] Specifically, the evaluation module 140 is configured to obtain a mental health evaluation result representing a mental health type label based on the comprehensive characterization of the micro-motion displacement state of the facial muscle points. Specifically, in the embodiment of the present application, the evaluation module is configured to: pass the comprehensive characterization vector of the micro-motion displacement state of the facial muscle points through a mental health evaluator based on a classifier to obtain the mental health evaluation result, and the mental health evaluation result is used to represent the mental health type label. That is, the comprehensive characterization vector of the micro-motion displacement state of the facial muscle points obtained by performing displacement comprehensive analysis on the micro-motion displacement state information transfer feature vectors of multiple facial muscle points is classified, so as to intelligently judge the mental health type. It should be understood that the classifier is an effective data classification tool, which can classify and judge these complex vector data based on existing training data and algorithms. Through a large number of sample data training, the classifier can learn the mapping relationship between different facial muscle movement feature vectors and various mental health types. For example, during the training process, the classifier will analyze the comprehensive characterization vectors of the micro-motion displacement states of the facial muscle points of a large number of people with known mental health conditions, so as to summarize the corresponding rules between specific vector patterns and mental health types. In this way, when a new comprehensive characterization vector is input, the classifier can quickly and accurately judge the mental health type to which it belongs based on the learned rules, and convert the complex data into intuitive and easy-to-understand mental health type labels. In particular, the mental health type labels here can be anxiety, inhibition, stress, etc. The mental health type labels can provide preliminary evaluation references for professional psychiatrists and psychological counselors, helping them quickly understand the general mental health status of the tested person, so as to decide whether further in-depth evaluation or corresponding intervention measures are needed. For institutions such as schools and enterprises, these labels can help them conduct preliminary screening of the mental health status of students or employees, timely discover individuals who may have mental problems, so as to arrange subsequent counseling or treatment.
[0075] In summary, the AI-based mental health evaluation system 100 according to the embodiment of the present application is described. It collects the facial data information (facial video stream) of the target tested object through a camera, uses AI-based image recognition and analysis algorithms to perform key frame sampling and facial recognition on the facial data information, and then performs muscle unit partitioning and micro-motion state feature extraction on each recognized facial area image, and then performs core muscle displacement state time series transfer on the time series set of the micro-motion state features of each facial muscle point after extraction, so as to intelligently judge the mental health type according to the comprehensive characterization between the information transfer features of the micro-motion displacement states of each transferred facial muscle point. The present application uses a camera to collect data and processes it through AI algorithms, which can quickly complete operations such as key frame sampling, realizing efficient detection. At the same time, it can accurately capture the subtle changes in facial muscles, avoid subjective interference of patients, and improve the accuracy of evaluation.
[0076] In particular, in another specific embodiment of the present application, in addition to the analysis of facial muscle movement, the optical imaging technology of blood flow information is also an important component. This technology captures the blood flow changes under the facial skin in a non-invasive manner, providing additional physiological index data for mental health assessment. It should be understood that optical imaging technology (such as near-infrared spectroscopy imaging, NIRS) utilizes the interaction between light and tissues to obtain physiological information within the living body. Specifically, this technology emits light of specific wavelengths (usually in the near-infrared range), and these lights penetrate the skin and are absorbed and scattered by the subcutaneous tissues. Different tissue components have different absorption and scattering characteristics for light, so the internal state of the tissues can be inferred by detecting the reflected or transmitted light signals. When an individual experiences emotional fluctuations, the blood flow under the facial skin will change. For example, in a state of anxiety or tension, the sympathetic nervous system will be activated, resulting in vasoconstriction and a decrease in blood flow; while in a state of relaxation or pleasure, the parasympathetic nervous system dominates, blood vessels dilate, and blood flow increases. Therefore, in the present application, through the optical imaging technology of blood flow information, different color signals reflected from the human facial cortex at different spectral wavelengths are collected by a high-frame-rate camera, the characteristics of facial blood information changes are studied, the changes that cannot be distinguished by the human eye are amplified by image temporal filtering, and then through a unique algorithm, various physiological indexes of a person, such as heart rate, blood oxygen, respiration, etc., are calculated.
[0077] Specifically, the calculation of physiological indexes through the optical imaging technology of blood flow information can be achieved in the following way:
[0078] First is the selection and setting of the device. The high-frame-rate camera is the key acquisition device for this technology and needs to have the ability to shoot at a high frame rate. Common frame rates can reach hundreds of frames per second or even higher to ensure that the rapid changes in facial blood information can be accurately captured. At the same time, its spectral response range should match the specific spectral wavelengths used for detection, generally concentrating in the near-infrared spectral region because light in this region can better penetrate skin tissues, and hemoglobin in the blood has obvious absorption and scattering characteristics for light in this region, which is conducive to obtaining rich blood information. In actual application scenarios, the parameters of the camera, such as aperture, shutter speed, and gain, also need to be adjusted according to specific situations to ensure that the collected images have sufficient clarity and contrast. In addition, in order to accurately collect signals of different spectral wavelengths, corresponding optical filters are equipped to screen out light within a specific wavelength range, enabling the camera to focus on the collection of the required spectrum.
[0079] Before collection, it is necessary to ensure that the person to be tested is in a suitable state and environment. The person to be tested should remain quiet and relaxed, avoiding factors such as strenuous exercise and emotional excitement that may cause abnormal fluctuations in facial blood flow and affect the accuracy of the test results. The test environment should be kept as stable as possible, avoiding direct sunlight and electromagnetic interference, to create good conditions for the stable operation of the equipment and signal collection.
[0080] During the collection process, the high-frame-rate camera starts to work, emitting light of a specific spectral wavelength towards the human face. This light penetrates the facial skin and interacts with the blood in the subcutaneous tissue. Hemoglobin in the blood has different absorption and scattering characteristics for light of different wavelengths. When light shines on the blood, part of the light is absorbed by hemoglobin, and part of the light is scattered and finally reflected back from the facial cortex. The high-frame-rate camera continuously collects these reflected light signals at the set high frame rate and converts them into different color signals for recording. Since the intensity and color of the light reflected back at different wavelengths are different, these differences contain information about facial blood, such as the content of hemoglobin in the blood and the blood flow velocity.
[0081] The original image data collected contains a large amount of complex information. Among them, the change characteristics of facial blood information are relatively weak and are mixed with various noise and interference factors. To extract useful blood information, it is necessary to use image temporal filtering technology for processing. Image temporal filtering is based on the continuity and regularity of the change of blood information in the time dimension. By designing a specific filter, filtering operations are performed on the continuously collected image sequence. For example, using median temporal filtering can effectively remove random noise in the image and retain the trend of blood information change; while temporal weighted filtering based on the Gaussian function can perform different degrees of weighting on the image according to the distance in time, highlighting the change of blood information in the recent period more. Through these filtering operations, the change characteristics of blood information can be strengthened, and other interference factors can be suppressed, making the originally weak and indistinguishable change of facial blood information more obvious.
[0082] On the basis of filtering, the image decomposition and magnification technology is used to further process the image. This technology is mainly based on mathematical algorithms. By analyzing and transforming the filtered image, the subtle changes hidden in the image are magnified and displayed. For example, using wavelet transform technology, the image can be decomposed into different frequency sub-bands, and the sub-band containing the change characteristics of blood information can be extracted separately and magnified by adjusting relevant parameters. In this way, those facial images that hardly seem to change to the naked eye can clearly show the dynamic changes of blood information, such as the surging of blood flow and the dilation and contraction of blood vessels, after image decomposition and magnification.
[0083] After filtering, decomposition, and amplification processing, a relatively clear and prominent image sequence of facial blood information changes is obtained. Next, specific algorithms are used to calculate various physiological indicators. Taking heart rate calculation as an example, it is mainly based on the correlation between blood volume changes and heartbeats. By analyzing the curve of blood volume change over time in specific facial regions (such as blood-vessel-rich areas like the tip of the nose and cheeks) in the image, and using signal processing algorithms such as Fourier transform, the blood volume change signal in the time domain is converted to the frequency domain, and the frequency component corresponding to the heart rate is extracted from it, and then the heart rate value is calculated.
[0084] Calculating blood oxygen content is based on the absorption difference of hemoglobin for light of different wavelengths. In the near-infrared spectral region, the absorption coefficients of oxyhemoglobin and deoxyhemoglobin for specific wavelengths of light are different. By collecting the intensity information of light reflected back at different wavelengths and using the Beer-Lambert law to establish a mathematical model, the relative contents of oxyhemoglobin and deoxyhemoglobin are calculated, and thus the blood oxygen saturation index is obtained.
[0085] For respiration detection, it mainly utilizes the relationship between the minute movements of the chest and face during respiration and the resulting changes in facial blood flow. By analyzing the subtle changes in the overall brightness or specific regions in the facial image sequence and combining the hemodynamic change laws caused by respiratory movements, pattern recognition algorithms such as support vector machine (SVM) in machine learning or convolutional neural network (CNN) in deep learning are used to classify and identify these change patterns, thereby inferring indicators such as respiratory rate and respiratory depth.
[0086] During the entire implementation process, it is also necessary to calibrate and verify the calculated physiological indicators. By conducting comparative tests with traditional physiological detection devices (such as electrocardiogram monitors, pulse oximeters, etc.), a large amount of experimental data is collected, a calibration model is established, and the physiological indicators calculated by the optical imaging technology of blood flow information are corrected to improve their accuracy and reliability.
[0087] In this way, the information provided by the two technologies can corroborate and complement each other. In a state of anxiety, not only will specific facial expressions (such as frowning) appear, but there will also be blood flow changes (such as facial blood vessel constriction). By combining these two types of information, the emotional state can be identified more comprehensively and accurately. Through parallel computing, this mutual relationship can be utilized to cross-validate and comprehensively analyze the data. Finally, based on the analysis results of various mental health and emotion-related indicators, the individual's occupational psychological qualities such as confidence, self-control, motivation, calmness, and anxiety are evaluated in an all-round and multi-angle manner.
[0088] Specifically, for the assessment of professional quality and ability, there can be the following application scenarios:
[0089] 1. College student employment career guidance: It can quickly analyze the students' career psychological quality indicators, evaluate their personality traits, and recommend suitable career directions accordingly.
[0090] 2. Enterprise recruitment and job matching: During the enterprise recruitment process, it can conduct a comprehensive career psychological quality assessment of the applicants according to the specific psychological quality requirements of different positions, so as to select the most suitable talents. At the same time, this system can also be applied to the promotion of cadres' positions or political evaluations to ensure that the candidates meet the psychological quality requirements of the new positions.
[0091] 3. Pre-job tests for high-risk special enterprises: For industries such as steelmaking, glass manufacturing, and chemical engineering, conduct daily emotional health stability tests for employees before they go to work to ensure safe production and avoid accident risks caused by personal emotional safety stability problems.
[0092] 4. New recruit job assignment: Conduct psychological assessments on the newly enlisted recruits. Based on the results of their career psychological quality evaluations, assign the recruits with strong self-discipline, high stress resistance, and teamwork spirit to combat troops with high requirements for discipline and teamwork; assign the recruits with high concentration and quick reaction to technical positions to achieve the matching of people and positions and improve the overall combat effectiveness of the military.
[0093] As described above, the AI-based mental health assessment system 100 according to the embodiments of the present application can be implemented in various wireless terminals, such as a server with an AI-based mental health assessment algorithm. In one possible implementation, the AI-based mental health assessment system 100 according to the embodiments of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the AI-based mental health assessment system 100 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the AI-based mental health assessment system 100 can also be one of the numerous hardware modules of the wireless terminal.
[0094] Alternatively, in another example, the AI-based mental health assessment system 100 and the wireless terminal can also be separate devices, and the AI-based mental health assessment system 100 can be connected to the wireless terminal through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0095] The various implementations of the present disclosure have been described above. The above description is exemplary and not exhaustive. Also, it is not limited to the disclosed implementations. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the various implementation manners disclosed herein.
Claims
1. An AI-based mental health assessment system, characterized in that, Including: A face acquisition module, configured to acquire face data information of a target object to be measured through a camera, where the face data information is a face video stream; An identification module, configured to perform face region identification based on key frames on the face data information to obtain a time series set of face region images; A displacement state characterization module, configured to perform comprehensive analysis of the micro-motion displacements of facial muscle points on the time series set of the face region images to obtain a comprehensive characterization of the micro-motion displacement states of facial muscle points. Among them, the displacement state characterization module includes: a muscle point feature extraction unit, configured to perform micro-motion feature extraction of facial muscle points on the time series set of the face region images to obtain a time series set of micro-motion state features of multiple facial muscle points; a displacement comprehensive transfer unit, configured to perform comprehensive transfer of the time series of the displacements of the core facial muscle points on the time series set of the micro-motion state features of multiple facial muscle points to obtain the comprehensive characterization of the micro-motion displacement states of facial muscle points; An evaluation module, configured to obtain a mental health evaluation result representing a mental health type label based on the comprehensive characterization of the micro-motion displacement states of facial muscle points; Among them, the displacement comprehensive transfer unit includes: A facial muscle point micro-motion displacement transfer subunit, configured to perform time series information transfer of muscle displacement states based on core information focusing on each time series set of micro-motion state feature vectors in the time series set of micro-motion state feature vectors of multiple facial muscle points to obtain multiple micro-motion displacement state information transfer feature vectors of facial muscle points; A facial muscle point micro-motion displacement global analysis subunit, configured to perform comprehensive analysis of the micro-motion displacements of all facial muscle points on the multiple micro-motion displacement state information transfer feature vectors to obtain a comprehensive characterization vector of the micro-motion displacement states of facial muscle points as the comprehensive characterization of the micro-motion displacement states of facial muscle points; Among them, the facial muscle point micro-motion displacement transfer subunit includes: A secondary subunit for coarse-grained encoding of the micro-motion states of facial muscle points, configured to perform coarse-grained aggregation of micro-motion state information on the time series set of micro-motion state feature vectors of facial muscle points to obtain a coarse-grained aggregation encoding vector of the micro-motion state features of facial muscle points; A secondary subunit for fine-grained compensation of the micro-motion state features of facial muscle points, configured to perform fine-grained weight compensation of the micro-motion states of facial muscle points on the time series set of micro-motion state feature vectors of facial muscle points based on the coarse-grained aggregation encoding vector of the micro-motion state features of facial muscle points to obtain a fine-grained compensation aggregation encoding vector of the micro-motion state features of facial muscle points; A secondary subunit for interactive analysis of coarse and fine grain sizes, configured to perform interactive analysis on the fine-grained compensation aggregation encoding vector of the micro-motion state features of facial muscle points and the coarse-grained aggregation encoding vector of the micro-motion state features of facial muscle points to obtain the micro-motion displacement state information transfer feature vector; Among them, the secondary subunit for fine-grained compensation of the micro-motion state features of facial muscle points includes: The three - level sub - unit for calculating the compensation factor of the micro - movement state of facial muscle points is used to deeply extract the micro - movement state feature vector of facial muscle points and the coarsely - aggregated encoded vector of the micro - movement state of facial muscle points respectively to obtain the deep - layer feature vector of the micro - movement state of facial muscle points and the coarsely - aggregated deep - layer encoded vector of the micro - movement state of facial muscle points; calculate the differential compensation encoded vector of the micro - movement state of facial muscle points between the deep - layer feature vector of the micro - movement state of facial muscle points and the coarsely - aggregated deep - layer encoded vector of the micro - movement state of facial muscle points; multiply the differential compensation encoded vector of the micro - movement state of facial muscle points by the weight matrix of the micro - movement state of facial muscle points and then perform element - by - element addition with the bias value of the micro - movement state of facial muscle points to obtain the differential compensation corrected vector of the micro - movement state of facial muscle points; multiply the differential compensation corrected vector of the micro - movement state of facial muscle points by the scoring weight vector to obtain the kernel - aggregated compensation factor of the micro - movement state of facial muscle points corresponding to the micro - movement state feature vector of facial muscle points; The three - level sub - unit for calculating the weight factor is used to perform explicit compensation of the gating function of the micro - movement state on the time - series set of the kernel - aggregated compensation factor of the micro - movement state of facial muscle points to obtain the time - series set of the kernel - aggregated compensation weight factor of the micro - movement state of facial muscle points; The three - level sub - unit for compensation and aggregation is used to perform fine - grained time - series dynamic compensation and aggregation of the micro - movement state of muscle points on the time - series set of the kernel - aggregated compensation weight factor of the micro - movement state of facial muscle points, the time - series set of the coarsely - aggregated encoded vector of the micro - movement state of facial muscle points, and the time - series set of the micro - movement state feature vector of facial muscle points to obtain the fine - grained compensation - aggregated encoded vector of the micro - movement state of facial muscle points.
2. The AI-based mental health assessment system according to claim 1, wherein The recognition module is used for: Performing key - frame sampling on the facial data information to obtain a time - series set of facial key - frames; Performing facial region recognition on each facial key - frame in the time - series set of facial key - frames respectively to obtain a time - series set of facial region images.
3. The AI-based mental health assessment system according to claim 2, wherein The muscle - point feature extraction unit includes: The muscle - unit partitioning sub - unit is used to partition each facial region image in the time - series set of facial region images into muscle units respectively to obtain a time - series set of multiple facial muscle - unit region images; The muscle - point micro - movement state feature extraction sub - unit is used to pass each facial muscle - unit region image in the time - series set of multiple facial muscle - unit region images through a muscle - point micro - movement feature extractor based on a dilated convolutional neural network model respectively to obtain a time - series set of multiple facial muscle - point micro - movement state feature vectors as the time - series set of the multiple facial muscle - point micro - movement states.
4. The AI-based mental health assessment system according to claim 3, wherein, In response to the two - norm of the deep - layer feature vector of the micro - movement state of facial muscle points being less than the two - norm of the coarsely - aggregated deep - layer encoded vector of the micro - movement state of facial muscle points, after adding a constant one to the ratio between the two - norm of the deep - layer feature vector of the micro - movement state of facial muscle points and the two - norm of the coarsely - aggregated deep - layer encoded vector of the micro - movement state of facial muscle points, taking the logarithmic function value of the logarithm with base 2 as the bias value of the micro - movement state of facial muscle points; In response to the fact that the two-norm of the deep feature vector of the facial muscle point micro-motion state is greater than or equal to the two-norm of the deep encoded vector of the coarse-grained aggregation of the facial muscle point micro-motion state features, the ratio obtained by dividing the two-norm of the deep feature vector of the facial muscle point micro-motion state by the two-norm of the deep encoded vector of the coarse-grained aggregation of the facial muscle point micro-motion state features is used as the bias value of the facial muscle point micro-motion state.
5. The AI-based mental health assessment system according to claim 4, wherein, The global analysis sub-unit of the facial muscle point micro-motion displacement is configured to: transmit the feature vectors of the micro-motion displacement states of the multiple facial muscle points through the LSTM-based comprehensive analysis module of the facial muscle point micro-motion displacement to obtain the comprehensive characterization vector of the facial muscle point micro-motion displacement state.
6. The AI-based mental health assessment system according to claim 5, characterized in that, The evaluation module is configured to: pass the comprehensive characterization vector of the facial muscle point micro-motion displacement state through a mental health evaluator based on a classifier to obtain the mental health evaluation result, and the mental health evaluation result is used to represent the mental health type label.
Citation Information
Patent Citations
Mental health monitoring and early warning method and system based on hyperspectral video analysis
CN118216914A