Virtual reality-based depression adjuvant therapy system
The virtual reality-based depression treatment system utilizes VR glasses and a data processing module to monitor patients' emotions and pupil movements in real time and dynamically adjust the virtual scene, solving the problem of high cost in traditional CBT treatment and achieving convenient depression treatment.
Patent Information
- Application Number
- CN202610056345.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-02-13
AI Technical Summary
Traditional cognitive behavioral therapy (CBT) for treating depression and anxiety is expensive, and many patients cannot afford regular treatment, resulting in delayed treatment of depression.
This invention provides a virtual reality-based auxiliary treatment system for depression. It utilizes VR glasses, a camera module, and a data processing module to monitor the patient's emotions and pupil movements in real time through an emotion recognition model and a pupil tracking algorithm, and dynamically adjusts the virtual interactive scene to alleviate psychological stress.
It enables convenient treatment of depression without the need for intervention from mental health professionals, reduces treatment costs, and dynamically adjusts virtual scenarios through real-time emotion and pupil monitoring to assist in the treatment of depression.
Smart Images

Figure CN121528447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of optoelectronic applications and virtual reality, and in particular to a virtual reality-based auxiliary treatment system for depression. Background Technology
[0002] In recent years, more than 264 million people of all ages have suffered from depression and anxiety, making it a common mental illness worldwide. Depression and anxiety are serious mental health disorders that severely affect the daily lives of hundreds of millions of people around the world.
[0003] Medications and therapies used to treat symptoms of depression and anxiety are common traditional treatments. Cognitive behavioral therapy (CBT) has proven to be one of the most effective treatments for depression and anxiety. However, receiving regular CBT from a licensed mental health professional can be very expensive. Therefore, many people with depression forgo CBT, resulting in delayed treatment of their depression. Summary of the Invention
[0004] The purpose of this application is to provide a virtual reality-based auxiliary treatment system for depression, which can assist in the treatment of depression, is easy to use, and reduces costs.
[0005] To achieve the above objectives, this application provides the following solution.
[0006] This application provides a virtual reality-based auxiliary treatment system for depression, comprising: VR glasses, a camera module, and a data processing module; the data processing module is integrated into the VR glasses. After the patient wears the VR glasses, the VR glasses are used to play virtual interactive scenes; the camera module is used to capture facial and pupil images of the patient in real time while watching the virtual interactive scenes; the data processing module is used to determine the patient's emotion in each frame based on the real-time captured facial images using an emotion recognition model, and simultaneously to track pupil position changes in each frame based on the real-time captured pupil images using a pupil tracking algorithm to obtain pupil movement change data for each frame, and then fuse and analyze the emotion and pupil movement change data of multiple consecutive frames to output the patient's comprehensive emotion index; the VR glasses are also used to dynamically adjust the parameters of the virtual interactive scene or switch the virtual interactive scene based on the patient's comprehensive emotion index to alleviate the patient's psychological stress and emotions.
[0007] Optionally, the camera module includes: three cameras; one camera for capturing images of the patient's face; and two other cameras embedded in the VR glasses for capturing images of the patient's pupils.
[0008] Optionally, the virtual reality-based depression auxiliary treatment system further comprises a head-mounted lens holder; a camera for shooting patient facial images is fixed on the head-mounted lens holder; and the head-mounted lens holder is worn on the patient's head.
[0009] Optionally, the virtual reality-based depression auxiliary treatment system further comprises a 3D printed shell; and the 3D printed shell is used for packaging the VR glasses and the camera module.
[0010] Optionally, the virtual reality-based depression auxiliary treatment system further comprises a UI interface; and the UI interface is used to display the patient's emotion and pupil movement change data of each frame in the form of text or icons.
[0011] Optionally, the data processing module is a Wildfire Lu Ban Cat 2-V3 development board embedded with an emotion recognition model and a pupil tracking algorithm.
[0012] Optionally, the data processing module comprises an emotion recognition module, a pupil movement recognition module, and a dual-modal data fusion module. The emotion recognition module is used to identify the patient's emotion according to each real-time shot facial image by using a YOLOv8n lightweight neural network model, and output the patient's emotion label and confidence; the YOLOv8n lightweight neural network model adopts a YOLOv8n architecture, is compressed after pruning channels and INT8 quantization of the YOLOv8n architecture, and is obtained after deep learning training; The pupil movement recognition module is used to track the pupil position change according to each real-time shot pupil image by using a Mediapipe pupil tracking algorithm, and obtain the pupil movement change data of each frame; the pupil movement change data comprises 3D pupil coordinates, pupil diameter change amount, and gaze direction; The dual-modal data fusion module is used to synchronize the emotion label and the pupil movement change data by time stamp, and fuse the emotion label and the pupil movement change data synchronized by time stamp by using a weighted decision mechanism, to obtain the fusion data of each frame; and the emotion trend of the patient is analyzed by using a state machine according to the fusion data of continuous multiple frames, to output the comprehensive emotion index of the patient.
[0013] Optionally, in the aspect of tracking the pupil position change according to each real-time shot pupil image by using the Mediapipe pupil tracking algorithm to obtain the pupil movement change data of each frame, the pupil movement recognition module comprises: converting the pupil image into an RGB format and performing normalization processing to obtain a preprocessed pupil image; detecting eye key points by using a Face Mesh model of Mediapipe according to the preprocessed pupil image, and constructing an eye region of interest according to the eye key points; The pupil key points are extracted from the eye region of interest by using the face_landmarks of Mediapipe, and a pupil region is constructed according to the pupil key points; The pupil center is located from the pupil region, and a 3D pupil coordinate is obtained; The key points of the upper, lower, left and right boundaries of the pupil are selected, and the horizontal diameter and the vertical diameter of the pupil are calculated; The horizontal diameter and the vertical diameter are corrected in scale by using the formula The scale-corrected horizontal diameter or vertical diameter is obtained; in the formula, The scale-corrected horizontal diameter or vertical diameter is The horizontal diameter or vertical diameter is The scale factor is The pupil diameter change amount is determined according to the scale-corrected horizontal diameter or vertical diameter of the adjacent frame; The eye contour key points are detected in the preprocessed pupil image; A spherical model is fitted by using the eye contour key points, and the center coordinates of the spherical model are determined as the 3D coordinates of the eyeball center; The line-of-sight vector is determined according to the 3D coordinates of the eyeball center and the 3D pupil coordinates; The line-of-sight vector is converted from the eye coordinate system to the world coordinate system, and the gaze direction is obtained.
[0014] Optionally, the data processing module is further configured to perform abnormal state early warning and trigger the handle vibration and voice guidance when the emotional labels of the patient in continuous multiple frames are all negative emotions and the pupil is abnormal; the pupil abnormality includes that the pupil diameter change amount is greater than a scaling threshold, or the gaze directions of continuous multiple frames are different.
[0015] Optionally, in terms of dynamically adjusting the virtual interaction scene parameters or switching the virtual interaction scene according to the comprehensive emotional index of the patient, the VR glasses comprise: if the comprehensive emotional index of the patient is greater than an emotional index threshold, the virtual interaction scene is switched to a positive scene; if the comprehensive emotional index of the patient is less than or equal to the emotional index threshold, the saturation of the virtual interaction scene is reduced and white noise is increased.
[0016] According to the specific embodiments provided in the application, the application has the following technical effects.
[0017] The application provides a virtual reality-based auxiliary treatment system for depression, wherein a patient with depression wears VR glasses, virtual scenes that can make the patient feel comfortable and warm are played, the system monitors the real-time tracking of facial expressions and pupils of the patient, adjusts virtual interactive scene parameters or switches virtual interactive scenes according to the feedback of the patient at any time, performs psychological counseling on the patient with depression, adjusts the emotion of the patient, and realizes auxiliary treatment of depression; the data processing module is integrated in the VR glasses, so that the intervention of a mental health expert is not needed, the system is convenient to use, and the cost is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0019] Figure 1 A signal transmission schematic diagram of a virtual reality-based auxiliary treatment system for depression provided by the embodiments of the present application.
[0020] Figure 2 A flowchart of a biofeedback mechanism provided by the embodiments of the present application.
[0021] Figure 3 A VR emotion interactive treatment flowchart provided by the embodiments of the present application.
[0022] Figure 4 A hardware interaction flowchart provided by the embodiments of the present application.
[0023] Figure 5 An algorithm processing flowchart provided by the embodiments of the present application.
[0024] Figure 6 A closed-loop interaction flowchart provided by the embodiments of the present application.
[0025] Figure 7 A training set loss function trend diagram provided by the embodiments of the present application.
[0026] Figure 8 A test set confusion matrix diagram provided by the embodiments of the present application.
[0027] Figure 9 An emotion recognition UI interface diagram provided by the embodiments of the present application. DETAILED DESCRIPTION
[0028] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described in order to make the above and other objectives, features and advantages of the present application more apparent. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0029] The above and other objectives, features and advantages of the present application will become more apparent from the following detailed description made with reference to the accompanying drawings and specific embodiments.
[0030] Currently, virtual reality technology can be an effective method for treating depression and anxiety symptoms, and can provide lasting relief for some patients with depression and anxiety symptoms. Virtual reality can combine behavioral activation, cognitive restructuring, social skill training and immersive psychological education to create virtual CBT interventions, which can be treated without the intervention of licensed mental health professionals. This can allow patients with depression and anxiety to receive CBT and its benefits without the intervention of mental health professionals.
[0031] Therefore, in one example embodiment, as shown in Figure 1 The present application provides a virtual reality-based depression auxiliary treatment system, which includes a VR glasses, a camera module and a data processing module. The data processing module is integrated in the VR glasses. After the patient wears the VR glasses, the VR glasses are used to play a virtual interactive scene; the camera module is used to capture real-time facial images and pupil images of the patient during the process of watching the virtual interactive scene; the data processing module is used to determine the emotion of each frame of the patient according to each frame of the real-time captured facial image by using an emotion recognition model, and to track the pupil position change by using a pupil tracking algorithm according to each frame of the real-time captured pupil image, to obtain the pupil motion change data of each frame, and then to fuse and analyze the emotion and the pupil motion change data of continuous multiple frames to output a comprehensive emotion index of the patient; the VR glasses are further used to dynamically adjust the virtual interactive scene parameters or switch the virtual interactive scene according to the comprehensive emotion index of the patient, so as to relieve the psychological pressure and emotion of the patient.
[0032] The VR glasses are used to play an immersive virtual interactive scene (such as natural scenery, social interaction, etc.) for virtual reality healing to relieve the psychological pressure and emotion of the patient.
[0033] The system of the present application is based on a biofeedback mechanism, which realizes the auxiliary treatment of patients with depression by combining virtual reality technology with physiological signal monitoring. Biofeedback is a technology that uses precise tools to explore and amplify physiological change information of the human body, converts these information into signals easy to understand, and trains under the guidance of medical personnel, so that patients learn to use the processed signals from themselves to consciously control various physiological and pathological processes in the body, promote functional recovery, and thus achieve the purpose of treating diseases. As shown in Figure 2 The biofeedback mechanism includes physiological processes, signal processing and feedback, learning and regulation loop. The specific process is: biological signal sensors detect physiological activities (such as heart rate, brain waves, muscle tension) and convert them into electrical signals; the signal processing device amplifies, analyzes and digitizes the electrical signals, and the "feedback" signals presented to the user include visual, auditory or tactile aspects; the user consciously tries to change their thinking or behavior; the brain sends new regulation instructions and affects physiological activities through the nervous system. The ultimate goal of the user consciously trying to change their thinking or behavior is self-regulation without device assistance. In the present application, this mechanism is applied to emotion monitoring and pupil movement recognition, by capturing the psychological changes of patients during virtual scene interaction, physiological data of emotional fluctuations and pupil responses are obtained.
[0034] The system collects the patient's facial expressions and pupil movement information in real time by installing a camera module. The data processing module uses an emotion recognition model to analyze facial expressions and identify eight core emotional states (angry, disgusted, afraid, happy, neutral, sad, surprised, and contemptuous); based on a pupil tracking algorithm, the pupil position changes are tracked to capture the user's attention and unconscious reactions. After processing these data, intuitive biofeedback signals are formed to dynamically adjust the content and intensity of the virtual scene. For example, when the system detects that the patient's mood is low or the pupil response is weakened, it will automatically switch to a more positively stimulating scene, such as natural scenery or warm interaction, to stimulate the patient's interest and pleasure.
[0035] This closed-loop biofeedback mechanism not only helps patients more intuitively perceive their emotional state, but also gradually improves their psychological regulation ability through continuous interactive training. At the same time, the data recorded by the system can provide objective treatment reference for doctors to assist in developing individualized intervention programs, thereby improving treatment effectiveness. In the future, with the optimization of algorithms and the deepening of clinical verification, this technology is expected to become an important tool for the auxiliary treatment of depression.
[0036] As an optional implementation, the camera module includes: three cameras. One camera is used to take pictures of the patient's face; the other two cameras are implanted in the VR glasses to take pictures of the patient's pupils.
[0037] In the implementation, the virtual reality-based depression auxiliary treatment system further comprises a head-mounted lens support, and a camera for shooting a patient's facial image is fixed on the head-mounted lens support; and the head-mounted lens support is worn on the patient's head.
[0038] As an optional implementation, the virtual reality-based depression auxiliary treatment system further comprises a 3D-printed shell, and the 3D-printed shell is used for packaging the VR glasses and the camera module.
[0039] As an optional implementation, the virtual reality-based depression auxiliary treatment system further comprises a UI interface, and the UI interface is used for displaying, in the form of text or icons, the patient's emotion of each frame and the pupil movement change data of each frame.
[0040] Based on the VR virtual reality + multi-modal emotion recognition + real-time biological feedback technology, the application constructs a wearable and intelligent depression auxiliary treatment system. The system is divided into three modules of a hardware layer, an algorithm layer and an interaction layer, and realizes emotion monitoring, virtual scene adjustment and remote medical collaboration, such as Figure 3 The VR emotion interaction treatment flowchart is shown. Figure 3 The flowchart is divided into five stages: 01, the patient wears the VR glasses to perform virtual reality healing, and the system detects the patient's expression changes in real time; 02, a pupil tracking camera is arranged on the inner side of the head-mounted device, the line of sight / pupil changes are collected, and emotion recognition is facilitated; 03, multi-source data fusion analysis, the system outputs the emotional state; 04, emotional result synchronous feedback, the doctor end monitors the results in real time; 05, the doctor adjusts the treatment content or the stimulation scene according to the results, and realizes personalized intervention.
[0041] 1. Hardware layer (1) The hardware device composition is shown in Table 1.
[0042] Table 1 Hardware device composition
[0043] The system of the application is composed of five core hardware components of VR glasses, a head-mounted support, an embedded development board, a camera module and a 3D-printed shell, and realizes the depression auxiliary treatment function through highly integrated design.
[0044] The model of the VR glasses is HTC Vive Pro 2, which mainly provides an immersive virtual treatment scene (such as natural scenery, social interaction, etc.), and realizes dynamic rendering of the scene through SteamVR SDK. At the same time, a high refresh rate display screen (120Hz) and a wide field of view (120 ° ) are built-in to reduce dizziness and improve patient comfort.
[0045] The data processing module is composed of a facial expression recognition model based on YOLOv8 and a pupil movement monitoring module based on FaceMesh in Mediapipe, and is deployed on a Wildfire Rubicon Cat 2-V3 development board, so the data processing module is a Wildfire Rubicon Cat 2-V3 development board embedded with an emotion recognition model and a pupil tracking algorithm. That is, the embedded development board uses a Wildfire Rubicon Cat 2-V3 development board for real-time running of the YOLOv8n emotion recognition model and the Mediapipe pupil tracking algorithm, while receiving camera data through the MIPI interface and outputting analysis results to the VR device and the doctor's end through USB 3.0.
[0046] The Wildfire Rubicon Cat 2-V3 development board uses the Sunway RK3568 chip, which provides powerful edge computing capabilities. It has perfect development documents and community support, significantly shortening the development cycle. It has rich interfaces (USB 3.0 / GPIO / MIPI, etc.), which is convenient for system integration. And has been widely used in education, industrial control and other fields, and its stability has been verified by the market. The RK3568 chip supports mainstream frameworks such as TensorFlow Lite / PyTorch, deploys built-in hardware codecs, and the typical power consumption is less than 3W, meeting the energy efficiency requirements of mobile devices.
[0047] The camera module is a 1200W 4K75 ° The camera module is a 1200W 4K75 ° The camera module is a 1200W 4K75
[0048] The technical specifications of the OV2720 camera in Table 1: the sensor is 1 / 3 inch CMOS, 5 million effective pixels, the resolution is 2592x1944@30fps (supports ROI region cropping), the optical property is 76 ° The camera module is a 1200W 4K75
[0049] The 3D printed shell is made of medical grade resin, which has undergone 5 iterations. The design adopts integrated packaging of all hardware, with reserved heat dissipation channels (temperature rise <5°C). And the surface is frosted to avoid glare interference with the camera work, weighing only 380g.
[0050] The headband features a lightweight aluminum alloy frame and adjustable straps, designed to secure three OV2720 cameras for accurate facial data capture. It supports quick assembly and disassembly, adapts to different head sizes, and provides even pressure distribution (less than 200g total load). The headband lens holder supports a 12MP 4K 75mm camera. ° The distortion-free autofocus lens module allows the patient to directly view a 12-megapixel 4K 75mm lens. ° One camera in the distortion-free autofocus lens module extends 20 centimeters directly in front of the patient's face, enabling a 12-megapixel 4K 75... ° One of the cameras in the distortion-free autofocus lens module has a field of view that covers the patient's entire face, thereby capturing the patient's facial expressions.
[0051] (2) Hardware interaction process The hardware interaction process achieves a complete closed loop from data acquisition to feedback, such as Figure 4 As shown, after the patient puts on the integrated device (VR glasses), the device's three high-precision cameras begin to capture facial and pupil images in real time (corresponding to...). Figure 4 (Data collected by the camera). These images are transmitted to the Wildfire Luban Cat embedded development board via the MIPI interface, where the onboard NPU accelerates the execution of the YOLOv8n emotion recognition model and the Mediapipe pupil tracking algorithm, completing the emotion state analysis within 50ms (corresponding to...). Figure 4 The development board processes the data in real time. The processing results are output through two channels: one is fed back to the VR system in real time, triggering automatic scene adjustments (such as switching to a soothing scene when anxiety is detected); the other is uploaded to the doctor's end (e.g., a doctor's monitoring platform) via Wi-Fi / 5G, supporting remote intervention.
[0052] 2. Algorithm Layer (1) Algorithm module composition The data processing module includes an emotion recognition module, a pupil movement recognition module, and a dual-modal data fusion module. By monitoring the patient's facial expressions and pupil changes in real time, the virtual scene is dynamically adjusted to achieve the purpose of emotion regulation and psychological counseling.
[0053] The algorithm modules are shown in Table 2.
[0054] Table 2 Algorithm Module Composition
[0055] (2) Algorithm processing flow 1) Emotion Recognition Module The YOLOv8n architecture is optimized in depth, and is trained on a 20,000 finely annotated expression dataset covering different races, lighting, and angles to ensure generalization. The model is compressed by 60% through channel pruning + INT8 quantization, and achieves 18fps real-time inference on RK3568 NPU using the TensorRT engine. The FPN (Feature Pyramid Network) is integrated to maintain an accuracy of greater than or equal to 90% within a range of 0.3 meters to 1.2 meters.
[0056] That is, the emotion recognition module is used to recognize the patient's emotion according to each frame of the real-time captured facial image, using a YOLOv8n lightweight neural network model, outputting the patient's emotion label and confidence; the YOLOv8n lightweight neural network model adopts the YOLOv8n architecture, and after the YOLOv8n architecture is compressed through channel pruning and INT8 quantization, it is obtained through deep learning training.
[0057] The emotion recognition module uses the YOLOv8n lightweight neural network model, and the development environment is a conda environment, with Python 3.10, CUDA 12.1, and PyTorch 2.3.1 framework. The model training uses a dataset containing 20,000 expression images, covering eight core human expressions: anger, disgust, fear, happiness, neutral, sadness, surprise, and contempt. Through deep learning training, the model can capture the changes in user facial expressions in real time and output recognition results with high confidence, with a confidence of more than 80%, meeting the recognition accuracy. During training, the dataset is carefully cleaned and labeled to ensure the model's ability to recognize different emotional states. After training, the model is integrated into a Python-based application, which detects facial expressions for each frame of image by calling the real-time camera image, realizes dynamic detection and feedback of emotion categories, and displays the recognition results and confidence in the form of text or icons on the UI interface. A simple and intuitive UI interface is designed, including a live view window, emotion category indicators, and a confidence bar chart, making it easy for users or doctors to observe emotional trends. The emotion recognition UI interface is shown in Figure 9 .
[0058] 2) Pupil movement recognition module (also known as pupil tracking module) The pupil movement recognition module is used to track the changes in pupil position according to each frame of the real-time captured pupil image, using the Mediapipe pupil tracking algorithm, to obtain the pupil movement change data for each frame; the pupil movement change data includes 3D pupil coordinates, pupil diameter change, and gaze direction.
[0059] Based on Google Mediapipe customized development, real-time extraction of 478 eye key points is realized by using lightweight CNN (Convolutional Neural Network), and adaptive ROI (Region of Interest) is used to cope with head movement. The calculation of 3D pupil coordinates (±2 pixel error), diameter change (ellipse fitting, ±5% error) and gaze direction (corneal reflection, less than 3 ° error), to ensure high-precision physiological parameter monitoring.
[0060] The detailed implementation process of the Mediapipe pupil tracking algorithm in the pupil tracking module is as follows: ① First, image preprocessing: convert the image captured by the camera to RGB format and perform normalization. ② Face detection and key point positioning: use the FaceMesh model of Mediapipe to detect the face. ③ Pupil region positioning: construct the eye ROI by the eye key points, and further accurately position the pupil center. ④ Key point extraction and tracking: use the face_landmarks output of Mediapipe to extract the pupil-related key points.
[0061] The implementation process of adaptive ROI to cope with head movement is as follows: ① Initial ROI setting: based on the face key points, the initial eye ROI is fixed. ② Dynamic ROI update: use the pupil position of the previous frame to predict the ROI of the current frame. ③ Adjust the ROI position in combination with head pose estimation (such as using the head pose module of Mediapipe); if the pupil is out of the ROI, expand the ROI range or reinitialize the tracking. ④ Stability guarantee: smooth the ROI movement by Kalman filtering or sliding window averaging method to avoid jitter.
[0062] The calculation process of 3D pupil coordinates is as follows: ① Camera calibration: obtain the camera intrinsic parameters (focal length, principal point) and distortion coefficients. ② 3D eye key points: use the 3D face key points output by Mediapipe (normalized to [0, 1]) to project to 3D space combined with camera parameters. ③ Pupil center 3D coordinates: take the 3D coordinates of the point representing the pupil center among the eye key points.
[0063] The calculation process of diameter change is as follows: the pupil diameter is estimated by the distance between the key points: ① Select key points: select the key points on the upper and lower boundaries of the pupil. ② Calculate the Euclidean distance: the horizontal diameter is equal to the distance between the left and right boundary key points, and the vertical diameter is equal to the distance between the upper and lower boundary key points. ③ Dynamic calibration: scale the diameter according to the distance of the head from the camera (Z value): actual diameter = pixel diameter x scale factor. ④ Output change trend: the system records the change of the diameter over time, which is used for emotional state analysis (such as pupil dilation indicating interest or tension).
[0064] The gaze direction determination process is: ① eye center estimation: a spherical model is fitted using eye contour key points to estimate the 3D coordinates of the eye center. The line-of-sight vector Calculation: =Ppupil-Cey, where Ppupil is the center of the pupil and Ceye is the center of the eye. ② Coordinate system conversion: convert the line-of-sight vector from the eye coordinate system to the head coordinate system, and then combine the head pose (rotation matrix) to get the gaze direction in the world coordinate system. ③ Output direction category: such as "straight view", "left view", "right view", "up view", "down view", etc., for judging the user's attention state.
[0065] Therefore, it can be concluded that, according to the real-time captured each frame of pupil image, the Mediapipe pupil tracking algorithm is used to track the change of pupil position, and the pupil motion change data of each frame is obtained. The pupil motion recognition module includes: converting the pupil image into RGB format and performing normalization processing to obtain the preprocessed pupil image; according to the preprocessed pupil image, using the Face Mesh model of Mediapipe to detect eye key points, and constructing an eye region of interest according to the eye key points; using the face_landmarks of Mediapipe to extract the pupil key points from the eye region of interest, and constructing the pupil region according to the pupil key points; positioning the pupil center from the pupil region to obtain the 3D pupil coordinates; selecting the key points of the upper, lower, left and right boundaries of the pupil, and calculating the horizontal diameter and vertical diameter of the pupil; using the formula , the horizontal diameter and the vertical diameter are scaled to obtain the scaled horizontal diameter or the scaled vertical diameter; in the formula, is the scaled horizontal diameter or the scaled vertical diameter, is the horizontal diameter or the vertical diameter, is the scale factor; the change amount of the pupil diameter is determined according to the scaled horizontal diameter or the scaled vertical diameter of the adjacent frame; the eye contour key points are detected in the preprocessed pupil image; a spherical model is fitted using the eye contour key points, and the center coordinates of the spherical model are determined as the 3D coordinates of the eye center; the line-of-sight vector is determined according to the 3D coordinates of the eye center and the 3D pupil coordinates; the line-of-sight vector is converted from the eye coordinate system to the world coordinate system to obtain the gaze direction.
[0066] The pupil movement recognition module is based on the Mediapipe framework open sourced by Google Research. The Mediapipe framework provides a wealth of functions, including face detection, pose estimation, hand tracking, and is suitable for a variety of application scenarios such as AR, game interaction, and health monitoring. The Face Mesh module in the Mediapipe framework is used to accurately track the position and movement of the pupils. This module can capture the pupil state of the user in real time (such as direct view, left view, right view, etc.), providing multi-modal data support for emotion recognition. The cross-platform nature of the Face Mesh module allows it to efficiently run on embedded devices or VR headsets, working in conjunction with the emotion recognition module to further enhance the overall performance of the system. By combining pupil change data, the system can more comprehensively analyze the user's unconscious reactions in a virtual environment, such as pupil dilation, which may indicate cognitive interest or emotional fluctuations, providing more accurate basis for scene adjustment.
[0067] The emotion recognition module and the pupil movement recognition module are integrated through a Python application, forming a complete software system. The system has been tested and optimized multiple times in a laboratory environment to ensure real-time performance and accuracy. During the training process, loss function trend graphs (such as Figure 7 ) were recorded, verifying the convergence and stability of the model. Figure 7 Part (a) of the figure shows the precision trend, Figure 7 Part (b) of the figure shows the recall trend, Figure 7 Part (c) of the figure shows the 50% precision mean trend, Figure 7 Part (d) of the figure shows the 50%-90% precision mean trend. The confusion matrix of the test set (such as Figure 8 ) further demonstrates the recognition performance of the model on various expressions.
[0068] 3) Bimodal data fusion module The bimodal data fusion module is used to synchronize the emotion labels and pupil movement change data through timestamps, and uses a weighted decision mechanism to fuse the emotion labels and pupil movement change data synchronized by timestamps to obtain the fusion data of each frame. According to the fusion data of consecutive multiple frames, the state machine is used to analyze the emotional trend of the patient, and the comprehensive emotional index of the patient is output.
[0069] Specifically, the dual-modal data fusion adopts a weighted decision mechanism. First, the expression and pupil data are synchronized by timestamps, and then the fusion is performed with a weight ratio of 70% expression features and 30% pupil features, and the emotion trend is analyzed and judged by combining the state machine of the last 5 frames. The final output includes real-time emotion labels (8 categories) and confidence, pupil dynamic parameters (position / diameter / gaze direction), comprehensive emotion index (0-100 scale), and abnormal state warning. The entire algorithm is implemented on an embedded development board with an end-to-end delay of less than 80 ms, meeting the strict requirements of clinical real-time interaction.
[0070] The algorithm processing flow of the emotion recognition module, the pupil movement recognition module, and the dual-modal data fusion module is shown in Figure 5 .
[0071] 3. Interaction layer (1) The composition of the interaction module is shown in Table 3.
[0072] Table 3 Composition of interaction module
[0073] (2) Interaction process For example, the data processing module is also used for abnormal state warning when the emotion labels of the patient for consecutive multiple frames are all negative emotions, and the pupil is abnormal, triggering the handle vibration and voice guidance; the pupil abnormality includes that the pupil diameter change is greater than the scaling threshold, or the gaze direction of consecutive multiple frames is different.
[0074] In terms of dynamically adjusting the virtual interaction scene parameters or switching the virtual interaction scene according to the comprehensive emotion index of the patient, the VR glasses include: if the comprehensive emotion index of the patient is greater than the emotion index threshold, the virtual interaction scene is switched to a positive scene; if the comprehensive emotion index of the patient is less than or equal to the emotion index threshold, the saturation of the virtual interaction scene is reduced and white noise is increased.
[0075] The interaction process realizes closed-loop regulation and control through the intelligent linkage of multi-modal data and virtual environment: the embedded development board transmits the processed emotion labels, pupil coordinates, and comprehensive emotion index to the Unity engine through USB 3.0, dynamically adjusts the scene parameters (such as reducing the saturation and increasing the white noise when anxious); when a continuous depression state is detected, the ROS 2 system automatically switches to a safe scene and triggers tactile feedback, while the doctor can view the data in real time and remotely intervene through the Web monitoring interface, and the patient receives voice guidance and vibration prompts, forming an optimized closed loop of "data acquisition-analysis-regulation-feedback". The specific process is shown in Figure 6 . In the closed-loop interaction flowchart of Figure 6 , when the comprehensive emotion index (referred to as emotion index) is greater than 70, the positive scene is activated, and when the comprehensive emotion index is less than or equal to 70, the relaxation mode is started.
[0076] The determination of depressive state includes: emotional level and physiological level. Emotional level: YOLOv8n model continuously outputs high confidence results such as "neutral" or "sad". This is interpreted in the system as a manifestation of low mood and anhedonia (lack of interest in things). Physiological level: Mediapipe pupil tracking module detects significant pupil diameter narrowing (compared to individual baseline), which may be accompanied by slow eye movement, scattered gaze (inability to concentrate).
[0077] "Sustained" is a judgment based on time and stability, to avoid false triggering of the system due to transient emotional fluctuations. If the patient is just "sad" for a moment, the system will not overreact. But if the system finds that the patient's emotional and physiological indicators are stably in a state of depression and lack of response for a long enough period of time, it will be determined as "sustained depressive state".
[0078] Changes in pupil size are closely related to the subconscious reaction of the brain to certain stimuli. When a depressed patient sees pictures related to their past or hears familiar voices, they may subconsciously trigger changes in pupil size, such as pupil dilation or narrowing, and through the detection of pupil changes, the patient's unconscious reaction to specific stimuli can be understood. For example, when the system detects a sudden pupil dilation, it may mean that the content of the stimulus has triggered a cognitive response in the patient, indicating that the brain has a certain degree of familiarity with this information. Then relevant stimuli can be further increased, such as playing a video clip of the patient's past warm life, to help heal the patient's mind.
[0079] The present application is based on the combination of facial expression recognition and real-time pupil tracking with VR technology to conduct psychological counseling for patients with depression and adjust the patient's mood. Facial emotion recognition and pupil movement recognition technology are implanted into VR glasses through a network module, and three small cameras are installed to capture changes in the patient's facial expressions. Patients with depression wear VR glasses and play virtual scenes that may make them feel comfortable and warm. The system monitors the patient's facial expressions and pupils in real time and changes the scene or deepens the scene stimulation according to the patient's feedback, so as to achieve the purpose of relaxing the patient's body and mind and healing the mind.
[0080] The present application provides a unique treatment plan for patients with anhedonia by designing virtual reality technology. The use of head-mounted displays makes users feel immersed in a virtual, computer-generated world, creating a realistic perception illusion. Interactivity in the virtual environment enhances the immersive experience of positive experiences that may not be accessible in real life. Repeated exposure to these positive virtual experiences can help alleviate the patient's anhedonia. This exposure can increase the patient's interest and enjoyment in daily life and greatly enjoy this feeling.
[0081] The present application is a potentially effective adjunctive treatment for patients with depression by monitoring facial expressions and pupil changes. The present application can help identify the emotional response of patients under specific virtual environment stimuli and use relevant biofeedback to support the recovery of patients with depression.
[0082] The technical features of the above embodiments can be combined in any manner. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as there is no contradiction, it should be considered within the scope of the present application.
[0083] The principles and implementation modes of the present application are described by applying specific examples herein, and the above embodiment descriptions are only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In summary, the content of the present application should not be understood as a limitation.
Claims
1. A virtual reality-based auxiliary treatment system for depression, characterized in that, The virtual reality-based auxiliary treatment system for depression includes: VR glasses, a camera module, and a data processing module; The data processing module is integrated into the VR glasses; After the patient wears VR glasses, the VR glasses are used to play virtual interactive scenes; the camera module is used to capture facial and pupil images of the patient in real time while watching the virtual interactive scenes; the data processing module is used to determine the patient's emotion in each frame based on the real-time captured facial images using an emotion recognition model, and simultaneously to track pupil position changes in each frame based on the real-time captured pupil images using a pupil tracking algorithm to obtain pupil movement change data for each frame, and then fuse and analyze the emotion and pupil movement change data of multiple consecutive frames to output the patient's comprehensive emotion index; the VR glasses are also used to dynamically adjust the parameters of the virtual interactive scene or switch the virtual interactive scene based on the patient's comprehensive emotion index to alleviate the patient's psychological stress and emotions.
2. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, The camera module includes three cameras; A camera is used to capture images of the patient's face; Two additional cameras are embedded in the VR glasses to capture images of the patient's pupils.
3. The virtual reality-based auxiliary treatment system for depression according to claim 2, characterized in that, The virtual reality-based auxiliary treatment system for depression also includes: a head-mounted lens support; A camera used to capture images of the patient's face is mounted on a head-mounted lens holder; the head-mounted lens holder is worn on the patient's head.
4. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, The virtual reality-based depression auxiliary treatment system also includes: a 3D-printed shell; 3D-printed shells are used to encapsulate VR glasses and camera modules.
5. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, The virtual reality-based auxiliary treatment system for depression also includes: a UI interface; The UI is used to display the patient's mood and pupil movement changes in each frame in the form of text or icons.
6. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, The data processing module is the Wildfire Luban Cat 2-V3 development board, which embeds an emotion recognition model and a pupil tracking algorithm.
7. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, The data processing module includes: an emotion recognition module, a pupil movement recognition module, and a dual-modal data fusion module; The emotion recognition module is used to identify the patient's emotions based on each frame of facial image captured in real time, using a lightweight YOLOv8n neural network model, and outputs the patient's emotion label and confidence level; the lightweight YOLOv8n neural network model adopts the YOLOv8n architecture, and is obtained by deep learning training after channel pruning and INT8 quantization compression of the YOLOv8n architecture. The pupil motion recognition module is used to track pupil position changes based on each frame of pupil image captured in real time, using the Mediapipe pupil tracking algorithm to obtain pupil motion change data for each frame; the pupil motion change data includes 3D pupil coordinates, pupil diameter change, and gaze direction; The dual-modal data fusion module is used to synchronize emotion labels and pupil movement change data through timestamps, and adopts a weighted decision mechanism to fuse the emotion labels and pupil movement change data after timestamp synchronization to obtain fused data for each frame; based on the fused data of multiple consecutive frames, the state machine is used to analyze the patient's emotional trend and output the patient's comprehensive emotion index.
8. The virtual reality-based auxiliary treatment system for depression according to claim 7, characterized in that, In terms of tracking pupil position changes using the Mediapipe pupil tracking algorithm based on each frame of pupil images captured in real time to obtain pupil motion change data for each frame, the pupil motion recognition module includes: The pupil image is converted to RGB format and normalized to obtain the preprocessed pupil image; Based on the preprocessed pupil image, Mediapipe's Face Mesh model is used to detect key points in the eye, and the region of interest in the eye is constructed based on these key points. The pupil key points are extracted from the region of interest of the eye using Mediapipe's face_landmarks, and the pupil region is constructed based on the pupil key points; Locate the pupil center from the pupil region to obtain 3D pupil coordinates; Select key points on the upper, lower, left, and right boundaries of the pupil, and calculate the horizontal and vertical diameters of the pupil; Using formula The horizontal and vertical diameters are scaled to obtain the scaled horizontal or vertical diameter; where, The horizontal or vertical diameter after scale correction. It can be either the horizontal diameter or the vertical diameter. Scale factor; The change in pupil diameter is determined based on the horizontal or vertical diameter after scale correction of adjacent frames; Detect key points of the eye contour in the preprocessed pupil image; A spherical model is fitted using key points of the eye contour, and the center coordinates of the spherical model are determined as the 3D coordinates of the eyeball center. Determine the gaze vector based on the 3D coordinates of the eyeball center and the 3D pupil coordinates; Transform the gaze vector from the eye coordinate system to the world coordinate system to obtain the gaze direction.
9. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, The data processing module is also used to issue an abnormal state warning and trigger handle vibration and voice guidance when the patient's emotion label is negative for multiple consecutive frames and the pupils are abnormal. The pupil abnormalities include pupil diameter changes greater than the scaling threshold, or different gaze directions in multiple consecutive frames.
10. The virtual reality-based auxiliary treatment system for depression according to claim 1, characterized in that, VR glasses include those for dynamically adjusting virtual interaction scene parameters or switching virtual interaction scenes based on the patient's overall emotional index: If the patient's overall emotional index is greater than the emotional index threshold, the virtual interaction scenario will be switched to a positive scenario. If the patient's overall emotional index is less than or equal to the emotional index threshold, the saturation of the virtual interactive scene is reduced and white noise is increased.
Citation Information
Patent Citations
Method for detecting depression based on video analysis
CN114209322A
Multi-modal data fused virtual reality assisted diagnosis and treatment method for depression
CN120131018A
Method for multi-modal and multi-dimensional analysis and early warning of psychological health of students by using AI (Artificial Intelligence)
CN120408371A
Emotion regulation and control system based on adaptive virtual reality scene
CN120412919A