Method and device for mental health monitoring and assisted diagnosis based on multi-modal recognition

By constructing a multimodal time window alignment and quality gating mechanism and a multi-agent evidence review system, the problems of information loss and security risks in multimodal data analysis in existing technologies are solved, enabling traceable collaborative analysis and high-risk identification in mental health assessment, and providing interpretable risk reports.

CN122369947APending Publication Date: 2026-07-10CHANGZHOU UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGZHOU UNIV
Filing Date
2026-05-19
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies in mental health assessment suffer from limitations such as single-modal data analysis leading to missing information or biased judgments, lack of deep integration and collaborative analysis of multimodal data, inability to accurately identify high-risk individuals and implement safe interventions, and lack of traceable decision-making links.

Method used

We construct a complete technical chain of data collection, processing, reasoning, review, and record keeping, implement multimodal time window alignment and quality gating mechanisms, introduce security calibration rules and multi-agent evidence review system, and generate interpretable risk reports through synchronous collection and fusion analysis of multimodal data.

Benefits of technology

It enables traceable and collaborative analysis of the mental health assessment process, effective fusion of multi-source data and bias correction, accurate identification of high-risk factors, and provides transparent and interpretable decision support, ensuring the robustness and safety of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369947A_ABST
    Figure CN122369947A_ABST
Patent Text Reader

Abstract

This invention relates to the field of mental health diagnosis and treatment technology. The invention provides a method and device for mental health monitoring and auxiliary diagnosis based on multimodal recognition. The method includes: a data processing unit receiving and preprocessing multimodal data from eye movement, electroencephalography (EEG), facial recognition, speech, scales, and text; constructing a comprehensive feature set reflecting attentional state, physiological arousal, emotional expression, and subjective risk by extracting quantifiable features from the multimodal data; concatenating the comprehensive feature set into an input vector, inputting it into a local multi-disease MLP, outputting five risk probabilities, and performing structured scale security calibration; calculating independent scores for each channel based on the multimodal features, and performing weighted fusion with the MLP output, combining risk thresholds and self-harm intention escalation rules to generate a risk level; and generating an auxiliary assessment report through a multi-agent collaborative review mechanism. This approach improves the accuracy, interpretability, and traceability of mental health auxiliary assessment results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of psychological diagnosis and treatment technology, and in particular to methods and equipment for psychological health monitoring and auxiliary diagnosis and treatment based on multimodal recognition. Background Technology

[0002] As mental health issues receive increasing social attention, traditional assessment models relying on self-report scales and clinical interviews are no longer sufficient to meet the practical needs of large-scale, routine monitoring. Scale assessments are susceptible to subjective interference, and their reliability and validity are often reduced due to cognitive biases or immediate emotional fluctuations. While clinical interviews are professional, they are limited by human resources and time constraints, making it impossible to achieve broad coverage and dynamic tracking. To overcome these limitations, automated assessment technologies based on artificial intelligence have emerged in recent years, attempting to construct a more objective and continuous monitoring system by collecting multi-dimensional data such as facial, voice, eye movements, and physiological data. However, existing technological approaches still face several key bottlenecks in the entire chain from data collection to the generation of reliable conclusions.

[0003] Most current methods focus only on single-modal data analysis, making it difficult to comprehensively capture the complex manifestations of psychological states across multiple channels. Psychological activity is a comprehensive reflection of behavior, physiology, and cognitive emotions, and its abnormal signals often exhibit asynchronous or complementary characteristics across different modalities. Single-dimensional analysis is prone to information gaps or biased judgments. More importantly, psychological states exhibit significant dynamics and temporal evolution during the assessment process. If precise time window alignment and quality filtering are not performed on data from different sources and with different sampling rates, cross-modal correlation analysis will lose its temporal consistency foundation, leading to a decrease in the reliability of feature extraction and state judgment.

[0004] Furthermore, existing methods often simplify multimodal data processing to independent feature extraction and simple concatenation, lacking mechanisms for deep fusion and collaborative analysis of heterogeneous data. Different modalities differ in quality, feature scale, and semantic level, requiring normalization and quality gating by the device to construct a unified representation reflecting the dynamic evolution of psychological states. In addition, most technical solutions treat the analytical model as a black box, directly outputting classification results without preserving a complete traceable link from raw data to the final decision. This lack of transparency makes it impossible to verify key information such as intermediate features, inference probabilities, and fusion confidence levels output by the model, hindering uncertainty quantification and management when data quality is poor or evidence is contradictory.

[0005] Crucially, existing methods generally lack tiered early warning and safety intervention mechanisms when identifying high-risk situations such as self-harm intentions. Devices fail to dynamically correlate the consistency of multimodal evidence, the confidence level of model outputs, and preset clinical safety thresholds, and cannot automatically trigger manual review processes when potential crises are identified, posing safety risks due to misjudgments or omissions. Therefore, there is an urgent need for an intelligent data processing method that is deeply integrated with physical, multi-sensor monitoring devices. Summary of the Invention

[0006] To address the above issues, this invention constructs a complete technical chain of data acquisition, processing, reasoning, review, and record keeping, implements a multimodal time window alignment and quality gating mechanism, and introduces security calibration rules and a multi-agent evidence review system. This enables traceable collaborative analysis of the mental health assessment process, effective fusion and bias correction of multi-source data, accurate identification of high-risk factors, and generation of interpretable reports, providing reliable technical support for professionals' clinical decision-making.

[0007] According to embodiments of the present invention, a method and device for mental health monitoring and assisted diagnosis based on multimodal recognition are provided.

[0008] In a first aspect of the invention, a method for mental health monitoring and assisted diagnosis based on multimodal recognition is provided. The method includes: S01: The data processing unit receives and preprocesses multimodal data from eye tracking, electroencephalography, facial recognition, speech, scales, and text. S02: Construct a comprehensive feature set reflecting attentional state, physiological arousal, emotional expression, and subjective risk by extracting quantifiable features from multimodal data; S03: Concatenate the comprehensive feature set into an input vector, input it into the local multi-disease MLP, output five risk probabilities, and perform structured scale security calibration; S04: Calculate independent scores for each channel based on multimodal features, and perform weighted fusion with MLP output. Combine risk thresholds and self-harm intention escalation rules to generate risk levels. Step S05: Generate an auxiliary evaluation report through a multi-agent collaborative review mechanism.

[0009] Furthermore, the preprocessing described in step S01 includes: adding a quality field to each modal data, removing data segments that do not meet quality standards; and mapping data with different sampling frequencies to a unified time reference using a preset time window as the unit.

[0010] Furthermore, the quantifiable features mentioned in step S02 include: Eye movement variance entropy, fixation stability, and saccade intensity extracted from eye movement data; The theta / alpha / beta power ratio, relative beta power, relative gamma power, and attention-relaxation difference were extracted from EEG data. The comprehensive score, classification results, facial feature vector, and facial image quality score extracted from facial images; Speech rate, pause ratio, average pitch, average energy, and spectral features extracted from speech signals; The raw scores of the depression scale, anxiety scale, stress scale, sleep quality, and number of recent stressful events were extracted from the scales and text.

[0011] Furthermore, the structured scale safety calibration mentioned in step S03 specifically involves taking the maximum value of the risk probability output by the MLP and the safety lower limit probability output based on the scale rules, to obtain the calibrated risk probability.

[0012] Furthermore, step S04 specifically includes: Calculate eye movement score, electroencephalogram score, model risk score, and structured scale score respectively; The four scores are weighted and fused according to preset weights to obtain a total fused score; The total score is compared with a preset risk threshold to generate low, medium, and high risk levels; When a self-harm intention is detected, a crisis-level warning is output directly, and the total score is raised to the preset crisis threshold.

[0013] Furthermore, the multi-agent collaborative review mechanism described in step S05 includes a data parsing agent, an eye-tracking analysis agent, an EEG analysis agent, a DNN interpretation agent, a fusion decision-making agent, a report generation agent, and a self-review agent that run sequentially; each agent node records inputs, tools, summaries, evidence, confidence levels, constraints, and outputs, forming a traceable decision-making chain.

[0014] In a second aspect of the invention, a device for mental health monitoring and assisted diagnosis based on multimodal recognition is provided. The device includes: A host base and a host chassis, wherein the host chassis is mounted on the host base and a data processing unit is provided inside the host chassis, the data processing unit being used to implement the method as described in the first aspect; The host base is equipped with an interactive component, which is used to interact with the patient to exchange information during the diagnosis and treatment process. The main unit is equipped with a data acquisition module, which includes a PTZ camera, a voice acquisition unit, and a physiological data acquisition unit. The PTZ camera is used to track the user's face and acquire facial expression images and eye movement image data. The voice acquisition unit is used to acquire the patient's voice signal. The physiological data acquisition unit is used to acquire the patient's physiological signs data.

[0015] Furthermore, the PTZ camera includes a rotating part, a tilting part, and a high-definition camera. The rotating part is mounted on the main unit, the tilting part is mounted on the rotating end of the rotating part, and the high-definition camera is mounted on the tilting end of the tilting part. The voice acquisition unit includes a microphone, which is mounted on the frame of the high-definition camera. The physiological data acquisition unit includes a finger-clip pulse oximeter and / or a smart bracelet, used to monitor the user's blood oxygen, heart rate, and / or pulse data in real time, and transmit the data to the data processing unit inside the main unit in real time.

[0016] Furthermore, the rotating part includes a fixed cylinder and a rotating cylinder. The rotating cylinder is rotatably installed inside the fixed cylinder. A steering motor is fixedly installed inside the rotating cylinder. A drive gear is fixedly installed at the drive end of the steering motor. A driven gear ring is fixedly installed inside the fixed cylinder. The drive gear and the driven gear ring are meshed and connected. When the drive end of the steering motor rotates, the rotating cylinder is driven to rotate relative to the fixed cylinder through the meshing motion of the drive gear and the driven gear ring. The pitch unit includes a pitch motor and a gear reducer. The pitch motor is fixedly installed inside the rotating cylinder, and the gear reducer is fixedly installed on the rotating cylinder. The high-definition camera is fixedly installed at the output end of the gear reducer, and the input end of the gear reducer is fixedly connected to the drive end of the pitch motor.

[0017] Furthermore, the main unit is equipped with a memory that stores a mental health assessment model and auxiliary treatment plan. The data processing unit is electrically connected to the data acquisition module, the memory, and the interactive component. The interactive component includes a keyboard and a display screen mounted on the main unit.

[0018] This invention constructs a complete technical chain of data collection, processing, reasoning, review, and record keeping, implements a multimodal time window alignment and quality gating mechanism, and introduces security calibration rules and a multi-agent evidence review system. This enables traceable collaborative analysis of the mental health assessment process, effective fusion and bias correction of multi-source data, accurate identification of high-risk factors, and generation of interpretable reports, providing reliable technical support for professionals' clinical decision-making.

[0019] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description.

[0020] The beneficial effects of this invention are: 1. It has achieved a complete closed loop from multimodal data acquisition, preprocessing, feature extraction, model inference to evidence review and security decision-making. Through specific steps such as clear time window alignment, quality gating, and feature normalization, it forms a reproducible and traceable data processing flow. 2. Through simultaneous acquisition and fusion analysis of multimodal data from eye movement, electroencephalography (EEG), facial expression, speech, scales, and local multi-disease MLP, the evidence from each modality corroborates each other: eye movement reflects attention allocation and gaze shift, EEG reveals physiological arousal and cognitive load, facial and speech capture overt emotions, scales and text provide subjective risk cues, and MLP outputs nonlinear risk estimates, effectively overcoming the limitations of a single modality and thus improving the robustness and accuracy of the overall assessment. 3. Introduce multi-level quality gating and safety calibration rules. By monitoring the data quality of each modality in real time, the evidence weight is dynamically adjusted. When crisis signals such as self-harm intentions are detected, the escalation process is automatically triggered to ensure that high-risk scenarios are not missed. At the same time, the complete processing link is recorded through structured logs to support post-event review and model iteration. 4. The equipment output not only includes the risk level, but also provides scores for each modality, key evidence, constraints, confidence level descriptions, and suggestions for manual review. Through a multi-agent evidence review mechanism, a transparent and traceable decision-making basis is formed, which makes it easy for professionals to quickly understand the equipment's judgment logic and intervene when necessary, so as to achieve refined risk management through human-machine collaboration. Attached Figure Description

[0021] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. Wherein: Figure 1 A flowchart illustrating a method for mental health monitoring and assisted diagnosis based on multimodal recognition according to an embodiment of the present invention is shown. Figure 2 A three-dimensional structural schematic diagram of a diagnostic and treatment device according to an embodiment of the present invention is shown; Figure 3 A side view of the diagnostic and treatment device according to an embodiment of the present invention is shown; Figure 4 A side view cross-sectional structural schematic diagram of a diagnostic and treatment device according to an embodiment of the present invention is shown; Figure 5A schematic front cross-sectional view of a diagnostic and treatment device according to an embodiment of the present invention is shown. In the figure: 1. Main unit base; 2. Main unit housing; 3. High-definition camera; 4. Microphone; 5. Fixing cylinder; 6. Rotating cylinder; 7. Steering motor; 8. Driving gear; 9. Driven gear ring; 10. Pitch motor; 11. Gear reducer; 12. Keyboard; 13. Display screen. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] According to embodiments of the present invention, a method and device for mental health monitoring and assisted diagnosis based on multimodal recognition are proposed. By constructing a complete technical chain of data collection, processing, reasoning, review, and record keeping, implementing multimodal time window alignment and quality gating mechanisms, and introducing security calibration rules and a multi-agent evidence review system, the method achieves traceable collaborative analysis of the mental health assessment process, effective fusion and bias correction of multi-source data, accurate identification of high-risk factors, and generation of interpretable reports, providing reliable technical support for professionals' clinical decision-making.

[0024] The principles and spirit of the present invention will be explained in detail below with reference to several representative embodiments.

[0025] Figure 1 This is a schematic flowchart of a method for mental health monitoring and assisted diagnosis based on multimodal recognition according to an embodiment of the present invention. The method includes: S01: The data processing unit receives and preprocesses multimodal data from eye tracking, electroencephalography, facial recognition, speech, scales, and text. S02: Construct a comprehensive feature set reflecting attentional state, physiological arousal, emotional expression, and subjective risk by extracting quantifiable features from multimodal data; S03: Concatenate the comprehensive feature set into an input vector, input it into the local multi-disease MLP, output five risk probabilities, and perform structured scale security calibration; S04: Calculate independent scores for each channel based on multimodal features, and perform weighted fusion with MLP output. Combine risk thresholds and self-harm intention escalation rules to generate risk levels. Step S05: Generate an auxiliary evaluation report through a multi-agent collaborative review mechanism.

[0026] It should be noted that although the operation of the method of the present invention has been described in a specific order in the above embodiments and figures, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0027] To provide a clearer explanation of the above-mentioned method for mental health monitoring and assisted diagnosis based on multimodal recognition, two specific embodiments are described below. However, it is worth noting that these embodiments are only for better illustrating the present invention and do not constitute an improper limitation of the present invention.

[0028] The following two specific examples will further illustrate the methods for mental health monitoring and assisted diagnosis based on multimodal recognition.

[0029] Example 1

[0030] This embodiment discloses a method for mental health monitoring and assisted diagnosis based on multimodal recognition, and the specific steps are as follows.

[0031] Step S01: The data processing unit receives and preprocesses multimodal data including eye movement, electroencephalogram (EEG), facial, speech, scale, and text.

[0032] Specifically, the data processing unit receives facial video frames and eye-tracking coordinate data output from the PTZ camera, receives voice signals output from the microphone, receives heart rate, blood oxygen, pulse, and EEG field data output from physiological data acquisition units such as fingertip pulse oximeters or smart bracelets, and receives scale scores and text data generated through keyboard or display screen interaction.

[0033] To ensure data traceability, the data processing unit uses a unified data recording format to save intermediate results. For example, eye-tracking data records include timestamps, lateral fixation coordinates, longitudinal fixation coordinates, pupil size, and validity indicators; EEG or physiological data records include timestamps, delta wave power, theta wave power, alpha wave power, beta wave power, gamma wave power, focus level, relaxation level, and quality indicators; scale and text data records include depression scale scores, anxiety scale scores, stress scale scores, sleep quality scores, number of recent stressful events, self-harm ideation indicators, and text summaries.

[0034] Quality fields such as quality, valid, and status are added to each modality of data, and data segments with missing faces, invalid eye movement points, poor EEG contact, low voice energy, or missing items on the scale are removed.

[0035] Using 2-second or 5-second analysis windows, data from different sampling frequencies are mapped to a unified time base, and the corresponding eye-tracking, EEG, facial, speech, scale, and text features are saved according to the window number to obtain window-level multimodal features.

[0036] Step S02: Construct a comprehensive feature set reflecting attentional state, physiological arousal, emotional expression, and subjective risk by extracting quantifiable features from multimodal data.

[0037] For each time window Quantifiable feature indicators were extracted from eye-tracking, electroencephalography, facial, speech, scale, and text data, respectively.

[0038] In this embodiment, eye movement features are calculated according to the following steps to extract three core, interpretable eye movement features: eye movement variance entropy, fixation stability, and saccade intensity. First, within a time window... Obtain the effective set of eye-tracking points internally. ,in Indicates the lateral gaze coordinates. Indicates the longitudinal gaze coordinate. This indicates the number of effective eye-tracking points.

[0039] Calculate the variance of eye movement coordinates: , .

[0040] Calculate the distance between adjacent eye-tracking points, and calculate the trajectory entropy based on the distance distribution: , , , In the formula, Indicates falling into the first The proportion of adjacent eye movement point distance samples in a given eye movement jump distance interval to the total number of adjacent eye movement point distance samples. The number of bins representing the distance of eye movement jumps. Represents the unnormalized trajectory entropy. This represents the normalized trajectory entropy.

[0041] Calculate the normalized variance and eye-tracking variance entropy: , , , , In the formula, The normalized scale representing the coverage area of ​​eye movement coordinates within the time window. This represents a very small constant that prevents division by zero. This represents the dispersion of the normalized coordinates. Represents the normalized variance. This represents the eye-tracking variance entropy.

[0042] Calculate fixation stability and saccade intensity: , , In the formula, Indicates gaze stability, This indicates the intensity of saccades. The above indicators are used to determine the dispersion, stability, and intensity of eye movement in a user's gaze trajectory.

[0043] In this embodiment, EEG features are calculated according to the following steps to obtain the theta / alpha / beta power ratio, beta relative power, gamma relative power, and attention-relaxation difference. First, the effective EEG samples within the time window are averaged to obtain delta wave power, theta wave power, alpha wave power, beta wave power, and gamma wave power. When the proportion of effective EEG samples is less than 80%, the EEG results of this time window are only recorded as data quality constraints.

[0044] Calculate the power of the frequency band: , , , , , In the formula, , , , and These represent the average power of the delta, theta, alpha, beta, and gamma frequency bands within the time window, respectively.

[0045] Calculate total power, power ratio, and relative power: , , , In the formula, express Overall power ratio express Relative power, express Relative power.

[0046] Calculate the difference between focus and relaxation: , In the formula, This represents the combined power ratio of theta / alpha / beta. Represents the relative power of beta. Indicates the relative power of gamma. It represents the difference between focus and relaxation levels.

[0047] In this embodiment, the facial image is first used to locate the face region and local regions such as the eyes, eyebrows, nose, and corners of the mouth through a face detection and alignment algorithm. Then, the full-face image and local images are uniformly scaled to 256×256 pixels and input into a DNet-like visual network. This network includes a feature extraction module (FEM), a Coordinate Attention mechanism, and ViTBlock, used to simultaneously learn the full-face pose and subtle local expressions. The output facial features include: a comprehensive score, possible classification results, facial feature vectors, and a facial image quality score.

[0048] After framing, denoising, endpoint detection, and energy normalization, speech rate, pause ratio, average pitch, average energy, and spectral features are extracted from the speech signal. For embodiments requiring deep fusion, a high-energy window can be extracted after audio downsampling, followed by spectral transformation and lightweight convolutional coding. Then, a global audio vector is obtained through channel attention, which is then fused together with visual or textual features.

[0049] The scale scores generated through keyboard or screen interaction are normalized to generate standardized features for subsequent model input. Specifically, this includes normalizing the raw scores of the Depression Scale (PHQ-9), Anxiety Scale (GAD-7), Stress Scale (PSS-10), sleep quality score, and number of recent stressful events to obtain corresponding normalized feature values. Self-harm ideation markers are retained as binary features. Semantic analysis is performed on the user-input free text to generate a text summary, which serves as part of the structured risk evidence.

[0050] Step S03: Concatenate the comprehensive feature set into an input vector, input it into the local multi-disease MLP, output the five-category risk probabilities, and perform structured scale security calibration.

[0051] The comprehensive feature set obtained in step S02 is concatenated into a 13-dimensional input vector: normalized values ​​of the depression scale (PHQ-9), the anxiety scale (GAD-7), the stress scale (PSS-10), the sleep quality, the stressful events, self-harm ideation markers, and eye-tracking variance entropy. gaze instability beta relative power Difference between focus and relaxation The normalized value of theta / alpha / beta power ratio gamma relative power and EEG quality .

[0052] The model uses a two-layer neural network for risk reasoning: , , , in, Represents the input vector. and Represents the model weight matrix. and Indicates the bias term; Represents the hidden layer vector; ReLU represents the rectified linear activation function; Represents the sigmoid function; Represents the MLP output vector; output vector ,in, Indicates the probability of depression. Indicates the probability of anxiety risk. Indicates the probability of stress risk. Indicates the probability of sleep risk. This represents the probability of cognitive load risk. To prevent high-scale risk or crisis signals from being suppressed by the model, the comprehensive analysis module also takes the maximum value of each item in the model output and the rule safety lower limit to obtain the calibrated risk probability.

[0053] To prevent the model from outputting excessively low probabilities under high-scale risk or crisis signals, this scheme incorporates a structured scale safety calibration layer. This safety calibration layer generates a rule-based safety lower bound vector, Yrule, based on PHQ-9, GAD-7, PSS-10, sleep quality score, number of recent stressful events, and self-harm intention markers. This vector is then compared with the MLP output, and the larger value is taken for each item. , In the formula, This represents the safety lower bound vector obtained from the structured scale and textual risk rules. This represents the calibrated probabilities of the five risk categories. If the self-harm intent flag is true, the crisis-level rule is triggered first, and the process proceeds to manual review or crisis intervention. This calibration allows the model to utilize the non-linear patterns learned during training while maintaining a safety lower limit when clear risk signals such as scales, self-harm, and decreased attention appear.

[0054] Step S04: Calculate independent scores for each channel based on multimodal features, and perform weighted fusion with MLP output. Combine risk threshold and self-harm intention escalation rules to generate risk level.

[0055] In this embodiment, the independent scores for each channel are calculated as follows: Eye tracking score The fixation point, pupil, and eye movement trajectory data were collected from a high-definition camera and processed by eye movement variance entropy. gaze stability and saccade intensity The result is obtained through conversion. higher lower or A higher score indicates a more discrete attentional trajectory, weaker fixation stability, or more pronounced eye movement, resulting in a correspondingly higher eye movement score.

[0056] EEG score EEG frequency data from the physiological data acquisition unit or EEG acquisition module, after... Overall power ratio , relative power , relative power Difference between focus and relaxation and EEG quality analysis The conversion is obtained. If... In short, EEG scores are only used as a limiting factor in the report description.

[0057] The model risk score Sm is derived from the output of the local multi-disease MLP and is calculated using five risk probabilities after safety calibration. .

[0058] Structured scale scoring The results are obtained from PHQ-9, GAD-7, PSS-10, sleep quality score, number of recent stressful events, self-harm ideation markers, and text summaries entered from the keyboard or interactive components, after normalization and rule-based conversion. When the self-harm ideation marker is true, According to crisis management rules, it should be handled directly with the highest priority.

[0059] The four types of scores were all converted to a value between 0 and 1, with higher values ​​indicating a higher risk level for that modality. A final fusion score was then calculated based on preset weights. , In the formula, Indicates the total score after integration. This indicates the eye movement score. Indicates EEG score, Indicates risk score, This indicates the score on the structured scale.

[0060] When the total fusion score is less than 0.42, the output is low risk; between 0.42 and 0.55, the output is medium risk; and greater than or equal to 0.55, the output is high risk. When the structured scale or text input contains self-harm intentions, a crisis-level prompt is output directly, and the total fusion score is raised to no less than 0.95.

[0061] Step S05: Generate an auxiliary evaluation report through a multi-agent collaborative review mechanism.

[0062] After the fusion results are generated, the data processing unit runs a multi-agent decision chain in node order, including a data parsing agent, an eye-tracking evidence agent, an EEG evidence agent, an MLP risk interpretation agent, a fusion decision agent, a report generation agent, and a self-censorship agent. The data parsing agent is responsible for verifying the source and quality fields of the input data; the eye-tracking evidence agent is responsible for interpreting eye-tracking variance entropy, fixation stability, and saccade intensity; the EEG evidence agent is responsible for interpreting... Overall power ratio relative power, The system considers relative power and the difference between focus and relaxation levels. The MLP risk interpretation agent is responsible for interpreting the probabilities of five risk categories and their main input characteristics. The fusion decision agent is responsible for determining the fusion total score and risk level. The report generation agent is responsible for generating a readable report. The self-censorship agent is responsible for checking the evidence chain, data quality constraints, non-diagnostic claims, and human review exits. Each agent node records the source of input data, the processing tools invoked, the processing summary, key evidence, confidence level, constraints, and output results, forming a traceable agent operation trajectory.

[0063] Example 2

[0064] Please see Figure 2-5 This embodiment provides a diagnostic and treatment device, including: a main unit base 1 and a main unit box 2, the main unit box 2 is installed on the main unit base 1, and a data processing unit is provided inside the main unit box 2; The host unit 1 is equipped with an interactive component, which is used to interact with the patient to exchange information during the diagnosis and treatment process; The main unit 2 is equipped with a data acquisition module, which includes a PTZ camera, a voice acquisition unit, and a physiological data acquisition unit. The PTZ camera is used to track the user's face and acquire facial expression images and eye movement image data. The voice acquisition unit is used to acquire the patient's voice signal. The physiological data acquisition unit is used to acquire the patient's physiological signs data.

[0065] This device uses main unit 1 as its supporting base and main unit chassis 2 as its core carrier, with a built-in data processing unit for centralized control. The interactive components are located on main unit 1, providing users with an intuitive operating interface and information feedback channels. A pan-tilt-zoom camera dynamically tracks and captures facial and eye features, working in conjunction with a voice acquisition unit and a physiological data acquisition unit to achieve multi-dimensional synchronous acquisition of visual, auditory, and physiological characteristics, providing a raw data foundation for subsequent psychological state assessment.

[0066] The PTZ camera includes a rotating part, a tilting part, and a high-definition camera 3. The rotating part is mounted on the main unit 2, the tilting part is mounted on the rotating end of the rotating part, and the high-definition camera 3 is mounted on the tilting end of the tilting part. The voice acquisition unit includes a microphone 4, which is mounted on the frame of the high-definition camera 3.

[0067] The pan-tilt camera, through the combination of its rotating and tilting components with the high-definition camera 3, forms a spatial positioning structure with both horizontal and vertical degrees of freedom. The high-definition camera 3 can adjust its pointing direction in real time according to the user's position changes, ensuring the acquisition of high-definition facial images. The microphone 4 is located on the camera's frame; its relatively close position to the user's head effectively reduces environmental noise interference and improves the signal-to-noise ratio and acquisition quality of the voice signal.

[0068] The rotating part includes a fixed cylinder 5 and a rotating cylinder 6. The rotating cylinder 6 is rotatably installed inside the fixed cylinder 5. A steering motor 7 is fixedly installed inside the rotating cylinder 6. A drive gear 8 is fixedly installed at the drive end of the steering motor 7. A driven gear ring 9 is fixedly installed inside the fixed cylinder 5. The drive gear 8 and the driven gear ring 9 are meshed and connected. When the drive end of the steering motor 7 rotates, the rotating cylinder 6 is driven to rotate relative to the fixed cylinder 5 through the meshing motion of the drive gear 8 and the driven gear ring 9.

[0069] The steering motor 7 drives the drive gear 8 to rotate. The meshing of the drive gear 8 with the driven gear ring 9 inside the fixed cylinder 5 converts the motor's rotational motion into a horizontal deflection motion of the rotating cylinder 6 relative to the fixed cylinder 5. This gear transmission structure provides a stable transmission ratio, ensuring smooth horizontal rotation and precise positioning of the camera.

[0070] The pitch unit includes a pitch motor 10 and a gear reducer 11. The pitch motor 10 is fixedly installed inside the rotating cylinder 6, and the gear reducer 11 is fixedly installed on the rotating cylinder 6. The high-definition camera 3 is fixedly installed at the output end of the gear reducer 11, and the input end of the gear reducer 11 is fixedly connected to the drive end of the pitch motor 10. The pitch unit drives the gear reducer 11 through the pitch motor 10, thereby causing the high-definition camera 3 to tilt vertically. In this process, the gear reducer 11 increases the output torque and acts as a self-locking mechanism, allowing the camera to adapt to users of different heights or sitting postures. The pitch mechanism, in conjunction with the horizontal rotation mechanism, achieves omnidirectional coverage of the camera's field of view.

[0071] The physiological data acquisition unit includes a finger clip pulse oximeter and / or a smart bracelet, used to monitor the user's blood oxygen, heart rate and / or pulse data in real time, and transmit the data to the data processing unit in the main unit 2 in real time.

[0072] The physiological data acquisition unit contacts the user's body via a clip-on pulse oximeter or smart bracelet to sense key physiological indicators such as blood oxygen saturation, heart rate, and pulse in real time. These physiological signals are collected by the data processing unit inside the main unit 2 via wired or wireless transmission, serving as important supplementary parameters for stress levels and autonomic nervous system status in the multimodal assessment system.

[0073] The main unit 2 is equipped with a memory that stores mental health assessment models and auxiliary treatment plans. The data processing unit is electrically connected to the data acquisition module, the memory, and the interactive components.

[0074] The memory serves as the device's static data center, pre-stored with trained mental health assessment models and multiple treatment protocols. The models within the memory process real-time data, and interactive components display the final results, ensuring a seamless workflow for data collection, analysis, and feedback.

[0075] The interactive components include a keyboard 12 and a display screen 13 mounted on the main unit 1.

[0076] The interactive components consist of a keyboard 12 and a display screen 13. The display screen 13 is responsible for presenting visual information such as diagnostic questions, digital human guidance interface, and assessment result reports; the keyboard 12 serves as the physical interface for users to actively input information, allowing users to answer psychological assessment questionnaires or conduct text communication in text form, thus expanding the device's ability to acquire semantic information.

[0077] The data processing unit is used to execute a method for mental health monitoring and assisted diagnosis based on multimodal recognition as described in Example 1, which will not be elaborated here.

[0078] like Figure 2-5As shown, when the power is turned on and the device is started, the data processing unit inside the main unit 2 performs a self-test on the data acquisition module and interactive components. The device enters standby mode, and the display screen 13 presents the initial digital human interaction interface. When the user is in front of the device, the data processing unit recognizes the user's facial position through the initial image acquired by the high-definition camera 3. The data processing unit sends control commands to the pan-tilt camera: driving the steering motor 7 in the rotating part, which drives the rotating cylinder 6 to rotate horizontally through the meshing transmission of the active gear 8 and the driven gear ring 9; driving the pitch motor 10 in the pitch part, which drives the high-definition camera 3 to adjust the vertical angle through the gear reducer 11. Through the combined adjustment of the horizontal and vertical dimensions, the high-definition camera 3 is always locked and aimed at the user's facial area. During the diagnosis and treatment interaction, the device simultaneously activates multi-channel data acquisition: the high-definition camera 3 captures the user's facial expression image stream in real time and records eye movement data such as the fixation point position and pupil diameter. The voice acquisition unit acquires the user's voice signal in real time. The user wears the physiological data acquisition unit, and the device reads vital signs data such as blood oxygen saturation, heart rate, and pulse in real time. Users input questionnaire answers or engage in text conversations with the digital human via keyboard 12 to generate text data.

[0079] The data processing unit performs steps S02-S05 on the acquired data, which will not be described in detail here.

[0080] In one data-supported embodiment, the model version stored in the memory is "Public EEG Supervised Training Multilayer Perceptron Version 0.2". The model input includes 13 normalized features, and the output includes five risk probabilities: depression, anxiety, stress, sleep, and cognitive load. The training table used for engineering validation is real_eeg_dnn_dataset.csv, which contains 522 sample records and 39 fields. The fields include sample number, source file, public dataset source, MDD / Healthy group label, number of EEG exported rows, number of EEG channels, scale placeholder field, eye movement features, EEG frequency band features, quality score, normalized input features, and five risk outputs.

[0081] The 522 sample records mentioned above were derived from feature tables converted from publicly available EEG data using a local engineering script. The depression risk head used the MDD / Healthy grouping labels from the publicly available EEG file as engineering supervision signals; anxiety, stress, sleep, and cognitive load outputs were used as auxiliary risk heads, and were combined with structured scales for safety calibration. The model file contained 105 test samples; in the label engineering validation of this publicly available EEG file, the depression binary classification engineering indices accuracy, precision, recall, and F1 were all 1.0. The mean absolute error is 0.0021.

[0082] In one publicly available data access embodiment, the system converts the publicly available EEG samples from figshare_4244171_mdd_healthy_controls into uploadable engineering examples. The index.csv file includes two examples: PUBLIC_H_S1_EC and PUBLIC_MDD_S1_EC. The PUBLIC_H_S1_EC source file is H S1 EC.edf, with 300 records, 23 EEG channels, and 60 rows of exported EEG features. The PUBLIC_MDD_S1_EC source file is MDD S1 EC.edf, with 303 records, 21 EEG channels, and 60 rows of exported EEG features. Both examples use EEG Fp1-LE, EEG F3-LE, EEG C3-LE, EEG P3-LE, EEG O1-LE, and EEG F7-LE channels for generation. , , , , Frequency band power characteristics.

[0083] It should be noted that the publicly available EEG dataset does not include eye-tracking channels or the system's mini-program scale fields. Therefore, in this publicly available data access embodiment, the EEG band power CSV is derived from real EDF signals, the eye-tracking CSV is neutral placeholder data, and the scale JSON and MLP risk output JSON are label adaptation fields used to verify the data reception, field validation, feature extraction, fusion decision-making, and report generation processes of this invention. The system will not directly use the MDD / Healthy label in the publicly available file name as the final risk level, but will instead store it as data source and engineering verification information in the evidence chain and constraints.

[0084] In one upload verification embodiment, the data processing unit receives four types of files: eye-tracking CSV, EEG CSV, MLP risk output JSON, and front-end structured scale JSON. It verifies the eye-tracking CSV fields, eye-tracking sample size, EEG CSV fields, EEG sample size, five-category risk probabilities, model name / version, scale fields, and probability value ranges. For a set of upload sessions that pass verification, the eye-tracking CSV contains 60 lines, the EEG CSV contains 60 lines, and the probabilities for all five categories of risks are between 0 and 1, and are recorded. and To facilitate traceability; for upload sessions lacking scale fields or model versions, the system writes them into the restriction conditions, without directly amplifying the risk level.

[0085] In a medium-risk operational example, the input data included 60 lines of eye-tracking data, 60 lines of EEG band power data, structured scale fields, and five risk categories of MLP outputs. The fusion results included evidence such as higher EEG beta relative power, significantly lower difference between focus and relaxation, and a higher number of recent stressful events; the EEG score was 0.65, the MLP model score was 0.639, the structured scale score was 0.625, the total fusion score was 0.485, the confidence level was 0.95, the system output was medium risk, and provided suggestions for follow-up within 1 to 2 weeks, re-measurement of the scale, tracking changes in stressful events, and retaining a manual review exit.

[0086] In a publicly available MDD example model output embodiment, the MLP risk output corresponding to PUBLIC_MDD_S1_EC includes: depression risk 0.9269, anxiety risk 0.6199, stress risk 0.6885, sleep risk 0.5570, and cognitive load risk 0.3406, with the highest risk label being depression risk. The same output file records the model_version as public-eeg-supervised-mlp-v0.2 and includes the limitation that "non-depression output is an auxiliary head, the model is only used for engineering validation and auxiliary screening, and does not constitute a clinical diagnosis." The above data is used to demonstrate that the present invention can preserve the original source, input fields, model version, probability output, limitations, and reporting evidence chain.

[0087] A visual mental health assessment report is output on display screen 13 for users or doctors to view. The auxiliary diagnosis and treatment module drives the digital human avatar on display screen 13, switching corresponding expressions, tone of voice, and dialogue based on the user's psychological state to provide personalized psychological guidance and suggestions. Based on the assessment conclusions, the system matches and recommends corresponding personalized treatment plans from its memory. The device records user feedback during the interaction process and stores the feedback results for subsequent targeted optimization of the assessment model and auxiliary diagnosis and treatment plans.

[0088] While the spirit and principles of the invention have been described with reference to several specific embodiments, it should be understood that the invention is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for ease of description. The invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

[0089] Regarding the limitation of the scope of protection of this invention, those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solution of this invention are still within the scope of protection of this invention.

Claims

1. A method for mental health monitoring and assisted diagnosis based on multimodal recognition, characterized in that, The method includes: S01: The data processing unit receives and preprocesses multimodal data from eye tracking, electroencephalography, facial recognition, speech, scales, and text. S02: Construct a comprehensive feature set reflecting attentional state, physiological arousal, emotional expression, and subjective risk by extracting quantifiable features from multimodal data; S03: Concatenate the comprehensive feature set into an input vector, input it into the local multi-disease MLP, output five risk probabilities, and perform structured scale security calibration; S04: Calculate independent scores for each channel based on multimodal features, and perform weighted fusion with MLP output. Combine risk thresholds and self-harm intention escalation rules to generate risk levels. Step S05: Generate an auxiliary evaluation report through a multi-agent collaborative review mechanism.

2. The method for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 1, characterized in that, The preprocessing described in step S01 includes: adding a quality field to each modal data and removing data segments that do not meet quality standards; and mapping data with different sampling frequencies to a unified time reference using a preset time window as the unit.

3. The method for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 1, characterized in that, The quantifiable features mentioned in step S02 include: Eye movement variance entropy, fixation stability, and saccade intensity extracted from eye movement data; The theta / alpha / beta power ratio, relative beta power, relative gamma power, and attention-relaxation difference were extracted from EEG data. The comprehensive score, classification results, facial feature vector, and facial image quality score extracted from facial images; Speech rate, pause ratio, average pitch, average energy, and spectral features extracted from speech signals; The raw scores of the depression scale, anxiety scale, stress scale, sleep quality, and number of recent stressful events were extracted from the scales and text.

4. The method for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 1, characterized in that, The structured scale safety calibration mentioned in step S03 specifically involves taking the maximum value of the risk probability output by the MLP and the safety lower limit probability output based on the scale rules, to obtain the calibrated risk probability.

5. The method for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 1, characterized in that, Step S04 specifically includes: Calculate eye movement score, electroencephalogram score, model risk score, and structured scale score respectively; The four scores are weighted and fused according to preset weights to obtain a total fused score; The total score is compared with a preset risk threshold to generate low, medium, and high risk levels; When a self-harm intention is detected, a crisis-level warning is output directly, and the total score is raised to the preset crisis threshold.

6. The method for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 1, characterized in that, The multi-agent collaborative review mechanism described in step S05 includes a data parsing agent, an eye-tracking analysis agent, an EEG analysis agent, a DNN interpretation agent, a fusion decision-making agent, a report generation agent, and a self-review agent that run sequentially. Each agent node records inputs, tools, summaries, evidence, confidence levels, constraints, and outputs, forming a traceable decision-making chain.

7. A device for mental health monitoring and assisted diagnosis based on multimodal recognition, characterized in that, The device includes: A host base (1) and a host chassis (2), wherein the host chassis (2) is mounted on the host base (1), and a data processing unit is provided inside the host chassis (2), wherein the data processing unit is used to implement the method for monitoring and assisting diagnosis and treatment of mental health based on multimodal recognition as described in any one of claims 1-6; An interactive component is provided on the host base (1), which is used to interact with the patient to exchange information during the diagnosis and treatment process; The host chassis (2) is equipped with a data acquisition module, which includes a PTZ camera, a voice acquisition unit and a physiological data acquisition unit. The PTZ camera is used to track the user's face and acquire the user's facial expression image and eye movement image data. The voice acquisition unit is used to acquire the patient's voice signal. The physiological data acquisition unit is used to acquire the patient's physiological signs data.

8. The device for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 7, characterized in that, The PTZ camera includes a rotating part, a tilting part, and a high-definition camera (3). The rotating part is mounted on the main unit (2), the tilting part is mounted on the rotating end of the rotating part, and the high-definition camera (3) is mounted on the tilting end of the tilting part. The voice acquisition unit includes a microphone (4), which is mounted on the frame of the high-definition camera (3). The physiological data acquisition unit includes a finger clip pulse oximeter and / or a smart bracelet for real-time monitoring of the user's blood oxygen, heart rate, and / or pulse data, and for real-time transmission of the data to the data processing unit inside the main unit (2).

9. The device for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 8, characterized in that, The rotating part includes a fixed cylinder (5) and a rotating cylinder (6). The rotating cylinder (6) is rotatably installed inside the fixed cylinder (5). A steering motor (7) is fixedly installed inside the rotating cylinder (6). A drive gear (8) is fixedly installed at the drive end of the steering motor (7). A driven gear ring (9) is fixedly installed inside the fixed cylinder (5). The drive gear (8) and the driven gear ring (9) are meshed and connected. When the drive end of the steering motor (7) rotates, the rotating cylinder (6) is driven to rotate relative to the fixed cylinder (5) through the meshing motion of the drive gear (8) and the driven gear ring (9). The pitch unit includes a pitch motor (10) and a gear reducer (11). The pitch motor (10) is fixedly installed inside the rotating cylinder (6), and the gear reducer (11) is fixedly installed on the rotating cylinder (6). The high-definition camera (3) is fixedly installed at the output end of the gear reducer (11), and the input end of the gear reducer (11) is fixedly connected to the drive end of the pitch motor (10).

10. The device for mental health monitoring and assisted diagnosis based on multimodal recognition according to claim 7, characterized in that, The main unit (2) is equipped with a memory that stores a mental health assessment model and auxiliary treatment plan. The data processing unit is electrically connected to the data acquisition module, the memory and the interactive component respectively. The interactive component includes a keyboard (12) and a display screen (13) set on the main unit (1).