Vision-text based mental disorder screening method, device, equipment and medium
By using a vision-text-based approach, combining facial video and PPG signals with a psychological assessment scale, a mental disorder screening method was achieved that does not rely on subjective reports, thus improving the accuracy of the screening.
Patent Information
- Application Number
- CN202511257638.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing methods for screening mental disorders rely on subjective self-reports from those being screened, resulting in low screening accuracy.
A vision-text-based approach is adopted, which collects facial videos and uses video feature extraction models and mental disorder classification models for screening. The PPG signal and psychological assessment scale are combined for training and testing to achieve screening without relying on subjective reports.
It improves the accuracy of mental disorder screening and can truly reflect the actual situation of the examinee.
Smart Images

Figure CN120747102B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to the technical field of mental illness diagnosis, and particularly relates to a mental disorder screening method and device based on vision-text, equipment and medium. BACKGROUND
[0002] Mental disorder (for example, dementia, attention deficit disorder, learning disorder, schizophrenia, mood disorder, depression and addiction disorder) screening is crucial for effective intervention and monitoring of diseases. Unlike many physiological diseases, mental disorder lacks clear biological diagnostic criteria. Traditional mental disorder screening mainly relies on structured interviews, clinical questionnaires and behavior observation, such as depression screening scale (PHQ-9), Beck depression scale (BDI) and generalized anxiety disorder scale (GAD-7). The above methods highly depend on the subjective self-report of the examinee in the evaluation process, and some examinees may have self-denial and deliberate concealment, resulting in that the mental disorder screening result cannot truly reflect the actual situation, and the accuracy of the screening is low. SUMMARY
[0003] The embodiment of the present application provides a mental disorder screening method and device based on vision-text, equipment and medium, and aims to solve the problem of low accuracy of existing mental disorder screening.
[0004] In a first aspect, the embodiment of the present application provides a mental disorder screening method based on vision-text, comprising:
[0005] A face video of the examinee when receiving detection is collected, and the face video is input into a video feature extraction model for feature extraction to obtain a visual feature, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model by using a sample mental disorder data set;
[0006] The visual feature is input into a mental disorder classification model for classification to obtain a classification result, wherein the mental disorder classification model is obtained by training and testing a classifier by using a sample psychological test.
[0007] In a second aspect, the embodiment of the present application further provides a mental disorder screening device based on vision-text, comprising:
[0008] A collection and extraction unit is configured to collect a face video of the examinee when receiving detection, and input the face video into a video feature extraction model for feature extraction to obtain a visual feature, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model by using a sample mental disorder data set;
[0009] The classification unit is configured to input the visual feature into a mental disorder classification model to obtain a classification result, wherein the mental disorder classification model is obtained by training and testing a classifier using a sample psychological test.
[0010] In a third aspect, an embodiment of the present application further provides a computer device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the above method when executing the computer program.
[0011] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program can implement the above method when being executed by a processor.
[0012] An embodiment of the present application provides a mental disorder screening method and device based on vision-text, equipment and medium. Wherein, the method comprises: collecting the face video of the detection object when receiving detection, and inputting the face video into a video feature extraction model to extract features to obtain visual features, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model using a sample mental disorder data set; inputting the visual features into a mental disorder classification model to classify and obtain a classification result, wherein the mental disorder classification model is obtained by training and testing a classifier using a sample psychological test. The technical scheme of the embodiment of the present application first inputs the face video into the video feature extraction model to extract features to obtain the visual features, and then inputs the visual features into the mental disorder classification model to classify and obtain the classification result, so as to screen the mental disorder. In the whole mental disorder screening process, it is not necessary to rely on the subjective self-report of the examinee, and the actual situation of the examinee can be truly reflected, so that the accuracy of the mental disorder screening is improved. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the technical scheme of the embodiment of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0014] Figure 1 A flowchart of a mental disorder screening method based on vision-text provided by an embodiment of the present application is shown in the figure;
[0015] Figure 2 A sub-flowchart of a mental disorder screening method based on vision-text provided by an embodiment of the present application is shown in the figure;
[0016] Figure 3Another sub-process schematic diagram of a vision-text-based mental disorder screening method provided by an embodiment of the present application;
[0017] Figure 4 Another sub-process schematic diagram of a vision-text-based mental disorder screening method provided by an embodiment of the present application;
[0018] Figure 5 A schematic block diagram of a vision-text-based mental disorder screening device provided by an embodiment of the present application;
[0019] Figure 6 A schematic block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present application.
[0021] It should be understood that, when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] It should also be understood that the terms used in the present application specification are only for the purpose of describing particular embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms as well.
[0023] It should be further understood that the term "and / or" used in the present application specification is intended to refer to any combination of one or more of the associated listed items and all possible combinations thereof.
[0024] As used in the present application specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted as meaning "upon determining" or "in response to determining" or "upon detecting [a described condition or event]" or "in response to detecting [a described condition or event]" depending on the context.
[0025] Referring to Figure 1 , Figure 1 is a flowchart of a vision-text-based mental disorder screening method provided by an embodiment of the present application. The vision-text-based mental disorder screening method will be described in detail below. As shown in Figure 1 , the method comprises the following steps S110-S120.
[0026] S110, a face video of a detection object when receiving detection is collected, and the face video is input into a video feature extraction model for feature extraction to obtain a vision feature, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model using a sample mental disorder data set.
[0027] In the embodiment of the present application, for the convenience of understanding, the training and testing process of the video feature extraction model will be described in detail first. In the present embodiment, the sample mental disorder data set comprises a sample face video, a sample PPG signal and sample text description information, as shown in Figure 2As shown, the step of training and testing the video feature model and the text feature model by using the sample mental disorder dataset to obtain the video feature extraction model can specifically include steps S111-S114: S111, collecting the sample facial video and the sample PPG signal of the tested person when receiving the experimental stimulus; S112, obtaining the sample text description information corresponding to the sample facial video and the sample PPG signal; S113, training and testing the video feature model by using the sample facial video and the sample PPG signal to obtain a target video feature model; and S114, training and testing the target video feature model and the text feature model by using the sample facial video and the sample text description information to obtain the video feature extraction model. It should be noted that in the present embodiment, the PPG (Photoplethysmography, blood volume pulse wave) signal is a biomedical signal for non-invasive monitoring of blood volume changes through optical principles, and is widely used in the detection of physiological parameters such as heart rate, blood oxygen saturation, blood pressure, etc. The sample description information is obtained by professional clinicians to describe the reaction of the tested person during the emotional picture / video stimulation / interview dialogue experiment, for example, the description of "mood depression", "low interest", etc. to obtain the text description of the professional clinician for the tested person in the clinical scene. It should also be noted that in the present embodiment, when the tested person performs the emotional picture / video stimulation / interview dialogue experiment, the facial video and PPG signal of the tested person are collected by an automatic collection module as sample facial video and sample PPG signal. Understandably, during the experiment, the tested person needs to wear a smart watch or a blood oxygen detection instrument throughout the experiment in order to obtain the blood volume pulse wave (PPG) signal of the human body. The automatic collection module includes a server, an industrial camera, a display module, a PPG collection module, and a data storage module. When each experiment starts, the central control program of the server will automatically enter the detection process and start real-time face detection and PPG signal detection. If the face and PPG signal are normal, the central control program of the server will automatically start signal synchronization collection. Once the face moves out of the target area or the experimental period ends, the central control program of the server will stop signal collection and store the signal storage path into the information library of the tested person, and mark the collection time and collection duration. In addition, the tested person also needs to complete the corresponding psychological evaluation table according to his own situation, such as the PHQ-9 (Patient Health Questionnaire-9), BDI (Beck Depression Inventory), and GAD-7 (Generalized Anxiety Disorder-7).
[0028] In an embodiment, for example, in the present embodiment, as shown in Figure 3 The step S113 can specifically include steps S1131-S1137:
[0029] S1131, connecting a linear layer behind the encoder in the video feature model as an estimator;
[0030] S1132, dividing the sample face video into a sample training face video and a sample test face video;
[0031] S1133, dividing the sample PPG signal into a sample training PPG signal and a sample test PPG signal;
[0032] S1134, inputting the sample training face video into the encoder for feature extraction to obtain a video feature, and inputting the video feature into the estimator to obtain a predicted rPPG signal;
[0033] S1135, fine-tuning parameters of the encoder according to the predicted rPPG signal and the sample training PPG signal to obtain an initial video feature model;
[0034] S1136, inputting the sample test face video into the encoder of the initial video feature model for feature extraction to obtain a test video feature, and inputting the test video feature into the estimator to obtain a test rPPG signal;
[0035] S1137, determining whether to take the initial video feature model as the target video feature model according to the test rPPG signal and the sample test PPG signal.
[0036] In the embodiment of the present application, the video feature model is, for example, a MaskFusionNet and an rPPG-MAE network model, wherein the MaskFusionNet is a dual-stream fusion model used for estimating an rPPG signal, and the rPPG-MAE is a self-supervised learning non-contact physiological signal monitoring. Feature extraction is performed by using an encoder in the video feature model, and a linear layer is connected behind the encoder as an estimator to fine-tune the parameters of the encoder. It should be noted that in the present embodiment, the sample face video is divided into a sample training face video and a sample test face video according to a first preset ratio; the sample PPG signal is divided into a sample training PPG signal and a sample test PPG signal according to a second preset ratio. The first preset ratio and the second preset ratio can be the same or different, and are set according to the time requirement. Understandably, when the first preset ratio and the second preset ratio are the same, the division ratio is, for example, 4:1, and when the first preset ratio and the second preset ratio are different, for example, the first preset ratio is 4:1 and the second preset ratio is 7:3. It should be further noted that in the present embodiment, if the test mean absolute error calculated according to the SDNN of the test rPPG signal and the SDNN of the sample test PPG signal is less than a preset mean absolute error, it indicates that the model precision is higher, and then the initial video feature model is taken as the target video feature model, otherwise, the initial video feature model is retrained until the test mean absolute error is less than the preset mean absolute error.
[0037] Further, in the present embodiment, step S1135 comprises: determining a predicted SDNN according to the predicted rPPG signal, and determining a true SDNN according to the sample training PPG signal; calculating a mean absolute error according to the predicted SDNN and the true SDNN; fine-tuning the parameters of the encoder according to the mean absolute error to obtain the initial video feature model. It should be noted that in the present embodiment, the SDNN (Standard Deviation of NN intervals) is the standard deviation of the continuous normal heartbeat interval (NN interval), which is a core index for evaluating heart rate variability (HRV) and reflects the function of the autonomic nervous system. The sample training PPG signal is used as a supervision signal, the mean absolute error is calculated according to the predicted SDNN determined according to the predicted rPPG signal and the true SDNN determined according to the sample training PPG signal, and the parameters of the encoder are fine-tuned based on the mean absolute error to obtain the initial video feature model. It should be further noted that in the present embodiment, the rPPG signal is a remote PPG signal, that is, the PPG signal can be obtained remotely and non-contactly by using the encoder in the video feature model.
[0038] In an embodiment, for example, as shown in FIG. 11, step S114 can specifically include steps S1141-S1148: Figure 4
[0039] S1141, dividing the sample face video into a sample training face video and a sample test face video;
[0040] S1142, dividing the sample text description information into sample training text information and sample test text information;
[0041] S1143, inputting the sample training face video into the target video feature model to extract training visual features;
[0042] S1144, inputting the sample training text information into the text feature model to extract training text features;
[0043] S1145, inputting the training visual features and the training text features into the MLP network to obtain reduced dimension visual features and reduced dimension text features;
[0044] S1146, calculating a loss value according to the reduced dimension visual features and the reduced dimension text features through a loss function, and fine-tuning parameters in the target video feature model and the text feature model based on the loss value to obtain a fine-tuned video feature model and a fine-tuned text feature model;
[0045] S1147, inputting the sample test face video and the sample test text information into the fine-tuned video feature model and the fine-tuned text feature model respectively to obtain test visual features and test text features;
[0046] S1148, determining whether to use the fine-tuned text feature model as a text feature extraction model and the fine-tuned video feature model as the video feature extraction model according to the test visual features and the test text features.
[0047] In the embodiment of the present application, the sample face video is divided into a sample training face video and a sample test face video according to a third preset ratio, and the sample text description information is divided into a sample training text information and a sample test text information according to a fourth preset ratio. The third preset ratio, the fourth preset ratio, the first preset ratio and the second preset ratio can be the same or different, and are determined according to actual requirements. It should be noted that in the embodiment, the text feature model is a BERT model, the loss value is calculated by the loss function according to the reduced visual feature and the reduced text feature, and the parameters in the target video feature model and the text feature model are fine-tuned based on the loss value, that is, the similarity between the reduced visual feature and the reduced text feature is maximized by the method of contrastive learning, that is, the noise data in the sample training face video and the sample training text information is removed, and finally the training of the cross-modal aligned visual-text model is realized, that is, the training of the target video feature model and the text feature model is completed to obtain a text feature extraction model and a video feature extraction model, and the video feature extraction model supports the adaptation to a specific downstream task. It should be further noted that in the embodiment, contrastive learning is a method of learning feature representation by pulling positive sample pairs (similar samples) and pushing negative sample pairs (dissimilar samples) apart. In a cross-modal task, the goal is to make the text features and video features of the same content as close as possible in the embedding space, while making the text features and video features of different contents as far apart as possible. Understandably, if the similarity value between the test visual feature and the test text feature is greater than a preset similarity value, the fine-tuned text feature model is used as the text feature extraction model and the fine-tuned video feature model is used as the video feature extraction model, otherwise, the target video feature model and the text feature model are continuously trained until the similarity value is greater than the preset similarity value.
[0048] Further, the loss function protects the first loss function and the second loss function, and the calculation of the loss value according to the reduced visual feature and the reduced text feature by the loss function includes: performing normalization processing on the reduced visual feature and the reduced text feature to obtain a normalized visual feature and a normalized text feature; calculating a first loss value and a second loss value according to the normalized visual feature and the normalized text feature by the first loss function and the second loss function; and calculating the sum of the first loss value and the second loss value to obtain the loss value. It should be noted that in the embodiment, the first loss function is shown in formula (1), and the second loss function is shown in formula (2). In formulas (1) and (2), v i and t i are the i-th normalized visual feature and normalized text feature, B is the number of selected samples for one training, As a hyperparameter, the value can be adjusted. It should be noted that in the embodiment, there are two kinds of samples in contrastive learning, one is a positive sample, and the other is a negative sample. The positive sample is a visual-text pair of the same semantics, that is, the sample training face video and the sample training text description information pair of the same semantics are taken as the positive sample. All unmatched visual-text pairs can be taken as negative samples, that is, the sample training face video and the sample training text description information pair of the unmatched are taken as negative samples. It should be noted that in the embodiment, the numerator of formulas (1) and (2) can constrain the model to pull the distance between the paired features, and the denominator can constrain the model to pull the distance between the non-matching features, so that the model improves the semantic judgment of the visual features from the learning of positive and negative samples. Understandably, the loss value is calculated to update the parameters in the target video feature model and the text feature model by back propagation;
[0049] (1);
[0050] (2);
[0051] In the embodiment, after the video feature extraction model is obtained by training and testing the video feature model and the text feature model using the sample mental disorder data set, the face video of the detection object during the detection is collected, and the face video is input into the video feature extraction model for feature extraction to obtain the visual feature.
[0052] S120, input the visual feature into the mental disorder classification model for classification to obtain a classification result, wherein the mental disorder classification model is obtained by training and testing the classifier using a sample psychological test form.
[0053] In the embodiment of the application, the visual feature is input into the mental disorder classification model for classification to obtain a classification result, wherein the classification result includes healthy, mild, moderate and severe. It should be noted that in the embodiment, the mental disorder classification model is obtained by training and testing the classifier using a sample psychological test form, including: obtaining sample visual features corresponding to the sample psychological test form;
[0054] The sample visual features are input into the classifier for training to obtain a mental disorder classification result; a classification accuracy is calculated according to the mental disorder classification result and an evaluation result corresponding to the sample psychological test; if the classification accuracy is greater than a preset classification accuracy, the classifier is taken as the mental disorder classification model. Understandably, if the classification accuracy is not greater than the preset classification accuracy, the classifier is retrained until the classification accuracy is greater than the preset classification accuracy. It should be further pointed out that in the embodiment, the classifier can be a deep residual neural network or a Softmax function, as long as classification can be realized. Understandably, if the examination object is a depressive patient, the sample psychological test is a depressive screening scale; if the examination object is an anxious patient, the sample psychological test is an anxious scale.
[0055] To sum up, in the embodiment, in the process of training the video feature model, the PPG signal is used as a supervision signal to fine-tune the video feature model to obtain a target video feature model; in the process of training the target video feature model and the text feature model, through contrast learning based on visual features and text features, the face video and the corresponding clinical text description information are mapped to a unified semantic space, so that the similarity of matched mental disorder visual features is maximized and the similarity of unmatched mental disorder visual features is minimized, to mine the potential representation mode shared by the mental disorder patients through visual modal face micro-movement and physiological characteristics, and on this basis, the target video feature model is fine-tuned to obtain the visual features extracted by feature extraction; the visual features of the face video of the examination object during detection are extracted based on the visual features extracted by feature extraction, the visual features are input into the mental disorder classification model for classification to obtain a classification result, and in the whole mental disorder screening process, the subjective self-report of the examination object is not needed, the actual situation of the examination object can be truly reflected, and the accuracy of mental disorder screening is improved.
[0056] Figure 5 is a schematic block diagram of a mental disorder screening device 200 based on vision-text provided by an embodiment of the present application. As shown in Figure 5 corresponding to the above mental disorder screening method based on vision-text, the present application further provides a mental disorder screening device 200 based on vision-text. The mental disorder screening device 200 based on vision-text includes units for executing the above mental disorder screening method based on vision-text, and the device can be configured in a computer device. Specifically, please refer to Figure 5 , the mental disorder screening device 200 based on vision-text includes a collection and extraction unit 201 and a classification unit 202.
[0057] The collection extraction unit 201 is configured to collect a face video of a detection object when the detection object is detected, and input the face video into a video feature extraction model to extract visual features, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model using a sample mental disorder data set.
[0058] In some embodiments, such as the present embodiment, the step of obtaining the video feature extraction model by training and testing the video feature model and the text feature model using the sample mental disorder data set includes a collection unit, a first acquisition unit, a first training and testing unit, and a second training and testing unit.
[0059] The collection unit is configured to collect the sample face video and the sample PPG signal of a testee when the testee is stimulated; the first acquisition unit is configured to acquire sample text description information corresponding to the sample face video and the sample PPG signal; the first training and testing unit is configured to train and test the video feature model using the sample face video and the sample PPG signal to obtain a target video feature model; and the second training and testing unit is configured to train and test the target video feature model and the text feature model using the sample face video and the sample text description information to obtain the video feature extraction model.
[0060] In some embodiments, such as the present embodiment, the first training and testing unit is specifically configured to: connect a linear layer as an estimator after an encoder in the video feature model; divide the sample face video into a sample training face video and a sample test face video; divide the sample PPG signal into a sample training PPG signal and a sample test PPG signal; input the sample training face video into the encoder to extract video features, and input the video features into the estimator to obtain a predicted rPPG signal; fine-tune parameters of the encoder according to the predicted rPPG signal and the sample training PPG signal to obtain an initial video feature model; input the sample test face video into the encoder of the initial video feature model to extract test video features, and input the test video features into the estimator to obtain a test rPPG signal; and determine whether to use the initial video feature model as the target video feature model according to the test rPPG signal and the sample test PPG signal.
[0061] In some embodiments, such as the present embodiment, the first training and testing unit is further configured to: determine a predicted SDNN according to the predicted rPPG signal, and determine a true SDNN according to the sample training PPG signal; calculate a mean absolute error according to the predicted SDNN and the true SDNN; and fine-tune parameters of the encoder according to the mean absolute error to obtain the initial video feature model.
[0062] In some embodiments, such as the present embodiment, the second training and testing unit is specifically configured to: divide the sample facial video into a sample training facial video and a sample testing facial video; divide the sample text description information into sample training text information and sample testing text information; input the sample training facial video into the target video feature model to extract training visual features; input the sample training text information into the text feature model to extract training text features; input the training visual features and the training text features into an MLP network to obtain reduced-dimension visual features and reduced-dimension text features; calculate a loss value according to the reduced-dimension visual features and the reduced-dimension text features through a loss function, and fine-tune parameters in the target video feature model and the text feature model based on the loss value to obtain a fine-tuned video feature model and a fine-tuned text feature model; input the sample testing facial video and the sample testing text information into the fine-tuned video feature model and the fine-tuned text feature model, respectively, to obtain testing visual features and testing text features; and determine whether to use the fine-tuned text feature model as a text feature extraction model and use the fine-tuned video feature model as the video feature extraction model according to the testing visual features and the testing text features.
[0063] In some embodiments, such as the present embodiment, the second training and testing unit is further configured to: normalize the reduced-dimension visual features and the reduced-dimension text features to obtain normalized visual features and normalized text features; calculate a first loss value and a second loss value according to the normalized visual features and the normalized text features through the first loss function and the second loss function; and calculate the loss value as a sum of the first loss value and the second loss value.
[0064] In some embodiments, such as the present embodiment, the step of training and testing the classifier using the sample psychological test to obtain the mental disorder classification model includes a second obtaining unit, a training unit, a calculating unit, and an unit.
[0065] The second acquisition unit is configured to acquire sample visual features corresponding to the sample psychological test; the training unit is configured to input the sample visual features into the classifier to obtain a mental disorder classification result; the calculation unit is configured to calculate a classification accuracy according to the mental disorder classification result and an evaluation result corresponding to the sample psychological test; and the as unit is configured to take the classifier as the mental disorder classification model if the classification accuracy is greater than a preset classification accuracy.
[0066] The vision-text-based mental disorder screening device can be implemented in the form of a computer program, which can run on a computer device as shown in the drawings. Figure 6 The computer device 300 is a device with a vision-text-based mental disorder screening function.
[0067] Please refer to Figure 6 , Figure 6 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 300 is a device with a vision-text-based mental disorder screening function.
[0068] Please refer to Figure 6 , the computer device 300 includes a processor 302, a memory and a network interface 305 connected through a system bus 301, wherein the memory can include a non-volatile storage medium 303 and an internal memory 304.
[0069] The non-volatile storage medium 303 can store an operating system 3031 and a computer program 3032. The computer program 3032, when executed, can enable the processor 302 to perform a vision-text-based mental disorder screening method.
[0070] The processor 302 is configured to provide computing and control capabilities to support the operation of the entire computer device 300.
[0071] The internal memory 304 provides an environment for the execution of the computer program 3032 in the non-volatile storage medium 303, and the computer program 3032, when executed by the processor 302, can enable the processor 302 to perform a vision-text-based mental disorder screening method.
[0072] The network interface 305 is configured to perform network communication with other devices. Those skilled in the art can understand that Figure 6 The structure shown in the drawings is only a block diagram of part of the structure related to the present application scheme, and does not constitute a limitation on the computer device 300 to which the present application scheme is applied. The specific computer device 300 can include more or fewer components than those shown in the drawings, or combine certain components, or have a different component arrangement.
[0073] The processor 302 is configured to run the computer program 3032 stored in the memory to implement any of the embodiments of the above-mentioned visual-text based mental disorder screening method.
[0074] It should be understood that, in the embodiments of the present application, the processor 302 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0075] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a storage medium, which is a computer readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments.
[0076] Therefore, the present application also provides a storage medium. The storage medium can be a computer readable storage medium. The storage medium stores a computer program. The computer program is executed by a processor to make the processor execute any of the above-mentioned embodiments of the visual-text based mental disorder screening method.
[0077] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, etc. Various computer readable storage media that can store program codes.
[0078] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0079] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. For example, the division of the units is merely a logical function division. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In this way, the actual division of the units can be different from the division in the embodiment.
[0080] The steps in the method embodiments of the present application can be adjusted, combined and deleted in sequence according to actual needs. The units in the apparatus embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0081] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in the embodiments of the present application.
[0082] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0083] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, these modifications and variations also belong to the scope of the claims of the present application and their equivalent technologies, and the present application is intended to include these modifications and variations.
[0084] The above description is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any modifications or replacements within the technical scope disclosed by the present application can be easily thought by those skilled in the art, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A visual-text based mental disorder screening method, characterized in that, The method comprises the following steps: Collecting a face video of a detection object when the detection object is being detected, and inputting the face video into a video feature extraction model to extract visual features, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model using a sample mental disorder dataset; Inputting the visual features into a mental disorder classification model to obtain a classification result, wherein the mental disorder classification model is obtained by training and testing a classifier using a sample psychological test form; The sample mental disorder dataset comprises sample face videos, sample PPG signals, and sample text description information, the video feature extraction model is obtained by training and testing the video feature model and the text feature model using the sample mental disorder dataset, and the method comprises the following steps: Collecting the sample face videos and the sample PPG signals of the test personnel when the test personnel are being tested; Obtaining the sample text description information corresponding to the sample face videos and the sample PPG signals; Training and testing the video feature model using the sample face videos and the sample PPG signals to obtain a target video feature model; Training and testing the target video feature model and the text feature model using the sample face videos and the sample text description information to obtain the video feature extraction model.
2. The method of claim 1, wherein, The method for training and testing the video feature model using the sample face videos and the sample PPG signals to obtain a target video feature model comprises the following steps: Connecting a linear layer to the encoder in the video feature model as an estimator; Dividing the sample face videos into sample training face videos and sample testing face videos; Dividing the sample PPG signals into sample training PPG signals and sample testing PPG signals; Inputting the sample training face videos into the encoder to extract video features, and inputting the video features into the estimator to obtain a predicted rPPG signal; Fine-tuning the parameters of the encoder according to the predicted rPPG signal and the sample training PPG signal to obtain an initial video feature model; Inputting the sample testing face videos into the encoder of the initial video feature model to extract test video features, and inputting the test video features into the estimator to obtain a test rPPG signal; Determining whether to use the initial video feature model as the target video feature model according to the test rPPG signal and the sample testing PPG signal.
3. The method of claim 2, wherein, The method for fine-tuning the parameters of the encoder according to the predicted rPPG signal and the sample training PPG signal to obtain an initial video feature model comprises the following steps: Determining a predicted SDNN according to the predicted rPPG signal, and determining a real SDNN according to the sample training PPG signal; Calculating a mean absolute error according to the predicted SDNN and the real SDNN; Fine-tuning the parameters of the encoder according to the mean absolute error to obtain the initial video feature model.
4. The method of claim 1, wherein, The training and testing of the target video feature model and the text feature model by using the sample face video and the sample text description information obtain the video feature extraction model, including: dividing the sample face video into sample training face video and sample test face video; dividing the sample text description information into sample training text information and sample test text information; inputting the sample training face video into the target video feature model for feature extraction to obtain training visual features; inputting the sample training text information into the text feature model for feature extraction to obtain training text features; inputting the training visual features and the training text features into an MLP network to obtain reduced dimension visual features and reduced dimension text features; calculating a loss value by a loss function according to the reduced dimension visual features and the reduced dimension text features, and fine-tuning parameters in the target video feature model and the text feature model based on the loss value to obtain a fine-tuned video feature model and a fine-tuned text feature model; inputting the sample test face video and the sample test text information into the fine-tuned video feature model and the fine-tuned text feature model respectively to obtain test visual features and test text features; determining whether to use the fine-tuned video feature model as the video feature extraction model according to the test visual features and the test text features.
5. The method of claim 4, wherein, The loss function includes a first loss function and a second loss function, and the calculation of the loss value by the loss function according to the reduced dimension visual features and the reduced dimension text features includes: normalizing the reduced dimension visual features and the reduced dimension text features to obtain normalized visual features and normalized text features; calculating a first loss value and a second loss value by the first loss function and the second loss function according to the normalized visual features and the normalized text features; calculating the sum of the first loss value and the second loss value to obtain the loss value.
6. The method according to any one of claims 1 to 5, characterized in that, The training and testing of the classifier by using the sample psychological test table obtain the mental disorder classification model, including: obtaining sample visual features corresponding to the sample psychological test table; training the classifier by inputting the sample visual features to obtain a mental disorder classification result; calculating a classification accuracy according to the mental disorder classification result and an evaluation result corresponding to the sample psychological test table; if the classification accuracy is greater than a preset classification accuracy, using the classifier as the mental disorder classification model.
7. A vision-text based mental disorder screening device, characterized in that, including: a collection and extraction unit configured to collect a face video of a detection object when the detection object is being detected, and input the face video into a video feature extraction model for feature extraction to obtain visual features, wherein the video feature extraction model is obtained by training and testing a video feature model and a text feature model by using a sample mental disorder dataset; a classification unit configured to input the visual features into a mental disorder classification model for classification to obtain a classification result, wherein the mental disorder classification model is obtained by training and testing a classifier by using a sample psychological test table; and The sample mental disorder dataset includes sample facial videos, sample PPG signals, and sample text description information, and the video feature extraction model is obtained by training and testing the video feature model and the text feature model using the sample mental disorder dataset, including: The acquisition unit is configured to acquire the sample facial videos and the sample PPG signals of the testee when the testee receives the experimental stimulus. The first acquisition unit is configured to acquire the sample text description information corresponding to the sample facial videos and the sample PPG signals. The first training and testing unit is configured to train and test the video feature model using the sample facial videos and the sample PPG signals to obtain a target video feature model. The second training and testing unit is configured to train and test the target video feature model and the text feature model using the sample facial videos and the sample text description information to obtain the video feature extraction model.
8. A computer device, comprising: The computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program can implement the method of any one of claims 1-6 when executed by a processor.
Citation Information
Patent Citations
Psychological distress evaluation instrument for tumor patient and evaluation method of instrument
CN110473630A