Depression early screening method based on machine learning
By constructing a multimodal fusion analysis framework, combining fuzzy logic and convolutional neural networks, integrating facial expressions, speech emotions and text emotions analysis, the accuracy and objectivity of traditional depression screening methods are solved, and more efficient early screening of depression is achieved.
Patent Information
- Application Number
- CN202510238831.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional depression screening methods rely on doctors’ communication with patients and are susceptible to doctors’ professional level and patient privacy concerns. Machine learning methods are susceptible to image noise and single modal data when screening facial expressions, resulting in low screening accuracy.
By analyzing a large amount of clinical data, early depression screening model is constructed, facial expressions, speech emotions and text emotions analysis are integrated, fuzzy logic is combined with convolutional neural network, noise suppression effect is optimized, and the weights of each model are adjusted through dynamic weighting functions to achieve multi-dimensional data collaborative analysis.
It improves the objectivity and accuracy of depression screening, enhances the ability to capture weak expression changes, and solves the problems of misjudgment and misjudgment caused by the limitations of single modal data.
Smart Images

Figure CN120148824A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of depression screening, and specifically refers to an early depression screening method based on machine learning. Background Art
[0002] Depression is a common mental disorder that affects patients' daily life, work, study, and social relationships. In severe cases, it can lead to the inability of patients to live normally. Detecting depressive symptoms in a timely manner and intervening can avoid various adverse consequences caused by patients' long-term depression, enabling patients to better return to normal life and re-establish a positive attitude towards life and a healthy lifestyle. Traditional depression screening methods mainly rely on communication between doctors and patients. The professional level, experience, and attitude of doctors can have a significant impact on the screening results. Moreover, patients may be reluctant to honestly disclose their symptoms and feelings due to concerns about privacy leakage, which also affects the screening results. When using machine learning to screen through facial expressions, it is prone to interference caused by image noise, uneven illumination, and blurred edges. At the same time, when traditional machine learning is used for depression screening, it usually relies on single-modal data and makes predictions based on one-sided information, resulting in relatively low screening accuracy. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides an early depression screening method based on machine learning. Aiming at the technical problem that traditional depression screening methods mainly rely on communication between doctors and patients, and the professional level, experience, and attitude of doctors can have a significant impact on the screening results, and patients may be reluctant to honestly disclose their symptoms and feelings due to concerns about privacy leakage, affecting the screening results, this solution constructs an early depression screening model by analyzing a large amount of clinical data, integrates various factors, can objectively evaluate the state of patients, and improve the efficiency of depression screening. Aiming at the problem that when using machine learning to screen through facial expressions, it is prone to interference caused by image noise, uneven illumination, and blurred edges, this solution combines fuzzy logic with a convolutional neural network, optimizes the noise suppression effect by correcting the membership degree value, improves the model's ability to capture weak facial expression changes, and can more accurately evaluate the emotional state of patients. Aiming at the technical problem that when traditional machine learning is used for screening, it usually relies on single-modal data and makes predictions based on one-sided information, resulting in relatively low screening accuracy, this solution constructs a multi-modal fusion analysis framework, designs a facial expression analysis model, a speech emotion analysis model, and a text emotion analysis model respectively, and designs a dynamic weighting function to adjust the weights of each model, realizing the collaborative analysis of multi-dimensional data, improving the comprehensive judgment ability of early depression screening, and solving the problems of misjudgment and missed judgment caused by the limitations of single-modal data.
[0004] The technical solution adopted by the present invention is as follows: The present invention provides a method for early screening of depression based on machine learning. The method for early screening of depression based on machine learning specifically includes the following steps:
[0005] Step S1: Clinical data collection. Video data during the patient's visit is collected, and the video data is cut into segments every 3 seconds. The frame images, audio, and transcribed text of each segment are clipped, and the clipped frame images, audio, and transcribed text are respectively input into the corresponding analysis models.
[0006] Step S2: Construct a facial expression analysis model. A facial expression analysis model is constructed by combining fuzzy logic and a convolutional neural network. The tendency of the patient to have depression is judged through the frame images, and the facial emotional state is output.
[0007] Step S3: Construct a voice emotion analysis model. By extracting the Mel-frequency cepstral coefficient features of the audio, a voice emotion analysis model is constructed using a long short-term memory network. The tendency of the patient to have depression is judged through the audio, and the audio emotional state is output.
[0008] Step S4: Construct a text emotion analysis model. This model is constructed by a pre-trained T5 model and fine-tuned using transfer learning. The tendency of the patient to have depression is judged through the transcribed text, and the text emotional state is output.
[0009] Step S5: Multi-model fusion screening. The facial emotional state, audio emotional state, and text emotional state are unified using a weighting function, and the prediction weights of the three models are updated by adding reward values. After fusion, an early depression screening model is obtained. The frame images, audio, and transcribed text are input into the early depression screening model to obtain the depression screening result.
[0010] Further, for step S2, constructing the facial expression analysis model specifically includes the following steps:
[0011] Step S21: Collect a facial image set. The facial image set includes facial images of depression patients and non-patients and their corresponding status labels. After preprocessing the facial image set, the facial image set is divided into a training set, a rule extraction set, a rule selection set, and a test set according to a ratio of 3:2:2:1.
[0012] Step S22: Establish and initialize a convolutional neural network. A facial expression analysis model is constructed by combining fuzzy logic and a convolutional neural network. A fuzzy logic layer is designed to process the facial image set, calculate the derivative of the fuzzy direction, use the directional gradient to distinguish the noise and edges of the facial image, calculate the membership degree value of the facial image, and define the membership function of the fuzzy logic. The formula used is as follows: ;
[0013] In the formula, is the membership value, and are the coordinates of the pixel points of the facial image, is the grayscale value of the facial image, is the maximum grayscale value of the facial image, and are the fuzzy logic parameters;
[0014] Step S23: Modify the membership value in the hidden layer of the neural network. The formula used is as follows: ;
[0015] In the formula, is the modified membership value;
[0016] Step S24: Find the inverse function of the membership function to generate a new gray level of the facial image. The formula used is as follows: ;
[0017] In the formula, is the new gray value of the facial image;
[0018] Step S25: The convolutional layer of the neural network uses the hyperbolic tangent function as the activation function to extract the features of the facial image, designs fuzzy rules according to the rule extraction set and the rule selection set, and adjusts the weights of the features of the facial image;
[0019] Step S26: Input the training set into the facial expression analysis model for training. The convolutional layer of the neural network outputs the facial expression state. The formula used is as follows: ;
[0020] In the formula, is the activation value of the th facial image at the th position in the th convolutional layer, is the activation function, that is, the hyperbolic tangent function, is the size of the convolutional kernel, is the weight of the convolutional kernel in the th convolutional layer, is the bias term, is the output of the th convolutional layer;
[0021] Step S27: After the training is completed, input the test set into the facial expression analysis model for testing and adjustment, and define the loss function of the facial expression analysis model. The formula used is as follows: ;
[0022] In the formula, is the loss function, is the test set, is the membership matrix composed of membership values, is the set of cluster centers of the facial expression states of the facial images, is the total number of facial images in the test set, is the index of the cluster center, is the index of the facial image, is the penalty coefficient, is the distance from the facial image to the cluster center, is the regularization parameter, is the constraint function;
[0023] Furthermore, in step S3, a speech emotion analysis model is constructed, which specifically includes the following steps:
[0024] Step S31: Collect historical clinical interview audio records, discard the audio segments of doctors in the historical clinical interview audio records, and only retain the audio segments of patients, denoted as the audio analysis set, and mark the audio of patients with depression and the audio of other personnel in the audio analysis set;
[0025] Step S32: Divide the audio in the spectrogram analysis set into audio frames at a fixed interval of 500 ms with a Hamming window of 2.5 s, calculate the discrete Fourier transform of each audio frame, retain the logarithm of the amplitude spectrum of the audio frame, smooth the logarithm of the amplitude spectrum, collect the spectral components on the Mel frequency scale, and perform discrete cosine transform on the spectral components to obtain the Mel frequency cepstral coefficient features of each audio frame;
[0026] Step S33: Normalize the Mel frequency cepstral coefficient features of each audio frame using the Z-core normalization method to obtain low-level standard features;
[0027] Step S34: Establish and initialize a long short-term memory network as the speech emotion analysis model, use softmax as the activation function of the speech emotion analysis model, input the audio analysis set into the long short-term memory network for training, and make the speech emotion analysis model output the audio emotion state;
[0028] Furthermore, in step S5, multi-model fusion screening is performed, which specifically includes the following steps:
[0029] Step S51: Construct a multi-model training set, where the multi-model training set is the artificially marked fragments of patients' medical records. A fragment of a patient's medical record is used as a training instance, and whether the patient has a depressive tendency is used as a label. After dividing the multi-model training set, input it into the facial expression analysis model, the speech emotion analysis model, and the text emotion analysis model for prediction;
[0030] Step S52: Preset the prediction weights of each model, calculate the weighted prediction scores of each model, use the prediction result with the highest weighted prediction score as the current prediction result, preset a reward value. If the current prediction result is the same as the label of the training instance, add the reward value to the prediction weight of the model that outputs the current prediction result, and update the prediction weights of the models;
[0031] If the current prediction result is not the same as the label of the training instance, but there are models with prediction results the same as the label, then use the following formula to add the reward value to the prediction weights of the models with prediction results the same as the label: ;
[0032] In the formula, is the prediction weight, is each training instance, is the digital index of the model that outputs the current prediction result, is the reward value;
[0033] Step S53: If the prediction results of all models are not the same as the label, then calculate the prediction accuracy of each model, and add half of the reward value to the prediction weight of the model with the highest prediction accuracy, using the following formula: ;
[0034] Step S54: Update the prediction weights of each model until all training instances in the multi-model training set are input. Integrate the facial expression analysis model, the voice emotion analysis model, and the text emotion analysis model into an early depression screening model, and output the depression screening result, where the depression screening result is the prediction result with the highest weighted prediction score.
[0035] The beneficial effects achieved by the present invention using the above solution are as follows:
[0036] (1) Aiming at the technical problems of traditional depression screening methods mainly through communication between doctors and patients, where the professional level, experience, and attitude of doctors will have a greater impact on the screening results, and patients may be reluctant to frankly tell their symptoms and feelings due to concerns about privacy leakage, affecting the screening results. This solution constructs an early depression screening model by analyzing a large amount of clinical data, integrates various factors, can more objectively evaluate the state of patients, and improve the efficiency of depression screening;
[0037] (2)When screening using machine learning through facial expressions, problems such as image noise, uneven illumination, and edge blurring interference are likely to occur. In this solution, fuzzy logic is combined with a convolutional neural network to optimize the noise suppression effect by correcting the membership values, improving the model's ability to capture weak expression changes and enabling a more accurate assessment of the patient's emotional state;
[0038] (3)Regarding the technical problem that when traditional machine learning is used for screening, it usually relies on single-modal data and makes predictions from one-sided information, resulting in low screening accuracy. This solution constructs a multi-modal fusion analysis framework, designs a facial expression analysis model, a speech emotion analysis model, and a text emotion analysis model respectively, and designs a dynamic weighting function to adjust the weights of each model, realizing the collaborative analysis of multi-dimensional data, improving the comprehensive judgment ability of early depression screening, and solving the problems of misjudgment and missed judgment caused by the limitations of single-modal data. Brief Description of the Drawings
[0039] Figure 1 It is a flowchart of the steps of a method for early screening of depression based on machine learning provided by the present invention.
[0040] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. Detailed Embodiments
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] Embodiment 1: Refer to Figure 1 , this embodiment provides a method for early screening of depression based on machine learning, which specifically includes the following steps:
[0043] Step S1: Clinical data collection, collect video data during the patient's visit, cut the video data into a segment every 3 seconds, edit the frame images, audio, and transcribed text of each segment, and input the edited frame images, audio, and transcribed text into the corresponding analysis models respectively;
[0044] Step S2: Construct a facial expression analysis model, combine fuzzy logic with a convolutional neural network to construct a facial expression analysis model, judge the patient's depression tendency through the frame images, and output the facial emotional state;
[0045] Step S3: Construct a speech emotion analysis model. By extracting the Mel-frequency cepstral coefficients features of the audio, use a long short-term memory network to construct a speech emotion analysis model, and judge the depression tendency of the patient through the audio, and output the audio emotion state;
[0046] Step S4: Construct a text emotion analysis model, which is constructed by a pre-trained T5 model and fine-tuned using transfer learning. Judge the depression tendency of the patient through the transcribed text, and output the text emotion state;
[0047] Step S5: Multi-model fusion screening. Use a weighted function to unify the facial emotion state, audio emotion state and text emotion state, and update the prediction weights of the three models by adding reward values. After fusion, an early depression screening model is obtained. Input the frame image, audio and transcribed text into the early depression screening model to obtain the depression screening result.
[0048] Through the above solution, aiming at the technical problem that the traditional depression screening method mainly communicates between doctors and patients, the professional level, experience and attitude of doctors will have a greater impact on the screening results, and patients may be reluctant to frankly tell their symptoms and feelings due to concerns about privacy leakage, affecting the screening results. This solution constructs an early depression screening model by analyzing a large amount of clinical data, integrates multiple factors, can objectively evaluate the state of patients, and improve the efficiency of depression screening.
[0049] Example 2: Refer to Figure 1 , based on the above example, in step S2, construct a facial expression analysis model, which specifically includes the following steps:
[0050] Step S21: Collect a facial image set, which includes facial images of depression patients and non-patients and their corresponding status labels. After preprocessing the facial image set, divide the facial image set into a training set, a rule extraction set, a rule selection set and a test set according to the ratio of 3:2:2:1;
[0051] Step S22: Establish and initialize a convolutional neural network, combine fuzzy logic with the convolutional neural network to construct a facial expression analysis model, design a fuzzy logic layer to process the facial image set, calculate the derivative of the fuzzy direction, use the directional gradient to distinguish the noise and edges of the facial image, calculate the membership degree value of the facial image, and define the membership function of fuzzy logic. The formula used is as follows: ;
[0052] In the formula, is the membership degree value, and are the coordinates of the pixel points of the facial image, is the gray value of the facial image, is the maximum gray value of the facial image, and are fuzzy logic parameters;
[0053] Step S23: Modify the membership value in the hidden layer of the neural network, and the formula used is as follows: ;
[0054] In the formula, is the modified membership value;
[0055] Step S24: Find the inverse function of the membership function to generate a new gray level of the facial image, and the formula used is as follows: ;
[0056] In the formula, is the new gray value of the facial image;
[0057] Step S25: The convolutional layer of the neural network uses the hyperbolic tangent function as the activation function to extract the features of the facial image, designs fuzzy rules according to the rule extraction set and the rule selection set, and adjusts the weights of the features of the facial image;
[0058] Step S26: Input the training set into the facial expression analysis model for training, and the convolutional layer of the neural network outputs the facial expression state, and the formula used is as follows: ;
[0059] In the formula, is the activation value of the th facial image at position in the th convolutional layer, is the activation function, that is, the hyperbolic tangent function, is the size of the convolutional kernel, is the weight of the convolutional kernel in the th convolutional layer, is the bias term, is the output of the th convolutional layer;
[0060] Step S27: After the training is completed, input the test set into the facial expression analysis model for test adjustment, and define the loss function of the facial expression analysis model, and the formula used is as follows: ;
[0061] In the formula, is the loss function, is the test set, is the membership matrix composed of membership values, is the set of cluster centers of the facial expression states of facial images, is the total number of facial images in the test set, is the index of the cluster center, is the index of the facial image, is the penalty coefficient, is the distance from the facial image to the cluster center, is the regularization parameter, is the constraint function.
[0062] Through the above solution, when screening through facial expressions using machine learning, problems such as image noise, uneven illumination, and edge blurring interference are likely to occur. This solution combines fuzzy logic with a convolutional neural network, optimizes the noise suppression effect by correcting the membership value, improves the model's ability to capture weak expression changes, and can more accurately evaluate the patient's emotional state.
[0063] Example 3: Based on the above example, in the specific implementation of step S2, it is necessary to preprocess the facial images, including face detection and cropping, grayscale conversion, and histogram equalization;
[0064] Use the MediaPipe BlazeFace model to locate the face and crop the facial area, use cv2.cvtColor to convert the facial image to a grayscale image, and use cv2.equalizeHist to enhance the contrast of the facial image to achieve histogram equalization;
[0065] Regarding the problem that the hyperbolic tangent function is used as the activation function for the convolutional layer of the neural network, in this solution, the neuron parameters are set to 1.7159 and 2 / 3 respectively.
[0066] Example 4: Refer to Figure 1 , this example is based on the above example, and in step S3, a speech emotion analysis model is constructed, which specifically includes the following steps:
[0067] Step S31: Collect historical clinical interview audio records, discard the audio segments of doctors in the historical clinical interview audio records, and only retain the audio segments of patients, which are recorded as the audio analysis set. Mark the audio of patients with depression and the audio of other people in the audio analysis set;
[0068] Step S32: Divide the audio in the spectrogram analysis set into audio frames at a fixed interval of 500 ms with a 2.5 s Hamming window, calculate the discrete Fourier transform of each audio frame, retain the logarithm of the amplitude spectrum of the audio frame, smooth the logarithm of the amplitude spectrum, collect the spectral components on the Mel frequency scale, and perform a discrete cosine transform on the spectral components to obtain the Mel frequency cepstral coefficient features of each audio frame;
[0069] Step S33: Normalize the Mel-frequency cepstral coefficient features of each audio frame using the Z-core normalization method to obtain low-level standard features;
[0070] Step S34: Establish and initialize a long short-term memory network as a speech emotion analysis model, use softmax as the activation function of the speech emotion analysis model, input the audio analysis set into the long short-term memory network for training, and enable the speech emotion analysis model to output the audio emotion state.
[0071] Example 5: Refer to Figure 1 , this example is based on the above Example 4. To avoid overfitting of the speech emotion analysis model, in actual application, it is necessary to perform data augmentation on the audio of depression patients and the audio of other people in the audio analysis set before marking. The techniques used include but are not limited to noise injection, pitch enhancement, shift enhancement, and speed enhancement.
[0072] Example 6: Refer to Figure 1 , this example is based on the above example, Step S5, multi-model fusion screening, which specifically includes the following steps:
[0073] Step S51: Construct a multi-model training set. The multi-model training set is artificially marked fragments of patient medical records. A fragment of a patient medical record is used as a training instance, and whether the patient has a depressive tendency is used as a label. After splitting the multi-model training set, input it into the facial expression analysis model, speech emotion analysis model, and text emotion analysis model for prediction;
[0074] Step S52: Preset the prediction weights of each model, calculate the weighted prediction scores of each model, and use the prediction result with the highest weighted prediction score as the current prediction result. Preset a reward value. If the current prediction result is the same as the label of the training instance, add the reward value to the prediction weight of the model that outputs the current prediction result, and update the prediction weight of the model;
[0075] If the current prediction result is not the same as the label of the training instance, but there is a model with a prediction result the same as the label, then use the formula below to add the reward value to the prediction weight of the model with a prediction result the same as the label: ;
[0076] In the formula, is the prediction weight, is each training instance, is the digital index of the model that outputs the current prediction result, is the reward value;
[0077] Step S53: If the prediction results of all models are different from the labels, calculate the prediction accuracy of each model, and add half of the reward value to the prediction weight of the model with the highest prediction accuracy. The formula used is as follows: ;
[0078] Step S54: Update the prediction weights of each model until all training instances in the multi-model training set are input. Then, fuse the facial expression analysis model, the voice emotion analysis model, and the text emotion analysis model into an early depression screening model, and output the depression screening result, which is the prediction result with the highest weighted prediction score.
[0079] Through the above solution, aiming at the technical problem that when traditional machine learning is used for screening, it usually relies on single-modal data and makes predictions from one-sided information, resulting in low screening accuracy. This solution constructs a multi-modal fusion analysis framework, designs a facial expression analysis model, a voice emotion analysis model, and a text emotion analysis model respectively, and designs a dynamic weighting function to adjust the weights of each model, realizing the collaborative analysis of multi-dimensional data, improving the comprehensive judgment ability of early depression screening, and solving the problems of misjudgment and missed judgment caused by the limitations of single-modal data.
[0080] Embodiment 7. It should be noted that this solution is generally applicable to all populations, especially perinatal women, for the following reasons: Perinatal women are prone to transient expression blurring or mood swings due to hormonal changes. In this solution, the fuzzy logic layer (Steps S22 - S24) effectively suppresses noise interference such as uneven illumination and facial swelling by correcting the membership values, accurately identifies weak expression changes, and avoids missed judgments. At the same time, perinatal depression may manifest as contradictory emotions, and this solution integrates three modalities of data, namely facial, voice, and text, and adjusts the model weights through a reward mechanism, avoiding "one-size-fits-all" misjudgments.
[0081] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0082] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0083] The above describes the present invention and its embodiments. Such a description is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In summary, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, design similar structural methods and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. A method for early screening of depression based on machine learning, characterized in that: The specific steps include: Step S1: Clinical data collection: collecting video data of patients during their medical consultations, cutting the video data into segments every 3 seconds, and editing the frame images, audio, and transcribed text of each segment; Step S2: construct a facial expression analysis model, combining fuzzy logic with convolutional neural network to construct a facial expression analysis model, judge the patient's depression tendency through frame images, and output the facial emotion state; Step S3: constructing a speech emotion analysis model by extracting the Mel-frequency cepstral coefficient features of the audio and using a long short-term memory network to construct a speech emotion analysis model, judging the patient's depression tendency through the audio and outputting the audio emotion state; Step S4: Construct a text sentiment analysis model, which is constructed by the pre-trained T5 model and fine-tuned using transfer learning. The model determines the patient's depression tendency by transcribing the text and outputs the text's emotional state. Step S5: Multi-model fusion screening, using a weighted function to unify facial emotional state, audio emotional state and text emotional state, and updating the prediction weights of the three models by adding reward values. After fusion, an early depression screening model is obtained, and the frame image, audio and transcribed text are input into the early depression screening model to obtain the depression screening results.
2. The method for early screening of depression based on machine learning according to claim 1, characterized in that: Step S2, constructing a facial expression analysis model, specifically comprises the following steps: Step S21: collecting a facial image set, pre-processing the facial image set, and dividing the facial image set into a training set, a rule extraction set, a rule selection set, and a test set; Step S22: Establish and initialize a convolutional neural network, combine fuzzy logic with the convolutional neural network to construct a facial expression analysis model, design a fuzzy logic layer to process the facial image set, calculate the derivative of the fuzzy direction, use the directional gradient to distinguish the noise and edge of the facial image, calculate the membership value of the facial image, and define the membership function of the fuzzy logic. The formula used is as follows: ; In the formula, is the membership value, and is the coordinate of the pixel point of the facial image, is the grayscale value of the facial image, is the maximum grayscale value of the facial image, and is the fuzzy logic parameter; Step S23: correcting the membership value in the hidden layer of the neural network; Step S24: finding the inverse function of the membership function to generate a new grayscale of the facial image; Step S25: The convolution layer of the neural network uses the hyperbolic tangent function as an activation function to extract features of the facial image, designs fuzzy rules based on the rule extraction set and the rule selection set, and adjusts the weights of the features of the facial image; Step S26: inputting the training set into the facial expression analysis model for training, and the convolution layer of the neural network outputs the facial expression state; Step S27: After the training is completed, the test set is input into the facial expression analysis model for test adjustment.
3. The method for early screening of depression based on machine learning according to claim 1, characterized in that: Step S3, constructing a speech emotion analysis model, specifically includes the following steps: Step S31: Collect historical clinical interview audio records, retain only the patient's audio clips, record them as an audio analysis set, and mark the audio of the depression patient and the audio of other people in the audio analysis set; Step S32: Segment the audio in the spectral analysis set into audio frames, calculate the discrete Fourier transform of each audio frame, retain the logarithm of the amplitude spectrum of the audio frame, and smooth the logarithm of the amplitude spectrum, collect spectral components on the Mel frequency scale, perform discrete cosine transform on the spectral components, and obtain the Mel frequency cepstrum coefficient features of each audio frame; Step S33: normalizing the Mel-frequency cepstral coefficient features of each audio frame using a Z-core normalization method to obtain a low-level standard feature; Step S34: Establish and initialize a long short-term memory network as a speech emotion analysis model, use softmax as the activation function of the speech emotion analysis model, input the audio analysis set into the long short-term memory network for training, and make the speech emotion analysis model output the audio emotion state.
4. The method for early screening of depression based on machine learning according to claim 1, characterized in that: The step S5, multi-model fusion screening, specifically includes the following steps: Step S51: construct a multi-model training set, wherein the multi-model training set is a manually labeled patient medical record segment, a patient medical record segment is used as a training instance, and whether the patient has a tendency to depression is used as a label. The multi-model training set is segmented and input into a facial expression analysis model, a voice emotion analysis model, and a text emotion analysis model for prediction; Step S52: Preset the prediction weight of each model, calculate the weighted prediction score of each model, take the prediction result with the highest weighted prediction score as the current prediction result, preset the reward value, and if the current prediction result is the same as the label of the training instance, add the reward value to the prediction weight of the model that outputs the current prediction result, and update the prediction weight of the model; If the current prediction result is different from the label of the training instance, but there is a model whose prediction result is the same as the label, the reward value is added using the prediction weight of the model whose prediction result is the same as the label; Step S53: If the prediction results of all models are different from the labels, the prediction accuracy of each model is calculated, and half of the reward value is added to the prediction weight of the model with the highest prediction accuracy; Step S54: Update the prediction weight of each model, merge the facial expression analysis model, the voice emotion analysis model and the text emotion analysis model into an early depression screening model, and output the depression screening results.