Multi-modal fusion-based depression identification method and system, and storage medium
Through multimodal fusion technology, video and text features are extracted and fusion is solved, and the problem of only single-modal data is considered in the prior art, which improves the accuracy of depression recognition.
Patent Information
- Application Number
- CN202510660507.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art only considers single-modal video data in depression recognition, and fails to make full use of multimodal data for more overall depression recognition.
The multimodal fusion method is adopted, by obtaining video data and segmenting it into short-term video units, the pre-trained model is used to extract the video feature vector and text feature vector, calculate the correlation between the two and perform feature fusion, and finally use a bidirectional long and short-term memory network for classification.
Through multimodal fusion, the model can learn depressive characteristics more carefully, improve the accuracy of depression recognition, and enhance the attention to video and text correlation characteristics.
Smart Images

Figure CN120182899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and depression recognition technology, and particularly relates to a depression recognition method, system and storage medium based on multimodal fusion. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technologies such as computer vision and deep learning, more and more researchers have begun to focus on automatic depression intensity recognition technology. The development of these technologies, especially the rise of large language models, has made it possible to identify more refined and accurate changes in depressive emotions by analyzing video and text data. These subtle feature changes may be closely related to depressive symptoms. By training a model to recognize these features, it can assist mental health professionals in making more accurate diagnoses and interventions, thereby improving the treatment effect.
[0003] For example, the Chinese patent application with the publication number CN118865463A discloses a depression detection system based on facial action units, including model, model, model and a logistic regression neural network. By designing a new loss function algorithm, it solves the problem of image recognition accuracy restricted by geometric constraints and data imbalance in deep learning, expands the range of face image recognition, and divides model and the results output by the model are encoded into model. Each face image is regarded as a group of independent regions, the face is divided into regions and key points in the face image are obtained, and it can identify the neural movement units in the local regions of the face at a finer granularity, thereby achieving higher accuracy and improving the accuracy of depression detection. This method achieves more fine-grained depression recognition and detection by dividing the face region, but only considers single-modal video data and does not consider data from other dimensions for more comprehensive depression recognition.
[0004] For another example, the Chinese invention patent application with the publication number CN118969202A discloses a depression detection and analysis method and device based on multi-modal data and large models. The method includes obtaining the first text data, audio data, and video data of a patient; performing text extraction based on the audio data to obtain second text data; inputting the first text data and the second text data into a large language model based on a preset prompt to obtain a first depression evaluation score; performing feature extraction based on the audio data and the video data to obtain audio-visual features; performing regression analysis based on the audio-visual features to obtain a second depression evaluation score; and performing data fusion based on the first depression evaluation score and the second depression evaluation score to obtain a final depression index. This method obtains a depression evaluation score through a large language model and then performs data fusion to obtain the final depression index, but it does not perform feature fusion at the feature level, which may cause the model to lose attention to video and text features. Summary of the Invention
[0005] The purpose of the present invention is to provide a depression recognition method, system, and storage medium based on multi-modal fusion to solve one or more technical problems existing in the prior art and at least provide a beneficial choice or create conditions.
[0006] The present invention adopts the following technical solutions to achieve the above-mentioned invention purpose: The present invention provides a depression recognition method based on multi-modal fusion, including: Obtaining video data; Performing short-time sequence time window partitioning on the obtained video data to divide the video data into multiple short-time sequence video units; Using the image encoder of the image-text contrast pre-training model to process each short-time sequence video unit and obtaining video feature vectors; Generating a description through a large language model, using the text encoder in the image-text contrast pre-training model to extract semantic information, and obtaining text feature vectors; Calculating the correlation between the video feature vectors and the text feature vectors, and then performing fusion of the video feature vectors and the text feature vectors to obtain feature fusion vectors; Using a bidirectional long short-term memory network to classify the feature fusion vectors and outputting recognition results.
[0007] Further, the method for performing short-time sequence time window partitioning on the obtained video data to divide the video data into multiple short-time sequence video units includes: Using a face alignment network to obtain multiple facial key points and extracting facial key points from each frame of video; Taking the center point of the interpupillary distance as a reference benchmark point and calculating the offset of the key points in the eyebrow, eye, nose, and mouth regions relative to the reference benchmark point; Take the average of all frame offset values in the first five seconds of the video as the initial offset value of the video segment; Calculate the relative offset value between each subsequent frame and the initial offset value to determine whether there are significant changes in the face; Identify the moments when the offset value peaks through a peak detection algorithm, and mark its start and end positions and duration; Use the K-means clustering algorithm to process the duration data of these time periods to determine the unit length of video segmentation with the cluster center; Adopt a centralized point search algorithm to determine the centralized point of facial movement, and evenly divide the input frame sequence into two subsequences; Calculate the sum of relative offset values of each subsequence, retain the subsequence with the larger sum, and remove the subsequence with the smaller sum until the centralized point where the peak frame is located is located; Use the landmark point data with a unit length before and after the centralized point as a short-time sequence video unit.
[0008] Furthermore, the method of using the image encoder of the text-image contrast pre-training model to process each short-time sequence video unit and obtain a video feature vector includes: Sample frames from the short-time sequence video unit, and the size of each frame is to form an input , where represents the number of frames, 3 represents the three color channels of red, yellow, and blue, and represent the height and width of the frame respectively; for each frame , use the shared image encoder to extract the feature vector , where , represents the length of the feature vector, and the formula is as follows: ; feature vectors are then fed into the temporal model to learn the temporal features of the facial depressive expression and obtain the finally output video feature vector, and the formula is as follows: ; Among them, represents a special learnable vector, representing the class token in the visual self-attention model, represents the learnable position embedding, represents the video feature vector, .
[0009] Further, the method of generating a description through a large language model, extracting semantic information using the text encoder in the image-text contrast pre-training model, and obtaining a text feature vector includes: Using the large language model to convert video content into an action description of facial depressive expressions; Using a description related to facial behavior instead of the class name as the input to the text encoder, and the description of each class is contextualized by a learnable prompt. The formula is as follows: ; Among them, represents a hyperparameter specifying the number of context tokens, , represents the number of depression classification categories, takes 4, divided into non-depressed, mildly depressed, moderately depressed, and severely depressed, each , , is a tokenizer that splits the facial expression description text generated by the large language model into basic units for model processing ; By passing the prompt to the text encoder , the classification weight vectors are obtained as the final output text feature vector. The formula is as follows: ; Among them, represents the text feature vector, .
[0010] Further, the method of calculating the correlation between the video feature vector and the text feature vector, and then fusing the video feature vector and the text feature vector to obtain a feature fusion vector includes: Calculating the correlation coefficient between the video feature vector and the text feature vector , the formula is as follows: ; Among them, the magnitude of represents the correlation between the video feature vector and the text feature vector, represents the activation function; Calculating the video and text correlation feature , the formula is as follows: ; Then, the video feature vector and the text feature vector are respectively added to the video and text correlation features to obtain an enhanced video feature vector and an enhanced text feature vector. The formula is as follows: ; ; Among them, represents the enhanced video feature vector, represents the enhanced text feature vector; Fuse the enhanced video feature vector and the enhanced text feature vector to obtain a feature fusion vector. The formula is as follows; ; Among them, represents the feature fusion vector.
[0011] The present invention provides a depression recognition system based on multimodal fusion, including: An acquisition unit for acquiring video data; A segmentation unit for dividing the acquired video data into short-time sequence time windows and segmenting the video data into multiple short-time sequence video units; A video feature vector generation unit for processing each short-time sequence video unit by using the image encoder of the text-image contrast pre-training model and obtaining a video feature vector; A text feature vector generation unit for generating a description through a large language model, extracting semantic information by using the text encoder in the text-image contrast pre-training model, and obtaining a text feature vector; A calculation fusion unit for calculating the correlation between the video feature vector and the text feature vector, and then fusing the video feature vector and the text feature vector to obtain a feature fusion vector; An output unit for classifying the feature fusion vector by using a bidirectional long short-term memory network and outputting an identification result.
[0012] The present invention provides a depression recognition system based on multimodal fusion, including a memory and a processor; The memory is used for storing instructions; The processor is used for operating according to the instructions to execute the steps of the above method.
[0013] The present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0014] The beneficial effects of the present invention are as follows: By extracting video features and simultaneously using a large language model to analyze and summarize the subtle changes in depressive emotions in the video to obtain text features, the model can learn depressive features more carefully. Calculate the correlation between the video features and the text features of the two, and then perform deep fusion at the feature level, effectively enhancing the model's attention to the associated features of the video and the text, and improving the accuracy of depression recognition. Brief Description of the Drawings
[0015] Figure 1 It is a flowchart of the first depression recognition method based on multimodal fusion provided according to an embodiment of the present invention; Figure 2 It is a flowchart of the second depression recognition method based on multimodal fusion provided according to an embodiment of the present invention; Figure 3 It is a schematic diagram of facial key points in a depression recognition method based on multimodal fusion provided according to an embodiment of the present invention; Figure 4 It is a schematic diagram of a feature fusion structure in a depression recognition method based on multimodal fusion provided according to an embodiment of the present invention. Detailed Embodiment
[0016] The present invention will be further described below in conjunction with specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0017] As Figures 1 to 2 , a depression recognition method based on multimodal fusion provided by the present invention includes the following steps: S100, Obtain video data; S200, Divide the obtained video data into short-time sequence time windows, and split the video data into multiple short-time sequence video units; S300, Use the image encoder of the text-image contrast pre-trained model to process each short-time sequence video unit, and obtain video feature vectors; S400, Generate descriptions through a large language model, extract semantic information using the text encoder in the text-image contrast pre-trained model, and obtain text feature vectors; S500, Calculate the correlation between the video feature vectors and the text feature vectors, and then perform the fusion of the video feature vectors and the text feature vectors to obtain feature fusion vectors; S600, Use a bidirectional long short-term memory network to classify the feature fusion vectors and output the recognition results.
[0018] Among them, in the S200, it specifically includes: The data preprocessing module uses a face alignment network to obtain 68 detailed facial key points, extracts facial key points from each frame of the video. As shown in Figure 3, with the center point of the interpupillary distance as the reference benchmark point, calculate the offset of the key points in the eyebrow, eye, nose, and mouth regions relative to the benchmark point, and take the average of the offset values of all frames in the first 5 seconds of the video as the initial offset value of this video segment. Subsequently, calculate the relative offset value between each subsequent frame and the initial offset value to determine whether there are significant changes in the face. Use the peak detection algorithm to identify the moments when the offset value peaks, and mark its start and end positions and duration. Use the K-means clustering algorithm to process the duration data of these time periods to determine the unit length of video segmentation with the cluster center. Preferably, to determine the concentration point of facial movement, a concentration point search algorithm is adopted. The algorithm process is as shown in Algorithm 1, and its working process is as follows: Divide the input frame sequence into two equal parts on average (for example, divide 40 frames into subsequences of frames 1-20 and 21-40). Calculate the sum of the relative offset values of each subsequence, retain the subsequence with the larger sum, and discard the subsequence with the smaller sum. Repeat the above segmentation and comparison process until the concentration point where the peak frame is located is located. Finally, use the landmark point data with a unit length before and after the concentration point as a short-time sequence video unit to construct a dataset for model training.
[0019] Algorithm 1 Concentration Point Search Algorithm Input Frame Sequence Set
[0020] Initialize Initialize Frame Output Set ,
[0021] REPEAT 1: ; / / Divide the input frames into two equal parts 2: ; / / Retain the segment with relatively larger offset ; / / When the length of the set is 1, end the search, and at this time, locate the peak frame with the largest offset value Output: Output the frame set where the concentration point is located .
[0022] Among them, in the S300, it specifically includes: The video branch uses the image encoder of the text-image contrast pre-training model to process each short-time sequence video unit. Specifically, given a short-time sequence video unit, the model samples frames, and the size of each frame is , forming an input , where represents the number of frames, 3 represents the three color channels of red, yellow, and blue, and represent the height and width of the frame respectively; For each frame , first use the shared image encoder to extract the feature vector , where , represents the length of the feature vector, and the formula for this process is as follows: ; After that, feature vectors are then fed into the temporal model to learn the temporal features of facial depressive expressions and obtain the video feature vector output by the final video branch. The formula for this process is as follows: ; Among them, represents a special learnable vector, representing the class token in the visual self-attention model, represents the learnable position embedding, represents the video feature vector, , add the learnable vector to a learnable position embedding . This position embedding provides the relative position information of each frame in the video time series for the model. In this way, the model can not only learn the visual features of each frame but also understand the relationship features of these frames in the time series.
[0023] Among them, in the S400, it specifically includes: Use the large language model to convert the video content into an action description of facial depressive expressions; Through the model of the text branch, use the description related to facial behavior to replace the category name as the input of the text encoder. The description of each category is used as the context by the learnable prompt. The formula is as follows:
[0024] Among them, represents a hyperparameter, specifying the number of context tokens, , represents the number of depression classification categories. In this method takes 4, divided into non-depressed, mildly depressed, moderately depressed, and severely depressed. Each , is a vector with the same dimension as the word embedding, is the tokenizer, which tokenizes the facial expression description text generated by the large language model Split into the basic units for model processing Here, category-specific context is adopted, and the context vector is independent of each description. By passing the prompt to the text encoder , we get classification weight vectors as the text feature vectors of the final output. The formula for this process is as follows: ; where represents the text feature vector, .
[0025] Preferably, for the selection of large language models, ChatGPT, MOSS (MegaScale Open Science) from Fudan University, LLaMA (Large Language Model Meta AI) from Meta AI, etc. can be used. GPT-3.5 is developed by OpenAI, with 175 billion parameters and adopts the Transformer architecture; LLaMA has 65 billion parameters. It has made innovations in the model structure, such as using the RMSNorm normalization function, SwiGLU activation function, and Rotary Position Encoding (RoPE) to improve the model performance and stability. The open-source nature of LLaMA makes it easy for developers to use and improve; MOSS is a conversational large language model developed by the Natural Language Processing Laboratory of Fudan University, with 16 billion parameters, supporting Chinese-English bilingual conversations. MOSS's language model has been pre-trained on approximately 700 billion Chinese, English, and code words, which provides it with strong language understanding capabilities, and resources such as the code, data, and model parameters of the MOSS model have been opened.
[0026] Among them, in the S500, it specifically includes: Through the feature fusion module, the video feature vector obtained from the video branch and the text feature vector obtained from the text branch are fused in features. First, the correlation between the video feature vector and the text feature vector is calculated, and then fusion is performed. The fusion structure is as Figure 4 shown. First, the correlation coefficient between the video feature vector and the text feature vector is calculated , and the formula is as follows:
[0027] where The magnitude of represents the correlation between the video feature vector and the text feature vector, represents the activation function; After that, calculate the video and text relevance features , and the formula is as follows:
[0028] Then, add the video feature vector and the text feature vector to the video and text relevance features respectively to obtain the enhanced video feature vector and the enhanced text feature vector. The formula is as follows:
[0029]
[0030] Among them, represents the enhanced video feature vector, represents the enhanced text feature vector; Finally, fuse the enhanced video feature vector and the enhanced text feature vector to obtain the output result of the feature fusion module, that is, the feature fusion vector. The formula is as follows;
[0031] Among them, represents the feature fusion vector.
[0032] Among them, in the S600, it specifically includes: Use a bidirectional long short-term memory network in the classifier module to classify the input feature fusion vector and output the classification result.
[0033] In the model training stage, since the text encoder has learned general text features, these features are beneficial for text information such as depression emotion description. By fixing the text encoder, the model can stably utilize these pre-trained features, and at the same time let the image encoder learn how to match visual information with text features, thereby improving the accuracy of depression recognition. Therefore, the parameters of the text encoder remain unchanged, while the parameters of the image encoder will be fine-tuned according to the video data.
[0034] Preferably, in the module structure and training process of the deep learning network, a loss function is constructed. The loss function used is the mean squared error loss (Mean Squared Error Loss, MSELoss), and its formula is as follows:
[0035] Among them is the true value of the th sample, is the predicted value of the model for this sample, is the number of samples.
[0036] During the module structure and training process of the deep learning network, the network is trained for a total of 800 iterations. The initial learning rate is 0.0001, and the learning rate is decreased by a factor of 10 every 20 iterations. The model is saved at the end of each iteration, and the model of the last iteration is selected as the network.
[0037] A depression recognition system based on multimodal fusion provided by the present invention includes: An acquisition unit for acquiring video data; A segmentation unit for dividing the acquired video data into short-time sequence time windows and segmenting the video data into multiple short-time sequence video units; A video feature vector generation unit for using the image encoder of the text-image contrast pre-training model to process each short-time sequence video unit and obtaining a video feature vector; A text feature vector generation unit for generating a description through a large language model, extracting semantic information using the text encoder in the text-image contrast pre-training model, and obtaining a text feature vector; A calculation and fusion unit for calculating the correlation between the video feature vector and the text feature vector, and then fusing the video feature vector and the text feature vector to obtain a feature fusion vector; An output unit for classifying the feature fusion vector using a bidirectional long short-term memory network and outputting a recognition result.
[0038] A depression recognition system based on multimodal fusion provided by the present invention may also be: including a memory and a processor; the memory is used for storing instructions; The processor is used to operate according to the instructions to execute the steps of the aforementioned pedestrian re-identification method based on motion information.
[0039] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the aforementioned pedestrian re-identification method based on motion information.
[0040] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0041] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0042] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0043] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0044] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for identifying depression based on multimodal fusion, characterized in that, including: Obtain video data; Perform short-time sequence time window division on the obtained video data, and split the video data into multiple short-time sequence video units; Use the image encoder of the image-text contrast pre-trained model to process each short-time sequence video unit, and obtain video feature vectors; Generate a description through a large language model, use the text encoder in the image-text contrast pre-trained model to extract semantic information, and obtain text feature vectors; Calculate the correlation between the video feature vectors and the text feature vectors, and then perform the fusion of the video feature vectors and the text feature vectors to obtain feature fusion vectors; Use a bidirectional long short-term memory network to classify the feature fusion vectors and output the recognition result.
2. The method for identifying depression based on multimodal fusion according to claim 1, characterized in that, The method of performing short-time sequence time window division on the obtained video data and splitting the video data into multiple short-time sequence video units includes: Use a face alignment network to obtain multiple facial key points, and extract facial key points from each frame of the video; Taking the center point of the interpupillary distance as the reference benchmark point, calculate the offset of the key points in the eyebrow, eye, nose, and mouth regions relative to the reference benchmark point; Take the average of all frame offset values in the first five seconds of the video as the initial offset value of the video segment; Calculate the relative offset value between each subsequent frame and the initial offset value to determine whether there is a significant change in the face; Use a peak detection algorithm to identify the moment when the offset value peak appears, and mark its start and end positions and duration; Use the K-means clustering algorithm to process the duration data of these time periods, and determine the unit length of video segmentation with the cluster center; Use a concentration point search algorithm to determine the concentration point of facial movement, and evenly divide the input frame sequence into two subsequences; Calculate the sum of the relative offset values of each subsequence, retain the subsequence with the larger sum, and remove the subsequence with the smaller sum until the concentration point where the peak frame is located is located; Use the landmark data of one unit length before and after the concentration point as a short-time sequence video unit.
3. The method for identifying depression based on multimodal fusion according to claim 1, characterized in that, The method of using the image encoder of the image-text contrast pre-trained model to process each short-time sequence video unit and obtain video feature vectors includes: Sampling from short-time sequence video units frames, each frame having a size of to form an input , where represents the number of frames, 3 represents the three color channels of red, yellow, and blue, and represent the height and width of the frame respectively; for each frame , use a shared image encoder to extract the feature vector , where , represents the length of the feature vector, and the formula is as follows: ; The eigenvector is then fed into the time series model Learn the time series features of facial depressive expressions to obtain the finally output video eigenvector. The formula is as follows: ; Among them, represents a special learnable vector, representing the class token in the visual self-attention model, represents the learnable position embedding, represents the video feature vector, .
4. The method for identifying depression based on multimodal fusion according to claim 1, characterized in that, The method of generating a description through a large language model, using the text encoder in the image-text contrast pre-trained model to extract semantic information, and obtaining text feature vectors includes: Use a large language model to convert the video content into an action description of facial depressive expressions; Use the description related to facial behavior to replace the category name as the input of the text encoder, and the description of each category is contextualized by a learnable prompt, and the formula is as follows: ; Among them, represents a hyperparameter that specifies the number of context tokens, , represents the number of depression classification categories, taking 4, divided into non - depression, mild depression, moderate depression, and severe depression each , , is a tokenizer that splits the facial expression description text generated by the large - language model into basic units for model processing ; By passing the prompt to the text encoder , we obtain a classification weight vector as the final output text feature vector, and the formula is as follows: ; Among them, represents the text feature vector, .
5. The method for identifying depression based on multimodal fusion according to claim 1, characterized in that, The method of calculating the correlation between the video feature vectors and the text feature vectors, and then performing the fusion of the video feature vectors and the text feature vectors to obtain feature fusion vectors includes: Calculate the correlation coefficient between the video feature vector and the text feature vector , and the formula is as follows: ; Among them, The magnitude represents the correlation between the video feature vector and the text feature vector, represents the activation function; Calculate the video and text relevance features , and the formula is as follows: ; Then add the video feature vectors and the text feature vectors to the video and text correlation features respectively to obtain enhanced video feature vectors and enhanced text feature vectors, and the formula is as follows: ; ; Among them, represents the enhanced video feature vector, represents the enhanced text feature vector; Fuse the enhanced video feature vectors and the enhanced text feature vectors to obtain feature fusion vectors, and the formula is as follows; ; Among them, represents the feature fusion vector.
6. A depression recognition system based on multimodal fusion, characterized in that, including: An obtaining unit for obtaining video data; A segmentation unit for dividing the acquired video data into short-time sequence time windows and segmenting the video data into multiple short-time sequence video units; A video feature vector generation unit for processing each short-time sequence video unit by using an image encoder of a text-image contrast pre-trained model and obtaining a video feature vector; A text feature vector generation unit for generating a description through a large language model, extracting semantic information by using a text encoder in the text-image contrast pre-trained model, and obtaining a text feature vector; A calculation and fusion unit for calculating the correlation between the video feature vector and the text feature vector, and then fusing the video feature vector and the text feature vector to obtain a feature fusion vector; An output unit for classifying the feature fusion vector by using a bidirectional long short-term memory network and outputting an identification result.
7. A depression recognition system based on multimodal fusion, characterized in that, It includes a memory and a processor; The memory is used for storing instructions; The processor is used for operating according to the instructions to execute the steps of the method according to any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Facial movement unit-based depression detection system
CN118865463A
Method and equipment for detecting and analyzing depression based on multi-modal data and large model
CN118969202A