Face image depression recognition system based on instance-level relation distillation

Through the face image depression recognition system with instance-level relationship distillation, the problems of insufficient generalization ability and difficulty in migration and adaptation of depression detection models in the prior art are solved, and the model is highly generalized and robust in a multi-task environment.

CN120032415AInactive Publication Date: 2025-05-23COMMUNICATION UNIVERSITY OF CHINA
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510502280.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Inadequate generalization ability of depression detection models and difficulty in model migration and adaptation in the prior art lead to challenges in identifying the symptoms of depression in different populations.

Method used

A face image depression recognition system for instance-level relationship distillation is proposed, including face image extraction and enhancement module, contrast learning common feature space module, buffer setting module, asymmetric supervision contrast loss module and multi-task instance-level relationship distillation module. Through the combination of these modules, the multi-task instance-level relationship distillation of the model is realized.

Benefits of technology

It significantly improves the generalization ability and robustness of the model, reduces the risk of catastrophic forgetting, and allows the model to maintain excellent performance and adaptability in a multitasking environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032415A_ABST
    Figure CN120032415A_ABST
Patent Text Reader

Abstract

The invention discloses a face image depression recognition system based on instance-level relation distillation, and relates to the field of depression recognition and prediction. The problems of insufficient generalization ability and difficult model migration adaptation in the depression detection technology in the prior art are solved. The system comprises a face image extraction and enhancement module, a contrast learning common feature space module, an asymmetric supervision contrast loss module, a buffer area setting module and a multi-task instance-level relation distillation module. It is ensured that the learned knowledge can be effectively migrated when the model faces a new task, and therefore the risk of disastrous forgetting is reduced. The training process of the model is optimized, and the generalization ability and the accuracy of the model in multi-task depression detection and recognition are remarkably improved. The method is also suitable for multi-task face image depression auxiliary diagnosis under different data sets, and is specifically applied to the field of depression tendency recognition and prediction in face images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of depression recognition prediction, and in particular to a facial image depression recognition system based on instance-level relationship distillation. Background Art

[0002] Given that depression symptoms vary significantly among different populations, traditional depression detection technologies based on facial images generally face the problem of insufficient generalization, which poses a major challenge to the migration and adaptation of models. In the existing technical solutions, there are still some problems such as how to improve the generalization and robustness of depression detection models and how to prevent catastrophic forgetting. This application proposes a facial image depression recognition system based on instance-level relation distillation. Summary of the invention

[0003] The present invention aims to solve the problems of insufficient generalization ability and difficulty in model migration and adaptation in the existing depression detection technology. To solve the above technical problems, the present invention is implemented through the following technical solutions: Solution 1: The present invention proposes a facial image depression recognition system based on instance-level relation distillation, the facial image depression recognition system comprising: The face image extraction and enhancement module is used to collect interview video data of at least one data set, perform face image cropping and extraction tasks, and construct a correspondence between face images and depression labels. , used to train and validate the multi-task image depression detection and recognition model; The common feature space module is used to compare the face image and label input by the face image extraction and enhancement module. Processing is performed to extract common depression identification features and output sample information in the common feature space; Buffer setting module, used to store face image and label pairs , and compare the face image and label stored in the buffer Make updates; Asymmetric supervised contrast loss module is used to compare the face image updated by the buffer setting module with the label pair The samples with the same label in the buffer are regarded as positive sample sets, and the face images updated by the buffer setting module are compared with the label pairs. Samples with different labels are considered as negative sample sets; the similarity between positive samples and negative samples is compared; The instance-level relation distillation module under multi-task is used to update the face image processed by the buffer and the label pair. The input is input into the multi-task image depression detection and recognition model before and after the multi-task image depression detection and recognition model to generate feature vector representations respectively, and the difference between the instance-level similarities of the positive samples and the negative samples in the buffer before and after the multi-task image depression detection and recognition model is calculated.

[0004] Furthermore, a preferred embodiment is provided, wherein the facial image extraction and enhancement module further includes a step of enhancing the correspondence between the facial image and the depression label when performing the facial image cropping and extraction task.

[0005] Furthermore, a preferred implementation is provided, wherein the contrastive learning common feature space module, the buffer setting module, the asymmetric supervised contrast loss module, and the multi-task instance-level relationship distillation module all use enhanced output face images and depression label pairs.

[0006] Furthermore, a preferred embodiment is provided, wherein the multi-task instance-level relationship distillation module also includes a step of normalizing the temperature parameter.

[0007] Further, a preferred embodiment is provided, in which the method for calculating the difference between the instance-level similarities of positive samples and negative samples in the buffer before and after the depression detection model in the instance-level relationship distillation module under the multi-task is:

[0008] in, is the co-spatial feature map representation output by the depression detection model, is a hyperparameter, p i The samples with the same label as anchor point i are regarded as positive sample sets, S is the number of samples in the buffer setting module, k represents the number of positive and negative sample sets in the current batch, and z p All positive sample sets in the current batch, z k Represents all positive and negative sample sets in the current batch, except for sample outside.

[0009] Furthermore, a preferred implementation is provided, wherein the buffer setting module further includes a step of performing comparative learning on positive samples and negative samples through a SupCon loss function.

[0010] Furthermore, a preferred implementation is provided, wherein the multi-task instance-level relationship distillation module also includes a step of aligning feature representations of a previous depression detection model and a current depression detection model.

[0011] Furthermore, a preferred implementation is provided, in which an instance relationship distillation method is performed using an IRD loss function to align feature representations of a previous depression detection model with a current depression detection model.

[0012] The present invention is beneficial in that: The facial image depression recognition system based on instance-level relation distillation described in this invention implements an innovative multi-task instance-level relation distillation framework, which is specifically built for multi-task image depression detection and recognition models. The workflow of the framework is divided into several key steps to ensure that the model can achieve optimal performance when dealing with complex multi-task environments.

[0013] The face image depression recognition system based on instance-level relation distillation of the present invention performs detailed face image cropping, label matching and data augmentation operations on each data set. This step not only ensures the accuracy and diversity of the input data, but also enhances the model's adaptability to various changes that may occur in actual application scenarios by applying a series of advanced image processing techniques such as random cropping, color jittering and blur processing.

[0014] The facial image depression recognition system of the present invention adopts an asymmetric supervised contrast loss strategy between batches to optimize the model training process. The core of this method is to enable the model to pay more attention to the learning of sample features within the batch in each iteration through the calculation of supervised contrast loss, while reducing excessive reliance on past task samples. This asymmetric design helps the model maintain sensitivity to key features during the learning process and effectively improves the model's ability to distinguish different tasks.

[0015] The facial image depression recognition system with instance-level relation distillation described in the present invention uses instance-level relation distillation technology under multi-task to further improve the generalization ability of the model. This technology captures and transmits the relationship between samples of different tasks, ensuring that the model can effectively transfer the learned knowledge when facing new tasks, thereby reducing the risk of catastrophic forgetting. This instance-level relation distillation not only improves the robustness of the model in a multi-task environment, but also ensures that the model can maintain excellent performance in the face of changing task requirements.

[0016] In summary, the facial image depression recognition system based on instance-level relation distillation described in the present invention not only optimizes the training process of the model through a series of carefully designed steps, but also significantly improves the generalization ability and accuracy of the model in multi-task depression detection and recognition.

[0017] The present invention is also applicable to multi-task facial image depression-assisted diagnosis under different data sets, specifically in the field of depression tendency recognition and prediction in facial images. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of the framework of a facial image depression recognition system based on instance-level relationship distillation as described in Implementation Method 1.

[0019] Figure 2 This is a schematic diagram of the overall process of the multi-task instance-level relationship distillation model described in the facial image depression recognition system using an instance-level relationship distillation method described in Implementation Method 1.

[0020] Figure 3 A schematic diagram of the asymmetric supervised contrast loss module between batches described in the facial image depression recognition system using instance-level relational distillation as described in Implementation Example 1.

[0021] Figure 4 This is a schematic diagram of the instance-level relationship distillation module under multiple tasks described in the facial image depression recognition system using an instance-level relationship distillation method described in Implementation Method 1. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the implementation methods of the present application clearer, the technical solutions in the implementation methods of the present application will be clearly and completely described below in conjunction with the drawings in the implementation methods of the present application. Obviously, the described implementation methods are only part of the implementation methods of the present application, not all of the implementation methods.

[0023] Embodiment 1: This embodiment provides a facial image depression recognition system based on instance-level relation distillation, and the facial image depression recognition system includes: The face image extraction and enhancement module is used to collect interview video data of at least one data set, perform face image cropping and extraction tasks, and construct a correspondence between face images and depression labels. , used to train and validate the multi-task image depression detection and recognition model; The common feature space module is used to compare the face image and label input by the face image extraction and enhancement module. Processing is performed to extract common depression identification features and output sample information in the common feature space; Buffer setting module, used to store face image and label pairs , and compare the face image and label stored in the buffer Make updates; Asymmetric supervised contrast loss module is used to compare the face image updated by the buffer setting module with the label pair The samples with the same label in the buffer are regarded as positive sample sets, and the face images updated by the buffer setting module are matched with the label pairs. Samples with different labels are considered as negative sample sets; the similarity between positive samples and negative samples is compared; The instance-level relation distillation module under multi-task is used to update the face image processed by the buffer and the label pair. The input is input into the multi-task image depression detection and recognition model before and after the multi-task image depression detection and recognition model to generate feature vector representations respectively, and the difference between the instance-level similarities of the positive samples and the negative samples in the buffer before and after the multi-task image depression detection and recognition model is calculated.

[0024] Implementation method 2. This implementation method is a further limitation of the facial image depression recognition system based on instance-level relationship distillation described in Implementation method 1. When performing facial image cropping and extraction tasks in the facial image extraction and enhancement module, it also includes a step of enhancing the correspondence between the facial image and the depression label.

[0025] Implementation method three. This implementation method is a further limitation of the facial image depression recognition system using instance-level relationship distillation described in implementation method two. The contrastive learning common feature space module, buffer setting module, asymmetric supervision contrast loss module, and multi-task instance-level relationship distillation module all use enhanced output facial images and depression label pairs.

[0026] Implementation method 4: This implementation method further limits the facial image depression recognition system of instance-level relationship distillation described in Implementation method 1, so that the instance-level relationship distillation module under multiple tasks also includes a step of normalizing the temperature parameter.

[0027] Implementation 5. This implementation is a further limitation of the facial image depression recognition system of the instance-level relationship distillation described in Implementation 1. The method for calculating the difference between the instance-level similarities of the positive samples and the negative samples in the buffer before and after the depression detection model in the instance-level relationship distillation module under the multi-task is:

[0028] in, is the co-spatial feature map representation output by the depression detection model, is a hyperparameter, p i The samples with the same label as anchor point i are regarded as positive sample sets, S is the number of samples in the buffer setting module, k represents the number of positive and negative sample sets in the current batch, and z p All positive sample sets in the current batch, z k Represents all positive and negative sample sets in the current batch, except for sample outside.

[0029] Implementation method six: This implementation method further limits the facial image depression recognition system using instance-level relationship distillation described in implementation method one, and the buffer setting module also includes a step of performing comparative learning on positive samples and negative samples using a SupCon loss function.

[0030] Implementation method seven. This implementation method is a further limitation of the facial image depression recognition system using instance-level relationship distillation described in implementation method one. The multi-task instance-level relationship distillation module also includes a step of aligning feature representations of a previous depression detection model with a current depression detection model.

[0031] Implementation method eight. This implementation method further limits the facial image depression recognition system using instance-level relationship distillation described in implementation method seven. The IRD loss function is used to perform instance relationship distillation method to align the feature representations of the previous depression detection model and the current depression detection model.

[0032] Embodiment 9: This embodiment provides an example, which is used to explain the above embodiments 1 to 8. The specific example is as follows: See also Figures 1 to 4 To illustrate this implementation, this embodiment proposes a method for displaying a target video label based on the installation position of a panoramic radar and its calibration parameters, and the method comprises the following steps: The multi-task facial image depression auxiliary diagnosis system with instance-level relationship distillation provided in this embodiment includes a facial image extraction and enhancement module, a contrastive learning common feature space module, an asymmetric supervision contrast loss module between batches, a buffer setting module and an instance-level relationship distillation module under multiple tasks.

[0033] 1. Face image extraction and enhancement module: This module focuses on processing interview video materials from multiple datasets, performs precise face image cropping and extraction tasks, and adopts a series of cutting-edge image enhancement technologies, such as RandomResizedCrop, ColorJitter, RandomGrayScale, and GaussianBlur. This application regards the depression detection tasks under these datasets as independent task units. In each dataset, that is, in each specific task, we construct a correspondence between face images and depression labels. , to facilitate model training and verification.

[0034] Among them, the RandomResizedCrop technology effectively simulates partial occlusion or changes in face position that may occur in actual scenes by randomly cropping and scaling images, thereby improving the robustness of the model. The ColorJitter technology further enhances the adaptability of the model by randomly adjusting the brightness, contrast, saturation, and hue of the image to simulate different lighting conditions and color distortion. The RandomGrayScale technology converts color images into grayscale images with a probability of 20%, which not only simulates the lack of color information that may be encountered in actual applications, but also helps to improve the model's generalization ability for color changes. In addition, the GaussianBlur technology simulates possible blurring during image acquisition by introducing Gaussian blur, making the model more robust. The comprehensive application of these technologies aims to comprehensively improve the model's adaptability to a variety of actual application scenarios and ensure the accuracy and reliability of the diagnostic system.

[0035] Here, see Figures 2 to 4 As shown in , we define x as the input face image, y as the depression label corresponding to x, and t as the index of the task set, where t ∈ {1, 2, …, T}, for example Figure 3 and Figure 4 In the task t1 and task t2, T is the total number of tasks. Here, different tasks are depression recognition under different data sets. At the same time, for each task, the face image-depression label pair Perform two random enhancement versions and ,in , the augmented dataset is , whose size is 2N, Among them, the innovative multi-task instance-level relation distillation (IRD) method described in this embodiment aims to enhance the generalization ability and robustness of the depression detection model in complex and changing environments, and realize the distillation migration of the model.

[0036] Instance-level relation distillation (IRD) technology focuses on extracting and transferring the relationships between different instances, which has significant advantages for multi-task learning. Through IRD, we can more effectively improve the generalization and robustness of the depression detection model and achieve model distillation migration. This is because IRD technology helps the model understand and grasp the intrinsic connections and differences between different tasks, so that the model will not forget the previously learned task information while absorbing new task knowledge. This ability is crucial to maintaining the model's long-term memory and adapting to new environments, thereby significantly improving the accuracy and reliability of depression detection under the multi-task learning framework.

[0037] The basic concept of contrastive learning loss is to maximize the similarity between different views of the same sample while minimizing the similarity between different samples. This method uses the data augmentation technique to increase the diversity of data by generating different versions of samples. In this way, the model can learn features that represent the data well. Contrastive learning includes positive sample contrast loss and negative sample contrast loss.

[0038] Positive Pair Contrast: This part of the loss is used to compare between enhanced versions (or different views of the same sample). The goal is to make these different views of the same sample as close as possible in the representation space, usually by maximizing their similarity.

[0039] Negative Pair Contrast: This part of the loss is used to compare different samples. The goal is to make different samples far away from each other in the representation space, usually by minimizing their similarity.

[0040] Assume that the original image , perform different transformations on A and B to obtain A( )) and B( ), at this time, the comparison loss hopes A( ) and B( ) is smaller than A( ) and any other images , The distance, that is, dist(A( ),B( )) dist(A( ), )), if the cosine distance is used, it is as follows:

[0041] in is the number of positive samples, which is 2N.

[0042] In this way, we ensure that the model can maintain high recognition accuracy and generalization ability in the face of diverse task environments. The design of this multi-task learning framework not only takes into account the differences between datasets, but also provides a more comprehensive and detailed perspective for the model to cope with the complex and ever-changing challenges of facial depression detection. All the image-label pairs used below are enhanced output face image-depression label pairs.

[0043] 2. Contrastive learning common feature space module: The core task of this module is to process the input image-label pair , through feature extraction and projection mapping operations, we extract efficient and representative universal depression recognition features. To this end, we define a parameter pair =( ,w), where w( ) represents a linear classifier, and, represents the feature representation. In addition, we use To represent the feature projection mapped to the common feature space, and It means feature extraction With feature map We use To represent the combined operation of feature extraction and projection mapping. On this basis, we conduct a detailed training process on the feature map and output the sample feature information in the common feature space, aiming to optimize the model's ability to extract and identify depression features.

[0044] The face image-depression label that has passed the face image extraction and enhancement module is used through feature extraction and projection mapping operations to extract efficient and representative universal depression recognition features.

[0045] 3. Buffer setting module: The core of this module is to set up a buffer to store samples of past tasks so that they can be reviewed and learned in subsequent tasks, thereby reducing catastrophic forgetting. The buffer setting and updating process is as follows.

[0046] The buffer size can be adjusted according to the specific situation of the task. Use 200 or 500 samples as the buffer size. The buffer stores samples of past tasks, including images and labels. The number of samples in the buffer should be balanced to avoid class imbalance problems. After the initial task training is completed, a part of the training samples are pushed into the buffer. The number of pushed samples can be adjusted according to the buffer size and the number of task samples. After the training of each new task is completed, a part of the training samples of the task is pushed into the buffer. At the same time, a part of the samples are deleted from the buffer to keep the buffer size unchanged. The strategy for deleting samples can be adjusted according to the category importance or the tendency of sample forgetting. During the subsequent task training, the samples in the buffer are trained together with the current task samples.

[0047] Buffer setting module: Set up a buffer to store samples of past tasks so that they can be retrospectively learned in subsequent tasks, thereby reducing catastrophic forgetting. After each new task training is completed, a portion of the training samples are pushed into the buffer. At the same time, a portion of the samples are deleted from the buffer to keep the buffer size unchanged. The strategy for deleting samples can be adjusted based on the category importance or the tendency of sample forgetting. During the training of subsequent tasks, the samples in the buffer are trained together with the current task samples. The SupCon loss function is used for contrastive learning, and the samples in the buffer are used as negative samples. The IRD loss function is used for instance relationship distillation to align the feature representation of the past model with the feature representation of the current model.

[0048] 4. Asymmetric supervised contrast loss module between batches: The core task of this module is to reduce the supervised contrast loss batch by batch during the model training process to achieve a specific effect. This design aims to enhance the model's ability to identify the characteristics of samples within a batch while avoiding excessive reliance on past task samples. In each batch of training, the current task sample is used as an anchor point, and the samples with the same label as anchor point i in the current batch are regarded as positive samples. Represents; samples in the buffer with different labels from the anchor point are regarded as negative sample sets, and Indicates, M indicates the number of buffer samples. Represents all sample sets in the current batch and the buffer, and the number of samples is 2N+M. We define the supervised contrast loss within a batch as:

[0049] in, is the co-spatial feature map representation output in module 2, is a hyperparameter used to control the similarity between positive and negative samples. This asymmetric supervised contrast loss design concept aims to prevent the model from overfitting to a small number of historical task samples during training, while prompting the model to pay more attention to the changes and differences in sample features between different batches. In this way, the model can more accurately capture the essential characteristics of the data, thereby improving its generalization ability and recognition performance in complex and changing environments.

[0050] Asymmetric supervised contrast loss module between batches: In the model training process, this module focuses on reducing the supervised contrast loss batch by batch, with the goal of achieving a series of specific training effects. The core of this strategy is to improve the model's ability to identify the characteristics of samples within a batch, ensuring that the model can extract key information from samples in each batch, while avoiding the model's excessive dependence on previously encountered task samples. In this way, the model can maintain high accuracy and robustness when facing new, unseen data.

[0051] Instance-level relation distillation module under multi-task: The core task of this module is to propose an instance-level relation distillation (IRD) method in the multi-task model training process of different depression datasets. IRD adjusts the changes in feature relationships between batch samples through self-distillation to achieve generalization of model capabilities.

[0052] The IRD module is mainly performed between Model 1 and Model 2, where Model 1 is the model of the past task and Model 2 is the model of the current task. Model 1 is the model after completing the training of the past task, and its parameters are fixed. Model 2 is the model of the current task, and its parameters are continuously updated during the training process. The buffer samples after the task update (including the randomly balanced sample set in the previous task) are input into Model 1 and Model 2 to generate feature vector representations respectively. The difference between the instance-level similarity of the buffer samples of Model 1 and Model 2 is calculated, that is, the dot product between the feature vectors, and normalized by the temperature parameter.

[0053] We define Represents the normalized instance-level similarity between anchor point (sample) i and sample j:

[0054] Given by Parameterized Representation and Temperature Hyperparameters In other words, the instance-level similarity vector p( ) is the normalized similarity between a sample and other samples in the batch.

[0055] Specifically, for each sample in a batch , we define the normalized similarity between anchor point i and other samples in the batch and buffer samples as:

[0056] The IRD loss function encourages Model 2 to keep the instance-level similarity of buffer samples consistent with Model 1. At the same time, IRD loss is used together with SupCon loss to update the parameters of Model 2. In this way, Model 2 can learn more discriminative feature representations and maintain the stability of feature representations during multi-task learning, thereby effectively reducing catastrophic forgetting and improving the model's continuous learning ability.

[0057] That is, the IRD loss quantifies the difference in instance-level similarity between the current representation and the past representation; the past representation refers to the snapshot of the model at the end of the previous task. and represents the parameters of the past / current model, and the IRD loss is defined as:

[0058] Among them, the logarithm and multiplication on the vector represent the element-wise logarithm and multiplication. We need to note that we use different hyperparameters k for the past and current similarity vectors; on the other hand, these two hyperparameters The number will remain constant throughout the mission.

[0059] Instance-level relation distillation module under multi-task: Under the framework of multi-task learning, this module uses instance-level relation distillation method to adjust the changes in feature relations between samples in different batches, so as to achieve generalization of model capabilities. The key to this method is to maintain the stability of feature representation, so that key knowledge can be effectively retained and transferred even in the learning process of multiple tasks. Instance-level relation distillation significantly reduces the problem of catastrophic forgetting by maintaining the model's understanding of the common features of different tasks, that is, the model will not forget the previously learned knowledge when learning new tasks. This not only improves the model's continuous learning ability, but also ensures that the model can maintain consistent efficiency and accuracy when handling multiple tasks. Through this strategy, the model can show better adaptability and generalization capabilities in a constantly changing and expanding task environment.

[0060] For the purposes of the present specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or apparatus or in conjunction with such instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) with one or N wirings, a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program may be printed, because the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting or processing in other suitable ways as necessary, and then stored in a computer memory. It should be understood that various parts of the present invention may be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, N steps or methods may be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0061] Those skilled in the art will appreciate that the above are only preferred embodiments of the present invention, and the various embodiments of the present disclosure and / or the features described in the claims may be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. It is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art may still modify the technical solutions described in the aforementioned embodiments, or perform equivalent substitutions on some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

[0062] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A facial image depression recognition system based on instance-level relation distillation, characterized in that: The facial image depression recognition system comprises: The face image extraction and enhancement module is used to collect interview video data of at least one data set, perform face image cropping and extraction tasks, and construct a correspondence between face images and depression labels. , used to train and validate the multi-task image depression detection and recognition model; The common feature space module is used to compare the face image and label input by the face image extraction and enhancement module. Processing is performed to extract common depression identification features and output sample information in the common feature space; Buffer setting module, used to store face image and label pairs , and compare the face image and label stored in the buffer Make updates; Asymmetric supervised contrast loss module is used to compare the face image updated by the buffer setting module with the label pair The samples with the same label in the buffer are regarded as positive sample sets, and the face images updated by the buffer setting module are matched with the label pairs. Samples with different labels are considered as negative sample sets; compare the similarity between positive samples and negative samples; The instance-level relation distillation module under multi-task is used to update the face image processed by the buffer and the label pair. The input is into the multi-task image depression detection and recognition model before and after the multi-task image depression detection and recognition model, and feature vector representations are generated respectively. The difference between the instance-level similarities of the positive samples and the negative samples in the buffer before and after the multi-task image depression detection and recognition model is calculated.

2. The facial image depression recognition system based on instance-level relation distillation according to claim 1, characterized in that: When executing the facial image cropping and extraction task in the facial image extraction and enhancement module, the step of enhancing the corresponding relationship between the facial image and the depression label is also included.

3. The facial image depression recognition system based on instance-level relation distillation according to claim 2, characterized in that: The contrastive learning common feature space module, buffer setting module, asymmetric supervision contrast loss module, and multi-task instance-level relationship distillation module all use enhanced output face images and depression label pairs.

4. The facial image depression recognition system based on instance-level relation distillation according to claim 1, characterized in that: The multi-task instance-level relationship distillation module also includes a step of normalizing the temperature parameter.

5. The facial image depression recognition system based on instance-level relation distillation according to claim 1, characterized in that: The method for calculating the difference between the instance-level similarities of positive samples and negative samples in the buffer before and after the multi-task image depression detection and recognition model in the multi-task instance-level relationship distillation module is: in, is the co-spatial feature map representation output by the depression detection model, is a hyperparameter, p i The samples with the same label as anchor point i are regarded as positive sample sets, S is the number of samples in the buffer setting module, k represents the number of positive and negative sample sets in the current batch, and z p All positive sample sets in the current batch, z k Represents all positive and negative sample sets in the current batch, except for sample outside.

6. The facial image depression recognition system based on instance-level relation distillation according to claim 1, characterized in that: The buffer setting module also includes a step of performing comparative learning on positive samples and negative samples through a SupCon loss function.

7. The facial image depression recognition system based on instance-level relation distillation according to claim 1, characterized in that: The multi-task instance-level relationship distillation module also includes a step of aligning feature representations before and after the multi-task image depression detection and recognition model.

8. The facial image depression recognition system based on instance-level relation distillation according to claim 7, characterized in that: The IRD loss function is used to perform instance relation distillation method to align the feature representations of the multi-task image depression detection and recognition model before and after the multi-task image depression detection and recognition model.

Citation Information

Patent Citations

  • Video behavior recognition method based on weighted fusion of multiple image tasks

    CN113536922A

  • Face attribute recognition method and system based on self-adaptive comparison knowledge distillation

    CN114299591A

  • Unsupervised ship re-identification method and system based on multistage contrast learning

    CN114821237A

  • Image target detection method based on deep supervision self-distillation

    CN114863248A

  • Generalized continuous classification method based on online contrast distillation network

    CN114972839A