Group emotion recognition method and system based on multi-scale context and noise suppression
By adopting multi-scale context and noise suppression techniques in group emotion recognition, the problem of noise individual impact and label incorrect annotation is solved, and the recognition performance and consistency are improved.
Patent Information
- Application Number
- CN202510230390.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
Existing group emotion recognition methods are insufficient in dealing with noisy individuals, resulting in reduced model prediction capabilities and inconsistent results, and over-reliance on vague group-level labels may mask key individual characteristics.
Using a multi-scale context and noise suppression method, the face, object and global features are extracted through object detection, the multi-scale interactive attention module and noise perception fusion module are used to reduce the impact of noise individuals, and the negative impact of the wrong label category is minimized through the multi-branch label re-labeling module.
It significantly improves the model's perceived ability of emotion characteristics at different scales, reduces the impact of noise on group emotion recognition, improves recognition performance and consistency, and is suitable for complex and noisy environments.
Smart Images

Figure CN120071069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of group emotion recognition in artificial intelligence technology, and particularly relates to a group emotion recognition method and system based on multi-scale context and noise suppression. Background Art
[0002] Group emotion refers to the overall emotional state shown by a group composed of multiple individuals in a specific scenario. This emotion is affected by both individual emotions and the comprehensive effects of interactions among individuals within the group and external environmental factors. Compared with individual emotion recognition, group emotion recognition requires integrating multi-dimensional information in complex scenarios to infer the overall emotional characteristics of the group. In recent years, with the rapid development of social media and the increasing demand for public opinion monitoring, relying solely on individual emotion recognition is no longer sufficient to predict potential risks in complex public environments. Therefore, group emotion recognition has gradually become an important research direction in the field of artificial intelligence, and its application scenarios cover multiple fields such as public safety monitoring, social behavior analysis, and military strategy.
[0003] The core of group emotion recognition lies in capturing the context interaction relationship between individuals in the scenario and integrating these interactions to model the overall emotional state of the group. Existing methods usually use long short-term memory networks, graph neural networks, and attention mechanism-based technologies to model the complex relationships between individuals. Before aggregating individual features into group features, these methods generally refine the individual features to improve the accuracy of emotion prediction. However, most of these methods rely on a single group-level emotion label, making it a weakly supervised problem. The single label ignores the diversity of individual emotion expressions in the group, which may not only introduce noise but also lead to inconsistent network classification results, thus affecting the overall prediction performance.
[0004] To address the weakly supervised problem, some studies have proposed new strategies to reduce the interference of noisy individuals on the model by assigning different weights. For example, some studies have adopted a cascaded attention mechanism to assign higher weights to key individuals to highlight their role in emotion recognition; another type of study focuses on significant individual features through the attention mechanism. On this basis, a new semi-supervised group emotion recognition framework has also been developed to alleviate the problem of insufficient labeled samples by using unlabeled image data. However, these methods still have the following two main limitations: 1) Existing methods mainly focus on reducing the influence of noisy individuals in the feature aggregation process but fail to directly address the actual interference of noisy individuals on emotion recognition; in the case of a high proportion of noisy individuals, even if adjusted by assigning lower weights, the prediction ability of the model will still be significantly weakened; 2) Over-reliance on fuzzy group-level labels may obscure key individual features related to group emotion. This limitation not only leads to inconsistent prediction results but also further restricts the improvement of model performance. Summary of the Invention
[0005] Aiming at the deficiencies in the prior art, the present invention provides a group sentiment recognition method and system based on multi-scale context and noise suppression, which integrates individual and scene context information at different scales, models fine-grained individual interaction relationships, and jointly suppresses noise at the individual level and label level to reduce the impact of noise on group sentiment recognition, thereby improving the performance of group sentiment recognition.
[0006] The present invention achieves the above technical objectives through the following technical means.
[0007] Group sentiment recognition method based on multi-scale context and noise suppression:
[0008] Perform object detection on the original group image to obtain face images and object images;
[0009] Extract face features, object features, and global features from the face images, object images, and original group images respectively;
[0010] Send the face features, object features, and global features into a face emotion classification model, an object emotion classification model, and a global emotion classification model for training respectively to obtain the group-level emotion probabilities of each branch, and fuse the emotion probabilities of each branch into the group emotion probability, that is, the final classification result of the group emotion;
[0011] The face emotion classification model and the global emotion classification model interact through a multi-scale interactive attention module;
[0012] When sending the face features into the face emotion classification model for training, suppress noise individuals through a noise-aware fusion module;
[0013] During the training process of each branch, use a multi-branch label relabeling module to minimize the negative impact of mislabeled categories in the original group image.
[0014] Furthermore, the face emotion classification model and the global emotion classification model interact through a multi-scale interactive attention module, specifically: the global feature of each scale is concatenated with all face features x f to explore the importance weights of faces under global features at different scales using an attention network, and model the interaction relationships of faces at different scales.
[0015] Even further, the exploration of the importance weights of faces under global features at different scales using the attention network is specifically:
[0016]
[0017] α i,j = sigmoid(W(x ij ))
[0018] where cat(·) represents the concatenation operation, and W Q , W K , W V respectively represent converting x ij into query vector, key vector, and value vector through linear transformation. d k is the dimension of the key vector, α i,j represents the weight of the face feature of individual i under the global feature at the j-th scale. W(·) is the attention network, and sigmoid(·) is the normalization operation.
[0019] Furthermore, the global feature and the face feature are fused under the guidance of the weight α i,j to obtain the feature
[0020] Furthermore, the noise individual is suppressed through the noise-aware fusion module. Specifically: the weighted feature is sent to the classifier to obtain the individual emotion prediction result According to the individual emotion prediction result, three sub-groups of positive, negative, and neutral are divided. The individual emotion prediction results in each sub-group are fused, and then the emotion prediction results of the three sub-groups are fused as the final face emotion prediction result.
[0021] Furthermore, while the classifier obtains the individual emotion prediction result, the feature is sent to the Conv1D layer for fusion, and then the group emotion trend
[0022] Furthermore, is regarded as the individual certainty degree β i , and the weight of individuals with obvious emotion expression is increased through :
[0023]
[0024] where c represents the category to which the individual belongs, represents the emotion prediction results of different sub-groups.
[0025] Furthermore, the face emotion prediction result is: where the group emotion probability [·,·,·] represents concatenation on the last dimension, average(·) is the balancing function, and y fRepresents the facial emotion prediction result, that is, the group-level emotion probability of the face branch.
[0026] Furthermore, the multi-branch label re-labeling module is used to minimize the negative impact of incorrectly labeled categories in the original group image, specifically:
[0027] For the face branch, object branch or global branch, the prediction results of each branch include the prediction results of three categories: positive, negative and neutral. In each branch, the category corresponding to the maximum value of the three category prediction results is the prediction category l of the branch. pred , the maximum value of the prediction result is the prediction category l of the branch pred The probability value when With the original label l ture The predicted probability value of the corresponding category The difference exceeds the threshold δ b , then modify the image label l to l pred , otherwise it remains the original label l ture :
[0028]
[0029] A group emotion recognition system based on multi-scale context and noise suppression, including:
[0030] Preprocessing module, which performs target detection on the original group image;
[0031] Feature encoder, extracts facial features, object features and global features;
[0032] Sentiment classification model, which inputs individual features, object features and global features, and outputs the group-level sentiment probability of each branch;
[0033] Multi-scale interactive attention module, which realizes the interaction between the face emotion classification model and the global emotion classification model;
[0034] Noise-aware fusion module, used to fuse noise-suppressed individuals;
[0035] The multi-branch label re-labeling module reduces the negative impact of incorrectly labeled categories in the original group images during the training process of each branch.
[0036] The beneficial effects of the present invention are:
[0037] (1) The present invention aims at the possible noisy individuals in a group. By fusing each individual's face features with global features of different scales and inputting them into an attention network, the model's ability to perceive emotional features of different scales is significantly improved, thereby obtaining the importance weights of individuals under global features of different scales, and initially weakening the influence of noisy individuals. Secondly, the individual-level noise and label-level noise are processed respectively, and the influence of noise in group emotion recognition is jointly suppressed through the two. Therefore, the group emotion recognition method based on multi-scale context and noise suppression has the ability to distinguish the importance differences of individuals under global features of different scales and suppress noise in real scenarios, making it applicable to group emotion recognition in complex and noisy environments.
[0038] (2) The present invention proposes a multi-level noise suppression method. At the individual level, it aggregates individual prediction results, enhances the model's ability to accurately interpret group emotions by effectively suppressing the influence of misleading or irrelevant instances, and adaptively updates noise labels at the label level to minimize their negative impact on robust emotion prediction. The group facial images in complex environments are blurred, making it difficult to accurately label individuals in the data, and there is only a unique group emotion label. However, there are noisy individuals with inconsistent individual emotion predictions in the group, and the existence of these noisy individuals may seriously confuse the model's judgment. At the same time, due to the inherent subjectivity of emotion labels, these labels are prone to introducing noise. In these cases, simply assigning low weights to noisy individuals cannot solve their direct impact on emotion recognition, and a high proportion of noisy individuals will significantly damage the prediction. At the same time, over-relying on vague group-level labels will obscure individual features related to group emotions, resulting in inconsistent predictions and a decline in model performance. The group emotion recognition method based on multi-scale context and noise suppression realizes the joint suppression of noise at the individual level and label level, effectively reducing the influence of noise on group emotion recognition, and thus being applicable to more complex real scenarios.
[0039] (3) The group emotion recognition method and system based on multi-scale context and noise suppression of the present invention are designed to be lightweight, plug-and-play, and are easy to efficiently embed into any emotion classification model to complete the group emotion recognition task, greatly alleviating the noise problem in the group emotion recognition task. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a framework diagram of the group emotion recognition based on multi-scale context and noise suppression according to the present invention;
[0041] Fig. 2(a) is a distribution diagram of visualized features of the basic model;
[0042] Fig. 2(b) is a distribution diagram of visualized features of the group emotion recognition method based on multi-scale context and noise suppression according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0043] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.
[0044] As Figure 1 shown, the group emotion recognition system based on multi-scale context and noise suppression of the present invention includes a preprocessing module, a feature encoder, an emotion classification model, a multi-scale interactive attention module, a noise-aware fusion module, and a multi-branch label relabeling module. The emotion classification model includes a face emotion classification model, an object emotion classification model, and a global emotion classification model, wherein the face emotion classification model and the global emotion classification model interact through a multi-scale interactive attention module; the emotion classification models of the present invention all adopt a classifier composed of a fully connected layer and an activation function. The preprocessing module performs object detection on the original group image to obtain the face images and object images of the individual branches. The feature encoder extracts the corresponding features of the face branch, object branch, and global branch respectively. The face features, object features, and global features are respectively sent into the corresponding emotion classification models for training to obtain the group-level emotion probabilities of each branch. When the face features are sent into the face emotion classification model for training, the noise-aware fusion module is used to suppress noise individuals. Finally, the emotion probabilities of each branch are fused into the group emotion probability. During the training process of each branch, the multi-branch label relabeling module is used to minimize the negative impact of mislabeled samples (the originally labeled categories).
[0045] The specific process of the group emotion recognition method based on multi-scale context and noise suppression is as follows:
[0046] (1) Use the MTCNN network (Multi-task Cascaded Convolutional Networks) to locate and crop the faces in the original group image. The obtained pictures contain the face information of all individuals, and use the pre-trained Faster-RCNN network (Region-based Convolutional Neural Networks) to identify the main objects in the cropped pictures. The cropped face pictures and object pictures are input into the pre-trained vgg16 network (a kind of feature encoder) to obtain the face features of the i-th individual and the k-th related object features For the object branch, after fusing all object features and inputting them into the object emotion classification model, the prediction result y o is obtained, that is, the group-level emotion probability of the object branch. For the overall picture containing global information, it is sent into the resnet50-fpn network (Feature Pyramid Network, a kind of feature encoder) to obtain global features at different scales Among them, j represents the j-th scale. Subsequently, rolalign(·) (Region of Interest Alignment) is used to align the global features of different scales. After splicing and fusion, the prediction result y of the global branch is finally obtained by the global sentiment classification model g , that is, the group-level sentiment probability of the global branch.
[0047] (2) Input the face features and the global features of different scales. The multi-scale interactive attention module effectively captures the rich emotional features of individuals in different backgrounds by using the global background features of different scales; specifically:
[0048] Concatenate the global features of each scale with all the face features x f to explore the importance weights of the face under the global features of different scales by using the attention network, and model the interaction relationship of the face at different scales; α i,j represents the weight of the face feature of individual i under the global feature of the j-th scale, which is obtained by the attention network W(·) and sigmoid(·):
[0049]
[0050] α i,j = sigmoid(W(x ij )) (3)
[0051] Among them, cat(·) represents the concatenation operation, and W Q , W K and W V represent converting x ij into query, key, and value vectors through three linear transformations; calculate the inner product between the query vector W Q (x ij ) and the key vector W K (x ij ) to measure the similarity. The similarity calculation result is scaled by dividing by , where d k is the dimension of the key vector; then use the softmax function to obtain the attention weights, representing the relationship strength between each individual face and all other individual faces, and multiply the attention weights by the value vector W V (x ij ) to model the interaction relationship of the face at different scales, and finally normalize the output result to the weight α i,j through the sigmoid function;
[0052] The concatenated global features and the face features are weighted by α i,jCalculated under the guidance of to fully identify important individuals under different-scale global backgrounds:
[0053]
[0054] (3) To clearly constrain the negative impact of noisy instances, based on the package-level prediction of face feature fusion, the output results of the multi-scale interaction attention module are fed into the noise-aware fusion module. According to the emotional classification results of individual faces, individuals are divided into positive, negative, and neutral sub-groups, and the interaction among individuals within the same sub-group is explored intensively. After reducing the impact of noisy individuals, the face prediction results are fused to further enhance the model's ability to understand group emotions.
[0055] Specifically manifested as:
[0056] The feature is fed into the Conv1D layer (belonging to the face emotion classification model) for fusion, and after passing through the fully connected layer, the group emotion trend is preliminarily predicted Meanwhile, to distinguish noisy individuals, the feature is input into the classifier to predict the emotional scores of different individual faces, and the maximum value of the predicted scores is taken as the prediction result of the current individual face where c represents the category (positive, negative, and neutral) to which the individual belongs.
[0057] Based on the individual prediction results, the whole is divided into three sub-groups: positive sub-group, negative sub-group, and neutral sub-group. The individual prediction results in each sub-group are fused, and then the emotional prediction results of the three sub-groups are fused as the final face emotion prediction result. During the fusion process of the sub-groups, the constraint network only explores the associations among individuals with consistent emotional expressions within the same sub-group to suppress the influence of other noisy individuals with inconsistent emotional expressions.
[0058] Individual prediction results to a certain extent represent the certainty degree of the face emotion classification model's prediction of individuals, the lower it is, the more likely the model is to make a wrong judgment on the emotion of individual i. Therefore, is regarded as the individual certainty degree β at the same time i , through increasing the weights of individuals with obvious emotional expressions and suppressing the negative impact brought by confusing noisy individuals, then:
[0059]
[0060] By fusing the emotional prediction results of different sub-groups the group emotion probability that suppresses the influence of noisy individuals is obtained where [·, ·, ·] represents concatenation in the last dimension:
[0061]
[0062] Then, the average(·) function is used to balance the group sentiment trend obtained from the preliminary prediction and the group sentiment probability after individual noise suppression fusion to obtain the prediction result y of the face branch f :
[0063]
[0064] (4) After training reaches a certain level, re-label the different branches independently. Specifically, evaluate the deviation between the predicted class probability of a sample (a face image, an object image, or an original group image) and its true class probability. When this deviation exceeds a predefined threshold, a new pseudo-label will be assigned to the sample during the training process of each respective branch.
[0065] Determine the threshold δ by observing the model training process b , where b represents different branches; taking the face branch as an example, for the prediction result y of this branch f = [y f1 , y f2 , y f3 , y f1 , y f2 , y f3 represent the prediction results corresponding to the three categories of positive, negative, and neutral respectively. The category corresponding to the maximum value among the prediction results of the three categories is the predicted category l of the face branch pred , and the maximum value of the prediction result is the probability value of the predicted category l of the face branch pred of When differs from the predicted probability of the corresponding category of the original label l ture (i.e., the labeled category of the branch image) by more than the threshold δ , then modify the picture label l to l b , otherwise it remains the original label l ture : fusion
[0066]
[0067] Among them, l represents the picture label, which is the annotation category (positive, negative, neutral) of the original group image in the training dataset; during the training process of each branch, the face picture label, object picture label, and global picture label are consistent with the annotation category of the original group image and are dynamically adjusted through the above label relabeling method.
[0068] (5) Finally, through the face branch, object branch, and global branch, obtain the emotion prediction results reflected by different information in the picture. To alleviate the emotion differences between different information, fuse the emotion features of each branch to obtain the feature fusion prediction result y fusion . For the three branches and the feature fusion prediction result, adopt the grid search method to obtain appropriate decision fusion parameters. The specific prediction formula is as follows:
[0069] y fusion = W fusion (cat(x f , x g , x o )) (10)
[0070] y last = α * y f + β * y g + γ * y o + δ * y fusion (11)
[0071] Among them, x f , x g , x o respectively represent the face feature, object feature, and global feature, W fusion (·) represents the classifier composed of the fully connected layer and the activation function, y last represents the group emotion probability, and α, β, γ, δ are the fusion weights of each branch, and the sum of the weights is 1.
[0072] The face branch emotion classification model, object branch emotion classification model, and global branch emotion classification model are all trained using cross-entropy loss.
[0073] Figures 2(a) and (b) are the comparison of the visualization feature distribution maps between the basic model and the group emotion recognition method based on multi-scale context and noise suppression described in the present invention. The basic model is obtained by deleting the multi-scale interactive attention module, the noise perception fusion module, and the multi-branch label relabeling module on the basis of the present invention. Among them, as shown in Figure 2(a), the basic model has difficulty in distinguishing different emotion categories, mainly affected by individual noise and label inconsistency, which often leads to incorrect emotion classification. In this case, the ability of the model to correctly capture group emotions is severely interfered by noise data, and the noise still exists even after basic classification processing. While in Figure 2(b), through the application of the group emotion recognition method based on multi-scale context and noise suppression, the overlap between emotion categories is significantly reduced, forming a more distinct cluster distribution. This clustering phenomenon indicates that the present invention has successfully suppressed the influence of noise and enhanced the distinguishability between emotion categories. Even in the case where individual emotion labels are unavailable, the model can still effectively capture the subtle differences between the emotions of group members.
[0074] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the substantial content of the present invention, any obvious improvements, substitutions, or variations that those skilled in the art can make all belong to the protection scope of the present invention.
Claims
1. A group emotion recognition method based on multi-scale context and noise suppression, characterized by: Perform target detection on the original group image to obtain face images and object images; Extracting facial features, object features and global features from facial images, object images and original group images respectively; The facial features, object features and global features are respectively sent to the facial emotion classification model, the object emotion classification model and the global emotion classification model for training, to obtain the group-level emotion probability of each branch, and the emotion probabilities of each branch are integrated into the group emotion probability, that is, the final classification result of the group emotion; The face emotion classification model and the global emotion classification model interact through a multi-scale interactive attention module; When the facial features are fed into the facial emotion classification model for training, the noise individuals are suppressed through the noise perception fusion module; During the training process of each branch, a multi-branch label re-labeling module is used to minimize the negative impact of incorrectly labeled categories in the original group images.
2. The group emotion recognition method based on multi-scale context and noise suppression according to claim 1 is characterized in that: The face emotion classification model and the global emotion classification model interact through a multi-scale interactive attention module, specifically: the global features of each scale are With all facial features x f Perform splicing and use the attention network to explore the importance weights of global features of faces at different scales, and model the interactive relationships of faces at different scales.
3. The group emotion recognition method based on multi-scale context and noise suppression according to claim 2 is characterized in that: The attention network is used to explore the importance weights of global features of faces at different scales, specifically: α i,j =sigmoid(W(x ij )) Among them, cat(·) represents the concatenation operation, W Q , W K , W V Respectively represent the linear transformation of x ij Converted into query vector, key vector and value vector, d k is the dimension of the key vector, α i,j represents the weight of the facial features of individual i under the global features of the jth scale, W(·) is the attention network, and sigmoid(·) is the normalization operation.
4. The group emotion recognition method based on multi-scale context and noise suppression according to claim 3 is characterized in that: The global features and face features are weighted by α i,j Under the guidance of 5. The method for group emotion recognition based on multi-scale context and noise suppression according to claim 4, characterized in that: The noise-aware fusion module is used to suppress noise individuals. Specifically, the weighted features Send it to the classifier to get the individual emotion prediction result According to the individual emotion prediction results, the face is divided into three subgroups: positive, negative and neutral. The individual emotion prediction results in each subgroup are fused, and then the emotion prediction results of the three subgroups are fused as the final face emotion prediction result.
6. The method for group emotion recognition based on multi-scale context and noise suppression according to claim 5, characterized in that: While the classifier obtains the individual emotion prediction results, the feature The data is sent to the Conv1D layer for fusion, and then the group sentiment trend is initially predicted by the fully connected layer.
7. The method for group emotion recognition based on multi-scale context and noise suppression according to claim 5, characterized in that: Will At the same time, it is regarded as the degree of individual certainty β i ,pass Increase the weight of individuals who express obvious emotions: Among them, c represents the category to which the individual belongs, Represents the sentiment prediction results of different subgroups.
8. The method for group emotion recognition based on multi-scale context and noise suppression according to claim 7, characterized in that: The facial emotion prediction result is: Among them, the group sentiment probability [·,·,·] means concatenation on the last dimension, average(·) is the balance function, y f Represents the facial emotion prediction result, that is, the group-level emotion probability of the face branch.
9. The method for group emotion recognition based on multi-scale context and noise suppression according to claim 1, characterized in that: The multi-branch label re-labeling module is used to minimize the negative impact of incorrectly labeled categories in the original group image, specifically: For the face branch, object branch or global branch, the prediction results of each branch include the prediction results of three categories: positive, negative and neutral. In each branch, the category corresponding to the maximum value of the three category prediction results is the prediction category l of the branch. pred , the maximum value of the prediction result is the prediction category l of the branch pred The probability value when With the original label l ture The predicted probability value of the corresponding category The difference exceeds the threshold δ b , then modify the image label l to l pred , otherwise it remains the original label l ture :
10. A system for implementing the group emotion recognition method based on multi-scale context and noise suppression according to any one of claims 1 to 9, characterized in that: include: Preprocessing module, which performs target detection on the original group image; Feature encoder, extracts facial features, object features and global features; Sentiment classification model, which inputs individual features, object features and global features, and outputs the group-level sentiment probability of each branch; Multi-scale interactive attention module, which realizes the interaction between the face emotion classification model and the global emotion classification model; Noise-aware fusion module, used to fuse noise-suppressed individuals; The multi-branch label re-labeling module reduces the negative impact of incorrectly labeled categories in the original group images during the training process of each branch.