Emotion consistency double-channel aggregation group image sentiment recognition method and system
Patent Information
- Application Number
- CN202611256673.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]综上所述,现有视觉群体情感识别技术中,个体线索与全局线索之间缺乏显式的关联性建模,导致模型难以区分与全局情绪一致的有效线索和偏离整体的差异线索,也难以量化群体内部情绪差异对最终识别结果的影响
(1)通过计算个体人脸情绪响应与全局情绪参考之间的相似度,获得一致性分数,据此将人脸线索划分至高一致性通道和低一致性通道,分别进行加权聚合。该方式能够有效区分与全局情绪参考相对一致的人脸线索和相对偏离的人脸线索,避免将主导情绪线索、个体情绪偏差线索及低质量噪声线索简单混合为单一表征,从而提升了群体情感识别的准确性。
Smart Images

Figure CN122799482A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and artificial intelligence, and in particular to a method and system for group image emotion recognition based on dual-channel aggregation of emotion consistency. Background Technology
[0002] Group emotion recognition is a core technology in the fields of emotion computing, intelligent perception, and social behavior analysis. It aims to automatically analyze and judge the overall emotional state, sentiment, and trends of a group based on multi-source sensory data related to the group. In visual group emotion recognition tasks, the system needs to infer the overall emotion category at the group level from images or video frames containing multiple people, integrating the emotional expressions of individual members, the group's activity status, and contextual information. This technology has broad application prospects in areas such as smart classroom analysis, public safety monitoring, human-computer interaction, and market research.
[0003] Currently, visual group emotion recognition mainly relies on two types of information sources: individual facial expression cues based on facial regions, and global contextual cues encompassing the spatial distribution of scenes, objects, and people. For the former, existing methods typically first extract multiple facial regions from an image using face detection algorithms, then obtain individual facial representations using expression feature extraction networks, and finally aggregate these individual features into a group facial representation using strategies such as average pooling, max pooling, voting mechanisms, or attention weighting. For the latter, some research attempts to incorporate global information such as scene layout, human pose, or environmental objects, combining facial features with global image features through feature concatenation or multimodal fusion to improve recognition performance in complex scenes.
[0004] However, the aforementioned existing technical solutions still face many challenges in practical applications, making it difficult to meet the requirements of stable and accurate recognition in complex multi-person scenarios. Firstly, existing methods primarily focus on engineering aspects such as "how to effectively aggregate multiple facial features" or "how to fuse facial features with global contextual features," generally neglecting the modeling of the intrinsic correlation between individual emotional responses and the overall emotional atmosphere of the group. Specifically, in real-world group images, the contribution of facial emotional cues from different individuals to the overall group emotion varies significantly: some faces express emotions highly consistent with the dominant emotion of the group, effectively reflecting the overall emotional tendency; while others may exhibit expression patterns inconsistent with or even contradicting the overall emotional atmosphere due to individual emotional differences, facial occlusion, image blurring, pose deviation, or uneven local lighting. If existing methods indiscriminately perform simple averaging or ordinary attention aggregation on all facial features, they are prone to mixing high-contribution dominant cues, low-contribution individual bias cues, and abnormal cues caused by noise into a single representation, thereby weakening the model's ability to perceive key information.
[0005] Secondly, while introducing global contextual features can mitigate the interference of local noise to some extent, existing fusion strategies mostly involve direct splicing or addition at the feature level, failing to explicitly characterize the consistency or deviation between individual facial cues and the global emotion reference. This "black box" fusion approach not only lacks interpretability but also makes it difficult for the model to effectively distinguish which individual cues should be consistent with the global reference and reinforced, and which cues should be retained as meaningful intra-group emotional differences. Due to the lack of quantification methods for the degree of difference between the two types of cues, existing methods often experience a significant decrease in recognition accuracy and stability when faced with scenarios involving significant intra-group emotional differentiation or a few abnormal expressions.
[0006] In summary, existing visual group emotion recognition technologies lack explicit modeling of the correlation between individual and global cues. This makes it difficult for models to distinguish between valid cues consistent with the overall emotion and discrepancies that deviate from the overall picture, and also makes it difficult to quantify the impact of intra-group emotional differences on the final recognition results. How to structurally distinguish and specifically model crowd emotion cues based on the relationship between individuals and the overall picture has become one of the most important problems that urgently need to be solved in this field. Summary of the Invention
[0007] To address the shortcomings of existing technologies, the present invention aims to provide a group image emotion recognition method and system based on dual-channel aggregation of emotion consistency, which performs structured differentiation and targeted modeling of group emotional cues to improve the accuracy of group emotion recognition.
[0008] To achieve the above objectives, this invention provides a group image emotion recognition method based on dual-channel aggregation of emotion consistency, comprising the following steps: Face preprocessing is performed on the group image to be identified to obtain face images; Extract individual facial features from each face in the face image and global context features from the group image. Map the individual facial features and global context features to a feature space of a unified dimension to obtain face projection features and global projection features. An individual facial emotion response is generated based on the individual facial features of each face, a global emotion reference is generated based on the global projection features, and a consistency score is calculated between the individual facial emotion response and the global emotion reference. All faces within the same group of images are divided into a high-consistency face set and a low-consistency face set according to their consistency scores. The face projection features in each set are then weighted and aggregated to obtain high-consistency aggregated features and low-consistency aggregated features. The difference between the highly consistent aggregation feature and the low consistent aggregation feature is calculated to obtain the conflict intensity; By integrating the high-consistency aggregation features, the low-consistency aggregation features, the conflict intensity, and the global projection features, a classifier is used to output the group sentiment recognition result.
[0009] Furthermore, the step of performing face preprocessing on the group image to be identified to obtain face images further includes: Face detection is performed on the group image to obtain the bounding boxes and key point information of each face region; the key point information includes the location information of local feature points with semantic identifiers in the face image; The corresponding face region is cropped from the group image based on the location box to obtain a face image; The cropped face image is aligned using the key point information.
[0010] Furthermore, the step of calculating the consistency score between the individual facial emotion response and the global emotion reference also includes calculating using cosine similarity: in, This represents the consistency score corresponding to the i-th face. This represents the cosine similarity operation. Indicates an individual's facial emotion response. This indicates a general sentiment reference. The consistency score is normalized to the [0,1] interval to obtain the normalized consistency score.
[0011] Furthermore, the step of dividing all faces within the same group of images into a set of high-consistency faces and a set of low-consistency faces based on their consistency scores also includes: Calculate the mean consistency score based on the consistency scores of all faces within the same group of images: in, The mean value is the consistency value, and N is the number of faces in the same group of images, where N is greater than or equal to 1. Let be the consistency score corresponding to the i-th face; Faces with a consistency score greater than or equal to the mean consistency value are classified into a high consistency face set, and faces with a consistency score less than the mean consistency value are classified into a low consistency face set.
[0012] Furthermore, the face is divided into a high-consistency face set and a low-consistency face set, and the face projection features in each set are weighted and aggregated to form a dual-channel aggregation including a high-consistency channel and a low-consistency channel.
[0013] Furthermore, in the high-consistency channel, the face projection features in the high-consistency face set are weighted and summed according to a first weight to obtain the high-consistency aggregated features. The formula is as follows: Among them, the first weight This represents the face weight of the i-th face in the high-consistency channel. Represents a set of highly consistent faces. Represents the face projection features of the i-th face; In the low-consistency channel, the face projection features in the low-consistency face set are weighted and summed according to the second weight to obtain the low-consistency aggregated features. The formula is as follows: Among them, the second weight This represents the weight of the i-th face in the low-consistency channel. This represents a set of faces with low consistency. Let i represent the facial projection features of the i-th face.
[0014] Furthermore, the high consistency channel assigns a corresponding first weight to each face based on the normalized consistency score; the normalized consistency score is obtained by normalizing the calculated consistency score between the individual face emotion response and the global emotion reference to the [0,1] interval. The low consistency channel assigns a corresponding second weight to each face based on the reverse consistency score; the reverse consistency score is equal to 1 minus the normalized consistency score.
[0015] Furthermore, the conflict intensity is calculated using cosine distance: Where c represents the intensity of the conflict; This represents the cosine similarity operation; This indicates a high degree of consistency in aggregation characteristics; This indicates a low-consistency aggregation characteristic.
[0016] Furthermore, the step of fusing the high-consistency aggregation feature, the low-consistency aggregation feature, the conflict intensity, and the global projection feature, and outputting the group emotion recognition result through a classifier, further includes: concatenating the high-consistency aggregation feature, the low-consistency aggregation feature, the conflict intensity, and the global projection feature: ; in, Indicates fusion characteristics, This indicates a high degree of consistency in aggregation features. represents low-consistency aggregation features, c represents conflict intensity, and g represents global projection features; The fused features are input into the classifier to obtain the group emotion recognition result.
[0017] Furthermore, if the set of highly consistent faces or the set of low-consistency faces is empty, then the average feature or zero vector of all faces is used as the aggregate feature of the set.
[0018] Furthermore, it also includes a training step: using group images with group sentiment labels as training samples, calculating the main classification loss based on the difference between the predicted results and the true labels, and updating the parameters.
[0019] Furthermore, the training step also includes setting one or more of the following auxiliary supervision terms: High consistency auxiliary classification loss constrains high consistency aggregated features to have sentiment discrimination ability; Low consistency auxiliary classification loss constrains low consistency aggregated features to retain discriminative information; Global auxiliary classification loss constrains global projection features to capture group sentiment tendencies; The consistency constraint loss constrains the consistency relationship between the global emotion reference and the average representation of the emotion responses of all faces within the same group of images.
[0020] To achieve the above objectives, the present invention also provides a group image emotion recognition system based on dual-channel aggregation of emotion consistency, used to implement the group image emotion recognition method based on dual-channel aggregation of emotion consistency as described above, comprising: The face preprocessing module is used to preprocess the faces of the group images to be identified, and obtain face images; The feature extraction module is used to extract individual facial features of each face from the face image, extract global context features from the group image, and map the individual facial features and the global context features to a feature space of a unified dimension to obtain face projection features and global projection features. The consistency calculation module is used to generate an individual facial emotion response based on the individual facial features, generate a global emotion reference based on the global projection features, and calculate the consistency score between the individual facial emotion response and the global emotion reference. The dual-channel aggregation module is used to divide the faces into a high-consistency face set and a low-consistency face set based on the consistency scores of all faces in the same group of images, and to perform weighted aggregation of the face projection features in each set to obtain high-consistency aggregation features and low-consistency aggregation features. The conflict modeling module is used to calculate the difference between the high-consistency aggregation feature and the low-consistency aggregation feature to obtain the conflict intensity; The emotion classification module is used to fuse the high consistency aggregation features, the low consistency aggregation features, the conflict intensity, and the global projection features, and output the group emotion recognition result through a classifier.
[0021] To achieve the above objectives, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the computer program stored in the memory to implement the group image emotion recognition method with dual-channel aggregation of emotion consistency as described above.
[0022] To achieve the above objectives, the present invention also provides a computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the group image emotion recognition method based on dual-channel aggregation of emotion consistency as described above.
[0023] The present invention provides a group image emotion recognition method based on dual-channel aggregation of emotion consistency. By extracting and projecting features from the face region and the global image in the group image respectively, the facial cues are divided into two groups and aggregated according to the consistency relationship between each face feature and the global features. This achieves structured modeling of the emotional differences within the group, thereby improving the accuracy of group emotion recognition.
[0024] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. Attached Figure Description
[0025] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a group image emotion recognition method based on dual-channel aggregation of emotion consistency according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a group image emotion recognition system based on dual-channel aggregation of emotion consistency according to an embodiment of the present invention; Figure 3 This is a schematic diagram of an electronic device structure according to an embodiment of the present invention. Detailed Implementation
[0026] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0027] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0028] The term "comprising" and its variations as used in this invention are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0029] It should be noted that the concepts of "first" and "second" may be mentioned in this invention only to distinguish different devices, components or parts, and are not used to limit the order of the functions performed by these devices, components or parts or their interdependence.
[0030] It should be noted that the terms "one" and "multiple" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless explicitly stated otherwise in the context, they should be understood as "one or more". "Multiple" should be understood as two or more.
[0031] Example 1 In embodiments of the present invention, a group image emotion recognition method based on dual-channel aggregation of emotion consistency is provided, which is applicable to visual image emotion analysis tasks involving multiple people.
[0032] Figure 1 The flowchart below shows a group image emotion recognition method based on dual-channel aggregation of emotion consistency according to an embodiment of the present invention. Figure 1 The embodiments of the present invention will be described in further detail.
[0033] Those skilled in the art should understand that the numbering of the following steps is for the purpose of clarity of description only and does not constitute an absolute limitation on the execution order. In actual implementation, some steps can be executed in parallel or the order can be adjusted, as long as the overall identification process described in this invention can be achieved.
[0034] In step 101: Obtain the group image to be identified, and perform face preprocessing on the group image to obtain an aligned face image sequence.
[0035] This step first acquires the group image to be identified. This group image can be a static image stored in a local image dataset, or a single frame image captured from a camera, monitoring system, or video stream. If the input is a video stream, keyframes can be extracted or single frames can be extracted at preset sampling intervals as processing objects, depending on the actual application requirements.
[0036] After acquiring the group images, face preprocessing is performed. This preprocessing specifically includes three sub-steps: face detection, cropping, and alignment. Face detection: Face detection algorithms are used to locate face regions in group images and obtain the bounding boxes and key point information of each face region.
[0037] Face cropping: Based on the detected face bounding boxes, crop the corresponding face images from the original group image.
[0038] Face alignment: Using the detected key point information (including the positional information of local feature points with semantic labels in the face image), the cropped face image is aligned so that the face is under a uniform pose reference in the image, thereby reducing the interference of pose changes on subsequent expression feature extraction.
[0039] Detection results with out-of-bounds bounding box coordinates, empty cropped regions, abnormal region sizes, or insufficient keypoints for face alignment can be excluded during preprocessing and not included in subsequent processing. After the above preprocessing, a group image yields several aligned face images. Further, the aligned face images are adjusted to the input size required by the feature extraction network (e.g., 224×224 pixels), and pixel normalization is performed (e.g., normalizing pixel values to the [0,1] interval or standardizing according to the mean and standard deviation of the ImageNet dataset) for input into the subsequent feature extraction network. For ease of subsequent description, let N be the number of valid faces retained in a group image after preprocessing (N is a non-negative integer; in this embodiment, when N=0, a preset fallback strategy is used, such as directly classifying based on global context features or returning the default sentiment category).
[0040] In embodiments of this invention, the face detection algorithm can employ MTCNN (Multi-task Cascaded Convolutional Networks), or RetinaFace, a YOLO-based face detection model, or other algorithms capable of face detection and keypoint localization. Keypoints refer to semantically labeled local feature points in a face image, including but not limited to eye keypoints (such as the left and right corners of the eyes and the center of the pupils), nose keypoints (such as the tip of the nose), mouth keypoints (such as the left and right corners of the mouth), and feature points in areas such as the eyebrows and jawline. Keypoints are automatically output by the face detection algorithm (e.g., MTCNN) while detecting a face, and are used for subsequent face alignment operations. Different face detection algorithms can output different numbers and locations of keypoints; for example, MTCNN typically outputs 5 keypoints (eyes, nose tip, and both corners of the mouth), while some algorithms can output 68, 98, or 106 keypoints. The basic function of keypoints is to provide spatial reference for alignment operations. As long as the goal of correcting the face pose to a unified reference system can be achieved, the specific number and distribution of keypoints do not affect the implementation of the method of this invention.
[0041] A valid face is a face sample that is determined to be usable for subsequent feature extraction after preprocessing. Specifically, a face detection result is considered a valid face if it meets the following conditions: (1) The detection box coordinates are within the image boundary, and the width and height of the box are greater than the preset minimum size threshold (e.g., 32×32 pixels) to ensure that the face region has sufficient visual information for feature extraction (the minimum size threshold can be adjusted according to the input requirements of the feature extraction network; 32×32 is only an example); (2) The cropped face region is not empty and has a normal size, and there is no empty region caused by the cropped region falling completely outside the image boundary; (3) The number of detected key points is sufficient and the confidence level is higher than the preset threshold, which can support subsequent face alignment operations (e.g., all 5 key points of MTCNN are valid, or at least three key points, namely the eyes and the tip of the nose, are required to complete basic alignment). For detection results that do not meet the above conditions, the system excludes them in the preprocessing stage and does not include them in subsequent processing. The remaining face detection results after the above screening are valid faces, and their number is denoted as N.
[0042] Step 102: Extract individual facial features from the aligned face image, extract global context features from the original group image, and map the two types of features to a feature space of the same dimension.
[0043] This step specifically includes three sub-steps: individual facial feature extraction, global contextual feature extraction, and feature projection. Individual Face Feature Extraction: The aligned face images obtained in step 101 are input into the face expression feature extraction network to obtain the individual face features corresponding to each face. In one specific implementation, the face expression feature extraction network adopts a DAN (Deep Alignment Network) face expression feature extraction network based on ResNet-18, which can output deep features with expression discrimination capabilities. If there are multiple valid faces in an image, these face images can be processed sequentially or in batches to obtain the corresponding set of individual face features. The face expression feature extraction network can also be replaced by other deep neural networks with face expression feature extraction capabilities, such as a face expression recognition network based on Visual Transformer (ViT) or other convolutional neural networks. Let the individual face features extracted by this network for the i-th face be... ,in .
[0044] Global contextual feature extraction: The original group image (i.e., the complete input image without face cropping) is input into the global visual encoding network to obtain global contextual features. This global contextual feature is used to characterize information such as scene background, distribution of people, group activity status, and overall visual atmosphere in the entire group image. In one specific implementation, the global visual coding network adopts the CLIP ViT (Contrastive Language-Image Pre-Training VisionTransformer) visual coding network, which can extract global image representations rich in semantic information. In another specific implementation, the global visual coding network can also adopt ResNet, Swin Transformer, ViT, or other visual coding networks capable of extracting global contextual features of the entire image.
[0045] Feature projection: due to individual facial features and global context features Because they originate from different sources, their feature dimensions and distributions may differ. A projection layer maps these two types of features to a feature space of unified dimension. Specifically, this maps individual facial features... Input the face projection layer to obtain face projection features. ; global context features The global projection layer is input to obtain the global projected features g. In one specific implementation, the projection layer may consist of a linear mapping layer (fully connected layer), a normalization layer (e.g., Layer Normalization or Batch Normalization), and a nonlinear activation function (e.g., ReLU or GELU). The projected features are then... It has the same feature dimensions as g (e.g., 256 or 512 dimensions) to facilitate subsequent consistency calculations and feature fusion.
[0046] Step 103: For each face, calculate the consistency score between its individual facial emotion response and the global emotion reference.
[0047] This step includes three sub-steps: individual facial emotion response generation, global emotion reference generation, and similarity calculation. Individual facial emotion response generation: The individual facial features obtained in step 102 are used to generate individual facial emotions. Input a facial emotion recognition head to obtain the individual facial emotion response of the i-th face. In one specific implementation, the facial emotion recognition head may consist of fully connected layers and / or nonlinear activation functions, used to map individual facial features into vectors or embedded representations of emotional responses. This emotional response is not required to be the final emotion classification probability, but rather a feature vector reflecting the distribution location or intensity of the expression of the face in the emotion space.
[0048] Global sentiment reference generation: Input the global projection features g obtained in step 102 into the global sentiment reference generation network to obtain the global sentiment reference. Global sentiment reference This serves as a reference point to characterize the overall emotional tendency or atmosphere perceived from a global perspective of the entire group image. In one specific implementation, the global emotion reference generation network consists of fully connected layers and nonlinear activation functions.
[0049] Similarity calculation: Calculate the individual facial emotion response for each face. Global sentiment reference The similarity between the faces is used to obtain the consistency score corresponding to the i-th face. : in, This represents the cosine similarity operation. The value range is [-1, 1]. Consistency score. This score quantifies the degree of matching between the individual emotional expression of the i-th face and the overall emotional reference of the group: a higher score indicates that the emotional response of the face is more consistent with the global emotional reference; a lower score indicates that the emotional response of the face deviates more from the global emotional reference. To facilitate subsequent weight calculation and processing, the consistency score is normalized to the [0,1] interval. in, This represents the normalized consistency score, with a value range of [0,1].
[0050] In embodiments of this invention, the consistency score is not a manually assigned face weight, nor is it determined directly based on face area, position, or detection confidence. Instead, it is automatically calculated based on the relationship between individual facial emotion responses and global emotion references. For multiple faces in the same group image, their corresponding consistency scores are obtained, forming a sequence of face consistency scores within the group image. The output of this step is the consistency score and the normalized consistency score, which are passed to a dual-channel aggregation program as the basis for dividing high-consistency channels and low-consistency channels.
[0051] It should be noted that the consistency score is not limited to cosine similarity. In other implementations, dot product similarity, Euclidean distance-based similarity measures (e.g., normalizing the distance after inverting it), learnable similarity networks, or attention mechanisms can also be used to obtain the consistency score. The above three similarity calculation methods can be used interchangeably and all fall within the scope of the present invention.
[0052] In step 104: Based on the consistency scores of all faces within the same group of images, multiple facial cues are divided into a set of high-consistency faces and a set of low-consistency faces.
[0053] In this step, the mean consistency score within the group of images is first calculated based on the consistency scores of all valid faces obtained in step 103. : Then, using this consensus mean As a threshold for dividing the current group of images, facial cues are divided into two sets: in, Represents a set of highly consistent faces. This represents a set of faces with low consistency. A set of faces with high consistency is used to represent facial cues that are relatively consistent with the global sentiment reference, while a set of faces with low consistency is used to represent facial cues that are relatively deviate from the global sentiment reference.
[0054] This partitioning method has the following characteristics: First, the partitioning threshold is automatically determined based on the consistency distribution within the current group of images, without relying on a manually set fixed threshold. Therefore, it can adapt to the differences in the distribution of consistency scores in different group images (for example, when the consistency scores of all faces in a certain group of images are high, partitioning by the mean consistency score can still preserve the relative high and low relationship. However, it should be noted that if all scores are high, the face scores in the low consistency face set may still be at an absolutely high level, but they are still considered "lower" relative to the current group of images. This reflects the design idea of "relative partitioning" rather than "absolute partitioning"). Second... and This constitutes a complete division of all valid faces in the current group image, with no overlap or omission.
[0055] It should be noted that in extreme cases, such as when N=1, the mean consistency value is the consistency score of a single face, determined according to the classification rules. Once established, this face will be included in the high-consistency face set. , If N=0 (i.e., no valid faces are detected in the group image), fallback strategies such as reporting an error, returning to the default category, or using pure global context features for classification can be adopted.
[0056] In step 105: The face projection features of the high-consistency face set and the low-consistency face set are weighted and aggregated respectively to obtain high-consistency aggregated features and low-consistency aggregated features.
[0057] This step uses dual-channel aggregation, including two independent weighted aggregation channels: High consistency channel: for a high consistency set of faces For each face in the dataset, based on its normalized consistency score... Generate face weights within the channel Then to Face projection features of all faces in the middle By performing weighted summation, highly consistent aggregated features are obtained. : in, The weight of the i-th face in the high-consistency channel is given by the following condition: In one specific implementation, It can be determined by the normalized consistency score. After Softmax normalization, we get: in, This is a temperature coefficient (which can be set to 1 or used as a learnable parameter) used to adjust the smoothness of the weight distribution.
[0058] In another specific implementation, the weights within the high consistency channel can also be generated by the weight generation network based on face projection features, global projection features, and consistency scores.
[0059] Low consistency channels: for low consistency sets of faces For each face in the dataset, based on its reverse consistency score... Generate face weights within the channel Then to Face projection features of all faces in the middle Weighted summation yields low-consistency aggregated features. : in, The weight of the i-th face in the low-consistency channel is given by the following condition: In one specific implementation, It can be obtained by normalizing the reverse consistency score using Softmax: in, This represents the temperature coefficient. The purpose of using the reverse consistency score is to indicate that within low consistency channels, the greater the deviation from the global sentiment reference (i.e., ... Smaller faces should receive a higher aggregation weight, thus increasing their chances of being considered as smaller faces. The typical characteristics of "deviation" clues are more prominently represented in the text.
[0060] It should be noted that the Softmax normalization described above is only one specific implementation of the intra-channel weight calculation in this invention and does not constitute the only limitation on the weight generation method. The weight generation for both high-consistency and low-consistency channels follows the same basic idea: the high-consistency channel assigns corresponding weights to each face based on the normalized consistency score (the larger the value, the more consistent with the global sentiment reference); the low-consistency channel assigns corresponding weights to each face based on the reverse consistency score (the larger the value, the greater the deviation from the global sentiment reference). Based on this basic idea, any implementation that generates intra-channel face weights based on consistency scores or reverse consistency scores can be used to implement the dual-channel aggregation of this invention, including but not limited to: generating weight distribution based on consistency scores using attention mechanisms (such as scaled dot product attention); adaptively calculating weights based on face projection features and consistency scores using gating networks; regressing and generating weights based on learnable multilayer perceptrons using consistency scores and / or face projection features as input; and other alternative methods that map consistency scores to normalized weights. All of the above alternative methods fall within the scope of "generating intra-channel face weights based on consistency scores" in this invention.
[0061] For boundary cases, if a certain channel set is empty (e.g.) If the channel is empty, the aggregated features corresponding to that channel can be processed using any of the following fallback strategies: (1) use the average features of all valid faces as the aggregated features of that channel; (2) use the zero vector as the aggregated features of that channel. The specific choice of the above fallback strategy does not affect the core method flow of this invention, and those skilled in the art can set it according to actual deployment needs.
[0062] Step 106: Calculate the conflict intensity based on the cosine distance between the high-consistency aggregation feature and the low-consistency aggregation feature.
[0063] In this step, the high consistency aggregation feature is calculated. With low consistency aggregation characteristics The cosine similarity between them is used to calculate the conflict intensity: Where c represents the conflict intensity, and its value ranges from [0,1]. and The more similar the cues (i.e., the closer the high-consistency cues are to the low-consistency cues in the feature space), the closer the conflict intensity is to 0; and The greater the difference, the closer the conflict intensity is to 1. Conflict intensity is used to characterize the degree of difference between two types of facial cues: when the conflict intensity is small (close to 0), it indicates that the facial cues within the group are relatively consistent, and there is no significant emotional differentiation between low-consistency faces and high-consistency faces; when the conflict intensity is large (close to 1), it indicates that there are obvious emotional differences or differentiations within the group, thus providing supplementary discriminative information for sentiment classification.
[0064] In embodiments of the present invention, the method for calculating conflict intensity is not limited to cosine distance; it can also be characterized using Euclidean distance, Manhattan distance, feature difference, bilinear interaction, multilayer perceptron, or attention interaction module. and The differences between them.
[0065] Step 107: Fuse high consistency aggregation features, low consistency aggregation features, conflict intensity and global projection features, and output the group emotion recognition result through the fusion classifier.
[0066] This step specifically includes: aggregating highly consistent features. Low consistency aggregation characteristics The conflict intensity c and the global projection feature g are concatenated to obtain the fused feature z: in,[ ] indicates a feature concatenation operation; Input the fused feature z into the fusion classifier to obtain the group emotion recognition result: Among them, C( Let y represent the fusion classifier, and y represent the output group emotion recognition result. In one specific implementation, the fusion classifier can be composed of a normalization layer (e.g., Batch Normalization or Layer Normalization), a fully connected layer, a non-linear activation function (e.g., ReLU), and an output layer in sequence. For scenarios where the emotion category is three-class (negative, neutral, and positive), the output layer is a fully connected layer with 3 neurons, followed by a Softmax activation function, outputting the predicted probabilities of the three categories, and taking the category with the highest probability as the final recognition result. If the actual application scenario requires more granular emotion categories (e.g., expanded to six basic emotion categories or more discrete categories), the dimension of the output layer can be adjusted according to the category definition. If the actual application scenario requires predicting continuous emotion dimensions (such as arousal and valence), the output layer can be adjusted to the corresponding regression output structure; this invention does not limit this.
[0067] This completes the entire processing flow from inputting a group image to outputting a group emotion recognition result. Steps 101 to 107 above together constitute the complete scheme of the group image emotion recognition method based on emotion consistency dual-channel aggregation of the present invention.
[0068] Furthermore, this method also includes step 108: model training and parameter optimization. Steps 101 to 107 above describe the processing flow of the method of the present invention in the inference stage (i.e., the trained model performs forward computation) to calculate the prediction results. Before being put into practical application, the model involved in the method of the present invention needs to be trained, and this step describes the training process.
[0069] During the training phase, group images with group sentiment labels are used as training samples. For each training sample, the prediction result is calculated following the forward processing steps 101 to 107 above. The main classification loss (e.g., using cross-entropy loss) is calculated based on the difference between the prediction result and the true label. To enhance training stability and improve the model's ability to distinguish different feature sources, one or more of the following auxiliary supervision terms can be set: High consistency auxiliary classification loss: constrains the high consistency aggregated features to have sentiment discrimination ability; Low consistency auxiliary classification loss: constrains low consistency aggregated features to retain some discriminative information, rather than being degraded as pure noise features; Global auxiliary classification loss: constrains the global projection features to capture basic group sentiment tendencies; Consistency constraint loss: For example, constraining the global emotion reference to maintain a reasonable consistency with the average representation of the emotion responses of all faces in the same group of images, in order to help the training convergence of the global emotion reference generation network.
[0070] During training, the above losses are jointly optimized according to preset weights (used to balance the relative importance of each loss term).
[0071] Regarding parameter update strategies, face feature extraction networks (such as DAN networks based on ResNet-18) and global visual encoding networks (such as CLIP ViT networks) can be completely frozen, partially fine-tuned, or fully involved in training, depending on the task requirements. Projection layers (face projection layers and global projection layers), face emotion recognition heads, global emotion reference generation networks, and fusion classifiers are usually included as trainable parts in the optimization.
[0072] After training, the parameters of the best-performing model are saved for subsequent inference. During inference, only the forward computation process described in steps 101 to 107 is executed; no further loss calculations or parameter updates are performed. This allows for the output of group emotion recognition results for any input group image. The entire inference process eliminates the need for manually specifying key faces or setting fixed consistency thresholds, achieving end-to-end automated group emotion recognition.
[0073] The following examples illustrate specific application scenarios of the method of the present invention: Example 1: Classroom Scene. Taking the group sentiment analysis of a classroom image as an example, after receiving a classroom image containing multiple students, the system first performs face detection, crops and aligns the face regions in the image, obtaining an aligned face image sequence. Then, the system extracts the facial features corresponding to each face and the global contextual features of the entire classroom image. In the consistency calculation stage, the system calculates the consistency score for each face based on the similarity between the emotional response of each face and the global sentiment reference. If the emotional responses of most students' faces are consistent with the overall classroom atmosphere, these faces are classified into the high consistency channel; if some students exhibit expressions that deviate from the overall classroom atmosphere due to individual reasons, they are classified into the low consistency channel. The system performs weighted aggregation of the facial features in the two channels to obtain high consistency aggregation features and low consistency aggregation features, and calculates the conflict intensity between them. Finally, the system fuses the high consistency aggregation features, low consistency aggregation features, conflict intensity, and global projection features, and outputs the group sentiment recognition result for this classroom scene through a fusion classifier, such as judging the overall classroom atmosphere as "positive," "neutral," or "negative."
[0074] Example 2: Social Gathering Scene. Taking the sentiment analysis of a social gathering image as an example, the system receives an image of a gathering with multiple people, detects and aligns the facial regions in the image, and extracts individual facial features and global contextual features. If most individuals exhibit positive emotions, while a few individuals have weaker expressions or are not entirely consistent with the overall atmosphere, the system uses consistency scores to aggregate faces that are more consistent with the global sentiment reference into high-consistency features, and relatively deviating faces into low-consistency features. The system also uses conflict intensity to describe the degree of difference between the two types of cues. Finally, the system fuses the dual-channel (high-consistency channel and low-consistency channel) facial aggregation features, conflict intensity, and global projection features to output the overall sentiment category of the gathering image.
[0075] The two examples above, corresponding to classroom and social gathering scenarios respectively, are intended to illustrate the application of the present invention in different multi-person scenarios and are not intended to limit the scope of application of the present invention. Those skilled in the art can apply the present invention to other group sentiment analysis tasks involving multiple people, such as surveillance scenarios, meeting scenarios, and sports event scenarios, according to actual needs.
[0076] The group image emotion recognition method based on dual-channel aggregation of emotion consistency provided by this invention has the following beneficial effects: (1) By calculating the similarity between individual facial emotion responses and global emotion references, a consistency score is obtained. Based on this score, facial cues are divided into high consistency channels and low consistency channels, and then weighted and aggregated separately. This method can effectively distinguish between facial cues that are relatively consistent with the global emotion reference and those that are relatively different, avoiding the simple mixing of dominant emotion cues, individual emotion deviation cues, and low-quality noise cues into a single representation, thereby improving the accuracy of group emotion recognition.
[0077] (2) Facial cues that deviate from the global emotion reference in the low consistency channel are retained and aggregated separately, so that the emotion difference information within the group is retained instead of being directly discarded; at the same time, by calculating the conflict intensity between the high consistency aggregated features and the low consistency aggregated features, supplementary discriminative information is provided for the final emotion classification, further enhancing the model's discriminative ability.
[0078] (3) Facial features caused by individual occlusion, pose shift, local blurring, or individual expression differences are classified into the low consistency channel so that they no longer interfere with the aggregation of dominant cues in the high consistency channel; at the same time, their difference information is retained through conflict intensity to participate in the final classification. The above methods reduce the interference of local noise and individual bias on the overall recognition results and improve the stability and robustness of group emotion recognition in complex multi-person scenarios.
[0079] (4) Calculate the mean consistency score based on the consistency scores of all faces within the same group of images, and use this mean score as the basis for dividing high consistency channels and low consistency channels. This division method is automatically determined by the consistency distribution within the current image, without relying on a fixed threshold set manually. It can adapt to the distribution differences of consistency scores in different scenarios and has good scene generalization ability.
[0080] (5) Simultaneously extract individual facial features and global contextual features. In the final classification stage, high consistency aggregation features, low consistency aggregation features, conflict intensity and global projection features are used together in the decision-making. This can utilize both the fine expression information of the facial region and the global visual information such as scene background and crowd distribution, thus achieving an organic combination of local fine features and global scene information.
[0081] (6) By using intermediate outputs such as consistency scores, dual-channel partitioning and conflict intensity, the model’s decision-making process has a clear logical chain, which enhances the interpretability of the model’s decision-making process.
[0082] Example 2 In embodiments of the present invention, a group image emotion recognition system with emotion consistency dual-channel aggregation is also provided, for performing the steps of the group image emotion recognition method with emotion consistency dual-channel aggregation described in Embodiment 1.
[0083] Figure 2 This is a schematic diagram of the structure of a group image emotion recognition system based on dual-channel aggregation of emotion consistency according to an embodiment of the present invention, as shown below. Figure 2 As shown, the system of the present invention includes a face preprocessing module 201, a feature extraction module 202, a consistency calculation module 203, a dual-channel aggregation module 204, a conflict modeling module 205, and an emotion classification module 206 connected in sequence. The data flow path between the modules is as follows: after the group image is processed by the face preprocessing module 201, it passes sequentially through the feature extraction module 202, the consistency calculation module 203, the dual-channel aggregation module 204, the conflict modeling module 205, and the emotion classification module 206, finally outputting the group emotion recognition result. Specifically, the output of the feature extraction module 202 is bypassed to the consistency calculation module 203, the dual-channel aggregation module 204, and the emotion classification module 206; the output of the consistency calculation module 203 is transmitted to the dual-channel aggregation module 204; the output of the dual-channel aggregation module 204 is transmitted to the conflict modeling module 205 and the emotion classification module 206; and the output of the conflict modeling module 205 is transmitted to the emotion classification module 206. The following will combine... Figure 2 The specific implementation methods of each module are explained.
[0084] The face preprocessing module 201 receives a group image to be identified, detects, crops, and aligns the face regions within it, and outputs an aligned face image sequence to the feature extraction module 202. The input to this module is a group image (a static image or a single frame from a video stream). Internally, the module calls a face detection algorithm to obtain the bounding boxes and keypoint information for each face region; it crops the corresponding face regions from the original image based on the bounding boxes and aligns the cropped face images using the keypoints. Detection results with out-of-bounds bounding box coordinates, empty cropped regions, abnormal region sizes, or insufficient keypoints for face alignment are excluded; the remaining detection results are considered valid faces. After preprocessing, a group image yields several aligned face images, which are adjusted to the input size and pixel normalization format required by the feature extraction module 202. The output is an aligned face image sequence, which is then sent to the feature extraction module 202.
[0085] The feature extraction module 202 extracts individual facial features and global contextual features respectively, and maps the two types of features to a feature space of the same dimension, outputting facial projection features and global projection features. Specifically, the input of the feature extraction module 202 includes two paths: the aligned facial image sequence from the facial preprocessing module 201, and the original group image. For the facial image sequence, the module calls the facial expression feature extraction network to extract the individual facial features corresponding to each face in each image; for the original group image, the module calls the global visual encoding network to extract global contextual features. The individual facial features are mapped by the facial projection layer to obtain facial projection features, and the global contextual features are mapped by the global projection layer to obtain global projection features, both of which have the same feature dimension. This module outputs the facial projection features to the consistency calculation module 203 and the dual-channel aggregation module 204; and outputs the global projection features to the consistency calculation module 203 and the sentiment classification module 206.
[0086] The consistency calculation module 203 calculates the similarity between the individual facial emotion response and the global emotion reference for each face, and outputs the consistency score and normalized consistency score for each face to the dual-channel aggregation module 204. The inputs to the consistency calculation module 203 include facial projection features and global projection features from the feature extraction module 202. The facial projection features are processed by a facial emotion recognition head to obtain the individual facial emotion response, and the global projection features are processed by a global emotion reference generation network to obtain the global emotion reference. The consistency calculation module 203 calculates the similarity between the two as the consistency score. Finally, it outputs the consistency score and normalized consistency score for each face and sends them to the dual-channel aggregation module 204.
[0087] The dual-channel aggregation module 204 is used to divide multiple facial cues into a high-consistency face set and a low-consistency face set based on the consistency score. It performs weighted aggregation of the facial projection features in each set, outputting high-consistency and low-consistency aggregated features to the conflict modeling module 205 and the sentiment classification module 206. Its inputs include the consistency score (and normalized consistency score) from the consistency calculation module 203 and the facial projection features from the feature extraction module 202. This module uses the average consistency of all faces within the current group image as the partitioning threshold to divide facial cues into high-consistency and low-consistency face sets. After channel partitioning, the module performs weighted aggregation of the facial projection features in each channel: the high-consistency channel generates face weights based on the normalized consistency score, and the low-consistency channel generates face weights based on the reverse consistency score. These weighted sums are then used to obtain the high-consistency and low-consistency aggregated features. If a channel set is empty, the module uses the average feature or zero vector of all valid faces as the aggregation feature for that channel.
[0088] The conflict modeling module 205 is used to calculate the difference between high-consistency aggregated features and low-consistency aggregated features, and outputs the conflict intensity to the sentiment classification module 206.
[0089] The sentiment classification module 206 is used to fuse high-consistency aggregation features, low-consistency aggregation features, conflict intensity, and global projection features to output a group sentiment recognition result. The inputs to the sentiment classification module 206 include: high-consistency and low-consistency aggregation features from the dual-channel aggregation module 204, conflict intensity from the conflict modeling module 205, and global projection features from the feature extraction module 202. These inputs are concatenated into a fusion feature, which is then input into the fusion classifier to obtain the group sentiment recognition result. This group sentiment recognition result can be used for further display, storage, or transfer to other business systems.
[0090] During the training phase, the system uses group images with sentiment labels as training samples. After performing calculations according to the aforementioned forward process, it calculates the main classification loss (e.g., using cross-entropy loss) based on the difference between the predicted results and the true labels, and updates the parameters. To enhance training stability, the system can also set one or more of the following auxiliary supervision terms: high consistency auxiliary classification loss, low consistency auxiliary classification loss, global auxiliary classification loss, and consistency constraint loss. These losses can be jointly optimized according to preset weights.
[0091] During the inference phase, the system receives the group image to be identified, performs forward computation sequentially according to the modules described above, and finally outputs the group emotion recognition result. The entire inference process does not require manual specification of key faces or manual setting of fixed consistency thresholds.
[0092] This system embodiment is based on the exact same inventive concept as the aforementioned method embodiment 1, and the functions performed by each module correspond to the steps in the method embodiment. The system can be implemented by software, hardware, firmware, or any combination thereof, and its specific implementation does not constitute a limitation on the present invention.
[0093] Example 3 In embodiments of the present invention, an electronic device is also provided. Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention, such as... Figure 3 As shown, the electronic device of the present invention includes a processor 301 and a memory 302, wherein, The memory 302 stores a computer program, which, when read and executed by the processor 301, performs the steps described above in the embodiment of the group image emotion recognition method with dual-channel aggregation of emotion consistency.
[0094] Example 4 In embodiments of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, wherein the computer program is configured to execute the steps in the embodiments of the group image emotion recognition method with emotion consistency dual-channel aggregation as described above when running.
[0095] In this embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0096] It will be understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A group image emotion recognition method based on dual-channel aggregation of emotion consistency, characterized in that, Includes the following steps: Face preprocessing is performed on the group image to be identified to obtain face images; Extract individual facial features from each face in the face image and global context features from the group image. Map the individual facial features and global context features to a feature space of a unified dimension to obtain face projection features and global projection features. An individual facial emotion response is generated based on the individual facial features of each face, a global emotion reference is generated based on the global projection features, and a consistency score between the individual facial emotion response and the global emotion reference is calculated. All faces within the same group of images are divided into a high-consistency face set and a low-consistency face set according to their consistency scores. The face projection features in each set are then weighted and aggregated to obtain high-consistency aggregated features and low-consistency aggregated features. The difference between the highly consistent aggregation feature and the low consistent aggregation feature is calculated to obtain the conflict intensity; By integrating the high-consistency aggregation features, the low-consistency aggregation features, the conflict intensity, and the global projection features, a classifier is used to output the group sentiment recognition result.
2. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, The step of performing face preprocessing on the group image to be identified to obtain face images further includes: Face detection is performed on the group image to obtain the bounding boxes and key point information of each face region; the key point information includes the location information of local feature points with semantic identifiers in the face image; The corresponding face region is cropped from the group image based on the location box to obtain a face image; The cropped face image is aligned using the key point information.
3. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, The step of calculating the consistency score between the individual facial emotion response and the global emotion reference further includes calculating using cosine similarity: in, This represents the consistency score corresponding to the i-th face. This represents the cosine similarity operation. Indicates an individual's facial emotional response. This indicates a general sentiment reference. The consistency score is normalized to the [0,1] interval to obtain the normalized consistency score.
4. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, The step of dividing all faces within the same group of images into a set of high-consistency faces and a set of low-consistency faces based on their consistency scores further includes: Calculate the mean consistency score based on the consistency scores of all faces within the same group of images: in, The mean value is the consistency value, and N is the number of faces in the same group of images, where N is greater than or equal to 1. Let be the consistency score corresponding to the i-th face; Faces with a consistency score greater than or equal to the mean consistency value are classified into a high consistency face set, and faces with a consistency score less than the mean consistency value are classified into a low consistency face set.
5. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, The face is divided into a high-consistency face set and a low-consistency face set. The face projection features in each set are weighted and aggregated to form a dual-channel aggregation including a high-consistency channel and a low-consistency channel.
6. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 5, characterized in that, In the high-consistency channel, the face projection features in the high-consistency face set are weighted and summed according to a first weight to obtain the high-consistency aggregated features. The formula is as follows: Among them, the first weight This represents the weight of the i-th face in the high-consistency channel. Represents a set of highly consistent faces. Represents the face projection features of the i-th face; In the low-consistency channel, the face projection features in the low-consistency face set are weighted and summed according to the second weight to obtain the low-consistency aggregated features. The formula is as follows: Among them, the second weight This represents the weight of the i-th face in the low-consistency channel. This represents a set of faces with low consistency. Let i represent the facial projection features of the i-th face.
7. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 6, characterized in that, The high consistency channel assigns a corresponding first weight to each face based on the normalized consistency score; the normalized consistency score is obtained by normalizing the calculated consistency score between the individual face emotion response and the global emotion reference to the [0,1] interval. The low consistency channel assigns a corresponding second weight to each face based on the reverse consistency score; the reverse consistency score is equal to 1 minus the normalized consistency score.
8. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, The conflict intensity is calculated using cosine distance: Where c represents the intensity of the conflict; This represents the cosine similarity operation; This indicates a high degree of consistency in aggregation characteristics; This indicates a low-consistency aggregation characteristic.
9. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, The step of fusing the high-consistency aggregation feature, the low-consistency aggregation feature, the conflict intensity, and the global projection feature, and outputting the group emotion recognition result through a classifier, further includes: concatenating the high-consistency aggregation feature, the low-consistency aggregation feature, the conflict intensity, and the global projection feature. ; in, Indicates fusion characteristics, This indicates a high degree of consistency in aggregation features. represents low-consistency aggregation features, c represents conflict intensity, and g represents global projection features; The fused features are input into the classifier to obtain the group emotion recognition result.
10. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, If the set of highly consistent faces or the set of low-consistency faces is empty, then the average feature or zero vector of all faces is used as the aggregate feature of the set.
11. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 1, characterized in that, It also includes a training step: using group images with group sentiment labels as training samples, calculating the main classification loss based on the difference between the predicted results and the true labels, and updating the parameters.
12. The group image emotion recognition method based on dual-channel aggregation of emotion consistency according to claim 11, characterized in that, The training steps also include setting one or more of the following auxiliary supervision items: High consistency auxiliary classification loss constrains high consistency aggregated features to have sentiment discrimination ability; Low consistency auxiliary classification loss constrains low consistency aggregated features to retain discriminative information; Global auxiliary classification loss constrains global projection features to capture group sentiment tendencies; The consistency constraint loss constrains the consistency relationship between the global emotion reference and the average representation of the emotion responses of all faces within the same group of images.
13. A group image emotion recognition system based on dual-channel aggregation of emotion consistency, characterized in that, A group image emotion recognition method for implementing the emotion consistency dual-channel aggregation as described in any one of claims 1 to 12 includes: The face preprocessing module is used to preprocess the faces of the group images to be identified, and obtain face images; The feature extraction module is used to extract individual facial features of each face from the face image, extract global context features from the group image, and map the individual facial features and the global context features to a feature space of a unified dimension to obtain face projection features and global projection features. The consistency calculation module is used to generate an individual facial emotion response based on the individual facial features, generate a global emotion reference based on the global projection features, and calculate the consistency score between the individual facial emotion response and the global emotion reference. The dual-channel aggregation module is used to divide the faces into a high-consistency face set and a low-consistency face set based on the consistency scores of all faces in the same group of images, and to perform weighted aggregation of the face projection features in each set to obtain high-consistency aggregation features and low-consistency aggregation features. The conflict modeling module is used to calculate the difference between the high-consistency aggregation feature and the low-consistency aggregation feature to obtain the conflict intensity; The emotion classification module is used to fuse the high consistency aggregation features, the low consistency aggregation features, the conflict intensity, and the global projection features, and output the group emotion recognition result through a classifier.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor is configured to execute the computer program stored in the memory to implement the group image emotion recognition method based on dual-channel aggregation of emotion consistency as described in any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which is loaded and executed by a processor to implement the group image emotion recognition method based on dual-channel aggregation of emotion consistency as described in any one of claims 1 to 12.