Emotion recognition method based on hierarchical feature fusion in digital image

Through the methods of hierarchical feature fusion and spatial alignment position embedding, top-down interaction paths and bottom-up inference paths are constructed, which solves the problem of insufficient accuracy in the existing emotion recognition methods and achieves a more efficient emotion recognition effect.

CN120375480APending Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340350.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing emotion recognition methods have insufficient accuracy in model recognition, and it is difficult to fully capture subtle changes in emotions. In dealing with data imbalance problem, the recognition results are not ideal.

Method used

The emotion recognition framework of hierarchical feature fusion is adopted to extract facial, body and background features through backbone networks, and feature interaction and reasoning is used to use the cross attention mechanism and self-attention module for feature interaction and inference. Combining the spatially aligned position embedding module and dynamic scaling factor to improve the focus loss function, building a top-down interaction path and bottom-up inference path.

Benefits of technology

It significantly improves the model performance in multi-category emotion recognition tasks, solves the problem of data imbalance, and improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375480A_ABST
    Figure CN120375480A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion recognition method based on hierarchical feature fusion in a digital image. The method comprises the steps that training set images are acquired, each training image is a digital image, and the digital image comprises a plurality of users; each training image is processed, three target windows of each training image are obtained, and the three target windows are the face window, the body window and the background window of each user; the three target windows are input into a backbone network in parallel, facial features, body features and background features of each user are obtained, and the backbone network comprises a facial feature extractor, a body feature extractor and a background feature extractor; a first cross-window interaction layer and a second cross-window interaction layer are constructed, and the first cross-window interaction layer and the second cross-window interaction layer comprise a cross attention mechanism, a feedforward network and jump connection. According to the invention, the technical problem of low recognition accuracy of a model for emotion recognition in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition, and more particularly, to an emotion recognition method for hierarchical feature fusion in digital images. Background Art

[0002] Emotion recognition technology has been widely applied in fields such as human-computer interaction, monitoring, augmented reality, and robotics. Its core goal is to perceive and understand the emotional attitudes of others. The importance of this technology in daily life is obvious, because the perceived emotions can not only directly affect an individual's behavioral decisions, but also indirectly affect the quality and effectiveness of interpersonal interactions. Traditional emotion recognition research mainly focused on the analysis of facial expressions. With the progress of technology, researchers have begun to explore the associations between emotions and other modalities such as gait and speech signals. In recent years, context-based emotion recognition models have gradually become a research hotspot. Such models significantly improve the accuracy of emotion recognition by integrating multimodal data such as body postures, facial expressions, background information, and depth maps in images.

[0003] However, despite the significant progress made in these studies, existing emotion recognition methods still have many deficiencies. (1) Emotion itself is a highly abstract and complex psychological and physiological state. It is difficult to comprehensively capture the subtle changes in emotions relying solely on single local information (such as facial expressions); research shows that humans usually combine multiple pieces of information for deductive reasoning when perceiving emotions, while most existing technologies only process data at different scales through simple feature splicing or element-wise fusion, lacking in-depth modeling of the reasoning process; (2) The imbalanced distribution of emotion data further exacerbates the training difficulty of the model, making it difficult for deep models to extract discriminative feature representations; when dealing with the problem of data imbalance, most existing emotion recognition frameworks fail to effectively model the correlations and coexistences between emotions and do not fully exploit the intrinsic structural information of the data, thus limiting the generalization ability of the model and resulting in less than ideal recognition results. Summary of the Invention

[0004] Embodiments of the present invention provide an emotion recognition method for hierarchical feature fusion in digital images to at least solve the technical problem of low recognition accuracy of the model for emotion recognition in the prior art.

[0005] According to one aspect of the embodiments of the present invention, there is provided an emotion recognition method for hierarchical feature fusion in digital images. The method may include: obtaining training set images, where each training image is a digital image and includes multiple users; processing each training image to obtain three target windows for each training image, where the three target windows are respectively the facial window, body window, and background window of each user; parallelly inputting the three target windows into a backbone network to obtain the facial features, body features, and background features of each user, where the backbone network includes a facial feature extractor, a body feature extractor, and a background feature extractor; constructing a first cross-window interaction layer and a second cross-window interaction layer, where the first cross-window interaction layer and the second cross-window interaction layer include a cross-attention mechanism, a feed-forward network, and a skip connection; inputting the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features; inputting the body features and facial features of each user into the second cross-window interaction layer to obtain optimized facial features; encoding the position information of the background window of each user in each training image to obtain a background position embedding; encoding the positions of the facial window and body window of each user in each training image to obtain the position information of the facial window and body window of each user in each training image; obtaining a facial position embedding based on the background position embedding and the position information of the facial window of each user in each training image; obtaining a body position embedding based on the background position embedding and the position information of the body window of each user in each training image; introducing a set of learnable emotion prototypes; constructing a first cross-window inference layer, a second cross-window inference layer, and a third cross-window inference layer, where each cross-window inference layer includes a self-attention module and a cross-attention module; inputting the set of learnable emotion prototypes, the optimized facial features, and the facial position embedding into the first cross-window inference layer to obtain the emotion features of the optimized facial features; inputting the emotion features of the optimized facial features, the optimized body features, and the body position embedding into the second cross-window inference layer to obtain the emotion features of the optimized body features; inputting the emotion features of the optimized body features, the background features, and the background position embedding into the third cross-window inference layer to obtain the emotion category of each user. Wherein, after obtaining the emotion category of each user, in each batch-size during the training process of the training set images, based on the target loss function, update the hyperparameters of the model that processes the training set images during the training process. When the target loss function is minimized, the model training ends.

[0006] Optionally, the processing each training image to obtain three target windows for each training image includes: cropping each training image using a pre-detected bounding box to obtain three different initial windows of each cropped training image; adjusting the three different initial windows to obtain the target window corresponding to each initial window.

[0007] Optionally, the expression for inputting the three target windows into the backbone network in parallel to obtain the facial features, body features, and background features of each user is as follows:

[0008]

[0009] where f F is the facial feature of each user, f B is the body feature of each user, f C is the background feature of each user, is the backbone network, F is the face of each user, B is the body of each user, C is the background of each user, θ1 is the parameter of the facial feature extractor, θ2 is the parameter of the body feature extractor, and θ3 is the parameter of the background feature extractor.

[0010] Optionally, inputting the body feature and background feature of each user into the first cross-window interaction layer to obtain the optimized body feature includes: inputting the body feature and background feature of each user into the cross-attention mechanism to obtain the enlarged body feature of each user; inputting the enlarged body feature of each user into the feed-forward network and connecting it with the body feature through a skip connection to obtain the optimized body feature.

[0011] Optionally, inputting the body feature and facial feature of each user into the second cross-window interaction layer to obtain the optimized facial feature includes: inputting the body feature and facial feature of each user into the cross-attention mechanism to obtain the enlarged facial feature of each user; inputting the enlarged facial feature of each user into the feed-forward network and connecting it with the facial feature through a skip connection to obtain the optimized facial feature.

[0012] Optionally, inputting a set of learnable emotion prototypes, the optimized facial feature, and the facial position embedding into the first cross-window inference layer to obtain the emotion feature of the optimized facial feature includes: processing a set of learnable emotion prototypes through the self-attention module to obtain the emotion prototype feature; processing the emotion prototype feature, the optimized facial feature, and the facial position embedding through the cross-attention module to obtain the emotion feature of the optimized facial feature.

[0013] Optionally, inputting the emotion feature of the optimized facial feature, the optimized body feature, and the body position embedding into the second cross-window inference layer to obtain the emotion feature of the optimized body feature includes: processing the emotion feature of the optimized facial feature through the self-attention module to obtain the facial emotion feature; processing the facial emotion feature, the optimized body feature, and the body position embedding through the cross-attention module to obtain the emotion feature of the optimized body feature.

[0014] Optionally, embedding the emotional features, background features, and background positions of the optimized body features into the input third cross-window inference layer to obtain the emotion category of each user, including: processing the emotional features of the optimized body features through a self-attention module to obtain body emotional features; processing the body emotional features, background features, and background position embeddings through a cross-attention module to obtain the emotion category of each user.

[0015] Optionally, the expression of the target loss function is:

[0016]

[0017] where is the target loss function, α t is the adjustment factor, t is a marker indicating whether the emotion category predicted by the model for each user is consistent with the true label of the emotion category of each user, is the true label of the emotion category of each user, α is the first hyperparameter, σ represents the sigmoid function, c is a constant, β is the second hyperparameter, γ is the third hyperparameter, p t is the probability that the predicted emotion category of each user is the correct category, and r is the fourth hyperparameter.

[0018] Advantages of the present invention:

[0019] (1) The present invention proposes a hierarchical feature fusion emotion recognition framework. This framework uses the Transformer decoder layer to construct a top-down interaction path and a bottom-up inference path, deeply models the inference process, and significantly improves the model performance in multi-class emotion recognition tasks.

[0020] (2) The present invention proposes a spatial alignment position embedding module to better align the position information between hierarchical features, thereby more effectively promoting interaction and inference; in addition, the focal loss function is improved by introducing a dynamic scaling factor to better solve the problem of multi-class emotion data imbalance. Description of the Drawings

[0021] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0022] Figure 1 is a flowchart of a method for emotion recognition with hierarchical feature fusion in a digital image according to an embodiment of the present invention;

[0023] Figure 2 is a framework diagram of a method for emotion recognition with hierarchical feature fusion according to an embodiment of the present invention;

[0024] Figure 3 Schematic diagram of the spatial alignment position embedding module according to an embodiment of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] Embodiment 1

[0028] According to an embodiment of the present invention, there is provided an emotion recognition method for hierarchical feature fusion in digital images. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system including at least one set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that here.

[0029] Figure 1 Is a flowchart of an emotion recognition method for hierarchical feature fusion in digital images according to an embodiment of the present invention, as Figure 1 shown, the method may include the following steps:

[0030] Step S101, obtain training set images, process each training image to obtain three target windows of each training image, where each training image is a digital image, and the digital image includes multiple users, and the three target windows are the facial window, body window, and background window of each user respectively.

[0031] In the technical solution provided in step S101 of the present invention, a training set of images is obtained, where each training image is a digital image, and each digital image includes multiple users. The size of each training image is H is the height of each training image, W is the width of each training image, and 3 is the number of channels of each training image; each training image is cropped and adjusted to obtain three target windows of each training image, and the three target windows are the face window, body window, and background window of each user respectively.

[0032] Step S102, input the three target windows into the backbone network in parallel to obtain the face features, body features, and background features of each user, where the backbone network includes a face feature extractor, a body feature extractor, and a background feature extractor.

[0033] In the technical solution provided in step S102 of the present invention, the face window is subjected to feature extraction through the face feature extractor to obtain the face features of each user; the body window is subjected to feature extraction through the body feature extractor to obtain the body features of each user; the background window is subjected to feature extraction through the background feature extractor to obtain the background features of each user.

[0034] Step S103, construct a first cross-window interaction layer and a second cross-window interaction layer, input the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features; input the body features and face features of each user into the second cross-window interaction layer to obtain optimized face features, where the first cross-window interaction layer and the second cross-window interaction layer include a cross-attention mechanism, a feed-forward network, and a skip connection.

[0035] In the technical solution provided in step S103 of the present invention, Figure 2 is a framework diagram of an emotion recognition method for hierarchical feature fusion according to an embodiment of the present invention, Figure 2 There are two cross-window interaction layers in it, and from Figure 2 from top to bottom are the first cross-window interaction layer and the second cross-window interaction layer respectively.

[0036] From Figure 2 it can be seen that the three rectangular boxes behind the image feature extractor are, from top to bottom, the background features, body features, and face features of each user respectively. Input the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features, and input the body features and face features of each user into the second cross-window interaction layer to obtain optimized face features.

[0037] Step S104: Encode the position information of the background window for each user in each training image to obtain a background position embedding, and encode the positions of the face window and body window for each user in each training image to obtain the position information of the face window and the position information of the body window for each user in each training image.

[0038] In the technical solution provided in step S104 of the present invention above, the position embedding module plays a key role in the transformer, enabling the model to perceive the position of features during self-attention or cross-attention; in the framework proposed by the present invention, the position embedding module is not only because of the application of the transformer structure, but also because there may be multiple users in the same environment; when interacting with the context, the foreground and background should be distinguished to avoid information confusion; for this reason, the present invention proposes a spatially aligned position embedding for the face, body, and background three-level windows of each user. Figure 3 It is a schematic diagram of the spatially aligned position embedding module according to an embodiment of the present invention, as Figure 3 shown. Considering that these three-level windows are all cropped from the input image, therefore, first encode the position information of the entire image (i.e., the background window); this is achieved through a widely used learnable position embedding method, and the formula is as follows:

[0039]

[0040] f col = MLP((0, 1, …, W - 1); θ col )

[0041] f row = MLP((0, 1, …, H - 1); θ row )

[0042] where, is the background position embedding, MLP represents a multi-layer perceptron defined by parameter θ, θ col is the parameter of the output row, θ row is the parameter of the output column, f col is the position of the background window column, f col is the position of the background window row, [·; ·] represents the broadcast connection operation.

[0043] Encode the positions of the face window and body window to obtain the position information of the face window (g F ) and the position information of the body window (g B ).

[0044] Step S105: Based on the background position embedding and the position information of the face window of each user in each training image, obtain the face position embedding; based on the background position embedding and the position information of the body window of each user in each training image, obtain the body position embedding.

[0045] In the technical solution provided in step S105 of the present invention above, the expression of the face position embedding is:

[0046]

[0047] where is the face position embedding of each user in each training image, BI represents bilinear interpolation, and g F is the position information of the face window of each user in each training image. Through this simple and effective design, the features in the hierarchical context can be well aligned, thus realizing cross-attention interaction and reasoning.

[0048] The expression of the body position embedding is:

[0049]

[0050] where is the body position embedding of each user in each training image, and g B is the position information of the body window of each user in each training image.

[0051] Step S106: Introduce a set of learnable emotion prototypes, and construct the first cross-window inference layer, the second cross-window inference layer, and the third cross-window inference layer. Each cross-window inference layer includes a self-attention module and a cross-attention module.

[0052] In the technical solution provided in step S106 of the present invention above, a set of learnable emotion prototypes adapts to capture the clues of various emotions through the supervision signal. Its core idea is to infer emotions starting from the existing knowledge (learned prototypes), rather than relying on possibly inaccurate observation information.

[0053] It can be seen from Figure 2 that there are 3 cross-window inference layers, which are the first cross-window inference layer, the second cross-window inference layer, and the third cross-window inference layer from bottom to top.

[0054] Step S107: Embed a set of learnable emotion prototypes, optimized facial features, and facial positions into the first cross-window inference layer to obtain the emotion features of the optimized facial features; embed the emotion features of the optimized facial features, optimized body features, and body positions into the second cross-window inference layer to obtain the emotion features of the optimized body features; embed the emotion features of the optimized body features, background features, and background positions into the third cross-window inference layer to obtain the emotion category of each user. Among them, after obtaining the emotion category of each user, in each batch-size during the training process of the training set images, based on the target loss function, update the hyperparameters of the model that processes the training set images during the training process. When the target loss function is minimized, the model training ends.

[0055] In the technical solution provided in step S107 of the present invention, the first cross-window inference layer processes the embedding of a set of learnable emotion prototypes, optimized facial features, and facial positions to obtain the emotion features of the optimized facial features.

[0056] The second cross-window inference layer processes the embedding of the emotion features of the optimized facial features, optimized body features, and body positions to obtain the emotion features of the optimized body features.

[0057] The third cross-window inference layer processes the embedding of the emotion features of the optimized body features, background features, and background positions to obtain the emotion category of each user. In each batch-size during the training process of the training set images, according to the target loss function, update the hyperparameters of the backbone network, the first cross-window interaction layer, the second cross-window interaction layer, the first cross-window inference layer, the second cross-window inference layer, and the third cross-window inference layer. When the target loss function is minimized, the model training ends.

[0058] The above method of this embodiment will be further introduced below.

[0059] As an optional embodiment, in step S101, the processing of each training image to obtain three target windows for each training image includes: cropping each training image using a pre-detected bounding box to obtain three different initial windows of each cropped training image; adjusting the three different initial windows to obtain the target window corresponding to each initial window.

[0060] In this embodiment, each training image is cropped using a pre-detected bounding box to obtain three different initial windows of each cropped training image; the three different initial windows are adjusted to obtain the target window corresponding to each initial window.

[0061] As an alternative embodiment, in step S102, the three target windows are input into the backbone network in parallel, and the expressions for the facial features, body features, and background features of each user are obtained as follows:

[0062]

[0063] where f F is the facial feature of each user, f B is the body feature of each user, f C is the background feature of each user, is the backbone network, F is the facial window of each user, B is the body window of each user, C is the background window of each user, θ1 is the parameter of the facial feature extractor, θ2 is the parameter of the body feature extractor, and θ3 is the parameter of the background feature extractor.

[0064] In this embodiment, the facial feature extractor extracts features from the facial window to obtain the facial features of each user; the body feature extractor extracts features from the body window to obtain the body features of each user; and the background feature extractor extracts features from the background window to obtain the background features of each user.

[0065] As an alternative embodiment, in step S103, inputting the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features includes: inputting the body features and background features of each user into a cross-attention mechanism to obtain the enlarged body features of each user; and inputting the enlarged body features of each user into a feed-forward network and connecting them to the body features through a skip connection to obtain the optimized body features.

[0066] In this embodiment, in the first cross-window interaction layer, the background and body features are first projected into three latent spaces, which are respectively represented as query, key, and value:

[0067] Q = f B W1, K = f C w2, V = f C W3

[0068] where Q is the query, K is the key, and V is the value, are the parameters of the linear projection. Subsequently, the cross-attention weights are calculated based on the dot product similarity between the key and the query. These attention weights can be regarded as the semantic relationship between the background features and the body features. d is the number of channels, and d2 is the number of samples. is the latent space size.

[0069] Therefore, as shown in the following formula:

[0070]

[0071] Among them, δ i,j is the cross-attention weight, where i and j represent the abscissa and ordinate of the elements in the background feature and the body feature, and all δ i,j are integrated into a matrix by position to obtain δ.

[0072] The values (value) are weighted and summed according to the attention weights, so as to enrich the body features from a broader perspective. The expression of the expanded body feature for each user is:

[0073]

[0074] Among them, K T is the transpose of K, h is the height, w is the width, f B1 is the expanded body feature for each user, and d is the number of channels.

[0075] A standard feed-forward network composed of two fully connected layers is added, and combined with the skip connection mechanism to generate the optimized body feature. The formula is as follows:

[0076]

[0077] Among them, FFN is the standard feed-forward network, is the optimized body feature.

[0078] As an alternative embodiment, in step S103, the inputting of the body feature and the facial feature of each user into the second cross-window interaction layer to obtain the optimized facial feature includes: inputting the body feature and the facial feature of each user into the cross-attention mechanism to obtain the expanded facial feature of each user; inputting the expanded facial feature of each user into the feed-forward network and connecting it with the facial feature through the skip connection to obtain the optimized facial feature.

[0079] In this embodiment, the process of inputting the body feature and the facial feature of each user into the cross-attention mechanism to obtain the expanded facial feature of each user is the same as the process of inputting the body feature and the background feature of each user into the cross-attention mechanism to obtain the expanded body feature of each user. The expression of inputting the expanded facial feature of each user into the feed-forward network and connecting it with the facial feature through the skip connection to obtain the optimized facial feature is:

[0080]

[0081] Among them, is the optimized facial feature, f F1The enlarged facial features for each user;

[0082] Since the background window has no relative global information, the identity mapping is directly used, and the optimized background feature is

[0083] As an optional embodiment, in step S107, the step of inputting a set of learnable emotion prototypes, the optimized facial features, and the facial position embedding into the first cross-window inference layer to obtain the emotion features of the optimized facial features includes: processing a set of learnable emotion prototypes through a self-attention module to obtain emotion prototype features; processing the emotion prototype features, the optimized facial features, and the facial position embedding through a cross-attention module to obtain the emotion features of the optimized facial features.

[0084] In this embodiment, the expression of a set of learnable emotion prototypes is: where n is the number of emotion types and d is the number of channels.

[0085] The expression for obtaining the emotion prototype features by processing a set of learnable emotion prototypes through a self-attention module is:

[0086]

[0087] where SA(·) represents the self-attention module, is the emotion prototype feature.

[0088] The expression for the emotion features of the optimized facial features is:

[0089]

[0090] where is the emotion feature of the optimized facial features, and CA(·) is the cross-attention module.

[0091] As an optional embodiment, in step S107, the step of inputting the emotion features of the optimized facial features, the optimized body features, and the body position embedding into the second cross-window inference layer to obtain the emotion features of the optimized body features includes: processing the emotion features of the optimized facial features through a self-attention module to obtain facial emotion features; processing the facial emotion features, the optimized body features, and the body position embedding through a cross-attention module to obtain the emotion features of the optimized body features.

[0092] In this embodiment, the process of processing the emotional features of the optimized facial features through the self-attention module to obtain the facial emotional features is the same as the process of processing a set of learnable emotional prototypes through the self-attention module to obtain the emotional prototype features; the process of processing the facial emotional features, the optimized body features, and the body position embedding through the cross-attention module to obtain the emotional features of the optimized body features is the same as the process of processing the emotional prototype features, the optimized facial features, and the facial position embedding through the cross-attention module to obtain the emotional features of the optimized facial features.

[0093] As an alternative embodiment, in step S107, inputting the emotional features of the optimized body features, the background features, and the background position embedding into the third cross-window inference layer to obtain the emotion category of each user includes: processing the emotional features of the optimized body features through the self-attention module to obtain the body emotional features; processing the body emotional features, the background features, and the background position embedding through the cross-attention module to obtain the emotion category of each user.

[0094] In this embodiment, the process of processing the emotional features of the optimized body features through the self-attention module to obtain the body emotional features is the same as the process of processing a set of learnable emotional prototypes through the self-attention module to obtain the emotional prototype features; the process of processing the body emotional features, the background features, and the background position embedding through the cross-attention module to obtain the emotion category of each user is the same as the process of processing the emotional prototype features, the optimized facial features, and the facial position embedding through the cross-attention module to obtain the emotional features of the optimized facial features.

[0095] As an alternative embodiment, in step S107, the expression of the target loss function is:

[0096]

[0097] where, where, is the target loss function, α t is the adjustment factor, t is a marker indicating whether the emotion category predicted by the model for each user is consistent with the true label of the emotion category of each user, is the true label of the emotion category of each user, α is the first hyperparameter, σ represents the sigmoid function, c is a constant, β is the second hyperparameter, γ is the third hyperparameter, p t is the probability that the predicted emotion category of each user is the correct category, and r is the fourth hyperparameter.

[0098] In this embodiment, when the target loss function is minimized, the network training ends.

[0099] In the embodiments of the present invention, by obtaining training set images, where each training image is a digital image and the digital image includes multiple users; processing each training image to obtain three target windows for each training image, where the three target windows are the face window, body window, and background window of each user respectively; inputting the three target windows into the backbone network in parallel to obtain the face features, body features, and background features of each user, where the backbone network includes a face feature extractor, a body feature extractor, and a background feature extractor; constructing a first cross-window interaction layer and a second cross-window interaction layer, where the first cross-window interaction layer and the second cross-window interaction layer include a cross-attention mechanism, a feed-forward network, and a skip connection; inputting the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features; inputting the body features and face features of each user into the second cross-window interaction layer to obtain optimized face features; encoding the position information of the background window of each user in each training image to obtain a background position embedding; obtaining a face position embedding based on the background position embedding and the position information of the face window of each user in each training image; obtaining a body position embedding based on the background position embedding and the position information of the body window of each user in each training image; introducing a set of learnable emotion prototypes; constructing a first cross-window inference layer, a second cross-window inference layer, and a third cross-window inference layer, where each cross-window inference layer includes a self-attention module and a cross-attention module; inputting a set of learnable emotion prototypes, the optimized face features, and the face position embedding into the first cross-window inference layer to obtain the emotion features of the optimized face features; inputting the emotion features of the optimized face features, the optimized body features, and the body position embedding into the second cross-window inference layer to obtain the emotion features of the optimized body features; inputting the emotion features of the optimized body features, the background features, and the background position embedding into the third cross-window inference layer to obtain the emotion category of each user. After obtaining the emotion category of each user, in each batch-size during the training process of the training set images, based on the target loss function, update the hyperparameters of the model that processes the training set images during the training process. When the target loss function is minimized, the model training ends, solving the technical problem of low recognition accuracy of the model for emotion recognition in the prior art, and achieving the technical effect of improving the accuracy of emotion recognition by fusing the emotion recognition framework and the spatial alignment position embedding module through the closed-layer features.

[0100] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0101] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0102] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in electrical or other forms.

[0103] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] In addition, each functional unit in various embodiments of the present invention can be integrated in a first processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0105] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for emotion recognition by fusing hierarchical features in digital images, characterized in that, Including: Obtain training set images, process each training image to obtain three target windows for each training image. Each training image is a digital image that includes multiple users. The three target windows are the facial window, body window, and background window for each user respectively. Parallelly input the three target windows into the backbone network to obtain the facial features, body features, and background features for each user. The backbone network includes a facial feature extractor, a body feature extractor, and a background feature extractor. Construct a first cross-window interaction layer and a second cross-window interaction layer. Input the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features. Input the body features and facial features of each user into the second cross-window interaction layer to obtain optimized facial features. The first cross-window interaction layer and the second cross-window interaction layer include a cross-attention mechanism, a feed-forward network, and a skip connection. Encode the position information of the background window of each user in each training image to obtain a background position embedding. Encode the positions of the facial window and body window of each user in each training image to obtain the position information of the facial window and body window of each user in each training image. Based on the background position embedding and the position information of the facial window of each user in each training image, obtain a facial position embedding. Based on the background position embedding and the position information of the body window of each user in each training image, obtain a body position embedding. Introduce a set of learnable emotion prototypes and construct a first cross-window inference layer, a second cross-window inference layer, and a third cross-window inference layer. Each cross-window inference layer includes a self-attention module and a cross-attention module. Input a set of learnable emotion prototypes, optimized facial features, and facial position embedding into the first cross-window inference layer to obtain the emotion features of the optimized facial features. Input the emotion features of the optimized facial features, optimized body features, and body position embedding into the second cross-window inference layer to obtain the emotion features of the optimized body features. Input the emotion features of the optimized body features, background features, and background position embedding into the third cross-window inference layer to obtain the emotion category of each user. After obtaining the emotion category of each user, in each batch-size during the training process on the training set images, based on the objective loss function, update the hyperparameters of the model that processes the training set images during the training process. When the objective loss function is minimized, the model training ends.

2. The method according to claim 1, wherein The process of processing each training image to obtain three target windows for each training image includes: Use pre-detected bounding boxes to crop each training image to obtain three different initial windows for each cropped training image. Adjust the three different initial windows to obtain the target window corresponding to each initial window.

3. The method according to claim 2, wherein The expression for parallelly inputting the three target windows into the backbone network to obtain the facial features, body features, and background features of each user is: Among them, f F is the facial feature of each user, f B is the body feature of each user, f C is the background feature of each user, is the backbone network, F is the face of each user, B is the body of each user, C is the background of each user, θ1 is the parameter of the facial feature extractor, θ2 is the parameter of the body feature extractor, and θ3 is the parameter of the background feature extractor.

4. The method according to claim 3, characterized in that, Inputting the body features and background features of each user into the first cross-window interaction layer to obtain optimized body features, including: Inputting the body features and background features of each user into the cross-attention mechanism to obtain the enlarged body features of each user; Inputting the enlarged body features of each user into the feed-forward network and connecting them with the body features through skip connections to obtain optimized body features.

5. The method according to claim 4, wherein Inputting the body features and facial features of each user into the second cross-window interaction layer to obtain optimized facial features, including: Inputting the body features and facial features of each user into the cross-attention mechanism to obtain the enlarged facial features of each user; Inputting the enlarged facial features of each user into the feed-forward network and connecting them with the facial features through skip connections to obtain optimized facial features.

6. The method according to claim 5, wherein Inputting a set of learnable emotion prototypes, optimized facial features, and facial position embeddings into the first cross-window inference layer to obtain the emotion features of the optimized facial features, including: Processing a set of learnable emotion prototypes through the self-attention module to obtain emotion prototype features; Processing the emotion prototype features, optimized facial features, and facial position embeddings through the cross-attention module to obtain the emotion features of the optimized facial features.

7. The method according to claim 6, characterized in that, Inputting the emotion features of the optimized facial features, optimized body features, and body position embeddings into the second cross-window inference layer to obtain the emotion features of the optimized body features, including: Processing the emotion features of the optimized facial features through the self-attention module to obtain facial emotion features; Processing the facial emotion features, optimized body features, and body position embeddings through the cross-attention module to obtain the emotion features of the optimized body features.

8. The method according to claim 7, wherein Inputting the emotion features of the optimized body features, background features, and background position embeddings into the third cross-window inference layer to obtain the emotion category of each user, including: Processing the emotion features of the optimized body features through the self-attention module to obtain body emotion features; Processing the body emotion features, background features, and background position embeddings through the cross-attention module to obtain the emotion category of each user.

9. The method according to claim 1, wherein The expression of the target loss function is: Among them, is the target loss function, α t is the adjustment factor, t is a marker indicating whether the emotion category predicted by the model for each user is consistent with the true label of the emotion category of each user, is the true label of the emotion category of each user, α is the first hyperparameter, σ represents the sigmoid function, c is a constant, β is the second hyperparameter, γ is the third hyperparameter, p t is the probability that the predicted emotion category of each user is the correct category, and r is the fourth hyperparameter.