Image processing method and device, electronic equipment and computer readable storage medium
By extracting and splicing expression and identity features in facial expression modeling, capturing joint information and minimizing mutual information, the shortcomings of independent modeling of expression and identity features in traditional methods are solved, and more accurate and robust expression recognition and identity authentication are achieved.
Patent Information
- Application Number
- CN202510678626.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Traditional facial expression modeling methods assume that expression features and identity features are independent, ignoring the potential correlation between the two, resulting in decreased expression classification accuracy or unstable identity recognition.
By extracting expression features and identity features from image data and performing feature splicing, the joint information between expression and identity is captured, and the expression features are adjusted by minimizing mutual information to reduce identity information, thereby achieving more accurate expression recognition and identity authentication.
The ability to model personalized expressions has been enhanced, and the accuracy and robustness of expression classification and identity verification have been improved.
Smart Images

Figure CN120673453A_ABST
Abstract
Description
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an image processing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] In the field of facial expression modeling, traditional methods typically decompose facial expression images into independent expression features and identity features, assuming that expression features and identity features are mutually independent, thereby ignoring the potential correlation between the two. However, in real-world scenarios, the performance of expression features is often affected by identity features. For example, different individuals may experience different facial muscle changes when expressing the same emotion. Therefore, modeling expression or identity features alone cannot fully capture the complex relationships in the real world, potentially leading to reduced expression classification accuracy or instability in identity recognition. Summary of the Invention
[0003] To solve the above technical problems, the embodiments of the present application provide an image processing method, device, electronic device, and computer-readable storage medium.
[0004] In a first aspect, an embodiment of the present application provides an image processing method, comprising:
[0005] Obtain image data of facial expressions;
[0006] extracting a first expression feature and a first identity feature from the image data;
[0007] performing feature concatenation on the first expression feature and the first identity feature to obtain a first joint feature, and extracting joint information between the first expression feature and the first identity feature from the first joint feature;
[0008] estimating mutual information between the first facial expression feature and the first identity feature, and adjusting the first facial expression feature to a second facial expression feature with the goal of minimizing the mutual information;
[0009] A classification result of the facial expression is predicted based on the second expression feature and the joint information.
[0010] In a second aspect, an embodiment of the present application provides an image processing device, which is applied to an electronic device and includes:
[0011] An acquisition module, used to acquire image data of facial expressions;
[0012] An extraction module, configured to extract a first expression feature and a first identity feature from the image data;
[0013] a processing module, configured to perform feature concatenation on the first expression feature and the first identity feature to obtain a first joint feature;
[0014] The extraction module is further configured to extract joint information between the first expression feature and the first identity feature from the first joint feature;
[0015] The processing module is also used to estimate the mutual information between the first expression feature and the first identity feature, adjust the first expression feature to the second expression feature with the goal of minimizing the mutual information; and predict the classification result of the facial expression based on the second expression feature and the joint information.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0017] a memory for storing computer-executable instructions or computer programs;
[0018] The processor is used to implement an image processing method provided in an embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements an image processing method provided in an embodiment of the present application.
[0020] In the technical solution of the embodiment of the present application, expression features and identity features are extracted from image data respectively, so as to realize the extraction of independent features of expression and identity; joint features are obtained by feature splicing of expression features and identity features, thereby effectively solving the deficiency of independent modeling of expression and identity features in traditional methods by capturing the joint information between expression and identity; the mutual information between expression features and identity features is estimated, and the expression features are adjusted to expression features containing less identity information with the goal of minimizing the mutual information, so as to realize more accurate and robust expression recognition and identity authentication; and the classification results of facial expressions are predicted by the joint information and expression features containing less identity information, so as to synergistically improve the expression classification, thereby not only enhancing the modeling ability of personalized expressions, but also improving the performance in multi-task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 1 is a schematic diagram of the architecture of an image processing system 100 provided in an embodiment of the present application;
[0022] Figure 2 is a structural diagram of the image processing method provided in an embodiment of the present application;
[0023] Figure 3 Schematic diagram of the image processing method provided in the embodiment of the present application;
[0024] Figure 4 Schematic diagram of the image processing method provided in the embodiment of the present application;
[0025] Figure 5 is a schematic diagram of the facial expression feature classification provided in an embodiment of the present application;
[0026] Figure 6 Schematic diagram of the image processing method provided in the embodiment of the present application;
[0027] Figure 7 Schematic diagram of the image processing method provided in the embodiment of the present application;
[0028] Figure 8 Schematic diagram of the structure of the image processing device provided in an embodiment of the present application;
[0029] Figure 9 This is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0031] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0032] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0033] In the following description, the terms "first\second\..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second\..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0035] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0036] 1) Facial expression modeling refers to the process of building a computational model that can simulate, analyze, or recognize changes in facial expressions. This model aims to capture dynamic features such as facial muscle movement and skin deformation, and translate human emotional states (such as joy, sadness, and anger) into quantifiable or visual digital representations.
[0037] 2) Expression features refer to dynamic information related to facial expression in facial images, namely short-term deformation caused by facial muscle movement (such as a raised corner of the mouth or a furrowed eyebrow). These features are directly related to an individual's emotional state (such as happiness, anger, sadness) and are independent of personal identity.
[0038] 3) Identity features refer to static information related to individual identity in facial images, namely, inherent facial attributes that are not affected by facial expressions (such as bone structure, facial proportions, skin color, etc.). These features are used to distinguish different people.
[0039] 4) The image encoder (also known as a visual encoder or feature extractor) is a component specifically designed for image processing. The image encoder receives an image as input and converts the raw pixel data into a set of dense vector representations through a series of neural network layers, such as the ResNet and Vision Transformer (ViT). During this process, the image encoder learns an abstract representation of the image content and semantic information. Ultimately, this set of vectors is mapped to a fixed-dimensional space as image features to facilitate subsequent cross-modal interaction.
[0040] In the facial expression modeling methods of related technologies, facial expression images are usually decomposed into two independent information parts: expression features and identity features. Among them, expression features represent dynamic information related to expression in facial images, that is, short-term deformation caused by facial muscle movement, such as raised corners of the mouth and wrinkled eyebrows. Expression features are directly related to an individual's emotional state, such as happiness, anger, and sadness, and are independent of personal identity. Identity features represent static information related to individual identity in facial images, that is, inherent facial attributes that are not affected by expression, such as bone structure, facial feature proportions, skin color, etc. Identity features can be used to distinguish different people.
[0041] Traditional facial modeling methods often assume that expression and identity are independent of each other, thus ignoring the potential correlation between them. However, in real-world scenarios, the representation of expression features is often influenced by identity features. For example, different individuals may experience different changes in facial muscles when expressing the same emotion. Therefore, modeling expression or identity features independently cannot fully capture the complex relationships in the real world, potentially leading to decreased expression recognition accuracy or instability in identity recognition. Consequently, traditional approaches suffer from the limitations of independent modeling. For example, when extracting expression and identity features, traditional methods often adopt a decoupling strategy, attempting to separate expression features from identity features and pursuing mutual independence between them. This modeling approach ignores the synergistic relationship between expression and identity features, resulting in incomplete feature representation and difficulty in accurately modeling individual expression variations. Furthermore, traditional modeling methods lack joint information. For example, the relationship between expression and identity features is dynamic and complex, and certain expressions may be expressed differently across individuals. This individualized variation requires capturing joint information. However, existing methods often lack joint representation mechanisms, making it difficult to effectively model the interactive relationship between expression and identity.
[0042] In order to solve the above problems, the embodiments of the present application provide an image processing method, device, electronic device, and computer-readable storage medium. By introducing collaborative feature modeling and joint information optimization mechanism, it effectively solves the shortcomings of independent modeling of expression and identity features in traditional methods, thereby achieving more accurate and robust expression recognition and identity authentication.
[0043] The following describes exemplary applications of the device provided in the embodiments of the present application. The electronic device provided in the embodiments of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and car terminals, and can also be implemented as servers.
[0044] See also Figure 1 , Figure 1 This is a structural diagram of the image anomaly detection system architecture for a virtual scene provided by an embodiment of the present application, for example, Figure 1 The server 100, the terminal device 200 and the network 300 are involved. The terminal device 200 is connected to the server 100 via the network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two.
[0045] In some embodiments, the embodiments of the present application can be implemented collaboratively by a server and a terminal device. For example, the terminal device 200 sends image data of facial expressions to the server 100. The server 100 uses the image processing method provided in the embodiments of the present application to predict the classification results of the facial expressions and sends the classification results of the facial expressions to the terminal device 200. The terminal device 200 uses the classification results of the facial expressions to implement facial expression recognition or identity verification.
[0046] In other embodiments, the present invention can be implemented on a terminal device alone. The terminal device 200 sends facial expression image data to the server. The server 100 receives the facial expression image data and sends the facial expression classification model provided in the present invention to the terminal device 200. The terminal device 200 receives the model sent by the server and downloads it locally. The terminal device 200 obtains the facial expression classification result based on the model and image data. The terminal device 200 uses the facial expression classification result to perform facial expression recognition or identity verification.
[0047] In some embodiments, the terminal device or server can implement the image processing method provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer program can be a native program or software module in the operating system. In short, the above-mentioned computer-executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module, or plug-in in any form. The terminal devices include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.
[0048] See also Figure 2 , Figure 2 This is a schematic diagram of the principle of the image processing method provided by an embodiment of the present application. First, image data of facial expressions is obtained, and a first expression feature and a first identity feature are extracted from the image data. Next, the first expression feature and the first identity feature are feature-joined to obtain a first joint feature, and joint information between the first expression feature and the first identity feature is extracted from the first joint feature. Then, the mutual information between the first expression feature and the first identity feature is estimated, and the first expression feature is adjusted to a second expression feature with the goal of minimizing the mutual information. Finally, the classification result of the facial expression is predicted based on the second expression feature and the joint information. See Figure 3 , Figure 3 This is a flow chart of the image processing method provided in the embodiment of the present application, which is described in detail below.
[0049] In step 301, image data of facial expressions is obtained.
[0050] In some embodiments, facial expression image data may be obtained from an expression dataset. For example, the source of facial expression data may be a public dataset containing basic and complex expressions, covering a wide range of ages, ethnicities, and occlusion scenarios. Another example is a self-constructed dataset, such as one obtained by recording photos using a capture device under controlled lighting conditions.
[0051] In some implementations, low-quality samples in facial expression datasets can be removed through blur detection and extreme pose filtering. For example, the Laplacian operator can be used to calculate an image clarity score, filtering out samples with scores below a threshold. Based on head pose estimation, images with pitch angles greater than 0 degrees or yaw angles greater than 45 degrees can be removed. Facial expression data can also be balanced by oversampling and undersampling. For example, oversampling can be used to generate synthetic images for minority class samples, followed by undersampling to randomly remove majority class samples to ensure that the sample sizes of each class are similar.
[0052] In other embodiments, after acquiring the image data of facial expressions, the image data may be preprocessed. For example, the image data of facial expressions may be subjected to face alignment processing to standardize the geometric structure of the face to ensure posture standardization. For another example, the image data may be subjected to data normalization processing by performing color space normalization and / or size normalization on the image data, i.e., grayscale conversion on the image data to reduce the amount of computation. Alternatively, the image data may be subjected to Z-score normalization to eliminate illumination differences, or alternatively, the image data may be subjected to histogram equalization to enhance the details of low-contrast images. The image data may be subjected to a uniform resolution, and the image may be scaled to the model input size, using bilinear interpolation to maintain smoothness.
[0053] In other embodiments, after acquiring facial expression image data, data augmentation can be performed on the image data. For example, geometric transformations can be performed on the image data to avoid excessive distortion of the facial structure through slight rotations, or random cropping can be performed to preserve the core facial area. Another example is photometric transformations, brightness adjustments, contrast adjustments, color dithering, etc. Data augmentation can accelerate model convergence and increase robustness.
[0054] In step 302, a first expression feature and a first identity feature are extracted from image data.
[0055] In the embodiments of the present application, the first expression feature is used to describe the person's attribute information in the image data, specifically the person's posture and expression information. This represents dynamic information related to expression in the facial image, namely, short-term deformation caused by facial muscle movement, such as a raised corner of the mouth or furrowed eyebrows. Expression features are directly related to the individual's emotional state and are independent of their identity. The first identity feature is used to describe the person's identity information in the image data. This represents static information related to the individual's identity in the facial image, namely, inherent facial attributes that are not affected by expression, such as bone structure, facial proportions, and skin color.
[0056] In some embodiments, extracting the first expression feature and the first identity feature from the image data includes: extracting the first expression feature from the image data through an expression encoder; extracting the first identity feature from the image data through an identity encoder; wherein the identity encoder is pre-trained.
[0057] Here, the pre-processed image data is subjected to expression feature extraction processing, for example, expression features in the image data are identified by an expression encoder to capture fine-grained information related to the expression. The expression encoder can use a deep neural network (DNN), a convolutional neural network (CNN) or other image processing technology to identify expression features in the image data, and encode this information into a high-dimensional vector to obtain a first expression feature. The pre-processed image data is subjected to identity feature extraction processing, for example, expression features in the image data are identified by an identity encoder to extract identity features related to the individual identity, such as facial geometry, muscle distribution, etc. The identity encoder can use a DNN, CNN or other image processing technology to identify identity features in the image data, and encode this information into a high-dimensional vector to obtain a first identity feature. Among them, the identity feature requires a face dataset for training.
[0058] It should be noted that the feature extraction methods for the expression encoder and identity encoder require different parameters but similar structures. For example, both the expression encoder and identity encoder may use the Transformer structure for feature extraction, or both may use convolutional neural network structures such as ResNet or VGG for feature extraction, and the identity encoder may be pre-trained on a large-scale face dataset. The application does not specify the specific feature extraction method used.
[0059] In some embodiments, the image data can be subjected to expression encoding processing by CNN, such as a residual neural network (Resnet), to obtain a first expression feature. The embodiments of the present application do not limit the specific implementation methods of identity encoding processing of image data. In a convolutional neural network, there are usually multiple fully connected layers or multiple global average pooling layers, so that a final fixed-size feature vector is generated at the end of the convolutional neural network. This is the first expression feature, which contains the expression feature information of the image data.
[0060] In this example, the ResNet convolutional architecture is used to extract facial expressions. Local features are extracted layer by layer using a sliding window of convolution kernels. The first convolutional layer downsamples the image to 112×112. The image is then passed through four groups of residual blocks, with channels ranging from 64 to 256, 128 to 512, 256 to 1024, and 512 to 2048. Each group of channels includes a bottleneck structure. Finally, global average pooling is performed to generate the final feature map, which is then compressed into a 2048-dimensional vector. This is followed by a fully connected layer to output a 512-dimensional facial expression feature.
[0061] For example, the expression encoder is used to extract the expression features of the image data to obtain the first expression feature, which can be expressed by formula (1):
[0062] F exp =CNN(I;θ CNN1 ) (1)
[0063] Among them, F exp represents the first expression feature, CNN represents convolutional neural network, I represents image data, θ CNN1 are the parameters of the convolutional neural network of the expression encoder.
[0064] It should be noted that the identity feature encoder can also extract identity features using the same ResNet convolutional architecture as the expression encoder, outputting a 512-dimensional identity feature to obtain the first identity feature, which can be expressed by formula (2). The difference is that the parameters of the identity encoder and the expression encoder are independent of each other, with no shared layers or weights.
[0065] F id =CNN(I;θ CNN2 ) (2)
[0066] Among them, F id represents the first identity feature, CNN represents convolutional neural network, I represents image data, θ CNN2 are the parameters of the convolutional neural network of the identity encoder.
[0067] For example, the image data can be processed by expression encoding through a network model such as DNN to obtain a first expression feature and a first identity feature, for example, an Additive Angular Margin Loss for Deep Face Recognition (ArcFace) network in deep face recognition.
[0068] In other embodiments, the image data may be segmented to obtain a plurality of pixel blocks; the plurality of pixel blocks may be sorted into a sequence of pixel blocks to be encoded; and the sequence of pixel blocks to be encoded may be encoded to obtain a first identity feature and a first expression feature, wherein the expression encoder and the identity encoder use the same segmentation and embedding process, but the identity encoder is trained using an identity data set.
[0069] For example, assuming the image data size is 224×224, it is divided according to the set pixel block size (for example, 16×16 pixel block size), and the result after division is (224 / 16). 2 = 196 pixel blocks (Pat ches), multiple pixel blocks are arranged (e.g., in order from left to right and from top to bottom) as a pixel block sequence to be encoded, the pixel block sequence to be encoded is used as an input unit for identity encoding processing, image data is divided into multiple pixel blocks according to the set length and width of the pixel blocks, and the multiple pixel blocks are arranged in order from left to right and from top to bottom to form the pixel block sequence to be encoded.
[0070] For example, the pixel block sequence to be coded can be embedded coded by an embedding layer (Embedding Layer) to obtain embedded features. For example, the pixel block sequence to be coded is convolved by a convolution operation in the embedding layer, and the pixel block sequence to be coded after the convolution is position coded (Position Embedding) to obtain embedded features. Before the embedded coding process, each pixel block to be coded in the pixel block sequence to be coded can be flattened into a one-dimensional vector. For example, for a 16×16×3 pixel block to be coded, the length of the flattened vector is 768. The position coding process can be generated by a fixed algorithm, for example, it can be implemented using a combination of sine and cosine functions. For each position (for example, the pixel position of each pixel point of each pixel block to be coded) pos and the dimension i of the feature, the position coding of the even dimension uses the sine function, while the position coding of the odd dimension uses the cosine function. The embodiment of the present application does not limit the specific implementation method of the position coding.
[0071] For example, before attention encoding, the embedded features can also be normalized (for example, layer normalization) to obtain normalized features. For the normalized features for attention encoding, a linear transformation is first performed to generate three matrices: query vector (Query, Q), key vector (Key, K), and value vector (Value, V). For each attention head (Head Attention), the dot product of Q and K is calculated to obtain the attention score. Next, the attention score is normalized, for example, by applying a normalization function (such as the Softmax function) to convert the attention score into a probability distribution. The attention probability distribution is used to weight V to generate a new feature representation. Finally, the outputs of all attention heads are connected and linearly transformed, and linearly projected into a 512-dimensional embedding. A rich representation that includes the correlation between different positions in the first normalized feature is obtained, namely the attention encoding feature.
[0072] For example, the attention encoding features can be mapped through a feed-forward neural network (FFN) layer to obtain a 512-dimensional first identity feature and a 512-dimensional first expression feature, respectively. The feed-forward neural network layer can, for example, adopt a multilayer perceptron (MLP) structure.
[0073] It should be noted that the expression encoder and identity encoder can use the same Transformer structure for feature extraction. The difference is that the parameters of the expression encoder and identity encoder are independent, and the identity encoder is trained using the identity dataset.
[0074] In step 303 , feature concatenation is performed on the first expression feature and the first identity feature to obtain a first joint feature, and joint information between the first expression feature and the first identity feature is extracted from the first joint feature.
[0075] Here, refer to Figure 5 Different individuals experience different facial muscle changes when expressing the same emotion. The relationship between facial expressions and identity features is dynamic and complex. For example, certain expressions may be expressed differently in different individuals. This individualized difference needs to be captured through joint information. By concatenating the first expression feature and the first identity feature, a first joint feature is generated. From this first joint feature, the joint information between the first expression feature and the first identity feature is extracted.
[0076] In some embodiments, the first expression feature and the first identity feature can be spliced together using a feature splicing method to obtain a spliced feature. For example, a 512-dimensional expression feature and a 512-dimensional identity feature can be spliced together to obtain a 1024-dimensional spliced feature.
[0077] In some embodiments, the 1024-dimensional concatenated features can be fused using a feature mapping method to obtain a first joint feature. For example, the 1024-dimensional concatenated features can be mapped to a fixed-size feature space using a multi-layer perceptron (MLP), and the mapped concatenated features can be fused to obtain the first joint feature. MLP is a feedforward neural network composed of multiple fully connected layers, each using an activation function. MLP can automatically learn nonlinear relationships between features and transform them into a form suitable for subsequent tasks.
[0078] For example, the first expression feature and the first identity feature are concatenated to obtain a first joint feature, which can be expressed by formula (3):
[0079] F MT =F(F exp ; F id θ F ) (3)
[0080] Among them, F MT represents the first joint feature, F represents the fusion process, F exp Indicates the first facial expression feature, F id represents the first identity feature, θ F Represents the parameters of the fusion process (such as the parameters of the MLP network when the fusion process is performed through MLP).
[0081] In step 304 , the mutual information between the first expression feature and the first identity feature is estimated, and the first expression feature is adjusted to the second expression feature with the goal of minimizing the mutual information.
[0082] In some embodiments, the extracted first expression feature and first identity feature are optimized by minimizing the mutual information between the expression and the identity, minimizing the mutual information between the first expression feature and the first identity feature, thereby achieving the independence of the expression feature and the identity feature, avoiding the inclusion of identity information in the expression feature, reducing the identity information contained in the expression feature, and the expression information contained in the identity feature, thereby ensuring the independence of the two.
[0083] For example, the dependency between the two random variables, the first expression feature and the first identity feature, can be measured by mutual information (MI), which can be defined by formula (4):
[0084] I(X;Y)=H(X)-H(X∣Y)=H(Y)-H(Y∣X) (4)
[0085] Where X represents the first facial expression feature; Y represents the first identity feature; H(X) is the information entropy of variable X; and H(X|Y) is the conditional entropy, representing the remaining uncertainty about X after Y is known. A larger I(X;Y) indicates more redundant information between X and Y. Minimizing I(X;Y) aims to reduce the identity information contained in facial expression features and the facial expression information contained in identity features.
[0086] In some embodiments, see Figure 4 , Figure 3 In step 304 shown, the mutual information between the first expression feature and the first identity feature is estimated, which can be achieved through the following steps 401 to 402, which are described in detail below.
[0087] Step 401: Calculate the joint distribution and marginal distribution between the first expression feature and the first identity feature through a mutual information estimator to obtain the joint distribution expectation and the marginal distribution logarithmic expectation.
[0088] Step 402: Calculate the mutual information between the first expression feature and the first identity feature based on the joint distribution expectation and the marginal distribution logarithmic expectation.
[0089] For example, I(X;Y) can be estimated by maximizing the variational lower bound of mutual information using the Mutual Information Neural Estimation (MINE) method based on a neural network. Specifically, it can be expressed as:
[0090] I(X;Y)≥E p(x,y) [T(x,y)]-logE p(x)p(y) [e T(x,y) ] (5)
[0091] The mutual information estimator can be expressed as T(x,y), which can include a neural network (called a statistical network) to distinguish the joint distribution p(x,y) and the marginal distribution product p(x)p(y). p(x,y) [T(x,y)] represents the joint distribution expectation, logE p(x)p(y) [e T(x,y) ] represents the marginal distribution logarithmic expectation. By calculating the mutual information of the first expression feature and the first identity feature through the joint distribution expectation and the marginal distribution logarithmic expectation, it can be obtained by formula (6):
[0092]
[0093] Among them, the mutual information loss function can be obtained by maximizing the mutual information lower bound, which is equivalent to minimizing the negative estimate.
[0094] In some embodiments, see Figure 6 ,exist Figure 4 Before step 401 shown, the following steps 601 to 603 may be performed to train the mutual information estimator, which will be described in detail below.
[0095] Step 601: Acquire a joint distribution sample pair, where the joint distribution sample pair includes a first expression feature sample and a first identity feature sample.
[0096] Step 602: Randomly arrange the first identity feature samples to obtain second identity feature samples, where the second identity feature samples include identity feature information that is unrelated to the first expression feature samples.
[0097] Step 603: Generate a marginal distribution sample pair based on the first expression feature sample and the second identity feature sample, where the marginal distribution sample pair includes the first expression feature sample and the second identity feature sample.
[0098] For example, the input data includes expression features extracted by the expression encoder, which can be expressed as X = {x1, ..., x N}, the input data includes the identity features extracted by the identity encoder, which can be expressed as Y = {y1,…,y N According to the input data, sample pairs are constructed to obtain joint distribution sample pairs. Joint distribution sample pairs can also be called true sample pairs or positive sample pairs, which can be expressed as (x i ,y i ), where x i Represents the first expression feature sample, y i Represents the first identity feature sample.
[0099] For example, the first identity feature sample is randomly arranged to obtain the second identity feature sample y j , the second identity feature sample includes identity feature information that is irrelevant to the first expression feature sample, that is, j≠i. At this time, the marginal distribution sample pair generated based on the first expression feature sample and the second identity feature sample can be expressed as (x i ,y j ).
[0100] In some embodiments, see Figure 7 ,exist Figure 4 The illustrated step 401 can also be implemented through the following steps 701 to 704, which are described in detail below.
[0101] Step 701: Calculate the joint distribution of the joint distribution sample pairs through the mutual information estimator to obtain a joint distribution score value.
[0102] For example, the mutual information estimator can be expressed as T(x,y), and the mutual information estimator is used to calculate the joint distribution of sample pairs (x i ,y i ), the joint distribution score can be expressed by formula (7):
[0103] T joint =T(x i ,y i )(i=1,2,…,N) (7)
[0104] Step 702: Calculate the marginal distribution of the marginal distribution sample pair using a mutual information estimator to obtain a marginal distribution score value.
[0105] For example, the marginal distribution sample pair (x i ,y j ), the marginal distribution score can be expressed by formula (8):
[0106] T marginal =T(x i ,y j )(j≠i) (8)
[0107] Step 703: Perform expectation estimation processing on the joint distribution score value to obtain the joint distribution expectation.
[0108] For example, the joint distribution expectation can be expressed as E p(x,y) [T(x,y)], the joint distribution expectation is calculated by the mutual information estimator, and the obtained joint distribution expectation can be expressed by formula (9):
[0109]
[0110] Step 704: perform expectation estimation processing on the marginal distribution score value to obtain the marginal distribution logarithmic expectation.
[0111] For example, the marginal logarithmic expectation can be expressed as logE p(x)p(y) [e T(x,y) ], the marginal distribution logarithmic expectation is calculated by the mutual information estimator, and the obtained marginal distribution logarithmic expectation can be expressed by formula (10):
[0112]
[0113] In some embodiments, Figure 3Step 304 shown includes: minimizing the joint distribution expectation and the marginal distribution logarithmic expectation based on the first expression feature and the first identity feature through a mutual information estimator to obtain a second expression feature; wherein the identity feature information contained in the second expression feature is less than the identity feature information contained in the first expression feature.
[0114] For example, a second expression feature is generated by minimizing the mutual information between the expression feature and the identity feature through a mutual information estimator. The second expression feature satisfies the characteristics of reducing identity information and retaining expression information. Specifically, the identity feature information contained in the second expression feature is less than that in the first expression feature, and the expression discrimination ability of the first expression feature is maintained.
[0115] For example, the first facial expression feature can be represented as X old , the first identity feature can be expressed as Y old , the second expression feature can be expressed as X new The estimation of joint distribution expectation and marginal distribution logarithmic expectation using the mutual information estimator T(x,y) based on the MINE network can be expressed by formula (11):
[0116]
[0117] In step 305, a classification result of the facial expression is predicted based on the second expression feature and the joint information.
[0118] In some embodiments, the second facial expression feature X new is the facial expression feature after mutual information minimization, and the joint information Z joint The first expression feature X old and the first identity characteristic Y old The first joint feature F generated by MLP fusion after splicing MT By combining the joint information Z joint Input to the expression classifier and output the expression classification result.
[0119] In some embodiments, the first loss function is determined based on a variational lower bound of the mutual information. For example, maximizing the mutual information lower bound is equivalent to minimizing the negative estimated value to obtain the first loss function, and the first loss function can be expressed by formula (12):
[0120]
[0121] In some embodiments, a second loss function is determined based on the facial expression classification result and the facial expression classification label. For example, the second loss function is obtained by directly optimizing the expression classification task through cross entropy loss. The second loss function can be expressed by formula (13):
[0122]
[0123] Among them, y i,c represents the true label of the i-th sample, p i,c Represents the category probability of the expression classification model.
[0124] In some embodiments, the first loss function and the second loss function are summed or weighted to obtain a target loss function. For example, by combining the classification loss and the mutual information minimization loss, the target loss function can be expressed by formula (14):
[0125] L=L cls +L mine (14)
[0126] In some embodiments, the expression encoder and the mutual information estimator are updated based on the target loss function.
[0127] Specifically, the expression encoder F can be optimized by the back propagation algorithm according to the loss calculated by the above formula (14). exp and mutual information estimator T(x,y), thereby improving the performance of expression recognition and identity feature extraction.
[0128] For example, before backpropagation, forward propagation is first performed to obtain the prediction results and loss value of the expression classification model. The preprocessed image data is input into the model. The convolutional layer extracts local features (such as edges and textures), and nonlinearity is introduced using an activation function (such as ReLU). Next, the feature map dimension is reduced using a pooling layer (such as max pooling), and then the features are mapped to the classification space through a fully connected layer, such as outputting 7-dimensional expression probabilities. Finally, the loss is calculated by comparing the predicted value with the true label.
[0129] For example, backpropagation is used to calculate the target loss function, that is, the gradient of the model parameters, to guide parameter updates. Specifically, the gradient calculation is first performed by calculating the gradient layer by layer based on the chain rule from the output layer. The output layer gradient includes the derivative of the loss with respect to the predicted probability, and the hidden layer gradient includes the gradient propagated back layer by layer (such as the fully connected layer, pooling layer, and convolutional layer). Next, the gradient can be automatically calculated through automatic differentiation.
[0130] For example, the model parameters are adjusted according to the gradient to minimize the loss function. Specifically, the parameters can be updated by using an optimization algorithm, gradient descent or other methods. The gradient descent formula can be expressed by formula (15):
[0131]
[0132] Here, θ represents the model parameters and η represents the learning rate. It should be noted that the historical gradient can be cleared before each backpropagation to avoid gradient accumulation.
[0133] In this way, forward propagation is responsible for generating predictions and calculating target losses, and backpropagation guides parameter optimization through gradient calculation, and the two are iterated until the model converges.
[0134] From the above, it can be seen that the embodiment of the present application provides an image processing method, which implements a collaborative modeling method of expression and identity through steps 301 to 305, steps 401 to 402, steps 601 to 603, and steps 701 to 704, and realizes a collaborative modeling method of expression and identity, by extracting expression features and identity features from image data respectively, thereby realizing the extraction of independent features of expression and identity; further, by feature splicing of expression features and identity features to obtain joint features, thereby effectively solving the deficiency of independent modeling of expression and identity features in traditional methods by capturing the joint information between expression and identity; further, by estimating the mutual information between expression features and identity features, the expression features are adjusted to expression features containing less identity information with the goal of minimizing the mutual information, thereby realizing more accurate and robust expression recognition and identity authentication; and then, the classification results of facial expressions are predicted by the joint information and expression features containing less identity information, and the expression classification is collaboratively improved, thereby not only enhancing the modeling ability of personalized expressions, but also improving the performance in multi-task scenarios (such as expression classification and identity recognition).
[0135] Figure 8 This is a schematic diagram of the structure of the image processing device provided in the embodiment of the present application. Figure 1 , used in electronic devices such as Figure 8 As shown, the image processing device 800 includes:
[0136] The acquisition module 801 is used to acquire image data of facial expressions.
[0137] The extraction module 802 is configured to extract a first expression feature and a first identity feature from the image data.
[0138] The processing module 803 is configured to perform feature concatenation on the first expression feature and the first identity feature to obtain a first joint feature.
[0139] The extraction module 802 is further configured to extract joint information between the first expression feature and the first identity feature from the first joint feature.
[0140] The processing module 803 is also used to estimate the mutual information between the first expression feature and the first identity feature, adjust the first expression feature to the second expression feature with the goal of minimizing the mutual information; and predict the classification result of the facial expression based on the second expression feature and the joint information.
[0141] In some embodiments, the extraction module 802 is further configured to extract a first expression feature from the image data via an expression encoder; and extract a first identity feature from the image data via an identity encoder; wherein the identity encoder is pre-trained.
[0142] In some embodiments, the processing module 803 is also used to calculate the joint distribution and marginal distribution between the first expression feature and the first identity feature through a mutual information estimator to obtain the joint distribution expectation and the marginal distribution logarithmic expectation; based on the joint distribution expectation and the marginal distribution logarithmic expectation, calculate the mutual information between the first expression feature and the first identity feature.
[0143] In some embodiments, before the processing module 803 calculates the joint distribution and marginal distribution between the first expression feature and the first identity feature, the acquisition module 801 is also used to obtain a joint distribution sample pair, and the joint distribution sample pair includes a first expression feature sample and a first identity feature sample.
[0144] The processing module 803 is further used to randomly arrange the first identity feature samples to obtain second identity feature samples, where the second identity feature samples include identity feature information that is unrelated to the first expression feature samples; and to generate a marginal distribution sample pair based on the first expression feature samples and the second identity feature samples, where the marginal distribution sample pair includes the first expression feature samples and the second identity feature samples.
[0145] In some embodiments, the processing module 803 is further used to calculate the joint distribution of the joint distribution sample pairs through the mutual information estimator to obtain a joint distribution score value; calculate the marginal distribution of the marginal distribution sample pairs through the mutual information estimator to obtain a marginal distribution score value; perform expectation estimation processing on the joint distribution score value to obtain the joint distribution expectation; perform expectation estimation processing on the marginal distribution score value to obtain the marginal distribution logarithmic expectation.
[0146] In some embodiments, the processing module 803 is also used to minimize the joint distribution expectation and the marginal distribution logarithmic expectation based on the first expression feature and the first identity feature through the mutual information estimator to obtain the second expression feature; wherein the identity feature information contained in the second expression feature is less than the identity feature information contained in the first expression feature.
[0147] In some embodiments, the processing module 803 is also used to determine a first loss function based on a variational lower bound of mutual information; determine a second loss function based on the classification result of facial expressions and the classification label of facial expressions; sum or weighted sum the first loss function and the second loss function to obtain a target loss function; and update the expression encoder and mutual information estimator based on the target loss function.
[0148] Those skilled in the art should understand that Figure 8 The implementation functions of each module in the image processing device shown can be understood by referring to the relevant description of the aforementioned method. Figure 8 The functions of the modules in the image processing device shown can be implemented by a program running on a processor or by a specific logic circuit.
[0149] Figure 9 900 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may be a terminal device or a server. Figure 9 The electronic device 900 shown includes a processor 910, which can call and run a computer program from a memory to implement the method in the embodiment of the present application.
[0150] Alternatively, as Figure 9 As shown, the electronic device 900 may further include a memory 920. The processor 910 may call and execute a computer program from the memory 920 to implement the method in the embodiment of the present application.
[0151] The memory 920 may be a separate device independent of the processor 910 , or may be integrated into the processor 910 .
[0152] Alternatively, as Figure 9 As shown, the electronic device 900 may further include a transceiver 930 , and the processor 910 may control the transceiver 930 to communicate with other devices. Specifically, the transceiver 930 may send information or data to other devices, or receive information or data sent by other devices.
[0153] The transceiver 930 may include a transmitter and a receiver. The transceiver 930 may further include an antenna, and the number of antennas may be one or more.
[0154] Optionally, the electronic device 900 may specifically be a server in an embodiment of the present application, and the electronic device 900 may implement the corresponding processes implemented by the server in each method in the embodiment of the present application. For the sake of brevity, they will not be repeated here.
[0155] Optionally, the electronic device 900 may specifically be a mobile terminal / terminal device of an embodiment of the present application, and the electronic device 900 may implement the corresponding processes implemented by the mobile terminal / terminal device in each method of the embodiment of the present application. For the sake of brevity, they will not be repeated here.
[0156] It should be understood that the processor of the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by hardware integrated logic circuits in the processor or software instructions. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly implemented as a hardware decoding processor, or can be implemented by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0157] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0158] It should be understood that the above-mentioned memories are exemplary but not restrictive. For example, the memories in the embodiments of the present application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM RAM (DR RAM), etc. In other words, the memories in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0159] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program.
[0160] Optionally, the computer-readable storage medium can be applied to the network device in the embodiments of the present application, and the computer program enables the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0161] Optionally, the computer-readable storage medium can be applied to the mobile terminal / terminal device in the embodiments of the present application, and the computer program enables the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0162] An embodiment of the present application also provides a computer program product, including computer program instructions.
[0163] Optionally, the computer program product can be applied to the network device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0164] Optionally, the computer program product can be applied to the mobile terminal / terminal device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0165] The embodiment of the present application also provides a computer program.
[0166] Optionally, the computer program can be applied to the network device in the embodiments of the present application. When the computer program runs on a computer, the computer executes the corresponding processes implemented by the network device in the various methods of the embodiments of the present application. For the sake of brevity, they are not described here.
[0167] Optionally, the computer program can be applied to the mobile terminal / terminal device in the embodiments of the present application. When the computer program runs on the computer, the computer executes the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0168] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0169] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0170] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0171] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0172] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0173] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0174] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Obtain image data of facial expressions; extracting a first expression feature and a first identity feature from the image data; performing feature concatenation on the first expression feature and the first identity feature to obtain a first joint feature, and extracting joint information between the first expression feature and the first identity feature from the first joint feature; estimating mutual information between the first facial expression feature and the first identity feature, and adjusting the first facial expression feature to a second facial expression feature with the goal of minimizing the mutual information; A classification result of the facial expression is predicted based on the second expression feature and the joint information.
2. The method according to claim 1, characterized in that The extracting the first expression feature and the first identity feature from the image data includes: extracting a first expression feature from the image data by an expression encoder; A first identity feature is extracted from the image data by an identity encoder; wherein the identity encoder is pre-trained.
3. The method according to claim 2, characterized in that The estimating the mutual information between the first expression feature and the first identity feature includes: Calculating the joint distribution and marginal distribution between the first expression feature and the first identity feature by a mutual information estimator to obtain the joint distribution expectation and the marginal distribution logarithmic expectation; Based on the joint distribution expectation and the marginal distribution logarithmic expectation, the mutual information between the first expression feature and the first identity feature is calculated.
4. The method according to claim 3, characterized in that Before calculating the joint distribution and marginal distribution between the first expression feature and the first identity feature by the mutual information estimator, the method further includes: Acquire a joint distribution sample pair, the joint distribution sample pair comprising a first expression feature sample and a first identity feature sample; Randomly arranging the first identity feature samples to obtain second identity feature samples, wherein the second identity feature samples include identity feature information unrelated to the first expression feature samples; A marginal distribution sample pair is generated based on the first expression feature sample and the second identity feature sample, where the marginal distribution sample pair includes the first expression feature sample and the second identity feature sample.
5. The method according to claim 4, characterized in that The calculating the joint distribution and marginal distribution between the first expression feature and the first identity feature by a mutual information estimator to obtain the joint distribution expectation and the marginal distribution logarithmic expectation includes: Calculating the joint distribution of the joint distribution sample pairs by the mutual information estimator to obtain a joint distribution score value; Calculating the marginal distribution of the marginal distribution sample pair by the mutual information estimator to obtain a marginal distribution score value; Performing expectation estimation processing on the joint distribution score value to obtain the joint distribution expectation; An expectation estimation process is performed on the marginal distribution score value to obtain the marginal distribution logarithmic expectation.
6. The method according to claim 4, characterized in that The estimating the mutual information between the first expression feature and the first identity feature, and adjusting the first expression feature to a second expression feature with the goal of minimizing the mutual information, includes: The second expression feature is obtained by minimizing the joint distribution expectation and the marginal distribution logarithmic expectation based on the first expression feature and the first identity feature through the mutual information estimator; wherein the identity feature information contained in the second expression feature is less than the identity feature information contained in the first expression feature.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Determine the first loss function based on the variational lower bound of mutual information; determining a second loss function based on the facial expression classification result and the facial expression classification label; Summing or weighted summing the first loss function and the second loss function to obtain a target loss function; The expression encoder and the mutual information estimator are updated based on the target loss function.
8. An expression classification model optimization device, characterized in that: The device is applied to electronic equipment, and includes: An acquisition module, used to acquire image data of facial expressions; An extraction module, configured to extract a first expression feature and a first identity feature from the image data; a processing module, configured to perform feature concatenation on the first expression feature and the first identity feature to obtain a first joint feature; The extraction module is further configured to extract joint information between the first expression feature and the first identity feature from the first joint feature; The processing module is also used to estimate the mutual information between the first expression feature and the first identity feature, and adjust the first expression feature to the second expression feature with the goal of minimizing the mutual information; and to predict the classification result of the facial expression based on the second expression feature and the joint information.
9. An electronic device, characterized in that: include: a memory for storing computer executable instructions or computer programs; The processor is configured to implement the image processing method according to any one of claims 1 to 7 when executing the computer-executable instructions or computer programs stored in the memory.
10. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the image processing method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Facial expression recognition method based on joint learning identity information and emotion information
CN109359599A
Cited By
Precision sensitive fine granularity identification method based on Posit driving symbol mode shaping super-dimensional calculation
CN122133515A
Precision-sensitive fine-grained recognition method based on posit-driven symbolic pattern shaping super-dimensional computation
CN122133515B