Facial expression recognition method and device of image, equipment and medium
By acquiring the natural state and current state facial images of the target object, and using a deep learning model to extract feature maps at multiple spatial scales and fuse the difference data, the problem of inaccurate prediction of single-frame images is solved, and high-precision facial expression recognition is achieved.
Patent Information
- Application Number
- CN202511062022.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-28
AI Technical Summary
Existing facial expression prediction methods based on single-frame images are affected by the facial expression baseline, resulting in inaccurate predictions and failing to meet the high-precision requirements of practical applications.
By acquiring the target face image in its natural state and the test face image in its current state, a pre-trained deep learning model is used to extract feature maps at multiple spatial scales, determine facial dynamic difference data, and fuse the difference data to generate fused features, ultimately obtaining the facial expression type.
It significantly improves the accuracy of facial expression prediction, and can accurately distinguish between facial expression features in an individual's natural state and facial expression features caused by emotional changes, meeting the high precision requirements of practical applications.
Smart Images

Figure CN121033907A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to online processing business scenarios such as finance, insurance, and healthcare. In particular, it relates to a method, device, equipment, and medium for facial expression recognition of images. Background Technology
[0002] In the cutting-edge and highly valuable field of facial expression analysis, facial expression prediction has attracted considerable research attention and demonstrated enormous development potential due to its wide range of applications. For example, in face-to-face interviews, accurate monitoring of the emotions of the parties involved can help interrogators better grasp their psychological state, improving interrogation efficiency and quality. In customer satisfaction surveys, analyzing customer expressions can provide more realistic and intuitive feedback, helping companies optimize services and products. In the field of mental health analysis and treatment, expression analysis provides crucial evidence for psychologists to diagnose conditions and assess treatment effectiveness. Facial expression analysis also has significant application value in the financial insurance sector. In the insurance claims process, claims personnel can analyze the insured's expressions to determine the truthfulness and credibility of their account of the incident, effectively preventing insurance fraud and reducing the operational risks of insurance companies. In the field of financial investment, analyzing investors' facial expressions can reveal their emotional reactions to market conditions, providing a reference for investment decisions and helping investors better seize market opportunities and avoid investment risks. In the medical field, for patients who cannot accurately express their feelings verbally, such as comatose patients, infants, or patients with speech disorders, facial expression analysis can help medical staff understand their pain levels, comfort levels, and other physical conditions, so as to adjust treatment plans in a timely manner, provide more attentive care, and improve the quality of medical care.
[0003] Facial expression baselines, representing an individual's facial expressions in a natural, stress-free, or habitual state, play a crucial role in facial expression analysis. They serve as an important reference standard for judging the authenticity of emotions and recognizing micro-expressions, providing a benchmark framework for accurately interpreting facial expressions. However, current industry products have significant shortcomings in facial image expression detection. Existing technologies primarily utilize single-frame images to predict facial expressions, a method with significant drawbacks. Because single-frame images cannot fully account for the influence of the facial expression baseline, it is difficult to accurately distinguish between facial expression features in a natural state and those arising from emotional changes during the prediction process. This leads to inaccurate prediction results, failing to meet the high-precision requirements of facial expression analysis in practical applications. Summary of the Invention
[0004] The purpose of this application is to propose a method, apparatus, computer device, and storage medium for facial expression recognition of images, in order to solve the problem that existing facial expression prediction methods based on single-frame images are affected by the facial expression baseline, resulting in inaccurate predictions and an inability to provide reliable expression analysis results for various application scenarios.
[0005] Firstly, a method for facial expression recognition in images is provided, which employs the following technical solution: The process involves acquiring the target face image and the test face image, where the target face image is the face of the target object in its natural state. A pre-trained deep learning model is used to extract features from both the target and test face images, resulting in feature maps at multiple spatial scales. Based on these feature maps, the facial dynamic difference data between the target and test face images is determined. The facial dynamic difference data is then fused to generate the fused features of the target object. Finally, feature processing is performed on the fused features and the features of the test face image to obtain the facial expression type of the test face image.
[0006] Secondly, a facial expression recognition device for images is provided, which adopts the following technical solution: The acquisition module is used to acquire the target face image and the face image to be tested of the target object. The target face image is the face image of the target object in its natural state. The extraction module is used to extract features from the target face image and the face image to be tested using a pre-trained deep learning model, and obtain feature maps at multiple spatial scales. The determination module is used to determine the facial dynamic difference data between the target face image and the face image to be tested based on the feature map; The fusion module is used to fuse facial dynamic difference data to generate fused features of the target object; The processing module is used to perform feature processing on the fused features and the features of the face image to be tested, so as to obtain the facial expression type of the face image to be tested.
[0007] Thirdly, a computer device is provided, which adopts the following technical solution: The process involves acquiring the target face image and the test face image, where the target face image is the face of the target object in its natural state. A pre-trained deep learning model is used to extract features from both the target and test face images, resulting in feature maps at multiple spatial scales. Based on these feature maps, the facial dynamic difference data between the target and test face images is determined. The facial dynamic difference data is then fused to generate the fused features of the target object. Finally, feature processing is performed on the fused features and the features of the test face image to obtain the facial expression type of the test face image.
[0008] Fourthly, a computer-readable storage medium is provided, which adopts the following technical solution: The process involves acquiring the target face image and the test face image, where the target face image is the face of the target object in its natural state. A pre-trained deep learning model is used to extract features from both the target and test face images, resulting in feature maps at multiple spatial scales. Based on these feature maps, the facial dynamic difference data between the target and test face images is determined. The facial dynamic difference data is then fused to generate the fused features of the target object. Finally, feature processing is performed on the fused features and the features of the test face image to obtain the facial expression type of the test face image.
[0009] Compared with existing technologies, the embodiments of this application have the following main advantages: By acquiring the target face image in its natural state and the test face image in its current state, and using a pre-trained deep learning model to extract feature maps at multiple spatial scales, it can comprehensively and accurately capture feature information at different levels of the face, fully considering the changes in facial expressions in different spatial dimensions. Based on these feature maps, facial dynamic difference data is determined, which can effectively distinguish between facial expression features in an individual's natural state and facial expression features caused by emotional changes, overcoming the defect of inaccurate prediction due to the influence of the facial expression baseline in existing single-frame image judgments. The dynamic difference data is fused to generate fused features, further integrating key information and enhancing the representational ability of the features. Finally, the fused features and the test face image features are processed to obtain the facial expression type, which can significantly improve the accuracy of facial expression prediction, providing reliable expression analysis results for many application scenarios and meeting the high-precision requirements of practical applications. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 A flowchart of an embodiment of the facial expression recognition method for images according to this application; Figure 3 This is a schematic diagram of the structure of an embodiment of a facial expression recognition device based on images according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0019] It should be noted that the facial expression recognition method for images provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the facial expression recognition device for images is generally set in the server / terminal device.
[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0021] Continue to refer to Figure 2 A flowchart of an embodiment of a facial expression recognition method for images according to this application is shown. The facial expression recognition method for images includes the following steps: Step S201: Obtain the target face image and the face image to be tested of the target object. The target face image is the face image of the target object in its natural state.
[0022] The target group refers to the individuals whose facial expressions need to be detected and identified in the facial expression analysis process. These individuals come from a wide range of sources, such as parties involved in face-to-face interviews, customers in customer satisfaction surveys, patients in the field of mental health analysis and treatment, insured persons and investors in financial insurance scenarios, and comatose patients, infants, or patients with speech disorders in medical scenarios.
[0023] The target facial image is a photograph of the target subject's face in a natural, stress-free, or habitual state. It characterizes the facial expression features of the target subject in a natural state, providing an important benchmark for facial expression analysis. This benchmark is used for subsequent comparative analysis with the test facial image to accurately determine the target subject's current emotional state. For example, in mental health analysis, the target facial image is a photograph of the patient's face taken in a calm and relaxed state, serving as a benchmark for judging their emotional changes.
[0024] Among them, the face image to be tested is the face image of the target object in the current state. It can be obtained by real-time shooting or by collecting the face information of the target object at a specific moment. It represents the facial expression features of the target object at the current moment and together with the target face image, it constitutes the basic data for analysis. By comparing the differences between the two, the emotional changes of the target object in the current state can be accurately identified.
[0025] Step S202: Using a pre-trained deep learning model, feature extraction is performed on the target face image and the face image to be tested to obtain feature maps at multiple spatial scales.
[0026] The deep learning model is a neural network model pre-trained on a large amount of facial image data. It is used to extract features from target facial images and test facial images. Through the structure of a multi-layer neural network, it deeply mines key information in facial images, improving the accuracy and comprehensiveness of feature extraction.
[0027] Feature extraction is the process of using deep learning models to process target face images and test face images to extract representative and discriminative feature information from the images.
[0028] Among them, multiple spatial scales refer to the analysis and processing of face images from different sizes and resolutions during the feature extraction process. By decomposing and sampling face images at different levels, feature information at different scales can be obtained, which can comprehensively capture the detailed features and overall structural features of face images.
[0029] Step S203: Based on the feature map, determine the facial dynamic difference data between the target face image and the face image to be tested.
[0030] Among them, the feature map is the image representation obtained by the deep learning model after extracting features from the target face image and the face image to be tested. It can represent the feature information of the face image at different spatial scales in the form of a matrix, and the value of each element reflects the feature intensity at the corresponding position.
[0031] Facial dynamic difference data, specifically, is data that reflects changes in facial expressions of the target subject, calculated using a specific algorithm based on the feature maps of the target face image and the test face image. It is used to determine the current emotional state of the target subject; for example, facial dynamic difference data is obtained by calculating the difference between corresponding elements in the feature maps of the target and test face images.
[0032] Step S204: Fuse the facial dynamic difference data to generate the fused features of the target object.
[0033] Fusion is the process of integrating facial dynamic difference data. It can be achieved by comprehensively analyzing different facial dynamic difference data and merging them into a unified feature representation.
[0034] Among them, the fusion feature is the feature representation of the target object generated after fusion processing. The fusion feature integrates multiple aspects of the target object's facial expression information in its natural state and current state, and has stronger representation ability and discriminative power, and can more accurately reflect the facial expression features of the target object.
[0035] Step S205: Perform feature processing on the fused features and the features of the face image to be tested to obtain the facial expression type of the face image to be tested.
[0036] Feature processing involves further processing and analyzing the fused features and the features of the face image to be tested. Operations such as feature fusion, pooling, and classification score calculation can extract more valuable information from the features, improving the accuracy and reliability of facial expression type judgment.
[0037] Facial expression type is the result of classifying and judging the current facial expression of the target subject. It can represent the target subject's current emotional state, such as happiness, sadness, anger, surprise, etc. By accurately judging the facial expression type, valuable reference information can be provided for many application scenarios such as face-to-face interviews, customer satisfaction surveys, mental health analysis and treatment, finance and insurance, and medicine.
[0038] In one example, firstly, when the target subject is in a natural, relaxed state, a high-definition camera is used to capture an image of their face, ensuring the image is clear, unobstructed, and fully presents facial features. The target subject's face image in its current state is captured in real-time. The acquired target and test face images are then input into a pre-trained deep learning model. Through model processing, feature maps at multiple spatial scales are obtained, reflecting facial feature information at different levels and resolutions. Based on the feature maps, a specific algorithm is used to calculate the facial dynamic difference data between the target and test face images. For example, by comparing changes in the shape and position of key features such as the eyes and mouth, expression differences are quantified. The obtained facial dynamic difference data is then fused, and a weighted average method is used to generate a fused feature, which integrates information from multiple aspects of expression changes. Next, the fused feature is fused with the features of the test face image to obtain the target fused feature. A pre-defined pooling layer is used to globally pool the target fused feature, reducing the feature dimensionality and obtaining a global feature vector. Then, a pre-defined fully connected layer is used to calculate the classification score of the global feature vector, outputting the facial expression type of the test face image, such as nervousness, calmness, or lying. By fully considering the facial expression baseline, the system accurately distinguishes the expression features of the target object in its natural state and current state, effectively improving the accuracy of expression prediction.
[0039] This application embodiment can acquire a target face image in its natural state and a test face image in its current state, and use a pre-trained deep learning model to extract feature maps at multiple spatial scales. This comprehensively and accurately captures feature information at different levels of the face, fully considering the changes in facial expressions across different spatial dimensions. Based on these feature maps, dynamic facial difference data is determined, effectively distinguishing between facial expression features in an individual's natural state and those resulting from emotional changes. This overcomes the shortcomings of existing single-frame image judgments, which are prone to inaccurate predictions due to the influence of the facial expression baseline. The dynamic difference data is fused to generate fused features, further integrating key information and enhancing the feature representation capability. Finally, the fused features and the test face image features are processed to derive the facial expression type, significantly improving the accuracy of facial expression prediction. This provides reliable expression analysis results for numerous application scenarios, meeting the high-precision requirements of practical applications.
[0040] In some optional implementations of this embodiment, step 202, using a pre-trained deep learning model, extracts features from the target face image and the face image to be tested, obtaining feature maps at multiple spatial scales, specifically including the following steps: The target face image and the face image to be tested are input into a pre-trained deep learning model. Features are extracted through multiple hidden layers and output layers of the deep learning model to generate multiple initial feature maps at different spatial scales. The initial feature maps corresponding to the target face image and the face image to be tested are compared to determine the feature differences at each spatial scale. Based on the feature differences, the initial feature maps are adjusted to obtain feature maps at multiple spatial scales.
[0041] Hidden layers are a crucial component of pre-trained deep learning models, located between the input and output layers. They are used for deep processing of the input target and test face images, extracting multi-level facial features from their pixel information. For example, the first hidden layer can extract low-level features such as edges and textures, while subsequent hidden layers can extract high-level features related to facial organ shapes and expressions as the layers deepen.
[0042] The output layer is the last layer of the pre-trained deep learning model. It originates from the top-level design of the entire model architecture and is located at the end of the model processing flow. It is used to integrate and transform the feature information obtained from multiple hidden layers to generate initial feature maps of multiple different spatial scales.
[0043] The initial feature map is an image representation generated by inputting the target face image and the face image to be tested into a pre-trained deep learning model, and then processing them through multiple hidden layers and output layers of the model. The initial feature map represents the preliminary feature information of the face image at different spatial scales and can be presented in the form of a matrix, where the value of each element reflects the feature intensity at the corresponding location. Initial feature maps at different spatial scales can display the features of the face from different resolutions and levels.
[0044] The comparison process involves analyzing and comparing the initial feature maps corresponding to the target face image and the test face image. This is used to determine feature differences at various spatial scales. For example, it quantifies the degree of difference in facial features under different conditions by calculating the difference between corresponding pixels in the two initial feature maps, the correlation coefficient, or using more complex similarity metrics.
[0045] Feature difference is the result data obtained by comparing the initial feature maps corresponding to the target face image and the test face image. Feature difference characterizes the specific changes in facial features of the target object in its natural state and current state, including increases or decreases in feature values, changes in feature shape, and changes in feature distribution. This difference data can reflect information such as changes in the target object's emotions and expressions. For example, changes in the degree of eye opening and closing, and the degree of upward or downward slant of the corners of the mouth.
[0046] The adjustment process involves modifying and optimizing the initial feature map based on the identified feature differences. Through specific algorithms and rules, the adjustment process alters the element values in the initial feature map according to the magnitude and direction of the feature differences. For example, it may increase or decrease the feature intensity of certain feature regions or adjust the distribution of features.
[0047] In one example, in an insurance claims scenario, the target is an insured person applying for vehicle insurance compensation. First, when the insured person is in a natural, relaxed state, such as during a normal consultation with the insurance company, a high-definition camera is used to capture their face, ensuring the image is clear and fully represents their facial features. Then, while the insured person is recounting the accident and making the claim, another face image is captured in real-time. These two images are then input into a pre-trained deep learning model, which is trained on a large dataset of facial images with different expressions from various insurance claims scenarios. The model extracts features from the images through multiple hidden layers, such as convolutional and pooling layers. The hidden layers progressively extract low-level to high-level features from the images, including edges, textures, and organ shapes. The output layer then generates multiple initial feature maps at different spatial scales. Large-scale feature maps reflect the overall facial contours, while small-scale feature maps capture subtle changes in facial expressions. Next, the initial feature maps corresponding to the target face image and the test face image are compared. The method of calculating the pixel differences at corresponding locations in the feature maps is used to determine the feature differences at various spatial scales. For example, feature differences in the area around the eyes may reflect the level of tension the insured is experiencing during their statement. Based on these feature differences, the initial feature maps are adjusted to strengthen features related to emotional changes and weaken irrelevant features, resulting in more accurate feature maps at multiple spatial scales.
[0048] In one example, within a medical setting, a child with a speech impairment who cannot accurately express their feelings is used as the target subject. When the child is quiet and natural, such as asleep, a professional medical camera is used to capture their face. When medical staff examine or treat the child, and the child's facial expressions change due to pain or discomfort, a test face image is captured in real time. These two images are then fed into a pre-trained deep learning model, which has been trained on a large number of facial images of children with different expressions in medical scenarios. The model's multiple hidden layers perform multi-level feature extraction on the images, and the output layer generates multiple initial feature maps at different spatial scales. The initial feature maps corresponding to the target face image and the test face image are then compared. By analyzing the differences in facial muscle movements, changes in facial features, etc., in the feature maps, the differences in features at each spatial scale are determined, such as the degree of mouth opening and the degree of eyebrow furrowing. Based on these differences, the initial feature maps are adjusted to obtain feature maps that accurately reflect the child's current state.
[0049] This application embodiment can input the target face image and the test face image into a pre-trained deep learning model. Utilizing its multiple hidden layers, it can perform multi-level, progressive feature mining on the image. From low-level edge and texture information to high-level facial organ shapes and expression-related features, it progressively extracts and generates multiple initial feature maps at different spatial scales via the output layer, comprehensively and meticulously representing the features of the face in different dimensions. By comparing the initial feature maps corresponding to the target and test face images, the feature differences at each spatial scale can be accurately located. Whether it's changes in the overall facial contour or subtle movements of local facial features, they can all be clearly captured. Adjusting the initial feature maps based on feature differences can eliminate interference caused by differences between the natural state and the current state, making the resulting feature maps at multiple spatial scales more closely match the current real expression features of the target object. This series of operations works together to improve the accuracy and comprehensiveness of facial expression feature extraction.
[0050] In some optional implementations, before step S202, which uses a pre-trained deep learning model to extract features from the target face image and the face image to be tested, and obtains feature maps at multiple spatial scales, the following steps are also included: Obtain the base loss function, which includes the difference between the label value and the predicted value; obtain the training set, and determine the weight of each type of expression by calculating the distribution ratio of each type of expression in the training set; according to the weight, weight the difference term to obtain the weighted loss function; use the weighted loss function to iteratively train the preset initial deep learning model to obtain the pre-trained deep learning model.
[0051] The base loss function is a mathematical expression used to measure the degree of difference between the model's predictions and the actual situation. This function reflects this difference by constructing a relationship between the label values and the predicted values.
[0052] The label value represents the actual expression category of the target object, which is a known and accurate piece of information. The predicted value represents the model's judgment on the expression category of the input data based on the current parameters.
[0053] The difference term comes from the quantitative description of the difference between the label value and the predicted value in the basic loss function, and it represents the degree of deviation between the model prediction result and the true label.
[0054] The training set refers to the collected and organized facial image data with labeled values, which represents the data set used to train the model and includes facial image samples with various expressions, angles, and lighting conditions.
[0055] The various types of facial expressions are derived from the classification and definition of human facial expressions, which represent the different appearance patterns of the face in different emotional states. These include happiness, sadness, anger, surprise, fear, disgust, and neutrality.
[0056] The distribution ratio represents the proportion of different expression types in the training set. By analyzing the distribution ratio, we can understand the distribution of various expressions in the training set, providing a basis for subsequent adjustments to the model training strategy.
[0057] The weights represent the degree of importance given to the differences between different types of facial expressions during model training. For example, assigning larger weights to facial expression types with a smaller distribution ratio can make the model pay more attention to these relatively few samples during training, avoiding inaccurate recognition of certain expressions due to an imbalance in the number of samples.
[0058] The weighted loss function is derived from an improvement on the basic loss function. It is obtained by weighting the difference terms in the basic loss function according to the weights of each type of expression. The weighted loss function allows the model to give different levels of attention to different expression types during training.
[0059] The initial deep learning model represents the starting state of the model training and possesses a certain framework for feature extraction and classification. It can include structures such as convolutional layers, pooling layers, and fully connected layers.
[0060] Iterative training refers to the process of repeatedly inputting training set data into the model, calculating the loss function, and updating the model parameters based on the loss function.
[0061] In one example, within a customer risk assessment scenario in the financial insurance industry, it's crucial to accurately identify a customer's true emotions during communication to determine the credibility of their statements. First, a base loss function is obtained. The label values in this function are the true expression categories determined through precise manual annotation of a large number of historical customer expressions, such as "relaxed," "nervous," and "anxious." The predicted values are the initial judgments of the model on the expressions in the input customer facial images. A difference term quantifies the deviation between the label values and the predicted values. For example, a cross-entropy loss function is used to calculate the difference. Next, a training set is obtained, collecting facial images of customers of different ages, genders, and regions during insurance transactions, and labeling them with their true expressions, forming a training set with rich samples. The distribution ratio of each expression type in the training set is calculated by statistically analyzing the sample size, such as "nervous" expressions accounting for 20% of the total samples and "relaxed" expressions accounting for 60%. Based on these distribution ratios, expression types with fewer samples are assigned larger weights, such as assigning a weight of 0.3 to "nervous" and 0.1 to "relaxed." Then, the difference term is weighted according to these weights to obtain a weighted loss function. This weighted loss function is used to iteratively train a pre-set initial deep learning model (such as a convolutional neural network model). In each iteration, the model predicts images in the training set based on the current parameters, calculates the weighted loss function value, and updates the model parameters through backpropagation, gradually reducing the loss function value. After multiple iterations of training, a pre-trained deep learning model is finally obtained.
[0062] In this embodiment, during model training, a fundamental loss function containing the difference between label and predicted values provides a key quantitative standard for measuring the accuracy of model predictions. This clearly reflects the degree to which predicted values deviate from the true labels, guiding the direction of model optimization. Weights are determined by acquiring the training set and calculating the distribution ratio of each type of expression, fully considering the frequency differences of different expressions in the dataset. Larger weights are assigned to expression types that appear less frequently in the training set and have a low distribution ratio, preventing the model from ignoring these important expression features due to data imbalance. A weighted loss function is obtained by weighting the difference terms according to the weights, allowing the model to more reasonably focus on various expressions during training. This weighted loss function is used to iteratively train the initial deep learning model. In each iteration, the model continuously adjusts its parameters based on the weighted loss, gradually improving its ability to recognize various expressions, ultimately resulting in a pre-trained deep learning model that can more accurately and comprehensively recognize different expressions, especially those that originally had a small proportion in the dataset.
[0063] In some optional implementations, step S203, determining the facial dynamic difference data between the target face image and the face image to be tested based on the feature map, specifically includes the following steps: The system calculates the difference between the feature maps of the target face image and the face image to be tested at each spatial scale; it normalizes the difference values using a preset difference calculation formula to obtain standardized difference information; it determines the dynamic change features of the feature maps at each spatial scale based on the difference information; and it performs weighted processing on the dynamic change features to generate facial dynamic difference data between the target face image and the face image to be tested.
[0064] The difference value represents the degree of deviation between corresponding features in two feature maps at the same spatial scale. During feature extraction, deep learning models capture feature information from face images at different spatial scales. Since the target face image is an image of the object in its natural state, while the test face image is an image of the current state, differences such as facial expressions lead to different feature maps. Calculating the difference value quantitatively measures the magnitude of this difference, providing fundamental data for subsequent analysis of facial dynamics. For example, at a certain spatial scale, if the feature value at a certain location in the target face image feature map is 0.5, and the corresponding feature value in the test face image is 0.8, the difference of 0.3 is the difference value at that location.
[0065] The label value accurately represents the facial expression category to which the face image belongs. During model training and expression recognition, the label value serves as a standard reference to measure the accuracy of the model's predictions. For example, labeling a face image showing a distinctly upturned mouth and slightly squinted eyes as a "happy" expression would be the label value corresponding to that image, used to guide the model's learning and judgment.
[0066] The difference calculation formula refers to the rules used to quantify and transform the difference values to obtain standardized results.
[0067] Normalization refers to the process of mapping difference values to a specific interval (such as [0, 1]) according to certain rules.
[0068] Among them, the difference information refers to the standardized data that accurately reflects the differences between the feature maps of the target face image and the face image to be tested.
[0069] In one example, in a financial insurance customer identity verification and risk assessment scenario, it is often necessary to judge the changes in a customer's facial expressions during communication to assess the credibility of their statements. First, obtain the target face image in a natural state (e.g., during initial registration) and the test face image during the current transaction. Using a deep learning model, feature extraction is performed on both the target and test face images to obtain feature maps at multiple spatial scales, such as local detail feature maps and overall contour feature maps. Next, the difference between the two feature maps at each spatial scale is calculated. For example, in the local detail feature map, comparing the feature differences of parts such as the eyes and mouth, if the eye feature value in the target face image is 0.3 and the corresponding position in the test face image is 0.6, then the difference value is 0.3. Then, using a preset difference calculation formula, the difference value is normalized to obtain standardized difference information, mapping the difference value to the [0, 1] interval. Based on the difference information, the dynamic change characteristics of the feature maps at each spatial scale are determined. For example, the difference information in the overall contour feature map shows that the facial contour gradually tightens, indicating that the customer may be in a tense state. Finally, the dynamic change features are weighted. Considering that local detail features are more critical for expression judgment, they are assigned a weight of 0.7, while the overall contour features are assigned a weight of 0.3. This generates facial dynamic difference data between the target face image and the face image to be tested, thereby helping to determine the authenticity of the customer's statement and reduce the risk of insurance business.
[0070] This application embodiment can accurately locate the deviation of facial features at different scales by calculating the difference value between the feature maps of the target and the tested face image at each spatial scale. Since the target face image is in its natural state and the tested image is in its current state, there are differences such as expressions between the two. This calculation can comprehensively capture feature changes from subtle to overall, providing a detailed data foundation for subsequent analysis. The difference calculation formula is used to normalize the difference values, which can eliminate the influence of differences in the range and dimensions of feature values at different spatial scales, making the difference information have a unified standard, facilitating comparison and analysis, and ensuring that the degree of feature difference can be accurately measured at different scales. Based on the difference information, the dynamic change features of the feature map at each spatial scale can be determined, which can deeply explore the laws of facial expression changes and clearly present the dynamic pattern of feature changes at different scales with expression. Weighting the dynamic change features can highlight the influence of key scale features on the final result. By reasonably allocating weights according to the importance of different scales in expression recognition, data that accurately reflects the dynamic differences between the target and the tested face images can be generated, effectively improving the accuracy and reliability of facial expression recognition.
[0071] In some alternative implementations, the step "determine the dynamic change characteristics of the feature maps at each spatial scale based on the difference information" specifically includes the following steps: Based on the scale parameters of each spatial scale, the difference information is divided to obtain the distribution data of the difference information at each spatial scale; using the preset analysis tools, the change characteristics of the distribution data are extracted to determine the dynamic change characteristics of the feature map at each spatial scale.
[0072] The scale parameter is derived from the quantitative definition of the spatial scale of the feature map. In facial image feature analysis, feature maps of different spatial scales capture facial information of different granularities, and the scale parameter is used to accurately describe these different scales. For example, when extracting facial features, a 32×32 pixel feature extraction window is set as the scale parameter, and the feature map extracted based on this window corresponds to a specific spatial scale.
[0073] In this context, distributed data refers to the results obtained by dividing the normalized differential information according to the scale parameters of each spatial scale. It characterizes the specific distribution of differential information at different spatial scales.
[0074] In this context, analytical tools refer to a set of software or algorithms specifically designed for processing and analyzing distributed data. They are used to uncover the changing characteristics within the distributed data in order to determine the dynamic patterns of the feature maps.
[0075] The change characteristics refer to how the distribution data changes under different conditions or over time. In facial dynamic difference analysis, the distribution data changes with the facial expressions of the target face image and the face image being tested. For example, analysis of the distribution data reveals that at a small spatial scale, the distribution of difference values gradually shifts towards higher value regions as the expression changes from calm to smiling; this shifting pattern is the extracted change characteristic.
[0076] Among them, dynamic change features can accurately describe the feature information of feature maps changing with facial expressions at various spatial scales, and it represents the dynamic evolution law of feature maps under different states.
[0077] In one example, in a remote customer identity verification and risk assessment scenario in the financial insurance industry, it is often necessary to analyze a customer's facial expressions to determine the veracity of their statements. First, the system acquires the target's face image in its natural state during initial registration, as well as the image of the face to be tested during the current remote verification. After extracting feature maps using a deep learning model and calculating and normalizing the values, standardized difference information is obtained. Next, the difference information is segmented according to scale parameters at various spatial scales. For example, small-scale parameters correspond to local facial details (such as the corners of the eyes and mouth), while large-scale parameters correspond to the overall facial contour. This segmentation reveals that at small scales, the difference information is concentrated in areas such as upward slant of the eyes and slight movement of the mouth, while at large scales, the difference information is reflected in changes in overall facial relaxation. Then, a pre-set analysis tool is used, based on a machine learning model trained on a large amount of facial expression data from financial scenarios. The distributed data is input into the analysis tool to extract change features. For example, at small scales, it is found that the difference value of upward slant of the eyes shows a fluctuating upward trend over time, while the difference value of slight movement of the mouth shows a significant peak after a specific question. At a large scale, the overall facial relaxation difference gradually decreases when presenting key information. Finally, based on these changes, the dynamic change characteristics of the feature maps at each spatial scale are determined. Considering the importance of different scales, a weight of 0.7 is assigned to the small scale (because it can capture subtle facial expressions and is crucial for judging authenticity), and a weight of 0.3 is assigned to the large scale. This weighted processing generates facial dynamic difference data to help assess the credibility of customer statements and reduce insurance business risks.
[0078] This application's embodiments divide the difference information according to the scale parameters of each spatial scale, accurately classifying the difference information according to different spatial ranges and feature granularities. Since different spatial scales focus on different aspects of the facial image, small scales can capture subtle local feature differences, while large scales can grasp overall contour changes. This division clearly presents the distribution of difference information at each scale, providing a detailed and targeted data foundation for subsequent analysis. Using analytical tools to extract the variation characteristics of the distributed data allows for in-depth exploration of the underlying patterns. Based on this, determining the dynamic change characteristics of the feature map at each spatial scale comprehensively and accurately reflects the dynamic process of facial features changing with expression or state, considering not only local detail changes but also overall feature evolution, providing a reliable basis for generating dynamic facial difference data.
[0079] In some optional implementations, step S204, fusing facial dynamic difference data to generate fused features of the target object, specifically includes the following steps: Facial dynamic difference data is input into multiple convolutional layers for feature transformation to obtain multiple transformed difference features; these multiple difference features are then concatenated to generate the fusion features of the target object.
[0080] Multiple convolutional layers refer to a set of core components used for feature extraction and transformation of input data. They are used to progressively transform and abstract the features of the input facial dynamic difference data. For example, in a deep learning model for facial expression recognition, there are five convolutional layers. The first convolutional layer can extract low-level features such as edges and textures from the facial dynamic difference data, while subsequent convolutional layers further extract higher-level and more abstract features, such as the shape change patterns of facial organs.
[0081] Feature transformation refers to the operation of mapping facial dynamic difference data from one feature space to another. It is used to uncover more discriminative features hidden within facial dynamic difference data.
[0082] Among these, multiple differential features are a set of feature vectors with specific meanings and discriminative power output from each convolutional layer. This feature information reflects the manifestation of facial dynamic differences in various aspects, providing a comprehensive and detailed description of all facets of dynamic differences.
[0083] Here, concatenation refers to the operation of merging multiple vectors or matrices with different feature information according to specific rules. It is used to integrate multiple differential features output by multiple convolutional layers to generate fused features of the target object.
[0084] In one example, within the financial insurance business, to accurately assess a customer's emotional state during insurance application to prevent fraud risk, the technical solution of this embodiment is adopted. First, the target face image of the customer in their natural state during insurance consultation, and the test face image when currently answering key questions, are acquired. Using a deep learning model, features are extracted from these two images to obtain feature maps at multiple spatial scales, thereby determining facial dynamic difference data. Next, the facial dynamic difference data is input into multiple convolutional layers. For example, three convolutional layers are set. The first convolutional layer uses a 3×3 kernel to capture subtle dynamic changes in local facial areas, such as twitches in the corners of the eyes and mouth. The second convolutional layer uses a 5×5 kernel to focus on the overall dynamic trends of larger facial areas, such as changes in the degree of relaxation of cheek muscles. The third convolutional layer uses a 7×7 kernel to extract global facial dynamic features, such as the coordinated changes in subtle head movements and facial expressions. After feature transformation by these three convolutional layers, multiple transformed difference features are obtained. Then, the multiple difference features output by these three convolutional layers are concatenated. Following a process from local to global, the subtle local differences from the first layer, the overall regional differences from the second layer, and the dynamic global differences from the third layer are sequentially concatenated to generate a fused feature of the target object. This fused feature comprehensively and meticulously reflects the dynamic changes of the customer's face at different levels, providing a reliable basis for accurately determining the type of facial expression the customer uses when answering key questions, thereby assessing their emotional state and the authenticity of the insurance application, and effectively reducing the risks in financial and insurance business.
[0085] This application's embodiments transform facial dynamic difference data by inputting it into multiple convolutional layers. Because different convolutional layers have kernels of different scales and unique computational logic, they can deeply mine feature information from multiple dimensions within the facial dynamic difference data. Small-scale convolutional kernels can capture subtle local dynamic changes on the face, such as blinking frequency and the amplitude of mouth twitching. Large-scale convolutional kernels can grasp the overall dynamic trend of the face, such as changes in facial contours, subtle head movements, and the coordination of facial expressions. Through processing by multiple convolutional layers, multiple transformed difference features are obtained, which are comprehensive and hierarchical. These multiple difference features are then concatenated to generate a fusion feature for the target object. This operation integrates feature information scattered across different convolutional layers, making the fusion feature contain rich facial dynamic information from local to global, and from subtle to macroscopic. This fusion feature can more accurately and comprehensively reflect the dynamic changes of the target object's face.
[0086] In some optional implementations, step S205 involves performing feature processing on the fused features and the features of the face image to obtain the facial expression type of the face image, specifically including the following steps: The fusion features and the features of the face image to be tested are fused to obtain the target fusion features of the target object; a preset pooling layer is used to perform global pooling on the target fusion features to obtain a global feature vector; a preset fully connected layer is used to calculate the classification score of the global feature vector and output the facial expression type of the face image to be tested.
[0087] Among them, target fusion features refer to the feature set obtained by fusing the fusion features of facial dynamic difference data with the features of the face image to be tested.
[0088] Pooling layers are special network layers used for dimensionality reduction and feature selection of input features. Specifically, they can be used to process target fusion features. Since target fusion features may contain a lot of redundant information, pooling layers can extract the most representative features, reducing computational complexity and improving the model's generalization ability.
[0089] Global pooling refers to pooling the entire feature map, unlike regular pooling which operates within a local window. It is used to extract global information from the target fusion features, compressing the information of the entire feature map into a global feature representation, thus capturing the overall feature distribution of the feature map.
[0090] The global feature vector originates from a set of feature values obtained after global pooling. It represents the global feature representation of the target fused features in vector form after global pooling.
[0091] A fully connected layer is a network layer in which each neuron is connected to all neurons in the layer above it. It represents the process of linearly transforming the input features using a weight matrix and bias vector, followed by a non-linear mapping using an activation function.
[0092] The classification score calculation refers to the process of assigning a score value to each expression category using a specific calculation method. For example, the fully connected layer outputs a vector containing 7 elements, corresponding to 7 basic expression categories. The classification score calculation method is used to calculate the score value of each expression category, and the category with the highest score value is the facial expression type of the face image being tested.
[0093] In one example, within the financial insurance business, to accurately assess a customer's true emotions during insurance application and reduce fraud risk, the technical solution of this embodiment is adopted. First, a target facial image of the customer in a natural state during insurance consultation and a test facial image of the customer answering key insurance questions are acquired. Using a deep learning model, feature maps at multiple spatial scales are extracted from both, facial dynamic difference data is determined and fused to generate fused features. Next, the fused features are fused with the features of the test facial image. For example, a concatenation fusion method is used, where information reflecting dynamic facial changes in the fused features, such as the frequency of mouth twitching and the degree of eyebrow raising, is sequentially concatenated with static facial information in the test facial image features, such as facial feature shapes and facial contours, to obtain a new feature vector, resulting in the target fused feature. This feature comprehensively covers the customer's facial state information. Then, a pooling layer is used to perform global pooling on the target fused feature. Taking max pooling as an example, the maximum value of each local region in the target fused feature map is selected, compressing the entire feature map into a global feature vector. This vector reflects the overall key features of the customer's facial state. Finally, a fully connected layer is used to calculate the classification score of the global feature vector. The fully connected layer performs a linear transformation on the global feature vector using a weight matrix and a bias vector, and then processes it through an activation function to output scores corresponding to different facial expression types (such as happy, nervous, angry, etc.). The expression type with the highest score is the facial expression type of the tested face image. Insurance personnel can use this to determine the true emotion of the customer when answering questions, effectively preventing insurance fraud.
[0094] This application embodiment obtains target fusion features by fusing fusion features with features of the face image to be tested. The fusion features contain information on the dynamic changes of the target object's face, while the features of the face image to be tested reflect the current static state of the face. The fusion of the two makes the target fusion features comprehensively and accurately cover the facial state of the target object, including both dynamic trends and retaining static details, providing a rich data foundation for subsequent accurate analysis. A pooling layer is used to perform global pooling on the target fusion features to obtain a global feature vector. The pooling operation reduces the feature dimensionality, removes redundant information, and retains the most representative global features, effectively reducing the amount of computation and enhancing the model's generalization ability, making the features more stable and discriminative. A fully connected layer is used to calculate the classification score of the global feature vector and output the facial expression type. The powerful nonlinear mapping capability of the fully connected layer can deeply explore the complex relationship between the global feature vector and the expression type. By calculating the classification score, the expression category of the face image to be tested can be accurately determined. In scenarios such as customer emotion assessment in finance and insurance, it can efficiently and accurately identify customer expressions and assist in business decision-making.
[0095] It should be emphasized that, in order to further ensure the privacy and security of the target face image, the face image to be tested, the feature map, the facial dynamic difference data, the fusion feature, and the facial expression type, the target face image, the face image to be tested, the feature map, the facial dynamic difference data, the fusion feature, and the facial expression type can also be stored in a blockchain node.
[0096] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0097] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0098] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0099] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0100] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0101] Further reference Figure 3 As a response to the above Figure 2 The present application provides an embodiment of an image facial expression recognition device, which is similar to the method shown. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0102] like Figure 3 As shown, the facial expression recognition device 400 for images in this embodiment includes: an acquisition module 401, an extraction module 402, a determination module 403, a fusion module 404, and a processing module 405. Wherein: The acquisition module 401 is used to acquire the target face image and the face image to be tested of the target object. The target face image is the face image of the target object in its natural state. The extraction module 402 is used to extract features from the target face image and the face image to be tested using a pre-trained deep learning model, and obtain feature maps at multiple spatial scales. The determination module 403 is used to determine the facial dynamic difference data between the target face image and the face image to be tested based on the feature map; The fusion module 404 is used to fuse facial dynamic difference data to generate fusion features of the target object; The processing module 405 is used to perform feature processing on the fused features and the features of the face image to be tested, so as to obtain the facial expression type of the face image to be tested.
[0103] In this embodiment, by acquiring the target face image in its natural state and the test face image in its current state, and using a pre-trained deep learning model to extract feature maps at multiple spatial scales, it can comprehensively and accurately capture feature information at different levels of the face, fully considering the changes in facial expressions across different spatial dimensions. Based on these feature maps, facial dynamic difference data is determined, which can effectively distinguish between facial expression features in an individual's natural state and those caused by emotional changes, overcoming the shortcomings of existing single-frame image judgments that are inaccurate due to the influence of the facial expression baseline. The dynamic difference data is fused to generate fused features, further integrating key information and enhancing the feature representation capability. Finally, the fused features and the test face image features are processed to obtain the facial expression type, which can significantly improve the accuracy of facial expression prediction, providing reliable expression analysis results for numerous application scenarios and meeting the high-precision requirements of practical applications.
[0104] In one embodiment, the extraction module 402 includes: The extraction submodule is used to input the target face image and the face image to be tested into a pre-trained deep learning model, extract features through multiple hidden layers and output layers of the deep learning model, and generate multiple initial feature maps of different spatial scales. The comparison submodule is used to compare the initial feature map corresponding to the target face image and the initial feature map corresponding to the face image to be tested, and to determine the feature differences at each spatial scale. The adjustment submodule is used to adjust the initial feature map based on feature differences to obtain feature maps at multiple spatial scales.
[0105] This application embodiment can input the target face image and the test face image into a pre-trained deep learning model. Utilizing its multiple hidden layers, it can perform multi-level, progressive feature mining on the image. From low-level edge and texture information to high-level facial organ shapes and expression-related features, it progressively extracts and generates multiple initial feature maps at different spatial scales via the output layer, comprehensively and meticulously representing the features of the face in different dimensions. By comparing the initial feature maps corresponding to the target and test face images, the feature differences at each spatial scale can be accurately located. Whether it's changes in the overall facial contour or subtle movements of local facial features, they can all be clearly captured. Adjusting the initial feature maps based on feature differences can eliminate interference caused by differences between the natural state and the current state, making the resulting feature maps at multiple spatial scales more closely match the current real expression features of the target object. This series of operations works together to improve the accuracy and comprehensiveness of facial expression feature extraction.
[0106] In one embodiment, the determining module 403 includes: The computation submodule is used to calculate the difference between the feature maps of the target face image and the face image to be tested at each spatial scale; The processing submodule is used to normalize the difference values using a preset difference calculation formula to obtain standardized difference information. The determination submodule is used to determine the dynamic change characteristics of the feature map at each spatial scale based on the difference information; The weighting submodule is used to weight the dynamically changing features to generate facial dynamic difference data between the target face image and the face image to be tested.
[0107] This application embodiment can accurately locate the deviation of facial features at different scales by calculating the difference value between the feature maps of the target and the tested face image at each spatial scale. Since the target face image is in its natural state and the tested image is in its current state, there are differences such as expressions between the two. This calculation can comprehensively capture feature changes from subtle to overall, providing a detailed data foundation for subsequent analysis. The difference calculation formula is used to normalize the difference values, which can eliminate the influence of differences in the range and dimensions of feature values at different spatial scales, making the difference information have a unified standard, facilitating comparison and analysis, and ensuring that the degree of feature difference can be accurately measured at different scales. Based on the difference information, the dynamic change features of the feature map at each spatial scale can be determined, which can deeply explore the laws of facial expression changes and clearly present the dynamic pattern of feature changes at different scales with expression. Weighting the dynamic change features can highlight the influence of key scale features on the final result. By reasonably allocating weights according to the importance of different scales in expression recognition, data that accurately reflects the dynamic differences between the target and the tested face images can be generated, effectively improving the accuracy and reliability of facial expression recognition.
[0108] In one embodiment, the determining submodule is further configured to divide the difference information according to the scale parameters of each spatial scale to obtain the distribution data of the difference information at each spatial scale; and to use a preset analysis tool to extract the change characteristics of the distribution data to determine the dynamic change characteristics of the feature map at each spatial scale.
[0109] This application's embodiments divide the difference information according to the scale parameters of each spatial scale, accurately classifying the difference information according to different spatial ranges and feature granularities. Since different spatial scales focus on different aspects of the facial image, small scales can capture subtle local feature differences, while large scales can grasp overall contour changes. This division clearly presents the distribution of difference information at each scale, providing a detailed and targeted data foundation for subsequent analysis. Using analytical tools to extract the variation characteristics of the distributed data allows for in-depth exploration of the underlying patterns. Based on this, determining the dynamic change characteristics of the feature map at each spatial scale comprehensively and accurately reflects the dynamic process of facial features changing with expression or state, considering not only local detail changes but also overall feature evolution, providing a reliable basis for generating dynamic facial difference data.
[0110] In one embodiment, the fusion module 404 includes: The transformation submodule is used to input facial dynamic difference data into multiple convolutional layers for feature transformation, and obtain multiple transformed difference features. The splicing submodule is used to splice multiple differential features to generate the fused features of the target object.
[0111] This application's embodiments transform facial dynamic difference data by inputting it into multiple convolutional layers. Because different convolutional layers have kernels of different scales and unique computational logic, they can deeply mine feature information from multiple dimensions within the facial dynamic difference data. Small-scale convolutional kernels can capture subtle local dynamic changes on the face, such as blinking frequency and the amplitude of mouth twitching. Large-scale convolutional kernels can grasp the overall dynamic trend of the face, such as changes in facial contours, subtle head movements, and the coordination of facial expressions. Through processing by multiple convolutional layers, multiple transformed difference features are obtained, which are comprehensive and hierarchical. These multiple difference features are then concatenated to generate a fusion feature for the target object. This operation integrates feature information scattered across different convolutional layers, making the fusion feature contain rich facial dynamic information from local to global, and from subtle to macroscopic. This fusion feature can more accurately and comprehensively reflect the dynamic changes of the target object's face.
[0112] In one embodiment, the processing module 405 includes: The fusion submodule is used to fuse the fusion features and the features of the face image to be tested to obtain the target fusion features of the target object; The pooling submodule is used to perform global pooling on the target fused features using a preset pooling layer to obtain a global feature vector; The classification submodule is used to calculate the classification score of the global feature vector using a preset fully connected layer and output the facial expression type of the face image to be tested.
[0113] This application embodiment obtains target fusion features by fusing fusion features with features of the face image to be tested. The fusion features contain information on the dynamic changes of the target object's face, while the features of the face image to be tested reflect the current static state of the face. The fusion of the two makes the target fusion features comprehensively and accurately cover the facial state of the target object, including both dynamic trends and retaining static details, providing a rich data foundation for subsequent accurate analysis. A pooling layer is used to perform global pooling on the target fusion features to obtain a global feature vector. The pooling operation reduces the feature dimensionality, removes redundant information, and retains the most representative global features, effectively reducing the amount of computation and enhancing the model's generalization ability, making the features more stable and discriminative. A fully connected layer is used to calculate the classification score of the global feature vector and output the facial expression type. The powerful nonlinear mapping capability of the fully connected layer can deeply explore the complex relationship between the global feature vector and the expression type. By calculating the classification score, the expression category of the face image to be tested can be accurately determined. In scenarios such as customer emotion assessment in finance and insurance, it can efficiently and accurately identify customer expressions and assist in business decision-making.
[0114] In one embodiment, the facial expression recognition device 400 for images further includes: The function acquisition module is used to obtain the basic loss function, which includes the difference between the label value and the predicted value. The calculation module is used to acquire the training set and determine the weight of each type of expression by calculating the distribution ratio of each type of expression in the training set. The weighting module is used to weight the difference terms according to the weights to obtain the weighted loss function; The training module is used to iteratively train a pre-set initial deep learning model using a weighted loss function to obtain a pre-trained deep learning model.
[0115] In this embodiment, during model training, a fundamental loss function containing the difference between label and predicted values provides a key quantitative standard for measuring the accuracy of model predictions. This clearly reflects the degree to which predicted values deviate from the true labels, guiding the direction of model optimization. Weights are determined by acquiring the training set and calculating the distribution ratio of each type of expression, fully considering the frequency differences of different expressions in the dataset. Larger weights are assigned to expression types that appear less frequently in the training set and have a low distribution ratio, preventing the model from ignoring these important expression features due to data imbalance. A weighted loss function is obtained by weighting the difference terms according to the weights, allowing the model to more reasonably focus on various expressions during training. This weighted loss function is used to iteratively train the initial deep learning model. In each iteration, the model continuously adjusts its parameters based on the weighted loss, gradually improving its ability to recognize various expressions, ultimately resulting in a pre-trained deep learning model that can more accurately and comprehensively recognize different expressions, especially those that originally had a small proportion in the dataset.
[0116] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0117] Computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only computer device 6 with memory 61, processor 62, and network interface 63 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0118] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0119] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for facial expression recognition methods. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.
[0120] In some embodiments, processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. Processor 62 is typically used to control the overall operation of computer device 6. In this embodiment, processor 62 is used to execute computer-readable instructions stored in memory 61 or to process data, such as computer-readable instructions for executing a facial expression recognition method for images.
[0121] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 6 and other electronic devices.
[0122] This application embodiment can acquire a target face image in its natural state and a test face image in its current state, and use a pre-trained deep learning model to extract feature maps at multiple spatial scales. This comprehensively and accurately captures feature information at different levels of the face, fully considering the changes in facial expressions across different spatial dimensions. Based on these feature maps, dynamic facial difference data is determined, effectively distinguishing between facial expression features in an individual's natural state and those resulting from emotional changes. This overcomes the shortcomings of existing single-frame image judgments, which are prone to inaccurate predictions due to the influence of the facial expression baseline. The dynamic difference data is fused to generate fused features, further integrating key information and enhancing the feature representation capability. Finally, the fused features and the test face image features are processed to derive the facial expression type, significantly improving the accuracy of facial expression prediction. This provides reliable expression analysis results for numerous application scenarios, meeting the high-precision requirements of practical applications.
[0123] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the facial expression recognition method for images as described above.
[0124] This application embodiment can acquire a target face image in its natural state and a test face image in its current state, and use a pre-trained deep learning model to extract feature maps at multiple spatial scales. This comprehensively and accurately captures feature information at different levels of the face, fully considering the changes in facial expressions across different spatial dimensions. Based on these feature maps, dynamic facial difference data is determined, effectively distinguishing between facial expression features in an individual's natural state and those resulting from emotional changes. This overcomes the shortcomings of existing single-frame image judgments, which are prone to inaccurate predictions due to the influence of the facial expression baseline. The dynamic difference data is fused to generate fused features, further integrating key information and enhancing the feature representation capability. Finally, the fused features and the test face image features are processed to derive the facial expression type, significantly improving the accuracy of facial expression prediction. This provides reliable expression analysis results for numerous application scenarios, meeting the high-precision requirements of practical applications.
[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0126] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
[0127] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A method for facial expression recognition in images, characterized in that, Includes the following steps: Acquire the target face image and the face image to be tested of the target object, wherein the target face image is the face image of the target object in its natural state; A pre-trained deep learning model is used to extract features from the target face image and the face image to be tested, resulting in feature maps at multiple spatial scales. Based on the feature map, determine the facial dynamic difference data between the target face image and the face image to be tested; The facial dynamic difference data are fused to generate the fused features of the target object; Feature processing is performed on the fused features and the features of the face image to be tested to obtain the facial expression type of the face image to be tested.
2. The method according to claim 1, characterized in that, The step of using a pre-trained deep learning model to extract features from the target face image and the face image to be tested, and obtaining feature maps at multiple spatial scales, specifically includes: The target face image and the face image to be tested are input into a pre-trained deep learning model. Features are extracted through multiple hidden layers and output layers of the deep learning model to generate multiple initial feature maps of different spatial scales. The initial feature map corresponding to the target face image and the initial feature map corresponding to the face image to be tested are compared to determine the feature differences at each spatial scale. Based on the aforementioned feature differences, the initial feature map is adjusted to obtain feature maps at multiple spatial scales.
3. The method according to claim 1, characterized in that, Before the step of using a pre-trained deep learning model to extract features from the target face image and the test face image to obtain feature maps at multiple spatial scales, the method further includes: Obtain the base loss function, which includes a difference term between the label value and the predicted value; Obtain a training set, and determine the weight of each type of expression by calculating the distribution ratio of each type of expression in the training set; Based on the weights, the difference terms are weighted to obtain a weighted loss function; The weighted loss function is used to iteratively train the preset initial deep learning model to obtain the pre-trained deep learning model.
4. The method according to claim 1, characterized in that, The step of determining the facial dynamic difference data between the target face image and the face image to be tested based on the feature map specifically includes: Calculate the difference between the feature maps of the target face image and the face image to be tested at each spatial scale; The difference values are normalized using a preset difference calculation formula to obtain standardized difference information. Based on the difference information, determine the dynamic change characteristics of the feature map at each spatial scale; The dynamic change features are weighted to generate facial dynamic difference data between the target face image and the face image to be tested.
5. The method according to claim 4, characterized in that, The step of determining the dynamic change characteristics of the feature map at each spatial scale based on the difference information specifically includes: Based on the scale parameters of each spatial scale, the difference information is divided to obtain the distribution data of the difference information at each spatial scale; Using preset analysis tools, the variation characteristics of the distribution data are extracted to determine the dynamic variation characteristics of the feature map at each spatial scale.
6. The method according to claim 1, characterized in that, The step of fusing the facial dynamic difference data to generate the fused features of the target object specifically includes: The facial dynamic difference data is input into multiple convolutional layers for feature transformation to obtain multiple transformed difference features. The multiple differential features are concatenated to generate the fused features of the target object.
7. The method according to claim 1, characterized in that, The step of performing feature processing on the fused features and the features of the face image to obtain the facial expression type of the face image to be tested specifically includes: The fusion features and the features of the face image to be tested are fused to obtain the target fusion features of the target object; A preset pooling layer is used to perform global pooling on the target fused features to obtain a global feature vector; A preset fully connected layer is used to calculate the classification score of the global feature vector and output the facial expression type of the face image to be tested.
8. A facial expression recognition device for images, characterized in that, include: The acquisition module is used to acquire the target face image and the face image to be tested of the target object, wherein the target face image is the face image of the target object in a natural state; The extraction module is used to extract features from the target face image and the face image to be tested using a pre-trained deep learning model, so as to obtain feature maps at multiple spatial scales. The determination module is used to determine the facial dynamic difference data between the target face image and the face image to be tested based on the feature map; The fusion module is used to fuse the facial dynamic difference data to generate the fusion features of the target object; The processing module is used to perform feature processing on the fused features and the features of the face image to be tested, so as to obtain the facial expression type of the face image to be tested.
9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the facial expression recognition method for an image as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions that, when executed by a processor, implement the steps of the facial expression recognition method for an image as described in any one of claims 1 to 7.