Facial expression feature extraction method, facial expression recognition method and electronic device
By extracting the blocks of multiple types of preset parts, building feature difference matrix and feature space of part, the problem of low precision in expression feature extraction in the prior art is solved, and a higher accuracy of expression recognition is achieved.
Patent Information
- Application Number
- CN202211513669.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-11-28
AI Technical Summary
The prior art ignores the differences between different types of expressions in the process of expression feature extraction, resulting in low accuracy of the extracted expression features, which in turn reduces the accuracy of expression recognition.
By obtaining multiple sample images, multiple blocks corresponding to each of the preset parts of multiple types are extracted from them, a feature difference matrix is constructed, and the site feature space is constructed using this matrix to obtain reference expression features corresponding to each preset expression category.
It improves the accuracy of expression feature extraction, increases the inter-class distance between different expression categories, and improves the accuracy of expression recognition.
Smart Images

Figure CN115830678B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a method for extracting expression features, a method for recognizing expressions, and an electronic device. Background Art
[0002] With the rise of the field of computer vision, expressions, as an important way of emotional communication, expression recognition, as a branch of computer vision, has received increasing attention. In the prior art, features are usually extracted from sample images corresponding to different types of expressions, and the extracted features are used as reference features. Then, the features corresponding to the expression to be recognized are compared with the reference features to perform expression recognition. However, in the prior art, the differences between different types of expressions are ignored during the feature extraction process, resulting in low accuracy of the extracted expression features, and thus the accuracy of expression recognition will also decrease. In view of this, how to improve the accuracy of expression feature extraction has become an urgent problem to be solved. Summary of the Invention
[0003] The main technical problem to be solved by the present application is to provide a method for extracting expression features, a method for recognizing expressions, and an electronic device, which can improve the accuracy of expression feature extraction.
[0004] To solve the above technical problem, a first aspect of the present application provides a method for extracting expression features. The method includes: obtaining a plurality of sample images, and extracting a plurality of sub-blocks corresponding to a plurality of preset parts of a plurality of types from the sample images; wherein, at least some of the sample images have different preset expression categories; based on the sub-blocks corresponding to the preset parts of the same type, determining the total feature and a plurality of sub-features corresponding to the preset parts; wherein, the total feature is related to all the sub-blocks corresponding to the preset parts, and each sub-feature is respectively related to the sub-blocks corresponding to one preset expression category on the preset parts; based on the differences between the plurality of sub-features relative to the total feature respectively, determining a feature difference matrix corresponding to the preset parts; using the feature difference matrix to construct a part feature space corresponding to the preset parts, and based on the sub-blocks corresponding to the preset parts and the part feature space, obtaining the reference expression features corresponding to each preset expression category on the preset parts
[0005] To solve the above technical problems, a second aspect of the present application provides an expression recognition method, which includes: obtaining an image to be recognized, and extracting corresponding recognition sub-blocks on each of multiple types of preset parts from the image to be recognized; based on the recognition sub-blocks corresponding to the preset parts of the same type and the part feature space corresponding to the preset parts, obtaining the recognition expression features corresponding to the preset parts, and determining the part similarity between the recognition expression features and the reference expression features corresponding to multiple preset expression categories; based on the part similarities corresponding to each of all types of preset parts, obtaining the target expression category corresponding to the image to be recognized; wherein, the part feature space and the reference expression features are obtained based on the expression feature extraction method described in the first aspect above.
[0006] To solve the above technical problems, a third aspect of the present application provides an electronic device, which includes: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described in the first aspect or the second aspect above.
[0007] In the above solution, after obtaining multiple sample images, multiple sub-blocks corresponding to each of multiple types of preset parts are extracted from all the sample images. That is to say, each sample image is divided into multiple types of preset parts, and each type of preset part corresponds to multiple sub-blocks obtained from multiple sample images. Among them, at least some of the sample images have different preset expression categories, so that all the sub-blocks corresponding to each type of preset part include different preset expression categories. Based on all the sub-blocks corresponding to the preset parts of the same type, the total features corresponding to the preset parts are determined. Based on the sub-blocks corresponding to each preset expression category on the preset parts of the same type, multiple sub-features corresponding to the preset parts are determined. Based on the differences between the multiple sub-features and the total features respectively, a feature difference matrix corresponding to the preset parts is determined. The part feature space corresponding to the preset parts is constructed by using the feature difference matrix. Based on the sub-blocks corresponding to the preset parts and the part feature space, the sub-blocks corresponding to different preset expression categories are dimensionally reduced to improve the efficiency of feature extraction, and the reference expression features corresponding to each preset expression category on the preset parts are obtained. Therefore, each type of preset part corresponds to reference expression features corresponding to different preset expression categories. And when constructing the feature difference matrix, each sub-feature corresponds to a preset expression category, fully considering the inter-class differences between the features corresponding to different preset expression categories, thereby increasing the inter-class distance between different preset expression categories, and finally improving the accuracy of the reference expression features corresponding to different preset expression categories on each type of preset part. Description of the Drawings
[0008] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. Among them:
[0009] Figure 1 is a schematic flowchart of an implementation manner of the expression feature extraction method of the present application;
[0010] Figure 2 is a schematic flowchart of another implementation manner of the expression feature extraction method of the present application;
[0011] Figure 3 is Figure 2 a schematic diagram of an application scenario corresponding to step S202 in;
[0012] Figure 4 is a schematic flowchart of an implementation manner of the expression feature extraction method of the present application;
[0013] Figure 5 is a schematic structural diagram of an implementation manner of an electronic device of the present application;
[0014] Figure 6 is a schematic structural diagram of an implementation manner of a computer-readable storage medium of the present application. Specific Embodiments
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0016] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article only describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0017] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the personal's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the personal's separate consent and at the same time meets the requirements of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consenting to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up information or asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0018] The expression feature extraction method provided by this application is used to extract the features corresponding to the face region in an image, and the expression recognition method is used to recognize the expression corresponding to the face region in the image. The execution subject of the expression feature extraction method and the expression recognition method provided by this application is a processor capable of calling the image.
[0019] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an implementation manner of the expression feature extraction method of this application. The method includes:
[0020] S101: Obtain a plurality of sample images, and extract a plurality of blocks corresponding to each of a plurality of types of preset parts from the sample images, where at least some of the sample images have different preset expression categories.
[0021] Specifically, after obtaining a plurality of sample images, extract a plurality of blocks corresponding to each of a plurality of types of preset parts from all the sample images, that is, divide each sample image into a plurality of types of preset parts, and each type of preset part corresponds to a plurality of blocks obtained from a plurality of sample images, where at least some of the sample images have different preset expression categories, so that all the blocks corresponding to each type of preset part include different preset expression categories.
[0022] In an application mode, obtain a plurality of sample images, where each sample image corresponds to a preset expression category, and the number of preset expression categories is a plurality. Sequentially extract the blocks on a plurality of types of preset parts from each sample image. After traversing all the sample images, arrange the blocks obtained on the sample images according to the preset part type to obtain a plurality of blocks corresponding to each type of preset part.
[0023] In another application mode, a plurality of sample images are obtained, where the plurality of sample images include a plurality of preset expression categories, and there are multiple sample images corresponding to each preset expression category. After traversing the images of all preset expression categories, multiple blocks on preset parts of the same type corresponding to different preset expression categories are obtained by extracting blocks on preset parts of multiple types from the sample images corresponding to each preset expression category.
[0024] In an application scenario, the sample images are face images, the preset expression categories include m categories, and the preset parts include eyes, nose, and mouth. Among them, the number of sample images for each preset expression category is n. Blocks corresponding to the eyes, nose, and mouth are respectively extracted from all the sample images to obtain block sets corresponding to the eyes, nose, and mouth respectively, and the dimension of each block set is n*m.
[0025] Optionally, before obtaining a plurality of sample images and extracting multiple blocks corresponding to multiple types of preset parts from the sample images, it includes: obtaining a plurality of initial sample images and performing preprocessing operations on the initial sample images to obtain sample images; where the preprocessing operations include at least one of angle correction, homomorphic filtering, and normalization processing.
[0026] Specifically, the initial sample images can be obtained from an image library or extracted from a video. Each initial sample image corresponds to a label of any preset expression category. After obtaining a plurality of initial sample images, preprocessing operations are performed on the initial sample images, where the preprocessing operations include at least one of angle correction, homomorphic filtering, and normalization processing to reduce geometric differences, size differences, and illumination differences and enhance the details of the obtained sample images.
[0027] In an application scenario, initial sample images are obtained, and the Viola-Jones object detection method is used to extract the region image corresponding to the face region from the initial sample images, and image angle normalization, homomorphic filtering, and data normalization processing are performed on the region image to reduce geometric differences, size differences, and illumination differences and enhance the details of the obtained sample images to obtain preprocessed sample images.
[0028] S102: Based on the blocks corresponding to the preset parts of the same type, determine the total feature and multiple sub-features corresponding to the preset parts, where the total feature is related to all the blocks corresponding to the preset parts, and each sub-feature is respectively related to the blocks corresponding to one preset expression category on the preset parts.
[0029] Specifically, based on all the sub - blocks corresponding to the preset part of the same type, determine the total feature corresponding to the preset part. Based on the sub - blocks corresponding to each preset expression category on the preset part of the same type, determine multiple sub - features corresponding to the preset part. Based on the differences between the multiple sub - features and the total feature respectively, determine the feature difference matrix corresponding to the preset part, where each sub - feature corresponds to a preset expression category.
[0030] In one application method, calculate the mean value of all the sub - blocks corresponding to the preset part of the same type to obtain the total mean value of the sub - blocks. Determine the total feature corresponding to the preset part based on the total mean value of the sub - blocks among all the sub - blocks corresponding to the preset part of the same type. Take all the sub - blocks corresponding to the same preset expression category on the preset part of the same type as the intra - class sub - blocks, calculate the mean value of each intra - class sub - block to obtain the intra - class mean value of the sub - blocks, and determine multiple sub - features corresponding to the preset part and matching the preset expression category based on the intra - class mean value of the sub - blocks among the intra - class sub - blocks corresponding to each preset expression category on the preset part of the same type.
[0031] In another application method, sum up all the sub - blocks corresponding to the preset part of the same type to obtain the total value of the sub - blocks. Determine the total feature corresponding to the preset part based on the total value of the sub - blocks among all the sub - blocks corresponding to the preset part of the same type. Take all the sub - blocks corresponding to the same preset expression category on the preset part of the same type as the intra - class sub - blocks, sum up each intra - class sub - block to obtain the intra - class total value of the sub - blocks, and determine multiple sub - features corresponding to the preset part and matching the preset expression category based on the intra - class total value of the sub - blocks among the intra - class sub - blocks corresponding to each preset expression category on the preset part of the same type.
[0032] S103: Determine the feature difference matrix corresponding to the preset part based on the differences between the multiple sub - features and the total feature respectively.
[0033] Specifically, use the differences between the multiple sub - features and the total feature respectively to construct the feature difference matrix corresponding to the preset part.
[0034] In one application method, the total feature is determined based on the total mean value of the sub - blocks among all the sub - blocks corresponding to the preset part of the same type, and the sub - features are determined based on the intra - class mean value of the sub - blocks among the intra - class sub - blocks corresponding to each preset expression category on the preset part of the same type. Use the difference between each intra - class mean value and the total mean value of the sub - blocks to calculate the covariance matrix to obtain the feature difference matrix corresponding to the preset part.
[0035] In another application method, the total feature is determined based on the total value of the sub - blocks among all the sub - blocks corresponding to the preset part of the same type, and the sub - features are determined based on the intra - class total value of the sub - blocks among the intra - class sub - blocks corresponding to each preset expression category on the preset part of the same type. Use the difference between each intra - class total value and the total value of the sub - blocks to calculate the covariance matrix to obtain the feature difference matrix corresponding to the preset part.
[0036] S104: Construct a part feature space corresponding to a preset part by using the feature difference matrix, and obtain the reference expression features corresponding to each preset expression category on the preset part based on the blocks corresponding to the preset part and the part feature space.
[0037] Specifically, construct a part feature space corresponding to a preset part by using the feature difference matrix, and perform dimensionality reduction on the blocks corresponding to different preset expression categories based on the blocks corresponding to the preset part and the part feature space to improve the efficiency of feature extraction, so as to obtain the reference expression features corresponding to each preset expression category on the preset part.
[0038] It can be understood that each type of preset part corresponds to its own part feature space, and the reference expression features corresponding to each preset expression category on the preset part.
[0039] In an application mode, sort the eigenvalues in the feature difference matrix from largest to smallest in value, select the first preset number of eigenvalues in the sorting to construct a projection matrix, determine the construction of the part feature space corresponding to the preset part based on the projection matrix, project the blocks corresponding to the preset part into the part feature space, perform dimensionality reduction on the blocks to achieve feature extraction, and obtain the reference expression features corresponding to each preset expression category on the preset part.
[0040] In another application mode, extract the eigenvalues exceeding the feature threshold from the feature difference matrix, construct a projection matrix based on the extracted eigenvalues, determine the construction of the part feature space corresponding to the preset part based on the projection matrix, project the blocks corresponding to the preset part into the part feature space, perform dimensionality reduction on the blocks to achieve feature extraction, and obtain the reference expression features corresponding to each preset expression category on the preset part.
[0041] It should be noted that the traditional principal component analysis technology (Principal Components Analysis, PCA) needs to convert the image into a one-dimensional vector, which increases the computational complexity, and the traditional PCA technology depends on the dispersion degree of a specific image and cannot guarantee the between-class separability of images corresponding to different preset expression categories. In this embodiment, when constructing the feature difference matrix, each sub-feature corresponds to a preset expression category, fully considering the between-class differences between the features corresponding to different preset expression categories, thereby increasing the between-class distance corresponding to different preset expression categories. When performing feature extraction on the image, the method of constructing the part feature space can avoid the consumption of converting into a vector, improve the efficiency of feature extraction, and ensure the between-class separability of features without increasing the computational workload.
[0042] In the above solution, after obtaining multiple sample images, multiple blocks corresponding to respective preset parts of multiple types are extracted from all the sample images. That is to say, each sample image is divided into preset parts of multiple types, and each type of preset part corresponds to multiple blocks obtained from multiple sample images. Among them, at least some of the preset expression categories of the sample images are different, so that all the blocks corresponding to each type of preset part include different preset expression categories. Based on all the blocks corresponding to the preset part of the same type, the total feature corresponding to the preset part is determined. Based on the blocks corresponding to each preset expression category on the preset part of the same type, multiple sub-features corresponding to the preset part are determined. Based on the differences between the multiple sub-features and the total feature respectively, a feature difference matrix corresponding to the preset part is determined. The part feature space corresponding to the preset part is constructed by using the feature difference matrix. Based on the blocks corresponding to the preset part and the part feature space, dimensionality reduction is performed on the blocks corresponding to different preset expression categories to improve the efficiency of feature extraction, and the reference expression features corresponding to each preset expression category on the preset part are obtained. Therefore, each type of preset part corresponds to reference expression features corresponding to different preset expression categories, and when constructing the feature difference matrix, each sub-feature corresponds to a preset expression category, fully considering the inter-class differences between the features corresponding to different preset expression categories, thereby increasing the inter-class distance corresponding to different preset expression categories, and finally improving the accuracy of the reference expression features corresponding to different preset expression categories on each type of preset part.
[0043] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the expression feature extraction method of the present application. The method includes:
[0044] S201: Classify multiple sample images based on the preset expression category corresponding to each sample image to obtain a subset of sample images corresponding to each preset expression category.
[0045] Specifically, classify multiple sample images according to the preset expression category corresponding to the sample images, so as to divide the multiple sample images into multiple subsets of sample images, and the multiple subsets of sample images correspond to one preset expression category.
[0046] Optionally, before classifying multiple sample images based on the preset expression category corresponding to each sample image to obtain a subset of sample images corresponding to each preset expression category, it includes: obtaining multiple initial sample images, and performing a preprocessing operation on the initial sample images to obtain sample images; wherein, the preprocessing operation includes at least one of angle correction, homomorphic filtering, and normalization processing.
[0047] Specifically, each initial sample image corresponds to a label of any preset expression category. After obtaining multiple initial sample images, preprocessing operations are performed on the initial sample images. Among them, the preprocessing operations include at least one of angle correction, homomorphic filtering, and normalization processing to reduce geometric differences, size differences, and illumination differences and enhance the details of the obtained sample images.
[0048] S202: For each type of preset part, multiple blocks of the preset part are respectively extracted from each sample image subset to obtain a subset of blocks corresponding to each preset expression category. Based on all subsets of blocks corresponding to the preset part, a set of blocks corresponding to the preset part is obtained.
[0049] Specifically, please refer to Figure 3 , Figure 3 is Figure 2 a schematic diagram of the application scenario of an implementation manner corresponding to step S202 in . Taking two preset expression categories and three types of preset parts as an example, among them, the preset parts are distinguished based on the same filling color. After the above step S201, the sample images are divided into two subsets of sample images according to the preset expression category. By traversing all subsets of sample images, the blocks on each type of preset part are sequentially extracted from the subsets of sample images. The sub-blocks corresponding to the same preset expression category are used as a subset of blocks, and all subsets of blocks corresponding to the preset part form a set of blocks corresponding to the preset part.
[0050] Furthermore, as Figure 3 shown, each set of blocks is distinguished according to the preset part, and the subsets of blocks within each set of blocks are distinguished according to the preset expression category, so as to facilitate determining the total features corresponding to all sub-blocks within different types of preset parts and the sub-features within the subset of blocks corresponding to each preset expression category.
[0051] In an application scenario, all sample images are divided into M categories according to the preset expression category, including ω1, ω2, …, ω M , and there are n i sample images in the subset of sample images corresponding to each category. C1, C2, …, C k , …, C K are all sample images, where the sample images can be divided into three sets of blocks: the eye part, the nose part, and the mouth part, which are respectively denoted as A k , B k , Dk, and each set of blocks is a matrix of m * n.
[0052] S203: Obtain the total block mean value corresponding to the preset part based on the mean values of all blocks in the block set corresponding to the preset part, and obtain the within-class block mean values corresponding to each preset expression category on the preset part based on the mean values of all blocks in each subset of blocks corresponding to the preset part.
[0053] Specifically, the preset part corresponds to a total feature and multiple sub-features. Calculate the mean value of all blocks in the block set corresponding to the preset part to obtain the total block mean value corresponding to the preset part. Calculate the mean value of all blocks in each subset of blocks corresponding to the preset part to obtain the within-class block mean values corresponding to each preset expression category on the preset part. Among them, the total feature corresponds to the total block mean value, the sub-feature corresponds to the within-class block mean value, and the within-class block mean value is a linear combination of within-class images. Therefore, the within-class block mean value corresponding to each subset of blocks retains a large amount of variation of a specific image and reduces the probability of loss of main image information.
[0054] In an application scenario, taking the mouth as an example of the preset part, the block set corresponding to the mouth is D k , calculate the total block mean value corresponding to all blocks in the block set to obtain the total block mean value matrix, and calculate the within-class block mean value corresponding to each subset of blocks in the block set to obtain the within-class block mean value matrix. Among them, the above process is expressed by the following formula:
[0055]
[0056]
[0057] Among them, is the total block mean value matrix, is the within-class block mean value matrix.
[0058] S204: Use the difference between each within-class block mean value and the total block mean value to calculate the covariance matrix, and obtain the feature difference matrix corresponding to the preset part.
[0059] Specifically, based on the difference between each within-class block mean value and the total block mean value, calculate the covariance matrix, so as to obtain the feature difference matrix corresponding to the preset part. The above process is expressed by the following formula:
[0060]
[0061] Among them, G is the feature difference matrix. In addition, it should be noted that the process of calculating the covariance matrix in the traditional method is expressed by the following formula:
[0062]
[0063] Among them, when calculating the covariance matrix, the average value of each class is used to replace the within-class specific image D in formula (4). k , which further increases the between-class distance, reduces the within-class difference, and increases the robustness and accuracy of facial expression recognition. In addition, as shown in formula (3), since the average value of each class is a linear combination of within-class images, the block within-class mean of each class retains a large amount of variation of the specific image, reduces the probability of loss of main image information, and increases the separability between classes.
[0064] S205: Construct a part feature space corresponding to the preset part by using the feature difference matrix, and obtain the reference expression features corresponding to each preset expression category on the preset part based on the block corresponding to the preset part and the part feature space.
[0065] Specifically, construct a part feature space corresponding to the preset part by using the feature difference matrix, and reduce the dimension of the blocks corresponding to different preset expression categories based on the block corresponding to the preset part and the part feature space to improve the efficiency of feature extraction, and obtain the reference expression features corresponding to each preset expression category on the preset part.
[0066] In an application mode, construct a part feature space corresponding to the preset part by using at least some of the eigenvalues in the feature difference matrix; project the blocks corresponding to each preset expression category on the preset part onto the part feature space to obtain the reference expression features corresponding to each preset expression category on the preset part.
[0067] Specifically, extract at least some of the eigenvalues from the feature difference matrix, thereby construct a part feature space corresponding to the preset part based on the extracted eigenvalues, and project the blocks corresponding to each preset expression category on the preset part onto the part feature space respectively to realize the dimensionality reduction of the image blocks, and obtain the reference expression features corresponding to each preset expression category on the preset part.
[0068] Furthermore, the extracted eigenvalues are related to the numerical magnitudes of the eigenvalues. Use the extracted eigenvalues to span a part feature space of the corresponding dimension, and perform dimensionality reduction by projecting the blocks onto the corresponding part feature space, thereby omitting the process of converting the image into a one-dimensional vector, reducing the computational workload and saving the time for vectorization, and retaining the structured information.
[0069] In an application scenario, construct a part feature space corresponding to the preset part by using at least some of the eigenvalues in the feature difference matrix, including: select a preset number of specified eigenvalues from the feature difference matrix based on the numerical values corresponding to the eigenvalues in the feature difference matrix; obtain a projection matrix based on the orthogonal eigenvectors corresponding to the specified eigenvalues, and construct a part feature space corresponding to the preset part by using the projection matrix; where the dimension of the part feature space corresponds to the preset number.
[0070] Specifically, sort the eigenvalues in the feature difference matrix in descending order of their numerical values, select a preset number of the top-ranked eigenvalues from the feature difference matrix as the specified eigenvalues, determine the orthogonal eigenvectors corresponding to the specified eigenvalues, construct a projection matrix based on the orthogonal eigenvectors, determine the space spanned by the projection matrix, and obtain the part feature space corresponding to the preset part.
[0071] Furthermore, the part feature space can be used to extract features from the blocks of the preset part that match the part feature space. Without continuously extracting features, the inter-class separability of information can be ensured. Finally, the feature dimension of the blocks is reduced, the feature information is optimized, thereby saving storage space and improving computational efficiency.
[0072] In a specific application scenario, r eigenvalues are extracted from the feature difference matrix, and the corresponding standard orthogonal eigenvectors of the r eigenvalues are X1, X2, …, X r , thus forming the projection matrix P = [X1, X1, …, X r . Project the blocks of the corresponding preset part into the subspace spanned by the projection matrix P, extract features from it, and obtain the projected block matrix D ′ k = D k P. Among them, the subspace spanned by the projection matrix P is the part feature space, and the block matrix includes reference expression features.
[0073] In this embodiment, for each type of preset part, a covariance matrix is constructed using the within-class mean of each block of each preset expression category to obtain a feature difference matrix. The within-class mean of the block is a linear combination of the within-class images. Therefore, the within-class mean of each block subset corresponding to each block retains a large amount of variation of a specific image, reducing the probability of loss of the main image information. Moreover, when constructing the part feature space, the inter-class differences between different preset expression categories are focused on, and the obtained feature space has stronger distinguishability. The part feature space is constructed based on the feature difference matrix, and dimensionality reduction is performed by projecting the blocks into the corresponding part feature space, thereby omitting the process of converting the image into a one-dimensional vector, reducing the computational workload and saving the time of vectorization, and retaining the structured information. Therefore, without increasing the computational cost, the inter-class separability is improved.
[0074] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an implementation manner of the expression feature extraction method of this application. The method includes:
[0075] S401: Obtain the image to be recognized, and extract the to-be-recognized blocks corresponding to each of multiple types of preset parts from the image to be recognized.
[0076] Specifically, after obtaining the image to be recognized, extract the to-be-recognized sub-blocks corresponding to multiple types of preset parts from the image to be recognized. That is to say, divide the image to be recognized into multiple types of preset parts.
[0077] S402: Based on the to-be-recognized sub-blocks corresponding to the preset parts of the same type and the part feature space corresponding to the preset parts, obtain the to-be-recognized expression features corresponding to the preset parts, and determine the part similarity between the to-be-recognized expression features and the reference expression features corresponding to multiple preset expression categories.
[0078] Specifically, project the to-be-recognized sub-blocks corresponding to the preset parts into the part feature space corresponding to the preset parts to obtain the to-be-recognized expression features for postgraduate study of the preset parts.
[0079] Furthermore, for each type of preset part, compare the to-be-recognized expression features with the reference expression features corresponding to multiple preset expression categories to obtain the part similarity between the to-be-recognized expression features and the reference expression features corresponding to multiple preset expression categories.
[0080] It should be noted that the part feature space and the reference expression features are obtained based on the expression feature extraction method in any of the above embodiments. For the description of related content, please refer to the detailed description of the above method embodiments and will not be elaborated here.
[0081] S403: Based on the part similarity corresponding to each type of preset part, obtain the target expression category corresponding to the image to be recognized.
[0082] Specifically, after obtaining the part similarity corresponding to each type of preset part, fuse the part similarity corresponding to all types of preset parts to obtain a fused similarity, and thus determine the target expression category corresponding to the image to be recognized based on the fused similarity. By detecting the sub-blocks of the image to be recognized and extracting the to-be-recognized expression features for different types of preset parts, the accuracy of expression recognition is improved. By fusing the part similarity of different types of preset parts and determining the expression category in the image to be recognized based on the fused similarity, the accuracy of expression recognition is improved.
[0083] In an application mode, based on the to-be-recognized sub-blocks corresponding to the preset parts of the same type and the part feature space corresponding to the preset parts, obtain the to-be-recognized expression features corresponding to the preset parts, and determine the part similarity between the to-be-recognized expression features and the reference expression features corresponding to multiple preset expression categories, including: projecting the to-be-recognized sub-blocks corresponding to the preset parts of the same type into the part feature space corresponding to the preset parts to obtain the to-be-recognized expression features corresponding to the preset parts; determining the part Euclidean distance between the to-be-recognized expression features and the reference expression features corresponding to each preset expression category; wherein, the part similarity is negatively correlated with the part Euclidean distance.
[0084] Specifically, the to-be-recognized chunks corresponding to the preset parts of the same type are projected onto the part feature space corresponding to the preset parts, so as to reduce the dimension of the to-be-recognized chunks, obtain the to-be-recognized expression features corresponding to the preset parts, omit the process of converting the image into a one-dimensional vector, reduce the computational workload and save the time for vectorization.
[0085] Furthermore, for the preset parts of the same type, calculate the Euclidean distance between the to-be-recognized expression features and the reference expression features corresponding to each preset expression category. After traversing all types of preset parts, obtain the Euclidean distances between all types of preset parts corresponding to the same preset expression category within the comparison dimension. The part similarity is negatively correlated with the Euclidean distance between parts. When the Euclidean distance between parts is smaller, the part similarity is higher.
[0086] In an application scenario, based on the part similarities corresponding to all types of preset parts, obtain the target expression category corresponding to the to-be-recognized image, including: within the comparison dimension corresponding to each preset expression category, perform a weighted sum of the Euclidean distances between parts corresponding to all types of preset parts to obtain the Euclidean distance of the image corresponding to the to-be-recognized image within each comparison dimension; take the preset expression category with the smallest Euclidean distance of the image as the target expression category corresponding to the to-be-recognized image.
[0087] Specifically, since the changes of different types of preset parts have different effects on expression recognition, in order to make full use of the differences of various types of preset parts under different expression types, different types of preset parts are assigned weight factors. Within the comparison dimension corresponding to each preset expression category, perform a weighted sum of the Euclidean distances between parts corresponding to all types of preset parts, so as to fuse the Euclidean distances between parts corresponding to all types of preset parts on the to-be-recognized image, obtain the Euclidean distance of the image corresponding to the to-be-recognized image within each comparison dimension, and thus combine the differences between different types of preset parts, distinguish the influence of different preset parts on expression recognition by adjusting the weight factors, and improve the accuracy of expression recognition.
[0088] Furthermore, take the preset expression category with the smallest Euclidean distance of the image as the target expression category corresponding to the to-be-recognized image, that is, take the preset expression category with the highest similarity as the target expression category corresponding to the to-be-recognized image.
[0089] In a specific application scenario, the preset parts include eyes, nose, and mouth. Within the comparison dimension corresponding to each preset expression category, respectively obtain the Euclidean distances between parts of the eyes, nose, and mouth as d1, d2, and d3, perform a weighted sum of the Euclidean distances between parts to obtain the Euclidean distance L of the image, where L = α1d1 + α2d2 + α3d3. Take the preset expression category with the smallest Euclidean distance L of the image as the target expression category corresponding to the to-be-recognized image.
[0090] In this embodiment, through block detection, for different types of preset parts, the expression features to be recognized are obtained by using the part feature space respectively. Corresponding weight factors are set for different types of preset parts. Within the comparison dimension corresponding to each preset expression category, the Euclidean distances of the parts corresponding to all types of preset parts are weighted and summed, so as to fuse the Euclidean distances of the parts corresponding to all types of preset parts on the image to be recognized, and obtain the Euclidean distance of the image corresponding to the image to be recognized within each comparison dimension. The preset expression category with the smallest Euclidean distance of the image is used as the target expression category corresponding to the image to be recognized. Thus, by combining the differences between different types of preset parts and adjusting the weight factors to distinguish the influence of different preset parts on expression recognition, the accuracy of expression recognition is improved.
[0091] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of an embodiment of an electronic device according to the present application. The electronic device 50 includes a memory 501 and a processor 502 that are coupled to each other. Among them, the memory 501 stores program data (not shown in the figure), and the processor 502 calls the program data to implement the method in any of the above embodiments. For the description of related content, please refer to the detailed description of the above method embodiments, and details are not described herein again.
[0092] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of an embodiment of a computer-readable storage medium according to the present application. The computer-readable storage medium 60 stores program data 600, and when the program data 600 is executed by a processor, it implements the method in any of the above embodiments. For the description of related content, please refer to the detailed description of the above method embodiments, and details are not described herein again.
[0093] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0095] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0096] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A method for extracting facial expression features, characterized in that, The method includes: Obtaining a plurality of sample images, and extracting a plurality of blocks corresponding to a plurality of types of preset parts from the sample images; wherein, at least some of the preset expression categories of the sample images are different; Based on the blocks corresponding to the preset parts of the same type, determining the total feature and a plurality of sub-features corresponding to the preset parts; wherein, the total feature is related to all the blocks corresponding to the preset parts, and each sub-feature is respectively related to the blocks corresponding to one preset expression category on the preset part; Based on the differences between the plurality of sub-features and the total feature respectively, determining a feature difference matrix corresponding to the preset part; Using the feature difference matrix to construct a part feature space corresponding to the preset part, and based on the blocks corresponding to the preset part and the part feature space, obtaining reference expression features corresponding to each preset expression category on the preset part.
2. The expression feature extraction method according to claim 1, wherein The obtaining a plurality of sample images, and extracting a plurality of blocks corresponding to a plurality of types of preset parts from the sample images includes: Classifying the plurality of sample images based on the preset expression category corresponding to each sample image, to obtain a subset of sample images corresponding to each preset expression category; For each type of preset part, respectively extracting a plurality of blocks of the preset part from each subset of sample images, to obtain a subset of blocks corresponding to each preset expression category, and based on all the subsets of blocks corresponding to the preset part, obtaining a set of blocks corresponding to the preset part.
3. The expression feature extraction method according to claim 2, wherein The determining the total feature and a plurality of sub-features corresponding to the preset part based on the blocks corresponding to the preset parts of the same type includes: Based on the mean value of all the blocks in the set of blocks corresponding to the preset part, obtaining the total mean value of the blocks corresponding to the preset part, and based on the mean value of all the blocks in each subset of blocks corresponding to the preset part, obtaining the within-class mean value of the blocks corresponding to each preset expression category on the preset part; wherein, the total feature corresponds to the total mean value of the blocks, and the sub-feature corresponds to the within-class mean value of the blocks; The determining the feature difference matrix corresponding to the preset part based on the differences between the plurality of sub-features and the total feature respectively includes: Using the difference between each within-class mean value of the blocks and the total mean value of the blocks to calculate a covariance matrix, to obtain the feature difference matrix corresponding to the preset part.
4. The facial expression feature extraction method according to claim 1, wherein The using the feature difference matrix to construct a part feature space corresponding to the preset part, and based on the blocks corresponding to the preset part and the part feature space, obtaining reference expression features corresponding to each preset expression category on the preset part includes: Using at least some of the eigenvalues in the feature difference matrix to construct a part feature space corresponding to the preset part; Projecting the blocks corresponding to each preset expression category on the preset part onto the part feature space, to obtain reference expression features corresponding to each preset expression category on the preset part.
5. The expression feature extraction method according to claim 4, wherein Constructing the part feature space corresponding to the preset part by using at least some of the eigenvalues in the feature difference matrix includes: Selecting a preset number of specified eigenvalues from the feature difference matrix based on the values corresponding to the eigenvalues in the feature difference matrix; Obtaining a projection matrix based on the orthogonal eigenvectors corresponding to the specified eigenvalues, and constructing the part feature space corresponding to the preset part by using the projection matrix; wherein, the dimension of the part feature space corresponds to the preset number.
6. The expression feature extraction method according to claim 1 or 2, characterized in that Before obtaining a plurality of sample images and extracting a plurality of sub-blocks corresponding to each of a plurality of types of preset parts from the sample images, it includes: Obtaining a plurality of initial sample images, and performing preprocessing operations on the initial sample images to obtain the sample images; wherein, the preprocessing operations include at least one of angle correction, homomorphic filtering, and normalization processing.
7. A method for facial expression recognition, characterized in that, The method includes: Obtaining an image to be recognized, and extracting sub-blocks to be recognized corresponding to each of a plurality of types of preset parts from the image to be recognized; Based on the sub-blocks to be recognized corresponding to the preset part of the same type and the part feature space corresponding to the preset part, obtaining the expression feature to be recognized corresponding to the preset part, and determining the part similarity between the expression feature to be recognized and the reference expression features corresponding to a plurality of preset expression categories; Based on the part similarities corresponding to each of all types of the preset parts, obtaining the target expression category corresponding to the image to be recognized; Wherein, the part feature space and the reference expression features are obtained by the expression feature extraction method described in any one of the above claims 1-6.
8. The facial expression recognition method according to claim 7, characterized in that Based on the sub-blocks to be recognized corresponding to the preset part of the same type and the part feature space corresponding to the preset part, obtaining the expression feature to be recognized corresponding to the preset part, and determining the part similarity between the expression feature to be recognized and the reference expression features corresponding to a plurality of preset expression categories, includes: Projecting the sub-blocks to be recognized corresponding to the preset part of the same type into the part feature space corresponding to the preset part to obtain the expression feature to be recognized corresponding to the preset part; Determining the Euclidean distance between the expression feature to be recognized and the reference expression features corresponding to each preset expression category; wherein, the part similarity is negatively correlated with the Euclidean distance between parts.
9. The expression recognition method according to claim 8, characterized in that Based on the part similarities corresponding to each of all types of the preset parts, obtaining the target expression category corresponding to the image to be recognized, includes: Within each comparison dimension corresponding to the preset expression category, performing a weighted sum of the Euclidean distances between parts corresponding to all types of the preset parts to obtain the Euclidean distance of the image corresponding to the image to be recognized within each comparison dimension; Taking the preset expression category with the smallest Euclidean distance of the image as the target expression category corresponding to the image to be recognized.
10. An electronic device, characterized in that, It includes: A memory and a processor that are coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described in any one of claims 1-6 or 7-9.
Citation Information
Patent Citations
Gait recognition method, device and system based on leg features and readable storage medium
CN111950418A
Facial expression recognition method and device, equipment and storage medium
CN113869234A