Facial expression recognition method based on multi-manifold learning

Through the multi-manifold learning method, significant areas of the face are detected and fractional Fourier transform and multi-manifold identification analysis algorithm are used to solve the problem of poor expression recognition effect in traditional methods, and efficient and robust facial expression recognition is achieved.

CN120544247APending Publication Date: 2025-08-26ZHENGZHOU UNIVERSITY OF AERONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510583875.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the existing facial expression recognition methods, the traditional single manifold learning hypothesis ignores the manifold distribution characteristics of different expressions, resulting in poor recognition effect and registration disasters, making it difficult to effectively distinguish different expressions.

Method used

Using a multi-manifold learning method, the significant areas of the face are detected through the active shape model, combined with fractional Fourier transform and multi-manifold discrimination analysis algorithm, the divergence between manifolds and the internal divergence of manifolds is maximized, the expression identification matrix is ​​extracted, and the error criterion of local linear embedding reconstruction is used for classification.

Benefits of technology

It significantly improves the accuracy and efficiency of facial expression recognition, can maintain robustness in complex environments, reduce calculation complexity, meet real-time application needs, and has a recognition rate of 95.83%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544247A_ABST
    Figure CN120544247A_ABST
Patent Text Reader

Abstract

The invention discloses a facial expression recognition method based on multi-manifold learning, and belongs to the field of computer vision and pattern recognition. The method comprises the following steps: firstly, detecting a facial salient region by using an active shape model and constructing an expression database; secondly, carrying out image preprocessing on the salient region by adopting fractional Fourier transform, and strengthening features; then, extracting an identification matrix corresponding to each expression category through a multi-manifold identification analysis method, so that the intra-manifold distance between samples of the same type is minimized, and the inter-manifold distance between samples of different types is maximized; and finally, projecting the test sample to the identification space of each manifold, calculating the distance between the test sample and each manifold, and selecting the expression category corresponding to the manifold with the minimum distance as the identification result of the test sample. The method focuses on a face salient region, effectively improves the accuracy and robustness of expression recognition through 2D-FrFT preprocessing and an M2DA optimization algorithm, and can be applied to the fields of intelligent human-computer interaction, sentiment analysis and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and pattern recognition, and in particular relates to a facial expression recognition method based on multi-manifold learning. Background Art

[0002] Facial expression recognition, a key tool for human-computer interaction, is widely used in a variety of real-life scenarios, including security monitoring, sentiment analysis, and virtual reality. Because facial expressions convey a wide range of human emotions and intentions, accurately recognizing facial expressions is crucial for understanding human emotions and enhancing the interactive capabilities of intelligent systems.

[0003] Current facial expression recognition methods can be roughly divided into two categories based on the different feature extraction methods: geometric feature-based and facial feature-based. The former mainly uses techniques such as template matching and image matching to extract local facial features for expression recognition. The latter is typically based on methods such as principal component analysis (PCA), linear discriminant analysis (LDA), PCA combined with LDA, active surface model (AAM), active shape model (ASM), independent component analysis (ICA), and their improved algorithms.

[0004] However, traditional feature dimensionality reduction methods such as Locality Preserving Projections (LPP) do not consider the sample's category information (using an unsupervised approach). Therefore, after embedding facial images into a low-dimensional manifold space, they cannot be well classified. To remedy this problem, supervised LPP introduces category information to construct a neighborhood graph, achieving significant results in facial expression recognition. Shan et al. conducted a comparative experiment comparing several commonly used facial expression recognition algorithms, including LPP, Supervised Locality Preserving Projections (SLPP), Locality Sensitive Discriminant Analysis (LSDA), PCA, and LDA. Experimental results show that SLPP achieves the best recognition results among these algorithms.

[0005] However, these methods generally assume that all expression samples are distributed in the same manifold space. Whether a single manifold space can effectively represent different expression features remains unproven. To address this issue, Xiao et al. proposed a multi-manifold facial expression recognition method (Facial Expression Recognition Based on Multiple Manifolds [J], Pattern Recognition, 2011, Vol. 44, No. 1: pp. 107-116). This method assumes that different expressions reside in different manifold spaces, allowing for the extraction of discriminative features for each expression (manifold). Experimental results demonstrate that the multi-manifold method outperforms the single-manifold method in facial expression recognition and indicate that local regions of the face contain more discriminative information about expressions than the entire face. Based on this finding, Chang et al. further constructed a training set using local image patches to verify the effectiveness of local facial regions in expression recognition. Experimental results show that, in most cases, the local method achieves better recognition performance than the global facial representation method. Kotsia et al. also demonstrated that even when local facial regions are occluded, local features can still provide stronger recognition capabilities. In addition, Lu et al. proposed the Discriminant Locality Preserving Projections (DLPP) method based on the difference criterion. This method maximizes the difference between the locally preserved inter-class scatter and the locally preserved intra-class scatter for feature extraction. Song et al. proposed a Multiple Maximum Scatter Difference (MMSD) criterion for multi-class classification. However, the computational complexity of MMSD remains high for high-dimensional datasets, making it unsuitable for practical applications.

[0006] Given the significant limitations of traditional single-manifold learning methods in revealing the structural information of facial expressions, these methods are based on the assumption that all samples lie in the same manifold space, assuming that all expressions reside on the same manifold. However, this assumption ignores the unique manifold distribution characteristics of each expression, resulting in suboptimal expression recognition performance. To overcome the limitations of traditional methods, emerging technologies such as subspace analysis-based feature extraction methods and multi-manifold discriminant analysis algorithms have emerged in recent years. These methods combine statistical analysis with manifold learning techniques to project facial images onto optimal subspaces or multiple manifolds, thereby achieving more effective extraction and classification of facial expression features. However, existing multi-manifold learning methods still face key issues such as the "registration curse" (misregistration). Designing more optimized multi-manifold learning expression feature extraction algorithms to effectively address data registration challenges is crucial for improving expression recognition performance. Summary of the Invention

[0007] The purpose of the present invention is to provide a facial expression recognition method based on multi-manifold learning, aiming to overcome the problems existing in existing expression recognition methods such as insufficient local feature extraction, high manifold learning complexity and poor classification effect. Based on feature extraction of facial salient areas and multi-manifold learning technology, this method can effectively distinguish different expressions and has high facial expression recognition accuracy and efficiency.

[0008] To achieve the above object, the technical solution adopted in the present invention is: a facial expression recognition method based on multi-manifold learning, comprising the following steps:

[0009] Step S1: using an active shape model to detect salient areas in a facial image, and classifying the detected salient areas to construct an expression database;

[0010] Step S2, using fractional Fourier transform to perform frequency domain conversion preprocessing on the salient area images in the expression database, so that the images are jointly expressed in the time-frequency plane, and a preprocessed training set is obtained;

[0011] Step S3, extracting the discriminant matrix corresponding to each expression category by a multi-manifold discriminant analysis algorithm, wherein the multi-manifold discriminant analysis algorithm optimizes the projection matrix with the optimization objectives of maximizing the inter-manifold divergence and minimizing the intra-manifold divergence;

[0012] Step S4: Project the salient area of ​​the test sample processed by steps S1 and S2 into the discriminant space of each manifold, calculate the distance between the test sample and each manifold, and select the expression category corresponding to the manifold with the smallest distance as the recognition result of the test sample.

[0013] Furthermore, in step S1, the salient areas include the left eye, the right eye, the left cheek, the right cheek and the mouth.

[0014] Furthermore, in step S1, the specific steps of detecting the facial image using the active shape model include:

[0015] S11. Define the manifold set of the training set The manifold set of the kth type of expression is The salient block of the i-th sample of the k-th class is Each sample The size of the salient block is a×b; where c is the number of expression categories in the expression library, n is the total number of samples, k = 1, 2, ..., c, i = 1, 2, ..., n k , n k is the number of samples of the k-th type of expression, is the i-th sample of the k-th expression, d is the original dimension of the salient block, t is the number of salient blocks, l k =t·n k and

[0016] S12. The active shape model locates facial key points, extracts five salient areas: the left eye, the right eye, the left cheek, the right cheek, and the mouth, and adjusts the size of the salient blocks in each area to a uniform size.

[0017] Furthermore, the specific steps of step S2 include: first performing a one-dimensional discrete fractional-order Fourier transform, and then performing a two-dimensional fractional-order Fourier transform, with the transform order parameters P1 and P2 both set to 0.5.

[0018] Furthermore, the specific implementation process of step S2 is:

[0019] Step S21: perform one-dimensional discrete fractional Fourier transform according to the following formula:

[0020] Where j is the imaginary unit, that is α represents the angle of rotation of the fractional-order transform in the time-frequency domain plane, and satisfies the relationship p is the transform order of FrFT, and p≠2n, n is an integer; F p represents the p-order fractional Fourier transform operation; f p(u) represents the function of variable u obtained by processing the function f(u′) by p-order fractional Fourier transform;

[0021] Step S22: perform a two-dimensional fractional Fourier transform using the following discrete form:

[0022] Where x(p,q) is the input original image, p and q are the row index and column index of the original image respectively; M and N are the number of rows and columns of the original image respectively; is the output image obtained after two-dimensional fractional Fourier transform, where m and n are the row and column coordinates of the output image respectively; is the two-dimensional fractional Fourier kernel function.

[0023] Furthermore, step S3 specifically includes: extracting multiple projection matrices from the training set of step S2 according to different expression categories, and projecting the salient blocks of each type of sample in the training set to the space; defining inter-manifold and intra-manifold neighbor sample sets, and under the conditions of maximizing the distance between manifolds and minimizing the distance within the manifold, successively calculating the projection matrix of each type of expression, and calculating the dimension of each category's identification matrix, based on the dimensionality requirement, screening out the most discriminative eigenvectors, and finally constructing a discrimination matrix that can effectively distinguish expressions.

[0024] Furthermore, in step S3, when extracting the discriminant matrix of each manifold through the multi-manifold discriminant analysis algorithm, define and They are The inter-manifold and intra-manifold neighbor sample sets, where j = 1, 2, ..., t, assume that a point for K b nearest neighbor, k b represents the number of neighbors of samples between manifolds, then in and Located on different manifolds; let a point for K w nearest neighbor, k w represents the number of neighbors of the sample in the manifold, then in and are on the same manifold;

[0025] The optimization goal is to maximize in is the inter-manifold scatter matrix of the k-th type of expression, is the scatter matrix within the manifold of the kth type of expression, and the discriminant vectors of each type satisfy The orthogonality relationship, and v,ε=1,2,…,d k .

[0026] Furthermore, the inter-manifold scatter matrix The calculation formula is:

[0027] Where n k represents the number of samples of the kth type of expression, t is the number of significant areas of the sample, k b is the number of neighbors of a sample between manifolds, is the jth significant block of the i-th sample of the k-th expression, for The nearest neighbor between manifolds, is the weight matrix between manifolds, M k is the correlation matrix with the k-th type of expression sample; ∑ k The entity is l k ×l b matrix, and is a diagonal matrix, and its entities are Columns and rows, that is and

[0028] The scatter matrix within the manifold The calculation formula is:

[0029] Where k w is the number of neighbors of the sample in the manifold, for The same manifold neighbors of is the weight matrix within the manifold; Represents the weight matrix, which is used to describe the connection weight between the sample and its neighbors; D k is a diagonal matrix, and its entity is Column vector of .

[0030] Furthermore, the weight matrix They are defined as:

[0031] Where ||·|| is the Euclidean norm and κ is a constant.

[0032] Furthermore, the design process of the classifier in step S4 is as follows:

[0033] Step S41: Let Φ be an expression test sample, and divide it into five significant regions after detecting facial key points using the active shape model. These five significant regions are used as a manifold M Φ =[φ1,…,φ t ];

[0034] Step S42: Make M Φ Projection to the discriminant space W k On the other hand, we get the projection matrix at the same time, M k In W k The projection matrix;

[0035] Step S43: Calculate M Φ and M k The distance d on the low-dimensional manifold Φk (M Φ ,M k ):

[0036] in,

[0037] Where, For Y Φj K l Neighbors, is the nearest neighbor reconstruction error factor;

[0038] Step S44: Set c * is the category label of the test sample Φ, through c * =arg min k=1,2,…,c d k (M Φ ,M k ) calculates the category information of the test sample.

[0039] The beneficial effects of the above scheme are:

[0040] (1) This invention is based on the assumption that different expressions are located in different manifold spaces, using M 2 The DA algorithm extracts discriminative matrices for each expression, maximizing inter-manifold distances and minimizing intra-manifold distances. This allows for effective differentiation of similar expressions (such as happiness and surprise). Because the algorithm deeply explores differences in the underlying geometric structure of expression data, it goes beyond surface features and understands the patterns of expression variation at a more fundamental level. This significantly improves the ability to discriminate expression features, expanding the depth and breadth of expression recognition applications. On the Cohn-Kanade database, the proposed method achieved a recognition rate of up to 95.83%.

[0041] (2) The present invention uses a two-dimensional fractional Fourier transform to pre-process the image of the salient area, which can jointly express the image from the time-frequency plane, strengthen the local structure and dynamic texture information, make the expression features richer and more discriminative, and help improve the recognition accuracy. In practical applications, since illumination changes and local occlusions can reduce the stability of traditional image features (such as grayscale and texture), 2D-FrFT can distinguish interference features from expression features in the frequency domain, remove interference through operations such as frequency domain filtering, retain expression-related features, effectively recognize expressions, and enhance the robustness of the algorithm in complex environments.

[0042] (3) The present invention uses the local linear embedding (LLE) reconstruction error criterion to calculate the distance between significant blocks and complete expression classification. The LLE reconstruction error criterion has good adaptability to expression data of different types and sources, can handle diverse expression features, and more accurately measure the similarity between the test sample and various expression samples (such as distinguishing between smiles and fake smiles). By accurately classifying by reasonably calculating the distance, the expression category of the test sample is determined. Compared with traditional classification methods, the accuracy of expression recognition is significantly improved, and it has obvious advantages in large-scale expression data testing.

[0043] (4) The present invention selects only five significant expression regions for feature extraction, abandoning the analysis of the entire facial structure, avoiding interference from non-expression areas of the face, and can more accurately extract key features related to expression. Compared with traditional methods, the selection of significant regions reduces the amount of data processing, reduces computational complexity, and improves algorithm operation efficiency, enabling expression recognition to be completed in a shorter time, meeting the requirements of real-time applications. Moreover, even if other parts of the face are blocked, it does not affect the extraction of effective expression features from these key regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flow chart of the present invention;

[0045] Figure 2 There are 24 samples in the original space and M 2 Distribution after DA projection. DETAILED DESCRIPTION

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] It should be noted that unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0048] like Figure 1 、 Figure 2 As shown, a facial expression recognition method based on multi-manifold learning includes the following steps:

[0049] Step S1: using an active shape model to detect salient areas in a facial image, and classifying the detected salient areas to construct an expression database;

[0050] Step S2, using fractional Fourier transform to perform frequency domain conversion preprocessing on the salient area images in the expression database, so that the images are jointly expressed in the time-frequency plane, and a preprocessed training set is obtained;

[0051] Step S3, extracting the discriminant matrix corresponding to each expression category by a multi-manifold discriminant analysis algorithm, wherein the multi-manifold discriminant analysis algorithm optimizes the projection matrix with the optimization objectives of maximizing the inter-manifold divergence and minimizing the intra-manifold divergence;

[0052] Step S4: Project the salient area of ​​the test sample processed by steps S1 and S2 into the discriminant space of each manifold, calculate the distance between the test sample and each manifold, and select the expression category corresponding to the manifold with the smallest distance as the recognition result of the test sample.

[0053] The following is a detailed description of the implementation process of each step:

[0054] Step S1: Assume that the expression database contains c categories of expressions, a total of n samples, and each sample is Indicates that k = 1, 2, ..., c, i = 1, 2, ..., n k , where n k represents the number of samples of the kth type of expression, Represents the i-th sample of the k-th expression. The Active Shape Model (ASM) method is used to extract the salient regions of each facial image. There are five salient regions: left eye, right eye, left cheek, right cheek, and mouth. The size of each salient block is adjusted to a×b. Assume is the manifold set of the training set, is the manifold set of the k-th type of expression, represents the salient block of the i-th sample of the k-th category, t=5 is the number of salient blocks, l k =t·n k , d represents the original dimension of the salient patch image features.

[0055] Step S2: First, perform one-dimensional discrete fractional Fourier transform (FrFT) processing according to formula (1).

[0056] Where j is the imaginary unit, that is α represents the angle of rotation of the fractional-order transform in the time-frequency domain plane, and satisfies the relationship p is the transform order of FrFT, and p≠2n, n is an integer; F p represents the p-order fractional Fourier transform operation; f p(u) represents the function of variable u obtained by processing the function f(u′) by p-order fractional Fourier transform;

[0057] Then, a two-dimensional fractional Fourier transform (2D-FrFT) is performed according to formula (2), specifically using the following discrete form:

[0058] Suppose the original image is x(p,q), p and q are the row index and column index of the original image respectively, M and N are the number of rows and columns of the input image respectively. After performing a two-dimensional fractional Fourier transform on it, the output image is obtained Where m and n are the row coordinates and column coordinates of the output image respectively. The two-dimensional FrFT transform expression is:

[0059] Where, is a two-dimensional fractional Fourier kernel function. Equivalently, the two-dimensional FrFT transform can be decomposed into two one-dimensional FrFT operations, expressed as:

[0060] Where, is the fractional order kernel function along the horizontal direction, is the fractional order kernel function along the vertical direction. The fractional order parameter is set to p1 = p2 = 0.5, that is, both are 1 / 2 order fractional Fourier transforms.

[0061] In order to effectively overcome the interference of illumination changes, local occlusion and geometric transformation on the expression of image features, the present invention uses fractional-order Fourier transform to perform frequency domain conversion on the salient areas of the image, thereby enhancing the local structure and texture information of the image and improving the discriminative ability of feature extraction.

[0062] In practice, the transform order of the one-dimensional discrete Fourier transform (FrFT) is flexible and adjustable, enabling precise adaptation to the time-frequency characteristics of different signals. In particular, for non-stationary signals, selecting the optimal transform order allows the signal energy to be highly concentrated within the fractional frequency domain, laying a solid foundation for subsequent analysis and processing. Furthermore, the two-dimensional fractional Fourier transform (2D-FrFT) can deeply explore the underlying geometric structure and distribution information of an image. Different expressions exhibit unique energy distributions and characteristic pattern differences in the two-dimensional fractional frequency domain. These features can serve as an important basis for expression classification, significantly improving the accuracy of expression recognition.

[0063] Step S31: Extract c projection matrices, W1, W2, ..., W from the training set after FrFT transformation c The salient blocks of each class of samples in the training set are projected into this space, expressed as: Make Y k Under a certain optimal criterion, M can be better represented k ,in d and d k They represent the original dimension and the projected dimension of the salient block image features respectively.

[0064] Step S32: Definition and They are The inter-manifold and intra-manifold neighbor sample sets, where j = 1, 2, ..., t. Assume that a point for K b nearest neighbor, k b represents the number of neighbors of samples between manifolds, then in and Located on different manifolds. Similar to the above definition, let a point for K w nearest neighbor, k w represents the number of neighbors of the sample in the manifold, then in and are located on the same manifold. When projecting features, the local structure within each manifold is preserved while the distance information between manifolds is maintained, that is, the distance between manifolds is maximized while the distance within the manifold is minimized. Finally, by solving the optimization problem of formula (4), the discriminant matrices W1, W2, ..., W c :

[0065] Where, trace() represents the trace of the matrix, are the weight matrix and the inter-manifold scatter matrix of the k-th expression respectively The intra-manifold scatter matrix of the kth type of expression They are defined as:

[0066] Where ||·|| is the Euclidean norm and κ is a constant.

[0067] From formula (4), we can see that the discriminant matrix of each expression needs to be maximized. Minimize at the same time In addition, in order to avoid the influence of different lengths of each discriminant vector and to ensure the linear independence of the vectors, the usual method is to make the discriminant matrix have an orthogonal relationship. For the above reasons, the discriminant vectors of each type should satisfy Under the constraints of and v,ε=1,2,…,d k , δ vε is the Kronecker symbol.

[0068] According to the above analysis, the original optimization problem (Formula (4)) is expressed in the following form:

[0069] In solving the discrimination matrix W1, W2,…, W c In the process, the orthogonality constraint of the discriminant vector is added to make the obtained discriminant matrix more effective and reasonable in distinguishing different expressions.

[0070] Step S33: Since it is difficult to find the optimal discriminant matrix for each expression using Equation (6) in practical applications, to solve this problem, the Fisher linear discriminant method is used to simplify the complex optimization problem to solving the projection matrix class by class. The projection matrix for each type of expression is found one by one using Equation (7).

[0071] Among them, the inter-manifold scatter matrix writing:

[0072] Where n k represents the number of samples of the kth type of expression, t is the number of significant areas of the sample, k b is the number of neighbors of a sample between manifolds, is the jth significant block of the i-th sample of the k-th expression, for The nearest neighbor between manifolds, is the weight matrix between manifolds, M k is the correlation matrix with the k-th type of expression sample; Σ k The entity is l k ×l b matrix, and is a diagonal matrix, and the entities are Columns and rows, that is and The degree of sample dispersion between different expression categories is calculated using formula (8).

[0073] Manifold scatter matrix writing:

[0074] Where k w is the number of neighbors of the sample in the manifold, for The same manifold neighbors of is the weight matrix within the manifold; Represents the weight matrix, which is used to describe the connection weight between the sample and its neighbors; D k is a diagonal matrix, and its entity is The column vector of . The discrete degree of samples within the same expression category is calculated by formula (9).

[0075] Then, the discriminant matrices of various types are extracted through formula (10), which enables the effective extraction and differentiation of expression features.

[0076] Where v = 1, 2, ..., d k , Respectively represent the corresponding to the first d k The eigenvector with the largest eigenvalue, is the corresponding eigenvalue, representing the eigenvector The ratio of the between-class divergence to the within-class divergence in the indicated direction.

[0077] Step S34: Solve the discriminant matrix of each manifold in turn by using the idea of ​​Directly Linear Discriminant Analysis (DLDA). At the same time, set the dimension d of each discriminant matrix k , the dimension of each type (manifold) discriminant matrix is ​​calculated with the help of the “trace ratio” method.

[0078] Given that and They are all semi-positive matrices. The eigenvector corresponding to the maximum eigenvalue is screened out by taking the maximum eigenvalue one by one, as shown in formula (11):

[0079] like This means that adding the eigenvector After that, the salient blocks of the same type (same manifold) are closer, while the salient blocks of different types (different manifolds) are farther away. Then the feature matrix of the kth type of expression is obtained:

[0080] The specific steps of classifier design in step S4 include:

[0081] Step S41: Let Φ be an expression test sample, detect facial key points by ASM method and divide it into five significant regions, perform FRFT preprocessing on the significant regions, and take the five significant regions as a manifold M Φ =[φ1,…,φ t ].

[0082] Step S42: Make M Φ Projection to the discriminant space W k On the other hand, we get the projection matrix at the same time, M k In W k The projection matrix, where

[0083] Step S43, M Φ and M k The distance on the low-dimensional manifold is dΦk (M Φ ,M k )express:

[0084] in,

[0085] Where, For Y Φj K l Neighbors, is the nearest neighbor reconstruction error factor. The optimization problem of Equation (13) is solved with the help of the reconstruction error criterion of LLE (local linear embedding), which measures the similarity between the test sample and various expression samples in the low-dimensional feature space.

[0086] Step S44: Set c * is the category label of the test sample Φ, and the expression category information of the test sample is calculated by the following discriminant, that is, the expression category with the smallest distance from the test sample is found as its classification result. c * =arg min k=1,2,…,c d k (M Φ ,M k ) (14)

[0087] In order to verify the Multiple Manifolds Discriminant Analysis (M 2 The advantages of the DA algorithm in terms of computational complexity are compared with the existing traditional expression recognition methods through simulation experiments, including PCA, PCA+LDA, Modular PCA and MMSD algorithms. The main configuration of the computer used in the experiment is: Intel Celeron CPU and 2GB RAM. The parameter settings involved in each method are as follows: the dimension of the PCA algorithm is set to nc to avoid the small sample problem; for Modular PCA, the size of each module is set to 16×16; the value of σ is selected by cross-validation; in order to fairly compare the above algorithms, the feature dimension corresponding to the best recognition effect is selected in the experiment. For M 2 DA algorithm, in the experiment, k is set based on experience b and k wThe recognition results of different traditional subspace methods on the three expression databases of Cohn-Kanade, Jaffe (The Japanese Female Facial Expression) and RML (Ryerson Multimedia Research Lab) are shown in Table 1. Table 1 Recognition results of different traditional subspace methods in three expression databases (%)

[0088] It can be clearly seen from the data in Table 1 that M 2 DA performs best in terms of recognition effect. This is because compared with other algorithms based on the whole face (such as PCA, PCA+LDA and MMSD), MD 2 DA can extract the discriminative features of specific expressions from local areas of the face, thus avoiding the influence of other non-expression areas on the recognition results. 2 DA is a supervised method that can accurately extract local expression features. Modular PCA is an unsupervised local feature extraction method. 2 The recognition rate of DA is significantly higher than that of Modular PCA.

[0089] To further evaluate M 2 To investigate the effectiveness of DA, we compare it with several major manifold learning algorithms, including LLE, LPP, DLPP, SLPP, Xiao's, Marginal Fisher Analysis (MFA), and S-OLPP (Supervised Orthogonal Locality Preserving Projections). With the exception of Xiao's, which is a multi-manifold learning method, all other algorithms are single-manifold learning methods. Simulation experiments were conducted on three facial expression databases. The parameters used in the experiments are as follows: for the LPP method, the heat kernel parameter was set to 5; for the MFA method, the number of inter-class and intra-class neighbors was set to 5 and 15, respectively; for the S-OLPP method, the dimensionality was reduced to 20; and for the SLPP method, the parameter was set to 1. For Xiao's method, the training set was first split into two parts: feature training and parameter adjustment. Specifically, 50% of the facial expression samples were selected for training the manifold learning model, 25% for parameter adjustment, and the remaining 25% for testing. In addition, to deal with the small sample problem, LPP, DLPP, MFA, S-OLPP and SLPP all use PCA to reduce the dimension to nc.

[0090] The experimental simulation results in Table 2 show that the multi-manifold learning algorithm is significantly better than the single-manifold learning algorithm. This shows that assuming that similar expressions are located in the same manifold, it can effectively reveal the underlying geometric structure of specific expressions, and this revelation is not affected by the specific individual. In other words, the goal of multi-manifold learning is to discover the low-dimensional multi-manifold structure of facial expressions in high-dimensional space. At the same time, M 2 The performance of DA method is always better than Xiao's method, which shows that in low-dimensional manifold space, M 2 DA can more effectively encode more discriminative information, thereby preserving the local characteristics of expression changes. In addition, compared with the global method, M 2 The DA method can also better reveal the local changes in facial expressions.

[0091] Furthermore, the recognition rates of the JAFFE and RML expression databases are lower than those of the Cohn-Kanade database. This may be because the samples collected from the former two databases have lower expression intensity and cannot clearly reflect the changing characteristics of facial expressions. Furthermore, the number of training samples in these two expression databases is far less than that of the Cohn-Kanade database, so they may not be able to extract enough effective identification information. Table 2 Recognition results of different manifold algorithms on Cohn-Kanade, JAFFE and RML libraries (%)

[0092] The experimental results show that, in the tests on Cohn-Kanade, Jaffe and RML expression databases, the proposed method can maintain good recognition performance compared with the traditional subspace analysis algorithm and manifold learning algorithm. 2 Core technological innovations such as the DA optimization algorithm have solved the shortcomings of traditional expression recognition methods in terms of feature discrimination, anti-interference ability, and computational efficiency, and achieved high-precision, high-robustness, and high-efficiency expression recognition effects. The technology has significant advantages and possesses practical application value and market competitiveness.

[0093] Finally, it should be noted that the parts of the present invention that are not described in detail are all prior art. Those skilled in the art will understand that the above description is only a preferred embodiment of the invention and is not intended to limit the invention. Although the invention has been described in detail with reference to the above examples, those skilled in the art can still modify the technical solutions described in the above examples or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, etc. made within the spirit and principles of the invention should be included in the scope of protection of the invention.

Claims

1. A facial expression recognition method based on multi-manifold learning, characterized in that: The following steps are involved: Step S1: using an active shape model to detect salient areas in a facial image, and classifying the detected salient areas to construct an expression database; Step S2, using fractional Fourier transform to perform frequency domain conversion preprocessing on the salient area images in the expression database, so that the images are jointly expressed in the time-frequency plane, and a preprocessed training set is obtained; Step S3, extracting the discriminant matrix corresponding to each expression category by a multi-manifold discriminant analysis algorithm, wherein the multi-manifold discriminant analysis algorithm optimizes the projection matrix with the optimization objectives of maximizing the inter-manifold divergence and minimizing the intra-manifold divergence; Step S4: Project the salient area of ​​the test sample processed by steps S1 and S2 into the discriminant space of each manifold, calculate the distance between the test sample and each manifold, and select the expression category corresponding to the manifold with the smallest distance as the recognition result of the test sample.

2. A facial expression recognition method based on multi-manifold learning according to claim 1, characterized in that, In step S1, the salient areas include the left eye, the right eye, the left cheek, the right cheek and the mouth.

3. A facial expression recognition method based on multi-manifold learning according to claim 2, characterized in that, In step S1, the specific steps of detecting the facial image using the active shape model include: S11. Define the manifold set of the training set The manifold set of the kth type of expression is The salient block of the i-th sample of the k-th class is Each sample The size of the salient block is a×b; where c is the number of expression categories in the expression library, n is the total number of samples, k = 1, 2, ..., c, i = 1, 2, ..., n k , n k is the number of samples of the k-th type of expression, is the i-th sample of the k-th expression, d is the original dimension of the salient block, t is the number of salient blocks, l k =t·n k and S12. The active shape model locates facial key points, extracts five salient areas: the left eye, the right eye, the left cheek, the right cheek, and the mouth, and adjusts the size of the salient blocks in each area to a uniform size.

4. A facial expression recognition method based on multi-manifold learning according to claim 1, characterized in that, The specific steps of step S2 include: first performing a one-dimensional discrete fractional-order Fourier transform, and then performing a two-dimensional fractional-order Fourier transform, with the transform order parameters P1 and P2 both set to 0.

5.

5. A facial expression recognition method based on multi-manifold learning according to claim 4, characterized in that, The specific implementation process of step S2 is: Step S21: perform one-dimensional discrete fractional Fourier transform according to the following formula: Where j is the imaginary unit, that is α represents the angle of rotation of the fractional-order transform in the time-frequency domain plane, and satisfies the relationship p is the transform order of FrFT, and p≠2n, n is an integer; F p represents the p-order fractional Fourier transform operation; f p(u) represents the function of variable u obtained by processing the function f(u′) by p-order fractional Fourier transform; Step S22: perform a two-dimensional fractional Fourier transform using the following discrete form: Where x(p,q) is the input original image, p and q are the row index and column index of the original image respectively; M and N are the number of rows and columns of the original image respectively; is the output image obtained after two-dimensional fractional Fourier transform, where m and n are the row and column coordinates of the output image respectively; is the two-dimensional fractional Fourier kernel function.

6. A facial expression recognition method based on multi-manifold learning according to claim 3, characterized in that, Step S3 specifically includes: extracting multiple projection matrices from the training set of step S2 according to different expression categories, and projecting the salient blocks of each category of samples in the training set into the space; defining inter-manifold and intra-manifold neighbor sample sets, and using the Fisher linear discriminant method to successively calculate the projection matrix of each category of expression under the conditions of maximizing the distance between manifolds and minimizing the distance within the manifold, and calculating the dimension of the discrimination matrix of each category, based on the dimensionality requirement, screening out the most discriminative eigenvectors, and finally constructing a discrimination matrix that can effectively distinguish expressions.

7. A facial expression recognition method based on multi-manifold learning according to claim 6, characterized in that, In step S3, when extracting the discriminant matrix of each manifold through the multi-manifold discriminant analysis algorithm, define and They are The inter-manifold and intra-manifold neighbor sample sets, where j = 1, 2, ..., t, assume that a point for K b nearest neighbor, k b represents the number of neighbors of samples between manifolds, then in and Located on different manifolds; let a point for K w nearest neighbor, k w represents the number of neighbors of the sample in the manifold, then in and are on the same manifold; The optimization goal is to maximize in is the inter-manifold scatter matrix of the k-th type of expression, is the scatter matrix within the manifold of the kth type of expression, and the discriminant vectors of each type satisfy The orthogonality relationship, and v,ε=1,2,…,d k .

8. A facial expression recognition method based on multi-manifold learning according to claim 7, characterized in that, The inter-manifold scatter matrix The calculation formula is: Where n k represents the number of samples of the kth type of expression, t is the number of significant areas of the sample, k b is the number of neighbors of a sample between manifolds, is the jth significant block of the i-th sample of the k-th expression, for The nearest neighbor between manifolds, is the weight matrix between manifolds, M k is the correlation matrix with the k-th type of expression sample; Σ k The entity is l k ×l b matrix, and is a diagonal matrix, and its entities are Columns and rows, that is and The scatter matrix within the manifold The calculation formula is: Where k w is the number of neighbors of the sample in the manifold, for The same manifold neighbors of is the weight matrix within the manifold; Represents the weight matrix, which is used to describe the connection weight between the sample and its neighbors; D k is a diagonal matrix, and its entity is Column vector of .

9. A facial expression recognition method based on multi-manifold learning according to claim 8, characterized in that, Weight Matrix They are defined as: Where ||·|| is the Euclidean norm and κ is a constant.

10. A facial expression recognition method based on multi-manifold learning according to claim 3, characterized in that, The design process of the classifier in step S4 is as follows: Step S41: Let Φ be an expression test sample, and divide it into five significant regions after detecting facial key points using the active shape model. These five significant regions are used as a manifold M Φ =[φ1,…,φ t ]; Step S42: Make M Φ Projection to the discriminant space W k On the other hand, we get the projection matrix at the same time, M k In W k The projection matrix; Step S43: Calculate M Φ and M k The distance d on the low-dimensional manifold Φk (M Φ ,M k ): in, Where, For Y Φj K l Neighbors, is the nearest neighbor reconstruction error factor; Step S44: Set c * is the category label of the test sample Φ, through c * =argmin k=1,2,…,c d k (M Φ ,M k ) calculates the category information of the test sample.