Image set recognition method integrating collaborative expression of nearest neighbor distance and Riemannian manifold

By combining the method of collaboratively expressing the nearest neighbor distance and Riemann manifold, the instability problem caused by illumination and angle changes in image set recognition is solved. Through the combination of European space, Grassmann manifold and SPD manifold, the Martinez distance matrix is optimized and the recognition rate is improved.

CN115496957BActive Publication Date: 2025-07-08GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210938573.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-05
Publication Date
2025-07-08
Estimated Expiration
2042-08-05

AI Technical Summary

Technical Problem

When the existing image set recognition methods deal with factors such as lighting and angle changes, the mean distance measurement is unstable, resulting in a decrease in recognition rate and failing to effectively integrate information from different modeling methods, affecting the recognition efficiency.

Method used

By integrating the method of collaboratively expressing nearest neighbor distance and Riemann manifold, including scale modeling of European space, Grassmann manifold and SPD manifold, combining regularized affine package collaborative expression method and nuclear method, we learn the Mahayana distance matrix, optimize the inter-class divergence and intra-class divergence, and obtain the optimal similarity measure between image sets.

Benefits of technology

The stability and recognition rate of image set recognition are improved, the influence of light and angle changes are overcome, and the current technology is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496957B_ABST
    Figure CN115496957B_ABST
Patent Text Reader

Abstract

The present invention discloses an image set recognition method that fuses and collaboratively expresses the nearest neighbor distance and the Riemannian manifold, comprising the following steps: S1: Obtain a training image set; S2: Perform Euclidean space scale modeling on the training image set using the nearest neighbor distance; S3: Perform Grassman manifold scale modeling on the training image set; S4: Perform SPD manifold scale modeling on the training image set; S5: Map the scales in different spaces to a unified Hilbert space, and then obtain a fused distance metric formula by learning the Mahalanobis distance matrix; S6: Learn the Mahalanobis distance matrix that can maximize the between-class scatter and minimize the within-class scatter through the kernel method to obtain the optimal similarity metric between image sets. The technical solution provided by the present invention has a higher recognition rate and robustness compared with the existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition, and particularly relates to an image set recognition method that fuses collaborative representation of nearest neighbor distance and Riemannian manifold. Background Art

[0002] Image recognition is an important application direction in current artificial intelligence research, and image set recognition is an extension of image recognition. Image set recognition refers to a recognition model that generally uses image sets as units for matching during both the training process and the recognition process. In contrast, traditional single-image recognition methods are characterized in that generally only the feature information of one image is used for image matching during the recognition stage. Image set recognition methods generally model image sets composed of images belonging to the same category. Using image sets as a whole for image recognition can effectively obtain more discriminative information; moreover, by modeling a large number of images with the same label and then performing model matching, it is not necessary to separately process the training and matching of single images, which will greatly improve the image recognition efficiency. Compared with traditional single-image recognition methods, image set recognition has greater advantages when dealing with a large amount of data with the same label, especially in complex situations such as poor image quality and multiple angles. With the rapid development of current mobile communications, social networks, video websites, video surveillance, etc., there are more and more ways to obtain image information, and the data volume is growing exponentially. Image set recognition will be able to play its advantages.

[0003] In the study of methods for comparing the similarity between image sets, many different methods have emerged. For example: probability models, linear subspaces, covariance matrices, affine / convex hull models, Riemannian manifolds, sparse representations, scale learning, etc. However, these methods often only focus on a certain modeling method when describing image sets. For example, they directly use the mean distance, nearest neighbor distance between different image sets, construct probability models, subspace models, covariance models, etc. However, the information obtained by different image set modeling methods is not the same, and they can complement each other through information to obtain better recognition effects. A natural idea is to effectively fuse different image set modeling information to integrate and obtain better classification effects.

[0004] Currently, there are related methods for fusing different modeling methods. For example, the method of fusing Euclidean distance and Riemannian manifold scale learning (HERML). This method combines the Euclidean distance (the mean distance of the image set), the covariance matrix, and the symmetric positive definite (SPD) matrix formed by the probability distribution model, and forms two different SPD Riemannian manifolds respectively. Finally, the kernel function is used to map the features of different metric spaces to the reproducing kernel Hilbert space, and the scale learning method based on information theory is used for fusion. Another multi-Riemannian manifold scale learning (MMML) method uses a discriminant method that maximizes the between-class scatter matrix and the within-class scatter matrix to learn the Mahalanobis distance. Through the obtained distance matrix, the kernel mapping is also used to fuse the Grassmann manifold composed of linear subspaces and the SPD Riemannian manifold composed of covariance matrices. Different modeling methods construct different effective information of the image set. The Euclidean distance plays a very important role in measuring the similarity between different images. In the early pattern recognition research on image matching problems, the Euclidean distance has always been the most important similarity metric criterion, such as classic algorithms like principal component analysis, linear discriminant analysis, K-means algorithm, support vector machine (SVM), etc. This shows that the Euclidean distance also plays an important role in measuring the similarity of the image set. However, in different feature fusion methods, MMML abandons the Euclidean distance and only uses linear subspaces and covariance matrices. In the improved version of MMML, the SPD manifold formed by the Gaussian distribution is added for fusion, which further improves the recognition rate of the image set, but this method still does not consider the Euclidean distance. Although the HERML method considers the role of the Euclidean distance, this method simply uses the Euclidean distance between the mean vectors of the image sets. Measuring the similarity between image sets using the mean distance of the image sets will have relatively large errors because factors such as illumination and object angle changes bring relatively large fluctuations to the mean of the image set. Therefore, using a better Euclidean distance measurement method will be able to better improve the multi-scale fusion effect. Summary of the Invention

[0005] In view of the existing problems, the purpose of the present invention is to provide an image set recognition method that fuses collaborative representation nearest neighbor distance and Riemannian manifold, and combines a new model and algorithm to solve the above problems.

[0006] The present invention provides the following technical solutions:

[0007] An image set recognition method that fuses collaborative representation nearest neighbor distance and Riemannian manifold, the method includes the following steps: S1: Obtain the training image set S2: Model the Euclidean space scale of the training image set using the nearest neighbor distance; S3: Model the Grassmann manifold scale of the training image set; S4: Model the SPD manifold scale of the training image set; S5: Map the scales in different spaces to a unified Hilbert space, and then obtain the fused distance metric formula by learning the Mahalanobis distance matrix; S6: Learn the Mahalanobis distance matrix that can maximize the between-class scatter and minimize the within-class scatter through the kernel method to obtain the optimal similarity metric between image sets.

[0008] Preferably, in step S2, the regularized affine hull collaborative representation method and the kernel regularized affine hull collaborative representation method are used to model the nearest neighbor Euclidean distance between training image sets between.

[0009] Preferably, the regularized affine hull collaborative representation method specifically includes the following: Use dictionary learning to perform sparse representation on the training image set to obtain an image set D composed of a few images i , and all compressed image sets are: where N is the number of the training image sets, and the test image set is denoted as X te , and establish the optimal function:

[0010]

[0011]

[0012]

[0013] where α and β are two parameter vectors to be optimized. ρ1 and ρ2 are constraint parameters, and τ is a parameter that determines the convex hull form of the model, τ ≤ 1, n q and n D are respectively the number of images in the sparse representation set D of the test image set X te and all training image sets, l p is the norm number, and the l1 norm in the RH-ISCRC method is used for regularization; make and to avoid extreme cases where α = 0 and β = 0; use the Lagrange multiplier method to obtain the dual form of the optimal function:

[0014]

[0015] where λ1 and λ2 are two Lagrange multipliers; use the alternating method to achieve parameter optimization: first fix the parameter α and update the optimal value of β, then fix β and update the optimal value of α; where the optimal value is optimized by the LARS algorithm with l1-minimization regularization, β = [β1,..., βN T each sub-vector β in i corresponds to a coefficient set D i , and the optimal value is denoted as and Through the formula:

[0016]

[0017] the nearest neighbor distance between the test image set X te and the i-th training image set is obtained; the mapping effect is obtained using a linear kernel: where k l (·) is a linear kernel function.

[0018] The kernel regularized affine hull collaborative representation method specifically includes the following: obtaining the dual form of the objective function of the model by the Gaussian mapping function and adopting the l2 norm regularization term:

[0019]

[0020]

[0021] The optimization process uses the alternating method to solve the above objective function to obtain the optimal parameters and where β = [β1,…,β N T the optimal parameters corresponding to each image set in the training set; using the Gaussian kernel function to calculate the nearest neighbor distance between the test image set X te and the i-th training image set in the high-dimensional space

[0022] The inner product expression of the nearest neighbor distance in the mapped high-dimensional Hilbert space is:

[0023]

[0024] Preferably, step S3 specifically includes the following: obtaining the linear subspace of each training image set X i through principal component analysis, thereby forming a Grassmann manifold; the linear subspace of the training image set X i is expressed using one of its orthonormal bases , where m is the dimension of the formed Grassmann manifold and also the number of vectors in the orthonormal basis, and d is the dimension of each image in the image set; among them, for N training image sets X i of C categories, there are N orthonormal bases Y = {Y1, Y2,..., Y N ​​} The Grassmann manifold composed of; by defining the Grassmann manifold mapping φ gr : G → H, the vector group expression of the orthogonal basis set Y in the Hilbert space is obtained:

[0025] Φ(Y) = [φ(Y1), φ(Y2),..., φ(Y N )]

[0026] Define the manifold distance between two sample points Y1, Y2 on the Grassmann manifold as:

[0027]

[0028] The corresponding Projection kernel function is calculated through the manifold distance:

[0029] where ||·|| F is the Frobenius norm.

[0030] Preferably, step S4 specifically includes the following content: Through the formula:

[0031]

[0032] The covariance matrix Z of the training image set is obtained; By Z i ; By Z i = Z i + λI to add perturbation to Z i to make it a symmetric positive definite matrix, where tr(Z i ) is to find the trace of the matrix Z i ; N training image sets are composed of N covariance matrices Z = {Z1, Z2,..., Z N} to form an SPD manifold; Through the LED distance function:

[0033] d LED (Z1, Z2) = ||log(Z1) - log(Z2)|| F

[0034] Calculate the distance metric method on the SPD manifold, and derive its kernel function formula from the LED distance function:

[0035] k LED (Z1, Z2) = tr[log(Z1)·log(Z2)]

[0036] where the symmetric positive definite matrix Z has the eigenvalue decomposition formula: Z = UΣU T , where log(Z) = Ulog(Σ)U T .

[0037] Preferably, step S5 specifically includes the following content: Define the distance between two image sets as:

[0038] where u q is the weight of the q-th fusion metric space, Q is the number of different fusion degree spaces and Q = 3, is the vector obtained by mapping the i-th image set in the q-th metric space to the Hilbert space in the corresponding metric space, P is the Mahalanobis distance matrix to be learned, and is obtained from the training image set; Decompose P: P = WW T , and transform the learning of P into learning the decomposition matrix W, then the distance between two image sets is rewritten as:

[0039] Preferably, step S6 optimizes this formula to obtain the optimal W matrix W * , through the formula:

[0040] to obtain the optimal W *

[0041] where

[0042]

[0043]

[0044] where, R w and R b respectively represent the within-class scatter matrix and between-class scatter matrix of the training image set after mapping in different metric spaces; Through the linear relationship between W and the training image set after mapping, the h-th column vector in W is obtained as: Then the formula: is transformed into:

[0045] where,

[0046]

[0047]

[0048] where, is the inner product of the i-th sample in the q-th fusion space and all training samples in the mapped Hilbert space; The formula: Through eigenvalue decomposition (R' w ) -1 R' b , and retaining the eigenvectors corresponding to the d z largest eigenvalues to obtain the optimal solution: The beneficial technical effects of the present invention are as follows:

[0049] The technical solution provided by the present invention, through the combination of new models and algorithms, overcomes the instability of the simple mean distance caused by changes in illumination, angle, etc. Compared with the prior art, the recognition rates of the GassianSPD+SPD+MEAN and Grassmann+SPD+MEAN algorithms using the MEAN distance have been greatly improved. Description of the Drawings

[0050] Figure 1 is a schematic flowchart of the image set recognition method that fuses collaborative representation nearest neighbor distance and Riemannian manifold provided by the present invention;

[0051] Figure 2 is the overall structural framework diagram of the image set recognition method that fuses collaborative representation nearest neighbor distance and Riemannian manifold provided by the present invention;

[0052] Figure 3 is a comparison image of the recognition rates between the image set recognition method that fuses collaborative representation nearest neighbor distance and Riemannian manifold provided by the present invention and the prior art solutions; Detailed Embodiments

[0053] The following will describe the embodiments of the present invention in detail. The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0054] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase does not necessarily refer to the same embodiment at every position in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that, without conflict, the embodiments described herein can be combined with other embodiments.

[0055] Embodiment

[0056] As Figure 1 、 2 shown, in a preferred embodiment of the present invention, the image set recognition method that fuses collaborative representation nearest neighbor distance and Riemannian manifold includes the following steps:

[0057] S1: Obtain the training image set S2: Perform Euclidean space scale modeling on the training image set using the nearest neighbor distance; S3: Perform Grassman manifold scale modeling on the training image set; S4: Perform SPD manifold scale modeling on the training image set; S5: Map the scales in different spaces to a unified Hilbert space, and then obtain the fused distance metric formula by learning the Mahalanobis distance matrix; S6: Learn the Mahalanobis distance matrix that can maximize the between-class scatter and minimize the within-class scatter through the kernel method to obtain the optimal similarity metric between image sets.

[0058] Among them, in step S2, the regularized affine hull collaborative representation method and the kernel regularized affine hull collaborative representation method are used to model the nearest neighbor Euclidean distance between training image sets between.

[0059] The regularized affine hull collaborative representation method specifically includes the following content: Use dictionary learning to perform sparse representation on the training image set to obtain an image set D composed of a few images i , and all the compressed image sets are: where N is the number of the training image sets, and the test image set is denoted as X te , and establish the optimal function:

[0060]

[0061]

[0062]

[0063] where α and β are two parameter vectors to be optimized. ρ1 and ρ2 are constraint parameters, τ is a parameter that determines the convex hull form of this model, τ ≤ 1, n q and n D are respectively the number of images in the sparse representation set D of the test image set X te and all the training image sets, l p is the norm number, and the l1 norm in the RH-ISCRC method is used for regularization; make and to avoid the extreme cases of α = 0 and β = 0; use the Lagrange multiplier method to obtain the dual form of the optimal function:

[0064]

[0065] where λ1 and λ2 are two Lagrange multipliers; use the alternating method to achieve parameter optimization: First, fix the parameter α and update the optimal value of β, then fix β and update the optimal value of α; among them, the optimal value is optimized through the LARS algorithm with l1-minimization regularization, β = [β1,..., β NT each sub-vector β i in the corresponding coefficient set D i , and denote the optimal value as and Through the formula:

[0066]

[0067] obtain the nearest neighbor distance between the test image set X te and the i-th training image set; use the linear kernel to obtain the mapping effect: where k l (·) is the linear kernel function.

[0068] The kernel regularized affine hull collaborative representation method specifically includes the following: obtain the dual form of the objective function of the model by the Gaussian mapping function, and adopt the l2-norm regularization term:

[0069]

[0070]

[0071] The optimization process uses the alternating method to solve the above objective function to obtain the optimal parameters and where β = [β1,..., β N T the optimal parameters corresponding to each image set in the training set; use the Gaussian kernel function to calculate the nearest neighbor distance between the test image set X te and the i-th training image set in the high-dimensional space

[0072] The inner product expression of the nearest neighbor distance in the mapped high-dimensional Hilbert space is:

[0073]

[0074] Step S3 specifically includes the following: obtain the linear subspace of each training image set X i through principal component analysis, so as to form a Grassmann manifold; the linear subspace of the training image set X i is expressed using one of its orthonormal bases , where m is the dimension of the formed Grassmann manifold and also the number of vectors in the orthonormal basis, and d is the dimension of each image in the image set; among them, for N training image sets X i of C categories, there is a Grassmann manifold composed of N orthonormal bases Y = {Y1, Y2,..., Y N}; by defining the Grassmann manifold mapping φ​​gr :G→H, obtain the vector group expression of the orthogonal basis set Y in the Hilbert space:

[0075] Φ(Y) = [φ(Y1), φ(Y2),..., φ(Y N )]

[0076] Define the manifold distance between two sample points Y1 and Y2 on the Grassmann manifold as:

[0077]

[0078] Calculate the corresponding Projection kernel function through the said manifold distance:

[0079] where ||·|| F is the Frobenius norm.

[0080] Preferably, step S4 specifically includes the following content: Through the formula:

[0081]

[0082] Obtain the covariance matrix Z of the training image set i ; Through Z i = Z i + λI to add perturbation to Z i to make it a symmetric positive definite matrix, where tr(Z i ) is to find the trace of the matrix Z i ; N training image sets consist of N covariance matrices Z = {Z1, Z2,..., Z N} to form an SPD manifold; Through the LED distance function:

[0083] d LED (Z1, Z2) = ||log(Z1) - log(Z2)|| F

[0084] Calculate the distance metric method on the SPD manifold, and deduce its kernel function formula from the said LED distance function:

[0085] k LED (Z1, Z2) = tr[log(Z1)·log(Z2)]

[0086] where the symmetric positive definite matrix Z has the eigenvalue decomposition formula: Z = UΣU T , where log(Z) = Ulog(Σ)U T .

[0087] Step S5 specifically includes the following: Define the distance between two image sets as:

[0088] where u q is the weight of the q-th fusion metric space, Q is the number of different fusion degree spaces and Q = 3, is the vector of the i-th image set in the q-th metric space mapped into the Hilbert space in the corresponding metric space, P is the Mahalanobis distance matrix to be learned, learned from the training image set; Decompose P: P = WW T , and transform the learning of P into learning the decomposition matrix W, then the distance between two image sets is rewritten as:

[0089] In step S6, optimize this formula to obtain the optimal W matrix W * , through the formula:

[0090]

[0091] where

[0092]

[0093]

[0094] where, R w and R b represent the within-class scatter matrix and between-class scatter matrix of the training image set after mapping in different metric spaces respectively; Through the linear relationship between W and the mapped training image set, the h-th column vector in W is obtained as: Then the formula: is deformed into:

[0095] where,

[0096]

[0097]

[0098] where, is the inner product of the i-th sample in the q-th fusion space and all training samples in the mapped Hilbert space; The formula: Through eigenvalue decomposition (R' w ) -1 R' b , and retaining the eigenvectors corresponding to the d z largest eigenvalues to obtain the optimal solution:

[0099] Specifically, for a set of images to be tested According to different spatial mapping functions, map it The projected sample vectors can be obtained: Using the formula Can get where is the inner product of the test sample set in the q-th fusion space and all training samples in the mapped Hilbert space. For the test image set and a certain prototype image set X g The final similarity metric between them is expressed by the following formula:

[0100]

[0101] To illustrate the effectiveness of the algorithm of the present invention, a comparison of the recognition rates of the algorithm proposed by the present invention on the YTC video face database is given. The algorithms to be compared are all closely related to the present invention, including the algorithms based on the nearest neighbor distance of collaborative representation: RH-ISCRC (2014) and KRH-ISCRC (2018), the fusion algorithm: the scale learning method HERML (2014) that fuses Euclidean distance and Riemannian manifold, and the multi-manifold scale learning algorithm MMML (2018). The present invention makes the YTC database into three test sets, namely YTC50, YTC100, and YTC200, which respectively represent that the number of images in each image set in the test set is less than 50 frames, less than 100 frames, and less than 200 frames. As shown in Table 1, the present invention gives the average recognition rate and recognition variance of different algorithms under the above three test sets, and the average recognition rate and variance are obtained in 10 different cross-validations.

[0102] The two algorithms proposed by the present invention have certain improvements in the YTC database compared with RH-ISCRC (in 2014) and KRH-ISCRC (in 2018) that only use the collaborative expression nearest neighbor distance (belonging to the Euclidean distance space). This shows that in the Euclidean space, further fusing the Grassmann manifold and the SPD manifold composed of the linear subspace and the covariance matrix can capture the change information and statistical information of the target object, thereby further improving the recognition rate. Compared with HERML (in 2014) and MMML (in 2018) that also use the fusion method in the present invention, the algorithms proposed by the present invention still have better recognition results. Although the HERML method takes the Euclidean distance into account in the fusion scheme, its Euclidean distance only uses the simple sample mean distance for measurement, which is greatly affected by many factors such as illumination, pose, and background. And the MMML method that abandons the Euclidean distance generally has better recognition results than the HERML method, because of the benefits brought by using the Grassmann manifold composed of linear subspaces; therefore, the present invention also uses the Grassmann manifold as one of the fusion parts.

[0103] It can be seen that the fusion nearest neighbor point and Riemannian manifold scale learning methods HNPRML-RHCR proposed by the present invention, that is, the fusion of the Grassmann manifold, the SPD manifold composed of the covariance matrix, and the regularized collaborative expression nearest neighbor distance (Grassmann+COVSPD+RHCR), and HNPRML-KRHCR, that is, the fusion of the Grassmann manifold, the SPD manifold composed of the covariance matrix, and the kernel regularized collaborative expression nearest neighbor distance (Grassmann+COVSPD+KRHCR), have better recognition results compared with the MMML method that only uses the Grassmann manifold and the covariance SPD manifold. This reflects that the collaborative expression nearest neighbor distance used in the present invention makes up for the defect that the MMML method only uses statistical features, and the nearest neighbor Euclidean distance information between image sets is also effectively utilized. The comparison of the average recognition rate and variance of the related methods of the present invention applied to the YTC database is shown in Table 1, Figure 3 as follows.

[0104] Table 1 Comparison of the average recognition rate and variance of the related methods of the present invention (unit: %)

[0105] Method YTC(50) YTC(100) YTC(200) KRH-ISCRC(2018) 83.2±1.1 84.8±0.9 85.7±0.8 RH-ISCRC(2014) 86.0±1.0 86.5±0.9 86.1±0.6 HERML(2014) 82.5±1.3 82.1±1.0 82.9±1.4 MMML(2018) 85.5±1.2 85.4±1.1 86.2±0.9 HNPRML-RHCR (This invention) 87.9±1.3 88.3±0.7 88.4±0.8 HNPRML-KRHCR (This invention) 87.5±0.8 87.1±0.8 87.9±0.6

[0106] In the technical solution provided by the present invention, the method for fusing the nearest neighbor distance and the scale learning of the Riemannian manifold and its application in image set recognition, abbreviated as HNPRML, where the nearest neighbor distance adopts two nearest neighbor distance models based on collaborative representation: namely, the regularized affine hull collaborative representation model (RHCR) and the kernel regularized affine hull collaborative representation model (KRHCR), and the two methods are represented by the abbreviations HNPRML-RHCR and HNPRML-KRHCR in Table 1. In Figure 3 it is expressed as Grassmann+COVSPD+RHCR (equivalent to HNPRML+RHCR) and Grassmann+COVSPD+KRHCR (equivalent to HNPRML-KRHCR), so that it can better represent which fusion models are adopted in the present invention, which is conducive to comparing with the Grassmann+COVSPD+MEAN and GaussianSPD+COVSPD+MEAN methods using the mean distance.

[0107] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning or limited experiments should be within the protection scope determined by the claims.

Claims

1. An image set recognition method that fuses and co-expresses the nearest neighbor distance and the Riemannian manifold, characterized in that, It includes the following steps: S1: Obtain a training image set S2: Perform Euclidean space scale modeling on the training image set using the nearest neighbor distance; S3: Perform Grassmann manifold scale modeling on the training image set; S4: Perform SPD manifold scale modeling on the training image set; S5: Map the scales in different spaces to a unified Hilbert space, and then obtain the fused distance metric formula by learning the Mahalanobis distance matrix; S6: Learn the Mahalanobis distance matrix that can maximize the between-class scatter and minimize the within-class scatter through the kernel method to obtain the optimal similarity metric between image sets; The step S2 models the nearest neighbor Euclidean distance between the training image set by using the regularized affine hull collaborative representation method and the kernel regularized affine hull collaborative representation method ; The specific steps of the regularized affine hull collaborative representation method are as follows: Use dictionary learning to perform sparse representation on the training image set to obtain an image set D composed of a small number of images i , and all compressed image sets are: where N is the number of the training image sets, and the test image set is denoted as X te , and establish an optimal function: where α and β are two parameter vectors to be optimized. ρ1 and ρ2 are constraint parameters, τ is a parameter determining the convex hull form of the model, τ ≤ 1, n q and n D are the number of images in the sparse representation sets D of the test image set X te and all training image sets, respectively, l p is the norm number, and the l1 norm in the RH-ISCRC method is used for regularization; make and to avoid extreme cases where α = 0 and β = 0; using the Lagrange multiplier method, the dual form of the optimal function is obtained: where λ1 and λ2 are two Lagrange multipliers; the parameter optimization is implemented by the alternating method: first fix the parameter α and update the optimal value of β, then fix β and update the optimal value of α; the optimal value is optimized by the LARS algorithm with l1-minimization regularization, and β = [β1, …, β N T in which each sub-vector β i corresponds to the coefficient set D i , and denote the optimal values as and through the formula:​ Obtain the test image set X te The nearest neighbor distance to the i-th training image set; obtain the mapping effect using the linear kernel: where k l (·) is the linear kernel function.

2. The method for image set recognition by fusing and synergistically expressing the nearest neighbor distance and the Riemannian manifold according to claim 1, wherein The specific content of the kernel regularized affine hull collaborative representation method includes the following: Obtain the dual form of the objective function of the model from the Gaussian mapping function and adopt the l2 norm regularization term: The optimization process uses the alternating method to solve the above objective function to obtain the optimal parameters and where β = [β1,…,β N T the optimal parameters corresponding to each image set in the training set; use the Gaussian kernel function to calculate the nearest neighbor distance between the test image set X te and the i-th training image set in the high-dimensional space ​ 3. The image set recognition method that fuses and synergistically expresses the nearest neighbor distance and the Riemannian manifold according to claim 2, wherein The inner product expression of the nearest neighbor distance in the mapped high-dimensional Hilbert space is:

4. The image set recognition method according to claim 1, characterized in that: The specific content of step S3 includes the following: Each training image set X is obtained by principal component analysis i The linear subspace of , thus forming a Grassmann manifold; the training image set X i The linear subspace of where m is the dimension of the Grassmann manifold, which is also the number of vectors in the orthogonal basis, and d is the dimension of each image in the image set; where, for a training image set X of C categories, i , there are N standard orthogonal bases Y = {Y1,Y2,...,Y N }; By defining the Grassmann manifold mapping φ gr :G→H, we get the vector group expression of the orthogonal basis set Y in the Hilbert space: Φ(Y) = [φ(Y1), φ(Y2), …, φ(Y N )] Define the manifold distance between two sample points Y1 and Y2 on the Grassmann manifold as: Calculate the corresponding Projection kernel function through the manifold distance: where ||·|| F is the Frobenius norm.

5. The image set recognition method that fuses and collaboratively expresses the nearest neighbor distance and the Riemannian manifold according to claim 1, wherein The specific content of step S4 includes the following: Through the formula: Obtain the training image set The covariance matrix Z of i ; Through Z i = Z i + λI to add perturbations to Z i to make it a symmetric positive definite matrix; The N training image sets consist of N covariance matrices Z = {Z1, Z2, …, Z N} to form an SPD manifold; Through the LED distance function: d LED (Z1, Z2) = ||log(Z1) - log(Z2)|| F Calculate the distance metric method on the SPD manifold, and deduce its kernel function formula from the LED distance function: k LED (Z1, Z2) = tr[log(Z1)·log(Z2)] Among them, the symmetric positive definite matrix Z has an eigenvalue decomposition formula: Z = UΣU T , where log(Z) = Ulog(Σ)U T , tr(Z i ) is to find the trace of matrix Z i .

6. The method for image set recognition that fuses and synergistically expresses the nearest neighbor distance and the Riemannian manifold according to claim 1, wherein The specific content of step S5 includes the following: Define the distance between two image sets as: where u q is the weight of the q-th fusion metric space, Q is the number of different fusion degree spaces and Q = 3, φ i q is the vector in the Hilbert space to which the i-th image set of the q-th metric space is mapped in the corresponding metric space, P is the Mahalanobis distance matrix to be learned, which is learned from the training image set; Decompose P: P = WW T , and transform the learning of P into learning the decomposition matrix W. Then the distance between the two image sets is rewritten as:

7. The method for image set recognition that fuses and collaboratively expresses the nearest neighbor distance and the Riemannian manifold according to claim 6, wherein The specific steps of step S6 include the following: For the formula: Optimize to obtain the optimal W matrix W * , through the formula: Obtain the optimal W * where Among them, R w and R b respectively represent the within-class scatter matrix and between-class scatter matrix of the training image set after mapping in different metric spaces; through the linear relationship between W and the mapped training image set, the h-th column vector in W is obtained as: Then the formula: is transformed into: Among them, Among them, is the inner product of the i-th sample in the q-th fusion space and all training samples in the mapped Hilbert space; formula: Through eigenvalue decomposition (R' w ) -1 R' b , and retaining the eigenvectors corresponding to the d z largest eigenvalues to obtain the optimal solution:

Citation Information

Patent Citations

  • Image similarity determination method and device, equipment and storage medium

    CN109272044A

  • Image set classification system and method based on manifold deep learning and an extreme learning machine

    CN109615005A