Image recognition method, system and equipment based on twin automatic encoders and medium

The geometric structure information of biometric features is extracted by a method based on a twin automatic encoder and the feature alignment is performed using the pose change insensitive loss function, which solves the problem of low recognition accuracy of biometric features during pose change, achieving higher recognition accuracy and better generalization ability.

CN120088816APending Publication Date: 2025-06-03INSPUR QILU SOFTWARE IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200278.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When the posture of biometric features changes, the discriminant ability of geometric structural features decreases, affecting the performance of biometric recognition algorithms.

Method used

The image recognition method based on the twin automatic encoder is adopted to extract geometric structure information of biometric features and use the pose change insensitive loss function to perform feature alignment to extract geometric features that are insensitive to pose change.

Benefits of technology

It enhances the ability to distinguish features, reduces the impact of posture changes on geometric feature extraction, and improves the recognition accuracy of biometric features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088816A_ABST
    Figure CN120088816A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method, system and device based on twin automatic encoders and a medium, belongs to the technical field of machine learning, deep learning and image processing, and aims to solve the technical problem of how to extract geometric features insensitive to posture changes. In order to reduce the influence of the posture change of the biological characteristics on the recognition performance of a geometric structure algorithm and improve the recognition precision of the biological characteristics, the adopted technical scheme is as follows: extracting the geometric structure information of the biological characteristics: carrying out preprocessing operations of background interference removal and tilt correction on the biological characteristics through a maximum between-class variance method, a binarization algorithm and Radon transformation; respectively extracting overall geometric structure information and local geometric structure information through a snake algorithm and a region generation algorithm; extracting geometric features: extracting the geometric features through a twin automatic encoder and carrying out feature alignment to reduce the difference between similar features of a hidden layer; extracting geometric features insensitive to attitude change; and identifying image features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of machine learning, deep learning and image processing, and specifically to an image recognition method, system, device and medium based on a twin autoencoder. Background Art

[0002] Since biometrics have advantages such as measurability and uniqueness, they are used to automatically verify or identify human identities. Biometric identification technology refers to the identification of a person through biological characteristics or behavioral characteristics. Compared with the identification method based on passwords or ID cards, it has great advantages. Since biometrics have better reliability and stability, biometric identification technology is more secure, and there is no need to worry about password leakage or forgotten passwords. At the same time, it is more convenient to use, and there will be no situation where it is difficult to authenticate identity due to forgetting to bring an ID card. Common biometrics include face, fingerprint, iris, human posture and signature.

[0003] The shape of an object is an interpretable and intuitive visual feature, and is often used to describe the structural information and posture information of an object. Therefore, shape features can be used as important features for image recognition or retrieval, while geometric structures can reflect the characteristics of biometric structure and morphology. Geometric features mainly include global morphological features and local morphological features. Global morphological features mainly include overall width, length, size and other features, which reflect the overall macroscopic shape structure information. Local morphological features mainly include the length and relative direction between local features, which reflect the local microscopic shape structure information. However, in actual application scenarios, biometrics will rotate axially. Biometric structures will be affected by posture changes, such as the length, width and absolute direction between local features, which will cause nonlinear deformation of the geometric structure, making it difficult for geometric features to play their effectiveness, thereby affecting the performance of biometric recognition algorithms based on geometric structure information. Therefore, the representation and recognition of geometric structure features that are insensitive to posture changes are of great research significance.

[0004] As the posture of the biometric feature changes, its geometric contour will produce nonlinear deformation, which leads to a decrease in the ability to distinguish geometric structure features and makes it difficult to effectively identify the biometric feature.

[0005] Therefore, how to extract geometric features that are insensitive to posture changes, reduce the impact of biometric posture changes on the recognition performance of geometric structure algorithms, and improve the recognition accuracy of biometrics is a technical problem that needs to be solved urgently. Summary of the invention

[0006] The technical task of the present invention is to provide an image recognition method, system, device and medium based on a Siamese autoencoder to solve the problems of how to extract geometric features insensitive to pose changes, reduce the influence of biometric pose changes on the recognition performance of geometric structure algorithms, and improve the recognition accuracy of biometric features.

[0007] The technical task of the present invention is realized in the following way. An image recognition method based on a Siamese autoencoder is as follows:

[0008] Extract the geometric structure information of biometric features: Through the maximum inter-class variance method, binarization algorithm and Radon transform, preprocessing operations of removing background interference and skew correction are performed on biometric features, and then the overall geometric structure information and local geometric structure information are extracted through the snake algorithm and region generation algorithm respectively;

[0009] Extract geometric features: Extract geometric features through a Siamese autoencoder and perform feature alignment to reduce the difference between similar features in the hidden layer and increase the difference between dissimilar features in the hidden layer; Among them, the Siamese autoencoder includes two convolutional autoencoders, an encoding loss function, a decoding loss function and a reconstruction loss function. The encoding loss and decoding loss are used to improve the encoding ability of the encoder in the convolutional autoencoder and the decoding ability of the decoder in the convolutional autoencoder, and the reconstruction loss is used to ensure that the encoder and decoder can reconstruct the input data;

[0010] Extract geometric features insensitive to pose changes: Perform feature alignment through a pose change insensitive loss function to extract geometric features insensitive to pose changes, thereby enhancing the discrimination ability of features and reducing the influence of pose changes on geometric feature extraction;

[0011] Image feature recognition: Use the Euclidean distance to measure the distance between the features of the image to be recognized and each feature in the feature library, and use the K-nearest neighbor algorithm for image feature recognition.

[0012] Preferably, the removal of background interference is as follows:

[0013] By converting the image color space from the RGB color space with relatively high color component correlation to the YCbCr color space with uncorrelated color components;

[0014] Since the plantar region has good clustering in the Cr component, the Otsu method is used to obtain the binarization threshold of the plantar region in the Cr component;

[0015] Binarize the plantar image according to the binarization threshold of the plantar region, retain the part of the plantar image where the pixel value is greater than the binarization threshold of the plantar region, discard the part of the plantar image where the pixel value is less than or equal to the binarization threshold of the plantar region, reduce the noise in the image, and more accurately extract the feature information of the plantar region.

[0016] Preferably, the skew correction is as follows:

[0017] Use the method of Radon transform to estimate the direction of the plantar image, and perform skew correction on the plantar image according to the estimated angle. The Radon transform formula is as follows:

[0018] RA(ρ,α) = ∫∫ D I(x,y)δ(ρ - xcosα - ysinα)dxdy;

[0019] where ρ is the pole of the image point (x,y); (ρ,α) is the polar coordinate of the point (x,y); D is the plane of the entire image; I(x,y) is the original image; δ is the Dirac function;

[0020] After Radon transform, the peak set RA(ρ,α) of the polar coordinate plane (ρ,α) is the integral value of the (x,y) plane along the line of angle α;

[0021] If the pixel points in any direction in the image are dense and there are peak points in the polar coordinate space, by finding the peak points (ρ max ,α max ) in the polar coordinate space, the skew angle can be obtained, that is, α max is the skew angle;

[0022] Correct the plantar image through the skew angle α max

[0023] Preferably, the siamese autoencoder includes two network branches. Each network branch includes a convolutional autoencoder and a decoder. The network structures of the convolutional autoencoders adopted by the two network branches are the same, and both include a convolutional layer, a max-pooling layer, and a Relu activation function; and the network structures of the convolutional autoencoder and the decoder adopted by each network branch are symmetric to each other;

[0024] Preferably, the parameters of each convolutional autoencoder of the siamese autoencoder are not shared, and the parameters are updated independently, as follows:

[0025] Take the geometric contour maps x 1 and x 2 as the input of the network;

[0026] Pass the input data x through the convolutional autoencoder respectively​1 and x 2 are encoded to obtain feature z 1 and z 2 ;

[0027] The features z 1 and z 2 are decoded by the decoder respectively to obtain output data y 1 and y 2 ;

[0028] Among them, the formula of the encoding loss function is:

[0029] The formula of the decoding loss function is:

[0030] If x 1 and x 2 are of the same class of samples, the smaller the distance between the features z 1 and z 2 obtained by the two samples after their respective convolutional autoencoders, that is, the smaller the encoding loss l 1 and z 2 ; After the encoded z c and z 1 and z 2 are decoded by the decoder respectively, the smaller the distance between the data y 1 and y 2 output by the decoder, that is, the smaller the decoding loss l d ; Align the features of the same class to reduce the difference between the features of the same class;

[0031] If x 1 and x 2 are of different classes of samples, the farther the distance between the features z 1 and z 2 obtained by the two samples after their respective convolutional autoencoders, that is, the larger the encoding loss l 1 and z 2 ; After the encoded z c and z 1 and z 2 are decoded by the decoder respectively, the farther the distance between the data y 1 and y 2 output by the decoder, that is, the larger the decoding loss l d ; Align the features of different classes to increase the difference between the features of different classes;

[0032] The formula of the reconstruction loss function is:

[0033] The smaller the reconstruction loss is, it indicates that the convolutional autoencoder can extract the core feature information of the input function, reduce the redundant information and interference information of the input data, enable the convolutional autoencoder to map the input image from the original three-dimensional space to a new multi-dimensional feature space, and the geometric features extracted in the new multi-dimensional feature space have better discrimination ability and distinctiveness.

[0034] Preferably, the pose change insensitive loss function formula is as follows:

[0035]

[0036] When there is a pose change problem in the same-class samples, the pose change insensitive loss constrains the convolutional autoencoder to ignore the influence of pose changes, further tightens the features of the same-class samples, reduces the distance between the in-class features, thereby enhancing the ability to capture stable features, reducing the influence of pose changes on the encoding ability of the convolutional autoencoder and the decoding ability of the decoder, enabling the network to learn discriminative features, thus obtaining an effective feature representation of the input image, and at the same time preventing the encoder and decoder from learning unhelpful identity mappings;

[0037] When the input is the same-class samples, the smaller the distance between the output data of the decoder of one branch and the input data of the convolutional autoencoder of the other branch, the less the influence of pose changes on feature extraction, that is, the smaller the pose change insensitive loss, and then geometric features insensitive to pose changes are extracted, alleviating the problem of pose changes on geometric feature extraction;

[0038] Combine the encoding loss l c , decoding loss l d , reconstruction loss l r and pose change insensitive loss l r ′ to obtain the overall loss function of the siamese autoencoder, and the formula is as follows:

[0039] l = λ(α 1 l c + α 2 l d + α 3 l r ′)+(1 - λ){max(0, m - α 4 l c - α 5 l d ) + α 6 l r};

[0040] Among them, m, α 1 , α 2 , α 3 , α 4 , α5 and α 6 are hyperparameters, and the parameter λ is the label indicating whether the sample pair is similar; when the sample pair belongs to the same class, λ = 1; when the sample pair does not belong to the same class, λ = 0; and in each training round, only samples of the same class or samples of different classes are included.

[0041] Preferably, the K-nearest neighbor algorithm is used for image feature recognition as follows:

[0042] Input the image to be recognized into the siamese autoencoder network to obtain the deep features of the sample to be recognized;

[0043] Calculate the Euclidean distance between the features of the sample to be recognized and all the features in the feature library;

[0044] Find the feature in the feature library with the smallest feature distance from the sample to be recognized, and determine the class to which the plantar image to be recognized belongs.

[0045] An image recognition system based on a siamese autoencoder, which is used to implement the image recognition method based on a siamese autoencoder as described above. The system includes:

[0046] A geometric structure information extraction module, which is used to perform preprocessing operations of removing background interference and tilt correction on biometric features through the maximum inter-class variance method, binarization algorithm, and Radon transform, and then extract the overall geometric structure information and local geometric structure information through the snake algorithm and region generation algorithm respectively;

[0047] A geometric feature extraction module, which is used to extract geometric features through a siamese autoencoder and perform feature alignment to reduce the difference between the same-class features in the hidden layer and increase the difference between the different-class features in the hidden layer; among them, the siamese autoencoder includes two convolutional autoencoders, an encoding loss function, a decoding loss function, and a reconstruction loss function. The encoding loss and decoding loss are used to improve the encoding ability of the encoder in the convolutional autoencoder and the decoding ability of the decoder in the convolutional autoencoder, and the reconstruction loss is used to ensure that the encoder and decoder can reconstruct the input data;

[0048] A geometric feature extraction module insensitive to pose changes, which is used to perform feature alignment through a pose change insensitive loss function to extract geometric features insensitive to pose changes, thereby enhancing the discrimination ability of the features and reducing the impact of pose changes on geometric feature extraction;

[0049] An image feature recognition module, which is used to measure the distance between the features of the image to be recognized and each feature in the feature library using the Euclidean distance, and use the K-nearest neighbor algorithm for image feature recognition.

[0050] An electronic device, including: a memory and at least one processor;

[0051] Among them, a computer program is stored on the memory;

[0052] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the image recognition method based on the Siamese autoencoder as described above.

[0053] A computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the image recognition method based on the Siamese autoencoder as described above.

[0054] The image recognition method, system, device and medium based on the Siamese autoencoder of the present invention have the following advantages:

[0055] (1) The present invention adopts a Siamese autoencoder network structure, which can obtain paired data information and constrain the network to learn more discriminative features; in order to alleviate the problem that pose changes affect geometric feature extraction, a pose change-insensitive loss is used to constrain the network to extract geometric features insensitive to pose changes, thereby enhancing the discrimination ability of features, having good generalization ability and recognition accuracy, and can significantly improve the recognition accuracy of the algorithm;

[0056] (2) The present invention proposes a Siamese autoencoder and a pose change-insensitive loss function, aiming to extract geometric features insensitive to pose changes to alleviate the impact of pose changes on the extraction of geometric structure information, and the Siamese autoencoder network can obtain paired data information, which enhances the encoding ability, decoding ability and reconstruction ability of the autoencoder by optimizing the encoding loss, decoding loss and reconstruction loss respectively, and constrains the network to learn more discriminative geometric features; aiming at the problem that pose changes affect geometric feature extraction, a pose change-insensitive loss function is used to effectively align features, further constrain the network, reduce the impact of pose changes on geometric feature extraction, and thus improve the recognition accuracy of biometric features;

[0057] (3) The present invention gives a machine learning algorithm for extracting the geometric contour map of biometric features, which can effectively represent the overall geometric structure and local geometric structure of biometric features and retain the original geometric structure information;

[0058] (4) The present invention constructs a two-branch autoencoder to extract geometric features and can effectively align features;

[0059] (5) The present invention reduces the impact of pose changes on geometric feature extraction through a pose change-insensitive loss, thereby enhancing the discrimination ability of geometric features and is widely applied to various biometric recognition technologies

[0060] (6) The present invention improves the recognition performance of biometric features with pose variations by extracting the geometric structure information of biometric features and using a twin autoencoder network and a pose-variation-insensitive loss function. Description of the Drawings

[0061] The present invention will be further described below in conjunction with the drawings.

[0062] Attached Figure 1 is a schematic structural diagram of a twin autoencoder. Detailed Embodiment

[0063] The image recognition method, system, device and medium based on the twin autoencoder of the present invention will be described in detail below with reference to the drawings of the specification and specific embodiments.

[0064] Embodiment 1:

[0065] This embodiment provides an image recognition method based on a twin autoencoder. The method is as follows:

[0066] S1. Extract the geometric structure information of biometric features: Perform preprocessing operations of removing background interference and skew correction on biometric features through the maximum inter-class variance method, binarization algorithm, and Radon transform, and then extract the overall geometric structure information and local geometric structure information through the snake algorithm and region generation algorithm respectively;

[0067] S2. Extract geometric features: Extract geometric features through a twin autoencoder and perform feature alignment to reduce the difference between similar features in the hidden layer and increase the difference between dissimilar features in the hidden layer; among them, the twin autoencoder includes two convolutional autoencoders, an encoding loss function, a decoding loss function, and a reconstruction loss function. The encoding loss and decoding loss are used to improve the encoding ability of the encoder in the convolutional autoencoder and the decoding ability of the decoder in the convolutional autoencoder, and the reconstruction loss is used to ensure that the encoder and decoder can reconstruct the input data;

[0068] S3. Extract geometric features insensitive to pose variations: Perform feature alignment through a pose-variation-insensitive loss function to extract geometric features insensitive to pose variations, thereby enhancing the discrimination ability of the features and reducing the influence of pose variations on geometric feature extraction;

[0069] S4. Image feature recognition: Use the Euclidean distance to measure the distance between the features of the image to be recognized and each feature in the feature library, and use the K-nearest neighbor algorithm for image feature recognition.

[0070] The removal of background interference in step S1 of this embodiment is specifically as follows:

[0071] ① By converting the color space of the image, it is transformed from the RGB color space with relatively high color component correlation to the YCbCr color space where the color components are mutually uncorrelated;

[0072] ② Since the plantar region has good clustering in the Cr component, the Otsu method is used to obtain the binarization threshold of the plantar region in the Cr component;

[0073] ③ According to the binarization threshold of the plantar region, the plantar image is binarized. The part of the plantar image with pixel values greater than the binarization threshold of the plantar region is retained, and the part with pixel values less than or equal to the binarization threshold of the plantar region is discarded, reducing the noise in the image and more accurately extracting the feature information of the plantar region.

[0074] Among them, removing background interference is to be able to comprehensively reflect the global shape information and local shape information of the biometric characteristics.

[0075] Since the biometric characteristics rotate around the z-axis, the geometric structure information and skin texture information will not undergo non-linear deformation. According to the characteristics of the image, it can be known that the gray value distribution in the vertical direction is the widest in space. The tilt correction in step S1 of this embodiment is specifically as follows:

[0076] ① The method of Radon transform is used to estimate the direction of the plantar image, and the plantar image is tilt-corrected according to the estimated angle. The Radon transform formula is as follows:

[0077] RA(ρ,α)=∫∫ D I(x,y)δ(ρ - xcosα - ysinα)dxdy;

[0078] Among them, ρ is the pole of the image point (x,y); (ρ,α) is the polar coordinate of the point (x,y); D is the plane of the entire image; I(x,y) is the original image; δ is the Dirac function;

[0079] ② After Radon transform, the peak set RA(ρ,α) of the polar coordinate plane (ρ,α) is the integral value of the straight line along the angle α in the (x,y) plane;

[0080] If the pixel points in any direction in the image are dense and there are peak points in the polar coordinate space, by finding the peak point (ρ max ,α max ) in the polar coordinate space, the tilt angle can be obtained, that is, α max is the tilt angle;

[0081] ③ The plantar image is corrected by the tilt angle α max .

[0082] When the biometric feature undergoes a pose change, the geometric contour map will also have a pose change problem, resulting in a non-linear deformation of the geometric structure. This will weaken the discrimination ability of the geometric structure information and make it difficult to play its due role. When using a general convolutional neural network for feature extraction, it is more susceptible to pose changes, resulting in an increase in the difference between similar features and making it difficult to extract geometric features that are insensitive to pose changes. Therefore, a siamese autoencoder network and a pose-change-insensitive loss function are designed to extract geometric features that are insensitive to pose changes.

[0083] The siamese autoencoder in step S2 of this embodiment includes two network branches. Each network branch includes a convolutional autoencoder and a decoder. The network structures of the convolutional autoencoders adopted by the two network branches are the same and both include a convolutional layer, a max-pooling layer, and a Relu activation function; and the network structures of the convolutional autoencoder and the decoder adopted by each network branch are symmetric to each other, as shown in the appendix Figure 1 shown;

[0084] The parameters of each convolutional autoencoder of the siamese autoencoder in step S2 of this embodiment are not shared and are independently updated. Specifically, as follows:

[0085] ① Use the geometric contour maps x 1 and x 2 as the inputs of the network;

[0086] ② Encode the input data x 1 and x 2 through the convolutional autoencoders respectively to obtain the features z 1 and z 2 ;

[0087] ③ Decode the features z 1 and z 2 through the decoders respectively to obtain the output data y 1 and y 2 ;

[0088] In order to improve the ability of each encoder of the siamese autoencoder to encode the original data and the ability of each decoder to decode the hidden layer features into the original data, an encoding loss and a decoding loss are given.

[0089] Among them, the formula of the encoding loss function is:

[0090] The formula of the decoding loss function is:

[0091] If x 1 and x 2 are of the same class of samples, x 1 and x2 The features z obtained after the two samples pass through their respective convolutional autoencoders 1 and z 2 The smaller the distance between them, that is, the encoding loss l c The smaller; after encoding, z 1 and z 2 After the two features are decoded by the decoder respectively, the data y output by the decoder 1 and y 2 The smaller the distance between them, that is, the decoding loss l d The smaller; align the features of the same category to reduce the difference between the features of the same category;

[0092] If x 1 and x 2 are heterogeneous samples, x 1 and x 2 The features z obtained after the two samples pass through their respective convolutional autoencoders 1 and z 2 The farther the distance between them, that is, the encoding loss l c The larger; after encoding, z 1 and z 2 After the two features are decoded by the decoder respectively, the data y output by the decoder 1 and y 2 The farther the distance between them, that is, the decoding loss l d The larger; align the features of different categories to increase the difference between the features of different categories;

[0093] Since the autoencoder learns an identity mapping, that is, the distance between the input data of the encoder and the output data of the decoder should be small, indicating that the autoencoder has good encoding and decoding capabilities. Therefore, the reconstruction loss function formula is:

[0094] The smaller the reconstruction loss, the more the convolutional autoencoder can extract the core feature information of the input function, reduce the redundant information and interference information of the input data, so that the convolutional autoencoder can map the input image from the original three-dimensional space to a new multi-dimensional feature space, and the geometric features proposed in the new multi-dimensional feature space have better discrimination ability and distinguishability.

[0095] Due to the change in posture, its geometric contour map will also deform, and the extracted geometric features are vulnerable to the influence of posture changes. Moreover, learning an identity mapping is the essence of a common autoencoder, that is, the input of the encoder is the same data as the output of the encoder. However, this identity mapping has a significant drawback. When the training data and the test data do not conform to the same distribution, such as the problem of posture changes in the samples, it will reduce the performance of the network model on the test dataset. Therefore, it is difficult for a common autoencoder to effectively solve the problem of posture changes.

[0096] The formula for the posture change insensitive loss function in step S3 of this embodiment is as follows:

[0097]

[0098] When there are posture change problems in the same-class samples, the posture change insensitive loss constrains the convolutional autoencoder to ignore the influence of posture changes, further tightens the features of the same-class samples, reduces the distance between the intra-class features, thereby enhancing the ability to capture stable features, reducing the influence of posture changes on the encoding ability of the convolutional autoencoder and the decoding ability of the decoder, enabling the network to learn discriminative features, thus obtaining an effective feature representation of the input image, and at the same time preventing the encoder and decoder from learning unhelpful identity mappings;

[0099] When the input is the same-class sample, the smaller the distance between the output data of the decoder of one branch and the input data of the convolutional autoencoder of the other branch, the less the influence of posture changes on feature extraction, that is, the smaller the posture change insensitive loss, and then geometric features insensitive to posture changes are extracted, alleviating the problem of posture changes on geometric feature extraction;

[0100] Combining the encoding loss l c , the decoding loss l d , the reconstruction loss l r and the posture change insensitive loss l r ′, the overall loss function of the siamese autoencoder is obtained, and the formula is as follows:

[0101] l = λ(α 1 l c + α 2 l d + α 3 l r ′)+(1 - λ){max(0, m - α 4 l c - α 5 l d ) + α 6 l r};

[0102] Among them, m, α1 , α 2 , α 3 , α 4 , α 5 and α 6 are hyperparameters, and the parameter λ is the label indicating whether the sample pair is similar; when the sample pair belongs to the same class, λ = 1; when the sample pair does not belong to the same class, λ = 0; and in each training round, only samples of the same class or different classes are included.

[0103] The specific process of using the K-nearest neighbor algorithm for image feature recognition in step S4 of this embodiment is as follows:

[0104] S401. Input the image to be recognized into the siamese autoencoder network to obtain the deep features of the sample to be recognized.

[0105] S402. Calculate the Euclidean distance between the features of the sample to be recognized and all the features in the feature library.

[0106] S403. Find the feature in the feature library with the smallest feature distance from the sample to be recognized, and determine the class to which the plantar image to be recognized belongs.

[0107] Embodiment 2:

[0108] This embodiment provides an image recognition system based on a siamese autoencoder. This system is used to implement the image recognition method based on a siamese autoencoder as in Embodiment 1. This system includes:

[0109] A geometric structure information extraction module, which is used to perform preprocessing operations of removing background interference and skew correction on biometric features through the maximum inter-class variance method, binarization algorithm, and Radon transform, and then extract the overall geometric structure information and local geometric structure information through the snake algorithm and region generation algorithm respectively;

[0110] A geometric feature extraction module, which is used to extract geometric features through a siamese autoencoder and perform feature alignment to reduce the difference between the same-class features in the hidden layer and increase the difference between the different-class features in the hidden layer; among them, the siamese autoencoder includes two convolutional autoencoders, an encoding loss function, a decoding loss function, and a reconstruction loss function. The encoding loss and decoding loss are used to improve the encoding ability of the encoder in the convolutional autoencoder and the decoding ability of the decoder in the convolutional autoencoder, and the reconstruction loss is used to ensure that the encoder and decoder can reconstruct the input data;

[0111] A geometric feature extraction module insensitive to pose changes, which is used to perform feature alignment through a pose change insensitive loss function to extract geometric features insensitive to pose changes, thereby enhancing the discrimination ability of the features and reducing the influence of pose changes on geometric feature extraction;

[0112] An image feature recognition module, which is used to measure the distance between the features of the image to be recognized and each feature in the feature library using the Euclidean distance, and perform image feature recognition using the K-nearest neighbor algorithm.

[0113] Embodiment 3:

[0114] This embodiment also provides an electronic device, including: a memory and a processor;

[0115] Wherein, the memory stores computer-executable instructions;

[0116] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the image recognition method based on the siamese autoencoder in any embodiment of the present invention.

[0117] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory may further include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, at least one magnetic disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0119] Embodiment 4:

[0120] This embodiment also provides a computer-readable storage medium, in which multiple instructions are stored. The instructions are loaded by the processor to make the processor execute the image recognition method based on the siamese autoencoder in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. Software program codes for implementing the functions of any one of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device is made to read and execute the program codes stored in the storage medium.

[0121] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0122] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0123] Furthermore, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0124] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part or all of the actual operations, thereby implementing the functions of any one of the above embodiments.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image recognition method based on twin autoencoders, characterized in that: The method is as follows: Extracting geometric structure information of biometric features: Preprocessing the biometric features to remove background interference and tilt correction is performed using the maximum inter-class variance method, binarization algorithm, and Radon transform, and then extracting overall geometric structure information and local geometric structure information using the snake algorithm and region generation algorithm respectively; Extract geometric features: Extract geometric features and align features through the twin autoencoder to reduce the differences between similar features in the hidden layer and increase the differences between heterogeneous features in the hidden layer. The twin autoencoder includes two convolutional autoencoders, an encoding loss function, a decoding loss function, and a reconstruction loss function. The encoding loss and the decoding loss are used to improve the encoding ability of the encoder in the convolutional autoencoder and the decoding ability of the decoder in the convolutional autoencoder. The reconstruction loss is used to ensure that the encoder and the decoder can reconstruct the input data. Extracting geometric features that are insensitive to posture changes: Using a posture-insensitive loss function to align features, extract geometric features that are insensitive to posture changes, and thus enhance the feature recognition capability; Image feature recognition: Use Euclidean distance to measure the distance between the features of the image to be identified and each feature in the feature library, and use the K nearest neighbor algorithm for image feature recognition.

2. The image recognition method based on twin autoencoder according to claim 1, characterized in that: To remove background interference, please do the following: By converting the image into a color space, the image is converted from the RGB color space to the YCbCr color space where the color components are unrelated. The maximum inter-class variance method was used in the Cr component to obtain the binary threshold of the plantar area; The plantar image is binarized according to the binary threshold of the plantar region, the part of the plantar image with a pixel value greater than the binary threshold of the plantar region is retained, and the part of the plantar image with a pixel value less than or equal to the binary threshold of the plantar region is discarded.

3. The image recognition method based on twin autoencoder according to claim 1, characterized in that: Tilt correction is as follows: The Radon transform method is used to estimate the direction of the plantar image, and the plantar image is tilted according to the estimated angle. The Radon transform formula is as follows: RA(ρ,α)=∫∫ D I(x,y)δ(ρ-x cosα-y sinα)dxdy; Where ρ is the pole of the image point (x, y); (ρ, α) are the polar coordinates of the point (x, y); D is the plane of the entire image; I(x, y) is the original image; δ is the Dirac function; After Radon transformation, the peak value set RA(ρ,α) of the polar coordinate plane (ρ,α) is the integral value of the (x,y) plane along the straight line of angle α; If the pixels in any direction of the image are dense, there is a peak point in the polar coordinate space. By finding the peak point in the polar coordinate space (ρ max ,α max ), we can get the tilt angle, i.e. α max is the tilt angle; By tilting angle α max Correction of plantar images.

4. The image recognition method based on twin autoencoder according to claim 1, characterized in that: The twin autoencoder includes two network branches, each of which includes a convolutional autoencoder and a decoder. The convolutional autoencoders used in the two network branches have the same network structure, including convolutional layers, maximum pooling layers, and Relu activation functions; and the network structure of the convolutional autoencoder used in each network branch is symmetrical with the decoder used in the corresponding network branch.

5. The image recognition method based on twin autoencoder according to claim 1, characterized in that: The parameters of each convolutional autoencoder of the twin autoencoder are not shared, and the parameters are updated independently, as follows: Take the geometric contour images x1 and x2 as the input of the network; The input data x1 and x2 are encoded respectively through the convolutional autoencoder to obtain features z1 and z2; The features z1 and z2 are decoded by the decoder to obtain output data y1 and y2; Among them, the encoding loss function formula is: The decoding loss function formula is: If x1 and x2 are samples of the same type, the smaller the distance between the features z1 and z2 obtained after the two samples x1 and x2 pass through their respective convolutional autoencoders, the smaller the encoding loss l c The smaller the encoding loss is, the smaller the distance between the decoder output data y1 and y2 is. d The smaller it is, the more features of the same category are aligned to reduce the differences between features of the same category; If x1 and x2 are heterogeneous samples, the distance between the features z1 and z2 obtained by the two samples x1 and x2 after passing through their respective convolutional autoencoders is greater, that is, the encoding loss l c The larger the encoding z1 and z2 are, the farther the distance between the decoder output data y1 and y2 is, that is, the decoding loss l d The larger the value, the more features of different categories are aligned, increasing the differences between features of different categories. The reconstruction loss function formula is: The smaller the reconstruction loss, the more the convolutional autoencoder can extract the core feature information of the input function and reduce the redundant information and interference information of the input data, so that the convolutional autoencoder can map the input image from the original three-dimensional space to a new multi-dimensional feature space. The geometric features proposed in the new multi-dimensional feature space have better recognition and differentiation.

6. The image recognition method based on twin autoencoder according to claim 1, characterized in that: The loss function formula for posture change insensitivity is as follows: When there is a posture change problem in samples of the same type, the posture change insensitive loss constrained convolutional autoencoder ignores the impact of posture changes and further tightens the features of samples of the same type, so that the distance between features within the class is reduced, thereby improving the ability to capture stable features; When the input is a sample of the same type, the smaller the distance between the decoder output data of one branch and the convolutional autoencoder input data of another branch, the smaller the impact of posture change on feature extraction, that is, the smaller the posture change insensitive loss, and then the geometric features insensitive to posture change are extracted, alleviating the problem of posture change affecting geometric feature extraction; The encoding loss l c , decoding loss l d , reconstruction loss l r And the posture change insensitive loss l r ′ to obtain the overall loss function of the twin autoencoder, the formula is as follows: l=λ(α1l c +α2l d +α3l′ r )+(1-λ){max(0,m-α4l c -a5l d )+α6l r }; Among them, m, α1, α2, α3, α4, α5 and α6 are hyperparameters, and the parameter λ is a label for whether the sample pair is similar; if the sample pair belongs to the same class, λ=1; if the sample pair does not belong to the same class, λ=0; and in each training round, only samples of the same class or samples of different classes are included.

7. The image recognition method based on twin autoencoder according to any one of claims 1 to 6, characterized in that: The K nearest neighbor algorithm is used for image feature recognition as follows: Input the image to be identified into the twin autoencoder network to obtain the deep features of the sample to be identified; Calculate the Euclidean distance between the sample feature to be identified and all features in the feature library; Find the feature library feature with the smallest feature distance of the sample to be identified, and determine the category to which the plantar image to be identified belongs.

8. An image recognition system based on a twin autoencoder, characterized in that: The system is used to implement the image recognition method based on the twin autoencoder according to any one of claims 1 to 7, and the system comprises: The geometric structure information extraction module is used to perform preprocessing operations on the biometric features to remove background interference and tilt correction through the maximum inter-class variance method, binarization algorithm and Radon transform, and then extract the overall geometric structure information and local geometric structure information through the snake algorithm and region generation algorithm respectively; A geometric feature extraction module is used to extract geometric features and perform feature alignment through a twin autoencoder, reduce the differences between similar features in the hidden layer, and increase the differences between heterogeneous features in the hidden layer; wherein the twin autoencoder includes two convolutional autoencoders, an encoding loss function, a decoding loss function, and a reconstruction loss function, wherein the encoding loss and the decoding loss are used to improve the encoding capability of the encoder in the convolutional autoencoder and the decoding capability of the decoder in the convolutional autoencoder, and the reconstruction loss is used to ensure that the encoder and the decoder can reconstruct the input data; The pose-insensitive geometric feature extraction module is used to align features through a pose-insensitive loss function, extract pose-insensitive geometric features, and thus enhance the feature discrimination capability; The image feature recognition module is used to use the Euclidean distance to measure the distance between the feature of the image to be recognized and each feature in the feature library, and adopt the K nearest neighbor algorithm to perform image feature recognition.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the image recognition method based on the twin autoencoder as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the image recognition method based on the twin autoencoder as described in any one of claims 1 to 7.