Footprint image retrieval method based on multi-feature dense connection attention enhancement network
By using a multi-feature densely connected attention-enhanced network, the problem of imperfect feature extraction in traditional footprint recognition methods is solved, achieving efficient footprint image retrieval and improving accuracy and speed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2026-04-07
AI Technical Summary
Traditional footprint recognition methods lack standardized feature description and extraction methods, resulting in low accuracy and slow speed in footprint retrieval.
A multi-feature dense connection attention enhancement network is adopted, including an initial feature extraction module, a multi-feature dense connection network module, a multi-scale feature fusion module, an attention enhancement module, and a feature output module. Combining image processing technology and deep learning networks, the network model is trained by gradient descent to extract and fuse multi-layer feature information of footprint images.
It improves the accuracy and speed of footprint image retrieval, saves human and material resources, enhances the correlation of footprint features, and improves the fit of the loss function and the training effect of the retrieval model.
Smart Images

Figure CN115687679B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and in particular to a method for retrieving footprint images based on a multi-feature densely connected attention-enhanced network. Background Technology
[0002] Footprint data is an important form of human information, containing crucial physiological and individual behavioral information. Footprints reflect the traces left by the feet in contact with the ground or other bearing surfaces during movement. Currently, biometric identification is widely used in recognition and retrieval fields, including DNA, iris, and facial recognition. However, to cope with complex recognition and retrieval environments, research into more human feature recognition technologies is becoming increasingly important. Footprint information is difficult to disguise and erase, giving its research irreplaceable advantages in the field of biometric identification. Therefore, research on footprint characteristics is of paramount importance.
[0003] Traditional footprint recognition methods focus on geometric information such as points, lines, and curvature, resulting in different and highly varied extracted human footprint features. A standardized feature description and extraction method has not been established. This incomplete feature extraction leads to low accuracy and slow speed in footprint retrieval. Summary of the Invention
[0004] This invention aims to address the shortcomings of existing technologies by proposing a footprint image retrieval method based on a multi-feature densely connected attention-enhanced network. This method seeks to delve deeper into the detailed features of footprint image information, thereby improving the accuracy and speed of footprint image retrieval.
[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0006] The footprint image retrieval method of the present invention using a multi-feature densely connected attention enhancement network is characterized by the following steps:
[0007] Step 1: Acquire a set of footprint images of T test subjects using an optical footprint acquisition instrument, and perform filtering, noise reduction, size scaling, horizontal flipping, and random affine transformation to obtain the processed sample set denoted as S = {S1, S2, ..., S...}. m S M}, where S t Let m represent the processed footprint sample set of the m-th collected object, where M represents the total number of collected objects, m ≤ M, and This represents the k-th footprint image after processing of the m-th acquired object; K represents the total number of footprint images.
[0008] Step 2: Establish a footprint image retrieval network based on a multi-feature dense connection attention enhancement network, including: an initial feature extraction module, a multi-feature dense connection network module, a multi-scale feature fusion module, an attention enhancement module, and a feature output module;
[0009] Step 2.1: Establish an initial feature extraction module, which includes: a convolutional layer, a batch normalization layer, an activation layer, and a pooling layer. The Gaussian random initialization method is used to initialize the weights of the convolutional layer in the initial feature extraction module.
[0010] The k-th footprint image The initial footprint features of the k-th footprint image are obtained after being processed by the initial feature extraction module and input into the footprint image retrieval network.
[0011] Step 2.2: Construct a multi-feature dense connection network module, including: R dense connection layers and R transformation layers, wherein any r-th dense connection layer sequentially contains n convolutional layers with kernels of X1×X1 and n convolutional layers with kernels of X2×X2; any r-th transformation layer module sequentially contains one convolutional layer with kernels of X3×X3 and one average pooling layer with kernels of X4×X4;
[0012] When r = 1, the initial footprint features of the k-th footprint image After processing by the r-th dense connection layer, the dense connection features of the r-th layer are obtained. After processing by the r-th transformation layer, the dense connection features of the r-th layer are obtained.
[0013] When r = 2, 3, ..., R, the (r-1)th dense connection feature After processing by the r-th dense layer and the r-th transformation layer, the r-th dense connection feature is obtained. Thus, the Rth dense connection feature is output by the Rth transformation layer.
[0014] Step 2.3: Construct a multi-scale feature fusion module, which includes: m convolutional layers with kernels of X5×X5 and m convolutional layers with kernels of X6×X6;
[0015] The R-th dense connection feature The input to the multi-scale feature fusion module is processed sequentially through m convolutional layers with kernels of X5×X5 and m convolutional layers with kernels of X6×X6 to obtain high-level features. Then use interpolation for high-rise buildings After upsampling, the upsampled features are obtained. Finally, the upsampled features are... Dense connection features of layer R-1 After fusion, the multi-scale fusion features of the k-th image are obtained.
[0016] Step 2.4: Construct the attention enhancement module, including: a global pooling layer, two fully connected layers, and a regression function; densely connect the features of the (R-1)th layer. Dense connection features of layer R and multi-scale fusion features The data are input into the attention enhancement module and processed sequentially through a global pooling layer and two fully connected layers. The results are then weighted using a regression function to obtain the k-th footprint image of the m-th test subject. The R-1 layer dense connection attention feature vector The k-th footprint image of the m-th test subject The Rth layer dense connection attention feature direction The k-th footprint image of the m-th test subject Multi-scale fusion features
[0017] Step 2.5: Establish the feature output module, including: average pooling layer and fully connected layer;
[0018] The feature is the Rth dense connection feature. The output features are obtained by sequentially passing through an average pooling layer and a fully connected layer. As the final representation, it is retrieved as a characterization vector during testing;
[0019] Step 3: Construct the loss function:
[0020] Step 3.1: Establish the loss function L1 using equation (1):
[0021]
[0022] In equation (1), This represents the k-th footprint image of the 0-th test subject. The R-1th layer densely connected attention feature vector, This represents the q-th footprint image of the m-th test subject. The R-1th layer densely connected attention feature vector; o≤M,q≤T;
[0023] Step 3.2: Use equation (2) to establish the loss function L2:
[0024]
[0025] In equation (2), This represents the R-th layer dense connection attention feature vector of the k-th image of the o-th test subject. This represents the R-th layer densely connected attention feature vector of the q-th image of the m-th test subject;
[0026] Step 3.3: Establish the loss function L3 using equation (3):
[0027]
[0028] In equation (3), This represents the multi-scale fusion attention feature vector of the k-th image of the o-th test subject. This represents the pyramid fusion attention feature vector of the q-th image of the m-th test subject;
[0029] Step 3.3: Use equation (4) to establish the total loss function L:
[0030] L=L1+L2+L3 (4)
[0031] Step 4: Based on the processed sample set S, the footprint image retrieval network is trained using the gradient descent method, and the total loss function L is calculated to update the network parameters until the total loss function L converges, thereby obtaining the trained footprint image retrieval model, which is used to retrieve and match footprint images to be retrieved.
[0032] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program that supports the processor in executing the footprint image retrieval method, and the processor is configured to execute the program stored in the memory.
[0033] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, performs the steps of the footprint image retrieval method.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] 1. This invention combines image processing technology, deep learning networks, and footprint image retrieval to design a complete footprint image retrieval framework. It analyzes the acquired footprint optical images, extracts and fuses multi-layered feature information from the images, and constructs a main network based on the concept of densely connected networks. Simultaneously, it applies an attention mechanism to focus on locally important information, collectively completing the feature extraction network. Compared with traditional methods, by fusing multi-feature footprint image information and focusing on more detailed footprint feature differences, it improves the accuracy and speed of footprint retrieval, while saving significant human and material resources.
[0036] 2. The dataset acquisition and data preprocessing of this invention include image denoising, image cropping, removal of the acquisition instrument ruler, label processing and other preprocessing methods, which are beneficial to the training of the network model and thus obtain better detection results.
[0037] 3. The multi-feature dense connection network module designed in this invention makes full use of multi-layer image features through dense connection networks, increases the correlation of footprint features, better mines the same feature information of the same person's footprints, and improves the accuracy of footprint retrieval.
[0038] 4. This invention incorporates a multi-scale feature fusion module and selects a loss function for measurement calculation based on the characteristics of features at different layers. By calculating the loss function for features at different depths, the fit of the loss function to the footprint image is improved, thus better training the retrieval model.
[0039] 5. The attention module of this invention helps the network training focus on key information, paying more attention to the key feature regions of footprint information, which can effectively assist the main network in supervised learning. Attached Figure Description
[0040] Figure 1 This is an overall flowchart of footprint image retrieval in this invention;
[0041] Figure 2 This is a diagram of the multi-feature dense connection attention enhancement network structure in this invention.
[0042] Figure 3 This is a structural diagram of the multi-feature dense connection module in this invention. Detailed Implementation
[0043] In this embodiment, a footprint image retrieval method based on a multi-feature densely connected attention enhancement network mainly utilizes the densely connected attention enhancement network to extract feature information from footprint images. By training a neural network, deep feature information of the footprint images is extracted and retrieved, which greatly improves the speed and accuracy of footprint image retrieval.
[0044] The dataset used in this invention was collected and processed in the laboratory, consisting of 6000 images across 300 classes. Each image name includes a category label with a person's ID. The training set comprises 3600 images across 180 classes, and the test set comprises 2400 images across 120 classes. The base dataset contains 700 images across 70 classes, and the query set also contains 700 images across 70 classes. Figure 1 As shown, the entire process can be divided into the following steps:
[0045] Step 1: Perform filtering, noise reduction, scaling, horizontal flipping, and random affine transformation on the collected data;
[0046] Step 2: Establish a footprint image retrieval network based on a multi-feature dense connection attention enhancement network, including: an initial feature extraction module, a multi-feature dense connection network module, a multi-scale feature fusion module, an attention enhancement module, and a feature output module;
[0047] The footprint dataset is fed into a footprint image retrieval network based on a multi-feature densely connected attention-enhanced network for training. After preprocessing, the footprint image dataset sequentially passes through an initial feature extraction module, a multi-feature densely connected network module, a multi-scale feature fusion module, an attention enhancement module, and a feature output module. The initial feature extraction module obtains initial feature information from the footprint images. This initial feature information is then processed by the multi-feature densely connected network. The obtained feature information is then processed by the multi-scale feature fusion module, the attention enhancement module, and the feature output module, and finally, a loss function is calculated to train the optimal network model, thereby improving the accuracy and speed of footprint image retrieval. Figure 2 As shown, the layer structure is as follows: Figure 3 As shown.
[0048] Step 2.1: Establish the initial feature extraction module, which includes: one convolutional layer, one batch normalization layer, one activation layer and one pooling layer. The Gaussian random initialization method is used to initialize the weights of the convolutional layer in the initial feature extraction module.
[0049] The k-th footprint image The input is fed into the footprint image retrieval network, and after processing by the initial feature extraction module, the initial footprint features of the k-th footprint image are obtained.
[0050] Step 2.2: Construct a multi-feature dense connection network module, including: R dense connection layers and R transformation layers, wherein any r-th dense connection layer sequentially contains n convolutional layers with kernels of X1×X1 and n convolutional layers with kernels of X2×X2; any r-th transformation layer module sequentially contains one convolutional layer with kernels of X3×X3 and one average pooling layer with kernels of X4×X4;
[0051] When r = 1, the initial footprint features of the k-th footprint image After processing by the r-th dense connection layer, the dense connection features of the r-th layer are obtained. After processing by the r-th transformation layer, the dense connection features of the r-th layer are obtained.
[0052] When r = 2, 3, ..., R, the (r-1)th dense connection feature After processing by the r-th dense layer and the r-th transformation layer, the r-th dense connection feature is obtained. Thus, the Rth dense connection feature is output by the Rth transformation layer.
[0053] Step 2.3: Construct a multi-scale fusion module, which includes: m convolutional layers with kernels of X5×X5 and m convolutional layers with kernels of X6×X6;
[0054] The Rth dense connection feature The input to the multi-scale fusion module is processed sequentially through m convolutional layers with kernels of X5×X5 and m convolutional layers with kernels of X6×X6 to obtain high-level features. Then use interpolation for high-rise buildings After upsampling, the upsampled features are obtained. Finally, the upsampled features are... Dense connection features of layer R-1 After fusion, the multi-scale fusion features of the k-th image are obtained.
[0055] Step 2.4: Construct the attention enhancement module, including: a global pooling layer, two fully connected layers, and a regression function; densely connect the features of the (R-1)th layer. Dense connection features of layer R and multi-scale fusion features The data are input into the attention enhancement module and processed sequentially through a global pooling layer and two fully connected layers. The results are then weighted using a regression function to obtain the k-th footprint image of the m-th test subject. The R-1 layer dense connection attention feature vector The k-th footprint image of the m-th test subject The Rth layer dense connection attention feature direction The k-th footprint image of the m-th test subject Multi-scale fusion features
[0056] Step 2.5: Establish the feature output module, including: average pooling layer and fully connected layer;
[0057] Feature R-th dense connection feature The output features are obtained by sequentially passing through an average pooling layer and a fully connected layer. As the final representation, it is retrieved as a characterization vector during testing;
[0058] Step 3: Construct the loss function:
[0059] Step 3.1: Establish the loss function L1 using equation (1):
[0060]
[0061] In equation (1), This represents the k-th footprint image of the 0-th test subject. The R-1th layer densely connected attention feature vector, This represents the q-th footprint image of the m-th test subject. The R-1th layer densely connected attention feature vector; o≤M,q≤T.
[0062] Step 3.2: Use equation (2) to establish the loss function L2:
[0063]
[0064] In equation (2), This represents the R-th layer dense connection attention feature vector of the k-th image of the o-th test subject. This represents the R-th layer densely connected attention feature vector of the q-th image of the m-th test subject;
[0065] Step 3.3: Establish the loss function L3 using equation (3):
[0066]
[0067] In equation (3), This represents the multi-scale fusion attention feature vector of the k-th image of the o-th test subject. This represents the pyramid fusion attention feature vector of the q-th image of the m-th test subject;
[0068] Step 3.3: Use equation (4) to establish the total loss function L:
[0069] L=L1+L2+L3 (4)
[0070] Step 4: Based on the processed sample set S, the footprint image retrieval network is trained using the gradient descent method, and the total loss function L is calculated to update the network parameters until the total loss function L converges, thereby obtaining the trained footprint image retrieval model, which is used to retrieve and match footprint images to be retrieved.
[0071] For retrieval problems, rank1 and map values are commonly used to evaluate model performance. Therefore, all images in the query set are treated as images to be retrieved, and the rank1 and map values on the test set are calculated by averaging the results. On the test set used in this invention, the optimal retrieval results were obtained with a rank1 accuracy of 95.9% and a map value of 86.2%.
[0072] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the aforementioned footprint image retrieval method. The processor is configured to execute the program stored in the memory.
[0073] In this embodiment, a computer-readable storage medium stores a computer program, which, when run by a processor, executes the steps of the aforementioned footprint image retrieval method.
Claims
1. A method for retrieving footprint images using a multi-feature densely connected attention-enhanced network, characterized in that, It is done according to the following steps: Step 1: Acquire footprint images of T test subjects using an optical footprint acquisition instrument, and perform filtering, noise reduction, size scaling, horizontal flipping, and random affine transformation to obtain the processed sample set, denoted as . ,in, This represents the processed footprint sample set of the m-th collected object, where M represents the total number of collected objects. ,and , This represents the k-th footprint image after processing of the m-th collected object; K Represents the total number of footprint images; Step 2: Establish a footprint image retrieval network based on a multi-feature dense connection attention enhancement network, including: an initial feature extraction module, a multi-feature dense connection network module, a multi-scale feature fusion module, an attention enhancement module, and a feature output module; Step 2.1: Establish an initial feature extraction module, which includes: a convolutional layer, a batch normalization layer, an activation layer, and a pooling layer. The Gaussian random initialization method is used to initialize the weights of the convolutional layer in the initial feature extraction module. The k-th footprint image The initial footprint features of the k-th footprint image are obtained after being processed by the initial feature extraction module and input into the footprint image retrieval network. ; Step 2.2: Construct a multi-feature densely connected network module, including: R densely connected layers and R transformation layers, wherein any r-th densely connected layer sequentially contains n convolutional kernels. The convolutional layer and n convolutional kernels are The convolutional layer; any r-th transformation layer module sequentially includes a convolutional kernel of... The convolutional layer and one convolutional kernel are The average pooling layer; When r=1, the initial footprint features of the k-th footprint image After processing by the r-th dense connection layer, the dense connection features of the r-th layer are obtained. After processing by the r-th transformation layer, the dense connection features of the r-th layer are obtained. ; When r = 2, 3, ..., R, the (r-1)th dense connection feature After processing by the r-th dense layer and the r-th transformation layer, the r-th dense connection feature is obtained. Thus, the Rth dense connection feature is output by the Rth transformation layer. ; Step 2.3: Construct a multi-scale feature fusion module, which includes m convolutional kernels. The convolutional layer has m convolutional kernels. Convolutional layers; The R-th dense connection feature The input to the multi-scale feature fusion module passes through m convolutional kernels sequentially. The convolutional layer has m convolutional kernels. After processing by the convolutional layer, high-level features are obtained. Then use interpolation for high-rise buildings After upsampling, the upsampled features are obtained. Finally, the upsampled features Dense connection features of layer R-1 After fusion, the multi-scale fusion features of the k-th image are obtained. ; Step 2.4: Construct an attention enhancement module, including: a global pooling layer, two fully connected layers, and a regression function; [This will] integrate the (R-1)th densely connected feature... The Rth dense connection feature and multi-scale fusion features The data are input into the attention enhancement module and processed sequentially through a global pooling layer and two fully connected layers. The results are then weighted using a regression function to obtain the k-th footprint image of the m-th test subject. The R-1 layer dense connection attention feature vector The k-th footprint image of the m-th test subject The densely connected attention feature vector of the Rth layer The k-th footprint image of the m-th test subject Multi-scale fusion feature vector ; Step 2.5: Establish the feature output module, including: average pooling layer and fully connected layer; The R-th dense connection feature The output features are obtained by sequentially passing through an average pooling layer and a fully connected layer. As the final representation, it is retrieved as a characterization vector during testing; Step 3: Construct the loss function: Step 3.1: Establish the loss function using equation (1) : (1) In equation (1), This represents the k-th footprint image of the 0-th test subject. The R-1th layer densely connected attention feature vector, This represents the q-th footprint image of the m-th test subject. The densely connected attention feature vector of the R-1th layer; , ; Step 3.2: Establish the loss function using equation (2) : (2) In equation (2), This represents the R-th layer dense connection attention feature vector of the k-th image of the o-th test subject. This represents the R-th layer densely connected attention feature vector of the q-th image of the m-th test subject; Step 3.3: Establish the loss function using equation (3) : (3) In equation (3), This represents the multi-scale fusion attention feature vector of the k-th image of the o-th test subject. This represents the pyramid fusion attention feature vector of the q-th image of the m-th test subject; Step 3.3: Establish the total loss function using equation (4) : (4) Step 4: Based on the processed sample set The gradient descent method was used to train the footprint image retrieval network, and the total loss function was calculated. To update network parameters until the total loss function is reached. The training continues until convergence, thus obtaining the trained footprint image retrieval model, which is used to retrieve and match footprint images to be retrieved.
2. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the footprint image retrieval method of claim 1, and the processor is configured to execute the program stored in the memory.
3. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the steps of the footprint image retrieval method of claim 1.