Footprint optical image retrieval method based on multi-feature expansion residual convolutional network
By combining multi-feature extended residual convolutional networks and attention mechanisms, the problems of large feature extraction errors and high costs in traditional footprint recognition methods are solved, achieving efficient footprint image retrieval and improving retrieval accuracy and speed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2023-05-16
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional footprint recognition methods rely on expert experience, have large feature extraction errors and high costs, and are difficult to form standardized feature descriptions. Furthermore, existing technologies are unable to effectively extract multi-layer features from footprint images, resulting in insufficient retrieval accuracy and speed.
A footprint optical image retrieval method based on multi-feature dilated residual convolutional network is adopted. Through data preprocessing and multi-layer feature extraction module, combined with attention mechanism and loss function, a footprint optical image retrieval model is constructed to extract and fuse multi-layer feature information.
It improves the accuracy and speed of footprint image retrieval, saves human and material resources, and can better identify and classify the footprint features of the same person, thus enhancing retrieval performance.
Smart Images

Figure CN117171377B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and in particular to a method for retrieving footprint optical images based on a multi-feature dilated residual convolutional network. Background Technology
[0002] Footprint optical images contain unique human biological characteristics. Compared with early identification technologies, biometric technology has the advantages of automation, standardization, easy feature collection, difficulty in forgery, and high reliability. Although biometric identification methods such as DNA, face, and iris recognition are widely used, footprints have irreplaceable advantages in the field of biometric identification due to their high stability, strong concealment, and difficulty in disguise. Therefore, conducting research on footprint feature recognition is of great significance.
[0003] In the medical field, doctors can use plantar pressure information to provide patients with precise and comprehensive treatment plans.
[0004] Traditional footprint recognition relies heavily on the subjective experience of experts. Furthermore, the extracted footprint features vary considerably, making it difficult to develop standardized feature descriptions and extraction methods. In summary, traditional footprint research methods place high demands on researchers, and manual partitioning and feature annotation inherently suffer from large errors, high costs, and a limited number of features, making them unsuitable for practical tasks. Summary of the Invention
[0005] This invention aims to address the shortcomings of existing technologies by proposing a footprint optical image retrieval method based on a multi-feature dilated residual convolutional network. The goal is to deeply mine various hidden features at both the shallow and deep levels of footprint images, thereby improving the accuracy and speed of footprint image retrieval.
[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0007] The present invention provides a footprint optical image retrieval method based on a multi-feature dilated residual convolutional network, characterized by the following steps:
[0008] Step 1: Data acquisition and preprocessing of footprint images:
[0009] An optical pressure footprint acquisition instrument was used to acquire a set of footprint images of M test subjects. These images were then subjected to median filtering for noise reduction, image scale removal, scaling, horizontal flipping, random affine transformation, and random erasure to obtain the processed sample set denoted as Y = {Y1, Y2, ..., Y}. m , ..., Y M}, where Y m Let m represent the footprint sample set after processing the m-th test subject, where M represents the total number of test subjects, m ≤ M, and This represents the nth footprint image of the m-th test subject; N represents the total number of footprint images collected for each test subject.
[0010] Step 2: Establish a multi-feature dilated residual convolutional network enhanced by attention mechanism, including: initial feature extraction module STAGE0, dilated residual convolutional network, multi-scale feature fusion module, and attention enhancement module;
[0011] Step 2.1: Establish the initial feature extraction module STAGE0, which includes: one convolutional layer, one batch normalization layer, one activation layer, and one max pooling layer.
[0012] The nth footprint image After processing by the initial feature extraction module STAGE0, the initial output features are obtained.
[0013] Step 2.2: Construct an extended residual convolutional network, including: a first feature extraction module STAGE1, a second feature extraction module STAGE2, a third feature extraction module STAGE3, a fourth feature extraction module STAGE4, and one transposed convolutional layer;
[0014] Each feature extraction module includes several BTNK units. Each BTNK unit includes, in sequence, several small blocks, three activation layers, and residual connections. Each small block consists of a convolutional layer and a batch normalization layer.
[0015] Initial output features The input is fed into a dilated residual convolutional network and processed sequentially through four feature extraction modules, resulting in four output features. in, This represents the i-th output feature obtained by the i-th feature extraction module STAGEi;
[0016] The fourth output feature obtained by the fourth feature extraction module STAGE4 After being processed in the transposed convolutional layer, the footprint features of the nth footprint image are obtained.
[0017] Step 2.3: Construct a multi-scale feature fusion module, which includes: k convolutional layers with kernel size of X5×X5 and k convolutional layers with kernel size of X6×X6;
[0018] The footprint features of the nth footprint image are input into the multi-scale feature fusion module, and then processed sequentially through k convolutional layers with kernel size X5×X5 and k convolutional layers with kernel size X6×X6 to obtain high-level features.
[0019] Upsampling method using interpolation for high-level features After processing, the upsampled features are obtained. Finally, the upsampled features are... Compared with the initial output features After fusion, the multi-scale fusion features of the nth footprint image are obtained.
[0020] Step 2.4: Construct the attention enhancement module, which includes: a global pooling layer, two fully connected layers, and a regression function;
[0021] Will Multi-scale fusion features The inputs are fed into the attention enhancement module and processed by the global pooling layer and two fully connected layers to obtain four attention features and one multi-scale fusion attention feature.
[0022] The four attention features are then weighted separately using the regression function to obtain the nth footprint image of the mth test subject. Four attention feature vectors in, This represents the i-th attention feature vector;
[0023] The multi-scale fused attention features are then weighted using the regression function to obtain the nth footprint image of the mth test subject. Multiscale fusion attention features
[0024] Step 3: Construct the loss function:
[0025] Step 3.1: Establish the first loss function L1 using equation (1):
[0026]
[0027] In equation (1), f represents the triplet loss. This represents the i-th attention feature vector output by the attention enhancement module from the b-th footprint image of the a-th test subject. This represents the i-th attention feature vector output by the attention enhancement module from the q-th footprint image of the p-th test subject. Denotes the 2-norm, [] + This represents the ReLU function, also called the corrected linear unit. When its input is greater than zero, the output is equal to the input; otherwise, the output is zero. 0 < a ≤ M; 0 < q ≤ N.
[0028] Step 3.2: Establish the second loss function L2 using equation (2):
[0029]
[0030] In equation (2), This represents the (i+1)th attention feature vector output by the attention enhancement module from the b-th footprint image of the a-th test subject. This represents the (i+1)th attention feature vector output by the attention enhancement module from the qth footprint image of the p-th test subject; i ≠ 4;
[0031] Step 3.3: Establish the third loss function L3 using equation (3):
[0032]
[0033] In equation (3), Let represent the multi-scale fusion attention feature vector of the b-th image of the a-th test subject. This represents the multi-scale fusion attention feature vector of the qth image of the p-th test object;
[0034] Step 3.3: Use equation (4) to establish the total loss function L:
[0035] L=L1+L2+L3 (4)
[0036] Step 4: Based on the processed sample set Y, the footprint image retrieval network is trained using the gradient descent method, and the total loss function L is calculated to update the network parameters until the total loss function L converges, thereby obtaining the trained footprint image retrieval model, which is used for footprint retrieval and identity matching of the footprint images to be retrieved.
[0037] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the footprint image retrieval method, and the processor is configured to execute the program stored in the memory.
[0038] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the footprint image retrieval method.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] 1. This invention combines image processing technology, neural network technology from deep learning, and image retrieval technology. It applies this to footprint optical images and designs a complete footprint optical image retrieval model. The acquired footprint optical images are processed, and multi-layered feature information is extracted and fused. A backbone neural network is built based on the idea of dilated residual convolutional networks. Simultaneously, an attention mechanism is applied to focus on locally important information in the footprint optical image from both channel and spatial perspectives, thus completing the construction of the feature extraction network. Compared with traditional methods, footprint optical images fused with multiple features can extract more detailed features of the footprints, improving the accuracy of footprint retrieval and the network's operating speed, while saving significant human and material resources.
[0041] 2. The dataset acquisition device of this invention is a footprint optical image acquisition device. Data preprocessing includes image denoising, image cropping, removal of the acquisition instrument scale, label processing, image correction, and image normalization. The purpose of preprocessing is to feed the footprint optical image into the network model for training, thereby obtaining the optimal model to achieve the best retrieval effect.
[0042] 3. The multi-feature dilated residual convolutional network module designed in this invention can fully extract features from multiple layers of images through dilated residual convolutional networks, and better identify, classify and retrieve footprint feature information belonging to the same person by associating the features of footprints in each layer, thereby improving the accuracy of footprint retrieval.
[0043] 4. The dilated residual convolutional network used in this invention focuses on the downsampling module of the original residual convolutional network. Since downsampling will cause the loss of details of image features, the concept of dilation is introduced. Specifically, after the footprint optical image passes through the transposed convolutional layer in the dilated residual convolutional network, the output features are enlarged many times, so that the footprint optical image retains more detailed features.
[0044] 5. This invention incorporates a multi-scale feature fusion module and selects an appropriate loss function for Euclidean distance measurement based on the characteristics of features at different layers. The loss function is calculated for features at different depths, which improves the matching degree of the loss function to the footprint optical image and better trains the retrieval model.
[0045] 6. The attention enhancement module of the present invention has the ability to be embedded into the dilated residual convolutional network. It can improve the feature extraction capability of the network model without significantly increasing the amount of computation and parameters. It processes features from both channel and spatial dimensions, realizes local feature extraction of footprint optical images, and improves the performance of the retrieval network. Attached Figure Description
[0046] Figure 1 This is an overall flowchart of footprint optical image retrieval in this invention;
[0047] Figure 2 This is a framework diagram of the footprint optical image retrieval system in this invention;
[0048] Figure 3 This is a structural diagram of the initial feature extraction module STAGE0 in this invention;
[0049] Figure 4 This is a structural diagram of the dilated residual convolutional network feature extraction module in this invention;
[0050] Figure 5a This is a structural diagram of the BTNK1 unit in this invention;
[0051] Figure 5b This is a structural diagram of the BTNK2 unit in this invention;
[0052] Figure 6 This is a structural diagram of the footprint optical image retrieval network in this invention. Detailed Implementation
[0053] In this embodiment, a footprint optical image retrieval method based on a multi-feature dilated residual convolutional network mainly extracts representational information from footprint optical images through a neural network for retrieval. It utilizes multi-feature fusion and attention mechanisms to enhance the feature extraction capability of the neural network, and amplifies the feature map by dilated residual convolutional networks, thereby enabling in-depth extraction of feature and detail information from footprint optical images. This improves various indicators of footprint optical image retrieval (running speed, retrieval accuracy). The method is performed according to the following steps:
[0054] Step 1: Data acquisition and preprocessing of footprint images. Data preprocessing includes operations such as filtering and denoising the acquired data, image scaling, horizontal flipping, and random affine transformation.
[0055] An optical pressure footprint acquisition instrument was used to acquire a set of footprint images of M test subjects. These images were then subjected to median filtering for noise reduction, image scale removal, scaling, horizontal flipping, random affine transformation, and random erasure to obtain the processed sample set denoted as Y = {Y1, Y2, ..., Y}. m , ..., Y M}, where Y m Let m represent the footprint sample set after processing the m-th test subject, where M represents the total number of test subjects, m ≤ M, and This represents the nth footprint image of the m-th test subject; N represents the total number of footprint images collected for each test subject.
[0056] The dataset used in this invention was independently collected and processed by the Footprint Laboratory of Anhui University. The dataset was divided into a training set and a test set. The training set contains optical images of footprints from 400 individuals, with 20 images of each person's left and right feet, totaling 8000 images. The test set consists of two parts: a gallery and a query set, each containing optical images of footprints from 80 individuals, but with a difference in data volume. The gallery contains 1600 footprint images, while the query set contains only 400: 10 images of each person's left and right feet, totaling 20 images. The query set contains 5 images of each person's left and right feet. The process of constructing the dataset and the overall workflow are as follows: Figure 1 and Figure 2 As shown.
[0057] Step 2: Establish a multi-feature dilated residual convolutional network enhanced by attention mechanism, including: initial feature extraction module STAGE0, dilated residual convolutional network, multi-scale feature fusion module, and attention enhancement module;
[0058] Step 2.1: Establish the initial feature extraction module STAGE0, which includes: one convolutional layer, one batch normalization layer, one activation layer, and one max pooling layer; as follows: Figure 3 The diagram shown is a structural diagram of the initial feature extraction module STAGE0 in this invention. STAGE0 has a relatively simple structure and can be regarded as a preprocessing step for the data input. The 224*224*3 footprint image is sequentially passed through a convolutional layer with a kernel size of X7×X7, a batch normalization layer, an activation layer, and a max pooling layer to complete STAGE0 and realize the data input.
[0059] for Figure 3 The structure of STAGE0, the nth footprint image The specifications are (3, 224, 224), representing the number of input channels, height, and width, respectively, i.e., (C, H, W). Since the network requires the input image size to be H = W, it is represented as (C, W, W). Step 1 of STAGE0 includes three sequential operations: convolution, batch normalization, and activation. Step 2 of STAGE0 is a max-pooling layer with a kernel size of X3 × X3.
[0060] The nth footprint image After processing by the initial feature extraction module STAGE0, the initial output features are obtained.
[0061] Step 2.2: Construct an extended residual convolutional network, including: a first feature extraction module STAGE1, a second feature extraction module STAGE2, a third feature extraction module STAGE3, a fourth feature extraction module STAGE4, and one transposed convolutional layer. For example... Figure 4 The diagram shown is a structural diagram of the feature extraction module of the dilated residual convolutional network in this invention.
[0062] Each feature extraction module includes several BTNK units. Each BTNK unit includes, in sequence, several small blocks, three activation layers, and residual connections. Each small block consists of a convolutional layer and a batch normalization layer.
[0063] like Figure 5a and Figure 5b As shown, the BTNK unit exists in two forms: BTNK1 and BTNK2. These two types of units correspond to two different situations: the number of input and output channels is the same (BTNK2), and the number of input and output channels is different (BTNK1).
[0064] Figure 5a The BTNK1 unit has four variable parameters: C, W, C1, and S. Compared to the BTNK2 unit, the BTNK1 unit has one more parallel convolutional layer, denoted as G(x). The BTNK1 unit corresponds to the case where the number of channels in the input x and the output F(x) are different. G(x) plays the role of matching the difference in the dimensions of the input and output (G(x) has the same number of channels as F(x)). Therefore, F(x) + G(x) can be summed, thereby changing the dimension of the network.
[0065] Figure 5b The BTNK2 unit has only two variable parameters, C and W, which are C and W in the shape (C, W, W) of the image input. Let the input with shape (C, W, W) be x, and let the three blocks of the BTNK2 unit be functions. Add the two (F(x)+x) and then pass it through a ReLU activation function to obtain the output of the BTNK2 unit. The shape of the output is still (C, W, W), which corresponds to the case where the number of input and output channels of the BTNK2 unit is the same.
[0066] Initial output features The input is fed into a dilated residual convolutional network and processed sequentially through four feature extraction modules, resulting in four output features. in, This represents the i-th output feature obtained by the i-th feature extraction module STAGEi;
[0067] The fourth output feature obtained by the fourth feature extraction module STAGE4 After processing in the transposed convolutional layer, the footprint features of the nth footprint image are obtained. The transposed convolutional layer is a convolutional layer with a kernel size of X4×X4, a stride of 4, and zero padding. This layer dilates the residual convolutional network to process the input features. Expanded to many times its original size;
[0068] Step 2.3: Construct a multi-scale feature fusion module, which includes: k convolutional layers with kernel size of X5×X5 and k convolutional layers with kernel size of X6×X6;
[0069] The footprint features of the nth footprint image are input into the multi-scale feature fusion module, and then processed sequentially through k convolutional layers with kernel size X5×X5 and k convolutional layers with kernel size X6×X6 to obtain high-level features.
[0070] Upsampling method using interpolation for high-level features After processing, the upsampled features are obtained. Finally, the upsampled features are... Compared with the initial output features After fusion, the multi-scale fusion features of the nth footprint image are obtained.
[0071] Step 2.4: Construct the attention enhancement module, which includes: a global pooling layer, two fully connected layers, and a regression function;
[0072] Will Multi-scale fusion features The inputs are fed into the attention enhancement module and processed by a global pooling layer and two fully connected layers to obtain four attention features and one multi-scale fusion attention feature.
[0073] The four attention features are then weighted separately using a regression function to obtain the nth footprint image of the mth test subject. Four attention feature vectors in, This represents the i-th attention feature vector;
[0074] Multi-scale fusion of attention features is then weighted using a regression function to obtain the nth footprint image of the mth test subject. Multiscale fusion attention features
[0075] Step 3: Construct the loss function:
[0076] Step 3.1: Establish the first loss function L1 using equation (1):
[0077]
[0078] In equation (1), f represents the triplet loss. This represents the i-th attention feature vector output by the attention enhancement module from the b-th footprint image of the a-th test subject. This represents the i-th attention feature vector output by the attention enhancement module for the q-th footprint image of the p-th test subject. Denotes the 2-norm, [] + This represents the ReLU function, also called the corrected linear unit. When its input is greater than zero, the output is equal to the input; otherwise, the output is zero. 0 < a ≤ M; 0 < q ≤ N.
[0079] Step 3.2: Establish the second loss function L2 using equation (2):
[0080]
[0081] In equation (2), This represents the (i+1)th attention feature vector output by the attention enhancement module from the b-th footprint image of the a-th test subject. This represents the (i+1)th attention feature vector output by the attention enhancement module for the q-th footprint image of the p-th test subject; i≠4;
[0082] Step 3.3: Establish the third loss function L3 using equation (3):
[0083]
[0084] In equation (3), Let represent the multi-scale fusion attention feature vector of the b-th image of the a-th test subject. This represents the multi-scale fusion attention feature vector of the qth image of the p-th test object;
[0085] Step 3.3: Use equation (4) to establish the total loss function L:
[0086] L=L1+L2+L3 (4)
[0087] Step 4: Based on the processed sample set Y, the footprint image retrieval network is trained using gradient descent, and the total loss function L is calculated to update the network parameters until the total loss function L converges, thus obtaining the trained footprint image retrieval model, which is used for footprint retrieval and identity matching of the footprint images to be retrieved. Figure 6 The diagram shows the overall structure of the footprint optical image retrieval network in this invention.
[0088] For the retrieval problem, the model's performance is evaluated using two metrics: Rank1 and mAP. All images in the target database are used as retrieval images to obtain the retrieval results, and the average value is calculated to obtain the Rank1 and mAP values on the test set. On the test set used in this invention, the best retrieval results were obtained with a Rank1 accuracy of 97.9% and a mAP value of 92.01%.
[0089] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the footprint image retrieval method, and the processor is configured to execute the program stored in the memory.
[0090] In this embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the footprint image retrieval method.
Claims
1. A method for retrieving footprint optical images based on a multi-feature dilated residual convolutional network, characterized in that, It is done according to the following steps: Step 1: Data acquisition and preprocessing of footprint images: Collected using an optical pressure footprint acquisition instrument A set of footprint images of the test subjects was processed by median filtering for noise reduction, removal of image rulers, scaling, horizontal flipping, random affine transformation, and random erasure. The resulting processed sample set is denoted as […]. ,in, This represents the footprint sample set after processing the m-th test object, where M represents the total number of test objects. ,and , This represents the nth footprint image of the m-th test subject; N represents the total number of footprint images collected for each test subject. Step 2: Establish a multi-feature dilated residual convolutional network enhanced by attention mechanism, including: initial feature extraction module STAGE0, dilated residual convolutional network, multi-scale feature fusion module and attention enhancement module; Step 2.1: Establish the initial feature extraction module STAGE0, which includes: one convolutional layer, one batch normalization layer, one activation layer, and one max pooling layer. The nth footprint image After processing by the initial feature extraction module STAGE0, the initial output features are obtained. ; Step 2.2: Construct an extended residual convolutional network, including: a first feature extraction module STAGE1, a second feature extraction module STAGE2, a third feature extraction module STAGE3, a fourth feature extraction module STAGE4, and one transposed convolutional layer; Each feature extraction module includes several BTNK units. Each BTNK unit includes, in sequence, several small blocks, three activation layers, and residual connections. Each small block consists of a convolutional layer and a batch normalization layer. Initial output features The input is fed into a dilated residual convolutional network and processed sequentially through four feature extraction modules, resulting in four output features. ;in, This represents the i-th output feature obtained by the i-th feature extraction module STAGEi; The fourth output feature obtained by the fourth feature extraction module STAGE4 After being processed in the transposed convolutional layer, the footprint features of the nth footprint image are obtained. ; Step 2.3: Construct a multi-scale feature fusion module, which includes k convolutional kernels of size [missing information]. The convolutional layer has k convolutional kernels of size [value missing]. Convolutional layers; The footprint features of the nth footprint image are input into the multi-scale feature fusion module and sequentially processed by k convolutional kernels of size [missing value]. The convolutional layer has k convolutional kernels of size [value missing]. After processing by the convolutional layer, high-level features are obtained. ; Upsampling method using interpolation for high-level features After processing, the upsampled features are obtained. Finally, the upsampled features Compared with the initial output features After fusion, the multi-scale fusion features of the nth footprint image are obtained. ; Step 2.4: Construct the attention enhancement module, which includes: a global pooling layer, two fully connected layers, and a regression function; Will Multi-scale fusion features The inputs are fed into the attention enhancement module and processed by the global pooling layer and two fully connected layers to obtain four attention features and one multi-scale fusion attention feature. The four attention features are then weighted separately using the regression function to obtain the nth footprint image of the mth test subject. Four attention feature vectors ;in, This represents the i-th attention feature vector; The multi-scale fused attention features are then weighted using the regression function to obtain the nth footprint image of the mth test subject. Multiscale fusion attention features ; Step 3: Construct the loss function: Step 3.1: Establish the first loss function using equation (1) : (1) In equation (1), Indicates the loss of the triplet. This represents the i-th attention feature vector output by the attention enhancement module from the b-th footprint image of the a-th test subject. This represents the i-th attention feature vector output by the attention enhancement module from the q-th footprint image of the p-th test subject. Represents the L2 norm, This represents the ReLU function, also called the corrected linear unit. When its input is greater than zero, the output is equal to the input; otherwise, the output is zero. ; ; Step 3.2: Establish the second loss function using equation (2) : (2) In equation (2), This represents the (i+1)th attention feature vector output by the attention enhancement module from the b-th footprint image of the a-th test subject. This represents the (i+1)th attention feature vector output by the attention enhancement module from the qth footprint image of the pth test subject. ; Step 3.3: Establish the third loss function using equation (3) : (3) In equation (3), Let represent the multi-scale fusion attention feature vector of the b-th image of the a-th test subject. This represents the multi-scale fusion attention feature vector of the qth image of the p-th test object; Step 3.3: Establish the total loss function using equation (4) : (4) Step 4: Based on the processed sample set The gradient descent method was used to train the footprint image retrieval network, and the total loss function was calculated. To update network parameters until the total loss function is reached. The training continues until convergence, thus obtaining the trained footprint image retrieval model, which is used to perform footprint retrieval and identity matching on the footprint images to be retrieved.
2. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the footprint image retrieval method of claim 1, and the processor is configured to execute the program stored in the memory.
3. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the steps of the footprint image retrieval method of claim 1.