Shoeprint retrieval method based on feature distribution reconstruction and multi-granularity branch fusion
Through the method of feature distribution reconstruction and multi-granularity branch fusion, the problem of insufficient local information extraction of convolutional neural networks in shoe print retrieval is solved, and efficient and accurate feature extraction and recognition of shoe print images are achieved.
Patent Information
- Application Number
- CN202510946755.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing convolutional neural networks fail to extract local information of on-site shoe prints in shoe print retrieval, resulting in reduced accuracy and completeness of feature extraction and inability to effectively utilize the sole texture details under complex backgrounds.
A method based on feature distribution reconstruction and multi-granularity branch fusion is adopted, through the Resnet50 network, TCCA module, DHSA module, DIGS module and MHFA module, to achieve efficient fusion of multi-scale features and accurate extraction of fine-grained regional information.
It significantly improves the accuracy and robustness of shoe print retrieval, can effectively capture subtle texture and shape differences in shoe print images, and improves the feature extraction capability and retrieval algorithm performance in complex scenarios.
Smart Images

Figure CN120448572B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning and target retrieval, and in particular to a shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion. Background Art
[0002] Shoe prints left at crime scenes hold significant forensic value, prompting researchers to continuously develop more accurate and efficient shoe print retrieval algorithms. Research into intelligent footprint analysis technology typically involves two phases, with shoe print retrieval being the core task of the first phase. This phase aims to analyze shoe print images from the scene and retrieve identical shoe prints from a database, laying the foundation for subsequent footprint identification.
[0003] With the rapid development of deep learning technology, shoeprint retrieval algorithms based on convolutional neural networks have achieved remarkable results. However, these algorithms still face many challenges in practical application, particularly the inadequate extraction of local information from shoeprints at the scene. At crime scenes, shoeprint images are often obscured by complex backgrounds, causing the sole texture to blend with the surroundings, resulting in partial loss of the sole pattern and increased noise. Therefore, analyzing the local details of the sole pattern becomes crucial for shoeprint retrieval. Although traditional convolutional neural networks excel at extracting image features, they primarily focus on overall features while ignoring local details. This processing approach can miss important details, reducing the accuracy and completeness of feature extraction, and limiting the effectiveness of the algorithm in practical applications.
[0004] However, there is currently no technical solution that can solve the above technical problems, and there is no shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion. Summary of the Invention
[0005] The present invention provides a shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion, which realizes shoe print retrieval of shoe print images to be processed by efficiently fusing multi-scale features and accurately extracting fine-grained regional information.
[0006] In a first aspect, the present invention provides a shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion, comprising:
[0007] Input the shoe print image to be processed into the preset Resnet50 front-end structure to obtain the first feature, and input the first feature into the TCCA module, DHSA module, and DIGS module to obtain large-scale features;
[0008] Input the large-scale features to the preset Resnet50 back structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, input the large-scale features, medium-scale features and small-scale features to the MHFA module to obtain the first output features;
[0009] Input the large-scale feature to the local feature extraction branch of the two-segmentation combination to obtain a second output feature, and input the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain a third output feature;
[0010] A target feature is determined based on the first output feature, the second output feature, and the third output feature. A similarity is calculated between the target feature and the image feature of each image in a preset shoe print database, and the type of shoes and brand corresponding to the image with the highest similarity is determined as the target retrieval result of the shoe print image to be processed.
[0011] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the preset Resnet50 includes a Resnet50 front part structure and a Resnet50 back part structure;
[0012] The front part of the Resnet50 structure includes a 7×7 convolution kernel, a second maximum pooling layer, a first residual block group consisting of three Bottleneck modules, a second residual block group consisting of four Bottleneck modules, and the first Bottleneck module in the third residual block group;
[0013] The latter part of the Resnet50 structure includes the second to sixth Bottleneck modules in the third residual block group, and a fourth residual block group consisting of three Bottleneck modules.
[0014] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the first feature is input to the TCCA module, the DHSA module and the DIGS module to obtain large-scale features, including:
[0015] Input the first feature to the TCCA module, perform pooling and inverse transform operations on the first feature in the height direction, multiply it element-wise with the original feature map to obtain a second feature, perform adaptive average pooling on the first feature, obtain a third feature through a transposition operation, and add the second and third features to obtain a fourth feature;
[0016] In the width direction, the first feature is processed to obtain the fifth feature, feature map pooling is performed on the first feature to obtain the sixth feature, the fourth feature, the fifth feature, and the sixth feature are averaged to obtain the seventh feature, and the seventh feature is point-wise multiplied by the first feature to obtain the eighth feature output by the TCCA module;
[0017] The eighth feature is input to the DHSA module to obtain the ninth feature, and the ninth feature is input to the DIGS module to obtain the large-scale feature.
[0018] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the eighth feature is input to the DHSA module to obtain the ninth feature, including:
[0019] The eighth feature is split into two channels, sorted along the height and width directions respectively, and the query, key and value are generated using the first 1×1 convolutional layer. The query and key are padded, and the feature map is rearranged using the Reshape_attn function. The similarity matrix is calculated to obtain the attention weight, and the attention weight is multiplied by the value to generate a context feature map.
[0020] The context feature map is processed using the output projection layer to obtain a processed feature map, the original order of the processed feature map is restored using the Torch.scatter operation, and the fused feature map is used to obtain the ninth feature using the second 1×1 convolutional layer.
[0021] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the ninth feature is input to the DIGS module to obtain the large-scale feature, including:
[0022] After inputting the ninth feature into the third 1×1 convolutional layer, upsampling and blocking operations are performed to obtain the first branch feature and the second branch feature;
[0023] After performing the Split operation on the first branch feature, the features are respectively input into the first 3×3 convolutional layer, the 1×11 convolutional layer, the 11×1 convolutional layer, and the Identity module to obtain the first output result, the second output result, the third output result, and the fourth output result. The first output result, the second output result, the third output result, and the fourth output result are merged to obtain the first branch output;
[0024] After inputting the second branch feature into the second 3×3 convolutional layer, the Mish activation function is used to obtain the second branch output;
[0025] After performing a dot multiplication operation on the first branch output and the second branch output, downsampling is performed to obtain downsampled features, and the downsampled features are input to the fourth 1×1 convolutional layer to obtain the large-scale features.
[0026] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the input of the large-scale features, the medium-scale features, and the small-scale features to the MHFA module to obtain the first output features includes:
[0027] The large-scale features are processed using downsampling and a fifth 1×1 convolutional layer to obtain a first processed feature, the medium-scale features are processed using a sixth 1×1 convolutional layer to obtain a second processed feature, the first processed features and the second processed features are added element-by-element to obtain a third processed feature, multi-scale feature extraction is performed on the third processed features using four 3×3 convolutional layers with different expansion rates, the third processed features are connected in series, and the third processed features are fused sequentially through a seventh 1×1 convolutional layer and a batch normalization layer to obtain a fourth processed feature;
[0028] Performing element-by-element addition on the fourth processed feature, the first processed feature, the second processed feature, and the attention-weighted feature map to obtain a fifth processed feature, processing the fifth processed feature using an adaptive average pooling layer to obtain a sixth processed feature, adding the sixth processed feature to the small-scale feature to obtain a seventh processed feature, performing self-attention processing on the seventh processed feature using two SelfAttentionBlock modules to obtain an eighth processed feature, performing element-by-element multiplication and addition operations on the eighth processed feature and the small-scale feature, and processing the result through an eighth 1×1 convolutional layer to obtain a ninth processed feature, processing the ninth processed feature using a first Reduction module to obtain the first output feature.
[0029] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the inputting of the large-scale feature to the local feature extraction branch combined with two-segmentation to obtain the second output feature includes:
[0030] Input the large-scale features into a preset module of the deep copy Resnet50 rear part structure to obtain a first extracted feature, spatially downsample the first extracted features through maximum pooling, split the first extracted features into 2:1 in height dimension and 1:2 in height dimension to obtain all second extracted features, and spatially downsample all the second extracted features through maximum pooling to obtain a first downsampled feature, a second downsampled feature, a third downsampled feature, a fourth downsampled feature, and a fifth downsampled feature;
[0031] The second down-sampled feature and the fourth down-sampled feature are added to obtain a first added feature, the third down-sampled feature and the fifth down-sampled feature are added to obtain a second added feature, and the first down-sampled feature, the first added feature, and the second added feature are processed by a second Reduction module to obtain a first key output feature, a second key output feature, and a third key output feature, and the first key output feature, the second key output feature, and the third key output feature are determined to be the second output feature.
[0032] According to the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the input of the large-scale feature to the local feature extraction branch combined with the three-segmentation to obtain the third output feature includes:
[0033] Input the large-scale features into a preset module of the deep copy Resnet50 rear part structure to obtain a third extracted feature, spatially downsample the third extracted features through maximum pooling, split them according to a height dimension of 4:3:3 and a height dimension of 3:3:4 to obtain all fourth extracted features, and spatially downsample all the fourth extracted features through maximum pooling to obtain a sixth downsampled feature, a seventh downsampled feature, an eighth downsampled feature, a ninth downsampled feature, and a tenth downsampled feature;
[0034] The seventh down-sampled feature and the ninth down-sampled feature are added to obtain a third added feature, the eighth down-sampled feature and the tenth down-sampled feature are added to obtain a fourth added feature, and the seventh down-sampled feature, the third added feature, and the fourth added feature are processed using a third Reduction module to obtain a fourth key output feature, a fifth key output feature, and a sixth key output feature, and the fourth key output feature, and the fifth key output feature, and the sixth key output feature are determined to be the third output feature.
[0035] In the second aspect, a shoe print device based on feature distribution reconstruction and multi-granularity branch fusion is provided, comprising:
[0036] A first input unit, the first input unit is used to input the shoe print image to be processed into a preset Resnet50 front part structure to obtain a first feature, and input the first feature into the TCCA module, the DHSA module and the DIGS module to obtain a large-scale feature;
[0037] A second input unit is used to input the large-scale features to the preset Resnet50 back structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, and input the large-scale features, medium-scale features, and small-scale features to the MHFA module to obtain first output features;
[0038] a third input unit, the third input unit being configured to input the large-scale feature to the local feature extraction branch of the two-segmentation combination to obtain a second output feature, and to input the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain a third output feature;
[0039] A determination unit is used to determine a target feature based on the first output feature, the second output feature, and the third output feature, calculate the similarity between the target feature and the image feature of each image in a preset shoe print database, and determine the shoe type and brand corresponding to the image with the highest similarity as the target retrieval result of the shoe print image to be processed.
[0040] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion is implemented.
[0041] First, the present invention designs a ternary coordinate coupled enhanced attention (TCCA) module. During feature extraction, this module first applies a multi-directional modeling mechanism to enhance the interaction and transfer between features, effectively capturing subtle texture and shape differences in shoe print images. Subsequently, a coordinate attention layer is added to further introduce global spatial dependencies and optimize the ability to extract key features. Through this multi-level attention mechanism, the TCCA module significantly improves the model's ability to express complex features in shoe print data, allowing the detailed information of small shoe prints to be fully utilized, thereby improving the accuracy and robustness of the retrieval algorithm.
[0042] Secondly, by introducing the dynamic range histogram self-attention DHSA module, we achieve enhanced feature expression and dynamic information interaction by combining dynamic range sorting, feature reconstruction, and multi-directional attention mechanism;
[0043] Thirdly, the present invention designs a deep multi-scale gated spatial reconstruction DIGS module, which realizes the multi-scale comprehensive representation of features by efficiently processing and dynamically interacting with features of different scales, combining the alternating fusion of high-resolution spatial information and low-resolution semantic information;
[0044] Furthermore, the present invention uses a new feature aggregation module for the global feature fusion branch, namely the Multi-Scale Hierarchical Feature Aggregation Module (MHFA). By combining multi-expansion rate convolution, attention mechanism, and multi-resolution feature fusion strategy, it comprehensively improves the ability to capture features of different scales and complexity, and realizes efficient interaction and unified modeling of multi-resolution features, thereby significantly improving the feature extraction capability and robustness in complex scenes.
[0045] Finally, the present invention proposes a local feature extraction branch that combines two-segmentation and three-segmentation, which adopts a multi-scale region division and feature interaction strategy: this strategy divides and aligns the regions of shoe print features, which helps to restore the feature details of local regions while unifying the feature space, ensuring that the uniqueness of different regions is preserved and avoiding the mixing of local features and irrelevant information. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 This is one of the flow diagrams of the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention;
[0048] Figure 2 This is the second flow chart of the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention;
[0049] Figure 3 This is a schematic diagram of the overall structure of Resnet50;
[0050] Figure 4 It is the overall network schematic diagram of the present invention;
[0051] Figure 5 Schematic diagram of the ternary coordinate coupling enhanced attention module TCCA of the present invention;
[0052] Figure 6 Schematic diagram of the dynamic range histogram self-attention module DHSA of the present invention;
[0053] Figure 7 Schematic diagram of the deep multi-scale gated spatial reconstruction module DIGS of the present invention;
[0054] Figure 8 Schematic diagram of the multi-scale hierarchical feature aggregation module MHFA of the present invention;
[0055] Figure 9 Schematic diagram of the structure of a shoe printing device based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention;
[0056] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0057] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0058] Figure 1 This is one of the flow charts of the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, combined with Figure 2 The greatest contribution of this invention is that it proposes a shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion to achieve a more accurate and robust recognition and classification method for shoe print data. In addition, this invention proposes a new attention mechanism TCCA, which further improves the network's sensitivity to key features of shoe prints in complex backgrounds by enhancing multi-directional information interaction, and shows significant advantages in processing fine-grained shoe print data. In order to further enhance the feature extraction capability, this invention introduces a DHSA module into the branch network, which optimizes the distribution of features and enables the network to pay more attention to key areas in shoe prints. At the same time, this invention designs a new gating module DIGS, which realizes the joint modeling of multi-scale information through PixelShuffle and depthwise separable convolution, and combines high-resolution and low-resolution features to ensure that the network can capture both local texture information and global structural information. Finally, the MHFA feature fusion module we designed realizes the adaptive fusion of multi-scale features through multi-expansion rate convolution and SIMAM attention mechanism, effectively improving the expression capability of shoe print regional features.
[0059] like Figure 1 As shown, the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion includes:
[0060] Step 101: Input the shoe print image to be processed into the preset Resnet50 front-end structure to obtain the first feature, and input the first feature into the TCCA module, the DHSA module, and the DIGS module to obtain the large-scale feature;
[0061] Step 102: Input the large-scale features to the preset Resnet50 back-end structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, input the large-scale features, medium-scale features, and small-scale features to the MHFA module to obtain the first output features;
[0062] Step 103: input the large-scale feature to the local feature extraction branch of the two-segmentation combination to obtain a second output feature, and input the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain a third output feature;
[0063] Step 104: Determine a target feature based on the first output feature, the second output feature, and the third output feature, calculate similarity between the target feature and the image feature of each image in a preset shoe print database, and determine the shoe type and brand corresponding to the image with the highest similarity as the target retrieval result of the shoe print image to be processed.
[0064] Before step 101, the present invention can preprocess the image. The dataset of the present invention uses a series of image enhancement operations, including uniformly adjusting the input image to a fixed size of 384×144×3 to ensure feature scale consistency, and enhancing the diversity of the dataset through random horizontal flipping. Subsequently, the image is normalized, and the pixel values are normalized to the standard distribution range using the mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225] to improve the network's adaptability to different image features. The preprocessed image set is the set of shoe print images to be processed, which has higher feature consistency and generalization, and helps to improve the training effect of the model. The final processed image is input into the backbone network of the model for feature extraction.
[0065] In step 101, Figure 3 This is a schematic diagram of the overall structure of Resnet50. The default Resnet50 includes the front part structure of Resnet50 and the back part structure of Resnet50.
[0066] The front part of the Resnet50 structure includes a 7×7 convolution kernel, a second maximum pooling layer, a first residual block group consisting of three Bottleneck modules, a second residual block group consisting of four Bottleneck modules, and the first Bottleneck module in the third residual block group;
[0067] The latter part of the Resnet50 structure includes the second to sixth Bottleneck modules in the third residual block group, and a fourth residual block group consisting of three Bottleneck modules.
[0068] Optional, early feature extraction and distribution reconstruction: The preprocessed image is passed through the first three stages of Resnet50 to extract basic image features. These layers focus on capturing low-level and mid-level information in the shoe print image, and then further optimize the feature expression and distribution through TCCA, DHSA, and DIGS.
[0069] like Figure 3 As shown in the figure, the overall structure of Resnet50 is as follows: the initial convolution layer first preprocesses the input image, including a large 7×7 convolution kernel (stride=2, padding=3) for preliminary feature extraction, and outputs a 64-channel feature map, followed by batch normalization (BatchNorm) and ReLU activation function, and then further downsampled through a 3×3 maximum pooling layer (stride=2) to reduce the feature map size to 1 / 4 of the input. Four residual block groups (Layer1-Layer4) constitute the main part of ResNet50, each group consists of multiple Bottleneck (bottleneck structure) and residual blocks stacked together. The Bottleneck structure adopts a 1×1 convolution dimensionality reduction → 3×3 convolution feature extraction → 1×1 convolution dimensionality increase design, which greatly reduces the amount of computation.
[0070] Optionally, this paper uses the first three stages of ResNet50 as the self.backbone for feature extraction. The feature map input to self.backbone has a shape of [40, 3, 384, 144] (40 samples, 3-channel RGB image, resolution 384×144). From top to bottom, the process can be roughly divided into five stages, based on the number of feature channels, each connected by a certain downsampling operation.
[0071] Stage 1: Consists of a 7×7 convolutional layer with a 64-channel kernel and a stride of 2, using a ReLU activation function for nonlinear mapping, and further downsampling through a 3×3 max pooling layer (with a stride of 2). The output feature map has a shape of [40, 64, 96, 36].
[0072] Stage 2: Consists of ResNet layer 1 and includes three Bottleneck modules. Each Bottleneck module consists of 1×1, 3×3, and 1×1 convolutional layers, with the number of output channels gradually increasing to 256. This stage primarily extracts low-level local texture and edge information. The output feature map has a shape of [40, 256, 96, 36].
[0073] Stage 3: Consists of ResNet layer 2 and includes four Bottleneck modules. Each module is downsampled using convolutions with a stride of 2, gradually expanding the number of channels to 512. This stage extracts mid-level structural information and regional texture patterns. The output feature map has a shape of [40, 512, 48, 18].
[0074] Stage 4: This consists of the first Bottleneck module in ResNet layer 3 and includes a downsampling operation, expanding the number of channels to 1024, further compressing the spatial resolution of the feature map. This stage extracts high-level semantic features, such as the overall outline and regional distribution of the shoe print. The output feature map has a shape of [40, 1024, 24, 9], which is denoted as the first feature F1.
[0075] Optionally, the inputting the first feature into the TCCA module, the DHSA module, and the DIGS module to obtain large-scale features includes:
[0076] Input the first feature to the TCCA module, perform pooling and inverse transform operations on the first feature in the height direction, multiply it element-wise with the original feature map to obtain a second feature, perform adaptive average pooling on the first feature, obtain a third feature through a transposition operation, and add the second and third features to obtain a fourth feature;
[0077] In the width direction, the first feature is processed to obtain the fifth feature, feature map pooling is performed on the first feature to obtain the sixth feature, the fourth feature, the fifth feature, and the sixth feature are averaged to obtain the seventh feature, and the seventh feature is point-wise multiplied by the first feature to obtain the eighth feature output by the TCCA module;
[0078] The eighth feature is input to the DHSA module to obtain the ninth feature, and the ninth feature is input to the DIGS module to obtain the large-scale feature.
[0079] Figure 5 This is a schematic diagram of the ternary coordinate coupled enhanced attention module TCCA of the present invention. The present invention designs a TCCA attention mechanism to further process the first feature F1 obtained at the end:
[0080] F1 is pooled in the height direction to capture the dependencies between channels. The generated weight map is then inversely transformed and element-wise multiplied with the original feature map to obtain the second feature F2. F1 is adaptively averaged pooled in the height direction to generate a feature description in the width direction. The shape is adjusted through a transpose operation to obtain the third feature F3. F2 and F3 are added to obtain the fourth feature F4 in the height direction. Similarly, the fifth feature F5 in the width direction is obtained. Spatial attention is introduced to directly pool the feature map to capture the dependencies in the spatial dimension to obtain the sixth feature F6. The seven features F7 are averaged by adding the three features F4, F5, and F6. The seventh feature F7 is then multiplied point by point with the original feature map F1 to strengthen the feature expression of the key area, resulting in the eighth feature F8 output by TCCA.
[0081] Furthermore, after obtaining the eighth feature F8, the eighth feature is input into the DHSA module to obtain the ninth feature, and the ninth feature is input into the DIGS module to obtain the large-scale feature.
[0082] Optionally, inputting the eighth feature into the DHSA module to obtain a ninth feature includes:
[0083] The eighth feature is split into two channels, sorted along the height and width directions respectively, and the query, key and value are generated using the first 1×1 convolutional layer. The query and key are padded, and the feature map is rearranged using the Reshape_attn function. The similarity matrix is calculated to obtain the attention weight, and the attention weight is multiplied by the value to generate a context feature map.
[0084] The context feature map is processed using the output projection layer to obtain a processed feature map, the original order of the processed feature map is restored using the Torch.scatter operation, and the fused feature map is used to obtain the ninth feature using the second 1×1 convolutional layer.
[0085] Figure 6 This is a schematic diagram of the dynamic range histogram self-attention module DHSA of the present invention. Then, the DHSA module is used for the eighth feature F8: it aims to significantly improve the network's ability to capture and express key features in shoe print images through dynamic range sorting and multi-directional self-attention mechanism.
[0086] During initialization, the dynamic range histogram self-attention module (DHSA) defines several key parameters, including the feature dimension Dim, the number of attention heads Num_heads, and a learnable temperature parameter, Temperature, used to adjust the scaling of attention weights. The module contains three main convolutional layers: a 1×1 convolutional layer that expands the input feature dimension to Dim × 5, a depthwise separable convolutional layer Qkv_dwconv that further processes the expanded feature map while maintaining the number of channels, and a final 1×1 convolutional layer, Project_out, that projects the attention output back to the original feature dimension.
[0087] During the forward pass, the input feature map x is first sorted by half its channels, sorting them in both the height and width directions to enhance the structured representation of the features. Next, a 1×1 convolutional layer generates the query, key, and value (q, k, v), which are further processed by a depthwise separable convolutional layer. The module then pads the query and key to ensure that the feature map size is divisible by the number of attention heads. The feature map is reshaped into a shape suitable for multi-head attention computation using the Reshape_attn function. After reshaping, the query and key are normalized, and the similarity matrix Attn between them is calculated. This matrix is normalized using a custom Softmax function to obtain the attention weights. The attention weights are then multiplied with the value features to generate a contextual feature map, which is further processed by the output projection layer Project_out. The module restores the original order of the feature maps using the Torch.scatter operation and multiplies the two attention outputs. The fused feature map is then projected back to its original number of channels using a 1×1 convolutional layer to produce the ninth feature F9. The entire process, through dense connections and a multi-directional attention mechanism, ensures the effective transfer and enhancement of features at different scales and directions, alleviating the vanishing gradient problem and improving the model's ability to represent both fine-grained and global features. Ultimately, the DHSA module outputs an optimized feature map, significantly enhancing the network's recognition and classification performance for shoeprint images in complex scenarios while maintaining high computational efficiency and a lightweight model.
[0088] Figure 7 Schematic diagram of the deep multi-scale gated spatial reconstruction module DIGS of the present invention, wherein the ninth feature is input to the DIGS module to obtain the large-scale feature, including:
[0089] After inputting the ninth feature into the third 1×1 convolutional layer, upsampling and blocking operations are performed to obtain the first branch feature and the second branch feature;
[0090] After performing the Split operation on the first branch feature, the features are respectively input into the first 3×3 convolutional layer, the 1×11 convolutional layer, the 11×1 convolutional layer, and the Identity module to obtain the first output result, the second output result, the third output result, and the fourth output result. The first output result, the second output result, the third output result, and the fourth output result are merged to obtain the first branch output;
[0091] After inputting the second branch feature into the second 3×3 convolutional layer, the Mish activation function is used to obtain the second branch output;
[0092] After performing a dot multiplication operation on the first branch output and the second branch output, downsampling is performed to obtain downsampled features, and the downsampled features are input to the fourth 1×1 convolutional layer to obtain the large-scale features.
[0093] Optionally, the ninth feature F9 is processed using DIGS. The DIGS module is a key feature enhancement component in this invention. It aims to further enhance the network's ability to represent both fine-grained and global features in shoeprint images through multi-scale information processing and dynamic feature fusion. This module first receives the feature map F9 output by the DHSA module.
[0094] The DIGS module consists of two main branches: a high-scale branch and a low-scale branch. The high-scale branch uses a series of two-scale convolution operations to capture multi-level feature information using receptive fields of different scales, ensuring that the network can simultaneously focus on local details and global structures; the low-scale branch uses a gating mechanism to dynamically adjust the feature flow to enhance attention to key feature areas and suppress interference from irrelevant or redundant information. Within each scale branch, the module designs multiple convolutional layers and gating units. These convolutional layers perform feature dimensionality reduction and expansion through depthwise separable convolution and 1×1 convolution. The gating unit dynamically adjusts the importance of features through learned weights, thereby achieving adaptive fusion of features and obtaining the large-scale feature F. 10 .
[0095] In step 102, it is necessary to realize global feature fusion, and the large-scale feature F 10 Processing: First, pass through the Res_conv4 module. The Res_conv4 module consists of all subsequent Bottleneck modules in the third stage (layer3) of ResNet50 except the first Bottleneck module. That is, input the large-scale features to the preset Resnet50 back structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, input the large-scale features, medium-scale features and small-scale features to the MHFA module to obtain the first output features.
[0096] Specifically, Resnet.layer3 usually contains 6 Bottleneck modules. Layer3[1:] extracts the second to sixth modules. Each Bottleneck module contains 1×1, 3×3 and 1×1 convolution layers, which are used for feature compression, spatial feature extraction and channel expansion respectively. Then, through the Res_g_conv5 module, the Res_g_conv5 module directly references the fourth stage (layer4) of ResNet50, which consists of multiple Bottleneck modules. It further extracts and integrates high-level semantic features to obtain a medium-scale feature G1 with a shape of [40, 2048, 12, 5]. The medium-scale feature G1 is then subjected to maximum pooling to obtain a small-scale feature G2 with a shape of [40, 2048, 1, 1]. The large-scale feature F 10 The first output feature G3 is obtained by performing feature aggregation with the medium-scale feature G1 and the small-scale feature G2 through the MHFA module.
[0097] Figure 8 Schematic diagram of the multi-scale hierarchical feature aggregation module MHFA of the present invention, wherein the large-scale features, medium-scale features, and small-scale features are input to the MHFA module to obtain the first output features, including:
[0098] The large-scale features are processed using downsampling and a fifth 1×1 convolutional layer to obtain a first processed feature, the medium-scale features are processed using a sixth 1×1 convolutional layer to obtain a second processed feature, the first processed features and the second processed features are added element-by-element to obtain a third processed feature, multi-scale feature extraction is performed on the third processed features using four 3×3 convolutional layers with different expansion rates, the third processed features are connected in series, and the third processed features are fused sequentially through a seventh 1×1 convolutional layer and a batch normalization layer to obtain a fourth processed feature;
[0099] Performing element-by-element addition on the fourth processed feature, the first processed feature, the second processed feature, and the attention-weighted feature map to obtain a fifth processed feature, processing the fifth processed feature using an adaptive average pooling layer to obtain a sixth processed feature, adding the sixth processed feature to the small-scale feature to obtain a seventh processed feature, performing self-attention processing on the seventh processed feature using two SelfAttentionBlock modules to obtain an eighth processed feature, performing element-by-element multiplication and addition operations on the eighth processed feature and the small-scale feature, and processing the result through an eighth 1×1 convolutional layer to obtain a ninth processed feature, processing the ninth processed feature using a first Reduction module to obtain the first output feature.
[0100] Optional, MHFA accepts F 10, G1 and G2 three features, first, F 10 After downsampling and 1×1 convolution layer processing, its spatial size and number of channels are adjusted to match the feature map size and number of channels of G1, ensuring that the subsequent element-by-element addition operation can be carried out smoothly. At the same time, G1 reduces the number of channels through 1×1 convolution layer to further optimize the feature expression. Next, after downsampling and channel number adjustment, F 10 _down and G1_reduced are element-wise added to form the fused feature map fused_features. This process effectively integrates key information from features at different scales. Fused_features then undergoes multi-scale feature extraction through four 3×3 convolutional layers (conv_1 to conv_4) with different dilation rates. The dilation rates of each convolutional layer are 1, 2, 3, and 4, respectively, to capture feature information across different receptive fields. The outputs of these multi-dilation rate convolutional layers are concatenated and fused through a 1×1 convolutional layer and a batch normalization (BN) layer to generate a comprehensive multi-scale feature representation, Fuse.
[0101] On this basis, the module introduces three SimamModule attention mechanisms, which are used to preprocess F 10 _down, G1_reduced and the fused Fuse feature map improve the expression ability of key feature areas through dynamic weighting. SimamModule effectively enhances the response of the salient area of the feature map by calculating the local attention weight of the feature map, while suppressing the interference of background noise. 10 _down, G1_reduced and the attention-weighted feature maps are added element-by-element to form the final fused feature Combined. The feature map is compressed to 1×1 through an adaptive average pooling layer to generate the global context feature Fused_MHFA.
[0102] Based on multi-scale feature fusion, Fused_MHFA is added to G2 to further integrate high- and low-order features, forming a comprehensive global feature representation. The fused features are then processed through two SelfAttentionBlock modules (Attention_sm and Attention_final). These modules use a multi-head attention mechanism to capture long-range dependencies in the feature map and optimize the representation of global features. Finally, the self-attentioned feature map is element-wise multiplied and added with G2 to generate the enhanced feature map Added_result. This feature map then passes through a final 1×1 convolutional layer to maintain the same number of channels, resulting in the first output feature G3. A Reduction module is applied to the first output feature G3 to generate feature G4, which is then processed through a fully connected layer Fc_id_2048_0 to produce the high-dimensional classification output G5.
[0103] In step 103, the input of the large-scale feature to the local feature extraction branch of the binary segmentation combination to obtain the second output feature includes:
[0104] Input the large-scale feature to the preset module of the deep copy Resnet50 back structure to obtain the first extracted feature H1, perform spatial downsampling on the first extracted feature H1 through maximum pooling, and obtain all second extracted features H2~H5 according to the height dimension 2:1 and the height dimension 1:2 split, and perform spatial downsampling on all second extracted features H2~H5 through maximum pooling to obtain the first down-sampled feature H6, the second down-sampled feature H7, the third down-sampled feature H8, the fourth down-sampled feature H9 and the fifth down-sampled feature H 10 ;
[0105] The second down-sampled feature H7 and the fourth down-sampled feature H9 are added to obtain the first added feature J1, and the third down-sampled feature H8 and the fifth down-sampled feature H 10 Add them to obtain the second added feature J2, and use the second Reduction module to process the first down-sampled feature H6, the first added feature J1 and the second added feature J2 to obtain the first key output feature J3, the second key output feature J4 and the third key output feature J5, and determine that the first key output feature, the second key output feature and the third key output feature are the second output feature.
[0106] Optional, two-segment combined local feature extraction branch: for F 10 Processing: By deeply copying the subsequent modules Res_conv4 and Res_p_conv5 of the backbone network, an independent feature extraction path is formed. This branch first uses the Res_conv4 module to further process F 10Multiple Bottleneck modules are used to extract higher-level semantic features, enhancing the depth of feature expression. Subsequently, the Res_p_conv5 module further integrates and extracts high-level semantic information. Three Bottleneck modules are used to expand the number of channels to 2048, and the output feature map H1 has a shape of [40, 2048, 12, 5]. These feature maps contain the complex patterns and regional distribution information of the shoe print image.
[0107] Next, H1 is spatially downsampled by the maximum pooling, and the features H2~H5 are obtained by splitting the height dimension 2:1 and the height dimension 1:2. H1~H5 are then spatially downsampled by the maximum pooling to obtain H6~H 10 , add H7 and H9 to get J1, H8 and H 10 Add them together to get J2. H6, J1, and J2 are passed through the Reduction module to reduce the number of channels from 2048 to 256 to reduce the computational complexity and retain the key information of the features to obtain features J3 to J5. Then, the fully connected layer Fc_id_2048_1 is used on J3 to generate the high-dimensional classification output J6. The fully connected layers Fc_id_256_1_0 and Fc_id_256_1_1 are used on J3 and J4 for classification prediction to obtain J7 and J8.
[0108] Optionally, inputting the large-scale feature into the local feature extraction branch of the three-segmentation combination to obtain a third output feature includes:
[0109] Input the large-scale features into a preset module of the deep copy Resnet50 rear part structure to obtain a third extracted feature, spatially downsample the third extracted features through maximum pooling, split them according to a height dimension of 4:3:3 and a height dimension of 3:3:4 to obtain all fourth extracted features, and spatially downsample all the fourth extracted features through maximum pooling to obtain a sixth downsampled feature, a seventh downsampled feature, an eighth downsampled feature, a ninth downsampled feature, and a tenth downsampled feature;
[0110] The seventh down-sampled feature and the ninth down-sampled feature are added to obtain a third added feature, the eighth down-sampled feature and the tenth down-sampled feature are added to obtain a fourth added feature, and the seventh down-sampled feature, the third added feature, and the fourth added feature are processed using a third Reduction module to obtain a fourth key output feature, a fifth key output feature, and a sixth key output feature, and the fourth key output feature, and the fifth key output feature, and the sixth key output feature are determined to be the third output feature.
[0111] Optional, three-segmentation combined local feature extraction branch: The overall structure is similar to the two-segmentation combined local feature extraction branch. The main difference is that the segmentation is performed according to the height dimension 4:3:3 and the height dimension 3:3:4. The features M1~M4 are obtained through the Reduction module, and M1 is then passed through the fully connected layer Fc_id_2048_2 to generate the high-dimensional classification output M5. The fully connected layers Fc_id_256_2_0, Fc_id_256_2_1 and Fc_id_256_2_2 are used to perform classification prediction on M2~M4 to obtain M6~M8.
[0112] In step 104, feature G4, features J3~J5 and features M1~M4, a total of 8 features, are concatenated using Concat to obtain feature Pre, that is, the target feature. The entire model returns the target feature Pre, G4, J3, M1, G5, J6, M5, J7, J8, M6, M7, M8, a total of 12 features, among which G4, J3, and M1 are used to calculate the ternary loss; G5, J6, M5, J7, J8, M6, M7, and M8 are used to calculate the cross entropy loss; and feature Pre is used to compare with all shoe print features in the entire shoe print database.
[0113] Furthermore, triplet loss is used to shorten the distance between similar samples and expand the distance between heterogeneous samples, thereby learning more discriminative feature representations and achieving more accurate shoe print retrieval. The Triplet Loss formula is as follows:
[0114]
[0115] The final plus sign indicates taking the positive part of the expression in the brackets. When the value in the brackets is less than or equal to 0, the loss is 0. 𝑁 is the number of samples in the training set. 𝑎, 𝑝, and 𝑛 represent anchor samples, positive samples, and negative samples, respectively. Here, positive samples refer to samples of the same pair of shoes as the anchor samples, and negative samples refer to samples of different shoes than the anchor samples. 𝑓 represents the Embedding model. 、 and Denote the anchor point, positive sample, and negative sample of the 𝑖th sample, respectively. This loss function focuses on calculating the maximum distance between the anchor sample and the positive sample, and the minimum distance between the anchor sample and the negative sample, that is, focusing on samples that are difficult to distinguish.
[0116] Optionally, the cross entropy loss is minimized to reduce the gap between the model's predicted category and the true category, thereby improving the accuracy of shoe print retrieval. The cross entropy loss calculation formula is as follows:
[0117]
[0118] Among them, 𝐶 represents the number of categories, 𝑁 represents the number of training samples, Representation sample The label belonging to the cth category, Indicates that the model is input In the case of , the probability of the 𝑐th category is predicted, and 𝜃 represents the model parameters. The features Pre obtained by the model are used to calculate the similarity with the features of all shoe prints in the entire shoe print database. The images with the highest similarity are sorted and selected as the retrieval results.
[0119] Figure 4 The overall network diagram of the present invention is as follows:
[0120] Shoeprint image preprocessing: The shoeprint image to be identified is converted into a grayscale image. Background separation techniques are used to remove irrelevant information, and data augmentation strategies are applied to improve the model's generalization capabilities. Specifically, the grayscale image is first adaptively scaled and cropped to a fixed size of 384×144 to ensure image scale consistency and facilitate more efficient feature extraction. The preprocessed shoeprint image is then used as input for further fine-grained feature extraction and modeling, thereby improving the accuracy and robustness of shoeprint recognition.
[0121] Early feature extraction and distribution reconstruction: The pre-processed shoe print image is passed through the first three layers of Resnet50, and only the first part is used in the third layer. The TCCA module is then used to enhance the global and detailed expression capabilities of the features through multi-directional modeling and coordinate attention mechanism, thereby optimizing the overall feature extraction of the shoe print. The DHSA module is then used to focus on the key areas of the feature map to enhance the relevance and importance of the features. The DIGS module is then used to further enhance the comprehensive representation capabilities of the local details and global dependencies of the shoe print through multi-scale modeling and dynamic feature fusion. In addition, in order to improve the network's fluidity and expression consistency of features, LayerNorm normalization operations and residual connections are added before DHSA and DIGS to ensure the stability of information transmission and the integrity of feature expression. After all the above layers, the early features are obtained; global feature fusion: the output early features are passed through the last five bottles of the third layer of Resnet50, and then through the fourth layer of Resnet50 to obtain large-scale features, and then through global maximum pooling to obtain medium-scale features, and then through global maximum pooling again to obtain small-scale features. The large-scale features, medium-scale features, and small-scale features are input into the MHFA feature aggregation module to obtain fused features. Finally, the fused features are passed through the fully connected layer to obtain global context features; local feature extraction: the output early features are segmented according to a specific ratio. The second segmentation branch finally obtains the global feature E1 and local features E2 and E3; the third segmentation branch obtains the global feature E4 and local features E5, E6, and E7; shoe print retrieval: the obtained global context features are further spliced with the obtained features E1 to E7 to obtain the final feature Pre, and the final feature Pre is compared with the shoe print database features.
[0122] The present invention is divided into two parts as a whole: early feature extraction and distribution reconstruction, and late feature extraction and multi-granularity branch fusion. The early part uses ResNet50 as the backbone network. Through its hierarchical convolutional structure, it gradually extracts features from low-level to high-level, which helps to model the local and global characteristics of shoe prints. In particular, the early layers can extract local features such as edges, and the subsequent layers capture global features. Subsequently, a newly designed ternary coordinate coupled enhanced attention module TCCA is introduced. This module uses a multi-level attention mechanism to comprehensively improve the expressive power of feature maps from different dimensions. In view of the diversity and complexity of features, it combines multi-directional and global information to more efficiently model and enhance feature maps.
[0123] In addition, the method introduces a dynamic range self-attention mechanism module (DHSA) for efficient processing of input feature maps, optimizing their expressiveness across multiple dimensions. This is particularly useful for scenarios with high-dimensional, complex structures and fine-grained features. Next, a deep multi-scale gated spatial reconstruction module (DIGS) is designed. This module, through multi-scale feature processing and dynamic feature fusion, enhances the network's expressiveness in high-dimensional data and its ability to model complex structures. The second part consists of a global feature fusion branch, a local feature extraction branch combining two-segmentation, and a local feature extraction branch combining three-segmentation. Within the global feature fusion branch, a newly designed multi-scale hierarchical feature aggregation module (MHFA) is used. This module combines multi-scale feature fusion, multi-dilation rate convolution, and a multi-head attention mechanism to efficiently fuse and enhance features at different levels, improving the network's representation of complex objects. The local feature extraction branches combining two-segmentation and three-segmentation fully exploit feature information through fine-grained segmentation and feature interaction, improving the representation of shoe print details and overall patterns. This method not only significantly enhances the ability to characterize fine-grained regional differences in shoe prints through multi-branch feature extraction and fusion mechanisms, but also utilizes multi-scale feature modeling and dynamic interaction strategies to improve the algorithm's adaptability to diverse shape and texture features in complex scenarios, thereby ensuring more stable and accurate retrieval results.
[0124] Figure 9 : is a structural diagram of a shoe print device based on feature distribution reconstruction and multi-granularity branch fusion provided by the present invention, the shoe print device based on feature distribution reconstruction and multi-granularity branch fusion comprises:
[0125] A first input unit 1 is used to input the shoe print image to be processed into a preset Resnet50 front-end structure to obtain a first feature, and input the first feature into a TCCA module, a DHSA module, and a DIGS module to obtain a large-scale feature;
[0126] A second input unit 2 is used to input the large-scale features to the preset Resnet50 back structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, and input the large-scale features, medium-scale features, and small-scale features to the MHFA module to obtain the first output features;
[0127] A third input unit 3 is used to input the large-scale feature to the local feature extraction branch of the two-segmentation combination to obtain a second output feature, and input the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain a third output feature;
[0128] Determination unit 4, the determination unit 4 is used to determine the target feature based on the first output feature, the second output feature and the third output feature, calculate the similarity between the target feature and the image feature of each image in the preset shoe print database, and determine the shoe type and brand corresponding to the image with the highest similarity as the target retrieval result of the shoe print image to be processed.
[0129] Figure 10 is a structural diagram of the electronic device provided by the present invention, Figure 10 Schematic diagram of the structure of the electronic device provided by the present invention. Figure 10 As shown, the electronic device may include: a processor 110 , a communications interface 120 , a memory 130 and a communication bus 140 , wherein the processor 110 , the communications interface 120 and the memory 130 communicate with each other via the communication bus 140 . The processor 110 can call the logic instructions in the memory 130 to execute a shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion, the method including: inputting the shoe print image to be processed into the preset Resnet50 front part structure to obtain a first feature, inputting the first feature into the TCCA module, the DHSA module and the DIGS module to obtain a large-scale feature; inputting the large-scale feature into the preset Resnet50 back part structure to obtain a medium-scale feature, inputting the medium-scale feature into the first maximum pooling layer to obtain a small-scale feature, inputting the large-scale feature, the medium-scale feature and the small-scale feature into the MHFA module to obtain a first output feature; inputting the large-scale feature into the local feature extraction branch of the two-segmentation combination to obtain a second output feature, inputting the large-scale feature into the local feature extraction branch of the three-segmentation combination to obtain a third output feature; determining the target feature according to the first output feature, the second output feature and the third output feature, performing similarity calculation between the target feature and the image feature of each image in the preset shoe print database, and determining the shoe type and brand corresponding to the image with the highest similarity as the target retrieval result of the shoe print image to be processed.
[0130] In addition, the logic instructions in the aforementioned memory 130 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0132] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion, characterized in that: include: Input the shoe print image to be processed into the preset Resnet50 front-end structure to obtain the first feature, and input the first feature into the TCCA module, DHSA module, and DIGS module to obtain large-scale features; Input the large-scale features to the preset Resnet50 back structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, input the large-scale features, medium-scale features and small-scale features to the MHFA module to obtain the first output features; Input the large-scale feature to the local feature extraction branch of the two-segmentation combination to obtain a second output feature, and input the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain a third output feature; Determining a target feature based on the first output feature, the second output feature, and the third output feature, calculating similarity between the target feature and the image feature of each image in a preset shoe print database, and determining the shoe type and brand corresponding to the image with the highest similarity as the target retrieval result for the shoe print image to be processed; The input of the first feature to the TCCA module, the DHSA module and the DIGS module to obtain large-scale features includes: Input the first feature to the TCCA module, perform pooling and inverse transform operations on the first feature in the height direction, multiply it element-wise with the original feature map to obtain a second feature, perform adaptive average pooling on the first feature, obtain a third feature through a transposition operation, and add the second and third features to obtain a fourth feature; In the width direction, the first feature is processed to obtain a fifth feature, feature map pooling is performed on the first feature to obtain a sixth feature, the fourth feature, the fifth feature, and the sixth feature are averaged to obtain a seventh feature, and the seventh feature is point-wise multiplied by the first feature to obtain an eighth feature output by the TCCA module; The eighth feature is input to the DHSA module to obtain the ninth feature, and the ninth feature is input to the DIGS module to obtain the large-scale feature.
2. The shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion according to claim 1 is characterized in that: The preset Resnet50 includes the front part structure of Resnet50 and the back part structure of Resnet50; The front part of the Resnet50 structure includes a 7×7 convolution kernel, a second maximum pooling layer, a first residual block group consisting of three Bottleneck modules, a second residual block group consisting of four Bottleneck modules, and the first Bottleneck module in the third residual block group; The latter part of the Resnet50 structure includes the second to sixth Bottleneck modules in the third residual block group, and a fourth residual block group consisting of three Bottleneck modules.
3. The shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion according to claim 1 is characterized in that: The inputting the eighth feature into the DHSA module to obtain the ninth feature includes: The eighth feature is split into two channels, sorted along the height and width directions respectively, and the query, key and value are generated using the first 1×1 convolutional layer. The query and key are padded, and the feature map is rearranged using the Reshape_attn function. The similarity matrix is calculated to obtain the attention weight, and the attention weight is multiplied by the value to generate a context feature map. The context feature map is processed using the output projection layer to obtain a processed feature map, the original order of the processed feature map is restored using the Torch.scatter operation, and the fused feature map is used to obtain the ninth feature using the second 1×1 convolutional layer.
4. The shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion according to claim 1 is characterized in that: Inputting the ninth feature into the DIGS module to obtain the large-scale feature includes: After inputting the ninth feature into the third 1×1 convolutional layer, upsampling and blocking operations are performed to obtain the first branch feature and the second branch feature; After performing the Split operation on the first branch feature, the features are respectively input into the first 3×3 convolutional layer, the 1×11 convolutional layer, the 11×1 convolutional layer, and the Identity module to obtain the first output result, the second output result, the third output result, and the fourth output result. The first output result, the second output result, the third output result, and the fourth output result are merged to obtain the first branch output; After inputting the second branch feature into the second 3×3 convolutional layer, the Mish activation function is used to obtain the second branch output; After performing a dot multiplication operation on the first branch output and the second branch output, downsampling is performed to obtain downsampled features, and the downsampled features are input to the fourth 1×1 convolutional layer to obtain the large-scale features.
5. The shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion according to claim 1 is characterized in that: The inputting the large-scale features, the medium-scale features, and the small-scale features into the MHFA module to obtain the first output features includes: The large-scale features are processed using downsampling and a fifth 1×1 convolutional layer to obtain a first processed feature, the medium-scale features are processed using a sixth 1×1 convolutional layer to obtain a second processed feature, the first processed features and the second processed features are added element-by-element to obtain a third processed feature, multi-scale feature extraction is performed on the third processed features using four 3×3 convolutional layers with different expansion rates, the third processed features are connected in series, and the third processed features are fused sequentially through a seventh 1×1 convolutional layer and a batch normalization layer to obtain a fourth processed feature; Performing element-by-element addition on the fourth processed feature, the first processed feature, the second processed feature, and the attention-weighted feature map to obtain a fifth processed feature, processing the fifth processed feature using an adaptive average pooling layer to obtain a sixth processed feature, adding the sixth processed feature to the small-scale feature to obtain a seventh processed feature, performing self-attention processing on the seventh processed feature using two SelfAttentionBlock modules to obtain an eighth processed feature, performing element-by-element multiplication and addition operations on the eighth processed feature and the small-scale feature, and processing the result through an eighth 1×1 convolutional layer to obtain a ninth processed feature, processing the ninth processed feature using the first Reduction module to obtain the first output feature.
6. The shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion according to claim 1 is characterized in that: The inputting the large-scale feature to the local feature extraction branch combined with the two-segmentation to obtain the second output feature includes: Input the large-scale features into a preset module of the deep copy Resnet50 rear part structure to obtain a first extracted feature, spatially downsample the first extracted features through maximum pooling, split the first extracted features into 2:1 in height dimension and 1:2 in height dimension to obtain all second extracted features, and spatially downsample all the second extracted features through maximum pooling to obtain a first downsampled feature, a second downsampled feature, a third downsampled feature, a fourth downsampled feature, and a fifth downsampled feature; The second down-sampled feature and the fourth down-sampled feature are added to obtain a first added feature, the third down-sampled feature and the fifth down-sampled feature are added to obtain a second added feature, and the first down-sampled feature, the first added feature, and the second added feature are processed by a second Reduction module to obtain a first key output feature, a second key output feature, and a third key output feature, and the first key output feature, the second key output feature, and the third key output feature are determined to be the second output feature.
7. The shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion according to claim 1 is characterized in that: The inputting the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain the third output feature includes: Input the large-scale features into a preset module of the deep copy Resnet50 rear part structure to obtain a third extracted feature, spatially downsample the third extracted features through maximum pooling, split them according to a height dimension of 4:3:3 and a height dimension of 3:3:4 to obtain all fourth extracted features, and spatially downsample all the fourth extracted features through maximum pooling to obtain a sixth downsampled feature, a seventh downsampled feature, an eighth downsampled feature, a ninth downsampled feature, and a tenth downsampled feature; The seventh down-sampled feature and the ninth down-sampled feature are added to obtain a third added feature, the eighth down-sampled feature and the tenth down-sampled feature are added to obtain a fourth added feature, and the seventh down-sampled feature, the third added feature, and the fourth added feature are processed using a third Reduction module to obtain a fourth key output feature, a fifth key output feature, and a sixth key output feature, and the fourth key output feature, and the fifth key output feature, and the sixth key output feature are determined to be the third output feature.
8. A shoe print device based on feature distribution reconstruction and multi-granularity branch fusion, characterized in that: include: A first input unit, the first input unit is used to input the shoe print image to be processed into a preset Resnet50 front part structure to obtain a first feature, and input the first feature into the TCCA module, the DHSA module and the DIGS module to obtain a large-scale feature; A second input unit is used to input the large-scale features to the preset Resnet50 back structure to obtain medium-scale features, input the medium-scale features to the first maximum pooling layer to obtain small-scale features, and input the large-scale features, medium-scale features, and small-scale features to the MHFA module to obtain first output features; a third input unit, the third input unit being configured to input the large-scale feature to the local feature extraction branch of the two-segmentation combination to obtain a second output feature, and to input the large-scale feature to the local feature extraction branch of the three-segmentation combination to obtain a third output feature; a determination unit configured to determine a target feature based on the first output feature, the second output feature, and the third output feature, calculate a similarity between the target feature and an image feature of each image in a preset shoe print database, and determine the shoe type and brand corresponding to the image with the highest similarity as a target retrieval result for the shoe print image to be processed; The input of the first feature to the TCCA module, the DHSA module and the DIGS module to obtain large-scale features includes: Input the first feature to the TCCA module, perform pooling and inverse transform operations on the first feature in the height direction, multiply it element-wise with the original feature map to obtain a second feature, perform adaptive average pooling on the first feature, obtain a third feature through a transposition operation, and add the second and third features to obtain a fourth feature; In the width direction, the first feature is processed to obtain a fifth feature, feature map pooling is performed on the first feature to obtain a sixth feature, the fourth feature, the fifth feature, and the sixth feature are averaged to obtain a seventh feature, and the seventh feature is point-wise multiplied by the first feature to obtain an eighth feature output by the TCCA module; The eighth feature is input to the DHSA module to obtain the ninth feature, and the ninth feature is input to the DIGS module to obtain the large-scale feature.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the shoe print retrieval method based on feature distribution reconstruction and multi-granularity branch fusion as described in any one of claims 1 to 7 is implemented.