Generalized few-sample target detection method based on differential multi-scale fusion
By proposing a differential convolution multi-scale feature fusion framework in radar small sample target detection, the problems of feature balance and object boundary modeling in radar small sample target detection are solved, and the high-precision and robust target detection effect is achieved.
Patent Information
- Application Number
- CN202510360244.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-17
AI Technical Summary
The existing radar small sample target detection method is difficult to adaptively dynamically balance and fuse local-global features extracted from radar echo data by convolution kernels of different scales. It is also difficult to accurately model the feature differences in the object boundary area at the pixel level during detection, while fully retaining the global information of the object.
A generalized small sample object detection method based on differential multi-scale fusion is proposed. By constructing a differential convolution multi-scale feature fusion framework, including a multi-scale feature fusion module and a differential convolution fusion module. The multi-scale feature fusion module uses five different scale convolution kernels to extract features and perform adaptive weighted fusion through the scale attention mechanism. The differential convolution fusion module fuses with the original feature map by establishing the direction sensitive differential feature map of the pixel neighborhood.
It realizes the precise positioning of object boundaries and the effective retention of object global information, and at the same time, it realizes the optimal balance extraction of local details and global semantic information, improving the detection accuracy and robustness of radar small sample target detection.
Smart Images

Figure CN120164072A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of object detection, and particularly to a generalized few-shot object detection method based on differential multi-scale fusion. Background Art
[0002] In the field of radar, few-shot object detection is a key technology dedicated to accurately identifying new-class objects through extremely few samples, providing an important solution idea for the problem of radar object detection under data-scarce conditions. Compared with general radar object detection, few-shot radar object detection faces two prominent challenges. First, how to learn features that can effectively represent different-class objects when the number of samples is extremely limited, which is crucial for the radar to accurately identify object classes; second, how to achieve a balance in detection accuracy between newly emerging object classes and the basic classes already familiar to the radar, avoiding overemphasis on new classes while neglecting basic classes, or vice versa.
[0003] Most existing radar object detection methods rely on fine-tuning of basic classes, but this approach is extremely likely to cause a large gap in detection performance between basic classes and new classes. In recent years, pre-trained models based on Vision Transformers (ViT) have emerged in the radar field and shown great application potential due to their powerful feature expression capabilities. For example, DE-ViT has made certain breakthroughs in using ViT features to process radar data, and CD-ViTO has further in-depth research based on DE-ViT, strongly promoting the development of radar cross-domain few-shot detection. However, there are still two key problems to be solved urgently in the application of these methods in radar. First, how to adaptively and dynamically balance and fuse local-global features extracted from radar echo data by different-scale convolutional kernels; second, when using radar to detect objects, how to accurately model the feature differences in the object boundary region at the pixel level while completely retaining the global information of the object. Summary of the Invention
[0004] Based on this, it is necessary to provide a generalized few-shot object detection method based on differential multi-scale fusion that can achieve precise positioning of the object boundary and effective retention of the object's global information while achieving the best balance extraction of local details and global semantic information for the above technical problems.
[0005] A generalized few-shot object detection method based on differential multi-scale fusion, the method includes:
[0006] Construct a few-shot object detection model; the few-shot object detection model includes a Region Proposal Network and a differential convolutional multi-scale feature fusion framework; the differential convolutional multi-scale feature fusion framework includes a multi-scale feature fusion module and a differential convolutional fusion module;
[0007] Generate few-shot initial region proposals according to the Region Proposal Network; extract the corresponding masked regions from the initial region proposals and optimize them by scale expansion to obtain the target regions.
[0008] Input the target regions into the multi-scale feature fusion module. By using five different-scale convolutional kernels to extract the features of the target regions, and using the scale attention mechanism to generate attention weights for adaptive weighted fusion, the fused multi-scale features are obtained.
[0009] Input the fused multi-scale features into the differential convolution fusion module. By establishing the directional sensitive differential feature map of the pixel neighborhood and fusing it with the original feature map, the final target features are obtained.
[0010] For the above generalized few-shot object detection method based on differential multi-scale fusion, the present application proposes a differential convolution multi-scale feature fusion framework, including two core designs: designing a multi-scale feature fusion module to adaptively weighted fuse the features extracted by five different-scale convolutional kernels using the scale attention mechanism to achieve the best balance extraction of local details and global semantic information; designing a differential convolution fusion module by establishing the directional sensitive differential feature map of the pixel neighborhood and fusing it with the original feature map. Description of the Drawings
[0011] Figure 1 It is a schematic flowchart of a generalized few-shot object detection method based on differential multi-scale fusion in an embodiment.
[0012] Figure 2 It is a framework diagram of a few-shot object detection model in an embodiment.
[0013] Figure 3 It is a schematic diagram of three pixel differential convolution instances, the pixel pair selection and convolution process of three differential convolutions in an embodiment. Detailed Embodiments
[0014] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application.
[0015] In one embodiment, as Figure 1 shown, a generalized few-shot object detection method based on differential multi-scale fusion is provided, including the following steps:
[0016] Step 102, construct a few-shot object detection model; the few-shot object detection model includes a Region Proposal Network and a differential convolution multi-scale feature fusion framework; the differential convolution multi-scale feature fusion framework includes a multi-scale feature fusion module and a differential convolution fusion module.
[0017] The overall architecture of the differential convolution multi-scale feature fusion framework (DiMiNet) proposed in this application is as shown in Figure 2 (a). DiMiNet realizes the detection of novel categories through efficient region proposal, multi-scale feature extraction, and the fusion of differential convolution and standard convolution. This framework is based on a category-agnostic region proposal network (RPN), and then the proposals are optimized through a differential multi-scale fusion layer. Two core innovations drive the performance improvement: (1) the differential convolution fusion module (DCFM) ( Figure 2 (b), green), which captures the fusion of pixel neighborhood contrast information and global object information; (2) the multi-scale feature fusion module (MFFM) ( Figure 2 (c), blue), which adaptively balances and fuses global and local representations. This process first extracts and scales the masked regions of the proposals generated by the RPN ( Figure 2 (a), yellow bounding box regions) for optimization. These regions are sequentially passed through three differential multi-scale fusion layers ( Figure 2 (a) the dashed boxes, combining MFFM and DCFM) for progressive feature optimization. The final output is converted into bounding boxes through a region-to-bounding box layer, and the class scores are mapped through spatial attention using the width-height average feature. The generalization of novel categories is achieved through the cosine similarity projection subspace (CSPS) between the ViT query features and the class prototypes ( Figure 2 (a), brown cuboid).
[0018] Step 104: Generate few-shot initial region proposals according to the region proposal network; extract the corresponding masked regions from the initial region proposals and perform scale expansion optimization to obtain the target regions.
[0019] As the starting link of the entire few-shot object detection model, the region proposal network can generate initial region proposals based on few-shot data, and accurately determine the target regions through the extraction and scale expansion optimization of the masked regions, providing a clear target range for subsequent feature extraction and fusion, enabling the multi-scale feature fusion module and the differential convolution fusion module to focus on in-depth processing of the target regions, improving the efficiency and accuracy of detection.
[0020] Step 106: Input the target regions into the multi-scale feature fusion module. Extract the features of the target regions using five different-scale convolutional kernels, generate attention weights using the scale attention mechanism, and perform adaptive weighted fusion to obtain the fused multi-scale features.
[0021] Existing methods are difficult to adaptively and dynamically balance and fuse the local-global features extracted from radar echo data using different-scale convolutional kernels. In the proposed few-shot object detection model, the multi-scale feature fusion module extracts the features of the target region using five different-scale convolutional kernels. Convolutional kernels of different scales have different receptive fields. Small-scale convolutional kernels can capture local detail information, while large-scale convolutional kernels can obtain global semantic information. Just like an intelligent regulator, it can deeply analyze the importance degree of features at each scale. Then, a scale attention mechanism is used to generate attention weights to adaptively weight and fuse the features extracted by different-scale convolutional kernels. This mechanism can automatically adjust the weights of features at different scales according to the importance degree of the features. When the target has unique local detail features that are crucial for accurately judging the target category, the scale attention mechanism will automatically assign higher weights to the features extracted by small-scale convolutional kernels, making them dominant in the fusion process. When the position and general shape of the target in the global scene are more critical for detection, the weights of the features extracted by large-scale convolutional kernels will be increased. Through this adaptive weighted fusion method, the multi-scale feature fusion module successfully achieves the dynamic balance and optimal fusion of features at different scales, thus realizing the best balance extraction of local details and global semantic information.
[0022] Step 108: Input the fused multi-scale features into the differential volume fusion module to fuse the direction-sensitive differential feature map of the pixel neighborhood with the original feature map to obtain the final target features.
[0023] In radar target detection, accurately modeling the feature differences in the boundary regions of objects and completely retaining the global information of objects are the keys to achieving accurate detection. The feature changes in the boundary regions of objects often contain rich information, which can help distinguish different targets and accurately outline the target contours. However, in traditional methods, it is difficult to accurately model these boundary features at the pixel level without losing the global information of the object. The differential convolution fusion module provides an effective way to solve this problem by establishing a direction-sensitive differential feature map of pixel neighborhoods. The direction-sensitive differential feature map can sensitively capture the subtle differences between pixel neighborhoods, especially in the boundary regions of objects, where this difference is more obvious. By accurately analyzing and modeling these differences, the boundary position of the object can be accurately located, and the shape and orientation of the boundary can be accurately depicted. For example, for a target with a complex shape in radar echoes, the differential feature map can clearly mark the changes of each pixel point on the target contour, thus achieving the precise positioning of the object boundary. At the same time, fusing the direction-sensitive differential feature map with the original feature map is a key step in retaining the global information of the object. The original feature map contains rich global information of the target, such as the overall shape, size of the target, and its position in the scene. By organically fusing the differential feature map with the original feature map, not only the feature differences in the boundary regions of the object are highlighted, but also the global information of the target is ensured to be completely retained. During the fusion process, the two feature maps complement each other, so that the final obtained target features can accurately reflect the detailed information of the object boundary and maintain a comprehensive understanding of the whole target, thus achieving the accurate modeling of the feature differences in the boundary regions of the object at the pixel level while completely retaining the global information of the object.
[0024] For the above-mentioned generalized few-shot target detection method based on differential multi-scale fusion, this application proposes a differential convolution multi-scale feature fusion framework, including two core designs: designing a multi-scale feature fusion module to adaptively weighted fuse the features extracted by five different-scale convolution kernels using a scale attention mechanism to achieve the best balance extraction of local details and global semantic information; designing a differential convolution fusion module to fuse the direction-sensitive differential feature map of pixel neighborhoods with the original feature map.
[0025] In one embodiment, the multi-scale feature fusion module includes a multi-scale feature extraction module and an adaptive scale fusion module.
[0026] In one embodiment, the target region is input into the multi-scale feature fusion module to adaptively weighted fuse the features of the target region extracted by five different-scale convolution kernels using a scale attention mechanism to obtain the fused multi-scale features, including:
[0027] The target region is input into the multi-scale feature fusion module, and the parallel small-scale convolutional layer, large-scale convolutional layer, and dilated convolutional layer in the multi-scale feature module are used for feature extraction to obtain multi-scale features;
[0028] In the adaptive scale fusion module, based on the attention mechanism, the adaptive fusion strategy dynamically fuses features by learning the importance weights of each scale branch to obtain fused multi-scale features.
[0029] In a specific embodiment, the multi-scale feature fusion module MFFM includes two core designs: multi-scale feature extraction and adaptive scale fusion. In terms of feature extraction, MFFM constructs a multi-branch parallel architecture, including 1×1, 3×3, 5×5, 7×7 convolutions and 3×3 dilated convolutions. Among them, the small-scale convolutions (1×1, 3×3) focus on capturing local fine features, while the large-scale convolutions (5×5, 7×7) and 3×3 dilated convolutions are used to extract more extensive context information. This multi-scale parallel design can still maintain effective modeling of features at different scales under the condition of scarce data. In terms of feature fusion, an adaptive fusion strategy based on the attention mechanism is designed. This mechanism dynamically fuses features by learning the importance weights of each scale branch. Specifically, the attention module first evaluates the discriminability of each scale feature, and then generates corresponding fusion weights, enabling features at different scales to be adaptively combined according to their importance. This dynamic fusion mechanism significantly improves the model's detection ability for targets at different scales, and at the same time enhances the robustness in the case of few samples and scarce data and complex scenarios.
[0030] In one embodiment, the target region is input into the multi-scale feature fusion module, and the parallel small-scale convolutional layer, large-scale convolutional layer, and dilated convolutional layer in the multi-scale feature module are used for feature extraction to obtain multi-scale features, including:
[0031] The target region is input into the multi-scale feature fusion module, and the parallel small-scale convolutional layer, large-scale convolutional layer, and dilated convolutional layer in the multi-scale feature module are used for feature extraction, and the multi-scale features obtained are
[0032] h1 = ReLU(BN(W 1×1 *h + b 1×1 ))
[0033] h2 = ReLU(BN(W 3×3 *h + b 3×3 ))
[0034] h3 = ReLU(BN(W 5×5 *h + b 5×5 ))
[0035] h4 = ReLU(BN(W7×7 *h + b 7×7 ))
[0036] h5 = ReLU(BN(W 3×3dilated *h + b 3×3dilated ))
[0037] F = Concat(h1, h2, h3, h4, h5)
[0038] Among them, * represents the convolution operation, BN represents batch normalization, b represents the bias term related to the corresponding convolution kernel, F represents the multi-scale feature, h represents the target region, and h1, h2, h3, h4, and h5 respectively represent the branches of five different convolution kernel sizes for extracting multi-scale features.
[0039] In one embodiment, the adaptive fusion strategy based on the attention mechanism in the adaptive scale fusion module performs dynamic fusion of features by learning the importance weights of each scale branch to obtain the fused multi-scale features, including:
[0040] The adaptive fusion strategy based on the attention mechanism in the adaptive scale fusion module performs dynamic fusion of features by learning the importance weights of each scale branch to obtain the fused multi-scale features as
[0041] m = AvgPool(F)
[0042] m = ReLU(Conv1(m))
[0043] s = σ(Conv2(m))
[0044] [α1, α2, α3, α4, α5] = s
[0045] h fusion = α1·h1 + α2·h2 + α3·h3 + α4·h4 + α5·h5
[0046] Among them, the features of different scales are concatenated together to obtain the multi-scale feature F ∈ R C×K×K , and then the fusion weights α1, α2, α3, α4, and α5 corresponding to different scale convolutions are generated through the scale attention mechanism. α1, α2, α3, α4, and α5 represent the fusion weights corresponding to different scale convolutions. AvgPool represents the global average pooling operation, Conv represents the convolution operation, σ represents the Sigmoid activation function, and s is the weight vector generated by the scale attention mechanism. h fusion represents the fused multi-scale feature, α i is the fusion weight generated by the scale attention mechanism, and h i respectively represent the scale features under different convolution kernels.
[0047] In one embodiment, the differential convolution fusion module includes a central pixel differential convolution variant, a standard convolution module, and an attention module; the fused multi-scale features are input into the differential convolution fusion module, and the final target features are obtained by fusing the direction-sensitive differential feature map of the pixel neighborhood with the original feature map, including:
[0048] Concatenate the input fused multi-scale features along the channel dimension to obtain a concatenated feature map, and process the concatenated feature map through an initial convolution layer, batch normalization, and ReLU activation function to obtain a processed feature map;
[0049] After independently processing the processed feature map using the standard convolution module and the central pixel differential convolution variant, local feature maps are obtained;
[0050] Concatenate the local feature maps along the channel dimension, calculate the attention weights through multiple 1×1 convolutions, and normalize the attention weights using Softmax to obtain the normalized attention weights;
[0051] Apply the normalized attention weights to the local feature maps to obtain weighted features; then add all the weighted features together to obtain the final target features.
[0052] In one embodiment, pixel difference calculation is inserted during the convolution process of the central pixel differential convolution variant, and the specific calculation process is as follows:
[0053]
[0054] Among them, x represents the input in the local area, w represents the weight of the convolution kernel, x i and x' i are the pixel values of the input respectively, and w i is the weight in the k×k convolution kernel.
[0055] In a specific embodiment, the present application designs three complementary pixel differential convolution (PDC) variants: 1. Central PDC (CPDC): Focuses on modeling the difference relationship between the central pixel and neighboring pixels. 2. Angular PDC (APDC): Focuses on capturing local structural changes along the circumferential direction. 3. Radial PDC (RPDC): Focuses on extracting intensity gradient information along the radial direction. As Figure 3 shown, these three PDC variants jointly construct a multi-view local structure representation framework through different pixel sampling and differential calculation strategies.
[0056] The PDC of the present application incorporates pixel difference calculation during the convolution operation, as shown in the following formula:
[0057]
[0058] Through systematic experimental evaluations on few-shot object detection tasks, it is found that CPDC exhibits significant performance advantages compared to APDC and RPDC. Specifically, CPDC achieves a significant improvement in detection accuracy compared to standard convolutions, while the performance of APDC and RPDC is comparable to that of standard convolutions. Based on this finding, a complementary combination design of CPDC and standard convolutions is adopted in the final differential convolution fusion module (DCFM). The core idea of DCFM is to fuse two complementary feature representations: the CPDC branch focuses on modeling local pixel differences and can effectively capture the fine-grained structural features of the object; the standard convolution branch focuses on extracting high-level semantic features such as color, contour, and texture. To achieve the optimal fusion of these two types of features, an attention mechanism is used to dynamically adjust the importance of each branch. Specifically, the attention module first evaluates the discriminability of the output feature maps of the two branches, then generates adaptive fusion weights, and finally weights and fuses the outputs of each convolution to obtain the final fused features. This attention-based dynamic fusion strategy enables the model to adaptively balance local structural information and high-level semantic information according to the characteristics of the input data.
[0059] In one embodiment, after independently processing the processed feature map using the standard convolution module and the central pixel differential convolution variant, a local feature map is obtained, including:
[0060] Performing convolution calculations on the processed feature map using the standard convolution module and the central pixel differential convolution variant, the obtained local feature maps are respectively
[0061] diff cv = PDC cv (x conv1 )
[0062] diff cd = PDC cd (x conv1 )
[0063] where x conv1 represents the processed feature map, PDC cv represents the standard convolution module, and PDC cd represents the central pixel differential convolution variant.
[0064] In one embodiment, the local feature maps are concatenated along the channel dimension, and attention weights are calculated through multiple 1×1 convolutions. Then, the attention weights are normalized using Softmax to obtain the normalized attention weights, including:
[0065] The local feature maps are concatenated along the channel dimension, and attention weights are calculated through multiple 1×1 convolutions. Then, the attention weights are normalized using Softmax to obtain the normalized attention weights as
[0066] diff stack = Concat(diff cv , diff cd )
[0067] attention_weights = Softmax(AttentionFC(diff stack ))
[0068] Among them, Concat represents the dimension concatenation operation.
[0069] In one of the embodiments, applying the normalized attention weights to the local feature map to obtain the weighted features includes:
[0070] Applying the normalized attention weights to the local feature map, the obtained weighted features are
[0071]
[0072] Among them, diffcv represents the output features of the standard convolution module, and diffcd represents the output features of the central pixel difference convolution variant; attentionweights[0] and attentionweights[1] are the attention weights of the standard convolution module and the central pixel difference convolution variant respectively.
[0073] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps do not necessarily execute in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0074] at least a part of the steps in
[0075] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A generalized few-shot target detection method based on differential multi-scale fusion, characterized in that: The method comprises: Constructing a few-shot target detection model; the few-shot target detection model includes a region proposal network and a differential convolution multi-scale feature fusion framework; the differential convolution multi-scale feature fusion framework includes a multi-scale feature fusion module and a differential convolution fusion module; Generate a few-sample initial region proposals according to the region proposal network; extract the corresponding mask region from the initial region proposal and perform proportional expansion optimization to obtain the target region; The target area is input into a multi-scale feature fusion module, and the features of the target area are extracted by using five different scale convolution kernels, and the attention weights are generated by a scale attention mechanism for adaptive weighted fusion to obtain fused multi-scale features; The fused multi-scale features are input into the differential volume fusion module and fused with the original feature map by establishing a direction-sensitive differential feature map of the pixel neighborhood to obtain the final target features.
2. The method according to claim 1, characterized in that: The multi-scale feature fusion module includes a multi-scale feature extraction module and an adaptive scale fusion module.
3. The method according to claim 2, characterized in that The target area is input into the multi-scale feature fusion module. The features of the target area are extracted by five different scale convolution kernels, and the scale attention mechanism is used to generate attention weights for adaptive weighted fusion to obtain fused multi-scale features, including: The target area is input into a multi-scale feature fusion module, and features are extracted using a small-scale convolution layer, a large-scale convolution layer and an expansion convolution layer in parallel in the multi-scale feature module to obtain multi-scale features; In the adaptive scale fusion module, the adaptive fusion strategy based on the attention mechanism dynamically fuses features by learning the importance weights of each scale branch to obtain fused multi-scale features.
4. The method according to claim 1, characterized in that The target area is input into the multi-scale feature fusion module, and features are extracted using the parallel small-scale convolution layer, large-scale convolution layer and dilated convolution layer in the multi-scale feature module to obtain multi-scale features, including: The target area is input into the multi-scale feature fusion module, and the parallel small-scale convolution layer, large-scale convolution layer and dilated convolution layer in the multi-scale feature module are used to extract features, and the multi-scale features are obtained as follows: h1=ReLU(BN(W 1×1 *h+b 1×1 )) h2=ReLU(UN(W 3×3 *h+b 3×3 )) h3=HeLU(BN(W 5×5 *h+b 5×5 )) h4=ReLU(BN(W 7×7 *h+b 7×7 )) h5=ReLU(BN(W 3×3dilated *h+b 3×3dilated )) F=Concat(h1, h2, h3, h4, h5) Among them, * represents the convolution operation, BN represents batch normalization, b represents the bias term associated with the corresponding convolution kernel, F represents multi-scale features, h represents the target area, and h1, h2, h3, h4 and h5 represent the branches of five different convolution kernel sizes to extract multi-scale features.
5. The method according to claim 1, characterized in that In the adaptive scale fusion module, the adaptive fusion strategy based on the attention mechanism dynamically fuses features by learning the importance weights of each scale branch to obtain fused multi-scale features, including: In the adaptive scale fusion module, the adaptive fusion strategy based on the attention mechanism dynamically fuses features by learning the importance weights of each scale branch, and obtains the fused multi-scale features as follows: m=AvgPool(F) m = ReLU(Conv1(m)) s=σ(Conv2(m)) [α1,α2,α3,α4,α5]=s h fusion =α1·h1+α2·h2+α3h3+α4·h4+α5·h5 Among them, α1, α2, α3, α4 and α5 represent the fusion weights corresponding to convolutions of different scales, AvgPool represents the global average pooling operation, Conv represents the convolution operation, σ represents the Sigmoid activation function, and s is the weight vector generated by the scale attention mechanism.
6. The method according to claim 1, characterized in that The differential convolution fusion module includes a central pixel differential convolution variant, a standard convolution module and an attention module; the fused multi-scale features are input into the differential volume fusion module and fused with the original feature map by establishing a direction-sensitive differential feature map of the pixel neighborhood to obtain the final target feature, including: The input fused multi-scale features are spliced along the channel dimension to obtain a spliced feature map, and the spliced feature map is processed by an initial convolution layer, batch normalization, and a ReLU activation function to obtain a processed feature map; After independently processing the processed feature map using a standard convolution module and a center pixel difference convolution variant, a local feature map is obtained; The local feature maps are spliced along the channel dimension, and the attention weights are calculated through multiple 1×1 convolutions, and the attention weights are normalized using Softmax to obtain the normalized attention weights; The normalized attention weights are applied to the local feature map to obtain weighted features; and all weighted features are added together to obtain the final target features.
7. The method according to claim 6, characterized in that The pixel difference calculation is inserted into the convolution process of the center pixel difference convolution variant. The specific calculation process is: Among them, x represents the input in the local area, w represents the weight of the convolution kernel, and x i and x′ i are the input pixel values, w i is the weight in the k×k convolution kernel.
8. The method according to claim 6, characterized in that After independently processing the processed feature map using a standard convolution module and a center pixel difference convolution variant, a local feature map is obtained, including: The processed feature maps are convolved using the standard convolution module and the center pixel difference convolution variant to obtain local feature maps: diff cv =PDC cv (x conv1 ) diff cd =PDC cd (x conv1 ) Among them, x conv1 Represents the processed feature map, PDC cv Represents the standard convolution module, PDC cd Denotes the center-pixel difference convolution variant.
9. The method according to claim 7, characterized in that: The local feature maps are concatenated along the channel dimension, and the attention weights are calculated through multiple 1×1 convolutions, and the attention weights are normalized using Softmax to obtain the normalized attention weights, including: The local feature maps are concatenated along the channel dimension, and the attention weights are calculated through multiple 1×1 convolutions. The attention weights are normalized using Softmax to obtain the normalized attention weights: diff stack =Concat(diff cv ,diff cd ) attention_weights=Softmax(AttentionFC(diff stack )) Among them, Concat represents the dimension concatenation operation.
10. The method according to claim 6, characterized in that Apply the normalized attention weights to the local feature map to obtain weighted features, including: The normalized attention weights are applied to the local feature map, and the weighted features are obtained as follows: Among them, diffcv represents the output features of the standard convolution module, diffcd represents the output features of the center pixel differential convolution variant; attentionweights[0] and attentionweights[1] are the attention weights of the standard convolution module and the center pixel differential convolution variant, respectively.
Citation Information
Cited By
Point cloud learning feature representation method and device based on neighborhood geometric embedding
CN120599434A
Iris real-time tracking and positioning method and device, equipment and storage medium
CN121259900A