Target detection method based on multi-scale feature weighted fusion
Through multi-scale feature weighted fusion and normalized Wasserstein distance label assignment strategy, the problems of insufficient feature fusion and position offset sensitivity in small target detection in remote sensing images are solved, and the detection accuracy and robustness are improved.
Patent Information
- Application Number
- CN202511011701.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional remote sensing image small target detection suffers from insufficient preservation of low-level detail information during feature fusion and the label allocation strategy is sensitive to the positional shift of small targets, resulting in insufficient detection accuracy and uneven anchor box allocation.
A multi-scale feature weighted fusion method is adopted to perform weighted fusion of multi-level feature maps by calculating adaptive fusion weights, and a target detection network is constructed using the label assignment strategy of normalized Wasserstein distance to improve the accuracy of small target detection.
It effectively enhances the detection accuracy and robustness of small targets and improves the target detection performance in complex backgrounds.
Smart Images

Figure CN120853012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to a target detection method based on multi-scale feature weighted fusion. Background Technology
[0002] With the rapid development of remote sensing technology, remote sensing images are increasingly widely used in environmental monitoring, disaster early warning, and urban planning. Especially in small target detection, as target scale decreases and background complexity increases, accurately identifying small targets in remote sensing images has become a major challenge in current research. Traditional small target detection methods mainly rely on rule-based algorithms, such as morphological processing and feature matching. However, these methods perform poorly in complex backgrounds, especially when the contrast between small targets and the background is low. These methods often fail to effectively distinguish small targets from complex backgrounds, easily leading to missed or false detections. Furthermore, while traditional deep learning methods based on convolutional neural networks (CNNs) have made some progress in large target detection, they still have significant shortcomings in small target detection. Small targets in remote sensing images often exhibit low resolution and blurred edges, and may be affected by noise, making it difficult for existing deep learning models to accurately detect these targets. Particularly in multi-scale feature fusion, existing techniques typically rely on bottom-up structures. This approach has limitations in preserving detailed information from low-level features, especially in the case of very finely detailed small targets, where these methods may fail to effectively extract key features. In terms of label assignment strategies, the traditional IoU (Intersection over Union) metric faces significant challenges in small object detection. Small objects are highly sensitive to minute positional shifts, causing the IoU value to drop rapidly, which in turn affects the matching of the object with the candidate bounding box. Since IoU-based label assignment strategies typically set a fixed threshold, this may result in some object boxes not being assigned suitable anchor boxes, thus impacting detection accuracy. Summary of the Invention
[0003] This application provides a target detection method based on multi-scale feature weighted fusion, aiming to solve the technical problems of insufficient detection accuracy caused by insufficient preservation of low-level detail information during feature fusion in traditional remote sensing image small target detection, and uneven anchor box allocation caused by the sensitivity of label allocation strategy to small target position offset. The method achieves the technical effect of improving the ability to express small target details through multi-scale feature adaptive fusion and improving the robustness of label allocation by utilizing normalized Wasserstein distance, thereby improving the detection accuracy of small targets.
[0004] This application provides a target detection method based on multi-scale feature weighted fusion. The method includes: inputting a remote sensing image, extracting multi-scale features from the remote sensing image to obtain multi-level feature maps; calculating a first adaptive fusion weight, upsampling and aligning the feature maps of adjacent levels in the multi-level features, and then using the first adaptive fusion weight to perform weighted fusion on the multi-level feature maps to output a first fused feature map; calculating a second adaptive fusion weight, downsampling the first fused feature map at multiple scales to obtain a downsampled fused feature map, and then using the second adaptive fusion weight to perform weighted fusion on the downsampled fused feature map to output a second fused feature; and constructing a target detection network using a label allocation strategy based on normalized Wasserstein distance, wherein the target detection network is used to receive the second fused feature corresponding to the detection head for target detection and output the target detection result.
[0005] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0006] The aforementioned target detection method based on multi-scale feature weighted fusion first inputs a remote sensing image and extracts its multi-scale features to generate feature maps at multiple levels. Then, it calculates the first adaptive fusion weights, aligns adjacent feature maps through upsampling, and performs weighted fusion to obtain the first fused feature map. Next, it calculates the second adaptive fusion weights and performs multi-scale downsampling on the first fused feature map to obtain a downsampled fused feature map, which is then weighted and fused again to output the second fused feature map. Finally, it constructs a target detection network using a normalized Wasserstein distance label assignment strategy. This network utilizes the second fused features to perform target detection and outputs the final detection result. This method effectively improves the detection accuracy of small targets in remote sensing images by dynamically fusing multi-scale features and optimizing label matching.
[0007] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating a target detection method based on multi-scale feature weighted fusion in one embodiment.
[0010] Figure 2 This is a flowchart illustrating the label allocation strategy of a target detection method based on multi-scale feature weighted fusion in one embodiment. Detailed Implementation
[0011] This application provides a target detection method based on multi-scale feature weighted fusion, which solves the technical problems of insufficient detection accuracy caused by insufficient preservation of low-level detail information during feature fusion in traditional remote sensing image small target detection, and uneven anchor box allocation caused by the sensitivity of label allocation strategy to small target position offset. It achieves the technical effect of improving the ability to express small target details through multi-scale feature adaptive fusion and improving the robustness of label allocation by using normalized Wasserstein distance, thereby improving the accuracy of small target detection.
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0013] It should be noted that the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product, or device.
[0014] Examples, such as Figure 1 As shown, this application provides a target detection method based on multi-scale feature weighted fusion, the method comprising:
[0015] Input a remote sensing image and extract multi-scale features from the remote sensing image to obtain a multi-level feature map.
[0016] In this embodiment, the input remote sensing image undergoes preprocessing, including pixel value normalization (e.g., standardizing the RGB three-channel values to the [0,1] range) and size unification (e.g., adjusting to 800×800 pixels). After preprocessing, features at different scales are extracted from the preprocessed remote sensing image using a convolutional neural network (CNN) or other feature extraction network. These features are information at different levels extracted from the remote sensing image through multiple convolutional operations. Lower-level features retain detailed information in the image, while higher-level features capture more semantic information. To extract these multi-scale features, the remote sensing image passes through a series of convolutional layers. The filter size of these convolutional layers can be gradually increased to capture information at different spatial scales. The output of each convolutional layer can be used as a feature map of the image. In this way, the image feature map not only has high spatial resolution (low-level feature map) but also captures higher-level semantic information (high-level feature map). These feature maps represent multi-dimensional information of the image at different scales, containing features at different levels from details to the global picture. In different feature maps, the lower the feature layer level, the more detailed information of small objects can be captured, with higher spatial resolution and more edge texture details. Conversely, the higher the feature layer level, the more global semantic information is captured, with lower spatial resolution and more strong semantic information. Finally, after multi-scale feature extraction, multiple layers of feature maps can be obtained. These feature maps contain multi-level, global, and local information of the image, becoming key data for subsequent fusion, weighted processing, and object detection.
[0017] Calculate the first adaptive fusion weight, upsample and align the feature maps of adjacent layers in the multi-level features, and then use the first adaptive fusion weight to perform weighted fusion on the multi-level feature maps to output the first fused feature map.
[0018] In one embodiment, multi-level feature maps extracted from remote sensing images are first processed to calculate the first adaptive fusion weights. Since feature maps at different levels have different resolutions, it is necessary to upsample adjacent feature maps to ensure they have the same spatial dimensions. Specifically, by using upsampling methods (such as deconvolution, bilinear interpolation, etc.), the low-resolution feature maps (usually high-level feature maps) are resized to align with the high-resolution feature maps (usually low-level feature maps). This process effectively combines the detailed information in low-level features with the semantic information in high-level features. Subsequently, an adaptive mechanism is introduced by concatenating adjacent feature maps, then performing global average pooling to pre-optimize the concatenated feature maps, and finally normalizing the weight coefficients calculated based on these global features using the Softmax activation function to obtain the first adaptive fusion weights. After obtaining the first adaptive fusion weight, the adjacent feature maps in the multi-level feature maps are weighted and fused using the first adaptive fusion weight. This fusion process is performed layer by layer. For example, after B2 (feature map) is fused with B3, it is then fused with B4. Finally, the first fused feature map containing multi-scale features is output. This first fused feature map contains detailed information of low-level features and global semantic information of high-level features, providing richer and more accurate feature representations for subsequent object detection tasks.
[0019] Furthermore, this application provides the first adaptive fusion weights for fusing feature maps from adjacent layers, and the method for calculating the first adaptive fusion weights includes:
[0020] The feature maps of adjacent layers are compressed and concatenated to obtain a concatenated feature map; global average pooling is performed on the concatenated feature map to extract global features; the global features of the concatenated feature map are received through two parallel channels, and the weight coefficients corresponding to the two parallel channels are calculated; the Softmax activation function is used to normalize the weight coefficients corresponding to multiple adjacent layers to obtain the first adaptive fusion weight.
[0021] Preferably, when calculating the first adaptive fusion weights, the upsampled feature maps of adjacent layers are first concatenated along the channel dimension to form a richer feature representation. This concatenated feature map contains feature information from multiple scales, including detailed information (low-level feature maps) and global semantic information (high-level feature maps), for subsequent weight calculation and weighted fusion. Then, a global average pooling operation is performed on the concatenated feature map; that is, for each channel, the pooling operation traverses the entire feature map and calculates the average value for each channel, resulting in a single value representing that channel. Through this global average pooling, the global features of the concatenated feature map can be extracted. These global features represent the overall expression of features at different scales and levels. Afterwards, the global features of the concatenated feature map are received through two parallel channels. Each channel independently processes the global features and calculates the corresponding weight coefficients. These two parallel channels can be two fully connected layers (or other forms of computational layers). Their function is to estimate the weight coefficients of each feature layer through learning, based on a learnable weight matrix and bias terms. These weight coefficients reflect the relative importance of each layer's features during the fusion process. The design of the parallel channels aims to understand global features from different perspectives, thereby improving fusion accuracy. After calculating the weight coefficients corresponding to the two parallel channels, the Softmax activation function is used to normalize these weight coefficients. The Softmax function compresses the weight coefficients to the range [0,1] and ensures that the sum of all weight coefficients equals 1, ensuring that each weight value is within a reasonable range, neither too large nor too small. Normalization helps enhance the stability of the network, making the fusion process smoother and ensuring that the weights and proportions of all layer features are reasonably adjusted. Finally, the weight coefficients processed by the Softmax activation function are the first adaptive fusion weights. The first adaptive fusion weights represent the relative importance of each adjacent layer's features during fusion. By dynamically adjusting the first adaptive fusion weights, the network can adaptively optimize the contributions of features from different layers, allowing low-level and high-level features to complement each other during the fusion process, thereby improving the detection performance of small targets.
[0022] Calculate the second adaptive fusion weights, perform multi-scale downsampling on the first fusion feature map to obtain a downsampled fusion feature map, and use the second adaptive fusion weights to perform weighted fusion on the downsampled fusion feature map to output the second fusion feature.
[0023] In one embodiment, a first fused feature map is first used as input, and a multi-scale downsampling operation is performed. The purpose of downsampling is to capture broader global semantic information by reducing the spatial resolution of the feature map. Specifically, downsampling operations (such as pooling or strided convolution) are used to reduce the spatial dimension of the first fused feature map, thereby generating a low-resolution downsampled fused feature map. This extracts more generalized global information, allowing the feature map to cover a larger receptive field and enhancing the perception of large-scale targets. Subsequently, an adaptive mechanism is introduced. The downsampled feature maps are concatenated, and then the weight coefficients of each feature map are calculated through fully connected layers or other forms of computational layers. This process is similar to the calculation of the first adaptive fusion weights. Then, these second adaptive fusion weights are used to perform weighted fusion of the downsampled fused feature maps. In this process, the contribution of each downsampled fused feature map is adjusted according to its calculated weight. Through weighted fusion, it is ensured that the more information-valuable parts of the downsampled feature maps receive higher weights, thus enabling the entire fused feature map to more accurately represent the global semantic information of the image. Finally, the output includes a second fusion feature map containing multi-scale information. This second fusion feature map is a feature representation after being weighted by the second adaptive fusion weight. It can achieve a balance between global semantics and local details, providing more comprehensive and richer feature support for subsequent object detection tasks. It can effectively improve the detection accuracy, especially in the detection of small objects, and enhance the model's recognition ability and robustness.
[0024] Furthermore, this application provides the second adaptive fusion weight for fusing feature maps downsampled at different scales, and the method for calculating the second adaptive fusion weight includes:
[0025] The downsampled fusion feature maps at different scales are concatenated to obtain a multi-scale concatenated feature map; the multi-scale concatenated feature map is weighted by a fully connected layer to output multiple weight coefficients; the multiple weight coefficients are normalized by the Softmax activation function to obtain the second adaptive fusion weight.
[0026] Optionally, downsampling to obtain downsampled fusion feature maps at different scales contains image information of varying sizes. By stitching these downsampled feature maps at different scales along the channel dimension, a richer multi-scale stitched feature map can be obtained. This multi-scale stitched feature map encompasses global and local information from multiple scales, providing more diverse feature inputs for subsequent weight mapping. Subsequently, the obtained multi-scale stitched feature map undergoes the same global average pooling as described above, compressing it into multi-scale global features. Then, similar processing is performed on the multi-scale global features through multiple parallel (e.g., four) fully connected layers to calculate the weight coefficients corresponding to these parallel fully connected layers. These weight coefficients determine the contribution of each scale's features during fusion. Finally, to ensure that the weight coefficients are within a reasonable range and can adapt to the subsequent weighted fusion process, a Softmax activation function is used for normalization, ensuring that each weight coefficient is within a stable and reasonable range, thus guaranteeing the stability and accuracy of feature fusion. Finally, the multiple weight coefficients normalized by the Softmax activation function are the second adaptive fusion weights. The second adaptive fusion weights represent the importance of downsampled fusion features at each scale in the fusion process. They can dynamically adjust the contribution ratio of features at different scales according to the features of the input image, ensuring that low-resolution and high-resolution features can be reasonably fused when fusing multi-scale features, thereby improving the performance of small target detection.
[0027] Furthermore, this application provides a downsampled fusion feature map obtained by downsampling the first fusion feature map at multiple scales, wherein the multiple scale downsampling includes 2x downsampling, 4x downsampling, and 8x downsampling.
[0028] Optionally, the first input is the first fused feature map after the first adaptive fusion process. This first fused feature map already possesses relatively rich details and semantic information. To further enhance the perception of global information and capture targets at a larger scale, different downsampling factors (e.g., 2x, 4x, 8x) are used to downsample the first fused feature map to expand the receptive field and capture a wider range of semantic information. In this process, the first fused feature map is first downsampled by 2x, that is, by pooling (such as max pooling or average pooling) or strided convolution operations, reducing the spatial resolution of the feature map, making the spatial size of the first fused feature map halved. This operation maintains relatively important semantic information and overall structural features while reducing resolution, thus preserving a wider range of contextual information. Subsequently, the first fused feature map is downsampled by 4x. This step further reduces the spatial size of the feature map through larger-scale pooling or strided convolution operations, making the spatial size of the first fused feature map reduced to one-quarter of its original size, thereby more effectively extracting global information, reducing excessive details, and helping the network focus on a wider range of semantic information. Finally, the first fused feature map is downsampled by 8 times to further reduce its resolution, shrinking its size to one-eighth of its original size. This captures a wider range of global information, particularly enhancing the perception of large-scale targets. While the feature map downsampled by 8 times has less detail, it provides stronger global contextual information, making it suitable for detecting overall targets. After multi-scale downsampling, three downsampled fused feature maps with different resolutions are obtained, corresponding to downsampling factors of 2x, 4x, and 8x, respectively. Each scale of feature map provides multi-layered information about the image from different perspectives. This information will be combined in the subsequent weighted fusion process to provide a richer and more comprehensive feature representation, supporting more efficient and accurate small target detection.
[0029] A target detection network is constructed using a label assignment strategy based on normalized Wasserstein distance. This target detection network receives the second fusion feature corresponding to the detection head, performs target detection, and outputs the target detection result.
[0030] In one embodiment, firstly, an object detection network is constructed based on a label assignment strategy using Normalized Wasserstein distance (NWD). NWD measures the normalized spatial distribution similarity between candidate anchor boxes and ground truth object boxes, thus exhibiting stronger robustness to small shifts in small objects and effectively avoiding mismatches caused by positional deviations. The label assignment strategy assigns the most suitable candidate anchor box to each object box based on the NWD between the object box and the candidate anchor box, ensuring more accurate label assignment. The object detection network is trained using a regression loss function based on positive and negative label samples, enabling more effective object identification and reducing detection errors caused by inaccurate label assignment. By inputting the second fusion feature provided by the detection head into the object detection network, the network outputs a target detection result including the detected object category and its corresponding bounding box location. This approach improves the accuracy of small object detection, especially in scenes with complex backgrounds or significant target scale variations, effectively enhancing the robustness and accuracy of object detection.
[0031] Furthermore, such as Figure 2 As shown, this application provides a label assignment strategy using normalized Wasserstein distance to construct an object detection network. The label assignment strategy using normalized Wasserstein distance includes:
[0032] Construct real bounding box samples and obtain any target bounding box from the real bounding box samples as the first target bounding box; obtain a first candidate anchor box set for the first target bounding box through a preset aspect ratio filtering strategy; expand the neighborhood grid of the first target bounding box and determine a second candidate anchor box set from the first candidate anchor box set; calculate the normalized Wasserstein distance between the first target bounding box and any candidate anchor box in the second candidate anchor box set, and output the normalized Wasserstein distance calculation value; based on the normalized Wasserstein distance calculation value, identify the first k candidate anchor boxes that are less than a preset normalized Wasserstein distance threshold to obtain the matching anchor box set corresponding to the first target bounding box.
[0033] Preferably, a set of real bounding box samples containing multiple real target boxes is first constructed. These samples typically come from image annotation data, representing actual targets in the image. Each real bounding box in each sample corresponds to an actual target and has a specific size and position. Then, a target box is arbitrarily selected from the real bounding box samples as the first target box to be detected. For the first target box, a preset aspect ratio filtering strategy is used. This strategy compares the aspect ratio of the original anchor box with a preset aspect ratio threshold, eliminating anchor boxes with excessively large differences from the threshold. This yields a first set of candidate anchor boxes, ensuring similarity in shape between the candidate boxes and the target boxes, thereby improving matching accuracy. Next, to further improve the coverage and matching accuracy of the candidate boxes, the first target box undergoes neighborhood grid expansion. This involves selecting multiple neighborhood grids around the target box and expanding them into new candidate regions. The original anchor boxes in these candidate regions are then compared according to the aspect ratio filtering strategy, generating a second set of candidate anchor boxes. These anchor boxes are generated based on the relative positions of the first target boxes and can cover potential target locations. Then, for each candidate box in the second set of candidate anchor boxes, the normalized Wasserstein distance (NWD) between the candidate anchor box and the first target box is calculated. NWD is an index that measures the spatial similarity between bounding boxes and can effectively reflect the spatial similarity between two boxes. The specific calculation method is as follows: Where NWD(G,A) is the normalized Wasserstein distance between the first target box G and the candidate anchor box A, and G x G y G w G h These are the center coordinates, width, and height of the first bounding box G, respectively. x 、A y 、A w 、A h Here, L1 represents the center point coordinates, width, and height of candidate anchor box A, respectively. L2 is the Euclidean distance between the first target box G and candidate anchor box A. C is a constant, typically set to the average size of the targets in the dataset. Finally, based on the calculated Normalized Wasserstein distance (NWD) values, the top k candidate anchor boxes with distance values less than a preset NWD threshold are selected. Candidate boxes with smaller NWD values have a more similar spatial distribution to the target box and a higher matching degree. Therefore, by selecting the k smallest NWD values, it is ensured that the candidate anchor box that best matches the target box is selected. Ultimately, the selected k candidate anchor boxes constitute the matching anchor box set corresponding to the first target box. The anchor boxes in this matching anchor box set are the most similar to the first target box, and they will serve as positive samples in subsequent object detection tasks, participating in object classification and bounding box regression, thereby improving detection performance.
[0034] Furthermore, after obtaining the set of matching anchor boxes corresponding to the first target box, this application provides a method that further includes:
[0035] Obtain the set of matching anchor boxes for each target box in the ground truth bounding box sample; analyze the set of matching anchor boxes for each target box to obtain overlapping matching anchor boxes, wherein the overlapping matching anchor boxes are anchor boxes corresponding to at least two target boxes; obtain multiple normalized Wasserstein distance calculation values for multiple target boxes corresponding to the overlapping matching anchor boxes; assign the overlapping matching anchor boxes to labeled target boxes according to the multiple normalized Wasserstein distance calculation values, wherein the labeled target boxes are the target boxes with the smallest normalized Wasserstein distance calculation values.
[0036] Optionally, for each ground truth bounding box, a corresponding set of matching anchor boxes is obtained through the same normalized Wasserstein distance calculation. These sets contain anchor boxes with high matching degrees to their corresponding ground truth bounding boxes, which are used as positive samples for object classification and bounding box regression in subsequent object detection tasks. Then, among the matching anchor box sets of all bounding boxes, overlapping matching anchor boxes are identified. Overlapping matching anchor boxes are those that match multiple ground truth bounding boxes simultaneously; that is, an anchor box may have high similarity matches with different ground truth bounding boxes. Next, for each overlapping matching anchor box, the normalized Wasserstein distance between this overlapping matching anchor box and all ground truth bounding boxes containing that anchor box is obtained. The multiple normalized Wasserstein distances for each overlapping matching anchor box are then compared, and the ground truth bounding box with the smallest normalized Wasserstein distance is selected as the labeled bounding box for each overlapping matching anchor box. The overlapping matching anchor box is then assigned to the corresponding labeled bounding box. This ensures that the anchor box is ultimately assigned to the most suitable bounding box, thereby improving the accuracy of object detection. The above methods can effectively resolve the conflict between anchor boxes and multiple target boxes, ensuring that each anchor box is ultimately assigned to the most suitable real target box, thereby improving the accuracy and robustness of small target detection tasks.
[0037] Furthermore, this application provides a method for constructing an object detection network using a label assignment strategy based on normalized Wasserstein distance, including:
[0038] After obtaining the set of matching anchor boxes for each target box in the ground truth box samples, the set of matching anchor boxes for each target box in the ground truth box samples is marked as a positive label sample; the set of unmatched anchor boxes for each target box in the ground truth box samples is marked as a negative label sample; a regression loss function is introduced to train the network on the positive label samples and the negative label samples to obtain the trained target detection network, wherein the regression loss function includes a regression term based on the normalized Wasserstein distance.
[0039] Optionally, after obtaining the set of matching anchor boxes for each real target box in the ground truth bounding box samples, the anchor boxes in the matching anchor box set are labeled as positive label samples. These positive label samples are those anchor boxes with high similarity to the target boxes. These anchor boxes will be used as positive samples for training in object detection and will be used to determine the location and category of the target. For the set of unmatched anchor boxes for each real target box, that is, those anchor boxes that fail to match the target box or have low similarity to the target box, these anchor boxes are labeled as negative label samples. These negative label samples correspond to background regions. These samples do not contain any target information and are usually used to train the model to identify background regions. After labeling the positive and negative samples, the object detection network is trained using a regression loss function. This function consists of two parts: a classification loss for the bounding box and a regression term. The classification loss measures the difference between the bounding box and the candidate anchor box, typically calculated using the cross-entropy loss function, ensuring the network correctly classifies the target from the background. The regression term calculates the difference between the predicted bounding box and the ground truth bounding box, using the normalized Wasserstein distance. The normalized Wasserstein distance measures the spatial similarity of the bounding boxes, better handling small targets or slight positional shifts, and is particularly suitable for dealing with minor spatial displacements of the bounding box. These two calculation methods make the regression loss function more adaptable to small target detection tasks, providing more accurate target localization. Subsequently, the positive and negative labeled samples are input into the object detection network built on YOLOv5s. The network learns how to identify targets and their locations in the image through forward propagation, loss calculation, backpropagation, and parameter optimization (such as stochastic gradient descent). During training, positively labeled samples help the network learn the features of the target and regress the accurate bounding box of the target; negatively labeled samples help the network learn the features of the background and avoid misidentifying the background area as the target, thus enabling the target detection network to adjust the bounding box position more accurately and improve detection accuracy.
[0040] Furthermore, this application provides a method for expanding the neighborhood grid of the first target box, wherein the neighborhood grid is a 3×3 grid area centered on the grid containing the center point of the first target box.
[0041] Optionally, during the neighborhood mesh expansion, the position (center point coordinates) and size information of the first target box are obtained. A 3×3 mesh region is created with the center point of the first target box as the reference. This mesh region is composed of the eight mesh cells surrounding the mesh where the center point of the first target box is located and the mesh cell where the center point is located. It covers the center point of the target box and the surrounding area to expand the possible location range of the target box, so as to consider potential target areas near the target box, thereby improving the robustness and accuracy of detection.
[0042] Furthermore, after outputting the second fusion feature, this application provides a method that further includes:
[0043] The first fusion feature map is upsampled at multiple scales to obtain an upsampled fusion feature map, wherein the upsampled fusion feature map introduces lateral connections; a third adaptive fusion weight is calculated, and the upsampled fusion feature map is weighted and fused using the third adaptive fusion weight to output a third fusion feature; the target detection network is used to receive the third fusion feature corresponding to the detection head to perform target detection and update the target detection result.
[0044] Preferably, after completing the semantic fusion of multi-scale features, to further improve the detection performance of small targets, the first fused feature map is upsampled at multiple scales (2x, 4x, 8x). The purpose of upsampling is to gradually increase the size of the feature map through methods such as deconvolution or bilinear interpolation, restoring it from low resolution to higher spatial resolution, thereby providing richer detail information. After multi-scale upsampling, an upsampled feature map is obtained. This upsampled feature map contains detailed information recovered from the high-resolution low-level feature map, which can enhance the model's accuracy in handling small targets and complex scenes. Subsequently, to further strengthen the information transfer between features of different scales, lateral connections are introduced. Lateral connections refer to connecting the upsampled high-resolution feature map with the lower-resolution feature map to form an upsampled fused feature map. This ensures that sufficient detail information and global semantic information are retained in the feature map, thereby enhancing the complementarity between feature maps and enabling better fusion of low-level detailed features and high-level semantic features, providing a more comprehensive feature representation for subsequent target detection. Next, the third adaptive fusion weights are calculated. This process is similar to the aforementioned adaptive fusion weight calculation method. Specifically, the upsampled fusion feature map is processed through a fully connected layer or other computational units to calculate the weight coefficient of each feature. Then, the Softmax activation function is used to normalize these weight coefficients to the range [0,1] to ensure the stability and accuracy of feature fusion. Ultimately, the obtained third adaptive fusion weights represent the relative importance of each feature in the fusion process. Then, the upsampled fusion feature map is weighted and fused using the third adaptive fusion weights. In this process, the contributions of different feature maps are adjusted according to the calculated weight coefficients, resulting in the fused third fusion feature. This third fusion feature incorporates information from multiple scales, enabling better capture of targets in the image, especially small and complex targets. Finally, the third fusion feature is input into the object detection network for classification and localization, outputting the final object detection result. This new object detection result is then used to update the original object detection result, thereby improving the accuracy of small object detection.
[0045] Furthermore, after providing the output target detection results, this application also includes the following method:
[0046] Calculate the detection performance index of the target detection result, which includes detection accuracy and detection recall; perform gradient network parameter feedback on the target detection network based on the detection performance index to obtain an updated target detection network.
[0047] Preferably, after obtaining the target detection results, a detection performance index is calculated based on these results. This index includes detection accuracy and detection recall, which measure the precision and comprehensiveness of the target detection network in the prediction results. Detection accuracy refers to the proportion of samples correctly predicted as targets out of all samples predicted as targets, while detection recall refers to the proportion of samples correctly predicted as targets out of all actual target samples. Subsequently, the calculated detection accuracy and recall are applied to the loss function, allowing the network to simultaneously optimize these two metrics during training. Then, based on the loss function, the gradient of the network under the current weights is calculated, and the gradient is propagated to each layer of the network using a backpropagation algorithm. The network parameters are then adjusted based on the gradient values, updating the network weights in a direction that reduces the loss function value, thereby improving target detection performance. By adjusting the network weights and biases, the network can exhibit better performance in the next training cycle, optimizing performance in small target detection tasks. Through this feedback mechanism, the target detection network can continuously adjust and optimize its parameters based on the detection performance index, thereby improving target detection accuracy.
[0048] In summary, the embodiments of this application have at least the following technical effects:
[0049] This embodiment first inputs a remote sensing image and extracts multi-scale features from the image to obtain multi-level feature maps. Then, it calculates a first adaptive fusion weight, upsamples and aligns the feature maps of adjacent levels in the multi-level features, and then uses the first adaptive fusion weight to perform weighted fusion of the multi-level feature maps, outputting a first fused feature map. Next, it calculates a second adaptive fusion weight, downsamples the first fused feature map at multiple scales to obtain a downsampled fused feature map, and then uses the second adaptive fusion weight to perform weighted fusion of the downsampled fused feature map, outputting a second fused feature. Finally, it constructs a target detection network using a label allocation strategy based on normalized Wasserstein distance. This target detection network receives the second fused feature corresponding to the detection head, performs target detection, and outputs the target detection result. These technical effects collectively solve the problems of insufficient detection accuracy caused by insufficient preservation of low-level detail information during feature fusion in traditional remote sensing image small target detection, and uneven anchor box allocation caused by the sensitivity of the label allocation strategy to small target position offset. It achieves the technical effect of enhancing the detail representation capability of small targets through multi-scale feature adaptive fusion and improving the robustness of label allocation using normalized Wasserstein distance, thereby improving the accuracy of small target detection.
[0050] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0051] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0052] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A target detection method based on multi-scale feature weighted fusion, characterized in that, The method includes: Input a remote sensing image, and extract multi-scale features from the remote sensing image to obtain a multi-level feature map; Calculate the first adaptive fusion weight, upsample and align the feature maps of adjacent levels in the multi-level features, and then use the first adaptive fusion weight to perform weighted fusion on the multi-level feature maps to output the first fused feature map; Calculate the second adaptive fusion weight, and obtain a downsampled fusion feature map by multi-scale downsampling of the first fusion feature map. Use the second adaptive fusion weight to perform weighted fusion on the downsampled fusion feature map and output the second fusion feature. A target detection network is constructed using a label assignment strategy based on normalized Wasserstein distance. This target detection network receives the second fusion feature corresponding to the detection head, performs target detection, and outputs the target detection result.
2. The target detection method based on multi-scale feature weighted fusion as described in claim 1, characterized in that, A target detection network is constructed using a label assignment strategy based on normalized Wasserstein distance. The label assignment strategy based on normalized Wasserstein distance includes: Construct a real bounding box sample, and obtain any target box in the real bounding box sample as the first target box; The first set of candidate anchor boxes for the first target box is obtained by using a preset aspect ratio filtering strategy; The neighborhood grid of the first target box is expanded, and a second set of candidate anchor boxes is determined from the first set of candidate anchor boxes; Calculate the normalized Wasserstein distance between the first target box and any candidate anchor box in the second candidate anchor box set, and output the normalized Wasserstein distance calculation value; Based on the normalized Wasserstein distance calculation value, the first k candidate anchor boxes that are less than the preset normalized Wasserstein distance threshold are identified, and the matching anchor box set corresponding to the first target box is obtained.
3. The target detection method based on multi-scale feature weighted fusion as described in claim 2, characterized in that, After obtaining the set of matching anchor boxes corresponding to the first target box, the method further includes: Obtain the set of matching anchor boxes for each target box in the real box sample; Analyze the set of matching anchor boxes for each target box to obtain overlapping matching anchor boxes, wherein the overlapping matching anchor boxes are anchor boxes that correspond to at least two target boxes; Obtain multiple normalized Wasserstein distance values for multiple target boxes corresponding to the overlapping matching anchor boxes; Based on the multiple normalized Wasserstein distance calculation values, the overlapping matching anchor boxes are assigned to the identified target boxes, and the identified target boxes are the target boxes with the smallest normalized Wasserstein distance calculation values.
4. The target detection method based on multi-scale feature weighted fusion as described in claim 2, characterized in that, Methods for constructing object detection networks using label assignment strategies based on normalized Wasserstein distance include: After obtaining the set of matching anchor boxes for each target box in the real box sample, the set of matching anchor boxes for each target box in the real box sample is marked as a positive label sample; The set of unmatched anchor boxes for each target box in the real box sample is marked as a negative label sample; A regression loss function is introduced to train the network on the positive label samples and the negative label samples to obtain a trained target detection network. The regression loss function includes a regression term based on the normalized Wasserstein distance.
5. The target detection method based on multi-scale feature weighted fusion as described in claim 2, characterized in that, The first target box is expanded by a neighborhood grid, which is a 3×3 grid area centered on the grid where the center point of the first target box is located.
6. The target detection method based on multi-scale feature weighted fusion as described in claim 1, characterized in that, The first adaptive fusion weight is used to fuse feature maps from adjacent layers. The method for calculating the first adaptive fusion weight includes: The feature maps of adjacent layers are compressed and stitched together to obtain a stitched feature map; Global average pooling is performed on the spliced feature map to extract the global features of the spliced feature map; The global features of the stitched feature map are received through two parallel channels, and the weight coefficients corresponding to the two parallel channels are calculated. The first adaptive fusion weights are obtained by normalizing the weight coefficients corresponding to multiple adjacent layers using the Softmax activation function.
7. The target detection method based on multi-scale feature weighted fusion as described in claim 6, characterized in that, The second adaptive fusion weight is used to fuse feature maps sampled at different scales. The method for calculating the second adaptive fusion weight includes: By stitching together downsampled fused feature maps at different scales, a multi-scale stitched feature map is obtained; The multi-scale stitched feature map is weighted by a fully connected layer, and multiple weight coefficients are output. The multiple weight coefficients are normalized using the Softmax activation function to obtain the second adaptive fusion weight.
8. The target detection method based on multi-scale feature weighted fusion as described in claim 7, characterized in that, The first fused feature map is downsampled at multiple scales to obtain a downsampled fused feature map. The multiple scale downsampling includes 2x downsampling, 4x downsampling, and 8x downsampling.
9. The target detection method based on multi-scale feature weighted fusion as described in claim 1, characterized in that, After outputting the second fused feature, the method also includes: The first fused feature map is upsampled at multiple scales to obtain an upsampled fused feature map, wherein the upsampled fused feature map introduces lateral connections; Calculate the third adaptive fusion weight, and use the third adaptive fusion weight to perform weighted fusion on the upsampled fusion feature map to output the third fusion feature; The target detection network is used to receive the third fusion feature corresponding to the detection head for target detection and update the target detection results.
10. The method as described in claim 1, characterized in that, After outputting the target detection results, the method also includes: Calculate the detection performance index of the target detection result, wherein the detection performance index includes detection accuracy and detection recall; Based on the detection performance indicators, the target detection network parameters are fed back using gradients to obtain an updated target detection network.
Citation Information
Cited By
A vehicle detection method, apparatus, equipment, and medium based on deformable convolution.
CN122416385A