Fine-grained Detection Method for Visible Light Remote Sensing Image Target Based on Hierarchical Knowledge Guidance

By constructing a multi-semantic hierarchical label system and feature fusion method, the problem of fine-grained detection in visible light remote sensing images is solved, the target is refinement recognition is achieved, and the detection performance of the model is improved.

CN119810429BActive Publication Date: 2025-07-11BEIJING SATELLITE INFORMATION ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510299055.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing rotary object detection methods are difficult to effectively perform fine-grained detection in visible remote sensing images, especially due to the poor model performance caused by deepening the target classification level, increasing similarity between different subclasses, data inhomogeneity and target number differences.

Method used

By constructing a multi-semantic hierarchical labeling system, ResNet50+FPN and Oriented RPN are used to extract multi-scale region of interest features, perform multi-level semantic feature extraction and adjacent hierarchical feature fusion, and use a loss function of dynamic weight allocation to balance multi-task losses, realizing target fine-grained detection of hierarchical knowledge-guided targets.

Benefits of technology

It improves the model's performance in fine-grained detection tasks, alleviates the problem of insufficient feature learning caused by sample size differences, and improves the accuracy and efficiency of target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810429B_ABST
    Figure CN119810429B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for fine-grained detection of targets in visible light remote sensing images based on hierarchical knowledge guidance, comprising: acquiring visible light remote sensing image data and constructing a multi-semantic hierarchical label system; extracting features of visible light remote sensing images and generating multi-scale region-of-interest features through an RPN network; adopting multi-branch extraction of multi-level semantic features for the multi-scale region-of-interest features; performing local-global feature fusion of adjacent levels on the multi-level semantic features to obtain enhanced features; using multiple level labels to supervise multi-level classification and supervising regression at the first semantic level; and streamlining the network structure in the inference stage to improve the inference speed. In the present invention, for multiple types of targets in visible light remote sensing images, the hierarchical relationship and information injection in the process of fine-grained target detection are realized. Combining multi-level feature fusion, the extraction and learning of the common features and fine-grained features of targets by the network are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visible light remote sensing image target detection and recognition, and particularly relates to a fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance. Background Art

[0002] The rise of deep learning has profoundly changed the traditional interpretation method of remote sensing images. Compared with visual interpretation and traditional manual feature extraction methods, the current target detection method based on rotated bounding boxes has achieved higher efficiency and better performance. With the continuous maturity and development of remote sensing technology, the spatial resolution of visible light remote sensing images has been continuously improved, making it possible to achieve fine-grained target detection in visible light remote sensing images. However, as the target classification level deepens, the similarity between different subclasses becomes greater. At the same time, due to problems such as scale, rotation, and illumination of the same subclass, it poses great challenges for the fine-grained detection task. And due to factors such as uneven distribution of ground objects, selection bias of data collection locations, and selection bias of target types, there may be large differences in the number of targets of different categories, further increasing the difficulty of the fine-grained detection task.

[0003] When the current rotated target detection method detects multiple fine-grained targets, it usually regards all target labels as labels of the same level and defaults to the same mutual exclusion relationship between them, which leads to the lack of hierarchical structure and information between targets, and to a certain extent deepens the problem of poor performance of the model in the fine-grained detection task. Most typical targets in remote sensing images are man-made targets, which have a strict classification system, and the classification level is closely related to the hierarchical features of the targets. Conducting research on them can better extract the differential features of target categories under weakly supervised conditions. At the same time, the features at the high semantic level can be transferred to the low semantic level, which can alleviate the problem of insufficient extraction of some target features caused by insufficient training samples to a certain extent.

[0004] Therefore, it is urgent to carry out research on the fine-grained target detection method guided by hierarchical prior knowledge to improve the performance of fine-grained target detection. Summary of the Invention

[0005] In order to solve the above technical problems existing in the prior art, the purpose of the present invention is to provide a fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance, so as to realize the refined recognition of typical targets in visible light remote sensing images, such as airplanes, ships, etc.

[0006] To achieve the above invention purpose, the present invention provides a fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance, including the following steps:

[0007] Step S1: Obtain visible light remote sensing image data and construct a multi-semantic level label system;

[0008] Step S2: Extract visible light remote sensing image features and generate multi-scale region of interest features through the RPN network;

[0009] Step S3: Adopt multi-branch to extract multi-level semantic features for the multi-scale regions of interest features;

[0010] Step S4: Perform local-global feature fusion for adjacent levels on the multi-level semantic features to obtain enhanced features;

[0011] Step S5: Use multiple level labels to supervise multi-level classification and supervise regression at the first semantic level;

[0012] Step S6: Simplify the network structure in the inference stage and only retain the classification branches at the finest level to improve the inference speed.

[0013] According to a technical solution of the present invention, the step S1 specifically includes:

[0014] Step S11: Obtain visible light remote sensing image data and the fine-grained category annotation and rotated bounding box annotation of the targets therein;

[0015] Step S12: Construct a label tree according to the target classification hierarchy relationship and generate a hierarchical semantic label from coarse to fine levels.

[0016] According to a technical solution of the present invention, the step S2 specifically includes:

[0017] Adopt ResNet50+FPN to extract multi-scale feature maps;

[0018] Based on the rotation characteristics of remote sensing targets, use Oriented RPN to extract a set of region of interest features with a fixed size , where .

[0019] According to a technical solution of the present invention, the step S3 specifically includes:

[0020] Input the features output by Oriented RPN into multiple parallel branches;

[0021] The first branch is directly output to the standard head module, and the remaining branches generate semantic features of levels except the first level through the feature re-extraction module levels ;

[0022] Among them, the feature re-extraction module is implemented by operations of 1×1 convolution for dimensionality reduction, 3×3 convolution for re-extraction, and 1×1 convolution for dimension recovery.

[0023] According to a technical solution of the present invention, in step S4, the features of adjacent levels are subjected to feature fusion, which specifically includes:

[0024] Step S41: Perform global pooling and concatenation on the feature and the feature to generate a global feature ;

[0025] Step S42: Directly add the feature and the feature , and then generate a local feature through a convolution operation;

[0026] Step S43: Add the global feature and the local feature and generate a weight through the Sigmoid function;

[0027] Step S44: Input the enhanced feature after fusion.

[0028] According to a technical solution of the present invention, in step S44, the enhanced feature is expressed as:

[0029]

[0030] According to a technical solution of the present invention, step S5 specifically includes:

[0031] Step S51: In the first semantic level, use the classification branch for coarse-grained category prediction and the regression branch for rotation bounding box coordinate regression;

[0032] Step S52: In the remaining levels, use the enhanced features for fine-grained classification;

[0033] Step S53: Balance the multi-task loss weights based on the loss function balancing method.

[0034] According to a technical solution of the present invention, the classification branch uses multi-class cross-entropy as the classification loss function, and the calculation formula of the classification loss function of multi-class cross-entropy is expressed as:

[0035]

[0036] Among them, is an indicator variable. When = 1, it means that the predicted category is the same as the actual category; when When it is = 0, it means that the predicted category is inconsistent with the actual category; It represents the probability that the th sample is correctly classified;

[0037] According to a technical solution of the present invention, the regression branch uses as the regression loss function to supervise the training of the supervision branch, and the loss calculation formula is expressed as:

[0038]

[0039] where represents the difference between the predicted value and the true value.

[0040] According to a technical solution of the present invention, the loss function balancing method includes:

[0041] Adopt the DWA dynamic weight allocation method to dynamically adjust the multi-task weights according to the loss ratio of adjacent training stages, specifically:

[0042] ,

[0043] where is the temperature coefficient, which is used to adjust the weight difference degree between different tasks; represents the loss ratio of the th task in the training stage and respectively represent the loss values of the th task in the training stage and ; represents the dynamic weight of the th task in the training stage

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] A fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance. By constructing a hierarchical multi-label supervised multi-layer semantic feature extraction and fusion feature network, it solves the problems of hierarchical relationship and information loss caused by flat labels. And through the feature interaction between the upper and lower levels, it alleviates to a certain extent the problem of insufficient feature learning of target categories with fewer samples caused by the difference in the number of different samples in the fine-grained detection task, thereby improving the performance of the model in the fine-grained detection task. This network can be improved and implemented based on multiple existing object detection algorithms and can be used as a general method for constructing a network for fine-grained target detection tasks in visible light remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 Schematically showing a flowchart of a fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance according to an embodiment of the present invention;

[0048] Figure 2 Schematically showing a network diagram of a fine-grained target detection method for visible light remote sensing images guided by hierarchical knowledge according to an embodiment of the present invention;

[0049] Figure 3 Schematically showing a structural diagram of a feature re-extraction module according to an embodiment of the present invention;

[0050] Figure 4 Schematically showing a structural diagram of a feature fusion module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0052] Such as Figure 1As shown in the figure, a fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance, when obtaining visible light remote sensing images and their related annotations, through its fine-grained level labels, according to the existing hierarchical knowledge, a hierarchical multi-label construction from coarse to fine is carried out. At the same time, a multi-level feature extraction network is designed to realize the extraction of different semantic feature information supervised by different semantic level labels. Then, a feature fusion method for adjacent levels is designed to realize feature fusion from shallow to deep. On the basis of maintaining common features, discriminative features are enhanced, and then the features at the fine-grained level are enhanced. Finally, in the training process, a loss summation method combined with dynamic weight allocation is used to make the model training more balanced. Compared with traditional methods, the present invention realizes the injection of hierarchical relationships and information in the process of fine-grained target detection, and through multi-level feature fusion, it is more conducive to the network to extract and learn the common features and fine-grained features of the target, further improving the accuracy of the model and realizing the dual drive of knowledge and data.

[0053] A fine-grained target detection method for visible light remote sensing images based on hierarchical knowledge guidance of the present invention includes the following steps:

[0054] Step S1, obtain visible light remote sensing image data and construct a label system with multiple semantic levels, specifically including:

[0055] Step S11, obtain visible light remote sensing image data and the fine-grained category annotation and rotated bounding box annotation of the targets therein;

[0056] Step S12, construct a label tree according to the target classification hierarchical relationship and generate hierarchical semantic labels from coarse to fine levels.

[0057] Step S2, extract visible light remote sensing image features and generate multi-scale region of interest features through the RPN network; among them, as Figure 2 shown, the feature extraction process includes:

[0058] Adopt a multi-scale feature extraction structure of ResNet50+FPN to obtain a multi-scale feature map of the image. Due to the target rotation characteristics in the remote sensing image, use Oriented RPN to perform region of interest extraction operations on the feature map to obtain the features of multiple regions of interest , where .

[0059] Step S3, adopt multi-branch extraction of multi-level semantic features for the multi-scale region of interest features, as Figure 3 shown, specifically including:

[0060] Output to multiple parallel branches after the previous Oriented RPN structure;

[0061] The first branch directly outputs the features obtained previously to the standard header module;

[0062] The remaining branches extract features through a feature re-extraction module to extract semantic features at different levels, obtaining semantic features at levels except the first level .

[0063] The process of the feature re-extraction module extracting features at different semantic levels includes:

[0064] For the features obtained previously use a convolutional module with a scale of to perform a dimensionality reduction operation, and use function for network activation to obtain features .

[0065] Use the convolutional module with a scale of to perform feature re-extraction operations on the features, and use to maintain the consistency of the length and width dimensions of the features during this process. Also use function for network activation. Through this process, obtain the features of the remaining levels

[0066] Use the convolutional module with a scale of to perform a dimensionality restoration operation on the re-extracted features, and use function for network activation to obtain features .

[0067] Step S4: Perform local-global feature fusion for adjacent levels of the multi-level semantic features to obtain enhanced features, that is, for the features of multiple levels obtained previously , perform adjacent-level feature fusion operations in the local-global feature fusion manner from shallow to deep to obtain enhanced level features .

[0068] The adjacent-level feature fusion operation includes:

[0069] For the features of three adjacent levels , the implementation method during their feature fusion process is:

[0070]

[0071]

[0072] Among them Represents a specific feature fusion method.

[0073] Such as Figure 4 shown, the local-global feature fusion method includes:

[0074] For two features to be fused, such as , respectively perform global pooling to obtain two features , and concatenate them. Then, perform convolution operations and function activation on the concatenated features in sequence to obtain the final global feature .

[0075] For two features to be fused, such as , directly use the addition operation to obtain a feature. Then, perform convolution operations, function activation, convolution operations on this feature in sequence to obtain the enhanced local feature .

[0076] Perform an addition operation on the global feature and the local feature, and then output the feature fusion weight through the function. Obtain the final enhanced feature through , that is, .

[0077] In summary, in step S4, the enhanced feature is expressed as:

[0078]

[0079] Step S5: Use multiple hierarchical labels to supervise multi-level classification and perform regression supervision at the first semantic level. The specific steps of step S5 include:

[0080] Step S51: In the first semantic level, after using the feature as input to the head module, construct a classification branch and a regression branch respectively. Among them, the classification branch is used to achieve the classification of the coarsest level of the target category; the regression branch is used to achieve the coordinate regression of the rotated bounding box of the target, and the five-point representation method is used to represent the horizontal and vertical coordinates of the target center point ( ), the target width ( ), the target height ( ), and the target rotation angle ( );

[0081] Step S52: For Semantic levels, using enhanced features After being input into the head modules of their respective levels, each level adopts a classification branch. The classification branch performs category classification of the target level of its respective level;

[0082] Step S53: Adopt a loss function balancing method to balance the relationship between the losses of each branch, making the training of each branch of the network more balanced.

[0083] In steps S51 and S52, the classification branch uses multi-class cross-entropy loss as the loss function, and the regression branch uses loss as the loss function, thereby supervising the training of the branch.

[0084] The calculation formula of the classification loss function of multi-class cross-entropy is:

[0085]

[0086] where, is an indicator variable. When = 1, it means that the predicted category is the same as the actual category; when = 0, it means that the predicted category is different from the actual category; represents the probability that the th sample is correctly classified;

[0087] The loss calculation formula is:

[0088]

[0089] where, represents the difference between the predicted value and the true value. For each parameter represented by the five-point form, its loss is calculated independently, and the final regression loss value is obtained by weighted summation.

[0090] In an embodiment of the present invention, preferably, in step S53, in order to balance the loss functions between multiple branches, the loss summation method is adopted to replace the ordinary loss addition method. It uses the loss ratio of different tasks in adjacent training to measure their learning speed, and then calculates the weights. For the th loss, its weight calculation method is as follows:

[0091] ,

[0092] where, is the temperature coefficient, which is used to adjust the weight difference degree between different tasks; Indicates the loss ratio of the th task in the training phase; and respectively indicate the loss values of the th task in the training phase and Indicates the dynamic weight of the th task in the training phase.

[0093] Step S6, streamline the network structure in the inference phase, and only retain the classification branches at the finest level to improve the inference speed.

[0094] It should be noted that although the embodiments described above of the present invention are illustrative, they are not limitations of the present invention. Therefore, the present invention is not limited to the above specific embodiments. Without departing from the principle of the present invention, any other embodiments obtained by those skilled in the art under the inspiration of the present invention are deemed to be within the protection scope of the present invention.

Claims

1. A fine-grained detection method for visible light remote sensing image targets based on hierarchical knowledge guidance, characterized in that It includes the following steps: Step S1, obtain visible light remote sensing image data and construct a multi-semantic level label system; Step S2, extract visible light remote sensing image features and generate multi-scale region of interest features through the RPN network; Step S3, use a multi-branch network to extract multi-level semantic features for the multi-scale region of interest features; Step S4, perform local-global feature fusion of adjacent levels on the multi-level semantic features to obtain enhanced features; Step S5, use the multi-semantic level labels to supervise multi-level classification and supervise regression at the first semantic level, specifically including: Step S51, in the first semantic level, use the classification branch to perform coarse-grained category prediction and the regression branch to perform rotation bounding box coordinate regression; Step S52, in the remaining levels, use the enhanced features for fine-grained classification; Step S53, balance the multi-task loss weights based on the loss function balancing method; the loss function balancing method includes: Adopt the DWA dynamic weight allocation method to dynamically adjust the multi-task weights according to the loss ratio of adjacent training stages, specifically: Among them, is the temperature coefficient, which is used to adjust the weight difference degree between different tasks; represents the -th task's loss ratio during the training phase; and respectively represent the -th task's loss values during the training phase and ; represents the -th task's dynamic weight during the training phase .​ Step S6, streamline the multi-branch network structure in the inference stage, and only retain the classification branch of the finest level to improve the inference speed.

2. The method for fine-grained detection of visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 1, characterized in that The specific content of step S1 includes: Step S11, obtain visible light remote sensing image data and the fine-grained category annotation and rotation bounding box annotation of the targets therein; Step S12: Construct a label tree according to the target classification hierarchy relationship to generate a multi-semantic level label system with levels from coarse to fine level.

3. The fine-grained detection method for visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 1, characterized in that, The specific content of step S2 includes: Adopt ResNet50+FPN to extract multi-scale feature maps; Based on the rotation characteristics of remote sensing targets, use Oriented RPN to extract a set of features of the region of interest with a fixed size , where .

4. The method for fine-grained detection of visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 1, wherein The specific content of step S3 includes: The features output by Oriented RPN are input to multiple parallel branches; The first branch is directly output to the standard header module, and the remaining branches generate semantic features of levels other than the first level through the feature re-extraction module levels ; Among them, the feature re-extraction module is implemented by operations of 1×1 convolution for dimensionality reduction, 3×3 convolution for re-extraction, and 1×1 convolution for dimensionality restoration.

5. The method for fine-grained detection of visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 4, wherein In step S4, for the features of adjacent levels feature fusion is performed, specifically including: Step S41. Perform global pooling on feature and feature and concatenate them to generate global feature ; Step S42: directly add Feature and Feature , and then generate a local feature through a convolution operation; Step S43: Add the global feature and the local feature and generate a weight through the Sigmoid function ; Step S44, input the fused enhanced features.

6. The method for fine-grained detection of visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 5, wherein In the step S44, the enhanced feature is expressed as: 。 7. The method for fine-grained detection of visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 1, characterized in that The classification branch uses multi-class cross-entropy as the classification loss function, and the calculation formula of the classification loss function of multi-class cross-entropy is expressed as: Among them, is an indicator variable. When = 1, it means that the predicted class is the same as the actual class; when = 0, it means that the predicted class is different from the actual class; represents the probability that the -th sample is correctly classified; represents the total number of classes of the classification target.

8. The method for fine-grained detection of visible light remote sensing image targets based on hierarchical knowledge guidance according to claim 1, wherein The regression branch uses as the regression loss function to train the supervision branch, and the loss calculation formula is expressed as: Among them, represents the difference between the predicted value and the true value.