Point labeling remote sensing target directional detection method and device
Through the improved ResNet50 model and feature fusion technology, the positive label allocation radius is dynamically adjusted and the rotation box is generated, which solves the efficiency and adaptability problems of the existing methods in remote sensing image detection, and achieves high-precision and efficient directional object detection.
Patent Information
- Application Number
- CN202510337791.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing point-supervised target detection method is inefficient in processing large-scale remote sensing images, has high memory requirements, and is difficult to adapt to different targets and complex scenarios, especially in dense object scenarios, where detection accuracy and insufficient memory are problems.
The improved ResNet50 model is used to combine feature extraction networks with hollow convolution and typical correlation analysis, feature fusion is performed through a hybrid channel attention mechanism, the positive label allocation radius is dynamically adjusted, and the target direction is determined using the eigenvalues of the covariance matrix to generate a rotation box.
It significantly improves detection accuracy and memory efficiency, improves adaptability to different targets and processing capabilities in complex scenarios, especially in dense object scenarios.
Smart Images

Figure CN120279410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a point annotation remote sensing target orientation detection method and device. Background Art
[0002] At present, with the booming development of remote sensing technology, remote sensing data such as aerial images and satellite images are showing explosive growth, which contains a vast amount of target information. The orientation remote sensing target detection technology has emerged under such a background. It plays a crucial role in accurately annotating small and densely arranged targets and has broad application prospects in many fields such as remote sensing image analysis, retail scene analysis, and scene text detection. However, traditional orientation target detection methods face a severe challenge, that is, the work of annotating the orientation bounding box (OBB) requires a large amount of human and time costs. This process is not only cumbersome but also error-prone, greatly limiting the large-scale application and development of target detection technology. To solve this problem, many weakly supervised methods have emerged in recent years, including horizontal box supervision and point supervision. Among them, the point supervision method has received extensive attention because it only needs to annotate one point and category of each target, significantly reducing the annotation cost. This method greatly improves the annotation efficiency on the premise of ensuring a certain detection accuracy, providing new ideas and directions for the development of orientation remote sensing target detection technology.
[0003] The existing point-supervised orientation target detection methods are mainly divided into three categories. The first category is based on the powerful SAM model. The P2RBox model proposed by Cao et al. uses a mask generator such as SAM to generate mask proposals, and a constraint module screens high-quality masks according to the centroid offset penalty. The checker module uses fine point sampling and complex loss calculation to deeply evaluate the mask quality. Finally, the symmetry axis estimation module converts the mask into a rotated box annotation. However, the SAM model is very effective on natural images, but when dealing with large-scale remote sensing images, this inefficiency is particularly obvious, seriously affecting the efficiency of target detection and making the SAM-based method difficult to meet the requirements of speed and memory in practical applications. Secondly, there are also challenges in the accuracy of direction estimation for asymmetric objects. The PointSAM model proposed by Liu et al. is based on a self-training framework, using the zero-shot ability of SAM, and training by iteratively generating pseudo-labels. The methods include prototype regularization (PBR), which generates target prototypes offline and dynamically updates the prediction prototypes, and uses the Hungarian algorithm to align the two to reduce the error accumulation in self-training; negative sample calibration (NPC) is also proposed, based on the assumption that instance masks do not overlap, using the hint of overlapping masks as a negative signal to adjust negative samples, thereby optimizing mask prediction. However, the double-branch structure of the self-training framework leads to a slow training speed, and the negative sample calibration has poor effects when dealing with sparsely distributed objects, lacking a strategy that can accurately complete the positive and negative sample allocation.
[0004] The second type of method is the method based on human prior knowledge. The method based on prior knowledge represented by Point2RBox proposed by Yu et al. needs to integrate human prior knowledge. However, different datasets often have different characteristics and distributions, which means that specific prior knowledge is required for each dataset, greatly reducing the generality of the method. It is difficult to combine with more powerful detectors, thus unable to fully utilize the performance improvement potential of advanced detectors, and seems powerless when facing complex and variable object detection tasks.
[0005] The third type of method is the method based on modular structure. PointOBB proposed by Luo et al. does not rely on manually designed priors and provides greater flexibility by decoupling pseudo-label generation from the detector, making it more suitable for efficient and scalable detection tasks. However, due to the teacher-student structure, the pseudo-label generation process is very slow, about 7-8 times longer than the subsequent detector training. In addition, due to multiple view conversions, its training requires a large amount of GPU memory. Moreover, the change in the number of regions of interest (Rols) may lead to out-of-memory problems, especially in dense object scenarios. Although restricting the number of Rols can alleviate this problem, it will lead to a performance decline. PointOBB-v2 proposed by Ren et al., as the current state-of-the-art point-supervised oriented detection method, generates class probability maps (CPMs) from point annotations and designs a new sample assignment strategy to capture the contours and directions of objects from the CPMs. Next, non-uniform sampling based on probability distribution is adopted, and principal component analysis (PCA) is used to determine the boundaries and directions of the targets. However, the method uses a fixed radius when assigning positive samples. When applied to other targets or datasets with different characteristics, adjustments may be required, otherwise the best performance may not be achieved.
[0006] In summary, there are various point-supervised oriented detection methods, each with its own advantages and disadvantages. The method based on SAM is powerful but faces challenges in dealing with cross-domain tasks such as aerial images. The method based on prior knowledge requires specific prior information, with limited generality and insufficient flexibility. The method based on the teacher-student structure, such as PointOBB, has defects such as slow pseudo-label generation, large memory requirements, and prone to memory problems. And although the PointOBB-v2 method has improved PointOBB to a certain extent, there are still limitations such as hyperparameters still depending on the targets and datasets. Summary of the Invention
[0007] To solve the problems existing in the above-mentioned prior art, the object of the present invention is to propose a point-annotation remote sensing target orientation detection method and device, which have significant improvements in aspects such as detection accuracy, memory efficiency, adaptability to different targets, and the ability to process complex scenes, providing a more effective solution for orientation remote sensing target detection.
[0008] To achieve the above object, the present invention provides the following solutions:
[0009] A point-annotation remote sensing target orientation detection method, comprising:
[0010] Obtain a point-annotation image, where the point-annotation image is an image with target center point annotations and target category annotations;
[0011] Input the point-annotation image into an improved ResNet50 model to obtain a class probability map CPM; wherein, the improved ResNet50 model includes: connecting a dilated convolutional layer to the output layer of the ResNet50 model to expand the receptive field, adding a feature extraction network based on canonical correlation analysis after the dilated convolutional layer for feature fusion, and introducing a secondary feature fusion based on a hybrid channel attention mechanism after the feature extraction network to obtain the class probability map CPM;
[0012] During the process of training the improved ResNet50 model, according to the original class probability map CPM0, obtain pseudo-labels of the target height and width, a dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, and combine the point-annotation information for positive and negative label assignment to determine the target morphology; the dynamic radius adjustment mechanism is to dynamically adjust the radius according to the different width and height pseudo-labels of different targets, the point-annotation information is the center point of the annotated target, during the process of positive and negative label assignment, draw a circle according to the annotated center point and the dynamic radius, the inside of the circle is used as the positive label, and draw a circle outside according to the distance between the center points of two adjacent identical targets and the center point of one of the targets, and the outside of the circle is used as the negative label;
[0013] Reduce the dimension of the class probability map CPM, assign weights to each data point after dimension reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, determine the width and height directions of the target according to the directions of the decomposed eigenvalues, and move outward along the width and height directions of the target. If the value at the current position is lower than the preset threshold, it indicates the target boundary, thereby generating a rotated bounding box.
[0014] Optionally, performing secondary feature fusion based on a hybrid channel attention mechanism to obtain the class probability map CPM includes:
[0015] Perform local average pooling and global average pooling on the input feature map, input the features after local pooling and global pooling into a 1D convolution for feature transformation. After 1D convolution and rearrangement, add the features after local pooling and the features after global pooling to obtain the first fusion result;
[0016] Perform unpooling on the first fusion result, multiply the first fusion result by the input feature map to obtain the second fusion result, that is, the class probability map CPM.
[0017] Optionally, obtaining the pseudo labels of the target height and width includes:
[0018] Perform convolution calculations on the horizontal and vertical directions of the original class probability map CPM0 to obtain the gradient value of each pixel point in the horizontal direction and the gradient value in the vertical direction in the original class probability map CPM0. Calculate the gradient magnitude based on the gradient values, set a first target threshold, and perform target boundary region judgment on the gradient magnitude and the first target threshold. If the pixel point with a gradient magnitude greater than the first target threshold, it is regarded as a boundary point of the target; if the pixel point with a gradient magnitude less than the first target threshold, it is regarded as a background or noise point;
[0019] Based on the boundary points of the target, obtain the outermost boundary points of the target in the horizontal and vertical directions. Take the difference between the abscissas of the leftmost and rightmost boundary points in the horizontal direction as the width pseudo label of the target, and take the difference between the ordinates of the uppermost and lowermost boundary points in the vertical direction as the height pseudo label of the target.
[0020] Optionally, obtaining the gradient value of each pixel point in the horizontal direction and the gradient value in the vertical direction in the original class probability map CPM0 includes:
[0021] Perform convolution calculation on the horizontal direction of the original class probability map CPM0, multiply the pixel values of the surrounding neighborhood pixels of each pixel point in the original class probability map CPM0 by the corresponding elements of the horizontal convolution kernel and sum them to obtain the gradient value of this pixel point in the horizontal direction;
[0022] Perform convolution calculation on the vertical direction of the original class probability map CPM0, multiply the pixel values of the surrounding neighborhood pixels of each pixel point in the original class probability map CPM0 by the corresponding elements of the vertical convolution kernel and sum them to obtain the gradient value of this pixel point in the vertical direction.
[0023] Optionally, determining the target shape includes:
[0024] Based on the pseudo labels of the target height and width, estimate the radius of the positive sample region:
[0025]
[0026] where r is the radius, k is the scaling factor, w is the width pseudo-label of the target, and h is the height pseudo-label of the target;
[0027] Adjust the radius of the positive sample region again according to the density of the target in the point-annotated image:
[0028] r adjusted = r × f(density)
[0029]
[0030] where r adjusted is the radius of the positive sample region after the second adjustment, r is the radius of the positive sample region, f(density) is the density factor, and d min is the distance between the centers of two adjacent nearest identical targets, i.e., the distance of the point annotation, and T is the density threshold.
[0031] Optionally, obtaining the target boundary includes:
[0032] Obtain multiple grid coordinates around the annotated points in the point-annotated image. According to the multiple grid coordinates, reduce the dimension of the class probability map CPM, assign weights to each data point after dimension reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, use the eigenvector corresponding to the largest eigenvalue as the main direction and its perpendicular position as the secondary direction, set a second target threshold, and move along the two directions. When the value at a certain position is lower than the second target threshold, it is regarded as the target boundary.
[0033] Optionally, constructing the covariance matrix includes:
[0034]
[0035] where C z is the covariance matrix, p i is the weight, z i is the data point, μ z is the mean of the data points, and T is the transpose.
[0036] Optionally, decomposing the eigenvalues of the covariance matrix includes:
[0037] C z v i = λ i v i
[0038] where v i is the eigenvector and λ i is the eigenvalue.
[0039] To achieve the above object, the present invention also provides a point-annotation remote sensing target orientation detection device, including:
[0040] A feature extraction module, which is used to obtain a point-annotation image. The point-annotation image is an image with a target center point annotation and a target category annotation. The point-annotation image is input into an improved ResNet50 model to obtain a class probability map CPM. Among them, the improved ResNet50 model includes: connecting a dilated convolutional layer to the output layer of the ResNet50 model to expand the receptive field, adding a feature extraction network based on canonical correlation analysis after the dilated convolutional layer for feature fusion, and introducing a secondary feature fusion based on a hybrid channel attention mechanism after the feature extraction network to obtain the class probability map CPM;
[0041] A dynamic radius adjustment module, which is used to obtain pseudo-labels of the target height and width according to the original class probability map CPM0 during the training of the improved ResNet50 model, a dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, and combining point-annotation information for positive and negative label assignment to determine the target morphology. The dynamic radius adjustment mechanism dynamically adjusts the radius according to different width and height pseudo-labels of different targets. The point-annotation information is the center point of the annotated target. During the positive and negative label assignment process, a circle is drawn according to the annotated center point and the dynamic radius. The inside of the circle is used as the positive label, and a circle is drawn outside according to the distance between the center points of two adjacent nearest same targets and the center point of one of the targets. The outside of the circle is used as the negative label;
[0042] A rotated bounding box generation module, which is used to reduce the dimension of the class probability map CPM, assign weights to each data point after dimension reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, determine the width and height directions of the target according to the directions of the decomposed eigenvalues, and move outward along the width and height directions of the target. If the value at the current position is lower than a preset threshold, it indicates the target boundary, thereby generating a rotated bounding box.
[0043] The beneficial effects of the present invention are:
[0044] The present invention proposes a new structure, which combines a feature optimization and fusion module of ResNet50, dilated convolution, and canonical correlation analysis, effectively enhancing the feature extraction ability and laying a foundation for generating high-quality class probability maps (CPMs). Secondly, a hybrid local channel attention mechanism is used to replace the projection layer, overcoming the problems of information loss and lack of dynamic adjustment ability in the projection layer, making the class probability maps CPMs richer and more accurate. Furthermore, the dynamic radius adjustment mechanism dynamically adjusts the size of the positive sample region according to the target density, reducing the overlap in high-density regions and false detections in low-density regions, significantly improving the detection accuracy, especially with obvious advantages when dealing with dense objects. At the same time, the mirror padding method for targets at the image edge also improves the detection accuracy and robustness. Finally, the t-SNE method is used to determine the target direction and boundary, which can more effectively capture the non-linear structure information of targets with complex shapes. Generally speaking, the method of the present invention has significant improvements in aspects such as detection accuracy, memory efficiency, adaptability to different targets, and the ability to handle complex scenarios, providing a more effective solution for directional remote sensing target detection. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0046] Figure 1 It is a flowchart of a point annotation remote sensing target directional detection method according to an embodiment of the present invention;
[0047] Figure 2 It is a network model structure diagram of the point annotation remote sensing target directional detection method according to an embodiment of the present invention;
[0048] Figure 3 It is a structure diagram of the feature optimization and fusion module of the point annotation remote sensing target directional detection method according to an embodiment of the present invention;
[0049] Figure 4 It is a structure diagram of the hybrid channel attention mechanism of the point annotation remote sensing target directional detection method according to an embodiment of the present invention;
[0050] Figure 5 It is a schematic diagram of the visualization results of each category of the point annotation remote sensing target directional detection method according to an embodiment of the present invention on the DOTA-V1.0 dataset. Detailed Embodiments
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0053] This embodiment discloses a point-annotation remote sensing target orientation detection method, including: obtaining a point-annotation image, which is an image with a target center point annotation and a target category annotation; inputting the point-annotation image into an improved ResNet50 model to obtain a class probability map CPM; wherein, the improved ResNet50 model includes: connecting an atrous convolution layer to the output layer of the ResNet50 model to expand the receptive field, adding a feature extraction network based on canonical correlation analysis after the atrous convolution layer for feature fusion, introducing a secondary feature fusion based on a hybrid channel attention mechanism after the feature extraction network to obtain the class probability map CPM; during the process of training the improved ResNet50 model, according to the original class probability map CPM0, obtaining pseudo-labels of the target height and width, a dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, and combining the point-annotation information for positive and negative label assignment to determine the target shape; the dynamic radius adjustment mechanism dynamically adjusts the radius according to the different width and height pseudo-labels of different targets, the point-annotation information is the center point of the annotated target, during the process of positive and negative label assignment, a circle is drawn according to the annotated center point and the dynamic radius, the inside of the circle is used as the positive label, and a circle is drawn outside according to the distance between the center points of the two nearest identical targets and the center point of one of the targets; the class probability map CPM is dimensionally reduced, weights are assigned to each data point after dimensional reduction, a covariance matrix is constructed, the eigenvalues of the covariance matrix are decomposed, the width and height directions of the target are determined according to the directions of the decomposed eigenvalues, and moving outward along the width and height directions of the target, if the value at the current position is lower than a preset threshold, it indicates the target boundary, thereby generating a rotated bounding box.
[0054] Specifically:
[0055] As Figure 1 shown, this embodiment discloses a point-annotation remote sensing target orientation detection method, including the following steps:
[0056] Obtaining a dataset picture with point annotations;
[0057] Input the image into the feature extraction network based on canonical correlation analysis, and obtain the class probability map (CPM) of the image based on the hybrid channel attention mechanism;
[0058] Based on the class probability map (CPM), obtain the pseudo-labels of the target height and width through the Sobel operator;
[0059] A dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, as well as the marked center point information, to complete the positive and negative label assignment;
[0060] Reduce the dimension of the class probability map (CPM) by the t-distributed stochastic neighbor embedding method, assign a weight to each data point after dimensionality reduction, calculate the covariance matrix, and perform eigenvalue decomposition on the covariance matrix;
[0061] Generate a rotated bounding box through the point annotation information, decomposition results, and set threshold.
[0062] Further, perform secondary feature fusion based on the hybrid channel attention mechanism to obtain the class probability map (CPM), including: performing local average pooling and global average pooling on the input feature map, inputting the features after local pooling and global pooling into a 1D convolution for feature transformation, after 1D convolution and rearrangement, adding the features after local pooling and the features after global pooling to obtain the first fusion result; performing an anti-pooling operation on the first fusion result, and multiplying the first fusion result with the input feature map to obtain the second fusion result, which is the class probability map (CPM).
[0063] Specifically:
[0064] Step 1: Before inputting the point annotation dataset into the trained PointOOD model, it is necessary to load the weights of the trained PointOOD model;
[0065] The point annotation remote sensing target orientation detection method provided by the present invention does not rely on the dataset with rotated bounding box annotations, but the dataset that annotates the target center point and category. Learn features from single-point annotations and generate rotated bounding boxes to achieve orientation detection. Before using the model, it is necessary to train the model according to a certain strategy, load the trained model weights during use, and then infer the orientation detection results based on the annotation information.
[0066] Step 2: The process of inputting the image into the feature extraction network based on canonical correlation analysis and obtaining the class probability map (CPM) of the image based on the hybrid channel attention mechanism includes: further fusing the features extracted by RESNET through the feature optimization and fusion network based on canonical correlation analysis; further fusing semantic information and expression ability through the hybrid channel attention mechanism based on global attention and local attention, and mapping it to the category probability space to generate the class probability map (CPM) of the image;
[0067] As a specific embodiment: for example Figure 3 As shown, each of the three output layers of ResNet50 is connected to three consecutive dilated convolutions. Dilated convolution is a special type of convolutional layer that allows the network to expand the receptive field without increasing the number of parameters. This means that the network can capture image information in a larger area without adding additional computational burden. Finally, the three-layer outputs are subjected to feature fusion, and the method used is canonical correlation analysis. Canonical correlation analysis can help fuse features from different levels by analyzing the correlations between features to select the most informative feature combinations, thereby improving the quality of feature representation, reducing the dimensionality of features, and removing redundant features. This helps reduce computational complexity and improve the efficiency of the model. By selecting the most relevant features, overfitting can also be reduced, which helps improve the generalization ability of the model, making the model perform better on new data. Canonical correlation analysis provides a statistical measure of the relationship between features, which helps understand the structure of the data and the relationship between features, thereby increasing the interpretability of the model. Let the feature vectors be:
[0068] X = (x1, x2, … xp)
[0069] Y = (y1, y2, … yq)
[0070] After being standardized, they are denoted as X * and Y * , and then let the linear combinations of the two sets of variables be respectively:
[0071]
[0072] where X * and Y * are the vectors of X and Y after being standardized, denoted as:
[0073]
[0074] ρ(U, V) represents the canonical correlation coefficient between the two sets of variables U and V, and its value range is [-1, 1]. When ρ(U, V) = 1, it means that there is a perfect positive correlation between U and V, that is, the change of U can be completely explained by the change of V, and vice versa. When ρ(U, V) = -1, it represents a perfect negative correlation. conv represents covariance, and var represents variance.
[0075] Then the optimization feature problem can be expressed as That is, in the image recognition task, one set of features is the color information of the image, and the other set is the texture information. Through canonical correlation analysis, we can find the linear combination of color features and texture features, discover the correlation between color changes and texture details, and provide a basis for optimizing features. In terms of feature selection, canonical correlation analysis can determine the correlation between feature pairs and help determine which features are more important for improving model performance. For highly correlated feature pairs, retaining one can represent this part of the information and remove redundant features. The two features must meet the constraints at the same time:
[0076] var(U)=a T ∑ XX a=1
[0077] var(V)=b T ∑ XX b=1
[0078] The linear correlation coefficient obtained by solving the optimization problem is set as a * and b * The fused feature is represented as U * =Xa * and V * =Yb * In this way, high-quality fusion features can be obtained, and then the three-layer outputs are fused in pairs. At this time, the output feature map contains features from low-level to high-level, rich semantic information and strong expression ability, and through the dimension reduction of canonical correlation analysis, the feature map is more compact and efficient. At this time, the class probability map CPM obtained by the attention mechanism will be richer and more accurate, which is convenient for subsequent estimation of boundaries and angles.
[0079] The projection layer usually uses a fixed linear transformation method to map high-dimensional features to low-dimensional space. This simple mapping strategy is prone to information loss. When processing complex image data, the rich details and semantic information in high-dimensional features may not be fully retained, which affects the subsequent accurate detection and positioning of target objects. Moreover, the projection layer lacks the ability to dynamically adjust different features. It uses a unified mapping method for all input features and cannot adaptively focus on important features according to the characteristics of the target object and scene changes. This is particularly limited when facing target objects with different feature distributions and complex backgrounds. In contrast, the hybrid channel attention mechanism has obvious advantages. The hybrid channel attention mechanism can effectively enhance the expression of local features by dividing the feature map into local areas for channel-level attention calculations, and can integrate global information so that the model can grasp the overall situation while paying attention to details. Its lightweight design can also improve performance without increasing too much computational burden. It can highlight key features and improve the adaptability of the model to complex scenes, which is expected to overcome the shortcomings of the projection layer and improve the accuracy and efficiency of target detection. The structure diagram of the hybrid channel attention mechanism, such asFigure 4 as shown
[0080] The hybrid channel attention mechanism first processes the input feature map through local average pooling and global average pooling. Local pooling focuses on the features of local regions, while global pooling captures the statistical information of the entire feature map. The features after local pooling and the features after global pooling both go through a 1D convolution (Conv1d) for feature transformation. 1D convolution is used to compress the feature channels while keeping the spatial dimensions unchanged. For the features after global pooling, after 1D convolution and rearrangement, they are combined with the local pooling features through an "addition" operation. This step fuses the global context information in the feature map. Finally, the feature map processed by local and global attention goes through an unpooling (UNAP) operation again to restore to the original spatial dimensions. Then it is combined with the original input features through a "multiplication" operation. This process is equivalent to a kind of feature selection, strengthening the attention to useful features.
[0081] Furthermore, obtaining the pseudo labels of the target height and width includes: performing convolution calculations on the horizontal and vertical directions of the original class probability map CPM0 to obtain the gradient values of each pixel point in the horizontal direction and the gradient values in the vertical direction in the original class probability map CPM0, calculating the gradient magnitude based on the gradient values, setting a first target threshold, and performing target boundary region judgment on the gradient magnitude and the first target threshold. If the pixel point with a gradient magnitude greater than the first target threshold, it is regarded as a boundary point of the target. If the pixel point with a gradient magnitude less than the first target threshold, it is regarded as a background or noise point; based on the boundary points of the target, obtain the outermost boundary points of the target in the horizontal and vertical directions, take the difference between the abscissas of the leftmost and rightmost boundary points in the horizontal direction as the width pseudo label of the target, and take the difference between the ordinates of the uppermost and lowermost boundary points in the vertical direction as the height pseudo label of the target.
[0082] Furthermore, obtaining the gradient values of each pixel point in the horizontal direction and the gradient values in the vertical direction in the original class probability map CPM0 includes: performing convolution calculations on the horizontal direction of the original class probability map CPM0, multiplying the pixel values of the surrounding neighborhood pixels of each pixel point in the original class probability map CPM0 with the corresponding elements of the horizontal convolution kernel and then summing to obtain the gradient value of this pixel point in the horizontal direction; performing convolution calculations on the vertical direction of the original class probability map CPM0, multiplying the pixel values of the surrounding neighborhood pixels of each pixel point in the original class probability map CPM0 with the corresponding elements of the vertical convolution kernel and then summing to obtain the gradient value of this pixel point in the vertical direction.
[0083] Specifically:
[0084] Step 3: The process of obtaining the pseudo - labels of the target height and width based on the class probability map CPM through the Sobel operator includes: The Sobel operator processes the class probability map CPM of the picture and calculates the pseudo - labels of the target height and width;
[0085] Apply the Sobel operator to the horizontal and vertical directions of the class probability map CPM for convolution calculation;
[0086] The convolution kernel of the Sobel operator in the horizontal direction is usually
[0087] For each pixel point in the class probability map CPM, multiply the pixel values in its surrounding neighborhood by the corresponding elements of the horizontal - direction convolution kernel and sum them to obtain the gradient value of this pixel point in the horizontal direction;
[0088] The convolution kernel of the Sobel operator in the vertical direction is usually
[0089] Similarly, for each pixel point in the class probability map CPM, multiply the pixel values in its surrounding neighborhood by the corresponding elements of the vertical - direction convolution kernel and sum them to obtain the gradient value of this pixel point in the vertical direction;
[0090] For each pixel point in the class probability map CPM, calculate the gradient magnitude G according to its gradient values in the horizontal and vertical directions. The formula for calculating the gradient magnitude is where G x is the gradient value in the horizontal direction, and G y is the gradient value in the vertical direction. This gradient magnitude reflects the severity of the image change at this pixel point, and there is usually a relatively large gradient magnitude at the target boundary;
[0091] Set a suitable threshold to perform threshold processing on the calculated gradient magnitude. Consider the pixel points with gradient magnitude greater than the threshold as the boundary points of the target, and the pixel points with gradient magnitude less than the threshold are regarded as background or noise points and are ignored. Through threshold processing, the boundary region of the target can be initially screened out to reduce the interference of non - target regions on subsequent calculations;
[0092] Among the boundary points after threshold processing, find the outermost boundary points of the target in the horizontal and vertical directions. The difference in the abscissas of the left - most and right - most boundary points in the horizontal direction is the width pseudo - label w of the target, and the difference in the ordinates of the top - most and bottom - most boundary points in the vertical direction is the height pseudo - label h of the target.
[0093] Furthermore, determining the target morphology includes: Based on the pseudo - labels of the target height and width, estimating the radius of the positive - sample region:
[0094]
[0095] Wherein, r is the radius, k is the scaling factor, w is the width pseudo-label of the target, and h is the height pseudo-label of the target;
[0096] According to the density of the target in the point-annotated image, the radius of the positive sample region is adjusted again:
[0097] r adjusted = r × f(density)
[0098]
[0099] Wherein, r adjusted is the radius of the positive sample region after secondary adjustment, and r is the radius of the positive sample region.
[0100] Specifically:
[0101] Step 4: Through the dynamic radius adjustment mechanism that can dynamically adjust the positive label assignment radius for different target pseudo-labels, and the annotated center point information, the process of positive and negative label assignment includes: Based on considering the differences in height and width of different targets and the target density, dynamically adjust the radius of positive label assignment based on pseudo-labels, and complete positive and negative label assignment based on point annotation information;
[0102] First, obtain the pseudo-labels w and h from the class probability map CPM through the Sobel operator to estimate the radius:
[0103]
[0104] Then consider the density of the target object in the image, and then adjust the radius of the positive sample region again based on the estimated density:
[0105] r adjusted = r × f(density)
[0106]
[0107] When the target density is high, reduce the radius to avoid overlap, and when the target density is low, increase the radius to cover more areas. Use the calculated dynamic radius to define the positive sample region to ensure that the positive sample region can accurately reflect the actual size and shape of the target object.
[0108] Furthermore, obtaining the target boundary includes: obtaining multiple grid coordinates around the annotated points in the point-annotated image, reducing the dimension of the class probability map CPM according to the multiple grid coordinates, assigning weights to each data point after dimension reduction, constructing a covariance matrix, decomposing the eigenvalues of the covariance matrix, taking the eigenvector corresponding to the largest eigenvalue as the main direction and its perpendicular position as the secondary direction, setting the second target threshold, moving along the two directions, and when the value at a certain position is lower than the second target threshold, it is regarded as the target boundary.
[0109] Furthermore, constructing the covariance matrix includes:
[0110]
[0111] where C z is the covariance matrix, p i is the weight, z i is the data point, and μ z is the mean of the data points.
[0112] Furthermore, decomposing the eigenvalues of the covariance matrix includes:
[0113] C z v i = λ i v i
[0114] where v i is the eigenvector and λ i is the eigenvalue.
[0115] Specifically:
[0116] Step 5: Dimension reduction of the class probability map CPM by the t-distributed stochastic neighbor embedding method, assigning a weight to each data point after dimension reduction, calculating the covariance matrix, and decomposing the eigenvalues of the covariance matrix; generating a bounding box through the point annotation information, decomposition results, and setting a threshold.
[0117] As a specific embodiment, after obtaining the positive and negative label assignment regions in Step 4, it is necessary to infer the direction and boundary of the target based on the labeled center point and the class probability map CPM. We first select a 7×7 grid around the labeled point, and the coordinates of the 49 points are respectively:
[0118] (x, y), x ∈ [-3, 3], y ∈ [-3, 3]
[0119] Then, use the t-SNE method to first reduce the dimension of the class probability map CPM. In the high-dimensional space, for the CPM data points, t-SNE calculates the similarity between data points through a specific formula to construct a joint probability distribution. The specific formula is:
[0120]
[0121] This formula is based on the Euclidean distance ||x i - x j || i - x j || 2 between data points x j and x i and, combined with the bandwidth parameter σ, calculates the conditional probability p j|i, which reflects the similarity degree among data points. By calculating for all point pairs, a joint probability distribution in the high-dimensional space is constructed, and this distribution embodies the local and global structural relationships of data points.
[0122] In the low-dimensional space, t-SNE uses the t-distribution to calculate the similarity between data points, and the formula is:
[0123]
[0124] Here, y i and y j are the mappings of the high-dimensional data points x i and x j in the low-dimensional space. The characteristics of the t-distribution make it easier for similar data points to be close in the low-dimensional space, and points in different clusters are relatively separated, which helps to reveal the internal structure of the data.
[0125] Finally, t-SNE optimizes the layout of the low-dimensional data points by minimizing a specific cost function, and this cost function is:
[0126]
[0127] where P and Q are the joint probability distributions in the high-dimensional space and the low-dimensional space respectively. Minimizing this cost function can make the similarity distributions of data points in the high-dimensional space and the low-dimensional space as close as possible, so as to retain the local and global structural relationships of the high-dimensional data in the low-dimensional space and achieve effective dimensionality reduction of the class probability map CPM.
[0128] Then, for the covariance matrix of the dimension-reduced low-dimensional data points, we assign a weight of p i to each point z i , and the mean of the data points is μ z , and the covariance matrix is defined as:
[0129]
[0130] For C z perform eigenvalue decomposition:
[0131] C z v i =λ i v i
[0132] Select the eigenvector v1 corresponding to the largest eigenvalue λ1 as the main direction. Since C z is a real symmetric matrix, it is ensured that the secondary direction is orthogonal to the primary direction. This orthogonality corresponds to the perpendicular relationship between two adjacent sides of the oriented frame. After determining the primary direction and the secondary direction, starting from the marked point, move along the two directions, and when the value at a certain position is lower than the set threshold, it indicates the target boundary.
[0133] As shown Figure 2 in the figure, the present invention also provides a point-annotated remote sensing target orientation detection device, including: a feature extraction module, configured to obtain a point-annotated image, where the point-annotated image is an image with a target center point annotation and a target category annotation, input the point-annotated image into an improved ResNet50 model to obtain a class probability map CPM; wherein, the improved ResNet50 model includes: connecting an atrous convolution layer to the output layer of the ResNet50 model to expand the receptive field, adding a feature extraction network based on canonical correlation analysis after the atrous convolution layer for feature fusion, introducing a secondary feature fusion based on a hybrid channel attention mechanism after the feature extraction network to obtain the class probability map CPM; a dynamic radius adjustment module, configured to obtain pseudo-labels of the target height and width according to the original class probability map CPM0 during the training of the improved ResNet50 model, a dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, and combining the point annotation information for positive and negative label assignment to determine the target shape; the dynamic radius adjustment mechanism is to dynamically adjust the radius according to different width and height pseudo-labels of different targets, the point annotation information is the center point of the annotated target, during the process of positive and negative label assignment, draw a circle according to the annotated center point and the dynamic radius, the inside of the circle is used as the positive label, and draw a circle outside according to the distance between the center points of the two adjacent nearest same targets and the center point of one of the targets, the outside of the circle is used as the negative label; a rotated bounding box generation module, configured to reduce the dimension of the class probability map CPM, assign weights to each data point after dimension reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, determine the width and height directions of the target according to the direction of the decomposed eigenvalues, and move outward along the width and height directions of the target. If the value at the current position is lower than a preset threshold, it indicates the target boundary, thereby generating a rotated bounding box.
[0134] Specifically:
[0135] This embodiment also provides a point-annotated remote sensing target orientation detection device. The usage method of the point-annotated remote sensing target orientation detection device can be mutually referred to in terms of effects with the point-annotated remote sensing target orientation detection method provided in the above embodiment. The point-annotated remote sensing target orientation detection device includes: a feature extraction module, a feature extraction optimization and fusion module, a hybrid channel attention mechanism, a dynamic radius adjustment mechanism, a prediction direction and boundary module;
[0136] The feature extraction module is used to extract the basic features of the target;
[0137] The feature extraction optimization and fusion module is used to extract the deep features of the target and optimize and fuse the features based on the method of canonical correlation analysis;
[0138] The hybrid channel attention mechanism is used to further fuse semantic information and spatial information, map the feature map to the category probability space, and generate the class probability map (CPM) of the picture;
[0139] The dynamic radius adjustment mechanism is used to dynamically adjust the radius assigned to the positive label based on the pseudo-labels of different target heights and widths, and better complete the positive and negative label assignment work;
[0140] The prediction direction and boundary module is used to predict the direction and boundary of each target based on the point annotation information for the high-quality class probability map (CPM), and generate a rotated bounding box.
[0141] The effects of the present invention will be further described below in conjunction with simulation experiments and the accompanying drawings.
[0142] As Figure 5 shown, the simulation conditions of the present invention are:
[0143] All models in the experiment were carried out on a server with 4 NVIDIA Tesla A100 40GB GPUs and 2 Intel Xeon 6330 CPUs. The experimental environment was Ubuntu 22.04, the programming language was Python 3.10, and the deep learning platform was PyTorch 2.0.1. To achieve fast convergence, the AdamW optimizer was used to train all models in the experiment. The basic learning rate was set to 0.005, and the cosine annealing strategy was used to adjust the learning rate. A total of 72 epochs were trained, and the learning rate decayed to 0.1 times the original every 6 epochs. The training batch size was set to 16.
[0144] The evaluation metric used the mean average precision (mAP) as the main metric to compare our method with existing alternative methods. To better evaluate the quality of the pseudo-labels generated from points, we also compared the mean intersection over union (mIoU) between the ground truth boxes and the predicted pseudo-oriented bounding boxes (obb).
[0145] Table 1 shows the influence of different fixed radii on different targets in the positive label assignment of the point annotation remote sensing target orientation detection method according to the embodiment of the present invention.
[0146] Table 1
[0147]
[0148] Table 2 shows the results of each category of the point annotation remote sensing target orientation detection method according to the embodiment of the present invention and other point annotation orientation detection methods on the DOTA-V1.0 dataset.
[0149] Table 2
[0150]
[0151] Table 3 shows the results of the point-annotation remote sensing target orientation detection method of the embodiments of the present invention and other point-annotation orientation detection methods on three datasets of DOTA-V1.0 / v1.5 / v2.0.
[0152] Table 3
[0153]
[0154]
[0155] Table 4 shows the ablation experiments of each module of the point-annotation remote sensing target orientation detection method of the embodiments of the present invention.
[0156] Table 4
[0157]
[0158] Table 5 shows the ablation experiments of the loss function weight coefficients of the point-annotation remote sensing target orientation detection method of the embodiments of the present invention.
[0159] Table 5
[0160]
[0161] Table 6 shows the ablation experiments of the scaling factor and density threshold of the point-annotation remote sensing target orientation detection method of the embodiments of the present invention.
[0162] Table 6
[0163]
[0164] Table 7 shows the ablation experiments of the t-SNE bandwidth parameters of the point-annotation remote sensing target orientation detection method of the embodiments of the present invention.
[0165] Table 7
[0166]
[0167] Table 8 shows the comparison of mIoU and detection accuracy between the point-annotation remote sensing target orientation detection method of the embodiments of the present invention and other point-annotation orientation target detection methods on three datasets of DOTA-V1.0 / v1.5 / v2.0.
[0168] Table 8
[0169]
[0170]
[0171] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A point annotation remote sensing target orientation detection method, characterized in that Including: Obtain a point-annotated image, which is an image with a target center point annotation and a target category annotation; Input the point-annotated image into the improved ResNet50 model to obtain a class probability map CPM; wherein, the improved ResNet50 model includes: connecting an atrous convolution layer to the output layer of the ResNet50 model to expand the receptive field, adding a feature extraction network based on canonical correlation analysis after the atrous convolution layer for feature fusion, and introducing a secondary feature fusion based on a hybrid channel attention mechanism after the feature extraction network to obtain the class probability map CPM; During the process of training the improved ResNet50 model, according to the original class probability map CPM0, obtain the pseudo-labels of the target height and width, a dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, and combine the point annotation information for positive and negative label assignment to determine the target morphology; the dynamic radius adjustment mechanism is to dynamically adjust the radius according to the different width and height pseudo-labels of different targets, and the point annotation information is the center point of the annotated target. During the process of positive and negative label assignment, draw a circle according to the annotated center point and the dynamic radius. The inside of the circle is used as the positive label, and draw a circle outside according to the distance between the center points of the two adjacent nearest same targets and the center point of one of the targets. The outside of the circle is used as the negative label; Reduce the dimension of the class probability map CPM, assign weights to each data point after dimension reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, determine the width and height directions of the target according to the directions of the decomposed eigenvalues, and move outward along the width and height directions of the target. If the value at the current position is lower than the preset threshold, it indicates the target boundary, thereby generating a rotated bounding box.
2. The point annotation remote sensing target orientation detection method according to claim 1, characterized in that Performing secondary feature fusion based on a hybrid channel attention mechanism to obtain the class probability map CPM includes: Perform local average pooling and global average pooling on the input feature map, input the features after local pooling and global pooling into a 1D convolution for feature transformation, after 1D convolution and rearrangement, perform an addition operation on the features after local pooling and the features after global pooling to obtain a first fusion result; Perform an anti-pooling operation on the first fusion result, and perform a multiplication operation on the first fusion result and the input feature map to obtain a second fusion result, that is, the class probability map CPM.
3. The point-annotation remote sensing target orientation detection method according to claim 1, wherein, Obtaining the pseudo-labels of the target height and width includes: Perform convolution calculations on the horizontal and vertical directions of the original class probability map CPM0 to obtain the gradient values of each pixel point in the horizontal direction and the vertical direction in the original class probability map CPM0, calculate the gradient amplitude based on the gradient values, set a first target threshold, and perform a target boundary region judgment on the gradient amplitude and the first target threshold. If the pixel point with a gradient amplitude greater than the first target threshold is regarded as the boundary point of the target, and if the pixel point with a gradient amplitude less than the first target threshold is regarded as a background or noise point; Based on the boundary points of the target, obtain the outermost boundary points of the target in the horizontal and vertical directions. Take the difference between the abscissas of the leftmost and rightmost boundary points in the horizontal direction as the width pseudo-label of the target, and take the difference between the ordinates of the uppermost and lowermost boundary points in the vertical direction as the height pseudo-label of the target.
4. The point annotation remote sensing target orientation detection method according to claim 3, characterized in that Obtaining the gradient value of each pixel point in the horizontal direction and the gradient value in the vertical direction in the original class probability map CPM0 includes: Perform convolution calculation on the horizontal direction of the original class probability map CPM0. Multiply the pixel values of the surrounding neighborhood pixels of each pixel point in the original class probability map CPM0 by the corresponding elements of the horizontal convolution kernel and sum them to obtain the gradient value of the pixel point in the horizontal direction; Perform convolution calculation on the vertical direction of the original class probability map CPM0. Multiply the pixel values of the surrounding neighborhood pixels of each pixel point in the original class probability map CPM0 by the corresponding elements of the vertical convolution kernel and sum them to obtain the gradient value of the pixel point in the vertical direction.
5. The point annotation remote sensing target orientation detection method according to claim 1, characterized in that, Determining the target morphology includes: Based on the pseudo-labels of the height and width of the target, estimate the radius of the positive sample region: Where r is the radius, k is the scaling factor, w is the width pseudo-label of the target, and h is the height pseudo-label of the target; According to the density of the target in the point annotation image, adjust the radius of the positive sample region again: r adjusted = r × f(density) Among them, r adjusted is the radius of the secondary adjustment positive sample region, r is the radius of the positive sample region, f(density) is the density factor, and d min is the distance between the centers of two adjacent nearest identical targets, that is, the distance of point annotation, and T is the density threshold.
6. The point annotation remote sensing target orientation detection method according to claim 1, characterized in that, Obtaining the target boundary includes: Obtain multiple grid coordinates around the annotation points in the point annotation image. According to the multiple grid coordinates, reduce the dimension of the class probability map CPM, assign weights to each data point after dimension reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, take the eigenvector corresponding to the largest eigenvalue as the main direction and its vertical position as the secondary direction, set the second target threshold, and move along the two directions. When the value at a certain position is lower than the second target threshold, it is regarded as the target boundary.
7. The point annotation remote sensing target orientation detection method according to claim 6, characterized in that, Constructing the covariance matrix includes: Among them, C z is the covariance matrix, p i is the weight, z i is the data point, μ z is the mean of the data points, and T is the transpose.
8. The point annotation remote sensing target orientation detection method according to claim 6, characterized in that Decomposing the eigenvalues of the covariance matrix includes: Czvi = λ i vi Among them, v i is the eigenvector, and λ i is the eigenvalue.
9. A point-annotation remote sensing target orientation detection device, characterized in that, Includes: A feature extraction module for obtaining a point annotation image, which is an image with the annotation of the target center point and the target category. Input the point annotation image into the improved ResNet50 model to obtain the class probability map CPM; among them, the improved ResNet50 model includes: connecting an atrous convolution layer to the output layer of the ResNet50 model to expand the receptive field, adding a feature extraction network based on canonical correlation analysis after the atrous convolution layer for feature fusion, and introducing a secondary feature fusion based on a hybrid channel attention mechanism after the feature extraction network to obtain the class probability map CPM; A dynamic radius adjustment module, which is used to obtain pseudo-labels of target height and width according to the original class probability map CPM0 during the training of the improved ResNet50 model, a dynamic radius adjustment mechanism for dynamically adjusting the positive label assignment radius for different target pseudo-labels, and combining point annotation information for positive and negative label assignment to determine the target form; the dynamic radius adjustment mechanism is to dynamically adjust the radius according to the different width and height pseudo-labels of different targets, the point annotation information is the center point of the annotated target, and during the process of positive and negative label assignment, a circle is drawn according to the annotated center point and the dynamic radius, and the inside of the circle is used as the positive label, and a circle is drawn outside according to the distance between the center points of the two nearest identical targets and the center point of one of the targets, and the outside of the circle is used as the negative label; A rotated bounding box generation module, which is used to reduce the dimension of the class probability map CPM, assign weights to each data point after dimensionality reduction, construct a covariance matrix, decompose the eigenvalues of the covariance matrix, determine the width and height directions of the target according to the directions of the eigenvalues obtained by the decomposition, and move outward along the width and height directions of the target. If the value at the current position is lower than the preset threshold, it indicates the target boundary, thereby generating a rotated bounding box.
Citation Information
Patent Citations
Weak supervision remote sensing target detection method based on hybrid hole convolution
CN112183414A
Salient target detection method based on adaptive feature fusion
CN117115601A
Anchor-frame-free remote sensing image rotating target detection method under attention mechanism
CN118379617A
General target detection method for adaptive attention guidance mechanism
WO2021139069A1