Remote sensing image target counting method based on sorting learning and double-branch network

Through the method based on sorting learning and dual-branch network, the annotation is automatically generated and the overall and local information of the image is paid attention to, which solves the problems of high annotation cost, large computing resources and low counting accuracy of the existing remote sensing image target counting method, and achieves efficient and accurate target counting.

CN120125503APending Publication Date: 2025-06-10SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510068532.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing remote sensing image target counting methods have problems such as high annotation cost, large computing resources and low counting accuracy, especially in complex scenarios and large-scale datasets.

Method used

Using the method based on sorting learning and dual-branch network, by constructing global and local sorting supervision information, automatically generating annotations, the network model can focus on the overall and local information of the image at the same time, and finally output count values.

Benefits of technology

Reduces annotation costs, improves counting accuracy, and enables efficient target counting in complex scenarios and large-scale datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125503A_ABST
    Figure CN120125503A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image target counting method based on sorting learning and a double-branch network, and the method comprises the following steps: constructing global sorting supervision information for a data set of remote sensing image target counting; constructing local sorting supervision information for a data set of remote sensing image target counting; constructing a double-branch network model according to the global sorting supervision information and the local sorting supervision information, and training the double-branch network model by using a training set in the data set until the double-branch network model is trained; inputting the images of the test set in the data set into the double-branch network model to obtain a predicted sorting sequence of the images in the data set; double-branch sorting loss is calculated, gain loss calculation is accumulated through approximate normalization depreciation, and a loss function of the double-branch network model is obtained in combination with the sum of the double-branch sorting loss; and performing least square regression fitting by using the image-level labels of the verification set in the data set, and mapping the prediction sequence of the double-branch network model to remote sensing image target counting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image target counting. Specifically, it particularly relates to a remote sensing image target counting method that globally combines a local sorting network with a dual-branch sorting learning model. Background Art

[0002] Remote sensing target counting is a challenging task in the field of computer vision. There are mainly three deep learning solutions for existing target counting tasks: (1) target detection-based methods, (2) density map regression-based methods, and (3) global regression-based methods. The target counting method based on target detection is relatively intuitive. The density map regression is the mainstream method for current target counting tasks. The global regression-based method directly regresses the total number of target values in the whole image.

[0003] The steps of the target detection-based method are as follows: First, a target detector is constructed by learning image features. Then, the image or video to be counted is input into the trained target detection model. Next, the bounding boxes and class information of the detected targets are output. Finally, the number of targets is determined by counting the number of bounding boxes. However, due to the complex background information contained in the image, a large number of target objects, and the overlap between targets, it is difficult for the target detection algorithm to accurately detect and count the targets. For example, in scenes with dense crowds in the monitoring area, and scenes with large-scale buildings and vehicles in remote sensing images.

[0004] The steps of the density map regression-based method are as follows: First, a true density map is obtained through two steps of manual annotation and Gaussian filtering processing. Then, a deep learning model is built to output a predicted density map, which can reflect the position distribution of the targets. Next, a pixel-by-pixel regression mapping is performed between the true density map and the predicted density map. Finally, the pixel values of the predicted density map are summed to obtain the final predicted number of targets.

[0005] The density map regression method usually has high accuracy. However, the biggest limitation of this method is that pixel-level labels are required in the generation process of the true density map, that is, manual annotation needs to label the positions of targets in the whole image one by one, which usually takes a lot of time and manpower, especially for complex scenes and large-scale data sets. In addition, density map regression usually requires high computing resources and time during the model training process.

[0006] The global regression-based method directly regresses the total number of target values in the whole image. The mainstream global regression-based method only requires image-level labels, that is, annotating the total number of targets in the image. Compared with the target detection-based method and the density map regression-based method, the global regression-based method greatly reduces the annotation cost. Summary of the Invention

[0007] The object of the present invention is to overcome the problems of the prior art, and provide a method for remote sensing image target counting based on ranking learning and a dual-branch network. With the help of ranking learning and large model technology, the network model automatically generates annotations, and the network model can simultaneously focus on the overall and local information of the image, and finally outputs the count value, reducing the annotation cost and improving the counting accuracy.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A method for remote sensing image target counting based on ranking learning and a dual-branch network includes the following steps:

[0010] For the dataset of remote sensing image target counting, construct global ranking supervision information;

[0011] For the dataset of remote sensing image target counting, construct local ranking supervision information;

[0012] According to the global ranking supervision information and the local ranking supervision information, construct a dual-branch network model, and use the training set in the dataset to train the dual-branch network model until the dual-branch network model is trained well;

[0013] Input the images in the test set of the dataset into the dual-branch network model to obtain the predicted ranking order of the images in the dataset;

[0014] Calculate the dual-branch ranking loss. Through the calculation of the approximate normalized discounted cumulative gain (NDCG) loss, combine the dual-branch ranking loss and to obtain the loss function of the dual-branch network model;

[0015] Use the image-level labels in the validation set of the dataset for least squares regression fitting to map the predicted ranking of the dual-branch network model to remote sensing image target counting.

[0016] Further, constructing the global ranking supervision information is specifically as follows:

[0017] The global ranking supervision information provides label information for the global network branch. The label-making process is obtained through manual annotation; instead of adopting image-level labels, that is, the label information is the sum of the number of all targets in the image, the global ranking supervision information is adopted, that is, only the size ranking relationship of the number of targets between images is labeled;

[0018] In the annotation process, the adopted annotation strategy is divided into two steps: determine the number of levels according to the number of targets in the image, and the annotator observes the number of targets and assigns a level to each image; then divide the images with the same level again, and one level corresponds to two labels, indicating two ranking labels of the images corresponding to one level.

[0019] Further, constructing the local ranking supervision information is specifically as follows:

[0020] The local sorting supervision information provides label information for the local network branch, and the label-making process is automatically obtained through annotation by the target segmentation model SAM; instead of using image-level labels, local sorting supervision information is adopted, and a whole image is segmented into multiple square sub-images of equal size, and the sorting relationship of the number of internal targets between the sub-images is labeled.

[0021] Obtaining local sorting supervision information through the target segmentation model SAM is divided into two steps: passing the image into the target segmentation model SAM, visualizing the output result of the Encoder layer to generate a neural network heat map; according to the average value of the heat map, separating the heat map into foreground and background, the foreground is the pixels greater than the average value, the background is the pixels less than the average value, setting the foreground value to 1 and the background value to 0.

[0022] After the foreground and background are separated, it is pre-assumed that within an image, the sizes of different instances of the same type of target are approximately equal, so the number of pixels occupied by each instance is also approximate, and the number of pixels in each sub-image represents the number of targets in that area. Sorting the pixel sums of the sub-images yields the required local sorting supervision information.

[0023] Furthermore, a dual-branch network model is constructed, specifically:

[0024] The dual-branch network model uses a backbone network based on Transformer, and after the image is feature-extracted, it is divided into a local network branch and a global network branch;

[0025] In the local network branch, the feature map is divided into multiple square sub-feature maps of equal size, and the output result is obtained through a fully connected layer; the number of neurons in the last layer of the fully connected layer is 1, and the sigmoid function is used for processing to control the output result within the interval [0, 1];

[0026] In the global network branch, the feature map is flattened, feature concatenation is performed with the output result of the local network branch, and then it passes through a fully connected layer. The number of neurons in the last layer of the fully connected layer is 1, and the sigmoid function is used for processing;

[0027] The final processing result of the dual-branch network model represents the predicted sorting order of the image in the entire dataset, ranging from [0, 1], and then the predicted result is linearly transformed and mapped to the true count through the least squares regression fitting method.

[0028] Furthermore, the calculation of the approximate normalized discounted cumulative gain NDCG loss is specifically:

[0029] The normalized discounted cumulative gain (NDCG) is used to measure the quality of ranking. The approximate NDCG loss function aims to construct an ideal ranking by optimizing the NDCG retrieval metric, using the concepts of "relevance" ψ rel and "similarity" ψ sim to build the ranking;

[0030] In a batch, for the training samples X = {x 1 ,..., x n}, where the retrieved image is The label value and predicted value of the retrieved image are y i and The query image corresponds to q = {x 1 ,..., x i-1 , x i+1 ,..., x n}. Taking the query image x q , the label value and predicted value of the query image are y q and

[0031] The calculation formulas for "relevance" ψ rel and "similarity" ψ sim are as follows:

[0032]

[0033] In the formula, ψ rel ∈[0,1], ψ sim ∈[0,1]; y max and y min represent the maximum label value and minimum label value in the entire dataset respectively; only the images with the largest and smallest label values in a batch are used as the retrieved images to avoid calculation conflicts.

[0034] Furthermore, sort "relevance" ψ rel and "similarity" ψ sim according to the relevance score. The sorted list of "relevance" corresponds to the ideal ranking, and the sorted list of "similarity" corresponds to the ranking predicted by the dual-branch network model. The calculation formula for the approximate NDCG is as follows:

[0035] score(x j ) = T × ψ rel (j);

[0036]

[0037] NDCG(x j) = T × ψ rel (j);

[0038] Where ψ(j) rank represents the ranking value of the query image j in the sorted list; score represents the correlation score between the query image and the retrieved image, which is calculated by multiplying the "relevance" ψ rel by the hyperparameter T.

[0039] Furthermore, in order to calculate the Discounted Cumulative Gain (DCG), the ranking values of each retrieved image in the two lists need to be calculated. Since the Normalized Discounted Cumulative Gain (NDCG) is non-differentiable and gradient descent cannot be directly used for optimization, the sigmoid function is used to approximate the ranking values predicted by the double-branch network model, enabling gradient calculation of the function through this transformation;

[0040] The "relevance" ranking value ψ rel (j) rank and the "similarity" ranking value ψ sim (j) rank are calculated as follows:

[0041]

[0042] Where j and k respectively represent the j-th and k-th images in the retrieved images; ψ rel (j) and ψ sim (j) respectively represent the values of the j-th retrieved image in the two sorted lists; is the indicator function, σ(b - a) is the approximation of the indicator function, and θ = 10 is set to control the approximation degree of the sigmoid function to the indicator function;

[0043] The definition of the approximate Normalized Discounted Cumulative Gain (NDCG) loss is as follows:

[0044] Loss ApproxNDCG = ∑ i 1 - NDCG(x i , q).

[0045] Furthermore, by combining the double-branch ranking loss and, the loss function of the double-branch network model is obtained, specifically:

[0046] Both the global network branch and the local network branch will generate ranking losses: for the local network branch, the output result after passing through the Sigmoid function is used to calculate the loss with the local ranking supervision information; in the global network branch, after feature concatenation with the local network branch and passing through the sigmoid function, its final result is used to calculate the loss with the global ranking supervision information;

[0047] The combination of the global ranking loss and the local ranking loss has a positive feedback effect on the dual-branch network model. The overall loss function of the dual-branch network model is defined as:

[0048] Loss = Loss global + λ * Loss local ;

[0049] In the formula, Loss global is the approximate normalized discounted cumulative gain (NDCG) loss value of the global ranking; Loss local is the approximate normalized discounted cumulative gain (NDCG) loss value of the local ranking; λ is the weight.

[0050] Furthermore, the least squares regression fitting is specifically as follows:

[0051] The output result of the dual-branch network model represents the ranking result of the target quantity of the image in this dataset; only by using the standard least squares regression fitting method to map the predicted ranking to the dual-branch network model count, the least squares regression fitting can be carried out using the image-level labels of the validation set. The formula is as follows:

[0052]

[0053] In the formula, η n represents the ranking order of the nth image predicted by the dual-branch network model, and the range is in the interval [0, 1]; represents the average value of the predicted ranking; y n represents the true label value; α and β are the two parameters in the linear transformation respectively; y′ is the predicted value of the true count; N represents the number of images in the training set.

[0054] Furthermore, the datasets for remote sensing image target counting include: RSOC_building dataset, VisDrone2019 People dataset, and VisDrone2019 Vehicle dataset;

[0055] Two evaluation metrics, the mean absolute error (MAE) and the root mean square error (RMSE), are used to evaluate the performance of remote sensing image target counting. The definitions of the two metrics are as follows:

[0056]

[0057] In the formula, K is the number of test images; is the predicted count of the ith image; C i is the true count of the ith image.

[0058] Compared with the prior art, the present invention utilizes ranking learning and large model technology. The network model automatically generates annotations, and can simultaneously focus on both the overall and local information of the image, and finally outputs a count value, reducing the annotation cost and improving the counting accuracy. The present invention uses combined-level labels, that is, ranking labels indicating the number of objects in a group of images, which are called ranking supervision information. This further reduces the annotation cost and achieves competitive counting accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a schematic flowchart of a remote sensing image object counting method based on ranking learning and a dual-branch network.

[0060] Figure 2 It is a visualization schematic diagram of the heatmap average separation method. (a) is the original image, (b) is the SAM heatmap, and (c) is the schematic diagram after the heatmap average is separated.

[0061] Figure 3 It is a network branch diagram of the dual-branch network model.

[0062] Figure 4 It is a ranking loss branch diagram of the dual-branch network model.

[0063] Figure 5 It is a schematic diagram of a dataset example. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following further describes the remote sensing image object counting method based on ranking learning and a dual-branch network of the present invention with reference to the accompanying drawings and specific embodiments.

[0065] Please refer to Figure 1 , the present invention discloses a remote sensing image object counting method based on ranking learning and a dual-branch network, including the following steps:

[0066] For the dataset of remote sensing image object counting, construct global ranking supervision information;

[0067] For the dataset of remote sensing image object counting, construct local ranking supervision information;

[0068] According to the global ranking supervision information and the local ranking supervision information, construct a dual-branch network model, and use the training set in the dataset to train the dual-branch network model until the dual-branch network model is trained well;

[0069] Input the images in the test set of the dataset into the dual-branch network model to obtain the predicted ranking order of the images in the dataset;

[0070] Calculate the dual-branch ranking loss. Through approximate normalized discounted cumulative gain (NDCG) loss calculation, combine the dual-branch ranking loss and, to obtain the loss function of the dual-branch network model;

[0071] Using the image-level labels of the validation set in the dataset for least squares regression fitting, mapping the prediction ranking of the dual-branch network model to the target count of remote sensing images.

[0072] In step S1, construct global ranking supervision information.

[0073] The global ranking supervision information provides label information for the global network branch, and the label-making process is obtained through manual annotation. Instead of adopting image-level labels, that is, the label information is the sum of the number of all targets in the image, this invention adopts global ranking supervision information, that is, only the size ranking relationship of the number of targets between images is annotated.

[0074] During the annotation process, the annotation cost of the global ranking supervision information is significantly reduced compared to that of image-level labels. The annotation strategy adopted in this invention is divided into two steps: (1) Roughly set 5 levels (very few, few, medium, many, very many) according to the number of targets in the image, and the annotator gives each image a level by roughly observing the number of targets; (2) Divide the images with the same level again, and one level corresponds to two labels. For example, the "very few" level corresponds to labels 0 and 1, the "few" level corresponds to labels 2 and 3, the "medium" level corresponds to labels 4 and 5, the "many" level corresponds to labels 6 and 7, and the "very many" level corresponds to labels 8 and 9, and a total of ten levels are divided, corresponding to the ranking labels of the images.

[0075] Compare the annotation complexity of image-level labels and global ranking supervision information, that is, the number of annotation times required to annotate the entire dataset. Suppose a dataset has m images, and on average each image has n targets, where n is often much larger than 2. The annotation complexity of image-level labels is n×m, while the annotation complexity of global ranking supervision information is 2×m. It can be seen that the more the number of targets in the dataset, relatively speaking, the more cost-effective it is to annotate the global ranking supervision information.

[0076] In step S2, construct local ranking supervision information.

[0077] The local ranking supervision information provides label information for the local network branch, and the label-making process is automatically obtained through the target segmentation model SAM (Segment Anything Model). Instead of adopting image-level labels, this invention adopts local ranking supervision information, splitting a whole image into multiple square sub-images of equal size, and annotating the size ranking relationship of the number of internal targets between the sub-images. For example, splitting an entire image with 32 targets into 4 sub-images, and the number of targets in the 4 sub-images is [16, 2, 8, 6], then the local ranking supervision information of this whole image is [4, 1, 3, 2].

[0078] As Figure 2 shown, the target segmentation model SAM provides an automated method for obtaining local ranking supervision information, and its process is divided into two steps: (1) Input the image into the target segmentation model SAM, visualize the output result of the Encoder layer, and generate a neural network heat map, as Figure 2 (b) shown. (2) According to the average value of the heat map, separate the heat map into foreground and background, where the foreground is the pixel greater than the average value, and the background is the pixel less than the average value. Set the foreground value to 1 and the background value to 0, as Figure 2 (c) shown.

[0079] The density of foreground pixels reflects the density of the corresponding real target. After separating the foreground and background, the present invention pre-assumes that within an image, the sizes of different instances of the same type of target are approximately equal. Therefore, the number of pixels occupied by each instance is also approximate. The number of pixels in each sub-image represents the number of targets in that area. Sorting the sum of pixels of the sub-images yields the required local ranking supervision information.

[0080] In step S3, a dual-branch network model is constructed.

[0081] As Figure 3 shown, the dual-branch network model first adopts a backbone network based on Transformer. After feature extraction, the image is divided into a local network branch and a global network branch.

[0082] In the local network branch, the feature map is divided into multiple square sub-feature maps of equal size, and the output result is obtained through a fully connected layer. The number of neurons in the last layer of the fully connected layer is 1, and the sigmoid function is used for processing to control the output result within the interval [0, 1].

[0083] In the global network branch, the feature map is flattened, feature concatenation is performed with the output result of the local network branch, and then it passes through a fully connected layer. The number of neurons in the last layer of the fully connected layer is 1, and the sigmoid function is used for processing.

[0084] The final processing result of the dual-branch network model represents the predicted ranking order of the image in the entire dataset, ranging from [0, 1]. Then, through the least squares regression fitting method, the predicted result is linearly transformed and mapped to the real count.

[0085] In step S5, the approximate normalized discounted cumulative gain (NDCG) loss is calculated.

[0086] The Normalized Discounted Cumulative Gain (NDCG) is used to measure the quality of ranking. The approximate NDCG loss function aims to construct an ideal ranking by optimizing the NDCG retrieval metric, using the concepts of "relevance" ψ rel and "similarity" ψ sim to build the ranking.

[0087] In a batch, for the training samples X = {x 1 ,..., x n}, where the retrieved image is The label value and predicted value of the retrieved image are y i and The query image corresponds to q = {x 1 ,..., x i-1 , x i+1 ,..., x n}. Taking the query image x q , the label value and predicted value of the query image are y q and

[0088] The calculation formulas for "relevance" ψ rel and "similarity" ψ sim are as follows:

[0089]

[0090] In the formula, ψ rel ∈[0,1], ψ sim ∈[0,1]; y max and y min represent the maximum label value and minimum label value in the entire dataset respectively. Using only the images with the maximum and minimum label values as the retrieved images in a batch can avoid calculation conflicts.

[0091] Sort the "relevance" ψ rel and "similarity" ψ sim according to the relevance scores. The sorted list of "relevance" corresponds to the ideal ranking, and the sorted list of "similarity" corresponds to the ranking predicted by the dual-branch network model. The calculation formula for the approximate NDCG is as follows:

[0092] score(x j ) = T × ψ rel (j);

[0093]

[0094] NDCG(x j ) = T × ψ rel (j);

[0095] Wherein, ψ(j) rank represents the ranking value of the query image j in the sorted list; score represents the correlation score between the query image and the retrieved image, which is calculated by multiplying the "relevance" ψ rel by the hyperparameter T.

[0096] To calculate the Discounted Cumulative Gain (DCG), it is also necessary to calculate the ranking values of each retrieved image in the two lists. At the same time, since the Normalized Discounted Cumulative Gain (NDCG) is non-differentiable, gradient descent cannot be directly used for optimization. Therefore, the sigmoid function is used to approximately represent the ranking values predicted by the dual-branch network model, and this transformation enables the function to perform gradient calculation.

[0097] The "relevance" ranking value ψ rel (j) rank and the "similarity" ranking value ψ sim (j) rank are calculated as follows:

[0098]

[0099] Wherein, j and k respectively represent the j-th and k-th images in the retrieved images; ψ rel (j) and ψ sim (j) respectively represent the values of the j-th retrieved image in the two sorted lists; is the indicator function, σ(b - a) is the approximate value of the indicator function, and θ = 10 is set to control the approximation degree of the sigmoid function to the indicator function.

[0100] Therefore, the definition of the approximate Normalized Discounted Cumulative Gain (NDCG) loss is as follows:

[0101] Loss ApproxNDCG = ∑ i 1 - NDCG(x i , q).

[0102] In step S5, the dual-branch sorting loss sum.

[0103] As Figure 4 shown, both the global network branch and the local network branch will generate sorting losses: the output result of the local network branch after passing through the Sigmoid function is calculated with the local sorting supervision information for loss; in the global network branch, after feature concatenation with the local network branch and passing through the sigmoid function, its final result is calculated with the global sorting supervision information for loss.

[0104] The combination of the global ranking loss and the local ranking loss with certain weights will have a positive feedback effect on the dual-branch network model. The overall loss function of the dual-branch network model is defined as:

[0105] Loss = Loss global + λ * Loss local ;

[0106] In the formula, Loss global is the approximate normalized discounted cumulative gain (NDCG) loss value of the global ranking; Loss local is the approximate normalized discounted cumulative gain (NDCG) loss value of the local ranking; λ is the weight.

[0107] In step S5, least squares regression fitting is performed.

[0108] The output result of the dual-branch network model represents the ranking result of the target quantity of the image in this dataset. Then, simply through the standard least squares regression fitting method, the predicted ranking is mapped to the dual-branch network model count, and the least squares regression fitting can be performed using the image-level labels of the validation set. The formula is as follows:

[0109]

[0110] y′ = αη + β;

[0111] In the formula, η n represents the ranking order of the nth image predicted by the dual-branch network model, ranging in the interval [0, 1]; represents the average value of the predicted ranking; y n represents the true label value; α and β are the two parameters in the linear transformation respectively; y′ is the predicted value of the true count; N represents the number of images in the training set.

[0112] Specific embodiments are given below to illustrate in detail the method for counting remote sensing image targets. Experiments are carried out on three datasets for remote sensing image target counting tasks, as Figure 5 shown, to compare the technical solutions of the present invention and the implementation effects of the prior art.

[0113] RSOC building dataset: The building subset of the RSOC dataset, which includes a total of 2468 images, where the training set includes 1205 images and the test set includes 1263 images.

[0114] VisDrone20 1 9People dataset: A dataset for counting crowds based on drones, which includes 2392 training samples, 329 validation samples, and 626 test samples.

[0115] VisDrone2019 Vehicle Dataset: A drone-based vehicle counting dataset that includes categories such as cars, vans, trucks, and buses. This dataset consists of 3,953 training samples, 364 validation samples, and 986 test samples.

[0116] The hardware conditions for completing the present invention are: NVIDIA GeForce GTX 1080 GPU and Intel(R) Core(TM) i7-7700 CPU 3.60GHz. The environment configuration is: Windows 10 operating system, Pytorch 1.7.1 deep learning framework, and Python 3.7 programming language.

[0117] The present invention uses two evaluation metrics, the Mean Absolute Error (MAE) and the Root Mean Square Error (RMSE), to evaluate the performance of remote sensing image target counting. The definitions of the two metrics are as follows:

[0118]

[0119] In the formula, K is the number of test images; is the predicted count of the i-th image; C i is the true count of the i-th image.

[0120] During the network training process, the present invention uses the Adam optimizer, and the settings of other hyperparameters are as follows: mini-batch size 8; number of iterations 500; learning rate of the linear transformation layer is set to 0.01, and the learning rate of other layers is set to 0.00001; weight decay 0.00005.

[0121] In terms of the dataset, on the six datasets of RSOC-building, VisDrone2019 People, and VisDrone2019 Vehicle, all image sizes in this experiment are adjusted to 512*512 to avoid insufficient GPU video memory during the experiment. The data augmentation used in the experiment is carried out by color enhancement methods such as adding noise and blurring to the images, rather than geometric methods, because in the local sorting branch, the positional relationships between sub-images within the image are annotated.

[0122] For the RSOC_building dataset, 10% of the images are randomly selected from the training set and set as the validation set. Validation sets are provided for the other two datasets. In the experiment, only global sorting supervision information and local sorting supervision information are required for the training set, while the validation set requires image-level labels, which can be used to calculate the two parameters α and β of the least squares regression fitting and can also verify the model accuracy.

[0123] The hyperparameter T can control the calculation of the correlation score. First, the local network branch part is removed to prevent the influence of the two hyperparameters, the loss weight coefficient λ and the number of local sub-images n, on the experiment. That is to say, the model only includes the global network branch.

[0124] Six values of T, namely {0.5, 1, 2, 5, 10, 20}, are taken, and the experimental accuracy is shown in Table 1. The results show that T = 2 is the most suitable value for the current label magnitude.

[0125] Table 1 Experimental results of different T values on RSOC - Building

[0126]

[0127]

[0128] Determine the correlation coefficient T = 2, and temporarily set the number of local sub-images to 16. First, find the optimal value in the interval {1, 0.5, 0.1, 0.05}. It is found from the results that the counting accuracy is close when λ = 0.5 and λ = 0.1. Then, the range is further narrowed to {0.3, 0.25, 0.2, 0.15}. Finally, the results show that the optimal counting accuracy is achieved when λ = 0.2. The specific experimental results are shown in Table 2.

[0129] Table 2 Experimental results of different λ on RSOC - Building

[0130]

[0131] After adjusting the correlation coefficient T and the weight coefficient λ between the two losses, seriously consider the number of locally sorted sub-images. If the number of sub-images is small, the local sorting does not exert its optimal effectiveness and there is a certain potential space; but if the number of sub-images is too large, the size of each sub-image is very small, and the number of targets inside a large number of sub-images is very small or even none, and the sorting difference of the target numbers between sub-images is lost. Therefore, it is not conducive to the model to sort between sub-images.

[0132] Table 3 Experimental results of different n values on RSOC - Building

[0133]

[0134]

[0135] Therefore, choose to crop the image into five types of the same length and width sizes of 2 * 2, 3 * 3, 4 * 4, 5 * 5, 6 * 6, that is, 4, 9, 16, 25, 36 sub-images. From the experimental results corresponding to the five cropping methods, it shows that when the number of sub-images is 16, the best benefit brought by local sorting can be exerted. The experimental results are shown in Table 3.

[0136] As shown in Table 4, the proposed global combined with local sorting network is compared with other state-of-the-art methods on the RSOC_building dataset. These include density map-based methods and global regression methods. For different methods, the required label types are also listed in Table 4. The results show that although only the global sorting supervision information with the least annotation difficulty is used, the method of the present invention achieves the most accurate counting accuracy on the RSOC_building dataset.

[0137] Table 4 Accuracy comparison on RSOC_building dataset

[0138]

[0139] Table L * Represents pixel-level label (Location), N * Represents image-level label (Number), R * Represents global ranking supervision information (Ranking).

[0140] As shown in Table 5, the proposed global combined with local sorting network is compared with other methods on the VisDrone2019 People dataset and VisDrone2019 Vehicle. The results show that the counting accuracy of the method of the present invention is more accurate than the density map-based method that requires pixel-level labels. Specifically, on the Vehicle dataset, compared with other optimal counting accuracies, the method of the present invention improves MAE and RMSE by 11.9% and 4.2% respectively, and also achieves competitive prediction performance on the People dataset.

[0141] Table 5 Accuracy comparison on VisDrone dataset

[0142]

[0143] In summary, the present invention uses ranking learning and large model technology, the network model automatically generates annotations, and the network model can pay attention to the overall and local information of the image at the same time, and finally outputs the count value, which reduces the annotation cost and improves the counting accuracy. The present invention uses a combination-level label, that is, a ranking label of the number of targets in a group of images, which is called ranking supervision information. This further reduces the annotation cost and achieves competitive counting accuracy.

[0144] The above description is a detailed description of the preferred feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. All equivalent changes or modified changes completed under the technical spirit disclosed by the present invention should fall within the patent scope covered by the present invention.

Claims

1. A remote sensing image target counting method based on ranking learning and dual-branch network, characterized in that: The following steps are involved: For the dataset of remote sensing image target counts, global ranking supervision information is constructed; For the dataset of remote sensing image target count, local sorting supervision information is constructed; According to the global sorting supervision information and the local sorting supervision information, a dual-branch network model is constructed, and the dual-branch network model is trained using the training set in the data set until the dual-branch network model is trained; Input the images of the test set in the dataset into the dual-branch network model to obtain the predicted sorting order of the images in the dataset; Calculate the dual-branch sorting loss, calculate the loss function of the dual-branch network model by approximating the normalized discounted cumulative gain NDCG loss, and combine the dual-branch sorting loss and ; The image-level labels of the validation set in the dataset are used for least squares regression fitting to map the predicted ranking of the dual-branch network model to the target counts of remote sensing images.

2. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 1 is characterized in that: Construct global sorting supervision information, specifically: The global ranking supervision information provides label information for the global network branch, and the labeling process is obtained through manual annotation. Instead of using image-level labels, that is, the label information is the sum of the number of all objects in the image, global ranking supervision information is used, that is, only the order relationship of the number of objects between images is annotated. During the labeling process, the labeling strategy adopted is divided into two steps: first, the number of levels is determined according to the number of targets in the image, and the annotator assigns a level to each image by observing the number of targets; images with the same level are divided again, with one level corresponding to two labels, indicating that one level corresponds to two sorting labels of the image.

3. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 1 is characterized in that: Construct local sorting supervision information, specifically: The local ranking supervision information provides label information for the local network branch. The labeling process is obtained through automatic labeling by the target segmentation model SAM. Instead of using image-level labels, local ranking supervision information is used to segment a whole image into multiple square sub-images of equal size, and the order of the number of internal objects in the sub-images is marked. The local sorting supervision information is obtained through the target segmentation model SAM, which is divided into two steps: the image is passed into the target segmentation model SAM, the output result of the Encoder layer is visualized, and the neural network heat map is generated; according to the average value of the heat map, the heat map is separated into foreground and background, the foreground is the pixels greater than the average value, and the background is the pixels less than the average value, the foreground value is set to 1, and the background value is set to 0; After separating the foreground and background, it is pre-assumed that in an image, the sizes of different instances of the same type of targets are approximately equal, so the pixels occupied by each instance are also approximately the same. The number of pixels in each sub-image represents the number of targets in the area. By sorting the pixels of the sub-images, the required local sorting supervision information is obtained.

4. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 1 is characterized in that: Construct a dual-branch network model, specifically: The dual-branch network model uses a Transformer-based backbone network. After feature extraction, the image is divided into a local network branch and a global network branch. In the local network branch, the feature map is divided into multiple square sub-feature maps of equal size, and the output result is obtained through the fully connected layer; the number of neurons in the last layer of the fully connected layer is 1, and the sigmoid function is used to control the output result in the [0,1] interval; In the global network branch, the feature map is flattened and concatenated with the output of the local network branch. Then, it passes through the fully connected layer. The number of neurons in the last layer of the fully connected layer is 1 and is processed using the sigmoid function. The final processing result of the dual-branch network model represents the predicted sorting order of the image in the entire data set, ranging from [0,1]. The predicted result is then linearly transformed and mapped to the actual count through the least squares regression fitting method.

5. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 1 is characterized in that: The approximate normalized discounted cumulative gain NDCG loss calculation is as follows: Normalized discounted cumulative gain NDCG is used to measure the quality of ranking. The approximate normalized discounted cumulative gain NDCG loss function optimizes the normalized discounted cumulative gain NDCG retrieval index to achieve the purpose of building an ideal ranking. The "relevance" ψ rel and "similarity" ψ sim Two concepts to construct the sorting; In a batch, for the training sample X = {x1,…,x n }, where x i ∈x,i∈{1,…,n}, take the retrieval image as x i , the label value and predicted value of the retrieved image are y i and The query image corresponds to q = {x1,…,x i-1 ,x i+1 ,…,x n }, take the query image x q , the label value and predicted value of the query image are y q and "Correlation" ψ rel and "similarity" ψ sim The calculation formula is as follows: In the formula, ψ rel ∈[0,1],ψ sim ∈[0,1]; y max and min They represent the maximum and minimum label values ​​in the entire data set respectively; in a batch, only the images with the largest and smallest label values ​​are used as retrieval images to avoid calculation conflicts.

6. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 5 is characterized in that: The "correlation" ψ rel and "similarity" ψ sim Sorting is performed according to the relevance score. The "relevance" sorting list corresponds to the ideal sorting, and the "similarity" sorting corresponds to the sorting predicted by the dual-branch network model. The calculation formula for the approximate normalized discounted cumulative gain NDCG is as follows: score(x j )=T×ψ rel (j); NDCG(x j )=T×ψ rel (j); Where ψ(j) rank represents the ranking value of the query image j in the sorted list; score represents the relevance score between the query image and the retrieved image, according to the "relevance" ψ rel It is calculated by multiplying it with the hyperparameter T.

7. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 6 is characterized in that: In order to calculate the discounted cumulative gain DCG, it is also necessary to calculate the ranking values ​​of each retrieved image in the two lists. Since the normalized discounted cumulative gain NDCG is not differentiable, it cannot be directly optimized using gradient descent. Therefore, the sigmiod function is used to approximate the ranking value predicted by the dual-branch network model. This transformation enables the function to perform gradient calculation. "Relevance" ranking value ψ rel (j) rank and the "similarity" ranking value ψ sim (j) rank The calculation formula is as follows: Where j and k represent the jth and kth images in the retrieval image, respectively; ψ rel (j) and ψ sim (j) represents the value of the jth retrieved image in the two sorted lists; is the indicator function, σ(ba) is the approximation of the indicator function, and setting θ=10 controls the approximation of the sigmoid function to the indicator function; The approximate normalized discounted cumulative gain NDCG loss is defined as follows: Loss ApproxNDCG =∑ i 1-NDCG(x i ,q)。 8. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 7 is characterized in that: Combining the dual-branch sorting loss and, we get the loss function of the dual-branch network model, specifically: Both the global network branch and the local network branch will generate sorting loss: the output result of the local network branch after being processed by the Sigmoid function is compared with the local sorting supervision information to calculate the loss; in the global network branch, the feature concatenation with the local network branch is processed by the sigmoid function, and the final result is compared with the global sorting supervision information to calculate the loss; The combination of global sorting loss and local sorting loss produces a positive feedback effect on the two-branch network model, and the overall loss function of the two-branch network model is defined as: Loss=Loss global +λ*Loss local ; In the formula, Loss global It is the approximate normalized discounted cumulative gain NDCG loss value of the global sorting; Loss local It is the approximate normalized discounted cumulative gain NDCG loss value of local sorting; λ is the weight.

9. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 1 is characterized in that: Least squares regression fitting, specifically: The output of the two-branch network model represents the ranking result of the number of targets in this image in this dataset. We only need to use the standard least squares regression fitting method to map the predicted ranking to the two-branch network model count, and use the image-level labels of the validation set to perform least squares regression fitting. The formula is as follows: y ′ =αη+β; Where η n Represents the ranking order of the nth image predicted by the dual-branch network model, ranging from [0,1]; represents the average of the predicted ranking; y n represents the true label value; α and β are two parameters in the linear transformation; y ′ is the true count prediction value; N is the number of images in the training set.

10. The remote sensing image target counting method based on ranking learning and dual-branch network according to claim 1, characterized in that: The datasets for remote sensing image target counting include: RSOC_building dataset, VisDrone2019 People dataset, and VisDrone2019 Vehicle dataset; The mean error (MAE) and the root mean square error (RMSE) are used to evaluate the performance of remote sensing image target counting. The definitions of the two indicators are as follows: Where K is the number of test images; is the predicted count of the i-th image; C i is the true count of the i-th image.

Citation Information

Cited By

  • Industrial weak semantic target detection method and device, equipment and medium

    CN122049345A