Remote sensing image rigid target fine granularity identification method and device
By smoothly labeling the remote sensing image dataset and introducing channel feature and spatial feature learning modules, the problems of insufficient utilization of scale information and limited feature extraction capabilities in the prior art are solved, and the accuracy of rigid target fine-grained recognition is significantly improved.
Patent Information
- Application Number
- CN202510620161.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, due to insufficient utilization of scale information by the model and limited feature extraction capabilities, the accuracy of rigid target fine-grainedness recognition is insufficient.
By smoothly annotating the remote sensing image dataset, soft labels based on the target length, width and aspect ratio are generated, and smooth labels are obtained in combination with hard labels. At the same time, a channel feature learning module and a spatial feature learning module are introduced to construct channel feature loss and spatial feature loss to improve the feature extraction capability of the feature encoder.
The accuracy of identifying rigid targets with obvious scale differences is improved, the model's ability to extract target discriminant features is enhanced, and the effect of fine-grained recognition of rigid targets in remote sensing images is significantly improved.
Smart Images

Figure CN120147894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular to a method and device for fine-grained recognition of rigid targets in remote sensing images. Background Art
[0002] With the rapid development of remote sensing imaging technology, the demand for fine-grained recognition of ground rigid targets (such as airplanes, ships, vehicles, etc.) is increasing. Existing deep learning-based methods for recognizing ground rigid targets usually directly divide the remote sensing image data in the dataset into a training set and a test set, scale the sizes of the remote sensing images to a unified scale and input them into the network to extract features, and use hard labels in the form of one-hot encoding to guide the model to complete training.
[0003] However, since the existing model training method directly uses a fixed-size input, ignoring the target scale distribution law, and the improved method based on the attention mechanism only focuses on multi-scale perception of the feature layer and does not combine the target scale statistical characteristics, the accuracy of the model for recognizing rigid targets is still low when recognizing different types of targets with similar apparent structures and large scale differences; moreover, rigid fine-grained targets have the characteristics of large intra-class differences and small inter-class variances, and the insufficient feature extraction ability of the model will also affect the effect of fine-grained recognition of rigid targets. Summary of the Invention
[0004] To this end, the technical problem to be solved by the present invention is to overcome the problem of insufficient accuracy of fine-grained recognition of rigid targets due to insufficient utilization of scale information of the model and limited extraction of discriminative features in the prior art.
[0005] To solve the above technical problem, the present invention provides a method for fine-grained recognition of rigid targets in remote sensing images, including: Performing smooth annotation on the remote sensing image dataset, including: Generating soft labels for each target in the remote sensing image according to the length, width and aspect ratio of each target; Obtaining smooth annotations for each target according to the hard labels and soft labels of each target in the remote sensing image; Passing the remote sensing images in the remote sensing image dataset through a feature encoder to obtain feature maps, and passing the feature maps through a regression layer to obtain the predicted probability distribution of each target in the remote sensing image; Constructing a classification loss according to the predicted probability distribution of each target in the remote sensing image and its corresponding smooth annotation; Inputting the feature maps into a channel feature learning module, and obtaining the predicted probability distribution of the feature channels of each target in the remote sensing image after channel masking, spatial average pooling and mapping; constructing a channel feature loss according to the predicted probability distribution of the feature channels of each target in the remote sensing image and its corresponding smooth annotation; The feature map is input into the spatial feature learning module, and after spatial information normalization and channel maximum pooling, the spatial feature loss is constructed; The total loss function is constructed with classification loss, channel feature loss and spatial feature loss to train feature encoder, regression layer, channel feature learning module and spatial feature learning module; A remote sensing target recognition model is constructed with the trained feature encoder and regression layer to perform target recognition on the remote sensing images to be detected.
[0006] Preferably, before smoothing and labeling the remote sensing image dataset, the method further includes: adopting a long-tail alignment training strategy to perform category balancing processing on the remote sensing image dataset.
[0007] Preferably, a soft label of each target is generated according to the length, width and aspect ratio of each target in the remote sensing image, including: Obtain the category of each target based on the hard label of each target in the remote sensing image; Calculate the mean and variance of the length, width, and aspect ratio of all objects in each category, and construct the Gaussian distribution of the length, width, and aspect ratio of each category; Calculate the probability that each target belongs to each category and obtain the soft label of each target, including: According to the Gaussian distribution of length, width and aspect ratio of each category, the probability that the length, width and aspect ratio of the current target belong to each category is calculated. , and ;in, Indicates the probability that the length of the current target belongs to the nth category, Indicates the probability that the width of the current target belongs to the nth category, Indicates the probability that the aspect ratio of the current target belongs to the nth category; , N represents the total number of categories; right , and After normalization, take the average to get the probability that the current target belongs to the nth category ; The soft label of the current target .
[0008] Preferably, before respectively calculating the average and variance of the length, width and aspect ratio of all objects in each category, the method further includes: The upper and lower limits of the length, width and aspect ratio of the targets in each category are calculated using the lower and upper quartiles of the length, width and aspect ratio of all targets in each category. The formulas include: ; ; ; Among them, and respectively represent the lower and upper limits of the length of the targets in the nth category, and respectively represent the lower and upper quartiles of the lengths of all targets in the nth category; and respectively represent the lower and upper limits of the width of the targets in the nth category, and respectively represent the lower and upper quartiles of the widths of all targets in the nth category; and respectively represent the lower and upper limits of the aspect ratio of the targets in the nth category, and respectively represent the lower and upper quartiles of the aspect ratios of all targets in the nth category; Remove the outliers of the lengths, widths, and aspect ratios of the targets in each category based on the upper and lower limits of the lengths, widths, and aspect ratios of the targets in each category.
[0009] Preferably, obtain the smooth annotation of each target according to the hard label and soft label of each target in the remote sensing image, and the formula is expressed as: ; Among them, represents the smooth annotation of the target, represents the hard label of the target, represents the soft label of the target, represents the proportionality coefficient.
[0010] Preferably, the feature encoder uses the ResNet18 network or the ViT network as the backbone network.
[0011] Preferably, input the feature map into the channel learning module, and obtain the feature channel prediction probability distribution of each target in the remote sensing image after channel masking, spatial average pooling, and mapping. The formula is expressed as: ; Among them, represents the feature channel prediction probability distribution of the target, represents the masking ratio, represents the feature map; represents the channel masking operation, represents the spatial average pooling operation, represents the mapping operation.
[0012] Preferably, after the feature map undergoes spatial information normalization and channel max pooling, a spatial feature loss is constructed, and the formula is expressed as: ; wherein, represents the spatial feature loss, represents the feature map, represents the operation of performing spatial information normalization on the feature map, represents the channel max pooling operation, represents a preset upper bound.
[0013] Preferably, the formula of the total loss function is expressed as: ; wherein, represents the total loss function; represents the classification loss, represents the predicted probability distribution of the target, represents the smoothed annotation of the target, represents the cross-entropy loss function; represents the channel feature loss, represents the predicted probability distribution of the feature channels of the target; represents the spatial feature loss, represents the feature map; represents the loss ratio coefficient.
[0014] The present invention also provides a remote sensing image rigid target fine-grained recognition device, including: An annotation module, used for performing smoothed annotation on the remote sensing image data set, including: Generating soft labels for each target in the remote sensing image according to the length, width, and aspect ratio of each target; Obtaining the smoothed annotation of each target according to the hard label and soft label of each target in the remote sensing image; A classification loss construction module, used for obtaining a feature map by passing the remote sensing image in the remote sensing image data set through a feature encoder, and obtaining the predicted probability distribution of each target in the remote sensing image by passing the feature map through a regression layer; constructing a classification loss according to the predicted probability distribution of each target in the remote sensing image and its corresponding smoothed annotation; A channel feature loss construction module, used for inputting the feature map into a channel learning module, and obtaining the predicted probability distribution of the feature channels of each target in the remote sensing image after passing through a channel mask, spatial average pooling, and mapping; constructing a channel feature loss according to the predicted probability distribution of the feature channels of each target in the remote sensing image and its corresponding smoothed annotation; A spatial feature loss construction module, used for constructing a spatial feature loss after the feature map undergoes spatial information normalization and channel max pooling; A training module, which is used to construct a total loss function with a classification loss, a channel feature loss, and a spatial feature loss, and train a feature encoder, a regression layer, and a channel learning module; An object recognition module, which is used to construct a remote sensing object recognition model with the trained feature encoder and regression layer, and perform object recognition on the remote sensing image to be detected.
[0015] The above technical solution of the present invention has the following beneficial effects compared with the prior art: For the method for fine-grained recognition of rigid objects in remote sensing images provided by the present invention, soft labels are generated based on the length, width, and aspect ratio of the object, and combined with hard labels to obtain smooth annotations of the object, so that the annotation information contains the scale information of the object. When using the smooth annotations to train the remote sensing object recognition model, it can guide the model to learn the scale differences between objects and improve the accuracy of recognizing rigid objects with obvious scale differences; moreover, when training the model, the present invention additionally constructs a channel feature loss and a spatial feature loss from two perspectives of channel features and spatial features to improve the feature extraction ability of the feature encoder and strengthen the model's ability to extract discriminative features of objects. The present invention is integrated on a general model network architecture, and introduces a small amount of parameter calculation during the training process without increasing the additional computational burden during inference, effectively improving the accuracy of fine-grained recognition of rigid objects in remote sensing images, and the effect is particularly prominent in categories with significant scale differences such as airplanes and ships.
[0016] Furthermore, the present invention adopts a long-tail alignment training strategy to construct a remote sensing image dataset, which can effectively alleviate the problem of unbalanced data categories in remote sensing images and improve the robustness of the method of the present invention. Description of the Drawings
[0017] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in combination with the drawings, where: Figure 1 is a flowchart of the method for fine-grained recognition of rigid objects in remote sensing images provided by the present invention; Figure 2 is a schematic diagram of the result of fine-grained recognition of a remote sensing image by the method of the present invention, where Figure 2 in (a) is the fine-grained recognition result and the ground truth label of an airplane, Figure 2 in (b) is the fine-grained recognition result and the ground truth label of a site, Figure 2 in (c) is the fine-grained recognition result and the ground truth label of a ship; Figure 3 is a structural diagram of the device for fine-grained recognition of rigid objects in remote sensing images provided by the present invention. Detailed Embodiments
[0018] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited do not limit the present invention.
[0019] Referring to Figure 1 as shown, the present invention proposes a method for fine-grained recognition of rigid targets in remote sensing images, including: S1: Obtain a remote sensing image dataset, and perform class balance processing on the remote sensing image dataset using a long-tail alignment training strategy.
[0020] Specifically, use the long-tail alignment training strategy to screen the remote sensing images in the remote sensing image dataset, and complete the supplementation of samples in the categories with fewer numbers through data augmentation and other means to ensure that the maximum number of targets in each category is N 2 = 3000, and the minimum is N 1 = 900. And in the class-balanced remote sensing image dataset, both the divided training set and test set use 5-fold cross-validation, that is, the remote sensing image dataset is randomly and evenly divided into 5 non-overlapping subsets, and the size of each subset is roughly equal. In each validation process, 4 of the subsets are selected as the training set for training the model, and the remaining 1 subset is used as the test set to evaluate the performance of the model on this subset. A total of 5 trainings and tests are performed, and different subsets are used as the test set each time to improve the robustness of the single training accuracy.
[0021] The present invention uses a long-tail alignment training strategy to construct a remote sensing image dataset, which can effectively alleviate the problem of uneven class distribution of remote sensing image data and improve the robustness of the method of the present invention.
[0022] S2: Based on the overall distribution of the image data, perform smooth annotation on the remote sensing image dataset through a smooth annotation generation method based on a Gaussian-rectangular distribution mixture function. Specifically, it includes the following steps: S21: Obtain the hard label of each target in the remote sensing image.
[0023] The hard label means that the label of each target is a clear category, usually an integer or a category name. In a classification task, the hard label indicates that the target belongs to a specific category, usually represented by one-hot encoding. Taking a three-classification task as an example, the categories are aircraft, ship, and vehicle. The one-hot encoding of aircraft (hard label 0) is [1, 0, 0]; the one-hot encoding of ship (hard label 1) is [0, 1, 0], and the one-hot encoding of vehicle (hard label 2) is [0, 0, 1].
[0024] S22: According to the length of each target in the remote sensing image , width and aspect ratio Generate soft labels for each target, including: S221: Obtain the category of each target according to the hard label of each target in the remote sensing image.
[0025] S222: Calculate the upper and lower limits of the length, width, and aspect ratio of the objects of each category, and remove abnormal values of the length, width, and aspect ratio of the objects in each category.
[0026] Specifically, the upper and lower limits of the length, width and aspect ratio of the objects in each category are calculated based on the lower and upper quartiles of the length, width and aspect ratio of all objects in each category. The formulas include: ; ; ; in, and They represent the lower and upper limits of the length of the target of the nth category, respectively. and denote the lower and upper quartiles of the lengths of all objects in the nth category, respectively; and Respectively represent the lower and upper limits of the width of the target of the nth category, and denote the lower and upper quartiles of the width of all objects in the nth category, respectively; and They represent the lower and upper limits of the aspect ratio of the target of the nth category, respectively. and They represent the lower and upper quartiles of the aspect ratio of all objects in the nth category, respectively.
[0027] The upper and lower limits of the length, width and aspect ratio of the targets in each category are used to eliminate outliers in the length, width and aspect ratio of the targets in each category.
[0028] S223: Calculate the mean and variance of the length, width and aspect ratio of all objects in each category after removing outliers, and construct the Gaussian distribution of the length, width and aspect ratio of each category.
[0029] S224: Calculate the probability that each target belongs to each category, and obtain a soft label for each target, including: According to the Gaussian distribution of length, width and aspect ratio of each category, the probability that the length, width and aspect ratio of the current target belong to each category is calculated. , and , the formula includes: ; ; ; Among them, represents the probability that the length of the current target belongs to the nth category, represents the probability that the width of the current target belongs to the nth category, represents the probability that the aspect ratio of the current target belongs to the nth category; 、 and respectively represent the average values of the lengths, widths, and aspect ratios of all targets in the nth category; 、 and respectively represent the variances of the lengths, widths, and aspect ratios of all targets in the nth category; , and N represents the total number of categories.
[0030] For 、 and , after normalizing and averaging using the softmax operation, the probability that the current target belongs to the nth category is obtained, and the formula is expressed as: ; The soft label of the current target .
[0031] S23: Obtain the smoothed annotation of each target according to the hard label and soft label of each target in the remote sensing image, and the formula is expressed as: ; Among them, represents the smoothed annotation of the target, represents the hard label of the target, represents the soft label of the target, represents the proportionality coefficient.
[0032] Set the adjustment proportionality coefficient , in this embodiment, takes values in [0, 1]. After analyzing the experimental results, taking 0.4 has the best effect.
[0033] The present invention calculates the soft label of the target based on the scale information of the target, further generates the smoothed annotation of the target in the remote sensing image, and uses the smoothed annotation to guide the model to learn the scale difference between targets during model training, thereby improving the recognition accuracy of the model.
[0034] S3: Obtain a feature map from the remote sensing images in the remote sensing image dataset through a feature encoder, and obtain the predicted probability distribution of each target in the remote sensing image from the feature map through a regression layer. 。
[0035] Preferably, the feature encoder uses a ResNet18 network or a ViT network as the backbone network.
[0036] S4: Construct a classification loss according to the predicted probability distribution of each target in the remote sensing image and its corresponding smoothed annotation.
[0037] S5: Input the feature map into a channel feature learning module to construct a channel feature loss.
[0038] The channel feature learning module includes channel masking, pooling and mapping, and channel feature loss calculation. The channel masking process aims to randomly mask the channel dimension information of the relevant feature map to obtain a masked feature; feed the masked feature into the pooling and mapping unit to obtain a result with the same dimension as the number of categories; finally, use the cross-entropy loss function to calculate the channel feature loss.
[0039] Specifically, after the feature map undergoes channel masking, spatial average pooling and mapping, the predicted probability distribution of the feature channels of each target in the remote sensing image is obtained, and the formula is expressed as: ; where represents the predicted probability distribution of the feature channels of the target, represents the masking ratio, represents the feature map; represents the channel masking operation, represents the spatial average pooling operation, represents the mapping operation.
[0040] Technically, the selection of the masking ratio needs to be selected according to the experimental results. The masking ratio in this embodiment has the best effect.
[0041] Construct a channel feature loss with the cross-entropy loss function according to the predicted probability distribution of the feature channels of each target in the remote sensing image and its corresponding smoothed annotation.
[0042] S6: Input the feature map into a spatial feature learning module to construct a spatial feature loss.
[0043] The spatial feature learning module includes spatial information normalization, channel max pooling, and spatial feature loss calculation.
[0044] Specifically, after the feature map undergoes spatial information normalization and channel max pooling, a spatial feature loss is constructed, and the formula is expressed as: ; Among them, represents the spatial feature loss, represents the feature map, represents the operation of normalizing the spatial information of the feature map, represents the channel max pooling operation, represents the preset upper bound.
[0045] In this embodiment, M = 3 to ensure that the spatial feature loss is greater than 0.
[0046] The present invention performs spatial information normalization and channel max pooling operations on the feature map, retains the discriminative channel features, further maximizes the response of the target local region using the spatial feature loss, and ensures the retention of valuable discriminative features.
[0047] S7: Construct a total loss function with the classification loss, channel feature loss, and spatial feature loss, which is expressed by the formula: ; Among them, represents the total loss function; represents the classification loss, represents the predicted probability distribution of the target, represents the smooth annotation of the target, represents the cross-entropy loss function; represents the channel feature loss, represents the predicted probability distribution of the feature channels of the target; represents the spatial feature loss; represents the loss ratio coefficient, which takes in this embodiment and has the best effect.
[0048] When constructing the total loss function, the present invention strengthens the ability of the model to extract discriminative features of the target from two perspectives of feature channels and spatial features respectively.
[0049] S8: Train the feature encoder, regression layer, channel feature learning module, and spatial feature learning module with the total loss function.
[0050] S9: Construct a remote sensing target recognition model with the trained feature encoder and regression layer; use the remote sensing target recognition model to perform target recognition on the remote sensing image to be detected.
[0051] A fine-grained recognition method for rigid targets in remote sensing images according to the present invention generates soft labels based on the length, width, and aspect ratio of the targets, and combines with hard labels to obtain smooth annotations of the targets, so that the annotation information contains the scale information of the targets. When training the remote sensing target recognition model using the smooth annotations, it can guide the model to learn the scale differences between targets and improve the accuracy of recognizing rigid targets with obvious scale differences. Moreover, when training the model, the present invention additionally constructs a channel feature loss and a spatial feature loss from two perspectives of channel features and spatial features to enhance the feature extraction ability of the feature encoder and strengthen the model's ability to extract discriminative features of the targets. The present invention is integrated on a general model network architecture, and introduces a small amount of parameter calculation during the training process without adding additional computational burden during inference, effectively improving the accuracy of fine-grained recognition of rigid targets in remote sensing images, and the effect is particularly prominent in categories with significant scale differences such as airplanes and ships.
[0052] Furthermore, the present invention adopts a long-tail alignment training strategy to construct a remote sensing image dataset, which can effectively alleviate the problem of unbalanced data categories in remote sensing images and improve the robustness of the method of the present invention.
[0053] To fully verify the effectiveness of the designed method in the fine-grained rigid object classification task, this embodiment conducts a comparative experiment between the method of the present invention and a baseline model, and conducts a detailed ablation study. The baseline model uses a standard classification network with a similar structure to the designed method, but does not apply the smooth annotation generation process based on the Gaussian-rectangular distribution mixture function, and does not introduce a channel feature learning module and a spatial feature learning module during the training process.
[0054] All experiments in this embodiment are carried out in an environment equipped with a single NVIDIA RTX4090 24-GB GPU, and the model is trained using the Adam optimizer. The model training adopts a training strategy of 30 rounds, the learning rate is initialized to 0.001, and decays to 1 / 10 of the original after the 15th round, and the batch size is set to 128. The experimental dataset selects a large-scale remote sensing fine-grained target recognition dataset, which covers remote sensing images with a resolution of 0.3-0.8m, including 4 major categories of airplanes, ships, vehicles, and sites, with a total of 37 fine-grained categories. For example, airplanes include categories such as Boeing 777, A220, A350, and Boeing 737, and sites include fine-grained categories such as tennis courts and basketball courts. After data augmentation (rotation, adding noise, etc.), the total number of instances in the dataset reaches 592345.
[0055] The parameter settings include: M = 3, β = 0.2, α = 0.4, λ = 0.3.
[0056] First, ResNet-18 or efficient-ViT (abbreviated as e-ViT) is used as the backbone network of the feature encoder, and the method of the present invention is applied to conduct fine-grained rigid object recognition experiments on the dataset. During the experiment, the input image of the model , and a single-level feature map is generated by the encoder . During training, the feature map is input into the subsequent channel feature learning module and spatial feature learning module to extract the discriminative features of the feature map respectively, and the prediction result is obtained through operations including flattening, mapping, and classifier, etc. Figure 2 is a schematic diagram of the result of fine-grained recognition of remote sensing images by the method of the present invention, where Figure 2 in (a) is the fine-grained recognition result and true label of the aircraft Figure 2 in (b) is the fine-grained recognition result and true label of the site Figure 2 in (c) is the fine-grained recognition result and true label of the ship
[0057] In this embodiment, the F1 score is also used to analyze the performance of the method of the present invention and the baseline model in terms of recognition accuracy. The formula for the F1 score is: ; ; ; where represents precision, represents recall, TP represents true positive, FP represents false positive, and FN represents false negative
[0058] At the same time, this embodiment also calculates the mF1-score, that is, the average value of the F1 scores of each category. Through the comparison of these indicators, the recognition accuracy of the model on different categories and overall samples can be comprehensively evaluated, and the higher the F1 index, the better the effect. Taking ResNet-18 as the backbone network, the comparison results of the F1 scores of the method of the present invention and the baseline model on different categories are shown in Table 1
[0059] Table 1. Comparison results of F1 scores of the method of the present invention and the baseline model on different categories
[0060] As can be seen from Table 1, the method of the present invention is superior to the baseline model in all indicators, and the average F1 score (mF1-score) has increased by 3.22. Based on the ResNet as the backbone network model, the method of the present invention can achieve stable performance improvement, and the more complex the model, the more significant the performance improvement. In different categories, the performance improvement ranges from large to small as aircraft, vehicles, ships, and sites, which is closely related to the density of the target scale distribution and the scale difference between various types of targets.
[0061] Referring to Figure 3 shown, based on the above-mentioned method for fine-grained recognition of rigid targets in remote sensing images, this embodiment also provides a device for fine-grained recognition of rigid targets in remote sensing images, including: An annotation module for performing smooth annotation on a remote sensing image data set, including: Generating soft labels for each target in the remote sensing image according to the length, width, and aspect ratio of each target in the remote sensing image; Obtaining smooth annotations for each target according to the hard labels and soft labels of each target in the remote sensing image; A classification loss construction module for obtaining a feature map by passing a remote sensing image in the remote sensing image data set through a feature encoder, and obtaining a predicted probability distribution of each target in the remote sensing image by passing the feature map through a regression layer; constructing a classification loss according to the predicted probability distribution of each target in the remote sensing image and its corresponding smooth annotation; A channel feature loss construction module for inputting the feature map into a channel learning module, and obtaining a predicted probability distribution of the feature channels of each target in the remote sensing image after channel masking, spatial average pooling, and mapping; constructing a channel feature loss according to the predicted probability distribution of the feature channels of each target in the remote sensing image and its corresponding smooth annotation; A spatial feature loss construction module for constructing a spatial feature loss after normalizing the spatial information and performing channel max pooling on the feature map; A training module for constructing a total loss function with the classification loss, channel feature loss, and spatial feature loss, and training the feature encoder, regression layer, and channel learning module; A target recognition module for constructing a remote sensing target recognition model with the trained feature encoder and regression layer, and performing target recognition on the remote sensing image to be detected.
[0062] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0063] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows Figure 1 one flow or multiple flows and / or blocks Figure 1 or one or more blocks.
[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows Figure 1 one flow or multiple flows and / or blocks Figure 1 or one or more blocks.
[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one flow or multiple flows and / or blocks Figure 1 or one or more blocks.
[0066] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for fine-grained recognition of rigid targets in remote sensing images, characterized in that: include: Smooth annotation of remote sensing image datasets, including: Generate a soft label for each target according to the length, width and aspect ratio of each target in the remote sensing image; Obtain smooth annotation of each target according to the hard label and soft label of each target in the remote sensing image; The remote sensing images in the remote sensing image dataset are passed through a feature encoder to obtain a feature map, and the feature map is passed through a regression layer to obtain the predicted probability distribution of each target in the remote sensing image; Construct the classification loss based on the predicted probability distribution of each target in the remote sensing image and its corresponding smooth annotation; The feature map is input into the channel feature learning module, and after channel masking, spatial average pooling and mapping, the feature channel prediction probability distribution of each target in the remote sensing image is obtained; the channel feature loss is constructed according to the feature channel prediction probability distribution of each target in the remote sensing image and its corresponding smooth annotation; The feature map is input into the spatial feature learning module, and after spatial information normalization and channel maximum pooling, the spatial feature loss is constructed; The total loss function is constructed with classification loss, channel feature loss and spatial feature loss to train feature encoder, regression layer, channel feature learning module and spatial feature learning module; A remote sensing target recognition model is constructed with the trained feature encoder and regression layer to perform target recognition on the remote sensing images to be detected.
2. According to claim 1, a remote sensing image rigid target fine-grained recognition method is characterized in that: Before smoothly annotating the remote sensing image dataset, the method also includes: adopting a long-tail alignment training strategy to perform category balancing processing on the remote sensing image dataset.
3. The method for fine-grained identification of rigid targets in remote sensing images according to claim 1, characterized in that: Generate a soft label for each target based on the length, width and aspect ratio of each target in the remote sensing image, including: Obtain the category of each target based on the hard label of each target in the remote sensing image; Calculate the mean and variance of the length, width, and aspect ratio of all objects in each category, and construct the Gaussian distribution of the length, width, and aspect ratio of each category; Calculate the probability that each target belongs to each category and obtain the soft label of each target, including: According to the Gaussian distribution of length, width and aspect ratio of each category, the probability that the length, width and aspect ratio of the current target belong to each category is calculated. , and ;in, Indicates the probability that the length of the current target belongs to the nth category, Indicates the probability that the width of the current target belongs to the nth category, Indicates the probability that the aspect ratio of the current target belongs to the nth category; , N represents the total number of categories; right , and After normalization, take the average to get the probability that the current target belongs to the nth category ; The soft label of the current target .
4. The method for fine-grained identification of rigid targets in remote sensing images according to claim 3 is characterized in that: Before calculating the mean and variance of the length, width, and aspect ratio of all objects in each category, it also includes: The upper and lower limits of the length, width and aspect ratio of the targets in each category are calculated using the lower and upper quartiles of the length, width and aspect ratio of all targets in each category. The formulas include: ; ; ; in, and They represent the lower and upper limits of the length of the target of the nth category, respectively. and denote the lower and upper quartiles of the lengths of all objects in the nth category, respectively; and Respectively represent the lower and upper limits of the width of the target of the nth category, and denote the lower and upper quartiles of the width of all objects in the nth category, respectively; and They represent the lower and upper limits of the aspect ratio of the target of the nth category, respectively. and Respectively represent the lower quartile and upper quartile of the aspect ratio of all objects in the nth category; The upper and lower limits of the length, width and aspect ratio of the targets in each category are used to eliminate outliers in the length, width and aspect ratio of the targets in each category.
5. The method for fine-grained identification of rigid targets in remote sensing images according to claim 1, characterized in that: According to the hard label and soft label of each target in the remote sensing image, the smooth annotation of each target is obtained, and the formula is expressed as: ; in, represents the smooth annotation of the target, represents the hard label of the target, represents the soft label of the target, Represents the proportionality factor.
6. The method for fine-grained identification of rigid targets in remote sensing images according to claim 1, characterized in that: The feature encoder uses a ResNet18 network or a ViT network as a backbone network.
7. The method for fine-grained identification of rigid targets in remote sensing images according to claim 1, characterized in that: The feature map is input into the channel learning module, and after channel masking, spatial average pooling and mapping, the feature channel prediction probability distribution of each target in the remote sensing image is obtained. The formula is expressed as: ; in, Represents the predicted probability distribution of the target’s feature channels, represents the mask ratio, Represents a feature map; Indicates a channel mask operation, represents the spatial average pooling operation, Represents a map operation.
8. The method for fine-grained identification of rigid targets in remote sensing images according to claim 1, characterized in that: After the feature map is normalized and the channel is max-pooled, the spatial feature loss is constructed, and the formula is expressed as: ; in, represents the spatial feature loss, represents the feature map, Indicates the normalization operation of spatial information on the feature map. represents the channel maximum pooling operation, Indicates the preset upper bound.
9. The method for fine-grained identification of rigid targets in remote sensing images according to claim 1, characterized in that: The formula of the total loss function is expressed as: ; in, represents the total loss function; represents the classification loss, represents the predicted probability distribution of the target, represents the smooth annotation of the target, represents the cross entropy loss function; represents the channel feature loss, Represents the predicted probability distribution of the target’s feature channels; represents the spatial feature loss, Represents a feature map; Represents the loss proportionality coefficient.
10. A remote sensing image rigid target fine-grained recognition device, characterized in that: include: The annotation module is used to smoothly annotate remote sensing image datasets, including: Generate a soft label for each target according to the length, width and aspect ratio of each target in the remote sensing image; Obtain smooth annotation of each target according to the hard label and soft label of each target in the remote sensing image; The classification loss construction module is used to obtain a feature map by passing the remote sensing image in the remote sensing image dataset through a feature encoder, and obtain the predicted probability distribution of each target in the remote sensing image by passing the feature map through a regression layer; the classification loss is constructed according to the predicted probability distribution of each target in the remote sensing image and its corresponding smooth annotation; The channel feature loss construction module is used to input the feature map into the channel learning module, and obtain the feature channel prediction probability distribution of each target in the remote sensing image after channel masking, spatial average pooling and mapping; the channel feature loss is constructed according to the feature channel prediction probability distribution of each target in the remote sensing image and its corresponding smooth annotation; The spatial feature loss construction module is used to construct the spatial feature loss after the feature map is normalized by spatial information and channel maximum pooling; The training module is used to construct a total loss function based on classification loss, channel feature loss, and spatial feature loss, and train feature encoders, regression layers, and channel learning modules; The target recognition module is used to build a remote sensing target recognition model with the trained feature encoder and regression layer to perform target recognition on the remote sensing image to be detected.
Citation Information
Patent Citations
Remote sensing image target fine identification method, electronic equipment and storage medium
CN114998748A
Long-tail target detection method based on monitoring scene
CN118097268A
Unsupervised domain adaptation target re-identification method
WO2022001489A1
Cited By
Fine-grained target detection and identification method and device and readable storage medium
CN120510456A