Remote Sensing Image Target Detection Method Based on Parameter Optimization

By adopting a parameter-based optimization method in remote sensing image object detection, using a preset remote sensing object detection model and context conversion module, the problems of poor generalization ability and lack of long-distance modeling in traditional methods are solved, and more efficient remote sensing image object detection is achieved.

CN115359366BActive Publication Date: 2025-06-10THE QUARTERMASTER RES INST OF THE GENERAL LOGISTICS DEPT OF THE CPLA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211001673.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-06-10
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

Traditional remote sensing image object detection methods have poor generalization capabilities for remote sensing images and poor detection effects. The existing methods lack long-distance modeling and perception capabilities, resulting in a low detection rate.

Method used

The remote sensing image object detection method based on parameter optimization is adopted, and the anchor box is determined by acquiring optical remote sensing image information, and the positioning and identification process is performed using a preset remote sensing object detection model. The model consists of the input end, the backbone feature extraction network, the feature fusion network and the output end. Through the context conversion module and the feature pyramid structure, the fusion and enhancement of features are achieved, and the model parameters are optimized through the newly proposed remote sensing object detection loss function.

Benefits of technology

The generalization ability and detection accuracy of remote sensing image object detection are improved, the long-distance modeling and perception ability of the model are enhanced, and the detection rate and efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359366B_ABST
    Figure CN115359366B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote sensing image target detection method based on parameter optimization, including: obtaining optical remote sensing image information; determining anchor boxes for remote sensing image target detection, and using a remote sensing target detection model to perform positioning and recognition processing on the optical remote sensing image information to obtain an output feature map information set; the output feature map information set includes a plurality of output feature map information; performing post-processing on the output feature map information set to obtain a target image detection information set. In the backbone feature extraction network and the feature fusion network of the present invention, a computing unit more suitable for computer vision is constructed through the operation of grouped convolution, and anchor boxes are selected through parameter optimization, solving the problem that the existing anchor box detection method causes deviation in the clustering result, resulting in a large deviation in the selection of anchor boxes and affecting the success rate of image detection, and optimizing the remote sensing target detection ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing, and in particular relates to a method for remote sensing image target detection based on parameter optimization. Background Art

[0002] With the development of remote sensing technology, remote sensing image detection has been widely used in military and civilian fields. Using satellite-taken remote sensing images for target detection can bring great convenience to fields such as maritime ship personnel search and rescue, military intelligence reconnaissance, and traffic flow monitoring. However, different from optical image detection in natural scenes, remote sensing image target detection faces the characteristics of drastic changes in the scales of detected targets, a large proportion of small targets, and complex image scenes, resulting in frequent problems of missed and false detections, which in turn seriously affect its detection accuracy and efficiency, and limit its application in satellite remote sensing technology to a certain extent.

[0003] Traditional remote sensing image target detection methods are usually based on digital image processing methods, that is, texture features are extracted first, and then methods such as template matching, shallow learning, and background modeling are used to detect and discriminate targets. However, these methods have poor generalization ability for remote sensing images and unsatisfactory detection effects.

[0004] Existing intelligent detection methods for remote sensing image targets mostly still use convolutional neural networks after improving classical target detection algorithms. Although they can improve the detection effect of remote sensing targets to varying degrees, the improved convolutional neural network can only model local information and lacks the ability of long-distance modeling and perception. Since remote sensing targets often have the characteristic of global dense distribution in images, remote sensing image target detection algorithms completely based on convolutional neural networks lack the ability of long-distance modeling and perception, have weak visual expression ability, and are prone to low detection rates of remote sensing targets. In addition, it is difficult for traditional artificial neural network models to balance detection accuracy and model lightweight, often sacrificing detection real-time performance to improve detection accuracy, or improving real-time performance but with insufficient detection accuracy. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that traditional remote sensing image target detection methods have poor generalization ability for remote sensing images and unsatisfactory detection effects, while existing intelligent detection methods for remote sensing image targets lack the ability of long-distance modeling and perception, have weak visual expression ability, and are prone to low detection rates of remote sensing targets.

[0006] To solve the above technical problems, a first aspect of an embodiment of the present invention discloses a method for remote sensing image target detection based on parameter optimization, and the method includes:

[0007] Obtain optical remote sensing image information; the optical remote sensing image information includes several optical remote sensing images;

[0008] Determine the anchor boxes for remote sensing image target detection, and use a preset remote sensing target detection model to perform positioning and recognition processing on the optical remote sensing image information to obtain an output feature map information set; the output feature map information set includes a number of output feature map information.

[0009] Perform post-processing on the output feature map information set to obtain a target image detection information set; the target image detection information set includes a number of target image detection information.

[0010] The remote sensing target detection model sequentially includes an input end (Input), a backbone feature extraction network (Backbone), a feature fusion network (Neck), and an output end (Head) from the input to the output direction.

[0011] The input end is used to receive the acquired optical remote sensing image information and perform preprocessing on it.

[0012] The preprocessing is to process the acquired optical remote sensing image information using a data augmentation method, and then use an adaptive image scaling method to unify the sizes of all optical remote sensing images.

[0013] There are two problems with the existing anchor box selection methods: (1) When the existing methods classify remote sensing targets in the dataset of this article, randomly select K data points from them as samples. If two points are in one cluster, this makes the clustering result not robust. (2) The existing methods use the distance between sample data as an index to divide K clusters in the remote sensing image dataset, and the centroid point of each cluster is obtained according to the mean value of all data points. This method regards the weights of different attributes in the distance formula as the same and does not consider the impact on the clustering effect under different attributes. When there are noise points or isolated points in the cluster, they will be far away from the centroid of the data, thus causing a large error in calculating the cluster centroid, having a great impact on the mean value calculation, and even causing the clustering centroid to deviate seriously from the dense area of the dataset, resulting in a deviation in the clustering result, leading to a large deviation in the selection of anchor boxes and affecting the success rate of image detection.

[0014] The above-mentioned determination of the anchor boxes for remote sensing image target detection, the anchor boxes are obtained by automatically learning the training dataset of the remote sensing target detection model, and the steps include:

[0015] S1, randomly select K points from the training dataset X as the initial clustering centers, each initial clustering center corresponds to a category and serves as the clustering center of its corresponding category, and the set of clustering centers is represented as C = {c 1 , c 2 ,..., c k}, where c i represents the clustering center of the i-th category;

[0016] S2. For each sample data x in the training data set X i , calculate the shortest distance D(x i ) between it and the current cluster centers, and classify the data sample x i into the category corresponding to the cluster center with the shortest distance to it.

[0017] S3. Calculate the probability that each sample data is selected as a cluster center next time. The calculation formula is:

[0018]

[0019] where P(x i ) is the probability that the sample data x i is selected as a cluster center next time; according to the values of all probabilities, divide the interval [0, 1] into several non-overlapping probability value intervals, and each probability value interval corresponds to a probability value for a sample data to be selected as a cluster center next time;

[0020] S4. Randomly generate a random number between [0, 1]. According to the probability value intervals obtained in step S3, determine the probability value interval to which the random number belongs, select the sample data corresponding to the belonging probability value interval, and use the data sample as the cluster center of its corresponding category;

[0021] S5. Repeat steps S2 to S4 until the change amount of the positions of the selected K cluster centers is less than a certain preset value, complete the clustering of the training data set X, and use the value boundaries of all data in each category as an anchor box.

[0022] The backbone feature extraction network includes a downsampling module (Focus layer), a feature extraction module (CBS layer), a residual module (C3), and a spatial pyramid pooling module (Spatial Pyramid Pooling, SPP). The backbone feature extraction network is used to extract the features of optical remote sensing image information.

[0023] The downsampling module is used to perform interval slicing operations on the preprocessed optical remote sensing image information longitudinally and horizontally in the image to obtain discrete slice information, then splice the discrete slice information, and finally perform convolution on the spliced information to obtain the first mapped feature.

[0024] The feature extraction module is used to perform two-dimensional convolution (Conv2d), normalization (BatchNorm), and activation layer operations on the first mapped feature in sequence to obtain the second mapped feature.

[0025] The residual module includes several classical residual structures (Bottlenecks). The residual module is used to perform convolutional layer operations on the second mapped feature input to it, add the value obtained after the convolutional layer operation to the original value of the second mapped feature to obtain the third mapped feature, so as to achieve residual feature transmission without increasing the output depth.

[0026] The spatial pyramid pooling module is used to perform several maximum pooling operations of different sizes on the third mapped feature, and then splice the results of the maximum pooling operations to obtain the image features of the optical remote sensing image.

[0027] The described feature fusion network includes a Feature Pyramid Networks (FPN) and a Path Aggregation Networks (PAN), which are used to realize the fusion of image features at different levels of the optical remote sensing image. The image features include class features and location features.

[0028] The described feature pyramid structure, from the input to the output end, sequentially includes a context conversion module, a feature extraction module, and an upsampling module; the feature output by the context conversion module passes through the feature extraction module and then through the upsampling module to obtain the fourth mapped feature, and the fourth mapped feature is spliced with the third mapped feature output by the residual module in the backbone feature extraction network to obtain the fifth mapped feature, and the fifth mapped feature is used as the output of the feature pyramid structure; the described path aggregation network structure, from the input to the output end, sequentially includes an input module, a residual module, a feature extraction module, and a context conversion module. The input module receives the fifth mapped feature output by the feature pyramid structure, and then passes it through the residual module and the feature extraction module respectively to obtain the sixth mapped feature, and the sixth mapped feature is spliced with the output of the feature extraction module in the feature pyramid structure to obtain the seventh mapped feature. After the seventh mapped feature passes through the residual module and the feature extraction module in sequence, the obtained feature is spliced with the output of the feature extraction module in the feature pyramid structure, and the spliced feature then passes through the context conversion module to obtain the output feature map information set;

[0029] The described context conversion module simultaneously realizes the functions of context information mining and self-attention learning integration, and promotes self-attention learning by making full use of the context information between the targets of adjacent remote sensing images, enhancing the expression ability of the output feature map.

[0030] The context conversion module first performs context encoding on all adjacent keys within the k×k picture grids obtained by segmenting the remote sensing image to obtain a feature matrix K with static context information 1 , K 1Concatenate with the Q space, and then perform two consecutive 1×1 convolution operations on the concatenation result to obtain the static context attention matrix A. The calculation process is as follows:

[0031] A = [K 1 , Q]W θ W δ ,

[0032] In the above formula, W θ is the first 1×1 convolution operation matrix, and W δ is the second 1×1 convolution operation matrix.

[0033] The context transformation module multiplies the context attention matrix A by the matrix V that has undergone 1×1 convolution to obtain the feature map matrix K 2 with dynamic context information. The calculation process is as follows:

[0034] K 2 = Conv 1×1 (V)A,

[0035] where Conv 1×1 (V) represents the matrix V that has undergone 1×1 convolution.

[0036] The context transformation module fuses K 2 and K 1 to obtain the output matrix Y with global and local information.

[0037] The output end is used to evaluate the difference between the output feature map information obtained by the feature fusion network and the real feature map information, and update the parameters of the remote sensing target detection model according to the evaluation result.

[0038] Use the remote sensing target detection loss function to evaluate the difference between the output feature map information obtained by the feature fusion network and the real feature map information. The remote sensing target detection loss function is obtained by calculating the overlap loss, calculating the center distance loss, and calculating the width and height loss. The formula is:

[0039]

[0040] In the formula, I EIOU represents the remote sensing target detection loss function, B represents the target ground truth box, and B i represents the target prediction box in the output feature map information. is the ratio of the intersection area to the union area of the target ground truth box and the target prediction box, b and b gt are the center points of the target prediction box and the target ground truth box respectively, ρ is the Euclidean distance between the above two center points, c is the diagonal distance of the smallest circumscribed rectangle covering the target prediction box and the target ground truth box, w and wgt are the lengths of the target prediction box and the target ground truth box, h and h gt are the widths of the target prediction box and the target ground truth box, C w and C h are the width and length of the minimum bounding rectangle covering the target prediction box and the target ground truth box respectively.

[0041] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0042] In the backbone feature extraction network and the feature fusion network of the present invention, the context conversion module is incorporated. It constructs an attention calculation unit more suitable for computer vision through grouped convolution and 1×1 convolution operations, effectively improving the correlation degree between K and Q, extracting static and dynamic context information of the input feature variables, and at the same time constructing a new loss function to optimize the remote sensing target detection ability of the model. The present invention selects the anchor box through parameter optimization, solves the problem that the existing anchor box detection method causes deviation in the clustering result, resulting in a large deviation in the selection of the anchor box and affecting the success rate of image detection, and optimizes the remote sensing target detection ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 is a schematic diagram of the composition of a remote sensing target detection model used in a remote sensing image target detection method disclosed in an embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of the composition of the context conversion module disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0047] In the description, claims, and above-mentioned drawings of the present invention, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or equipment.

[0048] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0049] Figure 1 It is a schematic diagram of the composition of a remote sensing target detection model used in a remote sensing image target detection method disclosed in an embodiment of the present invention; Figure 2 It is a schematic diagram of the composition of a context conversion module disclosed in an embodiment of the present invention.

[0050] The following will be described in detail respectively.

[0051] Embodiment 1

[0052] To solve the above technical problems, a first aspect of an embodiment of the present invention discloses a remote sensing image target detection method, the method comprising:

[0053] Obtain optical remote sensing image information; the optical remote sensing image information includes a plurality of optical remote sensing images;

[0054] Determine an anchor box for remote sensing image target detection, and use a preset remote sensing target detection model to perform positioning and recognition processing on the optical remote sensing image information to obtain an output feature map information set; the output feature map information set includes a plurality of output feature map information;

[0055] Perform post-processing on the output feature map information set to obtain a target image detection information set; the target image detection information set includes a plurality of target image detection information.

[0056] The remote sensing target detection model sequentially includes an input end (Input), a backbone feature extraction network (Backbone), a feature fusion network (Neck), and an output end (Head) from the input to the output direction.

[0057] The input end is used to receive the acquired optical remote sensing image information and preprocess it.

[0058] The preprocessing is to process the acquired optical remote sensing image information by using data augmentation methods such as Mosaic and flipping, and then use the adaptive image scaling method to unify the sizes of all optical remote sensing images.

[0059] The preprocessing also includes smoothing the acquired several optical remote sensing images to overcome the acquisition errors that occur within a certain period of time. Specifically, it includes gray processing the optical remote sensing images acquired within a period of time to obtain corresponding several gray matrices, calculating the eigenvectors of each gray matrix respectively to obtain an eigenvector group [x 1 ,x 2 ,…,x N , where N is the number of optical remote sensing images acquired within a period of time, calculating the cross-correlation matrix C of this eigenvector group, and performing eigenvalue decomposition on the cross-correlation matrix C to obtain:

[0060] C = VDV H ,

[0061] where V is the eigenvector matrix and D is the eigenvalue matrix. Normalize the diagonal elements of the matrix D and use them as the weight vector, perform weighted summation on the optical remote sensing images acquired within a period of time to obtain the smoothed value of the optical remote sensing images acquired within a period of time, which is used as the data after preprocessing.

[0062] The anchor boxes are obtained by automatically learning the training dataset of the remote sensing target detection model, and the steps include:

[0063] S1, randomly select K points from the training dataset X as the initial clustering centers. Each initial clustering center corresponds to a category and serves as the clustering center of its corresponding category. The set of clustering centers is represented as C = {c 1 ,c 2 ,...,c k}, where c i represents the clustering center of the i-th category;

[0064] S2, for each sample data x i in the training dataset X, calculate the shortest distance D(x i ) between it and the current respective clustering centers, and assign this data sample x i to the category corresponding to the clustering center with the shortest distance to it.

[0065] S3, calculate the probability that each sample data is selected as the clustering center next time, and its calculation formula is:

[0066]

[0067] where P(x i ) is the probability that the sample data x i is selected as the clustering center next time; according to the values of all probabilities, the interval [0, 1] is divided into several non-overlapping probability value intervals, and each probability value interval corresponds to the probability value of a sample data being selected as the clustering center next time; for example, there are four sample data, and the calculated probability values are 0.1, 0.2, 0.3, and 0.4 respectively, then the four divided probability value intervals are [0, 0.1), [0.1, 0.3), [0.3, 0.6), and [0.6, 1).

[0068] S4. Randomly generate a random number between [0, 1], and according to the probability value interval obtained in step S3, determine the probability value interval to which the random number belongs, select the sample data corresponding to the belonging probability value interval, and use this data sample as the clustering center of its corresponding category;

[0069] S5. Repeat steps S2 to S3 until the change amount of the positions of the selected K clustering centers is less than a certain preset value, complete the clustering of the training data set X, and use the value boundaries of all data in each category as an anchor box.

[0070] The backbone feature extraction network includes a downsampling module (Focus layer), a feature extraction module (CBS layer), a residual module (C3), and a spatial pyramid pooling module (Spatial Pyramid Pooling, SPP). The backbone feature extraction network is used to extract the features of optical remote sensing image information.

[0071] The downsampling module is used to perform interval slicing operations on the preprocessed optical remote sensing image information in the vertical and horizontal directions of the image to obtain discrete slice information, then splice the discrete slice information, and finally perform convolution on the spliced information to obtain the first mapped feature.

[0072] The feature extraction module is used to perform two-dimensional convolution (Conv2d), normalization (BatchNorm), and activation layer operations on the first mapped feature in sequence to obtain the second mapped feature. Among them, the role of two-dimensional convolution is to further extract target features, and the role of normalization is to make the input of each layer of neural network maintain the same distribution. The activation layer operation is implemented through the SiLU activation function.

[0073] The residual module includes several classical residual structures (Bottlenecks). The residual module is used to perform convolutional layer operations on the second mapped feature input to it, add the value obtained after the convolutional layer operation to the original value of the second mapped feature to obtain the third mapped feature, thereby achieving residual feature transmission without increasing the output depth.

[0074] The spatial pyramid pooling module is used to perform several maximum pooling operations of different sizes on the third mapped feature, and then splice the results of the maximum pooling operations to obtain the image features of the optical remote sensing image. The primary role of the spatial pyramid pooling module is to solve the problem of inconsistent input feature map sizes. In most object detection networks, a fully connected layer is generally used as the output layer at the end, which requires the size of the input feature map to be fixed. The SPP module, using fixed-block pooling operations, can achieve the same size output for inputs of different sizes, thus avoiding this problem. In addition, the fusion of features of different sizes in SPP is beneficial for the case where the target sizes in the image to be detected vary greatly.

[0075] The described feature fusion network includes a Feature Pyramid Networks (FPN) and a Path Aggregation Networks (PAN), and is used to achieve the fusion of image features at different levels of the optical remote sensing image. The image features include class features and location features.

[0076] The described feature pyramid structure, from the input end to the output end, successively includes a context conversion module, a feature extraction module, and an upsampling module; the feature output by the context conversion module passes through the feature extraction module and then the upsampling module to obtain the fourth mapped feature, and the fourth mapped feature is spliced with the third mapped feature output by the residual module in the backbone feature extraction network to obtain the fifth mapped feature; the described path aggregation network structure, from the input end to the output end, successively includes an input module, a residual module, a feature extraction module, and a context conversion module. The input module receives the fifth mapped feature output by the feature pyramid structure, and then passes it through the residual module and the feature extraction module respectively to obtain the sixth mapped feature. The sixth mapped feature is spliced with the output of the feature extraction module in the feature pyramid structure to obtain the seventh mapped feature. After the seventh mapped feature passes through the residual module and the feature extraction module in sequence, the obtained feature is spliced with the output of the feature extraction module in the feature pyramid structure, and the spliced feature then passes through the context conversion module to obtain the output feature map information set;

[0077] Specifically, the feature fusion network is composed of several feature extraction modules (CBS layers), residual modules (C3), upsampling modules, and C3_CoT modules, forming a Feature Pyramid Networks (FPN) and a Path Aggregation Networks (PAN). Among them, the feature pyramid structure is formed by the high-level features output by the context conversion module passing through the CBS module, then undergoing upsampling and concatenating with the features output by the third C3 structure (the 8th layer of the network structure) in the backbone feature extraction network. Then, it passes through the C3 and CBS modules respectively and then undergoes upsampling again. Finally, it is concatenated with the features generated by the second C3 module (the 4th layer of the network structure) in the backbone feature extraction network.

[0078] The path aggregation network is first composed of the features output by the feature pyramid structure, which are respectively passed through the C3 structure and the CBS module, concatenated with the feature maps output by the CBS module of the 16th layer of the network structure, then passed through the C3 structure and the CBS structure respectively, concatenated with the features output by the CBS structure of the 12th layer of the network structure, and then passed through the C3_CoT module. It is mainly used to achieve the fusion of features at different levels of the feature map.

[0079] In the convolutional network, as the number of convolutions increases, the feature levels change from low-level to high-level. Low-level features are closer to the visual content of the image itself. In low-level features, the position features of large objects and the category and position features of small objects are prominent. High-level features are more abstract and cannot be directly understood by humans. In high-level features, the category features of large objects are rich. Since in the convolutional network, as the number of convolutions increases, the feature levels change from low-level to high-level. Low-level features are closer to the visual content of the image itself. In low-level features, the position features of large objects and the category and position features of small objects are prominent. High-level features are more abstract and cannot be directly understood by humans. In high-level features, the category features of large objects are rich. Therefore, by using a feature fusion network, the image features are less likely to be lost. The feature pyramid structure realizes the transfer of the category features of medium and large objects in its high-level module to small objects in its low-level module. The path aggregation network structure realizes the transfer of the position features of large objects and the position and category features of small objects in its low-level module to medium objects in the high-level. The two complement each other and overcome their respective limitations, strengthening the feature extraction ability of the model. For the so-called large, medium, and small objects, an object with a size smaller than 32×32 pixels is considered a small object, an object with a size greater than or equal to 32×32 pixels and smaller than 96×96 pixels is considered a medium object, and an object with a size greater than or equal to 96×96 pixels is considered a large object. For the so-called high-level module, in the direction of information input to output, the module that first inputs information is the high-level module, and the module that later inputs information is the low-level module. The Head module is a detection structure that inputs features of three different sizes into the Detect module and respectively identifies remote sensing targets of large, medium, and small scales, thus well overcoming the limitations of the top features of the CNN network.

[0080] A context conversion module is introduced into the backbone feature extraction network and the feature fusion network to improve the model's global information acquisition ability while ensuring the local feature extraction ability, make full use of the input context information and guide the learning of the dynamic attention matrix, enhance the visual expression ability, and use the remote sensing target detection loss function to improve the precision of the remote sensing target recognition prediction box.

[0081] The structure of the context conversion module is as Figure 2 shown. It can be found that the original visual Transformer computing unit does not fully consider the connections between different spaces, and they are independent of each other, only learning the pairwise query-key relationship and ignoring the rich context between adjacent keys. Therefore, the present invention draws on CoTNet to improve the C3 structure and proposes a context conversion module. This module simultaneously realizes the functions of context information mining and self-attention learning integration, and promotes self-attention learning by making full use of the context information between the targets of adjacent remote sensing images, enhancing the expression ability of the output aggregated feature map.

[0082] For the input feature variable X of the context conversion module, it first performs context encoding on all adjacent keys within the k×k picture grids obtained by segmenting the remote sensing image to obtain a feature matrix K with static context information. 1 ; Then it concatenates K 1 with the Q space, and then performs two consecutive 1×1 convolution operations on the concatenated result to obtain the static context attention matrix A. The calculation process is as follows:

[0083] A = [K 1 , Q]W θ W δ ,

[0084] In the above formula, W θ is the first 1×1 convolution operation matrix, and W δ is the second 1×1 convolution operation matrix.

[0085] Subsequently, for the context attention matrix A, it multiplies with the matrix V that has undergone 1×1 convolution to obtain the feature map matrix K 2 with dynamic context information. The calculation process is as follows:

[0086] K 2 = Conv 1×1 (V)A,

[0087] where Conv 1×1 (V) represents the matrix V that has undergone 1×1 convolution.

[0088] Finally, it fuses K 2 with K 1 to obtain the output matrix Y with global and local information. The calculation process is expressed as:

[0089] Y = Fusion(K 1 , K 2 ).

[0090] The backbone network in the object detection model is a key part for extracting the hidden information of the input image. However, the original backbone network and the C3 structure in the feature fusion network are a fully convolutional structure. Although it has good local feature extraction ability, it lacks the ability to obtain global information. Therefore, in order to enable the model to further improve the global information acquisition ability while ensuring the local feature extraction ability, the present invention introduces the context conversion module into the original network model, improves the ResNet structure, and uses the idea of CoTNet to complete the construction of the C3 structure, forming a new context conversion module, enabling the network model to have the ability to obtain global information and improving the detection effect of remote sensing targets.

[0091] The output end is used to evaluate the difference between the output feature map information obtained by the feature fusion network and the real feature map information, and update the parameters of the remote sensing target detection model according to the evaluation result.

[0092] The difference between the output feature map information obtained by the feature fusion network and the real feature map information is evaluated by using the remote sensing target detection loss function. The loss function is positively correlated with the performance of the training model. However, the traditional CIOU loss function is too complex to measure the aspect ratio, resulting in a too slow convergence speed, and the aspect ratio cannot replace the individual length and width. Therefore, the present invention proposes a remote sensing target detection loss function, which solves the problem of large errors in the horizontal and vertical directions of the CIOU loss function, enhances the sensitivity to the width and height, and improves the convergence speed and regression accuracy. The remote sensing target detection loss function is realized by calculating the overlap loss, calculating the center distance loss, and calculating the width and height loss, and its formula is:

[0093]

[0094] In the formula, I EIOU represents the remote sensing target detection loss function, B represents the real target box, and B i represents the target prediction box in the output feature map information. is the ratio of the intersection area to the union area of the real target box and the target prediction box, b and b gt are the center points of the target prediction box and the real target box respectively, ρ is the Euclidean distance between the above two center points, c is the diagonal distance of the smallest circumscribed rectangle covering the target prediction box and the real target box, w and w gt are the lengths of the target prediction box and the real target box respectively, h and h gt are the widths of the target prediction box and the real target box respectively, C w and C h are the width and length of the smallest circumscribed rectangle covering the target prediction box and the real target box respectively.

[0095] The evaluation of the difference between the output feature map information obtained by the feature fusion network and the real feature map information includes:

[0096] Regarding the satellite remote sensing image data collected as a stationary random process, for the output feature map information and the real feature map information, autoregressive-moving average models, namely ARMA models, are respectively established to obtain the first ARMA model and the second ARMA model. For the coefficients of the two ARMA models, their cross-correlation matrix is calculated, and the maximum eigenvalue is calculated for this cross-correlation matrix. The maximum eigenvalue is used to discriminate the difference between the output feature map information and the real feature map information. At the same time, according to this maximum eigenvalue, the parameters of the remote sensing target detection model are updated.

[0097] The remote sensing target detection model is obtained through the following training steps:

[0098] Obtain the original image information set;

[0099] Perform annotation and data augmentation processing on the original image information set to obtain a first training image information set; the first training image information set includes a number of first training image information;

[0100] Determine the target training image information according to the first training image information set;

[0101] Use the target training image information and the target real image information to calculate the loss function, and use the loss function to train the first training model to obtain a second training model;

[0102] Judge whether the model training parameter information corresponding to the second training model meets the training termination condition to obtain a termination judgment result;

[0103] When the termination judgment result is no, use the second training model to update the first training model, and trigger the execution of determining the target training image information according to the third training image information set;

[0104] When the termination judgment result is yes, determine the second training model as the remote sensing target detection model.

[0105] The post-processing of the output feature map information set to obtain the target image detection information set includes:

[0106] Perform detection box decoding processing on the output feature map information set to obtain a detection box information set; the detection box information set includes a number of detection box information;

[0107] Perform category discrimination processing on the detection box information set to obtain an image category information set; the image category information set includes a number of image category information.

[0108] The detection box described above can be an anchor box.

[0109] The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0110] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each implementation can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0111] Finally, it should be noted that: The disclosed method for remote sensing image target detection based on parameter optimization according to the embodiments of the present invention is only the preferred embodiments of the present invention, and is only used to illustrate the technical solutions of the present invention, rather than limiting them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for remote sensing image target detection based on parameter optimization, characterized in that, the method includes: Obtain optical remote sensing image information; the optical remote sensing image information includes several optical remote sensing images; Determine the anchor boxes for remote sensing image target detection, and use the remote sensing target detection model to perform positioning and recognition processing on the optical remote sensing image information to obtain an output feature map information set; the output feature map information set includes several output feature map information; Perform post-processing on the output feature map information set to obtain a target image detection information set; the target image detection information set includes several target image detection information; The remote sensing target detection model sequentially includes an input end, a backbone feature extraction network, a feature fusion network, and an output end from the input to the output direction; The backbone feature extraction network includes a downsampling module, a feature extraction module, a residual module, and a spatial pyramid pooling module; the backbone feature extraction network is used to extract the features of the optical remote sensing image information; The downsampling module is used to perform interval slicing operations on the preprocessed optical remote sensing image information in the vertical and horizontal directions of the image to obtain discrete slice information, then splice the discrete slice information, and finally perform convolution on the spliced information to obtain the first mapped feature; The feature extraction module is used to perform two-dimensional convolution, normalization, and activation layer operations on the first mapped feature in sequence to obtain the second mapped feature; The residual module includes several classic residual structures, and the residual module is used to perform convolution layer operations on the input second mapped feature, and add the value obtained after the convolution layer operation to the original value of the second mapped feature to obtain the third mapped feature, so as to achieve residual feature transmission without increasing the output depth; The spatial pyramid pooling module is used to perform maximum pooling operations of several different sizes on the third mapped feature, and then splice the results of the maximum pooling operations to obtain the image features of the optical remote sensing image.

2. The method for remote sensing image target detection based on parameter optimization according to claim 1, characterized in that, the method includes: The determination of the anchor boxes for remote sensing image target detection is obtained by automatically learning the training data set of the remote sensing target detection model, and the steps include: S1. Randomly select K points from the training data set X as the initial cluster centers. Each initial cluster center corresponds to a category and serves as the cluster center for its corresponding category. The set of cluster centers is denoted as C = {c 1 , c 2 ,..., c k}, where c i represents the cluster center of the i-th category; S2. For each sample data x in the training data set X i , calculate the shortest distance D(x i ) between it and each current cluster center, and classify the sample data x i into the category corresponding to the cluster center with the shortest distance to it; S3. Calculate the probability that each sample data will be selected as the clustering center next time, and its calculation formula is: Among them, P(x i ) is the probability that the sample data x i is selected as the clustering center next time; according to the values of all probabilities, the interval [0, 1] is divided into several non-overlapping probability value intervals, and each probability value interval corresponds to the probability value that a sample data is selected as the clustering center next time; S4. Randomly generate a random number between [0, 1], and according to the probability value range obtained in step S3, judge the probability value range to which the random number belongs, select the sample data corresponding to the probability value range, and use this data sample as the clustering center of its corresponding category; S5. Repeat steps S2 to S4 until the change amount of the positions of the selected K clustering centers is less than a certain preset value, complete the clustering of the training data set X, and use the value boundaries of all data in each category as an anchor box.

3. The method for remote sensing image target detection based on parameter optimization according to claim 1, characterized in that, the method includes: The described feature fusion network includes a feature pyramid structure and a path aggregation network structure, and is used to achieve the fusion of image features at different levels of optical remote sensing images; The described feature pyramid structure, from the input end to the output end, sequentially includes a context transformation module, a feature extraction module, and an upsampling module; the features output by the context transformation module pass through the feature extraction module and then through the upsampling module to obtain the fourth mapped feature. The fourth mapped feature is concatenated with the third mapped feature output by the residual module in the backbone feature extraction network to obtain the fifth mapped feature, and the fifth mapped feature is used as the output of the feature pyramid structure; the described path aggregation network structure, from the input end to the output end, sequentially includes an input module, a residual module, a feature extraction module, and a context transformation module. The input module receives the fifth mapped feature output by the feature pyramid structure, and then passes it through the residual module and the feature extraction module respectively to obtain the sixth mapped feature. The sixth mapped feature is concatenated with the output of the feature extraction module in the feature pyramid structure to obtain the seventh mapped feature. After the seventh mapped feature passes through the residual module and the feature extraction module in sequence, the obtained feature is concatenated with the output of the feature extraction module of the feature pyramid structure, and the concatenated feature then passes through the context transformation module to obtain the output feature map information set.

4. The remote sensing image target detection method based on parameter optimization according to claim 3, characterized in that, the method includes: The described context conversion module simultaneously realizes the functions of context information mining and self-attention learning integration, and promotes self-attention learning by making full use of the context information between the targets of adjacent remote sensing images, enhancing the expression ability of the output feature map; the context conversion module performs context encoding on all adjacent keys within the k×k picture grids obtained by segmenting the remote sensing image to obtain a feature matrix K with static context information 1 , concatenate K 1 with the Q space, and then perform two consecutive 1×1 convolution operations on the concatenated result to obtain a static context attention matrix A, and its calculation process is as follows: A = [K 1 , Q]W θ W δ , In the above formula, W θ is the matrix of the first 1×1 convolution operation, and W δ is the matrix of the second 1×1 convolution operation; The context conversion module multiplies the context attention matrix A by the matrix V that has undergone 1×1 convolution to obtain the feature map matrix K with dynamic context information 2 , and its calculation process is as follows: K 2 = Conv 1×1 (V)A, Among them, Conv 1×1 (V) represents the matrix V after 1×1 convolution; The context conversion module combines K 2 with K 1 to obtain an output matrix Y with global and local information.

5. The remote sensing image target detection method based on parameter optimization according to claim 1, characterized in that, the method includes: The output end is used to evaluate the difference between the output feature map information obtained by the feature fusion network and the real feature map information, and update the parameters of the remote sensing target detection model according to the evaluation result.

6. The remote sensing image target detection method based on parameter optimization according to claim 5, characterized in that, the method includes: The difference between the output feature map information obtained by the feature fusion network and the real feature map information is evaluated by using a remote sensing target detection loss function.

7. The remote sensing image target detection method based on parameter optimization according to claim 6, characterized in that, the method includes: The remote sensing target detection loss function is obtained by calculating the overlap loss, calculating the center distance loss, and calculating the width and height loss, and its formula is: Where, I EIOU represents the loss function for remote sensing object detection, B represents the ground truth bounding box, and B i represents the predicted bounding box of the object in the output feature map information. is the ratio of the intersection area to the union area of the ground truth bounding box and the predicted bounding box of the object. b and b gt are the center points of the predicted bounding box and the ground truth bounding box of the object respectively. ρ is the Euclidean distance between the two center points. c is the diagonal distance of the smallest bounding rectangle covering the predicted bounding box and the ground truth bounding box. w and w gt are the lengths of the predicted bounding box and the ground truth bounding box of the object respectively. h and h gt are the widths of the predicted bounding box and the ground truth bounding box of the object respectively. C w and C h are the width and length of the smallest bounding rectangle covering the predicted bounding box and the ground truth bounding box respectively.

Citation Information

Patent Citations

  • Remote sensing image detection system and method

    CN114842001A