A method, device, equipment and medium for dynamic target recognition of remote sensing images
By introducing a convolutional intensive prediction network (CDN) and a dense prediction module (DPM) in remote sensing image processing, combined with a region generation network (RPN), the problem of large dynamic target recognition errors in remote sensing images is solved, and higher recognition accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202411256556.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-09
AI Technical Summary
The prior art has large errors in dynamic target recognition in remote sensing images, making it difficult to effectively deal with huge amount of information and high false alarm rate and misidentification rate.
Convolutional intensive prediction network (CDN) is used to combine dense prediction module (DPM) and region generation network (RPN). By training dynamic target recognition models, candidate areas are extracted, target areas are determined, dynamic target images are extracted, and dynamic target images are extracted, and dense predictions are performed on high-dimensional feature maps to improve recognition accuracy.
By narrowing the detection range and reducing background impact, the accuracy and efficiency of dynamic target recognition are improved, and the false alarm rate and misidentification rate are reduced.
Smart Images

Figure CN119206530B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition, and in particular relates to a method, device, equipment and medium for dynamic target recognition of remote sensing images. Background Art
[0002] Target detection and recognition in remote sensing images is performed in a large-scale scene with low signal-to-clutter ratio and high clutter. Targets are often mixed in the background containing a lot of information. This makes it difficult to extract the target of interest from the remote sensing image.
[0003] At present, there are two main applications of remote sensing image target detection and recognition: one is to establish a general target feature model for target detection and recognition; however, in remote sensing image processing, it is necessary to analyze image data in a large scene range, detect targets of interest and identify target categories. The amount of information that needs to be processed is huge, so the above general target recognition methods will lead to high false alarm rates and misidentification rates.
[0004] The other method is for the readers to directly analyze the images and use their long-term training experience to detect and interpret the targets. However, this type of image processing has a huge workload, low data processing efficiency, and the target detection and interpretation results are affected by the state of the readers, and there may still be high errors. Summary of the invention
[0005] In order to solve the problem of large errors in dynamic target recognition in remote sensing images in the prior art, the present invention provides a method, device, equipment and medium for dynamic target recognition in remote sensing images.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A method for dynamic target recognition of remote sensing images, the method comprising:
[0008] The dense prediction module DPM is introduced into CNN to construct a convolutional dense prediction network CDN, wherein the CDN includes an input layer, a feature extraction module, a dense prediction module and a classification layer; the CDN is trained to obtain a dynamic target recognition model;
[0009] Acquire multiple target remote sensing images taken continuously in the same area; extract multiple candidate regions from the multiple target remote sensing images through a region generation network RPN; determine the variance mean of the candidate regions, and determine the candidate regions with dynamic targets as target regions according to the variance mean; determine the foreground image through the maximum inter-class variance of the target region, and extract the image of the dynamic target in the foreground image;
[0010] The image of the dynamic target is input into the dynamic target recognition model, the input image is preprocessed through the input layer to obtain a preprocessed image, the feature extraction module extracts multi-scale features of the preprocessed image through multiple convolutional layers, the multi-scale features are enhanced through the dense prediction module to obtain enhanced features, and the enhanced features are classified through the classification layer to obtain a recognition result.
[0011] Optionally, determining the variance mean of the candidate region includes:
[0012] For each generated candidate region, the non-maximum suppression algorithm NMS is used to eliminate redundant candidate regions with high overlap and determine the variance mean of the remaining candidate regions.
[0013] Optionally, determining the candidate area where the dynamic target exists as the target area according to the variance mean includes:
[0014] For multiple candidate regions at the same position but at different time points, the pixel-level difference is calculated to determine the variance mean of the difference map, and the regions are sorted according to the variance mean to obtain a variance mean distribution histogram;
[0015] Determining the noise threshold of the target remote sensing image through the variance mean distribution histogram;
[0016] The candidate area where the dynamic target exists is determined as the target area according to the noise threshold.
[0017] Optionally, extracting an image of a dynamic target in a foreground image includes:
[0018] Using image segmentation technology to identify and extract the dynamic target image in the foreground image;
[0019] Determine an edge image of a dynamic target image, and then obtain image coordinates of the edge image;
[0020] The area corresponding to the image coordinates is extracted to obtain a dynamic target image.
[0021] Optionally, determining the edge image of the dynamic target image includes:
[0022] Performing Gaussian smoothing on the extracted dynamic target image to obtain a processed dynamic target image;
[0023] Determine the grayscale difference of each pixel in the processed dynamic target image in the horizontal direction and the vertical direction respectively, and determine the edge strength of the pixel through the grayscale difference;
[0024] According to the edge strength, each pixel is filtered by a non-maximum suppression algorithm;
[0025] Double threshold detection and edge connection are performed on the filtering result to determine the edge image of the dynamic target image.
[0026] Optionally, obtaining the image coordinates of the edge image includes:
[0027] Dividing the edge image into a plurality of grids, each grid having a plurality of bounding boxes;
[0028] Determining a confidence level of each bounding box, and adjusting the bounding box according to the confidence level to obtain adjusted bounding box coordinates;
[0029] The bounding box coordinates are transformed to obtain image coordinates of the edge image.
[0030] Optionally, the training of the CDN to obtain a dynamic target recognition model includes:
[0031] Acquire a training sample, wherein the training sample includes a sample image and its corresponding true result;
[0032] The sample image is input into the CDN to obtain a recognition result, and the CDN is trained with the goal of minimizing the difference between the recognition result and the corresponding true result to obtain a dynamic target recognition model.
[0033] A dynamic target recognition device for remote sensing images, the device comprising:
[0034] An acquisition module is used to acquire multiple continuously shot target remote sensing images of the same area;
[0035] A preprocessing module is used to extract multiple candidate regions from the multiple target remote sensing images through a region generation network RPN; determine the variance mean of the candidate regions, and determine the candidate regions with dynamic targets as target regions according to the variance mean; determine the foreground image through the maximum inter-class variance of the target region, and extract the image of the dynamic target in the foreground image;
[0036] A construction module is used to introduce a dense prediction module DPM into CNN to construct a convolutional dense prediction network CDN, wherein the CDN includes an input layer, a feature extraction module, a dense prediction module and a classification layer; the input image is preprocessed through the input layer to obtain a preprocessed image, the feature extraction module extracts multi-scale features of the preprocessed image through multiple convolutional layers, the multi-scale features are enhanced through the dense prediction module to obtain enhanced features, and the enhanced features are classified through the classification layer to obtain prediction results; the CDN is trained through the pre-acquired sample features to obtain a dynamic target recognition model;
[0037] The recognition module is used to input the image of the dynamic target into the dynamic target recognition model to obtain the recognition result.
[0038] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the dynamic target recognition method of the remote sensing image is implemented.
[0039] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the dynamic target recognition method of the remote sensing image is realized.
[0040] The dynamic target recognition method of remote sensing images provided by the present invention has the following beneficial effects:
[0041] First, by extracting candidate areas and determining the target areas where dynamic targets exist from the candidate areas, the detection range is narrowed and the amount of information that needs to be processed by the subsequent model is reduced, which is conducive to improving the accuracy of detection. Secondly, the image of the dynamic target is extracted to further reduce the influence of the background image. Finally, the extracted target image is input into the CDN. Through the CDN, the target image is directly predicted densely on the high-dimensional feature map, which can retain more detailed information of the dynamic target and use feature information at different levels for prediction, thereby improving the prediction accuracy and further improving the accuracy of dynamic target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0043] Figure 1 The present invention is a flowchart of a method for dynamic target recognition in remote sensing images provided according to an exemplary embodiment of the present invention.
[0044] Figure 2 The present invention is a block diagram of a dynamic target recognition device for remote sensing images according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0045] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.
[0046] The technical solutions provided by various embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.
[0047] First, the present invention provides a method for dynamic target recognition of remote sensing images, specifically, Figure 1 As shown, the following steps are included:
[0048] S101, acquiring a plurality of continuously captured target remote sensing images of the same area.
[0049] In one embodiment, satellite remote sensing can be performed, such as a SAR (synthetic aperture radar) satellite, which uses radar technology, is not obstructed by clouds, and has all-weather and all-day observation capabilities; or an optical satellite, which records visible light and infrared radiation through photosensitive elements to provide high-resolution, color images.
[0050] In another embodiment, aerial remote sensing can be performed, such as high-altitude remote sensing aircraft, which can obtain large-scale remote sensing image data; or small-scale, high-resolution remote sensing data can be obtained through drones; or helicopters, etc., the present invention is not limited to this.
[0051] S102: extracting multiple candidate regions from the multiple target remote sensing images through a region generation network RPN.
[0052] Specifically, after the target remote sensing image is grayed out preprocessed, there will be significant differences in grayscale between the dynamic target and the surrounding image. Using the ROI (regions of interest) grayscale histogram, these differences can be clearly displayed through the grayscale histogram, providing global grayscale information. Using the ROI regional grid map, regional block analysis or texture analysis is performed, and the region of interest is divided into several small squares to reveal local heterogeneity and detail features, and determine the ROI to mark the region of interest.
[0053] For example, the feature extraction of the target remote sensing image is performed through a dilated convolution with a dilation rate of 3x3 to obtain a feature map. At each feature point of the feature map, the RPN (Region Proposal Network) first generates multiple anchor boxes (Anchor Boxes), which serve as the initial shape of the candidate region. The number and shape (such as area and aspect ratio) of the anchor box can be adjusted according to the specific task. For example, the shape of the anchor box includes different aspect ratios such as 1:1, 1:2, and 2:1, as well as a variety of area sizes. Secondly, the RPN maps the anchor box back to the original image, and calculates the scaling ratio of the feature map to the original image to ensure that the position of the candidate region on the original image is accurate. For each anchor box, the RPN uses a small neural network for binary classification. For example, two parallel 1x1 convolutional layers can be used for binary classification to determine whether the target object is contained in the anchor box. At the same time, the RPN also performs bounding box regression on the anchor box to adjust the position and size of the anchor box to make it closer to the real object bounding box. Bounding box regression outputs the offset of the center point of the anchor box and the scaling ratio of width and height. Based on the results of classification and regression, RPN will filter out candidate regions containing objects and fine-tune them. Finally, RPN outputs a set of candidate regions that may contain target objects in the image. Each candidate region will have a corresponding score and precise location information.
[0054] S103: Determine a mean variance of the candidate region, and determine, according to the mean variance, that the candidate region where the dynamic target exists is the target region.
[0055] In this step, first, since there are multiple overlapping candidate regions, for each generated candidate region, NMS (Non-Maximum Suppression) is used to eliminate redundant candidate regions with high overlap and determine the variance mean of the remaining candidate regions.
[0056] Secondly, the pixel-level differences of multiple candidate areas can be determined to obtain a difference map; the distribution histogram of the variance mean of the difference map is determined, and the target area is determined based on the distribution histogram. Specifically, for candidate areas at the same location but different time points, the pixel-level differences are calculated, and then the variance mean of the difference map is determined, and the variance mean distribution histogram is obtained by sorting them according to the variance mean. Blocks with low and similar variance means are identified as static targets, while those with significant differences are regarded as dynamic targets. Candidate areas with dynamic targets are determined as target areas.
[0057] Furthermore, in this step, the noise threshold of the target remote sensing image can also be determined by the variance mean distribution histogram, and then the candidate area where the dynamic target exists is determined as the target area according to the noise threshold.
[0058] In this step, based on the noise threshold, the candidate regions can be subdivided into three categories: candidate regions with only static regions, candidate regions with regularly moving targets, and candidate regions with abnormally moving targets.
[0059] In addition, since there is a large amount of high-variance noise in remote sensing images, and this noise has obvious changes in time and space, amplitude and phase, and is random, it is difficult to accurately fit it with one or more probability distributions, resulting in a large amount of noise in the candidate area. In this step, the candidate area can also be filtered by the noise threshold.
[0060] S104, determining a foreground image by using the maximum inter-class variance of the target area, and extracting an image of the dynamic target in the foreground image.
[0061] Specifically, first, the optimal threshold can be found by calculating the inter-class variance value, and then the foreground image can be determined by the optimal threshold. Then, the dynamic target image in the foreground image can be identified and extracted by using image segmentation technology; the edge image of the dynamic target image can be determined, and then the image coordinates of the edge image can be obtained; the area corresponding to the image coordinates can be extracted to obtain the dynamic target image.
[0062] In one embodiment, a threshold value T is found to divide all pixels of the image into two categories. One category has pixel values less than or equal to T and is determined as the background area, and the other category has pixel values greater than T and is determined as the foreground area. When the inter-class variance of the two categories reaches the maximum value, the T value is considered to be the most appropriate threshold.
[0063] The percentage of background pixels is:
[0064]
[0065] The proportion of foreground pixels is:
[0066]
[0067] Calculate the average value of each category and the average gray value of the background:
[0068]
[0069] Average gray value of foreground:
[0070]
[0071] Find the threshold under the maximum inter-class variance in turn and return the calculated threshold. The inter-class variance is as follows:
[0072] g=w1*(μ-μ1) 2 +w2*(μ-μ2)2 ;
[0073] Among them, Sum is the sum of pixels, N1 is the number of all pixels in the background area, w1 is the ratio of the number of pixels in the background area to the total number of pixels in the image, and μ1 is the average pixel value; N2 is the number of all pixels in the foreground area, w2 is the ratio of the number of pixels in the foreground area to the total number of pixels in the image, and μ2 is the average pixel value; μ is the average pixel value of 0-255. Then analyze the shape, size, and connectivity of the target object to extract the feature set related to the target.
[0074] In another embodiment, the extracted dynamic target image is subjected to Gaussian smoothing to obtain a processed dynamic target image; the grayscale difference in the horizontal and vertical directions of each pixel in the processed dynamic target image is determined respectively, and the edge strength of the pixel is determined by the grayscale difference; each pixel is filtered by a non-maximum suppression algorithm according to the edge strength; the filtering result is subjected to double threshold detection and edge connection to determine the edge image of the dynamic target image. The edge image is divided into multiple grids, each grid having multiple bounding boxes; the confidence of each bounding box is determined, and the bounding box is adjusted according to the confidence to obtain the adjusted bounding box coordinates; the bounding box coordinates are converted to obtain the image coordinates of the edge image.
[0075] For example, the first step is to detect the boundary or area of the dynamic target image.
[0076] Image segmentation technology is used to identify and extract dynamic target images, and Gaussian smoothing is performed.
[0077]
[0078] Analysis shows that:
[0079] f(x,y)=f(x)*f(y);
[0080] Perform a two-dimensional Gaussian filter on the image and set the image template radius to r, then we have
[0081]
[0082] Among them, the coordinates of the center pixel are x and y are the offsets relative to the center point, * represents the convolution operator, f(x, y) represents the value of the two-dimensional Gaussian function at the point (x, y), σ is the standard deviation of the Gaussian function, and Gaussian blur is applied to remove noise and reduce the recognition of false edges.
[0083] The gradient magnitude and direction are then calculated using the edge detection method of the differential operator.
[0084] The pixel values within the region are calculated using a 3×3 template, and edge detection is achieved by leveraging the differences generated from the gray-scale values of the pixels within a specific region. The differential operator in the horizontal direction is denoted as Fx, and the differential operator in the vertical direction is denoted as Fy. After the convolution operation, the gray-scale differences in the horizontal and vertical directions can be obtained by using two templates respectively to calculate the gray-scale differences in the horizontal and vertical directions, and the gradient magnitudes in the horizontal and vertical directions can be obtained.
[0085] Template in the horizontal direction:
[0086]
[0087] Template in the vertical direction:
[0088]
[0089] For each pixel point in the input image, convolution operations are performed with the two templates respectively to calculate the gray-scale differences in the horizontal and vertical directions. Then, the sum of the absolute values of the gray-scale differences in the two directions is taken as the edge intensity of this pixel point.
[0090] Subsequently, non-maximum suppression is carried out. Non-maximum suppression is performed along the gradient orientation. In the first step, the orientation of the gradient is calculated first, and then it is compared whether the gradient value in the current direction of the current pixel point is the maximum value in this direction within the 3×3 region in order to better compare the magnitudes of the gradients. First, it is assumed that the gradient changes are continuous between adjacent pixel points along the vertical and horizontal directions. Then, the directions of the gradient magnitudes of this pixel point along the X and Y directions are compared. If Fx < Fy (and Fx and Fy are in the same direction), the weight is set and the two pairs of points in this direction in the relevant neighborhood are weighted. Finally, non-maximum suppression operation is performed on the gradient value of this point and the weighted gradient values of the other two pairs of points.
[0091] Subsequently, double-threshold detection and edge connection are carried out. Two thresholds are set, namely TL and TH. Among them, those greater than TH are all detected as edges, and those lower than TL are all detected as non-edges. For the intermediate pixel points, if they are adjacent to the pixel points determined to be edges, they are determined to be edges; otherwise, they are non-edges. The Fx in the partial x direction and the Fy in the partial y direction are calculated respectively, and the final edge image is obtained using the following formula:
[0092]
[0093] Second step: Analyze the boundaries or regions.
[0094] Resize the image and divide it into S×S grids. Each grid has B bounding boxes, and the center of the bounding box is in the corresponding grid. Each grid is responsible for predicting B bounding boxes. The size is reflected by the rectangular box, so each bounding box must have a width w and a height h. Knowing the center coordinates x and y of the bounding box is equivalent to the position of the bounding box.
[0095] Set each grid to generate B bounding boxes, that is, predict B prediction boxes, represented by Pr(Object). If there is an object in the bounding box, it is 1, otherwise it is 0. It is reflected by the intersection of the box where the object is actually located or the manually labeled box and the predicted bounding box. The closer the predicted box is to the manually labeled box, the larger the intersection of the two ratios is, and the more accurate it is. Set the confidence CS, which is expressed by the following formula:
[0096] CS = Pr(Object)*IoU;
[0097] When there is no object in this bounding box, Pr(Object) is 0 and the confidence is 0; when there is an object, Pr(Object) is 1 and the confidence is equal to the intersection-union ratio.
[0098] Step 3: Extract the boundary coordinate points of the area to obtain the precise boundary coordinate values of the dynamic target image.
[0099] The prediction box is fine-tuned to get the anchor box. The position of the anchor box is fixed. Set the coordinates of the center of the anchor box region:
[0100] center_x=cx+0.5;
[0101] center_y=cy+0.5;
[0102] The center coordinates of the prediction box are generated in the following way:
[0103] bx=cx+σ(tx);
[0104] by=cy+σ(ty);
[0105] Where tx and ty are real numbers, and σ(x) is the defined activation function, which is defined as follows:
[0106]
[0107] The activation function value is between 0 and 1. When tx=ty=0, bx=cx+0.5, by=cy+0.5, the center of the prediction box coincides with the center of the anchor box. The size of the anchor box is ph, pw, and the size of the prediction box is:
[0108] bh = pheth;
[0109] bw=pwetw;
[0110] Then the coordinates of the predicted box (tx, ty, tw, th) are obtained.
[0111] S105, introducing the dense prediction module DPM into CNN to build a convolutional dense prediction network CDN.
[0112] The CDN (Convolutional Dense Prediction Module Network) includes an input layer, a feature extraction module, a dense prediction module and a classification layer; the input image is preprocessed by the input layer to obtain a preprocessed image, the feature extraction module extracts multi-scale features of the preprocessed image through multiple convolutional layers, the multi-scale features are enhanced by the dense prediction module to obtain enhanced features, the enhanced features are classified by the classification layer to obtain recognition results; the CDN is trained to obtain a dynamic target recognition model.
[0113] Specifically, the input layer is used to receive input image data and perform necessary preprocessing, such as normalization, resizing, etc., to ensure the consistency of input data and the efficiency of network processing.
[0114] The feature extraction module includes convolutional layers, pooling layers, and fully connected layers. Through multiple convolutional layers, different convolution kernels are used to extract local features of the image. Each convolutional layer is followed by an activation function to increase the nonlinear ability of the network. Its output can be calculated by the following formula:
[0115] Y ij =∑ m ∑ n W mn X (i+m)(j+n) +b;
[0116] Among them, Y ij is the value of position (i, j) in the output feature map, X (i+m)(j+n) is the value of the corresponding position in the input feature map or image, W mn is the weight of the convolution kernel at position (m,n), and b is the bias term. By inserting a pooling layer between convolutional layers, the spatial dimension and number of parameters of the data are reduced while retaining important features. The features extracted by the convolutional layer and the pooling layer are globally integrated through the fully connected layer.
[0117] The dense prediction module takes the feature map output by the fully connected layer as input for further feature processing and prediction. First, the feature map output by the CNN is further processed, such as using additional convolutional layers or attention mechanisms to enhance the feature representation; second, dense prediction is performed, using the dense connection characteristics of DPM to independently predict or classify each position in the feature map. This can be achieved by applying a small fully connected layer or convolutional layer on the feature map. Its output can be calculated by the following formula:
[0118] FusedFeature=Conv(UpSample(F1)+F2);
[0119] Among them, F1 and F2 are feature maps from different levels, UpSample is an upsampling operation, and Conv is a convolution operation used to fuse features.
[0120] Finally, the output of this dense prediction module is passed to the classification layer to generate the overall classification result of the image. The classification layer is usually a fully connected layer followed by a softmax activation function to convert the output into a probability distribution.
[0121] After the CDN is built, it is necessary to train the CDN. Specifically, first obtain training samples, which include sample images and their corresponding real results; input the sample images into the CDN to obtain recognition results, and train the CDN to obtain a dynamic target recognition model with the goal of minimizing the difference between the recognition results and the corresponding real results.
[0122] S106: Input the image of the dynamic target into the dynamic target recognition model to obtain a recognition result.
[0123] The above method first extracts candidate regions and determines the target region where the dynamic target exists from the candidate regions to narrow the detection range and reduce the amount of information that the subsequent model needs to process, which is conducive to improving the accuracy of detection. Secondly, the image of the dynamic target is extracted to further reduce the influence of the background image. Finally, the extracted target image is input into the CDN. The target image is directly densely predicted on the high-dimensional feature map through the CDN, which can retain more detailed information of the dynamic target and use feature information at different levels for prediction, thereby improving the prediction accuracy and further improving the accuracy of dynamic target recognition.
[0124] Secondly, the present invention also provides a dynamic target recognition device for remote sensing images, such as Figure 2 As shown, including:
[0125] The construction module 201 is used to introduce the dense prediction module DPM into CNN to construct a convolutional dense prediction network CDN, wherein the CDN includes an input layer, a feature extraction module, a dense prediction module and a classification layer; the CDN is trained to obtain a dynamic target recognition model.
[0126] The acquisition module 202 is used to acquire a plurality of continuously captured target remote sensing images of the same area.
[0127] The preprocessing module 203 is used to extract multiple candidate regions from the multiple target remote sensing images through the region generation network RPN; determine the variance mean of the candidate regions, and determine the candidate regions with dynamic targets as target regions based on the variance mean; determine the foreground image through the maximum inter-class variance of the target region, and extract the image of the dynamic target in the foreground image.
[0128] The recognition module 204 is used to input the image of the dynamic target into the dynamic target recognition model, preprocess the input image through the input layer to obtain a preprocessed image, the feature extraction module extracts multi-scale features of the preprocessed image through multiple convolutional layers, enhances the multi-scale features through the dense prediction module to obtain enhanced features, and classifies the enhanced features through the classification layer to obtain a recognition result.
[0129] By using the above device, firstly, the candidate area is extracted, and the target area where the dynamic target exists is determined from the candidate area to narrow the detection range, thereby reducing the amount of information that needs to be processed by the subsequent model, which is beneficial to improving the accuracy of detection. Secondly, the image of the dynamic target is extracted to further reduce the influence of the background image. Finally, the extracted target image is input into the CDN, and the target image is directly densely predicted on the high-dimensional feature map through the CDN, which can retain more detailed information of the dynamic target and use feature information at different levels for prediction, thereby improving the prediction accuracy, and further improving the accuracy of dynamic target recognition.
[0130] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 The steps of a method for dynamic target recognition in remote sensing images are provided.
[0131] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The steps of a method for dynamic target recognition in remote sensing images are provided.
[0132] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0134] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0136] It should be noted that the above-described specific implementations can enable those skilled in the art to more fully understand the invention, but do not limit the invention in any way. Therefore, although the invention has been described in detail in this specification, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the protection scope of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved.
Claims
1. A method for dynamic target recognition in remote sensing images, characterized in that: The method comprises: The dense prediction module DPM is introduced into CNN to construct a convolutional dense prediction network CDN, wherein the CDN includes an input layer, a feature extraction module, a dense prediction module and a classification layer; the CDN is trained to obtain a dynamic target recognition model; Acquire multiple target remote sensing images taken continuously in the same area; extract multiple candidate regions from the multiple target remote sensing images through a region generation network RPN; determine the variance mean of the candidate regions, and determine the candidate regions with dynamic targets as target regions according to the variance mean; determine the foreground image through the maximum inter-class variance of the target region, and extract the image of the dynamic target in the foreground image; The image of the dynamic target is input into the dynamic target recognition model, the input image is preprocessed through the input layer to obtain a preprocessed image, the feature extraction module extracts multi-scale features of the preprocessed image through multiple convolutional layers, the multi-scale features are enhanced through the dense prediction module to obtain enhanced features, and the enhanced features are classified through the classification layer to obtain a recognition result.
2. The method for dynamic target recognition of remote sensing images according to claim 1, characterized in that: Determining the variance mean of the candidate region includes: For each generated candidate region, the non-maximum suppression algorithm NMS is used to eliminate redundant candidate regions with high overlap and determine the variance mean of the remaining candidate regions.
3. The method for dynamic target recognition of remote sensing images according to claim 1, characterized in that: The step of determining the candidate area where the dynamic target exists as the target area according to the variance mean value includes: For multiple candidate regions at the same position but at different time points, the pixel-level difference is calculated to determine the variance mean of the difference map, and the regions are sorted according to the variance mean to obtain a variance mean distribution histogram; Determining the noise threshold of the target remote sensing image through the variance mean distribution histogram; The candidate area where the dynamic target exists is determined as the target area according to the noise threshold.
4. The method for dynamic target recognition of remote sensing images according to claim 1, characterized in that: The method of extracting an image of a dynamic target in a foreground image comprises: Using image segmentation technology to identify and extract the dynamic target image in the foreground image; Determine an edge image of a dynamic target image, and then obtain image coordinates of the edge image; The area corresponding to the image coordinates is extracted to obtain a dynamic target image.
5. The method for dynamic target recognition of remote sensing images according to claim 4, characterized in that: Determining the edge image of the dynamic target image comprises: Performing Gaussian smoothing on the extracted dynamic target image to obtain a processed dynamic target image; Determine the grayscale difference of each pixel in the processed dynamic target image in the horizontal direction and the vertical direction respectively, and determine the edge strength of the pixel through the grayscale difference; According to the edge strength, each pixel is filtered by a non-maximum suppression algorithm; Double threshold detection and edge connection are performed on the filtering result to determine the edge image of the dynamic target image.
6. The method for dynamic target recognition of remote sensing images according to claim 4, characterized in that: Acquiring the image coordinates of the edge image includes: Dividing the edge image into a plurality of grids, each grid having a plurality of bounding boxes; Determining a confidence level of each bounding box, and adjusting the bounding box according to the confidence level to obtain adjusted bounding box coordinates; The bounding box coordinates are transformed to obtain image coordinates of the edge image.
7. The method for dynamic target recognition of remote sensing images according to claim 1, characterized in that: The training of the CDN to obtain a dynamic target recognition model includes: Acquire a training sample, wherein the training sample includes a sample image and its corresponding true result; The sample image is input into the CDN to obtain a recognition result, and the CDN is trained with the goal of minimizing the difference between the recognition result and the corresponding true result to obtain a dynamic target recognition model.
8. A dynamic target recognition device for remote sensing images, characterized in that: The device comprises: A construction module is used to introduce a dense prediction module DPM into CNN to construct a convolutional dense prediction network CDN, wherein the CDN includes an input layer, a feature extraction module, a dense prediction module and a classification layer; the CDN is trained to obtain a dynamic target recognition model; An acquisition module is used to acquire multiple continuously shot target remote sensing images of the same area; A preprocessing module is used to extract multiple candidate regions from the multiple target remote sensing images through a region generation network RPN; determine the variance mean of the candidate regions, and determine the candidate regions with dynamic targets as target regions according to the variance mean; determine the foreground image through the maximum inter-class variance of the target region, and extract the image of the dynamic target in the foreground image; The recognition module is used to input the image of the dynamic target into the dynamic target recognition model, preprocess the input image through the input layer to obtain a preprocessed image, the feature extraction module extracts multi-scale features of the preprocessed image through multiple convolutional layers, enhances the multi-scale features through the dense prediction module to obtain enhanced features, and classifies the enhanced features through the classification layer to obtain a recognition result.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.
Citation Information
Patent Citations
Dense target feature learning-based target detection method of optical remote-sensing image
CN108427912A
Method and system for constructing pine wood nematode disease recognition model in unmanned aerial vehicle remote sensing image for complex interference environment
CN118521887A