Remote sensing image detection method and device based on deep neural network
By segmenting high-resolution remote sensing images into sub-images and using deep neural network detection models, the problem of high-resolution remote sensing images is solved, and fast and efficient image detection and accurate land object category recognition are achieved.
Patent Information
- Application Number
- CN202510428192.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
When processing high-resolution remote sensing images, the calculation amount of the image detection model has increased significantly, resulting in a significant increase in the consumption of computing resources, which cannot meet the needs of rapid quality inspection.
By segmenting the remote sensing image to be detected into multiple sub-images, and inputting a pre-trained deep neural network image detection model for detection, the probability of each pixel in multiple geographic categories is calculated, and finally the regional proportion of the target geographic category in the remote sensing image is determined based on the proportion of the sub-image.
The calculation amount of the model when processing high-precision remote sensing images is reduced, the output rate is improved, and the accuracy of the detection results of the target object category area is improved through the calculation of the proportion of multiple sub-images.
Smart Images

Figure CN119942368A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image detection, and more specifically, to a remote sensing image detection method and device based on a deep neural network. Background Art
[0002] In the field of remote sensing image detection technology, it is necessary to further analyze the vegetation type, terrain information and other geographic information in the remote sensing image by detecting the content of various types of ground objects in the remote sensing image. Most of the time, two-thirds of the earth's area is covered by clouds, so the remote sensing images acquired by optical satellites are greatly affected by clouds. At the same time, as a remote sensing image, the cloud content and snow content of multispectral images are used as a criterion for determining the effective information content in the image. Remote sensing images with too high cloud content and snow content will not provide sufficient information for subsequent use, so it is necessary to detect the cloud content and snow content in remote sensing images.
[0003] Currently, due to the existence of multiple optical satellites working in orbit at the same time, the satellites upload a large amount of observation data every day and produce many high-resolution remote sensing images. However, manual inspection of high-resolution remote sensing images is time-consuming and inefficient, and cannot meet the requirements for rapid quality inspection of many remote sensing images.
[0004] As machine learning technology has achieved explosive growth in many fields, combining machine learning with remote sensing image detection tasks to develop automated, high-precision image detection models can effectively meet the needs of large-scale remote sensing image quality inspection. In related technologies, in order to obtain more accurate ground object category detection results, remote sensing images with high resolution are directly detected to achieve more accurate detection results. However, this also greatly increases the amount of calculation of the image detection model during the detection process, and the computing resources consumed also increase accordingly. Summary of the invention
[0005] In view of this, the present invention provides a remote sensing image detection method and device based on a deep neural network, which is used to solve the problem that the calculation amount of the image detection model increases significantly and the consumed computing resources also increase accordingly when detecting the object category of high-resolution remote sensing images.
[0006] A first aspect of the present invention provides a remote sensing image detection method based on a deep neural network, comprising: obtaining a remote sensing image to be detected, and dividing the remote sensing image into multiple sub-images; each sub-image includes at least one pixel; each sub-image is input into a pre-trained image detection model to obtain the probability that each pixel in at least one pixel of the sub-image belongs to a target object category among multiple preset object categories; the image detection model is constructed based on a deep neural network; according to the probability that each pixel in at least one pixel belongs to the target object category among multiple preset object categories, a first proportion of the area in each sub-image that meets the target object category in the sub-image is determined; according to multiple first proportions of multiple sub-images, a second proportion of the area with the target object category in the remote sensing image is determined.
[0007] According to an embodiment of the present invention, an image detection model includes 2N convolution modules, 2N first convolution layers, N transposed convolution modules, N second convolution layers and a transposed convolution layer; wherein the jth convolution module is connected to the j+1th convolution module through the jth first convolution layer; the 2Nth convolution module and the Nth transposed convolution module are connected through the 2Nth first convolution layer; the i-th transposed convolution module and the i-1th transposed convolution module are connected in series through the i-th second convolution layer; the 1st transposed convolution module and the transposed convolution layer are connected through the 1st second convolution layer; wherein, j=1,2,…,2N-1; i=2,3,…,N.
[0008] According to an embodiment of the present invention, the jth convolution module and the 2Nth convolution module are both used to perform a convolution operation on the input data and output a first feature map; the jth first convolution layer is used to perform channel compression on the first feature map input by the jth convolution module, and use the compressed first feature map as the input of the j+1th convolution module; the 2Nth first convolution layer is used to perform channel compression on the first feature map input by the 2Nth convolution module, and use the compressed first feature map as the input of the Nth transposed convolution module; the i-th transposed convolution module and the 1st transposed convolution module are used to perform channel compression on the input data in sequence. Perform transposed convolution processing and convolution processing to output a second feature map; the i-th second convolution layer is used to perform channel compression on the second feature map input by the i-th transposed convolution module, and use the compressed second feature map as the input of the i-1-th transposed convolution module; the first second convolution layer is used to perform channel compression on the second feature map input by the first transposed convolution module, and use the compressed second feature map as the input of the transposed convolution layer; the transposed convolution layer is used to process the input data and output the probability that each pixel in at least one pixel of the sub-image belongs to the target object category among the preset multiple object categories.
[0009] According to an embodiment of the present invention, in an image detection model, the 2i-1th first convolutional layer is connected to the ith second convolutional layer; the 1st first convolutional layer is connected to the 1st second convolutional layer; the ith second convolutional layer is used to merge the input data of the 2i-1th first convolutional layer and the ith transposed convolutional module, perform channel compression on the merged input data to obtain a compressed second feature map, and use the compressed second feature map as the input of the ith-1st transposed convolutional module; the 1st second convolutional layer is used to merge the input data of the 1st first convolutional layer and the 1st transposed convolutional module, perform channel compression on the merged input data to obtain a compressed second feature map, and use the compressed second feature map as the input of the transposed convolutional layer.
[0010] According to an embodiment of the present invention, the j-th first convolution layer is also used to sequentially perform convolution operations, normalization processing operations and activation function operations on the data input by the j-th convolution module to obtain a first feature map input to the j+1-th convolution module; the i-th second convolution layer is also used to sequentially perform convolution operations, normalization processing operations and activation function operations on the data input by the i-th transposed convolution module to obtain a second feature map input to the i-1-th transposed convolution module; the transposed convolution layer is also used to sequentially perform transposed convolution operations, normalization processing operations and function operation operations on the input data to output the probability that each pixel in at least one pixel of the sub-image belongs to the target object category among multiple preset object categories.
[0011] According to an embodiment of the present invention, any convolution module among the 2N convolution modules includes at least one convolution kernel of the first step size; any transposed convolution module among the N transposed convolution modules includes multiple convolution kernels of the first step size and at least one transposed convolution kernel of the second step size; wherein the first step size and the second step size are different.
[0012] According to an embodiment of the present invention, the i-th transposed convolution module is also used to: perform a transposed convolution operation of the second step length on the input data to obtain a third feature map; perform a convolution operation of the first step length on the third feature map to obtain a fourth feature map; perform a convolution operation of the first step length on the fourth feature map to obtain a fifth feature map; and use the fifth feature map as input to the i-1-th transposed convolution module.
[0013] According to an embodiment of the present invention, a second proportion of an area having a target object category in a remote sensing image is determined based on multiple first proportions of multiple sub-images, including: obtaining a pre-constructed database, the database storing interrelated third proportions and proportion intervals; determining in the database a target proportion interval that matches the first proportion of each sub-image, and a third proportion corresponding to the target proportion interval; and determining the second proportion of an area having a target object category in the remote sensing image based on multiple third proportions corresponding to multiple sub-images.
[0014] According to an embodiment of the present invention, an image detection model is trained in the following manner: obtaining historical remote sensing images containing multiple land object categories; performing land object category recognition on the historical remote sensing images to generate a mask map for each land object category in the historical remote sensing images; using the historical remote sensing images and the mask map for each land object category as sample data, preprocessing the sample data to obtain training data; and training a pre-constructed image detection model with the training data.
[0015] A second aspect of the present invention provides a remote sensing image detection device based on a deep neural network, comprising: an image acquisition module, used to acquire a remote sensing image to be detected, and divide the remote sensing image into a plurality of sub-images; each sub-image includes at least one pixel;
[0016] A model processing module, used to input each sub-image into a pre-trained image detection model to obtain the probability that each pixel in at least one pixel of the sub-image belongs to a target ground object category among a plurality of preset ground object categories; the image detection model is constructed based on a deep neural network;
[0017] The first proportion determination module is used to determine the first proportion of the area that meets the target object category in each sub-image according to the probability that each pixel in at least one pixel belongs to the target object category among multiple preset object categories; the second proportion determination module is used to determine the second proportion of the area with the target object category in the remote sensing image according to multiple first proportions of multiple sub-images.
[0018] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0019] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above method when executed.
[0020] A remote sensing image detection method and device based on a deep neural network provided by an embodiment of the present invention can achieve the following technical effects:
[0021] Through the method of this embodiment, after the remote sensing image to be detected is divided into multiple sub-images, the probability of each pixel in each sub-image belonging to a different ground object category is detected respectively; the amount of calculation of the model when processing high-precision remote sensing images is reduced, and the rate of output results is improved; at the same time, based on the first proportions of multiple sub-images, the second proportion of the area of the target ground object category in the remote sensing image is re-determined, further improving the accuracy of the regional detection results of the target ground object category. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0023] Figure 1 A flowchart of a remote sensing image detection method based on a deep neural network according to an embodiment of the present invention is schematically shown;
[0024] Figure 2 One of the structural diagrams of the image detection model according to an embodiment of the present invention is schematically shown;
[0025] Figure 3 The second structural diagram of the image detection model according to the embodiment of the present invention is schematically shown;
[0026] Figure 4A The structure diagram of the first convolutional layer according to an embodiment of the present invention is schematically shown;
[0027] Figure 4B Schematically shows a structural diagram of a second convolutional layer according to an embodiment of the present invention;
[0028] Figure 4C Schematically shows a structural diagram of a transposed convolutional layer according to an embodiment of the present invention;
[0029] Figures 5A to 5D The structural diagrams of four convolution modules according to the embodiments of the present invention are schematically shown respectively;
[0030] Figure 6 A schematic diagram of a structure of a transposed convolution module according to an embodiment of the present invention is shown;
[0031] Figure 7 A flowchart of a method for training an image detection module according to an embodiment of the present invention is schematically shown;
[0032] Figure 8 A block diagram of a remote sensing image detection device based on a deep neural network according to an embodiment of the present invention is schematically shown;
[0033] Fig. 9 A block diagram of an electronic device suitable for implementing the above description according to an embodiment of the present invention is schematically shown. DETAILED DESCRIPTION
[0034] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.
[0035] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0036] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0037] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0038] In the embodiment of the present invention, the remote sensing image to be detected may be a multispectral image with a variety of different types of objects, and the object categories may be mountains, fields, cities, wetlands, water bodies, oceans, snow, clouds, and forests. In the embodiment of the present invention, the remote sensing image may not provide sufficient information for subsequent use due to excessive cloud content and snow content. Therefore, by identifying the proportion of areas of the two object categories of snow and clouds in the remote sensing image, it is possible to provide support for users to understand the effective information content in the remote sensing image.
[0039] It should be noted that the embodiments of the present invention only exemplify the application scenario of detecting the types of objects in remote sensing images, and are not intended to limit the present invention.
[0040] Figure 1 The flowchart of a remote sensing image detection method based on a deep neural network according to an embodiment of the present invention is schematically shown.
[0041] like Figure 1As shown, the remote sensing image detection method based on deep neural network of this embodiment includes operation S101 to operation S104.
[0042] In operation S101, a remote sensing image to be detected is acquired, and the remote sensing image is divided into a plurality of sub-images; each sub-image includes at least one pixel.
[0043] Exemplarily, the remote sensing image to be detected may be a remote sensing image with a resolution of 8000×8000, and the remote sensing image with a resolution of 8000×8000 may be segmented into a plurality of sub-images with a resolution of 400×400.
[0044] In operation S102, each sub-image is input into a pre-trained image detection model to obtain the probability that each pixel in at least one pixel of the sub-image belongs to a target object category among a plurality of preset object categories; the image detection model is constructed based on a deep neural network.
[0045] In an embodiment of the present invention, each sub-image is input into a pre-trained image detection model. The probability of each individual pixel in each sub-image belonging to the target object category among the preset multiple object categories is obtained. Exemplarily, the preset multiple object categories include ocean, snow, cloud and forest. Correspondingly, for any pixel in the sub-image, the probability that the pixel belongs to the ocean category, snow category, cloud category and forest category can be obtained respectively; the image detection model in the embodiment of the present invention is constructed based on a deep neural network.
[0046] In operation S103, a first proportion of an area in each sub-image that meets the target ground object category is determined according to a probability that each pixel in the at least one pixel belongs to the target ground object category among a plurality of preset ground object categories.
[0047] Exemplarily, two ground object categories, snow and cloud, are used as target ground object categories, and the first proportion of the area that meets the target ground object category in each sub-image is determined. Exemplarily, for a pixel in a sub-image, the probability that the pixel belongs to the snow category and the cloud category are 10% and 15% respectively; the probability of each pixel in the sub-image belonging to the snow category and the cloud category is traversed. The sum of the probabilities that each pixel in the sub-image belongs to the snow category is obtained, and the sum of the probabilities is divided by the number of pixels in the sub-image to obtain the proportion of the snow area in the sub-image; similarly, the proportion of the cloud area in the sub-image is obtained; the sum of the proportion of the snow area and the proportion of the cloud area in the sub-image is taken as the first proportion.
[0048] In operation S104, a second proportion of the area having the target object category in the remote sensing image is determined according to the plurality of first proportions of the plurality of sub-images.
[0049] Through the method of this embodiment, after the remote sensing image to be detected is divided into multiple sub-images, the probability of each pixel in each sub-image belonging to a different ground object category is detected respectively; the amount of calculation of the model when processing high-precision remote sensing images is reduced, and the rate of output results is improved; at the same time, based on the first proportions of multiple sub-images, the second proportion of the area of the target ground object category in the remote sensing image is re-determined, further improving the accuracy of the regional detection results of the target ground object category.
[0050] Furthermore, the above operation S104 determines the second proportion of the area with the target object category in the remote sensing image based on the multiple first proportions of the multiple sub-images, and can also include: obtaining a pre-constructed database, the database storing interrelated third proportions and proportion intervals; determining in the database the target proportion interval that matches the first proportion of each sub-image, and the third proportion corresponding to the target proportion interval; and determining the second proportion of the area with the target object category in the remote sensing image based on the multiple third proportions corresponding to the multiple sub-images.
[0051] In order to make the cloud content calculation result of the image detection model for the remote sensing image closer to the subjective identification result of the quality inspector, a database can be pre-built, in which the database obtains the third proportion of the target object category in the historical remote sensing image obtained by the quality inspector through subjective judgment. It should be noted that due to the difference between the human eye recognition process and the machine learning process, for the same remote sensing image, when the human eye identifies the proportion of clouds in the remote sensing image, for a cloud area with sparse clouds, the entire cloud area will be considered to be all clouds; while the machine recognition result can be accurate to the pixel level, for a cloud area with sparse clouds, the entire cloud area will be considered to be a cloud area to a certain extent (such as 60%, depending on the sparseness of the clouds). Therefore, in order to make the multiple first proportions of multiple sub-images conform to the manual judgment results, it is necessary to construct interrelated third proportions and proportion intervals. Among them, all intervals from 0 to 1 are divided into multiple proportion intervals; each proportion interval corresponds to a third proportion; and each proportion interval corresponds to a first proportion.
[0052] Further, the target proportion interval that matches the first proportion of each sub-image and the third proportion corresponding to the target proportion interval are determined. According to the multiple third proportions corresponding to the multiple sub-images, the second proportion of the area with the target object category in the remote sensing image is determined by obtaining the average value of the multiple third proportions.
[0053] Figure 2 One of the structural diagrams of the image detection model according to an embodiment of the present invention is schematically shown.
[0054] like Figure 2As shown in FIG. 1 , the image detection model includes 2N convolution modules, 2N first convolution layers, N transposed convolution modules, N second convolution layers and transposed convolution layers. Among them, the jth convolution module is connected to the j+1th convolution module through the jth first convolution layer; the 2Nth convolution module is connected to the Nth transposed convolution module through the 2Nth first convolution layer; the ith transposed convolution module is connected to the i-1th transposed convolution module through the ith second convolution layer; the 1st transposed convolution module is connected to the transposed convolution layer through the 1st second convolution layer; among them, j=1,2,…,2N-1; i=2,3,…,N.
[0055] Exemplarily, the image detection model includes 8 convolution modules, eight first convolution layers, four transposed convolution modules, four second convolution layers and one transposed convolution layer. Figure 2 As shown on the left side of , the convolution modules from top to bottom are the first convolution module to the eighth convolution module; the convolution layers from top to bottom are the first convolution layer to the eighth convolution layer. Figure 2 As shown on the right side of , the transposed convolution modules from top to bottom are the first transposed convolution module to the fourth transposed convolution module; the convolution layers from top to bottom are the first second convolution layer to the fourth second convolution layer; among them, the eighth convolution module and the fourth transposed convolution module are connected through the eighth first convolution layer.
[0056] For example, eight levels of convolution modules and four levels of transposed convolution modules may be provided in the image detection model. It should be noted that the levels of the convolution modules and the transposed convolution modules in the present application may be specifically configured according to the image resolution size of the remote sensing image to be detected. The levels of the convolution modules and the transposed convolution modules may increase as the resolution of the remote sensing image increases. A higher level of image detection model can make the output results of the model more accurate, but it will also increase the amount of computation required by the image detection module.
[0057] exist Figure 2 In the image detection model shown, it is aimed at detecting remote sensing image types such as cloud content and snow content in remote sensing images; by configuring eight levels of convolution modules and four levels of transposed convolution modules, as well as the first convolution layer and the second convolution layer of the corresponding number of levels; it is able to ensure the accuracy of the image detection model output while avoiding excessive model calculation caused by too many levels.
[0058] In the following, in the embodiments of the present invention, Figure 2 The functions of the convolution modules, transposed convolution modules, first convolution layers, second convolution layers and transposed convolution layers shown are explained.
[0059] In the embodiment of the present invention, the j-th convolution module and the 2N-th convolution module are both used to perform a convolution operation on the input data and output a first feature map.
[0060] The j-th first convolution layer is used to perform channel compression on the first feature map input by the j-th convolution module, and use the compressed first feature map as the input of the j+1-th convolution module; the 2N-th first convolution layer is used to perform channel compression on the first feature map input by the 2N-th convolution module, and use the compressed first feature map as the input of the N-th transposed convolution module.
[0061] The i-th transposed convolution module and the first transposed convolution module are used to perform transposed convolution processing and convolution processing on the input data in sequence, and output a second feature map.
[0062] The i-th second convolution layer is used to perform channel compression on the second feature map input by the i-th transposed convolution module, and use the compressed second feature map as the input of the i-1-th transposed convolution module; the first second convolution layer is used to perform channel compression on the second feature map input by the first transposed convolution module, and use the compressed second feature map as the input of the transposed convolution layer.
[0063] The transposed convolution layer is used to process the input data and output the probability that each pixel in at least one pixel of the sub-image belongs to the target ground object category among multiple preset ground object categories.
[0064] Figure 3 The second structural diagram of the image detection model according to an embodiment of the present invention is schematically shown.
[0065] like Figure 3 As shown, in the image detection model, the 2i-1th first convolutional layer is connected to the i-th second convolutional layer; the 1st first convolutional layer is connected to the 1st second convolutional layer.
[0066] The cascade from top to bottom is: the 1st first convolutional layer is connected to the 1st second convolutional layer; the 3rd first convolutional layer is connected to the 2nd second convolutional layer; the 5th first convolutional layer is connected to the 3rd second convolutional layer; the 7th first convolutional layer is connected to the 4th second convolutional layer.
[0067] In the following, in the embodiments of the present invention, Figure 3 The functions of the four second convolutional layers after cascading are shown.
[0068] In the second convolution layer provided in the embodiment of the present invention, the i-th second convolution layer is used to merge the input data of the 2i-1-th first convolution layer and the i-th transposed convolution module, perform channel compression on the merged input data to obtain a compressed second feature map, and use the compressed second feature map as the input of the i-1-th transposed convolution module. The first second convolution layer is used to merge the input data of the first first convolution layer and the first transposed convolution module, perform channel compression on the merged input data to obtain a compressed second feature map, and use the compressed second feature map as the input of the transposed convolution layer.
[0069] In the embodiment of the present invention, Figure 3 As shown in the figure, the convolution module and the first convolution layer on the left constitute the encoder. The transposed convolution module and the second convolution layer on the right constitute the decoder. The encoder is used to extract spatial feature data from the sub-image, reduce the spatial dimension, and increase the number of channels. The decoder performs projection mapping on the extracted data to create a mask map containing the cloud and snow detection results at the original resolution of the sub-image.
[0070] Further, such as Figure 3 As shown, the feature maps in the encoding and decoding processes are connected by using a cascade method to achieve fusion between low-level and high-level features. Among them, the purpose of the first convolution layer and the second convolution layer is to extract translation-invariant features in the image. In each convolution layer, the convolution kernel is convolved with the input image to generate the corresponding feature map. Compared with the method of using a fully connected layer, the convolution kernels in the first convolution layer and the second convolution layer in the embodiment of the present invention share the same weights on the entire sub-image, which effectively reduces the weight parameters of the image detection model and learns the translation-invariant features on the sub-image.
[0071] Figure 4A A structural diagram of the first convolutional layer provided in an embodiment of the present invention.
[0072] like Figure 4A As shown, the j-th first convolutional layer is also used to perform convolution operations, normalization operations, and activation function operations on the data input by the j-th convolutional module in sequence to obtain the first feature map input to the j+1-th convolutional module.
[0073] Figure 4B This is a structural diagram of the second convolutional layer provided in an embodiment of the present invention.
[0074] like Figure 4B As shown, the i-th second convolutional layer is also used to perform convolution or transposed convolution operations, normalization operations, and activation function operations on the data input by the i-th transposed convolution module in sequence to obtain the second feature map input to the i-1-th transposed convolution module.
[0075] In the embodiment of the present invention, the first convolutional layer and the second convolutional layer may be configured as the same structure or as other different structures. Figure 4A and Figure 4B The structures of the first convolutional layer and the second convolutional layer in the present invention are only given as examples, and are not intended to limit the structures of the two convolutional layers.
[0076] Figure 4C A structural diagram of a transposed convolutional layer provided in an embodiment of the present invention.
[0077] like Figure 4C As shown, the transposed convolution layer is also used to perform a transposed convolution operation, a normalization operation, and a function operation on the input data in sequence, and output the probability that each pixel in at least one pixel of the sub-image belongs to the target object category among multiple preset object categories.
[0078] Exemplarily, the activation function in the embodiment of the present invention can be configured as a Relu function.
[0079] The structures of 2N convolution modules and N transposed convolution modules in an embodiment of the present invention are described below.
[0080] In an embodiment of the present invention, any convolution module among the 2N convolution modules includes at least one convolution kernel of the first step size; any transposed convolution module among the N transposed convolution modules includes multiple convolution kernels of the first step size and at least one transposed convolution kernel of the second step size; wherein the first step size and the second step size are different.
[0081] The structure of each convolution module in the embodiment of the present invention is described below.
[0082] Exemplarily, the eight convolution modules have four structural types: A, B, C, and D; among them, the types of the first convolution module, the third convolution module, the fifth convolution module, and the seventh convolution module are convolution module A, convolution module C, convolution module C, and convolution module D, respectively, and the types of the second convolution module, the fourth convolution module, the sixth convolution module, and the eighth convolution module are all convolution module B, and the structures of the four transposed convolution modules are the same.
[0083] like Figure 5A As shown, it is a structural diagram of a convolution module of type A provided in an embodiment of the present invention.
[0084] like Figure 5B As shown, it is a structural diagram of a convolution module of type B provided in an embodiment of the present invention.
[0085] like Figure 5C As shown, it is a structural diagram of a convolution module of type C provided in an embodiment of the present invention.
[0086] like Figure 5DAs shown, it is a structural diagram of a convolution module of type D provided in an embodiment of the present invention.
[0087] exist Figure 5A-5D In , the stride of the convolution kernels in the first convolution, the third convolution module, the fifth convolution module, and the seventh convolution module are all 1.
[0088] exist FIG. 5A to FIG. 5D In the second convolution, the fourth convolution module, the sixth convolution module, and the eighth convolution module, in addition to the convolution kernel with a stride of 1, two convolution kernels with a stride of 2 are also included. By using a convolution with a stride of 2 to downsample the sub-image, the size of the input image data is reduced and the input dimension is increased to obtain a more compact feature map. The second convolution, the fourth convolution module, the sixth convolution module, and the eighth convolution module correspond to four downsamplings, corresponding to different receptive fields, namely the pixel level, the small window around each pixel, the larger window around the pixel, and the maximum field of view of the model.
[0089] Figure 6 The structure diagram of the transposed convolution module according to an embodiment of the present invention is schematically shown.
[0090] like Figure 6 As shown, in the transposed convolution module provided in an embodiment of the present invention, the i-th transposed convolution module is used to perform a transposed convolution operation of the second step length on the input data to obtain a third feature map; perform a convolution operation of the first step length on the third feature map to obtain a fourth feature map; perform a convolution operation of the first step length on the fourth feature map to obtain a fifth feature map; and use the fifth feature map as the input of the i-1-th transposed convolution module.
[0091] The first transposed convolution module is used to perform a transposed convolution operation of the second step length on the input data to obtain a third feature map; perform a convolution operation of the first step length on the third feature map to obtain a fourth feature map; perform a convolution operation of the first step length on the fourth feature map to obtain a fifth feature map; and use the fifth feature map as the input of the transposed convolution module layer.
[0092] The following is an explanation of the process of the image detection module processing the sub-image in the embodiment of the present invention.
[0093] First, a 400×400×m sub-image is input into the image detection model and passed through the convolution module A, where m is the number of channels.
[0094] In convolution module A, the input image is convolved with a step size of 1 by 3×3 to obtain feature map a1; convolution is performed in the first convolution layer, the result is normalized, and then the Relu activation function is used; the image is convolved with a step size of 1 twice by 3×3 to obtain feature map a2; the image is convolved with a step size of 1 three times by 3×3 to obtain feature map a3; the input image is merged with the channels of a1, a2, and a3 to obtain A1; A1 is convolved with a step size of 1 by 1×1, and the channel is compressed to obtain a feature map of 400×400×m. Figure 1 ; The feature Figure 1 Input into the convolution module B.
[0095] In the convolution module B, a 3×3 convolution with a step size of 2 is performed to obtain the feature map b1_1;
[0096] Features Figure 1 Perform a 3×3 convolution with a step size of 1, and then perform a 3×3 convolution with a step size of 2 to obtain the feature map b1_2; merge the feature maps b1_1 and b1_2 to obtain B1; perform a 1×1 convolution with a step size of 1 on B1 and perform channel compression to obtain a 200×200×m feature map. Figure 2 ; The feature Figure 2 Input into the convolution module C.
[0097] In the convolution module C, a 3×3 convolution with a step size of 1 is performed to obtain the feature map c2_1; Figure 2 Perform two 3×3 convolutions with a step size of 1 to obtain feature map c2_2; Figure 2 Merge channels with c2_1 and c2_2 to get C2; perform 1×1 convolution with a step size of 1 on C2 and perform channel compression to get a feature of 200×200×m Figure 3 ; The feature Figure 3 Input into the convolution module B.
[0098] In the convolution module B, a 3×3 convolution with a step size of 2 is performed to obtain the feature map b3_1; Figure 3 Perform a 3×3 convolution with a step size of 1, and then perform a 3×3 convolution with a step size of 2 to obtain feature map b3_2; merge the feature maps b3_1 and b3_2 to obtain B3; perform a 1×1 convolution with a step size of 1 on B3, and perform channel compression to obtain a feature map 4 of 100×100×m; input feature map 4 into the convolution module C.
[0099] In the convolution module C, a 3×3 convolution with a step size of 1 is performed to obtain feature map c4_1; two 3×3 convolutions with a step size of 1 are performed on feature map 4 to obtain feature map c4_2; feature map 4 is channel-merged with c4_1 and c4_2 to obtain C4; a 1×1 convolution with a step size of 1 is performed on C4 to perform channel compression to obtain a feature map 5 of 100×100×m; feature map 5 is input into the convolution module B.
[0100] In convolution module B, a 3×3 convolution with a step size of 2 is performed to obtain feature map b5_1; a 3×3 convolution with a step size of 1 is performed on feature map 5, and then a 3×3 convolution with a step size of 2 is performed to obtain feature map b5_2; feature maps b5_1 and b5_2 are merged to obtain B5; a 1×1 convolution with a step size of 1 is performed on B5, and channel compression is performed to obtain a 50×50×m feature map. Figure 6 ; The feature Figure 6 Input into the convolution module D.
[0101] In the convolution module D, a 1×3 convolution with a step size of 1 is performed, and then a 3×1 convolution with a step size of 1 is performed to obtain the feature map d6_1; Figure 6 Perform a 1×5 convolution with a step size of 1, and then perform a 5×1 convolution with a step size of 1 to obtain the feature map d6_2; Figure 6 Perform a 1×7 convolution with a step size of 1, and then perform a 7×1 convolution with a step size of 1 to obtain the feature map d6_3;
[0102] The characteristics Figure 6 Merge channels with d6_1, d6_2, and d6_3 to obtain D6; perform 1×1 convolution with a step size of 1 on D6 and perform channel compression to obtain a 50×50×m feature. Figure 7 ; The feature Figure 7 Input into the convolution module B.
[0103] In convolution module B, a 3×3 convolution with a step size of 2 is performed to obtain feature map b7_1; a 3×3 convolution with a step size of 1 is performed on feature map 5, and then a 3×3 convolution with a step size of 2 is performed to obtain feature map b7_2; feature maps b7_1 and b7_2 are channel-merged to obtain B7; a 1×1 convolution with a step size of 1 is performed on B7 to perform channel compression to obtain a 25×25×m feature map. Figure 8 ; The feature Figure 8 Input into the fourth transposed convolution module.
[0104] In the fourth transposed convolution module, a 3×3 transposed convolution with a step size of 2 is performed to obtain the feature map u8_1; a 3×3 convolution with a step size of 1 is performed on the feature map u8_1 to obtain the feature map u8_2; a 3×3 convolution with a step size of 1 is performed on the feature map u8_2 to obtain the feature map u8_3; the feature map u8_3 is combined with the feature map Figure 7 Cascade to obtain feature map U8; perform 1×1 convolution with a step size of 1 on feature map U8 to obtain a 50×50×m feature Fig. 9 ; The feature Fig. 9 Input to the transposed convolution module.
[0105] In the third transposed convolution module, a 3×3 transposed convolution with a step size of 2 is performed in this module to obtain feature map u9_1; a transposed convolution is performed in the third transposed convolution layer, and the result is normalized, and then the Relu activation function is used; a 3×3 convolution with a step size of 1 is performed on the feature map u9_1 to obtain the feature map u9_2; a 3×3 convolution with a step size of 1 is performed on the feature map u9_2 to obtain the feature map u9_3; the feature map u9_3 is cascaded with the feature map 5 to obtain the feature map U9; a 1×1 convolution with a step size of 1 is performed on the feature map U9 to obtain the feature map 10 of 100×100×m; the feature map 10 is input into the transposed convolution module.
[0106] In the second transposed convolution module, a 3×3 transposed convolution with a step size of 2 is performed to obtain the feature map u10_1; a 3×3 convolution with a step size of 1 is performed on the feature map u10_1 to obtain the feature map u10_2; a 3×3 convolution with a step size of 1 is performed on the feature map u10_2 to obtain the feature map u10_3; the feature map u10_3 is combined with the feature map Figure 3 Cascade is performed to obtain feature map U10; a 1×1 convolution with a step size of 1 is performed on feature map U10 to obtain a 200×200×m feature map 11; feature map 11 is input into the transposed convolution module.
[0107] In the first transposed convolution module, a 3×3 transposed convolution with a step size of 2 is performed to obtain the feature map u11_1; a 3×3 convolution with a step size of 1 is performed on the feature map u11_1 to obtain the feature map u11_2; a 3×3 convolution with a step size of 1 is performed on the feature map u11_2 to obtain the feature map u11_3; the feature map u11_3 is combined with the feature map Figure 1 Cascade is performed to obtain feature map U11; a 1×1 convolution with a step size of 1 is performed on feature map U11 to obtain a 400×400×m feature map 12; feature map 12 is input into the transposed convolution layer.
[0108] Then, in the transposed convolution layer, a 1×1 transposed convolution with a step size of 1 is performed to obtain the feature value.
[0109] Finally, Softmax is used as an activation function, and the feature value is input into the Softmax function to calculate the first proportion of the area in each sub-image that meets the target object category in the sub-image, and then operations S103 to S104 are performed.
[0110] Figure 7 The flowchart of the training method of the image detection module according to the embodiment of the present invention is schematically shown.
[0111] like Figure 7 As shown, the training method of the image detection module of this embodiment includes operations S701 to S703.
[0112] In the embodiment of the present invention, the image detection model is trained in the following manner:
[0113] Operation S701: Acquire historical remote sensing images containing the multiple land object categories.
[0114] Operation S702: performing object category recognition on the historical remote sensing image to generate a mask image for each object category in the historical remote sensing image.
[0115] Operation S703: using the historical remote sensing image and the mask image of each land object category as sample data, preprocessing the sample data to obtain training data; and training a pre-built image detection model with the training data.
[0116] Based on the remote sensing image detection method based on deep neural network, the present invention also provides a remote sensing image detection device based on deep neural network. Figure 8 The device is described in detail.
[0117] Figure 8 A block diagram of a remote sensing image detection device based on a deep neural network according to an embodiment of the present invention is schematically shown.
[0118] like Figure 8 As shown, the remote sensing image detection device 800 based on a deep neural network of this embodiment includes an image acquisition module 810, a model processing module 820, a first proportion determination module 830 and a second proportion determination module 840.
[0119] The image acquisition module 810 is used to acquire the remote sensing image to be detected and divide the remote sensing image into a plurality of sub-images; each sub-image includes at least one pixel.
[0120] The model processing module 820 is used to input each sub-image into a pre-trained image detection model to obtain the probability that each pixel in at least one pixel of the sub-image belongs to the target object category among multiple preset object categories; the image detection model is constructed based on a deep neural network.
[0121] The first proportion determination module 830 is used to determine a first proportion of the area in each sub-image that meets the target object category in the sub-image according to the probability that each pixel in at least one pixel belongs to the target object category among multiple preset object categories.
[0122] The second proportion determination module 840 is used to determine a second proportion of the area having the target object category in the remote sensing image according to the multiple first proportions of the multiple sub-images.
[0123] According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits, or at least part of the functions of any one of them can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be at least partially implemented as hardware circuits, such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems on chips, systems on substrates, systems on packages, application specific integrated circuits (ASICs), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or by any one of the three implementation methods of software, hardware, and firmware, or by a proper combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, submodules, units, and subunits can be at least partially implemented as computer program modules, and when the computer program modules are run, the corresponding functions can be executed.
[0124] For example, any multiple of the image acquisition module 810, the model processing module 820, the first proportion determination module 830, and the second proportion determination module 840 can be combined in one module / unit / sub-unit for implementation, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present invention, at least one of the image acquisition module 810, the model processing module 820, the first proportion determination module 830, and the second proportion determination module 840 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the image acquisition module 810, the model processing module 820, the first proportion determination module 830 and the second proportion determination module 840 may be at least partially implemented as a computer program module, and when the computer program module is executed, a corresponding function may be performed.
[0125] It should be noted that the remote sensing image detection device based on a deep neural network in the embodiment of the present invention corresponds to the remote sensing image detection method based on a deep neural network in the embodiment of the present invention, and their specific implementation details are also the same, which will not be repeated here.
[0126] Fig. 9 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present invention is schematically shown. Fig. 9 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0127] like Fig. 9As shown, the electronic device 900 according to an embodiment of the present invention includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include an onboard memory configured for cache purposes. The processor 901 may include a single processing unit or multiple processing units configured to perform different actions of the method flow according to an embodiment of the present invention.
[0128] In RAM 903, various programs and data required for the operation of electronic device 900 are stored. Processor 901, ROM 902 and RAM 903 are connected to each other via bus 904. Processor 901 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 902 and / or RAM 903. It should be noted that the program can also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.
[0129] According to an embodiment of the present invention, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the input / output (I / O) interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage portion 908 as needed.
[0130] According to an embodiment of the present invention, the method flow according to an embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code configured to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiment of the present invention are executed. According to an embodiment of the present invention, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.
[0131] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.
[0132] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include, but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.
[0133] For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 902 and / or the RAM 903 described above and / or one or more memories other than the ROM 902 and the RAM 903 .
[0134] An embodiment of the present invention also includes a computer program product, which includes a computer program, which contains program code configured to execute the method provided by the embodiment of the present invention. When the computer program product runs on an electronic device, the program code is configured to enable the electronic device to implement the remote sensing image detection method based on a deep neural network provided by the embodiment of the present invention.
[0135] When the computer program is executed by the processor 901, the above functions defined in the system / device of the embodiment of the present invention are executed. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0136] In one embodiment, the computer program may be based on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0137] According to an embodiment of the present invention, the program code configured to execute the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions configured to implement the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features recorded in the various embodiments of the present invention can be combined and / or combined in various ways, even if such a combination or combination is not explicitly recorded in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recorded in the various embodiments of the present invention can be combined and / or combined in various ways. All these combinations and / or combinations fall within the scope of the present invention.
[0139] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A remote sensing image detection method based on a deep neural network, comprising: Acquire a remote sensing image to be detected, and divide the remote sensing image into a plurality of sub-images; Each of the sub-images comprises at least one pixel; Inputting each of the sub-images into a pre-trained image detection model to obtain the probability that each pixel in at least one pixel of the sub-image belongs to a target ground object category among a plurality of preset ground object categories; the image detection model is constructed based on a deep neural network; Determine, according to the probability that each pixel in the at least one pixel belongs to the target ground object category among the preset multiple ground object categories, a first proportion of the area in each of the sub-images that meets the target ground object category in the sub-image; According to the multiple first proportions of the multiple sub-images, a second proportion of the area having the target object category in the remote sensing image is determined.
2. According to the method of claim 1, the image detection model comprises 2N convolution modules, 2N first convolution layers, N transposed convolution modules, N second convolution layers and transposed convolution layers; in, The jth convolution module is connected to the j+1th convolution module through the jth first convolution layer; the 2Nth convolution module is connected to the Nth transposed convolution module through the 2Nth first convolution layer; the ith transposed convolution module is connected to the i-1th transposed convolution module through the ith second convolution layer; the 1st transposed convolution module is connected to the transposed convolution layer through the 1st second convolution layer; wherein, j=1,2,…,2N-1; i=2,3,…,N.
3. The method according to claim 2, wherein the j-th convolution module and the 2N-th convolution module are both used to perform a convolution operation on the input data and output a first feature map; The j-th first convolution layer is used to perform channel compression on the first feature map input by the j-th convolution module, and use the compressed first feature map as the input of the j+1-th convolution module; the 2N-th first convolution layer is used to perform channel compression on the first feature map input by the 2N-th convolution module, and use the compressed first feature map as the input of the N-th transposed convolution module; The i-th transposed convolution module and the first transposed convolution module are used to perform transposed convolution processing and convolution processing on the input data in sequence, and output a second feature map; The i-th second convolution layer is used to perform channel compression on the second feature map input by the i-th transposed convolution module, and use the compressed second feature map as the input of the i-1-th transposed convolution module; the first second convolution layer is used to perform channel compression on the second feature map input by the first transposed convolution module, and use the compressed second feature map as the input of the transposed convolution layer; The transposed convolution layer is used to process the input data and output the probability that each pixel in at least one pixel of the sub-image belongs to the target ground object category among multiple preset ground object categories.
4. The method according to claim 2, wherein in the image detection model, the 2i-1th first convolutional layer is connected to the i-th second convolutional layer; the 1st first convolutional layer is connected to the 1st second convolutional layer; The i-th second convolution layer is used to merge the input data of the 2i-1th first convolution layer and the i-th transposed convolution module, perform channel compression on the merged input data to obtain a compressed second feature map, and use the compressed second feature map as the input of the i-1th transposed convolution module; The first second convolutional layer is used to merge the input data of the first first convolutional layer and the first transposed convolutional module, perform channel compression on the merged input data to obtain a compressed second feature map, and use the compressed second feature map as the input of the transposed convolutional layer.
5. According to the method of claim 3, the j-th first convolutional layer is further used to sequentially perform a convolution operation, a normalization operation, and an activation function operation on the data input by the j-th convolutional module to obtain a first feature map input to the j+1-th convolutional module; The i-th second convolution layer is further used to sequentially perform convolution operations, normalization operations, and activation function operations on the data input by the i-th transposed convolution module to obtain a second feature map input to the i-1-th transposed convolution module; The transposed convolution layer is also used to perform a transposed convolution operation, a normalization processing operation and a function operation on the input data in sequence, and output the probability that each pixel in at least one pixel of the sub-image belongs to the target object category among multiple preset object categories.
6. The method according to claim 3, wherein any of the 2N convolution modules comprises at least one convolution kernel of the first step length; Any of the N transposed convolution modules includes a plurality of convolution kernels of a first step length and at least one transposed convolution kernel of a second step length; in, The first step length and the second step length are different.
7. The method according to claim 6, wherein the i-th transposed convolution module is further used for: Performing the transposed convolution operation of the second step length on the input data to obtain a third feature map; Performing the first step-length convolution operation on the third feature map to obtain a fourth feature map; Performing the first step-length convolution operation on the fourth feature map to obtain a fifth feature map; The fifth feature map is used as input of the i-1th transposed convolution module.
8. The method according to claim 1, wherein determining a second proportion of an area having the target object category in the remote sensing image according to the plurality of first proportions of the plurality of sub-images comprises: Acquire a pre-built database, wherein the database stores third proportions and proportion intervals that are associated with each other; Determine in the database a target proportion interval that matches the first proportion of each sub-image, and a third proportion corresponding to the target proportion interval; According to the multiple third proportions corresponding to the multiple sub-images, a second proportion of the area having the target object category in the remote sensing image is determined.
9. According to the method of claim 1, the image detection model is trained in the following manner: Acquire historical remote sensing images containing the multiple land feature categories; Identify the ground object categories of the historical remote sensing images and generate a mask map of each ground object category in the historical remote sensing images; Taking the historical remote sensing images and the mask images of each land object category as sample data, preprocessing the sample data to obtain training data; The pre-built image detection model is trained with the training data.
10. A remote sensing image detection device based on a deep neural network, comprising: An image acquisition module is used to acquire a remote sensing image to be detected and divide the remote sensing image into multiple sub-images; Each of the sub-images comprises at least one pixel; A model processing module, used for inputting each of the sub-images into a pre-trained image detection model to obtain the probability that each pixel in at least one pixel of the sub-image belongs to a target ground object category among a plurality of preset ground object categories; the image detection model is constructed based on a deep neural network; A first proportion determination module, configured to determine a first proportion of an area in each of the sub-images that meets the target ground object category in the sub-image according to a probability that each pixel in the at least one pixel belongs to the target ground object category among a plurality of preset ground object categories; The second proportion determination module is used to determine a second proportion of the area having the target object category in the remote sensing image according to the multiple first proportions of the multiple sub-images.
Citation Information
Patent Citations
Remote-sensing image detection method and device and computer equipment
CN108229261A
Remote sensing image identification method and apparatus, storage medium and electronic device
CN108304775A
Optical remote sensing image segmentation method based on multi-scale lightweight hole convolution
CN111080652A
Remote sensing image ground object classification method and system
CN111428781A
Full convolutional deep neural network remote sensing image classification method based on sparse point marking
CN115641513A