Monocular remote sensing image height data estimation method and device, equipment and storage medium
By using the hybrid pooling patch embedding module and the Transformer module in the monocular remote sensing image height estimation method, the problem of insufficient local information extraction capability in large-scale feature processing is solved, and more accurate and clear height estimation results are achieved.
Patent Information
- Application Number
- CN202510094675.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing monocular remote sensing image height estimation method lacks the local information extraction capability when processing large-scale features, which can easily lead to height estimation errors. The Transformer-based method lacks the ability to extract local details and blurs the details.
The image is initially extracted by the hybrid pooling patch embedding module, combined with the local information enhancement module (LIE module) in the Transformer module, the ability to pay attention to local information is enhanced, and the accuracy of high estimation is improved through image fusion.
It improves the local information extraction ability, reduces misjudgment of height estimation, can more accurately judge the height of the target, and at the same time enhances the detail information and texture clarity of the height map.
Smart Images

Figure CN120107330A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image recognition technology, and in particular, relates to a method, device, equipment and storage medium for estimating height data of a monocular remote sensing image. Background Art
[0002] Traditional methods of obtaining height often use multi-view methods and lidar-based methods, which are expensive and difficult to promote. Therefore, methods for estimating height based on monocular remote sensing images have been developed.
[0003] One method is the height estimation method of monocular remote sensing images based on convolution. It has good local information extraction ability, but when facing large-scale features, such as large-area building roofs connected in a block, this type of method cannot obtain enough global information due to the inherent nature of convolution, and it is easy to confuse it with other similar features, resulting in height estimation errors.
[0004] Another method is the Transformer-based monocular remote sensing image height estimation method. Compared with the convolution method, the Transformer network-based method can obtain more global information through the self-attention mechanism. When processing remote sensing images, by obtaining more global information, the height of pixels in the image can be judged more accurately.
[0005] However, the Transformer network-based method is not good at extracting local details. Although the height map extracted by the network is relatively accurate overall, it is relatively fuzzy and lacks details. It is easy to ignore some small objects, such as street lights, winter trees, tower buildings, and cars. In addition, the Transformer network requires greater computing power support. Summary of the invention
[0006] The embodiments of the present application provide a method, device, equipment and storage medium for estimating height data of a monocular remote sensing image, so as to improve the local information extraction capability and accurately judge the height of the target at the same time.
[0007] This application is implemented through the following technical solutions:
[0008] In a first aspect, an embodiment of the present application provides a method for estimating height data of a monocular remote sensing image, comprising:
[0009] Acquire a monocular remote sensing image of the target object.
[0010] The monocular remote sensing image is input into the first convolutional layer of the hybrid pooling patch embedding module to obtain the first feature map, and is input into the maximum pooling layer, the average pooling layer and the first depth-separable convolutional layer of the hybrid pooling patch embedding module to obtain the second feature map.
[0011] The second feature map is input into the Transformer module to obtain the global feature map.
[0012] The global feature map, the second feature map and the first feature map are input into the decoder for image fusion in sequence to obtain a first output image.
[0013] Based on the first output image, an estimated value of the height of the target object is obtained.
[0014] In combination with the first aspect, in some possible implementations, a maximum pooling layer, an average pooling layer, and a first depth-separable convolutional layer of a hybrid pooling patch embedding module are input to obtain a second feature map, including:
[0015] The first feature map is input into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling map and the average pooling map respectively.
[0016] Calculate the sum of the maximum pooling map and the average pooling map to get the pooling result map.
[0017] The pooling result map is input into the first depth-separable convolutional layer to obtain the second feature map.
[0018] In combination with the first aspect, in some possible implementations, the Transformer module includes: a Swin-LIEBlock unit and a Patch Merging unit.
[0019] The Swin-LIE Block unit is: the MLP module is completely replaced by the Swin Block unit of the local information enhancement module; wherein the local information enhancement module includes: a dimension increase layer, a second convolution layer, a second depth-separable convolution layer, a third convolution layer and a dimension reduction layer.
[0020] The second feature map is input into the Transformer module to obtain a global feature map, including:
[0021] The second feature map is input into the Swin-LIE Block unit and the Patch Merging unit to obtain the global feature map.
[0022] In combination with the first aspect, in some possible implementations, the Transformer module includes: a first Transformer module and a second Transformer module.
[0023] The first Transformer module includes: a first Swin-LIE Block unit and a first Patch Merging unit; the second Transformer module includes: a second Swin-LIE Block unit and a second Patch Merging unit.
[0024] The second feature map is input into the Transformer module to obtain a global feature map, including:
[0025] The second feature map is input into the first Swin-LIE Block unit and the first Patch Merging unit to obtain the first target image.
[0026] The first target image is input into the second Swin-LIE Block unit and the second Patch Merging unit to obtain the second target image.
[0027] The first target image and the second target image are input into the decoder for image fusion to obtain a global feature map.
[0028] In combination with the first aspect, in some possible implementations, inputting the second feature map into the first Swin-LIEBlock unit and the first Patch Merging unit to obtain a first target image includes:
[0029] The second feature map is input into the dimension increase layer to obtain a dimension increase map.
[0030] The dimension-increased graph is input into the second convolutional layer to obtain the first convolutional graph.
[0031] The first convolutional map is input into the second depth-wise separable convolutional layer to obtain a depth-wise separable convolutional map.
[0032] The first convolution map and the depth-separable convolution map are added and input into the third convolution layer to obtain the second convolution map.
[0033] The second convolutional graph is input into the dimensionality reduction layer to obtain a dimensionality reduced graph.
[0034] The dimension reduction map and the second feature map are added and input into the Patch Merging unit to obtain the first target image.
[0035] In combination with the first aspect, in some possible implementations, the method further includes:
[0036] The mean square error between the height estimate and the true height of the target object, the mean square error of the gradient in the horizontal direction of the height estimate and the true height of the target object, and the mean square error of the gradient in the vertical direction of the height estimate and the true height of the target object are calculated to obtain an estimated error value.
[0037] When the estimated error value is greater than or equal to the preset threshold, the hyperparameters of the hybrid pooling patch embedding module and the Transformer module are adjusted until the estimated error value is less than the preset threshold.
[0038] In combination with the first aspect, in some possible implementations, before the monocular remote sensing image is input into the hybrid pooling patch embedding module, the method further includes:
[0039] The monocular remote sensing image is preprocessed, and the preprocessed monocular remote sensing image is input into the hybrid pooling patch embedding module.
[0040] In a second aspect, an embodiment of the present application provides a monocular remote sensing image height data estimation device, comprising:
[0041] The data acquisition module is used to acquire the monocular remote sensing image of the target object.
[0042] The first processing module is used to input the monocular remote sensing image into the first convolution layer of the hybrid pooling patch embedding module to obtain a first feature map, and input the maximum pooling layer, the average pooling layer and the first depth separable convolution layer of the hybrid pooling patch embedding module to obtain a second feature map.
[0043] The second processing module is used to input the second feature map into the Transformer module to obtain a global feature map.
[0044] The first fusion module is used to input the global feature map, the second feature map and the first feature map into the decoder to perform image fusion in sequence to obtain a first output image.
[0045] The result output module is used to obtain a height estimation value of the target object based on the first output image.
[0046] In a third aspect, an embodiment of the present application provides a terminal device, comprising: a processor and a memory, the memory being used to store a computer program, and when the processor executes the computer program, the method for estimating height data of a monocular remote sensing image as described in any one of the first aspects is implemented.
[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for estimating height data of a monocular remote sensing image as described in any one of the first aspects is implemented.
[0048] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0049] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0050] This application processes the image based on the hybrid pooling patch embedding module, which can ensure better initial feature extraction of the image and reduce misjudgment of the height estimation result. The subsequent Transformer module adds the LIE module to enhance the ability to focus on local information. It combines the advantages of two estimation methods (convolution-based monocular remote sensing image height estimation method and Transformer-based monocular remote sensing image height estimation method), so that this scheme can improve the local information extraction capability while accurately judging the height of the target.
[0051] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0053] Figure 1 It is a flowchart of a method for estimating height data of a monocular remote sensing image provided by an embodiment of the present application;
[0054] Figure 2 It is a simplified flowchart of a method for estimating height data of a monocular remote sensing image provided by an embodiment of the present application;
[0055] Figure 3 This is a flow chart of processing an image by a hybrid pooling patch embedding module provided in an embodiment of the present application;
[0056] Figure 4 This is the structure of the Swin-LIE Block unit provided in one embodiment of the present application;
[0057] Figure 5 This is a flow chart of a local information enhancement module processing an image according to an embodiment of the present application;
[0058] Figure 6 is a simplified flow chart of a method for estimating height data of a monocular remote sensing image provided by another embodiment of the present application;
[0059] Figure 7 It is a structural schematic diagram of a monocular remote sensing image height data estimation device provided in one embodiment of the present application;
[0060] Figure 8 It is a structural diagram of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0061] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0062] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0063] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0064] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0065] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0066] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0067] The present invention provides a method for estimating height data of a monocular remote sensing image. Figure 1is a schematic diagram of the process of a method for estimating height data of a monocular remote sensing image provided by an embodiment of the present application, Figure 2 is a flowchart of a method for estimating height data of a monocular remote sensing image provided by an embodiment of the present application, referring to Figure 1 and Figure 2 , the detailed description of the monocular remote sensing image height data estimation method is as follows:
[0068] Step 101, obtaining a monocular remote sensing image of a target object.
[0069] Among them, remote sensing images have the following characteristics compared with other natural images: large coverage, more ground object information, and complex scenes and structures. Ground object features appear differently at different scales, such as cars, pedestrians, building roofs, and viaducts, which have different sizes and scales.
[0070] Step 102: Input the monocular remote sensing image into the first convolutional layer of the hybrid pooling patch embedding module to obtain a first feature map, and input the monocular remote sensing image into the maximum pooling layer, the average pooling layer and the first depth-separable convolutional layer of the hybrid pooling patch embedding module to obtain a second feature map.
[0071] In this embodiment, the hybrid pooling patch embedding module performs a preliminary extraction of feature information in the remote sensing image by first convolving the image instead of dividing it into blocks, so that while reducing the resolution and increasing the number of channels, the spatial information between the various parts of the image is also retained. Subsequently, a two-dimensional feature map with a further reduced resolution is obtained through maximum pooling and average pooling. These feature maps still contain more spatial information. The smooth output provided by average pooling can improve the stability of the model, and the peak value retained by maximum pooling can provide a stronger response. The two can be added together to perform complementary information fusion (by changing the convolution channel to 4 times the original, the image size to 1 / 2 of the original, and then performing pooling, the image size is changed to 1 / 2 of the original under the premise of keeping the channel unchanged. Through two transformations, the same effect of the original Transformer on image segmentation is achieved on the basis of ensuring good local information). In addition, due to the existence of pooling, the module has a certain noise resistance. Maximum pooling can ignore some unimportant details and noise, while average pooling reduces the influence of noise through smoothing processing, which has a beneficial effect on height estimation, because remote sensing images are shot at a very long distance and are easily affected by weather and often have a certain amount of noise. This module brings better initial feature extraction, which reduces misjudgment in height estimation and increases overall accuracy. The feature dimension and size of the module’s final output are the same as those of the Patch Embedding module, which makes it compatible with the following Transformer module.
[0072] Exemplarily, before the monocular remote sensing image is input into the hybrid pooling patch embedding module, the method further includes:
[0073] The monocular remote sensing image is preprocessed, and the preprocessed monocular remote sensing image is input into the hybrid pooling patch embedding module.
[0074] For example, Figure 3 As shown, the maximum pooling layer, average pooling layer and first depth separable convolution layer of the mixed pooling patch embedding module are input to obtain the second feature map, including:
[0075] The first feature map is input into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling map and the average pooling map respectively.
[0076] Calculate the sum of the maximum pooling map and the average pooling map to get the pooling result map.
[0077] The pooling result map is input into the first depth-separable convolutional layer to obtain the second feature map.
[0078] Step 103: input the second feature map into the Transformer module to obtain a global feature map.
[0079] Exemplarily, the Transformer module includes: a Swin-LIE Block unit and a Patch Merging unit.
[0080] like Figure 4 As shown in the figure, the Swin-LIE Block unit is: the MLP module is completely replaced by the Swin Block unit of the local information enhancement module; wherein the local information enhancement module (Local Information Enhancement, LIE) includes: a dimension increase layer, a second convolution layer, a second depth-separable convolution layer, a third convolution layer and a dimension reduction layer.
[0081] Exemplarily, the LIE module reshapes the input one-dimensional image sequence into a two-dimensional image format, uses specific convolution operations to increase the number of feature channels, then learns local information such as the two-dimensional neighboring relationship between pixels through deep separable convolution, and finally restores the original number of channels through dimensionality reduction operations, thereby achieving effective enhancement of local features. In this process, the LIE module adds the features before and after the deep separable convolution, combines the enhanced one-dimensional image sequence with the original input one-dimensional image sequence, and retains the original information while enhancing the information. These operations enable the original Transformer module to enhance its ability to focus on local information. After using this module, the height estimation results will be clearer, with clearer textures and edges, and can solve the problems of loss of detail information in the height estimation results, blurred height maps, and lack of necessary details and textures.
[0082] Exemplarily, the specific processing process of the LIE module on the data is:
[0083] ① Sequence to image: Expand the input one-dimensional image sequence data (format [B, L, C]) into a two-dimensional form (format [B, H, W, C]) and change the dimension to [B, C, H, W], where L = H*W, B is the number of images, L is the length of the image sequence in one-dimensional form, C is the number of image channels, and HW is the height and width of the image in two-dimensional form. For example, the input is (1, 65536, 3) and is expanded to (1, 3, 256, 256).
[0084] ②The B*C*H*W two-dimensional image data obtained in the previous step is subjected to 1*1 convolution to expand the number of channels to obtain a B*4C*H*W size feature map.
[0085] ③ Add the B*4C*H*W size feature map obtained in the previous step to the feature map of the same size obtained by a 3*3 depth-wise separable convolution to obtain a B*4C*H*W feature map containing position information.
[0086] ④The B*4C*H*W two-dimensional image data obtained in the previous step is subjected to 1*1 convolution to reduce the number of channels and obtain a B*C*H*W size feature map.
[0087] ⑤Image to sequence: Convert the two-dimensional B*C*H*W size feature map to B*C*(HW), and then exchange the dimensions to B*(HW)*C, that is, B*L*C sequence data.
[0088] ⑥Add the B*L*C data from the previous step and the input data from step ① to combine the information.
[0089] Composition: This module has a 1*1 convolutional layer, a 3*3 depth-separable convolutional layer, an image sequence conversion part, and a sequence-to-image conversion part.
[0090] Input and output: input sequence data, output sequence data.
[0091] Function: Convert the image sequence data processed by Transformer into two-dimensional data, and learn the local information of the image through convolution, so that the network can not only learn more global information through Transformer, but also enhance the two-dimensional local information in the feature map through this module.
[0092] The second feature map is input into the Transformer module to obtain a global feature map, including:
[0093] The second feature map is input into the Swin-LIE Block unit and the Patch Merging unit to obtain the global feature map.
[0094] Exemplarily, the Transformer module includes: a first Transformer module and a second Transformer module.
[0095] The first Transformer module includes: a first Swin-LIE Block unit and a first Patch Merging unit; the second Transformer module includes: a second Swin-LIE Block unit and a second Patch Merging unit.
[0096] The second feature map is input into the Transformer module to obtain a global feature map, including:
[0097] The second feature map is input into the first Swin-LIE Block unit and the first Patch Merging unit to obtain the first target image.
[0098] The first target image is input into the second Swin-LIE Block unit and the second Patch Merging unit to obtain the second target image.
[0099] The first target image and the second target image are input into the decoder for image fusion to obtain a global feature map.
[0100] For example, Figure 5 As shown, the second feature map is input into the first Swin-LIE Block unit and the first PatchMerging unit to obtain the first target image, including:
[0101] The second feature map is input into the dimension increase layer to obtain a dimension increase map.
[0102] The dimension-increased graph is input into the second convolutional layer to obtain the first convolutional graph.
[0103] The first convolutional map is input into the second depth-wise separable convolutional layer to obtain a depth-wise separable convolutional map.
[0104] The first convolution map and the depth-separable convolution map are added and input into the third convolution layer to obtain the second convolution map.
[0105] The second convolutional graph is input into the dimensionality reduction layer to obtain a dimensionality reduced graph.
[0106] The dimension reduction map and the second feature map are added and input into the Patch Merging unit to obtain the first target image.
[0107] For example, Figure 6As shown, the number of Transformer modules can also be three. Considering the accuracy requirement of the solution itself in the urban area, the solution with three Transformer modules can be defined as the optimal solution. The number of Transformer modules set here is not limited.
[0108] Step 104: input the global feature map, the second feature map, and the first feature map into a decoder for image fusion in sequence to obtain a first output image.
[0109] For example, based on Figure 6 It is not difficult to see that the process of sequential fusion of the global feature map, the second feature map and the first feature map input into the decoder for image fusion can be:
[0110] The global feature map is first fused with the second feature map, and the result is then fused with the first feature image.
[0111] Step 105: Obtain a height estimation value of the target object based on the first output image.
[0112] Exemplarily, the method further comprises:
[0113] The mean square error between the height estimate and the true height of the target object, the mean square error of the gradient in the horizontal direction of the height estimate and the true height of the target object, and the mean square error of the gradient in the vertical direction of the height estimate and the true height of the target object are calculated to obtain an estimated error value.
[0114] When the estimated error value is greater than or equal to the preset threshold, the hyperparameters of the hybrid pooling patch embedding module and the Transformer module are adjusted until the estimated error value is less than the preset threshold.
[0115] For example, when estimating the height of a remote sensing image, the estimated result is often uneven. Even if the overall height of these areas is accurate, the surface is uneven (with frequent height fluctuations). The following figure is specially processed to facilitate the observation of this uneven height map. Our local information enhancement module (LIE) enhances the network's ability to extract details, but also exacerbates this surface unevenness.
[0116] In the past, when estimating height, only the loss between the predicted value and the true value was calculated (for example, using MAE or MSE). Although this method can fit the overall height value well, it does not take into account the constraints on image details and edge information, and the predicted result will produce this uneven phenomenon. Therefore, the calculation of the above estimated error value is designed. The gradient constraint can be considered by estimating the error value. The introduction of gradient loss enhances the accuracy of the object edge while avoiding the uneven effect of the same height area.
[0117] The above-mentioned monocular remote sensing image height data estimation method processes the image based on the mixed pooling patch embedding module, which can ensure better initial feature extraction of the image and reduce misjudgment of the height estimation result. The subsequent Transformer module adds the LIE module to enhance the ability to pay attention to local information. It combines the advantages of two estimation methods (convolution-based monocular remote sensing image height estimation method and Transformer-based monocular remote sensing image height estimation method), so that this scheme can improve the local information extraction capability while accurately judging the height of the target.
[0118] In order to facilitate the understanding of this solution, the above solution is explained in the form of specific embodiment 1.
[0119] Embodiment 1:
[0120] (1) Data preprocessing
[0121] The dataset used in this method is the US3D public dataset, which contains satellite images and lidar-derived reference labels. The size of US3D images is 2048*2048 pixels, with a spatial resolution of 30-50cm. We downsampled them by a factor of two to get 1024*1024 pixel images for training and prediction.
[0122] The Potsdam dataset is a remote sensing image dataset for urban semantic segmentation. It contains 38 high-definition images with a resolution of 5 cm, and the size of each image is 6000*6000 pixels. The images in the figure are orthophotos, and 34 of them are randomly selected as training sets and 4 as validation sets. Each image is divided into 36 1024*1024 remote sensing images with overlapping areas by sliding windows with a window size of 1024*1024 and a step size of 995.
[0123] For other data, images with a size of twice 1024 or less are directly sampled to 1024. Images with a size of more than twice 1024 are divided into multiple images of 1024 using a sliding window.
[0124] (2) Network construction
[0125] A monocular remote sensing image height estimation network based on a hybrid pooling patch embedding module and a local information enhancement module designed by the present invention is constructed.
[0126] Network forward propagation process, such as Figure 6 As shown:
[0127] ① Input image: 1024*1024*3-channel monocular remote sensing image.
[0128] ②Stage 1: HPPE module (hybrid pooling patch embedding module):
[0129] The input image first passes through a 7×7 convolutional layer and outputs a feature map of size 512*512*64.
[0130] Next, the feature map is subjected to maximum pooling and average pooling operations to obtain two feature maps of 256*256*64 channels respectively.
[0131] Finally, the two feature maps are added together and fused through a 1×1 depthwise separable convolution (DWConv) to obtain a feature map of 256*256*64 channels.
[0132] ③Stage 2:
[0133] Two series-connected Swin-LIEBlocks receive a feature map of 256*256*64 channels as input and output a feature map of 256*256*64 channels after processing.
[0134] The above output 256*256*64 feature map is input into the PatchMerging module to halve the size of the feature map and double the number of channels to output a 128*128*128 feature map.
[0135] ④Stage 3:
[0136] Two series-connected Swin-LIE-TransformerBlocks receive a feature map of 128*128*128 channels as input and output a feature map of 128*128*128 channels.
[0137] The above output 128*128*128 feature map is input into the PatchMerging module to halve the size of the feature map and double the number of channels to output a 64*64*256 feature map.
[0138] ⑤Stage 4:
[0139] Ten series-connected Swin-LIE-TransformerBlocks receive the feature map of 64*64*256 channels in the previous step as input and output a feature map of 64*64*256 channels.
[0140] The above output 64*64*256 feature map is input into the PatchMerging module to halve the size of the feature map and double the number of channels to output a 32*32*512 feature map.
[0141] ⑥DecoderBlock:
[0142] The encoder finally outputs a 32*32*512 feature map, which is up-sampled layer by layer by four DecoderBlocks and fused with the skip connection feature map from the encoder, and finally outputs a 1024*1024*16 channel feature map.
[0143] ⑦HeightHead
[0144] Accept the final output feature map of size 1024*1024*16 from the decoder, and output a height map of size 1024*1024*1. The height map has the same size as the original image, and each pixel corresponds to the height value of the position object in the original image relative to the ground.
[0145] (3) Training network weights
[0146] ① Initialize all parameters of the network using random initialization
[0147] ② Prepare the data set and divide the preprocessed Potsdam or US3D data set into training set and validation set.
[0148] ③Define the loss function (calculation formula for estimating the error value)
[0149] ④Define the optimizer, Adam sets the initial learning rate, momentum and other hyperparameters
[0150] ⑤At the beginning of each epoch, divide the dataset into multiple batches.
[0151] For each batch of data, perform the following steps:
[0152] 1. Send the input image into the network to get the predicted height map.
[0153] 2. Calculate the height loss and height gradient loss between the predicted height map and the true height map.
[0154] 3. Perform back propagation and update network parameters.
[0155] ⑥Verify and save the model:
[0156] At the end of each epoch, use the validation set to evaluate the model performance. If the model performance on the validation set improves, save the model weights. Repeat step ⑤ until the model performance stabilizes and no longer increases, and finally select the model weights with the best performance.
[0157] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0158] Corresponding to the monocular remote sensing image height data estimation method described in the above embodiment, Figure 7 A structural block diagram of a monocular remote sensing image height data estimation device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0159] See also Figure 7 , the monocular remote sensing image height data estimation device in the embodiment of the present application may include:
[0160] The data acquisition module 201 is used to acquire a monocular remote sensing image of a target object.
[0161] The first processing module 202 is used to input the monocular remote sensing image into the first convolution layer of the hybrid pooling patch embedding module to obtain a first feature map, and input the maximum pooling layer, average pooling layer and first depth-separable convolution layer of the hybrid pooling patch embedding module to obtain a second feature map.
[0162] The second processing module 203 is used to input the second feature map into the Transformer module to obtain a global feature map.
[0163] The first fusion module 204 is used to input the global feature map, the second feature map and the first feature map into the decoder to perform image fusion in sequence to obtain a first output image.
[0164] The result output module 205 is used to obtain a height estimation value of the target object based on the first output image.
[0165] Exemplarily, the first processing module 202 may be used to:
[0166] The first feature map is input into the maximum pooling layer and the average pooling layer respectively to obtain the maximum pooling map and the average pooling map respectively.
[0167] Calculate the sum of the maximum pooling map and the average pooling map to get the pooling result map.
[0168] The pooling result map is input into the first depth-separable convolutional layer to obtain the second feature map.
[0169] Exemplarily, the Transformer module includes: a Swin-LIE Block unit and a Patch Merging unit.
[0170] The Swin-LIE Block unit is: the MLP module is completely replaced by the Swin Block unit of the local information enhancement module; wherein the local information enhancement module includes: a dimension increase layer, a second convolution layer, a second depth-separable convolution layer, a third convolution layer and a dimension reduction layer.
[0171] The second processing module 203 may be used to:
[0172] The second feature map is input into the Swin-LIE Block unit and the Patch Merging unit to obtain the global feature map.
[0173] Exemplarily, the Transformer module includes: a first Transformer module and a second Transformer module.
[0174] The first Transformer module includes: a first Swin-LIE Block unit and a first Patch Merging unit; the second Transformer module includes: a second Swin-LIE Block unit and a second Patch Merging unit.
[0175] The second processing module 203 may be used to:
[0176] The second feature map is input into the first Swin-LIE Block unit and the first Patch Merging unit to obtain the first target image.
[0177] The first target image is input into the second Swin-LIE Block unit and the second Patch Merging unit to obtain the second target image.
[0178] The first target image and the second target image are input into the decoder for image fusion to obtain a global feature map.
[0179] Exemplarily, the second processing module 203 may be used to:
[0180] The second feature map is input into the dimension increase layer to obtain a dimension increase map.
[0181] The dimension-increased graph is input into the second convolutional layer to obtain the first convolutional graph.
[0182] The first convolutional map is input into the second depth-wise separable convolutional layer to obtain a depth-wise separable convolutional map.
[0183] The first convolution map and the depth-separable convolution map are added and input into the third convolution layer to obtain the second convolution map.
[0184] The second convolutional graph is input into the dimensionality reduction layer to obtain a dimensionality reduced graph.
[0185] The dimension reduction map and the second feature map are added and input into the Patch Merging unit to obtain the first target image.
[0186] Exemplarily, the result output module 205 is further used for:
[0187] The mean square error between the height estimate and the true height of the target object, the mean square error of the gradient in the horizontal direction of the height estimate and the true height of the target object, and the mean square error of the gradient in the vertical direction of the height estimate and the true height of the target object are calculated to obtain an estimated error value.
[0188] When the estimated error value is greater than or equal to the preset threshold, the hyperparameters of the hybrid pooling patch embedding module and the Transformer module are adjusted until the estimated error value is less than the preset threshold.
[0189] Exemplarily, before the monocular remote sensing image is input into the hybrid pooling patch embedding module, the first processing module 202 may also be used to:
[0190] The monocular remote sensing image is preprocessed, and the preprocessed monocular remote sensing image is input into the hybrid pooling patch embedding module.
[0191] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0192] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0193] The present application also provides a terminal device, see Figure 8The terminal device 300 may include: at least one processor 310 and a memory 320, wherein the memory 320 is used to store a computer program 321, and the processor 310 is used to call and run the computer program 321 stored in the memory 320 to implement the steps in any of the above-mentioned method embodiments, for example Figure 1 Steps 101 to 105 in the illustrated embodiment. Alternatively, when the processor 310 executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 7 The functions of each module are shown.
[0194] Exemplarily, the computer program 321 may be divided into one or more modules / units, one or more modules / units are stored in the memory 320, and are executed by the processor 310 to complete the present application. The one or more modules / units may be a series of computer program segments that can complete specific functions, and the program segments are used to describe the execution process of the computer program in the terminal device 300.
[0195] Those skilled in the art will understand that Figure 8 It is only an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components, such as input and output devices, network access devices, buses, etc.
[0196] The processor 310 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.
[0197] The memory 320 may be an internal storage unit of the terminal device, or an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 320 is used to store the computer program and other programs and data required by the terminal device. The memory 320 may also be used to temporarily store data that has been output or is to be output.
[0198] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0199] The monocular remote sensing image height data estimation method provided in the embodiment of the present application can be applied to terminal devices such as computers, wearable devices, vehicle-mounted devices, tablet computers, laptops, netbooks, etc. The embodiment of the present application does not impose any restrictions on the specific type of terminal devices.
[0200] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the steps in each embodiment of the above-mentioned monocular remote sensing image height data estimation method.
[0201] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in each embodiment of the above-mentioned monocular remote sensing image height data estimation method when executing the computer program product.
[0202] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0203] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0204] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0205] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0206] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0207] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for estimating height data of a monocular remote sensing image, characterized in that: include: Acquire a monocular remote sensing image of the target object; Inputting the monocular remote sensing image into the first convolutional layer of the hybrid pooling patch embedding module to obtain a first feature map, and inputting the monocular remote sensing image into the maximum pooling layer, the average pooling layer and the first depth-separable convolutional layer of the hybrid pooling patch embedding module to obtain a second feature map; Inputting the second feature map into the Transformer module to obtain a global feature map; Inputting the global feature map, the second feature map and the first feature map into a decoder to perform image fusion in sequence to obtain a first output image; Based on the first output image, an estimated value of the height of the target object is obtained.
2. The method for estimating height data of a monocular remote sensing image according to claim 1, wherein: The input of the maximum pooling layer, the average pooling layer and the first depth-separable convolutional layer of the hybrid pooling patch embedding module to obtain a second feature map includes: Inputting the first feature map into the maximum pooling layer and the average pooling layer respectively to obtain a maximum pooling map and an average pooling map respectively; Calculate the sum of the maximum pooling map and the average pooling map to obtain the pooling result map; The pooling result map is input into a first depth-separable convolutional layer to obtain the second feature map.
3. The method for estimating height data of a monocular remote sensing image according to claim 1, wherein: The Transformer module includes: a Swin-LIE Block unit and a Patch Merging unit; The Swin-LIE Block unit is a Swin Block unit in which all MLP modules are replaced by local information enhancement modules; wherein the local information enhancement module includes a dimension increase layer, a second convolution layer, a second depth-separable convolution layer, a third convolution layer and a dimension reduction layer; The step of inputting the second feature map into a Transformer module to obtain a global feature map includes: The second feature map is input into the Swin-LIE Block unit and the Patch Merging unit to obtain a global feature map.
4. The method for estimating height data of a monocular remote sensing image as claimed in claim 3, characterized in that: The Transformer module includes: a first Transformer module and a second Transformer module; The first Transformer module includes: a first Swin-LIE Block unit and a first Patch Merging unit; the second Transformer module includes: a second Swin-LIE Block unit and a second Patch Merging unit; The second feature map is input into the Transformer module to obtain a global feature map, including: Inputting the second feature map into the first Swin-LIE Block unit and the first Patch Merging unit to obtain a first target image; Inputting the first target image into a second Swin-LIE Block unit and a second Patch Merging unit to obtain a second target image; The first target image and the second target image are input into a decoder for image fusion to obtain the global feature map.
5. The method for estimating height data of a monocular remote sensing image as claimed in claim 4, characterized in that: The step of inputting the second feature map into the first Swin-LIE Block unit and the first Patch Merging unit to obtain a first target image includes: Inputting the second feature map into the dimension increasing layer to obtain a dimension increasing map; Inputting the dimension-increased graph into the second convolutional layer to obtain a first convolutional graph; Inputting the first convolutional map into a second depth-separable convolutional layer to obtain a depth-separable convolutional map; Adding the first convolution map and the depth-separable convolution map and inputting the resultant convolution map into the third convolution layer to obtain a second convolution map; Inputting the second convolutional graph into the dimension reduction layer to obtain a dimension reduction graph; The dimension reduction map and the second feature map are added and input into a Patch Merging unit to obtain the first target image.
6. The method for estimating height data of a monocular remote sensing image according to claim 1, wherein: The method further comprises: Calculate the sum of the mean square error between the height estimate and the true height value of the target object, the mean square error of the gradient in the horizontal direction of the height estimate and the true height value of the target object, and the mean square error of the gradient in the vertical direction of the height estimate and the true height value of the target object to obtain an estimated error value; When the estimated error value is greater than or equal to a preset threshold, the hyper parameters of the hybrid pooling patch embedding module and the Transformer module are adjusted until the estimated error value is less than the preset threshold.
7. The method for estimating height data of a monocular remote sensing image according to claim 1, wherein: Before inputting the monocular remote sensing image into the hybrid pooling patch embedding module, the method further includes: The monocular remote sensing image is preprocessed, and the preprocessed monocular remote sensing image is input into a hybrid pooling patch embedding module.
8. A device for estimating height data of a monocular remote sensing image, characterized in that: include: A data acquisition module, used to acquire a monocular remote sensing image of a target object; A first processing module is used to input the monocular remote sensing image into the first convolution layer of the hybrid pooling patch embedding module to obtain a first feature map, and input the maximum pooling layer, the average pooling layer and the first depth-separable convolution layer of the hybrid pooling patch embedding module to obtain a second feature map; A second processing module, used for inputting the second feature map into a Transformer module to obtain a global feature map; A first fusion module is used to input the global feature map, the second feature map and the first feature map into a decoder to perform image fusion in sequence to obtain a first output image; A result output module is used to obtain a height estimation value of the target object based on the first output image.
9. A terminal device, comprising: A processor and a memory, wherein the memory stores a computer program that can be run on the processor, characterized in that when the processor executes the computer program, the method for estimating height data of a monocular remote sensing image as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for estimating height data of a monocular remote sensing image as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Single remote sensing image height information estimation method based on deep learning algorithm
CN114972989A
Monocular depth estimation method and system based on complete context information
CN116205962A
Monocular image depth estimation method based on attention mechanism
CN116630387A
Remote sensing image semantic segmentation method, system, equipment and medium
CN119251497A
Method and apparatus for depth estimation of monocular image, and storage medium
US20200226773A1