A method and apparatus for three-dimensional visualization of a construction area
By converting optical images into synthetic aperture radar images and generating affine transformation matrices, the problems of high acquisition costs and insufficient information in 3D modeling of construction areas are solved, achieving efficient and accurate 3D visualization.
Patent Information
- Application Number
- CN202510710806.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing technologies for 3D modeling of construction areas suffer from problems such as the time-consuming and expensive acquisition of synthetic aperture radar images, making it difficult to meet the needs of high-frequency monitoring, and the lack of depth information in optical images, which cannot be directly used for 3D modeling.
By acquiring initial synthetic aperture radar (SAR) images and optical images of the construction area, Fourier transform is used to convert the optical images into similar SAR images, and an affine transformation matrix is generated through a multilayer perceptron to achieve image matching and 3D visualization.
It significantly reduces data acquisition costs, improves the efficiency and accuracy of 3D modeling, and provides reliable 3D visualization support for real-time monitoring of construction sites.
Smart Images

Figure CN120612427B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of three-dimensional visualization, and particularly relate to a three-dimensional visualization method and device for a construction area. BACKGROUND
[0002] Constructing a three-dimensional model of a construction area can systematically improve the construction efficiency, safety and sustainability of the construction area, and has high practical value.
[0003] At present, a three-dimensional model of a construction area is mainly constructed by acquiring a SAR two-dimensional image of the construction area through a synthetic aperture radar, and then performing denoising, registration, fusion and other operations based on the SAR two-dimensional image to generate a three-dimensional model of the construction area. However, the construction area has strong personnel mobility and vehicle mobility, and image acquisition monitoring needs to be performed all day long.
[0004] There are two main defects in SAR-based two-dimensional image acquisition: on the one hand, the SAR image acquisition process is time-consuming and expensive, and it is difficult to meet the demand for high-frequency monitoring; on the other hand, although the traditional optical image is convenient to acquire, it lacks depth information and cannot be directly used for three-dimensional modeling. SUMMARY
[0005] The following is a summary of the subject matter of the detailed description. This summary is not intended to limit the scope of the claims.
[0006] The main purpose of the embodiments of the present application is to provide a three-dimensional visualization method and device for a construction area, which can fully exploit the correlation between different domain images, improve the accuracy of source domain image conversion to target domain image, and provide reliable three-dimensional visualization support for real-time monitoring of the construction site.
[0007] To achieve the above purpose, a first aspect of the embodiments of the present application provides a three-dimensional visualization method for a construction area, the method comprising:
[0008] acquiring an initial synthetic aperture radar image of the construction area and a current optical image, and Fourier transforming the current optical image into a pseudo synthetic aperture radar image;
[0009] stitching the initial synthetic aperture radar image and the pseudo synthetic aperture radar image as a to-be-input image, extracting global features of multiple scales and local features of multiple scales from the to-be-input image, fusing each global feature of the multiple scales and a local feature of a corresponding scale of the multiple scales to obtain fusion features of the multiple scales, and up-sampling the fusion features of the multiple scales to generate final features;
[0010] generating an affine transformation matrix according to the final feature based on a multi-layer perception, and converting the synthetic aperture radar image into a warped synthetic aperture radar image according to the affine transformation matrix;
[0011] generating a spatial point cloud of the construction area according to the warped synthetic aperture radar image and the initial synthetic aperture radar image to realize three-dimensional visualization of the construction area.
[0012] The three-dimensional visualization method of the construction area provided in the application has at least the following beneficial effects:
[0013] The method introduces an optical image as a source domain image by taking a synthetic aperture radar image as a target domain image, finds an affine transformation matrix to enable the source domain image to be converted to the target domain image and the corresponding points to be completely consistent in spatial position, and then realizes matching between the target domain image and the source domain image based on the affine transformation matrix, thereby performing three-dimensional modeling, significantly reducing data acquisition cost, significantly improving implementation efficiency on the premise of ensuring modeling accuracy. Moreover, the method is designed with double branches, and the global features and local features of the image obtained by splicing the target domain image and the source domain image are extracted in parallel at multiple scales, which can fully exploit the correlation between different domain images and improve the accuracy of converting the source domain image to the target domain image, thereby providing reliable three-dimensional visualization support for real-time monitoring of the construction site.
[0014] In some embodiments, the extracting the global features of multiple scales and the local features of multiple scales from the to-be-input image comprises:
[0015] inputting the to-be-input image into a first branch network to obtain the global features of multiple scales output by the first branch network; wherein the first branch network comprises a plurality of cascaded Transformer units and a down-sampling unit between every two Transformer units;
[0016] inputting the to-be-input image into a second branch network to obtain the local features of multiple scales output by the second branch network; wherein the second branch network comprises a plurality of cascaded CNN units and a down-sampling unit between every two CNN units.
[0017] In some embodiments, the process of fusing the global features of the target scale and the local features of the target scale to obtain the fusion features of the target scale comprises:
[0018] The global features at the target scale are subjected to global average pooling and global max pooling to obtain the first pooled features. The global features at the target scale are then subjected to global max pooling to obtain the second pooled features. The global features at the target scale are any one of the global features at the multiple scales, and the local features at the target scale are local features at the same scale as the global features at the target scale.
[0019] The local features at the target scale are subjected to global average pooling to obtain the third pooling feature, and the local features at the target scale are subjected to global max pooling to obtain the fourth pooling feature.
[0020] The first pooling feature and the second pooling feature are respectively used as the head and tail of the local feature at the target scale, and are concatenated to obtain the first concatenated feature. The third pooling feature and the fourth pooling feature are respectively used as the head and tail of the global feature at the target scale, and are concatenated to obtain the second concatenated feature.
[0021] The first and second splicing features are respectively input into the Transformer unit, and the features output by the Transformer unit are spliced together to obtain the first intermediate feature;
[0022] The global features at the target scale are input into the attention enhancement unit to obtain the second intermediate features output by the attention enhancement unit;
[0023] The third intermediate feature is obtained by multiplying the global features and local features of the target scale pixel by pixel.
[0024] The first intermediate feature, the second intermediate feature, and the third intermediate feature are concatenated to obtain the fused feature at the target scale.
[0025] In some embodiments, the attention enhancement unit is a CBAM convolutional block attention unit.
[0026] In some embodiments, the process of the CNN unit extracting local features includes:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] wherein, is input data of the CNN unit, is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a convolutional layer of is a SiLU activation function, is a channel max pooling, is a channel average pooling, is a convolutional layer of is a sigmoid activation function, is a mapping function of an element rearrangement operation, is a multi-layer perception, is a local feature extracted by the CNN unit. In some embodiments, before the fusing each global feature of the plurality of scales of global features and the corresponding local feature of the plurality of scales of local features to obtain a plurality of scales of fused features, the method further comprises:
[0036] calculating a difference image between the initial synthetic aperture radar image and the pseudo synthetic aperture radar image based on an LR operator;
[0037] extracting a plurality of scales of difference features of the difference image;
[0038] extracting a plurality of scales of difference features of the difference image;
[0039] the fusing each global feature of the plurality of scales of global features and the corresponding local feature of the plurality of scales of local features to obtain a plurality of scales of fused features, comprising:
[0040] fusing each global feature of the plurality of scales of global features, the corresponding difference feature of the plurality of scales of difference features, and the corresponding local feature of the plurality of scales of local features to obtain a plurality of scales of fused features.
[0041] To achieve the above object, a second aspect of the embodiment of the present application provides a three-dimensional visualization device of a construction area, the device comprising:
[0042] an image acquisition module, configured to acquire an initial synthetic aperture radar image and a current optical image of the construction area, and to perform Fourier transform on the current optical image to obtain an analog synthetic aperture radar image;
[0043] a feature extraction module, configured to splice the initial synthetic aperture radar image and the analog synthetic aperture radar image as a to-be-input image, to extract global features of multiple scales and local features of multiple scales from the to-be-input image, to fuse each global feature of the multiple scales and a corresponding local feature of the multiple scales to obtain fusion features of multiple scales, and to perform up-sampling fusion on the fusion features of multiple scales to generate final features;
[0044] an image affine transform module, configured to generate an affine transform matrix according to the final features by using a multi-layer perception, and to convert the analog synthetic aperture radar image into a distorted synthetic aperture radar image according to the affine transform matrix;
[0045] a three-dimensional visualization module, configured to generate a spatial point cloud of the construction area according to the distorted synthetic aperture radar image and the initial synthetic aperture radar image, so as to realize three-dimensional visualization of the construction area.
[0046] To achieve the above object, a third aspect of the embodiment of the present application provides an electronic device, comprising at least one control processor and a memory connected with the at least one control processor in communication; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the three-dimensional visualization method of the construction area.
[0047] To achieve the above object, a fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores computer executable instructions for enabling a computer to execute the three-dimensional visualization method of the construction area.
[0048] It can be understood that the beneficial effects of the second aspect to the fourth aspect and the related technologies are the same as the beneficial effects of the first aspect and the related technologies, and the related description can be referred to in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or related description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0050] Figure 1 is a flowchart of a three-dimensional visualization method of a construction area provided by an embodiment of the present application;
[0051] Figure 2 is a structural diagram of a model network provided by an embodiment of the present application;
[0052] Figure 3 is a unit structure diagram of a CNN provided by an embodiment of the present application;
[0053] Figure 4 is a structural diagram of a three-dimensional visualization system of a construction area provided by an embodiment of the present application;
[0054] Figure 5 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0056] As Figure 1 , an embodiment of the present application provides a three-dimensional visualization method of a construction area, the method comprising the following steps S110 to S140:
[0057] Step S110, acquiring an initial synthetic aperture radar image and a current optical image of a construction area, and Fourier transforming the current optical image into a pseudo synthetic aperture radar image.
[0058] In this step, the construction area refers to the area that needs to be three-dimensionally visualized. The initial synthetic aperture radar image refers to the synthetic aperture radar image acquired by using a synthetic aperture radar to collect the construction area. The present application uses the initial synthetic aperture radar image as a target domain image, and the subsequent all-weather collected optical image as a source domain image. The source domain image needs to be matched with the target domain image, so that the optical image can meet the standard of constructing a three-dimensional point cloud by the synthetic aperture radar image.
[0059] And the purpose of the present application is to match two or more images (source domain image and target domain image) of the same scene taken at different times, different optical sensors or different perspectives, by finding the spatial transformation rule (affine transformation matrix) to make the corresponding points of the two images completely consistent in spatial position.
[0060] Since the acquisition cost of synthetic aperture radar image is large, it is considered to acquire optical image all-weather, and based on Fourier transform, the acquired optical image is transformed into pseudo synthetic aperture radar image, so as to facilitate the subsequent matching process. The process of constructing pseudo synthetic aperture radar image is briefly introduced as follows:
[0061] For the image with the input size of , the Fourier transform is:
[0062] ;
[0063] Wherein, is the image after Fourier transform, is the variable in the frequency domain, is the variable in the spatial domain.
[0064] The center area of the image after Fourier transform is low frequency component, and the remaining area is high frequency component.
[0065] A mask can be used to assist the change of low frequency area, so that only the area near the center (0, 0) of the image is replaced:
[0066] ;
[0067] In the formula, is the ratio of the width and height of the center area of the image to the width and height of the overall area of the image, and here is 0.01.
[0068] The center low frequency component of the optical image in the source domain image is replaced by the low frequency component of the corresponding SAR image of the target domain, and then the frequency domain image is converted into spatial domain image by inverse Fourier transform, and the calculation formula of the pseudo SAR image generated by inverse Fourier transform is:
[0069] ;
[0070] In the formula, is the pseudo synthetic aperture radar image generated by the source domain optical image; , are the optical image of the source domain and the corresponding synthetic aperture radar image respectively; , correspond to the amplitude and phase of the image respectively. is an element-wise multiplication operation.
[0071] In this embodiment, the initial synthetic aperture radar image is taken as the target domain image, the current optical image is taken as the source domain image, and the source domain image is Fourier transformed into the pseudo synthetic aperture radar image. Finally, the spatial transformation rule (affine transformation matrix) is found to make the corresponding points of the source domain image and the target domain image completely consistent in spatial position.
[0072] In step S120, the initial synthetic aperture radar image and the pseudo synthetic aperture radar image are spliced as the to-be-input image. The global features of multiple scales and the local features of multiple scales are extracted from the to-be-input image. Each global feature of multiple scales and the corresponding local feature of multiple scales are fused to obtain the fusion features of multiple scales. The fusion features of multiple scales are up-sampled to generate the final features.
[0073] As Figure 2 In this step, the initial synthetic aperture radar image and the pseudo synthetic aperture radar image are spliced into the to-be-input image, which is used as the input data of the model. Here, the model refers to a model capable of generating the affine transformation matrix in steps S120 and S130. The main function of the model is to extract the global features and the local features in the to-be-input image, and to fuse the global features and the local features. Finally, the final features are used to generate the affine transformation matrix based on the multi-layer perception.
[0074] Further, the model includes:
[0075] (1) The first branch network includes four cascaded Transformer units, and a down-sampling unit is arranged between every two Transformer units. The Transformer unit is used to extract the global features of the image of the corresponding scale. The down-sampling unit is mainly used to down-sample the global features of the image. The Transformer unit refers to a neural network structure based on the self-attention mechanism, which can be implemented by a module containing a multi-head attention layer and a feedforward neural network, and is used to capture the long-distance spatial dependence in the image. The down-sampling unit refers to an operation module for reducing the spatial resolution of the feature map, which can be implemented by a convolution layer or a pooling layer with a step greater than 1. By gradually reducing the size of the feature map, the receptive field is expanded and the computational complexity is reduced.
[0076] (2) the second branch network, the first branch network comprises four cascaded CNN units and a down-sampling unit between each two CNN units. The CNN unit is used to extract image local features of a corresponding scale. The down-sampling unit is mainly used for down-sampling the image local features. The CNN unit refers to a convolutional neural network structure, which can be specifically implemented by a combination of stacked convolutional layers and a nonlinear activation function, and is used to extract texture and edge features of an image local region.
[0077] (3) the first fusion unit, which is mainly used to fuse each scale of global features in the multiple scales of global features and the corresponding scale of local features in the multiple scales of local features.
[0078] (4) the second fusion unit, which can be composed of four Transformer units and four up-sampling units, and is mainly used to up-sample and fuse the multiple scales of fused features to generate the final features.
[0079] (5) the multi-layer perceptron, which is mainly used to generate an affine transformation matrix from the final features.
[0080] Further, the input image is processed by cascaded Transformer units in the first branch network. Each Transformer unit models global context information through a self-attention mechanism, and the down-sampling unit compresses the spatial resolution of the feature map. Finally, multiple scales of global features are output. The second branch network extracts local features of different scales through cascaded CNN units. The CNN unit focuses on local neighborhood information aggregation in the convolution operation, and the down-sampling unit synchronously adjusts the feature map size. The multi-scale global features and local features output by the two networks are complementary. The global features represent the overall structure of the construction area, and the local features retain the construction details, such as equipment outlines or material textures.
[0081] Compared with the prior art, the traditional method usually uses CNN or Transformer structure alone to extract features, which makes it difficult to balance global information and local details. For example, a single CNN network may lose long-distance correlation information at a deep layer, and a structure using only Transformer has limitations in computational efficiency and local feature capture capability. Through the dual-branch design, global features and local features are extracted in parallel at multiple scales, effectively combining the advantages of the two structures.
[0082] In order to improve the matching degree of the source domain image (generated by a visible light image) and the target domain image (a synthetic aperture radar image), the method cooperatively extracts multi-scale global and local features, effectively combines the advantages of the two structures, improves the accuracy of the final feature extraction, and finally improves the accuracy of finding the affine transformation matrix between the source domain image and the target domain image.
[0083] In step S130, an affine transformation matrix is generated according to the multi-layer perception based on the final features, and the synthetic aperture radar image is converted into the warped synthetic aperture radar image according to the affine transformation matrix.
[0084] In step S130, an affine transformation matrix is generated according to the multi-layer perception based on the final features, and the synthetic aperture radar image is converted into the warped synthetic aperture radar image according to the affine transformation matrix.
[0085] In step S140, the spatial point cloud of the construction area is generated according to the warped synthetic aperture radar image and the initial synthetic aperture radar image, so as to realize the three-dimensional visualization of the construction area.
[0086] In step S130, an affine transformation matrix is generated according to the multi-layer perception based on the final features, and the synthetic aperture radar image is converted into the warped synthetic aperture radar image according to the affine transformation matrix.
[0087] The method has at least the following beneficial effects:
[0088] The traditional method relies on a single radar image data source and needs to be collected all day long, while the method introduces the synthetic aperture radar image as the target domain image, introduces the optical image as the source domain image, and introduces the optical image conversion mechanism, i.e. finding the spatial transformation rule (affine transformation matrix) to enable the source domain image to be converted to the target domain image, and the corresponding points of the source domain image and the target domain image are completely consistent in spatial position, and then the matching and three-dimensional modeling can be realized based on the target domain image and the source domain image, which significantly reduces the data acquisition cost, significantly improves the implementation efficiency under the premise of ensuring the modeling accuracy. Moreover, the method is designed with double branches, and the global features and the local features are extracted in multiple scales in parallel, which can fully exploit the correlation between images of different domains, effectively combine the advantages of the two structures to improve the accuracy of converting the source domain image to the target domain image. The method effectively solves the model lag problem caused by the low frequency of traditional radar image acquisition, and avoids the defect of missing depth information when directly using optical images, providing reliable three-dimensional visualization support for real-time monitoring of construction sites.
[0089] Further, the process of fusing the global feature of the target scale and the local feature of the target scale in step S120 to obtain the fused feature of the target scale includes steps S210 to S240, which are as follows:
[0090] In step S210, the global feature of the target scale is subjected to global average pooling and global maximum pooling to obtain a first pooled feature, and the global feature of the target scale is subjected to global maximum pooling to obtain a second pooled feature. The global feature of the target scale is any one of the global features of the multiple scales, and the local feature of the target scale is the local feature of the same scale as the global feature of the target scale among the multiple local features;
[0091] The local feature of the target scale is subjected to global average pooling to obtain a third pooled feature, and the local feature of the target scale is subjected to global maximum pooling to obtain a fourth pooled feature;
[0092] The first pooled feature and the second pooled feature are respectively taken as the head and the tail of the local feature of the target scale, and are spliced to obtain a first spliced feature, and the third pooled feature and the fourth pooled feature are respectively taken as the head and the tail of the global feature of the target scale, and are spliced to obtain a second spliced feature;
[0093] The first spliced feature and the second spliced feature are respectively input into a Transformer unit, and the features output by the Transformer unit are spliced to obtain a first intermediate feature.
[0094] In step S220, the global feature of the target scale is input into an attention enhancement unit to obtain a second intermediate feature output by the attention enhancement unit.
[0095] In step S230, the global feature of the target scale and the local feature of the target scale are multiplied pixel by pixel to obtain a third intermediate feature.
[0096] In step S240, the first intermediate feature, the second intermediate feature, and the third intermediate feature are spliced to obtain the fused feature of the target scale. The attention enhancement unit is a CBAM convolution block attention unit.
[0097] In step S210 of the present embodiment, the processing procedure includes:
[0098] (1) first, the global feature of the target scale is globally averaged and globally maximized to obtain the first pooled feature, and the global feature of the target scale is globally maximized to obtain the second pooled feature; the global average pooling refers to taking the average value of all pixels of the feature map to extract global information, which can be specifically implemented by average calculation across the entire spatial dimension, and is used to eliminate local noise interference. The global maximum pooling refers to taking the maximum value of all pixels of the feature map to capture significant features, which can be specifically implemented by maximum value calculation across the entire spatial dimension, and is used to retain key region information.
[0099] (2) the local feature of the target scale is globally averaged to obtain the third pooled feature, and the local feature of the target scale is globally maximized to obtain the fourth pooled feature.
[0100] (3) in a cross manner, they are added to the head and tail of the other feature to generate the first spliced feature and the second spliced feature.
[0101] (4) then the first spliced feature and the second spliced feature are input into the Transformer unit respectively, and the features output by the Transformer unit are spliced to obtain the first intermediate feature, which establishes the interaction and fusion of global and local information, and improves the expression ability of the feature.
[0102] In step S230 of the embodiment, the processing flow includes:
[0103] Because the synthetic aperture radar image is generated based on the optical image, and the target domain image is a synthetic aperture radar image, there is a large domain difference between the two, therefore, in order to improve the cross-domain correlation between the synthetic aperture radar image and the synthetic aperture radar image, and improve the consistency of feature processing between the synthetic aperture radar image and the synthetic aperture radar image, here the global feature of the target scale and the local feature of the target scale are multiplied pixel by pixel to obtain the third intermediate feature, which can highlight the correlation of significant features and suppress irrelevant information.
[0104] Among them, the pixel-by-pixel multiplication refers to multiplying the pixels at the corresponding positions of the two feature maps, which can be specifically implemented by element-level multiplication operation.
[0105] In step S220 of the embodiment, the processing flow includes:
[0106] The global feature of the target scale is input into the attention enhancement unit to obtain the second intermediate feature output by the attention enhancement unit, and through the enhancement of the attention mechanism, the model can better capture the feature correlation between images.
[0107] Among them, the attention enhancement unit refers to the weighting processing of the feature through the channel or spatial attention mechanism, which can be implemented by CBAM convolution block attention unit, and is used to enhance the weight of important feature region. Specifically, in the fusion process of global features and local features, the CBAM convolution block attention unit is applied to the processing of global features. First, the global features pass through the channel attention module, and the channel weight coefficient is calculated and applied to the input features, so as to enhance the response of important channels; then, the spatial attention module further processes the channel weighted features, and adjusts the contribution degree of different spatial positions through the spatial weight coefficient. By executing the channel and spatial attention mechanisms in series, the detailed information related to image registration in the global features is effectively strengthened, for example, the feature weight of high-frequency edge or texture region is dynamically enhanced, and the weight of background or noise region is suppressed. Therefore, the fused features can more accurately reflect the structural changes of the construction area, and reduce the registration error caused by feature redundancy or information loss.
[0108] Compared with the prior art, the method can effectively improve the accuracy of the fusion of global features and local features, so that the generation of the affine transformation matrix is more in line with the image change law of the actual construction scene. Therefore, the feature matching error in the image registration process is reduced, and the finally generated spatial point cloud data more accurately reflects the three-dimensional structure of the construction area, providing a reliable basis for dynamic monitoring and construction planning.
[0109] The application further proposes a process of fusing target scale global features and target scale local features to obtain target scale fused features, generating multi-dimensional feature expression by combining average pooling and maximum pooling, establishing cross-modal association by using splicing operation combined with Transformer unit, introducing attention mechanism to strengthen feature weight distribution, and enhancing local consistency by pixel-by-pixel multiplication, which significantly improves the comprehensiveness and accuracy of feature fusion.
[0110] Through the above technical solutions, the method effectively solves the problem of insufficient fusion accuracy caused by the feature difference between SAR (synthetic aperture radar) and optical images in the three-dimensional modeling process of the construction area, can fully extract and fuse multi-scale global structure information and local detail features, enhance the feature weight of key areas, and thus improve the quality of spatial point cloud generation, providing a more accurate data basis for three-dimensional visualization.
[0111] As Figure 3 , further, the process of extracting local features by the CNN unit includes:
[0112] ;
[0113] ;
[0114] ;
[0115] ;
[0116] ;
[0117] ;
[0118] ;
[0119] ;
[0120] in, For the input data of the CNN unit, for convolutional layers, for convolutional layers, for Convolutional layers, for Convolutional layers, for convolutional layers, for convolutional layers, The SiLU activation function is used. For channel max pooling, For channel average pooling, for Convolutional layers, It is the sigmoid activation function. For the mapping function of element rearrangement operation, It is a multilayer perceptron. Local features extracted for CNN units.
[0121] In this step, the input data is first fed in parallel into three convolutional layers with different kernel sizes, such as 3×3, 5×5, and 7×7 kernels, to extract local features from different receptive fields. The features from each branch after convolution are introduced with non-linearity through the SiLU activation function. Then, channel max pooling and average pooling operations are performed on the activation results to aggregate channel dimensionality information. The pooling results are then fed into a shared 1×1 convolutional layer to generate channel attention weights. These weights are normalized to the 0-1 range using the sigmoid function and then weighted channel-wise with the original convolutional output to enhance features. The weighted features are then rearranged element-wise to adjust the spatial and channel dimensional relationships, for example, mapping the positional information of the height and width dimensions to the channel dimensions. Finally, a multilayer perceptron performs dimensional transformation and non-linear mapping to output local features with strong representational capabilities.
[0122] Compared with the prior art, the method solves the problem of insufficient adaptability of the traditional feature extraction method in the scenario of dynamic change of the construction area, enhances the robustness to complex textures, moving objects and illumination changes through multi-scale feature fusion and attention mechanism, provides a high-precision local feature basis for subsequent image registration and three-dimensional reconstruction, and thus guarantees the accuracy and real-time performance of the three-dimensional visualization model of the construction area.
[0123] Further, before fusing each scale of the global features in the plurality of scales of global features and the corresponding scale of the local features in the plurality of scales of local features to obtain the plurality of scales of fusion features, the method further comprises the following steps:
[0124] calculating a difference image between the initial synthetic aperture radar image and the simulated synthetic aperture radar image based on an LR operator;
[0125] extracting a plurality of scales of difference features of the difference image;
[0126] fusing each scale of the global features in the plurality of scales of global features and the corresponding scale of the local features in the plurality of scales of local features to obtain the plurality of scales of fusion features, comprising:
[0127] fusing each scale of the global features in the plurality of scales of global features, the corresponding scale of the difference features in the plurality of scales of difference features, and the corresponding scale of the local features in the plurality of scales of local features to obtain the plurality of scales of fusion features.
[0128] In this step, the LR operator refers to a difference calculation method based on low-rank decomposition, and specifically, a matrix decomposition technique can be used to map the pixel difference between the two images into a low-rank matrix, thereby capturing the structured difference between the images. The difference image refers to a two-dimensional image generated by processing the LR operator, reflecting the pixel-level difference between the initial synthetic aperture radar image and the simulated synthetic aperture radar image.
[0129] Specifically, after the initial synthetic aperture radar image and the simulated synthetic aperture radar image are spliced into the input image, the LR operator is used to perform low-rank decomposition on the pixel difference between the two original images to generate a difference image. The difference image is subjected to a multi-level downsampling operation to extract difference features of different scales, such as using a three-layer convolutional network to extract feature maps at resolutions of 1 / 2, 1 / 4 and 1 / 8. In the feature fusion stage, each scale of the global features, the corresponding scale of the local features and the corresponding scale of the difference features are combined by channel splicing, such as superimposing the semantic information of the global features, the detailed information of the local features and the change information of the difference features, and then generating fusion features through convolution operation. This process makes the fusion features not only contain the spatial characteristics of the original image, but also strengthen the difference information between the images, thereby improving the estimation accuracy of the subsequent affine transformation matrix.
[0130] Compared with the prior art, the method can effectively improve the registration accuracy of three-dimensional point cloud reconstruction of the construction area, and in particular in the scene of light change or dynamic object interference, the affine transformation estimation error can be reduced through the feature fusion process enhanced by the differential features, so that a more accurate three-dimensional visualization model is generated.
[0131] Reference Figure 4 In one embodiment, a three-dimensional visualization device for a construction area is provided, and the device comprises:
[0132] The image acquisition module 1100 is configured to acquire an initial synthetic aperture radar image and a current optical image of the construction area, and to perform Fourier transformation on the current optical image to obtain a pseudo-synthetic aperture radar image;
[0133] The feature extraction module 1200 is configured to splice the initial synthetic aperture radar image and the pseudo-synthetic aperture radar image as a to-be-input image, to extract a plurality of scales of global features and a plurality of scales of local features from the to-be-input image, to fuse each scale of global features in the plurality of scales of global features and the corresponding scale of local features in the plurality of scales of local features, to obtain a plurality of scales of fusion features, and to perform up-sampling fusion on the plurality of scales of fusion features to generate a final feature,
[0134] The image affine transformation module 1300 is configured to generate an affine transformation matrix according to the final feature by using a multi-layer perception machine, and to convert the pseudo-synthetic aperture radar image into a warped synthetic aperture radar image according to the affine transformation matrix;
[0135] The three-dimensional visualization module 1400 is configured to generate a spatial point cloud of the construction area according to the warped synthetic aperture radar image and the initial synthetic aperture radar image, so as to realize three-dimensional visualization of the construction area.
[0136] It should be noted that the three-dimensional visualization device for the construction area provided in the embodiment is based on the same inventive concept as the three-dimensional visualization method for the construction area described above, and therefore the related content of the three-dimensional visualization method for the construction area described above is also applicable to the content of the three-dimensional visualization device for the construction area, and therefore, the details are not described herein.
[0137] As Figure 5 The embodiment of the present application further provides an electronic device, and the electronic device comprises:
[0138] at least one memory;
[0139] at least one processor;
[0140] at least one program;
[0141] The program is stored in the memory, and the processor executes the at least one program to implement the three-dimensional visualization method for the construction area described in the present disclosure.
[0142] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, and the like.
[0143] The electronic device of the embodiment of the present application is described in detail below.
[0144] The processor 1600 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiment of the present application.
[0145] The memory 1700 can be implemented in a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store an operating system and other application programs. When the technical solutions provided by the embodiment of the present application are implemented by software or firmware, the related program codes are stored in the memory 1700 and are called and executed by the processor 1600 to implement the three-dimensional visualization method of the construction area.
[0146] The input / output interface 1800 is configured to implement information input and output.
[0147] The communication interface 1900 is configured to implement communication interaction between the device and other devices. The communication can be implemented in a wired manner (for example, a USB, a network cable, and the like) or in a wireless manner (for example, a mobile network, WIFI, Bluetooth, and the like).
[0148] The bus 2000 is configured to transmit information between various components (for example, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900) of the device.
[0149] The processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are connected to each other in the device through the bus 2000.
[0150] The embodiment of the present application further provides a storage medium, which is a computer readable storage medium and stores computer executable instructions. The computer executable instructions are configured to enable a computer to execute the three-dimensional visualization method of the construction area.
[0151] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory that is remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0152] The embodiments described in the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0153] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0154] The device embodiments described above are only schematic, and units described as separate components can or can not be physically separate, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0155] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.
[0156] The terms "first", "second", "third", "fourth" and the like used in the description of the present application and the above-described drawings (if any) are used to distinguish similar objects, and do not necessarily have to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0157] It should be understood that, in the application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases of only A, only B and A and B existing at the same time, wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent a, b, c, "a and b", "a and c", "b and c", or "a and b and c", wherein a, b and c can be single or multiple.
[0158] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0159] The units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0160] In addition, the functional units in each embodiment of the application can be integrated into a processing unit, or each unit can be physically present, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0161] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program storage media.
[0162] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above implementation. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the embodiments of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the embodiments of the present application.
Claims
1. A three-dimensional visualization method for a construction area, characterized in that, The method includes: Acquire the initial synthetic aperture radar image and the current optical image of the construction area, and perform Fourier transform on the current optical image to simulate a synthetic aperture radar image; The initial synthetic aperture radar image and the simulated synthetic aperture radar image are stitched together as the input image. Global features and local features at multiple scales are extracted from the input image. The global features at each scale and the local features at the corresponding scale in the global features at multiple scales are fused to obtain fused features at multiple scales. The fused features at multiple scales are then upsampled and fused to generate the final features. The final features are used to generate an affine transformation matrix based on the multilayer perceptron, and the affine synthetic aperture radar image is converted into a distorted synthetic aperture radar image based on the affine transformation matrix. Based on the distorted synthetic aperture radar image and the initial synthetic aperture radar image, a spatial point cloud of the construction area is generated to achieve three-dimensional visualization of the construction area.
2. The three-dimensional visualization method for the construction area according to claim 1, characterized in that, The step of extracting global features and local features at multiple scales from the input image includes: The image to be input is input into the first branch network to obtain global features at multiple scales output by the first branch network; wherein, the first branch network includes multiple cascaded Transformer units and a downsampling unit is included between every two Transformer units; The image to be input is input into the second branch network to obtain local features at multiple scales output by the second branch network; wherein the second branch network includes multiple cascaded CNN units and a downsampling unit is included between every two CNN units.
3. The three-dimensional visualization method for the construction area according to claim 2, characterized in that, The process of fusing global and local features at the target scale to obtain fused features at the target scale includes: The global features at the target scale are subjected to global average pooling and global max pooling to obtain the first pooled features. The global features at the target scale are then subjected to global max pooling to obtain the second pooled features. The global features at the target scale are any one of the global features at the multiple scales, and the local features at the target scale are local features at the same scale as the global features at the target scale. The local features at the target scale are subjected to global average pooling to obtain the third pooling feature, and the local features at the target scale are subjected to global max pooling to obtain the fourth pooling feature. The first pooling feature and the second pooling feature are respectively used as the head and tail of the local feature at the target scale, and are concatenated to obtain the first concatenated feature. The third pooling feature and the fourth pooling feature are respectively used as the head and tail of the global feature at the target scale, and are concatenated to obtain the second concatenated feature. The first and second splicing features are respectively input into the Transformer unit, and the features output by the Transformer unit are spliced together to obtain the first intermediate feature; The global features at the target scale are input into the attention enhancement unit to obtain the second intermediate features output by the attention enhancement unit; The third intermediate feature is obtained by multiplying the global features and local features of the target scale pixel by pixel. The first intermediate feature, the second intermediate feature, and the third intermediate feature are concatenated to obtain the fused feature at the target scale.
4. The three-dimensional visualization method for the construction area according to claim 3, characterized in that, The attention enhancement unit is a CBAM convolutional block attention unit.
5. The three-dimensional visualization method for the construction area according to claim 2, characterized in that, The process of extracting local features by the CNN unit includes: ; ; ; ; ; ; ; ; in, The input data for the CNN unit, for convolutional layers, for convolutional layers, for convolutional layers, for convolutional layers, for convolutional layers, for convolutional layers, The SiLU activation function is used. For channel max pooling, For channel average pooling, for convolutional layers, It is the sigmoid activation function. For the mapping function of element rearrangement operation, It is a multilayer perceptron. The local features extracted by the CNN unit.
6. The three-dimensional visualization method for the construction area according to claim 1, characterized in that, Before fusing the global features at each scale of the multiple scales and the corresponding local features at the local scales to obtain the fused features at multiple scales, the method further includes: The difference image between the initial synthetic aperture radar image and the simulated synthetic aperture radar image is calculated based on the LR operator; Extract the difference features at multiple scales from the difference image; The step of fusing the global features of each scale from the global features of the multiple scales and the corresponding local features of the local features of the multiple scales to obtain the fused features of the multiple scales includes: The global features of each scale in the global features of the multiple scales, the difference features of the corresponding scale in the difference features of the multiple scales, and the local features of the corresponding scale in the local features of the multiple scales are fused to obtain the fused features of the multiple scales.
7. A three-dimensional visualization device for a construction area, characterized in that, The device includes: The image acquisition module is used to acquire the initial synthetic aperture radar image and the current optical image of the construction area, and to Fourier transform the current optical image into a simulated synthetic aperture radar image. The feature extraction module is used to stitch the initial synthetic aperture radar image and the simulated synthetic aperture radar image together as the input image, extract global features and local features at multiple scales from the input image, fuse the global features at each scale and the local features at the corresponding scale in the global features at multiple scales to obtain fused features at multiple scales, and upsample and fuse the fused features at multiple scales to generate the final features. An image affine transformation module is used to generate an affine transformation matrix from the final features based on a multilayer perceptron, and to convert the affine synthetic aperture radar image into a distorted synthetic aperture radar image based on the affine transformation matrix. The 3D visualization module is used to generate a spatial point cloud of the construction area based on the distorted synthetic aperture radar image and the initial synthetic aperture radar image, so as to realize the 3D visualization of the construction area.
8. The three-dimensional visualization device for the construction area according to claim 7, characterized in that, The device further includes: The difference feature extraction module is used to calculate the difference image between the initial synthetic aperture radar image and the simulated synthetic aperture radar image based on the LR operator; and to extract the difference features at multiple scales of the difference image. The feature extraction module is specifically used to fuse the global features of each scale in the global features of the multiple scales, the difference features of the corresponding scale in the difference features of the multiple scales, and the local features of the corresponding scale in the local features of the multiple scales to obtain the fused features of the multiple scales.
9. An electronic device, characterized in that, include: At least one control processor and a memory for communicatively connecting to the at least one control processor; The memory stores instructions that can be executed by the at least one control processor to enable the at least one control processor to perform the three-dimensional visualization method of the construction area according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the three-dimensional visualization method of the construction area as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-modal remote sensing image matching method for coupling phase structure and depth feature
CN118968249A