Three-dimensional visualization method and device for construction area

By converting optical images into simulated synthetic aperture radar images and generating affine transformation matrices, the problems of high acquisition cost and insufficient information in 3D modeling of construction areas are solved, and efficient and accurate 3D visualization support is achieved.

CN120612427AActive Publication Date: 2025-09-09CHINA RAILWAY 25TH BUREAU GRP +1
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510710806.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The existing technology for 3D modeling of construction areas has the following problems: synthetic aperture radar image acquisition is time-consuming and expensive, making it difficult to meet high-frequency monitoring needs; and optical images lack depth information and cannot be directly used for 3D modeling.

Method used

By acquiring synthetic aperture radar images and optical images of the construction area, the optical image is converted into a pseudo-synthetic aperture radar image using Fourier transform, and the affine transformation matrix is ​​generated through a multi-layer perceptron to achieve matching between images and three-dimensional visualization.

Benefits of technology

It significantly reduces data acquisition costs, improves 3D modeling efficiency and accuracy, and provides reliable 3D visualization support for real-time monitoring of construction sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612427A_ABST
    Figure CN120612427A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional visualization method and device for a construction area, and the method comprises the steps: enabling a synthetic aperture radar image to serve as a target domain image, introducing an optical image to serve as a source domain image, searching an affine transformation matrix, enabling the source domain image to be converted into the target domain image, and enabling corresponding points to be completely consistent in spatial position, the matching between the target domain image and the source domain image is realized based on the affine transformation matrix, and then three-dimensional modeling is executed, so that the data acquisition cost is remarkably reduced, and the implementation efficiency is remarkably improved on the premise of ensuring the modeling precision. According to the method, through the double-branch design, the global features and the local features of the image formed by splicing the target domain image and the source domain image are extracted in parallel on multiple scales, the correlation between different domain images can be fully mined, the accuracy of converting the source domain image into the target domain image can be improved, and the accuracy of converting the source domain image into the target domain image is improved. And reliable three-dimensional visualization support is provided for real-time monitoring of a construction site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of three-dimensional visualization technology, and in particular to a three-dimensional visualization method and device for a construction area. Background Art

[0002] Constructing a three-dimensional model of the construction area can systematically improve the construction efficiency, safety and sustainability of the construction area, and has high practical value.

[0003] Currently, constructing a 3D model of a construction area primarily involves acquiring 2D SAR images of the area using synthetic aperture radar (SAR). De-noising, registration, and fusion operations are then performed on the 2D SAR images to generate the 3D model. However, construction areas are subject to high levels of human and vehicle mobility, necessitating 24 / 7 image acquisition and monitoring.

[0004] SAR-based two-dimensional image acquisition has two major drawbacks: on the one hand, the SAR image acquisition process is time-consuming and expensive, making it difficult to meet high-frequency monitoring needs; on the other hand, although traditional optical images are easy to acquire, they lack depth information and cannot be directly used for three-dimensional modeling. Summary of the Invention

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] The main purpose of the embodiments of the present application is to propose a three-dimensional visualization method and device for construction areas, which can fully explore the correlation between images in different domains, improve the accuracy of converting source domain images to target domain images, and provide reliable three-dimensional visualization support for real-time monitoring of construction sites.

[0007] To achieve the above-mentioned objectives, a first aspect of an embodiment of the present application provides a three-dimensional visualization method for a construction area, the method comprising: Acquire an initial synthetic aperture radar image and a current optical image of the construction area, and perform a Fourier transform on the current optical image to obtain a simulated synthetic aperture radar image; splicing the initial synthetic aperture radar image and the simulated synthetic aperture radar image as an input image, extracting global features at multiple scales and local features at multiple scales from the input image, fusing the global features at each scale of the multiple global features with the local features at corresponding scales of the multiple local features to obtain fused features at multiple scales, and upsampling and fusing the fused features at multiple scales to generate final features; generating an affine transformation matrix from the final features according to a multi-layer perceptron, and converting the simulated synthetic aperture radar image into a distorted synthetic aperture radar image according to the affine transformation matrix; A spatial point cloud of the construction area is generated according to the distorted synthetic aperture radar image and the initial synthetic aperture radar image, so as to achieve three-dimensional visualization of the construction area.

[0008] The present application provides a three-dimensional visualization method for a construction area, which has at least the following beneficial effects: This method uses a synthetic aperture radar image as the target domain image, introduces an optical image as the source domain image, and searches for an affine transformation matrix that allows the source domain image to be converted to the target domain image so that corresponding points are completely consistent in spatial position. This method then matches the target domain image with the source domain image based on the affine transformation matrix, and then performs three-dimensional modeling. This significantly reduces data acquisition costs and significantly improves implementation efficiency while ensuring modeling accuracy. Furthermore, through a dual-branch design, this method extracts global and local features of the spliced ​​image of the target domain image and the source domain image at multiple scales in parallel, fully exploring the correlation between images in different domains and improving the accuracy of converting the source domain image to the target domain image, providing reliable three-dimensional visualization support for real-time monitoring of construction sites.

[0009] In some embodiments, extracting global features of multiple scales and local features of multiple scales from the input image includes: Inputting the image to be input into a first branch network to obtain global features of multiple scales output by the first branch network; wherein the first branch network includes a plurality of cascaded Transformer units and a downsampling unit is included between every two Transformer units; The image to be input is input into the second branch network to obtain local features of multiple scales output by the second branch network; wherein the second branch network includes multiple cascaded CNN units and a downsampling unit is included between every two CNN units.

[0010] In some embodiments, the process of fusing the global features of the target scale with the local features of the target scale to obtain the fused features of the target scale includes: Performing global average pooling and global maximum pooling on the global features of the target scale to obtain a first pooled feature, and performing global maximum pooling on the global features of the target scale to obtain a second pooled feature; wherein the global feature of the target scale is a global feature of any scale among the global features of the multiple scales, and the local feature of the target scale is a local feature of the same scale as the global feature of the target scale among the local features of the multiple scales; Performing global average pooling on the local features of the target scale to obtain a third pooling feature, and performing global maximum pooling on the local features of the target scale to obtain a fourth pooling feature; The first pooled feature and the second pooled feature are respectively used as the head and tail of the local feature of the target scale, and are concatenated to obtain a first concatenated feature. The third pooled feature and the fourth pooled feature are respectively used as the head and tail of the global feature of the target scale, and are concatenated to obtain a second concatenated feature. Inputting the first splicing feature and the second splicing feature into the Transformer unit respectively, and splicing the features respectively output by the Transformer unit to obtain a first intermediate feature; Inputting the global feature of the target scale into an attention enhancement unit to obtain a second intermediate feature output by the attention enhancement unit; Multiplying the global feature of the target scale and the local feature of the target scale pixel by pixel to obtain a third intermediate feature; The first intermediate feature, the second intermediate feature, and the third intermediate feature are concatenated to obtain a fused feature of the target scale.

[0011] In some embodiments, the attention enhancement unit is a CBAM convolutional block attention unit.

[0012] In some embodiments, the process of extracting local features by the CNN unit includes: ; ; ; ; ; ; ; ; in, is the input data of the CNN unit, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, is the SiLU activation function, is the channel maximum pooling, is channel average pooling, for The convolutional layer, is the sigmoid activation function, is the mapping function for element rearrangement operation, is a multi-layer perceptron, It is the local feature extracted by the CNN unit.

[0013] In some embodiments, before fusing the global features of each scale among the global features of the multiple scales with the local features of the corresponding scale among the local features of the multiple scales to obtain the fused features of the multiple scales, the method further includes: Calculating a difference image between the initial synthetic aperture radar image and the simulated synthetic aperture radar image based on an LR operator; extracting difference features of multiple scales of the difference image; The fusing of the global features of each scale in the global features of the multiple scales with the local features of the corresponding scale in the local features of the multiple scales to obtain fused features of the multiple scales includes: The global features of each scale among the global features of the multiple scales, the difference features of the corresponding scale among the difference features of the multiple scales, and the local features of the corresponding scale among the local features of the multiple scales are fused to obtain fused features of the multiple scales.

[0014] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a three-dimensional visualization device for a construction area, the device comprising: An image acquisition module is used to acquire an initial synthetic aperture radar image and a current optical image of the construction area, and Fourier transform the current optical image into a simulated synthetic aperture radar image; a feature extraction module, configured to splice the initial synthetic aperture radar image and the simulated synthetic aperture radar image as an input image, extract global features at multiple scales and local features at multiple scales from the input image, fuse the global features at each scale of the multiple global features with the local features at corresponding scales of the multiple local features to obtain fused features at multiple scales, and upsample and fuse the fused features at multiple scales to generate final features; an image affine transformation module, configured to generate an affine transformation matrix from the final features according to a multi-layer perceptron, and convert the simulated synthetic aperture radar image into a distorted synthetic aperture radar image according to the affine transformation matrix; A three-dimensional visualization module is used to generate a spatial point cloud of the construction area based on the distorted synthetic aperture radar image and the initial synthetic aperture radar image, so as to achieve three-dimensional visualization of the construction area.

[0015] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application provides an electronic device, comprising: at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the above-mentioned three-dimensional visualization method of the construction area.

[0016] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned three-dimensional visualization method of the construction area.

[0017] It can be understood that the beneficial effects of the second to fourth aspects compared with the relevant technologies are the same as the beneficial effects of the first aspect compared with the relevant technologies. Please refer to the relevant description in the first aspect and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 This is a flow chart of a three-dimensional visualization method for a construction area provided by one embodiment of the present application; Figure 2 This is a schematic diagram of the structure of a model network provided by an embodiment of the present application; Figure 3 This is a schematic diagram of the unit structure of a CNN provided by one embodiment of the present application; Figure 4 This is a schematic diagram of the structure of a three-dimensional visualization system for a construction area provided by an embodiment of the present application; Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0021] like Figure 1One embodiment of the present application provides a three-dimensional visualization method for a construction area, the method comprising the following steps S110 to S140: Step S110 , obtaining an initial synthetic aperture radar image and a current optical image of the construction area, and performing Fourier transform on the current optical image to obtain a simulated synthetic aperture radar image.

[0022] In this step, the construction area refers to the area that requires 3D visualization. The initial SAR image is a SAR image acquired using SAR of the construction area. This application uses the initial SAR image as the target domain image, and subsequent all-weather optical images as the source domain images. Both source domain images must be matched with the target domain image to ensure that the optical image meets the SAR image standard for constructing a 3D point cloud.

[0023] Moreover, the purpose of this application is to match two or more images (source domain images and target domain images) of the same scene taken at different times, from different optical sensors or from different perspectives, and to make the corresponding points of the two images completely consistent in spatial position by finding the spatial transformation rules (affine transformation matrix).

[0024] Since synthetic aperture radar image acquisition is expensive, we consider using all-weather optical image acquisition and transforming the acquired optical image into a simulated synthetic aperture radar image based on Fourier transform to facilitate the subsequent matching process. The following briefly describes the process of constructing a simulated synthetic aperture radar image: For input size Image , the Fourier transform is: ; in, is the image after Fourier transform, is a variable in the frequency domain, is a variable in the spatial domain.

[0025] The central area of ​​the image after Fourier transform is the low-frequency component, and the rest of the image is the high-frequency component.

[0026] A mask can be used Assist in the change of low-frequency areas to ensure that only the area near the image center (0,0) is replaced: ; Where: is the ratio of the width and height of the central area of ​​the image to the width and height of the entire image area, where The value is 0.01.

[0027] The central low-frequency component of the optical image in the source domain image is replaced by the low-frequency component of the SAR image in the corresponding target domain. Then the frequency domain image is converted into a spatial domain image by inverse Fourier transform. The calculation formula for generating a pseudo SAR image by inverse Fourier transform is: ; Where: is a pseudo synthetic aperture radar image generated from the source domain optical image; 、 are the optical image of the source domain and the corresponding synthetic aperture radar image; 、 They correspond to the amplitude and phase of the image respectively; It is an element-wise multiplication operation.

[0028] This embodiment treats the initial synthetic aperture radar as the target domain image, the current optical image as the source domain image, performs a Fourier transform on the source domain image to convert it into a simulated synthetic aperture radar image, and finally finds a spatial transformation rule (affine transformation matrix) to ensure that the corresponding points of the source domain image and the target domain image are completely consistent in spatial position.

[0029] In step S120, the initial synthetic aperture radar image and the simulated synthetic aperture radar image are spliced ​​together as an input image, global features at multiple scales and local features at multiple scales are extracted from the input image, global features at each scale in the multiple scales are fused with local features at corresponding scales in the multiple scales to obtain fused features at multiple scales, and the fused features at multiple scales are upsampled and fused to generate final features.

[0030] like Figure 2 In this step, the initial synthetic aperture radar image and the simulated synthetic aperture radar image are spliced ​​into an image to be input, which serves as input data for the model. The model here refers to a model that can realize the generation of the affine transformation matrix in steps S120 and S130. Its main function is to extract global features and local features from the image to be input, fuse them based on the global features and the local features, and finally generate an affine transformation matrix based on the final features using a multi-layer perceptron.

[0031] Furthermore, the model includes: (1) The first branch network includes 4 cascaded Transformer units and a downsampling unit is included between each two Transformer units. The Transformer unit is used to extract the global features of the image at the corresponding scale. The downsampling unit is mainly used to downsample the global features of the image. The Transformer unit refers to a neural network structure based on the self-attention mechanism. Specifically, it can be implemented by a module including a multi-head attention layer and a feedforward neural network to capture long-range spatial dependencies in the image. The downsampling unit refers to an operation module that reduces the spatial resolution of the feature map. Specifically, it can be implemented by a convolution layer or a pooling layer with a step size greater than 1. By gradually reducing the size of the feature map, the receptive field is expanded and the amount of calculation is reduced. (2) The second branch network, the first branch network includes 4 cascaded CNN units and a downsampling unit is included between each two CNN units. The CNN unit is used to extract local image features of the corresponding scale. The downsampling unit is mainly used to downsample the local features of the image. The CNN unit refers to a convolutional neural network structure, which can be implemented by combining stacked convolutional layers and nonlinear activation functions to extract texture and edge features of local image regions.

[0032] (3) The first fusion unit mainly realizes the fusion of the global features of each scale in the global features of multiple scales and the local features of the corresponding scale in the local features of multiple scales.

[0033] (4) The second fusion unit can be composed of 4 Transformer units and 4 upsampling units, which is mainly used to upsample and fuse the fusion features of multiple scales to generate the final features.

[0034] (5) Multilayer perceptron, mainly used to generate affine transformation matrix from the final features.

[0035] Furthermore, the input image is processed in the first branch network through cascaded Transformer units. Each Transformer unit models global contextual information through a self-attention mechanism, while the downsampling unit compresses the spatial resolution of the feature map, ultimately outputting global features at multiple scales. The second branch network extracts local features of different scales through cascaded CNN units. The CNN unit focuses on information aggregation in the local neighborhood during convolution operations, while the downsampling unit simultaneously adjusts the size of the feature map. The multi-scale global features output by the two networks complement the local features. The global features represent the overall structure of the construction area, while the local features retain construction details, such as equipment outlines or material textures.

[0036] Compared to existing technologies, traditional methods typically use either CNN or Transformer architectures alone to extract features, making it difficult to balance global information with local details. For example, a single CNN network may lose long-range correlation information in deep layers, while a Transformer-only architecture is limited in computational efficiency and ability to capture local features. This dual-branch design allows for the parallel extraction of global and local features at multiple scales, effectively combining the strengths of both architectures.

[0037] In order to improve the matching degree between the source domain image (generated by visible light image) and the target domain image (synthetic aperture radar image), this method effectively combines the advantages of the two structures through the collaborative extraction of multi-scale global and local features, improves the accuracy of the final feature extraction, and ultimately improves the accuracy of finding the affine transformation matrix between the source domain image and the target domain image.

[0038] Step S130 : generating an affine transformation matrix from the final features according to the multi-layer perceptron, and converting the simulated SAR image into a distorted SAR image according to the affine transformation matrix.

[0039] In this step, the multilayer perceptron is able to map the final features into the spatial transformation law between the source domain image and the target domain image - the affine transformation matrix. Then, based on the affine transformation matrix, the simulated SAR image can be converted into a distorted SAR image. Based on the affine transformation matrix and the simulated SAR image, a distorted SAR image is calculated. Here, the corresponding points of the distorted SAR image and the target domain image are completely consistent in spatial position, thus achieving matching between the distorted SAR image and the target domain image.

[0040] Step S140 : generating a spatial point cloud of the construction area according to the distorted SAR image and the initial SAR image, so as to achieve three-dimensional visualization of the construction area.

[0041] In this step, after a distorted SAR image is generated from the source image, the corresponding points in the distorted SAR image and the target image are completely spatially aligned. SAR image matching can be achieved based on the distorted SAR image and the initial SAR image, and a 3D point cloud can be constructed based on the images. This will not be described in detail here; constructing a 3D point cloud based on registered SAR images is a conventional technique in the art.

[0042] This method has at least the following beneficial effects: Traditional methods rely on a single radar image data source and require all-weather acquisition. However, this method utilizes synthetic aperture radar images as target domain images and introduces optical images as source domain images. This method also incorporates an optical image conversion mechanism, specifically finding spatial transformation patterns (affine transformation matrices) that enable the source domain image to be converted to the target domain image, ensuring that corresponding points in the source and target domain images are perfectly aligned in spatial position. This allows for matching and 3D modeling based on the target and source domain images, significantly reducing data acquisition costs and improving implementation efficiency while ensuring modeling accuracy. Furthermore, this method utilizes a dual-branch design to extract global and local features simultaneously at multiple scales, fully exploiting the correlations between images in different domains. This effectively combines the advantages of both architectures to enhance the accuracy of the source-to-target domain conversion. This method effectively addresses the model lag caused by the low acquisition frequency of traditional radar images while avoiding the lack of depth information associated with direct use of optical images, providing reliable 3D visualization support for real-time monitoring of construction sites.

[0043] Furthermore, the process of fusing the global features of the target scale and the local features of the target scale in step S120 to obtain the fused features of the target scale includes the following steps S210 to S240, wherein: Step S210: Perform global average pooling and global maximum pooling on the global features of the target scale to obtain a first pooled feature, and perform global maximum pooling on the global features of the target scale to obtain a second pooled feature. The global features of the target scale are global features of any scale among the global features of multiple scales, and the local features of the target scale are local features of the same scale as the global features of the target scale among the local features of multiple scales. Perform global average pooling on the local features of the target scale to obtain the third pooling feature, and perform global maximum pooling on the local features of the target scale to obtain the fourth pooling feature; The first pooling feature and the second pooling feature are respectively used as the head and tail of the local feature of the target scale, and the first splicing feature is obtained by splicing. The third pooling feature and the fourth pooling feature are respectively used as the head and tail of the global feature of the target scale, and the second splicing feature is obtained by splicing. The first splicing feature and the second splicing feature are respectively input into the Transformer unit, and the features respectively output by the Transformer unit are spliced ​​to obtain the first intermediate feature.

[0044] In step S220 , the global feature of the target scale is input into the attention enhancement unit to obtain the second intermediate feature output by the attention enhancement unit.

[0045] Step S230: Multiply the global feature of the target scale and the local feature of the target scale pixel by pixel to obtain a third intermediate feature.

[0046] In step S240, the first intermediate feature, the second intermediate feature, and the third intermediate feature are concatenated to obtain a fused feature of the target scale. The attention enhancement unit is a CBAM convolutional block attention unit.

[0047] In step S210 of this embodiment, the processing flow includes: (1) First, the global features of the target scale are subjected to global average pooling and global maximum pooling to obtain the first pooled features. The global features of the target scale are subjected to global maximum pooling to obtain the second pooled features. Global average pooling refers to taking the average value of all pixels in the feature map to extract global information. Specifically, it can be implemented by averaging across the entire spatial dimension to eliminate local noise interference. Global maximum pooling refers to taking the maximum value of all pixels in the feature map to capture significant features. Specifically, it can be implemented by maximizing across the entire spatial dimension to retain key area information.

[0048] (2) Perform global average pooling on the local features of the target scale to obtain the third pooling feature, and perform global maximum pooling on the local features of the target scale to obtain the fourth pooling feature.

[0049] (3) Add them to the head and tail of each other's features in a cross manner to generate the first splicing feature and the second splicing feature.

[0050] (4) The first concatenated feature and the second concatenated feature are then input into the Transformer unit respectively, and the features output by the Transformer unit are concatenated to obtain the first intermediate feature. The interaction and fusion of global and local information are established in the first intermediate feature, thereby improving the expressive power of the feature.

[0051] In step S230 of this embodiment, the processing flow includes: Because the simulated SAR image is generated based on the optical image, and the target domain image is a SAR image, there is a large domain difference between the two. Therefore, in order to improve the cross-domain correlation between the simulated SAR image and the SAR image, and to improve the consistency of feature processing between the SAR image and the simulated SAR image, the third intermediate feature is obtained by multiplying the global feature of the target scale and the local feature of the target scale pixel by pixel, which can highlight the correlation of significant features and suppress irrelevant information.

[0052] Among them, pixel-by-pixel multiplication refers to multiplying pixels at corresponding positions of two feature maps, which can be implemented by element-level multiplication operations.

[0053] In step S220 of this embodiment, the processing flow includes: The global features of the target scale are input into the attention enhancement unit to obtain the second intermediate features output by the attention enhancement unit. Through the enhancement of the attention mechanism, the model can better capture the feature correlation between images.

[0054] The attention enhancement unit, specifically implemented as a CBAM convolutional block attention unit, performs weighted feature processing via a channel or spatial attention mechanism. This can be achieved by using a CBAM convolutional block attention unit, which is used to enhance the weights of important feature regions. Specifically, during the fusion of global and local features, the CBAM convolutional block attention unit is applied to global feature processing. First, the global features pass through the channel attention module, where channel weight coefficients are calculated and applied to the input features, thereby enhancing the response of important channels. Subsequently, the spatial attention module further processes the channel-weighted features, adjusting the contributions of different spatial locations using the spatial weight coefficients. By cascading the channel and spatial attention mechanisms, detailed information relevant to image registration within the global features is effectively enhanced. For example, the feature weights of high-frequency edges or texture regions are dynamically enhanced, while the weights of background or noise regions are suppressed. As a result, the fused features can more accurately reflect structural changes in the construction area and reduce registration errors caused by feature redundancy or information loss.

[0055] Compared with existing technologies, this method effectively improves the accuracy of the fusion of global and local features, making the generated affine transformation matrix more consistent with the image variations of actual construction scenes. As a result, feature matching errors during image registration are reduced, and the resulting spatial point cloud data more accurately reflects the three-dimensional structure of the construction area, providing a reliable foundation for dynamic monitoring and construction planning.

[0056] This application further proposes a process of fusing the global features of the target scale and the local features of the target scale to obtain the fused features of the target scale. By combining average pooling and maximum pooling, multi-dimensional feature expression is generated, and cross-modal association is established by using splicing operations combined with Transformer units. The attention mechanism is introduced to strengthen the feature weight distribution, and local consistency is enhanced through pixel-by-pixel multiplication, which significantly improves the comprehensiveness and accuracy of feature fusion.

[0057] Through the above technical solution, this method effectively solves the problem of insufficient fusion accuracy caused by the difference in SAR (synthetic aperture radar) and optical image features during the three-dimensional modeling of the construction area. It can fully extract and fuse multi-scale global structural information and local detail features, enhance the feature weights of key areas, thereby improving the quality of spatial point cloud generation and providing a more accurate data foundation for three-dimensional visualization.

[0058] like Figure 3,Furthermore, the process of extracting local features by CNN unit includes: ; ; ; ; ; ; ; ; in, is the input data of the CNN unit, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, is the SiLU activation function, is the channel maximum pooling, is channel average pooling, for The convolutional layer, is the sigmoid activation function, is the mapping function for element rearrangement operation, is a multi-layer perceptron, Local features extracted by CNN units.

[0059] In this step, the input data is first fed in parallel into three convolutional layers with different kernel sizes. For example, 3×3, 5×5, and 7×7 kernels can be used to extract local features with different receptive fields. The convolutional features of each branch are then processed using the SiLU activation function to introduce nonlinearity. Channel-wise max pooling and average pooling are then performed on the activation results to aggregate channel-wise information. The two pooling results are fed into a shared 1×1 convolutional layer to generate channel-wise attention weights. The weights are normalized to the 0-1 range using the sigmoid function and then weighted channel-by-channel with the original convolution output to achieve feature enhancement. The weighted features are then resized to adjust the spatial and channel-wise relationship through element rearrangement operations, for example, mapping the height and width position information to the channel dimension. Finally, they are transformed and nonlinearly mapped through a multi-layer perceptron to output local features with strong representational capabilities.

[0060] Compared with existing technologies, this method solves the problem of insufficient adaptability of traditional feature extraction methods in dynamically changing construction area scenarios. Through multi-scale feature fusion and attention mechanism, it enhances the robustness to complex textures, moving objects and lighting changes, and provides a high-precision local feature foundation for subsequent image registration and three-dimensional reconstruction, thereby ensuring the accuracy and real-time performance of the three-dimensional visualization model of the construction area.

[0061] Furthermore, before fusing the global features of each scale in the global features of multiple scales with the local features of the corresponding scale in the local features of multiple scales to obtain the fused features of multiple scales, the method further includes the following steps: Calculate the difference image between the initial SAR image and the simulated SAR image based on the LR operator; Extracting difference features at multiple scales of the difference image; The global features of each scale in the global features of multiple scales are fused with the local features of the corresponding scale in the local features of multiple scales to obtain the fused features of multiple scales, including: The global features of each scale in the global features of multiple scales, the difference features of the corresponding scale in the difference features of multiple scales, and the local features of the corresponding scale in the local features of multiple scales are fused to obtain fused features of multiple scales.

[0062] In this step, the LR operator is a difference calculation method based on low-rank decomposition. Specifically, matrix decomposition techniques are used to map the pixel differences between the two images into a low-rank matrix, thereby capturing the structural differences between the images. The difference image is a two-dimensional image generated by the LR operator, reflecting the pixel-level differences between the original SAR image and the simulated SAR image.

[0063] Specifically, after the initial SAR image and the simulated SAR image are spliced ​​together to form the input image, the two original images are subjected to a low-rank decomposition of pixel differences using the LR operator to generate a difference image. This difference image undergoes multiple downsampling operations to extract differential features at different scales. For example, a three-layer convolutional network is used to extract feature maps at 1 / 2, 1 / 4, and 1 / 8 resolutions, respectively. During the feature fusion stage, the global features at each scale, the local features at the corresponding scale, and the differential features at the corresponding scale are combined through channel concatenation. For example, the semantic information of the global features, the detailed information of the local features, and the change information of the differential features are superimposed, and then a convolution operation is performed to generate the fused features. This process ensures that the fused features not only contain the spatial characteristics of the original images but also enhance the difference information between the images, thereby improving the accuracy of the subsequent affine transformation matrix estimation.

[0064] Compared with existing technologies, this method can effectively improve the registration accuracy of three-dimensional point cloud reconstruction in the construction area, especially in scenarios with changing lighting or interference from dynamic objects. The feature fusion process enhanced by differential features can reduce the affine transformation estimation error, thereby generating a more accurate three-dimensional visualization model.

[0065] Reference Figure 4 In one embodiment, a three-dimensional visualization device for a construction area is provided, the device comprising: The image acquisition module 1100 is used to acquire an initial synthetic aperture radar image and a current optical image of the construction area, and perform Fourier transform on the current optical image to convert it into a simulated synthetic aperture radar image; The feature extraction module 1200 is used to splice the initial synthetic aperture radar image and the simulated synthetic aperture radar image as the input image, extract global features of multiple scales and local features of multiple scales from the input image, fuse the global features of each scale in the multiple scales of global features with the local features of the corresponding scale in the multiple scales of local features to obtain fused features of multiple scales, and upsample and fuse the fused features of multiple scales to generate final features. The image affine transformation module 1300 is used to generate an affine transformation matrix from the final features according to the multi-layer perceptron, and convert the simulated synthetic aperture radar image into a distorted synthetic aperture radar image according to the affine transformation matrix; The 3D visualization module 1400 is used to generate a spatial point cloud of the construction area based on the distorted SAR image and the initial SAR image, so as to achieve 3D visualization of the construction area.

[0066] It should be noted that the three-dimensional visualization device of the construction area provided in this embodiment and the three-dimensional visualization method of the construction area mentioned above are based on the same inventive concept. Therefore, the relevant content of the three-dimensional visualization method of the construction area mentioned above is also applicable to the content of the three-dimensional visualization device of the construction area. Therefore, it will not be repeated here.

[0067] like Figure 5 , an embodiment of the present application further provides an electronic device, the electronic device comprising: at least one memory; at least one processor; at least one program; The programs are stored in the memory, and the processor executes at least one program to implement the three-dimensional visualization method of the construction area described above in the present disclosure.

[0068] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.

[0069] The electronic device according to the embodiment of the present application is described in detail below.

[0070] The processor 1600 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. Memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). Memory 1700 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 1700 and is called by processor 1600 to execute the three-dimensional visualization method of the construction area in the embodiments of this application.

[0071] Input / output interface 1800, used for information input and output; Communication interface 1900, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 ); The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .

[0072] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned three-dimensional visualization method of the construction area.

[0073] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0074] The embodiments described in this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0075] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0076] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0077] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0078] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0079] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0080] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0081] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0082] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0083] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0084] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the embodiments of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the embodiments of the present application.

Claims

1. A three-dimensional visualization method for a construction area, characterized in that: The method comprises: Acquire an initial synthetic aperture radar image and a current optical image of the construction area, and perform a Fourier transform on the current optical image to obtain a simulated synthetic aperture radar image; splicing the initial synthetic aperture radar image and the simulated synthetic aperture radar image as an input image, extracting global features at multiple scales and local features at multiple scales from the input image, fusing the global features at each scale of the multiple global features with the local features at corresponding scales of the multiple local features to obtain fused features at multiple scales, and upsampling and fusing the fused features at multiple scales to generate final features; generating an affine transformation matrix from the final features according to a multi-layer perceptron, and converting the simulated synthetic aperture radar image into a distorted synthetic aperture radar image according to the affine transformation matrix; A spatial point cloud of the construction area is generated according to the distorted synthetic aperture radar image and the initial synthetic aperture radar image, so as to achieve three-dimensional visualization of the construction area.

2. The three-dimensional visualization method of a construction area according to claim 1, characterized in that: The extracting global features of multiple scales and local features of multiple scales from the input image includes: Inputting the image to be input into a first branch network to obtain global features of multiple scales output by the first branch network; wherein the first branch network includes a plurality of cascaded Transformer units and a downsampling unit is included between every two Transformer units; The image to be input is input into the second branch network to obtain local features of multiple scales output by the second branch network; wherein the second branch network includes multiple cascaded CNN units and a downsampling unit is included between every two CNN units.

3. The three-dimensional visualization method of a construction area according to claim 2, characterized in that: The process of fusing the global features of the target scale and the local features of the target scale to obtain the fused features of the target scale includes: Performing global average pooling and global maximum pooling on the global features of the target scale to obtain a first pooled feature, and performing global maximum pooling on the global features of the target scale to obtain a second pooled feature; wherein the global feature of the target scale is a global feature of any scale among the global features of the multiple scales, and the local feature of the target scale is a local feature of the same scale as the global feature of the target scale among the local features of the multiple scales; Performing global average pooling on the local features of the target scale to obtain a third pooling feature, and performing global maximum pooling on the local features of the target scale to obtain a fourth pooling feature; The first pooled feature and the second pooled feature are respectively used as the head and tail of the local feature of the target scale, and are concatenated to obtain a first concatenated feature. The third pooled feature and the fourth pooled feature are respectively used as the head and tail of the global feature of the target scale, and are concatenated to obtain a second concatenated feature. Inputting the first splicing feature and the second splicing feature into the Transformer unit respectively, and splicing the features respectively output by the Transformer unit to obtain a first intermediate feature; Inputting the global feature of the target scale into an attention enhancement unit to obtain a second intermediate feature output by the attention enhancement unit; Multiplying the global feature of the target scale and the local feature of the target scale pixel by pixel to obtain a third intermediate feature; The first intermediate feature, the second intermediate feature, and the third intermediate feature are concatenated to obtain a fused feature of the target scale.

4. The three-dimensional visualization method of a construction area according to claim 3, characterized in that: The attention enhancement unit is a CBAM convolutional block attention unit.

5. The three-dimensional visualization method of a construction area according to claim 2, characterized in that: The process of extracting local features by the CNN unit includes: ; ; ; ; ; ; ; ; in, is the input data of the CNN unit, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, for The convolutional layer, is the SiLU activation function, is the channel maximum pooling, is channel average pooling, for The convolutional layer, is the sigmoid activation function, is the mapping function for element rearrangement operation, is a multi-layer perceptron, It is the local feature extracted by the CNN unit.

6. The three-dimensional visualization method of a construction area according to claim 1, characterized in that: Before fusing the global features of each scale among the global features of the multiple scales with the local features of the corresponding scale among the local features of the multiple scales to obtain fused features of the multiple scales, the method further includes: Calculating a difference image between the initial synthetic aperture radar image and the simulated synthetic aperture radar image based on an LR operator; extracting difference features of multiple scales of the difference image; The fusing of the global features of each scale in the global features of the multiple scales with the local features of the corresponding scale in the local features of the multiple scales to obtain fused features of the multiple scales includes: The global features of each scale among the global features of the multiple scales, the difference features of the corresponding scale among the difference features of the multiple scales, and the local features of the corresponding scale among the local features of the multiple scales are fused to obtain fused features of the multiple scales.

7. A three-dimensional visualization device for a construction area, characterized in that: The device comprises: An image acquisition module is used to acquire an initial synthetic aperture radar image and a current optical image of the construction area, and Fourier transform the current optical image into a simulated synthetic aperture radar image; a feature extraction module, configured to splice the initial synthetic aperture radar image and the simulated synthetic aperture radar image as an input image, extract global features at multiple scales and local features at multiple scales from the input image, fuse the global features at each scale of the multiple global features with the local features at corresponding scales of the multiple local features to obtain fused features at multiple scales, and upsample and fuse the fused features at multiple scales to generate final features; an image affine transformation module, configured to generate an affine transformation matrix from the final features according to a multi-layer perceptron, and convert the simulated synthetic aperture radar image into a distorted synthetic aperture radar image according to the affine transformation matrix; A three-dimensional visualization module is used to generate a spatial point cloud of the construction area based on the distorted synthetic aperture radar image and the initial synthetic aperture radar image, so as to achieve three-dimensional visualization of the construction area.

8. The three-dimensional visualization device of the construction area according to claim 7, characterized in that: The device further comprises: A difference feature extraction module is used to calculate a difference image between the initial synthetic aperture radar image and the simulated synthetic aperture radar image based on an LR operator; and extract difference features of multiple scales of the difference image; The feature extraction module is specifically used to fuse the global features of each scale among the global features of the multiple scales, the difference features of the corresponding scale among the difference features of the multiple scales, and the local features of the corresponding scale among the local features of the multiple scales to obtain fused features of multiple scales.

9. An electronic device, characterized in that: include: at least one control processor and a memory for communicatively coupling with the at least one control processor; The memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform the three-dimensional visualization method of the construction area according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the three-dimensional visualization method of the construction area according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Optical image and SAR image registration method based on multi-scale phase consistency

    CN114494371A

  • Multi-dimensional reinforcement learning synthetic aperture radar image target detection method

    CN114565860A

  • Optical and SAR image registration method, device and equipment based on position awareness

    CN116883466A

  • SAR (Synthetic Aperture Radar) image classification method based on CNN (Convolutional Neural Network) and Transform

    CN117237740A

  • Multi-modal image matching method and system, terminal equipment and storage medium

    CN117541833A