GIS digital twin scene construction method based on video mapping

Through the method of video mapping based on the construction of GIS digital twin scenarios, combined with sliding windows and deep learning technology, the problems of local and global coordinate systems separation, geometric consistency and cross-modal mapping accuracy bottlenecks are solved, and high-precision and high-efficiency scene digitization is achieved.

CN120107866AActive Publication Date: 2025-06-06SHANDONG FEIYUAN SPACE INFORMATION TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510571167.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing GIS digital twin scenario construction methods have the separation of local and global coordinate systems, the difficulty in modeling geometric consistency of dynamic scenarios, and the accuracy bottleneck of cross-modal mapping.

Method used

Using a video map-based method, the video stream is divided through sliding windows, the local coordinate system is defined, and combined with deep learning-driven multi-view feature interaction, high-precision local geometric structure modeling and global coordinate conversion are achieved.

Benefits of technology

It realizes the construction of high-precision and high-efficiency GIS digital twin scenarios, solves the problems of local and global coordinate systems, geometric consistency of dynamic scenes and cross-modal mapping accuracy, and is suitable for real-time digitalization needs of dynamic environments and large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107866A_ABST
    Figure CN120107866A_ABST
Patent Text Reader

Abstract

The invention provides a GIS digital twin scene construction method based on video mapping. Relates to the technical field of digital twinning scene construction. A real scene video is scanned and collected; segmenting the real scene video into a plurality of video slices; determining an intermediate video frame of each video slice, defining a local coordinate system of the video slice by using the intermediate video frame, and taking the local coordinate system of the first video slice as a global coordinate system; inputting the video slice into a pre-trained multi-view image-to-point cloud model, and producing a local geometric structure local coordinate point cloud of a corresponding local geometric structure based on the video slice; recording local coordinate point clouds of local geometric structures corresponding to all video frames, and inputting the local coordinate point clouds of the local geometric structures into the coordinate transformation model to obtain global coordinate point clouds of the local geometric structures; and splicing and combining the global coordinate point clouds of the local geometric structures together according to the global coordinates to obtain a GIS digital twinborn scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of GIS digital twin scene construction, and in particular to a GIS digital twin scene construction method based on video mapping. Background Art

[0002] With the widespread application of digital twin technology in smart cities, industrial inspection, autonomous driving and other fields, how to efficiently and accurately map real physical scenes into GIS digital twin scene models has become a key technical challenge.

[0003] In recent years, deep learning techniques (such as visual transformers) have shown potential in multi-view geometric modeling, and can implicitly model the geometric associations between multiple views through end-to-end training. However, existing methods still have the following problems: The separation of local and global coordinate systems: Most methods directly map multi-view images to the global coordinate system, but the continuous changes in local geometric structures in video streams can easily lead to uncertainty in global coordinate predictions. Geometric consistency of dynamic scenes: It is difficult to model the spatial correlation of local geometric structures between video frames, resulting in misalignment or holes when point clouds are stitched. Accuracy bottleneck of cross-modal mapping: The mapping of images to point clouds relies on manually designed features or fixed depth estimation networks, which are difficult to adapt to the geometric diversity of complex scenes.

[0004] To address the above problems, the present invention proposes a GIS digital twin scene construction method based on video mapping, which combines sliding window local geometric modeling with multi-view feature interaction driven by deep learning to achieve high-precision and high-efficiency scene digitization. Summary of the invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a GIS digital twin scene construction method based on video mapping.

[0006] In a first aspect, the present invention provides a GIS digital twin scene construction method based on video mapping, comprising: Scan and collect real scene videos; A sliding window of a set length divides the real scene video into multiple video slices according to a set step size; Determine the middle video frame of each video slice, define the coordinate system where the local geometric structure in the middle video frame is located as the local coordinate system of the video slice, wherein the local coordinate system of the video slice segmented by the first sliding window is used as the global coordinate system; Input any video slice into a pre-trained multi-view image to point cloud model, the multi-view image to point cloud model generates a local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, the local geometric structure local coordinate point cloud uses the local coordinate system coordinates of the video slice; Recording local geometric structure local coordinate point clouds corresponding to all video frames, taking the local geometric structure local coordinate point clouds as input, and inputting them into a coordinate transformation model, wherein the coordinate transformation model transforms the point cloud coordinates of all local geometric structure local coordinate point clouds into the global coordinate system to obtain the local geometric structure global coordinate point clouds; The global coordinate point clouds of each local geometric structure are spliced ​​together according to the global coordinates to obtain the GIS digital twin scene.

[0007] Furthermore, the local geometric structure in each video frame of the real scene video changes continuously with the continuous change of the time axis of the real scene video.

[0008] Furthermore, the multi-view image to point cloud model comprises: an image embedding layer, a position encoder, a multi-view image encoder based on a visual Transformer, a multi-view image interaction decoder based on a bidirectional visual Transformer, and a local point cloud prediction head; The image embedding layer processes the video frame to obtain the video frame patch embedding, and the position encoder uses the rotation position encoding to add the rotation position encoding to the video frame patch embedding to obtain the video frame feature; for each video slice, the visual Transformer of the multi-view image encoder encodes the video frame feature of the video frame in the video slice to obtain the encoded video frame feature representing the local geometric structure in the video frame of the video slice and the encoded position feature representing the local geometric structure of the video frame in the local coordinate system of the video slice; the multi-view image interactive decoder exchanges the encoded video frame features and the encoded position features of the intermediate video frame and the non-intermediate video frame through the multi-layer bidirectional visual Transformer to obtain the decoded video frame features and the decoded position features of the intermediate video frame and each non-intermediate video frame in multiple layers; the local point cloud prediction head fuses the multi-layer decoded video frame features and the decoded position features through the feature pyramid, and then the prediction head maps the fused features into the local point cloud coordinates and confidences per pixel; the local geometric structure local coordinate point cloud is used as input and input into the coordinate transformation model, and the coordinate transformation model transforms the point cloud coordinates of all local geometric structure local coordinate point clouds into the global coordinate system to obtain the local geometric structure global coordinate point cloud.

[0009] Furthermore, the multi-view image encoder includes multiple layers of visual Transformer. In each layer of visual Transformer, the visual Transformer performs self-attention on video frame features of any video frame, and cross-attention is performed on the video frame features after self-attention of the video frame with video frame features of other video frames respectively to obtain multiple groups of combined video frame features, and the multiple groups of combined video frame features are aggregated through pooling operation to obtain aggregated video frame features; the aggregated video frame features and the video frame features of the video frame after self-attention are added and combined, and then normalized and input into the feedforward network of the residual structure for further processing to obtain features combined with the overall local geometric structure information in the video slice, including: encoded video frame features and encoded position features of the local geometric structure of the video frame in the local coordinate system of the video slice.

[0010] Furthermore, the encoded video frame features and encoded position features of each video frame are mapped to the multi-view image interaction decoder through a linear layer. The multi-layer bidirectional visual Transformer of the multi-view image interaction decoder has the same structure as the visual Transformer of the multi-view image encoder, but the positions of the encoded video frame features and encoded position features of the intermediate video frames and non-intermediate video frames are exchanged during calculation to achieve bidirectional information exchange and obtain the intermediate video frames and each non-intermediate video frame's multi-layer decoded video frame features and decoded position features.

[0011] Furthermore, the local point cloud prediction head fuses the multi-layer encoded video frame features and the encoded position features through a feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidences; the prediction head includes: a convolution layer, an interpolation upsampling layer, a convolution layer, an activation function layer and a convolution layer; the pixel-by-pixel local point cloud coordinates are combined to form a local geometric structure local coordinate point cloud.

[0012] Furthermore, the multi-view image to point cloud model is trained using the local coordinates of the real scene point cloud, and the first loss function for training the multi-view image to point cloud model is: ,in, Video Frame The local coordinates of the corresponding real scene point cloud, Video frames predicted by the multi-view image-to-point cloud model The local coordinates of the corresponding local geometry local coordinate point cloud, It means to find the L1 distance between the predicted local coordinates and the true local coordinates. Represents the local coordinate prediction confidence matrix Entropy,training adjusts the multi-view image to point cloud model parameters to minimize the first loss function.

[0013] Furthermore, the coordinate transformation model is consistent with the structure of the multi-view image to point cloud model, including: a point cloud embedding layer, a position encoder, a multi-view point cloud encoder based on a visual Transformer, a multi-view point cloud interactive decoder based on a bidirectional visual Transformer, and a global point cloud prediction head; the coordinate transformation model uses the global coordinate system of the first group of local geometric structure local coordinate point clouds as a reference, and transforms the second group of local geometric structure local coordinate point clouds into the global coordinate system to generate a second group of local geometric structure global coordinate point clouds. Subsequently, the coordinate transformation model uses the global coordinates of the selected preceding local geometric structure global coordinate point cloud as a reference, and transforms the local coordinates of the current group of local geometric structure local coordinate point clouds into global coordinates, thereby transforming the local geometric structure local coordinate point clouds of all views into the global coordinate system.

[0014] Furthermore, when selecting a global coordinate system reference from the global coordinate point cloud of the local geometry of the preceding group, a confidence threshold matrix is ​​used to filter out the preceding views whose corresponding local coordinate prediction confidence matrices perform poorly.

[0015] Furthermore, the coordinate transformation model is trained using the global coordinates of the real scene point cloud, and the second loss function for training the coordinate transformation model is: ,in, Video Frame The global coordinates of the corresponding real scene point cloud, Video frames predicted by the coordinate transformation model The global coordinates of the local coordinate point cloud corresponding to the local geometry structure, It means to find the L1 distance between the two. Represents the global coordinate prediction confidence matrix The entropy of the coordinate transformation model is trained to minimize the second loss function.

[0016] In the second aspect, the present invention provides a GIS digital twin scene construction device based on video mapping, comprising: at least one processing unit, the processing unit, the storage unit and the acquisition unit are interconnected through a bus unit, the storage unit stores a computer program and data collected by the acquisition unit, and when the computer program is executed by the processing unit, the GIS digital twin scene construction method based on video mapping is implemented.

[0017] In a third aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the GIS digital twin scene construction method based on video mapping.

[0018] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art: The present application divides the scanned real scene video into multiple video slices; determines the intermediate video frame of each video slice, and defines the local coordinate system of the video slice according to the intermediate video frame. The local coordinate system is used for the multi-view image to point cloud model to construct a local geometric structure local coordinate point cloud. Among them, the local coordinate system of the first video slice cut by the sliding window is used as the global coordinate system, and the global coordinate system is used for the coordinate conversion model to convert the point cloud coordinates of the local geometric structure local coordinate point cloud to the global coordinate system to obtain the local geometric structure global coordinate point cloud; any video slice is input into the pre-trained multi-view image to point cloud model, and the multi-view image to point cloud model produces the local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, and the local geometric structure local coordinate point cloud uses the local coordinate system coordinates of the video slice; the local geometric structure local coordinate point cloud corresponding to all video frames is recorded, and the local geometric structure local coordinate point cloud is used as input and input into the coordinate conversion model, and the coordinate conversion model converts the point cloud coordinates of all local geometric structure local coordinate point clouds to the global coordinate system to obtain the local geometric structure global coordinate point cloud; splice and combine the global coordinate point clouds of each local geometric structure together according to the global coordinate to obtain the GIS digital twin scene. This application dynamically binds video slices to the local coordinate system: the video stream is cut into video slices with spatial correlation through a sliding window, and the intermediate frame is used as the local coordinate system benchmark to ensure the continuity and consistency of the local geometric structure.

[0019] This application implements the Transformer architecture for multi-view image point cloud conversion: a visual Transformer model including rotational position encoding and multi-view cross-attention mechanism is designed to explicitly model the geometric association of multiple frames in a video slice, and generate a local coordinate point cloud of a local geometric structure with high confidence. Progressive stitching strategy for global coordinate transformation: through a confidence-guided coordinate transformation model, the local point cloud is gradually aligned to the global coordinate system to avoid cumulative errors and improve the robustness of mapping in large-scale scenes. This method deeply integrates the real-time performance of SLAM and the high-precision advantages of deep learning at the technical level, providing a scalable, end-to-end solution for the construction of GIS digital twin scenes, especially for real-time digitization requirements of dynamic environments and large-scale scenes. By introducing rotational position encoding and bidirectional Transformer interactive decoding mechanisms, the accuracy of local geometric structure reconstruction and the stability of global stitching are significantly improved, providing a new technical path for the construction of high-precision maps in the fields of smart cities and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0022] Figure 1 A flowchart of a GIS digital twin scene construction method based on video mapping provided in an embodiment of the present invention; Figure 2 The overall architecture diagram of the model provided by the embodiment of the present invention; Figure 3 A schematic diagram of a multi-view image encoder provided by an embodiment of the present invention; Figure 4 A schematic diagram of a multi-view image interactive decoder provided by an embodiment of the present invention; Figure 5 A schematic diagram of a local point cloud prediction head provided by an embodiment of the present invention; Figure 6 A schematic diagram of a GIS digital twin scene construction device based on video mapping provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0025] Example 1 like Figure 1 As shown, the technology of the present invention realizes a GIS digital twin scene construction method based on video mapping, and the present invention includes the following steps: Scan and collect real scene videos Scanning makes real scene video The video frame has the following properties: the local geometric structures of the real scene in the video frame vary with the real scene video The continuous change of the time axis changes continuously, that is, the real scene video There are no spatial jumps in the local geometric structure.

[0026] The sliding window of set length converts the real scene video into Split into multiple video slices ,in, Real scene video The total number of video slices; any video slice The local geometric structures of all video frames in are spatially correlated.

[0027] Determine the middle video frame for each video slice , the coordinate system of the local geometric structure in the middle video frame is defined as the local coordinate system of the video slice, wherein the local coordinate system of the video slice segmented by the first sliding window is used as the global coordinate system.

[0028] Input any video slice into the pre-trained multi-view image to point cloud model. The multi-view image to point cloud model generates the local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice. The coordinates of the local geometric structure local coordinate point cloud are based on the local coordinate system corresponding to the video slice.

[0029] In the specific implementation process, the multi-view image to point cloud model includes: an image embedding layer, a position encoder, a multi-view image encoder based on a visual Transformer, a multi-view image interaction decoder based on a bidirectional visual Transformer, and a local point cloud prediction head.

[0030] The principle of the multi-view image to point cloud model is as follows: the image embedding layer processes each video frame of the video slice to obtain a video frame patch embedding, and the position encoder uses rotation position encoding to add rotation position encoding to the video frame patch embedding to obtain video frame features. For the i-th video slice The jth video frame in , and its video frame features are expressed as The superscript j represents the index of the video frame in the video slice, wherein the superscript mid represents the middle video frame.

[0031] For each video slice, the visual Transformer of the multi-view image encoder encodes video frame features of the video frames in the video slice to obtain encoded video frame features representing local geometric structures in the video frames of the video slice and encoded position features representing the local geometric structures of the video frames in the local coordinate system of the video slice.

[0032] In a specific implementation process, the multi-view image encoder includes a multi-layer visual Transformer.

[0033] In each layer of the visual transformer, first, the visual transformer performs self-attention on the video frame features of any input video frame, and cross-attention is performed on the video frame features of the video frame after self-attention with the video frame features of other video frames in the video slice to obtain multiple sets of combined video frame features, and the multiple sets of combined video frame features are aggregated through pooling operations to obtain the aggregated video frame features of the video frame; the aggregated video frame features of the video frame and the video frame features of the video frame after self-attention are added and combined and then normalized, and input into the feedforward network of the residual structure for further processing, and the feedforward network outputs the features of the video frame combined with the overall local geometric structure information in the video slice, including: encoded video frame features and encoded position features of the local geometric structure of the video frame in the local coordinate system of the video slice. The final visual transformer of the multi-view image encoder obtains the encoded video frame features that represent the local geometric structure of the video frame in the video slice and the encoded position features that represent the local geometric structure of the video frame in the local coordinate system of the video slice.

[0034] Taking the middle video frame as an example, the visual Transformer performs self-attention on the video frame features of the middle video frame: ; in, represents normalization, Represents self-attention. The principle of self-attention is as follows: The self-attention query, self-attention key, and self-attention value are obtained by linearly mapping the normalized video frame features of the intermediate video frame: ; Self-Attention The formula is as follows: ; in, is the softmax function.

[0035] Secondly, the visual Transformer performs cross-attention on the video frame features of the intermediate video frame after self-attention and the video frame features of each non-intermediate video frame: ; in, Indicates cross attention, Encode the rotational position of the video frame features of the intermediate video frames and non-intermediate video frames; the principle of cross attention is as follows: The cross-attention query is obtained by linearly mapping the normalized video frame features of the intermediate video frame after self-attention. ; The cross attention key is obtained by linearly mapping the normalized video frame features of non-intermediate video frames and cross attention value ; Cross-Attention The formula is as follows: ; in, Represents the pytorch cross_attn.rope(,) function that applies rotational position encoding to a cross-attention query.

[0036] Through cross attention, the video frame features of the intermediate video frame query features from the video frame features of each non-intermediate video frame, and combine the features of each non-intermediate video frame to obtain multiple sets of combined video frame features. ; Then, the visual Transformer combines multiple sets of video frame features Aggregating the aggregated video frame features of the intermediate video frames through a pooling operation; Finally, the aggregated video frame features of the intermediate video frame and the video frame features of the intermediate video frame after self-attention are added together and then normalized and input into the feedforward network of the residual structure for further processing.

[0037] The processing process of other video frames is the same, so that the encoding features of each video frame in the video slice are combined with the information of other video frames. The encoding process of the multi-view image encoder actually uses the associated local geometric structure features in the video slice to provide support for the local geometric structure feature encoding of the video frame in any video slice, ensuring the certainty and accuracy of the local geometric structure information represented by the encoding of each video frame.

[0038] like Figure 4As shown, the encoded video frame features and encoded position features of each video frame are mapped to the multi-view image interaction decoder through a linear layer. The multi-view image interaction decoder exchanges the encoded video frame features and encoded position features of the intermediate video frame and the non-intermediate video frame through a multi-layer bidirectional visual Transformer. The multi-layer bidirectional visual Transformer has the same structure as the visual Transformer of the multi-view image encoder, but when exchanging the positions of the encoded video frame features and encoded position features of the intermediate video frame and the non-intermediate video frame, bidirectional information exchange is achieved through bidirectional cross attention to obtain the decoded video frame features and decoded position features of the intermediate video frame and each non-intermediate video frame. Through this process, the spatial relationship between the non-intermediate video frame and the intermediate video frame is modeled. Provide accurate relative position association for the subsequent generation of the local coordinate point cloud of the local geometric structure.

[0039] like Figure 5 As shown, the local point cloud prediction head fuses the multi-layer decoded video frame features and the decoded position features through the feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidence. The prediction head includes: a convolution layer, an interpolation upsampling layer, a convolution layer, an activation function layer and a convolution layer. The pixel-by-pixel local point cloud coordinates form a local geometric structure local coordinate point cloud. The local point cloud prediction head generates a local geometric structure local coordinate point cloud based on the spatial position relationship between the non-intermediate video frame and the intermediate video frame.

[0040] In order to enable the multi-view image to point cloud model to achieve the above functions, the local coordinates of the real scene point cloud are used for training. The training loss is determined by the distance from the predicted point cloud point to the local coordinates of the real scene point cloud and the entropy of the local coordinate prediction confidence. The first loss function for training the multi-view image to point cloud model is: ,in, Video Frame The local coordinates of the corresponding real scene point cloud, Video frames predicted by the multi-view image-to-point cloud model The local coordinates of the corresponding local geometry local coordinate point cloud, It means to find the L1 distance between the predicted local coordinates and the true local coordinates. Represents the local coordinate prediction confidence matrix Entropy, minimizing the first loss function is to make the predicted local geometric structure local coordinate point cloud as close as possible to the real scene point cloud, and make the model prediction result more certain.

[0041] Record the local coordinate point cloud of the local geometric structure corresponding to all video frames.

[0042] The local coordinate point cloud of the local geometric structure is used as input and input into the coordinate transformation model, and the coordinate transformation model transforms the point cloud coordinates of all geometric structure point clouds into the global coordinate system. The structure of the coordinate transformation model is consistent with the structure of the multi-view image to point cloud model, and correspondingly includes: a point cloud embedding layer, a position encoder, a multi-view point cloud encoder based on a visual Transformer, a multi-view point cloud interactive decoder based on a bidirectional visual Transformer, and a global point cloud prediction head; the coordinate transformation model uses the global coordinate system of the first group of local geometric structure local coordinate point clouds as a reference to transform the second group of local geometric structure local coordinate point clouds into the global coordinate system to generate the second group of local geometric structure global coordinate point clouds, and subsequently, the coordinate transformation model uses the global coordinates of the selected previous local geometric structure global coordinate point cloud as a reference to transform the local coordinates of the current group of local geometric structure local coordinate point clouds into global coordinates, wherein when selecting the global coordinate system reference from the local geometric structure global coordinate point cloud of the previous group, the confidence threshold matrix is ​​used to filter out the previous view with poor performance of the corresponding local coordinate prediction confidence matrix. Through the above process, the coordinate transformation model transforms the local coordinate point cloud of the local geometric structure of all views into the global coordinate system.

[0043] The multi-view point cloud encoder has the same structure and function as the multi-view image encoder, and processes the local geometric structure and local coordinate point cloud in the point cloud slices divided by the same sliding window.

[0044] The multi-view point cloud interactive decoder has the same structure and function as the multi-view image interactive decoder. It processes local geometry structure local coordinate point clouds that have been converted into local geometry structure global coordinate point clouds and local geometry structure local coordinate point clouds that have not been converted into local geometry structure global coordinate point clouds, and performs two-way information exchange decoding between any combination of the two.

[0045] The global point cloud prediction head has the same structure and function as the local point cloud prediction head. Both generate point cloud coordinates based on spatial relationships.

[0046] It can be seen that the object is only programmed from the video frame to the point cloud frame, and the coordinate reference is changed from the local coordinate system to the global coordinate system. The working principles of the multi-view image to point cloud model and the coordinate transformation model are the same. Therefore, the two have the same model structure but different model parameters.

[0047] In order to enable the coordinate conversion model to achieve the above functions, end-to-end training is performed using the global coordinates of the real scene point cloud. The training loss is determined by the distance from the global coordinate prediction confidence weighted prediction point cloud point to the global coordinate of the real scene point cloud and the entropy of the global coordinate prediction confidence. The second loss function for training the coordinate conversion model is: ,in, Video Frame The global coordinates of the corresponding real scene point cloud, Video frames predicted by the coordinate transformation model The corresponding local geometry global coordinates of the point cloud, It means to find the L1 distance between all the coordinates of the two. Represents the global coordinate prediction confidence matrix The entropy of the coordinate transformation model is trained to adjust the parameters of the coordinate transformation model to minimize the second loss function, that is, to make the global coordinates of the predicted local geometric structure global coordinate point cloud as close as possible to the global coordinates of the real scene point cloud, and to make the prediction results of the coordinate transformation model more deterministic.

[0048] The global coordinate point clouds of each local geometric structure are spliced ​​together according to the global coordinates to obtain the GIS digital twin scene.

[0049] Example 2 See also Figure 6 As shown, an embodiment of the present invention provides a GIS digital twin scene construction device based on video mapping, including: at least one processing unit, the processing unit, the storage unit and the acquisition unit are interconnected through a bus unit, and the storage unit is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as the software programs, computer executable programs and modules corresponding to the GIS digital twin scene construction method based on video mapping in an embodiment of the present invention. The processing unit implements the above-mentioned GIS digital twin scene construction method based on video mapping by running the software programs, computer executable programs and modules stored in the storage unit, including: Scan and collect real scene videos; A sliding window of a set length divides the real scene video into multiple video slices according to a set step size; Determine the middle video frame of each video slice, define the coordinate system where the local geometric structure in the middle video frame is located as the local coordinate system of the video slice, wherein the local coordinate system of the video slice segmented by the first sliding window is used as the global coordinate system; Input any video slice into a pre-trained multi-view image to point cloud model, the multi-view image to point cloud model generates a local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, the local geometric structure local coordinate point cloud uses the local coordinate system coordinates of the video slice; Recording local geometric structure local coordinate point clouds corresponding to all video frames, taking the local geometric structure local coordinate point clouds as input, and inputting them into a coordinate transformation model, wherein the coordinate transformation model transforms the point cloud coordinates of all local geometric structure local coordinate point clouds into the global coordinate system to obtain the local geometric structure global coordinate point clouds; The global coordinate point clouds of each local geometric structure are spliced ​​together according to the global coordinates to obtain the GIS digital twin scene.

[0050] Of course, the computer program stored in the storage unit of the GIS digital twin scene construction device based on video mapping provided in an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in the GIS digital twin scene construction method based on video mapping provided in any embodiment of the present invention.

[0051] Example 3 An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed, the method for constructing a GIS digital twin scene based on video mapping is implemented, including: Scan and collect real scene videos; A sliding window of a set length divides the real scene video into multiple video slices according to a set step size; Determine the middle video frame of each video slice, define the coordinate system where the local geometric structure in the middle video frame is located as the local coordinate system of the video slice, wherein the local coordinate system of the video slice segmented by the first sliding window is used as the global coordinate system; Input any video slice into a pre-trained multi-view image to point cloud model, the multi-view image to point cloud model generates a local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, the local geometric structure local coordinate point cloud uses the local coordinate system coordinates of the video slice; Recording local geometric structure local coordinate point clouds corresponding to all video frames, taking the local geometric structure local coordinate point clouds as input, and inputting them into a coordinate transformation model, wherein the coordinate transformation model transforms the point cloud coordinates of all local geometric structure local coordinate point clouds into the global coordinate system to obtain the local geometric structure global coordinate point clouds; The global coordinate point clouds of each local geometric structure are spliced ​​together according to the global coordinates to obtain the GIS digital twin scene.

[0052] A computer-readable storage medium provided in an embodiment of the present invention stores a computer program that is not limited to the method operations described above, but can also execute related operations in a GIS digital twin scene construction method based on video mapping provided in any embodiment of the present invention.

[0053] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.

[0054] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0055] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0056] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A GIS digital twin scene construction method based on video mapping, characterized in that: include: Scan and collect real scene videos; A sliding window of a set length divides the real scene video into multiple video slices according to a set step size; Determine the middle video frame of each video slice, define the coordinate system where the local geometric structure in the middle video frame is located as the local coordinate system of the video slice, wherein the local coordinate system of the video slice segmented by the first sliding window is used as the global coordinate system; Input any video slice into a pre-trained multi-view image to point cloud model, the multi-view image to point cloud model generates a local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, the local geometric structure local coordinate point cloud uses the local coordinate system coordinates of the video slice; Recording local geometric structure local coordinate point clouds corresponding to all video frames, taking the local geometric structure local coordinate point clouds as input, and inputting them into a coordinate transformation model, wherein the coordinate transformation model transforms the point cloud coordinates of all local geometric structure local coordinate point clouds into the global coordinate system to obtain the local geometric structure global coordinate point clouds; The global coordinate point clouds of each local geometric structure are spliced ​​together according to the global coordinates to obtain the GIS digital twin scene.

2. The GIS digital twin scene construction method based on video mapping according to claim 1 is characterized in that: The local geometric structure in each video frame of the real scene video changes continuously with the continuous change of the time axis of the real scene video.

3. The GIS digital twin scene construction method based on video mapping according to claim 1 is characterized in that: The multi-view image to point cloud model includes: an image embedding layer, a position encoder, a multi-view image encoder based on a visual Transformer, a multi-view image interactive decoder based on a bidirectional visual Transformer, and a local point cloud prediction head; The image embedding layer processes the video frame to obtain the video frame patch embedding, and the position encoder uses the rotation position encoding to add the rotation position encoding to the video frame patch embedding to obtain the video frame feature; for each video slice, the visual Transformer of the multi-view image encoder encodes the video frame feature of the video frame in the video slice to obtain the encoded video frame feature representing the local geometric structure in the video frame of the video slice and the encoded position feature representing the local geometric structure of the video frame in the local coordinate system of the video slice; the multi-view image interactive decoder exchanges the encoded video frame features and the encoded position features of the intermediate video frame and the non-intermediate video frame through the multi-layer bidirectional visual Transformer to obtain the decoded video frame features and the decoded position features of the intermediate video frame and each non-intermediate video frame in multiple layers; the local point cloud prediction head fuses the multi-layer decoded video frame features and the decoded position features through the feature pyramid, and then the prediction head maps the fused features into the local point cloud coordinates and confidences per pixel; the local geometric structure local coordinate point cloud is used as input and input into the coordinate transformation model, and the coordinate transformation model transforms the point cloud coordinates of all local geometric structure local coordinate point clouds into the global coordinate system to obtain the local geometric structure global coordinate point cloud.

4. The GIS digital twin scene construction method based on video mapping according to claim 3 is characterized in that: The multi-view image encoder includes multiple layers of visual Transformer. In each layer of visual Transformer, the visual Transformer performs self-attention on the video frame features of any video frame, and cross-attention is performed on the video frame features of the video frame after self-attention with the video frame features of other video frames to obtain multiple groups of combined video frame features, and the multiple groups of combined video frame features are aggregated through pooling operation to obtain aggregated video frame features; the aggregated video frame features and the video frame features of the video frame after self-attention are added and combined, and then normalized and input into a feedforward network of a residual structure for further processing to obtain features combined with the overall local geometric structure information in the video slice, including: encoded video frame features and encoded position features of the local geometric structure of the video frame in the local coordinate system of the video slice.

5. The GIS digital twin scene construction method based on video mapping according to claim 4 is characterized in that: The encoded video frame features and encoded position features of each video frame are mapped to the multi-view image interaction decoder through a linear layer. The multi-layer bidirectional visual Transformer of the multi-view image interaction decoder has the same structure as the visual Transformer of the multi-view image encoder, but the positions of the encoded video frame features and encoded position features of the intermediate video frames and non-intermediate video frames are exchanged during calculation to achieve bidirectional information exchange and obtain the intermediate video frames and each non-intermediate video frame's multi-layer decoded video frame features and decoded position features.

6. The GIS digital twin scene construction method based on video mapping according to claim 3 is characterized in that: The local point cloud prediction head fuses the multi-layer coded video frame features and the coded position features through a feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidences; the prediction head comprises: a convolution layer, an interpolation upsampling layer, a convolution layer, an activation function layer and a convolution layer; the pixel-by-pixel local point cloud coordinates are combined to form a local coordinate point cloud of a local geometric structure.

7. The GIS digital twin scene construction method based on video mapping according to claim 3 is characterized in that: The multi-view image to point cloud model is trained using the local coordinates of the real scene point cloud, and the first loss function for training the multi-view image to point cloud model is: ,in, Video Frame The local coordinates of the corresponding real scene point cloud, Video frames predicted by the multi-view image-to-point cloud model The local coordinates of the corresponding local geometry local coordinate point cloud, It means to find the L1 distance between the predicted local coordinates and the true local coordinates. Represents the local coordinate prediction confidence matrix Entropy,training adjusts the multi-view image to point cloud model parameters to minimize the first loss function.

8. The GIS digital twin scene construction method based on video mapping according to claim 3 is characterized in that: The coordinate transformation model is consistent with the structure of the multi-view image to point cloud model, and includes: a point cloud embedding layer, a position encoder, a multi-view point cloud encoder based on a visual Transformer, a multi-view point cloud interactive decoder based on a bidirectional visual Transformer, and a global point cloud prediction head; the coordinate transformation model uses the global coordinate system of the first group of local geometric structure local coordinate point clouds as a reference, and transforms the second group of local geometric structure local coordinate point clouds into the global coordinate system to generate a second group of local geometric structure global coordinate point clouds. Subsequently, the coordinate transformation model uses the global coordinates of the selected preceding local geometric structure global coordinate point cloud as a reference, and transforms the local coordinates of the current group of local geometric structure local coordinate point clouds into global coordinates, thereby transforming the local geometric structure local coordinate point clouds of all views into the global coordinate system.

9. The GIS digital twin scene construction method based on video mapping according to claim 8 is characterized in that: When selecting a global coordinate system reference from the global coordinate point cloud of the local geometry of the predecessor group, a confidence threshold matrix is ​​used to filter out predecessor views whose corresponding local coordinate prediction confidence matrices perform poorly.

10. The GIS digital twin scene construction method based on video mapping according to claim 3 is characterized in that: The coordinate transformation model is trained using the global coordinates of the real scene point cloud, and the second loss function for training the coordinate transformation model is: ,in, Video Frame The global coordinates of the corresponding real scene point cloud, Video frames predicted by the coordinate transformation model The global coordinates of the local coordinate point cloud corresponding to the local geometry structure, It means to find the L1 distance between the global coordinates of the two. Represents the global coordinate prediction confidence matrix The entropy of the coordinate transformation model is trained to minimize the second loss function.

Citation Information

Patent Citations

  • Method and system for fusing multi-channel continuous video and three-dimensional twinborn scene on traffic road

    CN117152400A

  • Digital twinborn scene intelligent generation method based on multi-modal visual identification

    CN117456136A

  • Digital twinborn scene geometric modeling method and device

    CN118710846A

  • Airport terminal video digital twinning method, device, equipment and medium

    CN119135849A

  • Method and apparatus for providing spatial information using digital twin based images

    KR102767908B1