A method for constructing a GIS digital twin scene based on video mapping
Through the multi-view feature interaction combining local geometric modeling and deep learning on sliding window, the problem of local and global coordinate system separation and dynamic scene consistency in GIS digital twin scene construction is solved, and high-precision scene digitization is achieved, suitable for smart cities and autonomous driving and other fields.
Patent Information
- Application Number
- CN202510571167.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-06
AI Technical Summary
When existing methods map real physical scenes into GIS digital twin scene models, there are problems such as local and global coordinate systems separation, difficulty in modeling geometric consistency of dynamic scenes, and bottlenecks in cross-modal mapping accuracy, resulting in misalignment or holes during point cloud splicing.
Sliding window local geometric modeling is used to interact with multi-view features driven by deep learning. Through multi-view image-to-point cloud model and coordinate transformation model, combined with rotational position coding and bidirectional Transformer interactive decoding mechanism, geometric associations within video slices are explicitly modeled, and the local geometric structure is gradually transformed to the global coordinate system.
It realizes the construction of high-precision and high-efficiency GIS digital twin scenarios, improves the accuracy of local geometric structure reconstruction and the stability of global splicing, and is suitable for real-time digitization of dynamic environments and large-scale scenarios.
Smart Images

Figure CN120107866B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of GIS digital twin scene construction, and in particular to a method for constructing a GIS digital twin scene based on video mapping. Background Art
[0002] With the wide application of digital twin technology in fields such as smart cities, industrial inspection, and autonomous driving, how to efficiently and accurately map the real physical scene into a GIS digital twin scene model has become a key technical challenge.
[0003] In recent years, deep learning technologies (such as vision transformers) have shown potential in multi-view geometric modeling and can implicitly model the geometric associations between multi-views through end-to-end training. However, the existing methods still have the following problems: Disconnection between local and global coordinate systems: Most methods directly map multi-view images to the global coordinate system, but the continuous changes in local geometric structures in the video stream can easily lead to uncertainties in global coordinate prediction. Geometric consistency in dynamic scenes: It is difficult to model the spatial associations of local geometric structures between video frames, resulting in misalignment or holes during point cloud stitching. Precision bottleneck in cross-modal mapping: The mapping from images to point clouds relies on manually designed features or fixed depth estimation networks and is difficult to adapt to the geometric diversity of complex scenes.
[0004] To address the above problems, the present invention proposes a method for constructing a GIS digital twin scene based on video mapping, which combines sliding window local geometric modeling with deep learning-driven multi-view feature interaction to achieve high-precision and high-efficiency scene digitization. Summary of the Invention
[0005] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method for constructing a GIS digital twin scene based on video mapping.
[0006] In a first aspect, the present invention provides a method for constructing a GIS digital twin scene based on video mapping, including:
[0007] Scanning and collecting a video of a real scene;
[0008] A sliding window of a set length divides the real scene video into multiple video slices according to a set step size;
[0009] Determining the middle video frame of each video slice, and defining the coordinate system in which the local geometric structure is located in the middle video frame as the local coordinate system of the video slice, where the local coordinate system of the video slice divided by the first sliding window is used as the global coordinate system;
[0010] Input any video slice into a pre-trained multi-view image to point cloud model. The multi-view image to point cloud model generates local coordinate point clouds of corresponding local geometric structures based on the video slice. The local coordinate point clouds of the local geometric structures use the coordinates of the local coordinate system of the video slice.
[0011] Record the local coordinate point clouds of the local geometric structures corresponding to all video frames. Take the local coordinate point clouds of the local geometric structures as input and input them into a coordinate transformation model. The coordinate transformation model transforms the point cloud coordinates of all the local coordinate point clouds of the local geometric structures into the global coordinate system to obtain local geometric structure global coordinate point clouds.
[0012] Stitch and combine the local geometric structure global coordinate point clouds together according to the global coordinates to obtain a GIS digital twin scene.
[0013] Furthermore, the local geometric structures in each video frame of the real scene video change continuously with the continuous change of the real scene video time axis.
[0014] Furthermore, the multi-view image to point cloud model includes: an image embedding layer, a position encoder, a multi-view image encoder based on Vision Transformer, a multi-view image interaction decoder based on bidirectional Vision Transformer, and a local point cloud prediction head.
[0015] The image embedding layer processes the video frame to obtain video frame patch embeddings. The position encoder uses rotational position encoding to add rotational position encoding to the video frame patch embeddings to obtain video frame features. For each video slice, the Vision Transformer of the multi-view image encoder encodes the video frame features of the video frames in the video slice to obtain encoded video frame features representing the local geometric structures in the video frames of the video slice and encoded position features representing the local geometric structures of the video frames in the local coordinate system of the video slice. The multi-view image interaction decoder exchanges the encoded video frame features and encoded position features between the intermediate video frames and non-intermediate video frames through multiple layers of bidirectional Vision Transformer to obtain multi-layer decoded video frame features and decoded position features for the intermediate video frames and each non-intermediate video frame. The local point cloud prediction head fuses the multi-layer decoded video frame features and decoded position features through a feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidences. Take the local coordinate point clouds of the local geometric structures as input and input them into a coordinate transformation model. The coordinate transformation model transforms the point cloud coordinates of all the local coordinate point clouds of the local geometric structures into the global coordinate system to obtain local geometric structure global coordinate point clouds.
[0016] Furthermore, the multi-view image encoder includes multiple layers of vision transformers. In each layer of the vision transformer, the vision transformer performs self-attention on the video frame features of any video frame, and performs cross-attention on the video frame features after self-attention of this video frame and the video frame features of other video frames respectively to obtain multiple groups of combined video frame features. The multiple groups of combined video frame features are aggregated through a pooling operation to obtain aggregated video frame features; the aggregated video frame features and the video frame features of this video frame after self-attention are added and combined and then normalized and input into the feed-forward network of the residual structure for further processing to obtain the combined video slice's overall local geometric structure information features, including: encoded video frame features and encoded position features of the video frame local geometry in the local coordinate system of the video slice.
[0017] Furthermore, the encoded video frame features and encoded position features of each video frame are mapped by a linear layer to the multi-view image interaction decoder. The multi-layer bidirectional vision transformer of the multi-view image interaction decoder has the same structure as the vision transformer of the multi-view image encoder, but exchanges the positions of the encoded video frame features and encoded position features of the intermediate video frame and non-intermediate video frames during calculation to achieve two-way information exchange and obtain the decoded video frame features and decoded position features of multiple layers of the intermediate video frame and each non-intermediate video frame.
[0018] Furthermore, the local point cloud prediction head fuses the multiple layers of encoded video frame features and encoded position features through a feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidence levels; the prediction head includes: a convolutional layer, an interpolation upsampling layer, a convolutional layer, an activation function layer, and a convolutional layer; the pixel-by-pixel local point cloud coordinates are combined to form the local coordinate point cloud of the local geometric structure.
[0019] Furthermore, the multi-view image to point cloud model is trained using the local coordinates of the real-scene point cloud. The first loss function for training the multi-view image to point cloud model is: , where video frame corresponding local coordinates of the real-scene point cloud, is the local coordinate of the local coordinate point cloud of the local geometric structure predicted by the multi-view image to point cloud model for video frame , represents the L1 distance between the predicted local coordinates and the real local coordinates, represents the entropy of the local coordinate prediction confidence matrix , and the parameters of the multi-view image to point cloud model are trained and adjusted to minimize the first loss function.
[0020] Further, the coordinate transformation model has the same structure as the multi-view image to point cloud model, including: a point cloud embedding layer, a position encoder, a multi-view point cloud encoder based on Vision Transformer, a multi-view point cloud interaction decoder based on bidirectional Vision Transformer, and a global point cloud prediction head; the coordinate transformation model takes the global coordinate system where the local coordinate point cloud of the first set of local geometric structures is located as a reference, converts the local coordinate point cloud of the second set of local geometric structures to the global coordinate system to generate the global coordinate point cloud of the second set of local geometric structures. Subsequently, the coordinate transformation model takes the global coordinates of the previously selected global coordinate point cloud of the local geometric structures as a reference, converts the local coordinates of the local coordinate point cloud of the current set of local geometric structures into global coordinates, so as to convert the local coordinate point clouds of the local geometric structures of all views into the global coordinate system.
[0021] Further, when selecting the global coordinate system reference from the global coordinate point clouds of the local geometric structures of the previous group, a confidence threshold matrix is used to filter out the previous views with poor performance in the corresponding local coordinate prediction confidence matrix.
[0022] Further, the coordinate transformation model is trained using the global coordinates of the real-scene point cloud, and the second loss function for training the coordinate transformation model is: , where video frame the global coordinates of the corresponding real-scene point cloud, is the global coordinate of the local coordinate point cloud of the local geometric structure predicted by the coordinate transformation model for the video frame , denotes the L1 distance between the two, denotes the entropy of the global coordinate prediction confidence matrix , and the parameters of the coordinate transformation model are trained and adjusted to minimize the second loss function.
[0023] In a second aspect, the present invention provides a GIS digital twin scene construction device based on video mapping, including: at least one processing unit, which interconnects the processing unit, the storage unit, and the acquisition unit through a bus unit. The storage unit stores computer programs and the data acquired by the acquisition unit. When the computer program is executed by the processing unit, the GIS digital twin scene construction method based on video mapping as described above is implemented.
[0024] In a third aspect, the present invention provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, the GIS digital twin scene construction method based on video mapping as described above is implemented.
[0025] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:
[0026] This application slices the scanned real - scene video into multiple video slices; determines the middle video frame of each video slice, and defines the local coordinate system of the video slice according to the middle video frame. The local coordinate system is used for the multi - view image - to - point - cloud model to construct the local geometric structure local coordinate point cloud. Among them, the local coordinate system of the first video slice sliced by the sliding window is used as the global coordinate system. The global coordinate system is used for the coordinate transformation model to transform the point cloud coordinates of the local geometric structure local coordinate point cloud to the global coordinate system to obtain the local geometric structure global coordinate point cloud; inputs any video slice into the pre - trained multi - view image - to - point - cloud model. The multi - view image - to - point - cloud model produces the local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, and the local geometric structure local coordinate point cloud uses the coordinates of the local coordinate system of the video slice; records the local geometric structure local coordinate point cloud corresponding to all video frames, and takes the local geometric structure local coordinate point cloud as the input and inputs it into the coordinate transformation model. The coordinate transformation model transforms the point cloud coordinates of all the local geometric structure local coordinate point clouds to the global coordinate system to obtain the local geometric structure global coordinate point cloud; stitches and combines each local geometric structure global coordinate point cloud together according to the global coordinates to obtain the GIS digital - twin scene. This application dynamically binds the video slice and the local coordinate system: slices the video stream into video slices with spatial correlation through a sliding window, and uses the middle frame as the benchmark of the local coordinate system to ensure the continuity and consistency of the local geometric structure.
[0027] This application implements a Transformer architecture for multi - view image point - cloud conversion: designs a vision Transformer model including rotational position encoding and multi - view cross - attention mechanism to explicitly model the geometric correlation of multiple frames within the video slice and generate a local geometric structure local coordinate point cloud with high confidence. Progressive stitching strategy for global coordinate transformation: through a confidence - guided coordinate transformation model, gradually aligns the local point cloud to the global coordinate system, avoids cumulative errors, and improves the mapping robustness of large - scale scenes. This method deeply integrates the real - time performance of SLAM and the high - precision advantages of deep learning at the technical level, provides an extensible and end - to - end solution for the construction of GIS digital - twin scenes, and is especially suitable for the real - time digitalization requirements of dynamic environments and large - scale scenes. By introducing rotational position encoding and bidirectional Transformer interactive decoding mechanism, it significantly improves the accuracy of local geometric structure reconstruction and the stability of global stitching, providing a new technical path for the construction of high - precision maps in fields such as smart cities and autonomous driving. Description of the Drawings
[0028] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a flowchart of a method for constructing a GIS digital twin scene based on video mapping provided by an embodiment of the present invention;
[0031] Figure 2 It is an architecture diagram of the overall model provided by an embodiment of the present invention;
[0032] Figure 3 It is a schematic diagram of a multi-view image encoder provided by an embodiment of the present invention;
[0033] Figure 4 It is a schematic diagram of the principle of a multi-view image interaction decoder provided by an embodiment of the present invention;
[0034] Figure 5 It is a schematic diagram of a local point cloud prediction head provided by an embodiment of the present invention;
[0035] Figure 6 It is a schematic diagram of a device for constructing a GIS digital twin scene based on video mapping provided by an embodiment of the present invention. Detailed Embodiments
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0037] It should be noted that in this document, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the said element.
[0038] Embodiment 1
[0039] As Figure 1As shown in the figure, the technology of the present invention realizes a method for constructing a GIS digital twin scene based on video mapping. The present invention includes the following steps:
[0040] Scan and collect the real scene video . Scanning enables the real scene video The video frames of have the following properties: the local geometric structures of the real scene in the video frames change continuously with the continuous change of the time axis of the real scene video , that is, there is no spatial jump in the local geometric structure in the real scene video .
[0041] A sliding window of a set length divides the real scene video into multiple video slices according to a set step size , where is the total number of video slices of the real scene video ; all the local geometric structures of the video frames in any video slice are spatially correlated.
[0042] Determine the middle video frame of each video slice , and define the coordinate system in which the local geometric structure in the middle video frame is located as the local coordinate system of the video slice. Among them, the local coordinate system of the video slice divided by the first sliding window is used as the global coordinate system.
[0043] Input any video slice into a pre-trained multi-view image to point cloud model. The multi-view image to point cloud model generates the local coordinate point cloud of the corresponding local geometric structure based on the video slice. The coordinates of the local geometric structure local coordinate point cloud are referenced by the corresponding local coordinate system of the video slice.
[0044] In the specific implementation process, the multi-view image to point cloud model includes: an image embedding layer, a position encoder, a multi-view image encoder based on a vision Transformer, a multi-view image interaction decoder based on a bidirectional vision Transformer, and a local point cloud prediction head.
[0045] The principle of the multi-view image to point cloud model is as follows: the image embedding layer processes each video frame of the video slice to obtain the video frame patch embedding. The position encoder uses rotational position encoding to add rotational position encoding to the video frame patch embedding to obtain the video frame feature. For the jth video frame in the ith video slice , its video frame feature is expressed as . The superscript j in the upper right represents the index of the video frame in the video slice. Among them, when the superscript is mid, it represents the middle video frame.
[0046] For each video slice, the vision Transformer of the multi-view image encoder encodes the video frame features of the video frames in the video slice to obtain the encoded video frame features representing the local geometric structure in the video frames of the video slice and the encoded position features representing the local geometric structure of the video frame in the local coordinate system of the video slice.
[0047] In the specific implementation process, the multi-view image encoder includes multiple layers of vision Transformer.
[0048] In each layer of the vision Transformer, first, the vision Transformer performs self-attention on the video frame features of any input video frame, and performs cross-attention on the video frame features after self-attention of this video frame with the video frame features of other video frames in the video slice where it is located, to obtain multiple groups of combined video frame features. The multiple groups of combined video frame features are aggregated through a pooling operation to obtain the aggregated video frame features of this video frame; the aggregated video frame features of this video frame and the video frame features of this video frame after self-attention are added and combined and then normalized, and input into the feed-forward network of the residual structure for further processing. The feed-forward network outputs the features of this video frame that combines the overall local geometric structure information within the video slice, including: the encoded video frame features and the encoded position features of the local geometric structure of the video frame in the local coordinate system of the video slice. The final vision Transformer of the multi-view image encoder obtains the encoded video frame features representing the local geometric structure in the video frames of the video slice and the encoded position features representing the local geometric structure of the video frame in the local coordinate system of the video slice.
[0049] Taking the middle video frame as an example: The vision Transformer performs self-attention on the video frame features of the middle video frame:
[0050] ;
[0051] Among them, represents normalization, represents self-attention, and the principle of self-attention is as follows:
[0052] The video frame features of the middle video frame normalized through linear mapping are used to obtain its self-attention query, self-attention key, and self-attention value: ;
[0053] Self-attention represents the formula as follows:
[0054] ;
[0055] Among them, is the softmax function.
[0056] Secondly, the Vision Transformer performs cross-attention on the video frame features of the intermediate video frames after self-attention respectively with the video frame features of each non-intermediate video frame:
[0057] ;
[0058] Among them, represents cross-attention, is the rotational position encoding of the video frame features of the intermediate video frame and the non-intermediate video frame; the principle of cross-attention is as follows:
[0059] The video frame features of the intermediate video frame after self-attention normalized by linear mapping are used to obtain its cross-attention query ; the video frame features of the non-intermediate video frame normalized by linear mapping are used to obtain its cross-attention key and cross-attention value ;
[0060] The cross-attention is represented by the following formula:
[0061] ;
[0062] Among them, represents the pytorch's cross_attn.rope(,) function that applies rotational position encoding to the query of cross-attention.
[0063] Through cross-attention, the video frame features of the intermediate video frame query features from the video frame features of each non-intermediate video frame, and multiple sets of combined video frame features are obtained by combining the features of each non-intermediate video frame ;
[0064] Then, the Vision Transformer aggregates the multiple sets of combined video frame features through pooling operations to obtain the aggregated video frame features of the intermediate video frame;
[0065] Finally, the aggregated video frame features of the intermediate video frame and the video frame features of the intermediate video frame after self-attention are added and combined and then normalized and input into the feed-forward network of the residual structure for further processing.
[0066] The processing processes of other video frames are the same, so that the encoded features of each video frame in the video slice are combined with the information of other video frames. The encoding process of the multi-view image encoder actually uses the relevant local geometric structure features in the video slice to support the encoding of the local geometric structure features of the video frames in any video slice, ensuring the certainty and accuracy of the local geometric structure information represented by each video frame encoding.
[0067] Such asFigure 4 As shown, the encoded video frame features and encoded position features of each video frame are mapped by a linear layer to a multi-view image interaction decoder. The multi-view image interaction decoder exchanges the encoded video frame features and encoded position features of intermediate video frames and non-intermediate video frames through a multi-layer bidirectional vision Transformer. The structure of the multi-layer bidirectional vision Transformer is the same as that of the vision Transformer of the multi-view image encoder. However, when exchanging the encoded video frame features and encoded position features of intermediate video frames and non-intermediate video frames, bidirectional information exchange is achieved through bidirectional cross-attention, obtaining the decoded video frame features and decoded position features of multiple layers for intermediate video frames and each non-intermediate video frame. Through this process, the spatial relationship between non-intermediate video frames and intermediate video frames is modeled, providing an accurate relative position association for generating the local coordinate point cloud of the local geometric structure subsequently.
[0068] As Figure 5 shown, the local point cloud prediction head fuses the multi-layer decoded video frame features and decoded position features through a feature pyramid, and then the prediction head maps the fused features into per-pixel local point cloud coordinates and confidence. The prediction head includes: a convolutional layer, an interpolation upsampling layer, a convolutional layer, an activation function layer, and a convolutional layer. The per-pixel local point cloud coordinates form the local coordinate point cloud of the local geometric structure. The local point cloud prediction head generates the local coordinate point cloud of the local geometric structure based on the spatial position relationship between non-intermediate video frames and intermediate video frames.
[0069] To enable the multi-view image to point cloud model to achieve the above functions, the local coordinates of the real-scene point cloud are used for training. The training loss is determined by the distance between the predicted point cloud points weighted by the local coordinate prediction confidence and the local coordinates of the real-scene point cloud and the entropy of the local coordinate prediction confidence. The first loss function for training the multi-view image to point cloud model is: , where the video frame corresponding local coordinates of the real-scene point cloud, is the local coordinate of the local coordinate point cloud of the local geometric structure predicted by the multi-view image to point cloud model for the video frame , denotes calculating the L1 distance between the predicted local coordinates and the real local coordinates, denotes calculating the entropy of the local coordinate prediction confidence matrix . Minimizing the first loss function means making the predicted local coordinate point cloud of the local geometric structure as close as possible to the real-scene point cloud and making the model prediction results more deterministic.
[0070] Record the local coordinate point cloud of the local geometric structure corresponding to all video frames.
[0071] Taking the local coordinate point cloud of the local geometric structure as input and inputting it into the coordinate transformation model, the coordinate transformation model transforms the point cloud coordinates of the entire geometric structure point cloud into the global coordinate system. The structure of the coordinate transformation model is the same as that of the multi-view image to point cloud model, and correspondingly includes: a point cloud embedding layer, a position encoder, a multi-view point cloud encoder based on the Vision Transformer, a multi-view point cloud interaction decoder based on the bidirectional Vision Transformer, and a global point cloud prediction head; the coordinate transformation model takes the global coordinate system where the first set of local geometric structure local coordinate point clouds is located as a reference, transforms the second set of local geometric structure local coordinate point clouds into the global coordinate system to generate the second set of local geometric structure global coordinate point clouds. Subsequently, the coordinate transformation model takes the global coordinates of the selected previous local geometric structure global coordinate point clouds as a reference and converts the local coordinates of the current set of local geometric structure local coordinate point clouds into global coordinates. Among them, when selecting the global coordinate system reference from the previous set of local geometric structure global coordinate point clouds, the confidence threshold matrix is used to filter out the previous views with poor performance in the corresponding local coordinate prediction confidence matrix. Through the above process, the coordinate transformation model transforms the local coordinate point clouds of the local geometric structures of all views into the global coordinate system.
[0072] The multi-view point cloud encoder has the same structural function as the multi-view image encoder, and the processing object is the local coordinate point cloud of the local geometric structure in the point cloud slice divided by the same sliding window.
[0073] The multi-view point cloud interaction decoder has the same structural function as the multi-view image interaction decoder, and the processing objects are the local coordinate point cloud of the local geometric structure that has been converted into the local geometric structure global coordinate point cloud and the local coordinate point cloud of the local geometric structure that has not been converted into the local geometric structure global coordinate point cloud, and performs bidirectional information exchange decoding between any two combinations.
[0074] The global point cloud prediction head has the same structure and function as the local point cloud prediction head. Both generate point cloud coordinates based on the spatial relationship.
[0075] It can be seen that only the object is programmed from the video frame to the point cloud frame, and the coordinate reference changes from the local coordinate system to the global coordinate system. The working principles of the multi-view image to point cloud model and the coordinate transformation model are the same. Therefore, the two have the same model structure but different model parameters.
[0076] To enable the coordinate transformation model to implement the above functions, end-to-end training is performed using the global coordinates of the real-scene point cloud. The training loss is determined by the distance between the predicted point cloud points weighted by the global coordinate prediction confidence and the global coordinates of the real-scene point cloud and the entropy of the global coordinate prediction confidence. The second loss function for training the coordinate transformation model is: , where video frame The global coordinates of the corresponding real - world scene point cloud, are the video frames predicted by the coordinate transformation model The global coordinates of the global coordinate point cloud of the corresponding local geometric structure, indicating to calculate the L1 distance between all the coordinates of the two, indicating to calculate the entropy of the global coordinate prediction confidence matrix The entropy of, training to adjust the parameters of the coordinate transformation model to minimize the second loss function, that is, to make the global coordinates of the predicted local geometric structure global coordinate point cloud as close as possible to the global coordinates of the real - world scene point cloud, and to make the prediction results of the coordinate transformation model more deterministic.
[0077] Stitch and combine the global coordinate point clouds of each local geometric structure together according to the global coordinates to obtain the GIS digital twin scene.
[0078] Embodiment 2
[0079] Refer to Figure 6 As shown, the embodiment of the present invention provides a GIS digital twin scene construction device based on video mapping, including: at least one processing unit, which interconnects the processing unit, the storage unit, and the acquisition unit through a bus unit. The storage unit, as a computer - readable storage medium, can be used to store software programs, computer - executable programs, and modules, such as the software programs, computer - executable programs, and modules corresponding to a GIS digital twin scene construction method based on video mapping in the embodiment of the present invention. The processing unit realizes the above - mentioned GIS digital twin scene construction method based on video mapping by running the software programs, computer - executable programs, and modules stored in the storage unit, including:
[0080] Scan and acquire the real - world scene video;
[0081] A sliding window of a set length divides the real - world scene video into multiple video slices according to a set step size;
[0082] Determine the middle video frame of each video slice, and define the coordinate system where the local geometric structure is located in the middle video frame as the local coordinate system of the video slice. Among them, the local coordinate system of the video slice divided by the first sliding window is used as the global coordinate system;
[0083] Input any video slice into a pre - trained multi - view image - to - point - cloud model. The multi - view image - to - point - cloud model generates the local coordinate point cloud of the corresponding local geometric structure based on the video slice, and the local coordinate point cloud of the local geometric structure uses the coordinates of the local coordinate system of the video slice;
[0084] Record the local geometric structure local coordinate point cloud corresponding to all video frames, and use the local geometric structure local coordinate point cloud as input to the coordinate transformation model. The coordinate transformation model transforms the point cloud coordinates of all the local geometric structure local coordinate point clouds to the global coordinate system to obtain the local geometric structure global coordinate point cloud;
[0085] Stitch and combine the local geometric structure global coordinate point clouds of each part together according to the global coordinates to obtain the GIS digital twin scene.
[0086] Of course, the computer program stored in the storage unit of a GIS digital twin scene construction device based on video mapping provided by the embodiments of the present invention is not limited to the method operations described above, and can also execute the related operations in a method for constructing a GIS digital twin scene based on video mapping provided by any embodiment of the present invention.
[0087] Embodiment 3
[0088] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which when executed, implements the method for constructing a GIS digital twin scene based on video mapping, including:
[0089] Scan and collect the real scene video;
[0090] A sliding window of a set length divides the real scene video into multiple video slices according to a set step size;
[0091] Determine the middle video frame of each video slice, and define the coordinate system in which the local geometric structure is located in the middle video frame as the local coordinate system of the video slice. Among them, the local coordinate system of the video slice divided by the first sliding window is used as the global coordinate system;
[0092] Input any video slice into a pre-trained multi-view image to point cloud model. The multi-view image to point cloud model generates the local geometric structure local coordinate point cloud of the corresponding local geometric structure based on the video slice, and the local geometric structure local coordinate point cloud uses the coordinates of the local coordinate system of the video slice;
[0093] Record the local geometric structure local coordinate point cloud corresponding to all video frames, and use the local geometric structure local coordinate point cloud as input to the coordinate transformation model. The coordinate transformation model transforms the point cloud coordinates of all the local geometric structure local coordinate point clouds to the global coordinate system to obtain the local geometric structure global coordinate point cloud;
[0094] Stitch and combine the local geometric structure global coordinate point clouds of each part together according to the global coordinates to obtain the GIS digital twin scene.
[0095] A computer-readable storage medium provided by an embodiment of the present invention, the computer program stored therein is not limited to the method operations described above, and can also execute related operations in a method for constructing a GIS digital twin scenario based on video mapping provided by any embodiment of the present invention.
[0096] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of structures or units can be in electrical, mechanical or other forms.
[0097] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0099] The above are only specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for constructing a GIS digital twin scenario based on video mapping, characterized in that, Including: Scanning and collecting real-scene videos; Using a sliding window of a set length to slice the real-scene video into multiple video slices according to a set step size; Determining the middle video frame of each video slice, and defining the coordinate system where the local geometric structure is located in the middle video frame as the local coordinate system of the video slice. Among them, the local coordinate system of the video slice sliced by the first sliding window is used as the global coordinate system; Inputting any video slice into a pre-trained multi-view image to point cloud model. The multi-view image to point cloud model generates local coordinate point clouds of the corresponding local geometric structure based on the video slice, and the local coordinate point clouds of the local geometric structure use the coordinates of the local coordinate system of the video slice; Recording the local coordinate point clouds of the local geometric structure corresponding to all video frames, and taking the local coordinate point clouds of the local geometric structure as input and inputting them into a coordinate transformation model. The coordinate transformation model transforms the point cloud coordinates of all local coordinate point clouds of the local geometric structure into the global coordinate system to obtain local coordinate point clouds of the local geometric structure in the global coordinate system; Stitching and combining the local coordinate point clouds of each local geometric structure in the global coordinate to obtain a GIS digital twin scene.
2. The method for constructing a GIS digital twin scene based on video mapping according to claim 1, wherein The local geometric structures in each video frame of the real-scene video change continuously with the continuous change of the time axis of the real-scene video.
3. The method for constructing a GIS digital twin scene based on video mapping according to claim 1, wherein, The multi-view image to point cloud model includes: an image embedding layer, a position encoder, a multi-view image encoder based on a Vision Transformer, a multi-view image interaction decoder based on a bidirectional Vision Transformer, and a local point cloud prediction head; The image embedding layer processes the video frame to obtain video frame patch embeddings. The position encoder uses rotational position encoding to add rotational position encoding to the video frame patch embeddings to obtain video frame features. For each video slice, the Vision Transformer of the multi-view image encoder encodes the video frame features of the video frames in the video slice to obtain encoded video frame features representing the local geometric structures in the video frames of the video slice and encoded position features representing the local geometric structures of the video frames in the local coordinate system of the video slice. The multi-view image interaction decoder exchanges the encoded video frame features and encoded position features between the middle video frame and non-middle video frames through multiple layers of bidirectional Vision Transformers to obtain multi-layer decoded video frame features and decoded position features of the middle video frame and each non-middle video frame. The local point cloud prediction head fuses the multi-layer decoded video frame features and decoded position features through a feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidences. Taking the local coordinate point clouds of the local geometric structure as input and inputting them into a coordinate transformation model, the coordinate transformation model transforms the point cloud coordinates of all local coordinate point clouds of the local geometric structure into the global coordinate system to obtain local coordinate point clouds of the local geometric structure in the global coordinate system.
4. The method for constructing a GIS digital twin scene based on video mapping according to claim 3, wherein, The multi-view image encoder includes multiple layers of vision transformers. In each layer of the vision transformer, the vision transformer performs self-attention on the video frame features of any video frame, and performs cross-attention on the video frame features after self-attention of this video frame and the video frame features of other video frames respectively to obtain multiple groups of combined video frame features. The multiple groups of combined video frame features are aggregated through a pooling operation to obtain aggregated video frame features; the aggregated video frame features and the video frame features of this video frame after self-attention are added and combined and then normalized and input into the feed-forward network of the residual structure for further processing to obtain the feature combining the overall local geometric structure information in the video slice, including: the encoded video frame features and the encoded position features of the video frame local geometric structure in the local coordinate system of the video slice.
5. The method for constructing a GIS digital twin scene based on video mapping according to claim 4, wherein The encoded video frame features and encoded position features of each video frame are mapped to the multi-view image interaction decoder through a linear layer. The multi-layer bidirectional vision transformer of the multi-view image interaction decoder has the same structure as the vision transformer of the multi-view image encoder, but during calculation, the positions of the encoded video frame features and encoded position features of the intermediate video frame and non-intermediate video frames are exchanged to achieve two-way information exchange to obtain the decoded video frame features and decoded position features of multiple layers of the intermediate video frame and each non-intermediate video frame.
6. The method for constructing a GIS digital twin scene based on video mapping according to claim 3, wherein The local point cloud prediction head fuses the multi-layer encoded video frame features and encoded position features through a feature pyramid, and then the prediction head maps the fused features into pixel-by-pixel local point cloud coordinates and confidence levels; the prediction head includes: a convolutional layer, an interpolation upsampling layer, a convolutional layer, an activation function layer, and a convolutional layer; the pixel-by-pixel local point cloud coordinates are combined to form the local coordinate point cloud of the local geometric structure.
7. The method for constructing a GIS digital twin scenario based on video mapping according to claim 3, wherein The multi-view image to point cloud model is trained using the local coordinates of the real-scene point cloud. The first loss function for training the multi-view image to point cloud model is as follows: , where video frame the local coordinates of the corresponding real-scene point cloud, is the local coordinate of the local geometric structure local coordinate point cloud predicted by the multi-view image to point cloud model for the video frame , represents the L1 distance between the predicted local coordinates and the real local coordinates, represents the entropy of the local coordinate prediction confidence matrix The parameters of the multi-view image to point cloud model are trained and adjusted to minimize the first loss function.
8. The method for constructing a GIS digital twin scene based on video mapping according to claim 3, wherein, The coordinate transformation model has the same structure as the multi-view image to point cloud model, including: a point cloud embedding layer, a position encoder, a multi-view point cloud encoder based on a vision transformer, a multi-view point cloud interaction decoder based on a bidirectional vision transformer, and a global point cloud prediction head; the coordinate transformation model takes the global coordinate system where the first group of local geometric structure local coordinate point clouds are located as a reference, and transforms the second group of local geometric structure local coordinate point clouds into the global coordinate system to generate the second group of local geometric structure global coordinate point clouds. Subsequently, the coordinate transformation model takes the global coordinates of the previously selected local geometric structure global coordinate point clouds as a reference, and converts the local coordinates of the current group of local geometric structure local coordinate point clouds into global coordinates, so as to transform the local geometric structure local coordinate point clouds of all views into the global coordinate system.
9. The method for constructing a GIS digital twin scenario based on video mapping according to claim 8, characterized in that, When selecting the global coordinate system reference from the previous group of local geometric structure global coordinate point clouds, the corresponding previous views with poor performance in the local coordinate prediction confidence matrix are filtered out using the confidence threshold matrix.
10. The method for constructing a GIS digital twin scene based on video mapping according to claim 3, wherein, The coordinate transformation model is trained using the global coordinates of the real-scene point cloud. The second loss function for training the coordinate transformation model is as follows: , where video frame the global coordinates of the corresponding real-scene point cloud, is the global coordinate of the local geometric structure local coordinate point cloud corresponding to the video frame predicted by the coordinate transformation model , denotes the L1 distance between the two global coordinates, denotes the entropy of the global coordinate prediction confidence matrix . The parameters of the coordinate transformation model are trained and adjusted to minimize the second loss function.
Citation Information
Patent Citations
Digital twinborn scene geometric modeling method and device
CN118710846A
Airport terminal video digital twinning method, device, equipment and medium
CN119135849A