Dynamic scene 4D semantic map generation method and device and processing equipment

By designing a feedforward framework for generating 4D semantic maps of dynamic scenes, we have achieved joint processing of geometric perception and semantic alignment, which solves the problems of high computational cost and poor scalability in existing technologies and improves the practicality and generalization ability of dynamic scene understanding.

CN121600201APending Publication Date: 2026-03-03JIANGHAN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511600760.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing 4D vision-language models are computationally expensive in dynamic scenes, difficult to scale and deploy, lack the feasibility of real-time applications, and cannot effectively handle open and time-sensitive query tasks.

Method used

Design a feedforward framework for generating dynamic scene 4D semantic maps, including a streaming visual geometry transformer and a semantic bridging decoder, to achieve joint processing of geometric perception and semantic alignment, support training by merging multiple dynamic scenes, and apply directly during inference.

Benefits of technology

It improves computational efficiency and generalization capabilities, significantly enhancing the practicality of large-scale deployments and supporting scenario understanding tasks for applications such as embodied intelligence, metaverse, and digital twins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600201A_ABST
    Figure CN121600201A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic scene 4D semantic map generation method, a dynamic scene 4D semantic map generation device and processing equipment, and aims to realize geometric perception and semantic alignment combined processing in a single framework by designing a first feedforward framework for 4D semantic map generation. The framework comprises two core components, namely a streaming visual geometric converter for capturing space-time geometric features of a dynamic scene and a semantic bridging decoder for mapping the space-time geometric features to language aligned semantic spaces, so that the structural integrity is kept, and the semantic interpretability is improved. Different from a traditional method depending on time-consuming scene-level optimization, the method can effectively support multi-dynamic scene merging training, can be directly applied during reasoning, and is high in calculation efficiency and generalization ability. According to the design, the practicability of large-scale deployment is remarkably improved, a new thought of open vocabulary 4D scene understanding is developed, good data support can be provided for scene understanding tasks of applications such as intelligence, meta universe and digital twinning, and the method has good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of map generation, specifically to a method, apparatus, and processing device for generating dynamic scene 4D semantic maps. Background Technology

[0002] Scene understanding has become a core capability for modern applications such as embodied intelligence, metaverse, and digital twins. While recent advances in 3D vision-language learning have demonstrated strong performance in static scenes, they still fall short when extended to dynamic 4D scenes where both geometry and semantic features are constantly evolving.

[0003] Unlike static environments, real-world scenarios require temporal consistency, semantic coherence, and cross-frame alignment to handle open and time-sensitive query tasks. Directly applying 3D methods often leads to semantic drift and alignment instability, prompting the search for solutions from 4D vision-language models.

[0004] Recent research has begun to explore extending scene representation to the language-guided 4D domain. However, the inventors of this application have found that most existing 4D vision-language model construction methods still heavily rely on Gaussian sputtering processes. Although Gaussian sputtering exhibits good performance in controlled scenes, its fundamental drawback lies in the need for explicit scene-by-scene optimization. This requirement brings several key limitations: high computational costs, difficulty in achieving cross-video scalability, and the need to maintain independent models for different scenes, making large-scale deployment difficult. More importantly, the reliance on scene-by-scene training fundamentally weakens the feasibility of real-time applications, while efficiency and generalization ability are indispensable requirements for real-time applications. These limitations highlight the urgent need to develop breakthrough solutions to go beyond scene-specific processing flows. Summary of the Invention

[0005] This application provides a method, apparatus, and processing device for generating 4D semantic maps of dynamic scenes. By designing the first feedforward framework for 4D semantic map generation, it achieves joint processing of geometric perception and semantic alignment within a single architecture. This framework comprises two core components: a streaming visual geometric transformer that captures the spatiotemporal geometric features of dynamic scenes and a semantic bridging decoder that maps these features to a language-aligned semantic space. This approach maintains structural integrity while enhancing semantic interpretability. Unlike traditional methods that rely on time-consuming scene-level optimization, this method effectively supports training by merging multiple dynamic scenes, allowing for direct application during inference. It boasts high computational efficiency and strong generalization ability. This design significantly improves its practicality for large-scale deployment, opens up new avenues for open-vocabulary 4D scene understanding, and can provide excellent data support for scene understanding tasks in applications such as embodied intelligence, metaverse, and digital twins, demonstrating promising application prospects.

[0006] Firstly, this application provides a method for generating a dynamic scene 4D semantic map, the method comprising: Obtain the current video sequence from which the corresponding 4D semantic map is to be generated; The current video sequence is input into the 4D semantic map generation network, which includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder processes the input video sequence into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. Extract and output a time-independent 4D semantic map and a time-dependent 4D semantic map of a dynamic scene.

[0007] Secondly, this application provides a dynamic scene 4D semantic map generation device, the device comprising: The acquisition unit is used to acquire the current video sequence from which the corresponding 4D semantic map is to be generated; The generation unit is used to input the current video sequence into the 4D semantic map generation network. The 4D semantic map generation network includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder is used to process the video sequence input to the network into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. The output unit is used to extract and output the dynamic scene time-independent 4D semantic map and the dynamic scene time-dependent 4D semantic map.

[0008] Thirdly, this application provides a processing device, including a processor and a memory, wherein a computer program is stored in the memory, and when the processor invokes the computer program in the memory, it executes the method provided by the first aspect of this application or any possible implementation of the first aspect of this application.

[0009] Fourthly, this application provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform the method provided in the first aspect of this application or any possible implementation thereof.

[0010] From the above, it can be concluded that this application has the following beneficial effects: To address the goal of generating 4D semantic maps for dynamic scenes, this application designs the first feedforward framework for 4D semantic map generation. This framework achieves joint processing of geometric perception and semantic alignment within a single architecture. It comprises two core components: a streaming visual geometric transformer that captures the spatiotemporal geometric features of dynamic scenes, and a semantic bridging decoder that maps these features to a language-aligned semantic space. This approach maintains structural integrity while enhancing semantic interpretability. Unlike traditional methods that rely on time-consuming scene-level optimization, this framework effectively supports training across multiple dynamic scenes, allowing for direct application during inference. It boasts high computational efficiency and strong generalization capabilities. This design significantly improves the practicality of large-scale deployment, opens up new avenues for open-vocabulary 4D scene understanding, and can provide robust data support for scene understanding tasks in applications such as embodied intelligence, metaverse, and digital twins, demonstrating promising application prospects. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating a method for generating a dynamic scene 4D semantic map according to this application. Figure 2 A schematic diagram of the processing logic for the 4D semantic map generation network reasoning task of this application; Figure 3 This is a schematic diagram of a processing logic for the streaming visual geometry transformer of this application. Figure 4 This is a schematic diagram of the processing logic of the semantic bridging decoder in this application; Figure 5 A schematic diagram of a processing logic for the training task of the 4D semantic map generation network in this application; Figure 6 This is a schematic diagram of the processing logic of the supervised semantic generator in this application; Figure 7 A schematic diagram illustrating the processing logic of the object query service involved in this application; Figure 8 This is a schematic diagram of a dynamic scene 4D semantic map generation device according to this application. Figure 9 This is a schematic diagram of one type of processing equipment used in this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved.

[0015] The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between modules shown or discussed may be through some interfaces, and the indirect coupling or communication connection between modules may be electrical or other similar forms, none of which are limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules may be selected to achieve the purpose of the solution in this application according to actual needs.

[0016] Before introducing the dynamic scene 4D semantic map generation method provided in this application, we will first introduce the background content involved in this application.

[0017] The dynamic scene 4D semantic map generation method, apparatus, and computer-readable storage medium provided in this application can be applied to processing devices. By designing the first feedforward framework for 4D semantic map generation, it achieves joint processing of geometric perception and semantic alignment in a single architecture. This framework includes two core components: a streaming visual geometric transformer that captures the spatiotemporal geometric features of dynamic scenes and a semantic bridging decoder that maps spatiotemporal geometric features to a language-aligned semantic space. This maintains structural integrity while improving semantic interpretability. Unlike traditional methods that rely on time-consuming scene-level optimization, this approach can effectively support multi-dynamic scene merging training, can be directly applied during inference, and has high computational efficiency and strong generalization ability. This design significantly improves its practicality for large-scale deployment, opens up new avenues for open-vocabulary 4D scene understanding, and can provide good data support for scene understanding tasks in applications such as embodied intelligence, metaverse, and digital twins, showing promising application prospects.

[0018] The dynamic scene 4D semantic map generation method mentioned in this application can be executed by a dynamic scene 4D semantic map generation device, or by different types of processing devices such as servers, physical hosts, or user equipment (UE) that integrate the dynamic scene 4D semantic map generation device. The dynamic scene 4D semantic map generation device can be implemented in hardware or software. The UE can be a terminal device such as a smartphone, tablet, laptop, desktop computer, or personal digital assistant (PDA). The processing devices can be configured in a device cluster.

[0019] It is understandable that the solution in this application is usually based on existing data or data that has already been collected. Therefore, the processing device that executes the dynamic scene 4D semantic map generation method of this application or is equipped with the application service corresponding to the dynamic scene 4D semantic map generation method of this application usually only needs to meet the required data processing capabilities, and its specific device type and device deployment form are quite flexible.

[0020] If the direct acquisition of existing data mentioned above is also involved, then further hardware and software adaptations are needed for the processing equipment to enable it to acquire data. For example, if real-time acquisition of video sequences (composed of video frames) is required, the video acquisition equipment can be incorporated into the processing equipment cluster, or the processing equipment itself can be the control unit of the video acquisition equipment. Alternatively, a third-party call can be used to trigger the video acquisition equipment outside the processing equipment to perform real-time data acquisition.

[0021] The acquisition of video sequences is mainly achieved through cameras / cameras. The corresponding video acquisition devices can be either standalone camera / camera accessories or related devices that include camera / camera accessories.

[0022] In addition, if there is a need to display the processing progress (including the processing results), the processing device itself can be configured with the required display screen (including touch screen) to display the specific content. Of course, the processing device can also display the specific content through an external display device or other devices with a display screen.

[0023] The following section introduces the dynamic scene 4D semantic map generation method provided in this application.

[0024] First, refer to Figure 1 , Figure 1 This paper illustrates a flowchart of the dynamic scene 4D semantic map generation method of this application. The dynamic scene 4D semantic map generation method provided by this application may specifically include the following steps S101 to S103: Step S101: Obtain the current video sequence from which the corresponding 4D semantic map is to be generated; Understandably, in response to the need for generating 4D semantic maps for dynamic scenes, the first operation that this application solution can perform is to obtain the current video sequence for which the corresponding 4D semantic map needs to be generated under the current circumstances.

[0025] The acquisition and processing of the current video sequence can be either real-time acquisition, extraction and processing of pre-acquired ready-made data, or further processing of the original video sequence based on real-time acquisition / pre-acquired acquisition.

[0026] The latter, in addition to some routine preprocessing operations to improve video quality, can also involve the generation of video sequences. This can be assisted by relevant artificial intelligence (AI) technologies. Similarly, the real-time acquisition and processing mentioned above can, in some cases, be directly generated by relevant AI technologies to meet the flexible and ever-changing application needs in real-world situations.

[0027] In practice, the proposed solution typically initiates the dynamic scene 4D semantic map generation process in the form of a work task. Correspondingly, it may also involve the processing of obtaining dynamic scene 4D semantic map generation tasks. These tasks can be initiated manually, received from other devices, or initiated autonomously according to a corresponding autonomous initiation strategy. All of these are possible.

[0028] For the current video sequence, when it involves the task of generating a dynamic scene 4D semantic map, it can be carried in the task information, or it can be extracted from the corresponding storage location according to the instructions of the task information, or it can be collected and processed in real time according to the instructions of the task information, or it can be entered manually.

[0029] Furthermore, in specific applications, the original intention of this application is mainly to address the demand for generating high-precision dynamic scene 4D semantic maps for scene understanding tasks involved in three major application types: embodied intelligence, metaverse, and digital twins. Of course, in practical situations, as long as there is a demand for generating high-precision dynamic scene 4D semantic maps based on video sequences, this application can actually meet the requirements. Through the 4D semantic map generation network architecture specifically / targeted by this application, dynamic scene 4D semantic maps are provided with a balance of low application cost, high processing efficiency, high processing accuracy, and good generalization.

[0030] Step S102: Input the current video sequence into the 4D semantic map generation network. The 4D semantic map generation network includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder is used to process the video sequence input to the network into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. Understandably, as recent research in existing technologies begins to explore extending scene representation to the language-guided 4D domain, this application shifts to a feedforward 4D geometry reconstruction paradigm to address the scalability challenges brought about by scene-by-scene optimization. Streaming visual geometry transformers can achieve efficient reconstruction without specific scene optimization, demonstrating excellent performance in maintaining real-time performance and generalization capabilities. However, these methods only focus on geometry and motion modeling and lack semantic alignment mechanisms, making it difficult to support the 4D understanding requirements of open vocabularies. This technological gap highlights the necessity of building a new generation of frameworks that can achieve joint modeling of geometry and semantics in a single architecture.

[0031] As can be seen, this application designs a novel 4D semantic map generation network architecture, which is the first feedforward framework for 4D semantic map generation. It realizes joint processing of geometric perception and semantic alignment in a single architecture. In practical applications or network inference work, it specifically involves three major network structures / components: streaming visual geometric transformer, semantic bridging decoder, and point cloud coloring module. The streaming visual geometric transformer is specifically divided into a streaming visual geometric transformer encoder for the encoding task and a streaming visual geometric transformer decoder for the decoding task.

[0032] In layman's terms, the input of a semantic map generation network is a sequence of color videos of a dynamic scene, and the output is a sequence of 3D point clouds with semantic features (a specific manifestation of a 4D semantic map). The goal of a 4D semantic map generation network is to establish a general mapping relationship between the input and output through the design of the 4D semantic map generation network and training across dynamic scenes, so as to achieve fast, accurate and direct reasoning.

[0033] See Figure 2 The diagram shown illustrates a processing logic for the 4D semantic map generation network inference task of this application, for a T-frame dynamic scene test video sequence that has already been processed by the input network. When performing specific reasoning tasks, the 4D semantic map generation network of this application has the following characteristics: 2.1) The streaming visual geometric transformation encoder converts a T-frame dynamic scene test video sequence Processed into camera token sequence and geometric token sequence ; Among them, the streaming visual geometric transformer itself is an existing network concept based on the Transformer architecture. The streaming visual geometric transform encoder can be understood as a real-time spatiotemporal feature extraction model. Its core is to realize dynamic scene geometric perception by using alternating spatial attention and temporal causal attention mechanisms. Correspondingly, the streaming visual geometric transform decoder is responsible for decoding the spatiotemporal features extracted by the encoder into 4D geometric reconstruction output.

[0034] The camera tokens in the camera token sequence can be understood as the encoded result of the camera pose (such as rotation matrix and translation vector) of the current frame. Geometry tokens in a sequence of geometry tokens can be understood as representing the 3D geometric properties of the scene in the current frame, including point clouds and depth maps.

[0035] 2.2) The streaming visual geometry decoder will use the camera token sequence and geometric token sequence Processed into a reconstructed point cloud sequence ; The streaming visual geometric transform decoder mentioned above is responsible for decoding the spatiotemporal geometric features extracted by the encoder into 4D geometric reconstruction output, after inputting the preceding streaming visual geometric camera token sequence. and geometric token sequence Then, the corresponding reconstructed point cloud sequence can be obtained. .

[0036] At this point, it becomes clear that the main function of the streaming visual geometry transformer is to capture the spatiotemporal geometric features of dynamic scenes and output 4D geometric reconstruction results.

[0037] 2.3) The semantic bridging decoder will convert the geometric token sequence Process into time-independent predicted semantic sequences Predicting semantic sequences related to time series ; The main function of the semantic bridging decoder is to process the spatiotemporal geometric features, i.e., the geometric token sequence, output by the streaming visual geometric transform encoder. This is mapped to a language-aligned semantic space, thus maintaining structural integrity while improving semantic interpretability.

[0038] The processing results correspond to both time-independent semantic levels and time-related semantic levels, and include time-independent predicted semantic sequences. Predicting semantic sequences related to time series .

[0039] 2.4) The point cloud coloring module uses time-independent prediction of semantic sequences. Predicting semantic sequences related to time series For the reconstructed point cloud sequence Coloring is performed to obtain a dynamic scene time-independent 4D semantic map. 4D semantic map related to dynamic scene time sequence .

[0040] The point cloud coloring module is used for the final output, specifically for predicting semantic sequences in a time-independent manner. Predicting semantic sequences related to time series As a semantic information reference, the reconstructed point cloud sequence output by the streaming visual geometric decoder is used. By performing coloring and assigning corresponding semantic information, a time-independent 4D semantic map of a dynamic scene can be obtained. 4D semantic map related to dynamic scene time sequence .

[0041] For time-independent 4D semantic maps of dynamic scenes 4D semantic map related to dynamic scene time sequence These two are the dynamic scene 4D semantic maps generated by the network, and specifically, they facilitate better viewing and use through the two types of maps they contain.

[0042] At the same time, it is understandable that the 4D semantic map generation network constructed in this application may involve a preliminary configuration stage before it is put into practical use, i.e., for inference tasks, which mainly involves network training.

[0043] In short, the specific parameters of the components in the 4D semantic map generation network are trained by using training samples labeled with corresponding generated results (true values), i.e., dynamic scene training video sequences.

[0044] It is important to note that during the training process, only a few components may be trained, while some components may not need to be trained. This involves situations where the parameter iteration optimization of different components is not completed in the same training stage, and also situations where some components directly use pre-trained components (i.e., they do not need to be trained for this application).

[0045] Therefore, the 4D semantic map generation network architecture specifically designed in this application has the following main advantages at the overall level: (1) The first transformer-based feedforward framework unifies video reconstruction and semantic alignment tasks in one network. Through deep interaction of shared network base, the performance of video reconstruction and semantic alignment can be mutually enhanced. (2) A semantic bridging decoder was introduced, which maps dynamic scene perception features to a language-aligned semantic space, effectively bridging the gap between geometric perception and semantic prediction.

[0046] Step S103: Extract and output the dynamic scene time-independent 4D semantic map and the dynamic scene time-related 4D semantic map.

[0047] It's easy to understand that after the 4D semantic map generation network completes the current map generation process, it will process the dynamic scene time-independent 4D semantic map obtained by the point cloud coloring module. 4D semantic map related to dynamic scene time sequence To output the results, it is obvious that both can be extracted and output at this point.

[0048] In terms of specific output operations, the configuration is adaptively tailored to the scenario understanding tasks involved in applications such as embodied intelligence, metaverse, and digital twins.

[0049] For example, it can be used for local storage, off-site storage, result display, result push, output completion prompts, further map use, or other data processing and analysis, which can be flexibly adjusted according to actual needs.

[0050] Next, we will provide a more detailed explanation of the three major network structures of the 4D semantic map generation network mentioned above.

[0051] exist Figure 3 Based on the schematic diagram of a processing logic of the streaming visual geometry transformer of this application, it can be seen that the streaming visual geometry transformer encoder in the streaming visual geometry transformer can specifically include a label-free self-distillation image encoding module E, a merging module, and an alternating attention transformation module D. Correspondingly, during operation, the following can be achieved: 3.1) Dynamic scene video frames at time t The initial camera token at time t is obtained by processing the image encoding module E using self-distillation with no labels (DINO). and image tokens ; 3.2) Initial camera token and image tokens By merging modules and Cache of time The token sequences in the buffer are merged to obtain the cache at time t. , among which, initial cache Empty; 3.3) Cache at time t The token sequence in the image generates camera tokens at time t through the Alternating-Attention (AA) transform module D. and geometric tokens .

[0052] As is understandable, the above describes the specific operations of the streaming visual geometric transform encoder in the point cloud reconstruction process within the streaming visual geometric transformer. For each time step, the current cached camera token is combined with... Update the process.

[0053] The entire process can be represented by the following formula: , , .

[0054] Next, continue to refer to Figure 3The streaming visual geometry transform decoder in the streaming visual geometry transformer specifically includes a depth head. Camera head Corresponding to the inverse projection module P, during operation, the following can be observed: 3.4) Camera token at time t Through the camera lens Processing to obtain predicted camera parameters ; 3.5) Geometric token at time t via depth head Processing yields a reconstructed depth map ; 3.6) Predict camera parameters and reconstructed depth map The reconstructed point cloud at time t is obtained by processing the data using the inverse projection module P. .

[0055] As is understandable, the above describes the specific operations of the streaming visual geometric transform decoder in the point cloud reconstruction process within the streaming visual geometric transformer. For each time step, it combines two tokens: the camera token and the [unclear - possibly a typo]. and geometric tokens To output the corresponding specific point cloud reconstruction results, you can also start from... Figure 2 This has been confirmed.

[0056] The entire process can be represented by the following formula: , , .

[0057] From the specific working logic of the streaming visual geometry transformer above, it can be seen that, compared with the existing method of using Gaussian splashing as the network base, this application has the following advantages in selecting the streaming visual geometry transformer as the network base: (1) The streaming visual geometry transformer is a feedforward neural network based on a standard large model architecture. It adopts an alternating spatial attention and temporal causal attention architecture to achieve fast inference speed, uses multi-objective joint modeling and overcomplete prediction strategies to achieve high geometric accuracy, and uses 13 carefully selected large-scale scene datasets for fine-tuning to achieve strong generalization ability. Therefore, the streaming visual geometry transformer can provide a solid geometric foundation for the application to achieve the goal of fast and accurate direct inference of 4D semantic maps from dynamic scene videos. By generating reliable geometric spatiotemporal representations, it provides strong support for subsequent semantic alignment.

[0058] (2) A streaming visual geometric transformation decoder was used in the inference stage. By utilizing the generated camera intrinsic and extrinsic parameters, the powerful spatiotemporal features were mapped back to the 4D point cloud space, ensuring that semantic information was accurately injected and aligned at the point cloud level.

[0059] Next, continue reading Figure 4 The diagram shown illustrates a processing logic of the semantic bridging decoder in this application. Specifically, for the semantic bridging decoder in a 4D semantic map generation network, it may include a Dense Prediction Transformer (DPT). Time-independent semantic header Time-related semantic header and video frame reconstruction head (This applies to the network training process; during inference, the component may not work, its processing results may be ignored, or the component may be removed.) Correspondingly, during the working process, the following may occur: 4.1) The geometric token at time t Through dense prediction transformer Process to obtain context token ; 4.2) Context token Through time-independent semantic headers Semantic Header Related to Time The temporally independent prediction semantics at time t are obtained by processing them separately. Semantic prediction related to time series .

[0060] In addition, corresponding to the network training processing scheme settings mentioned later, context tokens will also be used during network training. Reconstructing the head from video frames To process and obtain the reconstructed video frame at time t .

[0061] Understandably, the above describes the specific operations of the semantic bridging decoder in the 4D semantic map generation network during the semantic decoding process of a normal inference task. At each time step, it incorporates geometric tokens. To process and obtain time-independent prediction semantics Semantic prediction related to time series The semantic reference information for these two aspects can also be found here. Figure 2 This was confirmed, and the specific operations in the semantic decoding process of the pre-training task also involve reconstructing video frames. Reference information in this regard.

[0062] The entire process can be represented by the following formula: , , , , Wherein, the context token at time t Time-independent predictive semantics Temporal correlation prediction semantics Reconstructing video frames , This indicates the spatial resolution of both the context token map and the semantic map. denoted by , b, c, and d represent the spatial resolution of the reconstructed video frame, respectively, and represent the spatial dimensions of the dense prediction transform, the temporally independent semantic embedding, and the temporally related semantic embedding.

[0063] From the specific working logic of the semantic bridging decoder above, it can be seen that the semantic bridging decoder of this application has the following advantages: (1) The semantic bridging decoder uses a context-aware dense predictive transformer, which combines the spatial sensitivity of local convolutional operations with the global modeling capability of the transformer. This allows it to capture long-range dependencies in both spatial and temporal dimensions. By stacking self-attention layers, it transforms geometric tokens into representations rich in contextual features, significantly improving its semantic discrimination capability. Furthermore, it remains trainable during the network training or optimization process, which will be discussed later, enabling continuous optimization to meet the requirements of semantic tasks.

[0064] (2) The semantic head and reconstruction head are easy to implement. They can be implemented using a simple multilayer perceptron structure or a U-Net structure. Actual experiments on the public dataset HyperNeRF have shown that the latter has a specific improvement of about 1% and 2% in terms of time-independent and time-dependent performance metrics, respectively.

[0065] Next, we come to the network training task mentioned earlier, which is involved in the process of putting the network into practical use, i.e., inference tasks.

[0066] As mentioned in the introduction to the semantic bridging decoder above, the semantic bridging decoder may also include a video frame reconstruction header. Used to transfer context tokens Reconstructing the head from video frames Processing to obtain the reconstructed video frame at time t .

[0067] In this case, please refer to Figure 5 The diagram shown illustrates one processing logic of the 4D semantic map generation network training task of this application. It can be seen that the 4D semantic map generation network of this application may also include a supervised semantic generator that serves the training of the network.

[0068] Correspondingly, the pre-training of the 4D semantic map generation network in this application can include: 5.1) Train the T-frame dynamic scene video sequence The geometric token sequence is obtained by processing the weight-frozen streaming visual geometric transformation encoder. ; It can be noted that the streaming visual geometric transformation encoder involved in this application can be a pre-trained streaming visual geometric transformation encoder, which does not require special training processing, thus having the feature of weight freezing. The streaming visual geometric transformation decoder is similar, or in other words, the overall streaming visual geometric transformation uses a pre-trained streaming visual geometric transformation, without the need for iterative optimization of the corresponding weight parameters during network training.

[0069] 5.2) Geometric token sequence The reconstructed video sequence is obtained through semantic bridging decoder processing. Time-independent prediction semantic sequences Predicting semantic sequences related to time series ; 5.3) Train the T-frame dynamic scene video sequence A time-independent supervised semantic sequence is obtained through processing by a supervised semantic generator. Temporally related supervised semantic sequences ; 5.4) Training video sequences using T-frame dynamic scenes Reconstructing video sequences Time-independent prediction semantic sequences Temporal correlation prediction semantic sequence Temporally independent supervised semantic sequences Temporally related supervised semantic sequences The total loss function is calculated to perform multi-objective joint optimization of the weights of the semantic bridging decoder.

[0070] From the specific training logic of the 4D semantic map generation network above, it can be seen that the 4D semantic map generation network of this application also has the following advantages: (1) Compared with the existing four-dimensional semantic map generation method based on Gaussian splashing, it can be trained on multiple dynamic scenes and applied directly during inference without scene-by-scene optimization, making large-scale practical deployment possible.

[0071] (2) Using a pre-trained streaming visual geometric transformation encoder during the training phase generates powerful spatiotemporal representations suitable for diverse scenarios, while avoiding redundant optimization and reducing computational costs, allowing the training process to focus on semantic alignment rather than relearning the geometric model from scratch.

[0072] (3) Since the semantic bridging decoder uses dual-head semantic and reconstruction decoding, the context token is projected onto the complementary semantic and visual subspaces through independent prediction heads for overcomplete training, which can significantly improve prediction performance.

[0073] For supervised semantic generators specifically designed for network training, please refer to [link / reference]. Figure 6 The diagram shown illustrates one processing logic of the supervised semantic generator of this application. Specifically, the supervised semantic generator may include a video segmentation module S, a temporally independent supervised semantic generation branch, and a temporally related supervised semantic generation branch. The temporally independent supervised semantic generation branch includes a Contrastive Language-Image Pre-training (CLIP) model. Along with the first alignment module, the temporally related supervised semantic generation branch includes the Multimodal Large Language Model (MLLM). Large Language Model (LLM) Corresponding to the second alignment module, during the working process, the following can be observed: 6.1) Train the T-frame dynamic scene video sequence The video segmentation module S uses the decoupled video segmentation approach (DEVA) to segment the video, resulting in a mask sequence of N video objects. ; 6.2) Train the T-frame dynamic scene video sequence By comparing language image pre-trained models Obtain the mask sequence of N video objects Corresponding image object semantic embedding vector ; Where i is the video object identifier, and the total number is N.

[0074] 6.3) Embed time-independent semantics into the vector using the first alignment module. With the mask sequence of N video objects Align all pixels within the region to obtain a time-independent supervised semantic sequence. ; It is understandable that 6.2) and 6.3) are the processing of the temporally independent supervised semantic generation branch.

[0075] 6.4) Train the T-frame dynamic scene video sequence Through multimodal large language model Obtain the mask sequence of N video objects The corresponding dynamic description of the video object; 6.5) Dynamically describe video objects using a large language model Obtain the sequence of time-related semantic embedding vectors ; 6.6) Embed time-related semantics into the vector sequence through the second alignment module. With the mask sequence of N video objects Align all pixels within the region to obtain a temporally correlated supervised semantic sequence. .

[0076] It is understandable that 6.4), 6.5) and 6.6) are the processing of the temporally related supervised semantic generation branch.

[0077] The entire process can be represented by the following formula: , , , , , Where i represents the object index when the total number of video objects is N, and the object mask. The value is 1 if the pixel belongs to an object, and 0 otherwise. ; .

[0078] From the specific processing logic of the supervised semantic generator above, it can be seen that, compared with the supervised semantic generators of existing methods, the supervised semantic generator of this application has the following advantages: (1) The video segmentation module uses DEVA, an extended version of the Segment Anything Model (SAM), which has strong video temporal coherence and dynamic target tracking capabilities, laying a good foundation for subsequent semantic alignment.

[0079] (2) A dual-branch structure of time-independent supervised semantics and time-dependent supervised semantics is adopted. The former provides static object-level constraints, while the latter captures semantics that evolve over time. The synergistic effect of the two significantly improves the model's ability to align semantics.

[0080] Meanwhile, regarding the total loss function or joint loss function specifically used in the multi-objective joint optimization mentioned above, combined with... Figure 5 Specifically, it can be expressed as follows: , , , , , in, Let α be the total loss function, and α be the total loss function for video reconstruction. The corresponding weights, where β is the total semantic alignment loss function. The corresponding weights are α>0 and β>0. To characterize the structural loss of video frames at each time point during reconstruction Norm loss function and characterization of pixel-level smoothness loss in video frames The weighted balancing factor of the norm loss function. , Let be the total loss function for the first semantic alignment. The total loss function for the second semantic alignment. To characterize the semantic alignment magnitude loss at each time step The weights corresponding to the norm loss function The weights corresponding to the cosine similarity loss function, which characterizes the semantic alignment direction loss at each time step, are used to define the weights. , .

[0081] Specifically, the total loss Total loss from video reconstruction Total loss for semantic alignment The weighted average is calculated, where α and β are the total video reconstruction losses, respectively. Total loss for semantic alignment The corresponding loss weights are α>0 and β>0. Total loss during video reconstruction Structural loss reconstructed from video frames at each moment Norm loss function and characterization of pixel-level smoothness loss in video frames The norm loss function is obtained by summing the weighted equilibrium values. To characterize the structural loss of video frames at each time point during reconstruction Norm loss function and characterization of pixel-level smoothness loss in video frames The weighted balancing factor of the norm loss function. ; Semantic alignment total loss function The total loss function for the first semantic alignment Second semantic alignment total loss function The sum, where the total loss function for the first semantic alignment is... (Second semantic alignment total loss function) The first semantic alignment magnitude loss (second semantic alignment magnitude loss) is characterized by each time step. The norm loss function is obtained by weighting and summing the cosine similarity loss function, which characterizes the first semantic alignment direction loss (second semantic alignment direction loss). and These represent the first semantic alignment magnitude loss (second semantic alignment magnitude loss) at each time step. The weights corresponding to the norm loss function and the cosine similarity loss function that characterizes the first semantic alignment direction loss (second semantic alignment direction loss) are... , .

[0082] As an example, experiments on the public datasets HyperNeRF and Neu3D confirm that... The value is 0.5. The value is 0.2. A value of 0.01 can achieve good video reconstruction and semantic alignment performance.

[0083] Compared with existing semantic alignment single-objective optimization methods, the multi-objective joint optimization method designed in this application has the following characteristics: (1) Overall, the dual-objective optimization of video reconstruction and semantic alignment can comprehensively improve the accuracy of semantic map generation. Experiments on the public dataset HyperNeRF have shown that, regardless of whether the semantic prediction is time-independent or time-related, the newly added video reconstruction optimization objective can bring an improvement of 5% and 1%-2% in the intersection-over-union ratio and accuracy, respectively.

[0084] (2) For the video reconstruction part, a hybrid approach is used. The target can effectively control structural losses. Norm loss function and pixel-level smoothness loss The balance between norm loss functions.

[0085] (3) For the semantic alignment part, a dual supervision scheme of time-independent and time-dependent methods is adopted, which enables the model to learn both static and dynamic semantics at the object level through supercomplete training, thereby improving the alignment effect in dynamic scenes.

[0086] Next, we will use a specific map query application service in the actual application of this application to better understand the above solution.

[0087] In the object query service, the object query goal from the user side is to extract the corresponding object point cloud from a specified video sequence, which is the foundation for applications such as surveillance video analysis, autonomous driving navigation, and robot path planning.

[0088] Therefore, such as Figure 7 The diagram shown illustrates a processing logic of the object query service involved in this application. The object query system based on the 4D semantic map configured in this application mainly includes a 4D semantic map generation module G and a dedicated multimodal large language model for query services. The system operates with two branches: time-independent object query and time-dependent object query. The process is as follows: 7.1) will Frame-based dynamic scene video sequence The data is pre-input into the 4D semantic map generation module G to obtain a time-independent 4D semantic map. 4D Semantic Map Related to Time Series This is for the purpose of user object query; Of course, for the preliminary map preparation work here, time-independent 4D semantic maps 4D Semantic Map Related to Time Series In some cases, the generation and processing can also be triggered in real time when the user initiates a related query service.

[0089] 7.2) When a user inputs a query statement in natural language into the system When initiating a query service, a dedicated multimodal large language model is used. For query statement Analyze and break down into Time-independent query objects and Time-series related query objects ; 7.3) Time-independent query branches use a query-specific contrastive language image pre-trained model. Generate each query object semantic embedding vector Then, a cosine similarity metric is used to measure the time-independent 4D semantic map. Matching is performed to obtain a time-independent object point cloud. ; Here, the processing in 7.3) and the processing in 7.4) can be performed simultaneously or one after the other, without any specific timing restrictions.

[0090] 7.4) Time-series related query branches use a dedicated large language model for the query service. Generate each query object semantic embedding vector Then, the cosine similarity metric is used to compare the data with a pre-generated temporally related 4D semantic map. Matching is performed to obtain a point cloud of time-related objects. This is understandable; it refers to a time-independent point cloud. Point clouds of time-related objects These are the two types of object query results that can be achieved based on the dynamic scene 4D semantic map. At this point, the query results can be fed back to the user to complete a query process.

[0091] The entire process can be represented by the following formula: , , , , , .

[0092] From the specific processing logic of the object query service above, it can be seen that, compared with the object query system based on the 4D semantic map generated by the Gaussian splashing method, the object query system based on this application has the following advantages: (1) Due to the use of a feedforward framework and multi-scene merging training, 4D semantic maps can be generated directly without repeatedly calling the supervised semantic generator to optimize model weights due to scene switching or changes in the video. While having a stronger adaptability to complex and ever-changing dynamic scenes, it greatly saves the time and computational costs of generating 4D semantic maps, laying a good foundation for its large-scale deployment on object query systems. (2) Due to the adoption of semantic alignment and video reconstruction dual-objective optimization, the generated 4D semantic map has higher accuracy, and thanks to this, the point cloud of the queried object also has higher accuracy.

[0093] (3) Due to the use of complementary dual semantic heads for supercomplete training, it can generate two types of 4D semantic maps with higher accuracy: time-independent and time-dependent. Thanks to this, the user query needs can be automatically decomposed into time-independent and time-dependent objects and queried in the corresponding dynamic scene 4D semantic maps with higher accuracy. Experiments on the public datasets HyperNeRF and Neu3D have also confirmed the above-mentioned excellent query effect.

[0094] In conclusion, regarding the above solutions, this application, for the goal of generating 4D semantic maps for dynamic scenes, designs the first feedforward framework for 4D semantic map generation. It achieves joint processing of geometric perception and semantic alignment within a single architecture. This framework comprises two core components: a streaming visual geometric transformer that captures the spatiotemporal geometric features of dynamic scenes and a semantic bridging decoder that maps these features to a language-aligned semantic space. This approach maintains structural integrity while enhancing semantic interpretability. Unlike traditional methods that rely on time-consuming scene-level optimization, this approach effectively supports training by merging multiple dynamic scenes, allowing for direct application during inference. It boasts high computational efficiency and strong generalization ability. This design significantly improves its practicality for large-scale deployment, opens up new avenues for open-vocabulary 4D scene understanding, and can provide excellent data support for scene understanding tasks in applications such as embodied intelligence, metaverse, and digital twins, demonstrating promising application prospects.

[0095] The above is an introduction to the dynamic scene 4D semantic map generation method provided in this application. In order to facilitate better implementation of the dynamic scene 4D semantic map generation method provided in this application, this application also provides a dynamic scene 4D semantic map generation device from the perspective of functional modules.

[0096] See Figure 8 , Figure 8 This is a schematic diagram of a dynamic scene 4D semantic map generation device according to this application. In this application, the dynamic scene 4D semantic map generation device 800 may specifically include the following structure: The acquisition unit 801 is used to acquire the current video sequence for which the corresponding 4D semantic map is to be generated; The generation unit 802 is used to input the current video sequence into the 4D semantic map generation network. The 4D semantic map generation network includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder is used to process the video sequence input to the network into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. Output unit 803 is used to extract and output dynamic scene time-independent 4D semantic map and dynamic scene time-related 4D semantic map.

[0097] In one exemplary embodiment, the streaming visual geometric transformation encoder includes a label-free self-distillation image encoding module E, a merging module, and an alternating attention transformation module D. During operation, the following occurs: Dynamic scene video frames at time t The initial camera token at time t is obtained by processing the image encoding module E without label. and image tokens and; Initial camera token and image tokens By merging modules and Cache of time The token sequences in the buffer are merged to obtain the cache at time t. , among which, initial cache Empty; Cache at time t The token sequence in the code generates camera tokens at time t through the alternating attention transformation module D. and geometric tokens .

[0098] In yet another exemplary embodiment, the streaming visual geometric transformation decoder includes a camera head. Depth head And the inverse projection module P, during operation, has: Camera token at time t Through the camera lens Processing to obtain predicted camera parameters ; Geometric token at time t via depth head Processing yields a reconstructed depth map ; Predict camera parameters and reconstructed depth map The reconstructed point cloud at time t is obtained by processing the data using the inverse projection module P. .

[0099] In yet another exemplary embodiment, the semantic bridging decoder includes a dense prediction transform. Time-independent semantic header Semantic Header Related to Time During the work process, there are: Geometric token at time t Through dense prediction transformer Process to obtain context token ; context token Through time-independent semantic headers Semantic Header Related to Time The temporally independent prediction semantics at time t are obtained by processing them separately. Semantic prediction related to time series .

[0100] In yet another exemplary embodiment, the semantic bridging decoder also includes a video frame reconstruction head. Used to transfer context tokens Reconstructing the head from video frames Processing to obtain the reconstructed video frame at time t ; The 4D semantic map generation network also includes a supervised semantic generator; The 4D semantic map generation network is pre-trained, including: T-frame dynamic scene training video sequence The geometric token sequence is obtained by processing the weight-frozen streaming visual geometric transformation encoder. ; geometric token sequence The temporally independent predicted semantic sequence is obtained through semantic bridging decoder processing. Temporal correlation prediction semantic sequence and reconstructing video sequences ; T-frame dynamic scene training video sequence A time-independent supervised semantic sequence is obtained through processing by a supervised semantic generator. Temporally related supervised semantic sequences ; Training video sequences using T-frame dynamic scenes Reconstructing video sequences Time-independent prediction semantic sequences Temporal correlation prediction semantic sequence Temporally independent semantic supervision sequences Semantic supervision sequences related to time series The total loss function is calculated to perform multi-objective joint optimization of the weights of the semantic bridging decoder.

[0101] In yet another exemplary embodiment, the supervised semantic generator includes a video segmentation module S, a temporally independent supervised semantic generation branch, and a temporally related supervised semantic generation branch. The temporally independent supervised semantic generation branch includes a contrastive language image pre-trained model. Along with the first alignment module, the temporally related supervised semantic generation branch includes a multimodal large language model. Large language model And the second alignment module, during the operation, has: T-frame dynamic scene training video sequence The video segmentation module S uses a decoupled video segmentation method to segment the video, resulting in a mask sequence of N video objects. ; T-frame dynamic scene training video sequence By comparing language image pre-trained models Obtain the mask sequence of N video objects Corresponding image object semantic embedding vector ; The time-independent semantics are embedded into the vector through the first alignment module. With the mask sequence of N video objects Align all pixels within the region to obtain a time-independent supervised semantic sequence. ; T-frame dynamic scene training video sequence Through multimodal large language model Obtain the mask sequence of N video objects The corresponding dynamic description of the video object; Dynamically describe video objects using a large language model. Obtain the sequence of time-related semantic embedding vectors ; The second alignment module embeds temporally relevant semantics into the vector sequence. With the mask sequence of N video objects Align all pixels within the region to obtain a temporally correlated supervised semantic sequence. .

[0102] In yet another exemplary embodiment, the total loss function is expressed as follows: , , , , , in, Let α be the total loss function, and α be the total loss function for video reconstruction. The corresponding weights, where β is the total semantic alignment loss function. The corresponding weights are α>0 and β>0. To characterize the structural loss of video frames at each time point during reconstruction Norm loss function and characterization of pixel-level smoothness loss in video frames The weighted balancing factor of the norm loss function. , Let be the total loss function for the first semantic alignment. The total loss function for the second semantic alignment. To characterize the semantic alignment magnitude loss at each time step The weights corresponding to the norm loss function The weights corresponding to the cosine similarity loss function, which characterizes the semantic alignment direction loss at each time step, are used to define the weights. , .

[0103] This application also provides a processing device from a hardware architecture perspective, see [link / reference]. Figure 9 , Figure 9 This diagram illustrates a structural schematic of the processing device of this application. Specifically, the processing device may include a processor 901, a memory 902, and an input / output device 903. The processor 901 executes the computer program stored in the memory 902 to implement, for example... Figure 1 The corresponding steps of the dynamic scene 4D semantic map generation method in the embodiment; or, when the processor 901 executes the computer program stored in the memory 902, it implements as follows: Figure 8 Corresponding to the functions of each unit in the embodiment, the memory 902 is used to store the functions executed by the processor 901 as described above. Figure 1 The computer program required for the dynamic scene 4D semantic map generation method in the corresponding embodiment.

[0104] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 902 and executed by processor 901 to complete this application. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a computer device.

[0105] The processing device may include, but is not limited to, processor 901, memory 902, and input / output device 903. Those skilled in the art will understand that the illustrations are merely examples of the processing device and do not constitute a limitation on the processing device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the processing device may also include network access devices, buses, etc., and processor 901, memory 902, input / output device 903, etc., are connected via a bus.

[0106] The processor 901 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting various parts of the device through various interfaces and lines.

[0107] The memory 902 can be used to store computer programs and / or modules. The processor 901 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 902 and by calling data stored in the memory 902. The memory 902 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the processing device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0108] When processor 901 executes a computer program stored in memory 902, it can specifically perform the following functions: Obtain the current video sequence from which the corresponding 4D semantic map is to be generated; The current video sequence is input into the 4D semantic map generation network, which includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder processes the input video sequence into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. Extract and output a time-independent 4D semantic map and a time-dependent 4D semantic map of a dynamic scene.

[0109] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the dynamic scene 4D semantic map generation device, processing equipment, and its corresponding units described above can be found in the following reference: Figure 1 The description of the dynamic scene 4D semantic map generation method in the corresponding embodiment will not be repeated here.

[0110] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0111] Therefore, this application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the present application. Figure 1 The steps of the dynamic scene 4D semantic map generation method in the corresponding embodiment can be found in the following example. Figure 1 The description of the dynamic scene 4D semantic map generation method in the corresponding embodiment will not be repeated here.

[0112] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0113] Because of the instructions stored in the computer-readable storage medium, the present application can be executed as described above. Figure 1 The steps of the dynamic scene 4D semantic map generation method in the corresponding embodiment can therefore achieve the results of this application. Figure 1The beneficial effects that the dynamic scene 4D semantic map generation method can achieve in the corresponding embodiment are detailed in the preceding description and will not be repeated here.

[0114] The above provides a detailed description of the dynamic scene 4D semantic map generation method, apparatus, processing device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating a dynamic scene 4D semantic map, characterized in that, The method includes: Obtain the current video sequence from which the corresponding 4D semantic map is to be generated; The current video sequence is input into a 4D semantic map generation network, which includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder processes the input video sequence into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. Extract and output the time-independent 4D semantic map and the time-dependent 4D semantic map of the dynamic scene.

2. The method according to claim 1, characterized in that, The streaming visual geometric transformation encoder includes a label-free self-distillation image encoding module E, a merging module, and an alternating attention transformation module D. During operation, the following occurs: Dynamic scene video frames at time t The image token at time t is obtained through the labelless self-distillation image encoding module E. and initial camera token ; The image token and the initial camera token Through the merging module and Cache of time The token sequences in the buffer are merged to obtain the cache at time t. , among which, initial cache Empty; The cache at time t The token sequence in the sequence generates geometric tokens at time t through the alternating attention transformation module D. and camera token .

3. The method according to claim 2, characterized in that, The streaming visual geometry transform decoder includes a depth head. Camera head And the inverse projection module P, during operation, has: The geometric token at time t Through the depth head Processing yields a reconstructed depth map ; The camera token at time t Through the camera head Processing to obtain predicted camera parameters ; The reconstructed depth map and the predicted camera parameters The reconstructed point cloud at time t is obtained through the inverse projection module P. .

4. The method according to claim 1, characterized in that, The semantic bridging decoder includes a dense prediction transform. Time-independent semantic header Semantic Header Related to Time During the work process, there are: Geometric token at time t Through the dense prediction transformer Process to obtain context token ; The context token Through time-independent semantic headers Semantic Header Related to Time The temporally independent prediction semantics at time t are obtained by processing them separately. Semantic prediction related to time series .

5. The method according to claim 1, characterized in that, The semantic bridging decoder also includes a video frame reconstruction head. Used to transfer the context token Reconstructing the head through the video frames Processing to obtain the reconstructed video frame at time t ; The 4D semantic map generation network also includes a supervised semantic generator; The 4D semantic map generation network was pre-trained with the following features: T-frame dynamic scene training video sequence The geometric token sequence is obtained by processing the streaming visual geometric transformation encoder with weight freezing. ; The geometric token sequence The reconstructed video sequence is obtained through processing by the semantic bridging decoder. Time-independent prediction of semantic sequences Predicting semantic sequences related to time series ; The T-frame dynamic scene training video sequence The time-independent supervised semantic sequence is obtained through the supervised semantic generator. Temporally related supervised semantic sequences ; The video sequence is trained using the T-frame dynamic scene. The reconstructed video sequence The time-independent predictive semantic sequence The temporal correlation prediction semantic sequence The time-independent semantic supervision sequence and the time-related semantic supervision sequence The total loss function is calculated to perform multi-objective joint optimization of the weights of the semantic bridging decoder.

6. The method according to claim 5, characterized in that, The supervised semantic generator includes a video segmentation module S, a time-independent supervised semantic generation branch, and a time-related supervised semantic generation branch. The time-independent supervised semantic generation branch includes a contrastive language image pre-trained model. Along with the first alignment module, the temporally related supervised semantic generation branch includes a multimodal large language model. Large language model And the second alignment module, during operation, has the following: The T-frame dynamic scene training video sequence The video segmentation module S uses a decoupled video segmentation method to segment the video, resulting in a mask sequence of N video objects. ; The T-frame dynamic scene training video sequence The contrastive language image pre-trained model Obtain the mask sequence of the N video objects. The corresponding time-independent semantic embedding vector ; The time-independent semantic embedding vector is obtained through the first alignment module. With the mask sequence of the N video objects Aligning all pixels within the region yields the time-independent supervised semantic sequence. ; The T-frame dynamic scene training video sequence Through the aforementioned multimodal large language model Obtain the mask sequence of the N video objects. The corresponding dynamic description of the video object; The video object is dynamically described using the large language model. Obtain the sequence of time-related semantic embedding vectors ; The time-related semantics are embedded into the vector sequence via the second alignment module. With the mask sequence of the N video objects Aligning all pixels within the region yields the temporally correlated supervised semantic sequence. .

7. The method according to claim 5, characterized in that, The total loss function is expressed as follows: , , , , , in, Let α be the total loss function, and α be the total loss function for video reconstruction. The corresponding weights, where β is the total semantic alignment loss function. The corresponding weights are α>0 and β>0. To characterize the structural loss of video frames at each time point during reconstruction Norm loss function and characterization of pixel-level smoothness loss in video frames The weighted balancing factor of the norm loss function , Let be the total loss function for the first semantic alignment. Let the total loss function be the second semantic alignment function. To characterize the semantic alignment magnitude loss at each time step The weights corresponding to the norm loss function The weights corresponding to the cosine similarity loss function, which characterizes the semantic alignment direction loss at each time step, are used to define the weights. , .

8. A dynamic scene 4D semantic map generation device, characterized in that, The device includes: The acquisition unit is used to acquire the current video sequence from which the corresponding 4D semantic map is to be generated; A generation unit is used to input the current video sequence into a 4D semantic map generation network. The 4D semantic map generation network includes a streaming visual geometry transformer, a semantic bridging decoder, and a point cloud coloring module. The streaming visual geometry transformer includes a streaming visual geometry transformer encoder and a streaming visual geometry transformer decoder. The streaming visual geometry transformer encoder processes the input video sequence into a camera token sequence and a geometry token sequence. The streaming visual geometry transformer decoder processes the camera token sequence and the geometry token sequence into a reconstructed point cloud sequence. The semantic bridging decoder processes the geometry token sequence into a time-independent predictive semantic sequence and a time-dependent predictive semantic sequence. The point cloud coloring module uses the time-independent predictive semantic sequence and the time-dependent predictive semantic sequence to color the reconstructed point cloud sequence, thereby obtaining a dynamic scene time-independent 4D semantic map and a dynamic scene time-dependent 4D semantic map. The output unit is used to extract and output the dynamic scene time-independent 4D semantic map and the dynamic scene time-related 4D semantic map.

9. A processing device, characterized in that, The method includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the method as described in any one of claims 1 to 7 when it invokes the computer program in the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 7.

Citation Information

Cited By

  • Test method and device for vehicle automatic driving planning algorithm, medium and product

    CN121807726A

  • Remote sensing rotating target self-attention mechanism construction method based on geometric compatible perception

    CN122176560A