Text semantic driven traffic scene point cloud generation method and system

By constructing point cloud and text alignment samples and introducing temporal condition injection and frequency modulation, the problems of insufficient semantic matching and geometric structure distortion in traffic scene point cloud generation are solved, realizing point cloud generation that conforms to traffic patterns and improving generation efficiency and effect.

CN121788705APending Publication Date: 2026-04-03CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies for generating point clouds of traffic scenes suffer from problems such as insufficient semantic matching, geometric distortion, and inconsistent temporal sequences. In particular, it is difficult to generate point clouds that conform to the physical laws of traffic in dynamic traffic scenes.

Method used

By constructing point cloud and text aligned samples and fine-tuning the model, temporal condition injection and temporal frequency modulation are introduced, a scoring function is designed to select reasonable frames, and temporal depth gradient correction and temporal coordinate fusion are adopted to generate point clouds that conform to traffic rules.

Benefits of technology

It achieves semantically controllable, geometrically faithful, and temporally coherent point cloud generation of traffic scenes, supports edge scene generation and simulation of intelligent vehicles, and reduces R&D costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788705A_ABST
    Figure CN121788705A_ABST
Patent Text Reader

Abstract

The invention relates to a traffic scene point cloud generation method driven by text semantics. The traffic scene point cloud generation method comprises a fine adjustment stage and a generation stage. In the fine tuning stage, traffic scene laser radar point cloud samples are collected, point cloud and text alignment samples are constructed, and a pre-training text coding model is finely tuned to obtain a time sequence pre-training model; in the generation stage, firstly, time sequence pre-training language coding is performed on traffic scene text description, time sequence point cloud semantic features and time sequence grammar features are extracted, and then a structured time sequence semantic vector sequence is generated through time sequence condition injection and time sequence frequency modulation; secondly, generating a plurality of candidate range graphs through dynamic resolution adaptation and inter-frame correlation correction, screening an optimal frame, and iteratively generating a continuous range graph sequence; and finally, sequentially executing time sequence depth gradient correction and time sequence coordinate fusion on each pixel of each frame in the range map sequence to generate a single-frame point cloud so as to obtain a continuous point cloud sequence. The intelligent vehicle edge scene generation and simulation can be effectively supported, and the research and development efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and artificial intelligence technology, and relates to a text semantic-driven method and system for generating point clouds of traffic scenes. Background Technology

[0002] Point clouds, as a core data form for representing three-dimensional spatial structures, have become a key technological support in the field of traffic scenario simulation and testing due to their high-precision and high-density geometric feature description capabilities. In application scenarios such as autonomous driving system development, intelligent transportation planning, and virtual scenario verification, point cloud data provides a high-fidelity real-world mapping for algorithm training and system testing by accurately reconstructing road topology, dynamic trajectories of traffic participants, and the distribution of environmental obstacles. However, the current methods for acquiring point cloud data still have significant limitations.

[0003] Current mainstream point cloud acquisition solutions heavily rely on active sensors such as LiDAR and depth cameras for on-site measurements. While these hardware devices offer millimeter-level measurement accuracy, their cost per unit is generally in the tens to hundreds of thousands of yuan range. Furthermore, they require high-precision inertial navigation systems (IMUs) and global positioning systems (GNSS) for spatial positioning, further raising the hardware barrier to data acquisition. Even more challenging are the dynamic interference factors in complex traffic scenarios (such as rain, snow, strong light reflection, and multipath effects), which significantly reduce sensor detection reliability, leading to problems such as noise pollution, uneven density, and even localized data loss in point cloud data.

[0004] For the coverage needs of long-tail scenarios (such as extreme weather, emergencies, and irregular obstacles), existing data collection methods reveal a dual dilemma of efficiency and cost. Taking autonomous driving testing as an example, to verify the system's response capability in rare scenarios, a test database containing tens of thousands of edge cases needs to be built. If relying entirely on on-site data collection, not only is it necessary to deploy multiple sensor arrays for long-term monitoring, but also to manually annotate key events in massive amounts of data, resulting in an exponential increase in time and economic investment. In addition, real-time data collection in some high-risk scenarios (such as traffic accident scenes) also poses safety risks, further limiting the feasibility of data acquisition. These bottlenecks severely restrict the large-scale application and iteration speed of 3D scene simulation technology, urgently requiring innovative solutions that break through the traditional data collection paradigm.

[0005] Current technologies have addressed certain technical problems to some extent, but they also have their own limitations. For example, the paper WU Y, ZHANG K, QIAN J, XIE J, YANG J. Text2LiDAR: Text-Guided LiDAR Point Cloud Generation via Equirectangular Transformer. In: LEONARDIS A, RICCI E, ROTH S, et al. (eds.), Computer Vision – ECCV 2024, Lecture Notes in Computer Science, vol. 15114, Springer, Cham, 2025, pp. 279–295. This paper uses a horizon-based latitude-longitude Transformer architecture combined with textual semantic control to generate LiDAR point clouds and supports continuous frame generation. However, this method lacks a mechanism to filter the validity of generated frames, which may lead to discontinuous jumps in the positions of dynamic targets in adjacent frames. The paper RAN H, GUIZILINI V, WANG Y. Towards Realistic Scene Generation with LiDAR Diffusion Models. In: 2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2024: 14738–14748. uses a range map as an intermediate representation and employs curve compression and point-by-point coordinate supervision to preserve the point cloud geometry. However, this method does not explicitly model the motion state of dynamic targets, potentially losing detailed features of moving objects during generation.

[0006] Chinese patent application "A 3D Point Cloud Generation Method Based on a Diffusion Model" proposes a U-Net-type diffusion model that combines a variational autoencoder and a multi-layer shared MLP structure to generate 3D point clouds of single objects. However, this method is only applicable to isolated object scenes and cannot generate realistic traffic environments containing multiple types of dynamic targets and complex geometric structures. Chinese patent application "A Method and Apparatus for Generating a 3D Point Cloud Model" proposes a method based on a neural radiation field model, trained using photometric consistency and depth smoothing loss, to generate a depth map which is then projected into a point cloud. However, this method relies on the input image during generation, lacks textual semantic driving capabilities, and cannot generate corresponding scene point clouds based on natural language descriptions. Summary of the Invention

[0007] In view of this, the purpose of this invention is to provide a text semantic-driven method and system for generating traffic scene point clouds. This method can generate traffic scene lidar point clouds that conform to traffic physics laws and geometric details based on user-input text descriptions of traffic scenes. Compared with existing technologies, this invention can effectively solve key problems such as insufficient semantic matching, geometric structure distortion, and temporal discontinuity in dynamic traffic scenes.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A text semantic-driven method for generating traffic scene point clouds is disclosed. This method generates traffic scene LiDAR point clouds that conform to traffic physics laws and geometric details based on user-input text descriptions of traffic scenes. The method includes a fine-tuning stage and a generation stage. The fine-tuning stage includes point cloud data processing and model fine-tuning. The generation stage includes traffic scene text semantic generation, traffic scene continuous range map generation, and traffic scene continuous point cloud generation.

[0009] The fine-tuning stage includes: point cloud data processing and model fine-tuning. LiDAR point cloud samples are collected for traffic scenes, and the shape, size and traffic dynamic attributes of traffic participants are labeled. Point cloud and text alignment samples are constructed based on the labeling results. The pre-trained text encoding model is fine-tuned using the alignment samples to obtain a time-series pre-trained model.

[0010] In the generation stage, the traffic scene text semantic generation includes: based on a temporal pre-trained model, temporal encoding is performed on the input traffic scene text description, and temporal point cloud semantic features and temporal grammatical features are extracted respectively; a temporal cross-attention layer that integrates point cloud semantic features and grammatical features is constructed through a temporal conditional injection mechanism, and the attention weight matrix of the previous frame is introduced to model inter-frame dynamic dependencies; then, through temporal frequency modulation, the semantic features are decomposed into low-frequency components corresponding to traffic stability structures and high-frequency components corresponding to dynamic behaviors, and the frequency domain gain is dynamically adjusted according to the scene complexity; finally, a structured temporal semantic vector sequence is generated to guide the generation of the range map. .

[0011] In the generation phase, the continuous range map generation of the traffic scene includes: dynamic resolution adaptation based on a structured temporal semantic vector sequence; extracting dynamic target motion attribute sub-vectors and mapping them to semantic guidance vectors, and performing inter-frame correlation correction on the channel weights of the current frame in conjunction with the channel weights of the previous frame; generating multiple candidate range maps that conform to traffic rules based on the corrected channel weights and the current frame resolution; constructing a scoring function containing three dimensions—target attributes, scene rules, and inter-frame consistency—to evaluate the rationality of the candidate range maps and select the optimal frame as the range map of the current frame; iteratively executing the above process to generate range maps for each frame in sequence, resulting in a continuous range map sequence. .

[0012] Furthermore, in the generation stage, the continuous point cloud generation of the traffic scene includes: based on the continuous range map sequence and LiDAR sensor parameters, performing temporal depth gradient correction and temporal coordinate fusion sequentially on each pixel in the single-frame range map to obtain the three-dimensional spatial coordinates of that pixel; collecting the three-dimensional coordinates of all pixels in a single frame to generate a single-frame LiDAR point cloud; performing the above process on each frame in the continuous range map sequence to finally obtain a continuous LiDAR point cloud sequence. .

[0013] Furthermore, the traffic scene text semantic generation step in the generation stage specifically includes: 1) Temporal pre-trained language encoding: Based on the temporal pre-trained model, the dynamic attributes of traffic scenes in the text are encoded to obtain temporal point cloud semantic features. Encoding the logical relationships of traffic scenes in the text yields temporal grammar features. ; in, For the first Frame-time point cloud semantic feature vectors The first The embedding vectors of the traffic participants in the frame, including the category embedding vector, lane affiliation embedding vector, unit motion direction vector, and speed level embedding vector; For the first Frame temporal syntax feature vectors The first Spatial relationship vectors and interaction relationship vectors among traffic participants in a frame; 2) Timing condition injection: using subset slice and subset slices in Construct a cross-attention layer of point cloud semantic and syntactic features adapted to traffic scenarios; then combine it with the attention weight matrix of the previous frame. Obtain the current frame weight matrix. ;

[0014] in, ; ; This is the normalization function; For feature dimensions; This is the time-series weight decay coefficient; For the first Semantic feature vectors of frame-time point clouds; For the first Frame temporal syntax feature vector; based on Fusion and The initial temporal semantic vector is obtained. ;

[0015] in, The semantic vector fusion coefficient of the previous frame; For the first Initial temporal semantic vector of a frame; 3) Timing-frequency modulation: (Referring to section 2.2) Perform discrete wavelet transform to decompose the traffic stability features into time-series low-frequency components. Corresponding high-frequency components of traffic dynamic characteristics in time series ;based on and Determine the complexity level of traffic scenarios ,right Apply and Matched dynamic gain Features are reconstructed through inverse discrete wavelet transform and fused with traffic feature constraints from the first two frames to output the first... Frame structured temporal semantic vector ;

[0016] in, This is an inverse discrete wavelet transform used to reconstruct high and low frequency features; For time series smoothing coefficients; The first Frame-structured temporal semantic vectors; 4) Generation of continuous semantic vectors: Repeat steps 1)-3) to generate the next continuous semantic vector in sequence. to Frame-based structured temporal semantic vectors are used to obtain a continuous sequence of structured temporal semantic vectors. .

[0017] Furthermore, the specific steps for generating the continuous range map of the traffic scene in the generation stage include: 1) Dynamic resolution adaptation: based on the first Frame structured temporal semantic vector Extracted target motion attribute sub-vectors ;Target motion attribute subvector The components are normalized and weighted to obtain a comprehensive exercise activity index. ;

[0018] in, for Norm; For component weights; for The function will Constraint to ; Combination and Establish linear mapping rules to generate the initial resolution. ;

[0019] in, These are the height and width of the base resolution, respectively; This is the complexity gain coefficient. ; The basic time-series weight decay coefficient; This is the complexity impact coefficient; Introducing the final resolution of the previous frame , generate the first Final frame resolution ;

[0020] in, This is the transition smoothing coefficient. When =1, =0; This is the floor function; 2) Inter-frame correlation correction: based on the first Final frame resolution and Category embedding vectors of traffic participants Constructing horizontal time series features Vertical time series characteristics ;Transform the dynamic target motion attribute subvector Mapped to semantic guidance vectors through fully connected layers Introducing a speed level adjustment factor and the channel weights of the previous frame For the current frame channel weights Perform smoothing correction;

[0021]

[0022] in, This is the normalization function; This is a vector concatenation function that combines horizontal time-series features. Vertical time series characteristics By concatenating along the feature dimensions, a joint feature vector is obtained; This is the speed influence coefficient; for Norm; 3) Multiple candidate frame generation and filtering: based on the range map of the previous frame. Current frame resolution and channel weights Generate baseline candidate frames Generate by adding minute noises that conform to traffic patterns. A total of diverse candidate frames were obtained. Candidate frames { Then, a scoring function is constructed that includes three dimensions: target attributes, scene rules, and inter-frame consistency. The candidate frame with the highest score is selected as the current frame range map. If the highest score does not meet the score requirement, then increase the score. Regenerate candidate frames until the score meets the requirements;

[0023] in, These are the weighting coefficients; For target attribute matching degree; For scene rule matching degree; Inter-frame consistency matching degree; Number of candidate range maps; For the first Index of candidate range maps; 4) Generation of continuous range plots: Repeat steps 3.1-3.3 to generate the next continuous range plot in sequence. to Frame range map, and then a continuous range map sequence is obtained. .

[0024] Furthermore, the continuous point cloud generation step for the traffic scene in the generation stage includes: 1) Temporal depth gradient correction: The first... Frame range map medium pixel Normalization yields And combined with the maximum ranging distance of the lidar sensor , obtained the Frame range map pixels original depth ;

[0025] Based on the horizontal angular resolution of LiDAR With vertical angular resolution The pixel is obtained using the center difference method. depth change rate ;

[0026] in, This represents the gradient of the original depth in the horizontal direction. The gradient of the original depth in the direction; , Representing pixels The horizontal column coordinates and vertical row coordinates in the range chart; Then introduce spatial correction coefficient , obtain pixels correction depth ;

[0027] in, , It is a constant; 2) Temporal coordinate fusion: Based on the horizontal field of view of LiDAR Vertical field of view With current range resolution , obtain pixels horizontal angle with vertical angle ;

[0028]

[0029] in, ; ; Combined with the correction depth Get pixels Three-dimensional spatial coordinates;

[0030] in, , , The order is number 1 Frame pixels Three-dimensional spatial coordinates; 3) Single-frame point cloud generation: For the first frame... Steps 1) and 2) are performed sequentially on all pixels in the frame range image to aggregate the three-dimensional spatial coordinates of each pixel, thus obtaining the first... Frame LiDAR point cloud ; 4) Continuous point cloud generation: For continuous range map sequences Steps 4.1 to 4.3 are executed frame by frame to obtain a continuous sequence of LiDAR point clouds of the traffic scene. .

[0031] The present invention also provides a text semantic-driven point cloud generation system for traffic scenes.

[0032] The beneficial effects of this invention are as follows: This invention improves semantic understanding and overcomes the ambiguity problem of general models by constructing point cloud and text-aligned samples and fine-tuning the model. It introduces temporal condition injection and temporal frequency modulation to separate static structure and dynamic behavior, ensuring the continuity of motion logic between frames. A scoring function with three dimensions—target attributes, scene rules, and inter-frame consistency—is designed to achieve reasonable selection and retrying of multiple candidate frames. In the point cloud reconstruction stage, temporal depth gradient correction and temporal coordinate fusion are employed to ensure rich geometric details and compliance with traffic physics rules. This invention achieves semantically controllable, geometrically faithful, temporally coherent, and rule-compliant traffic scene point cloud generation, improving the effectiveness of traffic scene point cloud generation. It can support the generation and simulation of edge scenes for intelligent vehicles, improving the efficiency of intelligent vehicle R&D and reducing R&D costs.

[0033] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 A flowchart of a text semantic-driven point cloud generation method for traffic scenes is provided as a preferred embodiment of the present invention; Figure 2 This is a flowchart of the traffic scene multi-candidate frame range map generation and filtering process of the present invention; Figure 3 This is a diagram of the architecture for generating a potential spatial diffusion reference frame based on temporal semantic guidance according to the present invention. Figure 4 This is a schematic diagram of the generation results of two consecutive frames of point cloud sequence in an embodiment of the present invention. Detailed Implementation

[0035] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0036] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0037] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0038] like Figure 1 As shown, the present invention provides a text semantic-driven traffic scene point cloud generation method, which includes a fine-tuning stage and a generation stage. The fine-tuning stage performs point cloud data processing and model fine-tuning, while the generation stage performs traffic scene text semantic generation, traffic scene continuous range map generation, and traffic scene continuous point cloud generation.

[0039] In the fine-tuning stage, the point cloud data processing and model fine-tuning steps are used to collect LiDAR point cloud samples of traffic scenes, and label the shape, size and traffic dynamic attributes of traffic participants; construct point cloud and text alignment samples based on the labeling results; and fine-tune the pre-trained text encoding model using the alignment samples to obtain the time-series pre-trained model. In the generation stage, the traffic scene text semantic generation step utilizes a temporally pre-trained model to temporally encode the input traffic scene text description, extracting temporal point cloud semantic features and temporal grammatical features respectively. A temporal cross-attention layer fusing point cloud semantic and grammatical features is constructed through a temporal conditional injection mechanism, and the attention weight matrix of the previous frame is introduced to model inter-frame dynamic dependencies. Then, through temporal frequency modulation, the semantic features are decomposed into low-frequency components corresponding to stable traffic structures and high-frequency components corresponding to dynamic behaviors, with the frequency domain gain dynamically adjusted according to scene complexity. Finally, a structured temporal semantic vector sequence is generated to guide the generation of the range map. ; In the generation phase, the continuous range map generation step utilizes a structured temporal semantic vector sequence for dynamic resolution adaptation; extracts dynamic target motion attribute sub-vectors and maps them to semantic guidance vectors; combines the channel weights of the previous frame to perform inter-frame correlation correction on the channel weights of the current frame; based on the corrected channel weights and the current frame resolution, generates multiple candidate range maps that conform to traffic rules; constructs a scoring function containing three dimensions—target attributes, scene rules, and inter-frame consistency—to evaluate the rationality of the candidate range maps and select the optimal frame as the current frame range map; iteratively executes the above process to generate range maps for each frame sequentially, resulting in a continuous range map sequence. ; In the generation stage, the continuous point cloud generation step for the traffic scene utilizes a continuous range map sequence and LiDAR sensor parameters to sequentially perform temporal depth gradient correction and temporal coordinate fusion on each pixel in a single frame range map, obtaining the three-dimensional spatial coordinates of that pixel in the world coordinate system; the three-dimensional coordinates of all pixels within a single frame are then combined to generate a single frame LiDAR point cloud; the above process is iteratively executed on each frame in the continuous range map sequence, ultimately yielding a continuous LiDAR point cloud sequence. ; Figure 2 This is a flowchart of the traffic scene multi-candidate frame range map generation and filtering process of the present invention. The process includes the following steps: 1. Generate a reference frame range map 1.1 When >1 hour Figure 3 This is a diagram of the potential spatial diffusion reference frame generation architecture based on temporal semantic guidance of the present invention.

[0040] First, view the range map of the previous frame. From pixel space via encoder Mapping to the latent space yields the initial latent representation. The initial potential representation It is fed into a Markov diffusion process; at each time step This process represents the current potential. Add a noise term that follows a normal distribution. This leads to the next step of obtaining the potential representation. ;go through After the first diffusion step, the latent representation is finally obtained and is completely covered by noise. ;

[0041] Among them, when hour, For the final noisy latent representation, This is the preset total number of diffusion steps; For the first The noisy latent representation after step diffusion; For the first The retention factor of the step; For the first Noise added step, dimension and latent representation Consistent; At the same time, the feature fusion module will use the current frame resolution and channel weight , fusion generation condition features ;

[0042] in, This represents a fully connected layer projection operator that projects discrete resolution parameters. Mapped to continuous feature vectors; This represents a layer normalization operator that standardizes the projected resolution features. This represents the element-wise multiplication operator, used to fuse normalized resolution features with channel weights corrected for inter-frame correlation. Then conditional features Injection denoising Guiding the network from Inverse denoising to generate intermediate latent representation The corrector is based on Temporal and conditional characteristics ,right Make corrections and enhancements to output an optimized latent representation. ; final, via decoder Inverse mapping back to pixel space to generate baseline candidate frames ; 1.2 When =1 For the first frame of the sequence, since there is no range map of the preceding frame. For reference, this invention ensures that the generation of the initial latent representation is entirely driven by textual semantics through an initialization strategy that includes clustering pre-computation, similarity calculation, and weighted fusion, and includes the following steps: 1.2.1 Scene Clustering Pre-computation Based on the text semantic labels corresponding to point cloud samples, the following approach is adopted: Clustering algorithms divide all samples into Traffic-related scenarios; for each scenario, all point cloud samples are processed one by one through an encoder. Mapping to the latent space yields the set of latent representations for this class of samples.

[0043] Then calculate the mean vector of this set, as the first... Latent representation of class scenarios ;

[0044] in, For the first Number of samples in the same scenario ; For the first The first in the class scenario The range map corresponding to each point cloud sample; For encoder functions, and generation phase At that time, the encoder that maps the range map to the latent representation is the same network; The mean coefficient; 1.2.2 Scene Similarity Calculation right For traffic-related scenarios, semantic features are extracted to construct scene semantic templates. Cosine similarity calculation is used. semantic templates for each type of scene Similarity score ;

[0045] in, The cosine similarity function is used. for Norm; And thus obtain and Class scene template corresponding A set of similarity scores ; 1.2.3 Scene-weighted fusion Based on similarity score ,right The latent representations corresponding to each traffic scenario are weighted and fused to generate... Initial latent representation when =1 ;

[0046] in, It is an exponential function; Finally As the initial latent representation of the diffusion process, subsequent execution and The process is completely identical when >1; 2 generation Candidate frame range map: generated by adding small noise that conforms to traffic patterns. A diverse range of candidate frames;

[0047]

[0048] in, The noise matrix has dimensions of ... Consistent, each element in the matrix independently follows a normal distribution. Adjusted by the dynamic target motion attributes; Noise intensity coefficient; Finally, a total of Candidate frames { }; 3D scoring mechanism: Construct a scoring function that includes three dimensions: target attributes, scene rules, and inter-frame consistency; Target attribute matching degree :

[0049] in, for The number of medium-speed transportation targets; for The Middle The attribute vector of each target; for The Middle The expected attribute vector of each target; Scene rule matching degree :

[0050] in, for The number of areas where traffic rules are violated; for The total number of regions that need to be checked by the rules; Inter-frame consistency matching :

[0051] in, for and Area of ​​the overlapping region; for and The total area of ​​the merged region; Then construct a three-dimensional scoring function. :

[0052] in, These are the weighting coefficients; 4. Obtain the highest-scoring range map: Select the candidate frame with the highest-scoring range map as the current frame range map. ;

[0053]

[0054] in, () is from Among the scores of candidate frames, the index of the candidate frame with the highest score is selected; The index of the candidate frame with the highest score; 5. If the highest score does not meet the score requirement, then the number of candidate range maps will be reduced. Increase to Regenerate the candidate range map until the highest score of the candidate frame in the generated range map meets the score requirement.

[0055] in, Number of candidate range maps; This represents the increment in the number of candidate range maps; 6. If the highest score meets the score requirement, output the candidate frames of the range map that match the highest score; Furthermore, in the step of generating the reference frame range map, the training of the entire generative model is divided into two main stages: first, an autoencoder is trained for data compression and reconstruction, and then, based on this, an end-to-end training is performed to... A diffusion-generating network with [the core of the network].

[0056] (1) Autoencoder training: The goal of this stage is to train the encoder. With decoder The self-encoder, composed of these components, enables accurate mapping from a high-dimensional range map to a low-dimensional latent space, while ensuring geometric fidelity in the inverse mapping from the latent representation back to the pixel space. a) Construction of training samples The point cloud sequence from the LiDAR point cloud constructed during the fine-tuning stage and the text alignment sample is converted into the corresponding range map sequence. ; Single frame range map As network input and using itself as the reconstruction target, self-supervised sample pairs are constructed for autoencoder training; b) Loss function design: use Loss function calculation and reconstruction range map Compared with the reference range diagram The pixel-level depth error is defined as:

[0057] in, Dimensions for the reference range diagram; For the reference range of pixels The depth value, To reconstruct the range map pixels The depth value; For reconstruction losses; The set of all pixel coordinates in the baseline range map; c) Training strategy design First, initialize the encoder. and decoder The network parameters are then determined, and the constructed self-supervised sample pairs are analyzed. Input network. Through forward propagation, the encoder... Input range map Compression into low-dimensional latent representation Then the decoder Reconstruct the range map from the latent representation. ; Calculate the reconstruction results Compared with the input benchmark Reconstruction losses between The encoder is jointly optimized using the backpropagation algorithm. and decoder The parameters are set. The training process uses mini-batch stochastic gradient descent, iteratively updating until the loss function converges; after this training phase is completed, the encoder... and decoder The parameters are fixed; (2) train: The goal at this stage is to train noise reduction. The modifier enables it to generate high-quality candidate range maps by utilizing the previous frame's historical range map, noise latent representation, and semantic conditions.

[0058] a) Construction of training samples Reuse the baseline range map sequence used for autoencoder training Based on the encoder trained in the previous step The true range map of the current frame is processed sequentially. Previous frame reference range map Encode to obtain the reference latent representation Initial latent representation ;

[0059]

[0060] Then Perform according to the Markov forward diffusion process of this invention. Step-by-step noise generation ; Structured temporal semantic vectors output by the temporal pre-trained model during the fine-tuning stage Perform dynamic resolution adaptation and inter-frame correlation correction to obtain the target resolution of the current frame. With the previous frame channel weights Then, conditional feature fusion is performed to obtain conditional features. ; Finally, for the first... Frame, constructing a baseline range map of the previous frame. Noisy latent representation Conditional characteristics The input signal, and the reference latent representation The sample set of supervisory signals; b) Loss Function Design To achieve synergistic optimization of denoising accuracy, temporal consistency, conditional matching degree, and reconstruction fidelity, a multi-task weighted total loss function is adopted, which includes diffusion denoising loss. Timing correction loss Conditional fusion loss and reconstruction losses The four sub-losses are defined as follows: Diffusion denoising loss Supervised noise reduction The error between predicted noise and actual added noise;

[0061] in, The number of channels, representing the potential channel count; These represent the height and width of the potential representation, respectively. For noise reduction The output of the first In the step-by-step prediction of potential representation, the first Channel, First line, number The characteristic values ​​of the column; For the first Step reference in the latent representation of the first Channel, First line, number The characteristic values ​​of the column; For the first The retention factor of the step; Timing correction loss Supervision and The dynamic target area error;

[0062] in, For dynamic regional supervision weights; The number of latent representation pixels for the dynamic target region; Position in the optimized latent representation output by the timing corrector eigenvalues; Position in the reference latent representation of the current frame eigenvalues; The set of coordinates of a dynamic target in its latent representation; Conditional fusion loss Supervision and Consistency:

[0063] in, For noise reduction Bottleneck layer fusion characteristics; The target fusion feature is generated from the reference latent representation and conditional features; Reconstruction losses Supervision via decoder The error between the generated baseline candidate frame and the baseline range map;

[0064] in, These are the height and width of the baseline range map, respectively. For decoder Optimize latent representation The baseline candidate frame obtained after decoding is in pixels The predicted depth value at that location; This is the absolute value operator; The set of all pixel coordinates in the baseline range map; Finally, the multi-task weighted total loss function is obtained:

[0065] Among them, among them, Differential loss Conditional loss Correlation loss Reconstruction losses Task weighting factors; c) Training strategy design Training in the encoder and decoder End-to-end joint optimization is performed after training is completed and parameters are fixed. During training, the reference latent representation is obtained by encoding the current frame's baseline range map. As the supervision target, the network input includes: a noisy latent representation generated by adding noise to the range map encoded from the previous frame. Previous frame true range map and conditional features Noise reduction Responsible for The noise is gradually predicted, and its output is enhanced for timing consistency by a corrector before being passed by the decoder. Generate candidate results. The entire model optimizes the total loss function across multiple tasks and updates it jointly. By combining the training parameters of the modifier, we can achieve synergistic optimization of the generated quality in terms of denoising accuracy, semantic control, dynamic coherence, and geometric fidelity.

[0066] Figure 4 This is a schematic diagram of a two-frame point cloud sequence generated in an embodiment of the present invention. Figure 4 (a) represents the current time. The point cloud distribution Figure 4 (b) is the next moment. The point cloud distribution is as follows. During the temporal evolution of the point cloud in two frames, the spatial geometry of background 1 and background 2 did not undergo significant abrupt changes; the spatial position of vehicle 1 remained unchanged in both frames, remaining stationary; vehicle 2 moved to the right according to preset traffic rules, and its spatial geometry and trajectory maintained continuous evolution in the time dimension. The stable continuity of the above scene elements in spatial layout and dynamic behavior effectively verifies that the point cloud sequence generated by this invention possesses spatiotemporal consistency in dynamic scenes.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A text semantic-driven point cloud generation method for traffic scenes, characterized in that, This method generates a traffic scene LiDAR point cloud that conforms to the physical laws and geometric details of traffic based on the text description of the traffic scene input by the user. It includes a fine-tuning stage and a generation stage. The fine-tuning stage includes point cloud data processing and model fine-tuning. The generation stage includes traffic scene text semantic generation, traffic scene continuous range map generation, and traffic scene continuous point cloud generation.

2. The text semantic-driven traffic scene point cloud generation method according to claim 1, characterized in that, The fine-tuning stage includes: point cloud data processing and model fine-tuning. LiDAR point cloud samples are collected for traffic scenes, and the shape, size and traffic dynamic attributes of traffic participants are labeled. Point cloud and text alignment samples are constructed based on the labeling results. The pre-trained text encoding model is fine-tuned using the alignment samples to obtain a time-series pre-trained model.

3. The text semantic-driven traffic scene point cloud generation method according to claim 2, characterized in that, In the generation stage, the traffic scene text semantic generation includes: based on a temporal pre-trained model, temporal encoding is performed on the input traffic scene text description, and temporal point cloud semantic features and temporal grammatical features are extracted respectively; a temporal cross-attention layer that integrates point cloud semantic features and grammatical features is constructed through a temporal conditional injection mechanism, and the attention weight matrix of the previous frame is introduced to model inter-frame dynamic dependencies; then, through temporal frequency modulation, the semantic features are decomposed into low-frequency components corresponding to traffic stability structures and high-frequency components corresponding to dynamic behaviors, and the frequency domain gain is dynamically adjusted according to the scene complexity; finally, a structured temporal semantic vector sequence is generated to guide the generation of the range map. .

4. The text semantic-driven traffic scene point cloud generation method according to claim 3, characterized in that, In the generation phase, the continuous range map generation of the traffic scene includes: dynamic resolution adaptation based on a structured temporal semantic vector sequence; extracting dynamic target motion attribute sub-vectors and mapping them to semantic guidance vectors, and performing inter-frame correlation correction on the channel weights of the current frame in conjunction with the channel weights of the previous frame; generating multiple candidate range maps that conform to traffic rules based on the corrected channel weights and the current frame resolution; constructing a scoring function containing three dimensions—target attributes, scene rules, and inter-frame consistency—to evaluate the rationality of the candidate range maps and select the optimal frame as the range map of the current frame; iteratively executing the above process to generate range maps for each frame in sequence, resulting in a continuous range map sequence. .

5. The text semantic-driven traffic scene point cloud generation method according to claim 4, characterized in that, In the generation stage, the continuous point cloud generation of the traffic scene includes: based on the continuous range map sequence and LiDAR sensor parameters, performing temporal depth gradient correction and temporal coordinate fusion on each pixel in the single-frame range map to obtain the three-dimensional spatial coordinates of that pixel; collecting the three-dimensional coordinates of all pixels in a single frame to generate a single-frame LiDAR point cloud; performing the above process on each frame in the continuous range map sequence to finally obtain a continuous LiDAR point cloud sequence. .

6. The text semantic-driven traffic scene point cloud generation method according to claim 5, characterized in that, The traffic scene text semantic generation steps in the generation phase specifically include: 1) Temporal pre-trained language encoding: Based on the temporal pre-trained model, the dynamic attributes of traffic scenes in the text are encoded to obtain temporal point cloud semantic features. Encoding the logical relationships of traffic scenes in the text yields temporal grammar features. ; in, For the first Frame-time point cloud semantic feature vectors The first The embedding vectors of the traffic participants in the frame, including the category embedding vector, lane affiliation embedding vector, unit motion direction vector, and speed level embedding vector; For the first Frame temporal syntax feature vectors The first Spatial relationship vectors and interaction relationship vectors among traffic participants in a frame; 2) Timing condition injection: using subset slice and subset slices in Construct a cross-attention layer of point cloud semantic and syntactic features adapted to traffic scenarios; then combine it with the attention weight matrix of the previous frame. Obtain the current frame weight matrix. ; in, ; ; This is the normalization function; For feature dimensions; This is the time-series weight decay coefficient; For the first Semantic feature vectors of frame-time point clouds; For the first Frame temporal syntax feature vector; based on Fusion and The initial temporal semantic vector is obtained. ; in, The semantic vector fusion coefficient of the previous frame; For the first Initial temporal semantic vector of a frame; 3) Timing-frequency modulation: (Referring to section 2.2) Perform discrete wavelet transform to decompose the traffic stability features into time-series low-frequency components. Corresponding high-frequency components of traffic dynamic characteristics in time series ;based on and Determine the complexity level of traffic scenarios ,right Apply and Matched dynamic gain Features are reconstructed through inverse discrete wavelet transform and fused with traffic feature constraints from the first two frames to output the first... Frame structured temporal semantic vector ; in, This is an inverse discrete wavelet transform used to reconstruct high and low frequency features; For time series smoothing coefficients; The first Frame-structured temporal semantic vectors; 4) Generation of continuous semantic vectors: Repeat steps 1)-3) to generate the next continuous semantic vector in sequence. to Frame-based structured temporal semantic vectors are used to obtain a continuous sequence of structured temporal semantic vectors. .

7. The text semantic-driven traffic scene point cloud generation method according to claim 6, characterized in that, The specific steps for generating the continuous range map of the traffic scene in the generation phase include: 1) Dynamic resolution adaptation: based on the first Frame structured temporal semantic vector Extracted target motion attribute sub-vectors ;Target motion attribute subvector The components are normalized and weighted to obtain a comprehensive exercise activity index. ; in, for Norm; For component weights; for The function will Constraints ; Combination and Establish linear mapping rules to generate the initial resolution. ; in, These are the height and width of the base resolution, respectively; This is the complexity gain coefficient. ; The basic time-series weight decay coefficient; This is the complexity impact coefficient; Introducing the final resolution of the previous frame , generate the first Final frame resolution ; in, This is the transition smoothing coefficient. When =1, =0; This is the floor function; 2) Inter-frame correlation correction: based on the first Final frame resolution and Category embedding vectors of traffic participants Constructing horizontal time series features Vertical time series characteristics ;Transform the dynamic target motion attribute subvector Mapped to semantic guidance vectors through fully connected layers Introducing a speed level adjustment factor and the channel weights of the previous frame For the current frame channel weights Perform smoothing correction; in, This is the normalization function; This is a vector concatenation function that combines horizontal time-series features. Vertical time series characteristics By concatenating along the feature dimensions, a joint feature vector is obtained; This is the speed influence coefficient; for Norm; 3) Multiple candidate frame generation and filtering: based on the range map of the previous frame. Current frame resolution and channel weights Generate baseline candidate frames Generate by adding minute noises that conform to traffic patterns. A total of diverse candidate frames were obtained. Candidate frames { Then, a scoring function is constructed that includes three dimensions: target attributes, scene rules, and inter-frame consistency. The candidate frame with the highest score is selected as the current frame range map. If the highest score does not meet the score requirement, then increase the score. Regenerate candidate frames until the score meets the requirements; in, These are the weighting coefficients; For target attribute matching degree; For scene rule matching degree; Inter-frame consistency matching degree; Number of candidate range maps; For the first Index of candidate range maps; 4) Generation of continuous range plots: Repeat steps 3.1-3.3 to generate the next continuous range plot in sequence. to Frame range map, and then a continuous range map sequence is obtained. .

8. The text semantic-driven traffic scene point cloud generation method according to claim 7, characterized in that, The continuous point cloud generation steps for the traffic scene in the generation phase include: 1) Temporal depth gradient correction: The first... Frame range map Medium pixel Normalization yields And combined with the maximum ranging distance of the lidar sensor , obtained the Frame range map pixels original depth ; Based on the horizontal angular resolution of LiDAR With vertical angular resolution The pixel is obtained using the center difference method. depth change rate ; in, This represents the gradient of the original depth in the horizontal direction. The gradient of the original depth in the direction; , Representing pixels The horizontal column coordinates and vertical row coordinates in the range chart; Then introduce spatial correction coefficient , obtain pixels correction depth ; in, , It is a constant; 2) Temporal coordinate fusion: Based on the horizontal field of view of LiDAR Vertical field of view With current range resolution , obtain pixels horizontal angle with vertical angle ; in, ; ; Combined with the correction depth Get pixels Three-dimensional spatial coordinates; in, , , The order is number 1 Frame pixels Three-dimensional spatial coordinates; 3) Single-frame point cloud generation: For the first frame... Steps 1) and 2) are performed sequentially on all pixels in the frame range image to aggregate the three-dimensional spatial coordinates of each pixel, thus obtaining the first... Frame LiDAR point cloud ; 4) Continuous point cloud generation: For continuous range map sequences Steps 4.1 to 4.3 are executed frame by frame to obtain a continuous sequence of LiDAR point clouds of the traffic scene. .

9. A text semantic-driven point cloud generation system for traffic scenes, characterized in that, The system employs the method described in any one of claims 1 to 8.