Traffic scene construction method and system, electronic equipment and storage medium

By combining the on-board surround-view camera and 3D Gaussian algorithm with a diffusion model, the problem of limited scene reconstruction range in autonomous driving scenarios is solved, and efficient and accurate construction of traffic scenes and flexible editing of traffic participant models are achieved.

CN120635343APending Publication Date: 2025-09-12WUHAN KOTEI INFORMATICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510599547.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In autonomous driving scenarios, existing technologies have limited scope for scene reconstruction based on the vehicle's perspective and cannot provide comprehensive and accurate scene information.

Method used

Traffic scene images are collected through the on-board panoramic camera, and masks of dynamic targets are generated using the target segmentation algorithm. Static scenes are reconstructed using the 3D Gaussian algorithm, and multi-angle images of traffic participants are generated through the diffusion model. Finally, three-dimensional reconstruction is performed using the 3D Gaussian algorithm and imported into the Unreal Engine to construct a virtual traffic scene.

Benefits of technology

It improves the efficiency of traffic scene construction, overcomes the problem of limited scene perspective, ensures the comprehensiveness and accuracy of scene information, and supports the editing and reuse of traffic participant models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635343A_ABST
    Figure CN120635343A_ABST
Patent Text Reader

Abstract

The invention provides a traffic scene construction method and system, electronic equipment and a storage medium, and is used for the field of automatic driving simulation testing, and the method comprises the steps: obtaining a traffic scene image collected by a vehicle-mounted panoramic camera; segmenting a dynamic target in a traffic scene through a target segmentation algorithm, generating a mask of the dynamic target, taking the mask as prior information, and performing static scene reconstruction through a 3D Gaussian algorithm to obtain a first pth model; inputting the traffic participant bounding box into a diffusion model, and generating a multi-angle picture of the traffic participant through reasoning of the diffusion model; based on the multi-angle picture, performing three-dimensional reconstruction on the traffic participants through a 3D Gaussian algorithm to obtain a second pth model; and respectively converting the first pth model and the second pth model into a ply format, and importing the ply format into a UE unreal engine to construct a virtual traffic scene. Through the scheme, the problem that a traditional scene construction view angle is limited can be solved, a road traffic scene can be truly restored, and comprehensiveness and accuracy of scene information are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving simulation testing, and in particular relates to a traffic scene construction method, system, electronic device and storage medium. Background Art

[0002] The construction of traffic scenes typically relies on manual modeling or point cloud collection, which has numerous drawbacks. For one thing, manual modeling is extremely inefficient and labor-intensive. Furthermore, when faced with complex scene structures, it struggles to accurately replicate fine features of the real world, resulting in low model accuracy. Furthermore, while point cloud collection can capture certain spatial information, it cannot effectively capture scene semantics or details like appearance and texture, such as complex traffic flows, unique terrain, or objects with unusual textures. Clearly, these scene construction algorithms cannot achieve satisfactory restoration results.

[0003] With the emergence of NeRF (Neural Radiance Field) and 3D Gaussian techniques, new perspective synthesis technologies based on these cutting-edge technologies are demonstrating impressive scene reconstruction capabilities and generalization potential. However, these technologies typically require the acquisition and input of multi-view data centered around the target object. However, in autonomous driving scenarios, these technologies are constrained by the strict installation and design requirements of automotive-grade cameras, resulting in a fixed perspective. This allows for only a very limited range of view translation (approximately 3 meters). This results in a very limited range of scene reconstruction based on the vehicle's perspective, making it incapable of providing comprehensive and accurate scene information for autonomous driving. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a traffic scene construction method, system, electronic device, and storage medium, which are used to solve the problem of limited range of current scene reconstruction based on vehicle perspective.

[0005] In a first aspect of an embodiment of the present invention, a method for constructing a traffic scene is provided, comprising: Obtain traffic scene images captured by a vehicle-mounted panoramic camera, and divide the traffic scene images into equal intervals according to the capture positions; Obtain traffic scene images within any distance range, segment dynamic targets in the traffic scene images using a target segmentation algorithm, and generate masks of the dynamic targets; The mask of the dynamic target is used as prior information, and the static scene is reconstructed through the 3D Gaussian algorithm to obtain the first PTH model; Traffic participants in traffic scene images are detected using a target detection algorithm, and their bounding boxes are input into a diffusion model. Multi-angle images of traffic participants are then generated through diffusion model reasoning. The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; Based on the multi-angle pictures of traffic participants, the 3D Gaussian algorithm is used to reconstruct the traffic participants in three dimensions to obtain the second PTH model; The first PTH model and the second PTH model are converted into PLY format respectively, and imported into the UE Unreal Engine to build a virtual traffic scene.

[0006] In a second aspect of an embodiment of the present invention, a traffic scene construction system is provided, including: An acquisition and division module is used to obtain traffic scene images captured by the on-board panoramic camera and divide the traffic scene images into equal intervals according to the acquisition positions; The target segmentation module is used to obtain traffic scene images within any distance range, segment dynamic targets in the traffic scene images through the target segmentation algorithm, and generate masks of dynamic targets; The first reconstruction module is used to use the mask of the dynamic target as prior information and reconstruct the static scene through the 3D Gaussian algorithm to obtain the first PTH model; The diffusion generation module is used to detect traffic participants in traffic scene images using a target detection algorithm, input the traffic participant bounding boxes into the diffusion model, and generate multi-angle images of traffic participants through diffusion model reasoning; The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; The second reconstruction module is used to perform three-dimensional reconstruction of the traffic participant based on the multi-angle pictures of the traffic participant by using the 3D Gaussian algorithm to obtain a second PTH model; The scene construction module is used to convert the first PTH model and the second PTH model into PLY format respectively, and import them into the UE Unreal Engine to construct a virtual traffic scene.

[0007] In a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the steps of the method described in the first aspect of the embodiment of the present invention when executing the computer program.

[0008] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect of the embodiment of the present invention are implemented.

[0009] In an embodiment of the present invention, a panoramic camera captures images of the traffic scene surrounding the vehicle, and a 3D Gaussian algorithm is used to reconstruct the static traffic scene in three dimensions based on a mask of dynamic targets. A diffusion model is then used to generate multi-angle images of traffic participants, which are then reconstructed in three dimensions using the 3D Gaussian algorithm. The 3D models are then imported into the Unreal Engine for rendering, thereby achieving automated construction of traffic scenes. This not only improves the efficiency of traffic scene construction but also overcomes the problem of limited viewing angles in scene construction, ensuring the comprehensiveness and accuracy of scene information. By individually modeling traffic participants, a more realistic restoration of the traffic scene is possible, and the editing and reuse of traffic participant models in various traffic scenarios is facilitated. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flow chart of a traffic scene construction method provided by one embodiment of the present invention; Figure 2 A schematic diagram of the structure of a traffic scene construction system provided by one embodiment of the present invention; Figure 3 The present invention provides a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0012] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0013] It should be understood that the terms "including" and similar expressions in the specification, claims, and drawings of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, or apparatus comprising a series of steps or units is not limited to the listed steps or units. Furthermore, the terms "first" and "second" are used to distinguish between different objects and are not intended to describe a specific order.

[0014] See also Figure 1 , a flow chart of a traffic scene construction method provided by an embodiment of the present invention includes: S101, obtaining a traffic scene image captured by a vehicle-mounted surround view camera, and dividing the traffic scene image into equal intervals according to the capture positions; The on-board surround-view camera is a panoramic camera used to capture the road scene in front of the vehicle, providing a panoramic view of the vehicle at various angles, such as 180 degrees, 270 degrees, and 360 degrees. Traffic scene images are captured by the on-board surround-view camera. Based on the different capture locations of traffic scene images, traffic scene images captured within a certain distance range can be grouped together.

[0015] For example, assuming that a vehicle is traveling at a normal speed in an urban area, and there are at least three panoramic cameras in the front, left, and right, and the cameras are set to collect data every 2 seconds, the traffic scene images within a distance of 100m can be taken as a group, and the scene images can be divided, and the traffic scene images of every 100 meters can be cut into a folder for saving.

[0016] S102, acquiring a traffic scene image within any distance range, segmenting dynamic targets in the traffic scene using a target segmentation algorithm, and generating a mask of the dynamic targets; The object segmentation algorithm is used to detect specific objects in a segmented image, such as a fully convolutional network (FCN), a U-Net network, or an R-CNN network. Object segmentation algorithms can also segment the sky, signs, vehicles, and other objects in traffic scene images.

[0017] The dynamic target refers to the moving vehicles in the traffic scene image, that is, the vehicles in motion. It is possible to determine whether each vehicle is a dynamic target based on the continuous frame images. The dynamic target can be segmented by the target segmentation algorithm and a corresponding mask can be generated for the dynamic target.

[0018] Masking is the process of partially occluding a processed image with a selected target, thereby controlling the image processing area or process. Masking can be used to block dynamic targets in traffic scenes, thereby constructing a static scene based on the blocked image.

[0019] S103, using the mask of the dynamic target as prior information, reconstructing the static scene through the 3D Gaussian algorithm to obtain the first PTH model; Using the mask as prior information, the 3D Gaussian algorithm is used to reconstruct the static scene of the input traffic scene image. The 3D Gaussian algorithm uses a 3D Gaussian distribution to represent objects in the scene. Each Gaussian distribution can be viewed as an ellipsoid, with parameters such as the center point coordinates, the covariance matrix (representing rotation and scaling), opacity, and color parameters. By optimizing these parameters, the algorithm can learn a scene representation that is close to reality.

[0020] A pth model is a model file saved using the PyTorch framework, with a .pth extension. The pth file only stores the model parameters and does not contain the model's structural information. This file format is relatively small, and generally requires redefining the model structure when loading.

[0021] S104, detecting traffic participants in the traffic scene image using a target detection algorithm, inputting the traffic participant bounding boxes into a diffusion model, and generating multi-angle images of the traffic participants through diffusion model reasoning; The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; The object detection algorithm detects specific objects in an image and outputs their location and detection bounding box. Its core is to learn a classifier using a training set. Then, in a test image, it slides windows of varying sizes across the entire image, performing a classification on each scan to determine whether the current window is the target. This object detection algorithm can detect all traffic participants in a traffic scene image—that is, all vehicles and pedestrians present—and generate bounding boxes for these participants, which are then fed into a diffusion model.

[0022] The diffusion model is a generative model based on probability theory. It defines a process for gradually transforming a data distribution into a Gaussian noise distribution (forward diffusion) and learns a reverse process to gradually recover the original data from the noise (backward diffusion), achieving high-quality generation. Traffic participant detection frames are fed into the diffusion model, which generates multi-angle images of the traffic participant. Multi-angle images refer to images of a traffic participant (such as a vehicle) from different perspectives and can be inferred based on the diffusion model.

[0023] The diffusion model is a pre-trained diffusion model. Pre-training is a training method in deep learning that uses unsupervised learning to train the model on large-scale unlabeled data to provide high-quality initial weights for subsequent tasks. A target diffusion model structure is constructed, including a base network and a target upsampling stack network. This structure is then initialized using the pre-trained model to obtain an initial diffusion model. This model is then further trained using the training set for image generation tasks to obtain a target diffusion model capable of generating images from different angles. Using a pre-trained diffusion model can effectively improve the model's generalization capabilities.

[0024] S105, based on the multi-angle images of the traffic participants, performing three-dimensional reconstruction of the traffic participants using a 3D Gaussian algorithm to obtain a second PTH model; For traffic participants in the traffic scene, three-dimensional reconstruction can be performed based on multi-angle images using a 3D Gaussian algorithm, which uses a 3D Gaussian distribution to represent objects in the scene and obtain a second PTH model corresponding to the traffic participant.

[0025] S106: Convert the first PTH model and the second PTH model into PLY format respectively, and import them into the UE Unreal Engine to construct a virtual traffic scene.

[0026] ‌PLY (Polygon File Format) is a commonly used polygon file format used to describe the geometric shape and surface properties of three-dimensional models. It uses a simple text or binary format and can store geometric elements such as points, lines, and surfaces, and contains optional attribute information such as normals, texture coordinates, and colors.

[0027] Converting the first PTH model and the second PTH model into the PLY format can facilitate the data exchange of three-dimensional models and their application in the Unreal Engine. It has the advantages of being easy to read, highly scalable, widely supported by software, and highly efficient in storage.

[0028] Among them, the XV3DGS-UEPlugin plug-in is used to import the ply format model into UE for use.

[0029] Optionally, each Gaussian distribution parameter in the first PTH model and the second PTH model is converted into a PLY format, where the Gaussian distribution parameters include at least the center point coordinates, spherical harmonic coefficients, and transparency.

[0030] UE (Unreal Engine) is a real-time 3D rendering engine that can render static traffic scenes in the first PTH model and traffic participants in the second PTH model.

[0031] In this embodiment, based on traffic scene images captured by a panoramic camera, 3D reconstruction of both the static scene and traffic participants is performed, and then rendered using Unreal Engine. This not only faithfully reproduces the road traffic scene but also overcomes the limited perspective inherent in traditional modeling based on fixed cameras, ensuring comprehensive and accurate scene information. Furthermore, 3D reconstruction of traffic participants using a diffusion model and 3D Gaussian algorithm facilitates editing and reuse of these models in various simulated traffic scenarios.

[0032] In one embodiment, a second PTH model is obtained, converted into a mesh model through a mesh conversion algorithm, and after rendering the texture and map, the mesh model is imported into the UE Unreal Engine and saved as a reusable model; the mesh conversion algorithm is used to map the polygon data corresponding to the PTH model to the mesh data of the mesh model.

[0033] The mesh conversion algorithm converts a PLY (Polygon File Format) file into mesh data. This is typically done by reading the vertex and face information in the PLY file and converting this information into a mesh structure suitable for rendering or further processing. The processing generally includes reading the PLY file, creating a mesh object for associated drawing, etc.

[0034] In this embodiment, by converting the traffic participant model into a mesh model and rendering the mapping map and texture and loading it into the Unreal Engine, not only can the accurate modeling and reproduction of the traffic participants be achieved, thereby improving the simulation effect, but also the 3D model can be used as a reusable resource for direct calling in other scenes.

[0035] In one embodiment, during the process of constructing a virtual traffic scene based on the Unreal Engine, the virtual traffic scene is edited, and traffic participants in the virtual traffic scene are added, modified, or deleted.

[0036] In this embodiment, by editing the traffic participants in the virtual traffic scene, corresponding traffic participants can be added, modified or deleted, which can improve the flexibility of scene construction and facilitate operations in various traffic scenes.

[0037] It should be understood that the sequence numbers of the steps in the above embodiments do not imply a specific order of execution; the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0038] Figure 2 A schematic diagram of a traffic scene construction system provided in an embodiment of the present invention includes: The acquisition and division module 210 is used to obtain the traffic scene image acquired by the vehicle-mounted surround view camera and divide the traffic scene image into equal intervals according to the acquisition positions; The target segmentation module 220 is used to obtain a traffic scene image within any distance range, segment the dynamic targets in the traffic scene image using a target segmentation algorithm, and generate a mask of the dynamic targets; A first reconstruction module 230 is configured to use the mask of the dynamic target as prior information and reconstruct the static scene using a 3D Gaussian algorithm to obtain a first PTH model; Diffusion generation module 240, configured to detect traffic participants in traffic scene images using a target detection algorithm, input the traffic participant bounding boxes into a diffusion model, and generate multi-angle images of the traffic participants through diffusion model reasoning; The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; The diffusion model is a pre-trained diffusion model.

[0039] A second reconstruction module 250 is used to perform three-dimensional reconstruction of the traffic participant using a 3D Gaussian algorithm based on the multi-angle images of the traffic participant to obtain a second PTH model; The scene construction module 260 is used to convert the first PTH model and the second PTH model into PLY format respectively, and import them into the UE Unreal Engine to construct a virtual traffic scene.

[0040] The converting of the first and second PTH models into PLY formats respectively includes: The Gaussian distribution parameters in the first PTH model and the second PTH model are converted into a PLY format, where the Gaussian distribution parameters at least include center point coordinates, spherical harmonic coefficients, and transparency.

[0041] Optionally, converting the first PTH model and the second PTH model into PLY format respectively further includes: Obtain a second PTH model, convert the second PTH model into a mesh model through a mesh conversion algorithm, render the texture and map, import the mesh model into the UE Unreal Engine, and save it as a reusable model; the mesh conversion algorithm is used to map the polygon data corresponding to the PTH model to the mesh data of the mesh model.

[0042] Preferably, the virtual traffic scene is edited to add, modify or delete traffic participants in the virtual traffic scene.

[0043] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0044] Figure 3 This is a schematic diagram of the structure of an electronic device provided by one embodiment of the present invention. The electronic device is used to construct a traffic scene. Figure 3 As shown, the electronic device 3 of this embodiment includes: a memory 310, a processor 320 and a system bus 330, and the memory 310 includes an executable program 3101 stored thereon. It can be understood by those skilled in the art that Figure 3 The electronic device structure shown in the figure does not constitute a limitation to the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0045] The following combination Figure 3A detailed introduction to the various components of electronic equipment: Memory 310 can be used to store software programs and modules. Processor 320 executes the software programs and modules stored in memory 310 to perform various functional applications and data processing of the electronic device. Memory 310 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback). The data storage area may store data generated based on the use of the electronic device (such as cached data). Memory 310 may also include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state memory device.

[0046] The memory 310 includes an executable program 3101 for the interface generation method. The executable program 3101 can be divided into one or more modules / units, which are stored in the memory 310 and executed by the processor 320 to implement, for example, automated construction of traffic scenes. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the executable program 3101 in the electronic device 3. For example, the executable program 3101 can be divided into functional modules such as an acquisition and segmentation module, an object segmentation module, a first reconstruction module, a diffusion generation module, a second reconstruction module, and a scene construction module.

[0047] The processor 320 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 310 and accessing data stored in the memory 310, it performs various functions of the electronic device and processes data, thereby monitoring the overall status of the electronic device. Optionally, the processor 320 may include one or more processing units; preferably, the processor 320 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, application programs, etc., and the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 320.

[0048] The system bus 330 connects the various functional components within the computer and can transmit data, address information, and control information. It can be a PCI bus, ISA bus, or CAN bus, for example. Instructions from the processor 320 are transmitted to the memory 310 via the bus, and the memory 310 feeds data back to the processor 320. The system bus 330 is responsible for the exchange of data and instructions between the processor 320 and the memory 310. Of course, the system bus 330 can also connect to other devices, such as network interfaces and display devices.

[0049] In an embodiment of the present invention, the executable program executed by the processing 320 included in the electronic device includes: Obtain traffic scene images captured by a vehicle-mounted panoramic camera, and divide the traffic scene images into equal intervals according to the capture positions; Obtain traffic scene images within any distance range, segment dynamic targets in the traffic scene images using a target segmentation algorithm, and generate masks of the dynamic targets; The mask of the dynamic target is used as prior information, and the static scene is reconstructed through the 3D Gaussian algorithm to obtain the first PTH model; Traffic participants in traffic scene images are detected using a target detection algorithm, and their bounding boxes are input into a diffusion model. Multi-angle images of traffic participants are then generated through diffusion model reasoning. The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; Based on the multi-angle pictures of traffic participants, the 3D Gaussian algorithm is used to reconstruct the traffic participants in three dimensions to obtain the second PTH model; The first PTH model and the second PTH model are converted into PLY format respectively, and imported into the UE Unreal Engine to build a virtual traffic scene.

[0050] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0051] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0052] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A traffic scene construction method, characterized in that: include: Obtain traffic scene images captured by a vehicle-mounted panoramic camera, and divide the traffic scene images into equal intervals according to the capture positions; Obtain traffic scene images within any distance range, segment dynamic targets in the traffic scene images using a target segmentation algorithm, and generate masks of the dynamic targets; The mask of the dynamic target is used as prior information, and the static scene is reconstructed through the 3D Gaussian algorithm to obtain the first PTH model; Traffic participants in traffic scene images are detected using a target detection algorithm, and their bounding boxes are input into a diffusion model. Multi-angle images of traffic participants are then generated through diffusion model reasoning. The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; Based on the multi-angle pictures of traffic participants, the 3D Gaussian algorithm is used to reconstruct the traffic participants in three dimensions to obtain the second PTH model; The first PTH model and the second PTH model are converted into PLY format respectively, and imported into the UE Unreal Engine to build a virtual traffic scene.

2. The method according to claim 1, characterized in that The diffusion model is a pre-trained diffusion model.

3. The method according to claim 1, characterized in that The converting of the first and second PTH models into PLY formats respectively includes: The Gaussian distribution parameters in the first PTH model and the second PTH model are converted into a PLY format, where the Gaussian distribution parameters at least include center point coordinates, spherical harmonic coefficients, and transparency.

4. The method according to claim 1, wherein The converting the first PTH model and the second PTH model into the PLY format respectively further includes: Obtain the second PTH model, convert the second PTH model into a mesh model through a mesh conversion algorithm, render textures and maps, import the mesh model into the UE Unreal Engine, and save it as a reusable model; The mesh conversion algorithm is used to map the polygon data corresponding to the PTH model into the mesh data of the mesh model.

5. The method according to claim 1, wherein The converting the first PTH model and the second PTH model into PLY format respectively and importing them into the UE Unreal Engine to construct the virtual traffic scene also includes: Edit the virtual traffic scene, add, modify or delete traffic participants in the virtual traffic scene.

6. A traffic scene construction system, characterized in that: include: An acquisition and division module is used to obtain traffic scene images captured by the on-board panoramic camera and divide the traffic scene images into equal intervals according to the acquisition positions; The target segmentation module is used to obtain traffic scene images within any distance range, segment dynamic targets in the traffic scene images through the target segmentation algorithm, and generate masks of dynamic targets; The first reconstruction module is used to use the mask of the dynamic target as prior information and reconstruct the static scene through the 3D Gaussian algorithm to obtain the first PTH model; The diffusion generation module is used to detect traffic participants in traffic scene images using a target detection algorithm, input the traffic participant bounding boxes into the diffusion model, and generate multi-angle images of traffic participants through diffusion model reasoning; The dynamic target is a moving vehicle, and the traffic participants are all vehicles and pedestrians in the traffic scene image; The second reconstruction module is used to perform three-dimensional reconstruction of the traffic participant based on the multi-angle pictures of the traffic participant by using the 3D Gaussian algorithm to obtain a second PTH model; The scene construction module is used to convert the first PTH model and the second PTH model into PLY format respectively, and import them into the UE Unreal Engine to construct a virtual traffic scene.

7. The system according to claim 6, characterized in that The diffusion model is a pre-trained diffusion model.

8. The system according to claim 6, wherein: The converting the first PTH model and the second PTH model into the PLY format respectively further includes: Obtain the second PTH model, convert the second PTH model into a mesh model through a mesh conversion algorithm, render textures and maps, import the mesh model into the UE Unreal Engine, and save it as a reusable model; The mesh conversion algorithm is used to map the polygon data corresponding to the PTH model into the mesh data of the mesh model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the traffic scene construction method according to any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the steps of a traffic scene construction method according to any one of claims 1 to 5 are implemented.