Driving scene generation model training method and device, equipment and storage medium

By using the scaling factor of Gaussian units to determine the target loss function for training in the driving scene generation model, overfitting is avoided, the visual quality problem of the generation model under sparse input conditions is solved, and the accuracy of driving scene geometry and visual quality are improved.

CN121600347APending Publication Date: 2026-03-03CHONGQING CHANGAN AUTOMOBILE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511799512.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing driving scene generation models are prone to overfitting during the generation process, which leads to a decline in the visual quality of the generated driving scenes, especially under sparse input conditions where the reconstruction performance bottleneck is obvious.

Method used

By obtaining the minimum and intermediate scaling factors corresponding to Gaussian elements in the training sample set, the target loss function is determined, and this function is used to train the initial driving scene generation model, applying thickness constraints to avoid overfitting.

Benefits of technology

Ensure the accuracy of the generated driving scene geometry and improve visual quality, especially under sparse input conditions to generate accurate driving scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600347A_ABST
    Figure CN121600347A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a driving scene generation model, a training device of the driving scene generation model, electronic equipment and a computer readable storage medium, and relates to the field of vehicle driving. The method comprises the following steps: acquiring a training sample set comprising a driving scene image; based on each training sample in the training sample set, a model is generated through the initial driving scene, and a first scaling factor and a second scaling factor are obtained; determining a target loss function based on the first scaling factor and the second scaling factor corresponding to each Gaussian primitive; training the initial driving scene generation model based on the target loss function to obtain a trained driving scene generation model; wherein the trained driving scene model performs thickness constraint on Gaussian primitives in the driving scene when generating the driving scene. Based on the scheme, the driving scene generation model can be prevented from being over-fitted when generating the driving scene, so that the visual quality of the generated driving scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle driving, and more particularly to a training method for a driving scene generation model, a training device for the driving scene generation model, an electronic device, and a computer-readable storage medium in the field of vehicle driving. Background Technology

[0002] Data collection of vehicle driving scenes requires extensive on-site data collection using numerous test vehicles and drivers, resulting in high acquisition costs. Therefore, to reduce these costs, driving scene generation models can be used. However, these models are prone to overfitting, which degrades the visual quality of the generated driving scenes. Summary of the Invention

[0003] This application provides a training method for a driving scene generation model, a training device for the driving scene generation model, an electronic device, and a computer-readable storage medium. This method can avoid overfitting when the driving scene generation model generates driving scenes, thereby ensuring the accuracy of the geometry when the driving scene generation model generates driving scenes and improving the visual quality of the generated driving scenes.

[0004] Firstly, this application provides a method for training a driving scene generation model, the method comprising: Obtain a training sample set including images of driving scenes; Based on each training sample in the training sample set, a first scaling factor and a second scaling factor are obtained by generating a model through an initial driving scenario; wherein, the first scaling factor and the second scaling factor are the scaling factors corresponding to the Gaussian elements of each training sample; the first scaling factor represents the minimum scaling factor among the multiple scaling factors corresponding to each Gaussian element, and the second scaling factor represents the scaling factor other than the maximum scaling factor and the minimum scaling factor among the multiple scaling factors; the directions corresponding to the multiple scaling factors are perpendicular to each other. The target loss function is determined based on the first and second scaling factors corresponding to each Gaussian element. The initial driving scene generation model is trained based on the objective loss function to obtain the trained driving scene generation model; wherein, the trained driving scene model applies thickness constraints to the Gaussian elements in the driving scene when generating the driving scene.

[0005] In this embodiment, when a training sample set including driving scene images is obtained, the initial driving scene generation model is not directly trained using this training sample set. Instead, based on each training sample in the training sample set, the initial driving scene generation model first obtains the minimum scaling factor (i.e., the first scaling factor) and the intermediate scaling factor (i.e., the second scaling factor) among multiple scaling factors corresponding to the Gaussian elements of each training sample. Then, a loss function (i.e., the target loss function) is determined using the minimum scaling factor and the intermediate scaling factor corresponding to each Gaussian element. The initial driving scene generation model is then trained using this loss function. This allows the trained driving scene generation model to impose thickness constraints on the Gaussian elements in the driving scene when generating the driving scene, thereby flattening the Gaussian elements. This enables the trained driving scene generation model to fit smoother and more realistic surface geometry when generating the driving scene, avoiding the presence of blurry "fog" or "cloud" Gaussian elements in the driving scene generated by the trained driving scene model. This prevents overfitting when the trained driving scene generation model generates the driving scene. Furthermore, the trained driving scene generation model can ensure the accuracy of the geometry of the driving scene when generating it, thereby improving the visual quality of the generated driving scene.

[0006] Furthermore, since the trained driving scene generation model can avoid overfitting when generating driving scenes and ensure the accuracy of the geometric shape of the driving scene, even with limited input data (i.e., sparse input), the trained driving scene generation model can still ensure the accuracy of the geometric shape of the generated driving scene when generating driving scenes corresponding to sparse input, thereby improving the visual quality of the generated driving scene.

[0007] In some embodiments, determining the target loss function based on the first scaling factor and the second scaling factor corresponding to each Gaussian element includes: Based on the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian element, the constraint terms corresponding to each Gaussian element are determined. Based on the constraints corresponding to each Gaussian element, the target loss function is determined; wherein the loss value of the target loss function is less than or equal to the preset loss value.

[0008] In this embodiment, the constraint terms corresponding to each Gaussian primitive are determined by the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian primitive. This ratio constraint avoids the trained driving scene generation model from overfitting the local scaling features of the training sample set, further preventing overfitting when generating driving scenes. Furthermore, by using a target loss function whose loss value is less than or equal to a preset loss value, a convergence target can be provided for the training of the initial driving scene generation model. This prevents the initial driving scene generation model from overfitting the appearance and sacrificing geometry during training, further preventing overfitting when generating driving scenes.

[0009] In some embodiments, the initial driving scene generation model includes a first generation module and a second generation module. The first generation module generates a static background in the driving scene, and the second generation module generates a dynamic foreground in the driving scene. The initial driving scene generation model is trained based on a target loss function to obtain a trained driving scene generation model, including: The first generation module is trained based on the target loss function and the first loss function to obtain the trained first generation module; wherein, the first loss function includes at least one of the color loss function corresponding to the static background, the depth loss function, and the loss function of the sky background opacity in the static background; the trained first generation module applies thickness constraints to the Gaussian elements in the static background when generating the static background; The second generation module is trained based on the target loss function and the second loss function to obtain the trained second generation module; wherein, the second loss function includes at least one of the color loss function, depth loss function and dynamic foreground opacity loss function corresponding to the dynamic foreground; the trained second generation module applies thickness constraints to the Gaussian units in the dynamic foreground when generating the dynamic foreground; The first generation module after training is combined with the second generation module after training to obtain the trained driving scene generation model.

[0010] In this embodiment, the first generation module (i.e., the static background module) is trained by using the target loss function and the loss functions corresponding to the static background. This allows the training of the first generation module to focus on the geometry without considering the influence of the dynamic foreground on the parameters of the first generation module. As a result, the trained first generation module can more accurately reproduce the geometry of the static background, thereby avoiding overfitting when generating the static background.

[0011] Furthermore, by training the second generation module (i.e., the dynamic foreground module) using the target loss function and the loss functions corresponding to the dynamic foreground, the training of the second generation module can focus on the modeling and motion posture of the dynamic target, without considering the influence of the static background on the parameters of the second generation module. This allows the trained second generation module to more accurately reproduce the modeling and motion posture of the dynamic target, thereby avoiding overfitting when generating the dynamic foreground.

[0012] In some embodiments, the first generation module is trained based on the target loss function and the first loss function to obtain the trained first generation module, including: The target loss function is weighted with the first preset weight, and the first loss function is weighted with the second preset weight to obtain the first weighted loss function; The first generation module is trained by using the first weighted loss function to obtain the trained first generation module.

[0013] In this embodiment of the application, the first generation module is trained by a weighted loss function obtained by weighting multiple loss functions (i.e., the target loss function and the first loss function). This avoids the situation where the first generation module can only fit single-dimensional features and cannot fit multiple-dimensional features when trained by a single loss function. This further avoids overfitting when the first generation module generates a static background.

[0014] In some embodiments, the second generation module is trained based on the target loss function and the second loss function to obtain the trained second generation module, including: The target loss function is weighted with the third preset weight, and the second loss function is weighted with the fourth preset weight to obtain the second weighted loss function. The second generation module is trained by using the second weighted loss function to obtain the trained second generation module.

[0015] In this embodiment of the application, the second generation module is trained by a weighted loss function obtained by weighting multiple loss functions (i.e., the target loss function and the second loss function). This avoids the situation where the second generation module can only fit single-dimensional features and cannot fit multiple-dimensional features when trained by a single loss function. This further avoids overfitting when the trained second generation module generates dynamic foregrounds.

[0016] In some embodiments, the initial driving scenario generation model described above further includes an adjustment module, and the method further includes: The adjustment module is trained based on the target loss function and the third loss function to obtain the trained adjustment module. The third loss function includes at least one of the following: color loss function and depth loss function corresponding to the driving scene, loss function of sky background opacity in the driving scene, loss function of driving scene opacity, and semantic loss function in the driving scene. The trained adjustment module performs appearance adjustment on the driving scene when generating the driving scene. The above-mentioned combination of the trained first generation module and the trained second generation module yields the trained driving scene generation model, including: The first generation module, the second generation module, and the adjustment module after training are combined to obtain the trained driving scene generation model.

[0017] In this embodiment, the adjustment module is trained by using the target loss function and various loss functions corresponding to the driving scene. This allows the trained adjustment module to adjust the appearance of both static background and dynamic foreground, avoiding only adjusting the appearance of static background or dynamic foreground. This further prevents the trained adjustment module from overfitting when generating driving scenes.

[0018] In some embodiments, the adjustment module is trained based on the target loss function and the third loss function to obtain the trained adjustment module, including: The target loss function is weighted with the fifth preset weight, and the third loss function is weighted with the sixth preset weight to obtain the third weighted loss function. The adjustment module is trained using a third weighted loss function to obtain the trained adjustment module.

[0019] In this embodiment of the application, the adjustment module is trained by a weighted loss function obtained by weighting multiple loss functions (i.e., the target loss function and the third loss function). This avoids the situation where the adjustment module can only adjust a single dimension feature and cannot adjust multiple dimensions feature when trained by a single loss function. This further avoids overfitting of the trained adjustment module when generating driving scenarios.

[0020] Secondly, this application provides a method for generating driving scenarios, the method comprising: Acquire input data including images of the driving scene; The input data is fed into the trained driving scene generation model to generate the target driving scene corresponding to the input data; wherein, the trained driving scene generation model is trained by the method described in the first aspect above; Output the target driving scenario.

[0021] In this embodiment, since the trained driving scene generation model can avoid overfitting when generating driving scenes, when the acquired input data including driving scene images is input into the trained driving scene generation model, the trained driving scene generation model can avoid overfitting when generating driving scenes corresponding to the input data, ensuring the accuracy of the geometric shape of the driving scene, thereby improving the visual quality of the generated driving scene.

[0022] Thirdly, this application provides a training device for a driving scene generation model, the device comprising: The acquisition module is used to acquire a training sample set including images of driving scenes; The processing module is used to obtain a first scaling factor and a second scaling factor based on each training sample in the training sample set through an initial driving scene generation model. The first and second scaling factors are the scaling factors corresponding to the Gaussian elements of each training sample. The first scaling factor represents the minimum scaling factor among the multiple scaling factors corresponding to each Gaussian element, and the second scaling factor represents the scaling factor other than the maximum and minimum scaling factors. The directions corresponding to each scaling factor are mutually perpendicular. Based on the first and second scaling factors corresponding to each Gaussian element, a target loss function is determined. The initial driving scene generation model is trained based on the target loss function to obtain a trained driving scene generation model. The trained driving scene model applies thickness constraints to the Gaussian elements in the driving scene when generating the driving scene.

[0023] Fourthly, this application provides a driving scene generation apparatus, the apparatus comprising: The acquisition module is used to acquire input data, including images of the driving scene. The processing module is used to input the input data into the trained driving scene generation model and generate the target driving scene corresponding to the input data; wherein, the trained driving scene generation model is trained by the method in the first aspect mentioned above; and outputs the target driving scene.

[0024] Fifthly, this application provides an electronic device including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the method of the first aspect above, or the method of the second aspect above.

[0025] Sixthly, this application provides a computer program product comprising: computer program code or instructions, which, when executed on a computer, cause the computer to perform the method described in the first aspect, or the method described in the second aspect.

[0026] In a seventh aspect, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect, or the method described in the second aspect. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0028] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0029] Figure 1 This is a schematic diagram of an autonomous driving scenario provided in an embodiment of this application.

[0030] Figure 2 This is another schematic diagram of an autonomous driving scenario provided in the embodiments of this application.

[0031] Figure 3 This is a flowchart illustrating a training method for a driving scene generation model provided in an embodiment of this application.

[0032] Figure 4 This is a schematic diagram of the architecture of an autonomous driving scene generation model provided in an embodiment of this application.

[0033] Figure 5 This is another schematic diagram of the training method for a driving scene generation model provided in the embodiments of this application.

[0034] Figure 6 This is another flowchart illustrating a training method for a driving scene generation model provided in this application embodiment.

[0035] Figure 7 This is a flowchart illustrating a method for generating a driving scene according to an embodiment of this application.

[0036] Figure 8 This is a schematic diagram of the structure of the training device for the driving scene generation model provided in this application embodiment.

[0037] Figure 9This is a schematic diagram of the structure of the driving scene generation device provided in the embodiments of this application.

[0038] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0039] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described below in conjunction with the accompanying drawings. The embodiments described below are only some embodiments of this application, not all embodiments. Therefore, the described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0041] In the following description, references to “some embodiments” or “other embodiments” describe a subset of all possible embodiments, but it is understood that “some embodiments” or “other embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0042] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0044] With the rapid development of computer vision technology, it has been widely applied to the 3D reconstruction of autonomous driving scenarios. However, under conditions of incomplete data coverage, few effective measurement points, and low information density from sparse onboard sensor data, existing 3D reconstruction models may suffer from significant performance bottlenecks when reconstructing autonomous driving scenarios. This is due to the limited modeling capabilities of existing 3D reconstruction models for large-scale and complex autonomous driving scenarios, as well as the inherent limitations of autonomous driving datasets. These bottlenecks make it difficult to achieve high-precision and robust autonomous driving scene reconstruction. This performance bottleneck is particularly pronounced when the vehicle is in a high-speed dynamic driving state. Furthermore, while image data acquired through onboard surround-view multi-camera systems can provide a wide field of view, it also suffers from low image overlap between different viewpoints and uneven lighting conditions across different viewpoints, further impacting the robustness and accuracy of autonomous driving scene reconstruction. Furthermore, the high complexity and dynamism of the geometric topology of autonomous driving scenarios, as well as the pose changes of highly dynamic targets at different timestamps, will lead to spatiotemporal inconsistencies in autonomous driving scenarios. This will result in discontinuities in the time dimension and inconsistencies in the spatial dimension of the reconstructed autonomous driving scenarios, further affecting the robustness and accuracy of autonomous driving scenario reconstruction.

[0045] The complexity of autonomous driving scenarios lies in the intricate interplay between rapidly changing static backgrounds and highly dynamic targets. These elements interact, with rapid changes in the static background interfering with vehicle sensors' recognition of highly dynamic targets, while the movement of highly dynamic targets can obscure or expose the static background, thus complicating the autonomous driving scenario. Consequently, given this complexity, reconstructing autonomous driving scenarios using existing 3D reconstruction models also faces challenges, resulting in significant performance bottlenecks. For example... Figure 1 As shown, tree 101 is a static background, and pedestrian 102 is a high-dynamic target; tree 101 obscures pedestrian 102. Alternatively, as... Figure 2 As shown, road sign 103 is a static background, while vehicle 104 in front is a highly dynamic target. Road sign 103 is revealed after vehicle 104 changes lanes.

[0046] The inherent limitations of autonomous driving datasets mean that when collecting autonomous driving data, it is often only possible to capture autonomous driving data from an extremely sparse and limited observation perspective, resulting in low redundancy of autonomous driving data.

[0047] It should be noted that although static backgrounds generally represent objects without the ability to move automatically (e.g., trees, traffic lights), highly dynamic targets and vehicles in motion are constantly changing, causing the static background to be affected by at least one of the following: vehicle status, road surface status, and environmental status. This results in a rapidly changing static background. For example, when a vehicle is traveling at 120 km / h, static roadside guardrails, road signs, and trees flash by on the left side of the field of vision due to the high speed, while the outline of a distant bridge rapidly magnifies from a blurry dot to a clear structure due to perspective, thus creating a rapidly changing static background. Furthermore, highly dynamic targets can represent objects other than vehicles that can move autonomously (e.g., other vehicles, pedestrians). Among them, vehicle status may include, but is not limited to, at least one of vehicle driving perspective, vehicle speed, vehicle steering, etc.; road surface status may include, but is not limited to, at least one of icy road surface, snowy road surface, wet and slippery road surface, dry road surface, etc.; environmental status may include, but is not limited to, the influence of at least one of ambient light, weather conditions, etc.; and weather conditions may include, but is not limited to, rainy and snowy weather, sunny weather, foggy and hazy weather, strong wind weather, etc.

[0048] Currently, to improve the robustness and accuracy of autonomous driving scene reconstruction, autonomous driving scene generation models such as the "3D Gaussian Splatting Model (3DGS Model)" can be used to reconstruct autonomous driving scenes. The 3DGS Model can reconstruct autonomous driving scenes relatively well from multi-view images. However, when reconstructing autonomous driving scenes, the 3DGS Model relies on dense multi-view image input. When the number of multi-view images is severely insufficient (e.g., less than or equal to 3 images), i.e., sparse input, the 3DGS Model is extremely prone to overfitting. This causes the 3DGS Model to only remember the autonomous driving scene information from the training views, but it cannot reconstruct a geometrically reasonable and visually consistent new perspective image of the autonomous driving scene. This results in numerous blurry, distorted, and / or unrealistic "floating" artifacts in the reconstructed autonomous driving scene image, leading to low visual quality. The main reason for the overfitting of the 3DGS Model is that when the number of multi-view images is severely insufficient, the 3DGS Model, in order to minimize training loss, will prioritize fitting the pixel details in these limited images rather than reproducing the real autonomous driving scene. For example, in an image, false dark spots on the road surface caused by lighting and shadows might be mistakenly modeled as real scene features by the 3DGS model. This results in the reconstructed autonomous driving scene only adapting to the training viewpoint and failing to generalize to other viewpoints, leading to overfitting. Furthermore, the limited number of multi-view images provides only a limited set of autonomous driving scene features. The 3DGS model cannot learn more autonomous driving scene features, causing the trained 3DGS model parameters to be unusable when faced with unlearned autonomous driving scene features. This prevents the 3DGS model from reconstructing unlearned autonomous driving scene features, resulting in a loss of generalization ability and overfitting.

[0049] In view of this, if overfitting of the model used to reconstruct autonomous driving scenes can be avoided when the input is sparse, the visual quality of the reconstructed autonomous driving images can be improved. Therefore, this application proposes a training method, a training device, an electronic device, and a computer-readable storage medium for a driving scene generation model. Through the embodiments of this application, when a training sample set including driving scene images is obtained, the initial driving scene generation model is not directly trained using this training sample set. Instead, based on each training sample in the training sample set, the minimum scaling factor and the intermediate scaling factor among multiple scaling factors corresponding to the Gaussian elements of each training sample are obtained through the initial driving scene generation model. Then, a loss function is determined using the minimum scaling factor and the intermediate scaling factor corresponding to each Gaussian element, and the initial driving scene generation model is trained using this loss function. This allows the trained driving scene model to impose thickness constraints on the Gaussian elements in the driving scene when generating driving scenes, which can prevent overfitting of the driving scene generation model when generating driving scenes. This ensures the accuracy of the geometry when the driving scene generation model generates driving scenes, thereby improving the visual quality of the generated driving scenes.

[0050] The following is combined Figures 3 to 6 The training method for the driving scene generation model provided in the embodiments of this application is described in detail.

[0051] Figure 3 This is a flowchart illustrating a training method for a driving scene generation model provided in an embodiment of this application. The method can be executed by an electronic device.

[0052] For example, such as Figure 3 As shown, the method 300 includes the following implementation process: S310, acquire a training sample set including driving scene images.

[0053] For example, to avoid overfitting of existing driving scene generation models (e.g., 3DGS Model), a training sample set including driving scene images can be obtained, and the existing driving scene generation model (which can be called the "initial driving scene generation model") can be trained using this training sample set, so that the trained driving scene generation model will not overfit or will reduce the probability of overfitting.

[0054] To make it easier to understand, you can combine... Figure 4 The training method for the driving scene generation model provided in the embodiments of this application will be described in detail. Figure 4 This is a schematic diagram of the architecture of an autonomous driving scene generation model provided in an embodiment of this application.

[0055] For example, such as Figure 4As shown, the autonomous driving scene generation model 400 may include a data processing module 401, a segmentation module 402, a static background module 403, a dynamic foreground module 404, an adjustment module 405, and a rendering module 406. At least one set of data collected by the autonomous vehicle (which may be referred to as a "training sample set") serves as input, and the generated autonomous driving scene serves as output.

[0056] Optionally, at least one piece of data collected by the vehicle may include at least one time-series image collected by an onboard camera and at least one 3D point cloud data collected by a LiDAR (Light Detection and Ranging) system.

[0057] It should be understood that the image mode of the time sequence image can be any of the following: Red Green Blue (RGB), Cyan Magenta Yellow Key (CMYK), Hue Saturation Value (HSV), etc., and the embodiments of this application do not limit it.

[0058] S320, based on each training sample in the training sample set, generates a model through an initial driving scenario to obtain the first scaling factor and the second scaling factor.

[0059] Wherein, the first scaling factor and the second scaling factor are the scaling factors corresponding to the Gaussian elements of each training sample; the first scaling factor represents the minimum scaling factor among the multiple scaling factors corresponding to each Gaussian element, and the second scaling factor represents the scaling factor other than the maximum and minimum scaling factors among the multiple scaling factors; the directions corresponding to the multiple scaling factors are perpendicular to each other.

[0060] For example, when obtaining the training sample set, each training sample in the training sample set can be input into the initial driving scene generation model. Through the initial driving scene generation model, the Gaussian units (which can be called "3D Gaussian units") corresponding to each training sample in the training sample set can be determined. When obtaining each Gaussian unit through the initial driving scene generation model, multiple scaling factors corresponding to each Gaussian unit can be determined. Then, the multiple scaling factors are sorted by size to obtain the smallest scaling factor (which can be called the "first scaling factor") among the multiple scaling factors, and the scaling factors other than the largest and smallest scaling factors (which can be called the "second scaling factors").

[0061] For example, when the training sample set is obtained, each training sample in the training sample set can be input into... Figure 4The static background module 403 determines the sparse point cloud corresponding to each training sample in the training sample set, constructs Gaussian elements corresponding to each training sample through the sparse point cloud, determines multiple scaling factors corresponding to each Gaussian element, sorts the multiple scaling factors by size, and obtains the first scaling factor and the second scaling factor among the multiple scaling factors.

[0062] Each 3D Gaussian element can be represented by an equivalent scaling matrix. S and rotation matrix R The covariance matrix formed First, the scaling factors for each 3D Gaussian primitive are extracted in three mutually perpendicular directions. For example, for an ellipsoid, the scaling factors for each of the three semi-axis directions (normal axis, major axis, and minor axis) are defined as follows: scaling factor s1 for the normal axis, scaling factor s2 for the major axis, and scaling factor s3 for the minor axis. Specifically, the smallest scaling factor among the scaling factors s1, s2, and s3 for each Gaussian primitive along the normal axis is taken as the first scaling factor, and the remaining scaling factors (excluding the largest and smallest scaling factors) among the scaling factors s1, s2, and s3 for each Gaussian primitive along the normal axis are taken as the second scaling factor. For example, if the scaling factor s1 in the normal direction of the i-th Gaussian element is less than the scaling factor s2 in the major semi-axis direction and the scaling factor s3 in the minor semi-axis direction, then the scaling factor s1 in the normal direction is the first scaling factor corresponding to the i-th Gaussian element, and the scaling factor s2 in the major semi-axis direction is the second scaling factor corresponding to the i-th Gaussian element.

[0063] S330, based on the first scaling factor and the second scaling factor corresponding to each Gaussian element, determines the target loss function.

[0064] For example, when the first scaling factor and the second scaling factor corresponding to each Gaussian unit are obtained, a loss function (which can be called the "target loss function") that can be used to avoid overfitting of the initial driving scenario generation model can be determined through the first scaling factor and the second scaling factor corresponding to each Gaussian unit. For example, the regularization loss function (which can be denoted as "...") (”).

[0065] For example, when the static background module 403 obtains the first scaling factor and the second scaling factor corresponding to each Gaussian element, it can determine the target loss function through the first scaling factor and the second scaling factor corresponding to each Gaussian element.

[0066] In one implementation, when determining the target loss function using the first and second scaling factors corresponding to each Gaussian element, the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian element can be determined first (i.e., Then, the constraint terms corresponding to each Gaussian element are determined by the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian element.

[0067] For example, a penalty rule is designed based on the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian unit. The deviation of the ratio of the first scaling factor to the second scaling factor from a preset loss value (e.g., 0) is used to form a constraint term. The lower the deviation of the ratio of the first scaling factor to the second scaling factor from the preset value, the smaller the loss value corresponding to the target loss function.

[0068] The preset loss value can represent the maximum loss value for flattening Gaussian elements, such as 0, 0.1, or 0.3, etc., and this application embodiment does not limit this.

[0069] Furthermore, after obtaining the constraint terms corresponding to each Gaussian element, the target loss function can be determined through these constraint terms. The loss value of the target loss function is less than or equal to a preset loss value.

[0070] For example, when determining the target loss function through the constraint terms corresponding to each Gaussian element, the constraint terms corresponding to each Gaussian element can be accumulated to obtain the accumulated constraint terms, and the accumulated constraint terms can be determined as the target loss function.

[0071] S340, the initial driving scene generation model is trained based on the objective loss function to obtain the trained driving scene generation model.

[0072] Among them, the trained driving scene model applies thickness constraints to the Gaussian elements in the driving scene when generating the driving scene.

[0073] For example, when the target loss function is obtained, the initial driving scene generation model can be trained using this target loss function to obtain a trained driving scene generation model. When generating driving scenes (e.g., the aforementioned autonomous driving scene), the trained driving scene model can impose thickness constraints on the Gaussian elements in the driving scene to flatten them.

[0074] It should be understood that the trained driving scenario model can generate autonomous driving scenarios for autonomous vehicles, as well as driving scenarios for manually driven vehicles. Whether a vehicle can achieve autonomous driving does not affect the driving scenarios generated by the trained driving scenario model.

[0075] For example, when the target loss function is obtained, the autonomous driving scene generation model 400 can be trained using the target loss function to obtain the trained autonomous driving scene generation model 400. This allows the trained autonomous driving scene generation model 400 to impose thickness constraints on the Gaussian elements in the driving scene when generating the driving scene, so as to flatten the Gaussian elements in the driving scene.

[0076] In such Figure 3 In the method 300 shown, when obtaining the training sample set including driving scene images, the initial driving scene generation model is not directly trained using this training sample set. Instead, based on each training sample in the training sample set, the initial driving scene generation model obtains the minimum scaling factor (i.e., the first scaling factor) and the intermediate scaling factor (i.e., the second scaling factor) among multiple scaling factors corresponding to the Gaussian elements of each training sample. Then, a loss function (i.e., the target loss function) is determined using the minimum scaling factor and the intermediate scaling factor corresponding to each Gaussian element. The initial driving scene generation model is then trained using this loss function. This allows the trained driving scene generation model to impose thickness constraints on the Gaussian elements in the driving scene when generating the driving scene, thereby flattening the Gaussian elements. This enables the trained driving scene generation model to fit smoother and more realistic surface geometry when generating the driving scene, avoiding the presence of blurry "fog" or "cloud" Gaussian elements in the driving scene generated by the trained driving scene model. This prevents overfitting when the trained driving scene generation model generates the driving scene. Furthermore, the trained driving scene generation model can ensure the accuracy of the geometric shape of the generated driving scene, thus improving the visual quality of the generated driving scene. Also, because the trained driving scene generation model can avoid overfitting and ensure the accuracy of the geometric shape of the driving scene, even with limited input data (i.e., sparse input), the trained driving scene generation model can still ensure the accuracy of the geometric shape of the generated driving scene when generating driving scenes corresponding to sparse input, thereby improving the visual quality of the generated driving scene.

[0077] It should be noted that S310~S340 above is a brief description of the training method of the driving scene generation model provided in the embodiments of this application. The following will further explain... Figure 3 The specific implementation methods shown in the embodiments are described in detail below: In step S340, when training the initial driving scene generation model using the objective loss function to obtain the trained driving scene generation model, the first generation module (i.e., the aforementioned static background module 403) and the second generation module (i.e., the aforementioned dynamic foreground module 404) included in the initial driving scene generation model can be trained first and second generation modules to obtain the trained first generation module and the trained second generation module. The first generation module is used to generate the static background in the driving scene, and the second generation module is used to generate the dynamic foreground in the driving scene.

[0078] When training the first generation module, it is necessary to train the first generation module together with the target loss function and the first loss function to obtain the trained first generation module, so that the trained first generation module can perform thickness constraints on the Gaussian elements in the static background when generating the static background.

[0079] The first loss function includes the color loss function corresponding to the static background (which can be denoted as "color loss function"). ), depth loss function (which can be denoted as "") The loss function for the opacity of the sky background in a static background (which can be denoted as "") and the loss function for the opacity of the sky background in a static background (which can be denoted as "") At least one of the following: .

[0080] In one implementation, when training the first generation module using the target loss function and the first loss function, the target loss function and the first loss function can be weighted first to obtain a weighted loss function (which can be called the "first weighted loss function"). The first generation module is then trained using the first weighted loss function to obtain the trained first generation module.

[0081] For example, the target loss function and its corresponding weights (which can be called "first preset weights") and the first loss function and its corresponding weights (which can be called "second preset weights") are weighted to obtain the first weighted loss function. First weighted loss function = target loss function × first preset weights + first loss function × second preset weights.

[0082] For example, when the target loss function is When, the corresponding first preset weight is And, in the first loss function is When the second preset weight is 1, and / or, in the first loss function, the corresponding second preset weight is 1. When, the corresponding second preset weight is And / or, in the first loss function is When, the corresponding second preset weight is For example, in the first loss function is , , The objective loss function is When, the corresponding first weighted loss function is denoted as .

[0083] Furthermore, when training the second generation module, it is necessary to train the second generation module together with the target loss function and the second loss function to obtain the trained second generation module, so that the trained second generation module can impose thickness constraints on the Gaussian elements in the dynamic foreground when generating the dynamic foreground.

[0084] The second loss function includes the color loss function corresponding to the dynamic foreground (which can be denoted as "color loss function"). ), depth loss function (which can be denoted as "") The loss function for dynamic foreground opacity (which can be denoted as "") and dynamic foreground opacity (which can be denoted as "") At least one of the following: .

[0085] In one implementation, when training the second generation module using the target loss function and the second loss function, the target loss function and the second loss function can be weighted first to obtain a weighted loss function (which can be called the "second weighted loss function"). The second generation module is then trained using the second weighted loss function to obtain the trained second generation module.

[0086] For example, a weighted average is applied to the target loss function and its corresponding weights (referred to as the "third preset weights"), and the second loss function and its corresponding weights (referred to as the "fourth preset weights") to obtain the second weighted loss function. The second weighted loss function = target loss function × third preset weights + second loss function × fourth preset weights. The third preset weights and the first preset weights can be the same or different; this embodiment does not limit this.

[0087] For example, when the target loss function is When, the corresponding third preset weight is And, in the second loss function is When the corresponding fourth preset weight is 1, and / or, in the second loss function, When, the corresponding fourth preset weight is And / or, in the second loss function is When, the corresponding fourth preset weight is For example, in the second loss function is , , The objective loss function is When, the corresponding second weighted loss function is denoted as .

[0088] Optionally, when training the first generation module and the second generation module, the order of the spherical harmonic coefficients of each Gaussian element can be set to be less than or equal to a preset order (e.g., order 0).

[0089] Furthermore, after obtaining the first generation module and the second generation module after training, the first generation module and the second generation module after training can be combined to obtain the trained driving scene generation model. That is, the trained driving scene generation model can be composed of the first generation module and the second generation module after training.

[0090] In one implementation, in addition to training the first generation module and the second generation module, the adjustment module (i.e., the aforementioned adjustment module 405) included in the initial driving scene generation model can also be trained to obtain the trained adjustment module.

[0091] When training the adjustment module, it is necessary to train the module by using both the target loss function and the third loss function to obtain the trained adjustment module, so that the trained adjustment module can make appearance adjustments to the driving scene when generating the driving scene.

[0092] The third loss function includes the color loss function corresponding to the driving scenario (which can be denoted as "color loss function"). ") and depth loss function (which can be denoted as "") The loss function for the sky background opacity in driving scenarios (which can be denoted as "") The loss function for driving scene opacity (which can be denoted as "") ") and the semantic loss function in the driving scenario (which can be denoted as "") At least one of the following: .

[0093] In one implementation, when training the adjustment module using the target loss function and the third loss function, the target loss function and the third loss function can be weighted first to obtain a weighted loss function (which can be called the "third weighted loss function"). The adjustment module is then trained using the third weighted loss function to obtain the trained adjustment module.

[0094] For example, the target loss function and its corresponding weights (which can be called the "fifth preset weights") and the third loss function and its corresponding weights (which can be called the "sixth preset weights") are weighted to obtain the third weighted loss function. The third weighted loss function = target loss function × fifth preset weights + third loss function × sixth preset weights.

[0095] For example, when the target loss function is At that time, the corresponding fifth preset weight is And, in the third loss function, When the corresponding sixth preset weight is 1, and / or, in the third loss function, At that time, the corresponding sixth preset weight is And / or, in the third loss function, At that time, the corresponding sixth preset weight is For example, in the third loss function... , , , , The objective loss function is When, the corresponding third weighted loss function is denoted as .

[0096] Optionally, when training the adjustment module, the order of the spherical harmonic coefficients of each Gaussian element can be set to be greater than a preset order (e.g., order 0). For example, the order of the spherical harmonic coefficients of each Gaussian element is 3.

[0097] Furthermore, after obtaining the first generation module, the second generation module, and the adjustment module after training, the first generation module, the second generation module, and the adjustment module after training can be combined to obtain the trained driving scene generation model. That is, the trained driving scene generation model can be composed of the first generation module, the second generation module, and the adjustment module after training.

[0098] In conjunction with the above embodiments, Figure 4 The various modules in this paper will be described in detail to further illustrate the training method of the driving scene generation model provided in the embodiments of this application.

[0099] The data processing module 401 can be used to preprocess each training sample in the training sample set (e.g., the aforementioned time-series images and 3D point cloud data), and to detect the 3D bounding boxes corresponding to high dynamic targets (which can be simply referred to as "dynamic targets"), track the motion pose of the high dynamic targets, thereby obtaining the preprocessed training sample set. The preprocessed training sample set is then input to the segmentation module 402.

[0100] For example, when obtaining time-series images captured by an in-vehicle camera, the data processing module 401 can process the captured time-series images using Structure-from-Motion (SfM) technology to generate the camera pose of the in-vehicle camera, thereby obtaining the image viewpoint corresponding to the time-series image. Furthermore, when processing the captured time-series images using SfM, image point cloud data (which can be called "SfM point cloud") corresponding to the time-series image can also be obtained. This image point cloud data is then fused with 3D point cloud data to obtain fused point cloud data (which can be called "sparse point cloud"), adding image point cloud data to the 3D point cloud data to obtain more comprehensive point cloud data. Moreover, when obtaining the sparse point cloud, object detection can be used to detect the 3D bounding boxes corresponding to dynamic targets, and a tracking algorithm can be used to obtain the motion pose of the dynamic targets.

[0101] The segmentation module 402 can be used to perform static and dynamic segmentation on the preprocessed training sample set to obtain the static background and dynamic foreground in the preprocessed training sample set, thus obtaining the scene segmentation result. The scene segmentation result is then input into the static background module 403 and the dynamic foreground module 404, respectively. For example, the static background can be at least one of trees, roads, and sky, and the dynamic foreground can be at least one of other vehicles, pedestrians, etc.

[0102] For example, upon obtaining the preprocessed training sample set, the preprocessed training samples can first undergo structured parsing to identify object types and boundaries, and to separate static and dynamic background regions. After parsing, the results are input to the depth estimation module to generate a depth image corresponding to the preprocessed training sample. This depth image is then input to the segmentation module 402 to generate a segmentation image (which can be called an "instance segmentation image") corresponding to the depth image, thus distinguishing between static background and dynamic foreground in the preprocessed training sample. The depth image transforms the "visual depth perception" in the preprocessed training sample into quantifiable "spatial distance information." The depth estimation module can be any of the following: Depth Anything V2, Hourglass network, or Marigold monocular depth estimation module. The segmentation model can be any of the following: Segment Anything Model (SAM), instance segmentation model, etc.

[0103] The static background module 403 can be used to construct the static geometric skeleton corresponding to the static background in the scene segmentation result, and input the static geometric skeleton to the adjustment module 405. Furthermore, the static background module 403 can obtain the initial geometric shape of the static background from the sparse point cloud. It also filters out the static point cloud corresponding to the static background from the sparse point cloud and initializes the 3D Gaussian parameters of the static point cloud to obtain the initialized 3D Gaussian parameters (which can be called the "original 3D Gaussian parameters"). The 3D Gaussian parameters of the static point cloud can include position coordinates. Rotation matrix Scaling matrix spherical harmonic coefficient Opacity At least one of the following. For example, if a sparse point cloud includes point clouds corresponding to other vehicles and point clouds corresponding to trees, then the point cloud corresponding to trees can be considered as a static point cloud.

[0104] It should be noted that, regarding the spherical harmonic coefficients During initialization, the spherical harmonic coefficients need to be... Initialize to a lower order, such as 0, to ensure that the color of the static background remains consistent under different viewing angles. This avoids color flickering or unrealistic situations caused by viewing angle changes introduced by higher-order spherical harmonic coefficients (SH), thereby improving the visual quality and stability of the generated autonomous driving scene.

[0105] For example, spherical harmonic coefficients The color value corresponding to order 0 can be obtained through the projection of the static point cloud. Specifically, first, the position coordinates of the static point cloud in each training sample are determined, then the color value corresponding to those position coordinates in each training sample is determined, and this color value is defined as the spherical harmonic coefficient. The color value corresponding to level 0.

[0106] To address the issue of significant ambiguity between geometry and appearance in the autonomous driving scene generation model 400 under sparse input, the model learns incorrect geometric shapes and compensates for this by using a complex higher-order SH (Structured View) to represent view-related appearances. This results in the generated autonomous driving scene conforming to the training view but not to the untrained view, failing completely on the untrained view. This is because the incorrect geometry learned by the autonomous driving scene generation model 400 results in an uneven, bumpy geometry. While a higher-order SH is used to forcibly correct this, generating a smoother geometry, this method violates the true geometric form and is only suitable for the training view. It fails to generate a smoother geometry when used on the untrained view (the "new view"). For example, if the model learns incorrect geometry and fits the side of the vehicle as "convex in the middle and concave at the top and bottom," a higher-order SH can compensate by making "convex areas darker and concave areas brighter from a frontal view," thus offsetting the geometric error and generating a scene that matches the training view. Figure 1 The result is a "flat car body". However, when switching to the new "right front" view, the compensation logic of the higher-order SH is only adapted to the trained front view. It can no longer offset geometric errors through the higher-order SH, and will generate a car body with "a bulge in the middle and a depression at the top and bottom", with chaotic color and brightness, which does not conform to the geometry of a real car at all.

[0107] To address the issue that the generated autonomous driving scene completely fails on untrained views, this application embodiment can solve the problem through three stages: static geometric skeleton construction, dynamic target specification modeling and posture optimization, and global appearance fine-tuning.

[0108] When constructing the static geometric skeleton, all collected training samples are required, but only the 3D Gaussian parameters of the static point cloud in the static background module 403 are optimized, and the spherical harmonic coefficients are... Initialized to order 0 (also known as "freezing higher-order appearance"), only basic color values ​​independent of image viewpoint are learned, thus avoiding the static background module 403 from fitting using higher-order SH. Furthermore, the aforementioned objective loss function is added, i.e., a planar prior is added. Since the planar prior is subject to significant planar geometric constraints, the static background module 403 can construct a static geometric skeleton with accurate geometry and consistent image viewpoint by utilizing the motion time difference between multiple training samples. It should be understood that during the construction of the static geometric skeleton, the 3D Gaussian parameters of the dynamic foreground in the dynamic foreground module 204 are frozen, i.e., the 3D Gaussian parameters of the dynamic foreground in the dynamic foreground module 204 are kept constant to prevent changes in the 3D Gaussian parameters of the dynamic foreground in the dynamic foreground module 204 as the 3D Gaussian parameters of the static point cloud change.

[0109] The main reason for overfitting in the static background module 403 with sparse input (which can be called a "sparse view") is that when the data is limited, the static background module 403 will abandon precise geometry and tend to fit "fog" or "cloud" Gaussian primitives with different directions but the same shape (which can be called "isotropic"), using blurry "fog" or "cloud" Gaussian primitives to replace the real plane. This causes the fitting result to overfit the limited data and deviate from the real scene (i.e., overfitting), thus failing to obtain accurate surface geometry. However, most object surfaces in autonomous driving scenarios (e.g., roads, vehicle bodies, building walls, road guardrails, etc.) generally exhibit planar characteristics locally, i.e., local planes. Using "fog" or "cloud" Gaussian primitives cannot obtain accurate surface geometry of the object. To fit an accurate surface geometry, a target loss function (e.g., a regularized loss function) can be added. This regularized loss function forces a fit on each 3D Gaussian primitive in each training sample, resulting in planes with different orientations and shapes (which can be called "anisotropy"), rather than isotropic "fog" or "cloud-like" Gaussian primitives. For example, a plane that is stretched long in its "extension direction" and compressed thin in its "perpendicular plane direction" becomes a flat plane.

[0110] Each 3D Gaussian element can be represented by an equivalent scaling matrix. S and rotation matrix R The covariance matrix formed The definition is as follows: First, the scaling factors corresponding to each 3D Gaussian element in three different directions are extracted. Then, the target loss function (also known as the "regularization loss term") is determined using these three scaling factors. The three different directions can represent mutually perpendicular directions, such as the three semi-major directions of an ellipsoid: the normal axis direction, the major semi-axis direction, and the minor semi-axis direction.

[0111] Specifically, the scaling factors corresponding to each 3D Gaussian element in the three semi-major directions of the ellipsoid are first extracted: the scaling factor s1 in the normal direction, the scaling factor s2 in the major semi-major direction, and the scaling factor s3 in the minor semi-major direction. When obtaining these three scaling factors, to obtain the plane corresponding to the 3D Gaussian element, the scaling factor s1 in the normal direction can be set much smaller than the scaling factors s2 and s3 in the major and minor semi-major directions, thus compressing the 3D Gaussian element into a flat plane that does not bulge towards the perpendicular plane.

[0112]

[0113] In formula (1), Let i represent the target loss function, and let i represent the i-th 3D Gaussian element among N 3D Gaussian elements. Let represent the minimum scaling factor after sorting the i-th 3D Gaussian elements. This represents the intermediate scaling factor after sorting the i-th 3D Gaussian elements. This represents the ratio of the first scaling factor to the second scaling factor of the i-th 3D Gaussian element (i.e., the ratio mentioned above). ).in, This represents the smallest scaling factor among the scaling factors s1 along the normal axis, s2 along the major semi-axis, and s3 along the minor semi-axis. Let represent the scaling factor that is neither the largest nor the smallest among the scaling factors s1 along the normal axis, s2 along the major axis, and s3 along the minor axis. For example, if the scaling factor s1 along the normal axis of the i-th 3D Gaussian element is less than the scaling factor s2 along the major axis and less than the scaling factor s3 along the minor axis, then the corresponding s3 is... Scaling factor s1 in the direction of the normal axis, corresponding to s² is the scaling factor along the major semi-axis. And, The corresponding semi-axis length direction is the direction to be flattened. Because It can represent the length of the thinnest direction in a 3D Gaussian unit. Represents the length in the thicker direction of a 3D Gaussian element, through and The ratio can reflect whether 3D Gaussian pixels are flattened. and The smaller the ratio, the more the i-th 3D Gaussian element is compressed. and The larger the ratio, the less the i-th 3D Gaussian element is compressed. This is because, in order to "flatten" the i-th 3D Gaussian element into a tiny planar projection (i.e., a flat plane), it is possible to... and The smaller the ratio, the better. The smaller, that is, the more... Minimize, by minimizing The static background module 403 can "flatten" 3D Gaussian primitives into tiny planar projections, thereby fitting a smoother and more realistic surface geometry, avoiding overfitting, and fundamentally suppressing problems such as blurry, distorted geometric noise and floating artifacts in autonomous driving scenes generated under sparse input, thus improving the visual quality of the generated autonomous driving scenes.

[0114] When adding the target loss function, the static background module 403 needs to first calculate the static background loss using all the collected training samples. Specifically, the static background in the scene segmentation result can be segmented using a mask to block out the dynamic foreground, resulting in a mask corresponding to the static background (which can be called the "static background mask"). The static background loss is then calculated using this static background mask.

[0115]

[0116] In formula (2), This represents the total loss between the rendered static background and the static background of each training sample. This represents the color loss between the rendered static background and the static backgrounds of each training sample. This color loss can include the mean absolute error loss of color (which can be denoted as "L1 Loss1") and the structural similarity loss (D-SSIM Loss1). L1 Loss1 measures the color difference between the rendered static background and the static backgrounds of each training sample, while D-SSIM Loss1 measures the differences in structure, texture, and brightness between the rendered static background and the static backgrounds of each training sample, so as to ensure the accuracy of the rendered static background in terms of visual appearance. This represents the depth loss (mean absolute error loss of depth) between the rendered static background and the static background of each training sample (i.e., LiDAR projection depth), to ensure the realism and accuracy of the geometry of the rendered static background, not just the accuracy of the visual appearance; where LiDAR projection depth is obtained by rendering 3D point cloud data onto the image plane through the camera projection matrix. The binary cross-entropy loss represents the opacity of the sky region in a static background. express The corresponding weights express The corresponding weights Representing a static background The corresponding weights.

[0117] It should be understood that All are pre-trained weights, and For a relatively large weight, for example, Greater than This is to provide strong geometric constraints for the static background module 403.

[0118] The dynamic foreground module 404 can be used for the canonical modeling and pose optimization of dynamic targets in the scene segmentation results, and input the pose-optimized dynamic target canonical modeling into the adjustment module 405. For each dynamic target, 3D point cloud data within the 3D bounding box of each dynamic target can be collected, and the 3D point cloud data can be transformed into the vehicle coordinate system (which can be called the "canonical space"), for example, a coordinate system with the starting time (e.g., 0 o'clock). Furthermore, the dynamic foreground module 404 can filter out the dynamic point cloud corresponding to the dynamic foreground from the sparse point cloud, and initialize the 3D Gaussian parameters of the dynamic point cloud to obtain the initialized 3D Gaussian parameters (which can be called the "original 3D Gaussian parameters"). The 3D Gaussian parameters of the dynamic point cloud can include position coordinates. Rotation matrix Scaling matrix spherical harmonic coefficient Opacity At least one of the following. For example, if a sparse point cloud includes point clouds corresponding to other vehicles and point clouds corresponding to trees, then the point cloud corresponding to other vehicles can be used as a dynamic point cloud.

[0119] It should be noted that, regarding the spherical harmonic coefficients During initialization, the spherical harmonic coefficients need to be... Initialize to a lower order, such as 0, to ensure that the color of the dynamic foreground remains consistent under different viewpoints, avoiding color flickering or unrealistic situations caused by viewpoint changes introduced by higher-order spherical harmonic coefficients, thereby improving the visual quality and stability of the generated autonomous driving scene.

[0120] When performing dynamic target specification modeling, all collected training samples are required, but only the 3D Gaussian parameters of the dynamic point cloud in the dynamic foreground module 404 are optimized, and the spherical harmonic coefficients are... Initialized to order 0, only basic color values ​​independent of the image viewpoint are learned; and, during dynamic target pose optimization, the rotation increment of each dynamic target's pose per frame is learned (which can be denoted as "..."). ) and translation increment (which can be denoted as " It should be understood that during dynamic target specification modeling and posture optimization, the 3D Gaussian parameters of the static background in the static background module 403 will be frozen, that is, the 3D Gaussian parameters of the static background in the static background module 403 will be kept unchanged to avoid the 3D Gaussian parameters of the static background in the static background module 403 changing with the 3D Gaussian parameters of the dynamic point cloud.

[0121] For example, when performing dynamic target specification modeling, it is necessary to first calculate the loss of the dynamic foreground using all collected training samples. Specifically, the dynamic foreground in the scene segmentation results can be masked to block out the static background, resulting in a mask corresponding to the dynamic foreground (which can be called the "dynamic foreground mask"). The loss of the dynamic foreground can then be calculated using this dynamic foreground mask.

[0122]

[0123] In formula (3), This represents the total loss between the rendered dynamic foreground and the dynamic foreground of each training sample. This represents the color loss between the rendered dynamic foreground and the dynamic foreground of each training sample. This color loss can include the mean absolute error loss of color (which can be denoted as "L1 Loss2") and the structural similarity loss (D-SSIM Loss2). L1 Loss2 measures the color difference between the rendered dynamic foreground and the dynamic foreground of each training sample, while D-SSIM Loss2 measures the differences in structure, texture, and brightness between the rendered dynamic foreground and the dynamic foreground of each training sample, so as to ensure the accuracy of the rendered dynamic foreground in terms of visual appearance. This represents the depth loss (mean absolute depth error loss) between the rendered dynamic foreground and the dynamic foreground of each training sample (i.e., LiDAR projection depth), to ensure the realism and accuracy of the rendered dynamic foreground geometry, not just the accuracy of the visual appearance. The binary cross-entropy loss represents the dynamic foreground opacity. express The corresponding weights express The corresponding weights Indicates dynamic foreground The corresponding weights.

[0124] It should be understood that All are pre-trained weights, and For a relatively large weight, for example, Greater than This is to provide strong geometric constraints for the dynamic foreground module 404.

[0125] For example, when optimizing the pose of a dynamic target, a dynamic foreground mask can be used to optimize the 3D Gaussian model corresponding to the dynamic target in the vehicle coordinate system. The rotation increment and translation increment of the 3D Gaussian model can be determined by the deviation between the position coordinates of the 3D Gaussian model in the vehicle coordinate system and the position coordinates in the world coordinate system. The global pose of the 3D Gaussian model can then be determined by the rotation increment and translation increment.

[0126] The adjustment module 405 can be used to adjust the appearance of the static geometric skeleton constructed by the static background module 403 and the 3D Gaussian model constructed by the dynamic foreground module 404 (i.e., the aforementioned global appearance fine-tuning). Specifically, when fine-tuning the appearance parameters of the static geometric skeleton and the 3D Gaussian model constructed by the dynamic foreground module 404 end-to-end, only the appearance of the static geometric skeleton and the 3D Gaussian model is adjusted, without changing their geometric shape, thus avoiding geometric deformation problems caused by global appearance fine-tuning, such as road surface deformation and vehicle contour misalignment. The appearance parameters may include at least one of the following: view-related reflection parameters and ambient lighting parameters. The adjustment of the appearance of the 3D Gaussian model can also be referred to as "canonical Gaussian base".

[0127] During global appearance fine-tuning, the 3D Gaussian parameters of the static point cloud and the dynamic point cloud are unfrozen, and the spherical harmonic coefficients are adjusted. With spherical harmonic coefficient All are subjected to order-up processing to increase the spherical harmonic coefficients. With spherical harmonic coefficient The order, for example, the spherical harmonic coefficients With spherical harmonic coefficient The order of all spherical harmonics is increased to the 3rd order, or, alternatively, the spherical harmonic coefficients can be... With spherical harmonic coefficient The order of the digits is increased to 4th or 5th order, but this application does not limit the implementation of the embodiments.

[0128] For example, when the adjustment module 405 performs global appearance fine-tuning, it can first be initialized. After initialization, the Neural Transient Module (MLP) within the adjustment module 405 can be trained. When training the MLP, the full-image loss of the autonomous driving scene needs to be calculated using all collected training samples, including static background loss and dynamic foreground loss.

[0129]

[0130] In formula (4), This represents the total loss between the rendered autonomous driving scene and the autonomous driving scenes of each training sample. This represents the color loss between the rendered autonomous driving scene and the autonomous driving scenes of each training sample. The color loss can include the mean absolute error loss of color (which can be denoted as "L1 Loss3") and the structural similarity loss (D-SSIM Loss3). L1 Loss3 measures the color difference between the rendered autonomous driving scene and the autonomous driving scenes of each training sample, while D-SSIM Loss3 measures the differences in structure, texture, and brightness between the rendered autonomous driving scene and the autonomous driving scenes of each training sample, so as to ensure the accuracy of the rendered autonomous driving scene in terms of visual appearance. This represents the depth loss (mean absolute depth error loss) between the rendered autonomous driving scene and the autonomous driving scenes of each training sample (i.e., LiDAR projection depth), to ensure the realism and accuracy of the geometry of the rendered autonomous driving scene, and not just the accuracy of the visual appearance. The binary cross-entropy loss represents the opacity of the autonomous driving scenario. express The corresponding weights express The corresponding weights Indicating autonomous driving scenarios The corresponding weights.

[0131] It should be understood that All are pre-trained weights, and For relatively small weights, for example, Less than This is to provide strong geometric constraints for the adjustment module 405.

[0132] Rendering module 406 can be used to render static geometric skeletons, 3D Gaussian models, and autonomous driving scenes from new perspectives at any given time. When rendering static geometric skeletons, it can render the corresponding Gaussian primitives (also known as "static Gaussian fields") according to the standard 3DGS rendering pipeline. When rendering a 3D Gaussian model, one can use the 3D Gaussian model (which can be called a "canonical Gaussian model"). The rotation matrix and translation vector in the vehicle coordinate system, along with the aforementioned rotation increment and translation increment, are used to obtain the global pose of the 3D Gaussian model. The 3D Gaussian model is then transformed to the world coordinate system through global pose transformation to obtain the 3D Gaussian model in the world coordinate system (which can be denoted as ""). (”).

[0133]

[0134] In formula (5), This represents the global pose of the 3D Gaussian model. This represents the rotation matrix of the 3D Gaussian model at time t. This represents the rotation increment of the 3D Gaussian model at time t. Let represent the translation vector of the 3D Gaussian model at time t. This represents the translation increment of the 3D Gaussian model at time t.

[0135] By from Calculation of higher order spherical harmonic coefficients The model uses the base color and base opacity, and then uses the timestamp (t) and view direction (d) of the new viewpoint as input to the MLP to retrieve the residual color and residual opacity corresponding to the timestamp and view direction of the new viewpoint. The base color and residual color, base opacity and residual opacity are then superimposed to obtain the superimposed color and opacity. The Gaussian elements in the autonomous driving scene are then dynamically projected into the 2D image space, and the rendering color and rendering opacity of each Gaussian element are calculated through alpha-blending. This rendering color and rendering opacity are then used to render the autonomous driving scene, improving the modeling capability and fidelity of the generated autonomous driving scene.

[0136]

[0137] In formula (6), Indicates the color after superposition. Indicates the basic color. Indicates the residual color.

[0138] In formula (7), Indicates the opacity of the overlay. This indicates the base opacity. This indicates the residual opacity.

[0139] Figure 5 This is another schematic diagram illustrating a training method for a driving scene generation model provided in an embodiment of this application. This method can be executed by an electronic device.

[0140] For example, such as Figure 5 As shown, the method 500 includes the following implementation process: S501, Data Input.

[0141] For example, in order to train an initial driving scene generation model, a training sample set including driving scene images can be obtained, and when the training sample set is obtained, it can be used as training data to input into the initial driving scene generation model.

[0142] S502, Data Preprocessing.

[0143] For example, the training sample set is input into the data processing module 401 for data preprocessing to obtain the preprocessed sample set.

[0144] S503, scene segmentation.

[0145] For example, the preprocessed sample set is input into the segmentation module 402 for scene segmentation to separate the static background and dynamic foreground in the preprocessed sample set, thereby obtaining the static background and dynamic foreground.

[0146] S504, Initialization of the driving scenario generation model.

[0147] For example, after scene segmentation is completed, the initial driving scene generation model can be initialized to obtain the initialized initial driving scene generation model.

[0148] S505, Static background module initialization, initializes the 0th order spherical harmonic coefficients.

[0149] For example, initializing the initial driving scene generation model may include initializing the static background module 403, and when initializing the static background module 403, the order of the spherical harmonic coefficients of each Gaussian element may be initialized to order 0, and order 0 is used to obtain the spherical harmonic coefficients.

[0150] S506, Static geometric skeleton construction.

[0151] For example, after the static background module 403 is initialized, the static geometric skeleton corresponding to each training sample can be constructed through the initialized static background module 403.

[0152] S507, Dynamic Foreground Module Initialization, Initializes 0th Order Spherical Harmonic Coefficients.

[0153] For example, initializing the initial driving scene generation model may include initializing the dynamic foreground module 404, and when initializing the dynamic foreground module 404, the order of the spherical harmonic coefficients of each Gaussian element may be initialized to order 0, and order 0 is obtained as the spherical harmonic coefficients.

[0154] It should be understood that S505 and S507 can be executed simultaneously or sequentially, and the embodiments of this application do not limit this.

[0155] S508, Dynamic Object Modeling.

[0156] For example, after the dynamic foreground module 404 is initialized, each dynamic object (i.e., the aforementioned dynamic target) can be modeled through the initialized dynamic foreground module 404.

[0157] S509, global appearance fine-tuning.

[0158] For example, when obtaining a static geometric skeleton through the initialized static background module 403 and a dynamic object model through the initialized dynamic foreground module 404, the modeling of the static geometric skeleton and the dynamic object can be finely adjusted globally through the adjustment module 405 to obtain a finely adjusted driving scene.

[0159] S510 generates the final driving scenario generation model.

[0160] For example, when the fine-tuned driving scene is obtained, the difference between the fine-tuned driving scene and the training sample set can be used to adjust the loss function (i.e., ...) in the static background module 403, the dynamic foreground module 404, and the adjustment module 405. , , Each of these steps is optimized until the difference between the rendered driving scene and the training sample set is small, and the rendered driving scene is basically similar to or exactly the same as the training sample set.

[0161] It should be noted that, Figure 5 All steps are in Figure 3 and Figure 4 The corresponding embodiments are described in detail, and will not be repeated here.

[0162] Figure 6 This is another flowchart illustrating a training method for a driving scene generation model provided in this application embodiment. This method can be executed by an electronic device.

[0163] For example, such as Figure 6 As shown, the method 600 includes the following implementation process: S601, Obtain a training sample set including driving scene images.

[0164] For example, in order to train an initial driving scene generation model, a training sample set including driving scene images can be obtained.

[0165] S602, input the training sample set into the data processing module for data preprocessing to obtain the preprocessed sample set.

[0166] For example, when the training sample set is obtained, it can be input into the data processing module 401 for data preprocessing to obtain the preprocessed sample set.

[0167] S603, input the preprocessed sample set into the segmentation module to perform scene segmentation and obtain static background and dynamic foreground.

[0168] For example, when the preprocessed sample set is obtained, the preprocessed sample set can be input into the segmentation module 402 for scene segmentation to segment the static background and dynamic foreground in the preprocessed sample set, thereby obtaining the static background and dynamic foreground.

[0169] S604: Input the static background into the static background module to construct a static geometric skeleton.

[0170] For example, when a static background is obtained, it can be input into the static background module 403 to construct the static geometric skeleton corresponding to the static background.

[0171] S605 inputs the dynamic foreground into the dynamic foreground module to construct a 3D Gaussian model.

[0172] For example, when a dynamic foreground is obtained, it can be input into the dynamic foreground module 404 to construct a 3D Gaussian model corresponding to each dynamic target in the dynamic foreground.

[0173] It should be understood that S604 and S605 can be executed simultaneously or sequentially, and the embodiments of this application do not limit this.

[0174] S606 inputs the static geometric skeleton and 3D Gaussian model into the adjustment module for global appearance fine-tuning, and obtains the global appearance fine-tuning result.

[0175] For example, when the static geometric skeleton and 3D Gaussian model are obtained, the static geometric skeleton and 3D Gaussian model can be input into the adjustment module 405 so that the global appearance of the static geometric skeleton and 3D Gaussian model can be finely adjusted by the adjustment module 405 to obtain the global appearance fine-tuning result.

[0176] S607 inputs the global appearance fine-tuning results into the rendering module to render the driving scene.

[0177] For example, when the global appearance fine-tuning result is obtained, the global appearance fine-tuning result can be input to the rendering module 406 so that the rendering module 406 can render the global appearance fine-tuning result and render the driving scene.

[0178] Optionally, when obtaining the static geometric skeleton, it can be first input into the rendering module 406 for rendering to obtain the rendered static geometric skeleton. Similarly, when obtaining the 3D Gaussian model, it can also be first input into the rendering module 406 for rendering to obtain the rendered 3D Gaussian model. Then, the rendered static geometric skeleton and the rendered 3D Gaussian model are input into the adjustment module 405 for global appearance fine-tuning to obtain the corresponding global appearance fine-tuning results. These global appearance fine-tuning results are then input into the rendering module 406 for rendering to produce the driving scene.

[0179] S608 outputs rendered driving scenes.

[0180] For example, upon obtaining a rendered driving scene, the rendered driving scene can be output. Furthermore, the difference between this rendered driving scene and the training sample set can be used to adjust the loss function in the static background module 403, the dynamic foreground module 404, and the adjustment module 405 (i.e.,...). , , Each of these steps is optimized until the difference between the rendered driving scene and the training sample set is small, and the rendered driving scene is basically similar to or exactly the same as the training sample set.

[0181] It should be noted that, Figure 6 All steps are in Figure 3 and Figure 4 The corresponding embodiments are described in detail, and will not be repeated here.

[0182] Figure 7 This is a flowchart illustrating a method for generating a driving scene according to an embodiment of this application. This method can be executed by an electronic device.

[0183] For example, such as Figure 7 As shown, the method 700 includes the following implementation process: S710 acquires input data including images of the driving scene.

[0184] For example, when generating a driving scene, input data including images of the driving scene can be acquired. This input data can be obtained in real time by the vehicle or it can be obtained from historical data collected by the vehicle.

[0185] S720 inputs the input data into the trained driving scene generation model to generate the target driving scene corresponding to the input data.

[0186] Among them, the trained driving scenario generation model is through Figure 3 and Figure 6 The method shown is used for training.

[0187] For example, the acquired input data is fed into the trained driving scene generation model so that the trained driving scene generation model can generate the driving scene (which can be called the "target driving scene") corresponding to the input data.

[0188] S730, output the target driving scenario.

[0189] For example, when the trained driving scenario generation model generates a target driving scenario, it can output the target driving scenario so that it can be used for at least one of the following: simulation testing, driving risk assessment, user driving training, etc.

[0190] In such Figure 7 In the method 700 shown, since the trained driving scene generation model can avoid overfitting when generating driving scenes, when the input data including the driving scene image is input into the trained driving scene generation model, the trained driving scene generation model can avoid overfitting when generating the driving scene corresponding to the input data, thus ensuring the accuracy of the geometric shape of the driving scene and improving the visual quality of the generated driving scene.

[0191] In summary, when obtaining a training sample set including driving scene images, the initial driving scene generation model is not directly trained using this set. Instead, based on each training sample in the set, the initial driving scene generation model first obtains the minimum and intermediate scaling factors among multiple scaling factors corresponding to the Gaussian elements of each training sample. Then, a loss function is determined using these minimum and intermediate scaling factors, and the initial driving scene generation model is trained using this loss function. This allows the trained model to impose thickness constraints on the Gaussian elements in the driving scene during generation, effectively flattening them. This enables the trained model to fit smoother and more realistic surface geometry, avoiding blurry "fog" or "cloudy" Gaussian elements and preventing overfitting. Consequently, the trained model ensures the accuracy of the geometric shape of the generated driving scene, improving its visual quality. Furthermore, training the static background module and the dynamic foreground module separately allows the trained static background module to more accurately reproduce the geometry of the static background, thus avoiding overfitting when generating static backgrounds. Similarly, the trained dynamic foreground module can more accurately model and reproduce the motion posture of dynamic targets, thus avoiding overfitting when generating dynamic foregrounds. Finally, when using the trained driving scene generation model, the accuracy of the driving scene's geometry is ensured, thereby improving the visual quality of the generated driving scene.

[0192] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values ​​or scenarios exemplified. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of this application.

[0193] The above text combined Figures 1 to 6 The training method for the driving scene generation model provided in the embodiments of this application is described in detail, and Figure 7 The method for generating driving scenarios provided in the embodiments of this application is described in detail below; the following will be combined with Figure 8 and Figure 10 The apparatus embodiments of this application are described in detail below. It should be understood that the apparatus in the embodiments of this application can perform the various methods described in the foregoing embodiments of this application, that is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.

[0194] Figure 8 This is a schematic diagram of the structure of the training device for the driving scene generation model provided in this application embodiment.

[0195] For example, such as Figure 8 As shown, the device 800 includes: The acquisition module 810 is used to acquire a training sample set including driving scene images; The processing module 820 is used to obtain a first scaling factor and a second scaling factor based on each training sample in the training sample set through an initial driving scene generation model; wherein the first scaling factor and the second scaling factor are the scaling factors corresponding to the Gaussian elements of each training sample; the first scaling factor represents the minimum scaling factor among the multiple scaling factors corresponding to each Gaussian element, and the second scaling factor represents the scaling factor other than the maximum and minimum scaling factors among the multiple scaling factors; the directions corresponding to each of the multiple scaling factors are mutually perpendicular; a target loss function is determined based on the first scaling factor and the second scaling factor corresponding to each Gaussian element; the initial driving scene generation model is trained based on the target loss function to obtain a trained driving scene generation model; wherein the trained driving scene model applies thickness constraints to the Gaussian elements in the driving scene when generating the driving scene.

[0196] In one possible implementation, the processing module 820 is specifically used for: Based on the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian element, the constraint terms corresponding to each Gaussian element are determined. Based on the constraints corresponding to each Gaussian element, the target loss function is determined; wherein the loss value of the target loss function is less than or equal to the preset loss value.

[0197] In one possible implementation, the initial driving scene generation model includes a first generation module and a second generation module. The first generation module generates the static background in the driving scene, and the second generation module generates the dynamic foreground in the driving scene. The processing module 820 is specifically used for: The first generation module is trained based on the target loss function and the first loss function to obtain the trained first generation module; wherein, the first loss function includes at least one of the color loss function corresponding to the static background, the depth loss function, and the loss function of the sky background opacity in the static background; the trained first generation module applies thickness constraints to the Gaussian elements in the static background when generating the static background; The second generation module is trained based on the target loss function and the second loss function to obtain the trained second generation module; wherein, the second loss function includes at least one of the color loss function, depth loss function and dynamic foreground opacity loss function corresponding to the dynamic foreground; the trained second generation module applies thickness constraints to the Gaussian units in the dynamic foreground when generating the dynamic foreground; The first generation module after training is combined with the second generation module after training to obtain the trained driving scene generation model.

[0198] In one possible implementation, the processing module 820 is specifically used for: The target loss function is weighted with the first preset weight, and the first loss function is weighted with the second preset weight to obtain the first weighted loss function; The first generation module is trained by using the first weighted loss function to obtain the trained first generation module.

[0199] In one possible implementation, the processing module 820 is specifically used for: The target loss function is weighted with the third preset weight, and the second loss function is weighted with the fourth preset weight to obtain the second weighted loss function. The second generation module is trained by using the second weighted loss function to obtain the trained second generation module.

[0200] In one possible implementation, the initial driving scene generation model also includes an adjustment module, and the processing module 820 is specifically used for: The adjustment module is trained based on the target loss function and the third loss function to obtain the trained adjustment module. The third loss function includes at least one of the following: color loss function and depth loss function corresponding to the driving scene, loss function of sky background opacity in the driving scene, loss function of driving scene opacity, and semantic loss function in the driving scene. The trained adjustment module performs appearance adjustment on the driving scene when generating the driving scene. The trained first generation module is combined with the trained second generation module to obtain the trained driving scene generation model, including: The first generation module, the second generation module, and the adjustment module after training are combined to obtain the trained driving scene generation model.

[0201] In one possible implementation, the processing module 820 is specifically used for: The target loss function is weighted with the fifth preset weight, and the third loss function is weighted with the sixth preset weight to obtain the third weighted loss function. The adjustment module is trained using a third weighted loss function to obtain the trained adjustment module.

[0202] Figure 9 This is a schematic diagram of the structure of the driving scene generation device provided in the embodiments of this application.

[0203] For example, such as Figure 9 As shown, the device 900 includes: The acquisition module 910 is used to acquire input data including driving scene images; Processing module 920 is used to input input data into the trained driving scene generation model to generate the target driving scene corresponding to the input data; wherein, the trained driving scene generation model is... Figure 3 and Figure 6 The method shown is used to train and output the target driving scenario.

[0204] It should be noted that the aforementioned device 800 or device 900 is embodied in the form of a functional module. The term "module" here can be implemented in software and / or hardware, without specific limitations.

[0205] For example, a "module" can be a software program, hardware circuit, or a combination of both that implements the above functions. Hardware circuits may include application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or combined processors) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.

[0206] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0207] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.

[0208] For example, such as Figure 10 As shown, the electronic device 1000 includes a memory 1010 and a processor 1020. The memory 1010 stores executable program code 1011, and the processor 1020 is used to call and execute the executable program code 1011 to perform a training method for a driving scene generation model, or a method for generating a driving scene.

[0209] It should be noted that the electronic device can be an intelligent device with an autonomous driving scenario generation model 400, including but not limited to: personal computers, tablets, handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem. The electronic device may have different names in different networks, such as: user equipment, access electronic device, user unit, user station, mobile station, mobile station, remote station, remote electronic device, mobile device, user electronic device, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, electronic device in a 5G network or future evolved network, etc. The comparison of embodiments in this application is not limited to these terms.

[0210] This application can divide electronic devices into functional modules based on the above method examples. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and represents only one logical functional division; in actual implementation, there may be other division methods.

[0211] When functional modules are divided according to their respective functions, the electronic device may include: an acquisition module and a processing module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0212] The electronic device provided in this application is used to execute the above-described method for training a driving scene generation model, or a method for generating a driving scene, and thus can achieve the same effect as the above-described implementation method.

[0213] When using integrated units, the electronic device may include a processing module and a storage module. The processing module is used to control and manage the operation of the electronic device. The storage module is used to support the execution of relevant program code and data by the electronic device.

[0214] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage module may be a memory.

[0215] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs (Digital Video Discs), CD-ROMs (Compact Disc Read-Only Memory), microdrives, magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read Only Memory), DRAMs (Dynamic Random Access Memory), VRAMs (Video Random Access Memory), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0216] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a training method for a driving scene generation model or a method for generating a driving scene as described in the above embodiments.

[0217] In addition, the electronic device provided in the embodiments of this application may specifically be a chip, component or module. The electronic device may include a connected processor and a memory. The memory is used to store instructions. When the electronic device is running, the processor may call and execute the instructions to make the chip execute a training method for a driving scene generation model in the above embodiments, or a driving scene generation method.

[0218] The electronic devices, computer-readable storage media, computer program products or chips provided in this application are all used to perform the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0219] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0220] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0221] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A training method for a driving scene generation model, characterized in that, The method includes: Obtain a training sample set including images of driving scenes; Based on each training sample in the training sample set, a first scaling factor and a second scaling factor are obtained through an initial driving scenario generation model; wherein, the first scaling factor and the second scaling factor are the scaling factors corresponding to the Gaussian elements of each training sample; the first scaling factor represents the minimum scaling factor among the multiple scaling factors corresponding to each Gaussian element, and the second scaling factor represents the scaling factor other than the maximum scaling factor and the minimum scaling factor among the multiple scaling factors; the directions corresponding to each of the multiple scaling factors are perpendicular to each other; Based on the first scaling factor and the second scaling factor corresponding to each Gaussian element, the target loss function is determined; The initial driving scene generation model is trained based on the target loss function to obtain a trained driving scene generation model; wherein, the trained driving scene model applies thickness constraints to the Gaussian elements in the driving scene when generating the driving scene.

2. The method according to claim 1, characterized in that, The step of determining the target loss function based on the first scaling factor and the second scaling factor corresponding to each Gaussian element includes: Based on the ratio of the first scaling factor to the second scaling factor corresponding to each Gaussian element, the constraint terms corresponding to each Gaussian element are determined. Based on the constraints corresponding to each Gaussian element, the target loss function is determined; wherein the loss value of the target loss function is less than or equal to a preset loss value.

3. The method according to claim 1 or 2, characterized in that, The initial driving scene generation model includes a first generation module and a second generation module. The first generation module is used to generate a static background in the driving scene, and the second generation module is used to generate a dynamic foreground in the driving scene. The step of training the initial driving scene generation model based on the target loss function to obtain the trained driving scene generation model includes: Based on the target loss function and the first loss function, the first generation module is trained to obtain the trained first generation module; wherein, the first loss function includes at least one of the color loss function, the depth loss function, and the loss function of the sky background opacity in the static background corresponding to the static background; the trained first generation module performs thickness constraints on the Gaussian elements in the static background when generating the static background; Based on the target loss function and the second loss function, the second generation module is trained to obtain the trained second generation module; wherein, the second loss function includes at least one of the color loss function, the depth loss function, and the loss function of the dynamic foreground opacity corresponding to the dynamic foreground; the trained second generation module performs thickness constraints on the Gaussian units in the dynamic foreground when generating the dynamic foreground; The trained first generation module and the trained second generation module are combined to obtain the trained driving scene generation model.

4. The method according to claim 3, characterized in that, The step of training the first generation module based on the target loss function and the first loss function to obtain the trained first generation module includes: The target loss function is weighted with the first preset weight, and the first loss function is weighted with the second preset weight to obtain the first weighted loss function; The first generation module is trained using the first weighted loss function to obtain the trained first generation module.

5. The method according to claim 3, characterized in that, The step of training the second generation module based on the target loss function and the second loss function to obtain the trained second generation module includes: The target loss function is weighted with the third preset weight, and the second loss function is weighted with the fourth preset weight to obtain the second weighted loss function; The second generation module is trained by the second weighted loss function to obtain the trained second generation module.

6. The method according to claim 3, characterized in that, The initial driving scenario generation model also includes an adjustment module, and the method further includes: Based on the target loss function and the third loss function, the adjustment module is trained to obtain the trained adjustment module; wherein, the third loss function includes at least one of the color loss function and depth loss function corresponding to the driving scene, the loss function of sky background opacity in the driving scene, the loss function of driving scene opacity, and the semantic loss function in the driving scene; the trained adjustment module performs appearance adjustment on the driving scene when generating the driving scene; The step of combining the trained first generation module with the trained second generation module to obtain the trained driving scene generation model includes: The trained first generation module, the trained second generation module, and the trained adjustment module are combined to obtain the trained driving scene generation model.

7. The method according to claim 6, characterized in that, The adjustment module is trained based on the target loss function and the third loss function to obtain the trained adjustment module, including: The target loss function is weighted with the fifth preset weight, and the third loss function is weighted with the sixth preset weight to obtain the third weighted loss function; The adjustment module is trained using the third weighted loss function to obtain the trained adjustment module.

8. A training device for a driving scene generation model, characterized in that, The device includes: The acquisition module is used to acquire a training sample set including images of driving scenes; The processing module is configured to obtain a first scaling factor and a second scaling factor based on each training sample in the training sample set through an initial driving scene generation model; wherein the first scaling factor and the second scaling factor are scaling factors corresponding to the Gaussian elements of each training sample; the first scaling factor represents the minimum scaling factor among the multiple scaling factors corresponding to each Gaussian element, and the second scaling factor represents the scaling factor other than the maximum scaling factor and the minimum scaling factor among the multiple scaling factors; the directions corresponding to each of the multiple scaling factors are mutually perpendicular; a target loss function is determined based on the first scaling factor and the second scaling factor corresponding to each Gaussian element; the initial driving scene generation model is trained based on the target loss function to obtain a trained driving scene generation model; wherein the trained driving scene model applies thickness constraints to the Gaussian elements in the driving scene when generating the driving scene.

9. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Plane prior guidance-based three-dimensional reconstruction method and system for scene in three-dimensional Gaussian splash chamber

    CN121982273A