Method and device for correcting vehicle automatic driving simulation data
By calculating the spatiotemporal consistency loss value and updating parameters through the meta-policy network model, simulation data that conforms to physical laws is generated, which solves the problems of data violation and synchronization in autonomous driving simulation and improves the accuracy and efficiency of simulation.
Patent Information
- Application Number
- CN202511545421.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-01-27
AI Technical Summary
In existing autonomous driving simulation technologies, simulation data violates physical laws under extreme scenarios, and data from multiple sensors are not synchronized, resulting in low simulation reliability and efficiency. Furthermore, long-term data errors accumulate, making it difficult to reproduce in the real physical world and resulting in insufficient accuracy.
The expert correction data is determined by the meta-policy network model, the spatiotemporal consistency loss value is calculated, and the model is updated using the target reward value to generate simulation data that conforms to physical laws and is spatiotemporally synchronized, thereby realizing the self-correction of simulation data.
It improves the accuracy, stability and efficiency of autonomous driving simulation, ensures the systematic self-correction of simulation data in the forward generation process, and enhances the characterization capability of multimodal sensor data.
Smart Images

Figure CN121413232A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle autonomous driving simulation technology, and in particular to a method and apparatus for correcting vehicle autonomous driving simulation data. Background Technology
[0002] To ensure that autonomous driving systems can cope with the safety challenges of complex real-world road environments, autonomous driving tests are conducted through simulation. In autonomous driving simulation, model tools such as world models and reinforcement learning correctors are generally used. Based on external real data, the autonomous driving simulation data of the vehicle is corrected by a discriminator and a corresponding processor.
[0003] In some extreme scenarios, when using simulation data generated by the above methods, phenomena that violate physical laws may occur due to the lack of constraints on dynamic characteristics, reducing the reliability and effectiveness of autonomous driving simulation. In addition, the data generated by multiple sensors in the simulation are not synchronized in time and space, which may lead to the failure of autonomous driving simulation. Furthermore, the simulation process is lengthy and costly, reducing the efficiency and stability of autonomous driving simulation.
[0004] Furthermore, the above methods are prone to generating accumulated biases. When generating long-term data from the world model, the error increases with the number of frames, leading to distortion of the trajectory of dynamic objects. Moreover, the long-tail scene generated by the world model has insufficient coverage, making it difficult to reproduce in the real physical world. This has significant limitations and reduces the accuracy of autonomous driving simulation data. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a method and apparatus for correcting vehicle autonomous driving simulation data. This method uses a meta-policy network model to determine the expert correction data corresponding to the original simulation scene parameters generated by the autonomous driving simulation test model, thereby determining the corrected vehicle simulation data. It calculates the spatiotemporal consistency loss value, compares this value with a preset loss threshold, and uses the calculated target reward value to update the parameters of the meta-policy network model. The result is the output of target vehicle simulation data that conforms to physical laws and is spatiotemporally synchronized, representing the multimodal sensor data of the vehicle's autonomous driving system. This achieves systematic self-correction of vehicle simulation data during the forward generation process, providing a data foundation for vehicle autonomous driving simulation and improving the accuracy, stability, and efficiency of autonomous driving simulation.
[0006] This application provides a method for correcting vehicle autonomous driving simulation data, the method comprising: Obtain the original vehicle simulation data generated by the preset autonomous driving simulation test model based on the original simulation scenario parameters; Based on the original vehicle simulation data, the expert correction data corresponding to the original simulation scenario parameters is determined using a preset meta-policy network model, and the autonomous driving simulation test model generates corrected vehicle simulation data based on the original simulation scenario parameters and the expert correction data. Based on the corrected vehicle simulation data, the preset real scene data, and the preset sensor parameters of the vehicle, the spatiotemporal consistency loss value is calculated, and it is determined whether the spatiotemporal consistency loss value is less than or equal to the preset loss threshold. If the spatiotemporal consistency loss value is greater than the preset loss threshold, then a target reward value is determined based on the corrected vehicle simulation data, the expert correction amount data and the spatiotemporal consistency loss value, and the parameters of the meta-policy network model are updated based on the target reward value to obtain the corrected vehicle simulation data generated by the meta-policy network model with updated parameters. If the spatiotemporal consistency loss value is less than or equal to the preset loss threshold, then the corrected vehicle simulation data is determined as the target vehicle simulation data.
[0007] Furthermore, the step of determining the expert correction data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model includes: The original vehicle simulation data is compressed and fused using a preset state encoder to obtain the original state vector; The original state vector is input into the meta-policy layer and execution layer of the preset meta-policy network model respectively to obtain expert policy information output by the meta-policy layer based on the original state vector; The expert policy information is input into the execution layer to obtain expert correction data corresponding to the original simulation scene parameters output by the meta-policy layer based on the original state vector and the expert policy information.
[0008] Furthermore, the expert strategy information is obtained through the following steps: The meta-strategy layer determines the scene gap feedback parameter between the original vehicle simulation data and the real scene data based on the original state vector and preset real scene data. Based on the scenario gap feedback parameters, at least one target expert strategy object is selected from the preset expert strategy library, and the expert strategy information corresponding to each target expert strategy object is output.
[0009] Furthermore, the expert correction data is obtained through the following steps: The execution layer determines the expert weight parameters corresponding to the target expert policy object based on the original state vector and the expert policy information. Based on the original state vector, the professional knowledge database set by the target expert strategy object, and the expert weight parameters, the expert correction data corresponding to the original simulation scene parameters is output.
[0010] Furthermore, the calculation of the spatiotemporal consistency loss value based on the corrected vehicle simulation data, preset real-world scene data, and preset vehicle sensor parameters includes: Based on the spatial projection data in the corrected vehicle simulation data, the real projection data in the real scene data, and the lidar parameters in the sensor parameters, calculate the spatial projection constraint value; Based on the optical flow prediction data in the corrected vehicle simulation data, the real projection data, and the motion compensation parameters in the sensor parameters, the spatiotemporal consistency constraint value is calculated. Based on the camera view data in the corrected vehicle simulation data, the real view data in the real scene data, and the camera parameters in the sensor parameters, the multi-view photometric consistency loss value is calculated. Based on the predicted mask data in the corrected vehicle simulation data and the real mask data in the real scene data, the effective mask loss value is calculated; The spatiotemporal consistency loss value is calculated based on the spatial projection constraint value, the spatiotemporal consistency constraint value, the multi-view photometric consistency loss value, and the effective mask loss value.
[0011] Furthermore, determining the target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value includes: Based on the corrected vehicle simulation data and the real scene data, the similarity value of the learned perception image patch is calculated, and the visual fidelity reward value corresponding to the similarity value of the learned perception image patch is determined. Based on preset physical constraint parameters and the modified vehicle simulation data, a physical compliance score is calculated using a preset physical compliance verification function, and the physical compliance reward value corresponding to the physical compliance score is determined. Based on the expert correction data, the correction magnitude value is determined; The target reward value is determined based on the spatiotemporal consistency loss value, the visual fidelity reward value, the physical compliance reward value, and the correction magnitude value.
[0012] Furthermore, after obtaining the corrected vehicle simulation data generated by the parameter-updated meta-policy network model, the correction method further includes: Based on the spatiotemporal consistency loss value corresponding to the corrected vehicle simulation data generated by the meta-policy network model after parameter updates, the parameters of the meta-policy network model are iteratively updated until the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0013] This application embodiment also provides a correction device for vehicle autonomous driving simulation data, the correction device comprising: The data acquisition module is used to acquire the original vehicle simulation data generated by the preset autonomous driving simulation test model based on the original simulation scene parameters; The expert correction module is used to determine the expert correction amount data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model, and the autonomous driving simulation test model generates corrected vehicle simulation data based on the original simulation scene parameters and the expert correction amount data. The spatiotemporal consistency judgment module is used to calculate the spatiotemporal consistency loss value based on the corrected vehicle simulation data, the preset real scene data and the preset sensor parameters of the vehicle, and to determine whether the spatiotemporal consistency loss value is less than or equal to the preset loss threshold. The model parameter update module is used to determine a target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value if the spatiotemporal consistency loss value is greater than the preset loss threshold, and to update the parameters of the meta-policy network model based on the target reward value, so as to obtain the corrected vehicle simulation data generated by the meta-policy network model with updated parameters. The data output module is used to determine the corrected vehicle simulation data as the target vehicle simulation data if the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0014] Furthermore, when the expert correction module is used to determine the expert correction amount data corresponding to the original simulation scene parameters based on the original vehicle simulation data and using a preset meta-policy network model, the expert correction module is used to: The original vehicle simulation data is compressed and fused using a preset state encoder to obtain the original state vector; The original state vector is input into the meta-policy layer and execution layer of the preset meta-policy network model respectively to obtain expert policy information output by the meta-policy layer based on the original state vector; The expert policy information is input into the execution layer to obtain expert correction data corresponding to the original simulation scene parameters output by the meta-policy layer based on the original state vector and the expert policy information.
[0015] Furthermore, when the expert correction module is used to obtain the expert strategy information, the expert correction module is used to: The meta-strategy layer determines the scene gap feedback parameter between the original vehicle simulation data and the real scene data based on the original state vector and preset real scene data. Based on the scenario gap feedback parameters, at least one target expert strategy object is selected from the preset expert strategy library, and the expert strategy information corresponding to each target expert strategy object is output.
[0016] Furthermore, when the expert correction module is used to obtain the expert correction amount data, the expert correction module is used to: The execution layer determines the expert weight parameters corresponding to the target expert policy object based on the original state vector and the expert policy information. Based on the original state vector, the professional knowledge database set by the target expert strategy object, and the expert weight parameters, the expert correction data corresponding to the original simulation scene parameters is output.
[0017] Furthermore, when the spatiotemporal consistency judgment module calculates the spatiotemporal consistency loss value based on the corrected vehicle simulation data, preset real-world scene data, and preset sensor parameters of the vehicle, the spatiotemporal consistency judgment module is used to: Based on the spatial projection data in the corrected vehicle simulation data, the real projection data in the real scene data, and the lidar parameters in the sensor parameters, calculate the spatial projection constraint value; Based on the optical flow prediction data in the corrected vehicle simulation data, the real projection data, and the motion compensation parameters in the sensor parameters, the spatiotemporal consistency constraint value is calculated. Based on the camera view data in the corrected vehicle simulation data, the real view data in the real scene data, and the camera parameters in the sensor parameters, the multi-view photometric consistency loss value is calculated. Based on the predicted mask data in the corrected vehicle simulation data and the real mask data in the real scene data, the effective mask loss value is calculated; The spatiotemporal consistency loss value is calculated based on the spatial projection constraint value, the spatiotemporal consistency constraint value, the multi-view photometric consistency loss value, and the effective mask loss value.
[0018] Furthermore, when determining the target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value, the model parameter update module is used to: Based on the corrected vehicle simulation data and the real scene data, the similarity value of the learned perception image patch is calculated, and the visual fidelity reward value corresponding to the similarity value of the learned perception image patch is determined. Based on preset physical constraint parameters and the modified vehicle simulation data, a physical compliance score is calculated using a preset physical compliance verification function, and the physical compliance reward value corresponding to the physical compliance score is determined. Based on the expert correction data, the correction magnitude value is determined; The target reward value is determined based on the spatiotemporal consistency loss value, the visual fidelity reward value, the physical compliance reward value, and the correction magnitude value.
[0019] Furthermore, the correction device also includes a model parameter optimization module, which is used for: Based on the spatiotemporal consistency loss value corresponding to the corrected vehicle simulation data generated by the meta-policy network model after parameter updates, the parameters of the meta-policy network model are iteratively updated until the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0020] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the vehicle autonomous driving simulation data correction method described above are performed.
[0021] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the vehicle autonomous driving simulation data correction method described above.
[0022] The present application provides a method and apparatus for correcting vehicle autonomous driving simulation data. The correction method includes: acquiring original vehicle simulation data generated by a preset autonomous driving simulation test model based on original simulation scene parameters; determining expert correction amount data corresponding to the original simulation scene parameters using a preset meta-policy network model based on the original vehicle simulation data, and generating corrected vehicle simulation data by the autonomous driving simulation test model based on the original simulation scene parameters and the expert correction amount data; calculating a spatiotemporal consistency loss value based on the corrected vehicle simulation data, preset real scene data, and preset sensor parameters of the vehicle, and determining whether the spatiotemporal consistency loss value is less than or equal to a preset loss threshold; if the spatiotemporal consistency loss value is greater than the preset loss threshold, determining a target reward value based on the corrected vehicle simulation data, the expert correction amount data, and the spatiotemporal consistency loss value, and updating the parameters of the meta-policy network model based on the target reward value to obtain the corrected vehicle simulation data generated by the parameter-updated meta-policy network model; if the spatiotemporal consistency loss value is less than or equal to the preset loss threshold, determining the corrected vehicle simulation data as the target vehicle simulation data.
[0023] Compared to existing technologies such as applied world models and reinforcement learning correctors, which correct vehicle autonomous driving simulation data based on external real data using discriminators and corresponding processors, this method uses a meta-policy network model to determine the expert correction data corresponding to the original simulation scene parameters generated by the autonomous driving simulation test model. This determines the corrected vehicle simulation data, calculates the spatiotemporal consistency loss value, compares it with a preset loss threshold, and uses the calculated target reward value to update the parameters of the meta-policy network model. The output is target vehicle simulation data that conforms to physical laws and is spatiotemporally synchronized, representing the multimodal sensor data of the vehicle's autonomous driving system. This achieves systematic self-correction of vehicle simulation data during the forward generation process, providing a data foundation for vehicle autonomous driving simulation and improving the accuracy, stability, and efficiency of autonomous driving simulation.
[0024] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is one of the flowcharts for a method of correcting vehicle autonomous driving simulation data provided in the embodiments of this application; Figure 2 A second flowchart illustrating a method for correcting vehicle autonomous driving simulation data provided in an embodiment of this application; Figure 3 This is one of the structural schematic diagrams of a vehicle autonomous driving simulation data correction device provided in the embodiments of this application; Figure 4 A second schematic diagram of a device for correcting vehicle autonomous driving simulation data provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.
[0028] Research has found that autonomous driving simulations typically employ modeling tools such as world models and reinforcement learning modifiers. Based on real-world data, these tools, along with discriminators and corresponding processors, correct the vehicle's autonomous driving simulation data. However, in some extreme scenarios, the simulation data generated using these methods may exhibit phenomena that violate physical laws due to unconstrained dynamic characteristics, thus reducing the reliability and effectiveness of autonomous driving simulations. Furthermore, the data generated by multiple sensors in the simulation are not synchronized in time and space, posing a risk of simulation failure. Additionally, the simulation process is lengthy and costly, further reducing the efficiency and stability of autonomous driving simulations.
[0029] Furthermore, the above methods are prone to generating accumulated biases. When generating long-term data from the world model, the error increases with the number of frames, leading to distortion of the trajectory of dynamic objects. Moreover, the long-tail scene generated by the world model has insufficient coverage, making it difficult to reproduce in the real physical world. This has significant limitations and reduces the accuracy of autonomous driving simulation data.
[0030] Among them, reinforcement learning correctors correct data based on a correction strategy that uses predefined thresholds (e.g., light intensity thresholds). This is a static rule correction that cannot adapt to dynamic scene changes and has a high misjudgment rate. The spatiotemporal decoupling planning of reinforcement learning correctors separates and optimizes the path and velocity curves, which sacrifices the optimality of the trajectory and increases the planning failure rate in complex scenarios.
[0031] Furthermore, the data generation model and the corrector operate independently, which leads to data latency and reduced system frame rate. In addition, the lack of prior physical knowledge from the world model results in rather blind data correction actions.
[0032] Based on this, this application provides a method for correcting vehicle autonomous driving simulation data. It uses a meta-policy network model to determine the expert correction amount data corresponding to the original simulation scene parameters generated by the autonomous driving simulation test model, thereby determining the corrected vehicle simulation data. The method calculates the spatiotemporal consistency loss value, compares it with a preset loss threshold, and uses the calculated target reward value to update the parameters of the meta-policy network model. This outputs target vehicle simulation data that conforms to physical laws and is spatiotemporally synchronized, representing the multimodal sensor data of the vehicle's autonomous driving system. This achieves systematic self-correction of vehicle simulation data during the forward generation process, providing a data foundation for vehicle autonomous driving simulation and improving the accuracy, stability, and efficiency of autonomous driving simulation.
[0033] Please see Figure 1 , Figure 1 This is one of the flowcharts for a method of correcting vehicle autonomous driving simulation data provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the method for correcting vehicle autonomous driving simulation data includes: S101. Obtain the original vehicle simulation data generated by the preset autonomous driving simulation test model based on the original simulation scene parameters.
[0034] In the embodiments of this application, the autonomous driving simulation test model may include a network model based on a world model architecture, such as a generative adversarial network (GAN) model and a neural radiation field (NeRF) model.
[0035] Among them, the world model is a type of generative or predictive model designed to learn the dynamics, structure, and observation generation mechanisms of the environment, enabling intelligent agents to "simulate" the world internally, thereby supporting planning, reasoning, and decision-making world models; the autonomous driving simulation test model includes a perception encoder, a dynamic model, and a generative encoder.
[0036] In this embodiment, the original simulation scene parameters include, but are not limited to, light transmission parameters, dynamic object trajectory parameters, and sensor noise parameters; the original vehicle simulation data includes, but is not limited to, original spatial projection data, original optical flow prediction data, original camera viewpoint data, original prediction mask data, original scene video frame data, and original LiDAR point cloud data.
[0037] Here, the original vehicle simulation data represents the multimodal sensor data generated in real time by the autonomous driving simulation test model based on the original simulation scene parameters when the vehicle's intelligent driving algorithm performs simulation testing.
[0038] S102. Based on the original vehicle simulation data, the expert correction data corresponding to the original simulation scenario parameters is determined using a preset meta-policy network model, and the autonomous driving simulation test model generates corrected vehicle simulation data based on the original simulation scenario parameters and the expert correction data.
[0039] In the embodiments of this application, the core advantage of the meta-policy network model lies in its meta-learning capability. By training it on a large number of different meta-tasks in advance, and configuring different weather, road conditions, physical parameters, etc. for each meta-task, the meta-policy network model learns quickly to obtain a good prior policy, thereby improving the generalization ability of the meta-policy network model.
[0040] Here, the objective function for training the meta-policy network model is shown below.
[0041] .
[0042] in, Represents the model parameters of the meta-policy network model; Indicates the first Meta-tasks (e.g., a specific combination of weather and road conditions); This represents the basic learner (execution layer). Indicates in the task Performance parameters evaluated on the test data; Indicates in the task The loss function on.
[0043] Here, the objective function for training the meta-policy network model is used to optimize the model parameters of the meta-policy network model, so that for each task... Basic learning device Having access to a small amount of test data for this task In any case, they can perform well, that is, they can minimize the losses.
[0044] In this step, in specific implementation, firstly, based on the original vehicle simulation data, the expert correction data corresponding to the original simulation scene parameters is determined using a preset meta-policy network model; then, based on the original simulation scene parameters and the expert correction data, the corrected simulation scene parameters are determined; finally, the autonomous driving simulation test model generates corrected vehicle simulation data based on the corrected simulation scene parameters and preset internal rendering parameters.
[0045] The corrected vehicle simulation data includes, but is not limited to, spatial projection data, optical flow prediction data, camera viewpoint data, prediction mask data, scene video frame data, and lidar point cloud data.
[0046] Furthermore, the autonomous driving simulation test model also outputs value estimation results, which are used to evaluate the current state of the meta-policy network model to determine the quality of the currently generated data, thereby guiding the update direction of the meta-policy network model.
[0047] In one possible implementation of this application, in specific implementation, step S102, which involves determining the expert correction data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model, may include: S1021. The original vehicle simulation data is compressed and fused using a preset state encoder to obtain the original state vector.
[0048] In this step, a preset state encoder is used to fuse and compress all heterogeneous original vehicle simulation data into a unified and low-dimensional original state vector.
[0049] For example, for image data, the state encoder may include a CNN encoder; for point cloud data, the state encoder may include a PointNet encoder or a Voxel encoder; for scalar data, the state encoder performs direct stitching or processes through a fully connected layer.
[0050] Here, this step involves transforming the raw vehicle simulation data into a feature representation that the meta-policy network model can "understand".
[0051] S1022. The original state vector is input into the meta-policy layer and execution layer of the preset meta-policy network model respectively to obtain expert policy information output by the meta-policy layer based on the original state vector.
[0052] In the embodiments of this application, the preset meta-policy network model may include a physically constrained meta-reinforcement learning corrector (PhysMeta-RL), which is an advanced agent training framework that combines physical prior knowledge, meta-learning, and reinforcement learning (RL). It aims to improve the generalization ability, sample efficiency, and security of agent network models in complex, dynamic, and physically constrained environments.
[0053] Specifically, physical constraints embed physical laws (e.g., Newtonian mechanics, conservation of energy, conservation of momentum, contact mechanics, etc.) as hard or soft constraints into the reinforcement learning framework; the goal of meta-reinforcement learning (Meta-RL) is to enable agent network models to adapt quickly when facing new tasks or new environments; and the corrector refers to the real-time correction of the original policy output during policy execution to ensure that physical constraints or safety boundaries are met.
[0054] In this embodiment of the application, the policy layer (Meta-Policy) acts as the "commander" of the meta-policy network model. Based on the current scene gap feedback parameters and physical constraint type, it selects a suitable target expert policy object from the preset expert policy library. For example, in rainy or foggy weather, it will select an expert who is proficient in "scattering physics", while in a congested scene, it may select an expert who is proficient in "trajectory prediction".
[0055] The Execution-Policy meta-policy network model acts as a "soldier," receiving expert policy information selected by the policy layer and outputting expert correction data representing fine-tuning actions based on the original state vector. This allows for rapid adaptation to new scenarios and dynamic adjustment of the generation parameters of the autonomous driving simulation test model.
[0056] In one possible implementation of this application, step S1022 may include: S10221. The meta-strategy layer determines the scene gap feedback parameter between the original vehicle simulation data and the real scene data based on the original state vector and the preset real scene data.
[0057] In the embodiments of this application, the scene gap feedback parameter is an indicator used to quantify the gap between the original vehicle simulation data and the real scene data, such as visual fidelity score, quantification of the dispersion of the gap between reality and simulation, scene complexity index, and weather label, etc.
[0058] S10222. Based on the scenario gap feedback parameters, select at least one target expert strategy object from the preset expert strategy library, and output the expert strategy information corresponding to each target expert strategy object.
[0059] Here, the expert strategy information may include a probability distribution or decision signal.
[0060] For example, suppose the expert strategy information is “[0, 1, 0]”, which means that the second expert strategy object is selected as the target expert strategy object in the expert strategy library.
[0061] For example, when the scenario gap feedback parameter indicates a high-complexity scenario and extreme weather, the target expert strategy object can be determined as a physics expert; when the scenario gap feedback parameter indicates a low-complexity scenario and urban congestion, the target expert strategy object can be determined as a trajectory expert.
[0062] S1023. Input the expert strategy information into the execution layer to obtain expert correction data corresponding to the original simulation scene parameters output by the meta-strategy layer based on the original state vector and the expert strategy information.
[0063] In one possible implementation of this application, step S1023 may include: S10231. The execution layer determines the expert weight parameters corresponding to the target expert strategy object based on the original state vector and the expert strategy information.
[0064] In this step, based on the original state vector and expert policy information, the expert weight parameters corresponding to the target expert policy object are loaded in the execution layer.
[0065] S10232 outputs the expert correction data corresponding to the original simulation scene parameters based on the original state vector, the professional knowledge database set by the target expert strategy object, and the expert weight parameters.
[0066] In the embodiments of this application, the expert correction data may include a continuous and multi-dimensional correction action vector, each dimension of which corresponds to a specific adjustable parameter in the autonomous driving simulation.
[0067] For example, the expert correction data may include corrections for light transmission parameters, dynamic object trajectory, and sensor noise parameters.
[0068] In this embodiment of the application, the expert correction data is output through the following expression.
[0069] .
[0070] in, This indicates the amount of data corrected by experts; Represents the original state vector; Indicates the target expert strategy object; This represents the expert weight parameters corresponding to the target expert strategy object and the professional knowledge database set for the target expert strategy object. Indicates a corrective action; "Indicates sampling from it."
[0071] S103. Based on the corrected vehicle simulation data, the preset real scene data and the preset sensor parameters of the vehicle, calculate the spatiotemporal consistency loss value, and determine whether the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0072] In this embodiment, the preset loss threshold can be specifically calibrated according to the actual correction requirements, simulation test requirements, and the architectural parameters of the meta-policy network model.
[0073] Here, calculating the spatiotemporal consistency loss value means calculating the consistency loss of corrected vehicle simulation data from multiple dimensions, aiming to solve the problem of mismatch between different sensor data in the spatiotemporal dimensions in vehicle autonomous driving simulation.
[0074] In one possible implementation of this application, in specific implementation, the step of calculating the spatiotemporal consistency loss value based on the corrected vehicle simulation data, preset real scene data, and preset sensor parameters of the vehicle in step S103 may include: S1031. Calculate the spatial projection constraint value based on the spatial projection data in the corrected vehicle simulation data, the real projection data in the real scene data, and the lidar parameters in the sensor parameters.
[0075] In this embodiment of the application, the spatial projection constraint value is calculated using the following formula.
[0076] .
[0077] in, Indicates the spatial projection constraint value; This indicates the number of lidar point clouds in the lidar parameters; This represents the total number of semantic categories in the spatial projection data; Indicates the first One lidar point; Points in spatial projection data Two-dimensional coordinates projected onto the camera image plane; Points in the actual projection data True semantic label encoding (when point) Belongs to semantic category When the point is 1, the value is 1; when the point is 1... Not belonging to semantic category (When, the value is 0). Indicates at the projection point The location-predicted semantic category The probability of.
[0078] Here, the spatial projection constraint value may include a cross-entropy loss to ensure that the three-dimensional geometric world and the two-dimensional visual appearance are semantically aligned.
[0079] Specifically, when a point in a 3D point cloud is projected onto the corresponding location on a 2D image, the semantic category prediction of that image must be consistent with the label of the 3D point; for example, a 3D point labeled as "car" should also have its image region of the projection point semantically classified as "car".
[0080] S1032. Calculate the spatiotemporal consistency constraint value based on the optical flow prediction data in the corrected vehicle simulation data, the real projection data, and the motion compensation parameters in the sensor parameters.
[0081] In this embodiment of the application, the consistency constraint value is calculated using the following formula.
[0082] .
[0083] in, Indicates the consistency constraint value; This indicates the number of valid 3D point cloud pairs in the optical flow prediction data; Represents the first in the real projection data A three-dimensional point in The coordinates projected onto the image at any given moment; The first parameter in the motion compensation parameters A three-dimensional point in The coordinates projected onto the image at any given moment; This indicates that the optical flow prediction data is generated by the optical flow prediction network. Points on the image at time The prediction is pointing to Optical flow vector at time; This indicates the calculation of the L1 norm, which is the sum of absolute values.
[0084] Here, the consistency constraint value ensures that the appearance of the same object is consistent under different camera views, in accordance with the constraints of the physical lighting model; the consistency constraint value is also used to compare the difference between "two-dimensional displacement calculated from the motion of three-dimensional points" and "optical flow displacement calculated from the two-dimensional image itself", while ideally, the two should be equal.
[0085] S1033. Based on the camera view data in the corrected vehicle simulation data, the real view data in the real scene data, and the camera parameters in the sensor parameters, calculate the multi-view photometric consistency loss value.
[0086] In this embodiment of the application, the multi-view photometric consistency loss value is calculated using the following formula.
[0087] .
[0088] in, This represents the loss value of photometric consistency across multiple viewing angles; Represents the source view in real-view data This corresponds to a two-dimensional pixel coordinate on the image. Representing the source perspective The image at the pixel point The color value at that location; Represents the viewpoint in camera viewpoint data This corresponds to a two-dimensional pixel coordinate on the image. Represents the viewpoint in camera viewpoint data The image at the pixel point The color value at that location; This represents the total number of elements in the camera view data, i.e., the total number of valid pixels (excluding points that are occluded, whose projections exceed the boundaries, etc.).
[0089] Here, the multi-view photometric consistency loss value is used to calculate the difference in appearance color value of the same 3D point under different viewpoints. The core purpose is to minimize the difference in appearance color of the same 3D point under two different camera viewpoints. The smaller the difference, the more consistent the lighting and texture rendering under different viewpoints, so as to ensure that the appearance of the same object is consistent under different camera viewpoints and conforms to the constraints of the physical lighting model.
[0090] S1034. Calculate the effective mask loss value based on the predicted mask data in the corrected vehicle simulation data and the real mask data in the real scene data.
[0091] In this embodiment of the application, the effective mask loss value is calculated using the following formula.
[0092] .
[0093] in, Indicates the effective mask loss value; Represents the actual mask data; This represents the predicted mask data; This indicates the calculation of binary cross-entropy loss.
[0094] Here, in the real mask data and the predicted mask data, a value of 1 indicates that the projection point is valid; a value of 0 indicates that the projection point is invalid (e.g., occluded or outside the boundary).
[0095] In this way, by calculating the effective mask loss value, it can be determined which regions have reliable multi-view or temporal constraints, thereby automatically ignoring invalid regions when calculating the consistency loss value, thus improving stability and effectiveness.
[0096] S1035. Calculate the spatiotemporal consistency loss value based on the spatial projection constraint value, the spatiotemporal consistency constraint value, the multi-view photometric consistency loss value, and the effective mask loss value.
[0097] In this embodiment of the application, the spatiotemporal consistency loss value is calculated using the following formula.
[0098] .
[0099] in, This represents the spatiotemporal consistency loss value; Indicates the spatial projection constraint value; Indicates the consistency constraint value; This represents the loss value of photometric consistency across multiple viewing angles; Indicates the effective mask loss value; , , , These represent the preset loss balance weight coefficients.
[0100] Here, the spatiotemporal consistency loss value transforms the hard constraints of the physical world (geometric consistency, motion consistency, appearance consistency) into an optimizable mathematical objective, ensuring the high fidelity of simulation test data from the source of data generation.
[0101] Specifically, the approach shifts from "passive inspection" to "active constraint," making consistency a hard objective for training and generating models, rather than an indicator for post-event verification, thereby improving data quality from the source. At the same time, constraints are imposed from three dimensions: space, time, and perspective, to ensure the authenticity of the simulation-generated data.
[0102] Furthermore, when training the autonomous driving simulation test model, the calculation of the spatiotemporal consistency loss is backpropagated together with the main loss of the autonomous driving simulation test model. In order to minimize the total loss, the autonomous driving simulation test model will be forced to adjust its parameters, thereby outputting data that is not only visually realistic, but also highly consistent in the spatiotemporal dimension.
[0103] S104. If the spatiotemporal consistency loss value is greater than the preset loss threshold, then based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value, a target reward value is determined, and based on the target reward value, the parameters of the meta-policy network model are updated to obtain the corrected vehicle simulation data generated by the parameter-updated meta-policy network model.
[0104] In this embodiment of the application, the step of calculating the target reward value will reflect the constraints of physical laws. The target reward value includes a physical compliance reward. If the vehicle simulation data is corrected in violation of physical laws, the physical compliance reward will give a corresponding penalty.
[0105] In this step, after the target reward value is calculated, it is fed back to the preset reinforcement learning optimizer. The reinforcement learning optimizer calculates the policy gradient and updates the network weights corresponding to the policy layer and execution layer in the meta-policy network model. This makes the meta-policy network model more likely to select the expert policy object that can obtain high rewards when encountering similar states, and outputs more accurate expert correction data.
[0106] In one possible implementation of this application, in specific implementation, the step of determining the target reward value based on the corrected vehicle simulation data, the expert correction amount data, and the spatiotemporal consistency loss value in step S104 may include: S1041. Based on the corrected vehicle simulation data and the real scene data, calculate the similarity value of the learned perception image blocks and determine the visual fidelity reward value corresponding to the similarity value of the learned perception image blocks.
[0107] In this embodiment of the application, the visual fidelity reward value is calculated using the following formula.
[0108] .
[0109] in, Indicates the visual fidelity bonus value; This indicates the correction of deep learning features in vehicle simulation data; Represents deep learning features in real-world scenario data; This represents the learning of image patch similarity values; This represents the weight coefficient corresponding to the similarity value of the image patches learned and perceived.
[0110] Here, the similarity value of the learned perceptual image patches is used to measure visual fidelity. The smaller the similarity value of the learned perceptual image patches, the higher the visual fidelity.
[0111] S1042. Based on the preset physical constraint parameters and the modified vehicle simulation data, calculate the physical compliance score using the preset physical compliance verification function, and determine the physical compliance reward value corresponding to the physical compliance score.
[0112] In this embodiment of the application, the physical compliance reward value is calculated using the following formula.
[0113] .
[0114] in, This represents the physical compliance reward value; This indicates a correction to the point cloud data in the vehicle simulation data; This indicates the physical compliance score. This represents the physical compliance verification function, which has preset physical constraint parameters; This represents the weighting coefficient corresponding to the physical compliance score.
[0115] Here, the physical compliance verification function is pre-set with physical constraint parameters, which represent the prior knowledge and rule parameters of the physical world.
[0116] In this embodiment, the physical compliance verification function can correspond to a pre-trained physical model, such as Mie scattering coefficient, material reflection parameter model, and vehicle kinematics model. The physical compliance verification function is used to output a score that measures the degree to which the point cloud data in the corrected vehicle simulation data conforms to physical laws, for example, calculating the KL divergence between the point cloud data in the corrected vehicle simulation data and the real point cloud data predicted according to physical laws.
[0117] S1043. Based on the expert correction data, determine the correction magnitude value.
[0118] In this step, the square norm is calculated for the expert correction data, and the calculated result is determined as the correction magnitude value.
[0119] Here, the adjustment magnitude is used to prevent the metapolicy network model from making overly aggressive and unstable adjustments in pursuit of rewards, thereby encouraging the metapolicy network model to make smooth parameter fine-tuning.
[0120] S1044. Determine the target reward value based on the spatiotemporal consistency loss value, the visual fidelity reward value, the physical compliance reward value, and the correction magnitude value.
[0121] In this embodiment of the application, the target reward value is calculated using the following formula.
[0122] .
[0123] in, Indicates the target reward value; Indicates the visual fidelity bonus value; This represents the physical compliance reward value; This indicates the amount of data corrected by experts; Indicates the correction magnitude value; This represents the weighting coefficient corresponding to the expert correction data; This represents the spatiotemporal consistency loss value.
[0124] Here, the larger the spatiotemporal consistency loss value, the smaller the target reward value will be.
[0125] S105. If the spatiotemporal consistency loss value is less than or equal to the preset loss threshold, then the corrected vehicle simulation data is determined as the target vehicle simulation data.
[0126] In this step, when the spatiotemporal consistency loss value is less than or equal to the preset loss threshold, it indicates that the corrected vehicle simulation data at this time meets the requirements of physical constraints and spatiotemporal consistency. The corrected vehicle simulation data is then identified as the target vehicle simulation data, which is used to conduct autonomous driving simulation tests.
[0127] Optional, please refer to Figure 2 , Figure 2 This is a second flowchart illustrating a method for correcting vehicle autonomous driving simulation data provided in an embodiment of this application. Figure 2 As shown in the figure, after obtaining the corrected vehicle simulation data generated by the parameter-updated meta-policy network model in step S104, the correction method for vehicle autonomous driving simulation data provided in this application further includes: S106. Based on the spatiotemporal consistency loss value corresponding to the corrected vehicle simulation data generated by the meta-policy network model after parameter update, the parameters of the meta-policy network model are iteratively updated until the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0128] In this step, in specific implementation, firstly, based on the corrected vehicle simulation data generated by the meta-policy network model with updated parameters, the corresponding spatiotemporal consistency loss value is calculated; then, the parameters of the meta-policy network model are iteratively updated through gradient backpropagation until the calculated spatiotemporal consistency loss value is less than or equal to a preset loss threshold; finally, the corrected vehicle simulation data obtained from the last parameter update is determined as the target vehicle simulation data.
[0129] The method for correcting vehicle autonomous driving simulation data provided in this application, compared with the method of correcting vehicle autonomous driving simulation data through a discriminator and a corresponding processor, determines the expert correction amount data corresponding to the original simulation scene parameters generated by the autonomous driving simulation test model through a meta-policy network model, and then determines the vehicle simulation data to be corrected, calculates the spatiotemporal consistency loss value, compares the spatiotemporal consistency loss value with a preset loss threshold, and uses the calculated target reward value to update the parameters of the meta-policy network model, outputting target vehicle simulation data that conforms to physical laws and is spatiotemporally synchronized, so as to characterize the multimodal sensor data of the vehicle's autonomous driving system, realizes the systematic self-correction of vehicle simulation data in the forward generation process, provides a data foundation for vehicle autonomous driving simulation, and improves the accuracy, stability and efficiency of autonomous driving simulation.
[0130] Please see Figure 3 , Figure 4 , Figure 3 This is one of the structural schematic diagrams of a vehicle autonomous driving simulation data correction device provided in the embodiments of this application. Figure 4 This is a second schematic diagram of a device for correcting vehicle autonomous driving simulation data provided in an embodiment of this application. Figure 3 As shown, the correction device 300 includes: The data acquisition module 310 is used to acquire the original vehicle simulation data generated by the preset autonomous driving simulation test model based on the original simulation scene parameters; The expert correction module 320 is used to determine the expert correction amount data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model, and the autonomous driving simulation test model generates corrected vehicle simulation data based on the original simulation scene parameters and the expert correction amount data. The spatiotemporal consistency judgment module 330 is used to calculate the spatiotemporal consistency loss value based on the corrected vehicle simulation data, the preset real scene data and the preset sensor parameters of the vehicle, and to determine whether the spatiotemporal consistency loss value is less than or equal to the preset loss threshold. The model parameter update module 340 is used to determine a target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value if the spatiotemporal consistency loss value is greater than the preset loss threshold, and to update the parameters of the meta-policy network model based on the target reward value, so as to obtain the corrected vehicle simulation data generated by the meta-policy network model after parameter update. The data output module 350 is used to determine the corrected vehicle simulation data as the target vehicle simulation data if the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0131] Furthermore, when the expert correction module 320 determines the expert correction amount data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model, the expert correction module 320 is used to: The original vehicle simulation data is compressed and fused using a preset state encoder to obtain the original state vector; The original state vector is input into the meta-policy layer and execution layer of the preset meta-policy network model respectively to obtain expert policy information output by the meta-policy layer based on the original state vector; The expert policy information is input into the execution layer to obtain expert correction data corresponding to the original simulation scene parameters output by the meta-policy layer based on the original state vector and the expert policy information.
[0132] Furthermore, when the expert correction module 320 obtains the expert strategy information, the expert correction module 320 is used to: The meta-strategy layer determines the scene gap feedback parameter between the original vehicle simulation data and the real scene data based on the original state vector and preset real scene data. Based on the scenario gap feedback parameters, at least one target expert strategy object is selected from the preset expert strategy library, and the expert strategy information corresponding to each target expert strategy object is output.
[0133] Furthermore, when the expert correction module 320 is used to obtain the expert correction amount data, the expert correction module 320 is used to: The execution layer determines the expert weight parameters corresponding to the target expert policy object based on the original state vector and the expert policy information. Based on the original state vector, the professional knowledge database set by the target expert strategy object, and the expert weight parameters, the expert correction data corresponding to the original simulation scene parameters is output.
[0134] Furthermore, when the spatiotemporal consistency judgment module 330 calculates the spatiotemporal consistency loss value based on the corrected vehicle simulation data, preset real-world scene data, and preset sensor parameters of the vehicle, the spatiotemporal consistency judgment module 330 is used to: Based on the spatial projection data in the corrected vehicle simulation data, the real projection data in the real scene data, and the lidar parameters in the sensor parameters, calculate the spatial projection constraint value; Based on the optical flow prediction data in the corrected vehicle simulation data, the real projection data, and the motion compensation parameters in the sensor parameters, the spatiotemporal consistency constraint value is calculated. Based on the camera view data in the corrected vehicle simulation data, the real view data in the real scene data, and the camera parameters in the sensor parameters, the multi-view photometric consistency loss value is calculated. Based on the predicted mask data in the corrected vehicle simulation data and the real mask data in the real scene data, the effective mask loss value is calculated; The spatiotemporal consistency loss value is calculated based on the spatial projection constraint value, the spatiotemporal consistency constraint value, the multi-view photometric consistency loss value, and the effective mask loss value.
[0135] Furthermore, when determining the target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value, the model parameter update module 340 is used to: Based on the corrected vehicle simulation data and the real scene data, the similarity value of the learned perception image patch is calculated, and the visual fidelity reward value corresponding to the similarity value of the learned perception image patch is determined. Based on preset physical constraint parameters and the modified vehicle simulation data, a physical compliance score is calculated using a preset physical compliance verification function, and the physical compliance reward value corresponding to the physical compliance score is determined. Based on the expert correction data, the correction magnitude value is determined; The target reward value is determined based on the spatiotemporal consistency loss value, the visual fidelity reward value, the physical compliance reward value, and the correction magnitude value.
[0136] Furthermore, such as Figure 4 As shown, the correction device 300 further includes a model parameter optimization module 360, which is used for: Based on the spatiotemporal consistency loss value corresponding to the corrected vehicle simulation data generated by the meta-policy network model after parameter updates, the parameters of the meta-policy network model are iteratively updated until the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
[0137] The vehicle autonomous driving simulation data correction device provided in this application, compared with the method of correcting vehicle autonomous driving simulation data through a discriminator and a corresponding processor, determines the expert correction amount data corresponding to the original simulation scene parameters generated by the autonomous driving simulation test model through a meta-policy network model, and then determines the vehicle simulation data to be corrected, calculates the spatiotemporal consistency loss value, compares the spatiotemporal consistency loss value with a preset loss threshold, and uses the calculated target reward value to update the parameters of the meta-policy network model, outputting target vehicle simulation data that conforms to physical laws and is spatiotemporally synchronized, so as to characterize the multimodal sensor data of the vehicle's autonomous driving system, realizes the systematic self-correction of vehicle simulation data in the forward generation process, provides a data foundation for vehicle autonomous driving simulation, and improves the accuracy, stability and efficiency of autonomous driving simulation.
[0138] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.
[0139] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 is running, the processor 510 and the memory 520 communicate via the bus 530. When the machine-readable instructions are executed by the processor 510, they can perform the operations described above. Figure 1 as well as Figure 2 The steps of the method for correcting vehicle autonomous driving simulation data in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0140] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 as well as Figure 2 The steps of the method for correcting vehicle autonomous driving simulation data in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0141] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0144] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0145] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for correcting vehicle autonomous driving simulation data, characterized in that, The correction method includes: Obtain the original vehicle simulation data generated by the preset autonomous driving simulation test model based on the original simulation scenario parameters; Based on the original vehicle simulation data, the expert correction data corresponding to the original simulation scenario parameters is determined using a preset meta-policy network model, and the autonomous driving simulation test model generates corrected vehicle simulation data based on the original simulation scenario parameters and the expert correction data. Based on the corrected vehicle simulation data, the preset real scene data, and the preset sensor parameters of the vehicle, the spatiotemporal consistency loss value is calculated, and it is determined whether the spatiotemporal consistency loss value is less than or equal to the preset loss threshold. If the spatiotemporal consistency loss value is greater than the preset loss threshold, then a target reward value is determined based on the corrected vehicle simulation data, the expert correction amount data and the spatiotemporal consistency loss value, and the parameters of the meta-policy network model are updated based on the target reward value to obtain the corrected vehicle simulation data generated by the meta-policy network model with updated parameters. If the spatiotemporal consistency loss value is less than or equal to the preset loss threshold, then the corrected vehicle simulation data is determined as the target vehicle simulation data.
2. The method according to claim 1, characterized in that, The step of determining the expert correction data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model includes: The original vehicle simulation data is compressed and fused using a preset state encoder to obtain the original state vector. The original state vector is input into the meta-policy layer and execution layer of the preset meta-policy network model respectively to obtain expert policy information output by the meta-policy layer based on the original state vector; The expert policy information is input into the execution layer to obtain expert correction data corresponding to the original simulation scene parameters output by the meta-policy layer based on the original state vector and the expert policy information.
3. The method according to claim 2, characterized in that, The expert strategy information is obtained through the following steps: The meta-strategy layer determines the scene gap feedback parameter between the original vehicle simulation data and the real scene data based on the original state vector and preset real scene data. Based on the scenario gap feedback parameters, at least one target expert strategy object is selected from the preset expert strategy library, and the expert strategy information corresponding to each target expert strategy object is output.
4. The method according to claim 3, characterized in that, The expert correction data is obtained through the following steps: The execution layer determines the expert weight parameters corresponding to the target expert policy object based on the original state vector and the expert policy information. Based on the original state vector, the professional knowledge database set by the target expert strategy object, and the expert weight parameters, the expert correction data corresponding to the original simulation scene parameters is output.
5. The method according to claim 1, characterized in that, The calculation of the spatiotemporal consistency loss value based on the corrected vehicle simulation data, preset real-world scene data, and preset vehicle sensor parameters includes: Based on the spatial projection data in the corrected vehicle simulation data, the real projection data in the real scene data, and the lidar parameters in the sensor parameters, calculate the spatial projection constraint value; Based on the optical flow prediction data in the corrected vehicle simulation data, the real projection data, and the motion compensation parameters in the sensor parameters, the spatiotemporal consistency constraint value is calculated. Based on the camera view data in the corrected vehicle simulation data, the real view data in the real scene data, and the camera parameters in the sensor parameters, the multi-view photometric consistency loss value is calculated. Based on the predicted mask data in the corrected vehicle simulation data and the real mask data in the real scene data, the effective mask loss value is calculated; The spatiotemporal consistency loss value is calculated based on the spatial projection constraint value, the spatiotemporal consistency constraint value, the multi-view photometric consistency loss value, and the effective mask loss value.
6. The method according to claim 1, characterized in that, The determination of the target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value includes: Based on the corrected vehicle simulation data and the real scene data, the similarity value of the learned perception image patch is calculated, and the visual fidelity reward value corresponding to the similarity value of the learned perception image patch is determined. Based on preset physical constraint parameters and the modified vehicle simulation data, a physical compliance score is calculated using a preset physical compliance verification function, and the physical compliance reward value corresponding to the physical compliance score is determined. Based on the expert correction data, the correction magnitude value is determined; The target reward value is determined based on the spatiotemporal consistency loss value, the visual fidelity reward value, the physical compliance reward value, and the correction magnitude value.
7. The method according to claim 1, characterized in that, After obtaining the corrected vehicle simulation data generated by the parameter-updated meta-policy network model, the correction method further includes: Based on the spatiotemporal consistency loss value corresponding to the corrected vehicle simulation data generated by the meta-policy network model after parameter updates, the parameters of the meta-policy network model are iteratively updated until the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
8. A device for correcting vehicle autonomous driving simulation data, characterized in that, The correction device includes: The data acquisition module is used to acquire the original vehicle simulation data generated by the preset autonomous driving simulation test model based on the original simulation scene parameters; The expert correction module is used to determine the expert correction amount data corresponding to the original simulation scene parameters based on the original vehicle simulation data using a preset meta-policy network model, and the autonomous driving simulation test model generates corrected vehicle simulation data based on the original simulation scene parameters and the expert correction amount data. The spatiotemporal consistency judgment module is used to calculate the spatiotemporal consistency loss value based on the corrected vehicle simulation data, the preset real scene data and the preset sensor parameters of the vehicle, and to determine whether the spatiotemporal consistency loss value is less than or equal to the preset loss threshold. The model parameter update module is used to determine a target reward value based on the corrected vehicle simulation data, the expert correction data, and the spatiotemporal consistency loss value if the spatiotemporal consistency loss value is greater than the preset loss threshold, and to update the parameters of the meta-policy network model based on the target reward value, so as to obtain the corrected vehicle simulation data generated by the meta-policy network model with updated parameters. The data output module is used to determine the corrected vehicle simulation data as the target vehicle simulation data if the spatiotemporal consistency loss value is less than or equal to the preset loss threshold.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the method for correcting vehicle autonomous driving simulation data as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for correcting vehicle autonomous driving simulation data as described in any one of claims 1 to 7.