Off-road automatic driving training data synthesis method based on reverse calibration and 3DGS

CN122839818APending Publication Date: 2026-09-29INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610996412.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0012]具体来说,本发明的目的是解决上述现有技术中存在的视觉合成与物理仿真解耦、本构反演通道狭窄、下游模型类型支持单一层次问题,提出一种基于物理仿真器逆向标定与三维高斯神经形变场的越野自动驾驶模型训练数据合成方法及系统,旨在解决以下具体技术问题:

Benefits of technology

[0066]本发明建立视觉合成与物理仿真之间的耦合映射关系,使合成图像中的车辙、压陷、植被倒伏等交互痕迹与真实物理过程的位移场对应;同一套前端计算(S1–S4)通过S5多模式适配器同时支持感知模型、世界模型、VLA、VLM 四类下游模型的训练数据合成且跨模式真值物理一致,无需为不同模型类型重复搭建流水线;填补现有NeRF / 3DGS方案无物理交互、游戏引擎方案无真实流变、潜在扩散方案物理因果断裂的多重短板,在越野场景下实现世界模型10 s长时多步推演的物理一致演化能力,并基于物理真值直接驱动VLA / VLM训练数据生成使其语义精度高于基于图像本身的LLM自动图像描述方案;支持沙、泥、雪、碎石、冰、植被六类典型越野材质的统一逆向标定与合成,相比实车采集,减少对实车采集和人工标注的依赖;通过H3 hex 空间索引 + 差分隐私的车-云联邦逆向标定使多车日志在云端联合逆向标定后,有利于新平台利用已有标定结果进行初始化,从而缩短标定收敛时间;通过主动困难样本挖掘与课程学习闭环基于实车失败日志反向生成对抗性样本,比静态数据集更针对性,且对感知/世界模型/VLA / VLM 四类下游模型同时生效,从产业部署视角支撑越野自动驾驶各类下游模型的持续闭环补强。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839818A_ABST
    Figure CN122839818A_ABST
Patent Text Reader

Abstract

The present application provides an off-road automatic driving training data synthesis method based on reverse calibration and 3DGS, comprising: extracting a triple consisting of sensor flow, control instruction flow and wheel-ground dynamics response flow from the off-road driving log of the vehicle; taking the wheel-ground dynamics response in the triple as the optimization target, reverse-calibrating the mechanical constitutive parameters of the virtual terrain material in the physical simulator, and mapping the wheel-ground contact displacement field differential output by the physical simulator to the Gaussian point cloud, thereby synthesizing the synthesized perception data after the tire is pressed on the ground; inputting the virtual vehicle model matched with the real vehicle, the posterior estimation, the virtual terrain, and the control instruction / expert action sequence into the physical simulator to obtain the physical true value including the settlement depth, the shear stress, the semantic mask, the passability, and the pixel-by-pixel alignment of the posterior estimation; taking each pixel in the synthesized perception data and the posterior estimation and the physical true value corresponding thereto as the off-road automatic driving training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous driving technology and simulation data synthesis technology, and particularly to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS. Background Technology

[0002] There are fundamental differences between the wilderness environment and urban paved roads in terms of terrain features, surface material properties, and vehicle-ground dynamics coupling methods. The high-quality data required for training downstream models of off-road autonomous driving (perception, world model, vision-language-action model VLA, vision-language model VLM, etc.) is still orders of magnitude lower than that for paved road scenarios. Perception datasets for paved road scenarios (such as KITTI, nuScenes, Waymo, Argoverse, etc.) have reached a total size of PB. While several publicly labeled datasets exist for off-road scenarios, key examples include RELLIS-3D (13,556 LiDAR scans + 6,235 RGB images, 20 classes of pixel-level semantic segmentation), RUGD (7,453 labeled frames, 8 types of terrain pixel-level segmentation), TAS500 (540 scene-level fine-grained vegetation and terrain segmentation), ORFD (12,198 frames of LiDAR + RGB pairing, 4 scenes × 4 weather × 4 lighting conditions for three types of drivability regions), GOOSE (10,000 frames of multimodal instances and point cloud segmentation), GOOSE-Ex (5,000 frames of cross-platform extension), ROUGH (the first off-road dataset containing dynamics), and ORAD-3D (57,808 frames of LiDAR + RGB). Public datasets such as pairing and annotation of 2D semantic segmentation, 3D occupancy, driving trajectory, and scene text description are available. Although these datasets have made significant progress in the past few years, they still have the following three limitations compared to the actual training needs of off-road autonomous driving models.

[0003] The first limitation is the narrow task coverage, lacking pairings between vehicle-terrain physical interaction traces and physical ground truths. The aforementioned public datasets are mainly geared towards passive perception tasks, and the recorded objects are limited to raw sensor data (RGB / LiDAR / GNSS / IMU) and manually labeled discrete categories. They almost entirely lack paired samples of vehicle-terrain physical interaction traces (rut deformation fields, indentation and subsidence depth, vegetation lodging deformation, and reflectivity gradients caused by humidity changes) and corresponding continuous physical ground truths (Bekker / Janosi constitutive parameter posteriors, element shear stress distribution, and compaction rolling resistance maps). However, these physical interaction traces constitute the prior signals for off-road perception and decision-making models to understand terrain physical characteristics. The second limitation is the sparse Cartesian product coverage, with severe fragmentation across geographical, seasonal, and platform dimensions. Single datasets are typically collected only from two or three fixed geographical locations. The Cartesian product dimension combinations across seasons (winter snow / spring thaw / summer wet / autumn dry), materials (sand / mud / snow / gravel / ice / meadow / moss), platforms (wheeled / tracked), and lighting conditions (strong sunlight / rain / twilight / night / snow blindness) are far from being covered by any single dataset. Directly merging multiple datasets leads to serious inconsistencies in the labeled ontology. The third limitation is that high-risk conditions such as getting stuck, deep slumping, skidding and instability, and rollover precursors represent the most critical decision boundaries for off-road autonomous driving. However, collecting data on these conditions in real vehicles would cause vehicle damage and endanger personnel, making extreme condition samples in real-world data extremely scarce—a hard constraint that is difficult to eliminate from both physical and ethical perspectives in real-world data collection. The common root of these three limitations lies in the fact that existing real-world datasets only record what sensors observe, failing to record the physical interaction between the vehicle and the terrain, and what physical traces this interaction produces. To fill this gap, this invention introduces a coupling mechanism between a physics simulator and a differentiable neural rendering in the data generation loop.

[0004] Furthermore, compared to the availability of publicly available datasets such as RELLIS-3D and ORAD-3D for perception models, physical data required for training end-to-end models and world models in off-road scenarios is extremely scarce. As of May 2026, no publicly available datasets provide the following two types of samples: (a) (image_seq, language_instruction, action_seq) triples in off-road scenarios (required for VLA training); (b) physically consistent data in off-road scenarios. Transfer trajectory (required for world model training, and must be guaranteed) Includes The resulting changes in the physical state of the terrain, including rut deformation Changes in soil compaction Lodging of vegetation The lack of physically consistent training data has led to a mismatch between existing mainstream world models (Wayve GAIA-1, NVIDIA Cosmos, etc.) and general VLAs (RT-2, OpenVLA). In off-road scenarios, all of them exhibited terrain evolution distortion and command-action semantic mismatch.

[0005] To address the shortcomings of the aforementioned datasets, various synthetic data augmentation schemes have been proposed in the industry. However, after comparison, the inventors believe that the existing four types of synthetic data schemes have the following problems.

[0006] The first category is procedural generation solutions based on game engines (Unreal, CARLA, AirSim, IsaacSim, etc.). These solutions utilize the game engine's physics and rendering pipeline to generate a large number of virtual scenes, but their tire-soil contact models are mostly simplified Pacejka / magic formulas or rigid body contact, lacking realistic soft terrain rheological properties (no Bekker pressure-sinking, no Janosi shear yielding), resulting in a performance degradation of downstream models trained on synthetic data on real soft terrain; at the same time, their texture and lighting rendering pipeline is decoupled from the physical response of real sensors, resulting in significant cross-domain bias.

[0007] The second category is image synthesis schemes based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS). Schemes such as NeRF-W, Instant-NGP, and SplatAD utilize real-world multi-view images to reconstruct scenes and synthesize new perspective images. These schemes are geared towards geometric / appearance reconstruction or general dynamic deformation, but do not use vehicle-terrain contact mechanics simulation outputs as deformation drivers and physical truth sources. Therefore, they cannot synthesize new appearances resulting from physical interactions such as rutted tracks, fallen vegetation, or snowdrifts. While US20240355047A1 (Three-dimensional Gaussian splatting initialization based on trained neuralradiance field representations) and US20250363723A1 (Method and apparatus for dynamic Gaussian splatting) involve 3DGS training and dynamic deformation, the former focuses on initialization methods, while the latter focuses on general 4D dynamic scenes; neither involves a closed-loop physical simulator for the terrain / vehicle domain.

[0008] The third category is robust training schemes based on domain randomization. This type of scheme randomly perturbs textures, lighting, and noise in the simulation environment to reduce the sim-to-real gap. However, the randomized objects have no corresponding relationship with real physics. It mainly improves robustness by expanding the visual distribution coverage rather than explicitly modeling the vehicle-terrain physical interaction. It is difficult to solve the problem of insufficient physical-visual consistency on its own, and it cannot specifically compensate for the lack of extreme working condition samples.

[0009] The fourth category is video-generated world model schemes based on latent diffusion techniques. Representative works include Wayve GAIA-1 / GAIA-2 / GAIA-3, NVIDIA Cosmos / Cosmos-Drive-Dreams / Cosmos Predict-2, and Wayve LINGO-1. These schemes use latent diffusion to generate seemingly realistic driving scenes and consistent multi-camera synthetic sequences in video space, which can be used to train end-to-end models and world models. However, their fundamental drawback is that visual likelihood ≠ physical causality, meaning that the generated... and , The causal relationships between these elements are entirely determined by the neural network's fit to the training distribution, rather than by physical laws. This leads to: (a) a lack of physical consistency in terrain evolution (the direction and magnitude of ruts, subsidence, and soil compaction may violate the Bekker / Janosi constitutive law); (b) potential generation quality degradation in rare extreme cases (the semantic rationality of the video space cannot be extrapolated to physical processes outside the training distribution); and (c) no interpretable, auditable physical ground truth aligned with the generated pixels. Therefore, while this type of solution performs exceptionally well on paved roads, it becomes unreliable in off-road scenarios due to broken physical causal chains.

[0010] After analyzing the mechanisms of the aforementioned existing technologies (including game engines, NeRF / 3DGS new perspective synthesis, domain randomization, and latent diffusion video world models), the inventors believe that existing synthetic data schemes have the following three levels of problems: First, visual synthesis schemes and physical simulation schemes lack explicit coupling for vehicle-terrain interaction in mathematics. Visual synthesis schemes usually aim at geometric reconstruction and appearance rendering, while physical simulation schemes usually aim at dynamic response and contact deformation. There is no differentiable coupling bridge between the two, which means that even if the physical simulation is accurate enough (such as the multibody dynamics of commercial ADAMS / Recurdyn), the deformation information output by it cannot be converted into visual training samples; even if the 3DGS reconstruction is realistic enough, its synthesized images do not contain physical traces of vehicle-terrain interaction. Secondly, the inversion of constitutive parameters relies on narrow channels. Existing 3DGS physical property estimation schemes (PhysGS / GaussianProperty / NeRF2Physics, etc.) are based on visual perception inversion (image → physical property). In off-road scenarios, visual information is severely interfered with by factors such as lighting, vegetation occlusion, and strong reflections, resulting in insufficient inversion confidence to drive safety decisions. While physical sensing inversion based on wheel-ground dynamic response in real vehicle logs is orthogonal to the visual path, it maintains recognizability in visually degraded scenarios but has not yet been integrated with 3DGS deformation mapping. Thirdly, existing schemes are geared towards specific tasks or specific output formats, making it difficult to simultaneously serve multiple downstream models under the same physical truth system. There is a lack of a unified framework that simultaneously supports the synthesis of training data for four types of downstream models: perception models, world models, VLA, and VLM, and it cannot guarantee the consistency of physical truth across model types. Summary of the Invention

[0011] To address the aforementioned issues, this invention proposes a method, system, device, and storage medium for synthesizing training data for off-road autonomous driving models based on inverse calibration of a physical simulator and a 3D Gaussian neural deformation field. Through inverse calibration of the physical simulator, the wheel-ground dynamic response from the real vehicle log is transformed into the constitutive parameter posterior of the virtual terrain (the information source is orthogonal to the visual field). Then, the displacement deformation field output by the simulator drives the 3DGS Gaussian point cloud to generate geometric and appearance-co-deformation, producing synthetic perception data with physical interaction traces. Furthermore, through a multi-mode output adapter for downstream model types (sharing front-end S1–S4 computation), the synthetic data is adapted to any one or a combination of four parallel output modes: perception model, world model, VLA, and VLM, in conjunction with physical compliance loss. The invention explicitly constrains the downstream model training objective to be consistent with the physical truth; it also supports multi-vehicle cloud-based federated reverse calibration (H3 hex spatial indexing + differential privacy), and proactive hard sample mining based on real vehicle failure logs applies to all four downstream modes simultaneously. This invention fundamentally repairs the causal break between visual synthesis and physical simulation, compensating for the deficiencies in world models and end-to-end models in existing open off-road datasets.

[0012] Specifically, the purpose of this invention is to address the problems of decoupling between visual synthesis and physical simulation, narrow constitutive inversion channels, and single-level support for downstream model types in the prior art. It proposes a method and system for synthesizing training data for off-road autonomous driving models based on inverse calibration of a physical simulator and a three-dimensional Gaussian neural deformation field, aiming to solve the following specific technical problems:

[0013] 1. Achieve differentiable coupling between physical simulation and neural rendering, and establish the displacement deformation field output from the simulator. Differentiable mapping to the 3DGS Gaussian ellipsoid parameters enables joint optimization of visual synthesis and physical simulation in end-to-end gradient backpropagation;

[0014] 2. Supports parallel reverse calibration of multiple materials, using a set of parallel constitutive equations (Bekker-Wong / Janosi-Hanamoto / Magic Formula / Pacejka / LuGre) to support a unified reverse calibration process for six typical off-road materials: sand, mud, snow, gravel, ice, and vegetation.

[0015] 3. Provides multiple parameter identification algorithm options, offering three parallel variants—conjugate gradient descent, Bayesian optimization, and co-evolutionary algorithm (CMA-ES)—to address different real-vehicle log scales, simulator differentiability, and objective function noise / modal number, covering different data conditions from small samples (approximately hundreds of segments) to large samples (approximately tens of thousands of segments), from differentiable to black-box, and from single-modal to multimodal (see claim 5 for specific selection criteria).

[0016] 4. Supports the synthesis of training data for multiple types of downstream models for off-road autonomous driving. The synthesized data can be used to train downstream vehicle decision-making models in off-road scenarios, including but not limited to (i) perceptual neural networks (semantic segmentation, accessibility prediction, depth estimation, physical parameter estimation, semantic occupancy grid prediction); (ii) visual-language model (VLM); (iii) visual-language-action model (VLA); (iv) world model.

[0017] 5. Supports vehicle-cloud federated reverse calibration, allowing real-world vehicle logs from multiple mobile platforms to be jointly reverse calibrated in the cloud and share constitutive parameter posteriors, which can be used for initialization or prior constraints of new platform calibration;

[0018] 6. Supports proactive hard sample mining and closed-loop course learning, and generates more adversarial synthetic samples based on the failure logs of deployed models in real-world scenarios to strengthen model weaknesses, and is effective simultaneously for four downstream modes: perception / world model / VLA / VLM;

[0019] 7. Supports trajectory-level training data synthesis for end-to-end decision models such as the Visual-Language-Action (VLA) model and the world model. By performing multi-step rollout on labeled virtual terrain, it outputs physically consistent data. Trajectory triple sequence (required for world model) Includes The resulting ruts, subsidence, and soil compaction evolution, output (image_seq, language_instruction, action_seq) triples (required for VLA) and (image, Q, A) pairs automatically generated based on physical truth (required for VLM) make up for the deficiencies of existing off-road datasets in the direction of world model and end-to-end model.

[0020] To address the shortcomings of existing technologies, such as Figure 5 As shown, this invention proposes a method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS, including:

[0021] Step 1: Extract a triplet consisting of sensor stream, control command stream, and wheel-ground dynamics response stream from the vehicle's off-road driving log; pre-classify the triplet according to the ground material type.

[0022] Step 2: Using the wheel-ground dynamic response in the ternary set as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physical simulator are calibrated in reverse. , to obtain the posterior estimate The physical simulator estimates a posteriori. The wheel-to-ground contact displacement field is obtained by performing forward integration of the parameters in the simulator. ;

[0023] Step 3, determine the contact displacement field of the wheel. Differential mapping is applied to the 3DGS Gaussian point cloud in the three-dimensional Gaussian neural deformation field to synthesize synthetic sensing data after a tire passes over the ground. This synthetic sensing data includes: ground RGB image, depth map, and normal map.

[0024] Step 4: Match the virtual vehicle model with the real vehicle and the posterior estimate. The virtual terrain and control commands / expert action sequences are input into the physics simulator to obtain data including subsidence depth. Shear stress Semantic masking, passability, posterior estimation The pixel-wise aligned physical ground truth; the posterior estimate of each pixel in the synthetic perceptual data. Corresponding to the physical truth value, it serves as training data for off-road autonomous driving.

[0025] The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS further includes:

[0026] Step 5: The output adapter organizes the autonomous driving training data into the training data required by the downstream model according to the data content and format required for downstream model training.

[0027] The downstream model is a perception model, and the training data is... Triples; or

[0028] The downstream model is a world model, and the training data consists of the world model executing preset action strategies on virtual terrain. Step trajectory triple sequence , Let t represent the vehicle state and terrain physical state at time t. For the control command at time t The resulting vehicle state and terrain physical state at time t+1; or

[0029] The downstream model is a VLA model, and the training data consists of triples composed of image_seq, language_instruction, and action_seq; image_seq is the observation sequence; language_instruction is generated by an instruction template library or a large language model based on trajectory fragments and physical ground truth; action_seq is a control sequence time-aligned with the above observation sequence; or

[0030] The downstream model is a VLM model, and the training data consists of paired samples containing synthetic perceptual data and their corresponding semantically rich text.

[0031] The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS also includes a training step:

[0032] The downstream model is trained and the 3DGS mapping and mechanical constitutive parameters are jointly optimized via differentiable links. Using loss function ;in Mechanical constitutive parameters are fed back to the physics simulator via a three-dimensional Gaussian neural deformation field mapping. In the formula Switch by mode: When the downstream model is a perception model Cross-entropy loss function for the perception task When the downstream model is a world model The regression loss function for world model state prediction When the downstream model is a VLA model For action regression loss function Language-conditional alignment loss function The sum; when the downstream model is a VLM model. Cross-entropy loss function for text generation or visual question answering ;

[0033] Losses due to physical compliance include: (i) subsidence site constraints (ii) Shear field constraints (iii) Energy conservation constraint (iv) Momentum conservation constraint (v) Multi-material boundary continuity constraints When the downstream model is a world model, Additional constraints are included regarding the physical consistency of trajectory-level terrain evolution to require that the world model predictions... , , Consistent with the output of the physics simulator;

[0034] The next-moment terrain physical state output by the world model; This is the true physical value for the next moment, output by the physics simulator.

[0035] The aforementioned method for synthesizing off-road autonomous driving training data based on inverse calibration and 3DGS involves encrypting and uploading the driving log to the cloud, and then using the cloud to perform joint inverse calibration of constitutive parameters posteriorly based on H3 hex spatial index and material type blocks. Cross-platform standardization To achieve heterogeneous platform reuse, a DP differential privacy mechanism is provided. Based on the failure logs of the deployed model in real-world scenarios, various adversarial perturbations are performed by splitting the traffic according to the trigger type, so as to generate adversarial synthetic samples in a preset ratio to strengthen the model's weaknesses.

[0036] The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS, in step 2 These are parameters configured in the virtual terrain material model of the physics simulator; It can be any one or a combination of the following: pressure-sinking model parameters, shear stress-shear displacement model parameters, tire model parameters, rigid contact friction parameters, and tribological parameters.

[0037] The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS, wherein step 2 includes:

[0038] Based on the number of driving logs, the differentiability of the physics simulator, and the modal number of the objective function, an optimization algorithm is selected. Using the wheel-ground dynamics response in the triplicate as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physics simulator are calibrated in reverse. ;

[0039] When the physical simulator is differentiable, the optimization algorithm uses conjugate gradient descent; when the number of driving logs is less than a threshold, the optimization algorithm uses Bayesian optimization; when the number of modes of the objective function is greater than a threshold, the optimization algorithm uses an evolutionary algorithm in conjunction with the physical simulator.

[0040] like Figure 6 As shown, this invention also proposes an off-road autonomous driving training data synthesis device based on reverse calibration and 3DGS, including:

[0041] Module 1 extracts a triplet from the vehicle's off-road driving log, consisting of sensor stream, control command stream, and wheel-ground dynamics response stream; and pre-classifies the triplet according to the ground material type.

[0042] Module 2, using the wheel-ground dynamics response in the ternary set as the optimization objective, reverse-calibrates the mechanical constitutive parameters of the virtual terrain material in the physical simulator. , to obtain the posterior estimate The physical simulator estimates a posteriori. The wheel-to-ground contact displacement field is obtained by performing forward integration of the parameters in the simulator. ;

[0043] Module 3 will determine the contact displacement field of the wheel. Differential mapping is applied to the 3DGS Gaussian point cloud in the three-dimensional Gaussian neural deformation field to synthesize synthetic sensing data after a tire passes over the ground. This synthetic sensing data includes: ground RGB image, depth map, and normal map.

[0044] Module 4 will match the virtual vehicle model with the real vehicle and the posterior estimate. The virtual terrain and control commands / expert action sequences are input into the physics simulator to obtain data including subsidence depth. Shear stress Semantic masking, passability, posterior estimation The pixel-wise aligned physical ground truth; the posterior estimate of each pixel in the synthetic perceptual data. Corresponding to this physical truth value, it serves as training data for off-road autonomous driving;

[0045] Or may also include:

[0046] Module 5, the output adapter, organizes the autonomous driving training data into the training data required by the downstream model based on the data content and format required for downstream model training;

[0047] The downstream model is a perception model, and the training data is... Triples; or

[0048] The downstream model is a world model, and the training data consists of the world model executing preset action strategies on virtual terrain. Step trajectory triple sequence , Let t represent the vehicle state and terrain physical state at time t. For the control command at time t The resulting vehicle state and terrain physical state at time t+1; or

[0049] The downstream model is a VLA model, and the training data consists of triples composed of image_seq, language_instruction, and action_seq; image_seq is the observation sequence; language_instruction is generated by an instruction template library or a large language model based on trajectory fragments and physical ground truth; action_seq is a control sequence time-aligned with the above observation sequence; or

[0050] The downstream model is a VLM model, and the training data consists of paired samples containing synthetic perceptual data and their corresponding semantically rich text.

[0051] Or may also include:

[0052] It also includes a training module:

[0053] The downstream model is trained and the 3DGS mapping and mechanical constitutive parameters are jointly optimized via differentiable links. Using loss function ;in Mechanical constitutive parameters are fed back to the physics simulator via a three-dimensional Gaussian neural deformation field mapping. In the formula Switch by mode: When the downstream model is a perception model Cross-entropy loss function for the perception task When the downstream model is a world model The regression loss function for world model state prediction When the downstream model is a VLA model For action regression loss function Language-conditional alignment loss function The sum; when the downstream model is a VLM model. Cross-entropy loss function for text generation or visual question answering ;

[0054] Losses due to physical compliance include: (i) subsidence site constraints (ii) Shear field constraints (iii) Energy conservation constraint (iv) Momentum conservation constraint (v) Multi-material boundary continuity constraints When the downstream model is a world model, Additional constraints are included regarding the physical consistency of trajectory-level terrain evolution to require that the world model predictions... , , Consistent with the output of the physics simulator;

[0055] The next-moment terrain physical state output by the world model; The next physical truth value output by the physics simulator;

[0056] Or may also include:

[0057] The driving log was encrypted and uploaded to the cloud. The cloud then used H3 hex spatial indexing and material type to jointly perform a post-calibration of constitutive parameters. Cross-platform standardization To achieve heterogeneous platform reuse, a DP differential privacy mechanism is provided; based on the failure logs of the deployed model in real-world scenarios, various adversarial perturbations are performed by traffic splitting according to trigger type, so as to generate adversarial synthetic samples in a preset ratio to strengthen the model's weaknesses.

[0058] In module 2 These are parameters configured in the virtual terrain material model of the physics simulator; It can be any one or a combination of the following: pressure-sinking model parameters, shear stress-shear displacement model parameters, tire model parameters, rigid contact friction parameters, and tribological parameters.

[0059] This module 2 may also be used for:

[0060] Based on the number of driving logs, the differentiability of the physics simulator, and the modal number of the objective function, an optimization algorithm is selected. Using the wheel-ground dynamics response in the triplicate as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physics simulator are calibrated in reverse. ;

[0061] When the physical simulator is differentiable, the optimization algorithm uses conjugate gradient descent; when the number of driving logs is less than a threshold, the optimization algorithm uses Bayesian optimization; when the number of modes of the objective function is greater than a threshold, the optimization algorithm uses an evolutionary algorithm in conjunction with the physical simulator.

[0062] The present invention also proposes an electronic device, including the aforementioned off-road autonomous driving training data synthesis device based on reverse calibration and 3DGS. The electronic device may be connected to an information display device, which is used to display the off-road autonomous driving training data with user-set display parameters, attributes, or through an artificial intelligence model.

[0063] The present invention also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any of the described methods for synthesizing off-road autonomous driving training data based on inverse calibration and 3DGS.

[0064] The present invention also proposes a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the described methods for synthesizing off-road autonomous driving training data based on inverse calibration and 3DGS.

[0065] As can be seen from the above solutions, the advantages of the present invention are:

[0066] This invention establishes a coupled mapping relationship between visual synthesis and physical simulation, enabling interactive traces such as tire tracks, indentations, and fallen vegetation in synthesized images to correspond to the displacement fields of real physical processes. The same front-end computation (S1–S4) simultaneously supports the synthesis of training data for four downstream models—perceptual models, world models, VLA, and VLM—through the S5 multi-mode adapter, ensuring cross-mode ground truth physical consistency without requiring repetitive pipeline construction for different model types. It addresses the shortcomings of existing NeRF / 3DGS solutions (lack of physical interaction), game engine solutions (lack of realistic rheology), and potential diffusion solutions (physical causality breaks). In off-road scenarios, it achieves physical consistency evolution of the world model over a long period of 10 seconds with multi-step deduction, and directly drives VLA / VLM training data generation based on physical ground truth, resulting in semantic accuracy higher than LLM automatic image description solutions based on the image itself. It supports unified reverse calibration and synthesis of six typical off-road materials: sand, mud, snow, gravel, ice, and vegetation, reducing reliance on real-vehicle data collection and manual annotation compared to real-vehicle data acquisition. It utilizes H3 hex spatial indexing + Differential privacy vehicle-cloud federated reverse calibration enables multi-vehicle logs to be jointly reverse calibrated in the cloud, which is beneficial for new platforms to use existing calibration results for initialization, thereby shortening the calibration convergence time. By actively mining difficult samples and learning a closed loop of courses, adversarial samples are generated in reverse based on real vehicle failure logs, which is more targeted than static datasets and works simultaneously on four types of downstream models: perception, world model, VLA, and VLM. From the perspective of industrial deployment, it supports the continuous closed-loop reinforcement of various downstream models of off-road autonomous driving. Attached Figure Description

[0067] Figure 1 The overall system architecture and data flow diagram of this invention includes seven major modules and their data flow directions: S1 real vehicle acquisition and preprocessing, S2 physical simulator reverse calibration, S3 three-dimensional Gaussian neural deformation field, S4 multi-channel synthesis, S5 multi-mode output adapter, S6 downstream model training, and S7 cloud federation + hard sample mining.

[0068] Figure 2 The flowcharts for three parallel algorithm variants (conjugate gradient descent / Bayesian optimization / CMA-ES) for the inverse parameter calibration of the S2 physics simulator of this invention, and their selection decision trees according to the three-dimensional criteria of [log size, simulator differentiability, and noise level];

[0069] Figure 3 This invention presents a flowchart of the S3 three-dimensional Gaussian neural deformation field—from physical deformation field to 3DGS Gaussian parameter differentiable mapping, including the collaborative update of four sets of attributes: position, covariance, opacity, and spherical harmonic color, and the end-to-end gradient backpropagation link.

[0070] Figure 4This is a schematic diagram of the multi-mode output adapter for the downstream model type S5 of the present invention (5a Sensing / 5b World Model / 5c VLA / 5d VLM four parallel output modes);

[0071] Figure 5 This is a schematic diagram illustrating the collaborative closed-loop of the S7 vehicle-cloud federated reverse calibration architecture and active hard sample mining for four types of downstream modes in this invention.

[0072] Figure 6 This is a flowchart of the method of the present invention;

[0073] Figure 7 This is a block diagram of the device of the present invention;

[0074] Figure 8 This is a schematic diagram of the structure of the first electronic device of the present invention;

[0075] Figure 9 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;

[0076] Figure 10 This is a schematic diagram of the structure of the second electronic device of the present invention.

[0077] Figure label:

[0078] A - First electronic device;

[0079] B-Off-road autonomous driving training data synthesis device based on reverse calibration and 3DGS;

[0080] C-Data acquisition equipment;

[0081] D-Information display device;

[0082] 1000 - Second electronic device;

[0083] Ⅰ-Computational Unit;

[0084] II-ROM;

[0085] III-RAM;

[0086] N-bus;

[0087] V-Interface;

[0088] VI - Input Unit;

[0089] VII - Output Unit;

[0090] VIII - Storage medium;

[0091] IX - Communication Unit. Detailed Implementation

[0092] It should be noted that, in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0093] In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0094] The processor described in this invention is the control center of an electronic device. It can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of this invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0095] Alternatively, the processor can perform various functions of the electronic device by running or executing software programs stored in memory and by calling data stored in memory.

[0096] In a specific implementation, as one example, the processor may include one or more CPUs. Each of these processors may be a single-core processor or a multi-core processor. Here, "processor" can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include servers, desktop computers, laptops, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.

[0097] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.

[0098] It should be noted that the structure of the electronic device shown in the accompanying drawings of this invention does not constitute a limitation thereof. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0099] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0100] It should also be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0101] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0102] It should also be understood that, in various embodiments of the present invention, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0103] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0106] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0107] This invention discloses a method, system, device, and storage medium for synthesizing training data for off-road autonomous driving models based on inverse calibration of a physical simulator and a 3D Gaussian Neural Deformation Field. This method extracts the wheel-ground dynamic response flow from field driving logs and pre-classifies it according to material type; in a multi-body dynamics physics simulator, the mechanical constitutive parameters of the virtual terrain material are inversely calibrated using the response flow as the target, and the posterior estimate is output; the displacement deformation field output by the simulator is applied to the pre-trained 3D Gaussian point cloud through a differentiable mapping, so that the position, covariance, opacity and spherical harmonic color of each Gaussian ellipsoid are co-deformed, and a synthetic RGB image with physical interaction traces such as ruts, vegetation collapse, and subsidence is generated and a multi-channel physical ground truth is generated; the above synthetic data is adapted to any one or a combination of four parallel modes: perception model, world model, visual-language-action model (VLA), and visual-language model (VLM) through an output adapter; a physical compliance loss is introduced into the total training loss to make the model output consistent with the physical ground truth; the method further supports multi-platform vehicle-cloud federated inverse calibration and active hard sample mining and course learning closed loop based on failure logs, and both are effective for the four downstream modes simultaneously. This invention fundamentally alleviates the problems of data scarcity, physical distortion, and cross-material generalization in four types of downstream models: off-road autonomous driving perception, world model, VLA, and VLM.

[0108] symbol Physical meaning unit Constitutive parameter vectors of material mechanics (any one or a combination of model parameters such as Bekker-Wong, Janosi-Hanamoto, Magic Formula, Pacejka, and LuGre) Depends on the model Posterior estimates of constitutive parameters after inverse calibration same Constitutive parameters posterior covariance matrix same Position in three-dimensional space At any moment displacement deformation field m Forward mapping function of multibody dynamics physics simulator (input vehicle state and control, constitutive parameters, output wheel-ground dynamic response at the next time step) — Wheel-to-ground dynamic response observation vectors in the vehicle log View Channel The five channels of the actual vehicle's four wheels are: longitudinal contact tangential force, normal contact force, motor torque, wheel speed, and suspension displacement. N / N / N·m / rad·s⁻¹ / m Pre-trained 3D Gaussian point cloud set (position / covariance / opacity / spherical harmonic color coefficients); Nɢ is the number of Gaussian elements. — Deformation field Jacobian matrix — Synthesize RGB images (expandable to multi-channel perceptual data such as depth maps, normal maps, and semantic masks). — Perceptual task labels (semantic mask / accessibility binary / depth / normal, etc.) — and Aligned pixel-wise physical truth tuples (sink depth per pixel) Shear stress Constitutive parameter lookup value, passability binary, hazard target mask) — World model trajectory triples: These represent the 6-DoF pose and terrain physical state at times t and t+1, respectively, including the vehicle's position. , , Representing the control actions The resulting change in the physical state of the terrain. The control commands at time t (throttle / brake / steering / torque of each wheel). =Change in rut depth / subsidence depth; =Change in soil compaction degree; =Change in vegetation lodging angle. — Original monitoring losses of downstream tasks (taken in S5 mode) / / / ) — Physical compliance loss (five mathematical forms available: subsidence field / shear field / energy conservation / momentum conservation / multi-material boundary continuity constraint) — Weight in total loss — Total training loss of downstream model —

[0109] Represents shear stress; , These represent the system's energy and momentum under the constraints of energy conservation and momentum conservation, respectively. This represents the number of log samples in the federally reverse-calibrated dataset (i.e., the sample size in the constitutive parameter posterior p(θ|D1:N)). This indicates the adversarial sample ratio in proactive hard sample mining. This represents the differential privacy budget. V represents the number of heterogeneous platforms (vehicles) participating in the federal reverse calibration; M represents the number of templates in the instruction template library (M≥50); K and T represent the number of starting points for the world model rollout and the number of time steps for a single rollout, respectively.

[0110] In this paper, the 3D Gaussian Neural Deformation Field refers to a dynamic Gaussian field representation that is end-to-end differentiated. It is formed by using the scene point cloud represented by 3D Gaussian Splatting (3DGS) as the carrier, the wheel-ground contact displacement deformation field output by the physical simulator as the driving source, and the geometric and appearance properties of each Gaussian element (Gaussian ellipsoid) are updated collaboratively with the displacement deformation field through differentiable deformation mapping. Specifically, for each Gaussian primitive in the scene, its position (mean) is updated by translation according to the displacement deformation field, its covariance matrix is ​​transformed by the Jacobian matrix of the displacement deformation field, and its opacity and spherical harmonic color coefficient are conditionally updated according to the local optical properties of the material in which the primitive is located (including but not limited to the change in reflectivity caused by changes in water content / humidity). All of the above update processes are differentiable to the terrain constitutive parameters calibrated by the physical simulator, and the gradient can be backpropagated end-to-end, so that physical simulation and neural rendering are jointly optimized in the same differentiable computation graph.

[0111] It should be noted that the three-dimensional Gaussian neural deformation field described in this paper is fundamentally different from the following two types of existing representations: (i) it is different from the three-dimensional Gaussian sputtering new perspective synthesis schemes for geometry / appearance reconstruction or general dynamic scenes (such as the 3DGS / NeRF scheme based on multi-view image reconstruction and the general 4D dynamic Gaussian scheme). The latter's Gaussian deformation is obtained by implicit fitting of time or latent variables through neural networks, and is not driven by the vehicle-terrain contact mechanics simulation output, nor does it produce continuous physical truth values ​​aligned pixel by pixel with the deformed pixels; (ii) it is different from the Gaussian physical property estimation scheme based on visual perception inversion of physical properties (such as PhysGS / GaussianProperty / NeRF2Physics, etc.). The latter estimates physical properties forward from the image, rather than driving Gaussian deformation backward from the physical displacement field. In this paper, the deformation of the three-dimensional Gaussian neural deformation field originates from the displacement deformation field output by the physical simulator (multibody dynamics / terrain mechanics solver) after inverse calibration of the actual vehicle's wheel-ground dynamics response, rather than being generated by the neural network fitting the visual training distribution. Therefore, its deformation has physical causal consistency in both geometry and appearance, and can simultaneously output the physical true value aligned with it pixel by pixel.

[0112] To achieve the above-mentioned technical effects, the present invention proposes the following key technical points:

[0113] Key Point 1: Reverse calibration of the physics simulator (Step B), using the actual vehicle's wheel-ground dynamics response (wheel speed, torque, suspension displacement, vehicle acceleration) as the optimization objective, and adjusting the constitutive parameters of the virtual terrain. Perform reverse calibration. It explicitly enumerates multiple constitutive equation combinations such as Bekker-Wong, Janosi-Hanamoto, Magic Formula, Pacejka, and LuGre, and the identification algorithm supports three parallel variants: conjugate gradient descent, Bayesian optimization, and CMA-ES. Technical effect: relying solely on physical sensing inversion, it maintains recognizability even in off-road scenarios with visual degradation (strong light, vegetation obstruction, snow blindness, etc.) due to its orthogonality with the visual source. Furthermore, the constitutive parameters are decoupled from vehicle dynamics and can be reused across heterogeneous platforms, which is different from the existing image-based inversion solutions such as PhysGS, GaussianProperty, and NeRF2Physics.

[0114] Key Point 2: Three-dimensional Gaussian neural deformation field (step C), which is the wheel-to-ground contact displacement deformation field output by the simulator. By applying differentiable mappings to a pre-trained 3D Gaussian point cloud, the position of each Gaussian ellipsoid is determined. according to Update, covariance according to Perform Jacobian transformation, opacity Color coefficients with spherical harmonics Updated based on local optical properties of the material; Technical effect: The synthesized image explicitly presents physical interaction traces such as ruts, indentations, and fallen vegetation, and the gradient of the entire deformation mapping process can be back-propagated to the constitutive parameters end-to-end. This allows for joint optimization of physical simulation and neural rendering, thereby addressing the dual shortcomings of existing 3DGS new perspective synthesis schemes that do not use vehicle-terrain contact mechanics simulation output as deformation drive and existing game engine schemes that lack realistic soft terrain rheological characteristics.

[0115] Key Point 3: Multi-mode output adapter for downstream model types (step E, including four parallel outputs of modes E1–E4), after sharing the calculations of front-end S1–S4, switches to the E1 perception model according to the target downstream model type (output). Triplet) / E2 World Model (output) Physically consistent trajectory triplet, At least include , , Four modes are supported: terrain evolution quantity, E3 VLA model (outputting triples of image_seq, language_instruction, action_seq), and E4 VLM model (outputting image-text pairing samples directly driven by physical truth). Technical benefits: The same front-end computing can simultaneously support the synthesis of training data for four types of downstream models and ensure physical consistency of truth across modes, avoiding the need to repeatedly build pipelines for different model types; it makes up for the shortcomings of existing open off-road datasets in the areas of world models and end-to-end models.

[0116] Key Point 4: Physical Compliance Losses End-to-end training of downstream models driven by the driver (step F), total loss of downstream models , Four forms are available in S5 mode ( / / / ), Five mathematical forms are explicitly supported: (i) Settlement field constraints (ii) Shear field constraints (iii) Energy conservation constraint; (iv) Momentum conservation constraint; (v) Multi-material boundary continuity constraint; Technical effect: The downstream model output is consistent with the physical truth in five independent dimensions, and The constitutive parameters are backpropagated end-to-end through a differentiable deformation map. This constitutes an end-to-end reverse supervision link from the downstream model output to the physical constitutive parameters, which is different from schemes that only perform forward physical property estimation or only perform visual reconstruction.

[0117] Key Point 5: Vehicle-Cloud Federated Reverse Calibration and Active Hard Sample Mining Closed Loop (Step G). Multi-vehicle logs are segmented by H3hex spatial index and material type for joint reverse calibration of constitutive parameter posterior in the cloud, with a corresponding differential privacy protection mechanism. Based on real vehicle failure logs, four adversarial perturbations (FGSM / PGD / CMA-ESadversarial / Bayesian uncertainty-aware) are applied according to trigger type to generate 15%–25% of adversarial synthetic samples. The closed loop is iterated periodically according to the course learning until the failure rate converges. Technical effects: Constitutive parameter posterior is reused between heterogeneous platforms, and the cold start convergence time of new platforms is significantly reduced. The failure rate decreases monotonically with the number of closed loop rounds. This closed loop is effective for four types of downstream modes: perception, world model, VLA, and VLM, supporting the continuous closed-loop reinforcement of various downstream models of off-road autonomous driving from the perspective of actual deployment.

[0118] Overall technical solution:

[0119] This invention provides a method for synthesizing training data for off-road autonomous driving models based on inverse calibration of a physical simulator and a three-dimensional Gaussian neural deformation field. The overall input of this invention is multi-source sensors from real vehicles / simulation and logs; the output is synthetic samples (images + physical ground truth / trajectory / image-text pairs) for downstream model training.

[0120] Its core workflow includes the following seven main steps: S1–S4 are shared front-end computations for different downstream model types; S5 is an output adapter for multiple downstream model types (switching output modes according to the target model type, each output mode can be selected independently or used in combination); S6 is downstream model training (selecting the corresponding task loss according to the output mode in S5); and S7 is cloud-based federated calibration and hard sample closure. For the overall system architecture and data flow between modules, please refer to the appendix. Figure 1 .

[0121] 1. Physical calibration data stream acquisition and preprocessing of real vehicle logs (S1): Extract triples [sensor stream, control command stream, wheel-ground dynamic response stream] from off-road driving logs of one or more mobile platforms; align them to the same reference time base by timestamp; pre-classify by ground material type, for example, the classification object is driving logs, based on manual annotation or weakly supervised clustering (such as moisture content, color histogram, wheel-ground response characteristics), typical material types include sand, mud, snow, gravel, ice, and vegetation. The purpose is to allow each medium block in S2 to be reverse-calibrated for constitutive parameters.

[0122] 2. Inverse parameter calibration of the physics simulator (S2): Using the wheel-ground dynamics response in the driving log as the optimization target, the mechanical constitutive parameters of the virtual terrain material in the physics simulator are calibrated in reverse. , to obtain the posterior estimate and covariance , These are parameters configured in the virtual terrain material model of the physics simulator; Including but not limited to Bekker-Wong pressure-sinking model parameters The parameters are any one or a combination of the following: Janosi-Hanamoto shear stress-shear displacement model parameters, Magic Formula tire model parameters, Pacejka rigid contact friction parameters, and LuGre tribological parameters. (After calibration) This is used in subsequent steps: S3 deformation, S4 rendering of truth values, and S5 rollout.

[0123] The physical simulator can employ a coupled multibody dynamics (MB) / discrete element method (DEM) solver for tire-soil contact and vehicle dynamics integration. Typical implementations include (but are not limited to): Adams / Car, RecurDyn, CarSim RT, and other multibody solvers; upper particles / gravel can be solved using DEM or equivalent simplified layers; if end-to-end gradient backpropagation is required, a differentiable MB solver path can be selected.

[0124] like Figure 2 As shown, the optimization algorithm supports three parallel variants: (a) conjugate gradient descent (suitable for analytically differentiable simulators); (b) Bayesian optimization (suitable for black-box simulators + small samples, Gaussian process posterior directly outputs parameter covariance matrix). (c) Co-evolutionary algorithm CMA-ES (applicable to multimodal objective functions); selection rules are based on three-dimensional criteria: [number of log samples, differentiability of simulator, number of objective function modes].

[0125] The key technical point of this invention lies not in proposing a new optimizer, but in presenting it as a parallel engineering variant of S2, which "uses the actual vehicle's wheel-ground dynamics response as the target and reverse-calibrates the constitutive parameters θ of the virtual terrain," and provides selection criteria based on "simulator differentiability, log size, and noise level." The resulting technical effect is: obtaining a simulation consistent with that of the actual vehicle. This allows the synthesized S3–S4 images to contain traces of real physical interactions, thereby improving the physical consistency of downstream perception / world model / VLA / VLM in off-road hold-out.

[0126] 3. Differentiable mapping from physical deformation field to 3DGS Gaussian parameters (S3): Based on the above-calibrated posterior estimate of the muddy ground. The simulator then replays the vehicle's movement on the muddy ground to obtain the wheel's force curve and the three-dimensional deformation of the ruts / subsidence on the ground, thus obtaining the wheel-ground contact displacement field. ;like Figure 3 As shown, the wheel-to-ground contact displacement field output by the simulator is... Applying differentiable mappings to pre-trained 3DGS Gaussian point clouds This results in coordinated deformation of both geometry and appearance. The simulator is then subjected to a ground contact displacement field. Differentiable mappings are applied to pre-trained 3DGS Gaussian point clouds to collaboratively update position, covariance, opacity, and spherical harmonic color, thereby synthesizing traces such as tire tracks, subsidence, and fallen vegetation. The 3DGS scene itself is first pre-trained using standard 3DGS reconstruction from real-world multi-view images. Training / joint optimization: to be discussed in subsequent S6 iterations. It can be backpropagated via S3 differentiable mapping to This invention does not simply compare the loss using "real-world parameters vs. synthetic traces", but rather aligns the downstream task loss + physical compliance loss with the simulation truth value.

[0127] Specifically, the position of each Gaussian ellipsoid is determined by... Update; covariance by Perform Jacobi transformation ( (For the deformation field Jacobian matrix); Opacity Color coefficients with spherical harmonics The conditions are updated based on the local optical properties of the material (such as changes in reflectivity caused by variations in humidity); the gradient of this update process can be propagated back to the constitutive parameters end-to-end. This allows for joint optimization of physical simulation and neural rendering.

[0128] 4. Synthetic perceptual image generation and multi-channel ground truth alignment (S4):

[0129] The virtual vehicle model matched with the real vehicle, and the constitutive parameters The virtual terrain / DEM, along with control commands / expert action sequences that align and play back or sample from S1, are input into the physical simulator to obtain the subsidence depth. Shear stress Accessibility binary labels can be generated by the simulator based on thresholding of mechanical quantities such as subsidence, slope, and friction, or output together with the rendering semantic channel.

[0130] The 3DGS rendering pipeline renders and synthesizes multi-channel sensor data, including RGB images, depth maps, and normal maps, under specified camera intrinsic and extrinsic parameters; each pixel is simultaneously aligned with its corresponding physical ground value: posterior estimation. Subsidence depth Shear stress The four sets of fields are implemented at the engineering level as a single dictionary or an equivalent multichannel tensor dataset, supporting random access by image number and slice query by physical quantity type.

[0131] 5. Multimodal Dataset Output Adapter for Downstream Model Types (S5): Adapts the synthetic sensing data and physical ground truth output from step S4 to any one or a combination of the following four parallel output modes according to the downstream model type (sharing all computations from frontend S1–S4, switching the output adapter only in S5 to achieve cross-mode physical consistency of ground truth):

[0132] Mode 5a is output by the perceptual model. Triples are used to train various perceptual networks, including semantic segmentation, accessibility prediction, depth estimation, physical parameter estimation, and semantic occupancy grid prediction.

[0133] Mode 5b is executed by the world model on predefined virtual terrain using preset action strategies (expert strategy / random strategy / adversarial perturbation strategy). Step physical rollout, output Trajectory triple sequence It includes both vehicle status (position, speed, suspension displacement) and terrain physical status (rut deformation field, subsidence depth map, soil compaction, vegetation lodging angle). Includes Changes in the physical state of the terrain , , Equal topographic evolution;

[0134] The term "calibrated" refers to the virtual off-road terrain that S2 has completed inverse calibration, specifically including: (1) matching the virtual vehicle geometry / inertia / suspension with the real vehicle; and (2) posterior constitutive parameters of each terrain block. (and covariance) (3) Optional preload DEM / scene geometry. 5b Perform T-step physical rollout on the calibrated terrain to ensure and The cause and effect are determined by f_sim(·; ) Guarantee; f_sim(·; ) indicates that the parameters have been calibrated. Forward mapping of the physics simulator: input state and Output the next state and contact / topographic evolution field.

[0135] Mode 5c uses the VLA model to align the visual sequence synthesized by S4 with the corresponding expert action sequence according to a time window, and uses an instruction template library (parametric scene description + task objective). The system can automatically generate natural language instructions for each trajectory segment using either a parameterized template or a large language model rewriting (multi-language / multi-style / multi-abstraction level), outputting a triplet of (image_seq, language_instruction, action_seq);

[0136] The expert action sequence is the aligned control instruction sequence `action_seq` used for training Mode 5c VLA, including throttle / braking / wheel torque, etc., aligned with the `image_seq` time window. Sources: ① Control instruction stream playback from the S1 real vehicle log; ② Rollout of preset strategies such as expert strategy / rule strategy / MPPI on calibrated virtual terrain; ③ Physically consistent trajectory slices from Example 7. The `language_instruction` is then automatically generated by the template library / LLM. `image_seq` is a multi-frame forward-facing camera / LiDAR observation sequence (synthetic perception data); `language_instruction` is a natural language task instruction automatically generated from the instruction template library (parametric scene description + task objective) or rewritten by LLM; `action_seq` is the expert / strategy control sequence (throttle, braking, steering, wheel torque) aligned with the above visual sequence time.

[0137] Mode 5d automatically generates semantically rich text from the physical ground truth (soil type, constitutive parameter estimation, subsidence prediction, accessibility level, detour cost estimation, etc.) output by the VLM model through structured templates or a large language model. The output consists of paired samples of (image, text) or (image, question, answer). For example, (image, text) is a synthetic wet mud RGB image + text "Wet clay mudflat, estimated subsidence 0.12 m, cohesion c=8 kPa, recommended speed limit 5km / h"; (image, question, answer) is an image of the same scene + Q "Is it safe to pass 15 m ahead?" + A "Not recommended: Simulated subsidence of 0.15 m exceeds the safety threshold, recommended to detour to the left, estimated detour time increased by 8 s". The semantic accuracy of the generated text is directly guaranteed by the physical ground truth, which is superior to the image-based LLM automatic image captioning scheme.

[0138] 6. Physical compliance losses Downstream model training driven by S6, such as Figure 4 As shown:

[0139] Used to train the downstream model selected by S5; Switch by mode: When mode is 5a for When the mode is 5b for When the mode is 5c for When the mode is 5D for The S1–S4 data synthesis pipeline itself is not a single neural network; it can optionally be used to synthesize data from 3DGS / sensor heads. Joint fine-tuning. Specifically:

[0140] The total training loss of the downstream models is uniformly set as ,in Take the output mode of S5 respectively (Perception tasks, cross-entropy / IoU / L1 / Focal Loss, etc.) (World Model State Prediction Regression) / (VLA motion regression + language conditional alignment) / (VLM text generation or visual question answering cross-entropy); For physical compliance losses, five mathematical forms are supported: (i) subsidence field constraints (ii) Shear field constraints (iii) Energy conservation constraint (iv) Momentum conservation constraint (v) Multi-material boundary continuity constraints ; Regarding the world model (5b model), Additional constraints may be included regarding the physical consistency of trajectory-level terrain evolution, thus requiring the world model to predict... , , Consistent with the output of the physics simulator, this is the key monitoring signal that distinguishes the off-road world model from the general world model.

[0141] The above Refers to two items in the state prediction regression loss of the world model (5b): =The world model's prediction of the state in the next moment; The next state of ground-truth obtained from the physical rollout in S5 mode 5b.

[0142] For the action regression loss of mode 5c VLA: Align the model's predicted actions with expert action_seq (in the form of L1 / L2 or diffusion / flow matching, depending on the VLA architecture).

[0143] Five kinds Both are auxiliary losses used to constrain the consistency between the downstream model output and the physical truth of the simulator during S6 training: (i) Settlement field constraint L=‖z_net−z_sim‖², The network-predicted subsidence map output by the downstream model. (ii) Shear field constraints, (i) Pixel-by-pixel sinking truth values ​​for the S4 simulator; / The images show the shear stress outputs from the downstream model and the simulator, respectively. , The energy output from the downstream model and simulator, respectively; and Output momentum for downstream model and simulator respectively; Multi-material boundary continuity: in the neighborhood of medium interface constitutive parameter field predicted by the network Apply total variation penalty L=∫_{ }‖ ‖1, suppress parameter jumps at boundaries such as sand-mud and mud-snow, which is achieved by calculating gradient penalty within the boundary expansion band of the semantic / material segmentation mask.

[0144] 7. Vehicle-Cloud Federated Reverse Calibration and Active Hard Sample Loop Closure (S7): (e.g., ...) Figure 5 As shown, multiple mobile platforms encrypt and upload logs to the cloud. The cloud then uses H3 hex spatial indexing (typically H3 res-9 hex is about 174 m) and material type blocks to jointly inversely calibrate the constitutive parameters posteriorly. Cross-platform standardization To achieve heterogeneous platform reuse, and provide supporting... -DP differential privacy mechanism; at the same time, based on the failure logs of the deployed model in real-world scenarios (including six trigger types: false detection / missed detection, control instability, long-term prediction divergence of the world model, VLA task failure, and VLM numerical QA error), four adversarial perturbations (FGSM / PGD / CMA-ESadversarial / Bayesian uncertainty-aware) are generated according to the trigger type to strengthen the model's weaknesses by generating adversarial synthetic samples in a preset proportion; this closed loop is effective for all four S5 output modes simultaneously.

[0145] To make the above-mentioned features and effects of the present invention clearer and easier to understand, specific embodiments are described below in conjunction with the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely illustrative. The scope of protection of the present invention is not limited to the disclosed embodiments, but is defined by the appended claims.

[0146] The overall workflow of this invention is shown in Figure 1, which includes seven core modules and their data flow: (1) Real vehicle log acquisition and preprocessing (S1 Multi-platform [sensor flow, control command flow, wheel-ground dynamic response flow] triplet alignment and clustering by timestamp); (2) Inverse parameter calibration of the physical simulator (S2 Inverse calibration of constitutive parameters of virtual terrain Bekker-Wong / Janosi-Hanamoto / Magic Formula / Pacejka / LuGre etc. in the multibody dynamics solver with log as the target) The optimization algorithm supports three parallel variants: conjugate gradient descent, Bayesian optimization, and CMA-ES. (3) Three-dimensional Gaussian neural deformation field (S3 updates the position, covariance, opacity, and color of the Gaussian ellipsoid according to the physical displacement field through differentiable deformation mapping). (4) Multi-channel synthesis and physical truth alignment (S4 renders multiple channels such as RGB, depth, normal, and semantics and pairs them pixel by pixel). (5) Downstream model type multi-mode output adapter (S5 switches to four parallel output modes according to the target model type: Perception 5a / World Model 5b / VLA 5c / VLM 5d); (6) Physical compliance loss Downstream model training driven by S6 total loss , (7) Vehicle-cloud federated reverse calibration and active hard sample closed loop (S7 multi-platform log encryption upload to the cloud joint reverse calibration + failure log driven hard sample generation simultaneously effective for four types of downstream modes). This application presents the technical solution from a semantic level through the following nine specific embodiments for different materials, different algorithm variants, and different The implementation process under different forms and downstream model types.

[0147] To facilitate reference to the subsequent specific embodiments, this section will first focus on describing the key data structures and shared front-end computing conventions used in the various embodiments described in this application.

[0148] 1. Key Data Structure (Pixel-by-Pixel Physical Truth Tuple): The physical truth paired for each synthesized image is carried in the engineering implementation as a single tuple, logically divided into four groups of fields: (a) Constitutive Parameter Group: Select any one or a combination of Bekker-Wong / Janosi-Hanamoto / Magic Formula / Pacejka / LuGre by material type; (b) Physical state group: depth of depression per pixel. Shear stress (c) Terrain evolution component group (only used in 5b world model mode): rut deformation Changes in soil compaction Changes in vegetation lodging angle (d) Uncertainty group: posterior covariance of constitutive parameters .

[0149] 2. Engineering conventions for shared front-end S1–S4 and multi-mode output adapter S5: Front-end S1–S4 calculations are fully shared across all downstream model types. The output adapter is switched only in step S5 according to the target downstream model type, with mode 5a output. Triplet; Mode 5b Output Trajectory triple sequence and Includes The resulting change in terrain physical state; Mode 5c outputs a triplet of (image_seq, lang_instr, action_seq); Mode 5d outputs a pair of (image, text) or (image, Q, A). Examples 1 to 6 mainly demonstrate engineered variants of the Mode 5a perception model applied to a shared front end, while Examples 7 to 9 demonstrate the three adapters for Modes 5b, 5c, and 5d.

[0150] 3. General computing power and software stack conventions: The training nodes adopt a general-purpose GPU cluster, the physics simulator is a general-purpose multibody dynamics solver, the 3DGS rendering framework is a general-purpose 3D Gaussian rendering pipeline, and the downstream model training framework is a general-purpose deep learning framework. This invention is not limited to a specific manufacturer or version.

[0151]

Example 1: Single-vehicle single-material (wet mud) basic closed loop

[0152]

Example 2: Multi-material hybrid reverse calibration and cross-material generalization

[0153]

Example 3: Engineering Selection of Three Parameter Identification Algorithm Variants

[0154] [Example 4: Multiple Mathematical Forms of Physical Compliance Losses and Their Adaptation to Multiple Downstream Tasks] Describes five types of physical compliance losses. The engineering matching process between mathematical forms and four types of downstream tasks: (1) Five types Mathematical enumeration: (i) Settlement field constraints (pixel by pixel) (i) Norm); (ii) Shear field constraint (pixel by pixel) (iii) Energy conservation constraints (Evaluation per physics step); (iv) Momentum conservation constraint (v) Includes inter-frame smoothing terms; (v) Multi-material boundary continuity constraints (1) Total variation norm of the material boundary neighborhood); (2) Default weights and engineering parameter tuning experience, the default values ​​of the five forms. To enable adjustable hyperparameters, during the initial training phase in engineering practice... Perform a linear warm start to avoid cold starts and Gradient direction conflict; (3) The adaptation rules of each physical compliance loss form to the four types of downstream tasks are as follows: the combination of form (i) + (ii) is most sensitive to semantic segmentation and passability prediction of the perception model; the combination of form (iii) + (iv) is most effective for the cumulative drift and inter-frame jitter of long-term extrapolation of the world model; form (iv) alone is most sensitive to the action coherence of the VLA model; the combination of form (i) + (ii) is most effective for improving the accuracy of numerical question answering of VLM; the weighted combination of each physical compliance loss form achieves peak improvement in all four types of downstream tasks; (4) In engineering implementation, due to the various The physical dimensions and numerical scales of the terms differ significantly (the energy term is approximately...). J-order, momentum term (on the order of kg·m / s), the weights need to be standardized before being unified. When performing joint search in multiple forms, the default weights are first used to train the baseline according to the task, and then Bayesian hyperparameter optimization is used in five dimensions. (5) Perform finite number of samplings in space; end-to-end reverse monitoring of link integrity, five forms All data is routed through the following path: [downstream model output → consistency with physical simulation → reverse propagation to constitutive parameters]. This end-to-end link differs from the open-loop architecture of existing 3DGS physical attribute estimation schemes [which only perform forward estimation and have no downstream loss coupling]. [Example 5: Multi-vehicle Cloud-based Federated Reverse Calibration Architecture] Description The engineering implementation of the heterogeneous platform sharing constitutive parameter posterior through cloud federated reverse calibration architecture: (1) The system consists of several heterogeneous off-road vehicles with curb weight as edge nodes, which communicate with the multi-copy cloud cluster through a wireless wide area uplink channel. The message middleware is responsible for the log channel from the vehicle end to the cloud and the bidirectional control plane. The data persistence layer uses time-series databases to store the original log stream, relational databases to store the reverse calibration results, and object storage to store the composite image. The specific component selection is determined by the deployment party according to the engineering preferences; (2) Privacy preprocessing pipeline at the edge node end: The GPS field of the real vehicle log is quantized into an H3 hex index (typical resolution is about 100 meters, balancing spatial resolution and single hex sample sufficiency) + Gaussian noise superposition. The vehicle ID hash is a pseudo identifier (the vehicle model information is retained for subsequent use). (3) When the cloud performs joint inverse calibration by H3 hex × material type blocks, the Bayesian optimization process is triggered to expand and shrink elastically when the single (hex, material) cumulative sample reaches the threshold, and cross-vehicle standardization is performed. All subsequent segments and vehicles share the same... Perform joint reverse calibration and output And update the model version number; (4) Differential privacy budget control and conflict resolution - Differential privacy mechanism with Gaussian noise addition, single hex reverse calibration rate limiting + differential privacy budget ledger auditing; local vehicle-side operation. With the cloud When the deviation exceeds the threshold, it is automatically accepted by the cloud; otherwise, an abnormal review queue is triggered. The model version number table retains historical versions to support single-step rollback as an emergency hot-fix method. (5) Engineering performance indicators include a significant reduction in the convergence standard deviation of multi-vehicle joint reverse calibration compared to single-vehicle independent calibration, and cross-vehicle Relative deviation After scaling and compression, the cold start convergence time of the new platform is greatly shortened (by directly pulling the existing hex posterior for warm start). This architecture makes up for the engineering deficiencies of existing off-road synthetic data solutions in the direction of vehicle-cloud federated reverse calibration.

[0155]

Example 6: Active Hard Sample Mining Based on Real Vehicle Failure Logs

[0156] [Example 7: Trajectory Synthesis for Training an Off-Road World Model (Mode 5b)] This describes the training and synthesis of complete tracks for an off-road world model on virtual off-road terrain that has been inversely calibrated via S2. The entire process of trajectory triplet: (1) Virtual terrain and starting point sampling, based on the wet clay mudflats reverse-calibrated in Example 1. Virtual terrain is loaded onto the actual digital elevation model, and several starting points are sampled on the terrain grid using the Latin hypercube method. And according to the terrain location (entrance / middle section / exit / boundary), the location is forced to be covered by buckets, and the initial vehicle speed is sampled in a uniform distribution; (2) Multi-material scale decomposition, the total trajectory is distributed across six typical off-road materials (sand, wet clay, snow, gravel, thin ice, meadow) according to the material category ratio to ensure Cartesian product coverage; (3) Three strategies are used for parallel action sampling, expert driving log playback + random uniform disturbance + adversarial disturbance (in Outside the boundary + CMA-ES maximizes the baseline world model prediction RMSE inversion (3) Strict upper limit on the weights of the adversarial strategy to prevent training distribution drift; (4) Physical deduction and multi-view visual synthesis, each step is executed Physical integral, outputting terrain evolution components / / When the state changes drastically, substep subdivision is triggered to ensure numerical stability; severe instability triggers a slide bar to terminate the round and truncate the data; multi-view cameras call S3-S4 to render and synthesize RGB + depth + normal, and backcheck each pixel. Physical truth; (5) Downstream world model training and results, packaged The training samples are used to train the general world model, and the total loss is ,in It is a five-component weighted combination; the world model synthesized through this trajectory outperforms the model without any prior knowledge in long-term terrain evolution prediction. The training baseline, with all three strategies being indispensable, enables the general latent diffusion world model to possess engineering-usable, physically consistent, long-term extrapolation capabilities in off-road scenarios, compensating for the lack of features in existing off-road datasets. The lack of trajectory triples.

[0157]

Example 8: Command-Action-Observation Triad Combination for Training Off-Road VLA Model (Mode 5c)

[0158]

Example 9: Physical Truth-Driven Image-Text Pairing Synthesis for Training Off-Road VLM Model (Mode 5d)

[0159] The following are system embodiments corresponding to the above method embodiments. This embodiment can be implemented in conjunction with the above embodiments. The relevant technical details mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.

[0160] like Figure 6 As shown, this invention also proposes an off-road autonomous driving training data synthesis device based on reverse calibration and 3DGS, including:

[0161] Module 1 extracts a triplet from the vehicle's off-road driving log, consisting of sensor stream, control command stream, and wheel-ground dynamics response stream; and pre-classifies the triplet according to the ground material type.

[0162] Module 2, using the wheel-ground dynamics response in the ternary set as the optimization objective, reverse-calibrates the mechanical constitutive parameters of the virtual terrain material in the physical simulator. , to obtain the posterior estimate The physical simulator estimates a posteriori. The wheel-to-ground contact displacement field is obtained by performing forward integration of the parameters in the simulator. ;

[0163] Module 3 will determine the contact displacement field of the wheel. Differential mapping is applied to the 3DGS Gaussian point cloud in the three-dimensional Gaussian neural deformation field to synthesize synthetic sensing data after a tire passes over the ground. This synthetic sensing data includes: ground RGB image, depth map, and normal map.

[0164] Module 4 will match the virtual vehicle model with the real vehicle and the posterior estimate. The virtual terrain and control commands / expert action sequences are input into the physics simulator to obtain data including subsidence depth. Shear stress Semantic masking, passability, posterior estimation The pixel-wise aligned physical ground truth; the posterior estimate of each pixel in the synthetic perceptual data. Corresponding to this physical truth value, it serves as training data for off-road autonomous driving;

[0165] Or may also include:

[0166] Module 5, the output adapter, organizes the autonomous driving training data into the training data required by the downstream model based on the data content and format required for downstream model training;

[0167] The downstream model is a perception model, and the training data is... Triples; or

[0168] The downstream model is a world model, and the training data consists of the world model executing preset action strategies on virtual terrain. Step trajectory triple sequence , Let t represent the vehicle state and terrain physical state at time t. For the control command at time t The resulting vehicle state and terrain physical state at time t+1; or

[0169] The downstream model is a VLA model, and the training data consists of triples composed of image_seq, language_instruction, and action_seq; image_seq is the observation sequence; language_instruction is generated by an instruction template library or a large language model based on trajectory fragments and physical ground truth; action_seq is a control sequence time-aligned with the above observation sequence; or

[0170] The downstream model is a VLM model, and the training data consists of paired samples containing synthetic perceptual data and their corresponding semantically rich text.

[0171] Or may also include:

[0172] It also includes a training module:

[0173] The downstream model is trained and the 3DGS mapping and mechanical constitutive parameters are jointly optimized via differentiable links. Using loss function ;in Mechanical constitutive parameters are fed back to the physics simulator via a three-dimensional Gaussian neural deformation field mapping. In the formula Switch by mode: When the downstream model is a perception model Cross-entropy loss function for the perception task When the downstream model is a world model The regression loss function for world model state prediction When the downstream model is a VLA model For action regression loss function Language-conditional alignment loss function The sum; when the downstream model is a VLM model. Cross-entropy loss function for text generation or visual question answering ;

[0174] Losses due to physical compliance include: (i) subsidence site constraints (ii) Shear field constraints (iii) Energy conservation constraint (iv) Momentum conservation constraint (v) Multi-material boundary continuity constraints When the downstream model is a world model, Additional constraints are included regarding the physical consistency of trajectory-level terrain evolution to require that the world model predictions... , , Consistent with the output of the physics simulator;

[0175] The next-moment terrain physical state output by the world model; The next physical truth value output by the physics simulator;

[0176] Or may also include:

[0177] The driving log was encrypted and uploaded to the cloud. The cloud then used H3 hex spatial indexing and material type to jointly perform a post-calibration of constitutive parameters. Cross-platform standardization To achieve heterogeneous platform reuse, a DP differential privacy mechanism is provided; based on the failure logs of the deployed model in real-world scenarios, various adversarial perturbations are performed by traffic splitting according to trigger type, so as to generate adversarial synthetic samples in a preset ratio to strengthen the model's weaknesses.

[0178] In module 2 These are parameters configured in the virtual terrain material model of the physics simulator; It can be any one or a combination of the following: pressure-sinking model parameters, shear stress-shear displacement model parameters, tire model parameters, rigid contact friction parameters, and tribological parameters.

[0179] This module 2 may also be used for:

[0180] Based on the number of driving logs, the differentiability of the physics simulator, and the modal number of the objective function, an optimization algorithm is selected. Using the wheel-ground dynamics response in the triplicate as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physics simulator are calibrated in reverse. ;

[0181] When the physical simulator is differentiable, the optimization algorithm uses conjugate gradient descent; when the number of driving logs is less than a threshold, the optimization algorithm uses Bayesian optimization; when the number of modes of the objective function is greater than a threshold, the optimization algorithm uses an evolutionary algorithm in conjunction with the physical simulator.

[0182] like Figure 8 As shown, in another embodiment of the present invention, a first electronic device A is also proposed, which includes the aforementioned off-road autonomous driving training data synthesis device B based on reverse calibration and 3DGS.

[0183] like Figure 9 As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D via a wired or wireless information transmission scheme. The data acquisition device C is used to collect and acquire the virtual vehicle model and the posterior estimate. The information display device D is used to display the off-road autonomous driving training data obtained by the present invention, including virtual terrain and control commands / expert action sequences.

[0184] Information display device D can process and organize the data output by the first electronic device A based on an information display mechanism to improve the readability of the data. This information display mechanism can be manually preset, for example, visualizing the data output by the first electronic device A. It can present the user with specified key information, such as news updates or system information, based on user-defined display parameters and / or attributes. Display parameters could be, for example, the data range to be displayed, and display attributes could be, for example, the font, color, or whether scrolling is enabled. This allows the user to access this information more quickly without having to navigate to secondary pages or scroll through pages, saving user effort. Alternatively, this information display mechanism can be an artificial intelligence (AI) display model, which can learn the user's key information interests based on previous usage habits, such as viewing time, number of clicks, and number of edits, and then automatically present the user with rich and necessary key information.

[0185] The present invention also provides a computer program product, which includes a computer program that can be stored on a readable storage medium. When the computer program is executed by a processor, the computer is able to execute the off-road autonomous driving training data synthesis method based on reverse calibration and 3DGS provided by the above methods.

[0186] In another embodiment, the present invention also proposes a storage medium VIII for storing a computer program that performs the aforementioned method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS. It should be understood that the storage medium in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0187] Figure 10 A schematic block diagram of a second electronic device 1000 that can be used to implement embodiments of the present invention is shown. The second electronic device 1000 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0188] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from storage medium VIII into random access memory (RAM) III. The RAM III may also store various programs and data required for the operation of the device 1000. The computing unit I, ROM II, and RAM III are interconnected via bus IV. An input / output (I / O) interface V is also connected to bus IV.

[0189] Multiple components in the second electronic device 1000 are connected to I / O interface V, including: input unit VI, such as a keyboard, mouse, etc.; output unit VII, such as various types of displays, speakers, etc.; storage medium VIII, such as a disk, optical disk, etc.; and communication unit IX, such as a network card, modem, wireless transceiver, etc. Communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0190] The computing unit I can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit I performs the various methods and processes described above, such as method steps S1-S4. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to perform methods by any other suitable means (e.g., by means of firmware).

[0191] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A method for synthesizing off-road autonomous driving training data based on inverse calibration and 3DGS, characterized in that, include: Step 1: Extract the triplet consisting of sensor stream, control command stream, and wheel-to-ground dynamics response stream from the vehicle's off-road driving log; Pre-classify the ternary group according to the type of ground material; Step 2: Using the wheel-ground dynamic response in the ternary set as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physical simulator are calibrated in reverse. , to obtain the posterior estimate The physical simulator estimates a posteriori. The wheel-to-ground contact displacement field is obtained by performing forward integration of the parameters in the simulator. ; Step 3, determine the contact displacement field of the wheel to the ground. Differential mapping is applied to the 3DGS Gaussian point cloud in the three-dimensional Gaussian neural deformation field to synthesize synthetic sensing data after a tire passes over the ground. This synthetic sensing data includes: ground RGB image, depth map, and normal map. Step 4: Match the virtual vehicle model with the real vehicle and the posterior estimate. The virtual terrain and control commands / expert action sequences are input into the physics simulator to obtain data including subsidence depth. Shear stress Semantic masking, passability, posterior estimation The pixel-wise aligned physical ground truth; the posterior estimate of each pixel in the synthetic perceptual data. Corresponding to this physical truth value, it serves as training data for off-road autonomous driving.

2. The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS as described in claim 1, characterized in that, Also includes: Step 5: The output adapter organizes the autonomous driving training data into the training data required by the downstream model according to the data content and format required for training the downstream model. The downstream model is a perception model, and the training data is... Triples; or The downstream model is a world model, and the training data consists of the world model executing preset action strategies on virtual terrain. Step trajectory triple sequence , Let t represent the vehicle state and terrain physical state at time t. For the control command at time t The resulting vehicle state and terrain physical state at time t+1; or The downstream model is a VLA model, and the training data consists of triples composed of image_seq, language_instruction, and action_seq; image_seq is the observation sequence; language_instruction is generated by an instruction template library or a large language model based on trajectory fragments and physical ground truth; action_seq is a control sequence time-aligned with the above observation sequence; or The downstream model is a VLM model, and the training data consists of paired samples containing synthetic perceptual data and their corresponding semantically rich text.

3. The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS as described in claim 1, characterized in that, It also includes training steps: The downstream model is trained and the 3DGS mapping and mechanical constitutive parameters are jointly optimized via differentiable links. Using loss function ;in Mechanical constitutive parameters are fed back to the physics simulator via a three-dimensional Gaussian neural deformation field mapping. In the formula Switch by mode: When the downstream model is a perception model Cross-entropy loss function for the perception task When the downstream model is a world model The regression loss function for world model state prediction When the downstream model is a VLA model For action regression loss function Language-conditional alignment loss function The sum; when the downstream model is a VLM model. Cross-entropy loss function for text generation or visual question answering ; Losses due to physical compliance include: (i) subsidence site constraints (ii) Shear field constraints (iii) Energy conservation constraint (iv) Momentum conservation constraint (v) Multi-material boundary continuity constraints When the downstream model is a world model, Additional constraints are included regarding the physical consistency of trajectory-level terrain evolution to require that the world model predictions... , , Consistent with the output of the physics simulator; The next-moment terrain physical state output by the world model; This is the true physical value for the next moment, output by the physics simulator.

4. The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS as described in claim 1, characterized in that, The driving log was encrypted and uploaded to the cloud. The cloud then used H3 hex spatial indexing and material type blocks to jointly perform a post-calibration of constitutive parameters. Cross-platform standardization Enables heterogeneous platform reuse, with supporting DP differential privacy mechanism; Based on the failure logs of the deployed model in real-world scenarios, various adversarial perturbations are performed by splitting the traffic according to the trigger type to generate adversarial synthetic samples in a preset ratio to reinforce the model's weaknesses.

5. The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS as described in claim 1, wherein in step 2... These are parameters configured in the virtual terrain material model of the physics simulator; It can be any one or a combination of the following: pressure-sinking model parameters, shear stress-shear displacement model parameters, tire model parameters, rigid contact friction parameters, and tribological parameters.

6. The method for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS as described in claim 1, characterized in that, Step 2 includes: Based on the number of driving logs, the differentiability of the physics simulator, and the modal number of the objective function, an optimization algorithm is selected. Using the wheel-ground dynamics response in the triplicate as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physics simulator are calibrated in reverse. ; When the physical simulator is differentiable, the optimization algorithm uses conjugate gradient descent; when the number of driving logs is less than a threshold, the optimization algorithm uses Bayesian optimization; when the number of modes of the objective function is greater than a threshold, the optimization algorithm uses an evolutionary algorithm in conjunction with the physical simulator.

7. A device for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS, characterized in that, include: Module 1 extracts a triplet from the vehicle's off-road driving log, consisting of sensor stream, control command stream, and wheel-to-ground dynamics response stream. Pre-classify the ternary group according to the type of ground material; Module 2, using the wheel-ground dynamics response in the ternary set as the optimization objective, reverse-calibrates the mechanical constitutive parameters of the virtual terrain material in the physical simulator. , to obtain the posterior estimate The physical simulator estimates a posteriori. The wheel-to-ground contact displacement field is obtained by performing forward integration of the parameters in the simulator. ; Module 3 will determine the contact displacement field of the wheel. Differential mapping is applied to the 3DGS Gaussian point cloud in the three-dimensional Gaussian neural deformation field to synthesize synthetic sensing data after a tire passes over the ground. This synthetic sensing data includes: ground RGB image, depth map, and normal map. Module 4 will match the virtual vehicle model with the real vehicle and the posterior estimate. The virtual terrain and control commands / expert action sequences are input into the physics simulator to obtain data including subsidence depth. Shear stress Semantic masking, passability, posterior estimation The pixel-wise aligned physical ground truth; the posterior estimate of each pixel in the synthetic perceptual data. Corresponding to this physical truth value, it serves as training data for off-road autonomous driving; Or may also include: Module 5, the output adapter, organizes the autonomous driving training data into the training data required by the downstream model based on the data content and format required for training the downstream model. The downstream model is a perception model, and the training data is... Triples; or The downstream model is a world model, and the training data consists of the world model executing preset action strategies on virtual terrain. Step trajectory triple sequence , Let t represent the vehicle state and terrain physical state at time t. For the control command at time t The resulting vehicle state and terrain physical state at time t+1; or The downstream model is a VLA model, and the training data consists of triples composed of image_seq, language_instruction, and action_seq; image_seq is the observation sequence; language_instruction is generated by an instruction template library or a large language model based on trajectory fragments and physical ground truth; action_seq is a control sequence time-aligned with the above observation sequence; or The downstream model is a VLM model, and the training data consists of paired samples containing synthetic perceptual data and their corresponding semantically rich text. Or may also include: It also includes a training module: The downstream model is trained and the 3DGS mapping and mechanical constitutive parameters are jointly optimized via differentiable links. Using loss function ;in Mechanical constitutive parameters are fed back to the physics simulator via a three-dimensional Gaussian neural deformation field mapping. In the formula Switch by mode: When the downstream model is a perception model Cross-entropy loss function for the perception task When the downstream model is a world model The regression loss function for world model state prediction When the downstream model is a VLA model For action regression loss function Language-conditional alignment loss function The sum; when the downstream model is a VLM model. Cross-entropy loss function for text generation or visual question answering ; Losses due to physical compliance include: (i) subsidence site constraints (ii) Shear field constraints (iii) Energy conservation constraint (iv) Momentum conservation constraint (v) Multi-material boundary continuity constraints When the downstream model is a world model, Additional constraints are included regarding the physical consistency of trajectory-level terrain evolution to require that the world model predictions... , , Consistent with the output of the physics simulator; The next-moment terrain physical state output by the world model; The next physical truth value output by the physics simulator; Or may also include: The driving log was encrypted and uploaded to the cloud. The cloud then used H3 hex spatial indexing and material type blocks to jointly perform a post-calibration of constitutive parameters. Cross-platform standardization To achieve heterogeneous platform reuse, a DP differential privacy mechanism is provided; based on the failure logs of the deployed model in real-world scenarios, various adversarial perturbations are performed by traffic splitting according to trigger type, so as to generate adversarial synthetic samples in a preset ratio to strengthen the model's weaknesses. In module 2 These are parameters configured in the virtual terrain material model of the physics simulator; It can be any one or a combination of the following: pressure-sinking model parameters, shear stress-shear displacement model parameters, tire model parameters, rigid contact friction parameters, and tribological parameters. This module 2 may also be used for: Based on the number of driving logs, the differentiability of the physics simulator, and the modal number of the objective function, an optimization algorithm is selected. Using the wheel-ground dynamics response in the triplicate as the optimization objective, the mechanical constitutive parameters of the virtual terrain material in the physics simulator are calibrated in reverse. ; When the physical simulator is differentiable, the optimization algorithm uses conjugate gradient descent; when the number of driving logs is less than a threshold, the optimization algorithm uses Bayesian optimization; when the number of modes of the objective function is greater than a threshold, the optimization algorithm uses an evolutionary algorithm in conjunction with the physical simulator.

8. An electronic device, characterized in that, The device for synthesizing off-road autonomous driving training data based on reverse calibration and 3DGS as described in claim 7 includes an electronic device or an information display device connected to it. The information display device is used to display the off-road autonomous driving training data using user-set display parameters, attributes, or through an artificial intelligence model.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for synthesizing off-road autonomous driving training data based on inverse calibration and 3DGS as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for synthesizing off-road autonomous driving training data based on inverse calibration and 3DGS as described in any of claims 1-6.

Citation Information

Patent Citations

  • Three dimensional gaussian splatting initialization based on trained neural radiance field representations

    US20240355047A1

  • Method and apparatus for dynamic gaussian splatting

    US20250363723A1