Virtual scene generation and domain adaptation migration method and device

CN122550865APending Publication Date: 2026-08-11BEIJING DUANDIAN INTELLIGENT MANUFACTURING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

这些因素共同导致在理想仿真环境中训练出的智能体策略,在迁移至实体机器人(如机器狗)时,往往会出现严重的性能退化甚至任务失败,难以达到预期的成功率

Benefits of technology

[0014]综上所述,本发明提供一种虚拟场景生成与域自适应迁移方法及装置,该方法包括:获取真实场地的激光扫描原始点云数据,并执行基于近邻距离计算的统计滤波与体素下采样操作,将处理后的目标点云输入表面重建算法生成封闭的三维网格模型,并实施多视图纹理透视投影,从而生成带有几何拓扑结构与高分辨率表面纹理的数字孪生场景模型;基于所述带有几何拓扑结构与高分辨率表面纹理的数字孪生场景模型中所划分的场景语义类别,将预先测量的物理标定参数与物理渲染材质的反射率参数映射至所述数字孪生场景模型的三维网格面片上,从而构建出融合表面纹理与真实物理交互特性的高保真仿真环境;通过所述融合表面纹理与真实物理交互特性的高保真仿真环境,对视觉感知参数与所述物理标定参数施加预设的随机扰动进行域随机化处理,并结合真实场景图像数据,进行对抗性特征对齐,从而生成域自适应训练场景;将待训练的具身智能体控制模型接入所述域自适应训练场景中,在经历动态气象与光照渲染条件的环境中进行持续的状态交互与动作反馈以完成强化学习,将完成强化学习的具身智能体控制模型部署至具身智能体,以在真实物理场地执行自主导航与动作交互任务。本申请的技术方案通过构建融合几何纹理与物理特性的高保真仿真环境,经过域随机化应与对抗对齐弥合虚实域差异,降低实地训练成本,提升具身智能体跨域泛化能力,实现虚拟训练向真实场地自主导航与动作交互的高效迁移。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550865A_ABST
    Figure CN122550865A_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for virtual scene generation and domain adaptive migration. The method includes: constructing a digital twin scene model based on real-world laser point clouds; mapping physical calibration and material reflectivity parameters to mesh patches according to semantic categories to establish a high-fidelity simulation environment that integrates texture and physical properties; subjecting the environment to domain randomization perturbations of visual and physical parameters, and performing adversarial feature alignment with real images to construct a domain adaptive training scene; placing an embodied intelligent agent control model within this environment, completing reinforcement learning under dynamic weather and lighting conditions, and then deploying it to the real-world site to achieve autonomous navigation and action interaction. The technical solution of this application, by constructing a high-fidelity digital twin scene and introducing a dual domain adaptive strategy, reduces data acquisition and annotation costs, effectively bridging the gap between virtual and reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of embodied intelligence, and more particularly to a method and apparatus for virtual scene generation and domain adaptive migration. Background Technology

[0002] With the rapid development of artificial intelligence technology, embodied intelligence has become a core research direction in the fields of robotics and intelligent control. In the research and development of embodied intelligence, training efficient and robust control strategies and visual models typically requires massive amounts of high-dimensional and diverse interactive data. However, conducting large-scale data acquisition directly in real physical environments faces severe challenges: on the one hand, the economic and time costs of on-site collection and manual annotation are extremely high; on the other hand, in complex, extreme, or dangerous scenarios such as industrial facilities and disaster sites, on-site operations are not only unsafe but also often difficult to implement due to physical limitations. Therefore, building simulation environments based on high-fidelity engines and conducting virtual training, followed by transferring the learned strategies to the real world (Sim2Real), has become an important technical path for the academic and industrial communities to achieve large-scale deployment of embodied intelligence.

[0003] However, existing simulation systems are generally constrained by the reality gap in practical applications, meaning that there are significant differences in distribution between the simulation environment and the real physical world across multiple dimensions. These differences manifest specifically as insufficient rendering accuracy at the visual level, large deviations between the simulation of physical properties (such as friction, mass distribution, and elastic coefficient) and actual working conditions, and distortion caused by noise modeling of sensors (such as LiDAR and depth cameras). These factors collectively lead to severe performance degradation or even task failure when the strategies of intelligent agents trained in ideal simulation environments are transferred to physical robots (such as robotic dogs), making it difficult to achieve the expected success rate. How to effectively reduce the transfer error between simulation and reality while ensuring high-fidelity generation of simulation scenes, and how to achieve high-success-rate zero-sample transfer in cross-domain environments, and how to properly solve these problems, have become urgent issues for the industry to address. Summary of the Invention

[0004] This invention provides a method and apparatus for virtual scene generation and domain adaptive migration, which is used to construct a high-fidelity digital twin scene and introduce a dual domain adaptive strategy, thereby reducing the cost of data acquisition and annotation and effectively bridging the gap between the virtual and the real world.

[0005] According to a first aspect of the present invention, a method for virtual scene generation and domain adaptive migration is provided, the method comprising: The system acquires raw point cloud data from laser scanning of the real site, performs statistical filtering and voxel downsampling based on nearest neighbor distance calculation, inputs the processed target point cloud into a surface reconstruction algorithm to generate a closed 3D mesh model, and implements multi-view texture perspective projection to generate a digital twin scene model with geometric topology and high-resolution surface texture. Based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, the pre-measured physical calibration parameters and the reflectivity parameters of the physical rendering material are mapped onto the three-dimensional mesh surface of the digital twin scene model, thereby constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics. Through the high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics, the visual perception parameters and the physical calibration parameters are subjected to preset random perturbation for domain randomization processing, and adversarial feature alignment is performed in combination with real scene image data to generate a domain adaptive training scene. The embodied intelligent agent control model to be trained is connected to the domain adaptive training scenario. It performs continuous state interaction and action feedback in an environment with dynamic weather and lighting rendering conditions to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

[0006] In one embodiment, generating a digital twin scene model with geometric topology and high-resolution surface texture includes: For the original laser scanning point cloud data, outliers are removed by using a set number of nearest neighbor points and a standard deviation threshold, and the overall point cloud density is unified by voxel downsampling to output the target point cloud; The three-dimensional spatial surface topology of the target point cloud is calculated using the Poisson surface reconstruction algorithm to reconstruct the closed three-dimensional mesh model, and the two-dimensional plane coordinate line unfolding is performed on the closed three-dimensional mesh model using an angle-based parameterization algorithm. By combining the visible light image data of the real site, the multi-view texture perspective projection is performed on the 3D mesh model unfolded by the automatic coordinate lines to generate a diffuse reflection map, and different levels of surface reduction calculation operations are automatically performed according to the camera observation distance threshold to generate a digital twin scene model with multiple levels of detail.

[0007] In one embodiment, constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics includes: The digital twin scene model with geometric topology and high-resolution surface texture is used as the data input. It is fed into a pre-trained semantic segmentation neural network for deep spatial feature extraction and patch classification, and outputs a set of mesh patches containing multiple scene semantic categories. Through on-site mechanical testing, the static friction coefficient, dynamic friction coefficient, mass distribution properties, and elastic recovery coefficient were obtained, and the bidirectional reflection distribution of various real material surfaces was extracted to construct the physical calibration parameters and the reflectivity parameters of the physically rendered material. The physical calibration parameters and the reflectivity parameters of the physically rendered material are injected into the material calculation assets of the physical simulation engine according to the specific semantic classification tags to which the mesh patch set belongs.

[0008] In one embodiment, the generation and access of the dynamic weather and lighting rendering conditions includes: Meteorological visual effects are configured within the simulation environment that integrates surface texture and real physical interaction characteristics. Dynamic lighting conditions, including the luminous intensity, absolute color temperature, and incident elevation angle of a global parallel light source, are superimposed. The meteorological visual effects include large-scale cloud density values ​​and local precipitation particle emissivity. The camera's exposure compensation and color grading correction parameters are dynamically adjusted in real time. Nonlinear photometric fusion calculations are performed on the meteorological visual effects and dynamic lighting conditions. The dynamic meteorological and lighting rendering conditions are output for the embodied intelligent agent control model to be trained to perform virtual visual signal perception.

[0009] In one embodiment, at the initial time point of reinforcement learning of the embodied intelligent agent control model to be trained, the surface texture color, global ambient light intensity, and the physical calibration parameters are uniformly and randomly perturbed according to a preset interval ratio. The gradient inversion layer is connected in series between the output of the visual encoder network used to extract high-dimensional feature vectors and the input of the environment domain discriminator responsible for binary classification, thus constructing an adversarial deep neural network learning architecture for cross-domain feature extraction. The environment domain discriminator performs binary classification to determine the authenticity of the image source between the virtual observation image output in real time by the high-fidelity simulation environment rendering engine and the unlabeled image data of the real scene. The negative gradient signal generated by the discrimination error is directly backpropagated to the visual encoder through the gradient inversion layer. In the continuous minimax adversarial game, the extracted environmental visual features approximate the consistent state of the data distribution between the domain adaptive training scene and the real application scene.

[0010] In one embodiment, it also includes: For the real visible light camera and depth detection sensor mounted on the physical robot platform, perform physical offline error calibration, extract static baseline data and dynamic shot characteristics under multiple working conditions, and fit heteroscedastic Gaussian mathematical equation and distance-related polynomial error compensation equation respectively, thereby generating a parameterized sensor noise model. The parameterized sensor noise model is used as an independent low-level noise intervention term and is superimposed on the virtual sensor output of the high-fidelity simulation environment in real time during each rendering cycle. This allows the observation signal carrying hardware-level noise factors to be combined with the visual feedback data after uniform probability sampling perturbation to form a matrix, which serves as the state input reference for the visual encoder network.

[0011] According to a second aspect of the present invention, a virtual scene generation and domain adaptive migration apparatus is provided, comprising: The acquisition module is used to acquire the original point cloud data of the real site by laser scanning, and perform statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation. The processed target point cloud is input into the surface reconstruction algorithm to generate a closed three-dimensional mesh model, and multi-view texture perspective projection is implemented to generate a digital twin scene model with geometric topology and high-resolution surface texture. The construction module is used to map pre-measured physical calibration parameters and reflectivity parameters of physically rendered materials onto the three-dimensional mesh surface of the digital twin scene model based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, thereby constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics. The generation module is used to apply a preset random perturbation to the visual perception parameters and the physical calibration parameters through the high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics, perform domain randomization processing, and combine it with real scene image data to perform adversarial feature alignment, thereby generating a domain adaptive training scene. The deployment module is used to connect the embodied intelligent agent control model to be trained into the domain adaptive training scenario, and to conduct continuous state interaction and action feedback in an environment experiencing dynamic weather and lighting rendering conditions to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

[0012] According to a third aspect of the present invention, an electronic device is provided, comprising: a communication interface, a processor, and a memory; The memory is used to store program instructions, which, when executed by the processor that is connected to the memory via the communication interface, implement any of the above-described virtual scene generation and domain adaptive migration methods.

[0013] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, which, when executed by a computer (e.g., a processor in the computer), implement any of the above-described virtual scene generation and domain adaptive migration methods.

[0014] In summary, this invention provides a method and apparatus for virtual scene generation and domain adaptive migration. The method includes: acquiring raw laser-scanned point cloud data of a real site, performing statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation, inputting the processed target point cloud into a surface reconstruction algorithm to generate a closed 3D mesh model, and implementing multi-view texture perspective projection to generate a digital twin scene model with geometric topology and high-resolution surface texture; based on the scene semantic categories defined in the digital twin scene model with geometric topology and high-resolution surface texture, mapping pre-measured physical calibration parameters and reflectivity parameters of physically rendered materials onto the 3D mesh patches of the digital twin scene model. This constructs a high-fidelity simulation environment that integrates surface texture and realistic physical interaction characteristics. Within this environment, pre-defined random perturbations are applied to visual perception parameters and physical calibration parameters for domain randomization. Combined with real-world scene image data, adversarial feature alignment is performed to generate a domain-adaptive training scenario. The embodied agent control model to be trained is then integrated into this domain-adaptive training scenario. Continuous state interaction and action feedback are conducted in an environment experiencing dynamic weather and lighting conditions to complete reinforcement learning. The reinforced agent control model is then deployed to the embodied agent to perform autonomous navigation and action interaction tasks in a real physical environment. This technical solution, by constructing a high-fidelity simulation environment that integrates geometric texture and physical characteristics, bridges the gap between virtual and real domains through domain randomization and adversarial alignment, reduces on-site training costs, enhances the cross-domain generalization ability of the embodied agent, and achieves efficient transfer from virtual training to autonomous navigation and action interaction in real-world environments.

[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and drawings.

[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 A flowchart of a virtual scene generation and domain adaptive migration method provided as an embodiment of the present invention; Figure 2 A flowchart of another virtual scene generation and domain adaptive migration method provided in an embodiment of the present invention; Figure 3 A flowchart of another virtual scene generation and domain adaptive migration method provided in the embodiments of the present invention; Figure 4 A flowchart of another virtual scene generation and domain adaptive migration method provided in the embodiments of the present invention; Figure 5 A flowchart of another virtual scene generation and domain adaptive migration method provided in the embodiments of the present invention; Figure 6 A flowchart of another virtual scene generation and domain adaptive migration method provided in the embodiments of the present invention; Figure 7 A structural diagram of a virtual scene generation and domain adaptive migration device provided in an embodiment of the present invention; Figure 8 A structural diagram of an electronic device provided as an embodiment of the present invention; Figure 9 A system framework diagram for virtual scene generation and domain adaptive migration provided as an embodiment of the present invention; Figure 10 A schematic diagram illustrating the point cloud to digital twin scene construction process provided for embodiments of the present invention; Figure 11 A schematic diagram of the environment rendering and physical calibration process provided for embodiments of the present invention; Figure 12 A schematic diagram of a domain adaptive training architecture provided for an embodiment of the present invention; Figure 13 A flowchart illustrating Embodiment 1 provided for the purposes of this invention; Figure 14 This is a schematic diagram of a high-fidelity virtual scene generation and Sim2Real domain adaptive migration architecture provided for embodiments of the present invention. Detailed Implementation

[0019] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0020] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0021] like Figure 1 As shown, the present invention provides a method for virtual scene generation and domain adaptive migration, which includes: In step S11, the original laser scanning point cloud data of the real site is acquired, and statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation are performed. The processed target point cloud is input into the surface reconstruction algorithm to generate a closed three-dimensional mesh model, and multi-view texture perspective projection is implemented to generate a digital twin scene model with geometric topology and high-resolution surface texture. In step S12, based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, the pre-measured physical calibration parameters and the reflectivity parameters of the physical rendering material are mapped onto the three-dimensional mesh surface of the digital twin scene model, thereby constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics. In step S13, through the high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics, a preset random perturbation is applied to the visual perception parameters and the physical calibration parameters to perform domain randomization processing, and adversarial feature alignment is performed in combination with real scene image data to generate a domain adaptive training scene. In step S14, the embodied intelligent agent control model to be trained is connected to the domain adaptive training scenario. In an environment experiencing dynamic weather and lighting rendering conditions, continuous state interaction and action feedback are carried out to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

[0022] In one embodiment, at the forefront of the current global convergence of artificial intelligence and robotics, Embodied Artificial Intelligence (EAI) is undergoing a paradigm shift from task-specific rule-based control to generalized control based on large-scale Vision-Language-Action (VLA) models. However, Deep Reinforcement Learning and VLA models, with their massive number of parameters, impose stringent requirements on the scale, quality, and scenario diversity of training data. In the real physical world, directly collecting massive amounts of diverse, high-quality interactive data through teleoperation or pre-programmed trajectories not only incurs extremely high costs in terms of capital, manpower, and time, but also faces serious safety hazards and real-world challenges in data annotation in edge scenarios such as industrial facility inspection, disaster site search and rescue, and even complex industrial waste sorting and material recycling facilities. Data scarcity has become a core bottleneck restricting embodied intelligence control strategies from the laboratory to the open physical world.

[0023] To overcome the data acquisition bottleneck in this physical dimension, shifting the training environment to low-marginal-cost, high-parallelism computer simulation platforms (such as Unreal Engine, Isaac Sim, and Gazebo), i.e., building a Sim2Real pipeline, has become a strategic consensus in academia and industry. Simulation platforms can generate hundreds of millions of interaction steps in a short time and automatically provide perfect pixel-level semantic segmentation, depth, and pose annotation. However, a seemingly insurmountable reality gap always exists between existing simulation technologies and the real physical world. This reality gap is not a single-dimensional error, but a complex system bias interwoven with differences in the visual and dynamic domains. At the visual level, the lighting, textures, reflectivity, and inherent sensor noise in simulation renderings deviate significantly from the real world. At the dynamic level, due to the simplification and distortion of physical interaction characteristics such as friction, contact models, mass distribution, and elastic collisions by physics engines, the optimal control strategies learned by embodied agents in simulations often exploit loopholes unique to the simulator.

[0024] When a robot completes policy training in a purely simulated environment without effective domain adaptation, it often encounters catastrophic performance degradation when transferring to the unknown real physical world with zero samples, sometimes even leading to task failure and equipment damage. Faced with this long-tailed and imbalanced Sim2Real regression problem, early industry explorations focused on simple domain randomization techniques. This involved large-scale random sampling and perturbation of the visual and physical parameters of the simulation model, attempting to include the real-world parameter distribution within a broad randomization range. However, simple domain randomization often results in overly conservative and suboptimal trained policies.

[0025] During real-world site scanning, the acquired raw laser point clouds are often accompanied by significant environmental noise and discrete, stray points. This solution introduces a statistical filtering algorithm based on K-nearest neighbor distance. By calculating the average distance from each point to its K nearest neighbors and setting a standard deviation threshold, outlier noise points deviating from the main distribution can be mathematically and accurately removed. To address the issue of uneven point cloud density in large-scale scenes, a fixed-size 3D voxel downsampling technique is employed. This not only significantly reduces the computational overhead of subsequent calculations but also ensures the spatial uniformity of the point cloud distribution.

[0026] This invention designs a dual-domain adaptive training architecture that combines domain randomization (DR) with adversarial feature alignment, as illustrated in the attached diagram. Figure 12 As shown.

[0027] From a logical depth perspective, these two mechanisms intervene in the gradient descent trajectory of the policy network from two distinct dimensions: the breadth of data used in model training and the depth of features. Firstly, parameter domain randomization, as a low-level data distribution enhancement method, forces a uniform probability perturbation of parameters within a preset range to the parameters in the high-fidelity simulation environment at each round's start node during the training of the embodied agent PPO (Proximal Policy Optimization) algorithm. This includes not only perturbations of visual appearance parameters (such as texture color ±15%, illumination intensity ±30%), but also random offsets of dynamic parameters (such as friction coefficient ±20%, object mass ±10%). Simultaneously, combined with UE5's built-in Volumetric Cloud plugin and the Niagara precipitation particle system, a seamless, continuous interpolation transition between seven weather conditions and three lighting conditions is simulated. Dynamic environment rendering and physical calibration are shown in the attached figure. Figure 11As shown. The mathematical essence of this high-intensity randomization perturbation is to inject high-variance noise into the joint probability distribution of the training data, forcing the convolutional kernels and multilayer perceptrons of the policy network to ignore non-essential feature edges that are easily subject to minor changes in the real environment, thereby converging the attention mechanism and anchoring it to the core geometric structures and topological logic that have cross-domain invariance to the environment.

[0028] Secondly, to address the long-tailed feature bias that randomization cannot completely eliminate, a feature alignment mechanism based on a DANN (Domain-Adversarial Neural Network) architecture is introduced. At the output layer of the visual feature encoder network (Encoder), a domain discriminator network is cascaded into the system, with a gradient reversal layer (GRL, weights progressively scheduled λ=0.1–1.0) connected between them. During the forward propagation phase of training, the domain discriminator makes every effort to perform binary classification of the input feature sources, distinguishing between real-time rendered images from a high-fidelity simulation environment and a small number of unlabeled images from the real-world environment that are additional input from the system. However, during backpropagation to calculate the gradient, the GRL directly inverts the sign of the classification error gradient generated by the domain discriminator and then propagates this negative gradient signal to the visual feature encoder. This mechanism causes the visual encoder to maximize the classification error of the domain discriminator rather than minimize it when updating its network weights. Through this continuous game of interaction, the visual encoder is gradually forced to evolve, and the extracted latent feature vectors reach a statistically significant state of high mixing and indistinguishability between the simulation domain and the real domain. This deep feature-level alignment endows the embodied agent with zero-shot generalization penetration capabilities in unseen real-world physical scenes.

[0029] In the specific implementation process, a vehicle-mounted or handheld LiDAR is used to perform a global scan of the target site to acquire raw point cloud data. The system performs statistical filtering based on nearest neighbor distance calculation (setting K=50, standard deviation threshold 1.0) to remove outliers and uses voxel downsampling (e.g., voxel size 5cm) to unify the data density. The processed target point cloud is input into a surface reconstruction algorithm (such as the Screened Poisson algorithm with an octree depth of 10–12 levels) to generate a closed 3D mesh model. Through multi-view texture perspective projection, the synchronously acquired 4K resolution RGB image is mapped onto this mesh, thereby generating a digital twin scene model with geometric topology and high-resolution surface texture. A schematic diagram of the point cloud to digital twin scene construction process is attached. Figure 10As shown. To ensure real-time interactive efficiency, this model automatically generates a 4-level LOD (Level of Detail) after being imported into the UE5 engine, ensuring that the rendering frame rate for large scenes is no less than 60fps.

[0030] The system employs a high-fidelity simulation environment to apply preset random perturbations to visual perception parameters (e.g., texture color ±15%, illumination intensity ±30%) and physical calibration parameters (e.g., friction coefficient ±20%) for domain randomization. Simultaneously, an adversarial feature alignment strategy is introduced, using a DANN architecture with a gradient inversion layer following the visual encoder. Adversarial training is conducted using a small amount of real-world image data (approximately 500 images), ensuring that the features extracted by the encoder are indistinguishable between the simulation and real-world domains. Furthermore, the system accurately simulates the noise characteristics of real sensors (e.g., IMU random walk, depth quantization error) to ensure that the domain-adaptive training scenario approximates the real environment as closely as possible in terms of data distribution.

[0031] In the training phase, the embodied agent control model to be trained (such as a VLA large model or reinforcement learning strategy) is integrated into the aforementioned domain-adaptive training scenario. This domain-adaptive training scenario utilizes Volumetric Cloud and the Niagara particle system to achieve real-time continuous transitions between dynamic weather conditions (such as rain, snow, and fog) and lighting rendering conditions (daytime, dusk, and nighttime). In this complex environment experiencing various visual and physical conditions, the embodied agent employs the PPO (Proximal Policy Optimization) algorithm for continuous state interaction and action feedback to complete reinforcement learning.

[0032] The embodied agent control model, having completed reinforcement learning, is quantized and deployed to embodied agents (such as physical robot dogs) using TensorRT. When performing autonomous navigation and action interaction tasks in chemical plant areas or earthquake disaster search and rescue sites, the model, having established adaptation to the real-world gap through domain randomization and feature alignment during training, achieves a zero-shot transfer success rate exceeding 85% in real-world environments. Taking an industrial inspection example, the embodied agent can successfully avoid obstacles such as pipes and storage tanks, and stably complete inspection tasks under complex lighting and slippery road conditions.

[0033] In a preferred embodiment, a high-fidelity virtual scene generation and Sim2Real transfer method for embodied intelligence training is proposed, and its overall system architecture is shown in the attached figure. Figure 9 As shown, the main technical solutions include the following: 1. High-fidelity virtual scene generation: Based on laser scanning data of real sites, high-precision digital twin scenes are built using Unreal Engine 5 (UE5).

[0034] The specific processing procedure is as follows: (a) Point cloud preprocessing: Statistical filtering was performed on the original lidar point cloud (density of about 1000 points / m²) to remove outliers (outlier removal based on K-nearest neighbor distance, K=50, standard deviation threshold 1.0), and the point cloud density was unified by voxel downsampling (voxel size 5cm). (b) Mesh reconstruction: The Screened Poisson Surface Reconstruction algorithm (octree depth 10–12 levels) is used to reconstruct the denoised point cloud into a closed triangular mesh, generating a mesh model with approximately 500,000–2,000,000 faces per scene; (c) Automatic UV unwrapping and texture mapping: The Mesh is automatically unwrapped using the angle-based parametric (ABF++) algorithm, and a 4K resolution diffuse texture is generated by multi-view texture projection using RGB images acquired synchronously by laser scanning. (d) UE5 import and LOD generation: Import the reconstructed FBX format model into UE5 and automatically generate 4-level LOD (Level of Detail). The original number of faces is retained in the foreground and reduced to 10% of the original number in the background to ensure that the real-time rendering frame rate of large scenes is not less than 60fps.

[0035] The system covers more than 20 typical environments, including city streets, woodlands, industrial facilities, and indoor warehouses, ensuring the diversity of training scenarios.

[0036] 2. Environmental Change Simulation: UE5 implements dynamic rendering support for 7 weather conditions (sunny, cloudy, overcast, light rain, heavy rain, fog, snow) and 3 lighting changes (daytime, dusk, nighttime).

[0037] The specific implementation is as follows: the weather system controls cloud density (0.0–1.0) and precipitation particle emissivity through the Volumetric Cloud plugin in UE5, and the rain and snow effects are simulated through the Niagara particle system; changes in illumination are achieved by adjusting the intensity (100,000 lux during the day, 5,000 lux at dusk, and 50 lux at night), color temperature (6,500K during the day, and 3,500K at dusk), and angle (solar altitude angle 10°–80°) of the Directional Light; the fog effect controls visibility (50m–5km) through the fog density (0.02–0.5) and scattering coefficient of the Exponential Height Fog; and the post-process volume dynamically adjusts exposure compensation, color grading, and lens flare parameters.

[0038] All weather and lighting parameters support continuous interpolation transitions during runtime, and can be automatically switched according to a preset schedule or random strategy during training, enabling the embodied agent to be trained under various visual conditions and improving its robustness to environmental changes.

[0039] 3. Physical Parameter Calibration: Accurate measurement and calibration of physical parameters from the real environment (such as friction coefficient, mass distribution, elastic recovery coefficient) and material reflectivity are performed, and then mapped to the UE5 Chaos physics engine through the following calibration process: (a) Friction coefficient calibration: The static friction coefficient μ_s and dynamic friction coefficient μ_d of the real ground material are measured using the inclined plane method and directly assigned values ​​through the StaticFriction and DynamicFriction properties of the Physical Material asset (e.g., μ_s=0.6, μ_d=0.4 for concrete ground, μ_s=0.3, μ_d=0.2 for wet and slippery metal). (b) Mass and inertia calibration: Set the Mass property of the rigid body based on the actual weighing data. The inertia tensor is automatically calculated by the Chaos engine based on the geometry of the colliding body. For objects with non-uniform density, manually set the center of mass offset. (c) Elastic recovery coefficient: The restitution coefficient (0.0–1.0) is measured by a falling ball experiment and mapped to the restitution property of the Physical Material; (d) Material reflectance calibration: The bidirectional reflectance distribution function (BRDF) of the real material in the visible light band was measured using a spectrophotometer, and the measurement data was fitted to the four-channel parameters of BaseColor, Metallic, Roughness and Specular of the UE5 PBR material.

[0040] In terms of semantic segmentation and automatic material assignment, the system adopts a point cloud semantic segmentation network based on PointNet++ (pre-trained on the ScanNet dataset with an mIoU of 72%), which automatically classifies the reconstructed mesh into 12 semantic categories such as ground, walls, metal pipes, and vegetation. Each category corresponds to a pre-labeled Physical Material and PBR material template, realizing a semi-automatic material assignment process (human review is only required for about 5% of the boundary areas), and maximizing the reproduction of the physical interaction characteristics of the real world.

[0041] 4. Introduction of Domain Adaptation Algorithm: A domain adaptation algorithm is introduced during training, employing a dual strategy combining domain randomization and feature alignment. (a) Domain randomization: In simulation training, visual appearance parameters (texture color ±15%, illumination intensity ±30%, camera intrinsic focal length ±5%) and physical parameters (friction coefficient ±20%, object mass ±10%) are uniformly and randomly perturbed. At the same time, Gaussian noise (σ=0.01–0.05) is injected into the virtual RGB sensor and structured noise (simulating the depth quantization error and edge flying points of RealSense D435i) is injected into the virtual depth sensor, so that the strategy is exposed to diverse perception conditions during the training phase. (b) Feature alignment: An adversarial domain adaptation method (DANN architecture) is adopted, and a gradient reversal layer (λ=0.1–1.0 progressive scheduling) is connected after the visual encoder to train the domain discriminator to distinguish between simulated images and a small number of real images (about 500 unlabeled images). Through adversarial training, the features extracted by the encoder are indistinguishable between the simulated domain and the real domain, thereby achieving feature-level domain transfer. (c) Sensor noise modeling: Offline calibration of real sensors, collection of noise samples and fitting of parametric noise models (RGB camera: heteroscedastic Gaussian model; depth sensor: distance-dependent multinomial noise model; IMU: random walk and bias instability parameters obtained from Allan variance analysis) to accurately reproduce the noise characteristics of real sensors in simulation.

[0042] 5. Policy Training and Transfer Validation: Large-scale training of the VLA model or reinforcement learning policy was performed in a constructed high-fidelity virtual scene. The PPO (Proximal Policy Optimization) algorithm was used for training, with 256 simulation instances running in parallel across more than 20 digital twin scenes. A single training session accumulated approximately 100 million interaction steps (approximately 48 hours, 8×A100 GPUs). During training, the domain randomization parameters were resampled at the beginning of each episode according to a preset distribution, and the feature alignment module updated the domain discriminator every 1000 steps. After training, the policy was deployed to the Jetson Orin NX platform of the physical robot dog via TensorRT quantization (FP16), and zero-shot transfer validation was performed in three real-world scenes that were not used in the training. Through the dual domain adaptation method of domain randomization and feature alignment, the success rate of transferring the policy from simulation training to the physical robot dog exceeded 85% (88% for navigation, 86% for obstacle avoidance, and 83% for object interaction), an improvement of approximately 25 percentage points compared to the baseline method without domain adaptation.

[0043] The technical solution in this embodiment has the following advantages compared to the prior art: I. Reduced data acquisition costs: By generating high-fidelity virtual scenes, the high cost and difficulty of data collection and annotation in real-world scenarios are effectively solved, providing a scalable and compliant training data source for VLA large models.

[0044] Second, the transfer success rate has been improved: by introducing a domain adaptive algorithm to calibrate physical parameters, material reflectivity and sensor noise, the success rate of transferring the strategy to the physical robot dog after training in simulation exceeds 85%, effectively overcoming the "reality gap".

[0045] III. Rich Scene and Environment Adaptability: More than 20 digital twin scenes of typical environments have been constructed, supporting dynamic rendering combinations of 7 weather types × 3 lighting types. UE5 Volumetric Cloud, Niagara particle system and PostProcess Volume are used to achieve continuous interpolation transition at runtime, which comprehensively improves the generalization ability of embodied intelligent agents.

[0046] IV. High-precision physical simulation: The friction coefficient, elastic recovery coefficient and BRDF parameters of real materials are calibrated by means of actual measurement such as tilting plane method, falling ball experiment, spectrophotometer, etc., and accurately mapped to UE5 Chaos physics engine and PBR material system. Combined with automatic semantic segmentation based on PointNet++ (mIoU 72%), semi-automatic material assignment is achieved.

[0047] V. Dual Domain Adaptive Strategy: Innovatively combining domain randomization (parameter perturbation) with adversarial feature alignment based on DANN architecture, the former enhances the robustness of the strategy to parameter changes, while the latter achieves feature distribution alignment between the simulated domain and the real domain through gradient inversion layer. The synergistic effect of the two improves the transfer success rate by about 25 percentage points compared to the baseline method.

[0048] Example 1: Industrial Facility Inspection Task In a chemical plant area inspection scenario, the system executes the following steps, and a flowchart of the process is attached. Figure 13 As shown: Step 1 (Point Cloud Acquisition and Scene Reconstruction): A vehicle-mounted LiDAR (such as Velodyne VLP-16) is used to perform a global scan of the chemical plant area, acquiring approximately 200 million raw point clouds. After statistical filtering for noise reduction and voxel downsampling, a triangular mesh (approximately 1.5 million faces) is reconstructed using the Screened Poisson algorithm. After automatic UV unwrapping and multi-view texture projection using ABF++, the mesh is imported into UE5 to generate a Level 4 LOD model.

[0049] Step 2 (Physical Parameters and Material Calibration): The PointNet++ semantic segmentation network automatically classifies the scene mesh into semantic categories such as concrete ground, metal pipes, steel railings, and tank exterior walls. Based on measured data, friction coefficients (concrete μ_s=0.6, metal μ_s=0.4), elastic recovery coefficients, and PBR material parameters are assigned to each category. After manual review of the boundary areas, the scene configuration is completed.

[0050] Step 3 (Dynamic Environment Configuration): Configure rainy weather (Niagara precipitation particle emissivity 500 / s, ground roughness reduced to 0.2) and twilight illumination (Directional Light intensity 5000 lux, color temperature 3500K, solar altitude angle 15°) in the simulation environment. At the same time, inject Gaussian noise (σ=0.03) into the virtual RGB sensor and distance-related noise into the depth sensor.

[0051] Step 4 (Domain Adaptive Training): The PPO algorithm is used for parallel training in this scene and 19 other scenes. Parameters are resampled for each episode with domain randomization. The DANN feature alignment module is trained adversarially using 500 real-world scene images. The policy converges after approximately 100 million interactions.

[0052] Step 5 (Transfer Validation): The trained policy was quantized using TensorRT FP16 and deployed to the robot dog's Jetson Orin NX platform. Inspection tests were conducted in a real-world factory environment during rain. The robot dog successfully avoided obstacles such as pipes, tanks, and forklifts, completing the entire inspection route with a transfer success rate of 88%.

[0053] Example 2: Search and Rescue Scenario at a Disaster Site In earthquake-induced building collapse search and rescue scenarios, the system rapidly reconstructs a digital twin of the collapsed area based on UAV aerial point cloud data. It simulates broken tiles, collapsed beams and columns, and a dusty environment in UE5, configuring nighttime lighting and fog conditions. An embodied agent is trained in this simulated scenario to develop strategies for navigating narrow spaces and searching for trapped personnel. Domain randomization randomly perturbs the tile stacking configuration, lighting angle, and dust concentration, while DANN feature alignment utilizes a small number of real collapse site images for adversarial training. After being transferred to a real robot dog, the navigation success rate in simulated collapsed building sites reaches 85%, demonstrating the generalization ability of this method in extreme environments.

[0054] The high-fidelity virtual scene generation and Sim2Real domain adaptive migration architecture in this embodiment are as follows: Figure 14As shown, it constructs a high-fidelity simulation environment that integrates geometric textures and physical properties. Through domain randomization and adversarial alignment, it bridges the gap between the virtual and real domains, reduces the cost of on-site training, enhances the cross-domain generalization ability of embodied intelligent agents, and realizes the efficient transfer of virtual training to autonomous navigation and action interaction in real-world environments.

[0055] In one embodiment, such as Figure 2 As shown, the generation of a digital twin scene model with geometric topology and high-resolution surface texture includes the following steps S21-S23: In step S21, for the original laser scanning point cloud data, outliers are removed by using a set number of nearest neighbor points and a standard deviation threshold, and the overall point cloud density is unified by voxel downsampling, and the target point cloud is output. In step S22, the three-dimensional spatial surface topology of the target point cloud is calculated using the Poisson surface reconstruction algorithm to reconstruct the closed three-dimensional mesh model, and the two-dimensional plane coordinate line unfolding is performed on the closed three-dimensional mesh model using an angle-based parameterization algorithm. In step S23, combined with the collected visible light image data of the real site, the multi-view texture perspective projection is performed on the three-dimensional mesh model unfolded by the automatic coordinate lines to generate a diffuse reflection map, and different levels of surface reduction calculation operations are automatically performed according to the camera observation distance threshold to generate a digital twin scene model with multiple levels of detail.

[0056] In one embodiment, for raw point cloud data of a real site obtained using LiDAR (such as vehicle-mounted or handheld devices), this embodiment performs a refined preprocessing procedure. The system uses a set number of nearest neighbor points (e.g., K=50) and a specific standard deviation threshold (e.g., 1.0 times the standard deviation) to perform statistical filtering calculations. By evaluating the local distribution characteristics of the point set, outliers and measurement noise are accurately identified and removed, thereby ensuring the smoothness of subsequent surface reconstruction. The filtered point cloud is resampled by voxel downsampling (e.g., setting a voxel size of 5cm), which unifies the overall point cloud density while preserving key geometric features, reduces data redundancy, and outputs a target point cloud that meets the modeling accuracy requirements.

[0057] Based on the target point cloud, the Poisson surface reconstruction algorithm is invoked. By solving the Poisson equation of the gradient field of the indicator function globally, the surface topology in three-dimensional space is calculated, thereby transforming the discrete sampling points into a closed three-dimensional mesh model with continuous manifold characteristics. To achieve high-precision texture mapping, the system further employs an angle-based parameterization algorithm to automatically unfold the coordinate lines of the two-dimensional plane onto the closed three-dimensional mesh model. This minimizes area and angle distortion while ensuring that the unfolded coordinate layout efficiently utilizes the texture space.

[0058] After completing topology reconstruction and coordinate unfolding, multi-view texture perspective projection is performed on the automatically unfolded 3D mesh model by combining real-world visible light image data and camera pose parameters. This process fuses and corrects real-world color information captured from multiple angles, generating a diffuse texture that accurately reflects the surface characteristics of the land and objects. To balance rendering accuracy and computational efficiency in simulation training, different levels of polygon reduction are automatically performed using mesh simplification algorithms such as edge collapse, based on camera observation distance thresholds. This generates a digital twin scene model with multi-level details. This multi-level detail structure allows the embodied agent to dynamically access resources of varying complexity based on visual distance during training, ensuring high fidelity of perceptual information while improving the rendering performance and real-time interaction of the simulation environment.

[0059] In one embodiment, such as Figure 3 As shown, the construction of a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics includes the following steps S31-S33: In step S31, the digital twin scene model with geometric topology and high-resolution surface texture is used as the data input end and fed into the pre-trained semantic segmentation neural network for deep spatial feature extraction and patch classification, and outputs a set of mesh patches containing multiple scene semantic categories. In step S32, static friction coefficient, dynamic friction coefficient, mass distribution properties and elastic recovery coefficient are obtained through on-site mechanical testing, and bidirectional reflection distribution of various real material surfaces is extracted to construct the physical calibration parameters and reflectivity parameters of the physically rendered material. In step S33, the physical calibration parameters and the reflectivity parameters of the physical rendering material are injected into the material calculation assets of the physical simulation engine according to the specific semantic classification tags to which the mesh patch set belongs.

[0060] In one embodiment, a digital twin scene model with geometric topology and high-resolution surface texture is used as the data input and fed into a pre-trained semantic segmentation neural network. In this embodiment, the neural network employs a deep learning architecture such as PointNet++, and achieves automated recognition and patch classification of different objects within the scene by extracting deep spatial features from the vertex spatial distribution and local geometric features of the 3D mesh. The process outputs a set of mesh patches containing multiple scene semantic categories, such as classifying the entire scene model into specific semantic labels like concrete ground, metal pipes, steel railings, and the outer wall of a storage tank.

[0061] To ensure the simulation environment possesses realistic physical response capabilities, key physical calibration parameters are obtained through on-site mechanical testing. Specifically, for various real materials corresponding to the aforementioned semantic classification labels, the static friction coefficient, dynamic friction coefficient, mass distribution properties, and elastic recovery coefficient are accurately measured using testing methods such as the tilted plane method, falling ball experiments, and Allan variance analysis. Simultaneously, the bidirectional reflectance distribution function (BRDF) parameters of various real material surfaces are extracted using a spectrophotometer, thereby constructing complete physical calibration parameters and reflectivity parameters for physically rendered materials. These obtained physical calibration parameters and reflectivity parameters for physically rendered materials are then precisely injected into the material computation assets of the physical simulation engine (such as the Chaos physics system in the UE5 engine) according to the specific semantic classification label to which the mesh face set belongs. By dynamically associating the geometric index of the face with the physical attribute database, the system can configure a corresponding physical material instance for each mesh face.

[0062] Through the aforementioned semantic mapping and parameter injection process, a high-fidelity simulation environment integrating surface texture and realistic physical interaction characteristics is constructed. In this environment, the embodied agent not only receives visual feedback with high-fidelity surface texture, but the friction, support force, and collision rebound effects generated when its feet or robotic arms interact with the scene model are also calculated in real-time by the injected physical calibration parameters. Since the material properties in the simulation environment are generated based on the mechanical calibration data of the real site, the gap in contact mechanics between the embodied agent and reality is bridged. This deep and consistent physical modeling, combined with a high-resolution rendered scene, provides the embodied agent control model with an extremely realistic reinforcement learning training ground.

[0063] In one embodiment, such as Figure 4 As shown, the generation and access of the dynamic weather and lighting rendering conditions includes the following steps S41-S42: In step S41, meteorological visual effects are configured within the simulation environment that integrates surface texture and real physical interaction characteristics. Dynamic lighting conditions, such as the luminous intensity, absolute color temperature, and incident elevation angle of the global parallel light source, are superimposed. The meteorological visual effects include large-scale cloud density values ​​and local precipitation particle emissivity. In step S42, the exposure compensation and color grading correction parameters of the camera are dynamically adjusted in real time, and nonlinear photometric fusion calculations are performed on the meteorological visual effects and dynamic lighting conditions. The dynamic meteorological and lighting rendering conditions are output for the embodied intelligent agent control model to be trained to perform virtual visual signal perception.

[0064] In one embodiment, for the generation and integration of dynamic weather and lighting rendering conditions, diverse weather visual effects are configured through the visual components of the graphics engine within a pre-constructed simulation environment that integrates surface textures and realistic physical interaction characteristics. The system simulates changes in sky obscurity from clear to overcast by adjusting the spatially large-scale cloud density value, and utilizes a particle system to control the local precipitation particle emissivity in real time to reproduce different precipitation intensities from drizzle to heavy rain. Furthermore, the system overlays dynamic lighting conditions from a global parallel light source, accurately simulating the natural light and shadow shifts and spectral evolution from sunrise, noon, to dusk by periodically correcting the luminous intensity, absolute color temperature, and incident elevation angle.

[0065] To ensure that the simulated visual signals conform to the physical imaging characteristics of a real camera, a real-time post-processing mechanism was further introduced. The system dynamically adjusts the exposure compensation and color grading correction parameters of the virtual camera in real time using algorithms, and performs nonlinear photometric fusion calculations on the aforementioned meteorological visual effects and dynamic lighting conditions. This simulates the scattering and refraction of light in complex meteorological media, as well as the reflection interaction on different material surfaces, ultimately outputting highly realistic dynamic meteorological and lighting rendering conditions. This rendering result is directly used by the embodied agent control model to be trained for virtual visual signal perception, enabling it to adapt to complex and changing visual environments during the reinforcement learning phase. The domain-adaptive training scene constructed through photometric fusion technology greatly enhances the robustness of the embodied agent control model to drastic changes in real-world lighting and severe weather.

[0066] In one embodiment, such as Figure 5 As shown, the generation domain adaptive training scenario includes the following steps S51-S53: In step S51, at the initial time node of reinforcement learning in the embodied intelligent agent control model to be trained, the surface texture color, global ambient light intensity, and the physical calibration parameters are uniformly and randomly perturbed according to a preset interval ratio; In step S52, the gradient inversion layer is connected in series between the output of the visual encoder network used to extract high-dimensional feature vectors and the input of the environment domain discriminator responsible for binary classification, thereby constructing an adversarial deep neural network learning architecture for cross-domain feature extraction; In step S53, the environment domain discriminator performs binary classification to distinguish the authenticity of the image source between the virtual observation image output in real time by the high-fidelity simulation environment rendering engine and the unlabeled image data of the real scene. The negative gradient signal generated by the discrimination error is directly backpropagated to the visual encoder through the gradient inversion layer. In the continuous minimax adversarial game, the extracted environmental visual features approximate the consistent state of the data distribution between the domain adaptive training scene and the real application scene.

[0067] In one embodiment, at the initial time point of reinforcement learning for the embodied agent control model to be trained, the system automatically triggers a domain randomization process. For the surface texture color, global ambient light intensity, and physical calibration parameters (such as friction coefficient, damping coefficient, etc.) in the aforementioned constructed high-fidelity simulation environment, uniform random perturbations are applied according to a preset range proportion (e.g., within ±20% of the baseline value). During the simulation phase, the system constructs a highly diverse perceptual and physical sample space, forcing the embodied agent control model to learn a universal control law that does not depend on specific visual colors or precise physical constants.

[0068] To achieve depth alignment between the simulation and real domains at the deep feature dimension, an adversarial deep neural network learning architecture for cross-domain feature extraction was constructed. Logically, it consists of three core components: a visual encoder responsible for transforming the input image into a high-dimensional feature vector; an environment discriminator responsible for performing binary classification; and a gradient inversion layer cascaded between the two. During the training phase, a high-fidelity simulation environment rendering engine outputs virtual observation images in real time, while the system incorporates pre-collected unlabeled real-scene image data as a target domain reference. The visual encoder extracts features from both types of images simultaneously and feeds the output high-dimensional feature vectors into the environment discriminator. The core function of the environment discriminator is to attempt to identify the source of the feature vector, i.e., to perform binary classification to determine the authenticity of the image source, trying to distinguish whether the feature originates from the virtual observation image or the real-scene image.

[0069] In this adversarial architecture, the gradient inversion layer plays a crucial role in domain alignment. When the environment domain discriminator generates a gradient signal based on the discrimination error, the gradient inversion layer directly inverts this signal during backpropagation, forming a negative gradient signal which is then fed back to the visual encoder. Through this continuous minimax adversarial game, the visual encoder is driven to eliminate domain-dependent features with significant "simulation" or "realism," instead extracting common features that are invariant between the simulation and real domains. Ultimately, the environmental visual features extracted by the visual encoder are driven to a consistent data distribution between the domain-adaptive training scenario and the real application scenario. This feature-level domain-adaptive transfer method enables the embodied agent control model to process real sensor signals indiscriminately when deployed in real physical environments.

[0070] In one embodiment, such as Figure 6 As shown, it also includes the following steps S61-S62: In step S61, for the real visible light camera and depth detection sensor mounted on the physical robot platform, physical offline error calibration is performed, static baseline data and dynamic shot characteristics under multiple working conditions are extracted, and heteroscedastic Gaussian mathematical equation and distance-related polynomial error compensation equation are fitted respectively, thereby generating a parameterized sensor noise model. In step S62, the parameterized sensor noise model is used as an independent low-level noise intervention term and is superimposed on the virtual sensor output of the high-fidelity simulation environment in real time during each rendering cycle. This allows the observation signal carrying the hardware-level noise factor to be combined with the visual feedback data after uniform probability sampling perturbation to form a matrix, which serves as the state input reference of the visual encoder network.

[0071] In one embodiment, to further reduce the distributional differences between the simulated environment and the real physical world at the perception level, refined offline error calibration was performed on the real visible light camera and depth sensing sensors mounted on the physical robot platform. By acquiring the raw sensor outputs under various environmental conditions, the system extracted static baseline data and dynamic shot characteristics under multiple conditions. Based on the extracted features, heteroscedastic Gaussian mathematical equations for visible light imaging characteristics and distance-dependent polynomial error compensation equations for depth sensing characteristics were fitted, thereby generating a parameterized sensor noise model encompassing the underlying hardware noise characteristics. This parameterized sensor noise model can reproduce the random electronic noise, quantum shot noise, and systematic detection bias caused by changes in measurement distance during the imaging process of real hardware.

[0072] Within each rendering cycle of reinforcement learning, the parameterized sensor noise model is used as an independent low-level noise intervention term, superimposed in real-time on the virtual sensor output of the high-fidelity simulation environment. Specifically, the system dynamically generates corresponding interference factors based on the real-time observation values ​​of the virtual sensor using the parameterized sensor noise model, resulting in the output observation signal carrying a hardware-level noise factor. Subsequently, the observation signal carrying the hardware-level noise factor is combined with the visual feedback data after the aforementioned uniform probability sampling perturbation for matrix synthesis. This synthesis matrix serves as the state input reference for the visual encoder network, allowing the embodied agent to be exposed to non-ideal perceptual signals highly consistent with real physical hardware during the training phase. Through this parameterized modeling and injection at the perception level, the robustness of the embodied agent control model to sensor errors is greatly enhanced, effectively reducing perceptual bias during the Sim2Real transfer process.

[0073] In one embodiment, Figure 7 This is a block diagram illustrating a virtual scene generation and domain adaptive migration apparatus according to an exemplary embodiment. Figure 7As shown, the virtual scene generation and domain adaptive migration device includes an acquisition module 71, a construction module 72, a generation module 73, and a deployment module 74.

[0074] The acquisition module 71 is used to acquire the original laser scanning point cloud data of the real site, and perform statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation. The processed target point cloud is input into the surface reconstruction algorithm to generate a closed three-dimensional mesh model, and multi-view texture perspective projection is implemented to generate a digital twin scene model with geometric topology and high-resolution surface texture. The building module 72 is used to map pre-measured physical calibration parameters and reflectivity parameters of physically rendered materials onto the three-dimensional mesh of the digital twin scene model based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, thereby constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics. The generation module 73 is used to apply a preset random perturbation to the visual perception parameters and the physical calibration parameters through the high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics, perform domain randomization processing, and combine it with real scene image data to perform adversarial feature alignment, thereby generating a domain adaptive training scene. The deployment module 74 is used to connect the embodied intelligent agent control model to be trained into the domain adaptive training scenario, and to perform continuous state interaction and action feedback in an environment experiencing dynamic weather and lighting rendering conditions to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

[0075] The acquisition module 71, the construction module 72, the generation module 73, and the deployment module 74 included in the block diagram of the virtual scene generation and domain adaptive migration device are controlled to execute the virtual scene generation and domain adaptive migration method described in any of the above embodiments.

[0076] like Figure 8 As shown, the present invention provides an electronic device 800, which includes: a communication interface, a processor 801, and a memory 802; The memory 802 stores program instructions. When executed by the processor 801, which is connected to the memory 802 via the communication interface, the program instructions acquire the original laser-scanned point cloud data of the real site, perform statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation, input the processed target point cloud into a surface reconstruction algorithm to generate a closed 3D mesh model, and implement multi-view texture perspective projection to generate a digital twin scene model with geometric topology and high-resolution surface texture. Based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, the pre-measured physical calibration parameters and the reflectivity parameters of the physically rendered materials are mapped to the digital twin scene. On the three-dimensional mesh of the scene model, a high-fidelity simulation environment integrating surface texture and real physical interaction characteristics is constructed. Through the high-fidelity simulation environment integrating surface texture and real physical interaction characteristics, a preset random perturbation is applied to the visual perception parameters and the physical calibration parameters for domain randomization processing. Combined with real scene image data, adversarial feature alignment is performed to generate a domain adaptive training scene. The embodied intelligent agent control model to be trained is connected to the domain adaptive training scene. In the environment experiencing dynamic weather and lighting rendering conditions, it performs continuous state interaction and action feedback to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

[0077] This invention provides a computer-readable storage medium storing computer program instructions. When executed by a processor, the computer program instructions acquire raw laser-scanned point cloud data of a real-world site, perform statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation, input the processed target point cloud into a surface reconstruction algorithm to generate a closed 3D mesh model, and implement multi-view texture perspective projection to generate a digital twin scene model with geometric topology and high-resolution surface texture. Based on the scene semantic categories defined in the digital twin scene model with geometric topology and high-resolution surface texture, pre-measured physical calibration parameters and reflectivity parameters of physically rendered materials are mapped to the digital twin scene model. On a three-dimensional mesh, a high-fidelity simulation environment integrating surface texture and real physical interaction characteristics is constructed. Through this high-fidelity simulation environment, preset random perturbations are applied to the visual perception parameters and the physical calibration parameters for domain randomization. Combined with real scene image data, adversarial feature alignment is performed to generate a domain adaptive training scene. The embodied intelligent agent control model to be trained is connected to the domain adaptive training scene. Under dynamic weather and lighting rendering conditions, it performs continuous state interaction and action feedback to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical environment.

[0078] It should be understood that the specific features, operations, and details described above regarding the method of the present invention can also be similarly applied to the apparatus and system of the present invention, or vice versa. Furthermore, each step of the method of the present invention described above can be performed by a corresponding component or unit of the apparatus or system of the present invention.

[0079] It should be understood that the various modules / units of the device of the present invention can be implemented wholly or partially through software, hardware, firmware, or a combination thereof. Each module / unit can be embedded in the processor of a computer device in hardware or firmware form or independent of the processor, or it can be stored in the memory of a computer device in software form for the processor to call to execute the operation of each module / unit. Each module / unit can be implemented as an independent component or module, or two or more modules / units can be implemented as a single component or module.

[0080] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores computer instructions executable by the processor, which, when executed by the processor, instruct the processor to perform steps of the methods of embodiments of the present invention. The computer device can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface and communication interface of the computer device can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the methods of the present invention.

[0081] This invention can be implemented as a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, causes the steps of the methods of embodiments of the invention to be performed. In one embodiment, the computer program is distributed across multiple network-coupled computer devices or processors, such that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be executed by one or more computer devices or processors, and one or more other method steps / operations may be executed by one or more other computer devices or processors. One or more computer devices or processors may execute a single method step / operation, or execute two or more method steps / operations.

[0082] It will be understood by those skilled in the art that the method steps of the present invention can be performed by a computer program instructing related hardware, such as a computer device or processor. The computer program may be stored in a non-transitory computer-readable storage medium, and its execution causes the steps of the present invention to be performed. Depending on the context, any references herein to memory, storage, databases, or other media may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0083] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for virtual scene generation and domain adaptation migration, characterized in that, include: The system acquires raw point cloud data from laser scanning of the real site, performs statistical filtering and voxel downsampling based on nearest neighbor distance calculation, inputs the processed target point cloud into a surface reconstruction algorithm to generate a closed 3D mesh model, and implements multi-view texture perspective projection to generate a digital twin scene model with geometric topology and high-resolution surface texture. Based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, the pre-measured physical calibration parameters and the reflectivity parameters of the physical rendering material are mapped onto the three-dimensional mesh surface of the digital twin scene model, thereby constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics. Through the high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics, the visual perception parameters and the physical calibration parameters are subjected to preset random perturbation for domain randomization processing, and adversarial feature alignment is performed in combination with real scene image data to generate a domain adaptive training scene. The embodied intelligent agent control model to be trained is connected to the domain adaptive training scenario. It performs continuous state interaction and action feedback in an environment with dynamic weather and lighting rendering conditions to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

2. The virtual scene generation and domain adaptive migration method as described in claim 1, characterized in that, The generation of a digital twin scene model with geometric topology and high-resolution surface texture includes: For the original laser scanning point cloud data, outliers are removed by using a set number of nearest neighbor points and a standard deviation threshold, and the overall point cloud density is unified by voxel downsampling to output the target point cloud; The three-dimensional spatial surface topology of the target point cloud is calculated using the Poisson surface reconstruction algorithm to reconstruct the closed three-dimensional mesh model, and the two-dimensional plane coordinate line unfolding is performed on the closed three-dimensional mesh model using an angle-based parameterization algorithm. By combining the visible light image data of the real site, the multi-view texture perspective projection is performed on the 3D mesh model unfolded by the automatic coordinate lines to generate a diffuse reflection map, and different levels of surface reduction calculation operations are automatically performed according to the camera observation distance threshold to generate a digital twin scene model with multiple levels of detail.

3. The virtual scene generation and domain adaptation migration method of claim 2, wherein, The high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics includes: The digital twin scene model with geometric topology and high-resolution surface texture is used as the data input. It is fed into a pre-trained semantic segmentation neural network for deep spatial feature extraction and patch classification, and outputs a set of mesh patches containing multiple scene semantic categories. Through on-site mechanical testing, the static friction coefficient, dynamic friction coefficient, mass distribution properties, and elastic recovery coefficient were obtained, and the bidirectional reflection distribution of various real material surfaces was extracted to construct the physical calibration parameters and the reflectivity parameters of the physically rendered material. The physical calibration parameters and the reflectivity parameters of the physically rendered material are injected into the material calculation assets of the physical simulation engine according to the specific semantic classification tags to which the mesh patch set belongs. 4.The virtual scene generation and domain adaptation migration method of claim 1, wherein, The generation and access of the dynamic weather and lighting rendering conditions include: Meteorological visual effects are configured within the simulation environment that integrates surface texture and real physical interaction characteristics. Dynamic lighting conditions, including the luminous intensity, absolute color temperature, and incident elevation angle of a global parallel light source, are superimposed. The meteorological visual effects include large-scale cloud density values ​​and local precipitation particle emissivity. The camera's exposure compensation and color grading correction parameters are dynamically adjusted in real time. Nonlinear photometric fusion calculations are performed on the meteorological visual effects and dynamic lighting conditions. The dynamic meteorological and lighting rendering conditions are output for the embodied intelligent agent control model to be trained to perform virtual visual signal perception.

5. The method of claim 4, wherein, The generated domain adaptive training scenario includes: At the initial time point of reinforcement learning in the embodied intelligent agent control model to be trained, the surface texture color, global ambient light intensity, and the physical calibration parameters are uniformly and randomly perturbed according to a preset interval ratio. The gradient inversion layer is connected in series between the output of the visual encoder network used to extract high-dimensional feature vectors and the input of the environment domain discriminator responsible for binary classification, thus constructing an adversarial deep neural network learning architecture for cross-domain feature extraction. The environment domain discriminator performs binary classification to distinguish the authenticity of the image source between the virtual observation image output in real time by the high-fidelity simulation environment rendering engine and the unlabeled image data of the real scene. The negative gradient signal generated by the discrimination error is directly backpropagated to the visual encoder through the gradient inversion layer. In the continuous minimax adversarial game, the extracted environmental visual features approximate the consistent state of the data distribution between the domain adaptive training scene and the real application scene.

6. The virtual scene generation and domain adaptation migration method of claim 5, wherein, Also includes: For the real visible light camera and depth detection sensor mounted on the physical robot platform, perform physical offline error calibration, extract static baseline data and dynamic shot characteristics under multiple working conditions, and fit heteroscedastic Gaussian mathematical equation and distance-related polynomial error compensation equation respectively, thereby generating a parameterized sensor noise model. The parameterized sensor noise model is used as an independent low-level noise intervention term and is superimposed on the virtual sensor output of the high-fidelity simulation environment in real time during each rendering cycle. This allows the observation signal carrying hardware-level noise factors to be combined with the visual feedback data after uniform probability sampling perturbation to form a matrix, which serves as the state input reference for the visual encoder network.

7. A virtual scene generation and domain adaptation migration apparatus, characterized by, include: The acquisition module is used to acquire the original point cloud data of the real site by laser scanning, and perform statistical filtering and voxel downsampling operations based on nearest neighbor distance calculation. The processed target point cloud is input into the surface reconstruction algorithm to generate a closed three-dimensional mesh model, and multi-view texture perspective projection is implemented to generate a digital twin scene model with geometric topology and high-resolution surface texture. The construction module is used to map pre-measured physical calibration parameters and reflectivity parameters of physically rendered materials onto the three-dimensional mesh surface of the digital twin scene model based on the scene semantic categories divided in the digital twin scene model with geometric topology and high-resolution surface texture, thereby constructing a high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics. The generation module is used to apply a preset random perturbation to the visual perception parameters and the physical calibration parameters through the high-fidelity simulation environment that integrates surface texture and real physical interaction characteristics, perform domain randomization processing, and combine it with real scene image data to perform adversarial feature alignment, thereby generating a domain adaptive training scene. The deployment module is used to connect the embodied intelligent agent control model to be trained into the domain adaptive training scenario, and to conduct continuous state interaction and action feedback in an environment experiencing dynamic weather and lighting rendering conditions to complete reinforcement learning. The embodied intelligent agent control model that has completed reinforcement learning is deployed to the embodied intelligent agent to perform autonomous navigation and action interaction tasks in a real physical site.

8. The virtual scene generation and domain adaptation migration apparatus of claim 7, wherein: The acquisition module, the construction module, the generation module, and the deployment module are controlled to execute the virtual scene generation and domain adaptive migration method according to any one of claims 2 to 6.

9. An electronic device, characterized in that, include: Communication interface, processor, memory; The memory is used to store program instructions, which, when executed by the processor connected to the memory via the communication interface, cause the electronic device to implement the virtual scene generation and domain adaptive migration method according to any one of claims 1 to 6.

10. A computer-readable storage medium having stored thereon program instructions, wherein, When the program instructions are executed by a computer, the computer causes the computer to implement the virtual scene generation and domain adaptive migration method according to any one of claims 1 to 6.