Virtual training scene construction method based on AI
By constructing an AI-based virtual training scenario and using deep learning and a graphics rendering engine to simulate dynamic home environments, the problem of high cost and low efficiency in robot training has been solved. This enables efficient training of robots in a virtual environment and improves their ability to perform tasks in complex home environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies suffer from high costs and low efficiency in robot training, especially in terms of insufficient task execution capabilities in complex home environments. Furthermore, existing virtual training solutions lack home-specific rule libraries and static models that cannot be parameterized, resulting in poor training outcomes.
By acquiring static scene data, designing training elements and binding dynamic behavior parameters and task rules, a virtual training scene is generated. Using deep learning and a graphics rendering engine, a multi-parameter configurable virtual training environment is constructed to simulate dynamic home scenes such as hinge dynamics, object disorder, and material properties, enabling the robot to undergo comprehensive training in a virtual environment.
It enables efficient and low-cost training of robots in a virtual environment, improving their task execution capabilities in complex home environments, shortening training time, and enhancing training effectiveness and adaptability.
Smart Images

Figure CN121659558A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an AI-based method for constructing virtual training scenarios. Background Technology
[0002] Virtualized home scenarios address core bottlenecks in robot learning by reducing costs and time, providing infinitely diverse environments, ensuring training safety, optimizing algorithm efficiency, and promoting skill transfer. Building diverse kitchen scenarios in the real world requires substantial physical resources (such as furniture and appliances), and space and materials are limited; robots require millions of training iterations, taking years to complete in reality. Virtual environments, accelerated by algorithms and AI, significantly shorten training time; and virtual environments can generate infinitely evolving scenarios based on data, overcoming physical limitations.
[0003] With the rapid development of the robotics market, the ability to perform tasks in complex home environments has become a core bottleneck for industrial application. Current robot training faces the dilemma of high cost and low efficiency: the cost of single-time debugging of physical prototypes in real-world scenarios is too high, and it is difficult to cover dynamic interference scenarios (such as cabinet door opening and closing and stain accumulation). Existing virtual training solutions have fundamental limitations—general-purpose engine training lacks a home-specific rule library, leading to a burden of secondary development; dedicated system training is limited by static models, which cannot parameterize key variables such as stain location and furniture movement trajectory.
[0004] The method of training robots by constructing virtual scenes emphasizes the importance of data and simulation of real-world scenarios. The massive amounts of detailed 3D models accumulated in the home decoration design field (including spatial layout, furniture structure, and material data) have long been limited to visual display purposes, and their dynamic training value has not been activated. Developing the value of existing data allows robots to be trained to operate under different lighting conditions, times, and home layouts to adapt to various environments and tasks, such as finding objects and opening / closing doors. Through more comprehensive training in virtual scenes, robots can better perform various tasks within the home. Summary of the Invention
[0005] The purpose of this invention is to provide an AI-based method for constructing virtual training scenarios to solve the technical problems mentioned in the background section.
[0006] The technical solution to achieve the objective of this invention is: an AI-based virtual training scene construction method, comprising the following steps:
[0007] S1. Obtain static scene data;
[0008] S2. Design training elements, generate dynamic behavior parameters, and set task rules;
[0009] S3. Bind and associate dynamic behavior parameters and task rules with static scene data to generate virtual training scenarios.
[0010] Furthermore, S1 specifically includes the following steps:
[0011] S11: Receive and parse the user's two-dimensional floor plan, and generate a preliminary three-dimensional model of the house structure through background filtering, polygon recognition, and separation of walls, doors and windows.
[0012] S12: A point cloud or mesh segmentation algorithm based on deep learning is used to perform fine identification of components at the preliminary model, generating an assembly relationship map that accurately describes the parent-child relationships and constraints between components.
[0013] S13: Based on the identified model, calculate the key spatial indicators related to the robot, output machine-readable spatial metadata, and directly serve the robot's path planning and task planning algorithms.
[0014] Furthermore, the design training elements include hinge dynamics simulation, object disorder simulation, state-driven system, and model material simulation.
[0015] Furthermore, the hinge dynamics simulation includes: binding hinge parameters to the cabinet door components and calculating the spatiotemporal occupancy matrix of the door opening and closing sweep area; the hinge parameters include rotation axis coordinates, maximum opening and closing angle, and angular velocity curve.
[0016] Furthermore, the object disorder simulation includes generating disordered states of objects that conform to human behavior, training the robot's spatial optimization decision-making ability, generating disordered states of objects based on probability distribution functions, and associating space utilization and stability constraints.
[0017] Furthermore, the state-driven system includes dynamically calculating the stain thickness based on the material's adsorption characteristics and the stain's exposure time, generating a visualized stain layer, and supporting adjustments to the thickness gradient and distribution density; matching cleaning tools based on the stain's chemical composition, and generating multi-level decision paths.
[0018] Furthermore, the model material simulation includes constructing a material-process mapping matrix, establishing a mandatory binding relationship between material ID and processing technology, and defining a linear equation between material hardness and tool rotation speed.
[0019] Furthermore, step S3 specifically includes:
[0020] S31: Receive static scene data, dynamic behavior parameters, and task rules; perform spatiotemporal binding of dynamic behavior parameters with the coordinates of movable parts in the static scene data; associate task rules with semantic tags in the static scene data; generate a virtual environment configuration file containing a timeline event instruction set;
[0021] S32: Based on the compiled virtual environment configuration file, call the graphics rendering engine and physics engine to instantiate the training scene.
[0022] By adopting the above technical solution, the present invention has the following beneficial effects:
[0023] (1) This invention transforms static home decoration data into a virtual training environment that supports dynamic rule injection and constructs a virtual training environment with multiple configurable parameters. It provides users with a customizable training scenario for training robots to perform complex tasks. Through more comprehensive training in the virtual scenario, the robot can better complete various tasks in the home. Attached Figure Description
[0024] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein...
[0025] Figure 1 This is a flowchart of the present invention.
[0026] Figure 2 This is a flowchart illustrating the specific process of step S1 of the present invention.
[0027] Figure 3 This is a flowchart illustrating the specific process of step S3 of the present invention. Detailed Implementation
[0028] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0030] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0031] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0032] In the description of the embodiments of the present invention, it should be understood that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of the invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used to facilitate the description of the present invention and to simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention.
[0033] In the description of the embodiments of the present invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances. The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the scope of protection of the present invention.
[0034] (Example 1)
[0035] See Figure 1 The AI-based virtual training scene construction method in this embodiment includes the following steps:
[0036] S1. Obtain static scene data; static scene data refers to the data set obtained from the "physical-semantic integrated data base" that describes the inherent attributes of components in the virtual scene, such as semantic information (type, material, function), physical attributes, and relationships.
[0037] See Figure 2 S1 specifically includes the following steps:
[0038] S11: Receives and parses the user's 2D floor plan, and through a series of steps such as background filtering, polygon recognition, and separation of walls, doors, and windows, initially generates a 3D model of the house structure (including the basic geometry and layout of walls, doors, and windows). After obtaining the initial 3D model...
[0039] S12: Based on deep learning point cloud or mesh segmentation algorithms, perform fine component-level identification on the preliminary model to generate an assembly relationship map, accurately describing the parent-child relationships and constraints between components, and elevating it from a "visual model" to a "physical-semantic integrated data base". This data base elevates the traditional geometric model into a machine-understandable and interactive digital entity, adding semantic labels after fine component identification; for example, "wall" includes "load-bearing wall" and "partition wall"; "door" includes the installation position and rotation axis of "door panel", "door frame" and "hinge".
[0040] S13: Based on the identified model, calculate key spatial indicators related to the robot, such as passage width and operating space height, and output machine-readable spatial metadata to directly serve the robot's path planning and task planning algorithms.
[0041] S2. Design training elements to generate dynamic behavior parameters; set task rules; dynamic behavior parameters refer to the rule parameters for movable parts to perform operations, such as motion trajectory, event triggering, and state changes. Task rules refer to the set training objectives and behavioral constraints that may exist in a specific scenario. Designing training elements to generate dynamic behavior parameters includes hinge dynamics simulation, object disorder simulation, state-driven systems, and model material simulation.
[0042] Hinge dynamics simulation includes: binding hinge parameters to cabinet door components, including rotation axis coordinates, maximum opening and closing angles, and angular velocity curves, and calculating the spatiotemporal occupancy matrix of the door's opening and closing sweep area. Hinge types include hinges, slide rails, and rotation axes.
[0043] The simulation of object disorder includes: generating disordered states of objects that conform to human behavior, training the robot's spatial optimization decision-making ability, generating disordered states of objects based on probability distribution functions, and associating space utilization and stability constraints. By setting an upper limit (e.g., ≤70%) for the space utilization parameter (i.e., the ratio of the total volume of objects to the usable volume of the container), the total number of objects generated is constrained, ensuring that the generated scenarios are both challenging and conform to physical stability and common sense in human daily behavior, reserving necessary space for training scenarios. For example, when the input furniture type is a wardrobe, a clothing stacking pattern is generated, with the item distribution generated according to a space utilization of ≤70%, where the clothing stacking pattern is generated according to 30% flat, 50% rolled, and 20% stacked. As another example, when the input furniture is a bookshelf, books are generated on the bookshelf, with the tilting angles of the books randomly distributed.
[0044] The state-driven system includes: providing the robot with a near-realistic stain treatment training environment, solving the problem of material-stain interaction logic that cannot be trained in static scenes. It dynamically calculates stain thickness based on material adsorption characteristics and stain exposure time, generating a visualized stain layer and supporting adjustments to thickness gradient and distribution density. Parameters are categorized into stain type (oil stains / tea stains / water scale, etc.), exposure time (stain duration), and material adsorption coefficient (e.g., wood 0.8, stainless steel 1.2, marble 1.5). Stain type determines cleaning agent selection, exposure time affects thickness and adhesion, and material adsorption coefficient changes cleaning difficulty. Cleaning tools are matched based on the stain's chemical composition, generating multi-level decision paths. The multi-level decision-making path includes: Level 1: Stain type identification and cleaning agent matching; Level 2: Enhanced physical parameters; Level 3: Cleaning action after assessing material compatibility and process selection; Level 4: Effect evaluation. If the stain residue is ≤5%, the task is completed; if 5% < stain residue ≤50%, return to Level 3 to continue cleaning; if the stain residue is >50%, return to Level 2, continue enhancing physical parameters, and then enter Level 3.
[0045] The model material simulation includes: constructing a material-process mapping matrix, establishing a mandatory binding relationship between material IDs and processing processes, and defining a linear equation between material hardness and tool rotation speed. For example, if the material type is wood flooring, the cleaning rule is wiping with the grain, and the pressure is ≤2 N / cm². 2 The material type is marble, with circumferential polishing and a rotation speed of ≥120rpm.
[0046] S3. Bind and associate dynamic behavior parameters and task rules with static scene data to generate virtual training scenarios.
[0047] Step S3 specifically includes:
[0048] See Figure 3S31: Receives static scene data, dynamic behavior parameters, and task rules. The static scene data contains movable and immovable parts; for example, a wall is immovable, while a door is movable. It spatiotemporally binds the dynamic behavior parameters to the coordinates of the movable parts in the static scene data; it associates the task rules with the semantic tags in the static scene data; and it generates a virtual environment configuration file containing a timeline event instruction set. Specifically, it spatiotemporally binds the dynamic behavior parameters to the coordinates of the movable parts in the static scene data, for example, constructing a four-dimensional spatiotemporal coordinate system and binding the dynamic behavior parameters (such as the curve of the cabinet door's opening angle changing over time) to the three-dimensional hinge coordinates in the hinge parameters, forming an executable instruction: "At time T, the door rotates N degrees." It also associates the task rules with the semantic tags in the static scene data, binding element designs to semantic tags; for example, when a surface with surface material = "wood" is identified, its cleaning rule is automatically associated with action = "wipe along the grain." Based on random event parameters, preset "trigger points" on the timeline. For example, there is a 30% probability of triggering the "cabinet door opening" event 120 seconds after the start of training. Establish causal chains, such as: "If the robot knocks over a water glass (event A), water stains will be generated on the floor (state change), and a cleaning task (new target) will be triggered."
[0049] S32: Based on the compiled virtual environment configuration file, the graphics rendering engine and physics engine are invoked to instantiate the training scene; among them, multi-level detail models are loaded on demand for objects in the scene, including fine models for high-definition rendering and simplified collider sets for real-time collision detection, and the compiled scene configuration file is run in the powerful graphics and physics engines.
[0050] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing a virtual training scene based on AI, characterized in that, Includes the following steps: S1. Obtain static scene data; S2. Design training elements and generate dynamic behavior parameters; Set task rules; S3. Bind and associate dynamic behavior parameters and task rules with static scene data to generate virtual training scenarios.
2. The method for constructing an AI-based virtual training scene according to claim 1, characterized in that, S1 specifically includes the following steps: S11: Receive and parse the user's two-dimensional floor plan, and generate a preliminary three-dimensional model of the house structure through background filtering, polygon recognition, and separation of walls, doors and windows. S12: A point cloud or mesh segmentation algorithm based on deep learning is used to perform fine identification of components at the preliminary model, generating an assembly relationship map that accurately describes the parent-child relationships and constraints between components. S13: Based on the identified model, calculate the key spatial indicators related to the robot, output machine-readable spatial metadata, and directly serve the robot's path planning and task planning algorithms.
3. The method for constructing an AI-based virtual training scene according to claim 1, characterized in that: The design training elements generate dynamic behavior parameters including hinge dynamics simulation, object disorder simulation, state-driven system, and model material simulation.
4. The method for constructing an AI-based virtual training scene according to claim 3, characterized in that: The hinge dynamics simulation includes: binding hinge parameters to the cabinet door components and calculating the spatiotemporal occupancy matrix of the door opening and closing sweep area; the hinge parameters include rotation axis coordinates, maximum opening and closing angle, and angular velocity curve.
5. The method for constructing an AI-based virtual training scene according to claim 3, characterized in that: The simulation of object disorder includes: generating disordered states of objects that conform to human behavior, training the robot's spatial optimization decision-making ability, generating disordered states of objects based on probability distribution functions, and associating space utilization and stability constraints.
6. The method for constructing an AI-based virtual training scene according to claim 3, characterized in that: The state-driven system includes: dynamically calculating the stain thickness based on the material's adsorption characteristics and the stain's exposure time, generating a visual stain layer, and supporting adjustments to the thickness gradient and distribution density; matching cleaning tools based on the stain's chemical composition, and generating multi-level decision paths.
7. The method for constructing an AI-based virtual training scene according to claim 3, characterized in that: The model material simulation includes: constructing a material-process mapping matrix, establishing a mandatory binding relationship between material ID and processing technology, and defining a linear equation between material hardness and tool rotation speed.
8. The method for constructing an AI-based virtual training scene according to claim 1, characterized in that: Step S3 specifically includes: S31: Receive static scene data, dynamic behavior parameters, and task rules; perform spatiotemporal binding of dynamic behavior parameters with the coordinates of movable parts in the static scene data; associate task rules with semantic tags in the static scene data; Generate a virtual environment configuration file containing a timeline event instruction set; S32: Based on the compiled virtual environment configuration file, call the graphics rendering engine and physics engine to instantiate the training scene.