Training data generation method and device, electronic equipment and storage medium

By generating natural language navigation instructions and verifying their accessibility in a 3D simulation environment, the problem of insufficient data diversity and path accessibility in existing technologies is solved. The generated training data is efficient, diverse and executable, and is suitable for visual language navigation tasks.

CN121959029APending Publication Date: 2026-05-01SHENZHEN XGRIDS-INNOVATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN XGRIDS-INNOVATION CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

While improving data diversity, existing visual language navigation training datasets struggle to guarantee the reachability of paths corresponding to navigation commands in a 3D environment, resulting in poor navigation performance of the model in unknown environments.

Method used

Generate natural language navigation instructions, construct a 3D simulation environment constrained by them, verify the correctness and accessibility of the construction in the simulation environment, verify the path accessibility through the navigation mesh, regenerate instructions after adjusting the reasons for inaccessibility, and ensure that the generated data is executable in the 3D environment.

Benefits of technology

It achieves improved data diversity while ensuring the 3D environmental accessibility of the path corresponding to the navigation command, improved consistency and reliability of the generated training data structure, significantly improved data scale and generation efficiency, and covers more diverse navigation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121959029A_ABST
    Figure CN121959029A_ABST
Patent Text Reader

Abstract

The invention provides a training data generation method and device, electronic equipment and a storage medium. The method comprises the following steps: constructing a three-dimensional simulation environment constrained by a natural language navigation instruction; executing the natural language navigation instruction in a simulation environment corresponding to the three-dimensional simulation environment to verify the construction correctness of the three-dimensional simulation environment; if the navigation route is correct, verifying whether the navigation route described by the natural language navigation instruction is reachable in the three-dimensional simulation environment or not based on a navigation grid corresponding to a ground walkable area in the three-dimensional simulation environment; and if yes, generating training data of a visual language navigation task for the preset intelligent agent according to the natural language navigation instruction and the three-dimensional simulation environment. According to the method and the device, the diversity of the VLN training data and the consistency of the three-dimensional space can be considered, and the accessibility of the path corresponding to the navigation instruction in the three-dimensional environment is ensured while the data diversity is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual language navigation, and more specifically, to a method, apparatus, electronic device, and storage medium for generating training data. Background Technology

[0002] Vision-and-Language Navigation (VLN) refers to the task of enabling an autonomous agent to navigate in an unknown 3D environment based on natural language instructions. Research on VLN tasks is severely constrained by data acquisition. Traditional VLN training datasets are constructed based on scans of real indoor environments, resulting in a limited number of scenes and a single route pattern, leading to poor model generalization ability and unsatisfactory navigation performance of the agent in unknown environments.

[0003] To alleviate data scarcity, existing technologies expand data through methods such as simulated environment collection, synthesis and editing of existing environmental data, or acquisition of internet data. However, these methods suffer from limitations in the number of scenarios, low data quality, and fragmentation. While generative models introduced in recent years can improve visual diversity, they neglect three-dimensional spatial constraints, resulting in inconsistent data and rendering the routes described by navigation instructions unusable in the corresponding environments. This leads to a severe disconnect between instructions and the actual environment.

[0004] In summary, existing technologies struggle to balance the diversity of VLN training data with consistency in 3D space. Ensuring the accessibility of paths corresponding to navigation commands in the 3D environment while improving data diversity remains a pressing technical challenge. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method, apparatus, electronic device and storage medium for generating training data, which can take into account both the diversity of VLN training data and the consistency of three-dimensional space, and while improving the diversity of data, ensure the accessibility of the path corresponding to the navigation command in the three-dimensional environment.

[0006] In a first aspect, embodiments of this application provide a method for generating training data, the method comprising: Generate natural language navigation instructions; the natural language navigation instructions are used to describe in natural language the navigation route from a pre-generated navigation starting point, through at least one pre-generated intermediate point, to a pre-generated navigation ending point; Construct a three-dimensional simulation environment constrained by the natural language navigation instructions; The natural language navigation command is executed in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; If the three-dimensional simulation environment is set up correctly, then based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment, verify whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment; If the navigation route described by the natural language navigation instruction is reachable in the three-dimensional simulation environment, then training data for the visual language navigation task is generated for the preset intelligent agent based on the natural language navigation instruction and the three-dimensional simulation environment.

[0007] In one possible implementation, if the navigation route described by the natural language navigation instructions is unreachable in the three-dimensional simulation environment, the method further includes: Based on the three-dimensional simulation environment and the navigation grid, determine the reasons why the navigation route described by the natural language navigation instructions is unreachable in the three-dimensional simulation environment; Based on the stated reasons for unreachability, the natural language navigation instructions are adjusted from at least one preset instruction adjustment level; Based on the adjusted natural language navigation instructions, the system jumps to the constructed three-dimensional simulation environment constrained by the natural language navigation instructions to continue execution.

[0008] In one possible implementation, generating natural language navigation instructions includes: Generate a navigation task that meets preset navigation conditions; the navigation task includes the navigation starting point, at least one intermediate transit point, and the navigation destination; The navigation task is converted into a natural language expression to obtain the natural language navigation instructions.

[0009] In one possible implementation, constructing the three-dimensional simulation environment constrained by the natural language navigation instructions includes: Generate a two-dimensional image constrained by the natural language navigation instructions; A three-dimensional simulation environment is constructed based on the two-dimensional image.

[0010] In one possible implementation, the natural language navigation instructions are executed in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment, including: In the simulation environment corresponding to the three-dimensional simulation environment, the virtual camera is controlled to move according to the navigation route described by the natural language navigation instructions, so as to capture the visual observation image sequence of the preset intelligent agent on the navigation route; The correctness of the construction of the three-dimensional simulation environment was verified based on the visual observation image sequence.

[0011] In one possible implementation, verifying the correctness of the construction of the three-dimensional simulation environment based on the visual observation image sequence includes: Based on the visual observation image sequence, extract the actual scene feature information corresponding to the three-dimensional simulation environment; If the actual scene feature information is consistent with the standard scene features corresponding to the natural language navigation command, then the three-dimensional simulation environment is correctly constructed.

[0012] In one possible implementation, verifying whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment includes: Find the shortest connected path in the navigation grid from the navigation start point corresponding to the natural language navigation command to the navigation end point; If the shortest connected path includes all intermediate points in the navigation route described by the natural language navigation instruction, and the order of passage of each intermediate point is consistent with the order in the navigation route, then the navigation route described by the natural language navigation instruction is determined to be reachable in the three-dimensional simulation environment.

[0013] Secondly, embodiments of this application also provide a training data generation apparatus, the apparatus comprising: A generation module is used to generate natural language navigation instructions; the natural language navigation instructions are used to describe in natural language the navigation route from a pre-generated navigation starting point, through at least one pre-generated intermediate point, to a pre-generated navigation endpoint. A construction module is used to construct a three-dimensional simulation environment constrained by the natural language navigation instructions; The verification module is used to execute the natural language navigation instructions in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; The verification module is also used to verify, if the three-dimensional simulation environment is set up correctly, whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment; The generation module is used to generate training data for the preset intelligent agent for a visual language navigation task, based on the natural language navigation instructions and the three-dimensional simulation environment, if the navigation route described by the natural language navigation instructions is reachable in the three-dimensional simulation environment.

[0014] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the training data generation method as described in any of the first aspects.

[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the training data generation method as described in any of the first aspects.

[0016] This application provides a method, apparatus, electronic device, and storage medium for generating training data. The method includes: constructing a three-dimensional simulation environment constrained by natural language navigation instructions; executing the natural language navigation instructions in a simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; if the three-dimensional simulation environment is correctly constructed, verifying whether the navigation route described by the natural language navigation instructions is reachable in the three-dimensional simulation environment based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment; if the navigation route described by the natural language navigation instructions is reachable in the three-dimensional simulation environment, generating training data for a preset intelligent agent for a visual language navigation task based on the natural language navigation instructions and the three-dimensional simulation environment. This application can balance the diversity of VLN training data with the consistency of three-dimensional space, improving data diversity while ensuring the reachability of the path corresponding to the navigation instructions in the three-dimensional environment. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart of a training data generation method provided in an embodiment of this application is shown; Figure 2 A flowchart of another method for generating training data provided in an embodiment of this application is shown; Figure 3 This paper shows a schematic diagram of the structure of a training data generation device provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0020] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0021] To enable those skilled in the art to utilize the content of this application, and in conjunction with the specific application scenario of "visual language navigation," the following implementation methods are provided. For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application. Although this application is primarily described in the context of "visual language navigation," it should be understood that this is merely an exemplary embodiment.

[0022] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0023] Vision-and-Language Navigation (VLN) refers to the task of enabling autonomous agents to navigate in unknown 3D environments based on natural language instructions. Research on VLN is severely constrained by data acquisition. Traditional VLN training datasets (such as Room-to-Room, R2R) are primarily constructed from scans of real indoor environments; for example, the Matterport3D dataset contains only about 90 house scenes (61 scenes in the training set). Researchers construct data by manually annotating navigation paths and instructions on man-made panoramic views, but manually acquiring panoramic images and writing instructions is extremely difficult and time-consuming. On the one hand, the limited data scale makes it difficult for existing VLN models to generalize to new environments. On the other hand, navigation instructions generated based on these limited scenes often cover a single route pattern; when the agent is deployed to an unknown environment, the mismatch between the training and inference scenarios significantly affects navigation performance.

[0024] To alleviate the problem of data scarcity, several methods for generating and enhancing VLN training data have emerged in recent years. One approach involves acquiring data in new simulated environments: utilizing additional 3D scenes provided by the simulator, randomly sampling visual observations, and generating instructions to augment the data. Another approach involves synthetically editing data from existing environments, such as changing the style, object appearance, viewpoint position, or texture of Matterport3D scenes. Some research has also attempted to obtain visual data with navigation commands from the internet. However, these approaches each have limitations: using new simulated environments is limited by the number of available scenes or remains confined to the statistical distribution of Matterport3D; using network data often results in poor quality, fragmentation, and a lack of structure, significantly differing from well-annotated VLN data. Recent work has also introduced generative models: for example, using diffusion models (such as Stable Diffusion) to recursively generate panoramic images based on text descriptions to increase visual diversity. This purely generative approach semantically enriches observations, but due to neglecting 3D spatial constraints, the generated data suffers from world inconsistency—i.e., chaotic geometric relationships and illogical layouts between images lead to a disconnect between navigation commands and the actual environment. In summary, how to improve data diversity while ensuring consistency between data and the real three-dimensional world (including path accessibility) remains a challenge that needs to be addressed.

[0025] The following section explains the existing methods for generating VLN training datasets and their limitations: One existing technique involves manually collecting real-world environment datasets: This is the earliest and most direct method in the VLN field. For example, the R2R dataset is built based on Matterport3D real-world indoor scenes. The specific process is as follows: using specialized equipment to scan real-world indoor environments to obtain panoramic image sets, manually selecting start and end points within these environments and planning navigation routes, then manually writing natural language instructions describing the routes. Each data entry contains information such as environment ID, start point location, target location, observation sequence along the path, and corresponding language instructions. To facilitate training, the environment is usually pre-represented as a navigation graph with nodes and edges. Nodes are observable viewpoints, and edges indicate traversability between viewpoints. The human annotator refers to the panoramic environmental image and, following the planned path on the navigation graph, describes the journey from the start point to the target location in colloquial language, including turns, passing objects or rooms, and finally stopping at the target location. For example: "Leave the bedroom, enter the kitchen, walk forward, turn left at the sofa, and stop at the window." This type of method ensures a high degree of correspondence between instructions and environmental paths because the data originates from real-world environments and its feasibility has been manually verified.

[0026] The shortcomings of existing technologies: The main limitations of manual dataset methods lie in the data scale and acquisition cost. First, due to limitations in real-world collection conditions, datasets such as Matterport3D have a very limited number of scenes (only about 61 houses for training). The routes and instructions in each environment are also limited, providing only tens of thousands of training samples in total, far from sufficient to cover diverse combinations of navigation instructions. Second, manual collection is inefficient: scanning new environments requires specialized equipment and time, and writing instructions requires manual intervention and a high degree of familiarity with the environment, making data expansion very difficult. Third, scene diversity is limited: these data mainly come from real residences, with limited layout types. The model may overfit to specific home arrangements during training, resulting in poor generalization ability to new environments. The lack of an automated mechanism to ensure coverage of more route types exacerbates the problems of insufficient and homogeneous training data. This invention addresses these shortcomings by aiming to significantly expand the scale and diversity of VLN training data without relying on manual annotation.

[0027] Existing technology two, data augmentation methods based on simulated environment synthesis and generation: With the development of simulation technology, some research has turned to the automatic generation of navigation data. One approach is to introduce additional virtual environments: for example, using hundreds of 3D scanning environments such as HM3D and Gibson provided by the Habitat simulator, placing the agent in these environments for random walks, and then automatically generating corresponding instructions. Specific methods include: constructing a navigation grid or topology map for each new environment, using a program to randomly sample start-end pairs and paths on the map, and then using a trained "narrator" model (such as EnvDrop Speaker) to generate instruction descriptions based on the paths. Another approach is to directly synthesize images and instructions using generative models: a typical example is the PanoGen method, which first extracts descriptive text based on the semantics of the original dataset, then calls a diffusion model to progressively generate a 360° panoramic image sequence that matches the description, and finally generates or reuses navigation instructions to match these synthesized visuals. Such methods attempt to increase data diversity by imagining new scenarios, rather than relying entirely on existing real-world scenes.

[0028] The shortcomings of existing technology 2: Although synthetic data augmentation methods have alleviated the data scarcity problem to some extent, there are still obvious shortcomings. (1) Environmental dependence limitation: The scheme of using additional simulated environments (such as HM3D, Gibson) is limited by the number of available environments. Even if there are more scenes than Matterport3D, the total number is still limited, and the cost of acquiring these scanned environments is high and it is not easy to expand infinitely. At the same time, new environments often come from similar sources (e.g., houses from the same data source), and the scene distribution pattern is still affected by the deviation of the original dataset, making it difficult to cover a sufficiently diverse range of navigation scenarios. (2) Lack of structural rationality: The method of completely random sampling of paths and generation of instructions may produce samples that lack logical coherence. For example, the image and text data scraped from the Internet are often fragmented, and the instructions and environment may not strictly correspond, lacking clear starting and ending semantics. Although some instruction sentences are grammatically correct, the locations or directions mentioned may not exist or be reasonable in the corresponding environment, resulting in large data noise. In other words, random generation lacks constraints on the environmental structure, and the generated tasks may not conform to common sense room connections or path planning habits. (3) No reachability verification: Existing synthetic methods usually do not strictly verify the executability of navigation. For example, the panoramic images generated by PanoGen are merely visual representations of a "new environment," but without real 3D spatial support, the agent cannot truly infer coherent spatial routes from the images. Even in simulator-based methods, impassable paths may occur during sampling (e.g., commands are forcibly spliced ​​together when there is no path between the start and end points). Due to the lack of path reachability checks, the final dataset may contain non-executable navigation commands. These commands are detrimental to training agents because they describe a task that does not exist or cannot be performed in the environment.

[0029] In view of this, embodiments of this application provide a method for generating training data, including: generating natural language navigation instructions; the natural language navigation instructions are used to describe, in natural language, a navigation route starting from a pre-generated navigation starting point, passing through at least one pre-generated intermediate transit point, and ending at a pre-generated navigation endpoint; constructing a three-dimensional simulation environment constrained by the natural language navigation instructions; executing the natural language navigation instructions in a simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; if the three-dimensional simulation environment is correctly constructed, verifying whether the navigation route described by the natural language navigation instructions is reachable in the three-dimensional simulation environment based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment; if the navigation route described by the natural language navigation instructions is reachable in the three-dimensional simulation environment, generating training data for a visual language navigation task for the preset intelligent agent based on the natural language navigation instructions and the three-dimensional simulation environment. The training data generation method provided by embodiments of this application can balance the diversity of VLN training data with the consistency of three-dimensional space, improving data diversity while ensuring the reachability of the path corresponding to the navigation instructions in the three-dimensional environment.

[0030] The following is a detailed description of a training data generation method provided in an embodiment of this application.

[0031] Reference Figure 1 The diagram shown is a flowchart illustrating a method for generating training data according to an embodiment of this application. The exemplary steps of this embodiment are described below: S101, Generate natural language navigation instructions.

[0032] In this embodiment, natural language navigation instructions are used to describe, in natural language, a navigation route starting from a pre-generated navigation starting point, passing through at least one pre-generated intermediate point, and ending at a pre-generated navigation destination. The specific generation process is as follows: Step 1: Generate a navigation task that meets the preset navigation conditions; the navigation task includes a navigation start point, at least one intermediate transit point, and a navigation destination.

[0033] In this application's implementation, the preset navigation conditions may include navigation scene type (such as residential indoor scene, office indoor scene, and school playground, etc.), navigation requirements under that navigation scene type (such as requiring cross-area navigation between rooms in a residential indoor scene, rather than local navigation within the same room), and navigation task difficulty parameters (such as path length, number of rooms involved, etc.). The navigation start point, intermediate points of interest (POIs), and navigation endpoint can be randomly selected based on the semantics of scene elements (living room, sofa, etc.) in each navigation scene. For example, the start point may be "next to the sofa in the living room," the endpoint may be "next to the dining table in the kitchen," and the intermediate point may be "corridor."

[0034] In this context, a Point of Interest (POI) refers to a location or object in the environment that is meaningful for the navigation task; it can be abstracted as the location of a point. For example, a room (living room, kitchen, etc.) or a specific object (such as a table or sofa) can serve as a POI to describe a reference point in the navigation target or path. POIs typically have attributes such as a name and location coordinates, facilitating the location of the corresponding target within the instructions and environment. When an agent executes a path from the navigation starting point to the navigation destination, it needs to pass through each intermediate point in sequence.

[0035] Step 2: Convert the navigation task into a natural language expression to obtain natural language navigation instructions.

[0036] In this embodiment, navigation tasks can be transformed into descriptive statements that conform to human habits, resulting in natural language navigation instructions, by using preset conversion templates or generating them using a large language model (LLM). For example, given a navigation start point and destination, the large language model will generate instructions such as "Start from the sofa in the living room, walk east through the hallway, enter the kitchen, and stop at the dining table." The natural language navigation instructions incorporate navigation task planning information to ensure that the instruction content strictly corresponds to the actual path, including details such as orientation (forward / left turn, etc.), landmarks, and room names.

[0037] Here, on the one hand, the embodiments of this application can use rule scripts to ensure terminological consistency and logical coherence of instructions; on the other hand, it also supports the introduction of pre-trained language models to enrich the diversity of instructions. After generating the natural language navigation instructions, the natural language navigation instructions are temporarily stored for use in the final training data packaging.

[0038] S102. Construct a three-dimensional simulation environment constrained by natural language navigation instructions.

[0039] In the embodiments of this application, the three-dimensional simulation environment constrained by natural language navigation instructions refers to a three-dimensional virtual environment constructed based on the navigation route described by the natural language navigation instructions. Its scene layout, spatial connectivity, navigation starting point position, intermediate passing point position, navigation ending point position, and object distribution all match the description of the navigation starting point, intermediate passing point, navigation ending point, and route direction in the instructions, ensuring that the semantic requirements of the environment and the navigation instructions are consistent.

[0040] Optionally, a 3D simulation environment constrained by natural language navigation instructions is constructed, including: the large language model generates parameter configuration instructions for each environmental component element corresponding to the natural language navigation instructions according to the semantic requirements of the natural language navigation instructions (such as matching the size parameters of "small living room" and the position parameters of "corridor connecting rooms"); based on the parameter configuration instructions, the corresponding basic element models in the scene library are driven to automatically adjust their element size, position and other attributes, and complete the assembly and adaptation of each environmental component element according to the combination rules, and finally build a 3D simulation environment that conforms to the constraints of natural language navigation instructions.

[0041] For example, if the instruction involves "living room" and "kitchen", the two room units, living room and kitchen, will be retrieved from the scene library and connected by a corridor or door to ensure that a path from the living room to the kitchen exists.

[0042] The scene library pre-stores parametric basic models of elements such as room type, walls, doors and windows, and furniture. Each model is associated with adjustable parameters such as size, scale, and material (e.g., wall height, door width, and furniture size).

[0043] Optionally, constructing a three-dimensional simulation environment constrained by natural language navigation instructions includes: generating a two-dimensional image constrained by natural language navigation instructions; and constructing a three-dimensional simulation environment based on the two-dimensional image.

[0044] For example, natural language navigation instructions can be parsed using a large language model to extract navigation nodes, spatial connectivity requirements, key objects, and scene layout semantics, which are then converted into two-dimensional layout parameters that can be recognized by procedural modeling (such as room dimensions, relative positions, door and window opening positions, and furniture placement coordinates). Two-dimensional vector models corresponding to elements such as room types, walls, doors, windows, and furniture are retrieved from the scene library. These models are associated one-to-one with the three-dimensional base models in the scene library, and parameterized adjustment interfaces for dimensions and positions are reserved. Using a parameter-driven approach, the attributes of each two-dimensional vector model are adjusted according to the aforementioned two-dimensional layout parameters. Planar assembly is completed according to conventional building structures and instruction semantics, generating a two-dimensional image containing room distribution, spatial connectivity relationships, and the planar layout of key objects. The layout parameters of this image are consistent with the parameter system of subsequent three-dimensional modeling. The scene layout information extracted from the two-dimensional image is analyzed, including room distribution, spatial connectivity, key object positions and size ratios, and converted into parametric configuration data that can be recognized by programmatic modeling. Based on the parametric configuration data, the attributes of the basic models in the scene library are adjusted in a parameter-driven manner to match the scene layout requirements in the two-dimensional image. According to the adjusted parameters and preset building structure rules, each basic model is automatically assembled and adapted, and the construction of walls, doors and windows, furniture placement and spatial connectivity are completed in sequence, finally completing the scene construction of the three-dimensional simulation environment.

[0045] Here, controlled generation means that the environment layout and content are constrained randomness: it introduces a certain degree of randomness to ensure diversity while adhering to conventional architectural structures and task semantics (e.g., avoiding illogical layouts like a living room directly connected to a bathroom). A combination of procedural modeling and a scene library is used: elements such as room types, walls, doors, windows, and furniture are predefined and assembled in a parameter-driven manner to create a scene that meets the requirements. Simultaneously, a navigation mesh (NavMesh) is calculated for the generated environment: walkable areas on the ground are automatically baked into NavMesh data for subsequent path planning. The output of this step is a navigable virtual 3D world, including environmental geometry, semantic labels, and navigation mesh information.

[0046] NavMesh (Navigation Mesh) is a polygonal mesh data structure used for path planning and navigation. It is often used by game AI and robots to mark which areas in a 3D environment are walkable. NavMesh represents a complex environment as several connected convex polygonal regions, providing path search algorithms with the ability to calculate feasible paths from the starting point to the destination.

[0047] S103. Execute natural language navigation commands in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment.

[0048] In this embodiment of the application, in the simulation environment corresponding to the three-dimensional simulation environment, the virtual camera is controlled to move according to the navigation route described by the natural language navigation instructions, so as to capture the visual observation image sequence of the preset intelligent agent on the navigation route; the correctness of the construction of the three-dimensional simulation environment is verified according to the visual observation image sequence.

[0049] The verification of the correctness of the construction of the 3D simulation environment based on the visual observation image sequence includes: extracting the actual scene feature information corresponding to the 3D simulation environment based on the visual observation image sequence; if the actual scene feature information is consistent with the standard scene features corresponding to the natural language navigation instructions, then the 3D simulation environment is correctly constructed.

[0050] Scene features can include the connectivity between various environmental elements (such as rooms) and the location of each environmental element.

[0051] Here, a 3D simulation environment is loaded into the simulation engine to achieve 3D environment rendering and simulate the visual observation of a preset intelligent agent (robot). A virtual camera is placed at the navigation starting point to obtain environmental images from the first-person perspective of the preset intelligent agent. The virtual camera moves along the navigation route described by the natural language navigation instructions, capturing key views along the way (e.g., panoramic frame sequences) to obtain a sequence of visual observation images. The simulation engine provides realistic rendering, making the visual information contained in the generated data as close as possible to the real world. Through simulation, we not only verify whether the environment is correctly built (e.g., whether rooms are connected, whether objects are in the expected positions), but also obtain the necessary perceptual input for subsequent verification steps. In many cases, simulation rendering can also be used to generate image data for training, such as image sequences along the path, depth maps, etc.; however, the main purpose of this embodiment is to ensure path feasibility, and the focus here is more on the interactive usability of the environment than visual quality.

[0052] It's important to note that robot simulation platforms like Unreal Engine or Gazebo can be used instead of the Unity engine, as long as they can generate scenes with navigation meshes. Similarly, in environment generation algorithms, probabilistic or machine learning-based methods can replace rule-based algorithms. For example, Generative Adversarial Networks (GANs) or diffusion models can directly generate 3D room layouts and object placements. These generative models, if well-trained, can also produce environments that meet constraints. Although the implementation methods differ, the essence is the generation of navigable virtual environments, and the output—the spatial relationships of rooms and walkable areas—has an equivalent effect on subsequent path planning.

[0053] S104. If the three-dimensional simulation environment is set up correctly, then based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment, verify whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment.

[0054] In this embodiment of the application, the shortest connected path between the navigation starting point and the navigation ending point corresponding to the natural language navigation instruction is found in the navigation grid; if the shortest connected path includes all intermediate points in the navigation route described by the natural language navigation instruction, and the passage order of each intermediate point is consistent with the order in the navigation route, then it is determined that the navigation route described by the natural language navigation instruction is reachable in the three-dimensional simulation environment.

[0055] Here, a preset navigation start point and destination are input into the intelligent agent. Classic path search and shortest path planning algorithms such as A* or Dijkstra are invoked to find the shortest path on the navigation grid. If the planning algorithm finds a connected path that matches the navigation route described in the instruction (e.g., passing through each intermediate point of interest (POI) described in the instruction), the reachability verification is considered successful. If the algorithm cannot find a path (e.g., the start and destination are blocked by a wall, or a required POI is unreachable), the instruction is deemed unexecutable, and a regeneration process is triggered. If necessary, a real navigation agent can be used to attempt actions according to the instruction in a simulation environment to more rigorously verify the executability of the instruction. However, static path search is usually sufficient to determine reachability. Through path verification, this embodiment of the application achieves a closed-loop reachability constraint: only navigation instructions that pass verification proceed to the next step, ensuring that every instruction in the output data is executable in the corresponding environment.

[0056] It should be noted that other path planning or verification methods can be used, as long as it can be determined whether the instruction is executable. For example, Dijkstra's algorithm based on grid maps can be used instead of NavMesh, or path search algorithms such as Probabilistic Road Map (PRM) and Rapid Expanding Random Tree (RRT) can be used for reachability analysis. Regarding reachability verification, a pre-trained navigation agent can be allowed to actually attempt to execute the instruction in a simulation environment; if it succeeds, feasibility is proven. These alternative methods are equivalent in principle to NavMesh+A*, all checking environmental connectivity and path existence, and therefore serving the same purpose in ensuring reachability.

[0057] S105. If the navigation route described by the natural language navigation instruction is reachable in the three-dimensional simulation environment, then generate training data for the visual language navigation task for the preset intelligent agent based on the natural language navigation instruction and the three-dimensional simulation environment.

[0058] In this embodiment, after the natural language navigation command passes the above verification steps, the data is formatted, packaged, and imported into the output training dataset. The standardization output module organizes information according to a predetermined training dataset format, such as: the identifier of the 3D simulation environment or the configuration file of the 3D simulation environment, the representation of the navigation starting point in environmental coordinates, the representation of the navigation ending point in environmental coordinates, the navigation path (which can be represented by a node sequence or coordinate sequence), the generated natural language navigation command, and possibly image sequences collected in the environmental simulation. All training data structures are kept consistent to facilitate downstream users in loading and training models for performing visual language navigation tasks, enabling the preset agent to perform visual language navigation tasks based on the trained model and improving the navigation performance of the preset agent.

[0059] The training dataset generated in this embodiment can be exported as JSON, CSV, or even a database, with environment files and NavMesh files included as needed. After standardization, each data entry clearly records the three key elements: "environment-path-instruction," and due to reachability guarantees, users no longer need to manually clean and filter the data. Once this step is completed, the automatically generated training data can be directly used by downstream algorithms or subjected to further quality checks and enhancements, thereby improving the navigation performance of the intelligent agent.

[0060] Among them, the pre-defined intelligent agent refers to an intelligent system that has a physical entity and can interact with the environment to perform visual language navigation tasks.

[0061] Furthermore, if the navigation route described by the natural language navigation instructions is unreachable in the 3D simulation environment, it is necessary to analyze the reasons for the unreachability and execute a corresponding regeneration strategy. This is an important part of closed-loop control. (Refer to...) Figure 2 As shown, another method for generating training data provided in this application embodiment is implemented as follows: S201. If the navigation route described by the natural language navigation instruction is unreachable in the three-dimensional simulation environment, then determine the reason why the navigation route described by the natural language navigation instruction is unreachable in the three-dimensional simulation environment based on the three-dimensional simulation environment and the navigation grid.

[0062] In the embodiments of this application, a large language model can be used to analyze the reasons why natural language navigation instructions are unreachable, in combination with a three-dimensional simulation environment and navigation grid, such as: unreasonable environment layout (e.g., lack of connecting doors), obstacles causing the path to be blocked, or instructions not matching the environment.

[0063] S202. Based on the reason for unreachability, adjust the natural language navigation instructions from at least one preset instruction adjustment level.

[0064] In the embodiments of this application, different measures are taken to automatically correct different reasons: for example, in the environmental level of the preset instruction adjustment layer, the room connection method is adjusted, doorways are added / removed, or obstacles are rearranged; in the task level of the preset instruction adjustment layer, if the environment is fixed, the location of the navigation start point / navigation end point is modified, or the intermediate points (POIs) mentioned in the natural language navigation instructions are filtered / adjusted.

[0065] S203. Based on the adjusted natural language navigation instructions, jump to the constructed three-dimensional simulation environment constrained by the natural language navigation instructions to continue execution.

[0066] In this embodiment, after adjustment, the natural language navigation instructions are reconstructed using a 3D simulation environment for subsequent verification of correctness and reachability. If the number of adjustments to a natural language navigation instruction exceeds a preset number, the instruction is deleted to avoid an infinite loop, and a new instruction is generated. This design ensures self-correction capabilities: iterating continuously until a consistent set of tasks, instructions, and environments is generated. The failure regeneration module creates a closed-loop feedback loop in data production, preventing invalid data from flowing into the final dataset.

[0067] In summary, the beneficial effects of the embodiments of this application are reflected in the following aspects: (1) Enhanced Data Quality and Consistency: In the generation process, the embodiments of this application impose semantic and reachability constraints, ensuring that the output data is highly consistent in structure and logic. On the one hand, each natural language navigation instruction corresponds to a real-world path, avoiding discrepancies between the statement and the environment; on the other hand, the 3D simulation environment, constrained by the natural language navigation instructions, ensures that the POIs involved in the natural language navigation instructions are reasonably distributed in the scene. This rule-guided training data generation ensures that the training data has good structured characteristics and authenticity, significantly reducing the occurrence of invalid or contradictory samples. Compared with unconstrained randomly generated datasets, the training data generated in the embodiments of this application is cleaner and more reliable, eliminating the need to spend effort to remove erroneous samples when training the agent, thus improving training effectiveness.

[0068] (2) Significantly improved training data scale and generation efficiency: Through a fully automated pipeline, the embodiments of this application can produce navigation data in batches at a much higher rate than manual methods. Without the need for manual annotation of each instruction and path, tens of thousands of diverse navigation training data samples can be generated in a short time. Especially after introducing procedural 3D world construction, the composable scene components and parameter space are extremely large, theoretically making the number of different environments generated unlimited. Compared to traditional datasets with only dozens of environments and tens of thousands of instructions, the embodiments of this application can achieve a leap in scale (e.g., generating millions of data points). The efficient generation capability also means that dataset versions can be iterated quickly as needed, promptly covering new scene types or instruction forms, thereby reducing the constraints of data bottlenecks on algorithm development.

[0069] (3) Improved Task Diversity and Realistic Alignment: Benefiting from controlled environment generation and rich text generation methods, the data generated in this application covers a wider range of navigation scenarios. By adjusting navigation conditions, paths of different lengths and complexities can be generated, covering various situations such as room switching, corridor turns, and multi-segment instructions; by changing environmental elements, various indoor layouts and home furnishings can be involved, thus approaching the diversity of the real world. More importantly, since each piece of data is guaranteed to be executable, these instructions can be regarded as valid navigation instructions that may occur in real life, which are reasonable and operable. This makes the generated data semantically and physically closer to real tasks, and the trained navigation agent can learn more practical skills, narrowing the gap between the game environment and the real environment.

[0070] (4) Engineering feasibility and scalability: The embodiments of this application fully utilize existing mature technologies (such as the physics and navigation modules of game engines, pre-trained language models, etc.), with clear decoupling of each module and a solid foundation for engineering implementation. It can be deployed on ordinary GPU workstations, requiring minimal human intervention in the generation process, meeting the requirements of labor-saving and efficient data acquisition in actual R&D. Furthermore, the modular design of the system allows for the replacement or upgrading of components to adapt to different needs. For example, new scene libraries can be added to expand the environment types, or more advanced language models can be switched to improve the naturalness of instructions. This flexible scalability gives the solution long-term viability, facilitating its integration into enterprise patent data production lines for batch production of dedicated datasets, and also enabling the generation of navigation data tailored to the sensor and motion characteristics of specific robot platforms. In summary, the embodiments of this application greatly facilitate the acquisition of high-quality VLN training data, laying a solid foundation for further research and application.

[0071] Based on the same inventive concept, this application also provides a training data generation device corresponding to the training data generation method. Since the principle of the device in this application is similar to the training data generation method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0072] Reference Figure 3 The diagram shown is a schematic of a training data generation device provided in an embodiment of this application. The training data generation device includes: The generation module 301 is used to generate natural language navigation instructions; the natural language navigation instructions are used to describe in natural language the navigation route from a pre-generated navigation starting point, through at least one pre-generated intermediate point, to a pre-generated navigation ending point; Construction module 302 is used to construct a three-dimensional simulation environment constrained by the natural language navigation instructions; The verification module 303 is used to execute the natural language navigation instructions in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; The verification module 303 is also used to verify, if the three-dimensional simulation environment is set up correctly, whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment; The generation module 304 is used to generate training data for the preset intelligent agent for a visual language navigation task based on the natural language navigation instruction and the three-dimensional simulation environment if the navigation route described by the natural language navigation instruction is reachable in the three-dimensional simulation environment.

[0073] In one possible implementation, the device further includes an adjustment module 305; if the navigation route described by the natural language navigation instruction is unreachable in the three-dimensional simulation environment, the adjustment module 305 is specifically configured to determine the reason why the navigation route described by the natural language navigation instruction is unreachable in the three-dimensional simulation environment based on the three-dimensional simulation environment and the navigation grid; adjust the natural language navigation instruction from at least one preset instruction adjustment level based on the reason for unreachability; and jump to the construction of the three-dimensional simulation environment constrained by the natural language navigation instruction to continue execution based on the adjusted natural language navigation instruction.

[0074] In one possible implementation, the generation module 301 is specifically used to: generate a navigation task that meets preset navigation conditions; the navigation task includes the navigation starting point, at least one intermediate transit point and the navigation endpoint; and convert the navigation task into a natural language expression to obtain the natural language navigation instructions.

[0075] In one possible implementation, the construction module 302 is specifically used to generate a two-dimensional image constrained by the natural language navigation instructions; and to construct a three-dimensional simulation environment based on the two-dimensional image.

[0076] In one possible implementation, the verification module 303 is specifically used to control a virtual camera to move according to the navigation route described by the natural language navigation instructions in the simulation environment corresponding to the three-dimensional simulation environment, so as to capture a sequence of visual observation images of the preset intelligent agent on the navigation route; and to verify the correctness of the construction of the three-dimensional simulation environment based on the sequence of visual observation images.

[0077] In one possible implementation, the verification module 303 is specifically used to extract actual scene feature information corresponding to the three-dimensional simulation environment based on the visual observation image sequence; if the actual scene feature information is consistent with the standard scene features corresponding to the natural language navigation command, then the three-dimensional simulation environment is correctly built.

[0078] In one possible implementation, the verification module 303 is further configured to: Find the shortest connected path in the navigation grid from the navigation start point corresponding to the natural language navigation command to the navigation end point; If the shortest connected path includes all intermediate points in the navigation route described by the natural language navigation instruction, and the order of passage of each intermediate point is consistent with the order in the navigation route, then the navigation route described by the natural language navigation instruction is determined to be reachable in the three-dimensional simulation environment.

[0079] The training data generation apparatus provided in this application embodiment can balance the diversity of VLN training data with the consistency of three-dimensional space, thereby improving data diversity while ensuring the accessibility of the path corresponding to the navigation command in the three-dimensional environment.

[0080] like Figure 4 As shown in the embodiment of this application, an electronic device 400 includes a processor 401, a memory 402, and a bus. The memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device is running, the processor 401 communicates with the memory 402 via the bus, and the processor 401 executes the machine-readable instructions to perform the steps of the training data generation method described above.

[0081] Specifically, the memory 402 and processor 401 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 401 runs the computer program stored in the memory 402, it can execute the above-mentioned training data generation method.

[0082] Corresponding to the above-described method for generating training data, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method for generating training data.

[0083] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0084] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0085] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0086] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0087] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating training data, characterized in that, The method includes: Generate natural language navigation instructions; the natural language navigation instructions are used to describe in natural language the navigation route from a pre-generated navigation starting point, through at least one pre-generated intermediate point, to a pre-generated navigation ending point; Construct a three-dimensional simulation environment constrained by the natural language navigation instructions; The natural language navigation command is executed in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; If the three-dimensional simulation environment is set up correctly, then based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment, verify whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment; If the navigation route described by the natural language navigation instruction is reachable in the three-dimensional simulation environment, then training data for the visual language navigation task is generated for the preset intelligent agent based on the natural language navigation instruction and the three-dimensional simulation environment.

2. The method for generating training data according to claim 1, characterized in that, If the navigation route described by the natural language navigation instructions is unreachable in the three-dimensional simulation environment, the method further includes: Based on the three-dimensional simulation environment and the navigation grid, determine the reasons why the navigation route described by the natural language navigation instructions is unreachable in the three-dimensional simulation environment; Based on the stated reasons for unreachability, the natural language navigation instructions are adjusted from at least one preset instruction adjustment level; Based on the adjusted natural language navigation instructions, the system jumps to the constructed three-dimensional simulation environment constrained by the natural language navigation instructions to continue execution.

3. The method for generating training data according to claim 1, characterized in that, The generation of natural language navigation instructions includes: Generate a navigation task that meets preset navigation conditions; the navigation task includes the navigation starting point, at least one intermediate transit point, and the navigation destination; The navigation task is converted into a natural language expression to obtain the natural language navigation instructions.

4. The method for generating training data according to claim 1 or 2, characterized in that, The construction of the three-dimensional simulation environment constrained by the natural language navigation instructions includes: Generate a two-dimensional image constrained by the natural language navigation instructions; A three-dimensional simulation environment is constructed based on the two-dimensional image.

5. The method for generating training data according to claim 1, characterized in that, Executing the natural language navigation commands in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment includes: In the simulation environment corresponding to the three-dimensional simulation environment, the virtual camera is controlled to move according to the navigation route described by the natural language navigation instructions, so as to capture the visual observation image sequence of the preset intelligent agent on the navigation route; The correctness of the construction of the three-dimensional simulation environment was verified based on the visual observation image sequence.

6. The method for generating training data according to claim 5, characterized in that, The step of verifying the correctness of the construction of the three-dimensional simulation environment based on the visual observation image sequence includes: Based on the visual observation image sequence, extract the actual scene feature information corresponding to the three-dimensional simulation environment; If the actual scene feature information is consistent with the standard scene features corresponding to the natural language navigation command, then the three-dimensional simulation environment is correctly constructed.

7. The method for generating training data according to claim 1, characterized in that, The verification of whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment, based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment, includes: Find the shortest connected path in the navigation grid from the navigation start point corresponding to the natural language navigation command to the navigation end point; If the shortest connected path includes all intermediate points in the navigation route described by the natural language navigation instruction, and the order of passage of each intermediate point is consistent with the order in the navigation route, then the navigation route described by the natural language navigation instruction is determined to be reachable in the three-dimensional simulation environment.

8. A training data generation apparatus, characterized in that, The device includes: A generation module is used to generate natural language navigation instructions; the natural language navigation instructions are used to describe in natural language the navigation route from a pre-generated navigation starting point, through at least one pre-generated intermediate point, to a pre-generated navigation endpoint. A construction module is used to construct a three-dimensional simulation environment constrained by the natural language navigation instructions; The verification module is used to execute the natural language navigation instructions in the simulation environment corresponding to the three-dimensional simulation environment to verify the correctness of the construction of the three-dimensional simulation environment; The verification module is also used to verify, if the three-dimensional simulation environment is set up correctly, whether the navigation route described by the natural language navigation command is reachable in the three-dimensional simulation environment based on the navigation grid corresponding to the walkable area on the ground in the three-dimensional simulation environment; The generation module is used to generate training data for a preset intelligent agent for a visual language navigation task, based on the natural language navigation instructions and the three-dimensional simulation environment, if the navigation route described by the natural language navigation instructions is reachable in the three-dimensional simulation environment.

9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the training data generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the training data generation method as described in any one of claims 1 to 7.