Robot simulation scene intelligent generation method and system based on large language model
Patent Information
- Application Number
- CN202610815759.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-06-08
AI Technical Summary
[0003]然而,在现有仿真平台的使用过程中,仿真场景的配置仍然面临以下主要困难:第一、场景配置流程复杂繁琐,耗时较长
传统方式需要用户手工在仿真平台的图形界面中编辑或通过编程方式创建场景,需要掌握场景描述文件格式、仿真平台的API、以及相关的脚本命令等专业知识。本发明中,用户无需掌握任何仿真平台专业知识,仅需用自然语言描述需求即可完成场景生成。具体而言,用户无需学习场景描述文件格式、无需学习仿真平台的API、无需了解脚本命令、无需掌握资产库的组织结构和命名规范、无需手动处理资源路径拼接、位置设置、父子关系管理等技术细节。这使得机器人仿真技术从专业人员专用工具转变为研究人员和工程师可便捷使用的通用工具,大幅降低了技术门槛。
Smart Images

Figure CN122413755B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of robot simulation technology, and specifically relates to a method and system for intelligent generation of robot simulation scenes based on a large language model. Background Technology
[0002] Robot simulation is a crucial tool in robot research and development, enabling the verification of algorithms and control strategies in a virtual environment, avoiding the wear and tear and safety risks associated with real hardware. Current mainstream robot simulation platforms (such as NVIDIA Isaac Sim, Gazebo, and Webots) offer advanced physics engines and high-fidelity sensor simulation capabilities. For example, NVIDIA Isaac Sim, as an enterprise-level robot simulation platform, provides GPU-accelerated AI computing capabilities.
[0003] However, in the use of existing simulation platforms, the configuration of simulation scenarios still faces the following major difficulties: First, the scenario configuration process is complex and cumbersome, and time-consuming. Traditional robot simulation scenario configuration requires multiple consecutive steps, including requirements analysis, asset search, environment configuration, robot configuration, and parameter debugging. The total time for the entire process reaches 2-3 hours, which is labor-intensive and prone to errors due to operational mistakes. Second, asset selection lacks intelligent and systematic methods. When the asset library contains hundreds of environmental resources, dozens of robot models, and hundreds of prop elements, it is difficult for users to quickly and accurately find the optimal combination that meets specific needs. The traditional method is for users to check candidate resources one by one in the asset library, making judgments based on limited information such as file names and folder classifications. This method has the following drawbacks: first, it is inefficient, and it may take 1-2 hours to find a suitable combination; second, the accuracy cannot be guaranteed, and users may make suboptimal or incorrect choices due to insufficient information; third, it cannot conduct multi-dimensional comprehensive evaluation and cannot consider complex factors such as environmental compatibility and capability matching. Third, sensor requirement extraction is difficult to accurately process users' natural language expressions. When describing simulation scenario requirements, users often use natural language rather than formal specifications. These natural language expressions are diverse and complex. Traditional keyword matching methods cannot accurately handle complex expressions. They are prone to errors in negative demands, cannot distinguish between background information and actual needs, have incomplete coverage of synonyms, and have limited ability to understand ambiguous expressions. A large number of erroneous extraction results will lead to incorrect asset selection. Fourth, existing scene generation methods lack standardization and automation. Users need to manually edit in the graphical interface of the simulation platform or create scenes through programming, requiring expertise in scene description file formats, platform APIs, and other areas. The operation is highly repetitive, prone to errors, and difficult to standardize. Fifth, asset evaluation is incomplete in multiple dimensions, and the evaluation methods are rudimentary. Existing asset selection methods usually only consider a single dimension and cannot comprehensively consider multiple important attributes of assets, resulting in unreliable recommendation quality.
[0004] Therefore, there is an urgent need for a method and system that can automatically and intelligently generate robot simulation scenarios to improve work efficiency, reduce learning costs, and enhance the accuracy of asset selection. Summary of the Invention
[0005] To address the aforementioned technical challenges, this invention proposes a method and system for intelligent generation of robot simulation scenes based on a large language model. By integrating innovative technologies such as large language models, multi-dimensional asset management, and intelligent recommendation, it achieves end-to-end automated conversion from natural language requirements to complete simulation scenes, significantly improving scene generation efficiency, enhancing the accuracy and intelligence of asset selection, and reducing the learning cost of system usage.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: Firstly, this invention proposes an intelligent generation method for robot simulation scenes based on a large language model, comprising the following steps: The system receives simulation scenario requirements described in natural language and processes them progressively through keyword matching, negative word detection, and semantic verification using a large language model to extract a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment. Based on the required virtual sensor list and the excluded virtual sensor list, the candidate robot models in the asset library are weighted and quantitatively scored across multiple dimensions, and the recommended target robot models are determined by ranking them according to the comprehensive scores. In the robot simulation platform, asset pre-verification, scene reset, resource loading, object configuration, and transaction commit operations are performed to generate a simulation scene. In the object configuration operation, the placement position of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is adopted to complete the instantiation of the target robot model and the mounting of its virtual sensors.
[0007] Furthermore, the simulation scenario requirements are progressively processed through keyword matching, negative word detection, and semantic verification using a large language model, extracting a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment, specifically: Based on a preset keyword mapping table, the simulation scenario requirements are matched to initially identify candidate virtual sensors and generate an initial list of required virtual sensors. Based on the initial list of required virtual sensors, negative words are detected in the context window. Candidate virtual sensors with negative contexts are removed from the required list and added to the list of excluded virtual sensors, resulting in the revised list of required virtual sensors and the list of excluded virtual sensors. Using the revised list of required virtual sensors and the list of excluded virtual sensors, as well as the requirements of the simulation scenario itself, as input, a large language model is invoked to perform semantic understanding and demand intent determination, which is used to distinguish between the user's real needs and background information, hypothetical statements or irrelevant content, and finally outputs a semantically verified list of required virtual sensors and a list of excluded virtual sensors.
[0008] Furthermore, the keyword mapping table adopts a dictionary data structure, with the key being the normalized virtual sensor type and the value being the corresponding list of Chinese and English synonyms.
[0009] Furthermore, the method for detecting negative words is as follows: For each candidate virtual sensor, locate its character position in the original user input, look forward to the preset character range before that position, use the preset character range as a context window, and detect negative words within the context window.
[0010] Furthermore, the method for detecting negative words is as follows: Dependency parsing is performed on the simulation scenario requirements to identify the predicate verb in the sentence. The scope of its dependency subtree is determined with the predicate verb as the center. The scope of the dependency subtree is used as a valid search window for negation words. Negation words are detected within the valid search window. The scope of the dependency subtree includes the subject, object, adverbial and nested modifiers of the predicate verb.
[0011] Furthermore, based on the required virtual sensor list and the excluded virtual sensor list, the candidate robot models in the asset library are weighted and quantitatively scored across multiple dimensions. Specifically: ; in, Indicates environmental similarity score; This indicates the virtual sensor matching score; This indicates the task suitability score; Indicates historical performance score; , , and These are the corresponding weighting coefficients.
[0012] Furthermore, environmental similarity scoring The calculation process is as follows: ; in, The set of labels representing candidate robot models; A set of tags representing the target scene; Virtual sensor matching score The calculation process is as follows: ; in, This represents the list of virtual sensors attached to the candidate robot model; This indicates the list of required virtual sensors; when the candidate robot model does not contain any sensors from the excluded virtual sensor list... ,otherwise ; Task suitability score The calculation process is as follows: ; in, This represents the set of capabilities of the candidate robot model; Indicates the total number of capabilities required for the task; Historical performance rating The calculation process is as follows: ; in, Indicates the first Historical performance indicators; This represents a function that standardizes the index to the range of 0-1. Indicates the first Weights of historical performance indicators; This represents the total number of historical performance metrics; the historical performance metrics include at least one or more of the following: simulation task success rate, average completion time ranking, and user rating.
[0013] Furthermore, during the simulation scene generation process: The asset pre-verification operation is used to confirm that all assets to be loaded exist in the asset library; The scene reset operation is used to delete all nodes created by historical generation operations in the current robot simulation platform scene, restoring the scene to its initial state; The resource loading operation is used to load environmental resources into the scene root node by reference. The object configuration operation adopts a delayed creation strategy, which first saves the configuration information of the target robot model in the internal data structure without immediately creating its scene description file reference, and automatically calculates the placement position of the props based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of the props. The transaction commit operation is used to package all commands to be executed accumulated in the object configuration operation into an atomic transaction and submit them to the simulation engine for execution.
[0014] Furthermore, the placement position of the prop is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources. Specifically, the navigation mesh in the environmental resource file is analyzed, the ground polygon is extracted as the reachable area, the candidate positions for prop placement are calculated based on the reachable area, and the physical query interface of the simulation platform is called to perform collision detection to ensure that the prop does not penetrate the static mesh of the environment. The candidate positions are calculated based on at least one of the following rules: salient position rules based on environmental semantics, directional layout rules based on task requirements, and optimization rules based on obstacle avoidance.
[0015] Secondly, this invention also proposes an intelligent generation system for robot simulation scenes based on a large language model, comprising: The requirements analysis module receives simulation scenario requirements described in natural language and processes them progressively through keyword matching, negative word detection, and semantic verification using a large language model to extract a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment. The asset scoring module is used to perform weighted quantitative scoring on candidate robot models in the asset library across multiple dimensions based on the required virtual sensor list and the excluded virtual sensor list, and to determine the recommended target robot model based on the comprehensive score ranking. The scene generation module is used to perform asset pre-verification, scene reset, resource loading, object configuration and transaction commit operations in the robot simulation platform to generate a simulation scene. In the object configuration operation, the placement position of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is adopted to complete the instantiation of the target robot model and the mounting of its virtual sensors.
[0016] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects: This invention proposes an intelligent generation method and system for robot simulation scenes based on a large language model, belonging to the field of robot simulation technology. The method includes: receiving simulation scene requirements described in natural language; progressively processing the simulation scene requirements through keyword matching, negative word detection, and semantic verification using a large language model to extract a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment; based on the list of necessary and excluded virtual sensors, weighted quantitative scoring of candidate robot models in an asset library across multiple dimensions, and determining the recommended target robot model based on the comprehensive score ranking; performing asset pre-verification, scene reset, resource loading, object configuration, and transaction commit operations in a robot simulation platform to generate the simulation scene; wherein, in the object configuration operation, the placement positions of props are automatically calculated based on the reachable areas and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is used to instantiate the target robot model and mount its virtual sensors. A corresponding system is also proposed based on this method. This invention integrates innovative technologies such as large language models, multi-dimensional asset management, and intelligent recommendation to achieve end-to-end automated conversion from natural language requirements to complete simulation scenarios, significantly improving scenario generation efficiency, enhancing the accuracy and intelligence of asset selection, and reducing the learning cost of system use.
[0017] (i) Significantly improve scene generation efficiency Traditional robot simulation scene configuration involves multiple consecutive steps, including requirements analysis, asset search, environment configuration, robot configuration, and parameter debugging. This process is time-consuming, labor-intensive, and prone to errors due to operational mistakes. This invention, through an end-to-end automated conversion mechanism, requires only a single natural language description from the user, and the system automatically completes all stages, including requirements understanding, asset selection, and scene generation. The entire process achieves an order-of-magnitude efficiency improvement compared to traditional methods, and the level of automation is increased from significant reliance on manual operation to a high degree of automation, reducing the number of manual interventions from multiple to a single input of the natural language requirement.
[0018] (ii) Significantly improve the accuracy of sensor demand identification Traditional keyword matching methods cannot accurately process users' natural language expressions, are prone to errors in negative requests, cannot distinguish between background information and actual needs, have incomplete coverage of synonyms, and have limited ability to understand ambiguous expressions. This invention employs a three-layer progressive sensor request extraction mechanism: the first layer performs basic identification through a comprehensive keyword mapping table, supporting Chinese and English synonyms; the second layer detects negative words within a context window, moving sensors with negative contexts from the required list to the exclusion list of virtual sensors; the third layer calls a large language model for semantic understanding and request intent determination, capable of distinguishing between actual needs and background information, hypothetical statements, or irrelevant content. This three-layer progressive design gives the system strong fault tolerance; even if an identification error occurs in one layer, subsequent layers still have the opportunity to correct it, significantly improving the accuracy of sensor request identification compared to traditional methods.
[0019] (III) Improve the accuracy and intelligence of asset selection Existing asset selection methods typically consider only a single dimension, such as simple label matching, failing to comprehensively consider multiple important asset attributes. This invention employs a multi-dimensional weighted quantitative scoring system to comprehensively evaluate candidate robot models from four dimensions: environmental similarity, virtual sensor matching degree, task adaptability, and historical performance. Environmental similarity is calculated using the Jaccard similarity coefficient to determine the degree of matching between robot labels and scene labels; sensor matching degree is based on sensor coverage and excluded sensor detections; task adaptability is based on capability matching rate; and historical performance is based on historical simulation data. This weighted comprehensive scoring across four dimensions upgrades asset selection from passive search to proactive intelligent recommendation, resulting in a significant improvement in asset selection accuracy compared to traditional methods. Furthermore, the weighting coefficients can be flexibly adjusted according to specific application scenarios to adapt to different application needs.
[0020] (iv) Significantly reduce the learning cost of using the system Traditional methods require users to manually edit or programmatically create scenes within the graphical interface of a simulation platform, necessitating expertise in scene description file formats, simulation platform APIs, and related script commands. In this invention, users do not need any simulation platform expertise; they can generate scenes simply by describing their requirements in natural language. Specifically, users do not need to learn scene description file formats, simulation platform APIs, script commands, the organizational structure and naming conventions of asset libraries, or manually handle technical details such as resource path concatenation, location settings, and parent-child relationship management. This transforms robot simulation technology from a tool for professionals to a universal tool easily used by researchers and engineers, significantly lowering the technical barrier.
[0021] (v) Achieve end-to-end automated conversion from natural language requirements to complete scenarios. Existing technologies lack a unified automated process for transforming user needs into complete scenarios, requiring users to perform manual operations at multiple stages, resulting in fragmented processes that are difficult to standardize and regulate. This invention achieves a complete automated conversion from unstructured natural language requirements to structured simulation scenarios. The entire process has the following characteristics: First, strong determinism—the same input always generates the same scenario, facilitating reproduction and verification; second, standardization—adopting a five-stage standardized process of verification-clearing-loading-configuration-commit, ensuring standardized operation; third, atomicity—transaction commits using an atomic mechanism to ensure the consistency and integrity of scenario generation; and fourth, full automation—requiring no manual intervention, automating the entire process from input to output.
[0022] (vi) Improve the level of intelligence in scene generation Regarding prop placement, existing technologies rely on manual operation, resulting in arbitrary placements and a lack of physical validation. This invention automatically calculates prop placement based on the reachability of environmental resources and spatial semantics: First, it analyzes the navigation grid in the environmental resource file, extracting ground polygons as reachable areas; then, it calculates candidate prop placement locations based on these reachable areas; finally, it calls the simulation platform's physics query interface to perform collision detection, ensuring that props do not penetrate the static environmental grid. Candidate locations can be calculated based on various rules: salient location rules based on environmental semantics (e.g., doorways, corners, in front of the control panel), directional layout rules based on task requirements (path-based distribution for navigation tasks, circular distribution for grasping tasks), and optimization rules based on obstacle avoidance.
[0023] In negation word detection, existing technologies use fixed-window detection, which cannot handle complex sentence structures. This invention employs a dynamic context window determination method based on dependency parsing: dependency parsing is performed on the natural language description to identify the predicate verb in the sentence. The scope of its dependency subtree is determined centered on the predicate verb, and this scope serves as the effective search window for negation words. The system supports a collaborative mechanism between fixed windows and dynamic dependency parsing windows. When dependency parsing successfully constructs a dependency tree, the dynamic window is used; when dependency parsing fails or the length of the natural language description falls below a preset threshold, it reverts to fixed-window detection, ensuring the system's robustness. Attached Figure Description
[0024] Figure 1 This is a flowchart of the intelligent generation method for robot simulation scenes based on a large language model proposed in Embodiment 1 of the present invention; Figure 2 This is an architecture diagram of the intelligent generation method for robot simulation scenes based on a large language model proposed in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the intelligent generation system for robot simulation scenes based on a large language model proposed in Embodiment 2 of the present invention. Detailed Implementation
[0025] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components, processing techniques, and processes are omitted in this invention to avoid unnecessarily limiting the invention.
[0026] Example 1 Embodiment 1 of this invention proposes an intelligent method for generating robot simulation scenes based on a large language model. It addresses the problems of cumbersome and time-consuming processes, lack of intelligent asset selection, inaccurate sensor requirement extraction, lack of standardized and automated scene generation, and single asset evaluation dimensions in the existing robot simulation scene configuration process. It provides an intelligent method that can automatically generate simulation scenes from natural language requirements end to end.
[0027] The execution entity of the intelligent generation method for robot simulation scene based on a large language model proposed in Embodiment 1 of this invention is a computer system, processor, or simulation platform.
[0028] The overall implementation idea of Embodiment 1 of the present invention is: adopting a three-stage process of "requirement analysis → asset scoring → scenario generation". The simulation scenario requirements described by users through natural language are received, then key information is automatically extracted through LLM analysis, multi-dimensional intelligent scoring and recommendation are performed from the asset library, and finally a complete 3D scenario is generated according to a standardized process.
[0029] Figure 1 is a flow chart of the intelligent generation method for robot simulation scenarios based on a large language model proposed in Embodiment 1 of the present invention; Figure 2 is a structural diagram of the intelligent generation method for robot simulation scenarios based on a large language model proposed in Embodiment 1 of the present invention; in combination with Figure 1 and Figure 2 the implementation process of Embodiment 1 of the present invention is described collectively.
[0030] In step S1, the simulation scenario requirements described in natural language are received, the simulation scenario requirements are processed progressively through keyword matching, negative word detection and semantic verification of the large language model in sequence, and the required virtual sensor list and the excluded virtual sensor list for the simulation environment are extracted; Keyword matching: A comprehensive sensor keyword mapping table is maintained in the execution subject first, which adopts a dictionary data structure, wherein keys are standardized sensor types (such as "lidar", "camera"), and values are corresponding Chinese and English synonym lists, that is, all possible keyword lists corresponding to the sensor.
[0031] The design of the keyword mapping table follows the following principles: first, it includes as many synonyms and near-synonyms as possible to cover various expression modes of users. For example, lidar may be expressed in various forms such as "激光雷达", "laser radar", "lidar", "激光", "点云", "laser scan", "rtx_lidar", "range_sensor", "距离传感器" and so on. Second, it supports both Chinese and English expressions to meet the needs of multilingual users. Third, it includes a mixture of common sayings and professional terms, for example, "摄像头" is both a common saying and a commonly used term, while "camera" is an English professional term.
[0032] Then word segmentation is performed on the user input, each word is compared with the keyword mapping table, the simulation scenario requirements are matched based on the preset keyword mapping table, candidate virtual sensors are preliminarily identified, and an initial required virtual sensor list is generated; for example: when a segmented word is found to completely match or be highly similar to a keyword (a similarity measure such as edit distance can be used), the corresponding sensor type is added to the detection list. The accuracy of this step is about 95%, and it can handle most direct, unambiguous user expressions.
[0033] Negative word detection: based on the initial required virtual sensor list, detecting negative words within a context window, moving candidate virtual sensors with negative context out of the required list and adding them to the excluded virtual sensor list, so as to obtain a corrected required virtual sensor list and a corrected excluded virtual sensor list; In the execution subject of the present application, a negative word list is established, including Chinese negative words (such as "bu", "wu", "bu yao", "pai chu", "chu wai", "mei you", "wu xu", etc.) and English negative words (such as "without", "no", "exclude", "do not want", "not need", etc.), Negative word detection adopts at least one of the following two schemes: Mode 1, fixed window detection: for each candidate virtual sensor, locating its character position in the original user input, looking forward to a preset character range before said position, taking said preset character range as the context window, and detecting negative words within said context window. Said preset character range is set to 20 characters for Chinese descriptions, and is set to the word number range corresponding to 20 characters for English descriptions.
[0034] For each candidate virtual sensor, the system locates its position in the original user input. Then it looks forward (to the left) to check the 20-character range before said position (this range is called context window). It checks whether a negative word appears in said window. If a negative word is found, it is determined that the sensor is not required by the user, and the sensor is removed from the required_sensors list (required virtual sensor list) and added to the excluded_sensors list (excluded virtual sensor list); otherwise, the sensor is retained in the required_sensors list.
[0035] Specific example description: the user input is "warehouse patrol, no camera needed, only lidar required". The first-level detection will find both "camera" and "lidar" keywords. Then negative word detection is performed on the camera keyword, the 20 characters before it are checked as "warehouse patrol, no need", and "no need" is found to be a negative word, so camera is moved into the excluded_sensors list. Negative word detection is performed on the lidar keyword, the 20 characters before it are checked as "only", and no negative word is found, so it is retained in the required_sensors list. Finally, the required_sensors list = ['lidar'] and the excluded_sensors list = ['camera'] are obtained.
[0036] Method 2: Dependency parsing dynamic window detection. Dependency parsing is performed on the simulated scenario requirements to identify the predicate verb in the sentence. The scope of its dependency subtree is determined centered on the predicate verb, and this scope serves as a valid search window for negation words. Negation words are detected within this valid search window. The dependency subtree scope includes the subject, object, adverbial, and nested modifiers of the predicate verb.
[0037] The method of using either method one or method two is dynamically selected based on the complexity of the simulation scenario requirements and the availability of dependency parsing. Specifically, when the length of the simulation scenario requirements is greater than a preset length threshold and dependency parsing can successfully construct a dependency parsing tree, method two is used; when the length of the simulation scenario requirements described by the natural language description is less than or equal to the preset length threshold, or when dependency parsing cannot successfully construct a dependency parsing tree, the method reverts to method one.
[0038] For short texts or scenarios where syntactic analysis fails, a fixed window mode is used: based on candidate virtual sensors, negative words are searched forward within a preset character range (the number of words corresponding to 20 Chinese characters or 20 English characters). For long texts or complex sentences, a dependency parsing mode is used: first, dependency parsing is performed on the user input to identify the predicate verb and its dependency subtree in the sentence, and negative words are searched only within the dependency subtree. The system defaults to prioritizing dependency parsing mode. When the text length is below a preset threshold (e.g., 10 words) or dependency parsing fails, it automatically switches to fixed window mode to ensure system robustness and processing efficiency.
[0039] LLM Semantic Validation: Taking the revised list of required virtual sensors and the list of excluded virtual sensors, as well as the requirements of the simulation scenario itself, as input, the large language model is invoked to perform semantic understanding and demand intent determination, which is used to distinguish the user's real needs from background information, hypothetical statements or irrelevant content, and finally outputs the semantically validated list of required virtual sensors and the list of excluded virtual sensors.
[0040] Some statements that seem to require a particular sensor are actually just background information or assumptions, rather than actual simulation requirements. For example, a user might input, "Research shows that mobile robots without LiDAR cannot navigate autonomously." The first and second layer processing might extract a list of required_sensors, s=['lidar'], but this statement is actually discussing the research background of a paper; the user has not explicitly requested the use of LiDAR in the simulation.
[0041] To handle such complex situations, the system sends the results extracted from the first two layers along with the user's original input to the LLM for secondary verification. Based on its language understanding capabilities, the LLM analyzes the user's actual intent and returns a revised list of requirements. This step significantly improves the accuracy of requirement identification, increasing it from around 90% in the first two layers to over 99%.
[0042] In step S2, based on the required virtual sensor list and the excluded virtual sensor list, the candidate robot models in the asset library are weighted and quantitatively scored in multiple dimensions, and the recommended target robot model is determined according to the comprehensive score ranking. This application uses a weighted linear synthesis method to calculate the asset compatibility score, specifically: ; in, Indicates environmental similarity score; This indicates the virtual sensor matching score; This indicates the task suitability score; Indicates historical performance score; , , and These are the corresponding weighting coefficients.
[0043] for example , , , .
[0044] This application , , and The values can be flexibly adjusted according to specific application scenarios. The choice of weights reflects the importance of different dimensions to the final decision. Environmental similarity has the highest weight, reflecting the importance of selecting a robot adapted to the scenario. Sensor matching has the second highest weight, reflecting that meeting the user's sensor needs is a basic requirement. Task adaptability has a lower weight, reflecting that this dimension can be compensated for through configuration and parameter adjustments. Historical performance has the lowest weight because this data may be incomplete or unavailable.
[0045] The environmental similarity score is calculated using the Jaccard set similarity coefficient. The calculation process is as follows: ; in, This represents the set of labels for candidate robot models. These labels are usually predefined in the asset library. For example, {mobile,outdoor,navigation} indicates that the robot is mobile, designed for outdoor use, and used for navigation tasks. The set of labels representing the target scene is inferred by the LLM based on user input or explicitly specified by the user. For example, {warehouse,indoor,large} represents a warehouse environment, an indoor scene, or a large space. In the example above, the intersection of the robot label and the environment label is {indoor} or {}, depending on the specific label of the robot. The union of the two sets is {mobile,outdoor,navigation,warehouse,indoor,large}.
[0046] Calculate the Jaccard coefficient as |intersection| / |union|. For example, if the intersection is {indoor} and the union has 6 elements, then Jaccard = 1 / 6 ≈ 16.7%.
[0047] Multiply the Jaccard coefficient by 100 to obtain the score for that dimension (0-100 points), and then multiply it by a weight of 40% to include it in the overall score.
[0048] The design philosophy behind this dimension is that robots and environments with more similar labels are more likely to be compatible. For example, robots that navigate indoors are more compatible with warehouses (which are typically indoors), while outdoor off-road robots are less compatible with warehouses.
[0049] Sensor matching score measures the degree to which a candidate robot meets the user's sensor requirements; virtual sensor matching score. The calculation process is as follows: ; in, This represents the list of virtual sensors attached to the candidate robot model, for example, ['lidar','camera','imu']. This represents a list of required virtual sensors, such as ['lidar', 'camera']; when the candidate robot model does not contain any sensors from the excluded virtual sensor list... ,otherwise ; In the example above, the coverage rate = 2 / 2 = 100%.
[0050] Also check if the robot contains any excluded sensors (excluded_sensors list). If the robot contains any excluded sensors, the score for this dimension should drop significantly.
[0051] Multiply the coverage rate by 100 to obtain the score for that dimension, and then multiply it by a weight of 35% to include it in the overall score.
[0052] The design philosophy for this dimension is: firstly, to meet the user's explicit needs (required sensors), and secondly, to avoid including sensors that the user has explicitly excluded.
[0053] Task suitability score measures whether the candidate robot's capabilities meet the requirements of the user's task. The calculation process is as follows: ; in, This represents the set of capabilities of a candidate robot model. For example, the "autonomous navigation" task may require two capabilities: {autonomous_navigation, obstacle_avoidance}. autonomous_navigation: autonomous navigation; obstacle_avoidance: obstacle avoidance.
[0054] This represents the total number of capabilities required for the task, and is usually predefined in the asset library, such as {autonomous_navigation,obstacle_avoidance,mapping}; mapping: map building.
[0055] In the example above, the matching rate = 2 / 2 = 100%.
[0056] Multiply the match rate by 100 to get the score for that dimension, and then multiply it by a weight of 15% to include it in the overall score.
[0057] This dimension has a lower weight because robots lacking certain capabilities may be able to compensate for this through configuration or simulation parameter adjustments.
[0058] Historical performance score is based on the candidate robot's performance data in historical simulations. The calculation process is as follows: ; in, Indicates the first Historical performance indicators; This represents a function that standardizes the index to the range of 0-1. Indicates the first Weights of historical performance indicators; This represents the total number of historical performance metrics; the historical performance metrics include at least one or more of the following: simulation task success rate, average completion time ranking, and user rating.
[0059] This dimension has the lowest weight because historical data may be insufficient for new assets, and historical performance may lack reference value for new simulation tasks.
[0060] The system determines the recommended target robot model based on a comprehensive score ranking; specifically, the top-ranked assets are the system's recommended solutions. The system can output the top N candidates (e.g., the top 3) for the user to make a final selection; or it can automatically select the top-ranked asset as the recommended solution.
[0061] In step S3, asset pre-verification, scene reset, resource loading, object configuration, and transaction commit operations are performed in the robot simulation platform to generate a simulation scene. In the object configuration operation, the placement position of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is adopted to complete the instantiation of the target robot model and the mounting of its virtual sensors.
[0062] The present invention automatically calculates the placement position of props based on the reachable area and spatial semantics of the loaded environmental resources. Specifically, it analyzes the navigation grid in the environmental resource file, extracts the ground polygon as the reachable area, calculates the candidate positions for prop placement based on the reachable area, and calls the physical query interface of the simulation platform to perform collision detection to ensure that the props do not penetrate the static environmental grid. The candidate positions are calculated based on at least one of the following rules: salient position rules based on environmental semantics, directional layout rules based on task requirements, and optimization rules based on obstacle avoidance.
[0063] The robot simulation platform used in this invention is the Isaac Sim platform.
[0064] In Embodiment 1 of this invention, the asset pre-verification operation is used to confirm that all assets to be loaded exist in the asset database; the specific steps include: 1. Iterate through the LLM-recommended asset list, including recommended environments, robots, and all props.
[0065] 2. For each asset, query it in the asset repository using its name or ID. Queries typically invoke interface functions of the asset repository management module.
[0066] 3. If the query returns valid asset information (including file path, metadata, etc.), the asset is marked as "verified".
[0067] 4. If any asset cannot be found in the asset database, an exception is thrown and detailed error information is logged, which is then returned to the user and LLM for processing.
[0068] 5. The system will only proceed to the next stage after all assets have been verified.
[0069] The scene reset operation deletes all nodes created by historical generation operations in the current robot simulation platform scene, restoring the scene to its initial state. The specific steps include: 1. Traverse all Prim nodes in the Isaac Sim scene graph. The scene graph is a tree structure, and the root node is usually " / ".
[0070] 2. Identify and record all nodes created by previous scene generation. These nodes typically follow specific naming conventions, such as " / Robot_", " / Prop_", " / Environment", etc.
[0071] 3. For each node that needs to be deleted, call the Isaac Sim delete API to completely remove the node and all its child nodes.
[0072] 4. The clearing operation is irreversible; in practical applications, it should be ensured that no nodes are accidentally deleted. The system can choose to retain certain global nodes (such as lights, cameras, etc.) and only delete nodes related to the simulation scene.
[0073] The resource loading operation is used to load environment resources into the scene root node by reference; the specific steps include: 1. Obtain the complete file path to the recommended environment from the asset repository. This path typically includes a network path (such as / omniverse: / / nucleus / ...) or a local path.
[0074] 2. Create an environment reference using Isaac Sim's stage API. In the Omniverse platform, references are an efficient resource loading mechanism that allows multiple scenes to share the same resource.
[0075] 3. Add the environment reference to the root node of the scene and create a Prim node named " / Environment" with its prim spec pointing to the USD file of the environment.
[0076] 4. Set the initial transformations of the environment (position, rotation, scaling). The environment is typically placed at the world coordinate origin, without rotation or scaling.
[0077] 5. Confirm that the environment has loaded successfully and check for any warnings or error messages.
[0078] The object configuration operation adopts a delayed creation strategy, first saving the configuration information of the target robot model in the internal data structure without immediately creating its scene description file reference, and automatically calculating the placement position of props based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props. Robot configuration section: 1. Obtain recommended robot information, including robot type, sensor configuration, initial position, etc.
[0079] 2. Store this configuration information in an internal data structure (usually a dictionary or object), but delay the creation of the robot's USD reference. This delayed creation strategy is a key design decision. The reason is that the delay allows the system to dynamically adjust the robot's position, orientation, and other parameters in later stages (such as the trajectory planning stage) without having to reload the entire robot USD.
[0080] 3. Record robot-related information in the created_objects dictionary, including prim_path (e.g., " / Robot_0"), asset_path (USD file path), metadata (metadata, such as sensor list, quality, etc.), and usd_created flag (initial value is False, indicating that USD has not yet been created).
[0081] Item configuration section: 1. Iterate through the list of recommended items and process each item.
[0082] 2. For each prop, first calculate its position and rotation in 3D space. Position calculation usually follows predefined layout rules, such as linear layout, mesh layout, etc. In a linear layout, the position of the j-th prop is calculated as follows: x-coordinate: (j-total_props / 2)*spacing (where j is numbered starting from 0, and spacing is the distance between adjacent props); y-coordinate: fixed value (e.g., 0); z-coordinate: fixed value (e.g., 0.5); This formula ensures that the props are evenly distributed in the scene, centered on the robot's starting position.
[0083] 3. For each item, the system retrieves its USD file path from the asset library, creates a reference to the item, and adds it to the scene.
[0084] 4. Set the transformation matrix of the prop, that is, apply its position, rotation, and scaling to the scene.
[0085] 5. The system records relevant information for each item in the created_objects dictionary.
[0086] The transaction commit operation is used to package all the commands to be executed accumulated in the object configuration operation into an atomic transaction and submit them to the simulation engine for execution.
[0087] 1. Collect all commands to be executed generated during the verification, clearing, loading, and configuration phases. These commands include node creation, attribute setting, transformation application, etc.
[0088] 2. Package all commands into an atomic transaction. This ensures that all commands either succeed completely or fail completely, avoiding inconsistencies in intermediate states caused by partial execution.
[0089] 3. Submit the transaction to the Isaac Sim simulation engine. The simulation engine executes all commands in the next frame.
[0090] 4. Obtain feedback on the execution results and check for any errors or warnings.
[0091] 5. If the execution is successful, the system returns the generated scene information (including a list of objects in the scene, their locations and attributes, etc.).
[0092] This invention proposes and implements a five-stage standard process: verification, clearing, loading, configuration, and commit. Each stage has clearly defined responsibilities and inputs / outputs. The verification stage ensures the integrity and validity of assets, preventing subsequent errors. The clearing stage employs a clear naming convention and a thorough deletion strategy to avoid interference from old objects. The loading stage uses a USD reference mechanism to improve resource loading efficiency. The configuration stage introduces a delayed creation strategy, allowing subsequent stages to dynamically adjust parameters. The commit stage employs an atomic transaction mechanism to guarantee consistency.
[0093] Example 2 Based on the intelligent generation method for robot simulation scenes based on a large language model proposed in Embodiment 1 of this invention, Embodiment 2 of this invention also proposes an intelligent generation system for robot simulation scenes based on a large language model. Figure 3 This is a schematic diagram of the intelligent generation system for robot simulation scenes based on a large language model proposed in Embodiment 2 of the present invention. The system includes: The requirements analysis module receives simulation scenario requirements described in natural language and processes them progressively through keyword matching, negative word detection, and semantic verification using a large language model to extract a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment. The asset scoring module is used to perform weighted quantitative scoring on candidate robot models in the asset library across multiple dimensions based on the required virtual sensor list and the excluded virtual sensor list, and to determine the recommended target robot model based on the comprehensive score ranking. The scene generation module is used to perform asset pre-verification, scene reset, resource loading, object configuration and transaction commit operations in the robot simulation platform to generate a simulation scene. In the object configuration operation, the placement position of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is adopted to complete the instantiation of the target robot model and the mounting of its virtual sensors.
[0094] In the requirements analysis module, the simulation scenario requirements are progressively processed through keyword matching, negative word detection, and semantic verification using a large language model, extracting a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment. Specifically: Based on a preset keyword mapping table, the simulation scenario requirements are matched to initially identify candidate virtual sensors and generate an initial list of required virtual sensors. Based on the initial list of required virtual sensors, negative words are detected in the context window. Candidate virtual sensors with negative contexts are removed from the required list and added to the list of excluded virtual sensors, resulting in the revised list of required virtual sensors and the list of excluded virtual sensors. Using the revised list of required virtual sensors and the list of excluded virtual sensors, as well as the requirements of the simulation scenario itself, as input, a large language model is invoked to perform semantic understanding and demand intent determination, which is used to distinguish between the user's real needs and background information, hypothetical statements or irrelevant content, and finally outputs a semantically verified list of required virtual sensors and a list of excluded virtual sensors.
[0095] The keyword mapping table uses a dictionary data structure, with the key being the normalized virtual sensor type and the value being a list of corresponding Chinese and English synonyms.
[0096] The first method for negative word detection is as follows: For each candidate virtual sensor, locate its character position in the original user input, look forward to the preset character range before that position, use the preset character range as a context window, and detect negative words within the context window.
[0097] The second method for detecting negation words is as follows: perform dependency parsing on the simulation scenario requirements, identify the predicate verb in the sentence, determine the scope of its dependency subtree centered on the predicate verb, use the scope of the dependency subtree as the effective search window for negation words, and detect negation words within the effective search window; wherein, the scope of the dependency subtree includes the subject, object, adverbial and their nested modifying components of the predicate verb.
[0098] This invention employs a three-layer progressive sensor requirement extraction mechanism: the first layer performs basic identification through a comprehensive keyword mapping table, supporting Chinese and English synonyms; the second layer detects negative words within a context window, moving sensors with negative contexts from the required list to the exclusion list of virtual sensors; the third layer utilizes a large language model for semantic understanding and requirement intent determination, capable of distinguishing between real requirements and background information, hypothetical statements, or irrelevant content. This three-layer progressive design gives the system strong fault tolerance; even if an identification error occurs in one layer, subsequent layers still have the opportunity to correct it, significantly improving the accuracy of sensor requirement identification compared to traditional methods.
[0099] In the asset scoring module, based on the required virtual sensor list and the excluded virtual sensor list, the candidate robot models in the asset library are weighted and quantitatively scored across multiple dimensions. Specifically: ; in, Indicates environmental similarity score; This indicates the virtual sensor matching score; This indicates the task suitability score; Indicates historical performance score; , , and These are the corresponding weighting coefficients.
[0100] Environmental similarity score The calculation process is as follows: ; in, The set of labels representing candidate robot models; A set of tags representing the target scene; Virtual sensor matching score The calculation process is as follows: ; in, This represents the list of virtual sensors attached to the candidate robot model; This indicates the list of required virtual sensors; when the candidate robot model does not contain any sensors from the excluded virtual sensor list... ,otherwise ; Task suitability score The calculation process is as follows: ; in, This represents the set of capabilities of the candidate robot model; Indicates the total number of capabilities required for the task; Historical performance rating The calculation process is as follows: ; in, Indicates the first Historical performance indicators; This represents a function that standardizes the index to the range of 0-1. Indicates the first Weights of historical performance indicators; This represents the total number of historical performance metrics; the historical performance metrics include at least one or more of the following: simulation task success rate, average completion time ranking, and user rating.
[0101] This invention employs a multi-dimensional weighted quantitative scoring system to comprehensively evaluate candidate robot models across four dimensions: environmental similarity, virtual sensor matching degree, task adaptability, and historical performance. Environmental similarity is calculated using the Jaccard similarity coefficient to determine the degree of matching between robot labels and scene labels; sensor matching degree is based on sensor coverage and excluded sensor detections; task adaptability is based on capability matching rate; and historical performance is based on historical simulation data. This weighted comprehensive scoring across four dimensions upgrades asset selection from passive search to proactive intelligent recommendation, resulting in a significant improvement in asset selection accuracy compared to traditional methods. Furthermore, the weighting coefficients can be flexibly adjusted according to specific application scenarios to adapt to different application needs.
[0102] In the scene generation module, the asset pre-verification operation is used to confirm that all assets to be loaded exist in the asset library; The scene reset operation is used to delete all nodes created by historical generation operations in the current robot simulation platform scene, restoring the scene to its initial state; The resource loading operation is used to load environmental resources into the scene root node by reference. The object configuration operation adopts a delayed creation strategy, which first saves the configuration information of the target robot model in the internal data structure without immediately creating its scene description file reference, and automatically calculates the placement position of the props based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of the props. The transaction commit operation is used to package all commands to be executed accumulated in the object configuration operation into an atomic transaction and submit them to the simulation engine for execution.
[0103] The placement of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources. Specifically, the navigation mesh in the environmental resource file is analyzed, the ground polygon is extracted as the reachable area, the candidate positions for prop placement are calculated based on the reachable area, and the physical query interface of the simulation platform is called to perform collision detection to ensure that the props do not penetrate the static mesh of the environment. The candidate positions are calculated based on at least one of the following rules: salient position rules based on environmental semantics, directional layout rules based on task requirements, and optimization rules based on obstacle avoidance.
[0104] This invention achieves a complete automated conversion from unstructured natural language requirements to structured simulation scenarios. The entire process has the following characteristics: First, it is highly deterministic, as the same input always generates the same scenario, which is easy to reproduce and verify; second, it is standardized, adopting a five-stage standardized process of verification-clearing-loading-configuration-commit, with standardized operation; third, it is atomic, with transaction commit using an atomic mechanism to ensure the consistency and integrity of scenario generation; and fourth, it is fully automated, requiring no manual intervention, with the entire process from input to output being automated.
[0105] The description of the relevant parts of the intelligent generation system for robot simulation scenes based on a large language model provided in Embodiment 2 of this invention can be found in the detailed description of the corresponding parts in the intelligent generation method for robot simulation scenes based on a large language model provided in Embodiment 1 of this application, and will not be repeated here.
[0106] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that the elements inherent in a process, method, article, or apparatus that includes a list of elements are included. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Additionally, portions of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.
[0107] While specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art can make other modifications or variations based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for intelligent generation of robot simulation scenes based on a large language model, characterized in that, Includes the following steps: The system receives simulation scenario requirements described in natural language and processes them progressively through keyword matching, negative word detection, and semantic verification using a large language model to extract a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment. The simulation scenario requirements are progressively processed through keyword matching, negative word detection, and semantic verification using a large language model. This process extracts a list of necessary virtual sensors and a list of excluded virtual sensors for the simulation environment. Specifically: the simulation scenario requirements are matched against a pre-defined keyword mapping table to initially identify candidate virtual sensors, generating an initial list of necessary virtual sensors; based on this initial list, negative words are detected within a context window, and candidate virtual sensors with negative contexts are removed from the necessary list and added to the excluded list, resulting in a revised list of necessary and excluded virtual sensors; using the revised list of necessary and excluded virtual sensors, along with the simulation scenario requirements themselves, as input, a large language model is invoked for semantic understanding and requirement intent determination. This distinguishes between the user's actual needs and background information, hypothetical statements, or irrelevant content, ultimately outputting a semantically verified list of necessary and excluded virtual sensors. Based on the required virtual sensor list and the excluded virtual sensor list, the candidate robot models in the asset library are weighted and quantitatively scored across multiple dimensions, and the recommended target robot models are determined by ranking them according to the comprehensive scores. In the robot simulation platform, asset pre-verification, scene reset, resource loading, object configuration, and transaction commit operations are performed to generate a simulation scene. In the object configuration operation, the placement position of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is adopted to complete the instantiation of the target robot model and the mounting of its virtual sensors. The placement of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources. Specifically, the navigation mesh in the environmental resource file is analyzed, ground polygons are extracted as reachable areas, candidate placement positions of props are calculated based on the reachable areas, and collision detection is performed by calling the physical query interface of the simulation platform to ensure that the props do not penetrate the static mesh of the environment. The calculation of the candidate positions is based on at least one of the following rules: salient position rules based on environmental semantics, directional layout rules based on task requirements, and optimization rules based on obstacle avoidance.
2. The intelligent generation method for robot simulation scenes based on a large language model according to claim 1, characterized in that, The keyword mapping table adopts a dictionary data structure, with the key being the normalized virtual sensor type and the value being the corresponding list of Chinese and English synonyms.
3. The intelligent generation method for robot simulation scenes based on a large language model according to claim 1, characterized in that, One method for detecting negation words is: For each candidate virtual sensor, locate its character position in the original user input, look forward to the preset character range before that position, use the preset character range as a context window, and detect negative words within the context window.
4. The intelligent generation method for robot simulation scenes based on a large language model according to claim 1, characterized in that, Another way to detect negation words is: Dependency parsing is performed on the simulation scenario requirements to identify the predicate verb in the sentence. The scope of its dependency subtree is determined with the predicate verb as the center. The scope of the dependency subtree is used as a valid search window for negation words. Negation words are detected within the valid search window. The scope of the dependency subtree includes the subject, object, adverbial and nested modifiers of the predicate verb.
5. The intelligent generation method for robot simulation scenes based on a large language model according to claim 1, characterized in that, Based on the required virtual sensor list and the excluded virtual sensor list, the candidate robot models in the asset library are weighted and quantitatively scored across multiple dimensions. Specifically: ; in, Indicates environmental similarity score; This indicates the virtual sensor matching score; This indicates the task suitability score; Indicates historical performance score; , , and These are the corresponding weighting coefficients.
6. The intelligent generation method for robot simulation scenes based on a large language model according to claim 5, characterized in that, Environmental similarity score The calculation process is as follows: ; in, The set of labels representing candidate robot models; A set of tags representing the target scene; Virtual sensor matching score The calculation process is as follows: ; in, This represents the list of virtual sensors attached to the candidate robot model; This indicates the list of required virtual sensors; when the candidate robot model does not contain any sensors from the excluded virtual sensor list... ,otherwise ; Task suitability rating The calculation process is as follows: ; in, This represents the set of capabilities of the candidate robot model; Indicates the total number of capabilities required for the task; Historical performance rating The calculation process is as follows: ; in, Indicates the first Historical performance indicators; This represents a function that standardizes the index to the range of 0-1. Indicates the first Weighting of historical performance indicators; This represents the total number of historical performance metrics; the historical performance metrics include at least one or more of the following: simulation task success rate, average completion time ranking, and user rating.
7. The intelligent generation method for robot simulation scenes based on a large language model according to claim 1, characterized in that, During the generation of simulation scenes: The asset pre-verification operation is used to confirm that all assets to be loaded exist in the asset library; The scene reset operation is used to delete all nodes created by historical generation operations in the current robot simulation platform scene, restoring the scene to its initial state; The resource loading operation is used to load environmental resources into the scene root node by reference. The object configuration operation adopts a delayed creation strategy, which first saves the configuration information of the target robot model in the internal data structure without immediately creating its scene description file reference, and automatically calculates the placement position of the props based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of the props. The transaction commit operation is used to package all the commands to be executed accumulated in the object configuration operation into an atomic transaction and submit them to the simulation engine for execution.
8. A robot simulation scene intelligent generation system based on a large language model, used to execute the robot simulation scene intelligent generation method based on a large language model as described in any one of claims 1 to 7, characterized in that, include: The requirements analysis module is used to receive simulation scenario requirements described in natural language, and progressively process the simulation scenario requirements through keyword matching, negative word detection and large language model semantic verification to extract the list of necessary virtual sensors and the list of excluded virtual sensors for the simulation environment. The asset scoring module is used to perform weighted quantitative scoring on candidate robot models in the asset library across multiple dimensions based on the required virtual sensor list and the excluded virtual sensor list, and to determine the recommended target robot model based on the comprehensive score ranking. The scene generation module is used to perform asset pre-verification, scene reset, resource loading, object configuration and transaction commit operations in the robot simulation platform to generate a simulation scene. In the object configuration operation, the placement position of props is automatically calculated based on the reachable area and spatial semantics of the loaded environmental resources to complete the intelligent layout of props, and a delayed creation strategy is adopted to complete the instantiation of the target robot model and the mounting of its virtual sensors.
Citation Information
Patent Citations
Three-dimensional body simulation environment generation method, device and equipment based on language guidance
CN120633132A
Simulation scene generation method and device based on large language model
CN121365503A