Somatosensory intelligent robot control method and system based on virtual-real fusion learning platform
By combining data collected by an optical positioning module in a virtual-real fusion learning platform, a simulation model of the target robot in a real scene is established, which solves the problem of insufficient synchronization between virtual and real environment data in existing platforms and improves the accuracy of robot control and the stability of task execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-26
AI Technical Summary
In existing virtual-real fusion learning platforms, simulation data and real-world data are processed through independent data channels, lacking a two-way real-time synchronization mechanism between the virtual and real environments. This leads to alignment deviations in time and space dimensions, resulting in performance degradation of simulation training models in real deployments. It is difficult to achieve real-time supervision and optimization of the virtual scene's execution process of the real robot, thus affecting the accuracy of robot control.
By using an embodied intelligent robot control method based on a virtual-real fusion learning platform, and combining optical positioning modules to collect actual posture and position data of the target robot in real scenarios, a simulation model of the target robot in the target simulation scenario is established. Data interaction and information synchronization are achieved using the ROS communication interface, and the training process is optimized to improve the accuracy of the control model.
This achieves spatial alignment between the training basis of the target robot control model and real-world scene data, reduces the deviation between simulation training results and real-world operating states, and improves the accuracy of robot control and the stability of task execution.
Smart Images

Figure CN122274944A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of virtual-real fusion simulation of mobile robots, robot closed-loop control and embodied intelligence training technology, and in particular to an embodied intelligent robot control method and system based on a virtual-real fusion learning platform. Background Technology
[0002] Currently, in the research and development and training of robots, to reduce the safety risks and costs brought by real-world testing, virtual scenarios are typically constructed using virtual-real fusion learning platforms such as Habitat, iGibson, and Arena. These virtual scenarios are then used for task training and algorithm verification. However, in existing virtual-real fusion learning platforms, simulation data and real-world data are usually processed through independent data channels, lacking a two-way real-time synchronization mechanism between the virtual and real environments. This results in alignment discrepancies in time and space dimensions, severely exacerbating the simulation-to-real gap, and causing a significant performance degradation of simulation-trained models in real-world deployments. Existing platforms rarely support real-time human-machine interaction and dynamic task switching, making it impossible to achieve real-time supervision and optimization of the real robot's execution process within the virtual scenario. This makes it difficult to support the training and deployment of robots for complex, long-cycle tasks. Furthermore, the dynamic models in simulation scenarios are usually built based on idealized physical assumptions, causing discrepancies between the motion state data and sensor data obtained by the robot in the simulation environment and the data in the real physical environment. Therefore, the control models trained on existing virtual-real fusion learning platforms often exhibit significant deviations in action execution when deployed in real-world environments, affecting the practical application effectiveness of the robot system.
[0003] Therefore, improving the accuracy of robot control has become an urgent technical problem to be solved. Summary of the Invention
[0004] The main objective of this application is to propose an embodied intelligent robot control method and scene simulation system based on a virtual-real fusion learning platform, aiming to improve the accuracy of robot control by optimizing the training process.
[0005] To achieve the above objectives, a first aspect of this application proposes a method for controlling an embodied intelligent robot based on a virtual-real fusion learning platform, applied to a terminal of a scene simulation system. The scene simulation system further includes a target robot and an optical positioning module. The terminal is electrically connected to both the target robot and the optical positioning module. The terminal is equipped with a virtual-real fusion learning platform, which has an interactive interface displaying preset task types and preset simulation scenarios. The virtual-real fusion learning platform is configured to allow the simulation model of the target robot to move within the preset simulation scenario. The method includes: In response to the selection of the preset task type and preset simulation scenario displayed on the interactive interface, the target task type and target simulation scenario are obtained, and the target simulation scenario is invoked through the virtual-real fusion learning platform, and the preset simulation model of the target robot in the target simulation scenario is loaded. Collect the actual posture data of the target robot in a real scene; The optical positioning module acquires the actual position data of the target robot in the real scene. The simulation modeling is driven by the actual posture data and the actual position data to obtain the simulation pose data of the target robot in the target simulation scene. The target control model is obtained by training a preset initial control model based on the target task type and the simulation pose data. The target robot is controlled based on the target control model.
[0006] In some embodiments, the optical positioning module includes a motion capture camera, the real-world scene has a first world coordinate system, and acquiring the actual position data of the target robot in the real-world scene through the optical positioning module includes: The motion capture camera acquires images containing the target robot. The image is used to identify the actual position of the target robot in the first world coordinate system. The target simulation scene has a reference coordinate system. The step of establishing a simulation model of the target robot in the target simulation scene based on the actual posture data and the actual position data, to obtain the simulation pose data of the target robot in the target simulation scene, includes: Determine the first transformation mapping relationship between the reference coordinate system and the first world coordinate system; The actual pose data and the actual position data are mapped to coordinates according to the first transformation mapping relationship to obtain the simulated pose data.
[0007] In some embodiments, the optical positioning module includes an LED position sensing screen, the target robot moves on the surface of the LED position sensing screen, the real scene has a second world coordinate system, and the acquisition of the actual position data of the target robot in the real scene through the optical positioning module includes: The actual position data of the target robot in the second world coordinate system is obtained through the LED position sensing screen; The target simulation scene has a reference coordinate system. The step of establishing a simulation model of the target robot in the target simulation scene based on the actual posture data and the actual position data, to obtain the simulation pose data of the target robot in the target simulation scene, includes: Determine the scaling ratio between the LED position sensing screen and the target simulation scene, and perform position calculation on the actual position data according to the scaling ratio to obtain the simulation position data of the target robot in the target simulation scene; Determine the second transformation mapping relationship between the reference coordinate system and the second world coordinate system; The actual attitude data is mapped to coordinates according to the second transformation mapping relationship to obtain the simulated attitude data; The simulated pose data is determined based on the simulated posture data and the simulated position data.
[0008] In some embodiments, the interactive interface includes a first display area and a second display area, the first display area including a task type sub-area, and the step of obtaining the target task type and target simulation scene in response to a selection operation of a preset task type and a preset simulation scene displayed on the interactive interface includes: In response to the selection operation of the task type sub-region on the interactive interface, a target type sub-region is selected from the task type sub-region, and the task type corresponding to the target type sub-region is taken as the target task type; The second display area is switched according to the target task type to obtain a scene preview sub-area; In response to the selection operation of the scene preview sub-region on the interactive interface, a target scene sub-region is obtained, and the preset simulation scene corresponding to the target scene sub-region is taken as the target simulation scene.
[0009] In some embodiments, after obtaining a target scene sub-region in response to a selection operation on the scene preview sub-region on the interactive interface, and using a preset simulation scene corresponding to the target scene sub-region as the target simulation scene, the method further includes: The user obtains instructions and operations input to the virtual-real fusion learning platform through the interactive interface, and obtains real-time tasks based on the instructions and operations. Control the target robot to execute the real-time task; During the execution of the real-time task, the simulation pose data is updated, and a real-time animation of the simulation model is generated based on the updated simulation pose data. The real-time animation is then displayed through the interactive interface.
[0010] In some embodiments, after obtaining a target scene sub-region in response to a selection operation on the scene preview sub-region on the interactive interface, and using a preset simulation scene corresponding to the target scene sub-region as the target simulation scene, the method further includes: If the target task type is data acquisition, then the first display area is updated to obtain a configuration parameter type sub-area; In response to the selection operation of the configuration parameter type sub-region on the interactive interface, and when the parameter type corresponding to the selected configuration parameter type sub-region is pedestrian configuration, the number of pedestrians is obtained; The pedestrian movement trajectory is randomly generated based on the number of pedestrians, and a simulation animation is generated based on the pedestrian movement trajectory. The simulation animation is then displayed through the interactive interface.
[0011] To achieve the above objectives, a second aspect of this application provides a terminal, the terminal including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.
[0012] This application proposes an embodied intelligent robot control method and terminal based on a virtual-real fusion learning platform. Through a scene simulation system, it acquires the target task type and target simulation scene in response to the selection of preset task types and preset simulation scenes displayed on the interactive interface. The target simulation scene is then invoked through the virtual-real fusion learning platform, which presets multiple simulation scenes to achieve diverse and reproducible training scenarios. Combining the actual posture data of the target robot in the real scene and the actual position data of the target robot in the real scene obtained through an optical positioning module, a simulation model of the target robot in the target simulation scene is established, and the simulation pose data of the target robot in the target simulation scene is obtained. Then, based on the target task type and simulation pose data, a preset initial control model is trained to obtain a target control model. Finally, the target robot is controlled based on the target control model. The method in this embodiment can introduce the actual posture data and actual position data of the target robot in the real scene into the control model training process in the target simulation scene. This makes the training basis of the target control model include not only the task conditions in the target simulation scene, but also the actual posture data and actual position data of the target robot in the real scene. This achieves spatial alignment between the real pose data and the virtual pose data in the virtual-real fusion learning platform, thereby reducing the deviation between the simulation training results of the target control model and the running state of the target robot in the real scene. Ultimately, this improves the accuracy of the target robot control and reduces the action execution deviation of the target robot in the actual deployment process.
[0013] To achieve the above objectives, a third aspect of this application proposes a scene simulation system, the system comprising: The scene simulation system includes a terminal, an optical positioning module, and a target robot as described in the second aspect. The terminal is equipped with a virtual-real fusion learning platform, which has an interactive interface. The terminal communicates with the target robot and the optical positioning module through a ROS communication interface.
[0014] In some embodiments, the optical positioning module includes a motion capture camera for acquiring images containing the target robot; the terminal is also used to identify the target robot's location using the images to obtain the target robot's actual location data.
[0015] In some embodiments, the optical positioning module includes an LED position sensing screen, the target robot is used to move on the surface of the LED position sensing screen, and the LED position sensing screen is used to acquire the actual position data of the target robot.
[0016] The scenario simulation system proposed in this application enables communication between the terminal, the target robot, and the optical positioning module through the ROS communication interface. This allows for data interaction and information synchronization between real-world scenario data and the virtual-real fusion learning platform within the same system architecture, achieving high alignment of data in the time dimension. The optical positioning module enables high-precision positioning of the target robot, achieving alignment of pose data in the spatial dimension. This avoids the data fragmentation problem caused by processing simulation data and real-world scenario data through independent data channels. Ultimately, the virtual-real fusion learning platform can perform task verification and model training based on real-world scenario data that more closely resembles actual operating conditions, improving data consistency between the virtual-real fusion learning platform and the real environment, and ultimately enhancing the control accuracy and task execution stability of the target robot in real-world scenarios. Attached Figure Description
[0017] Figure 1 This is a flowchart of the embodied intelligent robot control method based on a virtual-real fusion learning platform provided in the embodiments of this application; Figure 2 This is an optional scene diagram of the scene simulation system provided in this application embodiment; Figure 3 This is a schematic diagram of an optional interactive interface provided in an embodiment of this application; Figure 4 yes Figure 1 The flowchart of step S101 in the text; Figure 5This is another flowchart of the embodied intelligent robot control method based on a virtual-real fusion learning platform provided in the embodiments of this application; Figure 6 This is a schematic diagram of another optional interactive interface provided in the embodiments of this application; Figure 7 yes Figure 1 The flowchart of step S104 in the process; Figure 8 yes Figure 1 Another flowchart of step S104 in the process; Figure 9 This is another flowchart of the embodied intelligent robot control method based on a virtual-real fusion learning platform provided in the embodiments of this application; Figure 10 This is a schematic diagram of the hardware structure of the terminal provided in the embodiments of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Currently, in the process of robot research and training, theoretically, collecting data from robots in real-world scenarios and training them based on that data allows the training results to most closely approximate real-world expectations. However, data collection in real-world scenarios carries high costs and risks. Specifically, collisions between robots and obstacles in real-world environments can damage equipment and the environment, and even pose safety hazards to people. Furthermore, real-world scenarios are difficult to reproduce, making it challenging to obtain training samples covering a long-tail distribution (such as rare but crucial pedestrian obstacle avoidance scenarios). To mitigate the safety risks and costs associated with real-world testing, existing technologies typically utilize simulation platforms such as Habitat, iGibson, and Arena to construct virtual scenarios for task training and algorithm verification. However, in existing simulation platforms, simulation data and real-world data are usually processed through independent data channels, lacking a two-way real-time synchronization mechanism between virtual and real environments. This results in alignment discrepancies in time and space dimensions, severely exacerbating the simulation-to-real gap, and causing a significant performance degradation of simulation-trained models in real-world deployments. Existing platforms rarely support real-time human-robot in-loop intervention and dynamic task switching, making it impossible to achieve real-time monitoring and optimization of the robot's execution process in a virtual environment. This hinders the training and deployment of robots for complex, long-cycle tasks. Furthermore, the dynamic models in simulation scenarios are typically built based on idealized physical assumptions, leading to discrepancies between the motion state data and sensor data obtained by the robot in the simulation environment and the data from the real physical environment. Therefore, control models trained on existing simulation platforms often exhibit significant deviations in motion execution when deployed in real-world environments, impacting the practical application effectiveness of the robot system.
[0022] Therefore, improving the accuracy of robot control has become an urgent technical problem to be solved.
[0023] Based on this, embodiments of this application provide an embodied intelligent robot control method and scene simulation system based on a virtual-real fusion learning platform, aiming to improve the accuracy of robot control by optimizing the training process.
[0024] The embodied intelligent robot control method and scene simulation system based on the virtual-real fusion learning platform provided in this application relate to the fields of mobile robot virtual-real fusion simulation, robot closed-loop control and embodied intelligent training technology, which are specifically described through the following embodiments.
[0025] The embodied intelligent robot control method based on a virtual-real fusion learning platform provided in this application relates to the fields of mobile robot virtual-real fusion simulation, robot closed-loop control, and embodied intelligent training technology. The embodied intelligent robot control method based on a virtual-real fusion learning platform provided in this application is applied to the terminal of a scene simulation system. The scene simulation system also includes a target robot and an optical positioning module, and the terminal is electrically connected to both the target robot and the optical positioning module. The terminal is equipped with a virtual-real fusion learning platform, which has an interactive interface displaying preset task types and preset simulation scenes. The virtual-real fusion learning platform is configured to allow the simulation modeling of the target robot to move within the preset simulation scene.
[0026] Figure 1 This is an optional flowchart of the embodied intelligent robot control method based on a virtual-real fusion learning platform provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.
[0027] Step S101: In response to the selection operation of the preset task type and preset simulation scene displayed on the interactive interface, the target task type and target simulation scene are obtained, and the target simulation scene is called through the virtual-real fusion learning platform, and the preset simulation model of the target robot in the target simulation scene is loaded.
[0028] Step S102: Collect the actual posture data of the target robot in a real scene.
[0029] Step S103: Obtain the actual position data of the target robot in the real scene through the optical positioning module.
[0030] Step S104: Drive the simulation modeling based on the actual posture data and actual position data to obtain the simulation pose data of the target robot in the target simulation scene.
[0031] Step S105: Train the preset initial control model according to the target task type and simulation pose data to obtain the target control model.
[0032] Step S106: Perform robot control on the target robot based on the target control model.
[0033] Steps S101 to S106, as illustrated in this embodiment, involve using a scene simulation system to obtain the target task type and target simulation scene in response to the selection of preset task types and preset simulation scenes displayed on the interactive interface. The target simulation scene is then invoked through a virtual-real fusion learning platform, which presets multiple simulation scenes to diversify and reproducibly achieve the training scenarios. Combining the actual posture data of the target robot in the real scene and the actual position data of the target robot in the real scene obtained through the optical positioning module, a simulation model of the target robot in the target simulation scene is established, and the simulation pose data of the target robot in the target simulation scene is obtained. Then, based on the target task type and simulation pose data, a preset initial control model is trained to obtain a target control model. Finally, robot control is performed on the target robot based on the target control model. The method in this embodiment can introduce the actual posture data and actual position data of the target robot in the real scene into the control model training process in the target simulation scene. This makes the training basis of the target control model include not only the task conditions in the target simulation scene, but also the actual posture data and actual position data of the target robot in the real scene. This achieves spatial alignment between the real pose data and the virtual pose data in the virtual-real fusion learning platform, thereby reducing the deviation between the simulation training results of the target control model and the running state of the target robot in the real scene. Ultimately, this improves the accuracy of the target robot control and reduces the action execution deviation of the target robot in the actual deployment process.
[0034] Since the embodied intelligent robot control method based on the virtual-real fusion learning platform in this application relies on the aforementioned scene simulation system for hardware implementation, the scene simulation system of this application will be described in detail first. Please refer to... Figure 2 , Figure 2 This is an optional scene diagram of the scene simulation system provided in the embodiments of this application.
[0035] The scene simulation system in this embodiment includes a terminal, an optical positioning module, and a target robot. The terminal is equipped with a virtual-real fusion learning platform, which has an interactive interface. The terminal communicates with the target robot and the optical positioning module via a ROS communication interface (i.e., the ROS middleware in the figure).
[0036] The scenario simulation system illustrated in this application uses a ROS communication interface to enable communication between the terminal, the target robot, and the optical positioning module. This allows for data interaction and information synchronization between real-world scenario data and the virtual-real fusion learning platform within the same system architecture, achieving high data alignment in the time dimension. The optical positioning module enables high-precision positioning of the target robot, aligning pose data in the spatial dimension. This avoids the data fragmentation problem caused by processing simulation data and real-world scenario data through independent data channels. Ultimately, the virtual-real fusion learning platform can perform task verification and model training based on real-world scenario data that more closely resembles actual operating conditions, improving data consistency between the simulation scenario and the real environment. This ultimately enhances the control accuracy and task execution stability of the target robot in real-world scenarios.
[0037] In some embodiments, the optical positioning module includes multiple motion capture cameras for acquiring images containing the target robot. For example, in... Figure 2 The rectangles above the target robot represent motion capture cameras. The terminal is also used to identify the target robot's position through the images to obtain the target robot's actual position data. Specifically, in this embodiment, 14 motion capture cameras are set above the target robot. These are Nokov Mars 1.3H high-speed optical motion capture cameras, with each camera having a resolution of 1280×1024 and a frame rate of 240FPS, enabling sub-millimeter-level tracking within an effective capture space of 6m×5m.
[0038] In other embodiments, the optical positioning module includes an LED position sensing screen. The target robot moves on the surface of the LED position sensing screen, which acquires the actual position data of the target robot. Specifically, the LED position sensing screen measures 6m × 5m and is constructed by splicing together multiple P2.976 LED modules (each module measuring 250mm × 250mm). The 6m × 5m screen has a resolution of 2016 × 1680, with 128 interactive photosensitive points per square meter, and supports secondary development and integration.
[0039] It is understood that in some embodiments, the optical positioning module includes the aforementioned multiple motion capture cameras and LED position sensing screens, and the data collected by both can be used as a reference for data correction.
[0040] Furthermore, the terminal can be a smartphone, tablet, laptop, desktop computer, etc. The target robot can be the PuduBot2 with an omnidirectional chassis as a mobile robot platform, integrating multimodal sensors such as the Kinect V2 image sensor, inertial measurement unit (IMU), and LiDAR. The Kinect V2 is used to simultaneously acquire RGB and depth images, while the IMU and LiDAR assist in environmental perception and pose estimation. The target robot's pose is mainly calculated using the V-SLAM system, and chassis odometry data is fused to improve localization robustness in scenarios with missing textures or dynamic interference.
[0041] Further, please refer to Figure 3 , Figure 3 This is a schematic diagram of an optional interactive interface provided in an embodiment of this application. In this embodiment, the interactive interface includes a first display area and a second display area. The first display area may include a task type sub-area, and the second display area may include a scene preview sub-area. For example, Figure 3 The area outlined in red dashed lines is the first display area, and the area outlined in yellow dashed lines is the second display area. The first display area includes six task type sub-areas, and the second display area includes multiple scene preview sub-areas. It is understood that the position and size of the first and second display areas in the interactive interface, as well as the position, size, and number of the task type sub-areas and scene preview sub-areas, can be adjusted according to actual needs; this embodiment does not impose strict limitations on this.
[0042] In step S101 of some embodiments, the preset task type is a category of robot tasks pre-set within the virtual-real fusion learning platform of the terminal. In this embodiment, the preset task types include: path planning, virtual simulation, task planning, and data acquisition. The preset simulation scene is a three-dimensional virtual environment pre-built within the virtual-real fusion learning platform of the terminal. In this embodiment, the preset simulation scene is a simulation scene built based on Unity3D, including 13 high-fidelity 3D indoor scenes (covering typical environments such as homes, offices, supermarkets, and convenience stores), fully supporting physical rendering, dynamic lighting changes, and interactive object modeling. Relying on Unity's built-in physics engine, the virtual-real fusion learning platform can realistically simulate the dynamic behavior of the target robot in the real world, including key physical characteristics such as collision response, friction, and gravity.
[0043] It should be noted that a preset simulation model of the target robot is pre-constructed before step S101. This preset simulation model is the simulation model of the target robot in the target simulation scene; it is a pre-constructed digital twin. The simulation pose data obtained in real-time in subsequent steps is used to drive the preset simulation model to make synchronous adjustments, which are then displayed to the user through the interactive interface.
[0044] Please see Figure 4 In some embodiments, step S101 may include, but is not limited to, steps S401 to S403: Step S401: In response to the selection operation of the task type sub-region on the interactive interface, select the target type sub-region from the task type sub-region, and take the task type corresponding to the target type sub-region as the target task type.
[0045] Step S402: Switch the second display area according to the target task type to obtain the scene preview sub-area.
[0046] Step S403: In response to the selection operation of the scene preview sub-region on the interactive interface, the target scene sub-region is obtained, and the preset simulation scene corresponding to the target scene sub-region is taken as the target simulation scene.
[0047] In step S401 of some embodiments, the target task type is the task category to be executed selected by the user from the preset task types through the terminal's interactive interface. After the user clicks a specific target type sub-area in the task type sub-area by clicking with the mouse or touching the screen, the terminal detects and recognizes the operation and determines the corresponding target task type. For example, the task type sub-area of the terminal's interactive interface displays task types such as "path planning," "virtual simulation," "task planning," and "data acquisition." When the user clicks the target type sub-area corresponding to "path planning," the terminal determines that the current target task type is "path planning" and changes the color of the corresponding target type sub-area to indicate that the current area has been selected.
[0048] In step S402 of some embodiments, the terminal pre-stores scene preview data corresponding to different task types. After determining the target task type, the terminal automatically loads the scene preview data corresponding to the target task type into the second display area, thereby realizing the interface switching of the second display area. Figure 3 As shown, each scene preview sub-area includes the specific type name of the preset simulation scene and the corresponding preview image.
[0049] In step S403 of some embodiments, the target simulation scene is a specific three-dimensional virtual environment selected by the user from the above-mentioned preset simulation scene through the terminal's interactive interface.
[0050] Steps S401 to S403 as shown in the embodiments of this application reduce the complexity of the simulation task configuration process by setting a task type sub-area and a second display area in the interactive interface, and dynamically switching the second display area according to the target task type to form a scene preview sub-area. This improves the operational efficiency and intuitiveness of the robot simulation training system in the task configuration stage, and enables users to quickly complete the determination process of the target task type and target simulation scene, thereby improving the overall efficiency and operability of the robot simulation training process.
[0051] It should be noted that after step S403 (i.e., determining the target simulation scene) in some embodiments, various data of the current target simulation scene can be collected. Traditional simulation platforms (such as Habitat and AI2-THOR) rely on preset scripts or fixed paths to control pedestrian NPCs, which cannot simulate the randomness and sociality of real user behavior. The embodiments of this application can randomly generate the movement trajectory of pedestrians according to relevant algorithms during the configuration process, and use the random movement trajectory for subsequent robot training.
[0052] Please see Figure 5 After step S402 in some embodiments, the method of this application embodiment may also include, but is not limited to, steps S501 to S503: Step S501: If the target task type is data acquisition, then update the first display area to obtain the configuration parameter type sub-area.
[0053] Step S502: In response to the selection operation of the configuration parameter type sub-region on the interactive interface, and when the parameter type corresponding to the selected configuration parameter type sub-region is pedestrian configuration, obtain the number of pedestrians.
[0054] Step S503: Randomly generate pedestrian movement trajectories based on the number of pedestrians, generate simulation animations based on the pedestrian movement trajectories, and display the simulation animations through the interactive interface.
[0055] In step S501 of some embodiments, please refer to Figure 6 , Figure 6 This is another optional interactive interface diagram provided in this application embodiment. When the target task type is data acquisition, the task type sub-area of the first display area is switched to the configuration parameter type sub-area. For example... Figure 6As shown, the updated first display area includes six sub-areas with different configuration parameter types: view image, pedestrian configuration, sensor data, virtual reality peripheral network, and robot configuration. When the user selects a different parameter type sub-area, the second display area updates the corresponding interactive components. Taking view image as an example, the second display area updates different types of image data from the current robot's perspective, including images with 2D bounding boxes, images with 3D bounding boxes, depth images, etc., for the user to select.
[0056] In step S502 of some embodiments, when the parameter type corresponding to the parameter type sub-area selected by the user is pedestrian configuration, the second display area updates the interactive component accordingly. The form of obtaining the number of pedestrians can be by inputting or selecting a specific value in the input box or drop-down option corresponding to the pedestrian configuration. The terminal recognizes and obtains the value information, thereby determining the number of pedestrians that need to be generated in the current simulation scene.
[0057] In step S503 of some embodiments, several locations in the target simulation scene can be randomly selected as initial and final points, and random motion trajectories can be generated using a random path planning algorithm (such as fast random tree and its extended algorithm). Then, a walking speed is randomly set for each pedestrian, and the corresponding pedestrian walking animation resources are called according to each random motion trajectory and walking speed, ultimately forming a pedestrian simulation animation with random motion effects displayed on the interactive interface.
[0058] In other embodiments, the pedestrian movement trajectory can also be displayed via an LED position sensor screen.
[0059] Steps S501 to S503 as illustrated in this embodiment update the first display area and form a configuration parameter type sub-area when the target task type is data acquisition. This allows the interactive interface to enter the simulation scene parameter configuration state. After the user selects the pedestrian configuration parameter type, pedestrian quantity information is obtained. Then, pedestrian movement trajectories are randomly generated based on the number of pedestrians, and corresponding simulation animations are generated. This makes the pedestrian movement behavior in the simulation scene no longer dependent on a fixed script path, but dynamically changes based on randomly generated movement trajectories. This forms pedestrian behavior data with random characteristics in the virtual-real fusion learning platform, and the simulation animation is displayed through the interactive interface. This allows the virtual-real fusion learning platform to generate pedestrian movement data in the virtual environment that is closer to the distribution of behavior in the real environment. This provides more diverse data samples for the robot training process, improves the richness and authenticity of training data in the virtual-real fusion scene robot training process, enhances the adaptability of the robot control model in complex environments, and thus improves the control accuracy of the robot in real scenes.
[0060] In step S102 of some embodiments, the actual attitude data is the attitude information of the target robot when it moves in a real scene, including the pitch angle, roll angle and yaw angle of the target robot in the real scene. The three-axis rotation attitude data of the robot can be acquired and recorded in real time by an inertial measurement unit (IMU) set on the target robot body, and transmitted to the terminal for recording and processing through the ROS communication interface.
[0061] In step S103 of some embodiments, the actual position data is the real-time spatial position information of the target robot when it moves in a real scene, including the robot's real-time three-dimensional coordinate position. The actual position data needs to be collected in real time.
[0062] When the optical positioning module includes multiple motion capture cameras, the real scene has a first-world coordinate system, which is established by calibration using the extrinsic parameters of the motion capture cameras. In this case, step S103 includes the following steps: 1. Acquiring images containing the target robot using the motion capture cameras. Each motion capture camera can acquire two-dimensional image information containing optical markers or feature points on the target robot within its respective camera's viewpoint. 2. Performing position recognition on the images to obtain the actual position data of the target robot in the first-world coordinate system. Specifically, based on the pre-calibrated camera extrinsic parameters of each motion capture camera, including the position coordinates and orientation information of the motion capture camera itself in the real scene, the terminal calculates the real-time three-dimensional coordinate position of the target robot in the first-world coordinate system of the real scene using triangulation methods, based on the position information of the markers or feature points in the two-dimensional images acquired by each motion capture camera. This yields the actual position data of the target robot in the first-world coordinate system. Relying on the hardware provided in this embodiment, the actual position data in this embodiment can achieve sub-millimeter accuracy.
[0063] In other embodiments, the optical positioning module includes an LED position sensing screen. The target robot moves on the surface of the LED position sensing screen, and the real scene has a second-world coordinate system. The second-world coordinate system typically uses a specific point on the LED position sensing screen as the origin, and coordinate axes are established along the edges of the LED position sensing screen. For example, the LED position sensing screen is rectangular and placed on a horizontal surface. The origin is the upper left corner of the LED position sensing screen, and x-axis and y-axis are established along the two edges connected to the upper left corner. A z-axis is established in a direction perpendicular to the horizontal surface, thus obtaining the second-world coordinate system. In this case, the actual position data of the target robot in the second-world coordinate system is obtained through the LED position sensing screen. Since the LED position sensing screen has an array of position sensing points pre-set on the screen plane, the specific location of the target robot on the screen surface can be detected in real time through these position sensing points.
[0064] In step S104 of some embodiments, the simulation pose data is the three-dimensional spatial position and posture information of the target robot in the target simulation scene, and the target simulation scene has a preset reference coordinate system.
[0065] Please see Figure 7 When the optical positioning module includes multiple motion capture cameras, the simulation pose data acquisition step in step S104 may include, but is not limited to, steps S701 to S702: Step S701: Determine the first transformation mapping relationship between the reference coordinate system and the first world coordinate system.
[0066] Step S702: Perform coordinate mapping on the actual attitude data and actual position data according to the first transformation mapping relationship to obtain the simulated pose data.
[0067] In step S701 of some embodiments, the first transformation mapping relationship is used to realize the transformation relationship of spatial coordinates and attitude between the reference coordinate system and the first world coordinate system, specifically including rotation matrix and translation vector.
[0068] In step S702 of some embodiments, the actual pose data and the actual position data are first used to construct the actual pose matrix. The coordinate mapping process can be referred to the following analytical expression: (1), in, This represents the simulated pose data. This represents the rotation matrix used to transform from the first-world coordinate system to the reference coordinate system. This represents the actual pose matrix of the target robot in the first-world coordinate system. The vector represents the translation from the first-world coordinate system to the reference coordinate system, and the dot represents the multiplication operation.
[0069] Steps S701 to S702 shown in the embodiments of this application can accurately reproduce the actual motion state of the target robot in the real scene in the virtual-real fusion learning platform, thereby improving the consistency between the simulated pose data and the motion data of the target robot in the real scene, solving the problem of spatial position and posture data deviation caused by the difference in coordinate system, improving the training effect of the robot in the virtual-real fusion learning platform and the real reliability of the simulation model, and thus enhancing the control accuracy of the target robot control model in the real scene.
[0070] Please see Figure 8 When the optical positioning module includes an LED position sensing screen, the simulation pose data acquisition step in step S104 may include, but is not limited to, steps S801 to S804: Step S801: Determine the scaling ratio between the LED position sensing screen and the target simulation scene, and perform position calculation on the actual position data according to the scaling ratio to obtain the simulation position data of the target robot in the target simulation scene.
[0071] Step S802: Determine the second transformation mapping relationship between the reference coordinate system and the second world coordinate system.
[0072] Step S803: Perform coordinate mapping on the actual attitude data according to the second transformation mapping relationship to obtain the simulated attitude data.
[0073] Step S804: Determine the simulation pose data based on the simulation attitude data and simulation position data.
[0074] In step S801 of some embodiments, the size scaling ratio refers to the length ratio between the actual physical size of the LED position sensing screen and the virtual space size of the target simulation scene. For example, if the actual length of the LED position sensing screen is 5 meters and the width is 4 meters, and the corresponding virtual area in the target simulation scene is 50 meters long and 40 meters wide, then the terminal calculates a size scaling ratio of 1:10. Then, the actual two-dimensional coordinate data acquired by the target robot on the LED position sensing screen is converted into two-dimensional coordinate data in the target simulation scene according to the aforementioned size scaling ratio, and combined with the height parameter of the target robot, the three-dimensional simulation position data of the target robot in the target simulation scene is obtained. For example, assuming the height of the target robot is 0.5m, when the actual position data collected by the target robot on the LED position sensing screen is (3, 2), the terminal converts the actual position data into two-dimensional simulation position data in the target simulation scene as (30, 20) according to the size scaling ratio of 1:10, and then combines this with the height of the target robot to obtain the three-dimensional simulation position data as (30, 20, 0.5).
[0075] In step S802 of some embodiments, the second transformation mapping relationship is a rotation matrix used to realize the attitude rotation relationship between the reference coordinate system and the second world coordinate system.
[0076] In step S803 of some embodiments, the coordinate mapping process can refer to the following analytical expression: (2), in, This represents the simulated attitude data. This represents the rotation matrix used to transform from the second world coordinate system to the reference coordinate system. This represents the actual attitude data of the target robot in the second world coordinate system. This represents the translation vector from the second world coordinate system to the reference coordinate system.
[0077] In step S804 of some embodiments, the three-dimensional simulation position data and simulation posture data of the target robot in the target simulation scene are combined to form simulation pose data that simultaneously contains position and posture information.
[0078] Steps S801 to S804 as shown in the embodiments of this application can first determine the scaling ratio between the LED position sensing screen and the target simulation scene, and then use the scaling ratio to accurately obtain the simulation position data of the target robot in the target simulation scene. This allows the motion trajectory and posture of the target robot in the virtual simulation environment to accurately correspond to the actual motion state in the real environment, thereby improving the consistency of spatial and motion data between the virtual simulation environment and the real environment.
[0079] In step S105 of some embodiments, the target control model refers to a robot control algorithm model that has been trained and is suitable for the target task type. First, a preset initial control model corresponding to the target task type is invoked, using simulation pose data and sensor data required for the current target task type as input. The parameters of the initial control model are optimized and iteratively updated using machine learning methods such as supervised learning or reinforcement learning until the model can effectively complete the corresponding task in the target simulation scenario. It should be noted that the input data can be manually corrected through an interactive interface. A small amount of real data based on real-world scenarios can also be added during the training process to fine-tune the model parameters, thereby obtaining the decision model corresponding to the target task type, i.e., the target control model.
[0080] In step S106 of some embodiments, the terminal sends motion control commands output by the target control model to the target robot in real time through the ROS communication interface. The target robot adjusts its own motion state in real time based on the received control commands and the target control model to complete the corresponding task operation, thereby realizing real-time closed-loop control of the target robot. For example, in a real urban street scenario, when the target robot detects a pedestrian in front of it in real time during operation, the target robot automatically performs corresponding actions such as deceleration, detour, or stopping based on the pedestrian obstacle avoidance control model deployed in its own controller, so as to achieve precise avoidance of pedestrians.
[0081] Please see Figure 9 After step S403 in some embodiments, the method provided in this application embodiment may also include, but is not limited to, steps S901 to S903: Step S901: Obtain the command operation input to the virtual-real fusion learning platform through the interactive interface, and obtain the real-time task based on the command operation.
[0082] Step S902: Control the target robot to perform a real-time task.
[0083] Step S903: During the execution of the real-time task, the simulation pose data is updated, and a real-time animation of the simulation model is generated based on the updated simulation pose data. The real-time animation is then displayed through the interactive interface.
[0084] In step S901 of some embodiments, the instruction operation is an operation command input by the user to the virtual-real fusion learning platform through the interactive interface to specify the robot's behavior. The real-time task is the task content that the virtual-real fusion learning platform needs to execute immediately by the target robot, which is obtained from the instruction operation parsing. The instruction operation can be implemented by providing a task input area or control button area in the interactive interface. The user submits the task command to the virtual-real fusion learning platform through keyboard input, mouse click, or touch input. After receiving the instruction operation, the virtual-real fusion learning platform parses the instruction operation and determines the corresponding real-time task based on the parsing result. For example, if the user inputs the command "grab an apple" in the task input area of the interactive interface, the terminal receives the instruction operation and obtains the real-time task "control the target robot to grab an apple" through the virtual-real fusion learning platform parsing.
[0085] In step S902 of some embodiments, robot control instructions are generated according to the real-time task and sent to the target robot through the communication interface, thereby driving the target robot to perform the corresponding task action.
[0086] In step S903 of some embodiments, the real-time position and posture information of the target robot is continuously acquired during the execution of the real-time task, and the robot model state in the virtual-real fusion learning platform is updated according to the updated pose information, thereby generating a simulation animation consistent with the motion state of the real robot. Furthermore, all intervention events are fully recorded, including timestamp information, instruction content, and corresponding state data, thus forming a complete task process data record, providing reliable demonstration data for subsequent behavior analysis, strategy reproduction, and reinforcement learning training.
[0087] Steps S901 to S903 as shown in the embodiments of this application enable real-time synchronous dynamic display of the target robot's behavior in the real environment and the simulation modeling behavior in the virtual-real fusion learning platform. This allows users to observe the robot's status in real time and intervene in a timely manner during task execution, effectively improving the real-time performance and accuracy of data interaction during virtual-real fusion scenario training, and enhancing the robot's training effect and the reliability of actual control.
[0088] To comprehensively verify the effectiveness and generalization ability of the virtual-real fusion learning platform proposed in this application for realizing virtual-real fusion learning in complex dynamic tasks, systematic experiments on two types of typical robot tasks were designed and implemented, covering virtual-real intervention grasping and virtual-real assisted learning. Specifically, these included virtual-real intervention grasping tasks and verification of virtual-real fusion algorithm transfer (Real-to-Sim-to-Real).
[0089] Furthermore, the virtual-to-real intervention grasping task aims to verify the platform's ability to support online user intervention and dynamic task switching. In this embodiment, the target robot uses a seven-DOF Kinova Gen3 robotic arm as the execution unit, placing two types of target objects, apples and bananas, on a planar workbench. The initial command is "Pick up the apple" (Prompt A). When the robotic arm is halfway through execution, the current task is manually interrupted through the interactive interface, and a new command "Pick up the banana" (Prompt B) is injected. This intervention signal is synchronized to the target robot in real time via the ROS 2 communication interface, triggering the corresponding target control model to perform strategy replanning and action switching. Experimental results show that regardless of the initial command order (A→B or B→A), the success rate of the second stage task after intervention reaches 80%–100%, while the success rate of the first stage task without intervention is 0% due to forced termination. This result proves that the platform can achieve a millisecond-level human-machine collaborative intervention mechanism, ensuring that changes in commands on the virtual side can be seamlessly mapped to the physical execution end, significantly improving the robot's task adaptability and robustness in open environments.
[0090] Furthermore, the Real-to-Sim-to-Real algorithm transfer task aims to quantify the improvement effect of Real-to-Real fusion on sim-to-real transfer performance, specifically comparing two fine-tuning strategies: (1) fine-tuning using only pure virtual data collected from a virtual platform; and (2) joint fine-tuning combining virtual data and real-world feedback data. When the real-time task is "grabbing a specified fruit and placing it at a target point", pre-trained visual language action models such as OpenVLA-7B and RDT-1B are used for strategy deployment. Experimental results show that in the "grabbing an apple" task, the success rate of pure virtual fine-tuning is 60%, while the success rate of Real-to-Real fusion fine-tuning after introducing real feedback is increased to 90%. In the "grabbing a banana" task, the success rate increased from 40% to 80%. The results clearly demonstrate that by feeding back the execution state of the real robot (such as end-effector pose, contact force, and motion delay) to the simulation environment of the virtual-real fusion learning platform through the optical positioning module and the ROS2 communication link, the physical priors of the virtual model can be effectively calibrated, the gap between the simulation system and the real system can be significantly reduced, and the generalization ability of the strategy in the real world can be greatly improved.
[0091] In summary, through the verification of the above two types of tasks, the embodied intelligent robot control method, terminal and scene simulation system based on the virtual-real fusion learning platform provided in this application embodiment have shown significantly better comprehensive capabilities than existing simulation systems in terms of task intervention flexibility and virtual-real transfer effectiveness.
[0092] This application also provides a terminal, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned embodied intelligent robot control method based on a virtual-real fusion learning platform. This terminal can be any intelligent terminal, including tablet computers, in-vehicle computers, etc.
[0093] Please see Figure 10 , Figure 10 The hardware structure of a terminal according to another embodiment is illustrated. The terminal includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to execute the embodied intelligent robot control method based on a virtual-real fusion learning platform according to the embodiments of this application. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0094] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0095] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0098] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0099] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0100] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0101] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0103] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0104] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A somatic intelligent robot control method based on a virtual-real fusion learning platform, characterized in that, A terminal applied to a scene simulation system, the scene simulation system further including a target robot and an optical positioning module, the terminal being electrically connected to the target robot and the optical positioning module respectively, the terminal being equipped with a virtual-real fusion learning platform, the virtual-real fusion learning platform being provided with an interactive interface, the interactive interface displaying preset task types and preset simulation scenes, the virtual-real fusion learning platform being configured to allow the simulation modeling of the target robot to move within the preset simulation scene, the method including: In response to the selection of the preset task type and preset simulation scenario displayed on the interactive interface, the target task type and target simulation scenario are obtained, and the target simulation scenario is invoked through the virtual-real fusion learning platform, and the preset simulation model of the target robot in the target simulation scenario is loaded. Collect the actual posture data of the target robot in a real scene; The optical positioning module acquires the actual position data of the target robot in the real scene. The simulation modeling is driven by the actual posture data and the actual position data to obtain the simulation pose data of the target robot in the target simulation scene. The target control model is obtained by training a preset initial control model based on the target task type and the simulation pose data. The target robot is controlled based on the target control model.
2. The method of claim 1, wherein, The optical positioning module includes a motion capture camera, and the real-world scene has a first-world coordinate system. The process of acquiring the actual position data of the target robot in the real-world scene through the optical positioning module includes: The motion capture camera acquires images containing the target robot. The image is used to identify the actual position of the target robot in the first world coordinate system. The target simulation scene has a reference coordinate system. The step of establishing a simulation model of the target robot in the target simulation scene based on the actual posture data and the actual position data, to obtain the simulation pose data of the target robot in the target simulation scene, includes: Determine the first transformation mapping relationship between the reference coordinate system and the first world coordinate system; The actual pose data and the actual position data are mapped to coordinates according to the first transformation mapping relationship to obtain the simulated pose data.
3. The method of claim 2, wherein, The optical positioning module includes an LED position sensing screen. The target robot moves on the surface of the LED position sensing screen. The real-world scene has a second-world coordinate system. Obtaining the actual position data of the target robot in the real-world scene through the optical positioning module includes: The actual position data of the target robot in the second world coordinate system is obtained through the LED position sensing screen; The target simulation scene has a reference coordinate system. The step of establishing a simulation model of the target robot in the target simulation scene based on the actual posture data and the actual position data, to obtain the simulation pose data of the target robot in the target simulation scene, includes: Determine the scaling ratio between the LED position sensing screen and the target simulation scene, and perform position calculation on the actual position data according to the scaling ratio to obtain the simulation position data of the target robot in the target simulation scene; Determine the second transformation mapping relationship between the reference coordinate system and the second world coordinate system; The actual attitude data is mapped to coordinates according to the second transformation mapping relationship to obtain the simulated attitude data; The simulated pose data is determined based on the simulated posture data and the simulated position data.
4. The method according to any one of claims 1 to 3, characterized in that, The interactive interface includes a first display area and a second display area. The first display area includes a task type sub-area. The step of obtaining the target task type and target simulation scenario in response to a selection operation of a preset task type and a preset simulation scenario displayed on the interactive interface includes: In response to the selection operation of the task type sub-region on the interactive interface, a target type sub-region is selected from the task type sub-region, and the task type corresponding to the target type sub-region is taken as the target task type; The second display area is switched according to the target task type to obtain a scene preview sub-area; In response to the selection operation of the scene preview sub-region on the interactive interface, a target scene sub-region is obtained, and the preset simulation scene corresponding to the target scene sub-region is taken as the target simulation scene.
5. The method according to claim 4, characterized in that, After responding to the selection operation of the scene preview sub-region on the interactive interface, obtaining the target scene sub-region, and taking the preset simulation scene corresponding to the target scene sub-region as the target simulation scene, the method further includes: The user obtains instructions and operations input to the virtual-real fusion learning platform through the interactive interface, and obtains real-time tasks based on the instructions and operations. Control the target robot to execute the real-time task; During the execution of the real-time task, the simulation pose data is updated, and a real-time animation of the simulation model is generated based on the updated simulation pose data. The real-time animation is then displayed through the interactive interface.
6. The method according to claim 4, characterized in that, After responding to the selection operation of the scene preview sub-region on the interactive interface, obtaining the target scene sub-region, and taking the preset simulation scene corresponding to the target scene sub-region as the target simulation scene, the method further includes: If the target task type is data acquisition, then the first display area is updated to obtain a configuration parameter type sub-area; In response to the selection operation of the configuration parameter type sub-region on the interactive interface, and when the parameter type corresponding to the selected configuration parameter type sub-region is pedestrian configuration, the number of pedestrians is obtained; The pedestrian movement trajectory is randomly generated based on the number of pedestrians, and a simulation animation is generated based on the pedestrian movement trajectory. The simulation animation is then displayed through the interactive interface.
7. A terminal, characterized in that, The terminal includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 6.
8. A scene simulation system, characterized in that, The scene simulation system includes the terminal, optical positioning module, and target robot as described in claim 7, wherein the terminal is equipped with a virtual-real fusion learning platform, the virtual-real fusion learning platform is provided with an interactive interface, and the terminal communicates with the target robot and the optical positioning module respectively through the ROS communication interface.
9. The scene simulation system according to claim 8, characterized in that, The optical positioning module includes a motion capture camera, which is used to acquire images containing the target robot; the terminal is also used to identify the target robot's location through the images to obtain the target robot's actual location data.
10. The scene simulation system according to claim 9, characterized in that, The optical positioning module includes an LED position sensing screen, the target robot is used to move on the surface of the LED position sensing screen, and the LED position sensing screen is used to acquire the actual position data of the target robot.