Robot grabbing learning closed-loop optimization method and system based on visual language feedback
By introducing reinforcement learning into a simulation environment to generate diverse grasping actions and using a visual language model for realistic feedback optimization, the problem of single grasping actions and lack of feedback in existing technologies is solved, and adaptive optimization and performance improvement of robot grasping tasks are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-12
AI Technical Summary
Existing methods lack diversity in grasping actions and a feedback-driven closed-loop optimization mechanism when generating robot grasping demonstration data, resulting in insufficient generalization ability and robustness in real-world environments.
By introducing an active exploration reinforcement learning strategy into a simulation environment to generate diverse grasping actions, and using a visual language model to evaluate and provide feedback on real grasping tasks, a closed-loop optimization process is constructed to achieve dynamic interaction between data generation and strategy training.
It significantly improves the generalization ability and adaptability of the crawling strategy in real-world environments, achieving adaptive continuous learning and performance improvement, and is able to adapt to dynamic and unstructured environments.
Smart Images

Figure CN122008191A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotic arm control technology, and in particular to a robot grasping learning closed-loop optimization method and system based on visual language feedback. Background Technology
[0002] In the field of robotics, acquiring high-quality demonstration data is crucial for improving the performance of imitation learning or reinforcement learning algorithms. Traditional methods rely on manual teaching or repeated trials with real robots, which are costly and struggle to cover complex and diverse task scenarios. Therefore, generating demonstration data using simulation environments has become a mainstream research direction. Through simulation, labeled motion trajectories can be generated on a large scale and at low cost.
[0003] To bridge the visual gap between simulation and real-world environments (the Simulation to Real, Sim2Real problem), researchers have introduced high-fidelity techniques based on neural rendering, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). These techniques can reconstruct highly realistic 3D scenes with consistent lighting from multi-view images and render lifelike image sequences, thereby generating training data with a visual distribution closer to the real world.
[0004] However, existing data generation methods based on high-fidelity rendering have significant drawbacks: 1. Lack of diversity in grasping actions: Existing methods primarily focus on visual data augmentation, such as applying random perturbations to the object's pose, lighting, background, or camera viewpoint, but neglect the structural diversity of the grasping actions themselves. In the generated demonstration data, the relative grasping posture between the hand and the target object is usually fixed or has limited variation. This results in a limited grasping strategy space, leading to insufficient generalization ability and robustness when faced with minor changes in the object's pose and shape or different operational constraints in the real environment.
[0005] 2. Lack of a feedback-driven closed-loop optimization mechanism: Existing data generation processes are mostly unidirectional open-loop structures of "demonstration collection → high-fidelity rendering → strategy training". The execution effect of the strategy after deployment in the real environment (such as failure cases, success patterns, and human preferences) cannot be effectively fed back to the data generation stage. Therefore, the system cannot supplement weak links with data or optimize the distribution of the capture strategy based on feedback from the real world, making it difficult to achieve continuous adaptive learning and performance improvement. Summary of the Invention
[0006] To address the aforementioned issues, this invention proposes a closed-loop optimization method and system for robot grasping learning based on visual language feedback. This method can automatically generate diverse grasping demonstrations and perform closed-loop optimization based on real deployment feedback.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a robot grasping learning closed-loop optimization method based on visual language feedback, comprising the following steps: Perform high-precision 3D reconstruction of the target operation scene and objects; In a simulation environment, candidate grasping trajectories containing diverse grasping actions are generated through active exploration; Based on the 3D Gaussian splashing technology, the candidate grabbing trajectory executed in the simulation environment is rendered with high fidelity to generate corresponding visual demonstration data; The grasping strategy was trained using visual demonstration data; Deploy the trained grasping strategy in a real environment to perform grasping tasks; The visual language model is used to analyze the execution results of real grasping tasks, generate evaluation feedback signals, optimize the active exploration process in the simulation environment based on the evaluation feedback signals, and generate new visual demonstration data based on the optimized results for the next round of strategy training, forming a closed-loop optimization process.
[0008] As an alternative implementation, candidate grasping trajectories containing diverse grasping actions are generated through active exploration. Specifically, a reinforcement learning strategy is run in a simulation environment. The action space of the reinforcement learning strategy is designed to cover multiple poses of the end effector in order to explore and generate structurally different feasible grasping methods.
[0009] As an alternative implementation method, the visual language model is used to analyze the execution results of the real crawling task and generate evaluation feedback signals. Specifically, the visual observations during the real crawling process are input into the visual language model, and combined with preset text prompts, the score or reward signal output by the visual language model that represents the crawling quality is obtained.
[0010] As an alternative implementation, the active exploration process in the simulation environment is optimized based on the evaluation feedback signal. Specifically, the evaluation feedback signal is used as part of the reinforcement learning reward function to update the reinforcement learning strategy that generates candidate grasping trajectories, so that it tends to generate grasping actions that are more consistent with the evaluation feedback signal.
[0011] As an alternative implementation method, the closed-loop optimization process undergoes multiple iterations, and the visual demonstration data generated in each iteration is accumulated to build a demonstration database that is continuously expanded and optimized in both the grasping action space and the visual observation space.
[0012] As an alternative implementation, the score or reward signal is combined with the success rate of grasping, object stability, user preferences, or task-specific constraints to comprehensively evaluate the grasping quality.
[0013] Secondly, the present invention provides a robot grasping learning closed-loop optimization system based on visual language feedback, comprising: The 3D reconstruction module is configured to perform high-precision 3D reconstruction of the target operation scene and objects; The simulation exploration module is configured to generate candidate grasping trajectories containing diverse grasping actions through active exploration in a simulation environment. The high-fidelity rendering module is configured to: perform high-fidelity rendering of the scene of candidate grasping trajectories executed in the simulation environment based on three-dimensional Gaussian splashing technology, and generate corresponding visual demonstration data; The strategy training module is configured to train the grasping strategy using visual demonstration data; The real deployment module is configured to deploy the trained crawling strategy in a real environment to perform crawling tasks. The closed-loop control module is configured to: analyze the execution results of the real grasping task using a visual language model, generate evaluation feedback signals, optimize the active exploration process in the simulation environment based on the evaluation feedback signals, and generate new visual demonstration data based on the optimized results for the next round of strategy training, thus forming a closed-loop optimization process.
[0014] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0015] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0016] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a closed-loop optimization method for robot grasping learning based on visual language feedback. By introducing an active exploration reinforcement learning strategy into a simulation environment, it fundamentally solves the problem of the limited grasping methods in traditional demonstration data. This method overcomes the bottleneck of generating diverse grasping actions, enabling the systematic generation of various structurally feasible grasping actions. This significantly expands the coverage of strategy training data, thereby improving the generalization ability and adaptability of the final grasping strategy in real-world complex environments.
[0018] This invention proposes a closed-loop optimization method for robot grasping learning based on visual language feedback. It innovatively introduces a visual language model as a real-world "quality evaluator," transforming difficult-to-quantify grasping performance (such as stability and conformity to human intuition) into usable feedback signals. These signals are then injected back into the simulation data generation process, making data generation no longer static and blind, but dynamically and purposefully optimized based on real-world performance. This achieves a virtuous cycle of data generation and strategy performance improvement, and realizes closed-loop optimization based on real-world feedback.
[0019] This invention proposes a closed-loop optimization method for robot grasping learning based on visual language feedback. It constructs an adaptive continuous learning system framework that organically integrates candidate grasping generation, high-fidelity rendering, policy training, and realistic feedback into a closed-loop operating system. Through multiple iterations, this system continuously utilizes realistic feedback to correct the direction of simulation data generation, enabling the policy and generated data to evolve together. Ultimately, this achieves long-term, stable improvement in robot grasping performance, providing a solid technical foundation for reliable robot operation in dynamic, unstructured environments.
[0020] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0022] Figure 1 This is a diagram illustrating the overall architecture of the robot grasping learning closed-loop optimization method based on visual language feedback according to the present invention. Detailed Implementation
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed description is exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. Furthermore, it should be understood that the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but includes other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0027] Example 1 like Figure 1 As shown, this embodiment provides a closed-loop optimization method for robot grasping learning based on visual language feedback, including the following steps: Perform high-precision 3D reconstruction of the target operation scene and objects; In a simulation environment, candidate grasping trajectories containing diverse grasping actions are generated through active exploration; Based on the 3D Gaussian splashing technology, the candidate grabbing trajectory executed in the simulation environment is rendered with high fidelity to generate corresponding visual demonstration data; The grasping strategy was trained using visual demonstration data; Deploy the trained grasping strategy in a real environment to perform grasping tasks; The visual language model is used to analyze the execution results of real grasping tasks, generate evaluation feedback signals, optimize the active exploration process in the simulation environment based on the evaluation feedback signals, and generate new visual demonstration data based on the optimized results for the next round of strategy training, forming a closed-loop optimization process.
[0028] The specific solution of the present invention is as follows: This invention proposes a closed-loop optimization method for robot grasping learning based on visual language feedback, addressing core issues in existing methods such as limited grasping actions, lack of feedback in data generation, and insufficient adaptability. The overall framework consists of "scene and object reconstruction → candidate grasping generation → high-fidelity rendering → policy training → real-world deployment feedback → candidate grasping update," forming a continuously evolving closed-loop grasping learning system. First, addressing the issues of limited grasping actions and restricted policy space in traditional data generation, this invention introduces a candidate grasping generation module. This module, based on reinforcement learning, enables the robot to actively explore multiple feasible grasping methods in a simulation environment, significantly improving the coverage of grasping actions in demonstration data and making subsequent rendered demonstrations more diverse and exhibiting realistic human-like characteristics. Second, to address the shortcomings of existing data generation methods that cannot utilize real-world deployment performance and lack feedback, this invention designs a VLM evaluation feedback module. This module, based on a visual language model (VLM), evaluates the results after the robot actually performs a grasping action, then converts the evaluation results into preference or reward signals to update the reinforcement learning policy, ensuring that the policy continuously aligns with the needs of the real-world scenario during iteration. Finally, this invention constructs a closed-loop operating system that organically couples the aforementioned modules with the high-fidelity 3DGS rendering process, forming a dynamic loop between data generation, policy training, and real-world feedback. In this closed-loop structure, the policy not only learns from the demonstration data, but its resulting grasping actions also serve as upstream inputs to construct new demonstration data; the rendered high-fidelity data further enhances training stability; and real-world feedback, in turn, adjusts the policy distribution, enabling the system to continuously optimize through multiple iterations. Ultimately, this achieves enhanced grasping action diversity, real-world feedback-driven policy alignment, and a long-term adaptive closed-loop grasping learning process.
[0029] 3D Gaussian Splatting (3DGS): Uses multi-view images as input to construct a high-precision 3D representation of a scene. Each Gaussian point contains the following components: a 3D center point. Three-dimensional scale ,color Opacity and rotation quaternions These parameters are collectively referred to as , of which The parameters of a Gaussian are expressed as follows: During rendering, these Gaussian points are projected onto a two-dimensional image plane through a differentiable rasterization process. For the same pixel, multiple Gaussian points are superimposed sequentially according to their opacity and color in depth order, thereby achieving a blending rendering effect related to transparency. Through the above process, a high-quality rendered image containing rich geometric details and lighting effects can be generated, achieving a realistic reconstruction of the real scene.
[0030] In 3DGS, segmented objects in a scene can undergo rigid body transformations, such as translation and rotation, while maintaining high-quality rendering. An object is composed of several three-dimensional Gaussian points, and its overall motion can be achieved through a homogeneous transformation matrix. This is achieved by using a rotation matrix. Translation vector Together they constitute. For a location with a mean value... Covariance Matrix The new position of a 3D Gaussian point after rigid body transformation. With the new covariance matrix It can be obtained through the following formula: (1); (2); Through the above transformation, even if the object changes its pose in space, the Gaussian point representation can still achieve continuous, stable and accurate rendering effects, thereby ensuring the realism and consistency of dynamic scenes.
[0031] Candidate grab trajectory generation: In the field of robot learning and simulation data generation, existing high-fidelity rendering methods mainly focus on the diversity of visual aspects such as object pose, lighting, appearance, and camera viewpoint. However, in actual grasping tasks, the hand-object relative pose often remains fixed in demonstration data, resulting in highly concentrated data in the policy space. This single policy distribution limits the generalization ability of the trained grasping policy in real-world environments, especially when faced with slight changes in object shape, placement angle, and grasping constraints, making it difficult for the robot to stably complete the grasping operation.
[0032] To address this issue, this invention proposes introducing a candidate grasping generation module into a simulation environment, which actively explores various feasible grasping methods through reinforcement learning algorithms. Assume the robot's state space is... The action space is The reinforcement learning strategy can be represented as: (3); Where θ is the policy parameter. For a given state The strategy outputs the end effector action. Perform this action in the simulation environment and receive environmental rewards. Then, the policy parameters are updated using a reinforcement learning algorithm.
[0033] This module maps the grabbing actions obtained from policy execution after training to high-fidelity demonstration data: (4); in Indicates the state The following is a high-fidelity scene image obtained through 3D Gaussian splash rendering. In this way, the module can construct a policy dataset containing multimodal grasping actions. : (5); in This indicates the number of strategy attempts or the number of candidate grasps. The rich grasping samples cover the diversity of hand-object relative postures, achieving a structural expansion of the grasping strategy space. The trained strategy not only performs well in the simulation environment but also exhibits stronger robustness in real-world deployments, adapting to minor changes in object placement, shape, and environmental conditions, significantly reducing the grasping failure rate.
[0034] This module enables proactive exploration of grasping actions and the construction of diverse demonstration data, providing a solid foundation for subsequent closed-loop feedback optimization.
[0035] Visual language model evaluation feedback: After generating diverse grasping candidates in a simulation environment and rendering them as high-fidelity demonstration data, the strategy is trained based on the obtained data, and then the grasping task is executed in a real environment. However, due to the physical differences between the simulation and the real environment, the slight positional shifts of objects, and complex operational constraints, the strategy trained solely on the simulation demonstration data may not fully meet the requirements of the actual task. To solve the above problems, this invention proposes a VLM evaluation feedback module, which introduces a Vision-Language Model (VLM) to evaluate the grasping execution results and transforms the evaluation results into preference signals or rewards that can be used for reinforcement learning, thereby achieving closed-loop optimization of strategy performance.
[0036] The input to the Visual Language Model (VLM) is a key contact frame image at the time of grasping, which must include a semantic segmentation mask of the target object (such as a mug) and a clear task description text (such as "grab the mug to drink water").
[0037] The prompts passed to VLM are specifically designed to guide it through a structured reasoning process: First, they output the pixel coordinates of the current grab point directly on the provided mask image; second, they assign a quantitative score of 1-10 to the appropriateness of the grab point, based on the given task description. VLM is required to output in strict JSON format, containing two fields: current_grasp_coordinates and appropriateness_score.
[0038] During the deployment phase, the system collects VLM evaluation results from multiple crawling attempts and selects the one with the highest reasonableness score. For the pixel coordinates corresponding to this optimal crawl, the system initiates a coordinate transformation module based on 3D model registration to convert them into a 3D position in the coordinate system of the object being manipulated.
[0039] The core steps of this transformation process are as follows: First, the system utilizes a pre-acquired 3D model of the target object (obtained during the first step of scene and object reconstruction) to determine the object's 3D position and pose (i.e., 6D pose) relative to the camera in the contact frame using a model-based pose estimation algorithm. Then, using the 2D pixel coordinates provided by the VLM, the system directly queries the corresponding depth value from the depth image and, combined with camera intrinsic parameters, calculates the precise 3D coordinates of the point in the camera coordinate system through inverse perspective transformation. Finally, using the obtained object pose (rotation matrix R and translation vector t), the 3D coordinates are transformed from the camera coordinate system to the object model's own local coordinate system, thus obtaining the optimal grab point's 3D coordinates in the object model's local coordinate system.
[0040] Ultimately, the system defines this coordinate as the "ideal grasping point" that conforms to the task semantics and introduces it as a key parameter into the reward function of reinforcement learning. This guides the policy in future training not only to learn how to successfully grasp, but also to learn the optimal position on the object that conforms to the task function for grasping. This design achieves a precise and automated closed loop from high-level task instructions to low-level action space optimization.
[0041] Suppose the execution trajectory of the strategy in a real environment is as follows: (6); in, It represents the state in a real-world environment. The action that the strategy performs in this state. This marks the end of the task. The real deployment feedback module evaluates the crawling and execution process using VLM, generating preference scores or reward signals.
[0042] This scoring system comprehensively evaluates grasping quality by combining grasping success rate, object stability, user preferences, or task-specific constraints, thereby achieving policy calibration that better meets real-world needs. Subsequently, the preference signal output by the VLM is used for reinforcement learning updates. Through this feedback, the policy not only learns grasping actions from simulation-generated data but also dynamically optimizes based on the execution results of real-world deployments, thus gradually adapting to operational constraints and user preferences in the real environment.
[0043] The real deployment feedback update module of this invention achieves automated quality evaluation of grasping actions through VLM, directly feeding back the execution results in the real environment to the reinforcement learning training stage. This ensures that the strategy can continuously adapt to the needs of real tasks, achieving closed-loop optimization between simulation data generation and real deployment. Through continuous iteration, the robustness and generalization ability of the strategy in the real environment are significantly improved, while the quality and diversity of the generated demonstration data are also enhanced, providing a reliable guarantee for long-term adaptive learning.
[0044] Closed-loop operating system: This invention further proposes a closed-loop operating system module, which organically combines candidate capture generation, high-fidelity rendering, policy training, and real deployment feedback to construct a sustainable and adaptive closed-loop process for demonstration data generation and policy optimization. This closed-loop operating system aims to overcome the limitations of unidirectional static demonstration data in traditional simulation data generation methods, enabling real-time interaction and iterative optimization between demonstration data and policy training, thereby significantly improving the robustness and generalization ability of the policy in real-world environments. The core process of the closed-loop operating system includes five key steps.
[0045] Candidate Grasp Generation: In the simulation environment, the system first utilizes reinforcement learning to generate diverse grasping actions. This module not only considers the relative pose of the end effector and the object but also systematically covers different grasping patterns within the object's graspable region, forming a rich set of candidate grasping actions. Through this proactive exploration approach, the system can generate a large number of structurally diverse grasping actions, addressing the problems of traditional demonstration data being limited in policy space and coverage, thus providing ample action samples for subsequent data rendering and policy training.
[0046] High-fidelity rendering: Candidate grasping actions are generated into demonstration data using high-fidelity 3DGS rendering. The entire grasping action trajectory is rendered continuously, from the initial moment to the complete execution of the trajectory. The rendering process not only preserves the geometric, lighting, and texture details of the object but also visualizes the grasping action in visual space, ensuring high quality and diversity of training data at both the visual and grasping action levels. In this way, policy training is no longer limited to a single visual scene or fixed grasping posture but can learn grasping policies covering a broader state space.
[0047] Strategy Training: The grasping strategy is trained using high-fidelity demonstration data generated by rendering. During training, the strategy learns the selection and execution of grasping actions by imitating diverse demonstration data, thereby adapting to various object types, placement postures, and environmental conditions.
[0048] Real-world deployment feedback: The trained strategy is deployed in a real-world environment to perform grasping tasks. Due to physical differences, operational constraints, and minor object deviations between simulation and real-world environments, the strategy's performance in the real-world environment may be insufficient. Therefore, this invention introduces a Visual Language Model (VLM) to evaluate the grasping execution results, generating preference scores or reward signals for reinforcement learning updates. This enables dynamic optimization of the strategy parameters, allowing the strategy to gradually adapt to the operational constraints and task requirements of the real-world environment.
[0049] Candidate crawling update: Based on feedback from real deployments, the system adjusts the candidate crawling generation module, adding new crawling actions that meet real operational needs and user preferences, and reintegrates them into the simulation rendering process, forming a closed-loop iterative system of strategy-data-feedback. Through this mechanism, demonstration data and strategies can mutually reinforce each other, continuously optimizing strategy training results, while the diversity and quality of demonstration data are continuously improved, providing a reliable foundation for long-term adaptive learning.
[0050] The system adjusts the candidate crawling generation module as follows: the VLM feedback signal is used as the reward value to update the original reinforcement learning (RL) policy (π_θ). The reinforcement learning policy continues to evolve in the closed loop, thereby generating candidate crawls that are more in line with preferences. Further rendering is performed to generate demonstration data for training the policy.
[0051] The closed-loop operating system of this invention tightly integrates candidate grasping generation, high-fidelity rendering, policy training, and real-world deployment feedback, constructing an adaptive optimization closed loop from simulation to reality, from generation to feedback, and from data to policy. This closed loop not only ensures the robustness and generalization ability of the policy in complex real-world environments but also improves the coverage and diversity of demonstration data. It provides a solid technical foundation for robots to stably perform grasping tasks in multi-object, multi-pose, and dynamically changing environments, and lays theoretical and practical support for long-term continuous optimization and adaptive learning.
[0052] This invention addresses the shortcomings of existing 3DGS-based demonstration data generation methods, such as insufficient strategy diversity and lack of realistic feedback loops. It proposes a closed-loop data generation and strategy optimization method for robot grasping tasks, offering significant technical advantages over existing technologies. Existing methods primarily rely on fixed demonstration grasping postures, using visual perturbations (such as object pose, appearance, lighting, and camera viewpoint) to enhance the diversity of synthetic data. However, they neglect the structural diversity of the "grasping action itself," resulting in a limited strategy space and insufficient generalization ability. Furthermore, their data generation process is a unidirectional chain, with demonstration, rendering, and training independent of each other, making dynamic optimization based on real-world deployment results and targeted correction of poor grasping methods impossible.
[0053] To overcome the aforementioned shortcomings, this invention introduces a reinforcement learning-driven candidate grasping generation module, enabling the robot to actively explore feasible grasping methods in various poses during simulation, thereby enhancing the structural diversity of demonstration data from a policy perspective. Furthermore, it utilizes a visual language model (VLM) to evaluate the grasping execution results in real-world deployments, providing real-world feedback signals for updating the grasping generation strategy, thus achieving continuous alignment between the strategy and real-world requirements. In addition, this invention constructs a closed-loop framework of "candidate grasping generation—high-fidelity rendering—policy training—real-world deployment feedback—candidate grasping update," enabling policy learning to no longer rely on static demonstrations but instead continuously drive the adaptive evolution of demonstration generation and policy distribution through real-world feedback, significantly improving the robot's real-world deployment performance and long-term robustness in grasping tasks.
[0054] Example 2 This embodiment provides a robot grasping learning closed-loop optimization system based on visual language feedback, including: The 3D reconstruction module is configured to perform high-precision 3D reconstruction of the target operation scene and objects; The simulation exploration module is configured to generate candidate grasping trajectories containing diverse grasping actions through active exploration in a simulation environment. The high-fidelity rendering module is configured to: perform high-fidelity rendering of the scene of candidate grasping trajectories executed in the simulation environment based on three-dimensional Gaussian splashing technology, and generate corresponding visual demonstration data; The strategy training module is configured to train the grasping strategy using visual demonstration data; The real deployment module is configured to deploy the trained crawling strategy in a real environment to perform crawling tasks. The closed-loop control module is configured to: analyze the execution results of the real grasping task using a visual language model, generate evaluation feedback signals, optimize the active exploration process in the simulation environment based on the evaluation feedback signals, and generate new visual demonstration data based on the optimized results for the next round of strategy training, thus forming a closed-loop optimization process.
[0055] It should be noted that the above modules correspond to the steps in Embodiment 1, and the examples and application scenarios implemented by the above modules and their corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules can be executed in a computer system as part of the system.
[0056] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0057] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0058] A computer-readable storage medium for storing computer instructions that, when executed by a processor, perform the method of Embodiment 1.
[0059] The method in Example 1 can be directly executed by a hardware processor, or it can be executed by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0060] A computer program product includes a computer program that, when executed by a processor, implements the method in Embodiment 1.
[0061] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0062] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0063] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0064] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0065] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A closed-loop optimization method for robot grasping learning based on visual language feedback, characterized in that, Includes the following steps: Perform high-precision 3D reconstruction of the target operation scene and objects; In a simulation environment, candidate grasping trajectories containing diverse grasping actions are generated through active exploration; Based on the 3D Gaussian splashing technology, the candidate grabbing trajectory executed in the simulation environment is rendered with high fidelity to generate corresponding visual demonstration data; The grasping strategy was trained using visual demonstration data; Deploy the trained grasping strategy in a real environment to perform grasping tasks; The visual language model is used to analyze the execution results of real grasping tasks, generate evaluation feedback signals, optimize the active exploration process in the simulation environment based on the evaluation feedback signals, and generate new visual demonstration data based on the optimized results for the next round of strategy training, forming a closed-loop optimization process.
2. The robot grasping learning closed-loop optimization method based on visual language feedback as described in claim 1, characterized in that, By actively exploring and generating candidate grasping trajectories containing diverse grasping actions, specifically, a reinforcement learning strategy is run in a simulation environment. The action space of the reinforcement learning strategy is designed to cover multiple poses of the end effector in order to explore and generate structurally different feasible grasping methods.
3. The robot grasping learning closed-loop optimization method based on visual language feedback as described in claim 1, characterized in that, The visual language model is used to analyze the execution results of real crawling tasks and generate evaluation feedback signals. Specifically, the visual observations during the real crawling process are input into the visual language model and combined with preset text prompts to obtain the score or reward signal output by the visual language model that represents the crawling quality.
4. The robot grasping learning closed-loop optimization method based on visual language feedback as described in claim 3, characterized in that, Based on the evaluation feedback signal, the active exploration process in the simulation environment is optimized. Specifically, the evaluation feedback signal is used as part of the reinforcement learning reward function to update the reinforcement learning policy that generates candidate grasping trajectories, so that it tends to generate grasping actions that are more consistent with the evaluation feedback signal.
5. The robot grasping learning closed-loop optimization method based on visual language feedback as described in claim 1, characterized in that, The closed-loop optimization process undergoes multiple iterations, and the visual demonstration data generated in each iteration is accumulated to build a demonstration database that is continuously expanded and optimized in both the grasping action space and the visual observation space.
6. The robot grasping learning closed-loop optimization method based on visual language feedback as described in claim 3, characterized in that, The score or reward signal is combined with the success rate of crawling, object stability, user preferences or task-specific constraints to comprehensively evaluate the crawling quality.
7. A robot grasping learning closed-loop optimization system based on visual language feedback, characterized in that, include: The 3D reconstruction module is configured to perform high-precision 3D reconstruction of the target operation scene and objects; The simulation exploration module is configured to generate candidate grasping trajectories containing diverse grasping actions through active exploration in a simulation environment. The high-fidelity rendering module is configured to: perform high-fidelity rendering of the scene of candidate grasping trajectories executed in the simulation environment based on three-dimensional Gaussian splashing technology, and generate corresponding visual demonstration data; The strategy training module is configured to train the grasping strategy using visual demonstration data; The real deployment module is configured to deploy the trained crawling strategy in a real environment to perform crawling tasks. The closed-loop control module is configured to: analyze the execution results of the real grasping task using a visual language model, generate evaluation feedback signals, optimize the active exploration process in the simulation environment based on the evaluation feedback signals, and generate new visual demonstration data based on the optimized results for the next round of strategy training, thus forming a closed-loop optimization process.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-6.