Data-driven robot closed-loop joint optimization method and system

By adopting a closed-loop joint optimization method based on data-driven in robot simulation training, analyzing the task description and generating three-dimensional scene information, disassembling the task into a subtask and optimizing the execution strategy, the problem of independent processing of perception and manipulation in the existing technology is solved, and more efficient and robust robot skill training is achieved.

CN119328777BActive Publication Date: 2025-05-13ZHEJIANG YOULU ROBOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411896331.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-13
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The existing robot simulation training methods lack the robot's active perception and real-time feedback capabilities, resulting in reduced cumulative error and robustness, and perception and manipulation are independently processed and cannot be effectively jointly optimized.

Method used

The closed-loop joint optimization method of robots based on data-driven is adopted to generate three-dimensional scene information by analyzing the task description, disassembling the task into a subtask and generating execution strategies using a large language model, combining visual algorithms to identify target positions and poses, and optimizing execution strategies through reinforcement learning.

Benefits of technology

It significantly improves the robot's adaptability and task success rate in a noisy environment, enhances the joint optimization capability of perception and manipulation, reduces cumulative errors, and improves the robustness and control accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119328777B_ABST
    Figure CN119328777B_ABST
Patent Text Reader

Abstract

The present invention discloses a data-driven robot closed-loop joint optimization method and system, including: parsing the acquired specific task description to generate three-dimensional scene information; decomposing the specific task description into multiple subtasks and executing the subtasks after generating the subtask execution strategy through a large language model based on the type of the subtask; obtaining the type of the subtask and the three-dimensional scene information, identifying the position and posture of the target object through a visual algorithm; fusing the position and posture of the target object with the output information obtained by executing the subtask, and continuously optimizing and adjusting the subtask execution strategy using a reinforcement learning algorithm. The beneficial effects of the present invention are: greatly improving the training efficiency and strategy generalization ability, significantly improving the task completion rate, and reducing the dependence on the optimization of a single module while improving the accuracy of the strategy. The closed-loop system significantly reduces the cumulative error, allowing the robot to complete complex tasks more stably and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to a data-driven robot closed-loop joint optimization method and system. Background Art

[0002] With the continuous development of robotics and agent skill training, training in simulation environments provides a safe and efficient alternative for optimizing robot skills. Compared with the high-cost and low-security data collection methods in the real world, large-scale parallel computing in simulation environments accelerates data collection, which not only improves the efficiency of robot autonomous skill optimization, but also reduces the dependence on robot hardware. However, most existing robot skill training relies on manually designed tasks and scenarios, or creates simulation scenarios through generative models to train the robot's physical interaction capabilities. In addition, most robot simulations use an open-loop framework, which lacks the robot's active perception and real-time feedback capabilities, resulting in cumulative errors and reduced system robustness. They also ignore the correlation between robot perception and manipulation, treating the two as independent functional modules, without jointly optimizing the robot's perception and manipulation.

[0003] Therefore, the existing robot simulation training methods result in insufficient adaptability of the robot in noisy environments, affecting the success rate of actual task execution. Summary of the invention

[0004] The purpose of the present invention is to provide a data-driven robot closed-loop joint optimization method to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] In a first aspect, an embodiment of the present invention provides a data-driven robot closed-loop joint optimization method, comprising:

[0007] Parse the acquired specific task description to generate three-dimensional scene information;

[0008] Decompose a specific task description into multiple subtasks and generate a subtask execution strategy based on the subtask type through a large language model to execute the subtask;

[0009] Obtain the type of subtask and 3D scene information, and identify the position and posture of the target object through visual algorithms;

[0010] The position and posture of the target object are integrated with the output information obtained from executing the subtask, and the reinforcement learning algorithm is used to continuously optimize and adjust the subtask execution strategy.

[0011] Preferably, the acquired specific task description is parsed to generate three-dimensional scene information and the specific task is decomposed into multiple subtasks, including:

[0012] Understand the specific tasks;

[0013] Acquire 3D materials for task understanding based on large-scale datasets;

[0014] Use feature integration mechanism to obtain text and image embedding;

[0015] A scene domain generalization method based on reinforcement learning adjusts object layout by learning environmental characteristics and task requirements.

[0016] Preferably, the scene domain generalization method based on reinforcement learning adjusts the object layout by learning the environment characteristics and task requirements, including:

[0017] Based on a custom material library, an adaptive selection mechanism is used to randomize the textures of objects in the scene.

[0018] Prioritize compatible materials based on the spatial distribution of objects and contextual information;

[0019] Combine reinforcement learning to learn domain features and enhance scene generalization capabilities;

[0020] Dynamically adjust the complexity of the scene.

[0021] Preferably, the complexity of the scene is defined as follows: in, is the space occupancy rate, which represents the ratio of the object volume to the total volume of the scene, It is the layout orderliness, which is used to measure the uniformity of object distribution. and is the weight, is the total volume of objects in the scene, is the total volume of the scene, is the average distance, is the scene diagonal distance.

[0022] Preferably, the acquiring of the subtask type and three-dimensional scene information, and identifying the position and posture of the target object by a visual algorithm, includes:

[0023] When the robot fails to directly locate the target object in the initial stage, the object search control mode is triggered;

[0024] The object search control mode is used to control the robot to search for the task target in the scene;

[0025] After the target object is identified, path planning is initiated and the robot navigates to the target object location.

[0026] Preferably, when the robot fails to directly locate the target object in the initial stage, triggering the object search control mode includes:

[0027] Use the GroundingDINO model to detect the 2D bounding box of the target object;

[0028] Combined with the GraspNet model, the point cloud within the 2D bounding box is segmented into objects and grasp feasibility segments;

[0029] The three-dimensional bounding box, grasping points and grasping posture information of the target object are obtained from the information of object segmentation and grasping feasibility segmentation.

[0030] Preferably, the position and posture of the target object are integrated with the output information obtained by executing the subtask, and the subtask execution strategy is continuously optimized and adjusted using a reinforcement learning algorithm, including:

[0031] By querying the large language model, the reward function of the reinforcement learning subtask is generated according to the preset prompts, and the proximal policy optimization algorithm is used by default;

[0032] Directly generate effective control signals by optimizing the parameters of the strategy;

[0033] Predicting the best future actions for constrained proximal policy optimization algorithms via policies with a relative global perspective.

[0034] Preferably, the best future action of the proximal policy optimization algorithm constrained by policy prediction with a relatively global perspective includes:

[0035] Vision-based strategies use third-person RGB-D images of the current state as observations;

[0036] Generate the next action of the end effector based on the grasping position;

[0037] The difference between the observation and the next action is measured using the formula: Where T represents time synchronization, t represents the time synchronization index, α and β are weight coefficients, represents the divergence, represents the probability distribution of actions generated by the strategy under a given state, represents the probability of the vision-based strategy taking action α in state, represents the action generated at the time step based on the control strategy, Represents actions generated based on the vision policy.

[0038] In a second aspect, an embodiment of the present invention provides a data-driven robot closed-loop joint optimization system for implementing the method described in the above embodiment, the system comprising:

[0039] A generation module, the generation module is used to parse the acquired specific task description and generate three-dimensional scene information;

[0040] A manipulation module, which decomposes a specific task description into multiple subtasks and generates a subtask execution strategy based on the type of the subtask through a large language model and then executes the subtask;

[0041] A perception module, which is used to obtain the type of subtask and three-dimensional scene information, and recognize the position and posture of the target object through a visual algorithm;

[0042] The optimization module combines the position and posture of the target object with the output information obtained by executing the subtask, and uses the reinforcement learning algorithm to continuously optimize and adjust the subtask execution strategy.

[0043] In a third aspect, an embodiment of the present invention provides an electronic device, comprising at least one processor, which is communicatively connected to at least one memory, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the above embodiment.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. Data generation and scene diversification: The present invention generates a large number of three-dimensional simulation scenes through the scene generation module, which can provide sufficient data support for different types of tasks, greatly improving training efficiency and strategy generalization ability. Experiments show that the present invention improves the task success rate in the door opening task, with an average success rate of 83.2% in the rotating handle door opening task and an average success rate of 65.8% in the door pulling task.

[0046] 2. Perception-operation joint optimization: The perception-operation joint optimization module significantly improves the task completion rate through vision-based control correction. Experimental results show that the perception-operation joint optimization module improves the task success rate by 81.9% and shows strong robustness in different tasks.

[0047] 3. Closed-loop optimization and improved control accuracy: Through closed-loop joint optimization of perception and operation, while improving the accuracy of the strategy, it reduces the reliance on individual module optimization. The closed-loop system significantly reduces the cumulative error, allowing the robot to complete complex tasks more stably and accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0049] Figure 2 for Figure 1 Sub-flow chart of step S100;

[0050] Figure 3 for Figure 2 Sub-flow chart of step S140;

[0051] Figure 4 for Figure 1 Sub-flow chart of step S200;

[0052] Figure 5 for Figure 1 Sub-flow chart of step S400;

[0053] Figure 6 It is a schematic diagram of optimization of the optimization module in the embodiment system of the present invention;

[0054] Figure 7 It is a system module diagram of the present invention;

[0055] Figure 8 It is a system workflow diagram of the present invention;

[0056] Fig. 9 This is a module diagram of an electronic device of the present invention. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] See also Figure 1 The present invention provides a technical solution: a data-driven robot closed-loop joint optimization method for joint optimization of robot perception and manipulation. The method realizes closed-loop robot skill training by generating an infinite number of simulation environments. The method of the present invention proposes a perception-operation joint optimization training method, which optimizes the control strategy that only depends on egocentric control information through a visual strategy action sequence generated from a third-person global perspective, thereby forming a closed-loop skill learning system. The method comprises the following steps:

[0059] S100: parsing the acquired specific task description, generating three-dimensional scene information after parsing and decomposing the task into multiple subtasks;

[0060] S200: Identify the position and posture of the target object through a visual algorithm;

[0061] S300: Acquire the position and posture of the target object, and select a suitable strategy to execute the subtask through a large language model based on the type of the subtask;

[0062] S400: The position and posture of the target object are integrated with the output information obtained by executing the subtask, and the strategy of executing the subtask is continuously optimized and adjusted using the reinforcement learning algorithm.

[0063] See also Figure 2 In one embodiment of the present invention, step S100 further includes the following steps:

[0064] S110: Understand the specific task.

[0065] In this embodiment, the robot understands the task by obtaining a specific task description customized by the user. The robot first receives data related to the task. These data may include various types such as the target description of the task, environmental information, and object features. For example, in a warehouse handling task, the robot will receive the type of goods to be handled (such as the size and weight of the box), the target location (specific shelf coordinates in the warehouse), and the layout information of the warehouse (aisle location, shelf height and spacing, etc.). The robot's understanding of the task is the process of data processing. The data is processed through feature extraction. The processed data is used for task analysis and model matching in subsequent processes. After the match is successful, the corresponding strategy selection and optimization are performed.

[0066] S120: Acquire 3D materials for task understanding based on large-scale datasets.

[0067] In this embodiment, this step generates a 3D simulation scene suitable for embodied skill learning by configuring the scene and training. The 3D scene material of this module is based on existing large-scale data sets such as ParNet-Mobility and Objaverse. Furthermore, by randomization, different components of the asset are recombined to generate an extended asset data set.

[0068] S130: Obtain text and image embedding using feature integration mechanism.

[0069] In this embodiment, the text and image embeddings generated by the CLIP and SBERT models are integrated using a feature embedding mechanism to ensure that the objects in the scene meet the semantics of the task objectives. The feature embedding mechanism generates object numbers and detailed annotations through the GPT-4V large prophecy model, and compares this information with the BLIP-2 feature similarity of the objects in the database to select appropriate objects for scene layout. Furthermore, by dynamically optimizing the object scale and orientation, by calculating the relative spatial size of the object, strictly adjusting the size and spatial distribution of the generated object, and introducing hard constraints to avoid collisions and overlaps between objects, the physical rationality and visual consistency of the scene are ensured.

[0070] S140: A scene domain generalization method based on reinforcement learning, which adjusts the object layout by learning environmental characteristics and task requirements.

[0071] In this embodiment, a scene domain generalization method based on reinforcement learning is adopted. By learning environmental features and task requirements, the specific environment around the target object (such as walls and obstacles) is defined as a specific learning domain. In order to improve the generalization ability, the system randomly adjusts the lighting, shadows and textures in the scene, thereby increasing the complexity of visual perception. At the same time, by dividing the scene complexity into different levels (such as the complexity classification of doors), the generalization learning of the strategy is promoted, thereby improving the success rate and adaptability in actual tasks.

[0072] See also Figure 3 In one embodiment of the present invention, step S140 further includes:

[0073] S1401: Based on the custom material library, use an adaptive selection mechanism to randomize the texture of each object in the scene.

[0074] In this embodiment, the customized material library contains a variety of interior styles (such as modern, retro, minimalist, etc.) and is optimized for different room types.

[0075] S1402: Preferentially select compatible materials based on the spatial distribution and contextual information of the object.

[0076] In this embodiment, compatible materials are preferentially selected based on the spatial distribution of objects and context information to ensure consistency of style.

[0077] S1403: Combine reinforcement learning to learn domain features and enhance scene generalization capabilities.

[0078] In this embodiment, in order to enhance the scene generalization ability, this embodiment uses different room types and door locations, and combines reinforcement learning to learn domain features. By perturbing the texture, the model can not only learn surface features, but also learn the underlying spatial layout and geometric relationships, thereby improving its robustness in real-world applications.

[0079] S1404: Dynamically adjust the complexity of the scene.

[0080] In this embodiment, the complexity of the scene is defined as follows: in, is the space occupancy rate, which represents the ratio of the object volume to the total volume of the scene, It is the layout orderliness, which is used to measure the uniformity of object distribution. and is the weight, is the total volume of objects in the scene, is the total volume of the scene, is the average distance, is the scene diagonal distance.

[0081] Furthermore, in terms of scene diversity control, the module adopts a randomized control strategy to dynamically adjust the scene complexity factor. , which can generate different space ratios and layout orderliness scenes, thus achieving a variety of scene configurations from open spaces (such as large exhibition halls) to closed spaces (such as narrow corridors or small rooms). The system also introduces environmental elements such as wall decorations and columns, and randomly adjusts the position and height of objects, including furniture objects on the ground, hangings and walls, to create a 3D scene with a depth layer. This scene generation process not only makes the training environment of the visual system closer to the real life environment, but also improves the model's ability to adapt to complex and unpredictable environments during training through geometric complexity adjustment.

[0082] See also Figure 4 In one embodiment of the present invention, step S200 includes:

[0083] S210: When the robot fails to directly locate the target object in the initial stage, the object search control mode is triggered;

[0084] S220: Control the robot to search for a task target in the scene through an object search control mode;

[0085] S230: After identifying the target object, start path planning and navigate the robot to the target object location.

[0086] In the whole process of the above steps S210 to S230, the robot will perform collision detection between the environmental obstacles and the robot arm itself, and check whether it is reachable. When the task requires a grasping operation, the robot will further calculate the grasping evaluation index.

[0087] Specifically, step S210 of this embodiment uses the GroundingDINO model to detect the two-dimensional bounding box of the target object, and combines the GraspNet model to perform object segmentation and grasping feasibility segmentation on the point cloud within the two-dimensional bounding box, thereby obtaining the three-dimensional bounding box, grasping points and grasping posture information of the target object.

[0088] This embodiment integrates real-time perception capabilities, allowing the robot to perform tasks in realistic embodied scenarios and have the ability to perceive in unknown environments. It can also adaptively adjust the control strategy according to the complexity of the environment, thereby improving the task success rate and generalization ability.

[0089] In one embodiment of the present invention, step S300 includes:

[0090] Match each subtask to the most suitable execution method to improve the effect of skill learning. Match different motion planning and reinforcement learning for different task types. Motion planning is mainly used for tasks that require reaching the target object without collision, while reinforcement learning is more suitable for tasks that require frequent interaction with the environment. Decompose multi-step tasks into subtasks so that the corresponding execution method can be called for skill learning according to the type of subtask.

[0091] See also Figure 5 In one embodiment of the present invention, step S400 includes:

[0092] S410: Generate a reward function for each subtask based on the preset prompts through the large language model.

[0093] In this embodiment, the reward function for each subtask is generated according to the preset prompts by querying a large language model, and the Proximal Policy Optimization (PPO) algorithm is used by default. PPO is a control-based policy network with an observation space consisting of state variables related to the control task, such as joint angles and sensor readings.

[0094] In this embodiment, the preset prompts include the following: 1) natural language description of the subtask, 2) semantic annotations of the target object and its joints and links, 3) collision targets and constraints to be avoided, and 4) optional reward examples and adjustment directions.

[0095] In this embodiment, the GPT-4V model is used to parse the task requirements and generate a reward function based on the complexity of the environment and the robot's capabilities. For example, in the rotation task, the generated rewards include grasping rewards, posture rewards, joint angle rewards, and penalties for avoiding collisions.

[0096] S420: Directly generate an effective control signal by optimizing the parameters of the large language model.

[0097] In this embodiment, by optimizing parameters, PPO can directly generate effective control signals, ensure efficient control and stability in the action space, and is suitable for complex manipulation tasks that require precise and focused operations. The policy network is a multi-layer perceptron with a size of [128,128]. Each policy update optimizes the policy parameters through 50 stochastic gradient descent iterations. The generalized advantage estimate discount factor is set to 0.95 to adjust the smoothness of the advantage function estimate. The action space includes incremental translation and rotation.

[0098] S430: Predicting the best future action of the proximal policy optimization algorithm through a policy with a relatively global perspective.

[0099] In this embodiment, this embodiment is used as an optimization goal to help the robot reach the desired state more efficiently. In addition, this vision-based approach enhances the generalization ability of the strategy by using different state spaces, optimizing a more precise and focused control-based strategy.

[0100] Furthermore, in order to better integrate the control-based policy network and the pre-trained vision-based policy network and measure the difference between the actions generated by the two strategies, this application designs The loss function, which is used as part of the joint optimization policy update, combines the KL divergence and Euclidean distance to quantify the difference in the action probability distributions generated by the control policy and the vision policy.

[0101] Furthermore, The specific formula is as follows: in, is the action generated at the time step based on the control strategy, is an action generated based on a visual strategy, represents the probability distribution of actions generated by the policy under a given state. α and β are weight coefficients used to balance the contribution of KL divergence and Euclidean distance. KL divergence is used to measure the difference between the probability distributions of actions generated by the control and visual policies under a state. In order to further align actions, the Euclidean distance between actions is added as an additional constraint.

[0102] This embodiment constructs function, enabling the control-based policy to closely match the action probability distribution of the vision-based policy while maintaining a certain degree of proximity in the action space. By optimizing this total loss function as part of the control policy, the joint optimization can effectively learn the prior knowledge from the vision-based policy while maintaining the accuracy and specificity of the control-based policy.

[0103] See also Figure 7In an embodiment of the present invention, a data-driven robot closed-loop joint optimization system is provided to implement the method in the above embodiment. The system comprises:

[0104] A generation module 100, the generation module is used to parse the acquired specific task description and generate three-dimensional scene information;

[0105] A manipulation module 200, which decomposes a specific task description into multiple subtasks and generates a subtask execution strategy based on the type of the subtask through a large language model and then executes the subtask;

[0106] A perception module 300, which is used to obtain the type of subtask and three-dimensional scene information, and recognize the position and posture of the target object through a visual algorithm;

[0107] The optimization module 400 combines the position and posture of the target object with the output information obtained by executing the subtask, and uses a reinforcement learning algorithm to continuously optimize and adjust the subtask execution strategy.

[0108] See also Figure 8 , Figure 8 The workflow of each module of the system of the present invention is shown:

[0109] Step 1: The task starts, and the robot obtains the characteristic task description issued by the user;

[0110] Step 2: The generation module 100 understands the specific task description, generates three-dimensional scene information, and decomposes the specific task into multiple subtasks.

[0111] Based on the ParNet-Mobility and Objaverse datasets, this module integrates the text and image embeddings generated by the CLIP and SBERT models through a feature embedding mechanism to ensure that the generated 3D scenes meet the task requirements. The module uses randomized scene lighting, texture, and object distribution to optimize the layout and visual complexity of different environments.

[0112] Step 3: The manipulation module 200 first decomposes the user's specific task into multiple subtasks, and selects the corresponding algorithm to execute the subtask according to the type of subtask through the large model, decomposes complex tasks into subtasks, selects the optimal algorithm through the large language model, and integrates motion planning and reinforcement learning to ensure collision-free navigation and precise manipulation. Subtasks of different task types are matched with different methods to improve the effect of skill learning.

[0113] Step 4: The perception module 300 performs scene recognition on the three-dimensional scene, uses the GroundingDINO model to detect the two-dimensional bounding box, and combines GraspNet to segment the three-dimensional point cloud of the target object to obtain the grasping posture of the target object. The module supports the robot to find the target object in the adaptive embodied scene and realizes accurate task execution through path planning. And obtain the subtask type.

[0114] Step 5: The optimization module 400 obtains the position and posture of the perception module 300 and the information of the subtask execution of the manipulation module 200, integrates the control-based PPO strategy and the visual strategy through the joint optimization strategy, and uses KL divergence and Euclidean distance to measure the difference in the motion distribution generated by the two strategies. Determine whether the task success rate meets the standard. If the judgment result does not meet the standard, the module will learn prior knowledge from the third-person visual strategy while maintaining the accuracy of the control strategy, so as to further improve the robot's operation accuracy and task success rate. Please refer to Figure 6 , Figure 6 The optimization module 400 of the present invention is shown as a specific strategy model architecture based on the perception-manipulation joint optimization strategy. The entire model consists of multiple modules, and there is data flow and interaction between the modules. Among them, represents the received observation value, represents the original observation value, represents the displacement and rotation action in the Cartesian coordinate system, represents the actions produced by the visual strategy, represents the actions produced by the control strategy, represents the observation input of the control strategy, generated by the perception-operation joint optimization strategy, Represents the control loss function, which is used to connect the perception-operation joint optimization strategy and other modules. The loss function designed for this application is, and is the action for the next time step.

[0115] Table 1: Success rate of the robot in the door opening and rotation actions according to the number of optimizations Through the above steps, the method combined with the system provided in the embodiment of the present application can significantly improve the robot's action success rate.

[0116] Step 6: The manipulation module 200 continues to decompose the specific task into subtasks until there are no more subtasks that can be decomposed, and the task success rate of the optimization module 400 reaches a predetermined target value, and then the task ends.

[0117] See also Fig. 9 , Fig. 9A schematic diagram of an electronic device 20 that can implement an embodiment of the present invention is shown, and the electronic device is intended to represent various forms of control devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0118] The electronic device 20 includes at least one processor 21, and a memory connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 22 or the computer program loaded from the storage unit 28 to the random access memory (RAM) 13. In the RAM 23, various programs and data required for the operation of the electronic device 20 can also be stored. The processor 21, ROM 22 and RAM 23 are connected to each other through a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.

[0119] A number of components in the electronic device 20 are connected to the I / O interface 25, including: an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0120] The processor 21 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 21 performs the various methods and processes described above.

[0121] In some embodiments, the method of the above embodiment may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 28. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 20 via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the processor 21, one or more steps in the method described above may be performed. Alternatively, in other embodiments, the processor 21 may be configured to perform the method of the above embodiment in any other appropriate manner (e.g., by means of firmware).

[0122] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0123] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0124] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0125] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0126] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0127] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0128] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0129] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A data-driven robot closed-loop joint optimization method, characterized in that: include: Parse the acquired specific task description to generate three-dimensional scene information; Decompose a specific task description into multiple subtasks and generate a subtask execution strategy based on the subtask type through a large language model to execute the subtask; Obtain the type of subtask and 3D scene information, and identify the position and posture of the target object through visual algorithms; The position and posture of the target object are integrated with the output information obtained from executing the subtask, and the subtask execution strategy is continuously optimized and adjusted using the reinforcement learning algorithm, including: by querying the large language model, generating the reward function of the reinforcement learning subtask according to the preset prompts, using the proximal policy optimization algorithm by default, directly generating effective control signals by optimizing the parameters of the strategy, and using the strategy with a relatively global perspective to predict the best future action of the proximal policy optimization algorithm; The best future actions of the proximal policy optimization algorithm constrained by policy prediction with a relatively global perspective include: Vision-based strategies use third-person RGB-D images of the current state as observations; Generate the next action of the end effector based on the grasping position a t+1 ; Use L PMO Weighing observations and next actions t+1 The difference between them is expressed as follows: Where T represents time synchronization, t represents the time synchronization index, α and β are weight coefficients, and D KL represents the divergence, represents the probability distribution of actions generated by the strategy under a given state, represents the probability of the vision-based strategy taking action α in state, represents the action generated at the time step based on the control strategy, Represents actions generated based on the vision policy.

2. The data-driven robot closed-loop joint optimization method according to claim 1, characterized in that: The step of parsing the acquired specific task description to generate three-dimensional scene information includes: Conduct task understanding for specific tasks; Acquire 3D materials for task understanding based on large-scale datasets; Use feature integration mechanism to obtain text and image embedding; A scene domain generalization method based on reinforcement learning adjusts object layout by learning environmental characteristics and task requirements.

3. The data-driven robot closed-loop joint optimization method according to claim 2, characterized in that: A scene domain generalization method based on reinforcement learning adjusts the object layout by learning environmental characteristics and task requirements, including: Based on a custom material library, an adaptive selection mechanism is used to randomize the textures of objects in the scene. Prioritize compatible materials based on the spatial distribution of objects and contextual information; Combine reinforcement learning to learn domain features and enhance scene generalization capabilities; Dynamically adjust the complexity of the scene.

4. The data-driven robot closed-loop joint optimization method according to claim 3 is characterized in that: The complexity of the scenario is defined as follows: C scene =w1·R occupancy +w2·(1-O layout ) Among them, R occupcy is the space occupancy rate, which represents the ratio of the object volume to the total volume of the scene, O layout is the layout order, which is used to measure the uniformity of object distribution. w1 and w2 are weights, and v objects is the total volume of objects in the scene, v scene is the total volume of the scene, d avg is the average distance, d max is the scene diagonal distance.

5. The data-driven robot closed-loop joint optimization method according to claim 4, characterized in that: The obtaining of the subtask type and three-dimensional scene information and identifying the position and posture of the target object through a visual algorithm include: When the robot fails to directly locate the target object in the initial stage, the object search control mode is triggered; The object search control mode is used to control the robot to search for the task target in the scene; After the target object is identified, path planning is started to navigate the robot to the target object location.

6. The data-driven robot closed-loop joint optimization method according to claim 5, characterized in that: When the robot fails to directly locate the target object in the initial stage, the object search control mode is triggered, including: Use the GroundingDINO model to detect the 2D bounding box of the target object; Combined with the GraspNet model, the point cloud within the 2D bounding box is segmented into objects and grasp feasibility segments; The three-dimensional bounding box, grasping points and grasping posture information of the target object are obtained from the information of object segmentation and grasping feasibility segmentation.

7. A data-driven robot closed-loop joint optimization system, used to implement any of the methods described in claims 1 to 6, characterized in that: The system includes: A generation module, the generation module is used to parse the acquired specific task description and generate three-dimensional scene information; A manipulation module, which decomposes a specific task description into multiple subtasks and generates a subtask execution strategy based on the type of the subtask through a large language model and then executes the subtask; A perception module, which is used to obtain the type of subtask and three-dimensional scene information, and recognize the position and posture of the target object through a visual algorithm; The optimization module combines the position and posture of the target object with the output information obtained by executing the subtask, and uses the reinforcement learning algorithm to continuously optimize and adjust the subtask execution strategy.

8. An electronic device, characterized in that: The method comprises at least one processor, wherein the processor is communicatively connected to at least one memory, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Structure inspection agent navigation method based on damage driving and multi-mode multi-task learning

    CN116824303A

  • Multi-modal data fusion method and system based on traffic large model

    CN117058512A