A robot operation method for open instruction feedback reasoning and guidance optimization in a human-robot collaboration scenario

By combining large language models and visual language models, along with inverse reasoning and GraspNet models, the problem of diverse parsing of natural language commands and target recognition in human-machine collaboration scenarios is solved, enabling robots to perform high-precision operations and improve grasping success rates in complex tasks.

CN121157023BActive Publication Date: 2026-04-14HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2025-09-26
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies struggle to handle the open and diverse needs expressed by users in natural language in human-machine collaboration scenarios. Traditional language parsing methods lack contextual understanding and task abstraction capabilities, resulting in low recognition accuracy and poor operational accuracy for robots in complex tasks. Furthermore, traditional visual detection models have difficulty identifying and locating unknown targets, leading to a high failure rate in grasping.

Method used

A large language model is used to convert natural language instructions into task execution sequences. By combining a visual language model and a GraspNet model, a two-stage semantic guidance optimization strategy and a reverse reasoning mechanism are used to achieve accurate identification of target objects and determination of the optimal grasping posture. Category consistency loss and cross-modal attention consistency loss are introduced to improve the model's recognition accuracy and grasping success rate in complex industrial scenarios.

Benefits of technology

It improves the robot's recognition accuracy and operational accuracy in complex collaborative scenarios, significantly reduces target recognition errors and grasping failure rates, and enhances the robot's robustness in dynamic response and operational accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121157023B_ABST
    Figure CN121157023B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of collaborative robots, and discloses a robot operation method for open instruction feedback reasoning and guidance optimization in a human-machine cooperation scene. The method comprises the following steps: converting a natural language instruction into a task execution sequence by using a large language model, screening a target object corresponding to a multi-step task in the task execution sequence to obtain a target object corresponding to a current task; inputting an image of a scene where the obtained target object is located into a visual language model to obtain a bounding box of the target object; the visual language model is connected with an Adapter feature adapter and a LoRA low-rank adapter at the end; converting the bounding box into a target mask, solving an optimal pose of a robot for grabbing the target object according to the image of the scene where the target object is located and the target mask, and the robot grabs the target object according to the optimal pose. Through the application, the recognition accuracy, operation accuracy and robustness of the robot in a complex cooperation and multi-target scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of collaborative robots, and more specifically, relates to a robot operation method with open instruction feedback reasoning and guided optimization for human-robot collaboration scenarios. Background Technology

[0002] Driven by the rapid development of artificial intelligence technology, modern manufacturing is placing higher demands on the flexibility and adaptability of robot systems, especially when dealing with complex and dynamic task scenarios in personalized and unstructured production environments. These tasks impose a dual requirement on robot systems: they need to achieve high-precision operation and grasping execution, and also possess the ability to parse natural language commands and adaptively adjust to dynamic operational goals.

[0003] Current robot operation methods in human-robot collaborative scenarios mostly rely on structured instructions or pre-set task flows, making it difficult to handle the open and diverse needs expressed by users in natural language. Traditional language parsing methods lack contextual understanding and task abstraction capabilities, making it difficult to accurately identify user intentions and effectively decompose and plan tasks, especially in complex tasks involving multiple objects and multiple steps. This makes it difficult to achieve true human-robot interaction and autonomous task planning in complex collaborative tasks. In addition, during perception and prediction, industrial scenarios contain a large number of open-set targets (such as new tool models, parts, etc.), which traditional visual detection models and pose estimation methods struggle to identify and locate. This leads to problems such as target recognition errors and high grasping failure rates in actual operation, failing to meet the robot's requirements for operational accuracy and dynamic response. Summary of the Invention

[0004] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a robot operation method based on open instruction feedback reasoning and guided optimization for human-robot collaboration scenarios, solving the problems of low recognition accuracy, operational accuracy, and robustness in multi-target robot scenarios.

[0005] To achieve the above objectives, according to one aspect of the present invention, a robot operation method for open instruction feedback reasoning and guided optimization in human-robot collaboration scenarios is provided, characterized in that the method includes the following steps:

[0006] The natural language instructions are converted into a task execution sequence, and the target objects of the multi-step tasks in the task execution sequence are filtered to obtain the target object corresponding to the current task.

[0007] The image of the scene where the target object of the current task is located is input into the visual language model to obtain the bounding box of the target object; the visual language model is connected in series with an Adapter feature adapter and a LoRA low-rank adapter at the end;

[0008] The bounding box is converted into a target mask. Based on the image of the scene where the target object is located and the target mask, the optimal pose for the robot to grasp the target object is calculated. The robot then grasps the target object according to this optimal pose.

[0009] More preferably, the training process of the visual language model is as follows:

[0010] The image labeled with the bounding box of the target object is input into the initial visual language model for training, and the weights W1 between each network layer in the initial visual language model are obtained.

[0011] The visual language model is improved by concatenating one Adapter feature adapter and multiple LoRA low-rank adapters in the initial visual language model, while keeping the initial weights of the initial visual language model unchanged. The improved visual language model is trained using an image labeled with the bounding box of the target object, and the weights W2 between the one Adapter feature adapter and multiple LoRA low-rank adapters and the initial visual language model are obtained.

[0012] The weights W1 and W2 are combined to obtain weight W, which is the weight of the visual language model required after training, thereby realizing the training of the visual language model.

[0013] More preferably, during the training of the initial visual language model to obtain the weights W1, the loss function is calculated using the following formula:

[0014]

[0015]

[0016] in, It is the bounding box regression loss. , For loss weighting coefficients, It is the bounding box regression loss, which measures the difference between the predicted bounding box and the true bounding box. It is a generalized comparison of losses. This is the contrastive loss, where i is the training sample number. The total number of training samples, For predicting categories t The confidence probability, For category weight parameters, It is a moderating factor used to control the contribution of easily classified and difficult-to-classify samples to the loss.

[0017] More preferably, the loss function for training the improved visual language model is:

[0018]

[0019]

[0020]

[0021] in, It is a class consistency loss. It is the total number of samples. It is the first i The predicted class probability of each sample. It is the first i The true class label of each sample It is the cross-entropy loss function. It is cross-modal consistency loss. It refers to the number of attention layers. It is the number of attention heads per layer. It is the attention matrix of the h-th head in the k-th layer. It is the average attention moment of all heads in the k-th layer. It is the Frobenius norm. This is the total loss from the second phase of training. , , , These are the weight coefficients for bounding box regression loss, contrastive classification loss, class consistency loss, and cross-modal consistency loss, respectively. It is the bounding box regression loss. It is a comparative loss.

[0022] More preferably, the optimal pose for the robot to grasp the target object is determined using the GraspNet model, calculated as follows:

[0023] Objective function:

[0024]

[0025]

[0026]

[0027]

[0028] Constraints:

[0029] in, It is the final selection of the optimal grasping position and posture. It is the score of the j-th crawling candidate, reflecting the stability and feasibility of crawling. It is the angle between the grasping posture and the direction of gravity. , These are the weighting coefficients. It is the position of the grab point (X, Y, Z). It is the grasping posture (rotation matrix). It is the score or confidence level of the candidate grasping posture. The crawling candidates are sorted from highest to lowest score, and G is the set of crawling candidate poses. descending = True According to Sort the values ​​from highest to lowest. It is the grasping posture Mapped to rotation matrix form, The direction of g is the spindle direction of the end effector, and g is the direction of gravity. Detect candidate grasp poses in scene obstacles V The value indicates whether a collision has occurred; 1 indicates no collision, and 0 indicates no collision. The selected grasping pose must be optimized under collision-free constraints to ensure the feasibility and safety of grasping.

[0030] More preferably, the screening of target objects corresponding to multiple tasks in the task execution sequence adopts a reverse reasoning mechanism.

[0031] More preferably, the anti-reasoning mechanism is performed according to the following formula:

[0032]

[0033]

[0034]

[0035] in, Candidate target The score is based on the match between the task instructions and the task instructions. This is the current task instruction, and S is the task step number. It is based on the rating For candidate set Sort the sequences from highest to lowest matching degree. The ultimate operational goal is... It is the set of candidate targets.

[0036] More preferably, the process of converting natural language instructions into a task execution sequence using a large language model is as follows:

[0037]

[0038] in, It is a structured task execution sequence. It is a task planning and inference function based on a large language model. It is a scene RGB image. This is the current task instruction. This is a system prompt set by the large language model. It is the task step number. It refers to the type of the object being operated on.

[0039] More preferably, the conversion of the bounding box information of the target object into a target mask is achieved using the SAM segmentation model.

[0040] According to another aspect of the present invention, a robot operating system with open instruction feedback reasoning and guidance optimization for human-robot collaboration scenarios is provided. This system includes a target object acquisition module, a visual language model module, and an optimal pose acquisition module, wherein:

[0041] The target object acquisition module is used to determine the target object corresponding to the current task based on natural language instructions;

[0042] The visual language model is used to obtain the bounding box of the target object in an image of the scene in which the target object is located;

[0043] The optimal pose acquisition module is used to calculate the optimal pose for the robot to grasp the target object by using the bounding box of the target object and the image of the scene in which the target object is located.

[0044] In summary, compared with the prior art, the above-described technical solutions conceived by this invention have the following technical effects:

[0045] 1. This invention introduces a collaborative reasoning mechanism involving a large language model, a visual language model, and the determination of the optimal pose to construct a mapping process from natural language instructions to the machine's optimal pose. The purpose is to integrate the understanding and reasoning capabilities of the large language model with the spatial perception capabilities of the visual language model, thereby accurately understanding the diverse and open-ended instructions expressed by users and improving the robot's recognition accuracy, operational accuracy, and robustness in complex collaborative and multi-objective scenarios.

[0046] 2. This invention proposes a two-stage semantic-guided optimization strategy to optimize the visual language model. It introduces category consistency loss and cross-modal attention consistency loss, which effectively improves the model's recognition accuracy of unknown targets and cross-modal semantic alignment ability in complex industrial scenarios, and enhances the stability and generalization of detection results.

[0047] 3. In the process of calculating the optimal posture of the robot, this invention constructs a grasping candidate set based on the GraspNet model, and introduces a grasping scoring mechanism based on multiple dimensions such as grasping score, posture angle and collision detection, so as to realize the multi-objective optimal selection of grasping action and significantly improve the grasping success rate and execution robustness.

[0048] 4. This invention effectively alleviates the problem of ambiguity in operation targets through a reverse reasoning mechanism. After parsing the user's natural language instructions, if there are multiple candidate operation targets, the system will combine scene information to perform secondary semantic scoring on the candidate targets and confirm the final target through interaction with the user, ensuring the uniqueness and accuracy of the final operation target. This significantly improves the robot's instruction parsing ability and task execution targeting in complex human-machine collaborative tasks. Attached Figure Description

[0049] Figure 1 This is a robot operation method for open instruction feedback reasoning and guided optimization oriented to human-machine collaboration scenarios, constructed according to a preferred embodiment of the present invention.

[0050] Figure 2 This is a schematic diagram of the structure of a robot operation with open instruction feedback reasoning and guidance optimization for human-machine collaboration scenarios, constructed according to a preferred embodiment of the present invention.

[0051] Figure 3 This is a schematic diagram of the structure of the improved visual language model constructed according to a preferred embodiment of the present invention.

[0052] Figure 4 This is an example of a robot operation platform constructed according to a preferred embodiment of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0054] This invention provides a robot operation method for open command feedback reasoning and guided optimization in human-robot collaboration scenarios. First, a large language model accurately understands the user's open commands, generating a structured task planning sequence. This sequence is then further combined with scenario information for feedback reasoning, confirming the final target object through a feedback mechanism. Second, a visual language model identifies and locates target objects in the task steps. A two-stage semantic guidance optimization strategy further improves the accuracy and generalization performance of the visual language model in target recognition within industrial human-robot collaboration scenarios. Finally, using the obtained target mask and depth information, a multi-constraint grasping optimization module generates candidate grasping poses, which are then optimized using a grasping scoring function. The optimal grasping pose is output for the robot to execute the task.

[0055] Combination Figure 2Current robot operation methods in human-robot collaborative scenarios have limited ability to understand and execute natural language commands, lacking semantic understanding of open and diverse natural language expressions. This makes it difficult to achieve true human-robot interaction and autonomous task planning in complex collaborative tasks. Furthermore, during perception and prediction, industrial scenarios contain a large number of open-set targets (such as new tool models and parts), which traditional closed-set visual detection models struggle to identify. While open-set target detection models have strong generalization performance, their ability to identify targets in industrial scenarios is poor, resulting in a high failure rate in actual collaborative assembly scenarios and failing to meet the robot's requirements for operational accuracy and dynamic response. Therefore, this invention provides a natural language-guided robot operation method for human-robot collaborative scenarios.

[0056] like Figure 1 As shown, a robot operation method based on open instruction feedback reasoning and guided optimization for human-robot collaboration scenarios includes:

[0057] S1. Utilize a large language model to convert natural language instructions into task execution sequences; employ a feedback inference mechanism to analyze the objects being operated on within the task execution sequences. Perform the screening;

[0058] In this embodiment, a robot platform (see...) Figure 3 The depth camera captures an RGB-D image of the scene from the current viewpoint, and the user inputs open commands to the robot for the task at hand.

[0059] Based on the large language model GPT-4o, natural language commands from users in different scenarios are processed. The task is parsed, and its scope is limited by system-level prompts, enabling it to convert natural language instructions into a structured task execution sequence. Each step of the task Includes the objects that the robot needs to operate in the current step. Structured task sequence 1) Passing prior knowledge to the target detection module; 2) Feedback reasoning mechanism: during the task planning phase... The output contains multiple optional targets, which leads to ambiguity in task execution. That is, when the robot cannot confirm the target of the current step, the system will trigger a feedback reasoning mechanism to confirm the final target.

[0060] Furthermore, GPT-4o ensures the structured, logical, and practically executable nature of task sequences through predefined operators and objects, a process characterized as follows:

[0061]

[0062] in, Indicates the task step number. Indicates the type of object being operated on. Represents the RGB image of the scene. Indicates the current task instruction. This indicates a system prompt set for the large language model.

[0063] Feedback reasoning mechanism: for the candidate target set Each candidate target in The large language model further combines scene information to perform secondary reasoning and scoring, calculating the candidate target. With current task instructions Match score between .

[0064] After obtaining the rating set, sort it. The system will return the scoring results to the user in the form of prompts, and the user will select the final target object. Update the confirmed targets to the task sequence. Used to guide the subsequent target detection module in identification and localization;

[0065] In this embodiment, during the task planning phase... The output contains multiple optional targets, leading to ambiguity in task execution. Specifically, when the robot cannot determine the target for the current step, the system will trigger a feedback reasoning mechanism to confirm the final target. For the candidate target set... Each candidate target in The large language model further combines scene information to perform secondary reasoning and scoring, calculating the candidate target. With current task instructions Match score between After obtaining the rating set, sort it. The system will return the rating results to the user in a prompt format, and the user will select the final target category. Update the confirmed targets to the task sequence. This is used to guide the subsequent semantic object detection module in identification and localization. The feedback reasoning process can be represented as:

[0066]

[0067]

[0068]

[0069] in, Candidate targets The score is based on the match between the task instructions and the task instructions. Indicates based on rating For candidate set Sort the sequences from highest to lowest matching degree; This indicates the final operational goal.

[0070] S2. For the target object Perform identification and positioning to determine the target object. Location

[0071] Semantic object detection module: Based on the visual language model Grounding DINO, the input is the target object. The output is the bounding box of the target object, along with an RGB image. With category labels .

[0072] The model is trained using a two-stage semantic-guided optimization strategy. The first stage releases some parameters of the GroundingDINO model and uses bounding box regression loss. And comparative loss Obtain weights through training The second stage freezes the Grounding DINO backbone parameters and injects LoRA low-rank adapters into the cross-modal attention layer, self-attention layer, and feedforward network layer. This efficiently captures globally shared features between multiple tasks in the form of a low-rank matrix. At the same time, an Adapter feature adaptation module is added to compensate and optimize local feature channels, and class consistency loss is incorporated. And cross-modal attention consistency loss Obtain weights through training Finally, the two weights are combined to form the final weights of the model. The trained model receives structured task sequences output by a large language model. Identify and output the bounding box of the target object. With category labels .

[0073] The bounding box information of the target object is then input into the SAM segmentation model to obtain the target object mask. This guides the robot to accurately perform subsequent tasks; furthermore, after completing target detection, the segmentation model uses the SegmentAnyting model to generate a target region mask.

[0074] In this embodiment, based on the visual language model Grounding DINO, a two-stage semantic-guided optimization strategy is used to train the model. The first stage releases some parameters of the Grounding DINO model and uses bounding box regression loss. And comparative loss Obtain weights through training The loss in the first stage is expressed as:

[0075]

[0076]

[0077] in, It is the bounding box regression loss. , These are loss weighting coefficients used for adjustment. L1 Loss and GIoU The relative importance of losses in total losses. It is a comparison of losses. The total number of training samples, For predicting categories t The confidence probability, These are the category weight parameters.

[0078] The second stage freezes the Grounding DINO backbone parameters and injects LoRA low-rank adapters into the cross-modal attention layer, self-attention layer, and feedforward network layer. This efficiently captures globally shared features between multiple tasks in the form of a low-rank matrix. At the same time, an Adapter feature adaptation module is added to compensate and optimize local feature channels, and class consistency loss is incorporated. And cross-modal attention consistency loss Obtain weights through training The second-stage loss is expressed as:

[0079]

[0080]

[0081]

[0082] in, This represents the category consistency loss, which ensures consistency between the predicted category and the instruction category and improves semantic alignment. The total number of samples, For the first i The predicted class probability of each sample. For the first i The true class label of each sample Cross-entropy loss function measures the difference between the predicted distribution and the true distribution; Indicates cross-modal consistency loss. For the number of attention layers, The number of attention heads per layer, Let h be the attention matrix of the h-th head in the k-th layer. Let be the average attention moment of all heads in the k-th layer. The Frobenius norm measures the difference between matrices.

[0083] in, This represents the total losses from the second phase of training. , , , These are the weight coefficients for bounding box regression loss, contrastive classification loss, class consistency loss, and cross-modal consistency loss, respectively, used to balance the contribution of each sub-loss to the total loss.

[0084] Furthermore, the semantic object detection module optimizes the parameters of the Ground DINO model using a two-stage semantic-guided optimization strategy. The final model weights are expressed as follows:

[0085]

[0086] in, To obtain model weights for the first stage of training; To obtain model weights for the second stage of training.

[0087] The trained model receives structured task sequences output by a large language model. Identify and output the bounding box of the target object. With category labels Then, the bounding box information of the target object is input into the SAM segmentation model to obtain the target object mask. This guides the robot to accurately perform subsequent tasks.

[0088] The bounding box information of the target object is then input into the SAM segmentation model to obtain the target object mask. This guides the robot to accurately perform subsequent tasks; furthermore, after completing target detection, the segmentation model uses the SegmentAnyting model to generate a target region mask.

[0089] S3. Multi-constraint grasping optimization module: Target mask based on SAM generation and original RGB image and depth map Use the GraspNet model to generate a set of crawl locations. By combining the capture score, orientation, and collision detection, the optimal capture pose is determined within the restricted area based on the optimal capture selection function. Input the robot's pose to guide it in accurately performing the task.

[0090] Furthermore, the multi-constraint grasping optimization module uses the GraspNet model to generate grasping poses and generates target masks based on SAM. and original RGB image and depth map The optimal grasping pose is determined within a restricted area by comprehensively considering the grasping score, orientation, and collision detection, based on the optimal grasping selection function. In this embodiment, the target mask is generated based on SAM. and original RGB image and depth map The point cloud information of the target region is extracted, and the GraspNet model is used to generate grabbing candidates. GraspNet outputs a set of grabbing candidate poses based on the input RGB-D information.

[0091]

[0092] in, Indicates the position of the grab point (X, Y, Z); Indicates the grasping posture (rotation matrix); This indicates the score or confidence level of the candidate's grasping posture.

[0093] Sort the candidate set by score:

[0094]

[0095] To ensure the rationality of the grasping direction, the alignment between the direction of gravity and the tool tip is introduced as an evaluation metric, defined as follows:

[0096]

[0097] in, Indicates that the gesture will be captured. Mapped to rotation matrix form, This refers to the spindle direction of the end effector. This angle is used to prioritize gripping from "top to bottom" or in a direction consistent with gravity.

[0098] Further, a collision detection function is introduced:

[0099]

[0100] in, V This represents the model of all collidable objects in the scene. The pose is considered valid only if there is no collision.

[0101] Finally, a multi-objective weighted selection strategy is used to select the optimal grasping pose pair:

[0102]

[0103] in, The final optimal gripping position and posture; The score of the j-th crawling candidate reflects the stability and feasibility of crawling; To ensure a stable grip, the smaller the angle between the gripping posture and the direction of gravity, the better. , This is a weighting coefficient that balances the importance of the grab score and the grab angle constraint. The constraint ensures that the selected grab pose will not collide in the actual scene.

[0104] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A robot operation method for open instruction feedback reasoning and guided optimization in human-robot collaboration scenarios, characterized in that, The method includes the following steps: The natural language instructions are converted into a task execution sequence, and the target objects of the multi-step tasks in the task execution sequence are filtered to obtain the target object corresponding to the current task. Input the image of the scene where the target object of the current task is located into the visual language model to obtain the bounding box of the target object; The visual language model is connected in series with an Adapter feature adapter and a LoRA low-rank adapter at its end. The bounding box is converted into a target mask. The optimal pose for the robot to grasp the target object is calculated based on the image of the scene where the target object is located and the target mask. The robot grasps the target object according to the optimal pose. The training process of the visual language model is as follows: The image labeled with the bounding box of the target object is input into the initial visual language model for training, and the weights W1 between each network layer in the initial visual language model are obtained. The visual language model is improved by concatenating one Adapter feature adapter and multiple LoRA low-rank adapters in the initial visual language model, while keeping the initial weights of the initial visual language model unchanged. The improved visual language model is trained using an image labeled with the bounding box of the target object, and the weights W2 between the one Adapter feature adapter and multiple LoRA low-rank adapters and the initial visual language model are obtained. The weights W1 and W2 are combined to obtain weight W, which is the weight of the visual language model required after training, thereby realizing the training of the visual language model. The optimal pose for the robot to grasp the target object is determined using the GraspNet model, calculated as follows: Objective function: Constraints: in, It is the final selection of the optimal grasping position and posture. It is the score of the j-th crawling candidate, reflecting the stability and feasibility of crawling. It is the angle between the grasping posture and the direction of gravity. , These are the weighting coefficients. It is the position of the grab point (X, Y, Z). It is the grasping posture (rotation matrix). It is the score or confidence level of the candidate grasping posture. The crawling candidates are sorted from highest to lowest score, and G is the set of crawling candidate poses. descending = True According to Sort the values ​​from highest to lowest. It is the grasping posture Mapped to rotation matrix form, The direction of g is the spindle direction of the end effector, and g is the direction of gravity. Detect candidate grasp poses in scene obstacles V The value indicates whether a collision has occurred; 1 indicates no collision, and 0 indicates no collision. The selected grasping pose must be optimized under collision-free constraints to ensure the feasibility and safety of grasping. The process of filtering target objects corresponding to multiple steps in the task execution sequence employs a reverse reasoning mechanism. The anti-reasoning mechanism operates according to the following formula: in, Candidate target The score is based on the match between the task instructions and the task instructions. This is the current task instruction, and S is the task step number. It is based on the rating For candidate set Sort the sequences from highest to lowest matching degree. The ultimate operational goal is... It is the set of candidate targets.

2. The robot operation method for open instruction feedback reasoning and guided optimization in human-machine collaboration scenarios as described in claim 1, characterized in that, During the training of the initial visual language model to obtain the weights W1, the loss function is calculated using the following formula: in, It is the bounding box regression loss. , For loss weighting coefficients, It is the bounding box regression loss, which measures the difference between the predicted bounding box and the true bounding box. It is a generalized comparison of losses. This is the contrastive loss, where i is the training sample number. The total number of training samples, For predicting categories t The confidence probability, For category weight parameters, It is a moderating factor used to control the contribution of easily classified and difficult-to-classify samples to the loss.

3. The robot operation method for open instruction feedback reasoning and guided optimization in human-machine collaboration scenarios as described in claim 1, characterized in that, The loss function for training the improved visual language model is: in, It is a class consistency loss. It is the total number of samples. It is the first i The predicted class probability of each sample. It is the first i The true class label of each sample It is the cross-entropy loss function. It is cross-modal consistency loss. It refers to the number of attention layers. It is the number of attention heads per layer. It is the attention matrix of the h-th head in the k-th layer. It is the average attention moment of all heads in the k-th layer. It is the Frobenius norm. This is the total loss from the second phase of training. , , , These are the weight coefficients for bounding box regression loss, contrastive classification loss, class consistency loss, and cross-modal consistency loss, respectively. It is the bounding box regression loss. It is a comparative loss.

4. The robot operation method for open instruction feedback reasoning and guided optimization in human-robot collaboration scenarios as described in claim 1, characterized in that, The process of converting natural language instructions into task execution sequences using a large language model is as follows: in, It is a structured task execution sequence. It is a task planning and inference function based on a large language model. It is a scene RGB image. This is the current task instruction. This is a system prompt set by the large language model. It is the task step number. It refers to the type of the object being operated on.

5. The robot operation method for open instruction feedback reasoning and guided optimization in human-machine collaboration scenarios as described in claim 1, characterized in that, The conversion of the bounding box information of the target object into a target mask is achieved using the SAM segmentation model.

6. A robot operating system with open instruction feedback reasoning and guidance optimization for human-machine collaboration scenarios, wherein the system employs the operating method described in any one of claims 1-5, characterized in that, The system includes a target object acquisition module, a visual language model module, and an optimal pose acquisition module, wherein: The target object acquisition module is used to determine the target object corresponding to the current task based on natural language instructions; The visual language model is used to obtain the bounding box of the target object in an image of the scene in which the target object is located; The optimal pose acquisition module is used to calculate the optimal pose for the robot to grasp the target object by using the bounding box of the target object and the image of the scene in which the target object is located.

Citation Information

Patent Citations

  • Vision-language-action joint modeling-based disordered scene target object capturing method

    CN115861596A

  • Exhibition hall robot visual language navigation method based on large model

    CN119309580A