Humanoid robot wire harness operation large model training method based on meta-action space

By constructing a meta-action space and combining reinforcement learning with lifelong learning of human behavior intervention, the problem of lack of harness operation data of humanoid robots is solved, and their generalization ability and execution efficiency in complex harness operations are improved.

CN119885863BActive Publication Date: 2025-10-10TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411931976.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-10-10
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

In existing technologies, humanoid robots have insufficient generalization capabilities in complex wiring harness operation scenarios, lack sufficient training data, and long sequence operation tasks require high performance of large models, resulting in poor execution results.

Method used

A humanoid robot harness manipulation method based on meta-action space constructs a harness manipulation meta-action dataset, uses reinforcement learning to train a large model, and combines lifelong learning with human behavior intervention to generate more accurate harness manipulation strategies.

Benefits of technology

It improves the humanoid robot's understanding and execution efficiency of complex wiring harness operations, enhances the adaptability and execution efficiency of specific tasks, and realizes lifelong learning of wiring harness operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119885863B_ABST
    Figure CN119885863B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on meta-action space humanoid robot wire harness operation big model training, this method includes: based on human wire harness operation constructs wire harness operation meta-action;Wire harness operation meta-action data set is constructed;Based on the wire harness operation meta-action data set, utilize reinforcement learning to train humanoid robot wire harness operation big model;Input data is acquired, utilizes the humanoid robot wire harness operation big model that is trained to output joint motor parameters and is based on the joint motor parameters real-time update input data;Basic strategy is acquired, and residual strategy is acquired based on the basic strategy;Based on the basic strategy and residual strategy, humanoid robot wire harness operation is carried out.Compared with prior art, the present application improves the understanding and execution of model to complex wire harness operation task, solves the problem of lack of wire harness operation training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of embodied intelligence, and in particular to a method for training a large model of harness manipulation of a humanoid robot based on a meta-action space. Background Art

[0002] As a key component connecting various electronic devices, sensors, and control modules in a car, automotive wiring harnesses are an indispensable part of the vehicle's electrical system. Traditional assembly line production methods require a large number of people to perform heavy manual labor, which not only places a huge physical burden on workers but also leads to high labor and operating costs for factories. In recent years, with the development of technologies such as artificial intelligence, machine learning, and automation, the intelligence level of humanoid robots has been significantly improved, showing great application and development potential in industrial production, maintenance, medical care, and daily life services. Given the similar limb structure and movement patterns of humanoid robots to humans, research on humanoid robots to autonomously operate wiring harnesses to replace traditional assembly line methods shows great prospects.

[0003] At present, the methods for training humanoid robots at home and abroad focus on large-scale model technology, which has achieved efficient environmental perception, autonomous decision-making, and intelligent interaction. Although humanoid robots can complete a variety of tasks with the help of rich Internet knowledge learned by pre-trained large-scale models, and even show good generalization ability when facing some simple objects, scenes, and tasks that are not visible in training, they perform poorly in dexterous and complex tasks such as wire harness operation. There are some limitations: (1) The types and styles of wire harnesses in actual production are numerous and complex, and the difficulty of operation is much higher than the task scenarios when pre-training large-scale models. The method of directly deploying large-scale models without fine-tuning cannot achieve the expected results; (2) The datasets of wire harness production operations are relatively scarce, especially in specific tasks and detailed operations. There is a lack of sufficient training data to support the effective training and optimization of large-scale models; (3) The wire harness production process often involves long sequences of operation tasks, which puts higher requirements on the performance of large-scale models. Summary of the Invention

[0004] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a large-scale model training method for humanoid robot harness operation based on meta-action space, aiming to solve the generalization problem of humanoid robots in complex scenarios of harness operation, and provide a large-scale model training method for humanoid robot harness operation that can generate more accurate results.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] The present invention provides a method for training a large model of humanoid robot harness manipulation based on a meta-action space, the method comprising:

[0007] Constructing harness manipulation meta-actions based on human harness manipulation;

[0008] Acquire harness manipulation motions of a humanoid robot, and construct a harness manipulation meta-motion dataset based on the harness manipulation meta-motions;

[0009] Based on the harness manipulation meta-action dataset, using reinforcement learning to train a large model of humanoid robot harness manipulation, and defining a reward function of the reinforcement learning based on the output of the large model of humanoid robot harness manipulation;

[0010] Obtaining input data, using a trained humanoid robot harness manipulation model to output joint motor parameters and updating the input data in real time based on the joint motor parameters; the input data includes text instructions, visual observations, and robot perception data;

[0011] A basic strategy is obtained based on the joint motor parameters, and based on the basic strategy, a humanoid robot harness operation large model is trained to obtain a residual strategy in a human behavior intervention lifelong learning manner;

[0012] Humanoid robot harness manipulation is performed based on the basic strategy and residual strategy.

[0013] As a preferred technical solution, the types of the wire harness operation element actions include wire harness routing task operations, wire harness winding task operations, and wire harness detection task operations.

[0014] As a preferred technical solution, the harness routing task operation is to arrange multiple harnesses on the tooling plate as required and fix the harness position by a U-shaped clamp provided thereon, including: first grabbing, routing, first straightening and moving to the next routing position; wherein, the first grabbing is to grab two specific position points of the harness, lift the first preset height at a specified speed, and the harness grabbing point does not deviate during the first grabbing, routing and first straightening actions; the routing is performed on the basis of the first grabbing, and the robot moves the first grabbed harness and arranges it in the U-shaped clamp; the first straightening is performed on the basis of the first grabbing, and the robot arms move toward both ends along the direction of the harness to straighten the harness; the movement to the next routing position is performed after the routing is completed, and the robot arms release the harness and move to the next routing position;

[0015] The wire harness winding task operation is to wind a plurality of wire harnesses with a tape on a pre-wired tool plate, comprising: second grabbing, winding and tearing off the tape; wherein the second grabbing is to grab a position point of the tape by a robot double arm, to lift a second preset height at a preset speed, and the tape grabbing point does not deviate during the whole grabbing action; the winding is performed on the basis of the second grabbing, the robot holds the tape, and winds the plurality of wire harnesses on the tool plate; and the tearing off the tape is performed on the basis of the winding, and the wire harness after the winding is completed is cut off from the wound tape;

[0016] The wire harness detection task operation is to install the assembled wire harness on a detection table, and to connect the connector at the end of the wire harness with the connecting socket of the detection table, comprising: third grabbing and placing, plugging, second straightening and moving to the next plugging position; the third grabbing and placing is to grab two position points of the wire harness, to move to a third preset height above the detection table, to place the wire harness on the detection table and to release the wire harness; the plugging is to grab the end of the wire harness by one arm of the robot, to grab the end connector by the other arm of the robot, to move the two arms to plug the connector into the corresponding hole slot of the detection table; the second straightening is to tightly hold the wire harness by the robot double arms and to move to both ends along the direction of the wire harness to straighten the wire harness; and the moving to the next plugging position is performed after the plugging is completed, the robot double arms release the wire harness, and move to the next plugging position of the detection table.

[0017] As a preferred technical solution, the method for constructing the wire harness operation element action data set comprises:

[0018] Teaching the robot and collecting all robot actions in the teaching process;

[0019] Based on the wire harness operation element action, all robot actions are sorted and classified into different wire harness operation element actions, and a wire harness operation element action data set is obtained by integration.

[0020] As a preferred technical solution, the teaching method is virtual reality teaching.

[0021] As a preferred technical solution, the expression of the reward function is:

[0022]

[0023] Wherein, represents the instantaneous reward calculated according to the current state of the humanoid robot and the joint motor parameters of the executed action at time step t; γ is a discount factor; T is a time step; v t represents visual observation; p t represents perception data; q t represents the angle of the electric motor joint when the action is executed, represents the speed of the electric motor joint when the action is executed, Indicates the acceleration of the motor joint when performing an action.

[0024] As a preferred technical solution, the method further includes constructing a meta-action space based on the harness operation meta-action dataset, the steps of which include:

[0025] Extracting the joint dynamics representation corresponding to each data in the harness operation element action data set, encoding the joint dynamics representation, and obtaining the joint dynamics feature;

[0026] All joint dynamic features are combined to generate the meta-action space.

[0027] As a preferred technical solution, the method for outputting joint motor parameters is:

[0028] Discretizing the text instructions, visual observations, and perception data to obtain discretized data;

[0029] Based on the discretized data, the joint dynamics features are dynamically selected, and the joint motor parameters for generalized execution of the action are generated according to the joint dynamics features obtained by the dynamic selection. The expression is:

[0030] m=f1(v t ,l,p t ,M),

[0031]

[0032] Among them, m represents the joint dynamic characteristics obtained by dynamic selection, q t Indicates the angle of the motor joint when performing the action, Indicates the speed of the motor joint when performing the action, Indicates the acceleration of the motor joint when performing an action.

[0033] As a preferred technical solution, the method for obtaining the basic strategy is: performing strategy migration on the motor parameters and deploying them on the robot to obtain the basic strategy.

[0034] As a preferred technical solution, the method for obtaining the residual strategy is:

[0035] When the humanoid robot performs abnormally in executing the basic strategy, manual interruption and correction are performed;

[0036] Collect interruption data and correction data, and use reinforcement learning to generate residual policies.

[0037] Compared with the prior art, the present invention has the following advantages:

[0038] 1) The present invention provides a humanoid robot harness manipulation training method based on meta-action space. It plans complex harness manipulation tasks into relatively simple meta-action combinations, and collects a full-body dataset of humanoid robot harness manipulation guided by the collected meta-action combinations. This solves the problem of lack of harness manipulation training data and improves the model's understanding and execution of complex harness manipulation tasks.

[0039] 2) This invention uses reinforcement learning to fine-tune the pre-trained large model framework. Based on the large model output and the impact of the robot's execution of the output on the environment, a reinforcement learning reward function is defined based on the large model output, enabling the large model to generate harness actions that are more suitable for the current operating environment. This not only retains the large model's generalization ability and strong knowledge reserve for complex tasks, but also enhances its adaptability and execution efficiency in specific harness operation tasks.

[0040] 3) The present invention also adds lifelong learning of human behavior intervention. By assisting humanoid robots in correcting erroneous actions, the experience knowledge of wire harness operation and human behavior intervention instructions are continuously accumulated in a real wire harness operation environment, and residual strategies are trained, thereby realizing lifelong learning of wire harness operation. Together with the continuous progress of humanoid robots, it is expected to promote the application and promotion of humanoid robots in wire harness production lines. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of the method of the present invention;

[0042] Figure 2 This is a schematic diagram of the humanoid virtual reality teaching process of the present invention;

[0043] Figure 3 A schematic diagram of the training of autonomous harness manipulation of a humanoid robot based on a large model of the present invention;

[0044] Figure 4 Flowchart for obtaining residual strategy for the present invention. DETAILED DESCRIPTION

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0046] Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by a person of ordinary skill in the technical field to which this application belongs. The words "one", "a", "the" and the like used in this application do not indicate a limit on quantity and may indicate the singular or plural. The terms "include", "comprise", "have" and any variations thereof used in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units that are inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The word "multiple" used in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0047] This embodiment provides a large-scale model training method for humanoid robot harness operation based on meta-action space, including data collection, model training, sim-to-real migration and lifelong learning of humanoid robot harness operation, providing an effective solution for humanoid robots to perform complex harness operation skills, and is expected to promote the application and promotion of humanoid robots in harness production lines.

[0048] Specifically, the process of this method is as follows Figure 1 As shown, the following steps are included:

[0049] S1. Construct harness operation meta-actions based on human harness operation.

[0050] In detail, human harness operations include:

[0051] a. Wire harness routing task operation: This operation involves the robot's dual arms placing multiple wire harnesses on the tooling plate as required and securing the harnesses with U-shaped clips installed on the plate. The meta-actions involved include:

[0052] First grab (Pick_up): This action is to grab two specific position points of the harness and raise them to the first preset height at a specified speed. The harness grab points do not deviate during the first grab, routing and first straightening actions.

[0053] Routing: This action is performed based on the first grasping. The robot moves the first grasped harness and arranges it in a U-shaped clamp at a specific position.

[0054] Straighten: This action builds upon the initial grasping phase. The humanoid robot's arms move along the wire harness, straightening it. This meta-skill is useful for overcoming routing challenges caused by the soft and easily deformable nature of the wire harness.

[0055] Move to the next wiring position (Go_next): This action is performed after the wiring is completed. The robot arms release the wiring harness and move to the next wiring position.

[0056] b. Wire harness winding task operation: This operation involves grabbing tape and winding multiple wire harnesses on a pre-wired tooling board. The meta-actions involved include:

[0057] Second grab (Pick_up): This action is for the robot's arms to grab the position point of the tape and raise it to the second preset height at a preset speed. The tape grab point does not deviate during the entire grabbing process.

[0058] Wrap: This action is performed based on the second grasping. The robot holds the tape and wraps the multiple wire harnesses on the tooling board.

[0059] Tear off the tape (Tear_off): This action is performed on the basis of winding, and the winding tape is cut off after the winding is completed.

[0060] c. Wire harness inspection task operation: This operation involves installing the assembled wire harness on the inspection table and docking the connection socket equipped with the inspection table with the connector at the end of the wire harness. The meta-actions involved include:

[0061] The third grab and place action is to grab the wire harness at two positions, move it to the third preset height above the test table, place the wire harness on the test table and release it.

[0062] Insertion: This action involves one arm of the robot grabbing the end of the wire harness and the other arm grabbing the end connector, and moving both arms to insert the connector into the corresponding hole slot on the test bench.

[0063] Second, Straighten: This action involves the robot's arms grasping the wire harness and moving it in both directions to straighten it. This meta-skill is useful for overcoming the difficulties associated with wiring caused by the soft and easily deformable nature of the harness.

[0064] Move to the next plug-in position: This action is performed after the plug-in is completed. The robot arms release the wiring harness and move to the next plug-in position on the inspection table.

[0065] S2, acquire the human-shaped robot wire harness operation action, and construct a wire harness operation element action dataset based on the wire harness operation element action.

[0066] S21, teach the robot and collect all robot actions in the teaching process, and the flowchart is as shown in Figure 2

[0067] S211, construct a human-shaped robot teleoperation platform, including a wire harness operation platform, an RGB camera and a VR glasses.

[0068] S212, the teaching personnel synchronously observe the current observation picture of the human-shaped robot through the VR glasses and perform wire harness operation.

[0069] S213, estimate the body pose of the teaching personnel based on the shooting picture of the RGB camera by using a human pose estimation algorithm, and calculate the hand pose by the internal algorithm of the VR glasses.

[0070] S214, according to the body pose and hand pose obtained in step S213, redirect the whole body pose of the teaching personnel to the pose of the human-shaped robot by using inverse kinematics method.

[0071] S215, based on the imitation learning algorithm, make the human-shaped robot synchronously perform wire harness operation with the teaching personnel.

[0072] S216, the VR glasses return the observation picture of the human-shaped robot in real time.

[0073] Repeat steps S212 to S216 to obtain enough robot actions.

[0074] S22, based on the wire harness operation element action, classify all robot actions into different wire harness operation element actions, and integrate to obtain a wire harness operation element action dataset.

[0075] S3, based on the wire harness operation element action dataset, train a human-shaped robot wire harness operation large model by using reinforcement learning, and the flowchart of this step is as shown in Figure 3

[0076] According to the input data including visual observation and proprioceptive data and the output of the large model, define a reward function to enhance its adaptability and execution efficiency in specific wire harness operation tasks, and the expression is:

[0077]

[0078] wherein, represents the instantaneous reward calculated according to the current state of the human-shaped robot and the joint motor parameters of the executed action at time step t; γ is the discount factor; T is the time step; v t represents visual observation; p t represents perception data; q​​t angle of the electric motor joint when performing the action, velocity of the electric motor joint when performing the action, acceleration of the electric motor joint when performing the action.

[0079] S4, generating joint motor parameters.

[0080] S41, performing discrete tokenization processing on the text instruction l, visual observation v t and body perception data p t to obtain discretized data as input of the large model.

[0081] S42, using the humanoid robot wire bundle to operate the Transformer encoder in the large model to extract the joint dynamics representation corresponding to each data in the wire bundle operation primitive action data set, encode the joint dynamics representation to obtain joint dynamics features, and collect all joint dynamics features to generate a primitive action space M.

[0082] S43, the humanoid robot wire bundle operating large model dynamically selects joint dynamics features m e M in the primitive action space M based on the discretized data, and generates joint motor parameters for generalizing execution actions according to the joint dynamics features obtained by dynamic selection, and the expression is:

[0083] m = f1(v t ,l,p t ,M),

[0084]

[0085] wherein m represents the joint dynamics features obtained by dynamic selection, q t angle of the electric motor joint when performing the action, velocity of the electric motor joint when performing the action, acceleration of the electric motor joint when performing the action.

[0086] S44, collecting parameters of data changes caused by changes in the environment after the humanoid robot performs the wire bundle operation, including visual observation and body perception data, for generating joint motor parameters for the next wire bundle operation.

[0087] S5, obtaining a basic strategy and a residual strategy.

[0088] The joint motor parameters obtained in step S42 are transferred to the strategy and deployed on the robot to obtain the basic strategy. The robot then performs wiring harness operations based on the basic strategy. For example, if the text task instruction "perform wiring task" is input, the humanoid robot will autonomously grab the wiring harness and route it on the tooling board along the position of the U-shaped frame based on the tooling board and wiring harness scene observed by the head camera. For the failed operation of placing the wiring harness on the U-shaped clamp, the humanoid robot dynamically selects the wiring skill operation m from the meta-action space M based on the current state and the rich knowledge prior of the large model. route , re-perform the wiring operation.

[0089] Specifically, the process of obtaining the residual strategy is as follows Figure 4 Shown, including:

[0090] S51. Under real-time human monitoring, execute step S5. When an abnormality occurs in the robot's execution of the basic strategy, use remote operation to interrupt the humanoid robot's autonomous execution and perform online correction.

[0091] S52. Collect interruption data and correction data, and use reinforcement learning to generate residual strategies.

[0092] Through step S5, the experience knowledge of harness operation and human behavior intervention instructions are continuously accumulated in the real harness operation environment, the residual strategy is trained, and lifelong learning of harness operation is realized.

[0093] S6. Humanoid robot harness manipulation based on basic strategy and residual strategy.

[0094] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for training a large model of humanoid robot harness manipulation based on meta-action space, characterized in that: The method comprises: Constructing harness manipulation meta-actions based on human harness manipulation; Acquire harness manipulation motions of a humanoid robot, and construct a harness manipulation meta-motion dataset based on the harness manipulation meta-motions; Based on the harness manipulation meta-action dataset, using reinforcement learning to train a large model of humanoid robot harness manipulation, and defining a reward function of the reinforcement learning based on the output of the large model of humanoid robot harness manipulation; Obtaining input data, using a trained humanoid robot harness manipulation model to output joint motor parameters and updating the input data in real time based on the joint motor parameters; the input data includes text instructions, visual observations, and robot perception data; A basic strategy is obtained based on the joint motor parameters, and based on the basic strategy, a humanoid robot harness operation large model is trained to obtain a residual strategy in a human behavior intervention lifelong learning manner; Humanoid robot harness manipulation is performed based on the basic strategy and residual strategy.

2. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 1, characterized in that: The types of wire harness operation element actions include wire harness routing task operation, wire harness winding task operation, and wire harness detection task operation.

3. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 2, characterized in that: The wiring harness routing task is to arrange multiple wiring harnesses on the tooling plate as required and fix the wiring harness position by a U-shaped clamp provided thereon, including: first grabbing, routing, first straightening and moving to the next routing position; wherein, the first grabbing is to grab two specific position points of the wiring harness, and lift the first preset height at a specified speed, and the wiring harness grabbing point does not deviate during the first grabbing, routing and first straightening actions; the routing is performed on the basis of the first grabbing, and the robot moves the first grabbed wiring harness and arranges it in the U-shaped clamp; the first straightening is performed on the basis of the first grabbing, and the robot arms move toward both ends along the direction of the wiring harness to straighten the wiring harness; the movement to the next routing position is performed after the routing is completed, and the robot arms release the wiring harness and move to the next routing position; The harness winding task is to grab the tape and wind multiple harnesses on the pre-wired tooling board, including: second grabbing, winding and tearing off the tape; wherein, the second grabbing is the position point where the robot arms grab the tape, and raise it to a second preset height at a preset speed, and the tape grabbing point does not deviate during the entire grabbing process; the winding is carried out on the basis of the second grabbing, and the robot holds the tape and winds the multiple harnesses on the tooling board; the tearing off of the tape is carried out on the basis of the winding, and the wrapped harness is cut off with the winding tape; The wiring harness inspection task operation is to install the assembled wiring harness on the inspection table, and dock the connecting socket equipped on the inspection table with the connector at the end of the wiring harness, including: third grabbing and placing, plugging, second straightening and moving to the next plugging position; the third grabbing and placing is to grab the two position points of the wiring harness, move to the third preset height above the inspection table, place the wiring harness on the inspection table and release the wiring harness; the plugging is for one arm of the robot to grab the end of the wiring harness, and the other arm to grab the end connector, and move both arms to insert the connector into the corresponding hole groove of the inspection table; the second straightening is for the robot arms to hold the wiring harness tightly and move toward both ends along the direction of the wiring harness to straighten the harness; the movement to the next plugging position is carried out after completing the plugging, and the robot arms release the wiring harness and move to the next plugging position of the inspection table.

4. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 1, characterized in that: The method for constructing the harness operation meta-action dataset is as follows: Teach the robot and collect all robot movements during the teaching process; Based on the harness operation meta-actions, all the robot actions are sorted and classified into different harness operation meta-actions, and the harness operation meta-action dataset is obtained by integration.

5. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 4, characterized in that: The teaching method is virtual reality teaching.

6. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 1, characterized in that: The expression of the reward function is: in, represents the instantaneous reward calculated based on the current state of the humanoid robot and the parameters of the joint motors performing the action at time step t; γ is the discount factor; T is the time step; v t represents visual observation; p t Represents the perception data; q t Indicates the angle of the motor joint when performing the action, Indicates the speed of the motor joint when performing the action, Indicates the acceleration of the motor joint when performing an action.

7. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 1, characterized in that: The method further includes constructing a meta-action space based on the harness operation meta-action dataset, the steps of which include: Extracting the joint dynamics representation corresponding to each data in the harness operation element action data set, encoding the joint dynamics representation, and obtaining the joint dynamics feature; All joint dynamic features are combined to generate the meta-action space.

8. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 7, characterized in that: The method for outputting joint motor parameters is: Discretizing the text instructions, visual observations, and perception data to obtain discretized data; Based on the discretized data, the joint dynamics features are dynamically selected, and the joint motor parameters for generalized execution of the action are generated according to the joint dynamics features obtained by the dynamic selection. The expression is: m=f1(v t ,l,p t ,M), Among them, m represents the joint dynamic characteristics obtained by dynamic selection, q t Indicates the angle of the motor joint when performing the action, Indicates the speed of the motor joint when performing the action, Indicates the acceleration of the motor joint when performing an action.

9. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 1, characterized in that: The method for obtaining the basic strategy is: performing strategy migration on the motor parameters and deploying them on the robot to obtain the basic strategy.

10. The method for training a large model of humanoid robot harness manipulation based on meta-action space according to claim 1, characterized in that: The method for obtaining the residual strategy is: When the humanoid robot performs abnormally in executing the basic strategy, manual interruption and correction are performed; Collect interruption data and correction data, and use reinforcement learning to generate residual policies.

Citation Information

Patent Citations

  • Method and system for data processing for robot action expression learning

    CN105825268A

  • Mobile robot non-prior map navigation decision-making method based on DDPG

    CN114396949A