Method for Determining Reinforcement Learning Reward of Robot Arm and Storage Medium

By using visual language models to determine the sub-target sequence and particle state in the robotic arm operation task, and isolate the perception error, efficient environmental perception and dynamic adaptation for complex scenarios are achieved, and the problem of insufficient environmental perception accuracy and adaptability in the prior art is solved.

CN119952727BActive Publication Date: 2025-07-22BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510394932.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

In the prior art, the robot operating task feedback signal method based on the visual language model has low environmental perception accuracy, poor adaptability, and strong applicability to tasks with clear goals or single scenarios in high complexity visual information and dynamic and variable scenarios.

Method used

By obtaining the task data of the robotic arm, using the general visual language model to determine the sub-target sequence and hidden state, initialize the particles, determine the reward result based on the particle's state and weight parameters, isolate the perception error of the visual language model, and realize online real-time adaptation to complex scenarios.

Benefits of technology

It improves the accuracy of environmental perception and the adaptability of dynamic and variable scenarios, ensures the accuracy and sensitivity of reward results to the real sub-target completion situation, solves the problem of sparse rewards in traditional methods, and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952727B_ABST
    Figure CN119952727B_ABST
Patent Text Reader

Abstract

The present application provides a method for determining the reinforcement learning reward of a robotic arm and a storage medium. Specifically, the method includes: obtaining the current task data of the robotic arm; determining at least one sub-goal sequence and sub-goal hidden state corresponding to the current task data according to the current task data and a general vision-language model; determining the sub-goal input states of each particle at non-initial time according to the updated sub-goal hidden states of each particle at the previous moment, and determining the sub-goal completion state at non-initial time according to the sub-goal input states of each particle and the weight parameters of each particle at non-initial time; at the current decision-making moment, determining the reward result at the current decision-making moment according to the sub-goal completion state at the current decision-making moment and the sub-goal completion state at the previous decision-making moment. The present application can isolate the perception error of the vision-language model from the policy optimization process and reduce the requirements for the vision-language model in terms of understanding complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of robot control, and more particularly, to a method and storage medium for determining the reinforcement learning reward of a robotic arm. Background Art

[0002] In recent years, research on vision-language models has gradually emerged and found some applications in robot operations. A vision-language model (VLM) can jointly encode visual and language information, thereby playing a role in understanding scene semantics, detecting target states, and inferring operation intentions. On this basis, many studies have attempted to use VLM to provide feedback signals (e.g., reward signals and task completion degrees) for robot operation tasks.

[0003] In the prior art, VLM can be used to provide feedback signals for robot operation tasks by adopting an explicit generation reward function method. Specifically, before the robot executes an operation, the VLM is called from the image and task description to explicitly generate a reward function, and then a small part of expert trajectory and random policy trajectory are used to modify the robot operation task.

[0004] However, this processing method in the prior art has a relatively simple scope of application and has certain requirements for the complexity of visual information. It is usually only applicable to tasks with clear goals or relatively simple scenes, and there are problems such as low environmental perception accuracy and poor adaptability to dynamically changing scenes. Summary of the Invention

[0005] The purpose of the present application is to provide a method, device, equipment, and storage medium for determining the reinforcement learning reward of a robotic arm to solve the problems of low environmental perception accuracy and poor adaptability to dynamically changing scenes in the prior art.

[0006] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:

[0007] In a first aspect, an embodiment of the present application provides a method for determining the reinforcement learning reward of a robotic arm, the method including:

[0008] Obtain the current task data of the robotic arm, where the current task data includes a task image and a task statement corresponding to the task image;

[0009] According to the current task data and a general vision-language model, determine at least one sub-goal sequence corresponding to the current task data, and according to each sub-goal sequence and the vision-language model, determine a sub-goal hidden state, where each sub-goal sequence corresponds to each sub-task in the current task data;

[0010] Initialize multiple particles according to the sub-goal hidden state;

[0011] At the initial moment, obtain the initial values of each particle, and determine the sub-goal completion state at the initial moment according to the initial values of each particle and the weight parameters of each particle at the initial moment;

[0012] At each non-initial moment, determine the sub-goal input state of each particle at the non-initial moment according to the updated sub-goal hidden state of each particle at the previous moment, and determine the sub-goal completion state at the non-initial moment according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment;

[0013] At the current decision moment, determine the reward result at the current decision moment according to the sub-goal completion state at the current decision moment and the sub-goal completion state at the previous decision moment.

[0014] In a second aspect, another embodiment of the present application provides a robotic arm reinforcement learning reward determination device, and the device includes:

[0015] An acquisition module, configured to acquire the current task data of the robotic arm, where the current task data includes a task image and a task statement corresponding to the task image;

[0016] A hidden state determination module, configured to determine at least one sub-goal sequence corresponding to the current task data according to the current task data and a general vision-language model, and determine a sub-goal hidden state according to each sub-goal sequence and the vision-language model, where each sub-goal sequence corresponds to each sub-task in the current task data;

[0017] An initialization module, configured to initialize multiple particles according to the sub-goal hidden state;

[0018] An initial moment determination module, configured to, at the initial moment, obtain the initial values of each particle, and determine the sub-goal completion state at the initial moment according to the initial values of each particle and the weight parameters of each particle at the initial moment;

[0019] A non-initial moment determination module, configured to, at each non-initial moment, determine the sub-goal input state of each particle at the non-initial moment according to the updated sub-goal hidden state of each particle at the previous moment, and determine the sub-goal completion state at the non-initial moment according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment;

[0020] A decision-making moment module, configured to determine a reward result at the current decision-making moment according to the sub-goal completion status at the current decision-making moment and the sub-goal completion status at the previous decision-making moment.

[0021] In a third aspect, another embodiment of the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus. The processor executes the machine-readable instructions to perform the steps of any method described in the first aspect above.

[0022] In a fourth aspect, another embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of any method described in the first aspect above.

[0023] The beneficial effects of the present application are as follows: By obtaining the current task data of the robotic arm; determining at least one sub-goal sequence corresponding to the current task data according to the current task data and a general vision-language model, and determining a sub-goal hidden state according to each sub-goal sequence and the vision-language model, and initializing a plurality of particles according to the sub-goal hidden state; at the initial moment, obtaining the initial values of each particle, and determining the sub-goal completion status at the initial moment according to the initial values of each particle and the weight parameters of each particle at the initial moment; at each non-initial moment, determining the sub-goal input state of each particle at the non-initial moment according to the updated sub-goal hidden state of each particle at the previous moment, and determining the sub-goal completion status at the non-initial moment according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, so that at the current decision-making moment, a reward result at the current decision-making moment can be determined according to the sub-goal completion status at the current decision-making moment and the sub-goal completion status at the previous decision-making moment, so that the reinforcement learning policy of the robotic arm can be updated based on the reward result, and the responsibility of the vision-language model can be limited to the atomic sub-goal generation level, isolating the perception error of the vision-language model from the policy optimization process, thereby reducing the requirements for the vision-language model in the complex scene understanding level.

[0024] In addition, by determining at least one sub-goal sequence corresponding to the current task data, the current task data can be decomposed into sequential sub-goals, and independent hidden state monitoring and reward calculation can be performed for each sub-goal. This can naturally cover the key regions of the high-dimensional state space, making full use of the sequential information, enabling the sub-goal completion status to be gradually corrected in time series, effectively suppressing the interference caused by detection errors and information uncertainty, achieving the online real-time adaptation of the reward result to the dynamic scenario, improving the environmental perception accuracy, and ensuring that the reward result has higher accuracy and sensitivity to the actual sub-goal completion situation. This solves the problem of sparse rewards in long-cycle tasks in traditional methods and can also improve the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 A flowchart showing a method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0027] Figure 2 A flowchart showing a method for determining at least one sub-goal sequence corresponding to the current task data in the method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0028] Figure 3 A flowchart showing a method for determining the hidden state of a sub-goal in the method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0029] Figure 4 A flowchart showing a method for determining the weight parameter in the method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0030] Figure 5 A flowchart showing a method for determining the weight parameter of each particle at the current moment in the method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0031] Figure 6 A flowchart showing a method for determining the sub-goal completion status at non-initial moments in the method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0032] Figure 7 A flowchart showing a method for obtaining the updated sub-goal hidden state of each particle at non-initial moments in the method for determining the reinforcement learning reward of a robotic arm provided by an embodiment of the present application;

[0033] Figure 8 Schematic diagram of a device for determining reinforcement learning rewards of a robotic arm provided by an embodiment of the present application;

[0034] Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.

[0036] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and illustrated in the accompanying drawings herein may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.

[0037] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the presence of the subsequently claimed features, but does not exclude the addition of other features.

[0038] In the prior art, the VLM can be used to provide feedback signals for robot operation tasks by adopting a method based on explicitly generating a reward function. Specifically, before the robot executes an operation, the VLM is called from the image and task description to explicitly generate a reward function, and then a small part of expert trajectories and random policy trajectories are used to modify the robot operation task.

[0039] However, when the VLM explicitly generates the reward function, it relies on single-frame or finite-frame input images and text descriptions. Moreover, mainstream VLMs lack the ability for long-term temporal modeling. If the task involves moving objects, single-frame or finite-frame images cannot reflect the motion trend and capture the continuity of environmental changes. This causes the generated reward function to deviate from the true task objectives, meaning that the VLM's explicit generation of the reward function can only be applied to relatively simple static snapshot scenarios.

[0040] Meanwhile, VLMs are usually trained on public datasets (such as COCO and ImageNet), which mainly focus on discrete object classification and simple scene annotation, lacking complex relationship annotations such as spatial constraints and physical interactions. When there are multiple object occlusions and similar object interferences in the scene, the VLM's ability to understand complex visual relationships is limited, making it difficult to accurately associate task descriptions with visual entities. This requires that in the process of the VLM explicitly generating the reward function, the strong assumptions of the number of objects ≤ 3 and no overlap must be met, and the input images need to meet conditions such as low occlusion rate and high contrast. That is, it is usually only applicable to tasks with clear objectives or relatively simple scenes, and has certain requirements for the complexity of visual information.

[0041] Furthermore, in the process of modifying the robot operation task by relying on a small number of expert trajectories, the number of expert trajectories is usually limited, and the state space of the robot operation task may be extremely high-dimensional. There may be a problem that small samples are difficult to cover the high-dimensional state space, resulting in systematic deviations in the performance of the modified reward function in the uncovered areas and possibly generating incorrect reward signals, meaning that it has poor adaptability to scenarios with low environmental perception accuracy and dynamic changes.

[0042] In addition, in the process of modifying the robot operation task by relying on random policy trajectories, there is a sparse reward problem, that is, in complex tasks, random policies can hardly accidentally trigger key events. At the same time, there is also a risk of incorrect association. Specifically, random policies may achieve short-term "success" in the wrong state by coincidence, resulting in the reward function learning false associations, meaning that it has poor adaptability to scenarios with low environmental perception accuracy and dynamic changes.

[0043] In summary, this processing method in the existing technology has a relatively simple scope of application, has certain requirements for the complexity of visual information, is usually only applicable to tasks with clear objectives or relatively simple scenes, and has problems such as low environmental perception accuracy and poor adaptability to dynamically changing scenarios.

[0044] Based on the above problems, an embodiment of the present application proposes a method for determining the reinforcement learning reward of a robotic arm. The method includes obtaining the current task data of the robotic arm; determining at least one sub-goal sequence corresponding to the current task data according to the current task data and a general vision-language model, and determining the sub-goal hidden state according to each sub-goal sequence and the vision-language model, and initializing a plurality of particles according to the sub-goal hidden state; at the initial moment, obtaining the initial values of each particle, and determining the sub-goal completion state at the initial moment according to the initial values of each particle and the weight parameters of each particle at the initial moment; at each non-initial moment, determining the sub-goal input state of each particle at the non-initial moment according to the updated sub-goal hidden state of each particle at the previous moment, and determining the sub-goal completion state at the non-initial moment according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, so that at the current decision moment, the reward result at the current decision moment can be determined according to the sub-goal completion state at the current decision moment and the sub-goal completion state at the previous decision moment, thereby being able to update the reinforcement learning policy of the robotic arm based on the reward result, being able to limit the responsibility of the vision-language model to the atomic sub-goal generation level, isolating the perception error of the vision-language model from the policy optimization process, thereby reducing the requirements for the vision-language model in the complex scene understanding level, and improving the environmental perception accuracy and the adaptability to dynamic and changeable scenes.

[0045] It can be understood that when training the reinforcement learning policy of the robotic arm, the method for determining the reinforcement learning reward provided by the embodiment of the present application can be executed to determine the reward of the reinforcement learning policy of the robotic arm, and the reinforcement learning policy of the robotic arm can be trained based on the reward, so that after the robot deploys the reinforcement learning policy of the robotic arm, it can intelligently, efficiently and safely complete physical interaction tasks in a complex dynamic environment through autonomous decision-making and continuous optimization.

[0046] Exemplarily, the robotic arm can be the robotic arm of a robot, such as a humanoid robot or other robots; of course, the robotic arm can also be the robotic arm of any device that requires precise operation, repetitive work or work in a dangerous environment, such as: the welding robotic arm on an automobile production line, the assembly robotic arm on an electronic product production line, and the packaging robotic arm in a food processing production line, etc.

[0047] The method for determining the reinforcement learning reward of the robotic arm provided by the embodiment of the present application will be described in detail below in combination with multiple embodiments.

[0048] Figure 1 FIG. is a schematic flowchart of a method for determining the reinforcement learning reward of the robotic arm provided by the embodiment of the present application. Referring to Figure 1 As shown, the execution subject of the method can be any electronic device with processing capabilities. The method includes:

[0049] S101. Obtain the current task data of the robotic arm.

[0050] Among them, the current task data includes a task image and a task statement corresponding to the task image.

[0051] Optionally, the task image is a visual information carrier corresponding to the current task of the robotic arm. For example, the task image can be an initial Red-Green-Blue Image (RGB image) within the field of view corresponding to the robotic arm in the current environment. The task statement corresponding to the task image can be a task description language or a task objective description read from user input.

[0052] Exemplarily, taking a humanoid robot as an example, the task image of the robotic arm of the humanoid robot can be an RGB image obtained through the camera of the humanoid robot. Taking the welding robotic arm on an automobile production line as an example, the task image of the welding robotic arm can be obtained through a visual sensor in the same welding system as the welding robotic arm, or can be obtained through a camera set on the welding robotic arm.

[0053] Optionally, after obtaining the current task data of the robotic arm, the RGB image can also be normalized, and the task statement can be semantically formatted.

[0054] Exemplarily, taking a humanoid robot as an example, the task image can be an RGB image taken by the humanoid robot from the current perspective, and the task statement can be "assemble the gear onto the red shaft and tighten the screw".

[0055] S102. According to the current task data and a general vision-language model, determine at least one sub-goal sequence corresponding to the current task data, and according to each sub-goal sequence and the vision-language model, determine the sub-goal hidden state.

[0056] Optionally, after obtaining the current task data, the current task data can be disassembled into sequential sub-goals through a general vision-language model to obtain at least one sub-goal sequence corresponding to the current task data, and the completion status of each sub-goal sequence can be judged through the general vision-language model to generate a hidden state representation of each sub-goal sequence, obtaining the sub-goal hidden state, thereby solving the problem that a single-frame image cannot model the motion trend and reducing the requirements for the vision-language model in the aspect of complex scene understanding.

[0057] Among them, each sub-goal sequence can be understood as each sub-task in the current task data. The sub-goal sequence is used to indicate the description of the spatial position relationship between the items corresponding to a sub-task in the current task data. Each parameter in the sub-goal hidden state is used to indicate the completion situation of a sub-goal sequence.

[0058] Among them, a general vision-language model can be understood as a vision-language model that does not require additional training or fine-tuning, such as open-source vision-language models, etc.

[0059] Exemplarily, continuing with the above task image as an RGB image and the task statement "assemble the gear onto the red shaft and tighten the screw" as an example, by performing this step, the task can be decomposed into multiple subtasks, including: "locate the red shaft", "grab the gear", "align the center hole of the gear with the shaft", "press until fully fitted", and "pick up the screw and screw it in", and obtain multiple subtarget sequences, including: the description of the spatial position or logical relationship between the items corresponding to the subtask of "locate the red shaft", for example: the red shaft is above the first component in the field of view and below the second component, and the spatial position or logical relationship between other components in the field of view, the description of the spatial position or logical relationship between the items corresponding to the subtask of "grab the gear", for example: the spatial position or logical relationship between the red shaft, the gear, and other components in the field of view, the spatial position or logical relationship between the gear, the shaft, and other components in the field of view under the subtask of "align the center hole of the gear with the shaft", the spatial position or logical relationship between the gear, the shaft, and other components in the field of view under the subtask of "press until fully fitted", and the spatial position or logical relationship between the gear, the screw, and other components in the field of view under the subtask of "pick up the screw and screw it in".

[0060] Exemplarily, the subtarget hidden state can be a fusion vector, and each parameter in the fusion vector is used to indicate the completion status of a subtarget sequence. For example: in the subtarget hidden state, according to the order of each subtarget sequence, it includes multiple parameters. The first parameter in the subtarget hidden state is used to indicate whether the subtarget sequence corresponding to the subtask of "locate the red shaft" is completed, the second parameter in the subtarget hidden state is used to indicate whether the subtarget sequence corresponding to the subtask of "grab the gear" is completed, the third parameter in the subtarget hidden state is used to indicate whether the subtarget sequence corresponding to the subtask of "align the center hole of the gear with the shaft" is completed, the fourth parameter in the subtarget hidden state is used to indicate whether the subtarget sequence corresponding to the subtask of "press until fully fitted" is completed, and the fifth parameter in the subtarget hidden state is used to indicate whether the subtarget sequence corresponding to the subtask of "pick up the screw and screw it in" is completed.

[0061] By using the current task data and a general vision-language model, determining at least one subtarget sequence corresponding to the current task data, and determining the subtarget hidden state according to each subtarget sequence and the vision-language model, it is possible to convert multi-modal perception data into a semantic feature representation that can be understood by a machine, and this semantic feature representation can retain the most relevant task information in the original observation while filtering out irrelevant background interference.

[0062] S103. Initialize multiple particles according to the sub-goal hidden state.

[0063] Optionally, after obtaining the sub-goal hidden state, multiple particles can be initialized according to the sub-goal hidden state to convert the task progress into a probabilistic state hypothesis, so that each particle can represent a possible task state hypothesis, improve the response to the real-time environment, and gradually attenuate the influence of incorrect hypotheses.

[0064] Exemplarily, multiple particles can be initialized according to the number or values of the parameters in the sub-goal hidden state. Among them, a particle can be understood as a point in a high-dimensional state space implemented based on the framework of a Bayesian filter. Initializing particles can include initializing the number, shape, quantity, state, and motion model of the particles.

[0065] S104. At the initial moment, obtain the initial values of each particle, and determine the sub-goal completion state at the initial moment according to the initial values of each particle and the weight parameters of each particle at the initial moment.

[0066] It can be understood that after the initialization of each particle is completed, the sub-goal completion state can be updated based on the Bayesian filter in the high-dimensional state space, that is, the autonomous exploration of each particle is realized.

[0067] Specifically, in each first time period composed of the initial moment and each non-initial moment, autonomous exploration can be cycled to obtain the sub-goal completion state at each moment in each first time period, so that when it is necessary to update the reward of the reinforcement learning policy, the rewards at the current decision moment and the previous decision moment can be determined from all the first time periods, and then the reinforcement learning policy can be updated based on the rewards at the current decision moment and the previous decision moment.

[0068] Optionally, at the initial moment, the initial values of each particle can be obtained, the weight parameters of each particle at the initial moment can be determined, and the initial values of each particle can be updated according to the weight parameters of each particle at the initial moment to obtain the updated sub-goal hidden state of each particle at the initial moment, so as to determine the sub-goal completion state at the initial moment according to the updated sub-goal hidden state of each particle at the initial moment.

[0069] Exemplarily, at the initial moment, the initial values of each particle obtained can all be the sub-goal hidden state h0.

[0070] Exemplarily, at the initial moment, after obtaining the initial values of each particle, the weight parameters of each particle at the initial moment can be calculated through a vision-language model, and the initial values of each particle can be updated through the weight parameters of each particle at the initial moment to obtain the updated sub-goal hidden states of each particle at the initial moment. Then, based on the updated sub-goal hidden states of each particle at the initial moment and the weight parameters of each particle at the initial moment, the sub-goal completion state at the initial moment can be determined, realizing continuous tracking of the sub-goal completion state.

[0071] S105. At each non-initial moment, based on the updated sub-goal hidden states of each particle at the previous moment, determine the sub-goal input states of each particle at the non-initial moment, and based on the sub-goal input states of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, determine the sub-goal completion state at the non-initial moment.

[0072] It can be understood that after obtaining the sub-goal completion states of each particle at the initial moment, each particle can be propagated in the high-dimensional state space based on the Bayesian filter combined with the motion model in time series, so as to achieve autonomous exploration, obtain the sub-goal completion states at each non-initial moment, and realize continuous tracking of the sub-goal completion state.

[0073] Optionally, at each non-initial moment, based on the updated sub-goal hidden states of each particle at the previous moment, determine the sub-goal input states of each particle at the non-initial moment, determine the weight parameters of each particle at the non-initial moment, and based on the sub-goal input states of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, determine the sub-goal completion state at the non-initial moment.

[0074] Exemplarily, at each non-initial moment, based on the updated sub-goal hidden states of each particle at the previous moment, determine the sub-goal input states of each particle at the non-initial moment, calculate the weight parameters of each particle at the non-initial moment through a vision-language model, update the sub-goal input states of each particle at the non-initial moment according to the weight parameters of each particle at the non-initial moment to obtain the updated sub-goal hidden states of each particle at the non-initial moment, and based on the updated sub-goal hidden states of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, determine the sub-goal completion state at the non-initial moment, continuously tracking the sub-goal completion state at the non-initial moment, so as to realize the combination of visual information and temporal information inference, be able to accurately depict the task completion progress over a longer time span, and ensure that the reward results generated based on the sub-goal completion states of each particle can maintain consistency and reliability at each stage.

[0075] S106. At the current decision-making moment, determine the reward result at the current decision-making moment according to the sub-goal completion status at the current decision-making moment and the sub-goal completion status at the previous decision-making moment.

[0076] It can be understood that during the training process of the reinforcement learning policy, the reinforcement learning policy is updated based on the reward result. Specifically, the reward result of the reinforcement learning policy can be updated according to a preset update period, so as to update the reinforcement learning policy through the reward result of the reinforcement learning policy.

[0077] Optionally, when obtaining the sub-goal completion status at each moment, it can be judged whether the current decision-making moment is reached. When the current decision-making moment is reached, the sub-goal completion status corresponding to the current decision-making moment and the sub-goal completion status at the previous decision-making moment are obtained, and the reward result at the current decision-making moment is determined according to the sub-goal completion status at the current decision-making moment and the sub-goal completion status at the previous decision-making moment, so that the reinforcement learning policy of the robotic arm can be updated according to the reward result.

[0078] Optionally, the sub-goal completion status at the current decision-making moment and the sub-goal completion status at the previous decision-making moment can be input into a preset reward function to calculate the reward result at the current decision-making moment, and the reward result at the current moment is used as the reward of the reinforcement learning policy of the robotic arm to realize the update of the reinforcement learning policy of the robotic arm.

[0079] Among them, the current decision-making moment can be the moment when the reward of the reinforcement learning policy needs to be updated, the previous decision-making moment can be the moment when the reward of the reinforcement learning policy was last updated, and the current decision-making moment and the previous decision-making moment can be determined by a preset update period.

[0080] Exemplarily, at the current decision-making moment t, the sub-goal completion status at the current decision-making moment t can be calculated and the sub-goal completion status at the previous decision-making moment t - τ The difference is calculated, and the obtained difference result is input into a preset reward function σ(), and the reward result r at the current decision-making moment is calculated t , which can be specifically implemented with reference to the following formula (1):

[0081]

[0082] Among them, r t is the reward result at the current decision-making moment, σ() is the reward function, is the sub-goal completion status at the previous decision-making moment t - τ, τ is the update period, is the sub-goal completion status at the current decision-making moment t. Exemplarily, the reward function can be an absolute value reward function or a rounding reward function.

[0083] In this embodiment, by obtaining the current task data of the robotic arm; according to the current task data and a general vision-language model, determining at least one sub-goal sequence corresponding to the current task data, and according to each sub-goal sequence and the vision-language model, determining a sub-goal hidden state, and according to the sub-goal hidden state, initializing a plurality of particles; at the initial moment, obtaining the initial values of each particle, and according to the initial values of each particle and the weight parameters of each particle at the initial moment, determining the sub-goal completion state at the initial moment; at each non-initial moment, according to the updated sub-goal hidden state of each particle at the previous moment, determining the sub-goal input state of each particle at the non-initial moment, and according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, determining the sub-goal completion state at the non-initial moment, so that at the current decision-making moment, according to the sub-goal completion state at the current decision-making moment and the sub-goal completion state at the previous decision-making moment, determining the reward result at the current decision-making moment, so that based on the reward result, the reinforcement learning policy of the robotic arm can be updated, the responsibility of the vision-language model can be limited to the atomic sub-goal generation level, isolating the perception error of the vision-language model from the policy optimization process, thereby reducing the requirements for the vision-language model in the complex scene understanding level.

[0084] In addition, by determining at least one sub-goal sequence corresponding to the current task data, the current task data can be decomposed into sequential sub-goals, and independent hidden state monitoring and reward calculation can be performed on each sub-goal, which can naturally cover the key areas of the high-dimensional state space, making full use of the sequential information, making the sub-goal completion state gradually corrected in time series, and effectively suppressing the interference caused by detection errors and information uncertainty, realizing the online real-time adaptation of the reward result to the dynamic scene, improving the environmental perception accuracy, and also ensuring that the reward result has higher accuracy and sensitivity to the real sub-goal completion situation, solving the problem of sparse rewards in long-cycle tasks in traditional methods, and also improving the training efficiency.

[0085] In a possible implementation manner, Figure 2 FIG. is a schematic flowchart of a process for determining at least one sub-goal sequence corresponding to the current task data in the method for determining the reinforcement learning reward of the robotic arm provided in the embodiment of the present application. Referring to Figure 2 shown in, the above S102 determines at least one sub-goal sequence corresponding to the current task data according to the current task data and a general vision-language model, including:

[0086] S201. Input the current task data and a preset first prompt word into a general vision-language model for perception processing to obtain the item information corresponding to the current task data.

[0087] Optionally, the current task data and a preset first prompt can be input into a general visual language model for perception processing of object recognition to obtain object information corresponding to the current task data.

[0088] Among them, the preset first prompt can be used to instruct the visual language model to output object or item information related to the task statement in the task image. The object information includes the object name.

[0089] Exemplarily, continuing with the above task image being an RGB image and the task statement being "assemble the gear onto the red shaft and tighten the screw" as an example, the task image, the task statement of "assemble the gear onto the red shaft and tighten the screw", and the first prompt can be input into a general visual language model for perception processing of object recognition to obtain object information corresponding to the current task data. For example, "red shaft", "gear", "gear center hole", "shaft", and "screw" can be obtained.

[0090] S202: Input the object information, the current task data, and a preset second prompt into a general visual language model for perception processing to obtain at least one sub-goal sequence.

[0091] Optionally, the object information, the current task data, and a preset second prompt are input into a general visual language model for spatial or logical relationship perception processing to obtain at least one sub-goal sequence.

[0092] Among them, the preset second prompt can be used to instruct the visual language model to output a description of the spatial or logical relationship of the object information related to the task statement in the task image.

[0093] Exemplarily, continuing with the above task image being an RGB image and the task statement being "assemble the gear onto the red shaft and tighten the screw" as an example, the task image, the task statement of "assemble the gear onto the red shaft and tighten the screw", the object information including "red shaft", "gear", "gear center hole", "shaft", and "screw", and the second prompt can be input into a general visual language model for spatial or logical relationship perception processing to obtain at least one sub-goal sequence.

[0094] By inputting the current task data and a preset first prompt into a general vision-language model for perception processing to obtain the item information corresponding to the current task data, it is possible to focus on item perception, avoid interference from task statements on the detection accuracy, and input the item information, the current task data, and a preset second prompt into the general vision-language model for perception processing to obtain at least one sub-goal sequence. It can anchor the abstract task to the specific environment through the item information, reduce error accumulation, and improve the environmental perception accuracy. At the same time, the item information can also be reused for generating multiple sub-goal sequences, reducing the number of calls to the vision-language model and improving the training efficiency.

[0095] In addition, by using a general vision-language model, there is no need to perform any parameter-level fine-tuning or training on the general vision-language model, reducing the costs of model deployment and iteration.

[0096] In one possible implementation Figure 3 is a schematic flowchart of a process for determining the sub-goal hidden state in the method for determining the reinforcement learning reward of the robotic arm provided by the embodiments of the present application. Refer to Figure 3 as shown, the above S102 determines the sub-goal hidden state according to each sub-goal sequence and the vision-language model, including:

[0097] S301. Verify each sub-goal sequence according to the vision-language model to obtain the sub-goal completion status corresponding to each sub-goal sequence.

[0098] Optionally, a Visual Question Answering (VQA) can be constructed and the obtained sub-goal sequences are verified through a general vision-language model to obtain the sub-goal completion status corresponding to each sub-goal sequence.

[0099] Among them, the sub-goal completion status includes that the sub-goal sequence is determined by the vision-language model to be completed and that the sub-goal sequence is determined by the vision-language model to be uncompleted. Exemplarily, the sub-goal completion status can be represented by 0 or 1. When the sub-goal sequence is determined by the vision-language model to be completed, the sub-goal completion status can be 1, and when the sub-goal sequence is determined by the vision-language model to be uncompleted, the sub-goal completion status can be 0.

[0100] S302. Determine the sub-goal hidden state according to the sub-goal completion status corresponding to each sub-goal sequence.

[0101] Optionally, after obtaining the sub-goal completion status corresponding to each sub-goal sequence, the sub-goal completion status corresponding to all sub-goal sequences can be summarized, and an N-dimensional binary vector h0 is constructed as the sub-goal hidden state.

[0102] Exemplarily, the sub-goal hidden state h0 can be represented by the following formula (2):

[0103]

[0104] By using a vision-language model to verify each sub-goal sequence, obtaining the sub-goal completion status corresponding to each sub-goal sequence, and determining the sub-goal hidden state according to the sub-goal completion status corresponding to each sub-goal sequence, it is possible to cross-verify the completion status of the sub-goal by combining visual observation and semantic understanding, avoid misjudgment by a single sensor, significantly reduce misoperations, and at the same time filter out perceptual noise while retaining key task information, enabling the robotic arm to have both human understanding ability and machine computing accuracy when performing tasks.

[0105] In a possible implementation manner, the above S301 verifies each sub-goal sequence according to the vision-language model and obtains the sub-goal completion status corresponding to each sub-goal sequence, including:

[0106] Traverse each sub-goal sequence. For the currently traversed sub-goal sequence, input the current sub-goal sequence, the current task data, and a preset third prompt word into the vision-language model for visual question-answering verification to obtain the sub-goal completion status corresponding to the current sub-goal sequence. After the traversal is completed, obtain the sub-goal completion status corresponding to each sub-goal sequence.

[0107] Optionally, each sub-goal sequence can be traversed. For the currently traversed sub-goal sequence, input the current sub-goal sequence, the task image and task statement in the current task data, and a preset third prompt word into a general vision-language model for visual question-answering verification to obtain the sub-goal completion status corresponding to the current sub-goal sequence. After the traversal is completed, obtain the sub-goal completion status corresponding to each sub-goal sequence, which can convert the discrete sub-goal sequences into independently verifiable state units, enabling each sub-goal sequence to be independently evaluated through real-time visual question-answering, avoiding misjudgment of the overall task progress, and moreover, improving the detection accuracy and reducing the load of the overall Central Processing Unit (CPU).

[0108] Among them, the preset third prompt word is used to instruct the robotic arm to judge the spatial relationship between various objects in the input image and obtain a judgment result. Exemplarily, the third prompt word can be "You are responsible for assisting in controlling a highly intelligent robotic arm to judge the spatial relationship between various objects in the image. According to the given current task data and the queried sub-goal sequence, please output 'yes' or 'no'."

[0109] In a possible implementation manner, the sub-goal hidden state is represented by a target vector. The above S103 initializes multiple particles according to the sub-goal hidden state, including:

[0110] Generate a preset number of particles and assign the value of the target vector to each particle.

[0111] Optionally, the sub-goal hidden state can be represented by an N-dimensional binary vector h0, that is, there are N sub-goal sequences, and an N-dimensional binary vector is obtained as the sub-goal hidden state.

[0112] Optionally, a preset particle shape, the number of particles, and the motion model parameters of the particles can be obtained, particles of the number of particles are generated, and the values of the target vector are respectively assigned to each particle.

[0113] Exemplarily, the particle shape N, the number of particles K = 100, the motion model parameters α = 0.7, and β = 0.04 can be obtained, K particles are generated, and all particles (continuous vectors between 0 and 1 with the shape of N) in the hidden state space are set to the initial value h0, and the same initial weight is assigned to them, so that the filter starts from a unified state, avoiding premature convergence to a local optimal solution caused by uneven initial distribution. Especially when there is insufficient information at the initial stage of the task, the global exploration ability is retained, and it can also prevent some particles from monopolizing subsequent resampling due to high initial weights, reducing the risk of computational overflow, and improving the compatibility and robustness in a dynamic environment.

[0114] In a possible implementation manner, Figure 4 This is a schematic flowchart of a process for determining weight parameters in the method for determining the reinforcement learning reward of the robotic arm provided in the embodiments of the present application. Refer to Figure 4 As shown, the weight parameters are determined through the following process:

[0115] S401: Input the sub-goal hidden state and a preset fourth prompt word into the vision-language model to generate the first position equation and the second position equation corresponding to each particle.

[0116] It can be understood that there is a close connection between the sub-goal completion status of each sub-goal sequence in the sub-goal hidden state and the spatial position of the object. However, due to the possible multi-modal distribution, occlusion, and detection errors in the scene, the direct calculation method is often too complex. Therefore, based on the vision-language model, the first position equation corresponding to each particle and the second position equation

[0117] Optionally, the sub-goal hidden state can be combined with a preset fourth prompt word and input into a general vision-language model for perception processing to generate the first position equation and the second position equation corresponding to each particle.

[0118] Among them, the first position equation is used to indicate at least one target position when the sub-goal sequence is completed. The first position equation The input is the bounding boxes of all items at the current moment, and the first position equation The output is the multiple possible position coordinates of the item when the sub-goal sequence is completed, and the second position equation Used to indicate at least one target position when the sub-goal sequence is not completed, and the second position equation The input is the bounding boxes of all items at the current moment, and the second position equation The output is the multiple possible position coordinates of the item when the sub-goal sequence is not completed.

[0119] S402. Obtain the bounding box information of the task image at the current moment.

[0120] Optionally, the bounding box information of the task image at the current moment (abbreviated as bounding box) can be obtained, where the current moment is the initial moment or any non-initial moment.

[0121] Exemplarily, the task image can be processed by the basic vision model SAM2 first to obtain the bounding box I of the task image from the initial moment 0 to the last non-initial moment T 0:T , so as to obtain the bounding box information I of the task image at the current moment t t .

[0122] S403. Determine the weight parameters of each particle at the current moment according to each first position equation, each second position equation, and the bounding box information at the current moment.

[0123] Optionally, after obtaining the first position equation corresponding to each particle and the second position equation corresponding to each particle, the evolution of the sub-goal input state of each particle between adjacent moments can be simulated according to each first position equation, each second position equation, and the bounding box information at the current moment, and the matching degree between the sub-goal state estimation of each particle and the actually observed object detection box can be compared, so as to determine the weight parameters of each particle at the current moment.

[0124] By inputting the sub-goal hidden state and the preset fourth prompt word into the vision-language model, generating the first position equation and the second position equation corresponding to each particle, obtaining the bounding box information of the task image at the current moment, and determining the weight parameters of each particle at the current moment according to each first position equation, each second position equation, and the bounding box information at the current moment, it is possible to associate the abstract sub-goal hidden state with the specific visual bounding box coordinates and the positions predicted by the position equation, eliminate the single-source error through cross-validation, improve the accuracy of real-time environment response, and also enable the computing resources to be concentrated in the key state space, improving the computing efficiency.

[0125] In a possible implementation manner Figure 5A schematic flowchart for determining the weight parameters of each particle at the current moment in the method for determining the reinforcement learning reward of the robotic arm provided in the embodiments of the present application. Refer to Figure 5 As shown, the current moment is any non-initial moment. In step S403, according to each first position equation, each second position equation, and the bounding box information at the current moment, the weight parameters of each particle at the current moment are determined, including:

[0126] S501. Determine the first matching degree information according to the bounding box information at the current moment and the first position equation.

[0127] Optionally, taking a particle as an example, the bounding box information I at the current moment t t can be input into the first position equation corresponding to this particle for calculation to obtain the possible first position set of the object when the sub-goal sequence is completed Specifically, it can be implemented with reference to the following formula (3):

[0128]

[0129] where, for the sake of simplicity of expression, the subscript t is ignored, i is the sub-goal sequence, I t is the bounding box information at the current moment t, is the first position set.

[0130] Optionally, after obtaining the first position set , the minimum average distance strategy can be adopted to compare the sub-goal state estimation of this particle, that is, the first position set and the actually observed object detection box, that is, the bounding box information I of the task image at the current moment t t to obtain the first matching degree information of this particle, that is, the error when the sub-goal state estimation is completed Specifically, it can be implemented with reference to the following formula (4):

[0131]

[0132] where, I t is the bounding box information of the task image at the current moment t, is the first position set the j-th position point in, is the first matching degree information.

[0133] S502. Determine the second matching degree information according to the bounding box information at the current moment and the second position equation.

[0134] Optionally, taking a particle as an example, the bounding box information I at the current moment t t can be input into the second position equation corresponding to this particle Perform an operation to obtain a set of possible second positions of the object when the sub-goal sequence is completed Specifically, it can be implemented with reference to the following formula (5):

[0135]

[0136] Among them, for the sake of simplicity of expression, the subscript t is ignored, i is the sub-goal sequence, and I t is the bounding box information at the current moment t, is the set of second positions.

[0137] Optionally, after obtaining the set of second positions it is possible to adopt the minimum average distance strategy to compare the sub-goal state estimation of the particle, that is, the set of second positions with the actually observed object detection box, that is, the bounding box information I of the task image at the current moment t t to obtain the second matching degree information of the particle, that is, the error when the sub-goal state estimation is not completed Specifically, it can be implemented with reference to the following formula (6):

[0138]

[0139] Among them, I t is the bounding box information of the task image at the current moment t, is the set of first positions is the j-th position point in it, is the second matching degree information.

[0140] S503. Determine the weight parameter of the particle at the current moment according to the sub-goal input state, the first matching degree information, and the second matching degree information of the particle at the current moment.

[0141] Optionally, after obtaining the first matching degree information and the second matching degree information, it is possible to calculate the weight parameter of the particle at the current moment according to the sub-goal input state, the first matching degree information, and the second matching degree information of the particle at the current moment.

[0142] Exemplarily, taking a particle at a non-initial moment as an example, it is possible to calculate the product of the sub-goal input state of the particle at the non-initial moment, the first matching degree information, and the second matching degree information as the weight parameter of the particle at the current moment.

[0143] Exemplarily, it is also possible to calculate the weight parameter ω of the particle at the current moment with reference to the following formula (7) k :

[0144]

[0145] Where N is the number of sub-goal sequences, is the sub-initial completion state of particle k, is the second matching degree information, is the first matching degree information. It can be understood that the higher the matching degree, the smaller the distance error and the greater the obtained weight.

[0146] It should be noted that when the current moment is the initial moment, the weight parameters of each particle at the current moment determined in S403 above can be implemented with reference to S501 - S503 above. Specifically, at the initial moment, the weight parameters of the particles at the initial moment can be calculated according to the initial values of the particles at the initial moment, the first matching degree information, and the second matching degree information, in combination with the above formula (7).

[0147] By using the bounding box information at the current moment and the first position equation, the first matching degree information is determined, and by using the bounding box information at the current moment and the second position equation, the second matching degree information is determined. Thus, according to the sub-goal input state, the first matching degree information, and the second matching degree information of the particles at the current moment, the weight parameters of the particles at the current moment are determined, which can unify the first position equation, the second position equation, and the bounding box information into a probability framework, transcending the traditional single-source dependence, enabling the obtained weight parameters to achieve precise adaptation in a dynamic environment, and being able to maintain high precision and high robustness simultaneously in a complex environment.

[0148] The above is an exemplary description of the process of determining the weight parameters. It can be understood that after obtaining the weight parameters, the sub-goal completion states at each moment can be determined according to the weight parameters. The following takes each non-initial moment as an example to give an exemplary description of the process of determining the sub-goal completion state.

[0149] In a possible implementation manner, Figure 6 is a schematic flowchart of a process for determining the sub-goal completion state at a non-initial moment in the method for determining the reinforcement learning reward of a robotic arm provided in an embodiment of the present application. Referring to Figure 6 shown, S105 above determines the sub-goal completion state at the non-initial moment according to the sub-goal input states of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, including:

[0150] S601. Update the sub-goal input states of each particle at the non-initial moment according to the weight parameters of each particle at the non-initial moment to obtain the updated sub-goal hidden states of each particle at the non-initial moment.

[0151] It can be understood that after obtaining the weight parameters ω of each particle at the non-initial moment kAfter that, to solve the problem of particle degeneracy and ensure that the particle set can effectively represent the posterior distribution, the sub-goal input state of each particle at non-initial times can be updated.

[0152] Optionally, after obtaining the weight parameter ω of each particle at non-initial times k the sub-goal input state of each particle at non-initial times can be updated according to the weight parameter ω of each particle at non-initial times k to obtain the updated sub-goal hidden state of each particle at non-initial times.

[0153] S602. Determine the sub-goal completion state at non-initial times according to the updated sub-goal hidden state of each particle at non-initial times and the weight parameter of each particle at non-initial times.

[0154] Optionally, after obtaining the updated sub-goal hidden state of each particle at non-initial times, the sub-goal completion state at non-initial times can be calculated according to the updated sub-goal hidden state of each particle at non-initial times and the weight parameter of each particle at non-initial times.

[0155] Exemplarily, the sub-goal completion state at non-initial times can be calculated using the following formula (8)

[0156]

[0157] where is the sub-goal completion state at non-initial time t, ω k is the weight parameter of particle k at non-initial times, is the updated sub-goal hidden state of particle k at non-initial time t, and K is the number of particles.

[0158] By updating the sub-goal input state of each particle at non-initial times through the weight parameter of each particle at non-initial times to obtain the updated sub-goal hidden state of each particle at non-initial times, short-term perception noise can be eliminated and long-term trend information can be retained, and the sub-goal completion state at non-initial times can be determined according to the updated sub-goal hidden state of each particle at non-initial times and the weight parameter of each particle at non-initial times, enabling the state evolution to be driven by dynamic weights and realizing the full-process probabilistic collaboration of perception-decision-control, thereby improving the adaptability to dynamic scenarios and the accuracy of state estimation.

[0159] In a possible implementation manner, Figure 7 is a schematic flowchart when obtaining the updated sub-goal hidden state of each particle at non-initial times in the method for determining the reinforcement learning reward of the robotic arm provided in the embodiments of the present application. Refer to Figure 7As shown, in the above S601, according to the weight parameters of each particle at non-initial moments, the sub-goal input states of each particle at non-initial moments are updated to obtain the updated sub-goal hidden states of each particle at non-initial moments, including:

[0160] S701. Determine the sampling value of the current particle at the non-initial moment according to a preset resampling function.

[0161] Optionally, taking the update of the current particle at the current non-initial moment as an example, the particles can be resampled according to a preset resampling function to obtain the sampling value of the current particle at the current non-initial moment, which can be specifically implemented with reference to the following formula (9):

[0162]

[0163] where u is random noise, K is the number of particles, k is the current particle, and the current particle is any one of the particles. u k is the sampling value of the current particle k at the current non-initial moment. Specifically, the sampling value can be a value related to k between 0 and 1.

[0164] S702. Obtain the updated sub-goal hidden states of each particle at the previous moment.

[0165] Optionally, taking the current non-initial moment t as an example, the updated sub-goal hidden states of each particle at the previous moment t - 1 of the current non-initial moment can be obtained.

[0166] S703. Update the sub-goal input state of the current particle at the non-initial moment according to the sampling value of the current particle, the updated sub-goal hidden states of each particle at the previous moment, and the weight parameters of each particle at the non-initial moment to obtain the updated sub-goal hidden state of the current particle at the non-initial moment.

[0167] Optionally, according to the sampling value u of the current particle at the current non-initial moment k the updated sub-goal hidden states of each particle at the previous moment t - 1, and the weight parameters ω of each particle at the current non-initial moment k the sub-goal input state at the current non-initial moment is updated to obtain the updated sub-goal hidden state of the current particle at the current non-initial moment

[0168] Exemplarily, it can be implemented with reference to the following formula (10):

[0169]

[0170] where is the updated sub-goal hidden state of the current particle k at the current non-initial moment, is the updated sub-goal hidden state of the current particle i at the previous moment t - 1, ω j is the weight parameter of the j-th particle at the current non-initial moment, u k is the sampling value of the current particle k at the current non-initial moment.

[0171] By means of a preset resampling function, the sampling value of the current particle at the non-initial moment is determined, and the updated sub-goal hidden states of each particle at the previous moment are obtained. According to the sampling value of the current particle, the updated sub-goal hidden states of each particle at the previous moment, and the weight parameters of each particle at the non-initial moment, the sub-goal input state of the current particle at the non-initial moment is updated to obtain the updated sub-goal hidden state of the current particle at the non-initial moment, which can replace low-weight particles with mutants of high-weight particles, avoid the loss of positioning caused by the degradation of the particle swarm, enhance the anti-degradation ability. At the same time, it can also inherit the historical state intelligently, reduce the loss of effective information. In addition, it can also make the particle state shrink towards the high-probability region, ensuring that the newly generated particles meet both physical feasibility and task semantic correctness.

[0172] It should be noted that the process of determining the sub-goal completion state at the initial moment can be implemented with reference to the process of determining the sub-goal completion state at the non-initial moment in the above S601 - S602 and S701 - S703. Specifically, according to the weight parameters of each particle at the initial moment, the initial values of each particle can be updated to obtain the updated sub-goal hidden states of each particle at the initial moment, and according to the updated sub-goal hidden states of each particle at the initial moment and the weight parameters of each particle at the initial moment, the sub-goal completion state at the initial moment is determined, which is not elaborated in this application.

[0173] Based on the same inventive concept, an apparatus for determining the reinforcement learning reward of a robotic arm corresponding to the method for determining the reinforcement learning reward of a robotic arm is also provided in the embodiments of this application. Since the principle of solving problems by the apparatus in the embodiments of this application is similar to that of the above method for determining the reinforcement learning reward of a robotic arm, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0174] Figure 8 is a schematic diagram of an apparatus for determining the reinforcement learning reward of a robotic arm provided by an embodiment of this application. Refer to Figure 8 As shown, the apparatus includes: an acquisition module 801, a hidden state determination module 802, an initialization module 803, an initial moment determination module 804, a non-initial moment determination module 805, and a decision moment module 806;

[0175] The acquisition module 801 is configured to acquire the current task data of the robotic arm, where the current task data includes a task image and a task statement corresponding to the task image;

[0176] The hidden state determination module 802 is configured to determine at least one sub-goal sequence corresponding to the current task data according to the current task data and a general vision-language model, and determine the sub-goal hidden state according to each sub-goal sequence and the vision-language model;

[0177] The initialization module 803 is configured to initialize a plurality of particles according to the sub-goal hidden state;

[0178] The initial moment determination module 804 is configured to, at the initial moment, obtain the initial values of the particles, and determine the sub-goal completion state at the initial moment according to the initial values of the particles and the weight parameters of the particles at the initial moment;

[0179] The non-initial moment determination module 805 is configured to, at each non-initial moment, determine the sub-goal input state of each particle according to the updated sub-goal hidden state of each particle at the previous moment, and determine the sub-goal completion state at the non-initial moment according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of the particles at the non-initial moment;

[0180] The decision moment module 806 is configured to, at the current decision moment, determine the reward result at the current decision moment according to the sub-goal completion state at the current decision moment and the sub-goal completion state at the previous decision moment.

[0181] Optionally, the hidden state determination module 802 is specifically configured to:

[0182] Input the current task data and a preset first prompt word into the general vision-language model for perception processing to obtain the item information corresponding to the current task data, where the item information includes the item name;

[0183] Input the item information, the current task data, and a preset second prompt word into the general vision-language model for perception processing to obtain at least one sub-goal sequence.

[0184] Optionally, the hidden state determination module 802 is specifically configured to:

[0185] Verify each sub-goal sequence according to the vision-language model to obtain the sub-goal completion status corresponding to each sub-goal sequence;

[0186] Determine the sub-goal hidden state according to the sub-goal completion status corresponding to each sub-goal sequence.

[0187] Optionally, the hidden state determination module 802 is specifically configured to:

[0188] Traverse each sub-goal sequence. For the currently traversed sub-goal sequence, input the current sub-goal sequence, the current task data, and a preset third prompt word into the vision-language model for vision-question answering verification to obtain the completion status of the sub-goal corresponding to the current sub-goal sequence. After the traversal ends, obtain the completion status of the sub-goals corresponding to each sub-goal sequence.

[0189] Optionally, the sub-goal hidden state is represented by a target vector; the initialization module 803 is specifically used for:

[0190] Generate a preset number of particles;

[0191] Assign the value of the target vector to each particle.

[0192] Optionally, it further includes: a weight parameter determination module, which is used for:

[0193] Input the sub-goal hidden state and a preset fourth prompt word into the vision-language model to generate a first position equation and a second position equation corresponding to each particle. The first position equation is used to indicate at least one target position when the sub-goal sequence is completed, and the second position equation is used to indicate at least one target position when the sub-goal sequence is not completed;

[0194] Obtain the bounding box information of the task image at the current moment, where the current moment is the initial moment or any non-initial moment;

[0195] Determine the weight parameter of each particle at the current moment according to each first position equation, each second position equation, and the bounding box information at the current moment.

[0196] Optionally, when the current moment is any non-initial moment, the weight parameter determination module is specifically used for:

[0197] Determine the first matching degree information according to the bounding box information at the current moment and the first position equation;

[0198] Determine the second matching degree information according to the bounding box information at the current moment and the second position equation;

[0199] Determine the weight parameter of the particle at the current moment according to the sub-goal input state of the particle at the current moment, the first matching degree information, and the second matching degree information.

[0200] Optionally, the non-initial moment determination module 805 is specifically used for:

[0201] Update the sub-goal input state of each particle at the non-initial moment according to the weight parameter of each particle at the non-initial moment to obtain the updated sub-goal hidden state of each particle at the non-initial moment;

[0202] Determine the sub-goal completion status at a non-initial moment according to the updated sub-goal hidden states of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment.

[0203] Optionally, the non-initial moment determination module 805 is specifically configured to:

[0204] Determine the sampling value of the current particle at the non-initial moment according to a preset resampling function, where the current particle is any one of the particles;

[0205] Obtain the updated sub-goal hidden states of each particle at the previous moment;

[0206] Update the sub-goal input state of the current particle at the non-initial moment according to the sampling value of the current particle, the updated sub-goal hidden states of each particle at the previous moment, and the weight parameters of each particle at the non-initial moment, so as to obtain the updated sub-goal hidden state of the current particle at the non-initial moment.

[0207] The description of the processing flow of each module in the device and the interaction flow between the modules can refer to the relevant descriptions in the above method embodiments, and will not be elaborated here.

[0208] The embodiment of the present application also provides an electronic device, as Figure 9 shown, Figure 9 is a schematic structural diagram of the electronic device provided by the embodiment of the present application, including: a processor 901, a memory 902, and optionally, a bus 903 may also be included. The memory 902 stores machine-readable instructions executable by the processor 901 (for example, Figure 8 the execution instructions corresponding to the acquisition module 801, the hidden state determination module 802, the initialization module 803, the initial moment determination module 804, the non-initial moment determination module 805, and the decision moment module 806 in the device in

[0209] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above method for determining the reinforcement learning reward of the robotic arm are executed.

[0210] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the method embodiments, and will not be elaborated herein. In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.

[0211] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0212] The above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application.

Claims

1. A method for determining the reinforcement learning reward of a robotic arm, characterized in that, Including: Obtain the current task data of the robotic arm, where the current task data includes a task image and a task statement corresponding to the task image; According to the current task data and a general vision-language model, determine at least one sub-goal sequence corresponding to the current task data, and according to each sub-goal sequence and the vision-language model, determine the sub-goal hidden state; Initialize multiple particles according to the sub-goal hidden state; At the initial moment, obtain the initial values of each particle, and according to the initial values of each particle and the weight parameters of each particle at the initial moment, determine the sub-goal completion state at the initial moment; At each non-initial moment, according to the updated sub-goal hidden state of each particle at the previous moment, determine the sub-goal input state of each particle at the non-initial moment, and according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment, determine the sub-goal completion state at the non-initial moment; At the current decision moment, according to the sub-goal completion state at the current decision moment and the sub-goal completion state at the previous decision moment, determine the reward result at the current decision moment.

2. The method for determining the reinforcement learning reward of the robotic arm according to claim 1, wherein The step of determining at least one sub-goal sequence corresponding to the current task data according to the current task data and a general vision-language model includes: Input the current task data and a preset first prompt word into the general vision-language model for perception processing to obtain the item information corresponding to the current task data, where the item information includes the item name; Input the item information, the current task data and a preset second prompt word into the general vision-language model for perception processing to obtain the at least one sub-goal sequence.

3. The method for determining the reinforcement learning reward of the robotic arm according to claim 1, wherein The step of determining the sub-goal hidden state according to each sub-goal sequence and the vision-language model includes: Verify each sub-goal sequence according to the vision-language model to obtain the sub-goal completion status corresponding to each sub-goal sequence; Determine the sub-goal hidden state according to the sub-goal completion status corresponding to each sub-goal sequence.

4. The method for determining the reinforcement learning reward of the robotic arm according to claim 3, characterized in that The step of verifying each sub-goal sequence according to the vision-language model to obtain the sub-goal completion status corresponding to each sub-goal sequence includes: Traverse each sub-goal sequence. For the currently traversed sub-goal sequence, input the current sub-goal sequence, the current task data and a preset third prompt word into the vision-language model for visual question-answering verification to obtain the sub-goal completion status corresponding to the current sub-goal sequence, and after the traversal ends, obtain the sub-goal completion status corresponding to each sub-goal sequence.

5. The method for determining the reinforcement learning reward of the robotic arm according to claim 1, wherein The sub-goal hidden state is represented by a target vector; The step of initializing multiple particles according to the sub-goal hidden state includes: Generate a preset number of particles; Assign the value of the target vector to each particle.

6. The method for determining the reinforcement learning reward of the robotic arm according to claim 1, wherein The weight parameters are determined through the following process: Input the sub-goal hidden state and a preset fourth prompt word into the vision-language model to generate a first position equation and a second position equation corresponding to each particle. The first position equation is used to indicate at least one target position when the sub-goal sequence is completed, and the second position equation is used to indicate at least one target position when the sub-goal sequence is not completed; Obtain the bounding box information of the task image at the current moment, where the current moment is the initial moment or any non-initial moment; Determine the weight parameters of each particle at the current moment according to each first position equation, each second position equation, and the bounding box information at the current moment.

7. The method for determining the reinforcement learning reward of the robotic arm according to claim 6, wherein When the current moment is any non-initial moment, the determining the weight parameters of each particle at the current moment according to each first position equation, each second position equation, and the bounding box information at the current moment includes: Determine the first matching degree information according to the bounding box information at the current moment and the first position equation; Determine the second matching degree information according to the bounding box information at the current moment and the second position equation; Determine the weight parameters of the particle at the current moment according to the sub-goal input state of the particle at the current moment, the first matching degree information, and the second matching degree information.

8. The method for determining the reinforcement learning reward of the robotic arm according to claim 1, wherein The determining the sub-goal completion state at the non-initial moment according to the sub-goal input state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment includes: Update the sub-goal input state of each particle at the non-initial moment according to the weight parameters of each particle at the non-initial moment to obtain the updated sub-goal hidden state of each particle at the non-initial moment; Determine the sub-goal completion state at the non-initial moment according to the updated sub-goal hidden state of each particle at the non-initial moment and the weight parameters of each particle at the non-initial moment.

9. The method for determining the reinforcement learning reward of the robotic arm according to claim 8, wherein The updating the sub-goal input state of each particle at the non-initial moment according to the weight parameters of each particle at the non-initial moment to obtain the updated sub-goal hidden state of each particle at the non-initial moment includes: Determine the sampling value of the current particle at the non-initial moment according to a preset resampling function, where the current particle is any one of each particle; Obtain the updated sub-goal hidden state of each particle at the previous moment; Update the sub-goal input state of the current particle at the non-initial moment according to the sampling value of the current particle, the updated sub-goal hidden state of each particle at the previous moment, and the weight parameters of each particle at the non-initial moment to obtain the updated sub-goal hidden state of the current particle at the non-initial moment.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of the robotic arm reinforcement learning reward determination method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Exhibition hall robot visual language navigation method based on large model

    CN119309580A

  • Training reinforcement machine learning systems

    US20210334696A1