Diffusion strategy model optimization method combined with state memory mechanism and embodied agent
By introducing a state memory mechanism into the DP model, generating a global condition vector and combining it with a state estimator, the problem of inaccurate decision-making in complex scenarios of the DP model is solved, and the success rate of action execution is greatly improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2026-04-10
AI Technical Summary
Existing dynamic programming (DP) models lack long-term state memory and cannot make accurate decisions in complex, multi-step scenarios, especially when visual differences are subtle and their success rate drops significantly.
A state memory mechanism is introduced to generate a global condition vector by acquiring the state information and image information of the target object. Combined with a diffusion network model and a state estimator, the action sequence is predicted and the state is predicted, thereby improving the accuracy of action execution.
It significantly improves the success rate of action execution in complex, multi-step scenarios. For example, in the experiment of sorting building blocks of different colors, the success rate increased from 15% to 90%, and in the experiment of wiping a table with a scouring pad, the success rate increased from 0 to over 70%.
Smart Images

Figure CN120635985B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a diffusion policy model optimization method combining state memory mechanism and embodied intelligent agent. BACKGROUND
[0002] In recent years, diffusion policy (DP) models based on imitation learning have gained wide attention in embodied intelligent tasks. Such models can predict future multi-frame continuous action sequences through a history of two-frame RGB image encoder (such as a 224x224 pixel Vision Transformer, ViT-224) and a diffusion network, and show good robustness in simple grasping, carrying and other tasks.
[0003] However, the related art still has the following defects: lack of long-term state memory, the existing DP model is only based on very short time sequence visual information, and cannot capture the "completed", "failed", "taken out" and other states in the task process. Sequential dependence loss, in tasks that need to be completed strictly according to steps, once the visual information cannot distinguish the current stage, the model cannot make correct decisions. These defects lead to a significant decrease in success rate in complex multi-step scenarios, especially in scenarios with weak visual differences. SUMMARY
[0004] Therefore, the present application provides a diffusion policy model optimization method combining state memory mechanism and embodied intelligent agent.
[0005] In a first aspect, the embodiment of the present application provides a diffusion policy model optimization method combining state memory mechanism, comprising: obtaining state information and image information of a target object at a current time; generating a global condition vector based on the state information and the image information; inputting the global condition vector into a diffusion network model to obtain a predicted action sequence composed of action data of multiple time steps; and using a state estimator to perform state prediction based on the global condition vector and the action data of the first time step in the multiple time steps to obtain a state prediction result of the next time.
[0006] In a second aspect, the embodiment of the present application provides an embodied intelligent agent, comprising: a robot body and an actuator, wherein the actuator executes the diffusion policy model optimization method combining state memory mechanism as described in any of the implementations of the first aspect, so as to enable the robot body to complete a task.
[0007] The diffusion strategy model optimization method provided by the embodiment of the present application and the embodied intelligent agent with the state memory mechanism can introduce state information and image information in the constructed global condition vector, and can be used to represent the state features and visual features of the target object, and apply the state features and visual features in the process of generating the action sequence by the diffusion network model.
[0008] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0010] Figure 1 is an exemplary system architecture to which the present application can be applied;
[0011] Figure 2 The flowchart of the diffusion strategy model optimization method provided by the embodiment of the present application combined with the state memory mechanism;
[0012] Figure 3 The flowchart of another diffusion strategy model optimization method provided by the embodiment of the present application combined with the state memory mechanism;
[0013] Figure 4 The flowchart of another diffusion strategy model optimization method provided by the embodiment of the present application combined with the state memory mechanism;
[0014] Figure 5 The flowchart of another diffusion strategy model optimization method provided by the embodiment of the present application combined with the state memory mechanism;
[0015] Figure 6 The flowchart of another diffusion strategy model optimization method provided by the embodiment of the present application combined with the state memory mechanism;
[0016] Figure 7A and Figure 7BA flowchart of a diffusion strategy model optimization method provided by an embodiment of the present application in an application scenario is shown in FIG. 1.
[0017] Figure 7C A flowchart of a diffusion strategy model optimization method provided by another embodiment of the present application in an application scenario is shown in FIG. 2.
[0018] Figure 8 A structural block diagram of a diffusion strategy model optimization device provided by an embodiment of the present application is shown in FIG. 3.
[0019] Figure 9 A structural diagram of an electronic device suitable for executing the diffusion strategy model optimization method provided by an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION
[0020] The technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present application.
[0021] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0022] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements, or it can be wireless connection, or it can be wired connection. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0023] In addition, the technical features involved in different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0024] Figure 1An exemplary system architecture 100 of embodiments of the diffusion strategy model optimization method and embodied agent to which the binding state memory mechanism of the present application can be applied is shown.
[0025] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0026] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 and the server 105 can be installed with various applications for realizing information communication between the two, such as instant messaging applications, etc.
[0027] The terminal devices 101, 102, 103 and the server 105 can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-mentioned electronic devices, and can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 105 is software, it can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here.
[0028] The server 105 can provide various services through various built-in applications. The data or information required to provide the services can be obtained from the terminal devices 101, 102, 103 through the network 104, or can be pre-stored locally in the server 105 in various ways. Therefore, when the server 105 detects that the data has been stored locally, it can choose to obtain the data directly from the local, in which case the exemplary system architecture 100 can not include the terminal devices 101, 102, 103 and the network 104.
[0029] Since data or information processing can require more computing resources and stronger computing capacity, the diffusion strategy model optimization method provided by the subsequent embodiments of the present application generally is executed by a server 105 with stronger computing capacity and more computing resources, and accordingly, the diffusion strategy model optimization device is generally also arranged in the server 105. However, it should be noted that when the terminal devices 101, 102 and 103 also have computing capacity and computing resources that meet the requirements, the terminal devices 101, 102 and 103 can also complete the above-mentioned operations by the related application installed thereon, and then output the same results as the server 105. Especially in the case of multiple terminal devices with different computing capacities, but the related application judges that the terminal device has strong computing capacity and more remaining computing resources, the terminal device can be allowed to execute the above-mentioned operations, so as to appropriately reduce the computing pressure of the server 105, and accordingly, the diffusion strategy model optimization device can also be arranged in the terminal devices 101, 102 and 103. In this case, the exemplary system architecture 100 can also not include the server 105 and the network 104.
[0030] It should be understood that Figure 1 The number of terminal devices, networks and servers in the system architecture 100 is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.
[0031] It should be understood that Figure 2 , Figure 2 A flowchart of a diffusion strategy model optimization method provided by an embodiment of the present application is shown in FIG. 2, wherein the flowchart 200 includes the following steps:
[0032] Step 201: Obtain the state information and image information of a target object at a current time.
[0033] This step is to obtain the state information and image information of the target object at the current time by the execution subject of the diffusion strategy model optimization method (for example, the server 105 shown in FIG. 1). The target object can be a device or apparatus that can be controlled by the execution subject to perform corresponding actions or operations (for example, a robot, a mechanical arm, etc.). The state information of the target object can be some states of the target object or components thereon, such as action angle, action type, size, position information, etc., and the image information can be an image of the target object collected by an image collection device (for example, a camera, etc.). Figure 1
[0034] Exemplarily, in the application scenario where the target object is a robot, an end effector (such as a flexible gripper, a mechanical gripper, a suction gripper, etc.) is arranged at the end of the arm of the robot, and a visual / physical marker (Tag) and an inertial measurement unit (IMU) are arranged on the gripper. The Tag can be a two-dimensional code or a bar code, which is used to mark the positions of the two grippers, and according to the positions, the position information of the arm of the robot and the width information, size information, etc. of the two grippers can be obtained; the IMU can be arranged at the end of the arm of the robot, which is used to detect the swing angle and other information of the arm of the robot, so as to obtain the action angle, posture or trajectory of the arm of the robot. The image information can be obtained by using a GoPro camera to collect the image of the arm of the robot. The state information of the robot includes the width information, position information and angle information in this scene, which can be used as the data source of the continuous quantity parameter.
[0035] In the application scenario where the target object is a robot, an end effector (such as a flexible gripper, a mechanical gripper, a suction gripper, etc.) is arranged at the end of the arm of the robot, and a visual / physical marker (Tag) and an inertial measurement unit (IMU) are arranged on the gripper. The Tag can be a two-dimensional code or a bar code, which is used to mark the positions of the two grippers, and according to the positions, the position information of the arm of the robot and the width information, size information, etc. of the two grippers can be obtained; the IMU can be arranged at the end of the arm of the robot, which is used to detect the swing angle and other information of the arm of the robot, so as to obtain the action angle, posture or trajectory of the arm of the robot. The image information can be obtained by using a GoPro camera to collect the image of the arm of the robot. The state information of the robot includes the width information, position information and angle information in this scene, which can be used as the data source of the continuous quantity parameter.
[0036] Step 202: generating a global condition vector based on the state information and the image information.
[0037] This step aims to generate a global condition vector based on the obtained state information and image information by the above-mentioned execution subject. The global condition vector contains vectors corresponding to the state information and image information, which are used to represent the state features and image features, etc. of the target object at the current time.
[0038] Step 203: inputting the global condition vector into a diffusion network model to obtain a predicted action sequence composed of action data of multiple time steps.
[0039] The step is designed to output, by the execution subject, a predicted action sequence composed of action data of multiple time steps based on the global condition vector through a diffusion network model. The diffusion network model is a diffusion policy (DP) based on a diffusion model, which models the visual motion policy as a conditional denoising diffusion process to realize the generation and control of complex action sequences. In this embodiment, the input is the global condition vector, and the diffusion network model can output a predicted action sequence composed of action data of multiple time steps.
[0040] Step 204: Based on the global condition vector and the action data of the first time step in the multiple time steps, state prediction is performed using a state estimator to obtain a state prediction result of the next moment.
[0041] The step is designed to output, by the execution subject, a predicted action sequence composed of action data of multiple time steps based on the global condition vector through a diffusion network model. The diffusion network model is a diffusion policy (DP) based on a diffusion model, which models the visual motion policy as a conditional denoising diffusion process to realize the generation and control of complex action sequences. In this embodiment, the input is the global condition vector, and the diffusion network model can output a predicted action sequence composed of action data of multiple time steps.
[0042] Further, the obtained state prediction result of the next moment can also participate in generating the global condition vector as the current state of the next moment and participate in the subsequent processing process.
[0043] The diffusion strategy model optimization method combining the state memory mechanism provided in the embodiments of the present application introduces state information and image information in the constructed global condition vector, which can be used to represent the state features and visual features of the target object, and applies the state features and visual features in the process of generating action sequences by the diffusion network model. In addition, based on the current state features, visual features and action sequences, the state estimator is used to predict the hidden state of the next moment, so that the generated action instructions can be more suitable for the actions and states of the target object, more accurate control is realized, and the success rate of action execution is improved.
[0044] Please refer to Figure 3 , Figure 3 The flowchart of the diffusion strategy model optimization method combining the state memory mechanism provided in the embodiments of the present application is shown in the flowchart 300, which is a specific implementation of step 201 in the flowchart 200. Figure 2 The flowchart 300 includes the following steps:
[0045] Step 301: Obtain continuous quantity parameters and discrete quantity parameters of the target object, wherein the continuous quantity parameters are used to represent width information, position information and angle information of the target object, and the discrete quantity parameters are used to represent state information or task stage information of the target object.
[0046] This step aims to obtain the continuous quantity parameters and the discrete quantity parameters of the target object by the above-mentioned execution body. The continuous quantity parameters are used to represent the relevant information (such as width, position, angle, etc.) of the target object. For example, the continuous quantity parameters can be obtained by a visual / physical tag and an inertial measurement unit (IMU) arranged on the target object or one or more components of the target object. For example, the “angle” in the above-mentioned relevant information can include initial angle information and / or current angle information. The initial angle information is the initial pose of the end effector at the beginning of the task. The current angle information is the target angle in the subsequent process of the task execution. The “position” in the above-mentioned relevant information can refer to the three-dimensional space position of the end effector at the current time step, which can correspond to x, y and z coordinates in the Cartesian coordinate system. In robot motion planning, the model needs to determine whether the robot arm has reached the target position according to the position information, such as whether it has reached the “grabbing point”, the “placing point” and the like. In state prediction or reinforcement learning, the position information can be used as part of the input features to predict the next action, such as predicting the movement of the robot arm, the opening and closing operation of the gripper and the like.
[0047] The discrete quantity parameters are used to represent the state or task stage of the target object (for example, information used to represent the states or stages of “successful grabbing”, “the object has been placed in the box”, “the wiping has been completed”, “the button has been pressed” and the like). In this embodiment, the discrete quantity parameters can be a vector represented by a state machine and / or a One-Hot.
[0048] Step 302: Splice the continuous quantity parameters and the discrete quantity parameters to obtain a state vector.
[0049] This step aims to splice the continuous quantity parameters and the discrete quantity parameters by the above-mentioned execution body to obtain a state vector.
[0050] Through the above process, in the embodiment of the present application, the basic information and the state information of the target object are perceived and obtained to generate a global condition vector, which can provide a more comprehensive and perfect reference vector for subsequent action prediction and state prediction, so as to improve the accuracy and success rate of the subsequent processing process.
[0051] Step 303: Collect an image of the target object.
[0052] The step is to acquire the image of the target object collected by the visual sensor (such as a camera, a GoPro camera, etc.) by the execution subject.
[0053] Step 304: input the image into a visual encoder to obtain low-dimensional image features.
[0054] The step is to extract the image features of the image by the visual encoder and perform dimension reduction processing to obtain low-dimensional image features by the execution subject. The visual encoder is used to convert high-dimensional data such as images or videos into low-dimensional semantic representations to provide an understandable embedding space for downstream tasks such as classification, detection, and generation. For example, the resolution of the image collected by the visual sensor is 224x224, and the visual encoder can extract the image features of the image and perform dimension reduction processing to obtain image features with a dimension of 1536.
[0055] The step is to extract the image features by the visual encoder and perform dimension reduction processing to obtain low-dimensional image features, which can be used to generate a global condition vector, and can provide a more comprehensive and perfect reference vector for subsequent action prediction and state prediction, to improve the accuracy and success rate of the subsequent processing process.
[0056] Further, in some optional embodiments of the present embodiment, the step 202 of generating a global condition vector based on state information and image information mainly includes: splicing the state vector and the low-dimensional image features to obtain the global condition vector. Alternatively, the continuous parameter and the low-dimensional image features are spliced to obtain spliced features, the discrete parameter is parameter segmented to form scaling parameters and offset parameters, and the spliced features, the scaling parameters and the offset parameters are linearly modulated to obtain the global condition vector.
[0057] It should be noted that the step numbers involved in the above process description are not used to limit the implementation order between steps. The acquisition order of the state information and the image information can be adjusted according to actual needs.
[0058] In this embodiment, the related information of the target object is represented from different dimensions by combining the continuous parameter and the discrete parameter, and the cumulative error is reduced by the cooperation of the continuous action generation and the discrete state estimation, so that more accurate control is realized and the control success rate is improved.
[0059] In some optional embodiments of the embodiments of the present application, the step 204 of inputting the global condition vector and the action data of the first time step in the plurality of time steps into the state estimator to perform state prediction mainly includes: performing gradient separation operation on the global condition vector to cut off the gradient propagation path; performing gradient separation operation on the extraction process of the action data of the first time step; and inputting the global condition vector after gradient separation and the action data of the first time step after gradient separation into the state estimator to perform state prediction, wherein the global condition vector is obtained by splicing the state vector and the low-dimensional image feature.
[0060] This step aims to perform gradient separation operation (detach) on the global condition vector and the action data by the above-mentioned execution body before inputting the global condition vector and the action data into the state estimator, to obtain the global condition vector after gradient separation and the action data after gradient separation. In this embodiment, the global condition vector and the action data are separated from the calculation graph of the diffusion network model by gradient separation (detach) operation, and then input into the state estimator to perform state prediction to obtain the state prediction result of the next moment.
[0061] In the embodiments of the present application, through the above process, the global condition vector and the action data input into the state estimator are separated from the calculation graph of the diffusion network model by gradient separation operation, so that the generated gradient will not propagate to the diffusion network model during subsequent training of the state estimator, avoiding the misguidance of the training of the state estimator to the parameters of the diffusion network model, and ensuring that the diffusion network model can focus on its own goal, i.e. generating reasonable action sequence.
[0062] In some optional embodiments of the embodiments of the present application, the output of the above-mentioned state estimator can include various forms, wherein it can be the state prediction result of the current moment state vector and the next moment state vector, and for such state vectors, it is usually represented by two states (True / False) of the vector. Or, it can also be represented by a One-Hot form, for example, generating a One-Hot encoding vector [0, 0, 0, 1, 0,..., 0] corresponding to the "successful grasping" state. The state estimator can include a Multilayer Perceptron (MLP). Further, as shown in the step 204, based on the global condition vector and the action data of the first time step in the plurality of time steps, the state estimator is used to perform state prediction to obtain the state prediction result of the next moment, mainly including: Figure 4
[0063] Step 401: performing nonlinear mapping through the output layer of the multilayer perceptron to obtain neuron output value.
[0064] This step aims to perform a nonlinear mapping by the output layer of the state estimator, which contains neurons matching the discrete quantity parameters, by the execution subject.
[0065] Step 402: generating a probability distribution vector of the discrete quantity parameters based on the activation function and the neuron output value.
[0066] This step aims to apply an activation function to the neuron output value of the output layer to generate a probability distribution vector of the discrete quantity.
[0067] Step 403: generating a hot encoding used to represent the next state vector based on the probability distribution vector and the corresponding dimension.
[0068] This step aims to set the dimension corresponding to the neuron with the highest probability to 1 and the remaining dimensions to 0 based on the probability distribution vector, thereby forming a hot encoding (One-Hot) used to represent the next state vector.
[0069] Please refer to Figure 5 , Figure 5 The flowchart of another diffusion strategy model optimization method combined with a state memory mechanism provided by the embodiments of the present disclosure, wherein the flowchart 500 comprises the following steps:
[0070] Step 501: obtaining state information and image information of a target object at a current time.
[0071] Step 502: generating a global condition vector based on the state information and the image information.
[0072] Step 503: inputting the global condition vector into a diffusion network model to obtain a predicted action sequence at a next time.
[0073] Step 504: using a state estimator to perform state prediction based on the global condition vector and action data at a first time step in a plurality of time steps to obtain a state prediction result at the next time.
[0074] The above steps 501-504 are consistent with steps 201-204 as shown in Figure 2 , and the same part of the content is described in the corresponding part of the previous embodiment, which will not be described here.
[0075] Step 505: passing the state prediction result at the next time to a next diffusion step to form state information at a current time of the next diffusion step.
[0076] In this embodiment, the state prediction result obtained at the current diffusion step can be further used as the state information at the current time of the next diffusion step to perform state prediction of the next diffusion step.
[0077] Step 506: taking the state information of the current time of the next diffusion step as one of the input conditions of the next diffusion step, for the diffusion network model to make a prediction of the subsequent action sequence, and for the state estimator to make a prediction of the subsequent state, to obtain the state prediction result of the subsequent time, forming a closed-loop iteration of state prediction.
[0078] Further, the diffusion strategy model optimization method combined with the state memory mechanism can further include: controlling the target object to perform corresponding actions based on the predicted action sequence and the state prediction result.
[0079] This step aims to control the target object to perform corresponding actions based on the predicted action sequence and the state prediction result after obtaining the predicted action sequence and the state prediction result of the next time by the above-mentioned execution subject.
[0080] Through the above process, the obtained state prediction result is passed to the next time step, and is integrated into the diffusion reasoning process of each step in the subsequent execution process, capturing long-term dependencies, thereby realizing cross-step information transmission, enabling the obtained state prediction result to be more consistent with the actions and states of the target object, realizing more accurate control, and improving the success rate of action execution.
[0081] In some optional embodiments of the embodiments of the present application, the training process of the diffusion network model and the state estimator described in any of the above embodiments is also adjusted accordingly. Figure 4 is a flowchart of the model training method of the embodiments of the present application. As shown in Figure 6 , the training process mainly includes:
[0082] Step 601: using a first proportion of sample data for labeling as a training set, using a second proportion of sample data as a validation set, and adding pseudo-labels to the remaining unlabeled sample data in combination with an initial diffusion network model and an initial state estimator.
[0083] This step aims to select a first proportion of sample data for labeling as a training set, and use a second proportion of sample data as a validation set, and input the remaining unlabeled sample data into the initial diffusion network model and the initial state estimator to add pseudo-labels to each unlabeled sample data. The initial diffusion network model is an untrained diffusion network model, and the initial state estimator is an untrained state estimator. The first proportion and the second proportion can be set according to actual needs, which can be the same or different. Illustratively, 5% of the sample data can be selected for labeling to obtain a training set; 5% of the sample data is selected as a validation set; and pseudo-labels are added to the remaining 90% of the unlabeled sample data.
[0084] In some optional embodiments of the embodiments of the present application, the remaining unlabeled sample data is input into the initial diffusion network model and the initial state estimator for training, the state of the remaining unlabeled sample data is predicted, the uncertainty of the remaining unlabeled sample data is calculated, and the diffusion network model and the state estimator after preliminary training are obtained; the current remaining unlabeled sample data is input into the diffusion network model and the state estimator after preliminary training to obtain a first prediction result, wherein the current remaining unlabeled sample data includes data that has not been processed in the remaining unlabeled sample data after preliminary training; and pseudo labels are added to the current remaining unlabeled sample data based on the first prediction result.
[0085] In the above process, the prediction uncertainty of the sample in the unlabeled data is determined by calculating the entropy value of the prediction result or the difference in multi-model voting, the entropy value measures the degree of confusion of the state probability distribution of the discrete task, and the difference in multi-model voting measures the degree of divergence of different models in predicting the same state.
[0086] Step 602: Training the initial diffusion network model and the initial state estimator based on the sample data with added pseudo labels, the training set and the validation set, respectively, to obtain a diffusion network model and a state estimator.
[0087] This step aims to train the initial diffusion network model and the initial state estimator based on the sample data with added pseudo labels, the training set and the validation set, thereby obtaining a trained diffusion network model and a state estimator.
[0088] In some optional embodiments of the embodiments of the present application, the sample data with an uncertainty higher than a preset threshold is labeled to obtain real state label sample data, which is added to the validation set to obtain an updated validation set; the updated validation set is used to evaluate the diffusion network model and the state estimator after preliminary training, and the pseudo labels are adjusted to obtain a diffusion network model and a state estimator.
[0089] Through the above process, since the model has been preliminarily trained and evaluated and optimized by the validation set, the prediction result has a certain credibility. The uncertainty sampling, manual labeling, validation set construction, model evaluation and pseudo label expansion process described above will be repeated continuously. In each iteration, the model is evaluated using the updated validation set, the historical pseudo label data is filtered and corrected, and the quality of the pseudo labels is gradually improved with the iteration process.
[0090] Further, the diffusion network model and the state estimator can be fine-tuned on the full sample data for 50 epochs, with an initial learning rate of 1e-4 and an Adam optimizer for parameter updating. The diffusion network model and the state estimator can be trained on the full data set for 50 complete times to further improve the accuracy and generalization ability of the diffusion network model and the state estimator.
[0091] To deepen the understanding, the present application also gives a specific implementation scheme in combination with a specific application scenario, please see as shown in Figure 7A and Figure 7B .
[0092] In this embodiment, the target object is taken as a mechanical arm as an example. The end of the mechanical arm is provided with an end effector (such as a flexible gripper, a mechanical gripper, a suction type gripper, etc.), and the gripper is provided with a visual / physical marker (Tag) and an inertial measurement unit (IMU) on the gripper. The Tag can be a two-dimensional code or a bar code, which is used to mark the position of the two grippers, and according to the position, the gripper width of the two grippers can be obtained, and the 3D pose information (including position information and Euler angle) relative to the collection camera can be obtained by detecting the image coordinates of the Tag corner points, and the initial angle can be calibrated based on the relative pose of the Tag coordinate system and the collection camera; the IMU can be arranged at the end of the mechanical arm, which is used to detect the swing angle and other information of the mechanical arm, so as to obtain the pose or trajectory of the mechanical arm. A GoPro camera can be used to collect image information of the mechanical arm, the sampling rate of the GoPro camera is 30 Hz, and the image size is 224x224.
[0093] In specific implementation, step 701: obtaining the state information and image information of the mechanical arm. Correspondingly, the state information includes: gripper width, end effector position and Euler angle on three coordinate axes, initial angle and other information, and the current state vector (for example, the "grasping success" represented by a hot encoding, corresponding to the One-Hot encoding vector [0, 0, 0, 1, 0,..., 0]). Among them, the gripper width, end effector position and Euler angle on three coordinate axes, initial angle and other information form 16-dimensional continuous quantity parameters, and the state vector is 16-dimensional discrete quantity parameters, thereby forming 32-dimensional state features.
[0094] Step 702: generating a global condition vector based on the state information and image information. A visual encoder (such as ViT-224) outputs a 1536-dimensional low-dimensional image feature based on the image information. In actual application, a 2-frame X-dimensional vector representing the current state can also be included. The image feature, state feature and 2-frame X-dimensional vector are reshaped and spliced to obtain the global condition vector.
[0095] Step 703: inputting the global condition vector into the diffusion network model to obtain a predicted action sequence composed of action data of multiple time steps. The global condition vector is input into the U-Net diffusion network to predict the action sequence of multiple time steps. In this embodiment, the action sequence may, for example, be an action sequence containing 7 steps, each step being 16-dimensional.
[0096] In implementation, the global condition vector can also be aligned with the feature dimension through linear transformation to adapt to the input requirement of the U-Net diffusion network.
[0097] Step 704: Based on the global condition vector and the action data of the first time step in the plurality of time steps, state prediction is performed using a state estimator to obtain a state prediction result of the next moment. After gradient separation operation of the action sequence and the global condition vector, the state estimator is input, and a 32+16-dimensional state prediction result of the next moment is output.
[0098] Step 705: Based on the predicted action sequence and the state prediction result, the robot arm is controlled to perform corresponding actions. Based on the state prediction result, the robot arm can be controlled to perform corresponding actions at the next moment.
[0099] Further, the obtained state prediction result of the next moment can also be used to generate a global condition vector for the current state of the next time step and participate in the subsequent processing process.
[0100] As a possible implementation, the present application also provides another specific implementation scheme, which is described with reference to Figure 7C After obtaining image information through a visual sensor such as GoPro, a ViT-224 or other type of image encoder can be used to extract features to form low-dimensional visual / image features such as 1536 dimensions. It should be noted that the image encoder can also include other types of neural networks such as ResNet, etc.
[0101] The continuous quantity parameters and the low-dimensional image features can be spliced to obtain spliced features such as 1568 dimensions. As another example, the continuous quantity parameters (32 dimensions) can also be mapped and projected into dimension-transformed continuous quantity parameters through linear transformation, and then the dimension-transformed continuous quantity parameters and the low-dimensional image features are spliced to form spliced features.
[0102] The embodiment also includes parameter segmentation of the discrete quantity parameters to form scaling parameters (Gamma) and offset parameters (Beta), and feature-wise linear modulation (FiLM) of the spliced features to modulate the features (such as unet vector) as global condition input into the DP model. The spliced vector is input into the FiLM MLP, and the dimension of the input determines the input layer design of the MLP, ensuring that the state and task information can be mapped to Gamma / Beta to achieve strong association between feature modulation and the current scene.
[0103] Exemplarily, the modulation mechanism of FiLM satisfies: Gamma x (visual features) + Beta. The feature modulation parameters film_params in FiLM are divided into two parts along the last dimension, namely Gamma and Beta. Gamma is used for element-wise multiplication of target features (such as visual features) (visual feature x gamma), and Beta is used for element-wise addition of target features (visual feature + beta).
[0104] The FiLM mechanism realizes deep fusion of state information and visual information, can dynamically adjust the importance of visual features according to different states, increase the weight of state information in the model input, and improve the prediction ability of the model for different state-action combinations. In addition, the FiLM parameters provide explainability of model decision-making, while maintaining high efficiency, realizing state-aware visual feature modulation, and thus generating more accurate and context-related action prediction.
[0105] In the embodiment of the application, the control process realized by the diffusion strategy model optimization method combined with the state memory mechanism as described in any of the above embodiments can achieve the following specific effects:
[0106] In the different color block sorting experiment, the sorting experiment refers to controlling the robot to sort the target object according to the preset color or shape sequence. In this experiment, the success rate of the original DP model is 15%, and the success rate of the embodiment of the application can reach 90%.
[0107] In the experiment of wiping the table with a scouring pad (low color contrast), the success rate of the original DP model is almost 0, and the success rate of the embodiment of the application can stably reach more than 70%.
[0108] Further reference Figure 8 , as an implementation of the method shown in the above figures, the application provides an embodiment of a diffusion strategy model optimization device combined with a state memory mechanism. The device embodiment corresponds to the method embodiment shown in Figure 2 , and the device can be specifically applied to various electronic devices.
[0109] As Figure 8As shown, the diffusion strategy model optimization apparatus 800 of the embodiment of the state memory mechanism can include an information acquisition module 801, a global condition vector generation module 802, an action sequence generation module 803, and a state prediction result generation module 804. The information acquisition module 801 is configured to acquire state information and image information of a target object at a current time; the global condition vector generation module 802 is configured to generate a global condition vector based on the state information and the image information; the action sequence generation module 803 is configured to input the global condition vector into a diffusion network model to obtain a predicted action sequence composed of action data of multiple time steps; and the state prediction result generation module 804 is configured to use a state estimator to perform state prediction based on the global condition vector and the action data of a first time step in the multiple time steps to obtain a state prediction result at a next time.
[0110] In the embodiment, the specific processing of the information acquisition module 801, the global condition vector generation module 802, the action sequence generation module 803, and the state prediction result generation module 804 in the diffusion strategy model optimization apparatus 800 of the state memory mechanism and the technical effects brought by the specific processing can be respectively referred to the related descriptions of steps 201-204 in the embodiment. Figure 2 The related descriptions of steps 201-204 in the embodiment are not repeated here.
[0111] The embodiment of the diffusion strategy model optimization apparatus of the state memory mechanism is provided as a device embodiment corresponding to the above method embodiment. In the diffusion strategy model optimization apparatus of the state memory mechanism provided by the embodiment, the state information and the image information are introduced into the constructed global condition vector, which can be used to represent the state features and visual features of the target object, and the state features and visual features are applied in the process of generating the action sequence by the diffusion network model. In addition, the current state features, the visual features, and the action sequence are combined with the state estimator to predict the hidden state at the next time, so that the generated action instruction can be more consistent with the action and state of the target object, more accurate control is achieved, and the success rate of action execution is improved.
[0112] According to the embodiments of the present application, the present application further provides an electronic device, which comprises at least one executor and a memory connected in communication with the at least one executor; wherein the memory stores instructions executable by the at least one executor, and the instructions are executed by the at least one executor to enable the at least one executor to implement the diffusion strategy model optimization method of the state memory mechanism described in any of the above embodiments when the at least one executor is executed.
[0113] According to the embodiments of the present application, the present application further provides a readable storage medium, which stores computer instructions for enabling a computer to implement the diffusion strategy model optimization method of the state memory mechanism described in any of the above embodiments when the computer is executed.
[0114] According to an embodiment of the present application, the present application also provides a computer program product, which, when executed by an executor, can implement the diffusion strategy model optimization method with the state memory mechanism described in any of the above embodiments.
[0115] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0116] As shown in Figure 9 The electronic device 900 includes an executor 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 can also be stored in the RAM 903. The executor 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0117] Various components in the electronic device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc., a mechanical arm 907, a storage unit 908, such as a magnetic disk, an optical disk, etc., and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0118] The executor 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the executor 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate executor, micro-executor, etc. The executor 901 executes various methods and processes described above, such as the diffusion policy model optimization method with state memory mechanism. For example, in some embodiments, the diffusion policy model optimization method with state memory mechanism can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the executor 901, one or more steps of the diffusion policy model optimization method with state memory mechanism described above can be performed to implement the control of the robotic arm 907 to perform corresponding actions. Alternatively, in other embodiments, the executor 901 can be configured to perform the diffusion policy model optimization method with state memory mechanism by any other appropriate means, such as by means of firmware.
[0119] Those skilled in the art will appreciate that embodiments of the present application can be supplied as a method, a system, or a computer program product. Thus, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage media, etc.) having computer usable program code embodied in the medium.
[0120] The present application is described with reference to the flowchart and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, an embedded processing chip or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 The flowchart and / or block diagrams can include one or more flowcharts and / or block diagrams that illustrate the functions and / or operations that can be performed by the present application. Figure 1 The flowchart and / or block diagrams can include one or more flowcharts and / or block diagrams that illustrate the functions and / or operations that can be performed by the present application.
[0121] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0122] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0123] Obviously, the above-described embodiments are only examples and are not intended to limit the present application. Other variations and modifications can be made based on the above description by those skilled in the art. Here, it is not necessary or possible to exhaust all the embodiments. The obvious variations and modifications derived therefrom are still within the scope of the present application.
Claims
1. A diffusion strategy model optimization method combined with a state memory mechanism, characterized in that, The method comprises the following steps: obtaining state information and image information of a target object at a current time; generating a global condition vector based on the state information and image information; inputting the global condition vector into a diffusion network model to obtain a predicted action sequence composed of action data of multiple time steps; using a state estimator to perform state prediction based on the global condition vector and the action data of the first time step among the multiple time steps to obtain a state prediction result at a next time; wherein the method further comprises the formation of a next diffusion step, comprising: passing the state prediction result at the next time to the next diffusion step to form state information at the current time of the next diffusion step; taking the state information at the current time of the next diffusion step as one of the input conditions of the next diffusion step, so that the diffusion network model performs prediction of a subsequent action sequence and the state estimator performs prediction of a subsequent state to obtain a state prediction result at a subsequent time, forming a closed-loop iteration of state prediction.
2. The method of claim 1, wherein, The method comprises the following steps: obtaining state information and image information of a target object at a current time; obtaining continuous parameters and discrete parameters of the target object, wherein the continuous parameters are used to represent width information, position information and angle information of the target object, and the discrete parameters are used to represent the state or task phase of the target object; concatenating the continuous parameters and the discrete parameters to obtain a state vector; collecting an image of the target object; 3. The method of claim 2, wherein, inputting the image into a visual encoder to obtain low-dimensional image features. The formation of the global condition vector comprises: concatenating the state vector and the low-dimensional image features to obtain the global condition vector; or 4. The method of claim 3, wherein, concatenating the continuous parameters and the low-dimensional image features to obtain a concatenated feature, performing parameter segmentation on the discrete parameters to obtain a scaling parameter and an offset parameter, and performing linear modulation on the feature based on the concatenated feature, the scaling parameter and the offset parameter. The method comprises the following steps: performing a gradient separation operation on the global condition vector to cut off the gradient propagation path; performing the gradient separation operation on the extraction process of the action data of the first time step; inputting the global condition vector after gradient separation and the action data of the first time step after gradient separation into the state estimator to perform state prediction, 5. The method of claim 1, wherein, wherein the global condition vector is obtained by concatenating the state vector and the low-dimensional image features. The state estimator comprises a multi-layer perceptron, and the method comprises the following steps: performing nonlinear mapping through the output layer of the multi-layer perceptron to obtain neuron output values; generating a probability distribution vector of the discrete parameters based on an activation function and the neuron output values; 6. The method according to any one of claims 1 to 5, characterized in that, generating a hot encoding representing the next state vector based on the probability distribution vector and the corresponding dimension. The method further comprises the following steps: The first proportion of sample data is labeled as a training set, and the second proportion of sample data is used as a validation set. Pseudo labels are added to the remaining unlabeled sample data in combination with an initial diffusion network model and an initial state estimator. The initial diffusion network model and the initial state estimator are trained based on the sample data with the pseudo labels, the training set, and the validation set, respectively, to obtain the diffusion network model and the state estimator.
7. The method of claim 6, wherein, The pseudo labels are added to the remaining unlabeled sample data in combination with the initial diffusion network model and the initial state estimator, including: The remaining unlabeled sample data is input into the initial diffusion network model and the initial state estimator for training. The state of the remaining unlabeled sample data is predicted, the uncertainty of the remaining unlabeled sample data is calculated, and a preliminary trained diffusion network model and a preliminary trained state estimator are obtained. The current remaining unlabeled sample data is input into the preliminary trained diffusion network model and the preliminary trained state estimator to obtain a first prediction result, wherein the current remaining unlabeled sample data includes data that has not been processed from the remaining unlabeled sample data after preliminary training. Pseudo labels are added to the current remaining unlabeled sample data based on the first prediction result.
8. The method of claim 7, wherein, The initial diffusion network model and the initial state estimator are trained based on the sample data with the pseudo labels, the training set, and the validation set, respectively, including: The sample data with an uncertainty higher than a preset threshold is labeled to obtain true state label sample data, which is added to the validation set to obtain an updated validation set. The preliminary trained diffusion network model and the preliminary trained state estimator are evaluated using the updated validation set, and the pseudo labels are adjusted to obtain the diffusion network model and the state estimator.
9. A body-aware agent, comprising: It includes: A robot body; An executor configured to execute the diffusion strategy model optimization method with a state memory mechanism as claimed in any one of claims 1-8 to enable the robot body to complete a task.