A robot operation reward generation system and method based on video generation
Through a robot running a reward generation system based on video generation, the image encoding network and video generation diffusion model network predicts the robot trajectory video, combined with multiple rewards for reward computer robot states, the problem of complex and error-prone reward design in the existing technology is solved, and more efficient robot task execution is achieved.
Patent Information
- Application Number
- CN202510397101.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In the prior art, the reward design is complex and error-prone, resulting in poor learning strategies for agents and affecting the robot task execution effect.
The robot running reward generation system based on video generation is adopted to predict the robot running trajectory video through image encoding network and video generation diffusion model network, and combine distance rewards, progress rewards and exploration rewards, and rewards of computer robot status to guide the learning strategies of the agent.
It provides rich learning signals, reduces sample size requirements, improves learning speed and robustness, and is suitable for different environments and tasks, ensuring the correct execution of robot tasks.
Smart Images

Figure CN119919524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot task execution. Specifically, it relates to a robot operation reward generation system and method based on video generation. Background Art
[0002] In recent years, using reinforcement learning for robot task execution has been a hot topic in the field of robot research. In the framework of reinforcement learning, an agent usually obtains the state of the robot based on its running trajectory to interact with the environment, and adjusts the strategy of the robot state according to the obtained reward feedback, so as to optimize the execution of robot tasks.
[0003] To improve the performance of reinforcement learning, reward design is crucial. The design of the reward directly affects the learning efficiency and final performance of the agent. The existing reward design is based on the robot's running trajectory, depends on the robot's task scenario, and requires manual design of the reward function, which is usually very complex and requires in-depth understanding of the details of the robot task to determine the reward in the state space. The process is cumbersome and error-prone. Especially in robot tasks, inappropriate reward design may cause the agent to learn bad strategies, resulting in errors in robot task execution. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the present invention provides a robot operation reward generation system and method based on video generation. The method generates a predicted robot running trajectory video through a video generation diffusion model network, which is used to calculate the reward of the robot's visual observation image, provides rich learning signals for the reinforcement learning agent, effectively guides the reinforcement learning agent to learn the strategies of complex robot tasks, and better optimizes the process of robot task execution.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions: A robot operation reward generation system based on video generation, comprising: an image encoding network, a video generation diffusion model network, and a reward module;
[0006] The image encoding network is used to generate a latent space encoded image according to the visual observation image of the robot, including an initial latent space encoded image generated by the initial visual observation image and a current latent space encoded image generated by the current visual observation image;
[0007] The video generation diffusion model network predicts the robot running trajectory video according to the generated initial latent space encoded image and the initial visual observation image;
[0008] The reward module calculates the reward for each robot state on the current latent space encoded image based on the robot state on the current latent space encoded image and the robot state on the predicted robot running trajectory video. The agent learns a strategy to maximize the expected cumulative reward based on the calculated reward, obtains the reward for the current visual observation image, and guides the task execution of the robot.
[0009] Further, the image encoding network consists of an encoder and a decoder. The encoder decomposes the visual observation image into several discrete vector representations, and the decoder converts the vector representations back into a high-dimensional image space to obtain the latent space encoded image.
[0010] Further, before the image encoding network generates the latent space encoded image based on the visual observation image of the robot, the image encoding network needs to be trained using a generative adversarial network. The specific process is as follows:
[0011] i. Collect the visual observation images of the robot;
[0012] ii. Input the collected visual observation images of the robot into the encoder of the image encoding network to decompose them into several discrete vector representations;
[0013] iii. Input the discrete vector representations into the generator in the decoder. Through encoding and reconstruction, convert the vector representations back into a high-dimensional image space. Evaluate the image quality of the converted high-dimensional image space through the discriminator in the decoder, and update the parameters of the image encoding network;
[0014] iv. Repeat steps ii-iii until the evaluation function of the image quality of the converted high-dimensional image space converges, and complete the training of the image encoding network.
[0015] Further, the evaluation function of the image quality of the converted high-dimensional image space Specifically:
[0016] ,
[0017] Among them, represents the visual observation image of the robot, represents that the visual observation image of the robot is converted into a low-dimensional vector representation through the encoder, represents the discrete vector representation obtained by mapping to the vector set, represents converting back into a high-dimensional image space to obtain the latent space encoded image,
[0018] Further, the video generation diffusion model network adds noise to the image encoded in the initial latent space, and uses the initial visual observation image and the task language description corresponding to the initial visual observation image as the learning conditions of the video generation diffusion model network to learn to remove the noise on the image encoded in the initial latent space and predict the robot running trajectory video.
[0019] Further, before predicting the robot running trajectory video, the video generation diffusion model network optimizes the parameters of the video generation diffusion model network by maximizing the log-likelihood estimation of the robot running trajectory video.
[0020] Further, the reward module includes: a distance reward module, a progress reward module, an exploration reward module, and a reward synthesis module;
[0021] The distance reward module calculates the distance reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image and the robot state on the robot running trajectory video;
[0022] The progress reward module calculates the progress reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image and the robot state on the robot running trajectory video;
[0023] The exploration reward module determines the exploration reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image;
[0024] The reward synthesis module is used to sum up the distance reward, the progress reward, and the exploration reward to obtain the reward for each robot state on the current latent space encoded image.
[0025] Further, the reward for each robot state on the current latent space encoded image is calculated as follows:
[0026] ,
[0027] where represents the distance reward for the t th robot state on the current latent space encoded image, , represents the similarity function, represents the t th robot state on the current latent space encoded image, represents the th robot state on the predicted robot running trajectory video, , represents the number of robot states on the predicted robot running trajectory video,h Indicates the index of, which represents the h th robot state on the predicted robot running trajectory video; which represents the progress reward at the t th robot state on the current latent space encoded image, ; which represents the exploration reward at the t th robot state on the current latent space encoded image.
[0028] Furthermore, the calculation process of the expected cumulative reward is as follows:
[0029] ,
[0030] wherein, C represents the expected cumulative reward, represents the expectation, represents the discount factor corresponding to the t th robot state on the current latent space encoded image.
[0031] Furthermore, the present invention also provides a method for generating a robot running reward of the above-mentioned robot running reward generation system based on video generation, including the following steps:
[0032] Step S1. After the environment is initialized, the initial visual observation image of the robot in the environment is input into the image encoder network to generate an initial latent space encoded image;
[0033] Step S2. The initial latent space encoded image and the initial visual observation image are input into the video generation diffusion model network to predict the robot running trajectory video;
[0034] Step S3. The visual observation image of the robot in the current environment is input into the image encoder network to generate the current latent space encoded image;
[0035] Step S4. The robot state on the current latent space encoded image and the robot state on the predicted robot running trajectory video are input into the reward module to calculate the reward of each robot state on the current visual observation image. The agent learns a policy to maximize the expected cumulative reward according to the calculated reward, obtains the reward of the current visual observation image, and guides the task execution of the robot.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] The robot running reward generation system and method based on video generation of the present invention uses the initial visual observation image of the robot and the corresponding task language description to guide the video generation diffusion model network to predict the robot running trajectory video that conforms to the task language description, which is used to calculate the reward of the robot visual observation image. The reward calculation process is only applied to the image space, can adapt to the general robot architecture, and is not affected by the different spatial structures of heterogeneous robots, realizing a more general application scenario;
[0038] In the robot running reward generation system and method based on video generation of the present invention, the reward module sums up the distance reward, progress reward and exploration reward. Among them, the distance reward can guide the robot strategy to learn in a direction that is more matched between the robot observation state and the predicted robot running trajectory video; the progress reward can evaluate the completion progress of the robot task and ensure the correct execution of the robot task; the exploration reward encourages the agent of reinforcement learning to quickly learn how to recover from suboptimal or bad states, thereby improving the learning efficiency; the reward of each robot state on the current latent space encoded image is obtained by summing up the three rewards. The agent learns the strategy of maximizing the expected cumulative reward according to the calculated reward, thereby obtaining the reward of the robot visual observation image and guiding the task execution of the robot;
[0039] The robot running reward generation system and method based on video generation of the present invention designs a general reward in the reward module, provides a dense and stable reward signal for the agent, reduces the number of samples required by the agent during the exploration process, and speeds up the learning speed; at the same time, it does not need to rely on the environmental feedback of the robot task, can be applied to various robot tasks, and can provide a consistent learning signal in different environments, improving the robustness and adaptability of the agent when facing different tasks and environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic diagram of the robot running reward generation system based on video generation of the present invention;
[0041] Figure 2 is a schematic diagram of the reward module in the present invention;
[0042] Figure 3 is a flowchart of the robot running reward generation method based on video generation of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0043] The technical solutions of the present invention will be further explained below with reference to the accompanying drawings.
[0044] As Figure 1 is a schematic diagram of the robot running reward generation system based on video generation of the present invention. The robot running reward generation system includes: an image encoding network, a video generation diffusion model network and a reward module;
[0045] The image encoding network is used to generate a latent space encoded image based on the visual observation image of the robot, including an initial latent space encoded image generated from the initial visual observation image and a current latent space encoded image generated from the current visual observation image;
[0046] The video generation diffusion model network predicts the robot running trajectory video based on the generated initial latent space encoded image and the initial visual observation image;
[0047] The reward module calculates the reward for each robot state on the current latent space encoded image based on the robot state on the current latent space encoded image and the robot state on the predicted robot running trajectory video. The agent learns a strategy to maximize the expected cumulative reward based on the calculated reward, obtains the reward for the current visual observation image, and guides the task execution of the robot.
[0048] In the present invention, the reward calculation process only involves the visual observation image of the robot and the robot running trajectory video predicted from the visual observation image, which can avoid relying on environmental rewards and solve the difficulty of manual design of environmental rewards. At the same time, since the reward calculation process is only applied to the image space, it can adapt to a general robot architecture and is not affected by the different spatial structures of heterogeneous robots, realizing a more general application scenario.
[0049] It should be noted that the robot tasks applicable to the present invention can be free combination, order swapping or individual execution of multiple robots, and do not need to rely on or depend on a fixed execution order.
[0050] The image encoding network consists of an encoder and a decoder. The encoder decomposes the visual observation image into several discrete vector representations, aiming to extract the feature information of the visual observation image layer by layer. In a technical solution of the present invention, the image encoder network may include a convolutional layer and a pooling layer to gradually extract the features of the visual observation image and convert it into a low-dimensional discrete vector representation through an embedding layer or a fully connected layer. The decoder converts the vector representation back to the high-dimensional image space, aiming to restore the feature information of the image layer by layer to obtain a latent space encoded image, which has a similar content and structure to the input visual observation image.
[0051] Before the image encoding network generates the latent space encoded image based on the visual observation image of the robot, the image encoding network needs to be trained using a generative adversarial network. Among them, the generator can use transposed convolution or upsampling layers to gradually reconstruct the image, and the discriminator uses convolutional layers to judge the authenticity of the reconstructed image, and trains the image encoding network through a binary classification task. During the training process, the goal of the image encoder network is to minimize the image reconstruction error and the adversarial loss of the discriminator, while using vector quantization technology to ensure the discreteness and compactness of the encoding. The image reconstruction error can use the mean square error to measure the difference between the original image and the reconstructed image. The discriminator can use the cross-entropy loss as the loss function. The optimizer usually uses the Adam optimizer and is iteratively updated according to the historical visual observation images in the static dataset to gradually optimize the parameters of the image encoder network, so as to obtain an efficient and excellent-performing image encoder network.
[0052] The specific process of training the image encoding network is as follows:
[0053] i. Collect the historical visual observation images of the robot and the corresponding task language descriptions of the historical visual observation images to form a static dataset, and uniformly sample the visual observation images of the robot in the static dataset to obtain an image sequence of a fixed length;
[0054] ii. Input the image sequence of the fixed length into the encoder of the image encoding network and decompose it into several discrete vector representations;
[0055] iii. Input the discrete vector representations into the generator in the decoder, and through encoding reconstruction, convert the vector representations back to the high-dimensional image space. Evaluate the image quality of the image converted back to the high-dimensional image space through the discriminator in the decoder, and update the parameters of the image encoding network;
[0056] The evaluation function of the image quality of the image converted back to the high-dimensional image space in the present invention Specifically:
[0057] ,
[0058] Among them, represents the visual observation image of the robot, represents that the visual observation image of the robot is converted into a low-dimensional vector representation through the encoder, represents the discrete vector representation obtained by mapping to the vector set, represents converted back to the high-dimensional image space to obtain the latent space encoded image, represents the second norm;
[0059] iv. Repeat steps ii-iii until the evaluation function of the image quality in the high-dimensional image space converges, completing the training of the image encoding network.
[0060] The video generation diffusion model network generates a realistic video frame sequence by gradually adding and removing noise to the initially encoded images in the latent space. Specifically, the video generation diffusion model network gradually adds noise to the initially encoded images in the latent space through a multi-step forward diffusion process to generate a series of intermediate states; then through a reverse diffusion process, using the initial visual observation image and the corresponding task language description of the initial visual observation image as the learning conditions of the video generation diffusion model network, it learns to remove the noise on the initially encoded images in the latent space and predicts the robot's running trajectory video. The video generation diffusion model network ensures the spatial and temporal consistency of the predicted robot running trajectory video by capturing the temporal relationship between video frames, and can handle the multimodality and complexity of the image sequence in the video data, being applicable to video generation tasks in multiple scenarios.
[0061] Before the video generation diffusion model network predicts the robot's running trajectory video, by maximizing the log-likelihood estimation of the robot's running trajectory video to optimize the parameters of the video generation diffusion model network and improve the generation ability and accuracy of the video generation diffusion model network.
[0062] In one technical solution of the present invention, maximizing the log-likelihood estimation of the robot's running trajectory video can be converted into an iterative solution process that makes the loss function of the video generation diffusion model network converge. The loss function of the video generation diffusion model network is expressed as:
[0063] ,
[0064] where represents the parameters of the video generation diffusion model network, represents the visual observation images in the static dataset, represents the corresponding task language description, K represents the maximum diffusion step, represents Gaussian noise, represents the static dataset, represents after k diffusion steps of the visual observation image, represents the added original noise, represents the expectation, represents the noise term predicted by the video generation diffusion model network.
[0065] As shown in Figure 2 , the reward module includes: a distance reward module, a progress reward module, an exploration reward module, and a reward synthesis module.
[0066] The distance reward module calculates the distance reward for each robot state on the current latent space encoded image based on the robot state on the current latent space encoded image and the robot state on the robot running trajectory video. The distance reward can guide the robot policy to learn in a direction that is more matched between the robot observation state and the predicted robot running trajectory video. The calculation process of the distance reward is as follows: , where represents the distance reward for the t -th robot state on the current latent space encoded image, represents the similarity function, represents the t -th robot state on the current latent space encoded image, represents the -th robot state on the predicted robot running trajectory video, , represents the number of robot states on the predicted robot running trajectory video, h represents index of, represents the h -th robot state on the predicted robot running trajectory video.
[0067] The progress reward module calculates the progress reward for each robot state on the current latent space encoded image based on the robot state on the current latent space encoded image and the robot state on the robot running trajectory video, and can evaluate the completion progress of the robot task. The reward of reinforcement learning should guide the robot to execute each key step in order to ensure the correct execution of the robot task. In some embodiments, in order to reduce the tendency of the robot to skip important intermediate links and directly jump to the end of the trajectory, a concept of "reaching state" is introduced, and this mechanism uses a similarity threshold to evaluate the progress , in the sequence obtained by the reinforcement learning agent interacting with the environment to learn the predicted trajectory, where marks the robot state , represents the number of robot states on the predicted robot running trajectory video, h represents index of, represents the state farthest from the image along the predicted trajectory in the time trajectory of the h -th robot state on the predicted robot running trajectory video. In some embodiments, the progress reward is defined as , where Represents the progress reward in the t th robot state on the current latent space encoded image.
[0068] The exploration reward module determines the exploration reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image , encouraging the reinforcement learning agent to quickly learn how to recover from suboptimal or bad states, thereby improving the learning sample efficiency. In some embodiments, the exploration reward uses the Random Network Distillation (RND) technique. RND calculates the exploration reward through a fixed randomly initialized target network and a trainable prediction network. Specifically, when the agent visits a new, unseen robot state, the prediction error is large, resulting in a higher exploration reward; for familiar states, the prediction error is small, and the exploration reward is correspondingly reduced.
[0069] The reward synthesis module is used to sum the distance reward, progress reward, and exploration reward to obtain the reward for each robot state on the current latent space encoded image . Based on the calculated reward, the agent can reduce the number of samples required during the exploration process, speed up the learning speed, learn the policy that maximizes the expected cumulative reward, thereby obtaining the reward for the robot's visual observation image and guiding the robot's task execution.
[0070] In some embodiments, the Q-Learning algorithm is used to train the policy according to the reward function. Specifically, in a certain robot state, the reward of the robot state is calculated. Based on the calculated reward, the agent learns the policy that maximizes the expected cumulative reward. Among them, the calculation process of the expected cumulative reward is:
[0071] ,
[0072] where C represents the expected cumulative reward, represents the expectation, represents the t th discount factor corresponding to the robot state on the current latent space encoded image.
[0073] According to an embodiment of the present application, at least one of the distance reward module, the progress reward module, the exploration reward module, and the reward synthesis module can be at least partially implemented as a hardware circuit, such as an FPGA, a PLA, a system-on-chip, a system-on-substrate, a system-on-package, an ASIC, or any other reasonable manner that can integrate or package circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them.
[0074] Alternatively, at least one of the distance reward module, the progress reward module, the exploration reward module, and the reward synthesis module can be at least partially implemented as a computer program module, and when the computer program module runs, it can execute corresponding functions.
[0075] The robot running reward generation system based on video generation according to the present invention can be a hardware device independently set in a terminal, or a software program running in the terminal. For example, when the terminal is a desktop computer, the robot running reward generation system can be embodied as an application program such as software in the desktop computer.
[0076] As Figure 3 , the present invention also provides a method for generating a robot running reward based on video generation, including the following steps:
[0077] Step S1: After the environment is initialized, input the initial visual observation image of the robot in the environment into the image encoder network to generate an initial latent space encoded image;
[0078] Step S2: Input the initial latent space encoded image and the initial visual observation image into the video generation diffusion model network to predict the robot running trajectory video;
[0079] Step S3: Input the visual observation image of the robot in the current environment into the image encoder network to generate the current latent space encoded image;
[0080] Step S4: Input the robot state on the current latent space encoded image and the robot state on the predicted robot running trajectory video into the reward module, calculate the reward for each robot state on the current visual observation image, and the agent learns a strategy to maximize the expected cumulative reward based on the calculated reward, obtains the reward for the current visual observation image, and guides the task execution of the robot.
[0081] The method for generating a robot running reward based on video generation according to the present invention generates a predicted robot running trajectory video through the video generation diffusion model network, which is used to calculate the reward for the robot visual observation image, provides rich learning signals for the reinforcement learning agent, effectively guides the reinforcement learning agent to learn the strategy of complex robot tasks, and better optimizes the robot task execution process.
[0082] In one technical solution of the present invention, there is also provided a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the robot running reward generation method based on video generation as described above.
[0083] In one technical solution of the present invention, there is also provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the robot running reward generation method based on video generation as described above is implemented.
[0084] In one technical solution of the present invention, there is also provided a computer program product including a computer program, and when the computer program is executed by a processor, the robot running reward generation method based on video generation as described above is implemented.
[0085] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0086] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0087] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A robot operation reward generation system based on video generation, characterized in that Including: An image encoding network, a video generation diffusion model network, and a reward module; The image encoding network is used to generate a latent space encoded image based on the visual observation image of the robot, including an initial latent space encoded image generated from the initial visual observation image and a current latent space encoded image generated from the current visual observation image; The video generation diffusion model network predicts the robot running trajectory video based on the generated initial latent space encoded image and the initial visual observation image; The reward module calculates the reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image and the robot state on the predicted robot running trajectory video. The agent learns a strategy to maximize the expected cumulative reward based on the calculated reward, obtains the reward for the current visual observation image, and guides the task execution of the robot; The reward module includes: a distance reward module, a progress reward module, an exploration reward module, and a reward synthesis module; The distance reward module calculates the distance reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image and the robot state on the robot running trajectory video; The progress reward module calculates the progress reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image and the robot state on the robot running trajectory video; The exploration reward module determines the exploration reward for each robot state on the current latent space encoded image according to the robot state on the current latent space encoded image; The reward synthesis module is used to sum up the distance reward, the progress reward, and the exploration reward to obtain the reward for each robot state on the current latent space encoded image.
2. The robot operation reward generation system based on video generation according to claim 1, wherein The image encoding network consists of an encoder and a decoder. The encoder decomposes the visual observation image into several discrete vector representations, and the decoder converts the vector representations back to the high-dimensional image space to obtain the latent space encoded image.
3. The robot operation reward generation system based on video generation according to claim 2, wherein, Before the image encoding network generates the latent space encoded image based on the visual observation image of the robot, the image encoding network needs to be trained using a generative adversarial network. The specific process is as follows: i. Collect the visual observation images of the robot; ii. Input the collected visual observation images of the robot into the encoder of the image encoding network to decompose them into several discrete vector representations; iii. Input the discrete vector representations into the generator in the decoder. Through encoding and reconstruction, convert the vector representations back to the high-dimensional image space. Evaluate the image quality of the converted high-dimensional image space through the discriminator in the decoder, and update the parameters of the image encoding network; iv. Repeat steps ii - iii until the evaluation function of the image quality of the converted high-dimensional image space converges, and complete the training of the image encoding network.
4. The robot operation reward generation system based on video generation according to claim 3, wherein The evaluation function for the image quality of the conversion back to the high-dimensional image space Specifically: Among them, represents the visual observation image of the robot, represents that the visual observation image of the robot is converted into a low-dimensional vector representation through an encoder, represents the discrete vector representation obtained by mapping represents the latent space encoded image obtained by converting back to the high-dimensional image space, represents the two-norm.
5. A robot operation reward generation system based on video generation according to claim 1, characterized in that, The video generation diffusion model network adds noise to the initial latent space encoded image, and uses the initial visual observation image and the corresponding task language description of the initial visual observation image as the learning conditions of the video generation diffusion model network to learn to remove the noise on the initial latent space encoded image and predict the robot running trajectory video.
6. The robot operation reward generation system based on video generation according to claim 5, characterized in that, Before predicting the running trajectory video of the robot, the video generation diffusion model network optimizes the parameters of the video generation diffusion model network by maximizing the log-likelihood estimation of the running trajectory video of the robot.
7. The robot operation reward generation system based on video generation according to claim 1, wherein The reward for each robot state on the current latent space encoded image is calculated as follows: Among them, represents the distance reward in the t th robot state on the current latent space encoded image, , represents the similarity function, represents the t th robot state on the current latent space encoded image, represents the th robot state on the predicted robot running trajectory video, , represents the number of robot states on the predicted robot running trajectory video, h represents 's index, represents the h th robot state on the predicted robot running trajectory video; represents the progress reward in the t th robot state on the current latent space encoded image, ; represents the exploration reward in the t th robot state on the current latent space encoded image.
8. The robot operation reward generation system based on video generation according to claim 7, wherein The calculation process of the expected cumulative reward is as follows: Among them, C represents the expected cumulative reward, represents the expectation, represents the discount factor corresponding to the t th robot state on the current latent space encoded image.
9. A method for generating robot operation rewards of the robot operation reward generation system based on video generation according to any one of claims 1-8, characterized in that, It includes the following steps: Step S1: After the environment is initialized, the initial visual observation image of the robot in the environment is input into the image encoder network to generate an initial latent space encoded image; Step S2: The initial latent space encoded image and the initial visual observation image are input into the video generation diffusion model network to predict the running trajectory video of the robot; Step S3: The visual observation image of the robot in the current environment is input into the image encoder network to generate the current latent space encoded image; Step S4: The robot state on the current latent space encoded image and the robot state on the predicted running trajectory video of the robot are input into the reward module to calculate the reward for each robot state on the current visual observation image. The agent learns a strategy to maximize the expected cumulative reward based on the calculated reward, obtains the reward for the current visual observation image, and guides the task execution of the robot.
Citation Information
Patent Citations
Robot state prediction method and device, electronic equipment and storage medium
CN117733874A