Diffusion action prediction method and device based on target view angle generation and electronic equipment

Through the Diffusion action prediction method generated based on the target perspective, a multi-view angle generation neural network model is used to perform three-dimensional scene reconstruction and image generation, and combined with the diffusion strategy to perform action prediction, the problem that robot action prediction is limited by a fixed perspective is solved, and the convenience of action prediction and interaction accuracy are improved.

CN120108038APending Publication Date: 2025-06-06TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510254364.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, robot motion prediction is limited by a fixed perspective, resulting in poor convenience of motion prediction and low interaction accuracy.

Method used

A Diffusion action prediction method generated based on the target perspective is proposed. By acquiring images from multiple initial perspectives, a multi-view angle generation neural network model is used to perform three-dimensional scene reconstruction and uniform sampling, an image corresponding to the target perspective is generated, and action prediction is performed in combination with diffusion strategies.

Benefits of technology

This method can reduce the limitation on the viewing angle of the camera device, improve the convenience of robot motion prediction and the accuracy of interaction with the environment, and generate high-quality embodied intelligent movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108038A_ABST
    Figure CN120108038A_ABST
Patent Text Reader

Abstract

The invention provides a Diffusion action prediction method and device generated based on a target view angle and electronic equipment, and the method comprises the steps: obtaining a first image corresponding to each initial view angle in a plurality of initial view angles; performing scene reconstruction on the first image corresponding to each initial view angle by adopting a multi-view-angle generation neural network model to obtain three-dimensional scene information; uniformly sampling the three-dimensional scene information by adopting a multi-view generation neural network model and calculating multi-view image features of each sampling point in the sampling point set; according to the multi-view image features of the sampling points, a multi-view generation neural network model is adopted to generate a second image corresponding to a target view; and performing action prediction on the robot by adopting a diffusion Policy and the second image, and determining the intelligent action of the robot, thereby solving the problems of reducing the visual angle limitation on a camera device, improving the convenience of robot action prediction and improving the accuracy of interaction between the robot and the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, specifically to the field of three-dimensional vision and embodied intelligence in the field of vision, and in particular to a Diffusion action prediction method, device and electronic device based on purpose perspective generation. Background Art

[0002] With the development of science and technology, robots are widely used in all aspects of production and life. For example, robot prediction technology can be used to predict the next action of the robot. Among them, this technology can be applied to various robots, including but not limited to humanoid robots or four-group robots, to improve the autonomy and adaptability of the robot. For example, the robot's posture information and images corresponding to a fixed perspective can be used to train the model. When performing action prediction, it is necessary to obtain images corresponding to the fixed perspective, which makes the action prediction more limited and the convenience of action prediction is poor. Summary of the invention

[0003] The present application aims to solve one of the technical problems in the related art at least to some extent.

[0004] To this end, the first purpose of this application is to propose a Diffusion action prediction method based on purpose perspective generation to generate high-quality actions, which can reduce the perspective limitation of the camera device, improve the convenience of robot action prediction and improve the accuracy of the robot's interaction with the environment.

[0005] The second objective of this application is to propose a Diffusion action prediction device based on target perspective generation.

[0006] The third objective of the present application is to provide an electronic device.

[0007] A fourth objective of the present application is to provide a computer-readable storage medium.

[0008] A fifth object of the present application is to provide a computer program product.

[0009] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a Diffusion action prediction method based on target perspective generation, comprising the following steps:

[0010] Acquire a first image corresponding to each initial viewing angle of a plurality of initial viewing angles;

[0011] Using a multi-view generation neural network model to reconstruct the scene of the first image corresponding to each initial view to obtain three-dimensional scene information;

[0012] The multi-view generation neural network model is used to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in the sampling point set;

[0013] According to the multi-view image features of each sampling point, the multi-view generation neural network model is used to generate a second image corresponding to the target view;

[0014] The diffusion strategy Diffusion Policy and the second image are used to predict the action of the robot to determine the action of the embodied intelligence of the robot.

[0015] To achieve the above-mentioned purpose, the second embodiment of the present application proposes a Diffusion action prediction device based on target perspective generation, including:

[0016] An image acquisition unit, used to acquire a first image corresponding to each initial viewing angle of a plurality of initial viewing angles;

[0017] An information acquisition unit, configured to reconstruct the scene of the first image corresponding to each initial perspective by using a multi-perspective generation neural network model to acquire three-dimensional scene information;

[0018] A feature acquisition unit, configured to uniformly sample the three-dimensional scene information using the multi-view generation neural network model and calculate multi-view image features of each sampling point in a sampling point set;

[0019] The image acquisition unit is used to generate a second image corresponding to a target perspective using the multi-view generation neural network model according to the multi-view image features of each sampling point;

[0020] The action prediction unit is used to use the diffusion strategy Diffusion Policy and the second image to predict the action of the robot and determine the action of the embodied intelligence of the robot.

[0021] To achieve the above-mentioned purpose, the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0022] The memory stores computer-executable instructions;

[0023] The processor executes the computer-executable instructions stored in the memory to implement any method described in the first aspect above.

[0024] To achieve the above-mentioned purpose, the fourth aspect embodiment of the present application proposes a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement any one of the methods described in the first aspect above.

[0025] To achieve the above-mentioned purpose, the fifth aspect of the present application proposes a computer program product, including a computer program, which, when executed by a processor, implements any method described in the first aspect above.

[0026] The present application provides a Diffusion action prediction method, device and electronic device based on target perspective generation. The method obtains a first image corresponding to each initial perspective in multiple initial perspectives; uses a multi-perspective generation neural network model to reconstruct the scene of the first image corresponding to each initial perspective to obtain three-dimensional scene information; uses the multi-perspective generation neural network model to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in a sampling point set; uses the multi-perspective generation neural network model to generate a second image corresponding to the target perspective based on the multi-view image features of each sampling point; and uses a diffusion strategy Diffusion. Policy and the second image predict the robot's action and determine the robot's embodied intelligent action, which solves the problem of using the robot's posture information and images corresponding to a fixed perspective to train the model, and needing to obtain images corresponding to the fixed perspective when performing action prediction, which makes the action prediction more restricted and the action prediction less convenient. Since the three-dimensional scene information can be reconstructed through the image of the initial perspective and the neural network model, it can be established through the image of the initial perspective when the image of the target perspective cannot be obtained, thereby improving the accuracy of obtaining the image corresponding to the target perspective to generate high-quality robot embodied intelligent actions, reducing the restrictions on the perspective of the camera device, and being able to generate the generalization of the neural network model from multiple perspectives, which can improve the convenience of action prediction, and can predict the actions of the corresponding perspectives through the images of the robot's environment, so that the robot can interact with the environment and improve the accuracy of the interaction.

[0027] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0029] Figure 1 A flowchart of a Diffusion action prediction method based on target perspective generation provided in an embodiment of the present application;

[0030] Figure 2 A flowchart of a Diffusion action prediction method based on target perspective generation provided in an embodiment of the present application;

[0031] Figure 3 An example schematic diagram of a Diffusion action prediction method based on target perspective generation provided in an embodiment of the present application;

[0032] Figure 4 An example schematic diagram of a Diffusion action prediction method based on target perspective generation provided in an embodiment of the present application;

[0033] as well as Figure 5 A schematic diagram of the structure of a Diffusion action prediction device based on purpose perspective generation provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0035] According to some embodiments, multi-perspective generation of three-dimensional vision is a revolutionary technology for the future, the core of which is to reconstruct a complete three-dimensional scene through image data from multiple perspectives. This technology usually relies on the combination of deep learning and computer vision, and can extract rich spatial information from images at different angles, thereby achieving high-precision three-dimensional reconstruction. With the emergence of emerging deep neural network models such as NeRF (neural radiant field), high-quality three-dimensional visual effects can be generated from a small number of perspectives in some embodiments. NeRF uses deep learning algorithms to analyze and synthesize images, and can infer complete three-dimensional structures from sparse image data, significantly reducing the complexity of acquisition and calculation, thereby not only improving the efficiency of reconstruction, but also enhancing the details and realism of the generated model.

[0036] In some embodiments, embodied intelligence refers to the ability to acquire knowledge and skills through the interaction of the body with the environment, which has received widespread attention in the fields of robotics and artificial intelligence. This form of intelligence emphasizes the importance of the body in perception and cognition, especially in complex tasks and dynamic environments. Through advanced sensors and deep learning technology, embodied intelligent systems can analyze and understand their surroundings in real time, so as to respond flexibly. This ability not only improves the adaptability of robots, but also opens up new possibilities for human-machine collaboration, allowing robots to interact with humans and the environment more naturally.

[0037] The following describes the Diffusion action prediction method and device based on target perspective generation according to an embodiment of the present application with reference to the accompanying drawings.

[0038] Figure 1A flowchart of a Diffusion action prediction method based on target perspective generation provided in an embodiment of the present application.

[0039] To address this problem, the present application provides a Diffusion action prediction method based on target perspective generation to generate high-quality actions, which can reduce the perspective restriction on the camera device, improve the convenience of robot action prediction, and improve the accuracy of the robot's interaction with the environment. Figure 1 As shown, the Diffusion action prediction method based on the target perspective generation includes the following steps:

[0040] Step 101, obtaining a first image corresponding to each initial viewing angle among a plurality of initial viewing angles;

[0041] According to some embodiments, the execution subject of the embodiments of the present application may be, for example, an electronic device. The electronic device does not specifically refer to a fixed device, and the name of the electronic device is not limited. The electronic device may also be called, for example, a terminal, a mobile device, etc. For example, when the device identification of the electronic device changes, the electronic device may also change accordingly.

[0042] In some embodiments, the initial viewing angle may be used to indicate the viewing angle at which the corresponding image can be acquired by the camera device. The initial viewing angle does not specifically refer to a fixed viewing angle. For example, when the camera angle of the camera device changes, the initial viewing angle may also change accordingly. For example, when the angle corresponding to the initial viewing angle changes, the initial viewing angle may also change accordingly.

[0043] In some embodiments, the first image may be, for example, a first image corresponding to each initial viewing angle. The first in the first image is used to distinguish from the remaining images and does not specifically refer to a fixed image. For example, when the parameters corresponding to the camera device change, the first image may also change accordingly. For example, when the acquisition scene corresponding to the first image changes, the first image may also change accordingly. For example, when the initial angle corresponding to the first image changes, the first image may also change accordingly.

[0044] According to some embodiments, the number of initial viewing angles may be, for example, multiple. For example, the number of initial viewing angles may be two or three. The embodiments of the present application are not limited to this. Among them, the method for obtaining the first image corresponding to each initial viewing angle in the multiple initial viewing angles is not limited. For example, the same camera device may be used to capture images at multiple different initial viewing angles, or different camera devices may be used to capture images at different initial viewing angles.

[0045] In some embodiments, a first image corresponding to each initial viewing angle of a plurality of initial viewing angles may be acquired.

[0046] Step 102, using a multi-view generation neural network model to reconstruct the first image corresponding to each initial view to obtain three-dimensional scene information;

[0047] In some embodiments, the multi-view generation neural network model may be, for example, a model that has been trained and can be used for scene reconstruction. The multi-view generation neural network model does not specifically refer to a fixed model. For example, when the model type corresponding to the multi-view generation neural network model changes, the multi-view generation neural network model may also change accordingly. For example, when the model parameters corresponding to the multi-view generation neural network model change, the multi-view generation neural network model may also change accordingly.

[0048] According to some embodiments, the name of the multi-view generation neural network model is also not limited. For example, the multi-view generation neural network model can also be called a scene generation model, a scene generation network model, etc.

[0049] According to some embodiments, scene reconstruction may be, for example, a process of re-establishing a scene corresponding to a target perspective, or a process of acquiring a scene corresponding to a target perspective, or an image corresponding to a target perspective.

[0050] In some embodiments, the three-dimensional scene information may be used to indicate information obtained by reconstructing the three-dimensional scene. The three-dimensional scene information does not specifically refer to a fixed information. For example, when the method of obtaining the three-dimensional scene information changes, the three-dimensional scene information may also change accordingly. For example, when each initial perspective or the first image corresponding to each initial perspective changes, the three-dimensional scene information may also change accordingly.

[0051] In some embodiments, a multi-perspective generation neural network model is used to reconstruct the scene of the first image corresponding to each initial perspective to obtain three-dimensional scene information.

[0052] Step 103, using the multi-view generation neural network model to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in the sampling point set;

[0053] In some embodiments, uniform sampling can be used to indicate a method for sampling three-dimensional scene information. The solution of the present application may not be limited to this sampling method. For example, a sampling method corresponding to the current application scenario can be determined. Uniform sampling can be used to indicate that the probability of each sampling point being collected within the sampling range is the same.

[0054] According to some embodiments, the sampling point set may be, for example, a collection of at least one sampling point. The sampling point set does not specifically refer to a fixed set. For example, when the number of sampling points included in the sampling point set changes, the sampling point set may also change accordingly. For example, when a sampling point in the sampling point set changes, the sampling point set may also change accordingly. For example, when the first image acquired changes, the sampling point set may also change accordingly.

[0055] In some embodiments, the multi-view image feature can be used to indicate the image feature required to generate an image corresponding to the target perspective, or the key image feature required to generate an image corresponding to the target perspective. The multi-view image feature does not specifically refer to a fixed feature. For example, when the sampling method of the image feature changes, the multi-view image feature can also change accordingly. For example, when the three-dimensional scene information changes, the multi-view image feature can also change accordingly.

[0056] In some embodiments, for example, the multi-view generation neural network model may be used to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in the sampling point set.

[0057] Step 104, generating a second image corresponding to a target perspective using the multi-view generation neural network model according to the multi-view image features of each sampling point;

[0058] In some embodiments, the second image can be used to indicate an image corresponding to the target perspective, for example. The second in the second image is used to distinguish it from the remaining images and does not specifically refer to a fixed image. For example, when the multi-perspective generation neural network model changes, the second image may also change accordingly. For example, when the target perspective changes, the second image may also change accordingly. The second image may be, for example, a two-dimensional image. The embodiments of the present application are not limited to this.

[0059] In some embodiments, the target viewing angle may be used to indicate the viewing angle required for action prediction, and the target viewing angle may be a viewing angle different from each initial viewing angle, or may be a viewing angle among multiple initial viewing angles. This embodiment of the present application does not limit this.

[0060] In some embodiments, the multi-view generation neural network model may be used to generate a second image corresponding to the target perspective based on the multi-view image features of each sampling point.

[0061] Step 105: Use the diffusion strategy and the second image to predict the action of the robot, and determine the action of the embodied intelligence of the robot.

[0062] According to some embodiments, the Diffusion Policy may be, for example, a strategy for representing the robot's visual motion strategy as a conditional denoising diffusion process to generate the robot's behavior.

[0063] In some embodiments, the robot may be used to indicate a robot for which action prediction is to be performed. The robot does not specifically refer to a fixed robot. For example, when the identifier corresponding to the robot changes, the robot may also change accordingly.

[0064] According to some embodiments, embodied artificial intelligence (EAI) may be, for example, an intelligent system that senses and acts based on a physical entity. The actions of embodied intelligence may be, for example, actions that sense and acquire information for a robot.

[0065] In some embodiments, a diffusion policy and the second image may be used to predict the action of the robot to determine the action of the embodied intelligence of the robot.

[0066] The present application provides a Diffusion action prediction method, device and electronic device based on target perspective generation. The method obtains a first image corresponding to each initial perspective in multiple initial perspectives; uses a multi-perspective generation neural network model to reconstruct the scene of the first image corresponding to each initial perspective to obtain three-dimensional scene information; uses the multi-perspective generation neural network model to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in a sampling point set; uses the multi-perspective generation neural network model to generate a second image corresponding to the target perspective based on the multi-view image features of each sampling point; and uses a diffusion strategy Diffusion. Policy and the second image predict the robot's action and determine the robot's embodied intelligent action, which solves the problem of using the robot's posture information and images corresponding to a fixed perspective to train the model, and needing to obtain images corresponding to the fixed perspective when performing action prediction, which makes the action prediction more restricted and the action prediction less convenient. Since the three-dimensional scene information can be reconstructed through the image of the initial perspective and the neural network model, it can be established through the image of the initial perspective when the image of the target perspective cannot be obtained, thereby improving the accuracy of obtaining the image corresponding to the target perspective to generate high-quality robot embodied intelligent actions, reducing the limitation of the viewing angle of the camera device, and being able to generate the generalization of the neural network model from multiple perspectives, which can improve the convenience of action prediction, and can predict the actions of the corresponding perspectives through the images of the robot's environment, so that the robot can interact with the environment and improve the accuracy of the interaction.

[0067] This embodiment provides another Diffusion action prediction method based on target perspective generation. Figure 2 A flowchart of a Diffusion action prediction method based on target perspective generation provided in an embodiment of the present application.

[0068] like Figure 2 As shown, the Diffusion action prediction method based on the target perspective generation may include the following steps:

[0069] Step 201, obtaining a first image corresponding to each initial viewing angle among a plurality of initial viewing angles;

[0070] The specific process is as described above and will not be repeated here.

[0071] In some embodiments, the multiple initial viewing angles may be, for example, two initial viewing angles. The first image may be, for example, an RGB image.

[0072] In some embodiments, the technical solution of the embodiment of the present application can be applied to the field of robotics or augmented reality, etc. The embodiment of the present application is not limited to this. The embodiment of the present application can be applied to various embodied intelligence systems.

[0073] Step 202, using the RGB image feature extraction neural network sub-model to extract features from the first images corresponding to the initial viewing angles to obtain an image feature set;

[0074] The specific process is as described above and will not be repeated here.

[0075] The three-dimensional Gaussian reconstruction neural network model includes an RGB image feature extraction neural network sub-model and a Gaussian point cloud reconstruction neural network sub-model, wherein the names of the sub-models are not limited. For example, the RGB image feature extraction neural network sub-model can also be called an RGB image feature extraction neural network, etc.

[0076] In some embodiments, the first image corresponding to each initial perspective may be, for example, a color-depth image. For example, a task in RLBench may be used to render multi-viewpoint training data to train a 3D Gaussian reconstruction neural network model.

[0077] According to some embodiments, the RGB image feature extraction neural network sub-model includes an encoder of a pure visual transformer structure and a decoder of a pure visual transformer structure, and the RGB image feature extraction neural network sub-model is used to extract features of the first image corresponding to each initial perspective to obtain an image feature set, including:

[0078] According to first images corresponding to any two initial viewing angles of the multiple initial viewing angles and camera device parameters corresponding to the initial viewing angles, an encoder of the pure visual transformer structure is used to perform feature extraction to obtain three-dimensional Gaussian parameters;

[0079] According to the three-dimensional Gaussian parameters, a decoder using the pure visual transformer structure generates an image feature set corresponding to the target viewing angle. This can reduce the number of scenes that require a large amount of overlap between camera devices to perform scene reconstruction, reduce the limitation that scene reconstruction cannot be performed when the number of first images is small, and improve the convenience and accuracy of scene reconstruction. In addition, there is no need to use the predicted depth and camera intrinsic parameters to promote the predicted Gaussian parameters to Gaussian primitives in the local coordinate system of each individual frame, and there is no need to convert to the world coordinate system, which can reduce the complexity of scene reconstruction and improve the efficiency of scene reconstruction.

[0080] According to some embodiments, the training data set may include, for example, 10 tasks of the panda robot arm, and each frame of each task renders a color image of 20 viewpoints. The image resolution is designed according to the sensor resolution used in the actual acquisition system. For example, 128*128 can be used as the image rendering resolution. The encoder is used to generate the three-dimensional Gaussian, and the specific architecture is the encoder structure of ViT. Input the camera intrinsic parameters and any two viewpoints to obtain the three-dimensional Gaussian of the entire scene:

[0081]

[0082] Use the decoder to generate the RGB image and depth image of the specified viewpoint based on the three-dimensional Gaussian point cloud. The size of the output RGB image is 128*128*3, and the size of the depth image is 128*128*1, where 128 is the image length and width, and 3 and 1 represent the number of channels.

[0083] Step 203, using the Gaussian point cloud reconstruction neural network sub-model to perform scene reconstruction on the image feature set to obtain three-dimensional scene information;

[0084] The specific process is as described above and will not be repeated here.

[0085] In some embodiments, the name of the three-dimensional scene information is not limited, and may be called, for example, reconstruction space, three-dimensional Gaussian point cloud, etc.

[0086] Using the Gaussian point cloud reconstruction neural network sub-model to reconstruct the scene of the image feature set can, for example, be based on the image feature set, use NoPoSplat to perform three-dimensional Gaussian point cloud reconstruction without camera extrinsic parameters to obtain three-dimensional scene information.

[0087] In some embodiments, the method further comprises:

[0088] Acquire training sample images corresponding to multiple viewing angles and camera device training sample parameters corresponding to the multiple viewing angles, wherein the camera device training sample parameters include the number of viewing angles corresponding to the multiple viewing angles and the intrinsic parameter of the camera device corresponding to each viewing angle of the multiple viewing angles;

[0089] A feedforward network with learnable parameters is obtained by training the camera device, and when the feedforward network meets the model training requirements, the feedforward network is used as the Gaussian point cloud reconstruction neural network sub-model.

[0090] In some embodiments, the training of the camera device to obtain a feedforward network with learnable parameters includes:

[0091] Mapping the training sample image to a standard space using a target image in the training sample image to obtain three-dimensional Gaussian distribution information corresponding to the training sample image;

[0092] According to the three-dimensional Gaussian distribution information, the mapping relationship information between the training sample image and the three-dimensional Gaussian distribution information is obtained, and the mapping relationship information is used as the feedforward network with learnable parameters. Therefore, the Gaussian distribution can be predicted in the standard space through the feedforward network to represent the potential three-dimensional scene from the sparse image with uncertain posture, which can improve the accuracy of the acquisition of three-dimensional scene information and improve the accuracy of the robot action prediction. Among them, the standard space can also be called the standard space.

[0093] According to some embodiments, Figure 3 Another example schematic diagram of a Diffusion action prediction method based on the purpose perspective generation is provided, in which: Figure 3 The dual-view setting is used as an example, and the RGB shortcut connection is omitted. Figure 3 As shown, the imaging device may be, for example, a camera, which may receive sparse multi-view images with uncertain postures and corresponding camera intrinsic parameters. Where V is the number of input views, I represents the internal parameters, and learns a feed-forward network f with learnable parameters θ θ The feed-forward network maps the input non-extrinsic image to a three-dimensional Gaussian distribution in the standard space with the first image as the input, which can represent the potential scene geometry and appearance, and learns the following mapping:

[0094]

[0095] The left side represents the input of the camera internal parameters, and the right side represents the Gaussian parameters, that is, the three-dimensional Gaussian distribution information. Specifically, the center position μ∈R 3 , transparency α∈R, rotation factor r∈R expressed as a quaternion 4 , scale s∈R 3, and the spherical harmonics (SH) c∈R with k degrees of freedom k The first image may be, for example, a target image.

[0096] In some embodiments, the feedforward network can be trained through a large-scale data set, and the technical solution of the present application can be applied to various scenarios and untrained scenarios, thereby increasing the scope of application of the technical solution. Based on this, the technical solution of the present application can be applied, for example, to a) synthesize a new perspective given a target camera transformation relative to a first input perspective, i.e., synthesize the target perspective, and b) perform relative pose estimation between different input perspectives.

[0097] Step 204, using the multi-view generation neural network model to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in the sampling point set;

[0098] The specific process is as described above and will not be repeated here.

[0099] According to some embodiments, uniformly sampling the three-dimensional scene information may specifically be, for example, uniformly sampling the three-dimensional Gaussian point cloud.

[0100] Step 205, according to the multi-view image features of each sampling point, using the multi-view generation neural network model to generate a second image corresponding to the target view;

[0101] The specific process is as described above and will not be repeated here.

[0102] In some embodiments, the multi-view generation neural network model may include, for example, an image rendering model, which may be a multi-view generation neural network model of a ViT architecture.

[0103] Step 206: Use the diffusion strategy and the second image to predict the action of the robot, and determine the action of the embodied intelligence of the robot.

[0104] The specific process is as described above and will not be repeated here.

[0105] According to some embodiments, using a diffusion policy and the second image to predict the action of the robot and determine the action of the embodied intelligence of the robot includes:

[0106] Inputting the second image into a standard temporal Transformer network model to generate visual input information corresponding to the target perspective;

[0107] Obtain random gripper state information, diffusion time step information, and Gaussian distribution noise action information;

[0108] Using the visual input information, the random gripper states information, the diffusion time step information and the noise action information conforming to the Gaussian distribution to perform action prediction, and determine a multi-view prediction image corresponding to the robot;

[0109] The multi-view prediction image is processed in three dimensions for consistency, and the action of the robot's embodied intelligence is determined. Therefore, a standard temporal Transformer network model can be used to fuse information from various sources to improve the efficiency of information fusion. The fused information may include data at different time steps, multiple camera perspectives, and other conditional variables, such as the state of the gripper and the diffusion time step.

[0110] In some embodiments, performing motion prediction may be, for example, performing autoregressive motion prediction.

[0111] According to some embodiments, the second image is input into a standard temporal Transformer network model to generate visual input information corresponding to the target perspective, for example, tokens input information of the visual part can be obtained. The standard temporal Transformer network model can be used as a Dissusion denoising module, for example.

[0112] In some embodiments, performing three-dimensional consistency processing on the multi-view predicted images to determine the embodied intelligent action of the robot includes:

[0113] Performing three-dimensional consistency processing on the multi-view prediction image to obtain the gripper position of the robot;

[0114] According to the position of the robot's gripper, a change matrix corresponding to the robot's gripper is obtained, and according to the change matrix, an action of the robot's embodied intelligence is determined.

[0115] For example, the change matrix T of the gripper point cloud of the front and rear states can be calculated by SVD to obtain the action, and the action can be deployed to the real machine so that the robot can gradually complete the specified task.

[0116] like Figure 4 As shown, the blocks of image observations and rendered action representations are converted into tags, such as Figure 4 Figure 4As shown on the left. Among them, Shoulder Camer; Wrist Camer; Linear Projection; Positional Embedding; Gripper Actions; Noise Added to Actions; Per-Pixel Denoising Directions for each camera view. Among them, the embeddings of Gripper states, Diffusion Time Step, and NoisyAction are connected to the image tags and processed by multiple self-attention layers. In order to enable the network to distinguish information from multiple camera views and different time steps, different learnable position embeddings are added to the corresponding tags. Among them, the corresponding tag embeddings are decoded by linear projection to predict the denoising direction F and the gripper action a for each camera view. g , and the noise ε added to the real action, where the real action can be directly represented in the action space, and the dimensionality of the action is determined by the degrees of freedom of the actual robot arm and gripper used.

[0117] In some embodiments, the RGB image feature extraction neural network sub-model can be used to extract features from the first images corresponding to each initial perspective to obtain an image feature set; the Gaussian point cloud reconstruction neural network sub-model can be used to reconstruct the scene from the image feature set to obtain three-dimensional scene information. Therefore, Gaussian point cloud reconstruction can be performed on the image corresponding to the initial perspective, which can achieve action prediction at different perspectives, enhance the robustness and generalization ability of the action output, and generalize the initial perspective corresponding to the camera device.

[0118] In order to implement the above embodiment, the present application also proposes a Diffusion action prediction device based on target perspective generation.

[0119] Figure 5 A schematic diagram of the structure of a Diffusion action prediction device based on purpose perspective generation provided in an embodiment of the present application.

[0120] like Figure 5 As shown, the Diffusion action prediction device generated based on the target perspective includes:

[0121] An image acquisition unit 501 is used to acquire a first image corresponding to each initial viewing angle of a plurality of initial viewing angles;

[0122] An information acquisition unit 502 is used to reconstruct the first images corresponding to the initial perspectives by using a multi-perspective generation neural network model to acquire three-dimensional scene information;

[0123] A feature acquisition unit 503 is used to uniformly sample the three-dimensional scene information using the multi-view generation neural network model and calculate the multi-view image features of each sampling point in the sampling point set;

[0124] The image acquisition unit 501 is used to generate a second image corresponding to a target perspective using the multi-view generation neural network model according to the multi-view image features of each sampling point;

[0125] The action prediction unit 504 is used to use the diffusion strategy Diffusion Policy and the second image to predict the action of the robot and determine the action of the embodied intelligence of the robot.

[0126] Further, in a possible implementation of the embodiment of the present application, the three-dimensional Gaussian reconstruction neural network model includes an RGB image feature extraction neural network sub-model and a Gaussian point cloud reconstruction neural network sub-model, and the information acquisition unit 502 is used to use the three-dimensional Gaussian reconstruction neural network model to reconstruct the first image corresponding to each initial perspective, and when acquiring the three-dimensional scene information, it is specifically used to:

[0127] Using the RGB image feature extraction neural network sub-model to extract features from the first image corresponding to each initial perspective to obtain an image feature set;

[0128] The Gaussian point cloud reconstruction neural network sub-model is used to reconstruct the scene of the image feature set to obtain three-dimensional scene information.

[0129] Further, in a possible implementation of the embodiment of the present application, the RGB image feature extraction neural network sub-model includes an encoder of a pure visual transformer structure and a decoder of a pure visual transformer structure, and the information acquisition unit 502 is used to use the RGB image feature extraction neural network sub-model to extract features of the first image corresponding to each initial perspective, and when acquiring the image feature set, it is specifically used to:

[0130] According to first images corresponding to any two initial viewing angles of the multiple initial viewing angles and camera device parameters corresponding to the initial viewing angles, an encoder of the pure visual transformer structure is used to perform feature extraction to obtain three-dimensional Gaussian parameters;

[0131] According to the three-dimensional Gaussian parameters, a decoder of the pure visual transformer structure is used to generate an image feature set corresponding to the target viewing angle.

[0132] Furthermore, in a possible implementation of the embodiment of the present application, the information acquisition unit 502 is further specifically configured to:

[0133] Acquire training sample images corresponding to multiple viewing angles and camera device training sample parameters corresponding to the multiple viewing angles, wherein the camera device training sample parameters include the number of viewing angles corresponding to the multiple viewing angles and the intrinsic parameter of the camera device corresponding to each viewing angle of the multiple viewing angles;

[0134] A feedforward network with learnable parameters is obtained by training the camera device, and when the feedforward network meets the model training requirements, the feedforward network is used as the Gaussian point cloud reconstruction neural network sub-model.

[0135] Further, in a possible implementation of the embodiment of the present application, the information acquisition unit 502 is used to obtain a feedforward network with learnable parameters according to the training of the camera device, specifically for:

[0136] Mapping the training sample image to a standard space using a target image in the training sample image to obtain three-dimensional Gaussian distribution information corresponding to the training sample image;

[0137] According to the three-dimensional Gaussian distribution information, mapping relationship information between the training sample image and the three-dimensional Gaussian distribution information is obtained, and the mapping relationship information is used as the feedforward network with learnable parameters.

[0138] Further, in a possible implementation of the embodiment of the present application, the action prediction unit 504 is used to use the diffusion strategy Diffusion Policy and the second image to predict the action of the robot, and determine the action of the embodied intelligence of the robot, specifically for:

[0139] Inputting the second image into a standard temporal Transformer network model to generate visual input information corresponding to the target perspective;

[0140] Obtain random gripper state information, diffusion time step information, and Gaussian distribution noise action information;

[0141] Using the visual input information, the random gripper states information, the diffusion time step information and the noise action information conforming to the Gaussian distribution to perform action prediction, and determine a multi-view prediction image corresponding to the robot;

[0142] The multi-view predicted images are subjected to three-dimensional consistency processing to determine the actions of the robot's embodied intelligence.

[0143] Further, in a possible implementation of the embodiment of the present application, the action prediction unit 504 is used to perform three-dimensional consistency processing on the multi-view prediction image to determine the action of the embodied intelligence of the robot, specifically for:

[0144] Performing three-dimensional consistency processing on the multi-view prediction image to obtain the gripper position of the robot;

[0145] According to the position of the robot's gripper, a change matrix corresponding to the robot's gripper is obtained, and according to the change matrix, an action of the robot's embodied intelligence is determined.

[0146] It should be noted that the aforementioned explanation of the embodiment of the Diffusion action prediction method based on the target perspective generation is also applicable to the Diffusion action prediction device based on the target perspective generation of this embodiment, and will not be repeated here.

[0147] The present application provides a Diffusion action prediction device based on target perspective generation, which is provided with an image acquisition unit for acquiring a first image corresponding to each of a plurality of initial perspectives; an information acquisition unit for reconstructing a scene by using a multi-perspective generation neural network model for the first image corresponding to each of the initial perspectives to acquire three-dimensional scene information; a feature acquisition unit for uniformly sampling the three-dimensional scene information by using the multi-perspective generation neural network model and calculating the multi-view image features of each sampling point in a sampling point set; the image acquisition unit for generating a second image corresponding to the target perspective by using the multi-perspective generation neural network model according to the multi-view image features of each sampling point; and an action prediction unit for using a diffusion strategy Diffusion. Policy and the second image predict the robot's action and determine the robot's embodied intelligent action, which solves the problem of using the robot's posture information and images corresponding to a fixed perspective to train the model, and needing to obtain images corresponding to the fixed perspective when performing action prediction, which makes the action prediction more restricted and the action prediction less convenient. Since the three-dimensional scene information can be reconstructed through the image of the initial perspective and the neural network model, it can be established through the image of the initial perspective when the image of the target perspective cannot be obtained, thereby improving the accuracy of obtaining the image corresponding to the target perspective to generate high-quality robot embodied intelligent actions, reducing the restrictions on the perspective of the camera device, and being able to generate the generalization of the neural network model from multiple perspectives, which can improve the convenience of action prediction, and can predict the actions of the corresponding perspectives through the images of the robot's environment, so that the robot can interact with the environment and improve the accuracy of the interaction.

[0148] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.

[0149] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.

[0150] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.

[0151] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.

[0152] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.

[0153] This application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, this application is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, risks can be minimized by limiting data collection and deleting data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.

[0154] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0155] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0156] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0157] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0158] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0159] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0160] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0161] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A Diffusion action prediction method based on target perspective generation, characterized in that: include: Acquire a first image corresponding to each initial viewing angle of a plurality of initial viewing angles; Using a multi-view generation neural network model to reconstruct the scene of the first image corresponding to each initial view to obtain three-dimensional scene information; The multi-view generation neural network model is used to uniformly sample the three-dimensional scene information and calculate the multi-view image features of each sampling point in the sampling point set; According to the multi-view image features of each sampling point, the multi-view generation neural network model is used to generate a second image corresponding to the target view; The diffusion strategy Diffusion Policy and the second image are used to predict the action of the robot to determine the action of the embodied intelligence of the robot.

2. The method according to claim 1, characterized in that in, The three-dimensional Gaussian reconstruction neural network model includes an RGB image feature extraction neural network sub-model and a Gaussian point cloud reconstruction neural network sub-model. The three-dimensional Gaussian reconstruction neural network model is used to reconstruct the first image corresponding to each initial perspective to obtain three-dimensional scene information, including: Using the RGB image feature extraction neural network sub-model to extract features from the first image corresponding to each initial perspective to obtain an image feature set; The Gaussian point cloud reconstruction neural network sub-model is used to reconstruct the scene of the image feature set to obtain three-dimensional scene information.

3. The method according to claim 2, characterized in that in, The RGB image feature extraction neural network sub-model includes an encoder with a pure visual transformer structure and a decoder with a pure visual transformer structure. The RGB image feature extraction neural network sub-model is used to extract features of the first image corresponding to each initial perspective to obtain an image feature set, including: According to first images corresponding to any two initial viewing angles of the multiple initial viewing angles and camera device parameters corresponding to the initial viewing angles, an encoder of the pure visual transformer structure is used to perform feature extraction to obtain three-dimensional Gaussian parameters; According to the three-dimensional Gaussian parameters, a decoder of the pure visual transformer structure is used to generate an image feature set corresponding to the target viewing angle.

4. The method according to claim 2, characterized in that: The method further comprises: Acquire training sample images corresponding to multiple viewing angles and camera device training sample parameters corresponding to the multiple viewing angles, wherein the camera device training sample parameters include the number of viewing angles corresponding to the multiple viewing angles and the intrinsic parameter of the camera device corresponding to each viewing angle of the multiple viewing angles; A feedforward network with learnable parameters is obtained by training the camera device, and when the feedforward network meets the model training requirements, the feedforward network is used as the Gaussian point cloud reconstruction neural network sub-model.

5. The method according to claim 4, characterized in that The feedforward network with learnable parameters obtained by training the camera device includes: Mapping the training sample image to a standard space using a target image in the training sample image to obtain three-dimensional Gaussian distribution information corresponding to the training sample image; According to the three-dimensional Gaussian distribution information, mapping relationship information between the training sample image and the three-dimensional Gaussian distribution information is obtained, and the mapping relationship information is used as the feedforward network with learnable parameters.

6. The method according to claim 1, characterized in that The adopting the diffusion policy and the second image to predict the action of the robot and determine the action of the embodied intelligence of the robot includes: Inputting the second image into a standard temporal Transformer network model to generate visual input information corresponding to the target perspective; Obtain random gripper state information, diffusion time step information, and Gaussian distribution noise action information; Using the visual input information, the random gripper states information, the diffusion time step information and the noise action information conforming to the Gaussian distribution to perform action prediction, and determine a multi-view prediction image corresponding to the robot; The multi-view predicted images are subjected to three-dimensional consistency processing to determine the actions of the robot's embodied intelligence.

7. The method according to claim 6, characterized in that The performing three-dimensional consistency processing on the multi-view prediction images to determine the embodied intelligent action of the robot includes: Performing three-dimensional consistency processing on the multi-view prediction image to obtain the gripper position of the robot; According to the position of the robot's gripper, a change matrix corresponding to the robot's gripper is obtained, and according to the change matrix, an action of the robot's embodied intelligence is determined.

8. A diffusion action prediction device based on target perspective generation, characterized in that: include: An image acquisition unit, used to acquire a first image corresponding to each initial viewing angle of a plurality of initial viewing angles; An information acquisition unit, configured to reconstruct the scene of the first image corresponding to each initial perspective by using a multi-perspective generation neural network model to acquire three-dimensional scene information; A feature acquisition unit, configured to uniformly sample the three-dimensional scene information using the multi-view generation neural network model and calculate multi-view image features of each sampling point in a sampling point set; The image acquisition unit is used to generate a second image corresponding to a target perspective using the multi-view generation neural network model according to the multi-view image features of each sampling point; The action prediction unit is used to use the diffusion strategy Diffusion Policy and the second image to predict the action of the robot and determine the action of the embodied intelligence of the robot.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.