Hierarchical robot operation strategy generation method, device and equipment
By using a cyclic consistent variational autoencoder and generative new perspective image completion, the problems of incomplete robot environmental understanding and excessive computational resource consumption are solved, thereby improving the robot's ability to perform long-term tasks in complex environments.
Patent Information
- Application Number
- CN202511414494.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-29
Smart Images

Figure CN120886274A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of robot intelligent control and artificial intelligence, and in particular to a layered robot operation strategy generation method, device and equipment. BACKGROUND
[0002] With the continuous expansion of the application field of robots, multi-task robot operation based on natural language instructions has become an important direction of agent research. In long-time sequence, multi-stage operation tasks, robots not only need to understand task semantics, but also need to accurately perceive environmental changes and respond flexibly. However, there are two main challenges in current robot strategy generation methods: First, existing perception means are usually based on a single perspective, resulting in incomplete understanding of the environment, especially under conditions of occlusion and local observation, lacking the ability to reason about unobserved areas; Second, existing strategy learning methods usually fail to fully distinguish semantic features from spatial geometric features when modeling features, resulting in mutual interference between sub-task discrimination and action generation, affecting overall generalization ability and execution stability.
[0003] Some studies attempt to introduce large-scale generative models to enhance environmental perception, but due to the presence of artifacts or distortion in generated images, direct use may lead to strategy degradation. At the same time, large generative models have high inference overhead, making it difficult to meet real-time control requirements. Therefore, it is of great research value and application prospect to study a new robot operation strategy generation method that combines generative completion, feature decoupling, cross-domain adaptation, and inference optimization. SUMMARY
[0004] In view of the deficiencies in the prior art described above, the present application provides a layered robot operation strategy generation method, device and equipment, which can effectively improve the long-time sequence task execution capability of robots based on language instructions in complex environments.
[0005] To achieve the above purpose, the present application provides a layered robot operation strategy generation method, comprising the following steps: Step 1, collect the hand-eye camera image and external camera image of the robot, and obtain the natural language instruction; Step 2, take the hand-eye camera image and external camera image collected at the current time as image observation, and input the image observation into a recurrent consistency variational autoencoder for feature decoupling to obtain semantic features and spatial features; Step 3, determine whether a new view hand-eye camera image needs to be generated: If yes, generate a new view hand-eye camera image, and perform multi-view fusion on the new view hand-eye camera image and the image observation, update the semantic features, and then proceed to step 4; Otherwise, proceed to step 4; Step 4, generating a robot skill category based on the semantic feature and the natural language instruction, and generating a robot action based on the spatial feature and the robot skill category.
[0006] Compared with the prior art, the present application has the following beneficial technical effects: The present application uses a cyclic consistency variational autoencoder to decouple image observation features, and then optimizes hierarchical strategy reasoning, thereby improving sub-skill discrimination accuracy and action generation stability. In addition, the generated new perspective hand-eye camera image enhances and completes environmental observation, improving the three-dimensional perception ability of the robot in a complex environment. The success rate and reliability of the robot in performing long-time sequence and multi-task operations in a complex and changing environment can be significantly improved, and the present application has broad application prospects. BRIEF DESCRIPTION OF DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the structures shown in the drawings without creative labor.
[0008] Figure 1 The flow chart of the hierarchical robot operation strategy generation method in the embodiments of the present application; Figure 2 The structural block diagram of the hierarchical robot operation strategy generation device in the embodiments of the present application; Figure 3 The structural block diagram of the terminal device in the embodiments of the present application.
[0009] The implementation of the present application, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0010] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0011] In addition, the technical solutions of each embodiment of the present application can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it. When the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor is it within the scope of protection required by the present application.
[0012] Embodiment 1 As Figure 1 A hierarchical robot operation strategy generation method is disclosed in the embodiment, which mainly includes the following steps: Step 1, based on the camera carried by the robot and the externally fixed camera, the hand-eye camera image and the external camera image of the robot are collected, and the natural language instruction is obtained; Step 2, the hand-eye camera image and the external camera image collected at the current time are taken as image observation, and the image observation is input into the recurrent consistency variational autoencoder for feature decoupling to obtain semantic features and spatial features; Step 3, judge whether a new view hand-eye camera image needs to be generated at the current time: If yes, generate a new view hand-eye camera image, and perform multi-view fusion on the new view hand-eye camera image and the image observation, update the semantic features, and then proceed to step 4; Otherwise, proceed to step 4; Step 4, generate a robot skill category based on the semantic features and the natural language instruction, and generate a robot action based on the spatial features and the robot skill category.
[0013] In the embodiment, the generation process of the semantic features and the spatial features is: extracting the semantic feature vector and the spatial feature vector of the hand-eye camera image, and extracting the semantic feature vector and the spatial feature vector of the external camera image; splicing the semantic feature vectors of the hand-eye camera image and the external camera image to obtain the semantic features; and splicing the spatial feature vectors of the hand-eye camera image and the external camera image to obtain the spatial features.
[0014] Let the image observation be which contains information from visible light sensors (i.e. cameras), such as object color, background texture, scene layout, etc. Assuming that the robot needs to learn multiple skills or operate in different tasks in the same environment, a simple two-dimensional convolution or Transformer feature extraction on s can obtain certain semantic and geometric representations, but its internal usually does not explicitly distinguish between "high-level semantic abstraction" and "low-level spatiotemporal details". Once it needs to adapt to different types of downstream tasks in the strategy network, it will often happen that some sub-tasks are over-fitted and some sub-tasks are not well generalized. To this end, the embodiment introduces a recurrent consistency constraint through a recurrent consistency variational autoencoder, which implicitly "cleans up" the segmentation of latent variables, forming complementary semantic and spatial subspaces.
[0015] Let represent the complete latent variable, where the semantic features focus on expressing semantic information of the scene, such as object categories, object distribution relationships in the scene, etc.; and the spatial features focus on the specific spatial geometric details such as the geometric pose, position coordinates, occlusion relationship or grasping angle of the object. The traditional variational autoencoder only provides a latent vector , ignoring the need for independent modeling of semantics and geometry. The cyclic consistency variational autoencoder in the embodiment further makes the following assumptions , where is the joint prior distribution, is the independent prior of semantic features, is the independent prior of spatial features, the above assumptions represent that the semantic features and the spatial features are independent of each other in the prior distribution. For any image observation , its posterior distribution is also correspondingly split into , where , correspond to the parameters of the semantic encoder and the geometric encoder, respectively.
[0016] The cyclic consistency variational autoencoder in the embodiment, in the training process, while forcing the network to learn a useful representation through the reconstruction loss, also auxiliary adds a cyclic consistency constraint to ensure the correctness and separation of semantics and geometry. Specifically, the training of the cyclic consistency variational autoencoder is realized through a forward process and a backward process. In order to strengthen the decoupling effect under the same distribution scene, the operation of latent variable exchange is adopted, that is, given two image observations , sampled under the same distribution scene, respectively encode to obtain and , and then exchange their semantic features at the decoding end to check whether the generated reconstructed image still maintains the correct semantic and spatial geometric structure. If the two image observations come from the same task scene, but the specific positions of the objects have slight differences, but the high-level semantic distribution relationship between the objects remains unchanged, then the semantic feature information obtained after feature space decoupling should be consistent. After feature exchange, an image reconstruction with the same semantic layout information and mixed specific spatial geometric layout should be obtained. At this time, if the model decoupling is sufficient, there will be no large error in the cyclic consistency constraint after the reconstructed image is encoded again, indicating that the semantic subspace and the spatial subspace are indeed effectively isolated in the functional partition, improving the learning efficiency under multiple skills and multiple scenes.
[0017] In the specific implementation process, the training process of the cyclic consistency variational autoencoder is as follows: obtain two image observations , under the same scene, and image observations under different scenes; encode the image observations , 、 Input cycle-consistent variational autoencoder, obtain semantic features of image observation and spatial features , semantic features of image observation and spatial features , semantic features of image observation and spatial features , and semantic features of image observation and spatial features ; ; After combining semantic features and spatial features , the reconstructed image observation is decoded, and after combining spatial features and semantic features , the reconstructed image observation is decoded, and based on the reconstructed image observation 、 , that is: 、 wherein, is a probability distribution based on decoding , is sampled from , and the like; Since image observations 、 come from the same task scene, exchanging their semantic features should not change the layout relationship in the scene, that is, the image reconstructed after the cross-state semantic feature exchange of the image should still belong to the same semantic category, that is, the same object type and scene structure, and the specific position or pose depends on the spatial feature subspace, from which the forward process loss can be calculated as ; wherein, is a hyperparameter for balancing the reconstruction error and the prior constraint, is the reconstruction error after cross-reconstruction, is the prior constraint, so that the latent variable distribution in the semantic subspace and the spatial subspace does not deviate too much from the standard normal distribution, thereby preventing the model from concentrating too much information in a single subspace; Next, the reverse process loss is calculated using cycle consistency, a semantic feature is randomly sampled, spatial features and semantic features are combined to decode the reconstructed image observation , and spatial features and semantic features The reconstructed image observation is decoded ; The reconstructed image observation is decoded 、 The reconstructed image observation is decoded The semantic features of the reconstructed image observation The spatial features of the reconstructed image observation The semantic features of the reconstructed image observation The spatial features of the reconstructed image observation The semantic features of the reconstructed image observation The spatial features of the reconstructed image observation The semantic features of the reconstructed image observation The spatial features of the reconstructed image observation ; Wherein, p is the decoding process, q is the encoding process; by minimizing the above reverse process loss, the network will be difficult to utilize the information of another subspace to transfer the information of a certain subspace, thereby forcing the semantic encoder and the geometric encoder to truly focus on the information type corresponding to itself, forming a clear semantic-geometric decoupling structure; is the encoding process only leaving the semantic features, and the reverse loss process only calculates the difference between the reconstructed semantic features; Finally, the total loss is calculated based on the forward loss and the reverse loss , and the network parameters of the cycle-consistent variational autoencoder are updated based on the total loss.
[0018] In the visual-based hierarchical skill learning task, the robot needs to rely on visual information for environment perception and task decision. By introducing a pre-trained visual large model to generate images of new perspectives, the robot can obtain more environmental information, thereby improving the accuracy of robot high-level skill discrimination. However, due to the high computational cost of the pre-trained large model, if new perspective generation is introduced at all times, it will lead to excessive consumption of computing resources, and may also introduce too much redundant information. Therefore, a reasonable triggering strategy needs to be designed to introduce new perspective generation only at key moments, so as to balance the computational efficiency and perception integrity. In view of this, the embodiment proposes a new perspective triggering strategy based on inter-frame mutual information, which detects the moment when the visual information uncertainty is high, and then introduces the pre-trained large model for completion. That is, in the specific implementation process of step 3, whether to generate a new perspective hand-eye camera image is judged as follows: First, the mutual information between the current frame hand-eye camera image and the last frame hand-eye camera image is calculated ; Second, the average value of the hand-eye camera image of the previous frame is calculated , determine whether it is true, if yes, the current needs to generate a new perspective hand-eye camera image, otherwise it does not need, wherein is a hyperparameter.
[0019] Mutual information is a measure of the correlation between two random variables, which can be used to evaluate the degree of information redundancy between the current frame and the previous frame. Assuming the joint probability distribution of the current frame and the previous frame is , the marginal probability distribution is and , the mutual information between the two frames is defined as: ; The mutual information measure reflects the degree of information sharing between the current frame and the previous frame. If the mutual information between the two frames is high, it means can obtain a lot of information from , so there is no need to introduce new perspective completion. However, if the mutual information between the two frames is low, it means has a larger new information increment compared to , which may be due to a large change in the robot's perspective or due to the occlusion of the key target area, resulting in the loss of some information. In this case, new perspective generation should be triggered to complete the missing information.
[0020] Since the inter-frame mutual information varies greatly between different task scenarios, this embodiment uses a dynamic mutual information threshold method for calculation. For example, review the last 10 hand-eye camera images, calculate the average value and the standard deviation of the last 10 hand-eye camera images, and use the hyperparameter to control the threshold value of new perspective generation, which can be set to . If the mutual information of the current frame is lower than the dynamic mutual information threshold, it is considered that the visual information of the current frame has changed significantly, thereby triggering new perspective generation.
[0021] In this embodiment, a pre-trained GenWarp model is used to generate a new perspective hand-eye camera image, and the spatial relationship between the robot and the target is combined during the generation process to adaptively select the perspective, improving the semantic consistency and geometric structure rationality of the new perspective hand-eye camera image in the robot operation process. Specifically: Assuming the maximum available perspective of the target object is at a distance of , when the perspective , the projection deformation of the object will exceed the acceptable threshold , thereby affecting the semantic consistency. Define the structural similarity measure between the generated image and the real perspective image and assume that the relationship between the view angle and the camera-object distance obeys the following function: ; where is a distance-dependent correction factor, is a parameter to control the influence of view angle change on structural similarity, is the angular offset. Starting from the perspective projection transformation, as the rotation angle increases, the projected shape of the target object changes exponentially, therefore an exponential decay function is adopted to describe the similarity degradation trend. In the case of close distance, the object projection is large, even if a large view angle change occurs, the image can still maintain a high similarity, therefore is approximately constant. While in the case of far distance, the object projection is small, the background information occupies a large proportion, leading to significant shape change of the image when the view angle changes, therefore should decrease with . Through experimental observation, we have: ; where is the distance decay coefficient, then the structural similarity function can be expressed as: ; Set the similarity threshold , the optimal angle for new view generation is such that: ; After calculation, we have: ; This formula shows that the optimal angle for new view generation decreases with distance , indicating that in the case of far distance, even a small view angle change can lead to significant semantic distortion, and the azimuth angle needs to be selected more carefully, while in the case of close distance, a larger view angle change can be allowed. This conclusion provides a constraint criterion for new view generation, making the generated image have better semantic consistency in the largest possible range.
[0022] In the specific implementation of step 3, the multi-view fusion process of the novel perspective hand-eye camera image and image observation is as follows: First, the novel perspective hand-eye camera image is input into a cyclic consistency variational autoencoder for feature decoupling, obtaining the semantic feature vector and spatial feature vector of the novel perspective hand-eye camera image; then, the semantic feature vector of the novel perspective hand-eye camera image is concatenated with the semantic features obtained in step 2 to obtain the updated semantic features. Since the novel perspective image has slight geometric distortion but accurate semantic information, only high-level semantic features are retained for high-level skill discrimination to avoid the slight distortion affecting the robot's operation success rate.
[0023] In robot hierarchical skill learning, the concept of basic skills has been proposed and widely applied to effectively transform high-level task instructions into low-level action sequences. Basic skills refer to reusable atomic operation units, such as robot grasping, rotation, and translation, which constitute the basic components of complex tasks. In the specific implementation of step 4, the generation process of robot skill categories is as follows: First, the MiniLM language model is used as the encoder to encode natural language instructions into instruction features. The MiniLM language model is distilled from a large-scale Transformer model, and the training corpus covers general language datasets such as Wikipedia. Its output is a 384-dimensional vector, which can map natural language sentences of arbitrary length to a fixed-length semantic vector space, providing stable and compact language input features for the policy model. Then, the semantic features and instruction features are input into the basic skill discriminator to obtain the probability of each skill category of the robot. Specifically, the basic skill discriminator adopts a conditional classifier, for example, given a semantic feature and a natural language instruction... The probability output of the basic skill discriminator can be defined as: ; in, Represents a set of skill categories. It is determined by the parameters of the neural network Implemented feature transformation, The function is used to map the output of the feature transformation to a probability distribution of skill categories; Finally, the skill category with the highest probability is output as the robot's skill category.
[0024] In practice, the process of generating robot actions is as follows: First, spatial features are input into the diffusion generation model, which generates more generalizable geometric spatial features through a stepwise backsampling process. In the training phase, the diffusion generation model introduces geometric spatial feature reconstruction loss from real environment observation samples to reduce the difference in action prediction distribution between virtual and real environments and improve deployment reliability. Then, the geometric space features, instruction features, and robot skill categories are input into a conditioned action generator to generate robot actions.
[0025] The diffusion generative model is a generative model based on stochastic differential equations, which realizes the machine learning reconstruction from noise to data samples step by step by defining a forward diffusion process and a corresponding reverse diffusion process. Specifically, the DDPM method is used to gradually add Gaussian noise to the spatial features . By gradually increasing the noise, the original geometric features are gradually disturbed and finally tend to approach a standard Gaussian distribution. When , the distribution can be regarded as a standard Gaussian distribution. The reverse diffusion process then gradually recovers the original distribution by learning a denoising function through a neural network, which enables the diffusion model to effectively cover more diverse spatial geometric feature representations beyond the initial geometric feature distribution.
[0026] As a preferred embodiment, a domain consistency constraint is introduced in the diffusion training process to improve the continuity and executability of action generation in virtual training environments and actual deployment environments. In the framework of transfer learning, it is assumed that the source domain and the target domain respectively obey different probability distributions and , where represents the observed image. In the virtual-real transfer task, is determined by the image observation data distribution in the virtual environment, and is determined by the image data collected by the sensors or cameras in the real world. Since , the performance of the model trained in the source domain directly used in the target domain will usually decrease. Therefore, a domain adaptation method is needed to enable the model to learn a cross-domain shared feature representation, thereby reducing the impact of distribution bias.
[0027] Consider the domain adaptive encoder , where is the feature embedding space. The goal is to find an optimal mapping such that the feature distributions of the source domain and the target domain are aligned in the latent space, i.e., satisfy: ; To this end, the Maximum Mean Discrepancy (MMD) is used as a measure of the similarity of the feature distributions. MMD calculates the similarity of two distributions by comparing their mean embeddings in the kernel space, which is defined as follows: ; where is a feature mapping function that maps the original features to a high-dimensional reproducing kernel Hilbert space In order to make the data of different distributions comparable in the space, the embodiment adopts a Gaussian kernel function for mapping. If , the MMD value is close to zero, indicating that the feature distributions of the two domains have been highly aligned.
[0028] In order to optimize the domain adaptation capability, a domain adaptation encoder is adopted to extract feature embeddings , wherein represents a low-dimensional feature representation extracted from an input image. During the training process, the data of the source domain and the target domain are respectively input to obtain and , and the MMD loss is minimized. By optimizing the domain adaptation encoder through gradient descent, the distribution difference between the domains can be gradually reduced while maintaining the task performance. In the specific application process, the MMD loss is trained together with the loss function in the training process of the diffusion generation model and the action generator, so as to obtain a transferable robot control strategy.
[0029] The action generator in the embodiment is obtained by improving the unconditional action generator in the SPIL method, that is, introducing a skill category as an additional input to construct a conditional skill generator. The essence is to use the skill representation in the latent space and the explicit skill category label y to jointly generate the robot action matched with the current task requirement, while using the robot action in the teaching data as a supervision signal for training. Specifically, the training process of the action generator is consistent with the conditional variational autoencoder paradigm, the input is the action sequence x and the corresponding soft skill label y from the teaching data, the encoder maps it to the latent space, and the decoder is responsible for the reconstruction of the action sequence. The training objective function is: ; wherein is a category conditional prior distribution, which is used to realize the structured modeling of the latent space; is a prior distribution, and the embodiment takes a standard Gaussian distribution; is a prior constraint, is a structured constraint, is a weight coefficient, which is used to adjust the influence of each regular term. Through the above structured training objective function, the action generator can not only accurately generate robot actions, but also form clear skill category boundaries in the latent space.
[0030] It is worth noting that although each step in the embodiment Figure 1 is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise explicitly stated herein, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps in the embodiment Figure 1 may include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.
[0031] Embodiment 2 Based on the hierarchical robot operation strategy generation method in Embodiment 1, this embodiment discloses a hierarchical robot operation strategy generation device, which refers to Figure 2 The hierarchical robot operation strategy generation device includes an image acquisition unit, an instruction acquisition unit, a feature decoupling unit, a new view generation unit, a multi-view fusion unit, a skill discrimination unit, and an action generation unit, specifically: The image acquisition unit is configured to acquire the hand-eye camera image and the external camera image of the robot; The instruction acquisition unit is configured to acquire the natural language instruction of the robot; The feature decoupling unit is configured to take the hand-eye camera image and the external camera image acquired at the current time as image observation, and input the image observation into the recurrent consistency variational autoencoder for feature decoupling to obtain semantic features and spatial features; The new view generation unit is configured to generate a new view hand-eye camera image; The multi-view fusion unit is configured to perform multi-view fusion on the new view hand-eye camera image and the image observation to update the semantic features; The skill discrimination unit is configured to input the semantic features and the natural language instruction into the basic skill discriminator to generate a robot skill category; The action generation unit is configured to input the spatial features and the robot skill category into the action generator to generate a robot action.
[0032] In this embodiment, the specific working processes and working principles of the image acquisition unit, the instruction acquisition unit, the feature decoupling unit, the new view generation unit, the multi-view fusion unit, the skill discrimination unit and the action generation unit are the same as those in the method of embodiment 1, and therefore the above will not be described herein. The various unit modules can be realized by software, hardware and combinations thereof in whole or in part, and the various unit modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the various unit modules.
[0033] Embodiment 3 As shown in Figure 3 a terminal device disclosed in this embodiment, comprising a transmitter, a receiver, a memory and a processor. Among them, the transmitter is used to send instructions and data, the receiver is used to receive instructions and data, the memory is used to store computer execution instructions, and the processor is used to execute the computer execution instructions stored in the memory to realize the method in the above embodiment 1.
[0034] It should be noted that the above memory can be independent or integrated with the processor. When the memory is independently arranged, the terminal device further comprises a bus for connecting the memory and the processor.
[0035] The above is only the preferred embodiment of the present application, and does not limit the protection scope of the present application. Any equivalent structural transformation made under the inventive concept of the present application, or direct / indirect application in other related technical fields is included in the protection scope of the present application.
Claims
1. A method for generating hierarchical robot operation strategies, characterized in that, Includes the following steps: Step 1: Acquire images from the robot's hand-eye camera and external camera, and obtain natural language commands; Step 2: The hand-eye camera image and the external camera image acquired at the current moment are used as image observations, and the image observations are input into the cyclic consistency variational autoencoder for feature decoupling to obtain semantic features and spatial features; Step 3: Determine whether a new perspective hand-eye camera image needs to be generated: If so, generate a new perspective hand-eye camera image, and fuse the new perspective hand-eye camera image with the image observation from multiple perspectives. After updating the semantic features, proceed to step 4. Otherwise, proceed to step 4; Step 4: Generate robot skill categories based on the semantic features and the natural language instructions, and generate robot actions based on the spatial features and the robot skill categories.
2. The hierarchical robot operation strategy generation method according to claim 1, characterized in that, In step 2, the generation process of the semantic features and the spatial features is as follows: Extract the semantic feature vector and spatial feature vector of the hand-eye camera image, and extract the semantic feature vector and spatial feature vector of the external camera image; The semantic feature is obtained by concatenating the semantic feature vectors of the hand-eye camera image and the external camera image; The spatial features are obtained by concatenating the spatial feature vectors of the hand-eye camera image and the external camera image.
3. The hierarchical robot operation strategy generation method according to claim 2, characterized in that, The training process of the recurrent consistency variational autoencoder is as follows: Acquire two frames of images from the same scene , and image observation in different scenarios ; Image observation , , Inputting a cyclic consistent variational autoencoder yields image observations. semantic features Spatial features Image observation semantic features Spatial features and image observation semantic features Spatial features ; semantic features Spatial features After combining and decoding, the reconstructed image observation is obtained. and spatial features With semantic features After combining and decoding, the reconstructed image observation is obtained. And based on reconstructed image observations , Calculate the forward process loss; Randomly sample a semantic feature Spatial features With semantic features After combining and decoding, the reconstructed image observation is obtained. and spatial features With semantic features After combining and decoding, the reconstructed image observation is obtained. ; Reconstructing image observations , Inputting a cyclic consistency variational autoencoder yields reconstructed image observations. semantic features Spatial features and reconstructed image observation semantic features Spatial features And based on semantic features and Calculate the loss in the reverse process; The total loss is calculated based on the forward process loss and the reverse process loss, and the network parameters of the cyclic consistent variational autoencoder are updated based on the total loss.
4. The hierarchical robot operation strategy generation method according to claim 1, 2, or 3, characterized in that, In step 3, determining whether a new perspective hand-eye camera image needs to be generated specifically involves: Calculate the mutual information between the current frame hand-eye camera image and the previous frame hand-eye camera image. ; And calculate before Average value of frame hand-eye camera images with standard deviation ,judge If the condition is met, then a new perspective hand-eye camera image needs to be generated; otherwise, it is not necessary. This is a hyperparameter.
5. The hierarchical robot operation strategy generation method according to claim 1, 2, or 3, characterized in that, In step 3, a pre-trained GenWarp model is used to generate new perspective hand-eye camera images. During the generation process, the perspective is adaptively selected by combining the spatial relationship between the robot and the target, thereby improving the semantic consistency and geometric rationality of the new perspective hand-eye camera images during robot operation.
6. The hierarchical robot operation strategy generation method according to claim 1, 2, or 3, characterized in that, Step 3, specifically the process of multi-view fusion of the new perspective hand-eye camera image and the image observation, is as follows: The new perspective hand-eye camera image is input into the cyclic consistency variational autoencoder for feature decoupling, resulting in the semantic feature vector and spatial feature vector of the new perspective hand-eye camera image; The semantic feature vector of the new perspective hand-eye camera image is concatenated with the semantic features obtained in step 2 to obtain the updated semantic features.
7. The hierarchical robot operation strategy generation method according to claim 1, 2, or 3, characterized in that, In step 4, the process of generating the robot skill category is as follows: First, the MiniLM language model is used as an encoder to encode the natural language instructions into instruction features; Then, the semantic features and the instruction features are input into the basic skill discriminator to obtain the probability of each skill category of the robot, wherein the skill categories include grasping, rotation and translation; Finally, the skill category with the highest probability is output as the robot's skill category.
8. The hierarchical robot operation strategy generation method according to claim 7, characterized in that, In step 4, the process of generating the robot's actions is as follows: First, the spatial features are input into the diffusion generation model, and a more generalizable geometric spatial feature is generated through a stepwise backsampling process; Then, the geometric spatial features, the instruction features, and the robot skill category are input into a conditional motion generator to generate the robot motion.
9. A hierarchical robot operation strategy generation device, characterized in that, The hierarchical robot operation strategy generation device, using the method according to any one of claims 1 to 8, comprises: The image acquisition unit is used to acquire images from the robot's hand-eye camera and external camera. The instruction acquisition unit is used to acquire the robot's natural language instructions; The feature decoupling unit is used to take the hand-eye camera image acquired at the current moment and the external camera image as image observations, and input the image observations into the cyclic consistency variational autoencoder for feature decoupling to obtain semantic features and spatial features. New perspective generation unit, used to generate new perspective hand-eye camera images; A multi-view fusion unit is used to fuse the new viewpoint hand-eye camera image with the image observation from multiple perspectives and update the semantic features; A skill discrimination unit is used to input the semantic features and the natural language instructions into a basic skill discriminator to generate robot skill categories; The motion generation unit is used to input the spatial features and the robot skill category into the motion generator to generate robot motions.
10. A terminal device, characterized in that, The terminal device is equipped with: Memory, used to store programs; A processor for executing the program stored in the memory, wherein when the program is executed, the processor is configured to perform the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Structure retentivity enhancement method for unpaired eye fundus images based on feature decoupling
CN115482177A
Robot control method and system, electronic equipment and storage medium
CN119369412A
Robot visual identification decision control method based on deep learning
CN120244983A
Single-view unknown object 6D pose estimation method based on segmentation and new view angle synthesis
CN120472000A
Domain adaptation using simulation to simulation transfer
US20200167606A1
Cited By
Robot control method and device based on decision state machine and computer equipment
CN121340305A