Robot control method and device, electronic equipment and storage medium

By acquiring the target operation text and initial viewpoint image, a target prediction video is generated using a video generation model, and feature extraction and inverse dynamics modeling are performed. This solves the problem of poor robot control accuracy in existing technologies and achieves more efficient dynamic visual change capture and control.

CN120953887APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511104667.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing robot control strategies struggle to effectively capture key information about dynamic visual changes, resulting in poor control accuracy.

Method used

By acquiring the target operation text and initial viewpoint image, a video generation model is used to predict image changes, generate a target prediction video, and perform feature extraction and inverse dynamics modeling to control the robot to execute the predicted actions.

Benefits of technology

It improves the accuracy and efficiency of robot control and enhances the generalization ability of control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953887A_ABST
    Figure CN120953887A_ABST
Patent Text Reader

Abstract

The invention provides a robot control method and device, electronic equipment and a storage medium, relates to the technical field of robot control, and is suitable for the fields of financial science and technology and medical health. The method comprises the following steps: acquiring a target operation text, and performing text coding on the target operation text to obtain a target operation text feature; obtaining an initial visual angle image, and carrying out image coding on the initial visual angle image to obtain initial image features; image change prediction is carried out according to the target operation text features and the initial image features, a target prediction video is obtained, and the target prediction video comprises at least two prediction images; performing feature extraction according to the target prediction video to obtain a video spatial-temporal feature sequence; performing inverse dynamic modeling according to the video spatial-temporal feature sequence and the target operation text features to obtain a target prediction action of the prediction image; and controlling the target robot to execute the target prediction action. The key information of dynamic visual change can be effectively captured, and the control accuracy of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and is applicable to the fields of financial technology and healthcare. In particular, it relates to a robot control method and device, electronic device, and storage medium. Background Technology

[0002] Robots can interpret user-input text to understand the user's intent and then perform corresponding actions. For example, in the fintech field, in on-site inspection and risk control scenarios, inspection robots can use visual recognition technology to patrol the premises of banks, securities institutions, and other businesses, checking the status of facilities such as door locks and fire-fighting equipment to ensure safety and compliance. As another example, in the healthcare field, in drug management scenarios, medication dispensing robots use visual recognition technology to accurately locate the corresponding medication, retrieve the bottle or package from the storage shelf, and then deliver the medication to the patient or a designated location.

[0003] In related technologies, the visual encoders in robot control strategies mainly rely on single-image reconstruction or dual-image contrast learning pre-training methods. These methods focus on the extraction of static visual features and are difficult to effectively capture key information of dynamic visual transformations, resulting in poor robot control accuracy. Summary of the Invention

[0004] The main objective of this application is to propose a robot control method and device, electronic device, and storage medium that can effectively capture key information of dynamic visual changes and improve the accuracy of robot control.

[0005] To achieve the above objectives, a first aspect of this application provides a robot control method, the method comprising:

[0006] Obtain the target operation text used to control the target robot, and encode the target operation text to obtain the target operation text features;

[0007] Acquire an initial viewpoint image of the target robot, and encode the initial viewpoint image to obtain initial image features;

[0008] Image change prediction is performed based on the target operation text features and the initial image features to obtain a target prediction video; wherein, the target prediction video includes at least two prediction images;

[0009] Feature extraction is performed on the target predicted video to obtain a video spatiotemporal feature sequence;

[0010] Inverse dynamics modeling is performed based on the video spatiotemporal feature sequence and the target operation text features to obtain the target predicted action for each predicted image;

[0011] Control the target robot to perform the target predicted action.

[0012] Optionally, the step of predicting image changes based on the target operation text features and the initial image features to obtain the target prediction video includes:

[0013] Obtain a video generation model; wherein the video generation model includes a text and image attention layer, at least two cascaded feature extractors, and at least two cascaded feature decoders, wherein the at least two feature decoders include a first feature decoder and a second feature decoder;

[0014] The target text-image fusion feature is obtained by performing attention processing on the target operation text feature and the initial image feature through the image-text attention layer.

[0015] The target image-text fusion features are extracted by at least two cascaded feature extractors to obtain a predicted image feature sequence;

[0016] The predicted image features are decoded using the first feature decoder to obtain the first image decoded feature sequence;

[0017] The second image decoding feature sequence is obtained by performing feature decoding on the first image decoding feature sequence using the second feature decoder;

[0018] The target predicted video is obtained by reconstructing the image from the second image decoding feature sequence.

[0019] Optionally, the step of extracting features from the target predicted video to obtain video spatiotemporal features includes:

[0020] The target prediction video is video encoded to obtain an image encoded feature sequence;

[0021] The first fused image feature sequence is obtained by concatenating the channels according to the image encoding feature sequence and the second image decoding feature sequence.

[0022] The video spatiotemporal features are obtained by concatenating channels based on the first fused image feature sequence and the first image decoding feature sequence.

[0023] Optionally, obtaining the video generation model includes:

[0024] Acquire first sample data, which includes first sample operation text, first sample initial image, and reference image;

[0025] The first sample operation text is text encoded to obtain the first sample operation text features, and the first sample initial image is image encoded to obtain the first sample initial image features;

[0026] The reference image is noise-added to obtain a noisy image, and the noisy image is image-encoded to obtain noise image features;

[0027] The initial video generation model is used to generate a video from the features of the noisy image, the features of the first sample operation text, and the features of the first sample initial image to obtain a sample prediction image.

[0028] Loss calculation is performed based on the sample predicted image and the reference image to obtain video generation loss data;

[0029] The initial video generation model is adjusted based on the video generation loss data to obtain the video generation model.

[0030] Optionally, the step of performing inverse dynamics modeling based on the video spatiotemporal features and target operation text features to obtain the target prediction action sequence for each predicted image includes:

[0031] The video spatiotemporal features from at least two target perspectives are spliced ​​together to obtain multi-view video spatiotemporal features;

[0032] Spatiotemporal attention processing is performed on the spatiotemporal features of the multi-view video to obtain the target prediction action feature sequence;

[0033] The target predicted action is generated by using a pre-trained action generation model to generate actions from the target predicted action feature sequence and the target operation text features, thereby obtaining the target predicted action for each predicted image.

[0034] Optionally, the step of performing spatiotemporal attention processing on the spatiotemporal features of the multi-view video to obtain the target prediction action feature sequence includes:

[0035] Spatial attention processing is performed on the spatiotemporal features of the multi-view video to obtain key spatial features of the multi-view video.

[0036] Temporal attention processing is performed on the key spatial features of the multi-view video to obtain key spatiotemporal features of the multi-view video.

[0037] The key spatiotemporal features of the multi-view video are processed by a preset feedforward network to obtain the target prediction action feature sequence.

[0038] Optionally, before generating the target predicted action feature sequence using a pre-trained action generation model to obtain the target predicted action for each predicted image, the method further includes:

[0039] Pre-training the action generation model specifically includes:

[0040] Acquire second sample data, which includes second sample operation text, second sample initial image, and reference action sequence;

[0041] The second sample operation text is text encoded to obtain the second sample operation text features, and the second sample initial image is image encoded to obtain the second sample initial image features;

[0042] Based on the text features of the second sample operation and the initial image features of the second sample, image change prediction is performed to obtain the sample prediction video;

[0043] Feature extraction is performed on the sample predicted video to obtain the sample video spatiotemporal feature sequence, and attention processing is performed on the sample video spatiotemporal feature sequence to obtain the sample predicted action feature sequence.

[0044] Noise is added to the reference action sequence to obtain the sample noise action feature sequence;

[0045] The sample noise action feature sequence, the second sample operation text feature, and the sample predicted action feature sequence are denoised using a preset initial action generation model to obtain a sample denoised action sequence.

[0046] Loss calculation is performed based on the reference action sequence and the sample denoised action sequence to obtain action generation loss data;

[0047] The parameters of the initial action generation model are adjusted based on the action generation loss data to obtain the action generation model.

[0048] To achieve the above objectives, a second aspect of this application provides a robot control device, the device comprising:

[0049] The text encoding module is used to acquire the target operation text for controlling the target robot, and to encode the target operation text to obtain the target operation text features.

[0050] The image encoding module is used to acquire the initial view image of the target robot, and to encode the initial view image to obtain initial image features;

[0051] A video generation module is used to predict image changes based on the target operation text features and the initial image features to obtain a target prediction video; wherein the target prediction video includes at least two prediction images;

[0052] The feature extraction module is used to extract features from the target predicted video to obtain a video spatiotemporal feature sequence;

[0053] The action prediction module is used to perform inverse dynamics modeling based on the video spatiotemporal feature sequence and target operation text features to obtain the target predicted action for each predicted image.

[0054] The action execution module is used to control the target robot to perform the target predicted action.

[0055] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the robot control method described in the first aspect.

[0056] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the robot control method described in the first aspect.

[0057] The robot control method, apparatus, electronic device, and storage medium proposed in this application, after acquiring the target operation text, also acquire an initial viewpoint image of the target robot, and then generate a target prediction video based on the target operation text and the initial viewpoint image. This target prediction video is used to represent the dynamic visual change process of the target robot under the control of the target operation text. Furthermore, the target prediction action of the target robot can be deduced from the target prediction video, and finally, the target robot is controlled to execute the target prediction action. In this way, this application can effectively capture key information of dynamic visual changes and improve the accuracy of robot control. Attached Figure Description

[0058] Figure 1 This is a flowchart of the robot control method provided in the embodiments of this application;

[0059] Figure 2 yes Figure 1 The flowchart for step 103 in the text;

[0060] Figure 3 yes Figure 2 The flowchart for step 201 in the document;

[0061] Figure 4 yes Figure 1 The flowchart for step 104 in the document;

[0062] Figure 5 yes Figure 1 The flowchart for step 105 in the document;

[0063] Figure 6 yes Figure 5 Another flowchart of step 502 in the process;

[0064] Figure 7 This is a block diagram of the module structure of the robot control device provided in the embodiments of this application;

[0065] Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0067] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0069] First, let's analyze some of the terms used in this application:

[0070] Robots: Generally refers to mechanical devices with a certain degree of autonomy or semi-autonomy, capable of performing various tasks and imitating or surpassing certain human behaviors. Robots typically consist of mechanical parts, sensors, control systems, and software, enabling them to perceive their environment, make decisions, and execute actions.

[0071] Video Diffusion Model (VDM) is a generative model based on a diffusion process, designed to generate high-quality, continuous video content. It combines the powerful generative capabilities of diffusion models with the continuity of time series, generating realistic dynamic images while maintaining video coherence and rich detail. The core idea of ​​the diffusion model is to train a model by progressively adding noise, and then progressively removing the noise using the reverse process, thereby generating data. In the video diffusion model, the process involves not only the spatial dimension (the relationship between pixels and frames) but also the temporal dimension (the continuity and motion information of video frames). The model typically includes: a forward diffusion process: progressively adding noise to the real video; and a reverse generation process: the trained model progressively removing noise to generate continuous video.

[0072] Stable Video Diffusion Model (SVDM) is a video generation technique based on a diffusion model, designed to create high-quality, continuous, and realistic video content. It leverages the highly successful image diffusion model concept (such as Stable Diffusion) and extends it to the video domain to achieve more stable and coherent dynamic scene generation.

[0073] A visual encoder is a model module in deep learning specifically designed to transform images, videos, or other visual data into abstract features or representations that can be understood by the model. Its main task is to extract the visual information from the input and convert it into compact, discriminative feature vectors, providing the foundation for subsequent tasks such as classification, detection, generation, and understanding.

[0074] A text encoder is a model module in deep learning designed to convert natural language text into digital representations (usually vectors or features) that machines can understand and process. By understanding the semantic, structural, and contextual information in the text, a text encoder transforms sentences or paragraphs into fixed-length or variable-length feature representations for various tasks such as classification, generation, and understanding.

[0075] Bilinear interpolation is a commonly used image scaling and interpolation method used to estimate the value of an unknown point in two-dimensional space based on known discrete point data. It is widely used in image processing, computer graphics, and digital signal processing, especially for smoothing pixel values ​​during image scaling, rotation, and transformation.

[0076] Currently, visual encoders in robot control strategies primarily rely on single-image reconstruction or dual-image contrast learning pre-training methods. These methods focus on extracting static visual features and struggle to effectively capture crucial information about dynamic physical evolution. While video diffusion models demonstrate the ability to predict future frames, existing methods typically depend on a complete multi-step denoising process to generate high-resolution videos, resulting in enormous computational overhead and limited control frequency. Furthermore, related technologies lack effective feature aggregation mechanisms when utilizing the internal representation of video models, failing to efficiently combine predicted future dynamic information with real-time control strategies, thus limiting the generalization ability and execution efficiency of the strategies.

[0077] Based on this, embodiments of this application propose a robot control method, a robot control device, an electronic device, and a computer-readable storage medium. By effectively capturing key information about dynamic visual changes, the accuracy of robot control is improved. This application also efficiently combines predicted future dynamic information with real-time control strategies, improving the generalization ability of the control strategy and the execution efficiency of the robot.

[0078] The robot control method provided in this application can be applied to terminals and servers, or it can be software running on the server. The server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or it can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the robot control method, but it is not limited to the above forms.

[0079] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0080] This application provides a robot control method, a robot control device, an electronic device, and a computer-readable storage medium. The specific embodiments are described below. First, the robot control method in the embodiments of this application is described.

[0081] It should be noted that in each specific embodiment of this application, when it is necessary to process data related to user identity or characteristics, such as operation text data input by the user, the user's permission or consent will be obtained first. Moreover, the collection, use and processing of this data will comply with relevant laws, regulations and standards.

[0082] Reference Figure 1 , Figure 1This is an optional flowchart of the robot control method provided in the embodiments of this application, which may include, but is not limited to, steps 101 to 106.

[0083] Step 101: Obtain the target operation text used to control the target robot, and encode the target operation text to obtain the target operation text features;

[0084] Step 102: Obtain the initial view image of the target robot, and perform image encoding on the initial view image to obtain the initial image features;

[0085] Step 103: Based on the target operation text features and the initial image features, perform image change prediction to obtain the target prediction video; wherein, the target prediction video includes at least two prediction images;

[0086] Step 104: Extract features from the target predicted video to obtain the video spatiotemporal feature sequence;

[0087] Step 105: Perform inverse dynamics modeling based on the video spatiotemporal feature sequence and target operation text features to obtain the target predicted action for each predicted image;

[0088] Step 106: Control the target robot to perform the target prediction action.

[0089] In steps 101 to 106 of this embodiment, after acquiring the target operation text, an initial viewpoint image of the target robot is also acquired. A target prediction video is then generated based on the target operation text and the initial viewpoint image. This target prediction video represents the dynamic visual changes of the target robot under the control of the target operation text. Furthermore, the target prediction action of the target robot can be deduced from the target prediction video, and finally, the target robot is controlled to execute the target prediction action. Thus, this application can effectively capture key information about dynamic visual changes and improve the accuracy of robot control.

[0090] For example, in the field of fintech, the target robot in on-site inspection and risk control scenarios is the patrol robot. This patrol robot can use visual recognition technology to inspect the premises of banks, securities institutions, and other businesses, checking the status of facilities such as door locks and fire-fighting equipment to ensure safety and compliance.

[0091] For example, in the context of drug management in the healthcare field, the target robot is a medication retrieval robot. This robot uses visual recognition technology to accurately locate the corresponding medication, retrieve the bottle or package from the storage shelf, and then deliver the medication to the patient or a designated location.

[0092] In step 101 of some embodiments, target operation text for controlling the target robot is obtained, and the target operation text is text-encoded to obtain target operation text features. Target operation text refers to natural language instructions or descriptions that describe the specific tasks or operations that the robot needs to perform. Target operation text clearly indicates a target and the actions or steps required to achieve that target. For example, the target operation text might be "pick up the red cube." Or, for example, "put the orange cube into the blue bowl."

[0093] In one example, in a field inspection and risk control scenario in the fintech sector, the target operation text is "inspect the bank lobby," or "go to the bank entrance to check the status of the door locks," or "check the status of the securities institution's fire-fighting equipment," etc.

[0094] In another example, in the context of drug management in the healthcare field, the target operation text is "get drug B for patient A" or "put the drug on the left into the appropriate storage shelf", etc.

[0095] In one embodiment, the target operation text can be text-encoded using a text encoder. The text encoder can be selected from the text encoder in a Contrastive Language-Image Pretraining (CLIP) model. The CLIP model is a multimodal deep learning model designed to achieve cross-modal understanding and association by jointly learning representations of images and text. The CLIP model is able to "understand" the relationship between images and descriptions, thus exhibiting strong performance across a variety of tasks.

[0096] The role of target operation text features is to serve as a semantic representation of task instructions. It is a linguistic guidance signal for robot behavior. Through cross-modal attention mechanism and fusion with visual features, it guides the video generation model to generate future video frames that conform to semantic expectations, thereby helping the robot to infer the actions that should be performed.

[0097] In step 102 of some embodiments, an initial viewpoint image of the target robot is acquired, and the initial viewpoint image is image-encoded to obtain initial image features. The initial viewpoint image refers to the first image acquired from the robot's sensors (e.g., a camera) before the robot starts or begins operation, used to establish an initial perception of the environment. The initial viewpoint image can refer to an image seen / detected / observed within the robot's field of vision (line of sight) before the robot starts or begins operation.

[0098] In one example, within a fintech field, the initial perspective image for on-site inspections and risk control can be a panoramic view of the site environment, an equipment status image, or a security access control and entrance / exit monitoring image. The panoramic view depicts the overall scene of the inspected area, such as the internal environment of a bank branch, securities company office, ATM area, warehouse, or server room. Equipment status images include on-site photos of key equipment, such as trading terminals, ATMs, servers, monitoring systems, and security access control equipment. Security access control and entrance / exit monitoring images include perspectives from entrances / exits and restricted areas, used to monitor the entry and exit of visitors and employees.

[0099] In another example, within the healthcare sector's pharmaceutical management scenario, the initial perspective image might be a panoramic view of the pharmaceutical inventory or a detailed image of the pharmaceutical storage area. A panoramic view of the pharmaceutical inventory refers to a picture of the entire storage area, such as the warehouse, pharmacy, or cold storage, used to show the overall layout of the pharmaceuticals. Detailed images of the pharmaceutical storage area are used to indicate the location of key pharmaceuticals, such as interior photos of specific shelves, cabinets, or refrigerated boxes.

[0100] In one embodiment, the initial viewpoint image can be image encoded using an image encoder. The image encoder can be selected from the CLIP model.

[0101] In step 103 of some embodiments, image change prediction is performed based on the target operation text features and the initial image features to obtain a target prediction video. The target prediction video includes at least two prediction images.

[0102] In one embodiment, reference is made to Figure 2 Step 103 may include:

[0103] Step 201: Obtain the video generation model; wherein the video generation model includes a text and image attention layer, at least two cascaded feature extractors and at least two cascaded feature decoders, and the at least two feature decoders include a first feature decoder and a second feature decoder.

[0104] Step 202: Attention processing is performed on the target operation text features and the initial image features through the image-text attention layer to obtain the target image-text fusion features;

[0105] Step 203: Extract features from the target image-text fusion features using at least two cascaded feature extractors to obtain the predicted image feature sequence;

[0106] Step 204: Perform feature decoding on the predicted image features using the first feature decoder to obtain the first image decoded feature sequence;

[0107] Step 205: Perform feature decoding on the first image decoding feature sequence using the second feature decoder to obtain the second image decoding feature sequence;

[0108] Step 206: Reconstruct the image from the second image decoding feature sequence to obtain the target prediction video.

[0109] In step 201, the video generation model is a deep learning model based on images and text to generate videos. The video generation model includes an image-text attention layer, at least two cascaded feature extractors, and at least two cascaded feature decoders. The image-text attention layer is used to perform attention fusion on image and text features. The feature extractors are used to perform deep feature extraction on the input features to obtain latent features that better highlight the visual transformation process. The feature decoders are used to decode the input features. The number of feature extractors is generally consistent with the number of feature decoders. For example, at least two feature extractors include a first feature extractor and a second feature extractor, and at least two feature decoders include a first feature decoder and a second feature decoder. Alternatively, at least two feature extractors include a first feature extractor, a second feature extractor, and a third feature extractor, and at least two feature decoders include a first feature decoder, a second feature decoder, and a third feature decoder.

[0110] In step 202, the text-image attention layer is specifically a cross-attention layer, which is a neural network structure design designed to achieve mutual information exchange and collaborative understanding between text and images.

[0111] In one embodiment, step 202 may include: obtaining a text query vector from the target operational text features, and obtaining an image key vector and an image value vector from the initial viewpoint image features; calculating the similarity between the text query vector and the image key vector (usually using dot product or additive attention) to obtain an image attention weight; performing feature weighting based on the image attention weight and the image value vector to obtain an image salient feature; obtaining a text key vector and a text value vector from the target operational text features, and obtaining an image query vector from the initial viewpoint image features; calculating the similarity between the image query vector and the text key vector (usually using dot product or additive attention) to obtain a text attention weight; performing feature weighting based on the text attention weight and the text value vector to obtain a text salient feature; and concatenating the image salient feature and the text salient feature to obtain the target image-text fusion feature. In this way, image features and text features can be fully fused, thereby achieving efficient integration of predicted future dynamic information with real-time control strategies, improving the generalization ability and execution efficiency of the control strategy.

[0112] In step 203, taking at least two feature extractors, including a first feature extractor and a second feature extractor, as an example, the first feature extractor can be used to extract features from the target image-text fusion features to obtain a first predicted image feature sequence; the second feature extractor can be used to extract features from the first predicted image feature sequence to obtain a second predicted image feature sequence; and the second predicted image feature sequence can be used as the predicted image feature sequence.

[0113] In steps 204 to 205, the predicted image features are decoded by the first feature decoder to obtain the first image decoded feature sequence; then the first image decoded feature sequence is decoded by the second feature decoder to obtain the second image decoded feature sequence.

[0114] In step 206, the image reconstruction can choose a deconvolution operation to obtain the target predicted video.

[0115] The advantage of the embodiments of steps 201 to 206 described above is that they can improve the accuracy of video generation, which is beneficial to improving the control accuracy of the robot.

[0116] In one embodiment, reference is made to Figure 3 Step 201 may include:

[0117] Step 301: Obtain the first sample data, which includes the first sample operation text, the first sample initial image, and the reference image;

[0118] Step 302: Perform text encoding on the first sample operation text to obtain the first sample operation text features, and perform image encoding on the first sample initial image to obtain the first sample initial image features;

[0119] Step 303: Add noise to the reference image to obtain a noisy image, and encode the noisy image to obtain the noise image features;

[0120] Step 304: Use the initial video generation model to generate a video based on the features of the noisy image, the features of the first sample operation text, and the features of the first sample initial image, to obtain the sample prediction image;

[0121] Step 305: Calculate the loss based on the sample predicted image and the reference image to obtain video generation loss data;

[0122] Step 306: Adjust the parameters of the initial video generation model based on the video generation loss data to obtain the video generation model.

[0123] In step 301, the first sample data refers to the data used to train the video generation model. The first sample data can come from at least one of the following: a human operation dataset, a robot operation dataset, or a domain-specific dataset. Training the video generation model using the first sample data allows the model to fully learn the features of data from different sources and effectively generate videos. Since different data sources provide different perspectives and knowledge about robot operation tasks, this significantly improves the model's generalization ability and task performance.

[0124] Human Operation Dataset: This portion of the data comes from various human operation records on the internet. It includes video data of humans performing tasks such as moving objects and operating tools in daily life. This data provides exemplary behaviors of how humans handle various objects and tasks, especially in human-computer interaction and object manipulation. This data is crucial for understanding the routine processes of operations, but because it is not necessarily a professional or precise operation, it often contains high diversity and uncertainty.

[0125] Robot operation data: This data comes from records of the robot performing tasks. Data on robots performing precision tasks (such as grasping, stacking, and moving objects) includes the robot's motion sequences, sensor data, and feedback from actual execution. Compared to human operation data, this data is typically more precise and provides stronger task execution capabilities, especially in specific applications such as robotic arm manipulation and grasping control. By optimizing this type of data, models can learn how the robot performs operations under precise control.

[0126] Domain-specific data: This type of data consists of operational data collected for specific application scenarios (such as the use of a particular tool, or tasks involving grasping or moving a specific object). It supplements the domain-specific knowledge gaps in human and robot operation datasets. For example, if the model needs to handle the task of how a medical robot precisely operates during surgery, domain-specific data would include information related to surgical instruments, patient positioning, and environmental conditions. By optimizing this type of data, the model can perform more customized task learning for a specific domain.

[0127] Regardless of the dataset from which the first sample data originates, it includes the first sample operation text, the first sample initial image, and a reference video. The reference video includes multiple reference images. The first sample operation text is essentially the same as the target operation text, referring to natural language instructions or descriptions that describe the specific task or operation the robot needs to perform. The first sample initial image is essentially the same as the initial viewpoint image, referring to the image seen / detected / observed within the robot's viewpoint (line of sight) before startup or operation. The reference image refers to the actual image seen / detected / observed by a human or robot within their viewpoint after performing the specific task.

[0128] In step 302, the specific processes of text encoding and image encoding can be referred to steps 101 and 102 above, and will not be repeated here.

[0129] In step 303, the noisy image can be represented as: Where, x t This represents the image after noise has been added at time step t, where t is the time step and indicates the current stage in the diffusion process. x0 represents the reference image (i.e., the noise-free image). b represents the noise intensity, and ε is a Gaussian noise variable used to add to the image during the diffusion process.

[0130] In step 304, the sample prediction image can be represented as v = V θ (x t ,l emb ,s0), where v represents the sample predicted image, V θ Represents the initial video generation model, l emb s0 represents the first sample operation text feature, and s0 represents the initial video frame, which is usually the first frame of the video, or it can be the initial viewpoint image.

[0131] In step 305, the L2 norm error function can be used to calculate the loss to obtain the video generation loss data. The video generation loss data can be expressed as:

[0132]

[0133] Among them, L D This represents the video generation loss data, where D represents the dataset, including the human operation dataset D. H Robot operation dataset D R Domain-specific data D C E represents the expected value (mean) of the sampled data distribution.

[0134] In step 306, the parameters of the initial video generation model are adjusted based on the video generation loss data to obtain a trained video generation model.

[0135] In one embodiment, the video generation loss data can be represented as:

[0136]

[0137] Where, ω H ω represents the weight of human-operated data. R ω represents the weight of the robot's operational data. C Indicates domain-specific data weights. This represents the video generation loss data corresponding to the human-operated dataset. This represents the video generation loss data corresponding to the robot operation dataset. This represents the video generation loss data corresponding to a domain-specific dataset.

[0138] Specifically, through the human operation dataset D H Robot operation dataset D R Domain-specific data D C Through joint optimization, the video generation model can extract different knowledge from each type of data: human operation datasets help the model understand common and intuitive operation methods, especially supporting the richness and diversity of data on human-object interaction. Robot operation datasets provide precise and controllable operation demonstrations, particularly suitable for tasks requiring high-precision operation. Optimization with this type of data allows the model to accurately simulate robot movements, playing a crucial role in high-precision operations such as micrometer-level grasping or assembly tasks. Domain-specific datasets help the model perform customized learning for a specific domain, enabling adaptive training for specific tasks. For example, in specific industrial environments, domain-specific data provides invaluable training samples for how robots perform tasks in complex physical environments. Furthermore, by optimizing these three types of data, the model can achieve multi-task learning. For instance, in practical applications, the model may not only need to complete simple operation tasks but also handle challenging and complex tasks, and even make reasonable inferences and decisions when faced with tasks never seen before. By introducing different data sources, the model can gain a more comprehensive understanding of the operation tasks and utilize different data sources to supplement and improve task execution strategies.

[0139] The advantage of the embodiments of steps 301 to 306 described above is that they can improve the model performance of the video generation model, which is conducive to improving the accuracy of video generation, and thus helps to improve the control accuracy of the robot.

[0140] In step 104 of some embodiments, feature extraction can be performed on the target prediction video by an image encoder to obtain a video spatiotemporal feature sequence.

[0141] In one embodiment, reference is made to Figure 4 Step 104 may include:

[0142] Step 401: Perform video encoding on the target prediction video to obtain the image encoding feature sequence;

[0143] Step 402: Channel concatenation is performed based on the image encoding feature sequence and the second image decoding feature sequence to obtain the first fused image feature sequence;

[0144] Step 403: Channel splicing is performed based on the first fused image feature sequence and the first image decoding feature sequence to obtain video spatiotemporal features.

[0145] Specifically, this embodiment innovatively extracts intermediate layer features of SVDM to construct a spatiotemporal perception representation.

[0146] In one example, a multi-scale feature aggregation module is proposed: Let the feature sequence of the m-th layer be... The following is obtained by aligning the spatial dimensions using bilinear interpolation and then stitching them together along the channel axis:

[0147]

[0148] Among them, F p L′ represents the spatiotemporal features of the video. m L represents the result after bilinear interpolation m T represents the number of frames, C represents the number of channels, m represents the number of layers, and W represents the number of layers. p H represents the width. p Indicates altitude.

[0149] The image encoding feature sequence, the second image decoding feature sequence, and the first image decoding feature sequence contain spatiotemporal information and higher-order features related to physical evolution. These features provide deeper information about the evolution of the video sequence, such as the robot arm's motion trajectory and changes in object position. This information is crucial for generating the robot's control actions. The video's spatiotemporal features enable the model to better understand the movement of objects and the state changes of the robot arm in the video.

[0150] The benefit of the embodiments of steps 401 to 403 described above is that they can improve the representation of spatiotemporal information and physical evolution information in the spatiotemporal features of the video, which is conducive to improving the accuracy of subsequent inverse dynamics modeling, and thus conducive to improving the control accuracy of the robot.

[0151] In step 105 of some embodiments, inverse dynamics modeling is performed based on the video spatiotemporal feature sequence and the target operation text features to obtain the target predicted action for each predicted image.

[0152] In one embodiment, spatial attention processing is applied to the video's spatiotemporal features to obtain key spatial features; temporal attention processing is then applied to these key spatial features to obtain key spatiotemporal features; a pre-set feedforward network is used to process these key spatiotemporal features to obtain a target predicted action feature sequence; and a pre-trained action generation model is used to generate actions from the target predicted action feature sequence and the target operation text features to obtain the target predicted action for each predicted image. This allows for action sequence modeling, which is beneficial for improving the robot's control accuracy.

[0153] In one embodiment, reference is made to Figure 5 Step 105 may include:

[0154] Step 501: The spatiotemporal features of the video from at least two target perspectives are spliced ​​together to obtain multi-view video spatiotemporal features;

[0155] Step 502: Perform spatiotemporal attention processing on the spatiotemporal features of the multi-view video to obtain the target prediction action feature sequence;

[0156] Step 503: Generate the target predicted action feature sequence and target operation text features using a pre-trained action generation model to obtain the target predicted action for each predicted image.

[0157] In step 501, the target viewpoint refers to one of multiple viewpoints of the robot. The target viewpoint can include a static viewpoint and a wrist viewpoint, etc. The spatiotemporal characteristics of the static viewpoint video can be represented as follows: The spatiotemporal characteristics of a video from a wrist perspective can be represented as follows: and Each feature contains video information acquired from different perspectives. In the model, video features from different perspectives are fused to enhance the model's spatiotemporal awareness.

[0158] In one embodiment, reference is made to Figure 6 Step 502 may include:

[0159] Step 601: Perform spatial attention processing on the spatiotemporal features of the multi-view video to obtain the key spatial features of the multi-view video.

[0160] Step 602: Perform temporal attention processing on the key spatial features of the multi-view video to obtain the key spatiotemporal features of the multi-view video.

[0161] Step 603: The key spatiotemporal features of the multi-view video are processed by a preset feedforward network to obtain the target prediction action feature sequence.

[0162] Specifically, a video converter can be designed to achieve spatiotemporal attention fusion to obtain a sequence of target prediction action features. The key spatial features of multi-view videos can be represented as: Multi-view video spatiotemporal features can be represented as: Q″=FFN(Temp-Attn(Q′)). Here, Q is an initial learned query vector, typically a representation of the spatiotemporal information learned, containing visual information extracted from the video prediction model. This query vector is one of the inputs to the space-time attention mechanism, used to focus on various temporal steps and spatial regions of the video. Q[i] is the query vector element at the i-th time step. The query vector element at each time step i interacts with specific features to generate the processing result for that time step. and These two feature distributions represent video features from static and wrist perspectives. Spat-Attn() refers to Spatial Attention, used to weight the spatial dimension of the input (such as an image or video frame), enabling the model to focus on specific regions in the image. Q' is the key spatial feature of the multi-view video, specifically the query vector after spatial attention processing, which contains weighted spatiotemporal information. The query vector at each time step is adjusted through spatial attention to better reflect the key information in the current frame. Temp-Attn(t) refers to Temporal Attention, used to weight the temporal dimension (time steps in the video sequence), enabling the model to focus on important time steps in time sequence. FFN(): Feedforward Neural Network, further extracting and enhancing spatiotemporal information. Q″ refers to the target predicted action feature sequence, specifically the query vector after processing by temporal attention and the feedforward network (FFN). It represents the final feature representation after integrating spatial and temporal attention mechanisms. This vector will serve as input in subsequent stages (such as the action generation stage) to further generate the robot's control actions.

[0163] It's worth noting that for multi-view inputs (such as static and wrist views), the model's spatial perception of the scene is enhanced by aggregating video features from different perspectives. This allows the model to more accurately capture the dynamic changes of objects, regardless of the robot's perspective. Inverse dynamics modeling involves learning the physical processes from the current state to the target state (e.g., moving from one location to another). This process helps the robot predict and generate appropriate control actions to complete the task.

[0164] Action generation models can use diffusion Transformer.

[0165] In one embodiment, prior to step 503, the robot control method may further include: a pre-trained motion generation model, specifically including:

[0166] Acquire second sample data, which includes second sample operation text, second sample initial image, and reference action sequence;

[0167] Text encoding is performed on the second sample operation text to obtain the second sample operation text features, and image encoding is performed on the second sample initial image to obtain the second sample initial image features;

[0168] Image change prediction is performed based on the operational text features of the second sample and the initial image features of the second sample to obtain the sample prediction video;

[0169] Feature extraction is performed on the sample prediction video to obtain the spatiotemporal feature sequence of the sample video, and attention processing is applied to the spatiotemporal feature sequence of the sample video to obtain the sample prediction action feature sequence.

[0170] Noise is added to the reference action sequence to obtain the sample noise action feature sequence;

[0171] The sample noise action feature sequence, the second sample operation text feature, and the sample predicted action feature sequence are denoised using a preset initial action generation model to obtain the sample denoised action sequence.

[0172] Loss is calculated based on the reference action sequence and the sample denoised action sequence to obtain action generation loss data;

[0173] The parameters of the initial action generation model are adjusted based on the action generation loss data to obtain the action generation model.

[0174] Action generation loss data can be represented as: D ψ (a k ,Q″)=argminL diff (ψ;A), Here, a0 represents the reference action sequence, which is the model's actual action at time step k=0. The model optimizes to make the generated action as close as possible to a0. k It is a sample noise action feature sequence, used to represent the action sequence during the diffusion process, specifically the noisy action generated at the k-th step of the diffusion. D ψ It is an action generation model. A denoiser model can be chosen as the action generation model, which uses a neural network with parameter ψ to denoise the action sequence a. k The task of the action generation model is to convert a... k Transform it into an action sequence that approximates the actual action a0. D ψ The input includes a k , l emb (Used to guide the robot to complete the task), Q″. D ψ The output is a sequence of sample denoising actions. emb This is the second operational text feature, which can be obtained by encoding the second target operational text (such as "put the red square into the blue box") using the CLIP model to obtain a vector representation. diff () is the diffusion loss function, which measures the error between the generated action sequence and the reference action sequence. By minimizing this loss, the model optimizes its generated action sequence to be closer to the real reference action sequence.

[0175] The advantage of the above embodiments is that they can improve the performance of the motion generation model, thereby improving the control accuracy of the robot.

[0176] In step 106 of some embodiments, the target robot is controlled to perform a target prediction action.

[0177] In one example, the robot control method provided in this application may include:

[0178] The first stage, the text-guided video generation process, includes: 1. Task instruction parsing: The robot receives the target operation text (e.g., "Put the red square into the blue plate"). 2. Video generation: The TVP model generates future predicted video frames based on the target operation question and the initial viewpoint image, simulating the scene changes after the robot performs the task. The series of video frames generated by the model are used to demonstrate how the robot manipulates objects and predict the possible visual effects after the action.

[0179] The second stage, inverse dynamics modeling and task execution, includes: 1. Feature extraction: Extracting spatiotemporal features from the generated video in the first stage. 2. Inverse dynamics modeling: Based on these features, the inverse dynamics model predicts how the robot will perform actions (such as grasping and moving objects), enabling the robot to accurately perform these operations in the real environment. 3. Task execution: Based on the predicted actions, the robot generates control commands and executes the task.

[0180] In one example, suppose the robot needs to place a red cube into a blue plate: First stage: A text-guided video generation model predicts a target video that represents the visual process of the robot picking up the cube from the table and placing it into the plate. Second stage: An inverse dynamics model generates a specific sequence of actions based on the predicted scene information, instructing the robot how to control its arm to perform actions such as grasping, moving, and placing, and then controls the robot to execute these actions.

[0181] In summary, the present application can achieve at least the following beneficial effects: (1) It improves task execution efficiency, especially the success rate in 12-DOF dexterity hand operation tasks. (2) By replacing multi-step denoising with single-step forward encoding, the control frequency of the robot is improved, and memory usage is reduced. (3) It can effectively suppress irrelevant texture interference and reduce generalization errors in complex tasks such as tool use. (4) It reduces the robot teaching cost. (5) The spatiotemporal attention mechanism enhances the modeling ability of long-term dependencies and improves the success rate in sequence operations containing multiple (e.g., more than 5) subtasks.

[0182] Please see Figure 7 This application also provides a robot control device that can implement the above-described robot control method. Figure 7 This is a block diagram of the module structure of the robot control device provided in the embodiments of this application. The device includes:

[0183] The text encoding module 701 is used to acquire the target operation text used to control the target robot, and to encode the target operation text to obtain the target operation text features.

[0184] The image encoding module 702 is used to acquire the initial view image of the target robot, and to encode the initial view image to obtain the initial image features;

[0185] The video generation module 703 is used to predict image changes based on the target operation text features and the initial image features to obtain a target prediction video; wherein the target prediction video includes at least two prediction images;

[0186] The feature extraction module 704 is used to extract features from the target predicted video to obtain a video spatiotemporal feature sequence.

[0187] The action prediction module 705 is used to perform inverse dynamics modeling based on the video spatiotemporal feature sequence and the target operation text features to obtain the target predicted action for each predicted image.

[0188] The motion execution module 706 is used to control the target robot to perform the target predicted motion.

[0189] It should be noted that the specific implementation of the robot's control device is basically the same as the specific implementation of the robot's control method described above, and will not be repeated here.

[0190] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the robot control method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0191] Please see Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0192] The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0193] The memory 802 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called by the processor 801 to execute the robot control method of the embodiments of this application.

[0194] The 803 input / output interface is used to implement information input and output.

[0195] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0196] Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);

[0197] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0198] This application embodiment also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described robot control method.

[0199] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0200] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0201] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0203] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0204] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0205] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0206] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0207] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0208] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0209] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0210] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for controlling a robot, characterized in that, The method includes: Obtain the target operation text used to control the target robot, and encode the target operation text to obtain the target operation text features; Acquire an initial viewpoint image of the target robot, and perform image encoding on the initial viewpoint image to obtain initial image features; Image change prediction is performed based on the target operation text features and the initial image features to obtain a target prediction video; wherein, the target prediction video includes at least two prediction images; Feature extraction is performed on the target predicted video to obtain a video spatiotemporal feature sequence; Inverse dynamics modeling is performed based on the video spatiotemporal feature sequence and the target operation text features to obtain the target predicted action for each predicted image; Control the target robot to perform the target predicted action.

2. The method according to claim 1, characterized in that, The step of predicting image changes based on the target operation text features and the initial image features to obtain the target predicted video includes: Obtain a video generation model; wherein the video generation model includes a text and image attention layer, at least two cascaded feature extractors, and at least two cascaded feature decoders, wherein the at least two feature decoders include a first feature decoder and a second feature decoder; The target text-image fusion feature is obtained by performing attention processing on the target operation text feature and the initial image feature through the image-text attention layer. The target image-text fusion features are extracted by at least two cascaded feature extractors to obtain a predicted image feature sequence; The predicted image features are decoded using the first feature decoder to obtain the first image decoded feature sequence; The second image decoding feature sequence is obtained by performing feature decoding on the first image decoding feature sequence using the second feature decoder; The target predicted video is obtained by reconstructing the image from the second image decoding feature sequence.

3. The method according to claim 2, characterized in that, The step of extracting features from the target predicted video to obtain video spatiotemporal features includes: The target prediction video is video encoded to obtain an image encoded feature sequence; The first fused image feature sequence is obtained by concatenating the channels according to the image encoding feature sequence and the second image decoding feature sequence. The video spatiotemporal features are obtained by concatenating channels based on the first fused image feature sequence and the first image decoding feature sequence.

4. The method according to claim 2, characterized in that, The acquisition of the video generation model includes: Acquire first sample data, which includes first sample operation text, first sample initial image, and reference image; The first sample operation text is text encoded to obtain the first sample operation text features, and the first sample initial image is image encoded to obtain the first sample initial image features; The reference image is noise-added to obtain a noisy image, and the noisy image is image-encoded to obtain noise image features; The initial video generation model is used to generate a video from the features of the noisy image, the features of the first sample operation text, and the features of the first sample initial image to obtain a sample prediction image. Loss calculation is performed based on the sample predicted image and the reference image to obtain video generation loss data; The initial video generation model is adjusted based on the video generation loss data to obtain the video generation model.

5. The method according to any one of claims 1 to 4, characterized in that, The step of performing inverse dynamics modeling based on the video spatiotemporal features and target operation text features to obtain the target prediction action sequence for each prediction image includes: The video spatiotemporal features from at least two target perspectives are spliced ​​together to obtain multi-view video spatiotemporal features; Spatiotemporal attention processing is performed on the spatiotemporal features of the multi-view video to obtain the target prediction action feature sequence; The target predicted action is generated by using a pre-trained action generation model to generate actions from the target predicted action feature sequence and the target operation text features, thereby obtaining the target predicted action for each predicted image.

6. The method according to claim 5, characterized in that, The step of performing spatiotemporal attention processing on the spatiotemporal features of the multi-view video to obtain the target predicted action feature sequence includes: Spatial attention processing is performed on the spatiotemporal features of the multi-view video to obtain key spatial features of the multi-view video. Temporal attention processing is performed on the key spatial features of the multi-view video to obtain key spatiotemporal features of the multi-view video. The key spatiotemporal features of the multi-view video are processed by a preset feedforward network to obtain the target prediction action feature sequence.

7. The method according to claim 5, characterized in that, Before generating the target predicted action feature sequence using a pre-trained action generation model to obtain the target predicted action for each predicted image, the method further includes: Pre-training the action generation model specifically includes: Acquire second sample data, which includes second sample operation text, second sample initial image, and reference action sequence; The second sample operation text is text encoded to obtain the second sample operation text features, and the second sample initial image is image encoded to obtain the second sample initial image features; Based on the text features of the second sample operation and the initial image features of the second sample, image change prediction is performed to obtain the sample prediction video; Feature extraction is performed on the sample predicted video to obtain the sample video spatiotemporal feature sequence, and attention processing is performed on the sample video spatiotemporal feature sequence to obtain the sample predicted action feature sequence. Noise is added to the reference action sequence to obtain the sample noise action feature sequence; The sample noise action feature sequence, the second sample operation text feature, and the sample predicted action feature sequence are denoised using a preset initial action generation model to obtain a sample denoised action sequence. Loss calculation is performed based on the reference action sequence and the sample denoised action sequence to obtain action generation loss data; The parameters of the initial action generation model are adjusted based on the action generation loss data to obtain the action generation model.

8. A control device for a robot, characterized in that, The device includes: The text encoding module is used to acquire the target operation text for controlling the target robot, and to encode the target operation text to obtain the target operation text features. The image encoding module is used to acquire the initial view image of the target robot, and to encode the initial view image to obtain initial image features; A video generation module is used to predict image changes based on the target operation text features and the initial image features to obtain a target prediction video; wherein the target prediction video includes at least two prediction images; The feature extraction module is used to extract features from the target predicted video to obtain a video spatiotemporal feature sequence; The action prediction module is used to perform inverse dynamics modeling based on the video spatiotemporal feature sequence and target operation text features to obtain the target predicted action for each predicted image. The action execution module is used to control the target robot to perform the target predicted action.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the robot control method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the robot control method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Mechanical arm control method, device and equipment based on large visual model and storage medium

    CN118143940A