Robot control method and device, electronic equipment and storage medium

CN122807958APending Publication Date: 2026-09-25ASTRIBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611311560.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,上述架构在实际部署中面临显著的推理延迟问题,现有VLA模型在每一帧都独立执行完整的前向计算流程,未能利用相邻帧之间的时序冗余,导致每帧推理延迟高,难以满足机器人高频控制的实时性要求

Benefits of technology

本申请实施提供的机器人控制方法,通过获取控制指令流数据和图像流数据,并分别确定当前时刻与上一时刻之间图像帧和控制指令的差异情况,实现了对机器人输入变化的精确感知。在此基础上,仅在差异情况不符合预设条件时,根据实际的差异情况选择性地将图像帧和控制指令中的至少一者输入编码器进行编码,获得当前时刻的特征对,并更新预设存储空间中的缓存特征对。通过上述方式,本申请实施例避免了每一帧均对图像帧和控制指令进行完整编码所带来的重复计算,显著减少了编码器前向计算的计算量,从而降低了每帧推理延迟。同时,预设存储空间中的特征对始终保持为最新时刻的编码结果,解码器基于该特征对解码得到的动作信息始终对应于当前时刻的输入状态,确保了控制指令与视觉输入变化被及时响应,在保证机器人控制精度的前提下实现了推理加速。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807958A_ABST
    Figure CN122807958A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a robot control method and device, electronic equipment and storage medium, and relate to the technical field of computers. The method comprises: acquiring control instruction stream data of a robot and image stream data collected by a vision sensor of the robot in real time; determining differences of an image frame and control instructions at a current time compared with those at a previous time according to the control instruction stream data and the image stream data; decoding the feature pairs at the current time into the trained decoder to obtain action information of the robot to be executed at the current frame, and controlling the robot to act according to the action information of the current frame to be executed. In the above manner, the amount of calculation of the forward calculation of the encoder is significantly reduced, thereby reducing the inference delay of each frame. At the same time, the feature pairs in the preset storage space always remain the encoding results at the latest time, ensuring that the changes of the control instructions and the vision input are responded in time, and realizing inference acceleration on the premise of ensuring the control accuracy of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a robot control method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of embodied intelligence technology, VLA (Vision-Language-Action) models are widely used in the field of robot control. VLA models typically employ a three-layer architecture: a vision sensor acquires real-time environmental image stream data at a fixed frame rate, while the user issues control commands via natural language; a visual encoder encodes the current frame image into visual features, and a text encoder encodes the control commands into command features; both are input into the VLM backbone network for cross-modal fusion, resulting in fused feature pairs; the action decoder decodes and outputs the action information to be executed by the robot in the current frame based on these feature pairs, thus achieving end-to-end closed-loop control of vision, language, and action.

[0003] However, the above architecture faces significant inference latency issues in actual deployment. Existing VLA models execute the complete forward computation process independently in each frame, failing to utilize the temporal redundancy between adjacent frames, resulting in high inference latency per frame, which makes it difficult to meet the real-time requirements of high-frequency robot control. Summary of the Invention

[0004] The purpose of this application is to at least solve one of the aforementioned technical defects. The technical solution provided by the embodiments of this application is as follows: In a first aspect, embodiments of this application provide a robot control method, including: Acquire the robot's control command stream data and the real-time image stream data collected by the robot's vision sensors; The differences between the current image frame and the control command and the previous time frame are determined based on the control command stream data and the image stream data. If the difference does not meet the preset conditions, then at least one of the image frame and control command at the current moment is input into the trained encoder for encoding according to the difference to obtain the feature pair at the current moment; the feature pair includes the image features obtained by encoding the image frame and the command features obtained by encoding the control command. Update the feature pairs stored in the preset storage space to the feature pairs at the current time. The features at the current moment are decoded into the trained decoder to obtain the action information that the robot needs to perform in the current frame. The robot's actions are then controlled based on the action information to be performed in the current frame.

[0005] Secondly, embodiments of this application provide a robot control device, including: The data acquisition module is used to acquire the robot's control command stream data and the image stream data acquired in real time by the robot's vision sensors; The data comparison module is used to determine the differences between the current image frame and the control command and the previous time frame based on the control command stream data and the image stream data. The data encoding module is used to encode at least one of the image frame and control command at the current moment into the trained encoder if the difference does not meet the preset conditions, so as to obtain the feature pair at the current moment; the feature pair includes the image features obtained by encoding the image frame and the command features obtained by encoding the control command. The feature pair update module is used to update the feature pairs stored in the preset storage space with the feature pairs at the current time. The data decoding module is used to decode the features input to the trained decoder at the current moment to obtain the action information to be executed by the robot in the current frame, and to control the robot's actions based on the action information to be executed in the current frame.

[0006] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory; The processor executes a computer program to implement the method provided in the first aspect embodiment or any alternative embodiment of the first aspect.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method provided in the first aspect embodiment or any optional embodiment of the first aspect.

[0008] The beneficial effects of the technical solutions provided in this application are: The robot control method provided in this application acquires control command stream data and image stream data, and determines the differences between the image frames and control commands between the current moment and the previous moment, thereby achieving precise perception of changes in robot input. Based on this, only when the differences do not meet preset conditions, at least one of the image frames and control commands is selectively input into the encoder for encoding, obtaining the feature pair at the current moment, and updating the cached feature pair in the preset storage space. Through this method, the embodiments of this application avoid the repetitive calculations caused by fully encoding the image frames and control commands for each frame, significantly reducing the computational load of the encoder's forward calculation, thus reducing the inference latency per frame. Simultaneously, the feature pairs in the preset storage space always remain the latest encoding results, and the action information decoded by the decoder based on these feature pairs always corresponds to the input state at the current moment, ensuring that changes in control commands and visual input are responded to in a timely manner, achieving inference acceleration while maintaining robot control accuracy. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0010] Figure 1 A flowchart illustrating a robot control method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the overall process of a robot control method in one example of an embodiment of this application; Figure 3 A structural block diagram of a robot control device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0012] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0014] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0015] Figure 1 This application provides a flowchart illustrating a robot control method, the execution entity of which can be a robot core processor or an intelligent agent, such as... Figure 1 As shown, the method may include: Step S101: Obtain the robot's control command stream data and the image stream data collected in real time by the robot's vision sensor.

[0016] In the embodiments of this application, the control command stream data is a sequence of task commands issued by the user to the robot in natural language. In specific implementations, the robot may be equipped with a voice receiving module (e.g., a microphone array) and / or a text input interface (e.g., a touchscreen or remote communication interface). The user can speak commands, such as "move the red cup on the table to the left of the tray." After the voice receiving module collects the voice signal, it is converted into a natural language command in text form by a voice recognition engine. The user can also directly input natural language command text through the text input interface. The acquired control command stream data forms a continuous sequence of commands in time, where each moment corresponds to an independent control command, used to indicate the task objective that the robot needs to perform. The acquisition frequency of the control command stream data is determined by the frequency of user commands. In this embodiment, the user typically issues commands at a lower frequency, meaning that the control commands remain unchanged for most of the continuous time period, only changing when a task switch is required. The visual sensor is an RGB (Red-Green-Blue) camera mounted on the robot body, which continuously captures images of the work scene in front of the robot. The camera's frame rate is, for example, 30 frames per second, meaning it continuously captures 30 images per second, with a time interval of approximately 0.033 seconds between adjacent frames. The camera outputs one image frame at each capture moment; this image frame is an RGB three-channel color image with a resolution of, for example, 224×224 pixels. The image stream data forms a continuous sequence of image frames in time, where each frame corresponds to a capture moment and is used to characterize the spatial state information of the robot's environment at that moment, including but not limited to the position, shape, color, and pose of objects, as well as the relative spatial relationships between objects.

[0017] It should be noted that the control command stream data and image stream data have a temporal correspondence. The control command at each moment in the control command stream data is paired temporally with the image frame in the image stream data at that moment: the control command issued by the user at the current moment corresponds to the environmental state represented by the image frame acquired by the robot's vision sensor at that moment. Together, they constitute the robot's complete input information at the current moment. The image frame reflects "what is in the environment and where it is," while the control command reflects "what the robot needs to do." For example, when the user issues the command "pick up the red cup," this command corresponds to the image frame containing the red cup captured by the camera at the current moment.

[0018] Specifically, when a robot performs a control task, it first needs to acquire two types of input data: control command stream data and image stream data. The image stream data is acquired in real time by the robot's vision sensors. In the actual implementation, the control command stream data and the image stream data are transmitted to the robot's computing unit through different data channels. After receiving these two types of data, the computing unit aligns them according to timestamps to ensure that the control command at the current moment is paired and associated with the image frame at that moment, forming an input data pair of "image frame + control command" for use in subsequent steps.

[0019] Step S102: Determine the differences between the current image frame and the control command compared to the previous moment based on the control command stream data and the image stream data.

[0020] Specifically, in this step, the robot's computing unit determines the differences between the current image frame and the previous image frame, as well as the differences between the current control command and the previous control command, based on the acquired control command stream data and image stream data. For the image stream data, the computing unit acquires the current image frame and reads the previous image frame from the cache. The current image frame is input into the visual encoder to obtain the current visual features, and the visual features cached from the previous time step are read, and the cosine similarity between the two is calculated. The closer the similarity value is to 1, the more similar the two images are. The calculated similarity value is compared with a preset threshold of 0.95: if it is greater than 0.95, the image is determined to have not changed significantly; if it is less than or equal to 0.95, the image is determined to have changed significantly.

[0021] For control command stream data, the computing unit acquires the control command at the current moment and the control command at the previous moment. The current and previous control commands are input into the text encoder to obtain the current command embedding vector and the historical command embedding vector, respectively, and their cosine similarity is calculated. The calculated similarity value is compared with a preset threshold of 0.99: if it is greater than 0.99, the commands are considered semantically identical; if it is less than or equal to 0.99, the commands are considered substantially different.

[0022] Through independent comparison of the two dimensions mentioned above, the computing unit obtains the image difference and instruction difference respectively, which together constitute the complete difference information between the current moment and the previous moment, and are used for decision-making in subsequent steps.

[0023] Step S103: If it is determined that the difference does not meet the preset conditions, then at least one of the image frame and the control command at the current moment is input into the trained encoder for encoding according to the difference to obtain the feature pair at the current moment; the feature pair includes the image features obtained by encoding the image frame and the command features obtained by encoding the control command.

[0024] In the embodiments of this application, feature pairs can be composed of visual features of image frames and linguistic features of control commands, and the form of feature pairs can be represented in the form of KV key-value components, such as... ,in, Represents the visual features of an image frame. Language features that represent control commands.

[0025] Specifically, in this step, the calculation unit determines whether the determined difference meets a preset condition. In this embodiment, the preset condition is that there is no difference between the current moment and the previous moment, that is, the image has not changed significantly (visual feature similarity greater than 0.95) and the control command has not changed substantially (semantic similarity greater than 0.99). In this case, the current frame can directly reuse the feature pair of the previous frame without re-encoding. If the difference does not meet the preset condition, that is, the image has changed significantly and / or the control command has changed substantially, the calculation unit inputs at least one of the image frame and the control command at the current moment into the trained encoder for encoding to obtain the feature pair at the current moment.

[0026] "Inputting at least one of the data to the encoder based on the differences" means that the encoder includes a visual encoder and a text encoder, both of which are trained feature extraction networks. The computing unit decides to input only the changed data into the corresponding encoder for encoding based on whether the image and control command have changed, while the unchanged parts are directly read from the preset storage space for reuse without repeated encoding. Specifically, the visual encoder is used to encode image frames. Its input is the original RGB image, and its output is a fixed-dimensional image feature sequence. This sequence represents the environmental spatial information in the image frame in the form of visual features, including the position, shape, color of objects, and their spatial relationships. The text encoder is used to encode control commands. Its input is natural language text, and its output is a fixed-dimensional command feature sequence. This sequence represents the semantic information of the command in the form of linguistic features, including the task objective and interaction intent.

[0027] Image features and instruction features together constitute the feature pair at the current moment. This feature pair fully represents the environmental spatial state of the robot at the current moment and the task objective to be performed. It can be used as input for subsequent decoding steps to generate motion control information for the current frame.

[0028] By using the selective encoding method described above, only the data that has changed needs to be encoded during each inference, while the unchanged parts are directly read from the cache and reused. This ensures the integrity of the feature pairs at the current moment and reduces the repetitive calculations caused by fully encoding the image frames and control commands for each frame, thereby saving computational overhead in the encoding process and inference time.

[0029] Step S104: Update the feature pairs stored in the preset storage space to the feature pairs at the current time.

[0030] In the embodiments of this application, the preset storage space is a cache area (e.g., GPU memory or dedicated cache) in the computing unit, and its capacity is sufficient to store at least one set of feature pairs. After encoding is completed, the computing unit writes the feature pairs at the current moment into the preset storage space, overwriting the historical feature pairs stored at the previous moment, thereby completing the update of the cached data.

[0031] The updated feature pair serves as a complete representation of the current moment, containing both current image and command features, for use in subsequent decoding steps. Simultaneously, this feature pair will also serve as a historical feature pair for the next moment, used in the next round of inference to determine differences and selectively reuse data. If the image and control commands remain unchanged in the next moment, the decoder can directly read the current feature pair from this storage space for reuse, without re-encoding.

[0032] Through the above update mechanism, the preset storage space always saves the latest encoded feature pairs, realizing dynamic maintenance and fast access of feature pairs.

[0033] Step S105: Decode the features at the current moment into the trained decoder to obtain the action information to be executed by the robot in the current frame, and control the robot's actions according to the action information to be executed in the current frame.

[0034] In the embodiments of this application, the decoder is a trained action generation network, such as a neural network model based on Flow Matching or Diffusion Policy architecture. Its input is the fused feature pairs, and its output is the action information to be executed by the robot in the current frame. The image features in the feature pairs provide spatial state information of the current environment, and the instruction features provide semantic information of the task objective. The decoder maps the above-mentioned fused features to motion parameters of each joint of the robot through its internal attention mechanism and regression head. The action information includes, but is not limited to, the target angles and angular velocities of each joint of the robot, the target pose of the end effector, and the motion trajectory. For example, the action information can be represented as the expected angles of each joint of a 7-DOF robotic arm, or the target coordinates and pose of the end effector in Cartesian space, and its physical form matches the kinematic model of the robot.

[0035] Specifically, the computing unit decodes the current-moment features into the trained decoder. The computing unit then converts the decoded motion information into specific drive signals, which are sent to the robot's robotic arm actuators to drive the joints of the robotic arm to move according to the target angle or pose, thereby achieving real-time control of the robot. Since the current-moment feature pairs are obtained by selectively encoding based on differences or by updating from a cache, the unchanged parts of the feature pairs have been obtained through a reuse mechanism. Therefore, the motion information output by the decoder based on these feature pairs accurately reflects the current visual and control state, while balancing inference speed and motion execution accuracy. Subsequently, the robot enters the next moment's acquisition and control loop.

[0036] The robot control method provided in this application acquires control command stream data and image stream data, and determines the differences between the image frames and control commands between the current moment and the previous moment, thereby achieving precise perception of changes in robot input. Based on this, only when the differences do not meet preset conditions, at least one of the image frames and control commands is selectively input into the encoder for encoding, obtaining the feature pair at the current moment, and updating the cached feature pair in the preset storage space. Through this method, the embodiments of this application avoid the repetitive calculations caused by fully encoding the image frames and control commands for each frame, significantly reducing the computational load of the encoder's forward calculation, thus reducing the inference latency per frame. Simultaneously, the feature pairs in the preset storage space always remain the latest encoding results, and the action information decoded by the decoder based on these feature pairs always corresponds to the input state at the current moment, ensuring that changes in control commands and visual input are responded to in a timely manner, achieving inference acceleration while maintaining robot control accuracy.

[0037] Based on the above embodiments, as an optional embodiment, the method further includes: If the difference is determined to meet the preset conditions, the feature pairs stored in the preset storage space will be used as the feature pairs at the current moment.

[0038] Specifically, based on the calculation unit's judgment of the differences between the current image frame and the control command compared to the previous time step in the aforementioned steps, when the image has not changed significantly and the control command has not changed substantially, the calculation unit determines that the difference meets the preset conditions.

[0039] In this scenario, the current image frame and control commands are essentially the same as those of the previous frame, meaning the current input state is largely consistent with the previous one. Therefore, the feature pairs for the current frame do not need to be re-encoded; instead, the feature pairs stored in the preset storage space are directly used as the current frame's feature pairs. This storage space is the cache area updated in the previous step, which stores the complete feature pairs obtained after encoding in the previous frame, including the image and command features from that frame. Since the visual environment and task objectives have not changed substantially between the current and previous frames, the feature pairs from the previous frame can still accurately represent the current environmental state and task semantics.

[0040] The computing unit copies or maps cached feature pairs in the preset storage space to feature pairs at the current moment through read operations, for use in subsequent decoding steps. At this time, the encoder does not need to perform any forward computation, which saves the computational overhead of the encoding stage and avoids inference delays caused by repeated encoding.

[0041] This reuse mechanism achieves feature acquisition with "zero computational overhead" while ensuring that the feature pairs match the current input state. Subsequently, the computing unit inputs the reused feature pairs into the decoder to generate action information and control the robot's actions. Throughout the process, data paths that do not change are bypassed, and re-encoding is triggered only when changes occur, thereby achieving optimal inference efficiency while ensuring control accuracy.

[0042] It should be noted that the cached feature pairs in the preset storage space have been updated to the latest feature pairs after the previous round of processing. When the difference meets the preset conditions, the cached feature pairs can be directly reused without additional processing. When the difference changes to no longer meet the preset conditions at a later time, the computing unit will re-execute the encoding operation and update the preset storage space to ensure the synchronization of cached data with the actual input state.

[0043] Based on the above embodiments, as an optional embodiment, at least one of the current image frame and control command is input into the trained encoder for encoding according to the differences, in order to obtain the feature pair at the current moment, specifically including: If the similarity between the image frames at the current time and the previous time is not greater than a preset threshold, and the control commands at the current time and the previous time are the same, then the image frame at the current time is input into the encoder for encoding to obtain the image features at the current time. The feature pair at the current moment is obtained by combining the image features at the current moment with the instruction features in the feature pairs stored in the preset storage space.

[0044] Specifically, when the computing unit determines that the similarity between the image frames at the current moment and the previous moment is no greater than a preset threshold, and the control commands at the current moment and the previous moment are the same, it indicates that only the visual environment has changed significantly at the current moment (e.g., the robot's head rotates, causing a change in the field of view, or objects on the workbench are moved or replaced), but the task objective given by the user has not changed. In this case, the computing unit only inputs the image frame at the current moment into the visual encoder for encoding to obtain the image features at the current moment. These features represent the position, shape, color, and spatial relationships between objects in the current environment in the form of a visual feature sequence. The command features do not need to be re-encoded; they are directly read from the preset storage space and reused from the historical command features cached at the previous moment.

[0045] After obtaining the current image features, the computing unit combines them with the reused instruction features to form a complete feature pair for the current moment. The specific combination method is as follows: according to the index positions of the features in the visual-language model input sequence, the image features are placed at the beginning of the sequence (corresponding to feature positions 1 to 196), and the instruction features are placed at the end of the sequence (corresponding to feature positions 197 to 228), merging them into a unified feature sequence of length 228. This feature pair simultaneously contains the updated environmental spatial information and the unchanged task objective information, accurately representing the complete input state at the current moment.

[0046] In this way, when only the visual input changes while the instruction remains the same, only the image frame needs to be encoded. The instruction features reuse the cached results from the previous time step. Compared to encoding both the image and the instruction simultaneously, this saves the computational overhead of text encoding and further improves inference efficiency. The combined feature pairs are then updated to a preset storage space, replacing the feature pairs from the previous time step, and are also used in subsequent decoding steps to generate action information.

[0047] Based on the above embodiments, as an optional embodiment, at least one of the current image frame and control command is input into the trained encoder for encoding according to the differences, in order to obtain the feature pair at the current moment, specifically including: If the similarity between the image frames at the current time and the previous time is not greater than a preset threshold, and the control commands at the current time and the previous time are different, then the image frame at the current time is input into the encoder for encoding to obtain the image features at the current time, and the control command at the current time is input into the encoder for encoding to obtain the command features at the current time. The image features and instruction features at the current moment are combined to obtain the feature pair at the current moment.

[0048] Specifically, when the computing unit determines that the similarity between the image frames at the current moment and the previous moment is no greater than a preset threshold, and the control commands at the current moment and the previous moment are different, it indicates that the visual environment at the current moment has changed significantly, and the task objective given by the user has also changed substantially (e.g., from "picking up the red cup" to "putting down the blue plate"). In this case, neither the image frame nor the control command can directly reuse the cached data from the previous moment, and both need to be encoded simultaneously.

[0049] The computing unit inputs the current image frame into the visual encoder to obtain the image features at that moment. The visual encoder extracts features from the image frame and outputs the current environmental spatial information represented as a sequence of visual features. Simultaneously, the computing unit inputs the current control commands into the text encoder to obtain the command features at that moment. The text encoder performs semantic encoding on the natural language commands and outputs the current task objective information represented as a sequence of linguistic features.

[0050] After obtaining the current image features and current instruction features, the computing unit combines them to form a complete feature pair for the current moment. The combination method is consistent with the previous embodiment: according to the index position of the features in the visual-language model input sequence, the image features are placed at the beginning of the sequence, and the instruction features are placed at the end of the sequence, merging them into a unified feature sequence. This feature pair simultaneously contains the updated environmental spatial information and the updated task objective information, completely representing the input state at the current moment.

[0051] Subsequently, the computing unit updates the feature pair to a preset storage space, and it is also used in subsequent decoding steps to generate action information. This embodiment ensures that the feature pair can fully respond to changes in both visual and semantic dimensions by encoding both the image and the control command simultaneously when both change, thereby guaranteeing that the action information output by the subsequent decoder matches the complete input state at the current moment.

[0052] Based on the above embodiments, as an optional embodiment, at least one of the current image frame and control command is input into the trained encoder for encoding according to the differences, in order to obtain the feature pair at the current moment, specifically including: If the similarity between the image frames at the current time and the previous time is greater than a preset threshold, and the control commands at the current time and the previous time are different, the control command at the current time is input into the encoder for encoding to obtain the command features at the current time. The current instruction features are combined with the image features in the feature pairs stored in the preset storage space to obtain the current feature pairs.

[0053] Specifically, when the computing unit determines that the similarity between the image frames at the current moment and the previous moment is greater than a preset threshold, and the control commands at the current moment and the previous moment are different, it indicates that the visual environment at the current moment has not changed significantly compared to the previous moment, but the task objective issued by the user has been substantially switched. In this case, the environmental spatial information corresponding to the image frame is basically the same as that at the previous moment, and the image features encoded and cached in the preset storage space at the previous moment can be directly reused; while the control commands need to be re-encoded to obtain new semantic representations due to the change in task objectives.

[0054] The computing unit inputs the control command at the current moment into the text encoder for encoding, obtaining the command features at the current moment. The text encoder performs semantic encoding on the natural language command, outputting the current task target information represented in the form of a sequence of language features. Image features do not need to be re-encoded; the computing unit directly reads and reuses the historical image features cached from the previous moment from the preset storage space.

[0055] Subsequently, the computing unit combines the instruction features at the current moment with the reused image features to obtain the complete feature pair at the current moment. This feature pair contains the invariant environmental spatial information and the updated task objective information, which can accurately represent the complete input state at the current moment, where the environment remains unchanged but the task has changed.

[0056] In this embodiment, since image features are directly read and reused from the cache, the computing unit only performs encoding operations on control instructions, saving computational overhead for visual encoding compared to encoding both images and instructions simultaneously. The combined feature pairs are then updated to a preset storage space and used in subsequent decoding steps to generate motion information matching the current instruction. This scheme is suitable for application scenarios where user instructions change frequently but the robot's environment is relatively fixed, such as performing different operation sequences on the same object in a static desktop environment.

[0057] Based on the above embodiments, as an optional embodiment, the encoder and decoder are trained in the following manner: Multiple training samples are acquired. The training samples are synchronously acquired image frames and control commands. The training label of each training sample is the action information generated by the robot executing the control commands. For each training sample, the training sample is input into the encoder to obtain feature pairs, the feature pairs are input into the decoder to obtain predicted action information, and the loss value corresponding to the training sample is obtained based on the difference between the predicted action information and the training label of the training sample. The loss function value is obtained based on the function value corresponding to all training samples. The parameters of the encoder and decoder are updated based on the loss function value until the loss function value converges, thus obtaining the pre-trained encoder and decoder.

[0058] In the embodiments of this application, "synchronous acquisition" refers to the temporal pairing and correspondence between image frames and control commands. Image frames are acquired by the robot's vision sensor at a fixed frame rate, and control commands are issued synchronously by the user via voice or text input at the moment the image frame is acquired. The image frames and control commands together constitute the input portion of a training sample. The training label for each training sample is the actual motion information generated when the robot executes the control command. This actual motion information is pre-acquired baseline data, such as the joint angle trajectories of the robotic arm or the pose sequence of the end effector in Cartesian space, recorded through teleoperation or manual demonstration. The label data is strictly aligned temporally with the input sample to ensure that the label accurately corresponds to the action that should be generated when the control command is executed under the environmental state represented by the image frame. The overall loss function value is the average of the loss values ​​of all training samples; the smaller the value, the closer the predicted motion information output by the decoder is to the actual motion information. Based on this overall loss function value, the calculation unit calculates the gradient of each network parameter in the encoder and decoder using the backpropagation algorithm and updates each parameter along the gradient descent direction to reduce the overall loss function value. The loss function can be the mean squared error loss function, which is used to quantify the degree of deviation between the predicted action and the actual action.

[0059] Specifically, multiple training samples are first acquired. Each training sample consists of synchronously acquired image frames and control commands, along with a corresponding training label. For each training sample, the computational unit inputs it into the encoder for encoding, obtaining a feature pair for that sample. This feature pair includes image features and command features; the image features represent the spatial state of the current environment, and the command features represent the semantic information of the task objective. Then, this feature pair is input into the decoder for decoding, obtaining predicted action information. The predicted action information is an estimate of the robot's actions output by the decoder, which initially deviates significantly from the actual action information. Based on the difference between the predicted action information and the training label, the computational unit calculates the loss value corresponding to that training sample, and then updates the encoder parameters based on the loss value.

[0060] Repeat the above process of forward propagation, loss calculation and parameter update until the overall loss function value converges to below the preset threshold (e.g., the change in the loss value is less than the preset tolerance). At this point, the parameters of the encoder and decoder are basically stable, and they can accurately extract feature pairs from the input image frames and control commands and decode the corresponding action information. Training is complete, and the trained encoder and decoder are obtained.

[0061] In the training process described above, the encoder and decoder are trained end-to-end as a whole, which enables the feature pairs extracted by the encoder to retain the effective information required by the decoder to generate accurate actions to the greatest extent, thereby improving the accuracy of robot control.

[0062] Based on the above embodiments, as an optional embodiment, the training samples are input into the encoder to obtain feature pairs, specifically including: Image frames of other training samples adjacent to the image frames of the training samples are randomly selected based on a preset probability and input into the encoder to obtain feature pairs.

[0063] Specifically, during the actual training process, the computing unit determines whether to apply a perturbation to the current training sample based on a preset probability p (p=0.15 in this embodiment, i.e., a 15% probability). Specifically, for each training sample, when a perturbation is triggered, the computing unit randomly selects an image frame from other training samples adjacent to that training sample's image frame, replaces the current training sample's image frame, and keeps the control command and training label of that training sample unchanged. Then, the replaced image frame and the original control command are input together into the encoder for encoding to obtain the corresponding feature pairs.

[0064] It should be noted that "adjacent" here refers to temporal adjacency, that is, in the image stream data continuously acquired by the robot's vision sensor, the image located in the frame before or after the current image frame. For example, the current training sample contains image frame I. t Control command C t and action tag A t The computing unit randomly selects an I-frame adjacent to the image frame of the training sample. t 1 or I t+1 As a replacement frame, one of the two is selected with a 50% probability to replace I. t Later with C t and A t Paired input encoder.

[0065] Through the aforementioned perturbation mechanism, the encoder adapts to the condition of "a one-frame temporal misalignment between the input image and the control command" during the training phase—that is, the visual input received by the encoder may deviate from the current control command in terms of timing by one frame. When the same temporal misalignment occurs during the inference phase due to feature pair reuse, the encoder and decoder can maintain stable performance output and will not experience accuracy degradation due to changes in input distribution. This perturbation enhancement strategy does not require additional labeled data; data expansion can be achieved simply by replacing training samples between frames, effectively improving the model's robustness to temporal changes.

[0066] Based on the above embodiments, as an optional embodiment, the preset conditions include that the current frame is not a frame with a preset refresh cycle, and the preset refresh cycle is to perform a forced refresh once every preset number of frames.

[0067] Specifically, in the aforementioned embodiments, the preset conditions include that the image has not changed significantly (visual feature similarity is greater than a preset threshold) and the control command has not changed substantially (semantic similarity is greater than a preset threshold), in which case the cached feature pairs can be reused. Based on the above conditions, this embodiment further adds a timing condition as a component of the preset conditions. Specifically, the preset conditions also include: the current frame is not a frame with a preset refresh cycle, and the preset refresh cycle is one forced refresh performed every preset number of frames. In this embodiment, forced refresh means that regardless of whether the image frame and control command at the current moment have changed compared to the previous moment, it is considered that the difference does not meet the preset conditions, and the encoder's forward calculation is forcibly triggered. The image frame and control command at the current moment are input into the encoder for encoding to obtain the feature pairs at the current moment, and the cached feature pairs in the preset storage space are updated. The preset refresh cycle is one forced refresh performed every preset number of frames. In this embodiment, the preset number K is an empirical value, preferably K=2, that is, one forced refresh performed every 2 frames. The calculation unit maintains a frame counter t. When tmodK=0 (that is, the current frame number is an integer multiple of K), the current frame is determined to be a frame with a refresh cycle. When tmodK=0, even if the visual feature similarity is greater than the preset threshold and the control command semantic similarity is greater than the preset threshold, the encoding operation is forcibly triggered. The image frame and control command at the current moment are input into the encoder to obtain the complete feature pair and update the cache.

[0068] The purpose of this forced refresh mechanism is that, during the reuse of feature pairs, although the feature pairs can accurately reflect the environmental state and task semantics, even if the image and control commands do not change substantially, after repeated reuse, the feature pairs may develop imperceptible small offsets due to numerical truncation or accumulated errors (such as floating-point precision loss or noise accumulation in cache access). By forcibly refreshing the cache at fixed intervals, real computational data can be periodically introduced to reset the cache state, eliminate accumulated errors, and thus ensure the long-term accuracy of feature pairs and the long-term stability of robot control.

[0069] The following is combined Figure 2 The overall flow of the robot control method provided in the embodiments of this application will be introduced, such as... Figure 2 As shown, the robot control method provided in this application embodiment can be divided into the following stages: S1. In the model training phase, the encoder and decoder are jointly trained end-to-end. Training samples consist of synchronously acquired image frames and control commands, with labels representing the actual motion information generated by the robot executing the commands. During training, frames adjacent to the current sample are randomly selected with a preset probability to replace the current frame, while keeping the commands and labels unchanged. This allows the encoder to adapt to situations where there is a one-frame misalignment between the vision and control commands. This perturbation enhancement mechanism improves the model's robustness to temporal changes, ensuring that reusing cached features during the inference phase does not lead to a decrease in accuracy.

[0070] S2, Input Data Acquisition Phase: Acquire the robot's control command stream data and the image stream data acquired in real time by the vision sensor. Control commands are natural language task instructions given by the user via voice or text, and the image stream is a sequence of color images continuously captured by an RGB camera at a fixed frame rate. These two types of data are paired temporally; that is, the control command at the current moment corresponds to the image frame acquired at that moment. Together, they constitute the robot's complete input at the current moment: the image represents what is in the environment, and the command represents what needs to be done.

[0071] S3, the difference determination stage: Based on the control command stream and image stream data acquired in S2, the differences between the current time step and the previous time step for each image frame and control command are calculated. The current image frame is encoded as a visual feature, and its cosine similarity is calculated with the visual features of the previous frame in the cache. This cosine similarity is then compared with a preset threshold to determine if the image has undergone significant changes. The current control command and the command from the previous time step are encoded as text embedding vectors, and their cosine similarity is calculated. This cosine similarity is then compared with a preset threshold to determine if the command has undergone substantial changes. The comparison results of these two dimensions together constitute complete difference information.

[0072] S4. Condition Judgment Phase: Based on the differences determined in S3, determine whether the preset conditions are met. The preset conditions include: the image has not changed significantly, the control command has not changed substantially, and the current frame is not a frame with a preset refresh cycle. When all three conditions are met, the preset conditions are met, and S4.1 is executed; when any condition is not met, the preset conditions are not met, and S4.2 is executed.

[0073] S4.1 Reusing cached feature pairs: When the difference determined by S4 meets the preset conditions, the input state at the current moment is not substantially different from the previous moment, and no re-encoding is required. The computing unit directly reads the complete feature pair cached at the previous moment from the preset storage space as the feature pair for the current moment. This feature pair contains image features and instruction features, which can accurately represent the environmental state and task semantics at the current moment. In this step, the encoder does not need to perform any forward computation, realizing "zero computational overhead" feature acquisition and saving inference time to the greatest extent.

[0074] S4.2 When S4 determines that the difference does not meet the preset conditions, at least one of the image frame and control command at the current moment is input into the encoder for encoding according to the specific difference to obtain the feature pair at the current moment.

[0075] S5. In the preset storage space update stage, the current-moment feature pairs obtained in S4.2 are written to the preset storage space, overwriting the historical feature pairs from the previous moment. The preset storage space is a high-speed cache area with sufficient capacity to store a complete set of feature pairs. The updated feature pairs are used for subsequent decoding steps at the current moment, and also serve as the "previous-moment feature pairs" for the next frame, for use in the next round of difference judgment and selective reuse. Through this update mechanism, the cache always saves the latest encoded feature pairs, realizing dynamic maintenance and fast access to feature pairs.

[0076] In the S6 robot control phase, the current feature pairs are input to the trained decoder, which decodes them and outputs motion information. The decoder maps the fused feature pairs to motion parameters for each joint of the robot. The motion information is converted into drive signals and sent to the robotic arm actuators to drive the robot's movement. Since the unchanged parts of the feature pairs have been obtained through a multiplexing mechanism, the motion information output by the decoder accurately reflects the current input state, achieving inference acceleration while ensuring control accuracy.

[0077] Figure 3 A structural block diagram of a robot control device provided in an embodiment of this application is shown below. Figure 3 As shown, the robot control device 300 may include: a data acquisition module 301, a data comparison module 302, a data encoding module 303, a feature pair update module 304, and a data decoding module 305, wherein, The data acquisition module 301 is used to acquire the robot's control command stream data and the image stream data acquired in real time by the robot's vision sensor; The data comparison module 302 is used to determine the differences between the current image frame and the control command and the previous time frame based on the control command stream data and the image stream data. The data encoding module 303 is used to encode at least one of the image frame and control command at the current moment into the trained encoder according to the difference situation if the difference situation is determined to be inconsistent with the preset conditions, so as to obtain the feature pair at the current moment; the feature pair includes the image features obtained by encoding the image frame and the command features obtained by encoding the control command. Feature pair update module 304 is used to update the feature pairs stored in the preset storage space to the feature pairs at the current time. The data decoding module 305 is used to decode the features input to the trained decoder at the current moment to obtain the action information to be executed by the robot in the current frame, and to control the robot's actions according to the action information to be executed in the current frame.

[0078] The robot control method provided in this application acquires control command stream data and image stream data, and determines the differences between the image frames and control commands between the current moment and the previous moment, thereby achieving precise perception of changes in robot input. Based on this, only when the differences do not meet preset conditions, at least one of the image frames and control commands is selectively input into the encoder for encoding, obtaining the feature pair at the current moment, and updating the cached feature pair in the preset storage space. Through this method, the embodiments of this application avoid the repetitive calculations caused by fully encoding the image frames and control commands for each frame, significantly reducing the computational load of the encoder's forward calculation, thus reducing the inference latency per frame. Simultaneously, the feature pairs in the preset storage space always remain the latest encoding results, and the action information decoded by the decoder based on these feature pairs always corresponds to the input state at the current moment, ensuring that changes in control commands and visual input are responded to in a timely manner, achieving inference acceleration while maintaining robot control accuracy.

[0079] Based on the above embodiments, as an optional embodiment, the device further includes a feature pair reading module, specifically used for: If the difference is determined to meet the preset conditions, the feature pairs stored in the preset storage space will be used as the feature pairs at the current moment.

[0080] Based on the above embodiments, as an optional embodiment, the data encoding module is specifically used for: If the similarity between the image frames at the current time and the previous time is not greater than a preset threshold, and the control commands at the current time and the previous time are the same, then the image frame at the current time is input into the encoder for encoding to obtain the image features at the current time. The feature pair at the current moment is obtained by combining the image features at the current moment with the instruction features in the feature pairs stored in the preset storage space.

[0081] Based on the above embodiments, as an optional embodiment, the data encoding module is further configured to: If the similarity between the image frames at the current time and the previous time is not greater than a preset threshold, and the control commands at the current time and the previous time are different, then the image frame at the current time is input into the encoder for encoding to obtain the image features at the current time, and the control command at the current time is input into the encoder for encoding to obtain the command features at the current time. The image features and instruction features at the current moment are combined to obtain the feature pair at the current moment.

[0082] Based on the above embodiments, as an optional embodiment, the data encoding module can also be used for: If the similarity between the image frames at the current time and the previous time is greater than a preset threshold, and the control commands at the current time and the previous time are different, the control command at the current time is input into the encoder for encoding to obtain the command features at the current time. The current instruction features are combined with the image features in the feature pairs stored in the preset storage space to obtain the current feature pairs.

[0083] Based on the above embodiments, as an optional embodiment, the device further includes an encoder training module, specifically used for: Multiple training samples are acquired. The training samples are synchronously acquired image frames and control commands. The training label of each training sample is the action information generated by the robot executing the control commands. For each training sample, the training sample is input into the encoder to obtain feature pairs, the feature pairs are input into the decoder to obtain predicted action information, and the loss value corresponding to the training sample is obtained based on the difference between the predicted action information and the training label of the training sample. The loss function value is obtained based on the function value corresponding to all training samples. The parameters of the encoder and decoder are updated based on the loss function value until the loss function value converges, thus obtaining the pre-trained encoder and decoder.

[0084] Based on the above embodiments, as an optional embodiment, the encoder training module is further configured to: Image frames of other training samples adjacent to the image frames of the training samples are randomly selected based on a preset probability and input into the encoder to obtain feature pairs.

[0085] Based on the above embodiments, as an optional embodiment, the preset conditions include that the current frame is not a frame with a preset refresh cycle, and the preset refresh cycle is to perform a forced refresh once every preset number of frames.

[0086] The following is for reference. Figure 4 It illustrates an electronic device suitable for implementing embodiments of this application (e.g., performing...). Figure 1 The diagram shows the structure of the terminal device or server 400 of the method shown. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable devices, etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0087] The electronic device includes a memory and a processor. The memory stores a program for executing the methods described in the above-described method embodiments; the processor is configured to execute the program stored in the memory. The processor may be referred to as processing device 401 as described below, and the memory may include at least one of read-only memory (ROM) 402, random access memory (RAM) 403, and storage device 408 as described below, as follows: like Figure 4 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0088] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0089] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this application.

[0090] It should be noted that the computer-readable storage medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0091] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0092] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0093] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: The system acquires the robot's control command stream data and the real-time image stream data collected by the robot's vision sensors. Based on the control command stream data and image stream data, it determines the differences between the current image frame and the control command compared to the previous moment. If the determined differences do not meet preset conditions, at least one of the current image frame and the control command is input into the trained encoder for encoding to obtain the current feature pair. The feature pair includes image features obtained by encoding the image frame and command features obtained by encoding the control command. The feature pairs stored in the preset storage space are updated to the current feature pairs. The current feature pairs are input into the trained decoder for decoding to obtain the action information to be executed by the robot in the current frame. The robot's actions are controlled based on the action information to be executed in the current frame.

[0094] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more controllable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0096] The modules or units described in the embodiments of this application can be implemented in software or hardware. The names of modules or units do not necessarily limit the specific unit; for example, a first constraint acquisition module can also be described as a "module for acquiring the first constraint".

[0097] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0098] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0099] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0100] The above description is only a partial embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A robot control method, characterized in that, include: Acquire the robot's control command stream data and the real-time image stream data collected by the robot's vision sensors; Based on the control command stream data and image stream data, determine the differences between the current image frame and the control command compared to the previous moment; If it is determined that the difference does not meet the preset conditions, then at least one of the image frame and the control command at the current moment is input into the trained encoder for encoding according to the difference to obtain the feature pair at the current moment; the feature pair includes the image features obtained by encoding the image frame and the command features obtained by encoding the control command. Update the feature pairs stored in the preset storage space to the feature pairs at the current moment; The features at the current moment are decoded into the trained decoder to obtain the action information to be executed by the robot in the current frame, and the robot's actions are controlled according to the action information to be executed in the current frame.

2. The method according to claim 1, characterized in that, The method further includes: If the difference is determined to meet the preset conditions, then the feature pairs stored in the preset storage space are used as the feature pairs at the current moment.

3. The method according to claim 1, characterized in that, The step of inputting at least one of the current-time image frame and control command into the trained encoder for encoding based on the differences to obtain the current-time feature pair includes: If the similarity between the image frame at the current moment and the image frame at the previous moment is not greater than a preset threshold, and the control commands at the current moment and the image frame at the previous moment are the same, then the image frame at the current moment is input into the encoder for encoding to obtain the image features at the current moment. The feature pair at the current moment is obtained by combining the image features at the current moment with the instruction features in the feature pairs stored in the preset storage space.

4. The method according to claim 1, characterized in that, The step of inputting at least one of the current-time image frame and control command into the trained encoder for encoding based on the differences to obtain the current-time feature pair includes: If the similarity between the image frames at the current time and the previous time is not greater than a preset threshold, and the control commands at the current time and the previous time are different, then the image frame at the current time is input into the encoder for encoding to obtain the image features at the current time, and the control command at the current time is input into the encoder for encoding to obtain the command features at the current time. The image features and instruction features at the current moment are combined to obtain the feature pair at the current moment.

5. The method according to claim 1, characterized in that, The step of inputting at least one of the current-time image frame and control command into the trained encoder for encoding based on the differences to obtain the current-time feature pair includes: If the similarity between the image frames at the current time and the previous time is greater than a preset threshold, and the control commands at the current time and the previous time are different, the control command at the current time is input into the encoder for encoding to obtain the command features at the current time. The current instruction features and the image features in the feature pairs stored in the preset storage space are combined to obtain the current feature pairs.

6. The method according to claim 1, characterized in that, The encoder and the decoder are trained in the following manner: Multiple training samples are acquired, which are synchronously acquired image frames and control commands. The training label of each training sample is the action information generated by the robot executing the control commands. For each training sample, the training sample is input into the encoder to obtain a feature pair, the feature pair is input into the decoder to obtain predicted action information, and the loss value corresponding to the training sample is obtained based on the difference between the predicted action information and the training label of the training sample. The loss function value is obtained based on the function value corresponding to all training samples. The parameters of the encoder and decoder are updated based on the loss function value until the loss function value converges, thus obtaining the pre-trained encoder and decoder.

7. The method according to claim 6, characterized in that, The step of inputting the training samples into the encoder to obtain feature pairs includes: Image frames of other training samples adjacent to the image frames of the training samples are randomly selected based on a preset probability and input into the encoder to obtain feature pairs.

8. The method according to claim 1, characterized in that, The preset conditions include that the current frame is not a frame with a preset refresh cycle, and the preset refresh cycle is to perform a forced refresh once every preset number of frames.

9. A robot control device, characterized in that, include: The data acquisition module is used to acquire the robot's control command stream data and the image stream data acquired in real time by the robot's vision sensors; The data comparison module is used to determine the differences between the current image frame and the control command and the previous time frame based on the control command stream data and the image stream data. The data encoding module is used to encode at least one of the image frame and the control command at the current moment into a trained encoder according to the difference situation if it is determined that the difference situation does not meet the preset conditions, so as to obtain the feature pair at the current moment; the feature pair includes the image features obtained by encoding the image frame and the command features obtained by encoding the control command. The feature pair update module is used to update the feature pairs stored in the preset storage space to the feature pairs at the current time. The data decoding module is used to decode the features input to the trained decoder at the current moment to obtain the action information to be executed by the robot in the current frame, and to control the robot's actions according to the action information to be executed in the current frame.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-8.