Method, device, medium and program product for outputting sequence of action instructions

By using a visual language action model based on a generative diffusion strategy, robots can parse human commands and generate human-like action command sequences, solving the problem of stiff interaction caused by the lack of non-verbal information in existing technologies and achieving more natural human-computer interaction.

CN121733544APending Publication Date: 2026-03-27SHANGHAI MATRIX SUPER INTELLIGENT SYSTEM INTEGRATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511947768.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-27

Smart Images

  • Figure CN121733544A_ABST
    Figure CN121733544A_ABST
Patent Text Reader

Abstract

The invention aims to provide a method and device for outputting an action instruction sequence, a medium and a program product, and the method comprises the steps: inputting a first camera image which is collected by a robot and comprises a first human body instruction action into a visual language action model based on a generative diffusion strategy, obtaining a first task semantic code corresponding to the first human body instruction action output by a visual language model located on the upper layer in the visual language action model; enabling an action planning module in the visual language action model to perform reasoning according to the first task semantic code, and outputting a corresponding hidden action vector; and an action expert module in the visual language action model performs reasoning according to the hidden action vector and the ontology state code corresponding to the current state information of the robot, and outputs an action instruction sequence corresponding to the robot, so that the robot executes the action instruction sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and more particularly to a technique for outputting sequences of action instructions. Background Technology

[0002] Human-computer interaction (HCI) is a core area of ​​robotics development, aiming to achieve natural and efficient communication between humans and robots. Traditional HCI technologies mainly rely on audio or text input, which can effectively convey information in certain scenarios. However, compared to human-to-human interaction, the lack of processing of non-verbal information (such as gestures and facial expressions) results in an unnatural and insufficiently human-like interactive experience.

[0003] As human-computer interaction technology develops towards naturalness and emotionalization, existing single-modal interaction systems based on audio or text have revealed significant limitations: body language (such as gestures and spatial pointing) and facial expressions (such as nodding for confirmation and smiling for feedback), which are crucial in human communication, are systematically ignored, leading to a one-sided understanding of robot intentions and a mechanical interactive experience. This deficiency is particularly prominent in spatial task execution (such as "taking the cup to the right front") and services for special groups (such as sign language interaction for the deaf and mute). Summary of the Invention

[0004] One object of this application is to provide a method, apparatus, medium, and program product for outputting a sequence of action instructions.

[0005] According to one aspect of this application, a method for outputting a sequence of action instructions is provided, the method comprising:

[0006] The first camera image containing the first human body command action, collected by the robot, is input into the visual language action model based on the generative diffusion strategy to obtain the first task semantic code corresponding to the first human body command action output by the visual language model located at the upper layer of the visual language action model.

[0007] This enables the action planning module in the visual language action model to perform inference based on the semantic encoding of the first task and output the corresponding hidden action vector.

[0008] The motion expert module in the visual language motion model performs inference based on the hidden motion vector and the ontology state code corresponding to the robot's current state information, and outputs the motion instruction sequence corresponding to the robot so that the robot executes the motion instruction sequence.

[0009] According to one aspect of this application, a computer device for outputting a sequence of action instructions is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.

[0010] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0011] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.

[0012] According to one aspect of this application, a computer device for outputting a sequence of action instructions is provided, the device comprising:

[0013] The module is used to input the first camera image containing the first human body command action collected by the robot into the visual language action model based on the generative diffusion strategy, and obtain the first task semantic code corresponding to the first human body command action output by the visual language model located in the upper layer of the visual language action model.

[0014] The first and second modules are used to enable the action planning module in the visual language action model to perform reasoning based on the semantic encoding of the first task and output the corresponding hidden action vector.

[0015] The first and third modules are used to enable the motion expert module in the visual language motion model to perform reasoning based on the hidden motion vector and the ontology state code corresponding to the robot's current state information, and output the motion instruction sequence corresponding to the robot so that the robot can execute the motion instruction sequence.

[0016] Compared with existing technologies, this application proposes a visual language action model architecture based on a generative diffusion strategy. By inputting camera images containing human commands and actions collected by the robot into the visual language model, and inputting the task semantic encoding output by the visual language model into the visual language action model, the action planning module can parse specified human actions (including emotional actions, spatial pointing actions, sign language actions, etc.), achieving a leap in the accuracy of human intention recognition. This enhances the anthropomorphism of human-computer interaction. Through the visual language action model, the robot can understand and process human commands and actions, capture non-verbal information, and thus more comprehensively understand human intentions. The action expert module controlled by the generative diffusion strategy enables the robot to respond with anthropomorphic body movements (such as nodding responses and gesture guidance), making the interaction more vivid and anthropomorphic. This can overcome the problem of stiff interaction caused by the lack of non-verbal information, making the interaction between robots and humans closer to the natural communication between humans, and improving the user experience. Attached Figure Description

[0017] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0018] Figure 1 This diagram illustrates a method for outputting a sequence of action instructions according to an embodiment of the present application.

[0019] Figure 2 A flowchart illustrating a method for training a visual language action model according to an embodiment of this application is shown.

[0020] Figure 3 This diagram illustrates a computer device structure for outputting a sequence of action instructions according to an embodiment of this application.

[0021] Figure 4 Exemplary systems that can be used to implement the various embodiments described in this application are shown.

[0022] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0023] The present application will now be described in further detail with reference to the accompanying drawings.

[0024] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.

[0025] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.

[0026] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0027] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.

[0028] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0029] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.

[0030] Figure 1A flowchart illustrating a method for outputting a sequence of action commands according to an embodiment of this application is shown. The method includes steps S11, S12, and S13. In step S11, a computer device inputs a first camera image containing a first human action command acquired by a robot into a visual language action model based on a generative diffusion strategy to obtain a first task semantic code corresponding to the first human action command output by the upper-level visual language model in the visual language action model. In step S12, the computer device causes the action planning module in the visual language action model to perform inference based on the first task semantic code and output the corresponding hidden action vector. In step S13, the computer device causes the action expert module in the visual language action model to perform inference based on the hidden action vector and the ontology state code corresponding to the robot's current state information and output the action command sequence corresponding to the robot, so that the robot executes the action command sequence.

[0031] In step S11, the computer device inputs the first camera image containing the first human body command action captured by the robot into the visual language action model based on the generative diffusion strategy, and obtains the first task semantic code corresponding to the first human body command action output by the visual language model located at the upper layer of the visual language action model.

[0032] In some embodiments, the computer device may be the robot itself, or it may be a remote workstation responsible for model reasoning.

[0033] In some embodiments, the robot captures a first camera image in front of it containing a first human command action. The human command action includes, but is not limited to, emotional actions, pointing actions to spatial locations of objects, sign language actions, etc. This example embodiment does not impose any special limitations on this.

[0034] In some embodiments, a Vision-Language-Action Model (VLA) is an AI model that combines visual observation and language commands to directly generate control actions (such as robot joint movements) in the physical world. Generative diffusion strategy refers to treating policy learning as a conditional generation process, progressively denoising in the action space to generate a sequence of action commands that meet the task requirements. The diffusion model, through progressive generation of "adding noise and denoising," can handle multimodal distributions and is suitable for complex action planning. A visual-language-action model based on generative diffusion strategy refers to a visual-language-action model that uses the diffusion model as an "action decoder" (i.e., action expert module) to generate a sequence of robot action commands.

[0035] In some embodiments, a first camera image containing a first human body command action is input into a trained visual language action model, so that the upper-level visual language model (VLM), i.e., a sub-model in the visual language action model, generates a first task semantic code corresponding to the first human body command action. Here, VLM is an artificial intelligence model that can simultaneously understand and process image and text information. It combines the capabilities of computer vision (processing images) and natural language processing (processing text) through deep learning technology to achieve cross-modal interaction between images and language. In some embodiments, VLM includes an image encoder network, which performs task semantic embedding on the first camera image to generate the corresponding first task semantic code. Here, the first task semantic code refers to extracting, understanding, and encoding the core semantics of the user's intention from the first camera image and converting it into an "operation mode" or "response paradigm" that can be executed within the model.

[0036] In step S12, the computer device causes the action planning module in the visual language action model to perform reasoning based on the semantic encoding of the first task and output the corresponding hidden action vector.

[0037] In some embodiments, the semantic encoding of the first task is input into the action planning module of the visual language action model, enabling the action planning module to reason based on the semantic encoding of the first task and output the corresponding hidden action vector. The action planning module is responsible for understanding and transforming the semantic encoding of the task into an abstract action representation (i.e., the hidden action vector) that is not yet concrete but already contains the intention to execute. In some embodiments, the visual language action model adopts a fast-slow system architecture, decomposing the entire visual language action model into two closely cooperating subsystems with different operating rhythms and responsibilities: 1. Slow system (upper-level system, i.e., the action planning module): responsible for high-level planning, complex decision-making, and long-term goals. It "thinks deeply but runs slowly." 2. Fast system (lower-level system, i.e., the action expert module): responsible for rapid execution, immediate response, and low-level control. It "responds extremely quickly but is relatively simple." The core advantage of this architecture lies in "using the intelligence of the slow system to guide the fast system, and using the agility of the fast system to ensure the real-time performance and security of the system."

[0038] In step S13, the computer device causes the motion expert module in the visual language motion model to perform reasoning based on the hidden motion vector and the ontology state code corresponding to the robot's current state information, and outputs the motion instruction sequence corresponding to the robot so that the robot executes the motion instruction sequence.

[0039] In some embodiments, the motion expert module (i.e., motion decoder) in the visual language motion model infers the robot's current state based on the hidden motion vectors and the ontology state encoding corresponding to the robot's current state information, and outputs a sequence of motion instructions for the robot. The robot's current state information includes, but is not limited to, the robot's joint angles, angular velocities, torques, etc., at the current moment. This example embodiment does not impose any special limitations on this. The motion expert module is a neural network component responsible for converting high-level instructions or latent (hidden motion vectors in this example embodiment) representations into specific sequences of motion instructions. In some embodiments, the robot's ontology state encoding refers to using a set of data (usually a vector or matrix) to uniquely and accurately describe the robot's current state information. In some embodiments, the motion instruction sequence consists of multiple motion instructions arranged in chronological order. The motion instructions are used to specify the precise quantized signals of the target angle (or position) that each joint of the robot needs to reach in the next time step. The motion instructions include, but are not limited to, the target angle or target position of each joint of the robot. This example embodiment does not have any special limitations on this. In some embodiments, the motion expert module uses flow matching to characterize the motion distribution to determine the real-time inference of the model and the continuity and smoothness of the output motion instruction sequence. Flow matching characterizes the motion distribution during the training process of the motion expert module by teaching the model with demonstration trajectories: starting from "pure noise", learning a smooth path "from random to correct", feeding "current noise point + condition" each time, and the model predicts "the direction of the next push". In the inference process of the motion expert module, starting from pure noise, the model "push" for a number of steps (e.g., 10 to 20 steps) and then outputs a complete motion instruction sequence. In some embodiments, the robot executes one of the multiple action instructions sequentially at each time step according to the order of the multiple action instructions in the action instruction sequence, wherein the robot's underlying joint driver (e.g., servo motor) drives the robot to execute the action instruction at a certain frequency (e.g., 200-400 Hz).

[0040] This application proposes a visual language action model architecture based on a generative diffusion strategy. By inputting camera images containing human commands and actions captured by the robot into the visual language model, and inputting the task semantic encoding output by the visual language model into the visual language action model, the action planning module can parse specified human actions (including emotional actions, spatial pointing actions, sign language actions, etc.), achieving a leap in the accuracy of human intention recognition and enhancing the anthropomorphism of human-computer interaction. Through the visual language action model, the robot can understand and process human commands and actions, capture non-verbal information, and thus more comprehensively understand human intentions. The action expert module controlled by the generative diffusion strategy enables the robot to respond with anthropomorphic body movements (such as nodding responses and gesture guidance), making the interaction more vivid and anthropomorphic. This can overcome the problem of stiff interaction caused by the lack of non-verbal information, making the interaction between robots and humans closer to the natural communication between humans and improving the user experience.

[0041] In some embodiments, the step of inputting a first camera image containing a first human body command action captured by a robot into a visual language action model based on a generative diffusion strategy to obtain a first task semantic code corresponding to the first human body command action output by the upper-level visual language model in the visual language action model includes: inputting a first camera image containing a first human body command action captured by a robot and a first prompting information corresponding to the first human body command action into a visual language action model based on a generative diffusion strategy to obtain a first task semantic code corresponding to the first human body command action output by the upper-level visual language model in the visual language action model and a second task semantic code corresponding to the first prompting information, wherein the first prompting information includes first audio information or first text information; wherein, the step of causing the action planning module in the visual language action model to perform inference based on the first task semantic code and output a corresponding hidden action vector includes: causing the action planning module in the visual language action model to perform inference based on the first task semantic code and the second task semantic code and output a corresponding hidden action vector. In some embodiments, the first prompt information corresponding to the first human body command action is also input into the visual language action model. The first prompt information can be text information or audio information. For example, if the first human body command action is an emotional action, the corresponding first prompt information can be the content spoken by the person when performing the emotional action (in text or audio form). For example, if the emotional action is "to show a very disappointed expression, while lowering the head and spreading the hands", the corresponding first prompt information is "I was scolded by ×× again today". As another example, if the first human body command action is an object spatial positioning action, the corresponding first prompt information can refer to the instruction that the robot needs to grasp or operate, such as "Please get that bracelet for me". As yet another example, if the first human body command action is a sign language action, the corresponding first prompt information can be the text translation of the sign language (in text or audio form). In some embodiments, the visual language model at the upper layer of the visual language action model also includes a text encoder network (or audio bias coder network or speech recognition network). This text encoder network (or audio bias coder network or speech recognition network) performs task semantic embedding on the first prompt information (text or audio) to generate the corresponding second task semantic encoding. In some embodiments, if the first prompt is in audio format, it can be converted to text format before being input into the visual language model. Alternatively, the visual language model can first convert the audio prompt to text format, and then the text encoder network can embed the text prompt into task semantics. In some embodiments, the second task semantic encoding refers to extracting, understanding, and encoding the core semantics of the user's intent from the first prompt and converting it into an "operation mode" or "response paradigm" that can be executed within the model.In some embodiments, the first task semantic encoding and the second task semantic encoding are input together into the action planning module of the visual language action model. The action planning module infers based on the first task semantic encoding and the second task semantic encoding and outputs the corresponding hidden action vector. The action planning module is responsible for understanding and converting the task semantic encoding into an abstract action representation (i.e., a hidden action vector) that is not yet concretized but already contains the execution intention. The visual language action model in this application corresponds to multimodal input, that is, it simultaneously inputs a first camera image containing a first human command action and a first prompt information corresponding to the first human command action. This enables the visual language action model to simultaneously parse human command actions (emotional actions, object spatial positioning pointing actions, sign language actions) and prompt information in text or audio form, thereby achieving a leap in the accuracy of intent recognition. By integrating multimodal data and combining the multi-head attention mechanism of the Transformer architecture, this application enables the visual language action model to provide more complete contextual information, ultimately enabling the robot to automatically capture human hand movements and their underlying intentions. For example, "waving" means calling the robot over, and "pointing to the trash on the ground" combined with the language command "help me throw the trash on the ground into the trash can" will cause the robot to pick up the trash and put it in the trash can. For example, if a user inputs the first prompt to the robot, "Put this apple on that table," the robot can combine visually captured body movements to understand the spatial location of the apple the user is pointing to and the spatial location where it needs to be placed, and then complete the precise placement operation.

[0042] In some embodiments, the first human command action includes at least one of the following: emotional actions; pointing actions to a spatial location; and sign language actions. In some embodiments, emotional actions include, but are not limited to, physical actions performed by a human body to represent emotions (such as anger, surprise, thinking, cheering, joy, etc.); pointing actions to a spatial location include, but are not limited to, physical actions performed by a human body to point to an object or a spatial location with a finger; and sign language actions include, but are not limited to, physical actions performed by a human body to convey information through visual elements such as gestures, facial expressions, and body movements. These are mainly used for communication by people with hearing or speech impairments. For example, a robot can understand sign language and respond with operational actions based on a first camera image captured by its camera that includes the sign language actions performed by the user. For example, when the robot sees the user performing a sign language action to say "thank you" through the camera, the user will output a corresponding physical action to express joy or "you're welcome." This application utilizes a limb-driven spatial perception enhancement mechanism to transform the natural spatial indication characteristics of human limb movements (such as pointing and demarcation) into label-free training data, significantly reducing the model's reliance on manual annotation. Through a generative diffusion strategy-controlled motion expert module, the robot can dynamically respond with anthropomorphic limb movements (such as nodding responses and gesture guidance), especially in scenarios serving the deaf and mute, enabling bidirectional sign language interaction. For example, it can be applied to sign language translation work, enabling robots to autonomously translate sign language. This technical approach not only overcomes the problem of stiff interaction caused by the lack of non-verbal information, but also opens up new dimensions in terms of spatial task execution accuracy (such as "placing an apple on a designated table" combining gestures and voice) and inclusion of special groups. This application integrates multimodal data, including human commands and actions (emotional actions, spatial pointing actions, and sign language actions) and natural language prompts (audio and text), and leverages a visual language action model based on a generative diffusion strategy. This enables robots to more accurately understand the complex intentions of human commands and generate responsive physical actions, especially in scenarios involving spatial location, posture, or emotional expression. For example, by combining a gesture instruction to "pick up the cup on the right" with an audio command, the robot can perform tasks more accurately.

[0043] In some embodiments, the method further includes: determining the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence; if the robot executes the action instruction sequence to the next inference step, the motion expert module performs the next inference to obtain the latest action instruction sequence corresponding to the robot, so that the robot executes the latest action instruction sequence. In some embodiments, since the robot executes one of the multiple action instructions sequentially at each time step according to the order of the multiple action instructions in the action instruction sequence, the next inference step corresponding to the action instruction sequence can be obtained by subtracting a preset step threshold (a pre-set empirical value) from the total number of steps in the action instruction sequence (i.e., the number of action instructions in the action instruction sequence). If the robot executes the action instruction sequence to the next inference step (i.e., if the robot executes the action instruction sequence to the action instruction corresponding to the next inference step), the motion expert module performs the next inference. At this time, there are still remaining action instructions in the action instruction sequence that have not yet been executed. That is, while the robot executes the remaining action instructions in the action instruction sequence, the motion expert module performs the next inference to achieve decoupling of inference and execution, i.e., asynchronous inference. In some embodiments, the motion expert module performs the next inference to obtain the latest motion instruction sequence corresponding to the robot, so that the robot begins to execute the latest motion instruction sequence. If there are still unexecuted remaining motion instructions in the previous motion instruction sequence (i.e., the motion instruction sequence), the remaining motion instructions are not executed. In some embodiments, the robot starts executing from the first step of the latest motion instruction sequence, that is, the robot starts executing from the first motion instruction in the latest motion instruction sequence, or the robot starts executing from the Nth step of the latest motion instruction sequence, that is, the robot starts executing from the Nth motion instruction in the latest motion instruction sequence. Here, N = the number of steps (i.e., the number of motion instructions executed) executed by the previous motion instruction sequence (i.e., the motion instruction sequence) from the next inference step to the current time step, or N = the number of steps executed by the previous motion instruction sequence (i.e., the motion instruction sequence) from the next inference step to the current time step + 1. For example, if the motion expert module starts the next inference when the previous motion instruction sequence reaches a0, and the inference ends when the previous motion instruction sequence reaches a4, the latest output motion instruction sequence is from a0' to a... 10In this scenario, the robot has actually executed multiple action commands a0-a4. Then, the robot starts executing from action command a0' in the latest action command sequence, or it starts executing from action command a5' in the latest action command sequence. In other words, action commands a0'-a4' in the latest action command sequence have not actually been executed. This application improves the continuity of the model's output trajectory and its dynamic response to the environment by employing a reasoning-while-executing approach. It enhances the coherence of the action command sequence and ensures the temporal consistency of the output trajectory, thereby improving the robot's dynamic response to the environment.

[0044] In some embodiments, determining the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence includes: determining the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence and the inference delay step corresponding to the action expert module. In some embodiments, the total number of steps in the action instruction sequence is subtracted from the inference delay step corresponding to the action expert module, and the result is used as the next inference step corresponding to the action instruction sequence. Here, the inference delay step refers to the number of steps elapsed from when the action expert module starts inference to when the action expert module outputs the action instruction sequence. That is, determining the timing of starting the next inference based on a fixed trigger threshold of the inference delay step of the action expert module can ensure the smoothness and real-time performance of asynchronous execution. The action expert module starts the next inference in advance in the middle and late stages of executing the previous action instruction sequence (depending on the inference delay step), thereby compensating for the model inference time and avoiding robot pauses or jitter.

[0045] In some embodiments, the step of causing the motion expert module to perform the next inference to obtain the latest motion instruction sequence corresponding to the robot includes: freezing at least one motion instruction preceding the latest motion instruction sequence corresponding to the robot as at least one unexecuted remaining motion in the motion instruction sequence; and causing the motion expert module to perform the next inference based on the frozen latest motion instruction sequence to obtain the complete latest motion instruction sequence. In some embodiments, at least one motion instruction preceding a preset number of the latest motion instruction sequence is 100% frozen as at least one unexecuted remaining motion in the previous motion instruction sequence (i.e., the motion instruction sequence), where the preset number is equal to the inference delay step corresponding to the motion expert module. That is, at least one motion instruction preceding the preset number of the latest motion instruction sequence output by the motion expert module (i.e., at least one frozen motion instruction) remains unchanged, so that the first 6 motion instructions in the latest motion instruction sequence output by the subsequent motion expert module remain unchanged. In some embodiments, the motion expert module performs the next inference based on the frozen latest motion instruction sequence to obtain a complete latest motion instruction sequence. That is, the motion expert module fills in the gaps in the frozen latest motion instruction sequence to obtain a complete latest motion instruction sequence. In other words, the motion expert module only needs to infer the motion instructions following at least one frozen motion instruction. Since at least one frozen motion instruction has already been given, the motion expert module starts inference from the last motion instruction in the previous sequence and continues generating the next motion instruction. In some embodiments, after the robot executes the last motion instruction in the previous sequence, the robot begins executing the first unfrozen motion instruction in the latest sequence. This ensures that the latest motion instruction sequence output by the motion expert module reflects the motion characteristics of the new strategy while maintaining continuity with the original sequence. For example, at least one unexecuted remaining motion in the previous sequence is a10-a15. , The latest action instruction sequence after freezing is [a10, a11, a12, a13, a14, a15, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?]. That is, the action expert module only needs to reason about the last 10 steps of action instructions. The action instructions for the first 6 steps have been given. The action expert module continues to reason from the last state of a15 to generate the next 10 steps of action instructions, that is, to fill in the blanks of the latest action instructions after freezing.

[0046] In some embodiments, the method further includes: setting a freeze weight corresponding to each action in the preceding at least one action instruction, wherein the freeze weight gradually decreases from 1 to 0. In some embodiments, the at least one action instruction preceding a preset number of the latest action instruction sequence is not completely frozen, but is partially frozen using a freeze weight that gradually decreases from 1 to 0, allowing the action expert module to become "more and more free," thereby avoiding the problem of error accumulation caused by complete freezing due to noise and small errors. The freeze weight is a technique in deep learning model training, which refers to fixing the gradient calculation of certain parameters during training so that they do not participate in the update and keep the initial or pre-trained values ​​unchanged.

[0047] In some embodiments, the motion planning module employs a slow analytical reasoning strategy, while the motion expert module employs a fast intuitive reasoning strategy. In some embodiments, based on the Dual Process Theory derived from human cognition, the visual language motion model includes a fast, reactive low-level system (motion planning module) and a slow, reasoning high-level system (motion expert module). That is, the motion planning module employs a slow analytical reasoning strategy, while the motion expert module employs a fast intuitive reasoning strategy. The visual language motion model decouples from the motion planning module and the motion expert module, enabling real-time collaboration and zero-shot generalization (e.g., handling unseen household objects).

[0048] In some embodiments, the method further includes: training the visual language action model based on training data to obtain a trained visual language action model, wherein the training data includes second camera images containing second human body commands captured by the robot, and the training target corresponding to the training data includes a target action command sequence of the robot regarding the second human body commands. In some embodiments, the training data includes second camera images containing second human body commands captured by the robot, and the training target corresponding to the training data includes a target action command sequence of the robot regarding the second human body commands; the visual language action model outputs a corresponding action command sequence based on the training data of the input model, thereby training the visual language action model based on the action command sequence and the target action command sequence (i.e., the training labels corresponding to the action command sequence) to obtain a trained visual language action model.

[0049] In some embodiments, the training data further includes second prompt information corresponding to the second human body command action, the second prompt information including second audio information or second text information. In some embodiments, the training data further includes second prompt information corresponding to the second human body command action, the second prompt information may be manually annotated, or may be obtained from a search engine, or may be generated by a preset content generation engine based on the second human body command action or the second camera image; this example embodiment does not impose any special limitations on this.

[0050] In some embodiments, the method further includes: capturing human movements performed by a trainee in response to the second human body command action, and redirecting the captured human movement data to obtain the target movement command sequence. In some embodiments, a trainee (or data acquisition personnel) wearing a motion capture device (i.e., motion capture equipment) performs corresponding human movements in response to the second human body command action, captures the human movements through the motion capture device, and then redirects the captured human movement data to obtain the target movement command sequence. The motion capture device (MoCap) is a key device used to accurately record human motion data and convert it into a digital model to drive the robot to imitate the corresponding movements. These devices provide the robot with natural, flexible, and efficient motion trajectories by capturing human motion data (including but not limited to joint positions, postures, timing, and mechanical information). Redirection is a technology that maps, transforms, and adapts human movements to the robot's body. Its core objective is to enable the robot to naturally and imitatively reproduce human movements while considering the fundamental differences between the two in terms of physical structure, degrees of freedom, and movement capabilities.

[0051] In some embodiments, the training objective further includes target current state information of the robot regarding the second human body command action; wherein, the method further includes: obtaining the target current state information by causing the robot to execute the target action command sequence. In some embodiments, after the trainee wears a motion capture device, they perform corresponding human movements in response to the second human body command action. The motion capture device captures the human movements, and then the captured human movement data is redirected to obtain the target action command sequence. Then, the motion capture teleoperation device controls the movement of each joint of the robot, causing the robot to execute the target action command sequence, guiding the robot to make the same movements to imitate human movements. The current state information of the robot at each moment during the execution of the target action command sequence is collected and used as the target current state information of the robot regarding the second human body command action. In some embodiments, the visual language action model outputs a sequence of corresponding action instructions based on the training data of the input model and the current state information of the robot. This allows the visual language action model to be trained based on the action instruction sequence and the target action instruction sequence (i.e., the training label corresponding to the action instruction sequence), the current state information of the robot at each moment during the execution of the action instruction sequence, and the current state information of the target at each moment during the execution of the target action instruction sequence (i.e., the training label corresponding to the current state information). This results in a trained visual language action model.

[0052] In some embodiments, the training process of the visual language action model includes a first-stage training and a second-stage training. The training data used in the first-stage training corresponds to multiple different types of second human command actions, while the training data used in the second-stage training corresponds to only one type of second human command action. In some embodiments, the training process of the visual language action model includes two stages. The first-stage training, or pre-training stage, uses training data corresponding to multiple different types of second human command actions. An initial visual language action model is generated by collecting a large number of different data sources. In this stage, a large amount of data from different robot platforms can be input to allow the model to form a preliminary distribution. The second-stage training, or post-training stage, uses second cue information corresponding to second human command actions primarily derived from human annotations. The training data used in the second-stage training corresponds to only one type of second human command action. That is, the second stage is only trained for a specific task, specifically for one of the following tasks: emotional actions, object spatial pointing actions, or sign language actions.

[0053] Figure 2 The flowchart illustrates a method for training a visual language action model according to an embodiment of this application.

[0054] like Figure 2 As shown, a “body language dataset”, i.e. a training dataset, is first created. The dataset consists of three parts: (1) emotional actions, (2) object spatial positioning and pointing, and (3) sign language actions. The emotional actions dataset contains body actions such as anger, surprise, thinking, cheering, and joy, as well as their corresponding audio and text data. The object spatial positioning dataset contains various pictures and language descriptions of pointing to different objects with fingers. The sign language action dataset contains videos of some commonly used sign language words and the corresponding text translations of the sign language words. Then, the dataset is used to pre-train, fine-tune, and quantize the visual language action model (i.e., the VLA model). Finally, the trained visual language action model is deployed and tested.

[0055] Figure 3 This diagram illustrates a computer device structure for outputting a sequence of action instructions according to an embodiment of this application. The computer device includes a first module 11, a second module 12, and a third module 13. The first module 11 is used to input a first camera image containing a first human action instruction acquired by a robot into a visual language action model based on a generative diffusion strategy, to obtain a first task semantic code corresponding to the first human action instruction output by the upper-level visual language model in the visual language action model. The second module 12 is used to enable the action planning module in the visual language action model to perform reasoning based on the first task semantic code and output a corresponding hidden action vector. The third module 13 is used to enable the action expert module in the visual language action model to perform reasoning based on the hidden action vector and the ontology state code corresponding to the robot's current state information, and output a sequence of action instructions corresponding to the robot, so that the robot executes the sequence of action instructions.

[0056] Module 11 is used to input the first camera image containing the first human body command action collected by the robot into the visual language action model based on the generative diffusion strategy, and obtain the first task semantic code corresponding to the first human body command action output by the visual language model located in the upper layer of the visual language action model.

[0057] In some embodiments, the computer device may be the robot itself, or it may be a remote workstation responsible for model reasoning.

[0058] In some embodiments, the robot captures a first camera image in front of it containing a first human command action. The human command action includes, but is not limited to, emotional actions, pointing actions to spatial locations of objects, sign language actions, etc. This example embodiment does not impose any special limitations on this.

[0059] In some embodiments, a Vision-Language-Action Model (VLA) is an AI model that combines visual observation and language commands to directly generate control actions (such as robot joint movements) in the physical world. Generative diffusion strategy refers to treating policy learning as a conditional generation process, progressively denoising in the action space to generate a sequence of action commands that meet the task requirements. The diffusion model, through progressive generation of "adding noise and denoising," can handle multimodal distributions and is suitable for complex action planning. A visual-language-action model based on generative diffusion strategy refers to a visual-language-action model that uses the diffusion model as an "action decoder" (i.e., action expert module) to generate a sequence of robot action commands.

[0060] In some embodiments, a first camera image containing a first human body command action is input into a trained visual language action model, so that the upper-level visual language model (VLM), i.e., a sub-model in the visual language action model, generates a first task semantic code corresponding to the first human body command action. Here, VLM is an artificial intelligence model that can simultaneously understand and process image and text information. It combines the capabilities of computer vision (processing images) and natural language processing (processing text) through deep learning technology to achieve cross-modal interaction between images and language. In some embodiments, VLM includes an image encoder network, which performs task semantic embedding on the first camera image to generate the corresponding first task semantic code. Here, the first task semantic code refers to extracting, understanding, and encoding the core semantics of the user's intention from the first camera image and converting it into an "operation mode" or "response paradigm" that can be executed within the model.

[0061] Module 12 is used to enable the action planning module in the visual language action model to perform reasoning based on the semantic encoding of the first task and output the corresponding hidden action vector.

[0062] In some embodiments, the semantic encoding of the first task is input into the action planning module of the visual language action model, enabling the action planning module to reason based on the semantic encoding of the first task and output the corresponding hidden action vector. The action planning module is responsible for understanding and transforming the semantic encoding of the task into an abstract action representation (i.e., the hidden action vector) that is not yet concrete but already contains the intention to execute. In some embodiments, the visual language action model adopts a fast-slow system architecture, decomposing the entire visual language action model into two closely cooperating subsystems with different operating rhythms and responsibilities: 1. Slow system (upper-level system, i.e., the action planning module): responsible for high-level planning, complex decision-making, and long-term goals. It "thinks deeply but runs slowly." 2. Fast system (lower-level system, i.e., the action expert module): responsible for rapid execution, immediate response, and low-level control. It "responds extremely quickly but is relatively simple." The core advantage of this architecture lies in "using the intelligence of the slow system to guide the fast system, and using the agility of the fast system to ensure the real-time performance and security of the system."

[0063] Module 13 is used to enable the motion expert module in the visual language motion model to perform reasoning based on the hidden motion vector and the ontology state code corresponding to the robot's current state information, and output the motion instruction sequence corresponding to the robot so that the robot executes the motion instruction sequence.

[0064] In some embodiments, the motion expert module (i.e., motion decoder) in the visual language motion model infers the robot's current state based on the hidden motion vectors and the ontology state encoding corresponding to the robot's current state information, and outputs a sequence of motion instructions for the robot. The robot's current state information includes, but is not limited to, the robot's joint angles, angular velocities, torques, etc., at the current moment. This example embodiment does not impose any special limitations on this. The motion expert module is a neural network component responsible for converting high-level instructions or latent (hidden motion vectors in this example embodiment) representations into specific sequences of motion instructions. In some embodiments, the robot's ontology state encoding refers to using a set of data (usually a vector or matrix) to uniquely and accurately describe the robot's current state information. In some embodiments, the motion instruction sequence consists of multiple motion instructions arranged in chronological order. The motion instructions are used to specify the precise quantized signals of the target angle (or position) that each joint of the robot needs to reach in the next time step. The motion instructions include, but are not limited to, the target angle or target position of each joint of the robot. This example embodiment does not have any special limitations on this. In some embodiments, the motion expert module uses flow matching to characterize the motion distribution to determine the real-time inference of the model and the continuity and smoothness of the output motion instruction sequence. Flow matching characterizes the motion distribution during the training process of the motion expert module by teaching the model with demonstration trajectories: starting from "pure noise", learning a smooth path "from random to correct", feeding "current noise point + condition" each time, and the model predicts "the direction of the next push". In the inference process of the motion expert module, starting from pure noise, the model "push" for a number of steps (e.g., 10 to 20 steps) and then outputs a complete motion instruction sequence. In some embodiments, the robot executes one of the multiple action instructions sequentially at each time step according to the order of the multiple action instructions in the action instruction sequence, wherein the robot's underlying joint driver (e.g., servo motor) drives the robot to execute the action instruction at a certain frequency (e.g., 200-400 Hz).

[0065] In some embodiments, the step of inputting a first camera image containing a first human body command action captured by a robot into a visual language action model based on a generative diffusion strategy to obtain a first task semantic code corresponding to the first human body command action output by the upper-level visual language model in the visual language action model includes: inputting a first camera image containing a first human body command action captured by a robot and a first prompting information corresponding to the first human body command action into a visual language action model based on a generative diffusion strategy to obtain a first task semantic code corresponding to the first human body command action output by the upper-level visual language model in the visual language action model and a second task semantic code corresponding to the first prompting information, wherein the first prompting information includes first audio information or first text information; wherein, the step of causing the action planning module in the visual language action model to perform inference based on the first task semantic code and output the corresponding hidden action vector includes: causing the action planning module in the visual language action model to perform inference based on the first task semantic code and the second task semantic code and output the corresponding hidden action vector. Here, the related operations are similar to Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0066] In some embodiments, the first human command action includes at least one of the following: emotional gestures; pointing gestures to a spatial location; sign language gestures. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0067] In some embodiments, the computer device is further configured to: determine the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence; if the robot executes the action instruction sequence to the next inference step, cause the motion expert module to perform the next inference to obtain the latest action instruction sequence corresponding to the robot, so that the robot executes the latest action instruction sequence. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0068] In some embodiments, determining the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence includes: determining the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence and the inference delay step corresponding to the action expert module. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0069] In some embodiments, the step of causing the motion expert module to perform the next inference to obtain the latest motion instruction sequence corresponding to the robot includes: freezing at least one motion instruction preceding the latest motion instruction sequence corresponding to the robot as at least one unexecuted remaining motion in the motion instruction sequence; and causing the motion expert module to perform the next inference based on the frozen latest motion instruction sequence to obtain the complete latest motion instruction sequence. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0070] In some embodiments, the computer device is further configured to: set a freeze weight corresponding to each action in the preceding at least one action instruction, wherein the freeze weight gradually decreases from 1 to 0.

[0071] In some embodiments, the action planning module employs a slow analytical reasoning strategy, while the action expert module employs a fast intuitive reasoning strategy. Here, related operations and... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0072] In some embodiments, the computer device is further configured to train the visual language action model based on training data to obtain a trained visual language action model, wherein the training data includes second camera images containing a second human body command action acquired by the robot, and the training target corresponding to the training data includes a target action command sequence of the robot regarding the second human body command action. Here, related operations and... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0073] In some embodiments, the training data further includes second prompt information corresponding to the second human body command action, wherein the second prompt information includes second audio information or second text information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0074] In some embodiments, the computer device is further configured to: capture human movements performed by a trainee in response to the second human body instruction, and redirect the captured human movement data to obtain the target action instruction sequence.

[0075] In some embodiments, the training objective further includes target current state information of the robot regarding the second human body command action; wherein, the computer device is further configured to: obtain the target current state information by causing the robot to execute the target action command sequence. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0076] Figure 4 Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 4 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.

[0077] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.

[0078] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.

[0079] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0080] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.

[0081] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives).

[0082] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.

[0083] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.

[0084] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).

[0085] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0086] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.

[0087] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.

[0088] This application also provides a computer device, the computer device comprising:

[0089] One or more processors;

[0090] Memory, used to store one or more computer programs;

[0091] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.

[0092] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0093] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0094] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.

[0095] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.

[0096] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.

[0097] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. A method for outputting a sequence of action instructions, wherein, The method includes: The first camera image containing the first human body command action, collected by the robot, is input into the visual language action model based on the generative diffusion strategy to obtain the first task semantic code corresponding to the first human body command action output by the visual language model located at the upper layer of the visual language action model. This enables the action planning module in the visual language action model to perform inference based on the semantic encoding of the first task and output the corresponding hidden action vector. The motion expert module in the visual language motion model performs inference based on the hidden motion vector and the ontology state code corresponding to the robot's current state information, and outputs the motion instruction sequence corresponding to the robot so that the robot executes the motion instruction sequence.

2. The method according to claim 1, wherein, The step of inputting the first camera image containing the first human command action captured by the robot into a visual language action model based on a generative diffusion strategy, and obtaining the first task semantic code corresponding to the first human command action output by the upper-level visual language model in the visual language action model, includes: The first camera image containing the first human body command action and the first prompt information corresponding to the first human body command action are collected by the robot and input into the visual language model based on the generative diffusion strategy of the visual language action model. The first task semantic code corresponding to the first human body command action and the second task semantic code corresponding to the first prompt information are output by the visual language model at the upper layer of the visual language action model. The first prompt information includes first audio information or first text information. The step of enabling the action planning module in the visual language action model to perform inference based on the semantic encoding of the first task and output the corresponding hidden action vector includes: This enables the action planning module in the visual language action model to perform inference based on the semantic encoding of the first task and the semantic encoding of the second task, and output the corresponding hidden action vector.

3. The method according to claim 1 or 2, wherein the first human body command action includes at least one of the following: Emotional actions; The spatial orientation of an object; Sign language gestures.

4. The method according to claim 1, wherein, The method further includes: Based on the total number of steps in the action instruction sequence, determine the next inference step corresponding to the action instruction sequence; If the robot executes the action instruction sequence up to the next reasoning step, the action expert module performs the next reasoning to obtain the latest action instruction sequence corresponding to the robot, so that the robot executes the latest action instruction sequence.

5. The method according to claim 4, wherein, Determining the next inference step corresponding to the action instruction sequence based on the total number of steps in the action instruction sequence includes: The next inference step number corresponding to the action instruction sequence is determined based on the total number of steps in the action instruction sequence and the inference delay step number corresponding to the action expert module.

6. The method according to claim 4 or 5, wherein, The step of enabling the motion expert module to perform the next inference to obtain the latest motion instruction sequence corresponding to the robot includes: Freeze at least one action instruction preceding the latest action instruction sequence corresponding to the robot as at least one remaining action that has not been executed in the action instruction sequence; This allows the motion expert module to perform the next inference based on the frozen latest motion instruction sequence, thereby obtaining the complete latest motion instruction sequence.

7. The method according to claim 6, wherein, The method further includes: Set a freeze weight for each action in the preceding at least one action instruction, wherein the freeze weight gradually decreases from 1 to 0.

8. The method according to claim 1, wherein, The action planning module employs a slow analytical reasoning strategy, while the action expert module employs a fast intuitive reasoning strategy.

9. The method according to claim 1, further comprising: The visual language action model is trained based on the training data to obtain a trained visual language action model. The training data includes second camera images collected by the robot that contain second human body command actions. The training target corresponding to the training data includes the target action command sequence of the robot regarding the second human body command actions.

10. The method according to claim 9, wherein, The training data also includes second prompt information corresponding to the second human body command action, and the second prompt information includes second audio information or second text information.

11. The method according to claim 9 or 10, wherein, The method further includes: By capturing the human movements performed by the trainee in response to the second human body command, and redirecting the captured human body movement data, the target action command sequence is obtained.

12. The method according to claim 11, wherein, The training objective also includes the robot's current state information regarding the second human command action; The method further includes: The current state information of the target is obtained by having the robot execute the target action command sequence.

13. A computer device for outputting a sequence of action instructions, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 12.

14. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 12.