Task processing method and system for embodied robot based on heterogeneous separation model

CN122838113APending Publication Date: 2026-09-29WOCAO TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611343597.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-01
Publication Date
2026-09-29

AI Technical Summary

Benefits of technology

[0021]本申请实施例提供的基于异构分离式模型的具身机器人任务处理方法及系统,通过用户指令以及环境感知数据以确定任务处理模式,在标准处理模式下调用异构部署的第一视觉编码器、第一任务理解器和第一动作解码器、将第一视觉编码任务部署至本地的加速处理器、将第一任务理解器部署至本地的加速处理器与精细处理器协同执行任务理解、将第一动作解码任务分配至本地的精细处理器高精度执行动作解码,能够利用不同计算组件(如加速处理器和/或精细处理器)分别处理与其计算特性相匹配的模型子任务,从而提高具身机器人在本地生成任务执行动作的效率和及时性,降低任务处理过程对云端网络传输状态的依赖;同时,由精细处理器执行动作解码,有利于降低动作生成过程中的精度损失,提高任务执行动作的可靠性。可以理解的是,本申请能够实现标准任务处理模型在资源受限的边缘设备上的高效推理,从而在保证动作精度的前提下,解决相关技术中存在的云端推理延迟不可控、网络断连导致失控等问题,能够提高任务处理的实时性、保障断连环境下的自主运行能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838113A_ABST
    Figure CN122838113A_ABST
Patent Text Reader

Abstract

This application relates to a method and system for task processing of embodied robots based on a heterogeneous, discrete model. The method includes: determining a task processing mode based on user instructions and environmental perception data; if the task processing mode is a standard processing mode, inputting the user instructions and environmental perception data into a standard task processing model; visually encoding the environmental perception data using a first visual encoder deployed in an accelerator processor; performing task understanding on the visually encoded features and user instructions using a task understander deployed in both the accelerator and fine processors; and generating a first task execution action sequence based on task fusion features using a first action decoder deployed in the fine processor, controlling the embodied robot to execute the first task execution action sequence. This method can improve the efficiency and resource utilization of robot control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, and in particular to a method and system for processing embodied robot tasks based on a heterogeneous separable model. Background Technology

[0002] With the rapid development of artificial intelligence and robotics, vision-language-action (VLA) multimodal models have been widely applied in the field of embodied intelligence. A VLA model is an end-to-end model capable of simultaneously processing visual perception, natural language understanding, and robot motion generation. Users only need to input an image and a verbal command into a robot equipped with this model, and the model can directly output the robotic arm's motion trajectory. Due to its unified end-to-end architecture and powerful multimodal understanding capabilities, the VLA model is considered one of the core technological paths to achieving general embodied intelligence.

[0003] VLA models are typically deployed in the cloud. Specifically, the robot itself is only responsible for collecting environmental perception data (such as images) and receiving user commands. This data is uploaded to a cloud GPU server via a wireless network, where the cloud runs the complete inference process of the VLA model, generating action sequences which are then transmitted back to the robot for execution via the network. This cloud-based inference architecture fully leverages the powerful computing capabilities of the cloud, allowing VLA models with large parameter counts (typically 3B to 7B) to complete inference without significant model compression.

[0004] The cloud deployment method described above has at least the problems of low efficiency and difficulty for robots to obtain action execution results in a timely manner when the network is unstable. Summary of the Invention

[0005] Therefore, it is necessary to provide a holographic robot task processing method and system that can improve efficiency and timeliness of holographic robots in obtaining task execution actions, in order to address the above-mentioned technical problems.

[0006] Firstly, this application provides a method for embodied robot task processing based on a heterogeneous separable model, including:

[0007] Acquire user commands and environmental perception data collected by the embodied robot, and determine the task processing mode based on the user commands and environmental perception data;

[0008] If the task processing mode is the standard processing mode, then the user instructions and environmental perception data are input into the standard task processing model; the standard task processing model includes a first visual encoder, a task understander, and a first action decoder.

[0009] The environmental perception data is visually encoded by a first visual encoder deployed in an accelerator processor to obtain first visual encoded features.

[0010] Task fusion features are obtained by performing task understanding on the first visual encoded features and user instructions through task understanders deployed in accelerator processors and fine processors.

[0011] The first action decoder deployed in the fine processor generates a first task execution action sequence based on task fusion features, and controls the embodied robot to execute the first task execution action sequence.

[0012] Secondly, this application also provides a embodied robot task processing system based on a heterogeneous separable model, the system comprising:

[0013] The acquisition module is used to acquire user commands and environmental perception data collected by the embodied robot, and determine the task processing mode for the user commands based on the user commands and environmental perception data.

[0014] The input module is used to input user instructions and environmental perception data into the standard task processing model if the task processing mode is the standard processing mode; the standard task processing model includes a first visual encoder, a task understander, and a first action decoder.

[0015] The encoding module is used to visually encode the environmental perception data using a first visual encoder deployed in the accelerator processor to obtain first visual encoded features;

[0016] The understanding module is used to perform task understanding on the first visual encoded features and user instructions through a task understander deployed in the accelerator and fine processor to obtain task fusion features.

[0017] The generation module is used to generate a first task execution action sequence based on the task fusion features by using a first action decoder deployed in the fine processor, and to control the embodied robot to execute the first task execution action sequence.

[0018] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing method embodiments.

[0019] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0020] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0021] The embodied robot task processing method and system based on a heterogeneous, discrete model provided in this application determines the task processing mode through user instructions and environmental perception data. In the standard processing mode, it invokes a heterogeneously deployed first visual encoder, first task understander, and first action decoder; deploys the first visual encoding task to a local accelerator processor; the first task understander is deployed locally by the accelerator processor and a fine processor to collaboratively perform task understanding; and assigns the first action decoding task to the local fine processor for high-precision action decoding. This allows different computing components (such as accelerator processors and / or fine processors) to process model subtasks that match their computational characteristics, thereby improving the efficiency and timeliness of the embodied robot generating task execution actions locally and reducing the dependence of the task processing process on cloud network transmission status. Simultaneously, having the fine processor perform action decoding helps reduce accuracy loss during action generation and improves the reliability of task execution actions. It is understood that this application can achieve efficient inference of the standard task processing model on resource-constrained edge devices, thereby solving problems such as uncontrollable cloud inference latency and loss of control due to network disconnection in related technologies while ensuring action accuracy. This improves the real-time performance of task processing and ensures autonomous operation capabilities in disconnected environments. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is an application environment diagram of an embodied robot task processing method based on a heterogeneous separation model in one embodiment.

[0024] Figure 2 This is a flowchart illustrating a task processing method for an embodied robot based on a heterogeneous separation model in one embodiment.

[0025] Figure 3 This is a structural block diagram of an embodied robot task processing device based on a heterogeneous separation model in one embodiment.

[0026] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] Before describing the solutions of the embodiments of this application, the relevant technologies and their existing problems will be further explained:

[0029] In related technologies, cloud deployment methods have at least the following problems in practical applications:

[0030] First, the inference latency is uncontrollable, making it difficult to meet the requirements of real-time closed-loop control. The end-to-end latency of VLA model cloud inference typically includes three stages: image upload, cloud inference, and result feedback, with a cumulative time usually between 155 and 335 milliseconds. However, high-frequency closed-loop control of robotic arms typically requires a control cycle of less than 33 milliseconds (corresponding to a 30Hz control frequency). The latency of cloud inference far exceeds this requirement, forcing the robot to perform only open-loop control and preventing trajectory correction based on real-time visual feedback. When the position of objects in the environment changes slightly, open-loop control is prone to causing grasping failure.

[0031] Second, it is highly dependent on network connectivity, and completely loses control in scenarios where the connection is lost. In cloud-based reasoning mode, the robot's decision-making ability depends entirely on a stable network connection. Once the Wi-Fi or cellular network fluctuates or is interrupted, the robot will immediately lose its reasoning ability and be unable to continue performing tasks. In industrial or home service scenarios involving human-robot collaboration, this risk of disconnection could lead to serious safety incidents.

[0032] Third, privacy and cost issues are prominent. The continuous uploading of video footage from cameras to the cloud by the robot poses a risk of user privacy breaches. At the same time, the high operating costs of long-term rental of cloud GPU servers limit the commercialization of VLA models in budget-constrained SMEs and home environments.

[0033] Fourth, traditional edge deployment solutions struggle to balance model size and inference accuracy. To reduce reliance on the cloud, some attempts have been made to deploy VLA models directly on edge devices. However, limited by the edge devices' limited memory (typically 4GB) and computing power (typically around 6 TOPS), significant model compression is necessary. Common practices in related technologies involve using a uniform low-precision quantization scheme (such as INT8 or INT4 quantization) or directly distilling the entire VLA model into a small-parameter model. However, different functional modules within a VLA model exhibit significantly different sensitivities to numerical precision. Uniform low-precision quantization leads to substantial numerical errors in precision-sensitive motion decoding modules, manifesting as millimeter-level pose shifts at the robotic arm's end effector, severely reducing grasping success rates. Conversely, direct distillation into a small-parameter model results in a degradation of language understanding capabilities, making it unable to handle complex natural language commands and negating the core advantages of VLA models.

[0034] Therefore, how to maintain low inference latency while ensuring language understanding and action generation accuracy under the limited resource constraints of edge devices is a technical problem that urgently needs to be solved in the field of embodied intelligence.

[0035] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0036] The embodied robot task processing method based on a heterogeneous separable model provided in this application can be applied to, for example... Figure 1In the application environment shown, the robot body computing device 102 communicates with the edge server 104 via a network. The robot body computing device 102 is an embedded computing platform mounted on the embodied robot, internally deploying various heterogeneous processors, including at least an accelerator processor and a fine-grained processor. An accelerator processor is a computing unit that uses dedicated hardware circuits (such as matrix multiply-accumulate arrays or systolic arrays) to perform instruction-level parallel acceleration of common matrix multiplication operations in neural networks. It excels at handling intensive matrix multiplication and convolution operations, and can complete a large number of multiply-accumulate operations within a unit clock cycle, thereby significantly improving the execution efficiency of computationally intensive tasks. A fine-grained processor is a general-purpose computing unit with a complete instruction set, supporting high-bit-width floating-point operations and complex branch prediction. Although its peak computing power is usually lower than that of an accelerator processor, it can execute various operations with high precision and flexibly handle complex program control flows. In this embodiment, the accelerator processor is an NPU (Neural Processing Unit), and the fine-grained processor is a CPU (Central Processing Unit). It is understood that in other embodiments, the accelerator and the fine processor may also be other types of heterogeneous computing units. For example, the accelerator may also be a matrix multiplication acceleration unit in a GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), or FPGA (Field Programmable Gate Array), and the fine processor may also be a general-purpose processor with floating-point operation capabilities, such as a DSP (Digital Signal Processor) or MCU (Microcontroller).

[0037] The robot's computing device 102 collects environmental perception data through its onboard sensors (such as RGB cameras, depth cameras, force sensors, etc.) and obtains user commands through a microphone or text input interface. It completes the entire VLA model inference process locally; visual encoding, task understanding, and motion decoding are all executed collaboratively on different processors within the robot's computing device 102. For example, after generating a motion sequence, it sends it to the servo driver of the robotic arm for execution via a communication bus (such as a CAN bus or EtherCAT bus). During operation, the robot's computing device 102 can communicate with the edge server 104 via a network to obtain configuration information, report operating status, or download model updates.

[0038] Edge server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Edge server 104 can be deployed at the edge node of the local area network where the robot is located, close to the robot body, with low network latency. Edge server 104 can be used to store and manage global configuration information (such as the address list of point-to-point service nodes in each region, certificate information, etc.), and provide data services when the robot body computing device 102 needs to obtain configuration information or perform firmware / model upgrades.

[0039] It should be noted that the VLA model inference in this embodiment can be completed locally on the robot's computing device 102. The environmental perception data (such as images) and user commands involved do not leave the robot itself and do not need to be uploaded to the cloud or edge server, thus ensuring real-time performance while avoiding privacy risks. The robot's computing device 102 only communicates with the edge server 104 when non-real-time maintenance operations such as configuration synchronization, log reporting, or model upgrades are required. The edge server 104 does not participate in the real-time VLA model inference process; therefore, network latency and disconnections do not affect the robot's basic control capabilities. It is understood that the specific models and performance parameters of the accelerator processor and fine processor of the robot's computing device 102 can be selected according to actual needs.

[0040] In one exemplary embodiment, such as Figure 2 As shown, a method for embodied robot task processing based on a heterogeneous separable model is provided, which is then applied to... Figure 1 Taking the robot body computing device 102 or edge server 104 as an example, the explanation includes the following steps 202 to 206. Wherein:

[0041] Step 202: Obtain user instructions and environmental perception data collected by the embodied robot, and determine the task processing mode for the user instructions based on the user instructions and environmental perception data.

[0042] User commands are task descriptions given by the user to the embodied robot using natural language, such as "pick up the red cup on the table and place it on the tray on the left." User commands can be acquired through voice capture, text input, or other methods. Environmental perception data refers to the surrounding environment information collected by the embodied robot through its onboard sensors, which may specifically include RGB images, depth images, point cloud data, etc. In this embodiment, environmental perception data includes at least RGB images captured by the robot's camera.

[0043] An embodied robot is an intelligent robot system that possesses a physical entity and can perform tasks by interacting with its physical environment. Embodied robots need to perceive environmental changes in real time and respond with physical actions. Their decision-making process is highly sensitive to latency; excessive end-to-end latency from perception to action output will prevent the robotic arm from achieving effective closed-loop control in dynamic environments. Therefore, the task processing methods for embodied robots need to minimize inference latency while ensuring the accuracy of understanding.

[0044] Task processing modes need to adaptively select a balance between computational cost and accuracy based on the complexity of the task itself. Specifically, the amount of semantic information and spatial accuracy requirements of user instructions determine the processing difficulty of the task: multi-step operation instructions, such as "pick up A and put it in B, then pick up C and put it in D," involve complex long-term planning and require the language understanding capabilities of a complete model. High-precision spatial placement instructions, such as "put the cup in the upper left corner of the microwave oven," involve precise coordinate positioning and require high-precision motion generation capabilities. For these complex tasks, computing devices use a standard processing mode, calling a complete VLA model to ensure understanding accuracy and motion precision. For simple repetitive operations, such as continuously grasping the same object, or the fine-tuning stage of the visual servo end effector, where the semantic information is already determined and only rapid visual feedback is needed, the language understanding capabilities of a complete model are not required. A lightweight processing mode can meet the needs, while achieving lower latency and less resource consumption.

[0045] Task processing modes indicate how the robot will handle the current task. Specifically, task processing modes can include standard processing mode and lightweight processing mode. Standard processing mode is suitable for tasks that require precise understanding of complex language instructions, involve multi-step operations, or high-precision spatial placement. In this mode, the full vision-language-motion model will be invoked for processing. Lightweight processing mode is suitable for repetitive and simple operations or vision servoing stages with high latency requirements. In this mode, a lightweight model will be invoked for processing.

[0046] In some embodiments, determining the specific method of the task processing mode based on the user instruction and the environmental perception data further includes: determining the relative distance between the end effector of the embodied robot and the target object based on the environmental perception data; if the relative distance is less than or equal to a distance threshold, determining that the embodied robot is in the visual servo correction stage, and determining the task processing mode as the lightweight processing mode.

[0047] Specifically, the visual servoing correction stage refers to the stage where the embodied robot has completed semantic understanding and main motion planning of the user's instructions, and the end effector is close to the target object or target position. Only minor corrections to the position and posture of the end effector are needed based on real-time visual feedback. In this stage, the semantic content of the language instructions is usually already determined, eliminating the need to repeatedly call the complete task understander for cross-modal semantic fusion. Instead, it is more important to quickly generate local correction actions based on the current environmental perception data to improve the closed-loop control frequency and action response speed. A distance threshold is used to determine whether the embodied robot has entered the visual servoing correction stage. The distance threshold can be preset based on the end effector size, target object size, grasping tolerance, or task scenario. For example, the distance threshold can be 3cm. When the relative distance between the end effector and the target object is determined to be less than or equal to 3cm based on environmental perception data, it indicates that the end effector has entered the fine adjustment range near the target object. At this point, the robot can switch to lightweight processing mode, where a lightweight task processing model quickly generates local correction actions based on the current environmental perception data.

[0048] By switching to lightweight processing mode when the relative distance between the end effector and the target object is less than or equal to the distance threshold, the complete language understanding and motion decoding process can be repeatedly executed during the visual servoing correction stage, reducing inference latency and increasing the motion update frequency during the end effector correction stage. This improves the real-time control and grasping stability of the embodied robot when approaching the target object.

[0049] Step 204: If the task processing mode is the standard processing mode, then input the user instructions and environmental perception data into the standard task processing model; the standard task processing model includes a first visual encoder, a task understander, and a first action decoder.

[0050] The standard task processing model is a complete VLA (Vision-Language-Action) model pre-deployed on the robot's computing device. This model is functionally divided into three independent but collaborative sub-models: a first visual encoder, a task understander, and a first action decoder.

[0051] The first visual encoder extracts features from image information in the environmental perception data, converting raw pixels into high-dimensional visual feature representations. The task understander fuses visual features and textual features of user commands to understand the correspondence between the user's semantic intent and the current scene, outputting multimodal fused features. The first action decoder generates specific robot action sequences based on the fused features.

[0052] It should be noted that the three sub-models mentioned above are deployed on different computing units: the first visual encoder is deployed on an accelerator processor to utilize its efficient matrix operation capabilities; the task understander is deployed on both an accelerator processor and a fine processor to achieve a reasonable allocation of different computational types; and the first action decoder is deployed on a fine processor to ensure the high precision requirements of action generation. This is because: the core operation of the first visual encoder is dense matrix multiplication, which has high data parallelism and is suitable for execution on computing units that excel at accelerating integer matrix multiplication; the task understander includes two types of operations with different properties: matrix multiplication and exponentiation / normalization operations. The former is suitable for execution on matrix multiplication acceleration units, while the latter has high numerical precision requirements and involves complex control flow, making it suitable for execution on general-purpose computing units; and the first action decoder has a small number of parameters but stringent numerical precision requirements, making it suitable for execution on general-purpose computing units that support high-bit-width floating-point operations.

[0053] Accelerated processors are computing units that use dedicated hardware circuits, such as matrix multiply-accumulate arrays and systolic arrays, to perform instruction-level parallel acceleration of matrix multiplication operations commonly used in neural networks. They can complete a large number of multiply-accumulate operations per unit clock cycle, thereby significantly improving the execution efficiency of computationally intensive tasks. Fine-grained processors, on the other hand, are general-purpose computing units with a complete instruction set, supporting high-bit-width floating-point operations and complex branch prediction. Their peak computing power is lower than that of accelerated processors, but they can execute various operations with high precision and flexibly handle complex program control flows.

[0054] Step 206: Visually encode the environmental perception data using a first visual encoder deployed in the accelerator processor to obtain first visual encoded features.

[0055] An accelerator processor is a dedicated computing unit in a computing device, adept at handling intensive matrix multiplication and convolution operations. In this embodiment, the accelerator processor is an NPU (Neural Processing Unit), whose INT8 peak computing power can reach 6 TOPS. The first visual encoder adopts the ViT (Vision Transformer) architecture, whose core operations are self-attention calculation and feedforward network, all of which are intensive matrix multiplication operations, making it very suitable for execution on an accelerator processor.

[0056] Accelerated processors achieve their speed-up capabilities by integrating computational architectures specifically optimized for matrix multiplication. For example, an NPU contains a large array of computing units (such as a MAC array) capable of performing hundreds or even thousands of multiply-accumulate operations in parallel within a single clock cycle. Convolutional and fully connected layers in neural networks are essentially matrix multiplication operations. Accelerated processors can load the weight matrices and input feature matrices into on-chip caches, significantly reducing the number of data transfers between memory and computing units through data pipelines and computational parallelization. This overcomes the "memory wall" bottleneck of traditional processors, resulting in a significant speedup.

[0057] Visual encoding refers to the process of converting raw RGB image pixels into high-dimensional visual token feature vectors. Specifically, the first visual encoder receives a 640×480 pixel RGB image as input, divides the image into multiple image patches, and obtains an initial visual token after linear projection of each patch. Then, semantic information in the image is extracted through multi-layer self-attention computation, finally outputting 576 visual tokens, each of which is a 1152-dimensional vector, i.e., an output tensor with shape [576, 1152]. These visual tokens contain the spatial structure and semantic information of each region in the image, providing a visual foundation for subsequent task understanding.

[0058] The first visual encoding feature refers to the high-dimensional vector sequence extracted from the original image by the visual encoder, used to represent the semantic and spatial structural information of the image. Specifically, the original image exists in the form of a pixel matrix, where each pixel only contains color and brightness information and lacks high-level semantics. The visual encoder gradually aggregates local pixel information into global semantic features through a multi-layer self-attention mechanism and a feedforward network. Each visual token actually integrates the spatial location information of a certain region in the image and its association information with the global context. These visual encoding features exist in the form of a vector sequence (576 1152-dimensional vectors in this embodiment), which essentially maps the original image from pixel space to semantic feature space, enabling the subsequent language model to understand the image content in a way that "understands text".

[0059] In some embodiments, before deploying each compression sub-model, a memory budget is allocated based on the available memory of the edge device for model weights, KV cache, inference buffer, and runtime system usage. Each compression sub-model is deployed to its corresponding computing component only if the sum of these usages is less than the available memory of the edge device and a preset runtime redundancy is maintained. This avoids memory overflow caused by the overlap of cache and inference buffer during actual operation.

[0060] Step 208: Using a task understander deployed in an accelerator and a fine processor, perform task understanding on the first visual encoded features and user instructions to obtain task fusion features.

[0061] The task understander is the core language understanding module in a standard task processing model, employing a Transformer decoder architecture. The task understander typically has 2B parameters, making it the sub-model with the largest number of parameters in the entire model. The core computation of the task understander is the layered computation of the Transformer. Each layer includes matrix multiplication operations in the multi-head attention mechanism, linear transformations in the feed-forward network, as well as Softmax normalization and RMSnorm layer normalization operations.

[0062] Task understanding refers to the process of cross-modal semantic alignment and fusion of visual perception information and natural language instructions. The user instructions received by the embodied robot are abstract goal descriptions in natural language (such as "pick up the red cup"), while the visual encoder outputs specific visual features such as the position, shape, and color of objects in the image. The core challenge of task understanding lies in establishing a mapping relationship between the abstract linguistic concept ("red cup") and specific visual features (which area in the image corresponds to the red cup), i.e., "cross-modal alignment." The task understander establishes association weights between visual token sequences and text token sequences through an attention mechanism, enabling the model to know which areas in the image and which words in the user instruction should be focused on when generating actions.

[0063] The task understander employs a Transformer decoder architecture, with each layer containing two core computational stages: the first stage is multi-head self-attention computation, used to establish dependencies between elements within the input sequence; the main operations in this stage are matrix multiplication (QKV projection, attention score calculation); the second stage is a feedforward network and normalization, used to perform nonlinear transformations and numerical stabilization on the attention output, including matrix multiplication, exponentiation, and division. From a numerical computation perspective, matrix multiplication has a high tolerance for quantization errors because, during the multiplication and accumulation process, small numerical errors are distributed and cancel each other out across a large number of operations. However, the exponentiation operation in Softmax is more sensitive to the input values; even small deviations in the input can lead to several times the difference in the exponentiation result, thus requiring execution on a high-precision floating-point processor.

[0064] The fine-precision processor possesses a complete floating-point unit (FPU), natively supporting half-precision (FP16), single-precision (FP32), and even double-precision (FP64) floating-point operations under the IEEE 754 standard. Furthermore, the fine-precision processor exhibits complex control flow processing capabilities such as out-of-order execution and branch prediction, enabling efficient execution of programs containing conditional statements and loops. The fine-precision processor can flexibly control the computational flow while maintaining high precision, ensuring numerical accuracy while avoiding the waste of computational power by the accelerator processor in non-matrix operations.

[0065] The output of task understanding is task fusion features, which simultaneously encode "what the user wants to do" (intent information from language instructions) and "what objects are in the environment and where they are" (spatial information from visual perception), providing complete multimodal contextual information for subsequent action generation.

[0066] In this embodiment, the task understander is deployed across two types of processors: its matrix multiplication operations (including QKV projection, output projection, and linear layers of the feedforward network) are deployed on an accelerator processor, utilizing the NPU's INT8 computing power for efficient execution; its sensitive operations (including Softmax exponential normalization and RMSnorm layer normalization) are deployed on a fine processor, executing with FP16 high precision. The reason for deploying Softmax and RMSnorm on the fine processor is that these two types of operations involve exponential and division operations, and at low precision, numerical errors are significantly amplified, affecting the model's understanding accuracy.

[0067] The task understanding process includes: the task understander receives the visual tokens output by the first visual encoder and the text tokens output by the text segmenter, concatenates them as the input sequence, and processes them layer by layer through the Transformer decoder. In each layer, matrix multiplication operations such as QKV projection, attention score calculation, and feedforward network are performed on the NPU, while Softmax normalization (converting the attention score into a probability distribution) and RMSnorm layer normalization are performed on the CPU with FP16 precision. The final output is a multimodal feature that integrates visual and linguistic information, i.e., task fusion feature.

[0068] Step 210: Using the first action decoder deployed in the fine processor, a first task execution action sequence based on the task fusion features is generated, and the embodied robot is controlled to execute the first task execution action sequence.

[0069] Among them, the fine processor is a general-purpose computing unit in the computing device, which is good at handling control flow logic and high-bit-width floating-point operations. Its floating-point operation capability is sufficient to meet the computing requirements of motion decoding. The first motion decoder adopts the FlowMatching decoder architecture, with a parameter size of about 300M, and is responsible for mapping the task fusion features into specific robotic arm motion sequences.

[0070] The motion decoder is deployed on a fine processor rather than an accelerator processor because motion generation involves coordinate prediction in continuous space, which is sensitive to numerical accuracy. For a 7-DOF robotic arm with an arm length of about 0.8m, a 0.1-degree error in the joint angle will result in a displacement deviation of about 1.4mm at the end effector. If the error accumulates in each of the 16 steps of the motion, the end effector offset can reach more than 5.6mm, which is enough to cause the grasping to fail.

[0071] Specifically, the first motion decoder receives the task fusion features and performs iterative denoising processing using an 8-layer Transformer plus denoising MLP to generate a 16-step 6-DOF motion sequence (x / y / z displacement plus roll / pitch / yaw rotation), i.e., an output motion sequence matrix of shape [16, 6] as the first task execution motion sequence. The computing device converts the first motion sequence into drive commands for each joint of the robotic arm, controlling the robotic arm to perform the corresponding movements.

[0072] The aforementioned embodied robot task processing method based on a heterogeneous discrete model obtains device address information to determine the task processing mode of user instructions, calls heterogeneously deployed visual encoders, task understanders, and action decoders in the standard processing mode, assigns visual encoding tasks to accelerators, assigns task understanding tasks to accelerators and fine processors for collaborative execution, and assigns action decoding tasks to fine processors for high-precision execution. This enables efficient inference of the VLA model on resource-constrained edge devices, thereby solving problems such as uncontrollable cloud inference latency, loss of control due to network disconnection, high privacy costs, and severe degradation of unified quantization accuracy in related technologies while ensuring action accuracy. It can improve the real-time performance of task processing, ensure autonomous operation capability in disconnected environments, reduce privacy risks and operating costs, and take into account both language understanding ability and action generation accuracy.

[0073] In some embodiments, after determining the task processing mode regarding user instructions, the method further includes:

[0074] If the task processing mode is lightweight processing mode, then user commands and environmental perception data are input into the lightweight task processing model; the lightweight task processing model includes a second visual encoder and a second action decoder.

[0075] The environmental perception data is visually encoded by a second visual encoder deployed in an accelerator processor to obtain second visual encoded features.

[0076] The second action decoder, deployed in the fine processor, generates a second task execution action sequence based on the second visual encoding features, thereby controlling the embodied robot to execute the second task execution action sequence.

[0077] Lightweight processing mode is a simplified task processing approach compared to standard processing mode. It skips the core language understanding module and directly maps visual perception into action output, without performing cross-modal semantic fusion on user commands. In lightweight processing mode, user commands only serve as trigger or confirmation signals; their specific semantic content does not participate in the model's inference process. The model generates actions solely based on the current visual input, without parsing semantic information such as object names and spatial relationships from user commands.

[0078] In the actual operation of embodied robots, a large number of operations are repetitive and programmatic actions. For example, in an assembly line scenario, the robot continuously performs the action of "grabbing a workpiece from the conveyor belt and placing it into a packaging box." Although the user instruction is the same each time it is grabbed (or no new instruction is needed), the position of the workpiece may shift slightly each time, requiring the robot to adjust its grasping posture based on the current visual perception. At this time, the semantic content of the user instruction is already determined, eliminating the need for repeated language understanding; the robot can quickly generate the action based on visual perception. Therefore, the lightweight task processing model is designed to contain only two sub-models: a second visual encoder and a second action decoder. The second visual encoder is used to extract visual features from image information in the environmental perception data, and its function is the same as the first visual encoder in the standard mode. The second action decoder is used to directly generate robot action sequences based on visual encoded features. The difference between the second action decoder and the first action decoder in the standard mode is that the second action decoder does not receive fused features from the language backbone, but directly predicts actions from visual features. Therefore, its network structure can be more lightweight and its inference latency is lower.

[0079] In this embodiment, the applicable conditions of the lightweight processing mode are not simply determined by the complexity of the task, but are based on the criterion of semantic determinism: when the robot has clearly defined "what to do" (the semantics have been completed by the standard processing of the previous stage, or the task itself is a known repetitive operation), and only needs to solve "how to do it" (generate specific action trajectories according to the current visual environment), the lightweight processing mode can be adopted without repeatedly calling the complete language understanding backbone.

[0080] The second visual encoder can share the same network architecture and weight parameters as the first visual encoder in the aforementioned embodiments. In this embodiment, the second visual encoder reuses the model structure and compressed training weights of the first visual encoder. Its core operations are the self-attention calculation and feedforward network of the ViT architecture, all of which are dense matrix multiplication operations, and are deployed on an accelerated processor for execution.

[0081] The specific process of visual encoding is consistent with the aforementioned embodiments and will not be repeated here. It should be noted that since the second visual encoder shares the same network structure as the first visual encoder, its execution time on the accelerated processor is also basically the same, about 25ms. However, since the lightweight processing mode skips the task understander, which is the longest-running module (about 35ms), the overall inference latency is significantly reduced.

[0082] The second action decoder can employ the same decoding architecture as the first action decoder in the aforementioned embodiments, or it can use a more lightweight network structure. In this embodiment, the second action decoder uses a lightweight action network trained by knowledge distillation. This network is trained by distillation using a complete VLA model as the teacher model, and has approximately 100M parameters, only one-third of the number of parameters in the first action decoder (approximately 300M). The lightweight design of the second action decoder enables it to generate actions with lower latency and less memory usage.

[0083] Similar to the first motion decoder, the second motion decoder is deployed and executed on a fine processor. The second motion decoder receives the visually encoded features output from the second visual encoder and directly performs iterative denoising processing using a lightweight Transformer with a denoised MLP to generate a 16-step, 6-DOF motion sequence (x / y / z displacement plus roll / pitch / yaw rotation). This outputs a motion sequence matrix of shape [16, 6] as the second task execution motion sequence, which is then used to control the embodied robot to execute the second task execution motion sequence. Due to the significantly reduced number of parameters in the lightweight motion network, its inference latency is reduced from approximately 18ms for the first motion decoder to approximately 10ms.

[0084] In the aforementioned lightweight processing mode, by skipping the language understanding backbone module, reusing the visual encoder, and combining it with a lightweight motion decoder, inference latency can be significantly reduced in semantically clear task scenarios while maintaining a high grasping success rate, meeting the real-time requirements of high-frequency control scenarios such as visual servoing. The lightweight processing mode complements the standard processing mode, enabling the robot to adaptively achieve an optimal balance between accuracy and speed when facing tasks of varying complexity, ensuring both accurate understanding of complex tasks and efficient execution of simple tasks.

[0085] In some embodiments, if the task processing mode is a standard processing mode, the method further includes, before inputting user instructions and environment-aware data into the standard task processing model:

[0086] Obtain the original visual language action model, and split the original visual language action model into multiple sub-models according to the module function type; the multiple sub-models include a first visual encoder, a task understander, and a first action decoder.

[0087] The target compression strategy for each sub-model is determined based on its compression sensitivity; compression sensitivity characterizes the degree to which the model output accuracy is affected by compression.

[0088] Based on the target compression strategy of each sub-model, the corresponding sub-model is compressed to obtain the compressed sub-model corresponding to each sub-model.

[0089] Based on the computational resource requirements of each compression sub-model and the computational resource configuration characteristics of multiple candidate computational components, the target computational components corresponding to each compression sub-model are determined, and each compression sub-model is deployed on the corresponding target computational components; the target computational components include accelerated processors and fine processors.

[0090] The original visual language action model (i.e., the original VLA model) refers to an end-to-end model pre-trained on the cloud or a high-performance computing cluster, containing complete visual encoding, language understanding, and action decoding functions. While the original VLA model is functionally complete and highly accurate, its number of parameters and computational demands far exceed the capabilities of embedded devices, making it difficult to run directly on resource-constrained edge devices. Therefore, the deployment process described in this embodiment is necessary to convert it into an efficient inference model adapted for edge devices.

[0091] In some embodiments, after decomposing the original visual language action model into a first visual encoder, a task understander, and a first action decoder, the data transfer interface between adjacent sub-modules is standardized. The standardized definition includes at least one of the following: input tensor dimension, output tensor dimension, data type, quantization parameters, and cache identifier. Through the standardized interface, the first visual encoder, task understander, and first action decoder can perform model conversion, compilation optimization, quantization processing, and deployment verification respectively, enabling each sub-module to be compiled, quantized, and deployed independently to different computing components, reducing the coupling of the complete VLA model when deployed on edge devices.

[0092] Based on the different functional responsibilities of each sub-model in the VLA model, the original model is divided into multiple independent and deployable sub-models along the functional boundaries: a first visual encoder responsible for visual perception, a task understander responsible for language understanding and cross-modal fusion, and a first action decoder responsible for action generation. These three are connected through standardized data interfaces. The interface definitions clearly define the shape and data type of the tensors passed between the modules. For example, the shape of the visual token tensor output by the visual encoder to the task understander is [576, 1152], and the fusion feature tensor output by the task understander to the action decoder is the corresponding multimodal feature representation.

[0093] Compression sensitivity is a quantitative metric that measures how well a submodel tolerates compression operations (such as quantization). It characterizes the degree to which the output accuracy of a submodel decreases relative to its original accuracy after a specific compression strategy is applied. Lower compression sensitivity indicates a higher tolerance for compression, allowing for more aggressive compression strategies to achieve higher storage and computational efficiency; higher compression sensitivity indicates a lower tolerance for compression, requiring more conservative compression strategies to protect its output accuracy.

[0094] It should be noted that the sensitivity to compression varies significantly among sub-models of different functional types. This difference stems from the different data types and task natures handled by each sub-model. Specifically, the first visual encoder performs a dimensionality reduction mapping from pixel space to semantic feature space, and its output is a high-dimensional sparse semantic feature vector. A small amount of numerical error is diluted during the dimensionality reduction process, thus it has a high tolerance for compression. The task understander performs cross-modal semantic alignment, and its core operation is attention weight allocation. There is some redundancy among discrete semantic symbols (text and visual tokens), so it has a moderate tolerance for compression. The first action decoder performs coordinate prediction in continuous action space, and its output directly controls the physical movement of the robotic arm. Numerical errors accumulate and amplify during the iterative generation of action sequences, so it has the lowest tolerance for compression and requires the most conservative compression strategy.

[0095] The target compression strategy is a specific compression scheme determined based on compression sensitivity for each sub-model. The target compression strategy specifically includes quantization precision (i.e., the bit width of the compressed numerical representation) and quantization method (i.e., the method of performing the compression operation). By configuring differentiated compression strategies for sub-models with different compression sensitivities, an optimal balance can be achieved between overall compression ratio and accuracy preservation.

[0096] Model compression refers to the operation of reducing model size and computational complexity through specific techniques. Model compression can be implemented through various techniques, including but not limited to model quantization, parameter pruning, low-rank decomposition, and knowledge distillation. In this embodiment, model compression employs model quantization, which maps model weights and activation values ​​from high-bit-width floating-point numbers (such as FP32) to low-bit-width integers or floating-point numbers, thereby reducing the number of bits required to represent each parameter, achieving the goal of reducing model storage size and accelerating inference computation. The compressed sub-model is significantly smaller, enabling it to be loaded and run on memory-constrained edge devices.

[0097] A compressed sub-model refers to an executable model that is smaller in size but functionally equivalent, obtained by applying a target compression strategy to the original sub-model. The compressed sub-model retains the network structure and inference logic of the original sub-model, but the numerical representation of weight parameters and / or activation values ​​is changed, enabling it to run efficiently on heterogeneous computing units.

[0098] Computational resource requirement characteristics refer to the description of the computational resource requirements of the compressed sub-model during inference. Specifically, this may include at least one of the following: operation type characteristics, computational power requirement parameters, and precision preservation requirement parameters. Operation type characteristics describe the main categories of computational operations during the sub-model's inference process. Different types of operations have different hardware requirements: matrix multiplication operations are characterized by high data parallelism and regular computational patterns, making them suitable for execution on hardware with large-scale multiply-accumulate arrays; exponential normalization operations involve exponential function calculations, have high numerical precision requirements, and complex data dependencies, making them suitable for execution on general-purpose processors that support high-precision floating-point operations. For example, the main operation of the first visual encoder is matrix multiplication, the task understander includes both matrix multiplication and exponential normalization operations, and the main operation of the first action decoder is continuous-space iterative optimization operations.

[0099] The computational power requirement parameter describes the computational throughput requirements of the sub-model. Sub-models with high computational power requirements need to be deployed on processors with high peak computing power to avoid becoming bottlenecks in overall inference performance. For example, the first visual encoder and task understander have a large number of parameters and high computational load, thus requiring high computational power.

[0100] The accuracy preservation requirement parameter describes how sensitive a submodel is to numerical precision. Submodels with high accuracy preservation requirements need to maintain high bit-width floating-point precision during inference to avoid unacceptable deviations in the output due to numerical errors. For example, the first action decoder involves iterative prediction of continuous action coordinates, where numerical errors accumulate and amplify during iterations, thus requiring the highest level of accuracy preservation.

[0101] Candidate computing components refer to different types of processor units that can be deployed in a computing device; in this embodiment, they include accelerator processors and fine processors. Each candidate computing component has its specific computing resource configuration characteristics, which may include at least one of the following: available computing power parameters, supported computing precision types, and preferred operation types.

[0102] Available computing power parameters describe the upper limit of computation that a computing unit can complete per unit of time. This parameter is typically measured in TOPS (trillions of integer operations per second) or GFLOPS (billions of floating-point operations per second). Computing units with higher peak computing power are more suitable for handling computationally intensive tasks. For example, the accelerator processor in this embodiment has a peak INT8 computing power of 6 TOPS.

[0103] Supported computation precision types describe the data precision formats that the computing unit natively supports at the hardware level. Different computing units may support different precision formats. Some computing units are specifically optimized for low-bit-width integer operations (such as INT8 and INT4), enabling them to perform integer matrix multiplication with high throughput; others support the full IEEE 754 floating-point standard (such as FP16, FP32, and FP64), enabling them to perform various floating-point operations with high precision. The precision types supported by a computing unit directly determine which sub-models can be deployed on that unit for efficient operation.

[0104] The term "operation type of expertise" describes the optimizations a computing unit makes at the hardware architecture level for specific types of operations. For example, the NPU integrates a large-scale matrix multiply-accumulate array (MAC array), capable of performing hundreds of multiply-accumulate operations in parallel within a single clock cycle; therefore, its expertise lies in intensive matrix multiplication. The CPU, on the other hand, possesses a complete instruction set and out-of-order execution capabilities, enabling it to efficiently handle complex control flow programs containing branching, looping, and function calls; therefore, its expertise includes control flow logic and high-bit-width floating-point operations.

[0105] Deploying the compressed sub-models to the corresponding target computing components means allocating appropriate computing units to each compressed sub-model to perform its inference task based on the degree of matching between the computing resource requirements of each sub-model and the computing resource configuration characteristics of each computing component. In this embodiment, the first visual encoder performs intensive matrix multiplication and has high computing power requirements, so it is deployed on an accelerated processor that is good at matrix multiplication; the task understander includes two types of operations: matrix multiplication and exponential normalization, with its matrix multiplication part deployed on an accelerated processor and its exponential normalization part deployed on a fine processor; the first action decoder has high precision requirements and is deployed on a fine processor that supports high-bit-width floating-point operations.

[0106] In some embodiments, the target compression strategy includes at least one of quantization precision and quantization method; determining the target compression strategy for each sub-model based on the compression sensitivity of each sub-model includes:

[0107] If the compression sensitivity of the first visual encoder is less than the first threshold, then the first quantization precision is determined as the quantization precision of the first visual encoder, and the post-training quantization method is determined as the quantization method of the first visual encoder.

[0108] If the compression sensitivity of the task understander is greater than the first threshold and less than the second threshold, the first quantization precision is determined as the quantization precision of the feature calculation layer in the task understander, and the second quantization precision is determined as the quantization precision of the sensitive layer in the task understander, and the group quantization method is used as the quantization method of the task understander; the second quantization precision is greater than the first quantization precision.

[0109] If the compression sensitivity of the first action decoder is greater than the second threshold, then the half-precision floating-point number is determined as the quantization precision of the first action decoder.

[0110] The compression sensitivity value can be obtained through pre-tested sensitivity. Specifically, different degrees of quantization compression are applied to each sub-model, and the change in output accuracy of the sub-model on the standard test set before and after compression is measured. The accuracy attenuation is used as the quantization index of compression sensitivity. The first threshold is a preset sensitivity threshold value used to distinguish between sub-models that are highly tolerant of compression and those that are moderately sensitive to compression. When the compression sensitivity of a sub-model is less than the first threshold, it indicates that the sub-model has a high tolerance for compression operations and can use a lower bit width quantization accuracy and a simpler quantization method.

[0111] The first visual encoder performs a dimensionality reduction mapping from pixel space to semantic feature space, outputting a high-dimensional semantic feature vector. During this mapping process, a large amount of detailed information in the input image (such as texture and lighting variations) is gradually abstracted and discarded by the network layers, retaining only the key information relevant to semantic understanding. Therefore, a small amount of numerical error is diluted by this information compression during propagation layer by layer, ultimately having a very limited impact on the output semantic features. Based on this principle, the first visual encoder can employ a relatively aggressive compression strategy.

[0112] The first quantization precision refers to the target numerical bit width used when performing quantization operations on the first visual encoder, which in this embodiment is 8-bit integer precision (INT8). In this embodiment, post-training quantization means that when quantizing the model, there is no need to retrain or fine-tune the model. Instead, the numerical distribution range of the weights of each layer is statistically analyzed based on a small amount of calibration data. Quantization mapping parameters (such as scaling factors and zeros) are determined based on this distribution range, and then the weights are mapped from floating-point numbers to integers. The advantages of post-training quantization are its simplicity, low computational cost, and lack of additional training resources and data, making it suitable for sub-models with high quantization tolerance.

[0113] The second threshold is higher than the first threshold and is used to distinguish between moderately sensitive and highly sensitive sub-models. When the compression sensitivity of a sub-model is between the first and second thresholds, it indicates that the sub-model has a moderate tolerance for compression and needs to seek a balance between compression ratio and accuracy preservation, adopting a compromise compression strategy.

[0114] The compression sensitivity of task understanders falls between that of visual encoders and action decoders because they handle cross-modal semantic alignment tasks. In this task, the model needs to establish attention weight allocation relationships between visual and text token sequences, a computational process involving numerous correlation calculations between discrete semantic symbols. The discrete semantic space possesses a degree of redundancy, allowing the model to tolerate minor numerical errors in weight parameters and activation values. However, unlike the dimensionality reduction operation of visual encoders, the semantic alignment results in task understanders directly affect the correctness of subsequent action generation (if semantic alignment is incorrect, such as focusing on "red" on "tray," the entire task fails). Therefore, its tolerance is lower than that of visual encoders but higher than that of action decoders.

[0115] In this embodiment, group quantization refers to a strategy that divides the weight tensor into multiple subgroups, with each subgroup independently calculating quantization parameters (scaling factor and zero point) and performing quantization. Specifically, for the weight matrix in the task understander, it is divided into M weight parameter groups according to a preset grouping size (e.g., 128 weights per group), where M is a positive integer. For each group, the maximum and minimum values ​​of the weights within that group are calculated, and the scaling factor and zero point corresponding to that group are determined based on the range of these values. Then, each weight within that group is mapped from a floating-point number to an integer. It is understandable that the advantage of group quantization is that the numerical distribution in different regions of the same weight tensor may vary significantly. Global quantization can lead to insufficient quantization accuracy in regions with small numerical ranges and numerical overflow in regions with large numerical ranges. Group quantization, on the other hand, can adaptively adjust the mapping parameters for the numerical distribution of each local region, thereby achieving higher representation accuracy with the same bit width and effectively reducing quantization errors.

[0116] When the compression sensitivity of a sub-model exceeds the second threshold, it indicates that the sub-model is highly sensitive to compression. Even small numerical errors under low bit-width quantization can lead to severe degradation of the output results, thus requiring a high-precision preservation strategy. The first action decoder has the highest compression sensitivity because action decoding generates a sequence of coordinates in a continuous action space. Unlike visual coding and language understanding, where visual coding errors are one-off (misidentification of an object may occur but can be corrected in subsequent processing) and semantic errors have a tolerance space in discrete space, action errors accumulate and amplify during iterative denoising. Small errors at each step become the input for the next step, and errors propagate and accumulate along the time axis. Therefore, the action decoder has the highest requirement for numerical precision and requires the most conservative compression strategy.

[0117] The third precision refers to the target numerical bit width used when performing quantization operations on the first action decoder. Its precision level is higher than that of the first precision (INT8) and the second precision (INT4). In this embodiment, the third precision is 16-bit floating-point precision (FP16). The first action decoder is stored and run with FP16 precision, without performing low-bit-width integer quantization. Compared with INT8, FP16 has a larger dynamic range and higher numerical resolution, which can guarantee the numerical precision of each iteration in the action decoding process, so that the precision loss of the final generated action sequence is controlled within an acceptable range. When the first action decoder is stored with FP16 precision, it occupies approximately 600MB (the original FP32 is 1.2GB), which is slightly larger than the INT8 quantization size, but the precision retention rate can reach over 99.9%. The reason for choosing FP16 instead of FP32 is that FP16 has half the storage volume of FP32, executes faster on processors that support FP16 hardware acceleration, and its accuracy is sufficient to meet the needs of robotic arm motion control. FP16 has an error of less than 0.01 degrees (end-effector offset less than 0.5mm), which is far lower than the typical joint angle error of about 0.2 degrees (end-effector offset more than 10mm) introduced by INT8 quantization. Therefore, storing the motion decoder with FP16 accuracy is the optimal balance between storage efficiency and motion accuracy.

[0118] In some embodiments, the feature calculation layer includes weight parameters and activation parameters, the first quantization precision includes four-bit quantization and eight-bit quantization, and the second quantization precision is half-precision floating-point; the corresponding sub-models are compressed according to the target compression strategy of each sub-model to obtain the compressed sub-models corresponding to each sub-model, including:

[0119] For the task understander, the weight parameters are grouped based on the group quantization method to obtain M weight parameter groups; the number of weight parameters in each weight parameter group is less than or equal to the preset number, and M is a positive integer;

[0120] Based on four-bit quantization, the scaling factor and zero point corresponding to each weight parameter group are determined;

[0121] Based on the scaling factor and zero point corresponding to each weight parameter group, the weight parameters in the corresponding weight parameter group are quantized to obtain the quantized weight parameters.

[0122] Based on 8-bit quantization, the activation parameters are quantized to obtain quantized activation parameters. Based on the quantized weight parameters and quantized activation parameters, the quantized feature calculation layer is determined.

[0123] The parameters in the sensitive layer are quantized using half-precision floating-point numbers to obtain the quantized sensitive layer. Based on the quantized feature calculation layer and the quantized sensitive layer, the compressed sub-model corresponding to the task understander is determined.

[0124] The weight parameters refer to the trainable parameters of each linear transformation layer in the task understander (including the QKV projection layer, output projection layer, and linear layers of the feedforward network), stored in matrix form. In the Transformer decoder architecture, the weight parameters account for the vast majority of the total parameters of the task understander, and their storage volume determines the model's memory footprint. Therefore, efficient compression of the weight parameters is key to reducing the memory consumption of the task understander.

[0125] Group quantization refers to a strategy that divides the weight tensor into multiple independent subgroups, with each subgroup calculating its own quantization parameters and performing the quantization operation independently. Considering that the numerical distribution of weight parameters in different layers of a neural network is not globally uniform, weights in different layers may have different numerical ranges, and even weights in different channels or regions within the same layer may have different distribution characteristics. If a uniform quantization parameter is applied to the entire weight tensor (i.e., global quantization), regions with smaller numerical ranges will suffer from insufficient resolution after quantization, while regions with larger numerical ranges may experience numerical overflow. Group quantization, by adapting independent quantization parameters to each local region, can achieve higher representation accuracy with the same bit width.

[0126] All weight parameters in the task understander are divided according to a preset group size. The preset group size is a predetermined constant representing the number of weights in each weight parameter group. In this embodiment, the preset group size can be 128, meaning that every 128 weight parameters are divided into one group. Other group sizes (such as 64, 256, etc.) can also be selected according to actual needs. Smaller group sizes result in higher quantization accuracy but also higher storage overhead for quantized parameters, while larger group sizes result in lower storage overhead but may reduce quantization accuracy. M is the number of groups obtained by dividing the total number of weight parameters by the group size. For example, if the total number of weight parameters in the task understander is approximately 2B (2 billion parameters), and the group size is 128, then M is approximately 15.625 million groups. In actual implementation, for cases where the result is not divisible, zeros can be added to the last group or the remaining parameters can be directly retained as a group.

[0127] Four-bit quantization (INT4 quantization) refers to mapping each weight parameter from its original floating-point value (FP32 or FP16) to a 4-bit integer (representing 16 discrete values ​​from 0 to 15). 4-bit integers offer a higher compression ratio than 8-bit integers; the storage size after INT4 quantization is only one-eighth that of FP32 and half that of INT8. However, 4-bit integers can only represent 16 discrete values, meaning that all floating-point weight values ​​within each group must be mapped to one of these 16 discrete integers, and the mapping precision depends on the actual distribution range of the weight values ​​within that group.

[0128] The scaling factor and zero point are used to establish a one-to-one correspondence between floating-point and integer values. The scaling factor determines the floating-point interval corresponding to adjacent integer units after quantization. If the weight values ​​have a wide distribution range, the scaling factor is larger, resulting in coarser quantization resolution; if the weight values ​​have a narrow distribution range, the scaling factor is smaller, resulting in finer quantization resolution. The zero point determines the integer position corresponding to the zero value of the floating-point number after quantization, and is used to handle cases where the weight distribution is asymmetric about the zero point.

[0129] Weighted quantization refers to the operation of converting the floating-point value of each weight parameter into its corresponding 4-bit integer representation based on a determined scaling factor and zero point. Specifically, for each weight parameter, the floating-point value is divided by the scaling factor, the zero point is added, and then rounded to the nearest integer to obtain the corresponding INT4 quantized value. Because weighted quantization involves division and rounding operations, the quantized weight cannot be completely restored to its original floating-point value, introducing a certain amount of quantization error. The core value of grouped quantization lies in reducing quantization error by narrowing the range of values ​​covered. The narrower the range of values ​​covered by each group, the smaller the scaling factor, the higher the quantization resolution, and the smaller the quantization error.

[0130] Activation parameters refer to the activation values ​​output by each layer of the task understander during the forward propagation process, including attention scores and intermediate feature vectors. Unlike weight parameters, the numerical range and distribution of activation parameters vary with the input data, exhibiting greater dynamism and thus greater sensitivity to quantization errors. In this embodiment, activation parameters employ 8-bit quantization (i.e., INT8 quantization). Activation parameter quantization refers to the quantization operation that maps the floating-point activation values ​​(FP16 or FP32) generated during the forward propagation process to an 8-bit integer representation. Unlike weight quantization, activation quantization cannot pre-calculate the global distribution because the activation values ​​depend on the specific input image and text content. Therefore, activation quantization typically employs dynamic quantization or static quantization based on a calibration set. Before deployment, a set of representative calibration data (including multiple robot operation scene images and corresponding text instructions) is input into the task understander to collect the numerical distribution of activation values ​​at each layer, statistically determine the dynamic range of activation values ​​at each layer, and then determine the corresponding quantization parameters (scaling factor and zero point) for each layer. During inference, the activation values ​​at each layer are quantized into INT8 in real time according to the pre-determined quantization parameters for subsequent calculations.

[0131] The quantized feature computation layers refer to the various network layers in the task understander. The weight parameters in these layers are stored in INT4 format, and the activation parameters participate in the computation in INT8 format. This achieves dual optimization of storage and computation during inference. Storing weights in INT4 significantly reduces the model's memory footprint and loading bandwidth, while using INT8 for matrix multiplication calculations fully utilizes the INT8 computing power of the accelerator processor. During computation, the INT4 weight parameters are dequantized to INT8 (or directly used in the computation as INT4, depending on hardware support) when loaded into the computation unit. Matrix multiplication is then performed with the INT8 activation parameters, and the output is further quantized before being used as input for the next layer.

[0132] The layers containing exponentiation and division operations specifically include the Softmax exponentiation normalization layer and the RMSnorm normalization layer. A common characteristic of these two types of operations is that exponentiation or division is highly sensitive to small changes in the input value. In Softmax, the exponentiation function amplifies the output by a factor of e for every 1 increase in the input; in RMSnorm, the division operation exhibits a sharp increase in numerical instability when the denominator is small. In low-bit-width integer arithmetic, this sensitivity leads to the exponential amplification of quantization errors, severely impacting the model's accuracy in understanding the data.

[0133] In some embodiments, each sub-model is compressed according to a target compression strategy to obtain a compressed model corresponding to each sub-model, including:

[0134] Each sub-model is compressed according to the target compression strategy to obtain the initial compressed model corresponding to each sub-model.

[0135] Obtain a quantized perception training dataset for quantized perception training of the first visual encoder and the task understander; wherein, the quantized perception training dataset includes robot running scene images, text instructions corresponding to the robot running scene images, and robot running scene images with action labels marked by action sequence labels;

[0136] Simulated quantization nodes are inserted into the forward propagation paths of the initial compressed model corresponding to the first visual encoder and the initial compressed model corresponding to the task understander, respectively, to obtain the first inserted model corresponding to the first visual encoder and the second inserted model corresponding to the task understander; wherein, the simulated quantization nodes are used to simulate the execution of quantization mapping operation and dequantization mapping operation.

[0137] The samples in the robot's running scene image quantization perception training dataset are input into the first post-insertion model to obtain visual perception features. The visual perception features and the text features corresponding to the text instructions are input into the second post-insertion model to obtain fused multimodal features.

[0138] The fused multimodal features are input into the initial compressed model corresponding to the first action decoder to obtain the predicted action sequence. The loss value is calculated based on the predicted action sequence and the ground truth value of the action sequence label.

[0139] Backpropagation is performed on the first and second post-insertion models based on the loss values ​​to update the weight parameters in the first and second post-insertion models, thereby obtaining the compressed sub-model corresponding to the first visual encoder and the compressed sub-model corresponding to the task understander. The initial compressed model corresponding to the first action decoder is then determined as the compressed sub-model corresponding to the first action decoder.

[0140] The fused multimodal features are input into the initial compressed model corresponding to the first action decoder to obtain the predicted action sequence for forward propagation. The simulated quantization output is obtained through the simulated quantization node. The loss value is calculated based on the predicted action sequence output and the true value of the action sequence label corresponding to the sample.

[0141] Backpropagation is performed on the first and second post-insertion models based on the loss values ​​to update the weight parameters in the initial compressed models of the first and second post-insertion models, thereby obtaining the updated compressed sub-models corresponding to the first visual encoder and the task understander. The initial compressed model corresponding to the first action decoder is then determined as the compressed sub-model corresponding to the first action decoder. The updated compressed sub-model is used to replace the initial compressed model and is deployed to the corresponding target computing component to perform inference tasks.

[0142] The initial compressed model refers to the model obtained directly after quantization compression of each sub-model, without further accuracy optimization. Specifically, the first visual encoder is trained and quantized using INT8 to obtain the initial compressed version; the task understander is grouped and quantized using INT4 for weights and INT8 for activation to obtain the initial compressed version; and the first action decoder retains FP16 accuracy (no accuracy loss occurs, no quantization-aware training is required). Although these initially compressed models are significantly smaller and can be loaded and run on edge devices, the output accuracy of the models is somewhat reduced compared to the original FP32 model due to the numerical discretization error introduced by the quantization operation. For example, the accuracy retention rate of the task understander after W4A8 quantization is approximately 97.8%. The purpose of quantization-aware training is to recover this lost accuracy through subsequent fine-tuning training.

[0143] The quantization perception training dataset is a collection of samples specifically designed for quantization perception training. Each sample in this dataset contains an RGB image of a robot's operating scene, along with corresponding ground truth values ​​for action labels and text instructions. Ground truth values ​​for action labels refer to the sequence of target actions that the robotic arm's end effector should execute in that scene, which may include annotations such as end-effector pose, joint angles, or operational commands. Text instructions can be natural language descriptions, such as "Place the red square on the blue tray," used for task instruction.

[0144] This dataset is typically much smaller than the pre-training dataset for the original VLA model because the purpose of quantization-based perceptual training is only to fine-tune the model to adapt to the changes in numerical distribution brought about by quantization, rather than learning visual semantic understanding or language understanding capabilities from scratch. Therefore, large-scale training data is not required. The dataset can be sourced from imitation learning trajectory data collected during actual robot operation, or synthetic data generated through simulation environments. The dataset should cover various typical scenarios that the robot may encounter, including different lighting conditions, different object placements, and different background environments, to ensure that the quantization-based perceptual training model maintains good generalization performance in various real-world scenarios.

[0145] Simulated quantization nodes are special operation nodes inserted into the forward propagation computation graph of the model. Their function is to simulate the effect of quantization operations in the forward propagation path. Quantization mapping operations are used to map input floating-point values ​​to low-precision integer representations according to quantization parameters (scaling factor and zero point); dequantization mapping operations map the quantized integer values ​​back to floating-point values.

[0146] During training, the model's forward propagation encounters values ​​with quantization errors, enabling it to perceive the impact of quantization operations on its output. In the backpropagation path, since the simulated quantization nodes are inherently non-differentiable, a straight-through estimator (STE) technique is typically used. This approximates the gradient of the quantization node as 1 (assuming the quantization operation has no effect on the gradient), allowing the gradient to bypass the quantization node and propagate back to previous layers, thus updating the weight parameters. In this way, the model can progressively adjust its weight values ​​during training, ensuring that even with quantization errors, the model's output accuracy remains at a high level.

[0147] The visual perception features are the feature representations output by the first post-insertion model after visually encoding the input robot operation scene image. Specifically, they are 576 visual tokens, each token being a 1152-dimensional vector with a shape of [576, 1152]. The generation process of these features is as follows: the image is input into the first post-insertion model, and after passing through a multi-layer self-attention computation and feedforward network of the ViT architecture, the final output is a sequence of visual tokens.

[0148] Text features are the text token sequences and their corresponding embedded representations obtained after the text instructions are processed by a tokenizer. Visual perception features and text features are jointly input into a second post-insertion model. This second post-insertion model (i.e., the post-insertion version of the task understander) performs cross-modal semantic alignment and fusion through multi-layer computation of the Transformer decoder in a simulated quantization environment, outputting fused multimodal features. These fused features simultaneously encode information about "what the user wants to do" and "what objects are in the environment and where they are located."

[0149] The initial compressed model corresponding to the first action decoder refers to the compressed version of the first action decoder saved with FP16 precision (no quantization precision loss occurs, and no simulated quantization node needs to be inserted). The fused multimodal features are input into this model, and after iterative denoising processing by the Flow Matching decoder, the predicted action sequence is output, such as an action sequence matrix of shape [16, 6], representing 16 steps and 6 degrees of freedom actions. The loss value is a quantitative metric that measures the difference between the model's prediction and the true label. In this embodiment, the loss function uses action regression loss, specifically the mean squared error (MSE) between the model's predicted action sequence and the true action label. By minimizing the loss value, the model can gradually learn weight parameters that maintain accuracy in a quantized environment.

[0150] Backpropagation refers to the process of calculating the gradients (partial derivatives) of each trainable parameter (weight parameter) in the model with respect to the loss value based on the loss value, and then adjusting the parameter values ​​in the opposite direction of the gradients to reduce the loss value. Specifically, the gradient is calculated layer by layer from the output layer to the input layer using the chain rule. Each layer calculates the gradient of its parameters based on its input and output, and then propagates the gradient forward. For the simulated quantization node, due to the use of the pass-through estimator technique, the quantization operation is treated as an identity mapping, and the gradient can continue to propagate forward through this node without decay.

[0151] The weight parameters are updated using an optimizer (such as AdamW) that adjusts each parameter based on the calculated gradient values ​​and the learning rate. In this embodiment, quantization-aware training uses a learning rate much smaller than the original pre-training rate (approximately one percent of the original VLA model's pre-training learning rate, i.e., on the order of 1e-5) to ensure that the model does not change drastically during fine-tuning, and only minor adjustments are made to the weights to adapt to the changes in numerical distribution caused by quantization.

[0152] The compressed sub-models corresponding to the first visual encoder and the task understander replace their respective initial compressed models and are deployed to their respective target computing components (accelerated processor, and accelerated processor and fine processor, respectively) to perform the actual inference tasks. At this point, the compression, training, and deployment of all three sub-models are complete, and they are ready to be called upon in standard processing mode to perform inference tasks.

[0153] In some embodiments, the computing resource requirement characteristics include at least one of operation type characteristics, computing power requirement parameters, and precision preservation requirement parameters; the computing resource configuration characteristics include at least one of supported computing precision types and processable operation types; based on the computing resource requirement characteristics of each compression sub-model and the computing resource configuration characteristics of multiple candidate computing components, the target computing component corresponding to each compression sub-model is determined, and each compression sub-model is deployed to the corresponding target computing component, including:

[0154] The computational resource requirement characteristics of each compression sub-model are matched with the computational resource configuration characteristics of multiple candidate computational components to obtain the matching results.

[0155] If the matching result indicates that the operation type feature of the first visual encoder matches the operation type that the accelerator processor can process, the accelerator processor is identified as the target computing component corresponding to the first visual encoder.

[0156] If the matching result indicates that the operation type characteristics of the matrix multiplication operation part in the task understander match the operation type that the accelerator can process, and the precision retention requirement parameters of the precision-sensitive operation part in the task understander match the computational precision type supported by the fine processor, then the accelerator and the fine processor are identified as the target computing components corresponding to the task understander.

[0157] If the matching result indicates that the accuracy retention requirement parameter of the first action decoder matches the type of computational accuracy supported by the fine processor, the fine processor is identified as the target computational component corresponding to the first action decoder.

[0158] The first visual encoder employs the ViT (Vision Transformer) architecture, with core operations consisting of matrix multiplication in the multi-head self-attention mechanism and linear transformations in the feedforward network. Matrix multiplication operations exhibit high data parallelism; the multiplication between the weight matrix and the input feature matrix can be decomposed into a large number of independent multiply-add operations. These operations are independent of each other and can be executed in parallel across multiple computational units. The inference process of the first visual encoder is essentially composed entirely of this dense matrix multiplication, and its operation type is defined as dense matrix multiplication.

[0159] The accuracy preservation requirement parameter describes the sub-model's sensitivity to numerical accuracy. As mentioned earlier, visual encoding is a dimensionality reduction operation that maps the original image from pixel space to semantic feature space. Small numerical errors are diluted during dimensionality reduction, ultimately having a limited impact on the accuracy of semantic features. Therefore, the accuracy preservation requirement parameter for the first visual encoder is set to medium.

[0160] Multiple candidate computing components refer to different types of processor units that can be deployed in a computing device, including accelerators and fine processors in this embodiment. The types of operations that the accelerator can handle include fast matrix multiplication, while the types of computational precision supported by the fine processor include high-bit-width floating-point operations, and the types of operations that can be handled include control flow processing.

[0161] Matching refers to comparing the operation type characteristics, computing power requirements, and accuracy preservation requirements of each compression sub-model with the processing operation types and supported computing accuracy types of each candidate computing component to determine whether they are compatible. The criteria for suitability are: if the computing resource requirements of the sub-model are within the capabilities of the computing component, and the specialized operation type of the computing component happens to match the main operation type of the sub-model, then the match is successful; otherwise, the match fails.

[0162] The first-order visual encoder's computational type is characterized by dense matrix multiplication, while the accelerated processor can handle computations including fast matrix multiplication. Dense matrix multiplication exhibits high data parallelism and can be decomposed into a large number of independent multiply-accumulate operations executed in parallel. The accelerated processor integrates a large-scale array of computing units (such as a MAC array), capable of performing hundreds of multiply-accumulate operations in parallel within a single clock cycle, and its hardware architecture is highly compatible with the computational pattern of matrix multiplication. Therefore, when the matching result indicates that the two are a match, the accelerated processor is identified as the target computing component for the first-order visual encoder.

[0163] The specific method for deploying the first visual encoder to the accelerator is as follows: The compressed model file corresponding to the first visual encoder (containing network structure description and INT8 quantized weight parameters) is loaded into the dedicated memory space of the accelerator. A corresponding inference session is created on the accelerator, configuring the shape of the input tensor ([1, 3, 640, 480]) and the shape of the output tensor ([1, 576, 1152]), and setting the quantization parameters. When the accelerator receives the input RGB image data, it performs self-attention calculation and feedforward network operations according to the configured computation graph, outputting the visual encoded feature tensor.

[0164] The task understander employs a Transformer decoder architecture, and its computation process involves two distinct categories of operations. The matrix multiplication operation includes QKV projection (projecting input features into query, key, and value matrices through a linear transformation), attention score calculation (multiplying the transpose of the query matrix and the key matrix), output projection (restoring the attention-weighted result to its original dimensions through a linear transformation), and linear transformation of the feedforward network (two fully connected layers). These matrix multiplication operations account for approximately 90% of the overall computation of the task understander, exhibiting high data parallelism and regular computational patterns, matching the computational strengths of the accelerated processor.

[0165] The precision-sensitive operations include Softmax exponential normalization and RMSNorm layer normalization. Softmax involves calculating the exponential function e^x for each input element, summing all exponent values, and dividing each exponent by the sum to obtain the probability distribution. RMSNorm involves calculating the root mean square of the input tensor, dividing each element by the root mean square, and then multiplying by a scaling factor. Both types of operations involve complex mathematical functions such as exponentiation, division, and square root operations, making them extremely sensitive to numerical precision. Even minute changes in the input score in Softmax can lead to significant differences in the exponential results, thus altering the allocation of attention weights. Therefore, the precision requirement is set to "high," and the fine-grained processor supports FP16 high-bit-width floating-point operations, natively supporting various floating-point operation hardware instructions under the IEEE 754 standard, matching this precision requirement. Simultaneously, the fine-grained processor possesses a complete instruction set and out-of-order execution capabilities, enabling efficient processing of control-flow programs containing branching, looping, and function calls, making it suitable for executing operations like Softmax and RMSNorm that involve complex mathematical functions and conditional judgments.

[0166] Based on the above dual matching results, the matrix multiplication part is matched with the accelerated processor, and the precision-sensitive operation part is matched with the fine processor. The accelerated processor and the fine processor are jointly determined as the target computing components of the task understander. The deployment method is as follows: the network structure of the task understander is divided into a linear transformation part and a nonlinear normalization part. The linear transformation part is deployed to the accelerated processor, and the nonlinear normalization part is deployed to the fine processor. Data is transferred between the two parts through shared memory.

[0167] The accuracy requirement for the first motion decoder is set to "extremely high" because motion generation involves coordinate prediction in continuous space. For a 7-DOF robotic arm with an arm length of approximately 0.8m, a 0.1-degree error in the joint angle can result in a displacement deviation of approximately 1.4mm at the end effector. If errors accumulate in each of the 16 motion steps, the end effector offset can reach over 5.6mm, which is sufficient to cause grasping failure. The fine processor supports FP16 high-bit-width floating-point operations, enabling it to perform floating-point operations with high precision and possessing the control flow capability to handle iterative algorithms. Therefore, the accuracy requirement of the first motion decoder matches the accuracy type supported by the fine processor, making the fine processor the target computing component for the first motion decoder.

[0168] In the aforementioned heterogeneous deployment method, by matching the computational resource requirements of each compression sub-model with the computational resource configuration characteristics of candidate computing components, and based on the matching results, the visual encoder is deployed to the accelerator processor, the matrix multiplication part and precision-sensitive operation part of the task understander are deployed to the accelerator processor and the fine processor respectively, and the action decoder is deployed to the fine processor, thus achieving optimal matching between the three types of sub-models and the two types of processors. Compared with the scheme of deploying all modules on the same processor, this method can fully utilize the strengths of each computing unit without increasing hardware costs, achieving the optimal level of overall inference latency, while ensuring task understanding accuracy and action generation accuracy. In some embodiments, the task understander includes a linear transformation layer and a nonlinear normalization layer, with the linear transformation layer deployed in the accelerator processor and the nonlinear normalization layer deployed in the fine processor;

[0169] Task fusion features are obtained by using task understanders deployed in accelerators and fine processors to perform task understanding on the first visual encoded features and user instructions, including:

[0170] By using a linear transformation layer deployed in an accelerator processor, linear transformation processing is performed on the first visual encoded features and user instructions to obtain multimodal semantic features for action decision-making;

[0171] By deploying a nonlinear normalization layer in a fine processor, exponential normalization and layer normalization operations are performed on the multimodal semantic features to obtain task fusion features.

[0172] The linear transformation layer refers to the network layer in the task understander that performs matrix multiplication operations. This includes the QKV projection layer in the multi-head self-attention mechanism (projecting the input features into a query matrix Q, a key matrix K, and a value matrix V), the attention score calculation layer (multiplying the transposes of Q and K), the output projection layer (restoring the dimensionality of the attention weighted result through linear transformation), and the two fully connected layers in the feedforward network. A common characteristic of these linear transformation layers is that their computation process can be uniformly represented as a matrix multiplication operation Y = X·W + b, where X is the input feature matrix, W is the weight matrix, and b is the bias term.

[0173] The linear transformation layer, deployed in the accelerator processor, takes as input the first visual encoded features (576 visual tokens, each 1152-dimensional) output from the first visual encoder in the preceding steps, and the text token sequence processed by the tokenizer. The first visual encoded features and the text token sequence are concatenated into a unified input sequence, serving as the input to the first layer of the task understander. In the computation of each Transformer layer, the linear transformation layer receives the output of the previous layer (or the first layer receives the concatenated input sequence) as input and performs matrix multiplication operations. The matrix multiply-accumulate array of the accelerator processor can perform a large number of multiply-accumulate operations in parallel within a single clock cycle, enabling these computationally intensive operations to be completed within approximately 25ms (accounting for the majority of the overall inference time of the task understander). The intermediate results and final output of these linear transformations are stored in shared memory in INT8 or FP16 format for use by subsequent nonlinear normalization layers.

[0174] A nonlinear normalization layer refers to a network layer in a task interpreter that performs nonlinear normalization operations. This includes the Softmax exponential normalization layer in multi-head self-attention mechanisms and the RMSnorm normalization layers before and after each Transformer layer. These nonlinear normalization layers share the common characteristic of involving complex mathematical functions such as exponential operations (e^x), division operations, square root operations, and mean operations. They are highly sensitive to numerical precision, and errors are significantly amplified at low precision.

[0175] While accelerated processors (APCs) offer high efficiency for matrix multiplication at INT8 precision, they lack native hardware support for exponentiation, division, and square root operations. These operations require simulation on APCs, resulting in slow speed and inconsistent accuracy. Fine-grained processors, on the other hand, possess a complete floating-point unit (FPU) that natively supports various floating-point instructions under the IEEE 754 standard. Their precision and efficiency in performing exponentiation, division, and square root operations are superior to the software simulations on APCs. Furthermore, the exponentiation function in Softmax operations is highly sensitive to input values; even small deviations in the input can lead to significant differences in the exponentiation result, altering the entire attention probability distribution and affecting the accuracy of cross-modal semantic alignment. Therefore, it is necessary to execute these operations at high precision (FP16) on a fine-grained processor.

[0176] The output of the nonlinear normalization layer, after multi-head self-attention weighted fusion and layer normalization, becomes the task fusion feature vector. This feature encodes both visual and linguistic information, specifically including: which image regions correspond to which objects in the user's command (represented by attention weights), the spatial relationships between objects, and the type of operation the user intends to perform. The task fusion feature is stored in shared memory in FP16 format for the first action decoder deployed on the fine-grained processor to read and generate action sequences.

[0177] In some embodiments, the environmental perception data is multiple frames, and the first task execution action sequence includes task execution actions corresponding to the multiple frames of environmental perception data. Controlling the embodied robot to execute the first task execution action includes:

[0178] During the processing of continuous multi-frame environmental perception data, the first visual encoder, task understander and first action decoder are controlled to execute in a parallel inter-frame pipeline mode; the parallel inter-frame pipeline mode is a way to process different sub-models of different frames and execute them in parallel on different computing components.

[0179] The output interval between adjacent frames is determined based on the longest time value among the first visual encoder, task understander, and first motion decoder.

[0180] Based on the output interval between adjacent frames, the task execution actions corresponding to each frame of environmental perception data are output sequentially to control the embodied robot to execute the first task execution action sequence; wherein, the first visual encoder, task understander and first action decoder transmit frame data between adjacent units through a lock-free data queue.

[0181] In this context, continuous multi-frame environmental perception data refers to a sequence of RGB images continuously acquired by the embodied robot at a fixed frequency (e.g., 30Hz) during its continuous operation. Each frame of the image needs to undergo complete reasoning through a standard task processing model, including visual encoding, task understanding, and action decoding, before generating the corresponding action instructions. In the traditional serial processing method, each frame must wait for all three sub-models of the previous frame to complete before processing can begin. In this method, the processing delay of each frame is the sum of the times taken by the three sub-models.

[0182] The core idea of ​​inter-frame pipeline parallelism lies in leveraging the deployment characteristic of the three sub-models located in different computational units. While the subsequent module of a given frame is executing, the preceding module of the next frame is started in advance. Specifically, when the first action decoder of frame N performs action decoding on the fine processor, the first visual encoder of frame N+1 can simultaneously perform visual encoding on the accelerator processor. This parallelism is possible because the first visual encoder and the first action decoder are deployed on different types of processors—the former on the accelerator processor and the latter on the fine processor. They do not compete for the same hardware resources during computation, and therefore can execute independently within the same time period.

[0183] Once the visual encoding and task understanding of a frame are completed, the action decoding of that frame begins execution on the fine processor, triggering the start of visual encoding for the next frame on the accelerator processor. Since the accelerator processor enters an idle state after completing the matrix multiplication part of the task understanding for the current frame, and is originally idle during the action decoding phase of frame N (executed by the fine processor), this idle computing power can be effectively utilized by controlling the accelerator processor to begin processing the visual encoding of frame N+1 during this period.

[0184] The adjacent frame output interval refers to the time interval between the output of two consecutive action sequences after the pipeline enters a steady state. In an ideal pipeline, the interval between the action decoding output time of frame N+1 and the action decoding output time of frame N depends on the time consumed by the slowest stage (i.e., the bottleneck stage) in the pipeline. Because after the pipeline reaches stable operation, each stage is busy in each clock cycle (i.e., each frame cycle), and the system throughput is determined by the slowest stage.

[0185] A lock-free data queue is a data structure used to transfer frame data between sub-models. In this embodiment, a first lock-free data queue (SPSC queue, i.e., single producer-single consumer queue) is established between the first visual encoder and the task understander, and a second lock-free data queue is established between the task understander and the first action decoder. The characteristic of a lock-free data queue is that no locks are required during data enqueueing and dequeueing operations in a multi-threaded or multi-processor environment, avoiding the performance overhead and deadlock risk caused by lock contention. Because the queue adopts an SPSC design, it only supports one write end and one read end, eliminating the need for complex synchronization mechanisms. Enqueueing and dequeueing operations are atomic operations, enabling efficient data transfer between different stages of the pipeline.

[0186] Specifically, after completing visual encoding, the first visual encoder writes the visual token tensor into the first lock-free data queue. The task understander reads data from this queue, performs task understanding, and then writes the fused features into the second lock-free data queue. The first action decoder reads data from the second queue and performs action decoding. These three sub-models each run independently in a pipeline rhythm, asynchronously transferring frame data through the lock-free data queues.

[0187] In this embodiment, a first lock-free data queue is established between the first visual encoder and the task understander, and a second lock-free data queue is established between the task understander and the first action decoder. After completing visual encoding, the first visual encoder writes the visual token tensor into the first queue. The task understander reads data from this queue, performs task understanding, and then writes the fused features into the second queue. The first action decoder reads data from the second queue and performs action decoding. Because the queues adopt a lock-free design, they only support one write end and one read end, eliminating the need for complex synchronization mechanisms. Both enqueue and dequeue operations are atomic operations, enabling efficient data transfer between different stages of the pipeline.

[0188] In some embodiments, after deploying each compression sub-model to its corresponding target computing component, the method further includes:

[0189] Acquire the historical visual information processed by the task understander and the cached data corresponding to the historical language instructions, and quantize and compress the cached data for storage;

[0190] During the task understanding process of first visual encoded features and user instructions by task understanders deployed in accelerator processors and fine processors, the current scene image and current language instructions are acquired.

[0191] The current scene image is compared with historical visual information. If the comparison results are consistent, the cached data corresponding to the historical visual information is reused.

[0192] The current language instruction is compared with the historical language instruction. If the comparison results are consistent, the cached data corresponding to the historical language instruction is reused.

[0193] Here, cached data refers to the Key-Value cache (KV Cache) generated by the task understander (Transformer decoder architecture) during inference, used to store the attention key-value pair information of processed tokens. In the self-attention mechanism of the Transformer decoder, for each token in the input sequence, the model needs to calculate its corresponding Key vector and Value vector, which are repeatedly used in the attention calculation of subsequent layers. Without caching, each time an output token is generated, the Key and Value of all previous tokens need to be recalculated, resulting in a computational complexity of O(n²) and low inference efficiency. By caching the Key and Value of already calculated tokens, subsequent inference only needs to calculate the Key and Value of newly added tokens, reducing the computational complexity to O(n).

[0194] In some embodiments, the KV Cache generated by the task understander is compressed and stored. Specifically, the FP16 format key cache and value cache are quantized into INT8 format for storage. When the task understander needs to read cache data for subsequent calculations, the INT8 format cache data is dequantized into FP16 format as needed according to the corresponding quantization parameters to participate in attention calculations. This reduces the cache memory usage of the task understander during continuous frame inference.

[0195] The cached data corresponding to historical visual information refers to the key-value cache generated by the 576 visual tokens output by the first visual encoder during the robot's processing of historical frame images, across various layers of the task understander. The cached data corresponding to historical language commands refers to the key-value cache generated by the text tokens of the user's current language command after word segmentation, across various layers of the task understander. Quantization and compression storage of the cached data refers to the operation of compressing the FP16 format KV cached data into INT8 format for storage.

[0196] Since the task understander needs to repeatedly read and write to the KV cache during each frame of inference, and the memory of the edge device (4GB) is relatively limited, compressing the KV cache is of great significance. The specific method for quantization compression is as follows: for each value in the KV cache, its numerical distribution range is statistically analyzed, and quantization parameters (scaling factor and zero point) are determined based on this range. Then, the FP16 values ​​are mapped to INT8 integers for storage. In INT8 format, each element occupies only 1 byte, which is half the storage size of FP16's 2 bytes. The KV cache corresponding to the visual token is compressed from approximately 47.7MB to approximately 23.8MB. During inference, when the cached KV values ​​need to be read for attention calculation, the INT8 cache is dequantized to FP16 precision as needed, thus trading a small loss of precision for a significant reduction in memory usage.

[0197] The current scene image refers to the environmental RGB image captured by the embodied robot in the current frame, which is processed by the first visual encoder to generate corresponding first visual encoded features. The current language command refers to the natural language command text currently input by the user, which is processed by word segmentation to generate a corresponding text token sequence.

[0198] Comparison refers to comparing the visual content of the current frame scene with the visual content of historical frames scene to determine whether they are identical or highly similar. Specifically, the comparison method involves calculating hash values ​​(such as MD5 hashes or lighter hash algorithms) for the 576 visual tokens output by the first visual encoder to obtain the visual fingerprint of the current scene. Hash values ​​have the characteristic that the same input produces the same output, and different inputs are likely to produce different outputs, making them efficient for determining whether two sets of visual tokens are consistent. The hash value of the visual token in the current frame is compared with the hash values ​​of the historically stored visual tokens. If the comparison results are consistent, it means that the current scene and the historical scene are the same, and there is no need to recalculate the key-value cache corresponding to the visual tokens.

[0199] Reuse refers to skipping the recalculation of the KV cache corresponding to the current frame's visual token when the comparison results match, and directly using the previously stored historical visual cache data. By reusing the historical visual cache, this part of the calculation can be completely skipped, thereby significantly reducing the inference latency of the task understander.

[0200] Similar to visual cache reuse, the proportion of KV cache computation corresponding to text tokens in the task understander's inference computation varies depending on the text length. When the user instruction is long (such as a complex multi-step instruction), the number of text tokens is large, and the acceleration effect brought by cache reuse is more significant.

[0201] In some embodiments, acquiring user commands and environmental perception data collected by the embodied robot, and determining a task processing mode based on the user commands and environmental perception data, includes:

[0202] Parse the user commands to determine the number of action steps and spatial precision requirements corresponding to the user commands;

[0203] If the number of action steps exceeds the step number threshold or the spatial accuracy requirement is higher than the accuracy threshold, the task processing mode is determined to be the standard processing mode.

[0204] If the number of action steps is less than or equal to the step number threshold and the spatial accuracy requirement is less than or equal to the accuracy threshold, the task processing mode is determined to be lightweight processing mode.

[0205] This process involves semantic analysis of the task descriptions given by users in natural language to extract quantitative indicators related to task complexity. Parsing user instructions can be achieved using natural language understanding technology, specifically including the following sub-steps: segmenting and tagging the user instruction text to identify verbs (representing action types), nouns (representing the objects of operation), and locative and quantifier words (representing spatial location and precision requirements); decomposing the instruction into several atomic action steps based on the identified action types and objects of operation; and determining the spatial precision requirements of the task based on the locative words, quantifier words, and the action type itself.

[0206] The number of motion steps refers to the number of atomic actions contained in the user instruction. An atomic action is the smallest unit of motion that a robot can complete in a single operation. Spatial accuracy requirement refers to the positioning accuracy required for the robot's end effector target position in the user instruction. Methods for determining spatial accuracy requirements may include: identifying whether the user instruction contains precise spatial descriptive words, which imply a higher spatial accuracy requirement; identifying the relative relationship between the size of the target container and the size of the target object in the user instruction; when the container size is only slightly larger than the object size, the operation requires higher spatial accuracy; and identifying quantifiers and locative words in the user instruction, such as explicit spatial numerical descriptions. The result of judging the spatial accuracy requirement can be a discrete level or a continuous numerical value.

[0207] The accuracy threshold is used to distinguish between tasks with general accuracy requirements and those with high accuracy requirements. In this embodiment, the accuracy threshold can be set according to the physical characteristics of the robotic arm and the tolerance requirements of typical tasks. For example, it can be set to 10mm. When the task requires end-effector positioning accuracy higher than 10mm (i.e., tolerance less than 10mm), the standard processing mode is triggered.

[0208] In some embodiments, the process of determining a lightweight task processing model includes:

[0209] The standard task processing model is used as the teacher model, which includes at least a first visual encoder, a task understander, and a first action decoder.

[0210] The first visual encoder in the teacher model is reused as the second visual encoder, and a lightweight action network is constructed as the second action decoder; the number of parameters in the lightweight action network is smaller than the number of parameters in the task understander.

[0211] Using the teacher model's processing results on the training samples as a supervision signal, and freezing the weights of the second visual encoder, knowledge distillation training is performed on the lightweight action network to obtain a lightweight task processing model.

[0212] In this embodiment, the teacher model serves as the source model providing supervisory signals during the knowledge distillation process. In this case, it is the standard task processing model that has been compressed and deployed in the aforementioned embodiments. The teacher model possesses complete visual encoding, language understanding, and action decoding capabilities, enabling it to process complex multi-step language instructions and generate accurate action sequences. The network structure of the first visual encoder in the teacher model and its trained (including compressed and quantized perceptual training) weight parameters can be directly used as the second visual encoder without retraining. The first visual encoder has approximately 400M parameters, a moderate number, and after being compressed, deployed, and quantized perceptually trained in the aforementioned embodiments, it can complete visual encoding on an accelerated processor at a speed of approximately 25ms.

[0213] Lightweight action networks are newly constructed small-scale neural networks whose function is to receive visually encoded features and directly generate action sequences, essentially acting as a second action decoder. The number of parameters in a lightweight action network is designed to be significantly smaller than that of a task understander. Task understanders need to align the semantic information of natural language instructions with visual features across modalities, requiring a large number of parameters to store language knowledge and visual-language mapping knowledge; while lightweight action networks skip the language understanding stage and directly map visual features to actions, thus significantly reducing the number of parameters.

[0214] The training samples consist of consecutive multi-frame image sequences, corresponding language commands for each frame, and ground truth values ​​of the target action sequences that the robotic arm's end effector should execute for each frame. The sample size is approximately 50,000 robot imitation learning trajectory data.

[0215] Knowledge distillation training refers to the process of training a lightweight task processing model using high-quality supervision signals provided by a teacher model. The loss function of knowledge distillation training measures the differences between the lightweight task processing model and the teacher model in terms of action output, feature representation, and action distribution. The loss function of knowledge distillation training includes at least one of action regression loss, intermediate feature alignment loss, and action distribution divergence loss. Unlike traditional supervised learning, the supervision signal in distillation training is the soft output of the teacher model, rather than hard labels provided by humans.

[0216] The action regression loss is used to constrain the action sequence output by the lightweight task processing model to be close to the action sequence output by the teacher model. Specifically, the mean squared error between the first action sequence output by the teacher model and the second action sequence output by the lightweight task processing model can be calculated as the action regression loss. The intermediate feature alignment loss is used to constrain the intermediate features obtained by the lightweight task processing model during action generation to be consistent with the intermediate features of the corresponding layer or stage of the teacher model. Specifically, the distance between the task fusion features or action decoding intermediate features in the teacher model and the action generation intermediate features in the lightweight task processing model can be calculated as the intermediate feature alignment loss. The action distribution divergence loss is used to constrain the action distribution generated by the lightweight task processing model to be close to the action distribution generated by the teacher model. Specifically, the action probability distribution or action sampling distribution output by the teacher model can be used as the target distribution, and the action probability distribution or action sampling distribution output by the lightweight task processing model can be used as the prediction distribution. The divergence between the two can be calculated as the action distribution divergence loss. The action regression loss, intermediate feature alignment loss, and action distribution divergence loss can be weighted and summed to obtain the total loss value for knowledge distillation training, and the model parameters of the lightweight action network can be updated based on the total loss value; among them, the weights of the second visual encoder remain frozen and do not participate in the parameter update.

[0217] In some embodiments, after generating a first task execution action sequence about user instructions based on task fusion features using a first action decoder deployed in a fine processor, the method further includes:

[0218] Obtain actual execution feedback data after the embodied robot performs the action sequence of the first task;

[0219] The sequence of actions for the first task is corrected based on the actual execution feedback data to obtain the corrected task execution actions;

[0220] Control the embodied robot to perform the revised task actions.

[0221] Actual execution feedback data refers to the actual state information collected by the embodied robot through its various sensors after executing the action sequence generated by the first action decoder. This feedback data is used to characterize the actual pose, velocity, torque, and other information of the robotic arm's end effector after executing the action. In this embodiment, actual execution feedback data may include: the actual angles of each joint of the robotic arm, the actual three-dimensional spatial position of the end effector, and the relative distance between the end effector and the target object.

[0222] The actual frequency of feedback data acquisition is usually higher than the inference frequency of the task processing model. For example, the servo control loop of a robotic arm typically reads encoder data and performs position closed-loop control at a frequency of 1 kHz, while the inference frequency of the VLA model is approximately 28 Hz. Therefore, there are multiple opportunities to acquire high-frequency feedback data between two model inferences, which can be used to make local corrections to the current execution action.

[0223] Based on the deviations observed during actual execution, adjustments are made to the remaining motion sequences that have not yet been completed to compensate for the execution errors that have occurred. The specific implementation method of the correction is as follows: calculate the deviation between the actual feedback data and the expected motion state (such as end-effector position deviations Δx, Δy, Δz, and posture deviations Δroll, Δpitch, Δyaw), then map the deviation into compensation amounts for each joint angle through inverse kinematics, and superimpose the compensation amounts onto the joint angle commands of the remaining motion sequence.

[0224] In some embodiments, the training samples include a trajectory sample set, each trajectory sample in the trajectory sample set containing a sequence of consecutive multi-frame images, language instructions corresponding to each frame image, and the true value of the target action sequence that the end effector of the embodied robot should execute in each frame image.

[0225] Using the teacher model's processing results on the training samples as supervision signals, knowledge distillation training is performed on the second visual encoder and the lightweight action network to obtain a lightweight task processing model, including:

[0226] The trajectory sample set is input into the teacher model and the lightweight task processing model respectively to obtain the first action sequence output by the teacher model and the second action sequence output by the lightweight task processing model.

[0227] Update the model parameters of the lightweight task processing model based on the loss value between the first action sequence and the second action sequence.

[0228] The trajectory sample set is a dataset used for knowledge distillation training. Its sources can be imitation learning trajectory data collected during actual robot operation or synthetic data generated through a simulation environment. Each trajectory sample records a complete robot task execution process, specifically including three mutually aligned data dimensions: The first dimension is a continuous multi-frame image sequence, i.e., a sequence of RGB image frames collected at fixed time intervals during robot task execution, recording the continuous changes in the scene during task execution. The second dimension is the language command corresponding to each frame image, i.e., the user's natural language command received by the robot at that frame moment (in a complete trajectory, the language command usually remains unchanged, but to maintain the universality of the data format, each frame of each trajectory records the corresponding language command). The third dimension is the true value of the target action sequence that the embodied robot's end effector should execute under each frame image, i.e., given the image scene and language command, the target action sequence that the robotic arm's end effector should execute, specifically a 16-step, 6-DOF action sequence (x / y / z displacement plus roll / pitch / yaw rotation). This true value can be collected through manual teaching (recording action data when a human operator remotely controls the robotic arm to perform the task) or generated through motion planning algorithms.

[0229] The loss value is a quantitative metric that measures the difference between the first action sequence output by the teacher model and the second action sequence output by the student model. In this embodiment, the loss function is action regression loss, namely the mean squared error (MSE) between the first and second action sequences. The mean squared error is calculated as follows: for each training sample, the square of the difference between each action dimension in the action sequence output by the teacher model and the corresponding position value in the student model is calculated, and the average of the squared differences of all 96 values ​​is taken to obtain the loss value of that sample.

[0230] Based on the calculated loss value, the gradient of each trainable parameter in the lightweight task processing model with respect to the loss value is calculated using the backpropagation algorithm. Then, the parameter values ​​are adjusted along the gradient descent direction to reduce the loss value. Specifically, the AdamW optimizer is used with a learning rate of 5e-4 to update the parameters of the lightweight action network. During the update process, the parameters of the second visual encoder do not participate in the update because they reuse the weights of the first visual encoder of the teacher model and remain frozen. Each training round traverses all approximately 50K trajectory samples. After each round, the model's accuracy on the validation set is calculated. The training is conducted for a total of 100 epochs until the model accuracy converges.

[0231] By inputting trajectory sample sets into the teacher model and the lightweight task processing model respectively, the high-quality action sequences output by the teacher model are obtained as soft supervision signals. The loss value between the teacher output and the student output is calculated, and the student model parameters are updated by backpropagation based on the loss value. This enables the lightweight task processing model to achieve performance close to that of the teacher model with lower inference cost.

[0232] In some embodiments, the lightweight task processing model is trained through knowledge distillation. Specifically, a standard task processing model is used as the teacher model, and a model including a second visual encoder and a lightweight action network is used as the student model. The second visual encoder reuses the weights of the first visual encoder and keeps them frozen, while the lightweight action network generates a sequence of actions for the second task based on the visually encoded features output by the second visual encoder. During training, the lightweight action network is supervised and trained based on the sequence of actions for the first task output by the teacher model, the intermediate features of the teacher model, and the action distribution. This allows the lightweight task processing model to learn the action generation capabilities of the standard task processing model even without using the full task understander.

[0233] In some embodiments, the first visual encoder and / or the second visual encoder can be implemented using different visual feature extraction networks. For example, the first visual encoder can employ a SigLIP architecture or a DINOv2 architecture. The DINOv2 architecture acquires visual representation capabilities through self-supervised pre-training and has good feature representation capabilities in 3D spatial understanding, object boundary perception, or scene structure modeling. As another example, the first visual encoder can include a SigLIP encoding branch and a DINOv2 encoding branch. The SigLIP encoding branch is used to extract visual features with a high degree of alignment with linguistic semantics, while the DINOv2 encoding branch is used to extract visual features related to spatial structure and geometric relationships. The visual features output from the two encoding branches are then concatenated, weighted, or attention-based to obtain the visual encoded features used as input to the task understander.

[0234] In some embodiments, the quantization method of the task understander is not limited to the aforementioned group quantization method; weighted quantization methods such as GPTQ or AWQ can also be used. GPTQ quantization can compensate for quantization errors based on the second-order error information of the weight matrix, thereby reducing the impact of low-bit weighted quantization on the model output accuracy. AWQ quantization can identify weight channels that have a significant impact on the model output based on the activation distribution and employ more refined scaling or protection strategies for these weight channels, thus reducing the loss of task understanding accuracy while maintaining a high compression ratio. By using GPTQ or AWQ to quantize the weight parameters of the task understander, the accuracy preservation rate of the task understander can be improved under the same or similar storage overhead.

[0235] In some embodiments, the target computing component for the first action decoder can be determined based on the floating-point operation support capability of the candidate computing components. If the accelerator processor in the candidate computing component supports half-precision floating-point operations, the first action decoder can be deployed to the accelerator processor to perform half-precision floating-point operations. For example, in an edge computing platform employing an NPU that supports FP16 operations, the first action decoder can be deployed to the NPU for execution, thereby freeing up the computing resources of the fine-grained processor while ensuring the accuracy of action generation.

[0236] In some embodiments, the lightweight task processing model may further include a small language module. The small language module performs lightweight semantic parsing of user commands to obtain lightweight text features; the second action decoder generates a second task execution action sequence based on the second visual encoded features output by the second visual encoder and the lightweight text features. By retaining the small language module in the lightweight task processing model, the lightweight processing mode can handle some complex language commands while maintaining a low number of parameters and low inference latency.

[0237] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0238] Based on the same inventive concept, this application also provides a heterogeneous separation model-based avatar robot task processing device for implementing the above-mentioned heterogeneous separation model-based avatar robot task processing method. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the heterogeneous separation model-based avatar robot task processing device provided below can be found in the limitations of the heterogeneous separation model-based avatar robot task processing method described above, and will not be repeated here.

[0239] In one exemplary embodiment, such as Figure 3As shown, a 300 embodied robot task processing system based on a heterogeneous separation model is provided, comprising:

[0240] The acquisition module 302 is used to acquire user commands and environmental perception data collected by the embodied robot, and determine the task processing mode for the user commands based on the user commands and environmental perception data.

[0241] The input module 304 is used to input user instructions and environmental perception data into the standard task processing model if the task processing mode is the standard processing mode; the standard task processing model includes a first visual encoder, a task understander, and a first action decoder.

[0242] Encoding module 306 is used to visually encode environmental perception data using a first visual encoder deployed in the accelerator processor to obtain first visual encoded features.

[0243] The understanding module 308 is used to perform task understanding on the first visual encoded features and user instructions through a task understander deployed in the accelerator processor and the fine processor to obtain task fusion features.

[0244] The generation module 310 is used to generate a first task execution action sequence based on the task fusion features by using a first action decoder deployed in the fine processor, and to control the embodied robot to execute the first task execution action sequence.

[0245] The modules in the aforementioned embodied robot task processing system based on a heterogeneous, separable model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.

[0246] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a task processing method for embodied robots based on a heterogeneous, discrete model. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0247] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0248] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing method embodiments.

[0249] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0250] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0251] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0252] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0253] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0254] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for processing embodied robot tasks based on a heterogeneous separable model, characterized in that, The method includes: Acquire user commands and environmental perception data collected by the embodied robot, and determine the task processing mode for the user commands based on the user commands and the environmental perception data; If the task processing mode is the standard processing mode, then the user instruction and the environmental perception data are input into the standard task processing model; the standard task processing model includes a first visual encoder, a task understander, and a first action decoder. The environmental perception data is visually encoded by the first visual encoder deployed in the accelerator processor to obtain the first visual encoded features. The task understander deployed in the accelerator and the fine processor performs task understanding on the first visual encoded features and the user instructions to obtain task fusion features. The first action decoder deployed in the fine processor generates a first task execution action sequence based on the task fusion features, and controls the embodied robot to execute the first task execution action sequence.

2. The method according to claim 1, characterized in that, After determining the task processing mode for the user instruction, the method further includes: If the task processing mode is a lightweight processing mode, then the user command and the environmental perception data are input into the lightweight task processing model; the lightweight task processing model includes a second visual encoder and a second action decoder. The environmental perception data is visually encoded by the second visual encoder deployed in the acceleration processor to obtain second visual encoded features. The second action decoder deployed in the fine processor generates a second task execution action sequence based on the second visual encoding features, and controls the embodied robot to execute the second task execution action sequence.

3. The method according to claim 1, characterized in that, Before inputting the user instruction and the environmental awareness data into the standard task processing model if the task processing mode is a standard processing mode, the method further includes: Obtain the original visual language action model, and split the original visual language action model into multiple sub-models according to the module function type; the multiple sub-models include the first visual encoder, the task understander, and the first action decoder. The target compression strategy for each sub-model is determined based on the compression sensitivity of each sub-model; the compression sensitivity characterizes the degree to which the model output accuracy is affected by compression. Based on the target compression strategy of each sub-model, the corresponding sub-model is compressed to obtain the compressed sub-model corresponding to each sub-model. Based on the computational resource requirements of each compression sub-model and the computational resource configuration characteristics of multiple candidate computational components, the target computational component corresponding to each compression sub-model is determined, and each compression sub-model is deployed on the corresponding target computational component; the target computational component includes the acceleration processor and the fine processor.

4. The method according to claim 3, characterized in that, The target compression strategy includes at least one of quantization precision and quantization method; determining the target compression strategy for each sub-model based on the compression sensitivity of each sub-model includes: If the compression sensitivity of the first visual encoder is less than the first threshold, then the first quantization precision is determined as the quantization precision of the first visual encoder, and the post-training quantization method is determined as the quantization method of the first visual encoder. If the compression sensitivity of the task understander is greater than the first threshold and less than the second threshold, the first quantization precision is determined as the quantization precision of the feature calculation layer in the task understander, and the second quantization precision is determined as the quantization precision of the sensitivity layer in the task understander, and the group quantization method is used as the quantization method of the task understander; the second quantization precision is greater than the first quantization precision. If the compression sensitivity of the first action decoder is greater than the second threshold, then the half-precision floating-point number is determined as the quantization precision of the first action decoder.

5. The method according to claim 4, characterized in that, The feature calculation layer includes weight parameters and activation parameters. The first quantization precision includes four-bit quantization and eight-bit quantization. The second quantization precision is the half-precision floating-point number. The step of compressing the corresponding sub-model according to the target compression strategy of each sub-model to obtain the compressed sub-model corresponding to each sub-model includes: For the task understander, the weight parameters are grouped according to the group quantization method to obtain M weight parameter groups; the number of weight parameters in each weight parameter group is less than or equal to a preset number, and M is a positive integer; Based on the four-bit quantization, the scaling factor and zero point corresponding to each weight parameter group are determined; Based on the scaling factor and zero point corresponding to each weight parameter group, the weight parameters in the corresponding weight parameter group are quantized to obtain the quantized weight parameters. Based on the eight-bit quantization, the activation parameters are quantized to obtain quantized activation parameters. Based on the quantized weight parameters and the quantized activation parameters, the quantized feature calculation layer is determined. The parameters in the sensitive layer are quantized based on the half-precision floating-point number to obtain the quantized sensitive layer. Based on the quantized feature calculation layer and the quantized sensitive layer, the compressed sub-model corresponding to the task understander is determined.

6. The method according to claim 3, characterized in that, The step of compressing the corresponding sub-model according to the target compression strategy of each sub-model to obtain the compressed model corresponding to each sub-model includes: Each of the sub-models is compressed according to the target compression strategy to obtain the initial compressed model corresponding to each of the sub-models. Obtain a quantized perception training dataset for quantized perception training of the first visual encoder and the task understander; wherein, the quantized perception training dataset includes robot running scene images, text instructions corresponding to the robot running scene images, and action sequence labels; Simulated quantization nodes are inserted into the forward propagation paths of the initial compressed model corresponding to the first visual encoder and the initial compressed model corresponding to the task understander, respectively, to obtain the first inserted model corresponding to the first visual encoder and the second inserted model corresponding to the task understander; wherein, the simulated quantization nodes are used to simulate the execution of quantization mapping operations and dequantization mapping operations. The robot's running scene image is input into the first inserted model to obtain visual perception features. The visual perception features and the text features corresponding to the text instructions are input into the second inserted model to obtain fused multimodal features. The fused multimodal features are input into the initial compressed model corresponding to the first action decoder to obtain the predicted action sequence, and the loss value is calculated based on the predicted action sequence and the ground truth value of the action sequence label. Backpropagation is performed on the first post-insertion model and the second post-insertion model based on the loss value, and the weight parameters in the first post-insertion model and the second post-insertion model are updated to obtain the compressed sub-model corresponding to the first visual encoder and the compressed sub-model corresponding to the task understander. The initial compressed model corresponding to the first action decoder is determined as the compressed sub-model corresponding to the first action decoder.

7. The method according to claim 3, characterized in that, The computing resource requirement characteristics include at least one of operation type characteristics and precision preservation requirement parameters; the computing resource configuration characteristics include at least one of supported computing precision types and processable operation types; determining the target computing component corresponding to each compression sub-model based on the computing resource requirement characteristics of each compression sub-model and the computing resource configuration characteristics of multiple candidate computing components includes: The computational resource requirement characteristics of each compression sub-model are matched with the computational resource configuration characteristics of multiple candidate computational components to obtain the matching results. If the matching result indicates that the operation type feature of the first visual encoder matches the operation type that the accelerator processor can process, the accelerator processor is identified as the target computing component corresponding to the first visual encoder. If the matching result indicates that the operation type characteristics of the matrix multiplication operation part in the task understander match the operation type that the accelerator can process, and the precision retention requirement parameters of the precision-sensitive operation part in the task understander match the computational precision type supported by the fine processor, then the accelerator and the fine processor are identified as the target computing components corresponding to the task understander. If the matching result indicates that the accuracy retention requirement parameter of the first action decoder matches the type of computational accuracy supported by the fine processor, the fine processor is identified as the target computational component corresponding to the first action decoder.

8. The method according to claim 1, characterized in that, The task understander includes a linear transformation layer and a nonlinear normalization layer, wherein the linear transformation layer is deployed in the acceleration processor and the nonlinear normalization layer is deployed in the fine processor; The task fusion features are obtained by performing task understanding on the first visual encoded features and the user instructions through the task understander deployed in the accelerator and the fine processor, including: The first visual encoded features and the user instructions are subjected to linear transformation processing by the linear transformation layer deployed in the accelerator processor to obtain multimodal semantic features for action decision-making. By deploying a nonlinear normalization layer in the fine processor, exponential normalization and layer normalization operations are performed on the multimodal semantic features to obtain task fusion features.

9. The method according to claim 1, characterized in that, The environmental perception data consists of multiple frames, and the first task execution action sequence includes task execution actions corresponding to each of the multiple frames of environmental perception data. Controlling the embodied robot to execute the first task execution action sequence includes: During the processing of environmental perception data in multiple consecutive frames, the first visual encoder, the task understander, and the first action decoder are controlled to execute in an inter-frame pipeline parallel manner; the inter-frame pipeline parallel manner is a way to process different sub-models of different frames and execute them in parallel on different computing components. The output interval between adjacent frames is determined based on the longest time value among the first visual encoder, the task understander, and the first action decoder. According to the adjacent frame output interval, the task execution actions corresponding to each frame of environmental perception data are output sequentially to control the embodied robot to execute the first task execution action sequence; wherein, the first visual encoder, the task understander and the first action decoder transmit frame data between adjacent pairs through a lock-free data queue.

10. The method according to claim 3, characterized in that, After deploying each of the compression sub-models to the corresponding target computing component, the method further includes: The historical visual information processed by the task understander and the cached data corresponding to the historical language instructions are obtained, and the cached data is quantized, compressed and stored. During the process of performing task understanding on the first visual encoded features and the user instructions by the task understander deployed in the accelerator and the fine processor, the current scene image and the current language instruction are acquired. The current scene image is compared with the historical visual information. If the comparison results are consistent, the cached data corresponding to the historical visual information is reused. The current language instruction is compared with the historical language instruction. If the comparison results are consistent, the cached data corresponding to the historical language instruction is reused.

11. The method according to claim 2, characterized in that, The process of acquiring user commands and environmental perception data collected by the embodied robot, and determining the task processing mode for the user commands based on the user commands and the environmental perception data, includes: The user instruction is parsed to determine the number of action steps and spatial precision requirements corresponding to the user instruction; If the number of action steps is greater than the step number threshold or the spatial accuracy requirement is higher than the accuracy threshold, the task processing mode is determined to be the standard processing mode. If the number of action steps is less than or equal to the number of steps threshold and the spatial accuracy requirement is less than or equal to the accuracy threshold, the task processing mode is determined to be the lightweight processing mode.

12. The method according to claim 2, characterized in that, The process of determining the lightweight task processing model includes: The standard task processing model is used as the teacher model, and the teacher model includes at least the first visual encoder, the task understander, and the first action decoder. The first visual encoder in the teacher model is reused as the second visual encoder, and a lightweight motion network is constructed as the second motion decoder; the number of parameters of the lightweight motion network is smaller than the number of parameters of the task understander. Using the processing results of the training samples by the teacher model as a supervision signal, and with the weights of the second visual encoder frozen, knowledge distillation training is performed on the lightweight action network to obtain the lightweight task processing model.

13. A embodied robot task processing system based on a heterogeneous separation model, characterized in that, The system includes: The acquisition module is used to acquire user commands and environmental perception data collected by the embodied robot, and determine the task processing mode for the user commands based on the user commands and the environmental perception data. An input module is used to input the user instruction and the environmental perception data into a standard task processing model if the task processing mode is a standard processing mode; the standard task processing model includes a first visual encoder, a task understander, and a first action decoder. The encoding module is used to visually encode the environmental perception data using the first visual encoder deployed in the accelerator processor to obtain a first visual encoded feature; The understanding module is used to perform task understanding on the first visual encoded features and the user instructions through the task understander deployed in the accelerator and the fine processor to obtain task fusion features; The generation module is used to generate a first task execution action sequence based on the task fusion features using the first action decoder deployed in the fine processor, and to control the embodied robot to execute the first task execution action sequence.