Trajectory distillation low-latency inference method and system for embodied robot policy model
Patent Information
- Application Number
- CN202611017313.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-29
AI Technical Summary
该类方案没有从输入签名、降噪步数、内部缓存状态和输出回传路径进行协同重构,因此容易产生胖接口、动态shape、重复graph launch或额外数据重排,端到端收益有限
本发明针对具身机器人智能模型的实时控制需求,在保留扩散式或流匹配式策略模型训练稳定性和动作分布表达能力的同时,将教师模型(teacher)的关键动作修正能力蒸馏到学生模型(student)中,并将学生模型的推理过程进一步重构为输入签名固定、缓存状态内聚、步数模板可选择、执行图可回放、最终动作低同步开销的部署路径,从而在任务精度、动作稳定性和端到端推理时延之间取得更优平衡。
Smart Images

Figure CN122840273A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology for embodied robots, and more specifically to a low-latency inference method and system for trajectory distillation of embodied robot strategy models. Background Technology
[0002] Embossed robot intelligent models typically need to establish mappings between multimodal information such as vision, language, robot body state, and historical actions, and output executable action sequences. Visual-language action models, diffusion strategy models, and flow matching strategy models are gradually becoming important technical approaches for this type of task. These models typically receive multimodal inputs such as multi-camera images, natural language commands, and robot joint states, and generate future action sequences through several steps of noise reduction or flow matching iterations.
[0003] The advantages of diffusion-based or flow-matching action generation mechanisms lie in their stable training process, strong ability to model action distribution, and the ability to gradually correct action noise through multiple preset iterations. This preset number of iterations can be determined based on the model training settings, such as 10 steps, 20 steps, or other longer steps. However, in real-time control deployments, longer iteration trajectories can also lead to significant inference redundancy. There is often high trajectory correlation and repetitive conditional calculations between adjacent denoising steps, and the longer generation process introduced by the model to achieve training stability is not always equivalent to necessary computation during the inference phase.
[0004] For robot control tasks, there is an inherent tension between stability and low latency. On the one hand, the policy model needs to maintain smooth, semantically consistent, and robust action sequences; on the other hand, the control closed loop requires mapping perceived actions within tens of milliseconds. Therefore, how to preserve the stability of diffusion-based policy training while compressing redundant noise-reducing trajectories during the inference phase is a key issue in deploying this type of model.
[0005] In existing implementations, robot policy models often run using the original dynamic graph or the default execution mode of a general framework. The input side typically organizes multiple images, language tokens, masks, and state tensors using dynamic structures such as dictionaries, lists, and Observation objects; the action generation side includes a loop-based denoise process; and there are multiple Python boundaries between prefix / cache, denoise step, and post-processing. The Python boundaries referred to in this paper are those in the model's hot path involving data structure parsing, loop scheduling, function call encapsulation, tensor synchronization, or cache propagation boundaries handled by the Python interpreter. Specifically, Python boundaries include: the boundary of parsing business-side Observation, dict, or list inputs; the boundary generated by Python scheduling for each iteration step in the denoise loop; the interface boundary formed by prefix / cache as external parameters passed in and out; and the boundary of action output post-processing and host synchronization. These boundaries can lead to unstable dynamic graph capture, increased graph launches, GPU execution interruptions, or increased host-side waiting.
[0006] While the aforementioned implementation is feasible in offline verification, it suffers from issues such as high end-to-end latency, unstable graph capture, difficulty in reusing fixed execution graphs by the compiler, and significant host-side preparation and backhaul overhead in real-time robot control scenarios. More importantly, the original, relatively long-step diffusion inference path typically retains redundant iterative forms from the training phase, requiring deployment to pay the full inference cost for intermediate trajectories serving stable training. Simply using a general-purpose compiler, CUDA Graph, or low-precision quantization often fails to reliably address the scheduling overhead caused by the dynamic nature of multimodal inputs and iterative action generation.
[0007] Some technical solutions directly use a compilation backend on the original policy model, or optimize the denoising loop, attention operator, and matrix multiplication operator separately. These solutions do not coordinate the reconstruction of input signature, denoising steps, internal cache state, and output backhaul path, thus easily leading to fat interfaces, dynamic shapes, repeated graph launches, or additional data rearrangements, resulting in limited end-to-end benefits.
[0008] Furthermore, existing knowledge distillation schemes typically focus on the consistency of output actions, feature distributions, or prediction results, with little attention paid to task-aware compression of denoise time-step redundancy in diffusion-based or stream-matching action generation processes. Existing deployment optimization schemes usually focus on operator-level acceleration, graph compilation, or low-precision quantization, failing to collaboratively design multimodal input signature, denoise step selection, prefix / cache state organization, static template playback, and the final action synchronization path as a single end-to-end control link.
[0009] Therefore, in real-time control scenarios of embodied robots, how to simultaneously ensure motion stability, task accuracy, and low-latency deployment requirements is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0010] In view of the above problems, the present invention proposes a low-latency inference method and system for trajectory distillation of embodied robot strategy models, aiming to overcome or at least partially solve the above problems.
[0011] To achieve the above objectives, the present invention adopts the following technical solution:
[0012] In a first aspect, embodiments of the present invention provide a low-latency inference method for trajectory distillation of an embodied robot strategy model, comprising: In the offline training phase: the diffusion trajectory of the teacher model is used as a training reference to perform task-aware trajectory distillation training on the student model; the noise reduction inference process in the trained student model is expanded and recorded as multiple corresponding static execution templates according to different candidate iteration steps. During the online inference phase: Receive multimodal robot observation data and normalize it into a fixed tensor input signature; select a target template from multiple static execution templates based on the task delay budget and accuracy requirements; perform forward inference playback according to the target template based on the fixed tensor input signature, and generate and output the final action tensor.
[0013] Furthermore, the process of using the diffusion trajectory of the teacher model as a training reference to perform task-aware trajectory distillation training on the student model specifically includes: Obtain the teacher model in N 教 The correction amount output by each iteration step under each iteration step; N in the student model 学 In each iteration step, the output of the student model at each iteration step is supervised by using the correction amount output by the teacher model at the corresponding iteration step as a reference; where 1≤N 学 <N 教 ; The parameters of the student model are iteratively updated using the final action error between the final action sequence of the student model and the final action sequence of the teacher model as the training loss.
[0014] Furthermore, the task-aware trajectory distillation training also includes: Based on the final action error between the teacher model and the student model, the error at different iteration steps, gradient stability, changes in validation set error, or changes in task metrics, the learning rate, loss weight, or time step weight of the student model during the training process are dynamically adjusted.
[0015] Furthermore, the static execution template has a fixed step sequence, fixed time step parameters, fixed cache usage path, and fixed action projection path.
[0016] Furthermore, the fixed tensor input signature includes a fixed number, fixed order, and fixed shape range of multiple images, image masks, language instructions, language masks, robot state tensors, and noise tensors.
[0017] Furthermore, in the forward inference playback, the fixed cache in the target template is built and reused within the student model. The fixed cache is not passed into the student model as an external input parameter, nor is it output from the student model as an intermediate result.
[0018] Furthermore, it also includes: synchronizing only the final action tensor to the CPU or controller side, skipping the return of intermediate cache or intermediate noise reduction results.
[0019] Furthermore, it also includes: hardware metric feedback optimization steps: The hardware performance metrics collected during the end-to-end inference process include at least one of the following: graph startup count, host synchronization location, computing core hotspot, and data transfer volume. Based on hardware performance indicators, identify latency bottlenecks: when the number of graph startups exceeds a preset value or the graph startup scheduling latency exceeds a corresponding preset proportion, prioritize template merging, loop expansion, or graph boundary convergence; when the input preparation latency exceeds a corresponding preset proportion, prioritize fixed input signature normalization, tensor pre-allocation, or input normalization; when the backbone computation latency exceeds a corresponding preset proportion, prioritize selecting static execution templates with fewer iteration steps or performing operator fusion; when the host synchronization latency or output feedback latency exceeds a corresponding threshold, prioritize pruning intermediate state feedback and retain only final action synchronization.
[0020] Secondly, the present invention provides a low-latency inference system for trajectory distillation of an embodied robot strategy model. Using the above-described method, the system includes: The distillation training module is used to perform task-aware trajectory distillation training on the student model during the offline training phase, using the diffusion trajectory of the teacher model as a training reference. The template recording module is used to expand and record the denoising inference process in the trained student model into multiple corresponding static execution templates according to different candidate iteration steps during the offline training phase. A fixed signature input module is used to receive multimodal robot observation data during the online inference phase and normalize it into a fixed tensor input signature; The template selection module is used to select a target template from multiple static execution templates during the online inference phase based on the task latency budget and accuracy requirements. The replay execution module is used to perform forward inference replay based on the fixed tensor input signature during the online inference stage, according to the target template, to generate and output the final action sequence tensor.
[0021] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a low-latency inference method and system for trajectory distillation of embodied robot strategy models, which has the following beneficial effects: This invention addresses the real-time control requirements of embodied robot intelligent models. While preserving the training stability and action distribution representation capabilities of diffusion-based or flow-matching strategy models, it distills the key action correction capabilities of the teacher model into the student model. Furthermore, it reconstructs the inference process of the student model into a deployment path with fixed input signatures, cohesive cached states, selectable step templates, replayable execution graphs, and low synchronization overhead for the final action. This achieves a better balance between task accuracy, action stability, and end-to-end inference latency. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the overall framework of the trajectory distillation low-latency inference method for the embodied robot strategy model provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the recording and static playback reasoning architecture for various step templates provided in the embodiments of the present invention; Figure 3 This is a schematic diagram illustrating the optimization effect of H800 end-to-end inference latency provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] This invention discloses a low-latency inference method for trajectory distillation of embodied robot policy models, such as... Figure 1 As shown, it includes the following steps: In the offline training phase: the diffusion trajectory of the teacher model is used as a training reference to perform task-aware trajectory distillation training on the student model; the noise reduction inference process in the trained student model is expanded and recorded as multiple corresponding static execution templates according to different candidate iteration steps. During the online inference phase: Receive multimodal robot observation data and normalize it into a fixed tensor input signature; select a target template from multiple static execution templates based on the task delay budget and accuracy requirements; perform forward inference playback according to the target template based on the fixed tensor input signature, and generate and output the final action tensor.
[0026] Next, the above content will be explained in detail.
[0027] I. Offline training phase.
[0028] 1. Task-aware trajectory distillation training.
[0029] This invention starts from the task generation process of embodied robot intelligent models. The number of denoising or flow matching iterations for the teacher model is denoted as N. 教 Let N be the number of iterations for the student model. 学 And restrict 1≤N 学 <N 教 The teacher model achieves stable training and a better representation of action distribution through a relatively longer denoised trajectory; the student model meets the low latency requirements of real-time control with a smaller number of iterations. The deployment phase truly requires quickly obtaining executable actions under given observation conditions and task constraints, rather than completely reproducing every intermediate denoised state. Therefore, this invention first uses the stable diffusion trajectory or flow matching trajectory of the teacher model as a training reference, transferring the correction directions that contribute more to the final task actions to the student model, reducing redundant corrections during the inference phase.
[0030] In distillation training, the student model is not required to mechanically replicate all the intermediate trajectories of the teacher model, but rather to learn around the correction directions of the final action error and key denoise time steps. Specifically, the teacher model's performance at N... 教 The correction amount output at each iteration step in the N iteration steps; in the student model 学 In each iteration step, the output of the student model is supervised using the correction amount output by the teacher model at the corresponding iteration step as a reference; let the i-th iteration step of the student model template be... The iteration step corresponds to the first iteration step of the teacher model. There are iteration steps, among which =1,2,…, , This represents the total number of iterations. The student model's iteration number... The time step error of each iteration step relative to the corresponding iteration step in the teacher model It can be represented as:
[0031] in, The dimension representing the predicted action value or intermediate correction amount; Indicates the dimension index; The student model is represented by the first... The output of the nth iteration step Dimensional motion prediction value or intermediate correction amount; This indicates that the teacher model is in the same position as the student model in the first... The output of the corresponding iteration step in the iteration step is the first iteration step. Dimensional motion prediction value or intermediate correction amount; This represents the time step mapping function, used to map the student model's first step. Each iteration step is mapped to the corresponding iteration step in the teacher model; Represents norm operations.
[0032] Simultaneously, the final action error between the student model's final action sequence and the teacher model's final action sequence is used as the training loss to iteratively update the student model's parameters. This final action error... Represented as:
[0033] in, Indicates the time domain length of action prediction; Indicates a time index; Indicates the dimension of a single-step action; This indicates that the final action sequence of the student model is in the 1st... The time point, the first Action value of the action dimension; This indicates that the teacher model's final action sequence is in the 1st... The time point, the first Action value of the action dimension; It represents the absolute value.
[0034] To prevent motion distribution drift in the student model due to denoising during the distillation process, this invention introduces a dynamic perceptual training adjustment mechanism. This mechanism adjusts the training based on the final motion error between the teacher and student models. The learning rate, loss weights, or time-step weights are dynamically adjusted during the student model training process to account for changes in denoising time-step errors, gradient stability, validation set errors, or task metrics. This ensures the student model learns the key action correction directions, rather than completely replicating all redundant intermediate trajectories from the teacher model. Based on the aforementioned final action error... The first The time step weights corresponding to each iteration of denoising are represented as follows:
[0035] in, Indicates the weight of each time step; Indicates the total number of iterations; The student model is represented by the first... The time step error of each iteration step relative to the corresponding iteration step in the teacher model; The student model is represented by the first... The time step error of each iteration step relative to the corresponding iteration step in the teacher model.
[0036] When the error at a specific time step or stage is large, increase the optimization intensity of the corresponding stage; when the validation set error rises continuously, the gradient fluctuation index exceeds the threshold, or the model is close to convergence, reduce the global learning rate or reduce the weight of the corresponding time step.
[0037] 2. Static execution template recording.
[0038] The denoising inference process in the trained student model is expanded and recorded as multiple corresponding static execution templates according to different candidate iteration steps. The term "recording" in this paper does not simply refer to saving model files or exporting parameters, but rather to expanding the originally dynamically loop-controlled denoising inference process into static execution templates with fixed step order, fixed input signatures, fixed time step parameters, fixed prefix / cache usage paths, and fixed action projection paths. Because the step order within the static execution template is fixed, the compiler can see the static control flow and generate a stable execution graph.
[0039] For different control requirements, 3-step, 5-step, 7-step, or other candidate iteration steps can be recorded as static execution templates. Taking a 5-step template as an example, the loop is expanded into five defined forward computation segments: step0, step1, step2, step3, and step4. Taking a 7-step template as an example, the corresponding computation segments include steps0 to 6. Each computation segment includes time step embedding, suffix embedding, mask / position encoding preparation, backbone forward propagation, cache reading or updating, and action correction output. The denoising sequence, prefix / cache usage, and output projection path within each static execution template are determined, and therefore can be stably captured by mechanisms such as the compilation graph backend and CUDAGraph.
[0040] II. Online Reasoning Stage.
[0041] 1. Fixed tensor input signature is well-organized.
[0042] During the online inference phase, multimodal robot observation data is received and normalized into a fixed tensor input signature. The multimodal robot observation data includes multiple camera images, image masks, language commands (tokens), language masks, robot states, and random noise vectors. The corresponding normalized fixed tensor input signature includes a fixed number, fixed order, and fixed shape range of multiple images, image masks, language commands (tokens), language masks, robot state tensors, and noise tensors.
[0043] Specifically, dynamic inputs from the business side, in the form of observations, dictionaries, or lists, are parsed. Multiple images, masks, language tokens, robot states, and noise are regularized into a fixed number, fixed order, and fixed shape range of tensors before entering the model's hot path. For example, three images are used as inputs img0, img1, and img2, and the image masks are fixed as mask0, mask1, and mask2. Language tokens, language masks, robot states, and noise are fixed as langTokens, langMasks, state, and noise, respectively. This fixed contract does not require the business side to be completely static, but it requires the tensors entering the compilation hot path to be morphologically stable. Therefore, the compilation backend faces a morphologically stable, boundary-converged, and cache-cohesive inference path, rather than a temporary computation graph pieced together from Python objects and fat interfaces.
[0044] Furthermore, by transforming dynamic Observation, dict, and list inputs into a fixed number of tensor parameters, and establishing input contracts with a defined order and shape for multiple images, masks, language tokens, states, and noise, the compilation backend can obtain a stable graph capture entry point.
[0045] 2. Select the target template.
[0046] Based on the task latency budget and accuracy requirements, the target template is selected from multiple static execution templates instead of reinterpreting the dynamic loop in the hot path.
[0047] Specifically, during runtime, a target template is selected from multiple static execution templates based on the control cycle, acceptable action error, and hardware profile results. If the task prioritizes low latency, a template with fewer iteration steps (such as a 3-step template) can be selected; if the task prioritizes action accuracy or scene complexity, a template with more iteration steps (such as a 7-step or 10-step template) can be selected.
[0048] 3. Forward reasoning replay.
[0049] After selecting the target template, the compilation graph backend or static graph capture mechanism performs the following: based on the fixed tensor input signature, forward inference replay is performed according to the target template to generate and output the final action tensor. Here, "replay" refers to performing forward inference according to the expanded static template during online inference, that is, entering the model with a fixed input signature, completing prefix / cache construction and reuse within the model, and performing noise calculation and action projection in the order of step0 to stepN defined by the template.
[0050] During forward playback, the fixed cache (prefix / cache) in the target template is built and reused within the student model. It is neither passed to the student model as an external input parameter via a fat interface, nor is it output from the student model as an intermediate result, thus avoiding the external fat interface from disrupting graph capture stability. Specifically, the prefix / cache is generated and consumed internally by the model, avoiding its transmission as a large number of external parameters. During template playback, the cache remains in its internal model state, preventing it from being flattened into a fat external interface. This cache cohesion mechanism avoids flattening the prefix / cache into a large number of external parameters, thereby reducing runtime parameter management and graph input complexity, and reserving optimization space for the compiler across prefixes and noise reduction.
[0051] After playback is complete and the final action sequence tensor is generated and output, only the final action sequence tensor is synchronized to the CPU or controller side, skipping the return of intermediate buffers or intermediate noise reduction results, and feeding back the actual profiling results to the template selection and subsequent optimization process. The system only performs host return when the final action needs to be delivered to the controller, skipping the synchronization of useless states and intermediate tensors.
[0052] III. Hardware performance feedback optimization.
[0053] This embodiment also includes a hardware performance feedback optimization step. Hardware performance metrics are collected during the end-to-end inference process. These metrics include at least one of the following: graph launch count, host synchronization location, kernel hotspots, and data transfer volume. Based on these performance metrics, a hardware feedback loop is formed to determine whether the next step should be to compress input preparation, template boundaries, backbone computation, or output synchronization. During this determination, the total end-to-end latency can be decomposed into:
[0054] in, This represents the total latency of a single end-to-end inference operation. This indicates the delays in input parsing, tensor warping, and necessary preprocessing. This represents the delay in the kernel startup scheduling diagram; This indicates the computational delay of the model backbone and denoise template; This indicates the latency during which the host and device are waiting to synchronize. This represents the latency for final action feedback and control-side delivery. Based on this, the proportion of each part can be calculated:
[0055] in, This represents the proportion of the i-th type of delay; This represents the delay of any of the above sub-items; i can be input, start, calculation, synchronization, or output.
[0056] Based on hardware performance indicators, identify latency bottlenecks: when the number of graph startups exceeds a preset value or the graph startup scheduling latency exceeds a corresponding preset proportion, prioritize template merging, loop expansion, or graph boundary convergence; when the input preparation latency exceeds a corresponding preset proportion, prioritize fixed input signature normalization, tensor pre-allocation, or input normalization; when the backbone computation latency exceeds a corresponding preset proportion, prioritize selecting static execution templates with fewer iteration steps or performing operator fusion; when the host synchronization latency or output feedback latency exceeds a corresponding threshold, prioritize pruning intermediate state feedback and retain only final action synchronization.
[0057] As a result, the original dynamic, segmented, and multi-boundary policy inference chain has been reconstructed into a low-latency GPU inference system with selectable templates, stable playback, and continuous optimization.
[0058] IV. Experimental proof.
[0059] The technical solution of this embodiment will be described in detail using the DROID robot strategy model pi0-droid as an example.
[0060] The model input includes three 224×224 RGB images, a language command token, the robot's state, and a noise tensor; the model output is a sequence of future actions. The original teacher model uses N iterations. 教 The generated actions, after distillation, can produce an iteration step N. 学 Less than N 教 Multiple deployment templates are available, such as 3-step, 5-step, 7-step, or other candidate templates. The teacher model serves as the carrier for stable training and high-quality action distribution modeling, while the student model acts as the deployment carrier, absorbing key correction information from the teacher model's denoised trajectory. In this example, the deployment template actually verified is a 5-step denoise; under other deployment conditions, 3-step, 7-step, or other step templates can also be recorded.
[0061] The distillation stage employs a dynamic learning rate strategy: when the output deviation between the teacher model and the student model or the error in a specific denoise stage is large, the optimization intensity of the corresponding stage is increased; when the validation error worsens or the gradient fluctuation increases, the learning rate is reduced or the loss weight is adjusted, thereby preventing the student model from experiencing action distribution drift after reducing the number of iterations.
[0062] The deployment phase employs morphologically constrained inputs and cache cohesion mechanisms to better suit the boundaries of images, language, state, noise, and internal cache for compilation backend capture. An example deployment device is an NVIDIA H800 GPU.
[0063] It should be noted that pi0-droid, DROID data or task format, NVIDIA H800 GPU, and 5-step denoise template are all implementations and verification methods in this example, and do not constitute a limitation on the scope of protection of this invention. The same ideas of task-aware trajectory distillation, fixed input signature, template playback, and output synchronous pruning can also be used for other visual language action models, diffusion strategy models, stream matching strategy models, other GPUs, or other static graph capture backends.
[0064] like Figure 2 As shown, the offline phase first completes trajectory compression from the teacher model to the student model. The system does not require the student model to reproduce all intermediate denoised states of the teacher model point by point; instead, it focuses more on the correction direction of the final action error and key time steps. This leverages the stability gained by the teacher model through training on a longer diffusion trajectory while reducing the computational dependence on redundant intermediate trajectories during the deployment phase.
[0065] After distillation, the system records static templates for different candidate steps. For example, a 5-step template contains steps 0 to 4; a 7-step template contains steps 0 to 6. Each step includes time step embedding, suffix embedding, mask / position preparation, trunk forwarding, cache reading or updating, and action output projection. Because the step sequence within the template is fixed, the compiler can see the static control flow and generate a stable execution graph.
[0066] During the online phase, the system takes three images as inputs (img0, img1, and img2) and fixes the corresponding masks as mask0, mask1, and mask2. The language token, language mask, robot state, and noise are also taken as fixed tensor inputs. This fixed contract does not require the business side to be completely static, but it does require the tensors entering the compilation hot path to be morphologically stable.
[0067] Subsequently, the system performs prefix embedding on the image token and text token, constructing prefix-related intermediate states and a cache. This cache is not exposed to external interfaces but is reused internally by subsequent denoise steps, thus avoiding the disruption of graph capture stability by external fat interfaces.
[0068] At runtime, a template is selected based on the control cycle, acceptable action error, and hardware profiling results. If the task prioritizes low latency, a template with fewer iteration steps can be selected; if the task prioritizes action accuracy or scene complexity, a template with more iteration steps can be selected. Template playback involves performing a complete forward inference according to the selected static template. After playback, the system only sends the final action tensor back to the CPU or controller, and no longer sends back intermediate states unrelated to control.
[0069] The system also uses profiling results to determine whether optimizations have truly reduced end-to-end latency. If a graph launch has converged to a single operation or its latency percentage is below a threshold, the system will not continue to perform ineffective optimizations around the launch, but will instead focus on more rewarding aspects such as input preparation, backbone computation, or output synchronization. If host synchronization or output backhaul latency percentages are high, intermediate state backhauls will be prioritized. If backbone computation percentages are high, the denoise step template will be adjusted or local operator fusion will be performed.
[0070] In the H800 single-card experiment, the end-to-end p50 latency is shown in Table 1. This result is used to illustrate the technical effectiveness of each stage of this solution.
[0071] Table 1: End-to-end p50 latency results
[0072] The above data was obtained based on an experiment with an H800 single card. The input included three 224×224 RGB images, a language token, a language mask, robot state, and a noise tensor. The student deployment template was a 5-step denoise.
[0073] In terms of performance, such as Figure 3 As shown, the final form constraint input path end-to-end p50 is approximately 26.06 ms, indicating that this link can gradually converge the high scheduling and iteration overhead in ordinary dynamic graph execution into a low-latency deployment path.
[0074] In terms of accuracy, on the 5-step student template verified in this example, under the held-out dataset and the same evaluation configuration, the mean absolute error of offline actions decreased from 0.1453 for the teacher to 0.1169 for the student, an error reduction of approximately 0.8046x, or about 19.5%. The arithmetic relationship for this value is 0.1169 / 0.1453 ≈ 0.8046; this assumes that both are from the same dataset, the same action error caliber, and the same evaluation configuration. This arithmetic relationship can be expressed as:
[0075] in, This represents the mean absolute error of the student model's offline actions. This represents the offline action mean absolute error of the teacher model.
[0076] This improved accuracy primarily stems from task-aware distillation training from the teacher model to the student model. Execution path reconstruction, including denoised template replay, morphological input constraints, buffered cohesion, and synchronous output pruning, is mainly used to reduce end-to-end inference latency without altering model weights, while maintaining near-lossless output. A dynamic perceptual training adjustment mechanism balances the learning difficulty of different denoise time steps during the distillation phase, enabling the student model to not only reduce inference steps but also maintain stable performance in held-out action error.
[0077] These results demonstrate that compressible inference redundancy exists within the longer diffusion trajectory of the teacher model. While the longer iteration process of the teacher model is valuable for training stability, key correction capabilities can be transferred to the student model with fewer iterations during deployment via distillation, thereby achieving both low latency and low offline action error.
[0078] For execution path optimization with equal weights, denoising template replay and morphological constraint input mainly change the execution structure without altering the model parameters. The output difference between the morphological constraint input path and the default template replay path is small, with an average absolute error of approximately 0.00398 and a maximum absolute error of approximately 0.0231, indicating that execution path reconstruction has little impact on action output.
[0079] The morphological constraint input and output synchronous pruning reduces object construction, dictionary resolution, useless tensor backpropagation, and host-side waiting in the hot path, thus enabling the path to continue reducing end-to-end latency without changing the model weights.
[0080] The cache cohesion mechanism avoids flattening the prefix / cache into a large number of external parameters, thereby reducing runtime parameter management and graph input complexity, and reserving optimization space for the compiler across prefixes and denominations.
[0081] The hardware feedback filtering mechanism enables the system to distinguish between true benefit points and nominal optimization points. For example, in this instance, low-precision paths such as general quantization and FA3 FP8 attention-only significantly impact the stability of the current VLA model's action output and its sensitivity to task error; excessively long prefix padding introduces additional redundant computation. Therefore, this invention uses the above-mentioned routes as optional extensions or contrast schemes, while taking task-aware trajectory distillation, denoising template playback, morphological constraint input, and graph boundary convergence as the main implementation path.
[0082] Experiments also show that components such as mask, position, time-conditioning, and final action output in the current VLA model are more sensitive to numerical precision. Directly using coarse-grained quantification or global low-precision attention replacement can easily disrupt the smoothness of the action sequence and task consistency. Therefore, this invention prioritizes near-lossless execution path reconstruction and limits low-precision quantification to a direction that can be carefully verified on a module-by-module basis for subsequent expansion.
[0083] Based on the same inventive concept, embodiments of the present invention also provide a low-latency inference system for trajectory distillation of an embodied robot strategy model. Using the above-described method, the system includes: The distillation training module is used to perform task-aware trajectory distillation training on the student model during the offline training phase, using the diffusion trajectory of the teacher model as a training reference. The template recording module is used to expand and record the denoising inference process in the trained student model into multiple corresponding static execution templates according to different candidate iteration steps during the offline training phase. A fixed signature input module is used to receive multimodal robot observation data during the online inference phase and normalize it into a fixed tensor input signature; The template selection module is used to select a target template from multiple static execution templates during the online inference phase based on the task latency budget and accuracy requirements. The replay execution module is used to perform forward inference replay based on a fixed tensor input signature during the online inference phase, according to the target template, to generate and output the final action sequence tensor.
[0084] Since the principle behind the problem solved by this system is similar to the trajectory distillation low-latency inference method of the aforementioned embodied robot strategy model, the implementation of the device and the client can refer to the implementation of the aforementioned method, and the repeated parts will not be repeated.
[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0086] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A low-latency inference method for trajectory distillation of an embodied robot strategy model, characterized in that, include: In the offline training phase: the diffusion trajectory of the teacher model is used as a training reference to perform task-aware trajectory distillation training on the student model; The noise reduction inference process in the trained student model is expanded and recorded as multiple corresponding static execution templates according to different candidate iteration steps. During the online inference phase: Receive multimodal robot observation data and normalize it into a fixed tensor input signature; Based on the task latency budget and accuracy requirements, a target template is selected from multiple static execution templates; Based on the fixed tensor input signature, forward reasoning playback is performed according to the target template to generate and output the final action tensor.
2. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 1, characterized in that, The process of using the diffusion trajectory of the teacher model as a training reference to perform task-aware trajectory distillation training on the student model specifically includes: Obtain the teacher model in N 教 The correction amount output by each iteration step under each iteration step; N in the student model 学 In each iteration step, the output of the student model at each iteration step is supervised by using the correction amount output by the teacher model at the corresponding iteration step as a reference; where 1≤N 学 <N 教 ; The parameters of the student model are iteratively updated using the final action error between the final action sequence of the student model and the final action sequence of the teacher model as the training loss.
3. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 2, characterized in that, The task-aware trajectory distillation training also includes: Based on the final action error between the teacher model and the student model, the error at different iteration steps, gradient stability, changes in validation set error, or changes in task metrics, the learning rate, loss weight, or time step weight of the student model during the training process are dynamically adjusted.
4. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 1, characterized in that, The static execution template has a fixed step sequence, fixed time step parameters, fixed cache usage path, and fixed action projection path.
5. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 1, characterized in that, The fixed tensor input signature includes a fixed number, fixed order, and fixed shape range of multiple images, image masks, language commands, language masks, robot state tensors, and noise tensors.
6. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 1, characterized in that, In the forward inference playback, the fixed cache in the target template is built and reused inside the student model. The fixed cache is not passed into the student model as an external input parameter, nor is it output from the student model as an intermediate result.
7. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 1, characterized in that, Also includes: Only the final action tensor is synchronized to the CPU or controller side, skipping the return of intermediate cache or intermediate noise reduction results.
8. The trajectory distillation low-latency inference method for embodied robot strategy models as described in claim 1, characterized in that, Also includes: Hardware metrics feedback optimization steps: The hardware performance metrics collected during the end-to-end inference process include at least one of the following: graph startup count, host synchronization location, computing core hotspot, and data transfer volume. Based on hardware performance indicators, identify latency bottlenecks: when the number of graph startups exceeds a preset value or the graph startup scheduling latency exceeds a corresponding preset proportion, prioritize template merging, loop expansion, or graph boundary convergence; when the input preparation latency exceeds a corresponding preset proportion, prioritize fixed input signature normalization, tensor pre-allocation, or input normalization; when the backbone computation latency exceeds a corresponding preset proportion, prioritize selecting static execution templates with fewer iteration steps or performing operator fusion; when the host synchronization latency or output feedback latency exceeds a corresponding threshold, prioritize pruning intermediate state feedback and retain only final action synchronization.
9. A low-latency inference system for trajectory distillation of an embodied robot strategy model, characterized in that, The system comprising the method of any one of claims 1-8, wherein the method is: The distillation training module is used to perform task-aware trajectory distillation training on the student model during the offline training phase, using the diffusion trajectory of the teacher model as a training reference. The template recording module is used to expand and record the denoising inference process in the trained student model into multiple corresponding static execution templates according to different candidate iteration steps during the offline training phase. A fixed signature input module is used to receive multimodal robot observation data during the online inference phase and normalize it into a fixed tensor input signature; The template selection module is used to select a target template from multiple static execution templates during the online inference phase based on the task latency budget and accuracy requirements. The replay execution module is used to perform forward inference replay based on the fixed tensor input signature during the online inference stage, according to the target template, to generate and output the final action sequence tensor.