Video generation type vla model acceleration method and system based on time semantic continuity

CN122807858APending Publication Date: 2026-09-25INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610845001.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明提供一种基于时间语义连续性的视频生成式VLA模型加速方法及系统,用以解决现有技术中扩散模型迭代去噪导致的推理延迟极高、静态缓存策略引发闭环控制误差累积以及忽视跨决策步语义冗余造成的严重算力浪费的缺陷,实现机器人在闭环控制任务中的高频实时响应与高精度决策协同

Benefits of technology

[0017]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述视频生成式VLA模型加速方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807858A_ABST
    Figure CN122807858A_ABST
Patent Text Reader

Abstract

The application provides a video generation type VLA model acceleration method and system based on time semantic continuity, which comprises the following steps: acquiring a visual observation image and a language instruction of a current decision step; based on time semantic continuity, extracting an intermediate latent feature or a denoising trajectory generated by a previous decision step and pre-stored in a latent feature cache, and injecting the intermediate latent feature or the denoising trajectory after step number alignment as an initial prior for diffusion generation of the current decision step; dynamically adjusting a feature cache reuse frequency of a neural network block in a diffusion model according to a state perception result of the current visual observation image; performing denoising iteration combined with the initial prior and the adjusted reuse frequency to obtain a predicted video feature and perform action decoding, and outputting a robot control action. The application breaks the limitation of the diffusion model generated from zero, significantly reduces the video generation delay, and adaptively reduces redundant calculation by using action feedback and visual redundancy, so that efficient and real-time control is realized under the premise of ensuring the smoothness of the robot action and the accuracy of the decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and embodied intelligence, and in particular to a method and system for accelerating video-generative VLA models based on temporal semantic continuity. Background Technology

[0002] With the development of embodied intelligence technology, the Vision-Language-Action Model (VLA) combines high-dimensional visual observation with natural language commands to achieve end-to-end control of robots. Among these advancements, the introduction of video generation modules (such as diffusion models) to simulate future physical states has become an important paradigm for improving the generalization ability of robots for complex, long-range tasks.

[0003] However, video-generative VLA models face severe computational efficiency bottlenecks when deployed to physical edge devices. The generation process of the diffusion model follows a Markov inverse denoising trajectory, requiring dozens or even hundreds of forward computations of a massive neural network to generate a predicted video. In continuous decision-making scenarios for robots, the system must continuously call the diffusion model within high-frequency control cycles (typically 10Hz-30Hz). This contradiction between "high cost per generation" and "high-frequency calls for closed-loop control" results in extremely high end-to-end inference latency, making real-time deployment difficult.

[0004] To alleviate inference latency, the general computer vision field has attempted to introduce feature caching techniques. However, existing caching schemes are mainly designed for offline, isolated image or video generation tasks. When directly applied to embodied agents in closed-loop control, they suffer from the following serious drawbacks: First, it disrupts the temporal semantic continuity across decision steps. In continuous robot operations, the environmental background of adjacent decision steps typically undergoes small, gradual changes, while the higher-level instructions remain unchanged. Existing acceleration methods treat the video prediction for each decision step as a completely independent, isolated event, forcing the model to recalculate highly similar backgrounds and object structures from pure Gaussian noise at every moment, resulting in a significant waste of computational resources.

[0005] Second, static caching strategies lead to error accumulation and action collapse. Existing technologies often employ fixed, static scheduling schemes, failing to perceive the dynamic needs of robots as they transition from "macroscopic movement" to "fine manipulation" in terms of feature sensitivity. In the continuous autoregressive "planning-execution" loop, this forced reuse, decoupled from the physical task, generates minute state-awareness biases. These biases amplify over time, producing covariate shifts, ultimately causing action strategy collapse and severely impairing task success rates. Summary of the Invention

[0006] This invention provides a video-generated VLA model acceleration method and system based on temporal semantic continuity, which solves the defects of existing technologies such as extremely high inference latency caused by diffusion model iterative denoising, accumulation of closed-loop control errors caused by static caching strategies, and serious waste of computing power caused by ignoring semantic redundancy across decision steps. It enables robots to achieve high-frequency real-time response and high-precision decision collaboration in closed-loop control tasks.

[0007] This invention provides a method for accelerating video generative VLA models, comprising: Acquire visual observations and verbal commands for the current decision step; Based on temporal semantic continuity, intermediate latent features or denoising trajectories generated in the previous decision step during the diffusion model denoising process are extracted from the latent feature cache and injected after step alignment as the initial prior for the diffusion generation of the current decision step. Based on the state perception results of the current visual observation image, the feature cache reuse frequency of the neural network block in the diffusion model is dynamically adjusted. The denoising iteration is performed by combining the initial prior and the adjusted reuse frequency to obtain the predicted video features; The predicted video features are decoded to generate and output robot control actions.

[0008] According to the video generative VLA model acceleration method provided by the present invention, the step of extracting intermediate latent features generated in the diffusion model denoising process of the previous decision step and pre-stored in the latent feature cache, and injecting them after step alignment as the initial prior for diffusion generation in the current decision step, specifically includes: Calculate the image similarity between the visual observation images of the previous decision step and the current decision step; Based on a preset mapping relationship, the denoising step size position corresponding to the extracted features is determined according to the image similarity; wherein, the higher the image similarity, the closer the denoising step size position corresponding to the extracted features is to the output end in the denoising trajectory; Based on the determined denoising step size position, the intermediate latent features are extracted from the latent feature cache and injected into the denoising step corresponding to the current decision step as the initial prior for diffusion generation in the current decision step.

[0009] According to the video generative VLA model acceleration method provided by the present invention, the step of dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model specifically includes: The motion amplitude or semantic offset of the visual observation image at the current decision step is evaluated through a state-aware mechanism. When the robot is in the fine operation phase, reduce the frequency of feature cache reuse to perform complete forward computation; When the robot is in the simple movement stage, increase the frequency of feature cache reuse to reuse features from historical network layers.

[0010] According to the video generative VLA model acceleration method provided by the present invention, the step of dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model further includes: Acquire the robot's historical motion trajectories and assess their rate of change; The similarity threshold is dynamically set based on the rate of change: when the rate of change of the historical action trajectory is higher than the preset first threshold, the current state is determined to be critical, and the similarity threshold is increased; when the rate of change of the historical action trajectory is lower than the preset second threshold, the current state is determined to be safe, and the similarity threshold is decreased. Calculate the feature similarity between the feature map of the current denoising step and the feature map of the corresponding denoising step of the previous decision step in real time; If the feature similarity is higher than the similarity threshold, the calculation of the current neural network block is skipped, and the cached feature result is directly output. If the feature similarity is lower than the similarity threshold, the neural network block is activated to perform a complete forward computation, and the calculated new features are updated in the latent feature cache.

[0011] According to the video generative VLA model acceleration method provided by the present invention, the step of performing denoising iteration by combining the initial prior and the adjusted reuse frequency to obtain predicted video features specifically includes: The feature transfer path from predicted video features to the action decoder is optimized, and the smoothness of the output action sequence on the time axis is maintained by reusing the feature cache.

[0012] According to the video generative VLA model acceleration method provided by the present invention, after performing action decoding on the predicted video features to generate and output robot control actions, the method further includes: Control the robot to perform the control action; The intermediate latent features or denoised trajectories generated in the current decision step are stored in the latent feature cache to form an autoregressive accelerated closed loop.

[0013] This invention also provides a video-generative VLA model acceleration system, comprising the following modules: The acquisition module is used to acquire the visual observation image and language instructions for the current decision step; The prior injection module is used to extract intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step based on temporal semantic continuity. After step-aligned injection, these features are used as the initial priors generated by diffusion in the current decision step. The scheduling and adjustment module is used to dynamically adjust the feature cache reuse frequency of the neural network block in the diffusion model based on the state perception result of the current visual observation image. The video generation module is used to perform denoising iterations by combining the initial prior and the adjusted reuse frequency to obtain predicted video features; The motion decoding module is used to perform motion decoding based on the predicted video features and output robot control actions.

[0014] According to the video generative VLA model acceleration system provided by the present invention, the prior injection module is specifically used for: Calculate the image similarity between visual observation images of two adjacent decision steps; Based on a preset mapping relationship, the denoising step size position corresponding to the extracted features is determined according to the image similarity; wherein, the higher the image similarity, the closer the denoising step size position corresponding to the extracted features is to the output end in the denoising trajectory; Based on the determined denoising step size position, the intermediate latent features are extracted from the latent feature cache and injected into the denoising step corresponding to the current decision step as the initial prior for diffusion generation in the current decision step.

[0015] According to the video generative VLA model acceleration system provided by the present invention, the scheduling and adjustment module is specifically used for: Acquire the robot's historical motion trajectories and assess their rate of change; The similarity threshold is dynamically set based on the rate of change: when the rate of change of the historical action trajectory is higher than the preset first threshold, the current state is determined to be critical, and the similarity threshold is increased; when the rate of change of the historical action trajectory is lower than the preset second threshold, the current state is determined to be safe, and the similarity threshold is decreased. Real-time calculation of the feature similarity between the feature map state of the current denoising step and the feature map state of the corresponding denoising step in the previous decision step; If the feature similarity is higher than the dynamically set similarity threshold, the calculation of the current neural network block is skipped and the cached feature result is directly output. If the feature similarity is lower than the dynamically set similarity threshold, the neural network block is activated to perform a complete forward computation, and the calculated new features are updated in the latent feature cache.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video generative VLA model acceleration method as described above.

[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video-generating VLA model acceleration method as described above.

[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video generative VLA model acceleration method as described above.

[0019] This invention provides a method and system for accelerating video-generative VLA models based on temporal semantic continuity. By utilizing temporal semantic continuity, it extracts intermediate latent features or denoised trajectories from the previous decision step and injects them as the initial prior for the current decision step after step alignment. This breaks the traditional paradigm that diffusion models must start from pure noise. By reusing historical denoising paths, it greatly reduces the number of denoising iterations in the current step, fundamentally reducing the instantaneous latency of video generation. At the same time, it dynamically adjusts the feature cache reuse frequency of neural network blocks through state perception results, enabling the model to adaptively allocate computing resources according to the drastic changes in the environment. While significantly reducing redundant computing overhead and improving inference speed, it ensures the real-time updating of key perception features and effectively suppresses covariate shift and error accumulation in the continuous planning process. Thus, while ensuring the smoothness and accuracy of the robot's underlying action output, it achieves efficient real-time deployment of video-generative VLA models on edge devices. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 This is one of the flowcharts of the video generative VLA model acceleration method provided by the present invention.

[0022] Figure 2This is the second flowchart of the video generative VLA model acceleration method provided by the present invention.

[0023] Figure 3 This is a schematic diagram of the structure of the video-generating VLA model acceleration system provided by the present invention.

[0024] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] Current video generative VLA models primarily rely on diffusion models as their underlying generation architecture. The core of the diffusion model lies in its Markov backward denoising trajectory, which uses dozens or even hundreds of forward computations of a massive neural network to gradually restore pure Gaussian noise into a structured predicted video. In continuous robot control tasks, the system typically needs to call the model at a high frequency of 10Hz to 30Hz for planning-execution loops. This means that within each extremely short control cycle, the diffusion model must execute a complete denoising iteration from scratch. This inherent contradiction between the high cost of single-cycle generation and the high-frequency closed-loop calls results in significant instantaneous latency when the system is deployed on the edge, making it difficult to meet the requirements of real-time interaction.

[0027] While academia has proposed feature caching techniques such as DeepCache, their performance in embodied scenarios has been unsatisfactory. These existing techniques typically employ fixed static scheduling strategies, forcibly reusing historical features at preset denoising step lengths. However, in embodied intelligence scenarios, the evaluation criterion is not simply video quality, but task execution success rate. Because static caching ignores the dynamic semantic evolution during task execution (e.g., the visual feature sensitivity of a robotic arm approaching an object is much higher than during open movement), forced reuse can introduce subtle state perception biases. These biases are amplified in the autoregressive closed loop, producing covariate shifts and ultimately causing the action strategy to collapse. Furthermore, existing acceleration schemes treat the generation of each decision step as an isolated event, completely severing the inherent temporal semantic continuity in embodied operations. In practice, high-level instructions remain unchanged, and the environmental background is highly redundant; repeatedly regenerating these known structures from pure noise results in a huge waste of computational resources.

[0028] To address the aforementioned issues, this invention provides an adaptive caching acceleration scheme for video-generative VLA models based on temporal semantic continuity. The core logic of this scheme lies in the synergy of macroscopic step-by-step reuse and microscopic adaptive caching, shifting the denoising starting point of the diffusion model from pure noise to historical priors injected with step alignment, thereby skipping a large number of early denoising steps. Simultaneously, the system introduces a state-aware mechanism to dynamically adjust the network block-level caching frequency based on the fineness of the robot's current operation (e.g., distinguishing between fine grasping and macroscopic displacement). This not only significantly improves inference speed by utilizing semantic redundancy on the time axis but also overcomes the error accumulation problem caused by traditional acceleration techniques by optimizing the smoothness of the feature transfer path, enabling high-frequency, stable real-time closed-loop control of the large VLA model on the edge physical device.

[0029] Before providing a detailed description of the present invention, the following terms used in the embodiments of the present invention will be explained: Embodied Intelligence: Embodied intelligence refers to artificial intelligence systems that possess physical entities (such as robotic arms, humanoid robots, and self-driving cars). Unlike traditional AI (such as ChatGPT running on servers or image recognition software), embodied intelligence can not only process information in the digital world but also perceive the real physical environment through sensors (such as cameras, radar, and tactile sensors) and interact with the physical world through actuators (such as motors and grippers). Its core lies in "learning and recognizing through physical interaction," which is a crucial pathway towards Artificial General Intelligence (AGI).

[0030] The Vision-Language-Motion Model (VLA) is a multimodal large-scale model specifically designed for robot control. Traditional robot control typically requires processing vision, language, and motion in multiple independent modules, while the VLA model achieves end-to-end processing. It can simultaneously receive and understand visual input (such as an image of the current environment captured by a camera) and language commands (such as a human saying "put the red cup on the table"). After internal calculations by the neural network, it directly outputs the low-level physical motion commands required by the robot (such as the rotation angle of each joint of the robotic arm, the movement speed, or the opening and closing state of the gripper).

[0031] Diffusion Model: The diffusion model is currently the most mainstream and powerful underlying algorithm architecture in generative AI, especially in the fields of image and video generation, such as Sora and Midjourney. Its working principle involves two processes: first, a "noise-adding" process, which gradually adds random noise to a clear image until it becomes completely indistinguishable pure noise; second, a "denoising" process (i.e., the generation process), which trains the neural network to learn how to gradually restore the pure noise back to a clear image. In this invention, the diffusion model is used to predict and "generate" future video footage after a robot performs an action.

[0032] Cache reuse is a core optimization technique in computer science and AI acceleration. In complex AI model inference, the computational load is extremely large. The core idea of ​​cache reuse is to temporarily save (cache) the intermediate results (such as feature maps and latent variables) calculated by the model in the previous time step; when the system finds that the current input is highly similar to the previous input (e.g., a static background in a video), it directly retrieves and uses these saved results (reuses them), thus skipping redundant and time-consuming repetitive calculations, significantly reducing latency, and improving the system's real-time response speed.

[0033] Temporal semantic continuity refers to the phenomenon in continuous robot control tasks or video sequences where the visual images and physical environment states corresponding to adjacent time steps (such as step n and step n+1) typically undergo only minor changes. Specifically, temporal semantic continuity manifests as follows: between adjacent decision steps, the rate of change of environmental background features is below a first threshold, while the language instructions remain unchanged.

[0034] The following is combined Figure 1 and Figure 2 This embodiment provides a detailed explanation of the video generative VLA model acceleration method and the data flow of each functional module in the system.

[0035] In some specific embodiments of the present invention, such as Figure 1 As shown, this solution provides a method for accelerating video generative VLA models, including: Step S100: Obtain the visual observation image and language instructions for the current decision step; Step S200: Based on temporal semantic continuity, extract the intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step in the latent feature cache, and inject them after step alignment as the initial priors for the diffusion generation of the current decision step. Step S300: Based on the state perception result of the current visual observation image, dynamically adjust the feature cache reuse frequency of the neural network block in the diffusion model; Step S400: Combine the initial prior and the adjusted reuse frequency to perform denoising iteration to obtain the predicted video features; Step S500: Decode the predicted video features to generate and output robot control actions.

[0036] The following is through Figure 2 The specific embodiments will provide a detailed description of the above steps.

[0037] Step S100: Obtain the visual observation image and language instructions for the current decision step; In this step, at the current decision step (time step) The system uses an input processing module (such as...) Figure 2 (As shown in component 1) it receives multimodal input information. Specifically, this includes: the current visual observation image acquired through the robot's front-facing camera. (or denoted as the current frame) ) and natural language commands issued by the user. (e.g., "open the door").

[0038] Step S200: Based on temporal semantic continuity, extract the intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step in the latent feature cache, and inject them after step alignment as the initial priors for the diffusion generation of the current decision step. This step is used to perform straddle feature reuse.

[0039] In traditional video generative VLA models, the diffusion model for each decision step performs a complete denoising iteration starting from pure Gaussian noise. However, in continuous robotic tasks, visual observations between adjacent decision steps exhibit high similarity (i.e., temporal semantic continuity), while high-level language instructions remain unchanged. This step fundamentally alters the generative paradigm of the diffusion model by reusing intermediate latent features computed in the previous decision step. By shifting the denoising starting point of the current decision step from pure noise to historical priors injected with step alignment, 40%–60% of the early denoising steps are safely skipped, significantly reducing the video generation latency per decision.

[0040] In this embodiment, this step can be performed by Figure 2 The skip selector and latent feature cache module work together to complete this task.

[0041] In some possible embodiments of the present invention, the step of extracting intermediate latent features generated in the diffusion model denoising process of the previous decision step and pre-stored in the latent feature cache, and then injecting them after step alignment as the initial prior for diffusion generation in the current decision step, specifically includes: Calculate the image similarity between the visual observation images of the previous decision step and the current decision step; Based on a preset mapping relationship, the denoising step size position corresponding to the extracted features is determined according to the image similarity; wherein, the higher the image similarity, the closer the denoising step size position corresponding to the extracted features is to the output end in the denoising trajectory; Based on the determined denoising step size position, the intermediate latent features are extracted from the latent feature cache and injected into the denoising step corresponding to the current decision step as the initial prior for diffusion generation in the current decision step.

[0042] Specifically, this embodiment provides an implementation method for generating initial priors in the current decision step diffusion. For example, it can adopt... Figure 2 In one specific embodiment, the skip selector performs the following steps to complete the extraction and injection of the initial prior: First, calculate the visual observation image from the previous decision step. Visual observation images of the current decision step Image similarity between In practice, this can be achieved by calculating the cosine similarity of the feature vectors of two image frames, the structural similarity index (SSIM), or a negative correlation function of pixel-level differences.

[0043] Secondly, the system has a built-in mapping function. ,in, This represents the denoising step size position corresponding to the feature extracted from the latent feature cache. The mapping function is configured such that: when... When it exceeds the first threshold (e.g., 0.95), Smaller values ​​extract features closer to the output of the denoising trajectory (which is already close to a clear image); when... When it falls below the second threshold (e.g., 0.6), A larger value means extracting features closer to the input end of the denoised trajectory (the pure noise end). Through this dynamic mapping, the system can adaptively determine the depth of historical feature reuse based on the drastic changes in the environment.

[0044] Specifically, this embodiment also provides an implementation method for prior extraction and alignment. For example, it employs... Figure 2 If the latent feature cache module determines continuity, the system directly extracts the intermediate latent features or denoising trajectories generated during the diffusion model denoising process in the previous decision step from the latent feature cache (e.g., ...). Figure 2 (The intermediate representation marked as "reuse").

[0045] Because the embodied task involves extremely small displacements within a very short time step, the diffusion model possesses the semantic continuity to handle these minute displacements. This invention achieves model convergence simply by directly injecting features. Furthermore, this step transforms historical features into initial priors for the diffusion generation in the current decision step, breaking away from the traditional diffusion model's requirement to start from pure Gaussian noise (…). Figure 2 Steps in By eliminating the limitations of the initial generation process, the early redundant denoising steps are skipped directly (typically 40%-60% of the early steps can be skipped), significantly shortening the generation time.

[0046] Step S300: Based on the state perception result of the current visual observation image, dynamically adjust the feature cache reuse frequency of the neural network block in the diffusion model; This step is used to perform adaptive cache scheduling. Specifically, after obtaining the initial prior, the system enters the adaptive cache scheduling phase, and together with the feature generation phase of step S400, accelerates video generation. Steps S300 and S400 are respectively... Figure 2 The similarity checker 4 and the adaptive block cache module 5 are completed.

[0047] In some possible embodiments of the present invention, the step of dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model includes: The motion amplitude or semantic offset of the visual observation image at the current decision step is evaluated through a state-aware mechanism. When the robot is in the fine operation phase, reduce the frequency of feature cache reuse to perform complete forward computation; When the robot is in the simple movement stage, increase the frequency of feature cache reuse to reuse features from historical network layers.

[0048] Specifically, this embodiment provides a state-aware implementation method in which the system evaluates the motion amplitude or semantic offset of the currently observed visual image through a state-aware mechanism.

[0049] The state-aware mechanism is key to the adaptive acceleration of this invention. Its principle is to quantify the drastic changes in the current environment by analyzing optical flow information or feature cosine distance between consecutive frames, thereby inferring the robot's current operational stage (such as macroscopic path planning or fine grasping). In the fine operation stage, visual features are extremely sensitive to minute positional changes; excessive reuse of historical features at this stage can lead to accumulated motion deviations. In the simple movement stage, background and object structure changes slowly, making high-frequency reuse suitable. On-demand computation is possible: ensuring accuracy in critical stages and maximizing acceleration in non-critical stages, thus achieving a dynamic balance between generation speed and control precision.

[0050] When the robot is in the "fine operation stage" (such as approaching an object), the reuse frequency is automatically reduced; when it is in the "simple movement stage", the reuse frequency is automatically increased.

[0051] In some possible embodiments of the present invention, the step of dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model further includes: Acquire the robot's historical motion trajectories and assess their rate of change; The similarity threshold is dynamically set based on the rate of change: when the rate of change of the historical action trajectory is higher than the preset first threshold, the current state is determined to be critical, and the similarity threshold is increased; when the rate of change of the historical action trajectory is lower than the preset second threshold, the current state is determined to be safe, and the similarity threshold is decreased. Calculate the feature similarity between the feature map of the current denoising step and the feature map of the corresponding denoising step of the previous decision step in real time; If the feature similarity is higher than the similarity threshold, the calculation of the current neural network block is skipped, and the cached feature result is directly output. If the feature similarity is lower than the similarity threshold, the neural network block is activated to perform a complete forward computation, and the calculated new features are updated in the latent feature cache. Specifically, this embodiment provides a block-level dynamic reuse implementation method, which acquires the robot's historical motion trajectory and evaluates its rate of change; dynamically sets a similarity threshold based on the rate of change; calculates the feature similarity between the feature state of the current denoising step and the feature map of the corresponding denoising step in the previous decision step in real time; and executes the function when the feature similarity is higher than the similarity threshold. Figure 2 The dashed reuse path shown skips self-attention and feedforward computations; conversely, the solid path is executed, and the cache is updated using a summation operation after computation.

[0052] In a preferred embodiment, the similarity threshold of the adaptive block cache module is not a fixed value, but is dynamically adjusted based on the rate of change of the robot's historical motion trajectories. The system maintains a sliding window to record past... Action instructions for each decision step And calculate its rate of change. .when When the similarity threshold exceeds a preset first threshold, the robot is determined to be in a critical state requiring high-precision control (such as approaching a target or a sudden change in speed direction). At this point, the system dynamically increases the similarity threshold (e.g., from 0.7 to 0.9). This means that reuse is only triggered when the feature similarity is extremely high (i.e., the change is minimal), thus forcing the model to perform complete calculations in most cases to ensure accuracy. Conversely, when... When the similarity threshold is below a preset second threshold, it is determined to be a safe state (such as uniform linear motion). The system then lowers the similarity threshold (for example, to 0.5), thereby allowing more feature reuse and maximizing the acceleration effect.

[0053] It is worth noting that the smoothness is evaluated by calculating the rate of change of historical action trajectories. The lower the rate of change, the higher the smoothness, and vice versa, it indicates that the operation is in a critical state.

[0054] In addition, in specific embodiments, besides considering the rate of change of motion, it is sometimes necessary to consider the distance from the target. When approaching the target, the reuse rate can be reduced.

[0055] More specifically, in the neural network of the diffusion model, for each computational block, the feature similarity between the current denoising step feature and the corresponding feature map of the previous decision step is calculated in real time. If it is higher than a preset similarity threshold, the reuse path is executed, and the cached features are directly output; if it is lower than the threshold, the computational path is executed, the network block is activated to perform a complete forward computation, and the cache is updated.

[0056] For example, see further. Figure 2 2. Adaptive Block Caching Logic: The neural network (NN) within the diffusion model is divided into multiple computational blocks. During execution, the system exhibits two logical paths: First, the dashed reuse path: When the similarity checker determines that the similarity between the current feature and the cache is higher than a safety threshold, the system directly shuts down the NN block or skips its computation instructions, directly passing the feature map in the cache to the next layer via a bypass. Second, the solid computation path: When the similarity is low, the NN block is activated to perform a complete matrix multiplication operation, and the calculation result is processed via the summation operation shown in the diagram (…). (Operation) Update to the latent feature cache to ensure that the stored content is always dynamically updated with the environment.

[0057] This embodiment achieves a dynamic balance between generation speed and control precision through a dynamic routing mechanism, effectively avoiding error accumulation or covariate offset caused by static caching, and ensuring the generation accuracy of key action nodes.

[0058] Step S400: Combine the initial prior and the adjusted reuse frequency to perform denoising iteration to obtain the predicted video features; In the denoising steps that are not skipped, each neural network block of the diffusion model may still generate a significant amount of redundant computation, especially those blocks corresponding to static background regions. Before entering each block, the similarity checker quickly compares the current feature with the cached features from the same denoising step in the previous decision step. When the similarity is above a threshold, it indicates that the semantics of the region are almost unchanged, and the cached results can be directly reused, skipping the complex computations of the self-attention and feedforward networks. Conversely, when the similarity is below the threshold (such as the minute movements of a robotic arm's end effector during fine manipulation), the complete forward computation is forced to ensure accuracy. This mechanism further reduces redundant computation without losing key details, while simultaneously preventing the accumulation of errors in the autoregressive closed loop by updating the cache.

[0059] For example, see Figure 2 The system injects the initial prior obtained in step S200 into the diffusion model. This differs from the traditional method of injecting from pure noise. Initially, the model directly starts from the intermediate steps corresponding to historical features. Enter the denoising trajectory. At the same time, combined with the block-level reuse strategy determined in step S300, cached features are preferentially called when performing neural network forward computation, thereby completing fine denoising iteration and outputting high-quality predicted video features.

[0060] Step S500: Decode the predicted video features to generate and output robot control actions.

[0061] This step is used to perform action decoding and autoregressive loop closure. Unlike single-step optimization techniques, this embodiment employs an autoregressive loop closure method. After the robot performs the current action, the environmental state evolves. At this point, the latent features generated in the current step are stored in a cache as initial priors for the next decision step. This configuration in this embodiment enables the system to continuously utilize temporal semantic continuity, avoiding recalculation at the beginning for each decision step. Furthermore, as the task progresses, the historical feature library accumulates, the effectiveness of cross-step reuse gradually increases, and the overall inference latency tends to stabilize, thereby achieving high-frequency real-time control of the video-generative VLA model on the edge device.

[0062] It's worth noting that directly injecting features from the previous time step into an intermediate step is equivalent to forcibly inserting a foreign variable into the Markov chain of the diffusion model. This can lead to discontinuous jumps in the generated video on the timeline. If such jumps are passed to the motion decoder, the robot may produce instantaneous high-frequency tremors.

[0063] In some possible embodiments of the present invention, the step of performing denoising iterations by combining the initial prior and the adjusted reuse frequency to obtain predicted video features specifically includes: The feature transfer path from predicted video features to the action decoder is optimized. The smoothness of the output action sequence on the time axis is maintained by the feature cache reuse, ensuring that the smoothness and accuracy of the output action sequence are maintained while eliminating redundant calculations through feature reuse.

[0064] In possible implementations, feature interpolation or time smoothing operators can be used to mitigate the abrupt jumps caused by this injection.

[0065] In a possible embodiment, by establishing a jump link or residual path between the video generation module and the motion decoder, the intermediate latent features extracted during the acceleration process are directly fed to the decoder, reducing visual feature conversion loss, thereby reducing hierarchical transmission delay, lowering end-to-end communication delay, and ensuring the real-time performance of motion response.

[0066] Specifically, in order to ensure the smoothness of the action while implementing feature reuse acceleration, this embodiment provides an implementation method for feature transfer and decoding. Before performing action decoding, the predicted video features can be weighted and fused between time steps by introducing a smoothing filter or trajectory interpolation algorithm to eliminate feature abrupt changes caused by block-level reuse, avoid high-frequency vibration of the robot actuator, and ensure the continuity of output actions and the executability of physical operations in the accelerated state.

[0067] In a possible implementation, before passing the accelerated-generated predicted video features to the motion decoder, the system calculates the motion vector residual between the current step features and historical features in the latent feature cache. If non-physical jumps occur due to feature reuse, the system applies an exponentially weighted moving average algorithm to interpolate and smooth the predicted trajectory, ensuring that the output underlying physical motion commands (such as the angular velocity of a robotic arm joint) conform to the motion constraints of the physical entity and avoiding sudden tremors.

[0068] For example, see still Figure 2 The target video sequence output module outputs accelerated-generated predicted video features. The system optimizes the transmission path from video features to the action decoder. The action decoder receives these features, generates and outputs the robot control action for the current decision step. .

[0069] In some possible embodiments of the present invention, after performing motion decoding on the predicted video features to generate and output robot control actions, the method further includes: Control the robot to perform the control action; The intermediate latent features or denoised trajectories generated in the current decision step are stored in the latent feature cache to form an autoregressive accelerated closed loop.

[0070] Specifically, this embodiment provides an implementation method for physical execution and closed-loop update to drive the robot to perform control actions. This enables it to interact with the physical environment, updating the environment state to... Simultaneously, the intermediate latent features or denoised trajectories generated in the current step ( Figure 1 The middle is recorded as Store in the latent feature cache, such as Figure 2 The terms "1. Latent Feature Cache" and "2. Adaptive Block Cache" are shown in the table.

[0071] This embodiment stores the features of the current step back into the cache, making prior preparations for accelerating the next decision step, forming an efficient autoregressive closed loop, and ensuring that the system maintains smooth and accurate action output in continuous tasks.

[0072] This invention leverages the temporal semantic continuity in embodied intelligence tasks to construct a dual-acceleration architecture that combines macro-level cross-decision-step feature reuse with micro-level adaptive block cache frequency adjustment. Figure 1 and Figure 2 As can be seen, the system no longer processes each time step in isolation, but forms an autoregressive closed loop through latent feature caching.

[0073] To further illustrate the core content of this invention, the following will still be combined with... Figure 2 The macroscopic straddle multiplexing and microscopic adaptive block caching mechanisms of the present invention are described in detail with several specific embodiments.

[0074] Example 1: Macroscopic Accelerated Implementation Based on Step Feature Reuse In this embodiment, the system utilizes the continuity of environmental evolution in the robot's continuous operation tasks to achieve acceleration at the macro level.

[0075] First, after acquiring the visual observation image and language instructions in the current decision step, the system does not immediately perform full denoising. Instead, it receives continuous visual observations by skipping the selector and evaluates the continuity of the current task by analyzing semantic changes and motion amplitudes between adjacent frames (e.g., calculating pixel differences or feature cosine distance). If the difference value is below a preset threshold, it is determined that the semantics are highly continuous, and the latent feature caching module is activated.

[0076] This module no longer allows the diffusion model to proceed from the steps Instead of generating (pure noise), it directly extracts the previous decision step from the latent feature cache in the intermediate denoising step (such as step...). The system generates intermediate latent features or denoised trajectories. To compensate for the viewpoint shift caused by the robot's own motion, the system performs a specific number of step-aligned injections, transforming the extracted historical features into the initial prior injection model for the current step. Thus, the model directly skips a large number of redundant early background generation steps, significantly shortening the generation time.

[0077] It is worth noting that this embodiment uses visually observed images to calculate similarity. Alternatively, robot-generated perception information (such as data from joint angle sensors and inertial measurement units, IMUs) can be directly used to assess motion amplitude. When IMU data or the rate of change of joint angular velocity is below a threshold, the environment can also be determined to be quasi-static, thereby triggering step feature reuse.

[0078] Example 2: Micro-acceleration Implementation of State-Aware Adaptive Block Cache: In the specific execution of the denoising iteration, this embodiment further implements micro-acceleration based on the neural network block level.

[0079] The system introduces a similarity checker to monitor the feature state of the current denoising step in real time. When the state awareness mechanism assesses that the robot is in a "simple movement" stage (such as the robotic arm moving in an open area), the system determines that the feature difference is low at this time, automatically increases the reuse frequency of the adaptive block cache module, and directly skips the calculation of non-critical network layers.

[0080] Conversely, when the robot enters the "fine manipulation" stage (such as when a finger is about to touch an object), the state perception module detects an increase in semantic offset. The system then automatically reduces the reuse frequency and activates the neural network block to perform complete forward computation to ensure the accuracy of the underlying action output. This dynamic routing mechanism, while ensuring the success rate of the task, achieves on-demand allocation of computing resources based on the stage-specific characteristics of the embodied task.

[0081] Example 3: End-to-end collaboration and autoregressive closed-loop construction: This embodiment further constructs an end-to-end "video generation-motion prediction" collaborative framework. The motion decoder receives the predicted video features generated by the aforementioned dual acceleration mechanism. To prevent feature reuse during the acceleration process from affecting the smoothness of the motion, the system optimizes the feature transmission path to ensure that the generated video feature stream maintains logical coherence on the timeline.

[0082] Once the action decoder outputs the current control action, the robot executes the action and interacts with the environment. Simultaneously, the system stores the latest latent features or denoised trajectory generated in the current decision step back into the latent feature cache, preparing for acceleration in the next decision step. This autoregressive closed-loop mode enables video-generative VLA models to achieve up to 4.4 times speedup on edge devices, while improving the task success rate by approximately 8.2%, truly promoting the deployment of generative large models in real-world robotic scenarios.

[0083] To verify the technical effectiveness of this invention, the method of this invention was compared with a benchmark method (without acceleration) and existing accelerated methods such as Deepcached on a holographic device equipped with an NVIDIA Orin platform. The experiment employed continuous operation tasks such as "opening a door" and "grabbing a moving object," and the results are shown in Table 1. Table 1 shows the actual test results on the holographic device equipped with an NVIDIA Orin platform. The average video generation time of the proposed solution was reduced from 5.49s to 1.26s compared to the benchmark method, achieving a speedup of 4.4 times. Furthermore, in the "opening a door" task, compared to the risk of motion crashes in Deepcached, the success rate of this application remained stable at 44.6%, with only minimal loss compared to the baseline method, achieving Pareto optimality in both accuracy and speed.

[0084] Table 1 Task success rate Average video generation time acceleration ratio Baseline method 46.6% 5.49s 1.0x Deepcached acceleration methods 36.4% 3.37s 1.6x Baseline + Latent Feature Cache 51.2% 2.47s 2.2x Baseline + Adaptive Block Cache 42.0% 5.02s 1.1x This invention 44.6% 1.26s 4.4x Experimental data shows that, while maintaining a high task success rate (44.6%), this invention achieves an average video generation time acceleration of approximately 4.4 times. This solution not only solves the bottleneck of high computational cost in diffusion models but also effectively suppresses error accumulation caused by cache reuse, truly promoting the real-time deployment of video-generative VLA models in real-world robotic scenarios.

[0085] This invention targets semantic features rather than pure pixel alignment; it is suitable for high-frequency control, such as when the displacement is extremely small within a very short time step, and the denoising capability of the diffusion model itself is sufficient to handle such tiny displacement deviations.

[0086] The video-generating VLA model acceleration system provided by this invention is described below. The video-generating VLA model acceleration system described below can be referred to in correspondence with the video-generating VLA model acceleration method described above.

[0087] In some specific embodiments of the present invention, such as Figure 3 As shown, this solution provides a video-generative VLA model acceleration system, including: The acquisition module 10 is used to acquire the visual observation image and language instructions for the current decision step; The prior injection module 20 is used to extract intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step based on temporal semantic continuity, and after step-aligned injection, they are used as the initial priors generated by diffusion in the current decision step. The scheduling and adjustment module 30 is used to dynamically adjust the feature cache reuse frequency of the neural network block in the diffusion model based on the state perception result of the current visual observation image. The video generation module 40 is used to perform denoising iteration by combining the initial prior and the adjusted reuse frequency to obtain predicted video features; The motion decoding module 50 is used to perform motion decoding based on the predicted video features and output robot control actions.

[0088] Possibly, the prior injection module is specifically used for: Calculate the image similarity between visual observation images of two adjacent decision steps; Based on a preset mapping relationship, the denoising step size position corresponding to the extracted features is determined according to the image similarity; wherein, the higher the image similarity, the closer the denoising step size position corresponding to the extracted features is to the output end in the denoising trajectory; Based on the determined denoising step size position, the intermediate latent features are extracted from the latent feature cache and injected into the denoising step corresponding to the current decision step as the initial prior for diffusion generation in the current decision step.

[0089] Possibly, the scheduling adjustment module is specifically used for: Acquire the robot's historical motion trajectories and assess their rate of change; The similarity threshold is dynamically set based on the rate of change: when the rate of change of the historical action trajectory is higher than the preset first threshold, the current state is determined to be critical, and the similarity threshold is increased; when the rate of change of the historical action trajectory is lower than the preset second threshold, the current state is determined to be safe, and the similarity threshold is decreased. Real-time calculation of the feature similarity between the feature map state of the current denoising step and the feature map state of the corresponding denoising step in the previous decision step; If the feature similarity is higher than the dynamically set similarity threshold, the calculation of the current neural network block is skipped and the cached feature result is directly output. If the feature similarity is lower than the dynamically set similarity threshold, the neural network block is activated to perform a complete forward computation, and the calculated new features are updated in the latent feature cache.

[0090] In a possible embodiment, the scheduling adjustment module is further configured to: Based on real-time dynamic changes in the environment, the allocation of computing resources for each neural network computing block in the diffusion model is dynamically adjusted to achieve a dynamic balance between generation speed and control precision.

[0091] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a video generative VLA model acceleration method. This method includes: acquiring the visual observation image and language instructions of the current decision step; based on temporal semantic continuity, extracting intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step and pre-stored in the latent feature cache, injecting them after step alignment as the initial prior for diffusion generation of the current decision step; dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model according to the state perception result of the current visual observation image; performing denoising iterations in combination with the initial prior and the adjusted reuse frequency to obtain predicted video features; and performing action decoding on the predicted video features to generate and output robot control actions.

[0092] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0093] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the video generative VLA model acceleration method provided by the above methods. The method includes: acquiring the visual observation image and language instructions of the current decision step; based on temporal semantic continuity, extracting intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step and pre-stored in the latent feature cache, and injecting them after step alignment as the initial prior for diffusion generation of the current decision step; dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model according to the state perception result of the current visual observation image; performing denoising iteration in combination with the initial prior and the adjusted reuse frequency to obtain predicted video features; and performing action decoding on the predicted video features to generate and output robot control actions.

[0094] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the video generative VLA model acceleration method provided by the above methods. The method includes: acquiring the visual observation image and language instructions of the current decision step; based on temporal semantic continuity, extracting intermediate latent features or denoising trajectories generated in the diffusion model denoising process of the previous decision step and pre-stored in the latent feature cache, injecting them after step alignment as the initial prior for diffusion generation of the current decision step; dynamically adjusting the feature cache reuse frequency of the neural network block in the diffusion model according to the state perception result of the current visual observation image; performing denoising iteration in combination with the initial prior and the adjusted reuse frequency to obtain predicted video features; and performing action decoding on the predicted video features to generate and output robot control actions.

[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0096] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for accelerating a video-generative VLA model, characterized in that, include: Acquire visual observations and verbal commands for the current decision step; Based on temporal semantic continuity, intermediate latent features or denoising trajectories generated in the previous decision step during the diffusion model denoising process are extracted from the latent feature cache and injected after step alignment as the initial prior for the diffusion generation of the current decision step. Based on the state perception results of the current visual observation image, the feature cache reuse frequency of the neural network block in the diffusion model is dynamically adjusted. The denoising iteration is performed by combining the initial prior and the adjusted reuse frequency to obtain the predicted video features; The predicted video features are decoded to generate and output robot control actions.

2. The video-generating VLA model acceleration method according to claim 1, characterized in that, The extraction of intermediate latent features generated during the diffusion model denoising process in the previous decision step and pre-stored in the latent feature cache, which are then injected after step alignment, serves as the initial prior for diffusion generation in the current decision step. Specifically, this includes: Calculate the image similarity between the visual observation images of the previous decision step and the current decision step; Based on a preset mapping relationship, the denoising step size position corresponding to the extracted features is determined according to the image similarity; wherein, the higher the image similarity, the closer the denoising step size position corresponding to the extracted features is to the output end in the denoising trajectory; Based on the determined denoising step size position, the intermediate latent features are extracted from the latent feature cache and injected into the denoising step corresponding to the current decision step as the initial prior for diffusion generation in the current decision step.

3. The video generative VLA model acceleration method according to claim 1, characterized in that, The step of dynamically adjusting the feature cache reuse frequency of neural network blocks in the diffusion model specifically includes: The motion amplitude or semantic offset of the visual observation image at the current decision step is evaluated through a state-aware mechanism. When the robot is in the fine operation phase, reduce the frequency of feature cache reuse to perform complete forward computation; When the robot is in the simple movement stage, increase the frequency of feature cache reuse to reuse features from historical network layers.

4. The video generative VLA model acceleration method according to claim 3, characterized in that, The step of dynamically adjusting the feature cache reuse frequency of neural network blocks in the diffusion model further includes: Acquire the robot's historical motion trajectories and assess their rate of change; The similarity threshold is dynamically set based on the rate of change: when the rate of change of the historical action trajectory is higher than the preset first threshold, the current state is determined to be critical, and the similarity threshold is increased; when the rate of change of the historical action trajectory is lower than the preset second threshold, the current state is determined to be safe, and the similarity threshold is decreased. Calculate the feature similarity between the feature map of the current denoising step and the feature map of the corresponding denoising step of the previous decision step in real time; If the feature similarity is higher than the similarity threshold, the calculation of the current neural network block is skipped, and the cached feature result is directly output. If the feature similarity is lower than the similarity threshold, the neural network block is activated to perform a complete forward computation, and the calculated new features are updated in the latent feature cache.

5. The video generative VLA model acceleration method according to claim 1, characterized in that, The step of performing denoising iterations by combining the initial prior and the adjusted reuse frequency to obtain predicted video features specifically includes: The feature transfer path from predicted video features to the action decoder is optimized, and the smoothness of the output action sequence on the time axis is maintained by reusing the feature cache.

6. The video generative VLA model acceleration method according to claim 1, characterized in that, After decoding the predicted video features to generate and output robot control actions, the process further includes: Control the robot to perform the control action; The intermediate latent features or denoised trajectories generated in the current decision step are stored in the latent feature cache to form an autoregressive accelerated closed loop.

7. A video-generating VLA model acceleration system, characterized in that, include: The acquisition module is used to acquire the visual observation image and language instructions for the current decision step; The prior injection module is used to extract intermediate latent features or denoising trajectories generated in the previous decision step during the diffusion model denoising process, which are pre-stored in the latent feature cache based on temporal semantic continuity. After being injected with step alignment in the time dimension, these features are used as the initial priors generated by the diffusion of the current decision step. The scheduling and adjustment module is used to dynamically adjust the feature cache reuse frequency of the neural network block in the diffusion model based on the state perception result of the current visual observation image. The video generation module is used to perform denoising iterations by combining the initial prior and the adjusted reuse frequency to obtain predicted video features; The motion decoding module is used to perform motion decoding based on the predicted video features and output robot control actions.

8. The video-generating VLA model acceleration system according to claim 7, characterized in that, The prior injection module is specifically used for: Calculate the image similarity between visual observation images of two adjacent decision steps; Based on a preset mapping relationship, the denoising step size position corresponding to the extracted features is determined according to the image similarity; wherein, the higher the image similarity, the closer the denoising step size position corresponding to the extracted features is to the output end in the denoising trajectory; Based on the determined denoising step size position, the intermediate latent features are extracted from the latent feature cache and injected into the denoising step corresponding to the current decision step as the initial prior for diffusion generation in the current decision step.

9. The video-generating VLA model acceleration system according to claim 7, characterized in that, The scheduling and adjustment module is specifically used for: Acquire the robot's historical motion trajectories and assess their rate of change; The similarity threshold is dynamically set based on the rate of change: when the rate of change of the historical action trajectory is higher than the preset first threshold, the current state is determined to be critical, and the similarity threshold is increased; when the rate of change of the historical action trajectory is lower than the preset second threshold, the current state is determined to be safe, and the similarity threshold is decreased. Real-time calculation of the feature similarity between the feature map state of the current denoising step and the feature map state of the corresponding denoising step in the previous decision step; If the feature similarity is higher than the dynamically set similarity threshold, the calculation of the current neural network block is skipped and the cached feature result is directly output. If the feature similarity is lower than the dynamically set similarity threshold, the neural network block is activated to perform a complete forward computation, and the calculated new features are updated in the latent feature cache.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the video-generating VLA model acceleration method as described in any one of claims 1 to 6.