Visual motion prediction acceleration method and system directly based on pre-training model

By using the denoising processing method of UNet architecture in the field of robot control, combining feature difference calculation and Fourier spectrum energy analysis, the noise prediction characteristics are dynamically adjusted, and the problems of slow inference speed and poor noise distribution are solved, and the combination of fast inference and high accuracy is achieved.

CN120029162AActive Publication Date: 2025-05-23SHANGHAI MAJIKE IND INTELLIGENCE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510510837.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The prior art is difficult to simultaneously improve the inference speed of the diffusion model and maintain prediction performance in the field of robot control, resulting in slow inference speed and poor noise distribution.

Method used

Using the UNet-based denoising processing method, by calculating the feature difference of the output characteristics of the encoder between adjacent denoising steps, we judge whether the current denoising step is a non-critical step or a critical step, and the encoder characteristics and intermediate block characteristics of the most recent critical step are used as inputs of the decoder blocks of subsequent non-critical steps, dynamic adjustment of the noise prediction characteristics is performed.

Benefits of technology

It realizes that the inference process is significantly accelerated by reusing the encoder features without training the student model, while maintaining performance similar to the original model, ensuring the accuracy of the generated trajectory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029162A_ABST
    Figure CN120029162A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and discloses a visual motion prediction acceleration method and system directly based on a pre-training model, and the method comprises the steps: carrying out the denoising of an input signal based on a UNet which comprises an encoder, an intermediate block and a decoder; calculating the characteristic difference of the output characteristics of the encoder between the adjacent denoising steps; judging whether the current denoising step is a non-key step or a key step according to the characteristic difference; and taking the encoder feature of the nearest key step and the intermediate block feature as the input of the decoder block of the subsequent non-key step to obtain a noise prediction feature. The invention provides a new method, namely a fast strategy, which does not need to be retrained, the fast strategy can be regarded as a powerful and accelerated alternative scheme for learning a diffusion strategy controlled by a visual motion robot, and compared with an existing acceleration method, a comparison result shows that the fast strategy has the highest success rate in the aspect of visual motion reasoning speed; and the effectiveness and superiority are proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a visual motion prediction acceleration method and system that can be directly based on a pre-training model. Background Art

[0002] Diffusion models have become the mainstream paradigm for generative AI, and have achieved significant breakthroughs in multiple applications ranging from low-level computer vision to high-level vision-language tasks. In addition, some recent studies have also demonstrated the excellent performance of diffusion models in imitation learning for robot control. However, diffusion strategies often have slow reasoning speeds due to the large number of sampling steps, which limits the application of diffusion models in various generative applications and resource-constrained robotic systems.

[0003] In order to improve the inference speed of diffusion models, a variety of acceleration methods have been proposed in the field of text-to-image generation in recent years. These methods can be roughly divided into two categories: reducing the number of sampling steps and knowledge distillation. The first type of method reduces the number of sampling steps to achieve acceleration by minimizing the first-order discretization error of ordinary differential equations, such as implicit denoising diffusion models and design space methods based on diffusion-based generative models. Although this type of method can reduce the sampling steps to a certain extent, it is accompanied by the problem of reduced prediction performance caused by factors such as modifying the Markov property, which will lead to poor distribution of predicted noise and reduced prediction performance. The second type of method gradually distills a student model with faster inference from a teacher model with slower inference speed, and then achieves fast inference based on the trained student model. However, this type of method needs to further add the training cost of the student model to the training cost of the pre-trained teacher model, which takes more time and is difficult to generalize to other tasks.

[0004] In the field of robotics, there are currently only a few methods such as consistency strategies to achieve the acceleration of diffusion strategies. However, these methods often cannot take into account both reasoning speed and accuracy at the same time. Summary of the invention

[0005] The main purpose of the present invention is to solve the technical problem that the prior art cannot take into account both the speed and accuracy of reasoning. The visual motion prediction acceleration method based on the pre-trained model directly includes the following steps: Based on UNet, the input signal is denoised, and the UNet includes an encoder, an intermediate block, and a decoder; Calculate the feature differences of the encoder's output features between adjacent denoising steps; According to the feature difference, determining whether the current denoising step is a non-critical step or a critical step; The encoder features and intermediate block features of the most recent critical step are used as the input of the decoder block of the subsequent non-critical step to obtain the noise prediction features; Visual motion prediction is completed based on noise prediction features.

[0006] The calculating feature difference of output features of the encoder between adjacent denoising steps includes: Calculate the mean square error of the output features of each module of the one-dimensional convolutional UNet between adjacent denoising steps; The calculated mean square error is normalized according to the batch size, dimension and field of view length of the feature to obtain the feature difference. The formula is as follows: Among them, b is the batch size, d is the dimension, h is the field of view length, and i represents the input of the current module. is the feature output by the kth module in the nth denoising step, is the feature output by the kth module in the n-1th denoising step, Represents the feature difference between the kth module in the nth denoising step and the n-1th denoising step.

[0007] The step of judging whether the current denoising step is a non-critical step or a critical step according to the feature difference is specifically as follows: The first denoising step is the key step, and the subsequent denoising steps are judged as follows: The decoder output features are mapped to the frequency domain through Fourier transform, and the spectrum energy E is calculated. The formula is: Where F(·) represents the component amplitude, and (u,v) represents the position in the spectrum diagram; Constructing the scoring function : Among them, s1 and s2 are skip connection features, b is the output feature of the intermediate module; n represents the nth denoising step interval, M represents the mean square error of the output feature of the corresponding denoising step interval, T represents different tasks, and w is the weight; Set the number of critical steps; According to the set number of key steps, the steps corresponding to the step interval with the highest score are selected as key steps, and the others are non-key steps.

[0008] The noise prediction feature f of the non-critical step is scaled by factor c d2 Perform dynamic adjustment and output the corrected frequency domain feature f′ d2 , the formula is: .

[0009] The scaling factor c is calculated based on the ratio of the spectral energy change between the output features of the non-critical step and the corresponding critical step, and the formula is: Among them, E nonkey is the spectral energy of the output feature of the non-critical step, E key is the spectral energy of the output feature of the key step.

[0010] The scaling factor c is obtained by the following steps: The correction parameter dictionaries are pre-built, denoted as dict{DDIM} and dict{EDM}, which contain the scaling factors for each task.

[0011] The second aspect of the present invention provides a visual motion prediction acceleration system that can be directly based on a pre-trained model, comprising: A denoising processing unit, used for performing denoising processing on an input signal based on UNet, wherein the UNet includes an encoder, an intermediate block, and a decoder; A feature difference calculation unit, used to calculate the feature difference of the output features of the encoder between adjacent denoising steps; The judging unit is used to judge whether the current denoising step is a non-critical step or a critical step according to the feature difference, specifically: The first denoising step is the key step, and the subsequent denoising steps are judged as follows: The decoder output features are mapped to the frequency domain through Fourier transform, and the spectrum energy E is calculated. The formula is: Where F(·) represents the component amplitude, and (u,v) represents the position in the spectrum diagram; Building a scoring function : Among them, s1 and s2 are skip connection features, b is the output feature of the intermediate module; n represents the nth denoising step interval, M represents the mean square error of the output feature of the corresponding denoising step interval, T represents different tasks, and w is the weight; Set the number of critical steps; According to the set number of key steps, the step corresponding to the step interval with the highest score is selected as the key step, and the others are non-key steps; The noise prediction feature unit is used to use the encoder features and intermediate block features of the most recent key step as inputs to the decoder block of the subsequent non-key step to obtain noise prediction features.

[0012] The third aspect of the present invention provides an electronic device, comprising: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected via lines; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned visual motion prediction acceleration method that can be directly based on a pre-trained model as described above.

[0013] A fourth aspect of the present invention provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned visual motion prediction acceleration method that can be directly based on a pre-trained model as described above.

[0014] The present invention has the following beneficial effects: This paper proposes a novel fast inference robot control strategy, which achieves inference acceleration and maintains performance close to the original model by reusing encoder features without training a student model.

[0015] The present invention designs a dynamic key step selection strategy based on Fourier spectrum energy, which can adaptively select key steps according to different robot tasks, rather than using a fixed uniform sampling method, to more accurately retain task performance. This mechanism can directly screen key steps based on the characteristics of the noise prediction network itself for different tasks, without the need for additional training of the neural network to select the optimal key steps.

[0016] The present invention designs a Fourier spectrum energy noise correction strategy, which adjusts the noise of non-critical steps by calculating the change ratio of non-critical step energy relative to critical step energy, thereby effectively avoiding performance degradation and ensuring the accuracy of generated trajectories. At the same time, a method of constructing a correction parameter dictionary is proposed to avoid the time loss caused by repeated calculation of correction parameters during the inference process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a comparison result diagram of the diffusion strategy (DP), the consistency strategy (CP) and the fast strategy (FP) of the present invention in the simulation experiment.

[0018] Figure 2 It is the feature difference map between the denoising steps of each module.

[0019] Figure 3 This is the spectrum diagram of the second decoder module after Fourier transform under the PushT task. DETAILED DESCRIPTION

[0020] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0021] The purpose of the present invention is to solve the problem that the diffusion strategy in the field of robotics cannot realize real-time and accurate control of the robot due to its slow reasoning, and to provide a visual motion prediction acceleration strategy that can be directly based on a pre-trained model.

[0022] Visual motion prediction acceleration strategy division steps: Characteristic analysis of the diffusion strategy based on the UNet (U-shaped network) architecture: In the denoising process of a specific task, the encoder features change little and are highly similar, while the decoder features fluctuate greatly in different denoising steps. Reuse the key encoder features for non-critical steps within a certain time range to avoid repeatedly calculating these non-critical encoder features with small changes to reduce computational overhead; Key step selection: In view of the requirements of different robot control tasks, the noise prediction network parameters of the pre-trained diffusion strategy are different, but the frequency domain analysis method in the image field can be essentially used. Therefore, the present invention designs a key step selection strategy based on Fourier spectrum energy. The influence of the encoder (input decoder through jump connection) and the intermediate module in UNet on the final noise is quantified according to the spectrum energy as the weight, and the weighted combination with the mean square error of the corresponding module feature change obtained in the feature analysis is used to adaptively screen more critical denoising steps; Prediction noise correction: In order to compensate for the performance degradation and prediction trajectory deviation that may be caused by reusing the key step encoder features in non-critical steps, the present invention further designs a noise correction strategy, that is, to correct the noise of non-critical steps according to the change ratio of the non-critical step energy relative to the key step energy, so that the noise distribution predicted by the non-critical step after reusing the features of the previous key step is close to the noise distribution when the non-critical step itself does not reuse the key step features. At the same time, it is found that the energy changes of the same task are relatively consistent under different initial settings, so a correction parameter dictionary is directly constructed to avoid repeated calculation of correction parameters during inference.

[0023] Note: DP (diffusion strategy), CP (consistency strategy), DDPM (probabilistic denoising diffusion model), EDM (design space method based on diffusion-based generative model), DDIM (implicit denoising diffusion model).

[0024] For ease of understanding, the specific process of the embodiment of the present invention is described below. The first embodiment of the visual motion prediction acceleration method that can be directly based on the pre-training model in the embodiment of the present invention includes: Based on UNet, the input signal is denoised, and the UNet includes an encoder, an intermediate block, and a decoder; Calculate the feature differences of the encoder's output features between adjacent denoising steps; The step of judging whether the current denoising step is a non-critical step or a critical step according to the feature difference is specifically as follows: The first denoising step is the key step, and the subsequent denoising steps are judged as follows: The decoder output features are mapped to the frequency domain through Fourier transform, and the spectrum energy E is calculated. The formula is: Where F(·) represents the component amplitude, and (u,v) represents the position in the spectrum diagram; Building a scoring function : Among them, s1 and s2 are skip connection features, b is the output feature of the intermediate module; n represents the nth denoising step interval, M represents the mean square error of the output feature of the corresponding denoising step interval, T represents different tasks, and w is the weight; Set the number of critical steps; According to the set number of key steps, the steps corresponding to the step interval with the highest score are selected as key steps, and the others are non-key steps.

[0025] The encoder features and intermediate block features of the most recent critical step are used as the input of the decoder block of the subsequent non-critical step to obtain the noise prediction features; Visual motion prediction is completed based on noise prediction features.

[0026] Preferably, the step of calculating the feature difference of the output features of the encoder between adjacent denoising steps comprises: Calculate the mean square error of the output features of each module of the one-dimensional convolutional UNet between adjacent denoising steps; The calculated mean square error is normalized according to the batch size, dimension and field of view length of the feature to obtain the feature difference. The formula is as follows: Among them, b is the batch size, d is the dimension, h is the field of view length, and i represents the input of the current module. is the feature output by the kth module in the nth denoising step, is the feature output by the kth module in the n-1th denoising step, Represents the feature difference between the kth module in the nth denoising step and the n-1th denoising step.

[0027] The noise prediction feature f of the non-critical step is scaled by factor c d2 Perform dynamic adjustment and output the corrected frequency domain feature f′ d2 , the formula is: .

[0028] The scaling factor c is calculated based on the ratio of the spectral energy change between the output features of the non-critical step and the corresponding critical step, and the formula is: .

[0029] Among them, E nonkey is the spectral energy of the output feature of the non-critical step, E key is the spectral energy of the output feature of the key step.

[0030] The scaling factor c is obtained by the following steps: The correction parameter dictionaries are pre-built, denoted as dict{DDIM} and dict{EDM}, which contain the scaling factors for each task.

[0031] The above describes the visual motion prediction acceleration method that can be directly based on the pre-training model in the embodiment of the present invention. The following describes the visual motion prediction acceleration device that can be directly based on the pre-training model in the embodiment of the present invention: A denoising processing unit, used for performing denoising processing on an input signal based on UNet, wherein the UNet includes an encoder, an intermediate block, and a decoder; A feature difference calculation unit, used to calculate the feature difference of the output features of the encoder between adjacent denoising steps; The judging unit is used to judge whether the current denoising step is a non-critical step or a critical step according to the feature difference, specifically: The first denoising step is the key step, and the subsequent denoising steps are judged as follows: The decoder output features are mapped to the frequency domain through Fourier transform, and the spectrum energy E is calculated. The formula is: Where F(·) represents the component amplitude, and (u,v) represents the position in the spectrum diagram; Constructing the scoring function : Among them, s1 and s2 are skip connection features, b is the output feature of the intermediate module; n represents the nth denoising step interval, M represents the mean square error of the output feature of the corresponding denoising step interval, T represents different tasks, and w is the weight; Set the number of critical steps; According to the set number of key steps, the step corresponding to the step interval with the highest score is selected as the key step, and the others are non-key steps; The noise prediction feature unit is used to use the encoder features and intermediate block features of the most recent key step as inputs to the decoder block of the subsequent non-key step to obtain noise prediction features.

[0032] The present invention is a visual motion prediction acceleration strategy that can be directly based on a pre-trained model, comprising the following steps: 1. The typical denoising network uses a UNet-based architecture, which consists of an encoder E, an intermediate module M, and a decoder D. The noise prediction network based on the convolutional neural network still retains the core elements of UNet, and the features of different levels extracted by the encoder E are still input into the decoder D through jump connections. In the subsequent denoising steps, this experiment reuses the encoder or decoder features in the UNet structure of the diffusion strategy for noise prediction. By ignoring the calculation of the encoder and intermediate modules of the UNet in the non-critical step of the noise prediction process, the computational redundancy is reduced to achieve acceleration.

[0033] The present invention first calculates the mean square error of the output features of each module of the one-dimensional convolutional UNet between adjacent denoising steps, and then normalizes the calculated feature differences according to the batch size, dimension and field of view length of the features to obtain the final feature differences.

[0034] b is the batch size, d is the dimension (number of channels), h is the field of view length, and i represents the input of the current module. represents the feature difference between the kth module at the nth denoising step and the n-1th denoising step. The feature difference between each module in the denoising step under DDPM sampling is as follows: Figure 2 As shown in (a) in . The present invention finds that when using DDPM sampling steps, the feature differences between the encoder and decoder modules of the one-dimensional UNet in the diffusion strategy change little between the later denoising steps. However, for the DDPM sampling strategy, the number of denoising steps of 100 steps will result in a large span between key steps when a small number of key steps is set, resulting in excessive error accumulation, which in turn leads to setting too many key steps under DDPM sampling, and thus requires a lot of time to find the best key step setting among many key step combinations. Figure 2 As shown in (b), the feature difference of the encoder module is smoother and smaller than that of the decoder module. Although DDIM sampling changes the noise characteristics in the denoising process, the feature difference between the denoising steps is small, and since the number of denoising steps is small, there will not be too much error accumulation. Similarly, the analysis of EDM sampling also shows similar results, that is, with a small number of sampling steps, the encoder feature changes very little between denoising steps. Therefore, the method of reusing encoder features in the present invention can also be applied to DDIM sampling and EDM sampling.

[0035] 2. According to the previous feature analysis results, it is an ideal method to reuse the features of the key step in the non-critical step. The core idea of ​​the FP strategy (FastPolicy) of the present invention is to set the key denoising step T and reuse the encoder features in these steps. Since the encoder features transmitted through the jump connection and the output of the intermediate block change very little during the denoising process, the present invention directly uses the encoder features and intermediate block features of the current key step as the input of the corresponding decoder block in the subsequent non-critical step. In this way, the computational redundancy of the encoder and the intermediate block in the non-critical step is eliminated, thereby accelerating the reasoning process.

[0036] 3. Previous studies usually only focus on a single task (such as image generation) or design neural networks to find better key steps, but these methods often have poor generalization performance or require extra time to train the model. The present invention proposes a dynamic key step selection mechanism based on Fourier spectrum energy for adaptively selecting key steps. Each task of the diffusion strategy needs to train the corresponding model, so the key step designed for a certain task may not be well generalized to other categories of tasks, such as the key step of the PushT task is not necessarily the best key step for the subtask Square under Robomimic. The goal of the present invention is to dynamically select key steps based on the inherent characteristics of the noise prediction network for each task. For the one-dimensional UNet in the diffusion strategy, the final predicted noise is determined by three variables: the first jump connection, the second jump connection, and the intermediate module output, which are recorded as s1, s2, and b respectively. Denoising steps with large feature differences should be prioritized as key steps to reduce error accumulation. However, the final noise is affected by multiple variables, and their effects on noise prediction are different. UNet was originally designed for image processing, and noise feature analysis usually uses Fourier transform to study amplitude in the frequency domain. Although the diffusion strategy predicts action sequences, UNet still acts as a noise prediction network. Therefore, the analysis methods used for UNet in the image field are also applicable here. The present invention combines Fourier transform to convert the denoising process to the frequency domain for analysis. The final 1D convolution is only used for scaling, so the present invention performs frequency domain analysis on the output of the second decoder block (denoted as d2). The converted spectrum diagram is shown in the figure below. Figure 3 The energy calculation formula is: , where F(·) represents the component amplitude and (u,v) represents the position in the spectrum diagram, similar to the pixel coordinates. The present invention can quantify the contribution of s1, s2 and b to the energy of d2 respectively to evaluate their impact. To this end, the present invention constructs the following scoring function: Among them, n represents the nth denoising step interval, M corresponds to the mean square error of the step interval, and T represents different tasks (such as PushT). w is the weight, which is calculated based on the proportion of s1, s2 and b's respective contributions to energy. Then, according to the set number of key steps (set to 3 in this experiment), the step corresponding to the step interval with the highest score is selected as the key step. It should be noted that the first denoising step must be used as a key step.

[0037] For 1D UNet, since the final noise prediction depends on the output of the skip connection s1, s2 and the intermediate module, the scoring function of each denoising step should be composed of these three elements; the influence of these elements on the final result is visualized through the spectrum energy; the spectral energy contribution ratio is normalized as the weight. After obtaining the weight, the score is calculated by combining the feature differences corresponding to each module between the denoising steps obtained in the previous analysis, and the highest score is selected as the key step.

[0038] 4. Reusing features may cause deviations in prediction noise, thus affecting the accuracy of the final trajectory prediction. To solve this problem, the present invention calculates a scaling factor c based on the ratio of spectral energy changes between the d2 output features of the non-critical step and the corresponding critical step, and the formula is: It should be noted that in the denoising process, the critical steps always precede the non-critical steps. Then, the noise correction formula is , where f d2 and f′ d2 Respectively represent the frequency domain noise characteristics before and after correction. To avoid repeated calculation of scaling factors during the inference process, the present invention pre-constructs a correction parameter dictionary, denoted as dict{DDIM} and dict{EDM}, which contains the scaling factors of each task. In particular, for a second-order sampler with two noise predictions, since the second noise prediction is based on the result of the first noise prediction, the present invention only compensates for the first noise prediction to minimize the additional inference time.

[0039] It should be noted that the method of the present invention aims to reduce the inference time while maintaining a certain task completion rate, and there is no need to retrain the student model. However, the sampling of the important baseline methods of the present invention (diffusion strategy and consistency strategy) is completely different, namely DDIM and EDM. Therefore, in order to fairly compare the performance, the present invention adopts the same sampling strategy as the baseline method in the strategy of the present invention, which proves the generalization ability of the present invention to a certain extent. In addition to the UNet architecture, the present invention maintains the same module configuration (such as image encoder and normalization method) as the diffusion strategy and consistency strategy in all experiments and methods. In addition, the present invention also takes the observation data within two frames (including wrist camera images, third-person camera images, and end effector postures) as input, and outputs a continuous string of end effector postures.

[0040] Experimental results: All other metrics are calculated on an NVIDIA GPU L40 graphics card using the same settings.

[0041] Figure 1 The results of the diffusion strategy (DP), consistency strategy (CP) and the fast strategy (FP) of the present invention are demonstrated in simulation experiments, showing that the method of the present invention significantly accelerates the reasoning speed while maintaining the success rate.

[0042] See Figure 2 , showing the feature differences between denoising steps of each module for PushT task under sampling of denoising diffusion probability model (DDPM), implicit denoising diffusion model (DDIM) and design space method (EDM) based on diffusion generative model (discrete step number bins=80 and bins=8 respectively). The zoomed-out figure highlights the area with less feature change.

[0043] Table 1 shows a comprehensive comparison of the performance of the proposed strategy and the diffusion model based on DDIM. It can be seen that the proposed method can maintain comparable accuracy with the baseline method (diffusion strategy) in both single-arm and dual-arm tasks, and even slightly improves accuracy in some cases. More importantly, the reasoning time of the proposed method is about half of that of the baseline method.

[0044] Notes to Table 1: Shows the inference time (ms) and success rate (%) of the latest and best checkpoints during training. All models were trained for 1000 epochs. w / o NC: without noise correction; w / NC: with noise correction.

[0045] Table 1 Table 2 shows the performance comparison between the strategy of the present invention and the diffusion strategy and consistency strategy under EDM. It is worth noting that, except for the student model of the consistency strategy which was trained for 500 rounds, all pre-trained diffusion models were only trained for 400 rounds, in order to make a fair comparison under the configuration of the consistency strategy. All experimental results are based on data collected by professionals. The present invention observes that the accuracy performance of the present invention is even better than that of the diffusion strategy, which may be attributed to the better noise distribution in non-critical steps. In fact, the present invention is also more advantageous than the diffusion strategy in terms of reasoning speed. Compared with the consistency strategy, the reasoning speed of the method of the present invention is relatively slow, because the student model of the consistency strategy achieves further reasoning by distilling the knowledge of the teacher model to itself. However, in more complex tasks, the consistency strategy performs worse than the method of the present invention, which may be related to its training instability. More importantly, compared with the consistency strategy, the advantage of the present invention is that there is no need to retrain the student model, which further improves the deployment capability of the method of the present invention on resource-constrained real robot devices.

[0046] Table 2 shows the comparison of the simulation experimental performance of the fast strategy, consistency strategy and diffusion strategy based on EDM sampling. The present invention shows the inference time (milliseconds) and the average success rate (%) of the best checkpoint during training. All models were trained for 400 cycles except the student model of the consistency strategy (CP) which was trained for 500 cycles. All tasks were evaluated on PH data. In the setting of the consistency strategy, when a second-order solver is used and the key step is set to 3, the number of inference function calls of the fast strategy of the present invention is 7.

[0047] Table 2 An embodiment of the present invention also provides an electronic device, which may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) (for example, one or more processors) and memories, and one or more storage media for storing applications or data (for example, one or more mass storage devices). Among them, the memory and the storage medium may be short-term storage or permanent storage. The program stored in the storage medium may include one or more modules, each module may include a series of instruction operations in the electronic device. Furthermore, the processor may be configured to communicate with the storage medium and execute a series of instruction operations in the storage medium on the electronic device.

[0048] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input and output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that the electronic device structure in this embodiment does not constitute a limitation on the electronic device, and may include more or fewer components, or combine certain components, or arrange components differently.

[0049] An embodiment of the present invention provides a structure of an electronic device, which may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) (for example, one or more processors) and memories, and one or more storage media for storing applications or data (for example, one or more mass storage devices). Among them, the memory and the storage medium can be short-term storage or permanent storage. The program stored in the storage medium may include one or more modules, each module may include a series of instruction operations in the electronic device. Furthermore, the processor can be configured to communicate with the storage medium and execute a series of instruction operations in the storage medium on the electronic device.

[0050] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input and output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that the electronic device structure does not constitute a limitation on the electronic device, and may include more or less components than the aforementioned, or combine certain components, or arrange the components differently.

[0051] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are executed on a computer, the computer executes the steps of the aforementioned method.

[0052] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0053] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0054] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual motion prediction acceleration method that can be directly based on a pre-trained model, characterized in that: The following steps are involved: Based on UNet, the input signal is denoised, and the UNet includes an encoder, an intermediate block, and a decoder; Calculate the feature differences of the encoder's output features between adjacent denoising steps; According to the feature differences, it is determined whether the current denoising step is a non-critical step or a critical step, specifically: The first denoising step is the key step, and the subsequent denoising steps are judged as follows: The decoder output features are mapped to the frequency domain through Fourier transform, and the spectrum energy E is calculated. The formula is: Where F(·) represents the component amplitude, and (u,v) represents the position in the spectrum diagram; Building a scoring function : Among them, s1 and s2 are skip connection features, b is the output feature of the intermediate module; n represents the nth denoising step interval, M represents the mean square error of the output feature of the corresponding denoising step interval, T represents different tasks, and w is the weight; Set the number of critical steps; According to the set number of key steps, the step corresponding to the step interval with the highest score is selected as the key step, and the others are non-key steps; The encoder features of the most recent critical step and the intermediate block features are used as the input of the decoder block of the subsequent non-critical step to obtain the noise prediction features.

2. The visual motion prediction acceleration method directly based on a pre-trained model according to claim 1, characterized in that: The calculating feature difference of output features of the encoder between adjacent denoising steps includes: Calculate the mean square error of the output features of each module of the one-dimensional convolutional UNet between adjacent denoising steps; The calculated mean square error is normalized according to the batch size, dimension and field of view length of the feature to obtain the feature difference. The formula is as follows: Among them, b is the batch size, d is the dimension, h is the field of view length, and i represents the input of the current module. is the feature output by the kth module in the nth denoising step, is the feature output by the kth module in the n-1th denoising step, Represents the feature difference between the kth module in the nth denoising step and the n-1th denoising step.

3. The visual motion prediction acceleration method directly based on a pre-trained model according to claim 1, characterized in that: The noise prediction feature f of the non-critical step is scaled by factor c d2 Perform dynamic adjustment and output the corrected frequency domain feature f′ d2 , the formula is: 。 4. The visual motion prediction acceleration method directly based on a pre-trained model according to claim 3, characterized in that: The scaling factor c is calculated based on the ratio of the spectrum energy change between the output features of the non-critical step and the corresponding critical step, and the formula is: Among them, E nonkey is the spectral energy of the output feature of the non-critical step, E key is the spectral energy of the output feature of the key step.

5. The visual motion prediction acceleration method directly based on a pre-trained model according to claim 4, characterized in that: The scaling factor c is obtained by the following steps: The correction parameter dictionaries are pre-built, denoted as dict{DDIM} and dict{EDM}, which contain the scaling factors for each task.

6. A visual motion prediction acceleration system that can be directly based on a pre-trained model, characterized in that: The system comprises: A denoising processing unit, used for performing denoising processing on an input signal based on UNet, wherein the UNet includes an encoder, an intermediate block, and a decoder; A feature difference calculation unit, used to calculate the feature difference of the output features of the encoder between adjacent denoising steps; The judging unit is used to judge whether the current denoising step is a non-critical step or a critical step according to the feature difference, specifically: The first denoising step is the key step, and the subsequent denoising steps are judged as follows: The decoder output features are mapped to the frequency domain through Fourier transform, and the spectrum energy E is calculated. The formula is: Where F(·) represents the component amplitude, and (u,v) represents the position in the spectrum diagram; Building a scoring function : Among them, s1 and s2 are skip connection features, b is the output feature of the intermediate module; n represents the nth denoising step interval, M represents the mean square error of the output feature of the corresponding denoising step interval, T represents different tasks, and w is the weight; Set the number of critical steps; According to the set number of key steps, the step corresponding to the step interval with the highest score is selected as the key step, and the others are non-key steps; The noise prediction feature unit is used to use the encoder features and intermediate block features of the most recent key step as inputs to the decoder block of the subsequent non-key step to obtain noise prediction features.

7. An electronic device, comprising a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the electronic device executes the various steps of the visual motion prediction acceleration method that can be directly based on a pre-trained model as described in any one of claims 1-5.

8. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the various steps of the visual motion prediction acceleration method directly based on the pre-trained model as described in any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Industrial visual inspection method based on denoising diffusion implicit model and FNet auto-encoder

    CN117496251A

  • Texture synthesis method based on diffusion model and reweighting strategy

    CN118196227A

  • Multi-modal personalized diffusion model video generation and acceleration device and method based on importance evaluation

    CN118612509A

  • Image generation method and system based on diffusion model adaptive reasoning

    CN119671894A

Cited By

  • Vehicle control method and device, vehicle, storage medium, program product and chip

    CN121375835A

  • Vehicle control method and device, vehicle, storage medium, program product and chip

    CN121375835B