Action generation method and device based on dynamic feature cache, equipment and medium
By constructing a teacher-student mechanism for behavioral target proxy coefficients and action-perception collaboration, the problem of inconsistency between caching strategies and action targets in robot action generation models is solved, achieving efficient action generation and response, and improving the performance and robustness of embodied intelligence systems.
Patent Information
- Application Number
- CN202511189566.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-11
AI Technical Summary
Existing robot motion generation models suffer from problems such as a disconnect between training and inference, and inconsistencies between caching strategies and motion generation goals, leading to error accumulation and insufficient response performance.
We construct behavioral target proxy coefficients, build a target loss function based on relative entropy and learnable routing parameters, train the action generation model using a teacher-student mechanism that coordinates action and perception, generate real-time environmental perception through multi-step denoising and feature fusion, and dynamically reuse historical cached features to generate target actions.
This approach achieves improved response efficiency while ensuring decision stability in the action generation model, balances caching efficiency and task rewards, and enhances the overall performance and robustness of the embodied intelligence system.
Smart Images

Figure CN120932307A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an action generation method, apparatus, device, and medium based on dynamic feature caching. Background Technology
[0002] Currently, with the rapid development of artificial intelligence technology, embodied intelligent vision-language-action (VLA) models are widely used in various fields.
[0003] For example, in the financial sector, service robots that interact with users are typically set up in financial service halls to answer user questions or guide users through business transactions; in the medical field, auxiliary robots are typically set up in operating rooms to handle surgical instruments according to the doctor's instructions, thereby improving surgical efficiency.
[0004] However, existing action generation models used by robots still suffer from problems such as a disconnect between training and inference, and inconsistencies between caching strategies and action generation goals. Specifically, traditional learning-based caching ignores cross-temporal feature dependencies during multi-step denoising, leading to error accumulation. Furthermore, its training objective focuses on noise prediction errors rather than the quality of the final action decision, which is significantly mismatched with the environment-interaction-oriented optimization requirements of VLA models, making it difficult to balance caching efficiency with the real-time response performance of embodied intelligence systems. Summary of the Invention
[0005] In view of the above, it is necessary to provide an action generation method, apparatus, device and medium based on dynamic feature caching, in order to solve the problem that the action generation model cannot accurately generate actions.
[0006] An action generation method based on dynamic feature caching, the action generation method based on dynamic feature caching includes:
[0007] Construct behavioral target proxy coefficients, and construct a target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficients;
[0008] Based on the action-perception collaborative teacher-student mechanism, the target loss function is used to train an action generation model including a teacher model and a student model;
[0009] In response to an action generation instruction based on target data, the action generation instruction is parsed to obtain a sequence of selectable actions in the language context and the current environment.
[0010] The student model is used to perform multi-step denoising on the target data according to the target time step to obtain a denoised feature sequence.
[0011] The student model is used to generate real-time environmental awareness based on the denoised feature sequence and the language context.
[0012] The student model is used to fuse the real-time environment perception with the corresponding dynamic historical cache based on the learnable routing parameters to obtain a hybrid representation sequence;
[0013] The hybrid representation sequence is input into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence;
[0014] The target action is selected from the sequence of available actions based on the target action probability distribution.
[0015] An action generation device based on dynamic feature caching, the action generation device based on dynamic feature caching includes:
[0016] A construction unit is used to construct behavioral target proxy coefficients and construct a target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficients;
[0017] The training unit is used to train an action generation model, including a teacher model and a student model, based on a teacher-student mechanism that is based on action-perception collaboration, using the target loss function.
[0018] The parsing unit is used to respond to the action generation instruction based on the target data, and parse the action generation instruction to obtain the language context and the optional action sequence under the current environment state;
[0019] The denoising unit is used to perform multi-step denoising on the target data according to the target time step using the student model to obtain a denoised feature sequence.
[0020] A generation unit is used to generate real-time environmental awareness based on the denoised feature sequence and the language context using the student model.
[0021] The fusion unit is used to fuse the real-time environment perception and the corresponding dynamic historical cache based on the learnable routing parameters using the student model to obtain a hybrid representation sequence;
[0022] An input unit is used to input the hybrid representation sequence into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence;
[0023] The selection unit is used to select a target action from the sequence of available actions based on the target action probability distribution.
[0024] A computer device, the computer device comprising:
[0025] Memory, storing at least one instruction; and
[0026] The processor executes the instructions stored in the memory to implement the action generation method based on dynamic feature caching.
[0027] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the action generation method based on dynamic feature caching.
[0028] As can be seen from the above technical solutions, this invention can construct behavioral target proxy coefficients, breaking the limitation of traditional caching that only optimizes noise prediction errors, and binding caching strategies with actual task rewards; constructing a target loss function based on relative entropy, learnable routing parameters, and target proxy coefficients can effectively constrain decision consistency, protect task rewards, and balance caching efficiency; training the action generation model based on the action-perception collaborative teacher-student mechanism and target loss function enables the student model to effectively learn from the teacher model; using the student model to fuse real-time environmental perception with corresponding dynamic historical caching based on learnable routing parameters, and generating a target action probability distribution of optional action sequences based on the hybrid representation sequence to select the best target action, it can dynamically reuse historical features to reduce redundant calculations when generating actions, improving response efficiency while ensuring decision stability. Attached Figure Description
[0029] Figure 1 This is a flowchart of a preferred embodiment of the action generation method based on dynamic feature caching of the present invention.
[0030] Figure 2 This is a functional block diagram of a preferred embodiment of the action generation device based on dynamic feature caching of the present invention.
[0031] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the action generation method based on dynamic feature caching according to the present invention. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the action generation method based on dynamic feature caching according to the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0034] The action generation method based on dynamic feature caching is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0035] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0036] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0037] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0038] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0039] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0040] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0041] S10, construct the behavioral target proxy coefficients, and construct the target loss function based on the relative entropy (Kullback-Leibler divergence, also known as KL divergence), learnable routing parameters, and the target proxy coefficients.
[0042] In this embodiment, the target proxy coefficient is a quantitative indicator of the performance gap after the model performs actions in the simulated environment.
[0043] Specifically, the constructed behavioral target proxy coefficients include:
[0044] Calculate the decrease in task reward caused by cache errors under the constraints of environmental conditions and actions;
[0045] The expected value of the decrease in task reward is calculated based on the real-time action probability distribution at each time step, thus obtaining the target proxy coefficient.
[0046] The larger the decrease in task reward, the more severe the impact of caching error on the task.
[0047] The target proxy coefficient is used to characterize the decrease in average reward caused by caching errors, and is a quantification of the degree of damage to task performance caused by the current caching strategy.
[0048] The real-time action probability distribution at each time step can be approximated by the action probability distribution output by the student model or the teacher model.
[0049] By using the target proxy coefficients, the student model not only mimics the teacher model's action preferences but also focuses on the impact of these preferences on actual task rewards (such as navigation success rate and pickup accuracy). By using the expected value of the reward decrease caused by caching at each step as a weight, caching errors that would severely damage task performance can be automatically amplified during training, while biases with small impact on the final reward can be suppressed. In this way, the model can learn to efficiently utilize historical features while more accurately maintaining and improving overall task performance.
[0050] The above embodiments can break through the limitation of traditional caching that only optimizes noise prediction errors, bind the caching strategy to the actual task rewards, and enable the model to automatically identify high-impact caching errors (even if the target proxy coefficient is large) and prioritize correcting deviations that greatly damage task performance.
[0051] In this embodiment, constructing the target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficient includes:
[0052] Calculate the relative entropy loss between the action probability distribution output by the student model and the action probability distribution output by the teacher model;
[0053] The action distribution alignment loss is obtained by multiplying the target proxy coefficient by the relative entropy loss.
[0054] The learnable routing parameters are regularized to obtain the learnable routing parameters and;
[0055] Obtain the configured regularization hyperparameters and calculate the product of the regularization hyperparameters and the sum of the learnable routing parameters to obtain the cache loss;
[0056] The target loss function is obtained by summing the action distribution alignment loss and the cache loss.
[0057] By minimizing the relative entropy loss, the student model can replicate the teacher model's action preferences as closely as possible at each step, forcing the student model's action decisions to align with the teacher model. This ensures that when selecting actions, the fusion strategy based on cached and perceptual features (i.e., the learnable routing parameters) can produce decision outputs as reliable as the teacher model. After training in this way, the student model can both efficiently utilize feature caches and maintain action quality comparable to the teacher model.
[0058] The relative entropy loss avoids the problem of caching strategy optimization becoming disconnected from action goals, making the action distribution of the student model approximate that of the teacher model, thus ensuring consistency between the caching strategy and action goals. The relative entropy loss also accurately quantifies distribution differences, ensuring the directionality of routing parameter updates (i.e., adjusting routing parameters improves the quality of action decisions).
[0059] The learnable routing parameters are the learnable weights of the model at step t for the i-th block of historical features (i.e., cached fused features) and the current real-time environment perception trust level.
[0060] Regularizing the learnable routing parameters can prevent them from being too large (i.e., over-reliance on current perception, wasting cache) or too small (i.e., over-reliance on cache, leading to error accumulation).
[0061] The objective loss function integrates multiple losses, enabling the student model to learn and mimic the teacher model's optimal action distribution at each step, maintaining consistency in action decisions (i.e., by minimizing the knowledge distillation loss of KL divergence). It also guides the student model to focus on decision biases that actually affect environmental interaction rewards due to caching errors (i.e., through the behavioral objective proxy coefficient). Furthermore, it regularizes the learnable routing parameters to prevent over-caching. This combination ensures both efficient cache utilization and optimal task performance.
[0062] In the above embodiments, the action distribution alignment loss enables the student model to mimic the teacher's action distribution, ensuring decision consistency. Combined with the target proxy coefficient, it can prioritize the correction of high-impact cache errors, protecting task rewards. Furthermore, the cache loss avoids over-reliance on a single feature, thereby balancing cache efficiency.
[0063] S11, Based on the teacher-student mechanism of action-perception collaboration, the action generation model including the teacher model and the student model is trained using the target loss function.
[0064] In this embodiment, the action-perception collaborative teacher-student mechanism, which uses the target loss function to train an action generation model including a teacher model and a student model, includes:
[0065] The target loss function is minimized through backpropagation so that the output of the student model approaches the output of the teacher model.
[0066] Training stops when the value of the target loss function no longer decreases.
[0067] The currently obtained model is determined as the action generation model.
[0068] Through backpropagation, the learnable routing parameters of the student model (to balance cache utilization and action decision quality) and other parameters of the feature fusion module can be continuously updated.
[0069] Through the above embodiments, the model can be ensured to find the optimal balance between efficient cache reuse and improved decision quality based on the objective loss function, without additional GPU memory burden, thereby efficiently training the action generation model.
[0070] For example, in the financial field, the motion generation model can be used to generate the response voice of service robots; in the medical field, the motion generation model can be used to generate the response actions of assistive robots; and in the logistics field, the motion generation model can be used to generate the handling actions of handling robots.
[0071] This embodiment employs a consistent multi-step cache update strategy during both the training and subsequent inference phases to achieve a unified approach to feature reuse and dynamic routing.
[0072] S12, in response to the action generation instruction based on the target data, the action generation instruction is parsed to obtain the language context and the selectable action sequence under the current environment state.
[0073] In this embodiment, the action generation instruction can be automatically triggered when the target data upload is detected.
[0074] In this embodiment, the target data can be data from various modalities such as the user's voice or actions.
[0075] In this embodiment, the language context can be the context that restricts the action generation instruction, which is used to guide the model to prioritize what when generating actions (such as "long-term value investment" in a financial scenario) or to restrict the model (such as "avoiding damage to blood vessels" in a medical scenario).
[0076] In this embodiment, the current environmental state is a holistic description of the environmental state, including current visual observation, language command context, and other perceptual information.
[0077] In this embodiment, the optional action sequence may include "forward", "grab", or "say a word", etc., for the model to select.
[0078] S13, using the student model to perform multi-step denoising on the target data according to the target time step to obtain a denoised feature sequence.
[0079] In this embodiment, the student model can perform multi-step denoising on the target data based on the Diffusion Transformer architecture.
[0080] In this embodiment, the denoised feature sequence is the deep visual-language fusion feature obtained by the student model after the t-th denoising iteration (t is the time step index of the denoising iteration, where t is a positive integer), used to represent the current environmental observation. The denoised feature sequence carries the internal semantic representation of the target data such as sensor input (e.g., images, language) after multiple steps of gradual refinement.
[0081] Through the above embodiments, features can be refined through multi-step denoising to reduce the interference of noise in the original data on action decisions and improve the accuracy of feature representation.
[0082] S14, using the student model to generate real-time environmental perception based on the denoised feature sequence and the language context.
[0083] In this embodiment, the student model can generate a real-time environment-aware representation based on real-time denoising features and the language context.
[0084] S15, the student model is used to fuse the real-time environment perception and the corresponding dynamic historical cache based on the learnable routing parameters to obtain a hybrid representation sequence.
[0085] In this embodiment, the step of fusing the real-time environment perception and the corresponding dynamic historical cache based on the learnable routing parameters using the student model to obtain a hybrid representation sequence includes:
[0086] For each real-time environment perception, the product of the real-time environment perception and the learnable routing parameters is calculated to obtain the first feature;
[0087] Calculate the difference between 1 and the learnable routing parameters to obtain the complementary parameters of the learnable routing parameters;
[0088] The second feature is obtained by multiplying the complementary parameter with the corresponding dynamic history cache.
[0089] The sum of the first feature and the second feature is calculated to obtain the hybrid representation of the real-time environment perception;
[0090] The hybrid representation sequence is obtained by integrating all real-time environment-aware hybrid representations.
[0091] In this process, each fusion step utilizes dynamic historical cached features. Specifically, the fusion step t requires features from the previous t-1 steps, addressing the problem of traditional caching ignoring cross-temporal feature dependencies. This allows for dynamic reuse of historical features to reduce redundant computation and improve computational efficiency. Furthermore, the learnable routing parameters allow for flexible adjustment of feature trust levels, preventing error accumulation due to over-reliance on a single feature (e.g., trusting the cache when the environment is stable, and trusting the current perception when the environment changes drastically).
[0092] In the above embodiments, historical features can be utilized more smoothly in multi-round environmental interactions, achieving consistency within the decision-making process and effectively mitigating decision bias caused by error accumulation. The utilization efficiency of feature caching is improved without increasing additional GPU memory burden, enabling the model to respond more quickly to environmental changes in complex scenarios and maintain high decision-making stability.
[0093] In this embodiment, after obtaining the hybrid representation of the real-time environment perception, the method further includes:
[0094] The hybrid representation of the real-time environment awareness is stored in a cache.
[0095] For example, the hybrid representation of the real-time environment awareness can be stored in the cache according to the feature block index for quick subsequent retrieval.
[0096] In the above embodiments, updating cached features in real time can improve feature reuse efficiency while ensuring action response speed (no redundant calculation) and decision stability during the inference phase.
[0097] S16, the mixed representation sequence is input into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence.
[0098] In this embodiment, the target action probability distribution is a confidence assignment of each possible action (such as moving, grabbing, executing instructions, etc.), which is used to reflect the relative priority or selection tendency of each action in the current environmental state.
[0099] In this embodiment, after optimizing and training the parameters of the student model based on the teacher model, the student model is used to perform inference to generate the target action probability distribution of the optional action sequence. This can achieve a unified framework for training and inference, avoid the problem that the caching strategy optimized during training fails during inference in traditional methods, significantly enhance the generalization ability of the model across different tasks, and significantly improve the overall performance and robustness of the embodied intelligence system.
[0100] S17, Select a target action from the available action sequence according to the target action probability distribution.
[0101] In this embodiment, the target action probability distribution is used to guide the selection of actions.
[0102] Specifically, selecting the target action from the sequence of available actions based on the target action probability distribution includes:
[0103] Based on the target action probability distribution, the action with the highest confidence level is selected from the sequence of available actions as the target action.
[0104] Through the above embodiments, the optimal action rate can be accurately selected from the selectable action sequences based on the target action probability distribution output by the model.
[0105] In this embodiment, if the task reward drops below a threshold after the corresponding action execution entity performs the target action, an emergency cache update can be triggered to ensure the accuracy of action generation.
[0106] For example, when a financial service robot outputs inappropriate voice commands, or when an auxiliary robot in a medical operating room performs actions with deviations, an emergency cache update can be triggered to ensure the accuracy of the generated actions.
[0107] As can be seen from the above technical solutions, this invention can construct behavioral target proxy coefficients, breaking the limitation of traditional caching that only optimizes noise prediction errors, and binding caching strategies with actual task rewards; constructing a target loss function based on relative entropy, learnable routing parameters, and target proxy coefficients can effectively constrain decision consistency, protect task rewards, and balance caching efficiency; training the action generation model based on the action-perception collaborative teacher-student mechanism and target loss function enables the student model to effectively learn from the teacher model; using the student model to fuse real-time environmental perception with corresponding dynamic historical caching based on learnable routing parameters, and generating a target action probability distribution of optional action sequences based on the hybrid representation sequence to select the best target action, it can dynamically reuse historical features to reduce redundant calculations when generating actions, improving response efficiency while ensuring decision stability.
[0108] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the action generation device based on dynamic feature caching of the present invention. The action generation device 11 based on dynamic feature caching includes a construction unit 110, a training unit 111, a parsing unit 112, a denoising unit 113, a generation unit 114, a fusion unit 115, an input unit 116, and a selection unit 117. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0109] The construction unit 110 is used to construct behavioral target proxy coefficients and construct a target loss function based on relative entropy (Kullback-Leibler divergence, also known as KL divergence), learnable routing parameters, and the target proxy coefficients.
[0110] In this embodiment, the target proxy coefficient is a quantitative indicator of the performance gap after the model performs actions in the simulated environment.
[0111] Specifically, the target proxy coefficient for the construction behavior of the construction unit 110 includes:
[0112] Calculate the decrease in task reward caused by cache errors under the constraints of environmental conditions and actions;
[0113] The expected value of the decrease in task reward is calculated based on the real-time action probability distribution at each time step, thus obtaining the target proxy coefficient.
[0114] The larger the decrease in task reward, the more severe the impact of caching error on the task.
[0115] The target proxy coefficient is used to characterize the decrease in average reward caused by caching errors, and is a quantification of the degree of damage to task performance caused by the current caching strategy.
[0116] The real-time action probability distribution at each time step can be approximated by the action probability distribution output by the student model or the teacher model.
[0117] By using the target proxy coefficients, the student model not only mimics the teacher model's action preferences but also focuses on the impact of these preferences on actual task rewards (such as navigation success rate and pickup accuracy). By using the expected value of the reward decrease caused by caching at each step as a weight, caching errors that would severely damage task performance can be automatically amplified during training, while biases with small impact on the final reward can be suppressed. In this way, the model can learn to efficiently utilize historical features while more accurately maintaining and improving overall task performance.
[0118] The above embodiments can break through the limitation of traditional caching that only optimizes noise prediction errors, bind the caching strategy to the actual task rewards, and enable the model to automatically identify high-impact caching errors (even if the target proxy coefficient is large) and prioritize correcting deviations that greatly damage task performance.
[0119] In this embodiment, the construction unit 110 constructs the target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficient, including:
[0120] Calculate the relative entropy loss between the action probability distribution output by the student model and the action probability distribution output by the teacher model;
[0121] The action distribution alignment loss is obtained by multiplying the target proxy coefficient by the relative entropy loss.
[0122] The learnable routing parameters are regularized to obtain the learnable routing parameters and;
[0123] Obtain the configured regularization hyperparameters and calculate the product of the regularization hyperparameters and the sum of the learnable routing parameters to obtain the cache loss;
[0124] The target loss function is obtained by summing the action distribution alignment loss and the cache loss.
[0125] By minimizing the relative entropy loss, the student model can replicate the teacher model's action preferences as closely as possible at each step, forcing the student model's action decisions to align with the teacher model. This ensures that when selecting actions, the fusion strategy based on cached and perceptual features (i.e., the learnable routing parameters) can produce decision outputs as reliable as the teacher model. After training in this way, the student model can both efficiently utilize feature caches and maintain action quality comparable to the teacher model.
[0126] The relative entropy loss avoids the problem of caching strategy optimization becoming disconnected from action goals, making the action distribution of the student model approximate that of the teacher model, thus ensuring consistency between the caching strategy and action goals. The relative entropy loss also accurately quantifies distribution differences, ensuring the directionality of routing parameter updates (i.e., adjusting routing parameters improves the quality of action decisions).
[0127] The learnable routing parameters are the learnable weights of the model at step t for the i-th historical feature (i.e., the cached fused feature) and the current real-time environment perception trust level.
[0128] Regularizing the learnable routing parameters can prevent them from being too large (i.e., over-reliance on current perception, wasting cache) or too small (i.e., over-reliance on cache, leading to error accumulation).
[0129] The objective loss function integrates multiple losses, enabling the student model to learn and mimic the teacher model's optimal action distribution at each step, maintaining consistency in action decisions (i.e., by minimizing the knowledge distillation loss of KL divergence). It also guides the student model to focus on decision biases that actually affect environmental interaction rewards due to caching errors (i.e., through the behavioral objective proxy coefficient). Furthermore, it regularizes the learnable routing parameters to prevent over-caching. This combination ensures both efficient cache utilization and optimal task performance.
[0130] In the above embodiments, the action distribution alignment loss enables the student model to mimic the teacher's action distribution, ensuring decision consistency. Combined with the target proxy coefficient, it can prioritize the correction of high-impact cache errors, protecting task rewards. Furthermore, the cache loss avoids over-reliance on a single feature, thereby balancing cache efficiency.
[0131] The training unit 111 is used to train an action generation model including a teacher model and a student model using the target loss function based on a teacher-student mechanism based on action-perception collaboration.
[0132] In this embodiment, the training unit 111, based on an action-perception collaborative teacher-student mechanism, trains an action generation model including a teacher model and a student model using the target loss function, including:
[0133] The target loss function is minimized through backpropagation so that the output of the student model approaches the output of the teacher model.
[0134] Training stops when the value of the target loss function no longer decreases.
[0135] The currently obtained model is determined as the action generation model.
[0136] Through backpropagation, the learnable routing parameters of the student model (to balance cache utilization and action decision quality) and other parameters of the feature fusion module can be continuously updated.
[0137] Through the above embodiments, the model can be ensured to find the optimal balance between efficient cache reuse and improved decision quality based on the objective loss function, without additional GPU memory burden, thereby efficiently training the action generation model.
[0138] For example, in the financial field, the motion generation model can be used to generate the response voice of service robots; in the medical field, the motion generation model can be used to generate the response actions of assistive robots; and in the logistics field, the motion generation model can be used to generate the handling actions of handling robots.
[0139] This embodiment employs a consistent multi-step cache update strategy during both the training and subsequent inference phases to achieve a unified approach to feature reuse and dynamic routing.
[0140] The parsing unit 112 is used to respond to the action generation instruction based on the target data, and parse the action generation instruction to obtain the language context and the optional action sequence under the current environment state.
[0141] In this embodiment, the action generation instruction can be automatically triggered when the target data upload is detected.
[0142] In this embodiment, the target data can be data from various modalities such as the user's voice or actions.
[0143] In this embodiment, the language context can be the context that restricts the action generation instruction, which is used to guide the model to prioritize what when generating actions (such as "long-term value investment" in a financial scenario) or to restrict the model (such as "avoiding damage to blood vessels" in a medical scenario).
[0144] In this embodiment, the current environmental state is a holistic description of the environmental state, including current visual observation, language command context, and other perceptual information.
[0145] In this embodiment, the optional action sequence may include "forward", "grab", or "say a word", etc., for the model to select.
[0146] The denoising unit 113 is used to perform multi-step denoising on the target data according to the target time step using the student model to obtain a denoised feature sequence.
[0147] In this embodiment, the student model can perform multi-step denoising on the target data based on the Diffusion Transformer architecture.
[0148] In this embodiment, the denoised feature sequence is the deep visual-language fusion feature obtained by the student model after the t-th denoising iteration (t is the time step index of the denoising iteration, where t is a positive integer), used to represent the current environmental observation. The denoised feature sequence carries the internal semantic representation of the target data such as sensor input (e.g., images, language) after multiple steps of gradual refinement.
[0149] Through the above embodiments, features can be refined through multi-step denoising to reduce the interference of noise in the original data on action decisions and improve the accuracy of feature representation.
[0150] The generation unit 114 is used to generate real-time environmental perception based on the denoised feature sequence and the language context using the student model.
[0151] In this embodiment, the student model can generate a real-time environment-aware representation based on real-time denoising features and the language context.
[0152] The fusion unit 115 is used to fuse the real-time environment perception and the corresponding dynamic historical cache based on the learnable routing parameters using the student model to obtain a hybrid representation sequence.
[0153] In this embodiment, the fusion unit 115 uses the student model to fuse the real-time environment perception and the corresponding dynamic historical cache based on the learnable routing parameters to obtain a hybrid representation sequence including:
[0154] For each real-time environment perception, the product of the real-time environment perception and the learnable routing parameters is calculated to obtain the first feature;
[0155] Calculate the difference between 1 and the learnable routing parameters to obtain the complementary parameters of the learnable routing parameters;
[0156] The second feature is obtained by multiplying the complementary parameter with the corresponding dynamic history cache.
[0157] The sum of the first feature and the second feature is calculated to obtain the hybrid representation of the real-time environment perception;
[0158] The hybrid representation sequence is obtained by integrating all real-time environment-aware hybrid representations.
[0159] In this process, each fusion step utilizes dynamic historical cached features. Specifically, the fusion step t requires features from the previous t-1 steps, addressing the problem of traditional caching ignoring cross-temporal feature dependencies. This allows for dynamic reuse of historical features to reduce redundant computation and improve computational efficiency. Furthermore, the learnable routing parameters allow for flexible adjustment of feature trust levels, preventing error accumulation due to over-reliance on a single feature (e.g., trusting the cache when the environment is stable, and trusting the current perception when the environment changes drastically).
[0160] In the above embodiments, historical features can be utilized more smoothly in multi-round environmental interactions, achieving consistency within the decision-making process and effectively mitigating decision bias caused by error accumulation. The utilization efficiency of feature caching is improved without increasing additional GPU memory burden, enabling the model to respond more quickly to environmental changes in complex scenarios and maintain high decision-making stability.
[0161] In this embodiment, after obtaining the hybrid representation of the real-time environment perception, the hybrid representation of the real-time environment perception is stored in a cache.
[0162] For example, the hybrid representation of the real-time environment awareness can be stored in the cache according to the feature block index for quick subsequent retrieval.
[0163] In the above embodiments, updating cached features in real time can improve feature reuse efficiency while ensuring action response speed (no redundant calculation) and decision stability during the inference phase.
[0164] The input unit 116 is used to input the mixed representation sequence into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence.
[0165] In this embodiment, the target action probability distribution is a confidence assignment of each possible action (such as moving, grabbing, executing instructions, etc.), which is used to reflect the relative priority or selection tendency of each action in the current environmental state.
[0166] In this embodiment, after optimizing and training the parameters of the student model based on the teacher model, the student model is used to perform inference to generate the target action probability distribution of the optional action sequence. This can achieve a unified framework for training and inference, avoid the problem that the caching strategy optimized during training fails during inference in traditional methods, significantly enhance the generalization ability of the model across different tasks, and significantly improve the overall performance and robustness of the embodied intelligence system.
[0167] The selection unit 117 is used to select a target action from the sequence of available actions based on the target action probability distribution.
[0168] In this embodiment, the target action probability distribution is used to guide the selection of actions.
[0169] Specifically, the selection unit 117 selects a target action from the sequence of available actions based on the target action probability distribution, including:
[0170] Based on the target action probability distribution, the action with the highest confidence level is selected from the sequence of available actions as the target action.
[0171] Through the above embodiments, the optimal action rate can be accurately selected from the selectable action sequences based on the target action probability distribution output by the model.
[0172] In this embodiment, if the task reward drops below a threshold after the corresponding action execution entity performs the target action, an emergency cache update can be triggered to ensure the accuracy of action generation.
[0173] For example, when a financial service robot outputs inappropriate voice commands, or when an auxiliary robot in a medical operating room performs actions with deviations, an emergency cache update can be triggered to ensure the accuracy of the generated actions.
[0174] As can be seen from the above technical solutions, this invention can construct behavioral target proxy coefficients, breaking the limitation of traditional caching that only optimizes noise prediction errors, and binding caching strategies with actual task rewards; constructing a target loss function based on relative entropy, learnable routing parameters, and target proxy coefficients can effectively constrain decision consistency, protect task rewards, and balance caching efficiency; training the action generation model based on the action-perception collaborative teacher-student mechanism and target loss function enables the student model to effectively learn from the teacher model; using the student model to fuse real-time environmental perception with corresponding dynamic historical caching based on learnable routing parameters, and generating a target action probability distribution of optional action sequences based on the hybrid representation sequence to select the best target action, it can dynamically reuse historical features to reduce redundant calculations when generating actions, improving response efficiency while ensuring decision stability.
[0175] like Figure 3 The diagram shown is a structural schematic of a computer device that implements a preferred embodiment of the action generation method based on dynamic feature caching according to the present invention.
[0176] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as an action generation program based on dynamic feature caching.
[0177] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0178] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0179] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as code for action generation programs based on dynamic feature caching, but also to temporarily store data that has been output or will be output.
[0180] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing action generation programs based on dynamic feature caching) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0181] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-described embodiments of the action generation method based on dynamic feature caching, for example... Figure 1 The steps are shown.
[0182] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a construction unit 110, a training unit 111, a parsing unit 112, a denoising unit 113, a generation unit 114, a fusion unit 115, an input unit 116, and a selection unit 117.
[0183] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, a computer device, or a network device, etc.) or processor to execute portions of the action generation method based on dynamic feature caching described in the various embodiments of this invention.
[0184] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0185] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0186] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0187] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0188] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0189] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0190] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.
[0191] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0192] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0193] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0194] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement an action generation method based on dynamic feature caching, and the processor 13 can execute the multiple instructions to achieve the following:
[0195] Construct behavioral target proxy coefficients, and construct a target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficients;
[0196] Based on the action-perception collaborative teacher-student mechanism, the target loss function is used to train an action generation model including a teacher model and a student model.
[0197] In response to an action generation instruction based on target data, the action generation instruction is parsed to obtain a sequence of selectable actions in the language context and the current environment.
[0198] The student model is used to perform multi-step denoising on the target data according to the target time step to obtain a denoised feature sequence.
[0199] The student model is used to generate real-time environmental awareness based on the denoised feature sequence and the language context.
[0200] The student model is used to fuse the real-time environment perception with the corresponding dynamic historical cache based on the learnable routing parameters to obtain a hybrid representation sequence;
[0201] The hybrid representation sequence is input into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence;
[0202] The target action is selected from the sequence of available actions based on the target action probability distribution.
[0203] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0204] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0205] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0206] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0207] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0208] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0209] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0210] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0211] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An action generation method based on dynamic feature caching, characterized in that, The action generation method based on dynamic feature caching includes: Construct behavioral target proxy coefficients, and construct a target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficients; Based on the action-perception collaborative teacher-student mechanism, the target loss function is used to train an action generation model including a teacher model and a student model. In response to an action generation instruction based on target data, the action generation instruction is parsed to obtain a sequence of selectable actions in the language context and the current environment. The student model is used to perform multi-step denoising on the target data according to the target time step to obtain a denoised feature sequence. The student model is used to generate real-time environmental awareness based on the denoised feature sequence and the language context. The student model is used to fuse the real-time environment perception with the corresponding dynamic historical cache based on the learnable routing parameters to obtain a hybrid representation sequence; The hybrid representation sequence is input into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence; The target action is selected from the sequence of available actions based on the target action probability distribution.
2. The action generation method based on dynamic feature caching as described in claim 1, characterized in that, The constructed behavior target proxy coefficient includes: Calculate the decrease in task reward caused by cache errors under the constraints of environmental conditions and actions; The expected value of the decrease in task reward is calculated based on the real-time action probability distribution at each time step, thus obtaining the target proxy coefficient.
3. The action generation method based on dynamic feature caching as described in claim 1, characterized in that, The construction of the target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficient includes: Calculate the relative entropy loss between the action probability distribution output by the student model and the action probability distribution output by the teacher model; The action distribution alignment loss is obtained by multiplying the target proxy coefficient by the relative entropy loss. The learnable routing parameters are regularized to obtain the learnable routing parameters and; Obtain the configured regularization hyperparameters and calculate the product of the regularization hyperparameters and the sum of the learnable routing parameters to obtain the cache loss; The target loss function is obtained by summing the action distribution alignment loss and the cache loss.
4. The action generation method based on dynamic feature caching as described in claim 1, characterized in that, The teacher-student mechanism based on action-perception collaboration, which uses the objective loss function to train an action generation model including a teacher model and a student model, includes: The target loss function is minimized through backpropagation so that the output of the student model approaches the output of the teacher model. Training stops when the value of the target loss function no longer decreases. The currently obtained model is determined as the action generation model.
5. The action generation method based on dynamic feature caching as described in claim 1, characterized in that, The step of fusing the real-time environment perception with the corresponding dynamic historical cache based on the learnable routing parameters using the student model to obtain a hybrid representation sequence includes: For each real-time environment perception, the product of the real-time environment perception and the learnable routing parameters is calculated to obtain the first feature; Calculate the difference between 1 and the learnable routing parameters to obtain the complementary parameters of the learnable routing parameters; The second feature is obtained by multiplying the complementary parameter with the corresponding dynamic history cache. The sum of the first feature and the second feature is calculated to obtain the hybrid representation of the real-time environment perception; The hybrid representation sequence is obtained by integrating all real-time environment-aware hybrid representations.
6. The action generation method based on dynamic feature caching as described in claim 5, characterized in that, After obtaining the hybrid representation of the real-time environment perception, the method further includes: The hybrid representation of the real-time environment awareness is stored in a cache.
7. The action generation method based on dynamic feature caching as described in claim 1, characterized in that, The step of selecting a target action from the available action sequence based on the target action probability distribution includes: Based on the target action probability distribution, the action with the highest confidence level is selected from the sequence of available actions as the target action.
8. An action generation device based on dynamic feature caching, characterized in that, The action generation device based on dynamic feature caching includes: A construction unit is used to construct behavioral target proxy coefficients and construct a target loss function based on relative entropy, learnable routing parameters, and the target proxy coefficients; The training unit is used to train an action generation model, including a teacher model and a student model, based on a teacher-student mechanism that is based on action-perception collaboration, using the target loss function. The parsing unit is used to respond to the action generation instruction based on the target data, and parse the action generation instruction to obtain the language context and the optional action sequence under the current environment state; The denoising unit is used to perform multi-step denoising on the target data according to the target time step using the student model to obtain a denoised feature sequence. A generation unit is used to generate real-time environmental awareness based on the denoised feature sequence and the language context using the student model. The fusion unit is used to fuse the real-time environment perception and the corresponding dynamic historical cache based on the learnable routing parameters using the student model to obtain a hybrid representation sequence; An input unit is used to input the hybrid representation sequence into the action decision head of the student model to obtain the target action probability distribution of the optional action sequence; The selection unit is used to select a target action from the sequence of available actions based on the target action probability distribution.
9. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the action generation method based on dynamic feature caching as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the action generation method based on dynamic feature caching as described in any one of claims 1 to 7.