Robot control method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202611318584.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-25
AI Technical Summary
策略必须等待完整观测包到齐才能推理,动作生成频率被卡在最慢模态的频率上,无法实现高频闭环反应,接触/柔顺类任务表现受限
[0008]根据本申请实施例的一个方面,提供了一种计算机程序产品,包括计算机程序,计算机程序被处理器执行时实现上述机器人的控制方法。
Smart Images

Figure CN122807962A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial robot technology, and more specifically, to a robot control method, device, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] Vision-Language-Action (VLA) models transfer the general semantic knowledge of pre-trained Vision-Language Models (VLMs) to robot manipulation tasks, and have become the mainstream paradigm in the field of embodied intelligence. These models typically consist of three parts: a visual encoder, a language model backbone, and a motion head, and are obtained through motion-supervised fine-tuning on robot teaching data.
[0003] Existing VLAs generally inherit the synchronous single-clock structure of the vision-language pre-training paradigm, which aligns observations from all modalities on the same time grid, packages them at the same frequency, and then sends them uniformly into the backbone network to generate actions. The policy must wait for the complete observation package to arrive before inference, and the action generation frequency is stuck at the frequency of the slowest modality, making it impossible to achieve high-frequency closed-loop responses, thus limiting performance on contact / compliance tasks. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for controlling a robot, which can solve the above-mentioned problems of the prior art. The technical solution is as follows: According to one aspect of the embodiments of this application, a method for controlling a robot is provided, the robot including multiple modal sensors, the method comprising: Receive control commands; According to the control command, at least one control step is executed at a preset refresh frequency until a new control command is received. Wherein, the preset refresh frequency is greater than or equal to the maximum value among the refresh frequencies of all modal sensors; each control step includes: The current latent features stored in the latent feature buffers corresponding to each mode are fused to obtain the token for the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the sensor based on its own refresh frequency for the corresponding mode. The token input has a linear recursive model and is recursively processed step by step. The hidden state of a fixed size is updated to obtain the hidden state of the current control step. The hidden state is a vector obtained by compressing all tokens from the effective date of the control command to the current control step in a time sequence. Based on the hidden state and token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics; The linear recursive model is a selective state-space model, and the single-step recursion includes: Based on the token of the current control step, recursive parameters for the current control step are generated. The recursive parameters include a recursive step size parameter, an input projection coefficient, and a historical transfer matrix. The recursive step size parameter is used to control the forgetting rate of the historical hidden states. The input projection coefficient is a matrix that maps the token to the hidden state update space. The historical transfer matrix is a fixed-dimensional continuous system matrix. The historical transfer matrix is discretized based on the recursive step size parameter to obtain the discretized transfer matrix of the current control step. The hidden state of the previous control step is linearly transformed based on the discretized transfer matrix to obtain the historical transfer components in the hidden state update. The hidden state of the current control step is obtained by adding the historical transmission component to the input component obtained by mapping the token of the current control step through the input projection coefficient.
[0005] According to another aspect of the embodiments of this application, a control device for a robot is provided, the robot including a plurality of modal sensors, the device comprising: The control command receiving module is used to receive control commands; The control step execution module is used to execute at least one control step according to the control command at a preset refresh frequency until a new control command is received. The preset refresh frequency is greater than or equal to the maximum value among the refresh frequencies of all modes of sensors. Each control step includes: The current latent features stored in the latent feature buffers corresponding to each mode are fused to obtain the token for the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the sensor based on its own refresh frequency for the corresponding mode. The token input has a linear recursive model and is recursively processed step by step. The hidden state of a fixed size is updated to obtain the hidden state of the current control step. The hidden state is a vector obtained by compressing all tokens from the effective date of the control command to the current control step in a time sequence. Based on the hidden state and token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics; The linear recursive model is a selective state-space model, and the single-step recursion includes: Based on the token of the current control step, recursive parameters for the current control step are generated. The recursive parameters include a recursive step size parameter, an input projection coefficient, and a historical transfer matrix. The recursive step size parameter is used to control the forgetting rate of the historical hidden states. The input projection coefficient is a matrix that maps the token to the hidden state update space. The historical transfer matrix is a fixed-dimensional continuous system matrix. The historical transfer matrix is discretized based on the recursive step size parameter to obtain the discretized transfer matrix of the current control step. The hidden state of the previous control step is linearly transformed based on the discretized transfer matrix to obtain the historical transfer components in the hidden state update. The hidden state of the current control step is obtained by adding the historical transmission component to the input component obtained by mapping the token of the current control step through the input projection coefficient.
[0006] According to another aspect of the embodiments of this application, an electronic device is provided, the electronic device including a memory, a processor and a computer program stored in the memory, the processor executing the computer program to implement the above-described robot control method.
[0007] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-described robot control method.
[0008] According to one aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described robot control method.
[0009] The beneficial effects of the technical solution provided in this application are as follows: By establishing independent latent feature buffers for sensors of each modality and maintaining the latest values with a zero-order hold mechanism, the forced alignment constraint of traditional synchronous single-frequency clocks on cross-modal sampling is broken. This allows slow modalities such as vision to be refreshed at low frequency according to semantic change rate, while fast modalities such as ontology / force can be sampled at high frequency according to physical needs. This avoids redundant calculations in slow modalities and prevents undersampling of information in fast modalities. Secondly, the maximum refresh rate of all modalities is used as the system control beat. Combined with the mechanism of generating state representative words from the fusion buffer features at each step, this ensures that the linear recursive model can continue to advance at a frequency much higher than the visual refresh rate. This fundamentally removes the physical limitation that the action generation frequency is capped by the slowest modality. In actual tests, high-frequency closed-loop control at the 200Hz level can be achieved. Furthermore, a linear recursive model with a fixed-size hidden state is used to causally compress the representative lexical units of the entire historical state, achieving the memory capability of long-term non-Markov states with O(L) linear complexity. This avoids the memory explosion problem caused by the O(L²) complexity of full-history attention while retaining decision information from early key events, enabling differentiated and correct actions to be output even in scenarios with the same observation but different task histories. Finally, by completely decoupling the process of generating actions from action instructions from the modal refresh time, the system can output actions at each control step without waiting for a complete observation packet. This reduces the action latency from hundreds of milliseconds of waiting for multimodal data synchronization to microseconds of single-step recursion, significantly improving the execution stability and success rate of tasks with extremely high real-time requirements, such as contact operations and compliant control. At the same time, it achieves dynamic optimization allocation of computing power according to the modal semantic change rate. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0011] Figure 1 A schematic diagram illustrating the application environment for implementing the robot control method provided in the embodiments of this application; Figure 2 A flowchart illustrating a robot control method provided in an embodiment of this application; Figure 3 A flowchart illustrating a robot control method based on multimodal asynchronous buffering and SSM provided in this application embodiment; Figure 4 A flowchart of selectively parameterized SSM calculation is provided for embodiments of this application; Figure 5 A schematic diagram of a multi-frequency asynchronous buffer timing refresh process provided in an embodiment of this application; Figure 6 A flowchart illustrating a robot control method provided in another embodiment of this application; Figure 7 This is a schematic diagram of the structure of a robot control device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0013] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0015] First, let's introduce and explain several terms used in this application: The modality set is denoted as M, where m∈M represents a certain modality (such as language, vision, proprio, force / touch).
[0016] latentbuffer: A latent feature buffer maintained independently for each mode, asynchronously refreshed according to the sensor's own frequency for that mode, maintaining the most recent value between two refreshes, i.e., zero-order hold (ZOH). The latentbuffer for the m-th mode is defined as Z... m Buffer set B={Z m |m∈M}.
[0017] f mThe refresh rate of mode m, f ctrl This indicates the highest control frequency, the highest clock cycle on which the system action output is based; the SSM and action head advance step by step at this frequency, decoupled from the refresh time of each mode.
[0018] t represents the discrete physical control step with the highest control frequency as the beat.
[0019] The state at step t represents the token.
[0020] h t Let represent the hidden state at step t of the SSM.
[0021] y t This represents the output of the SSM at step t.
[0022] a t This represents the action generated in step t.
[0023] VLA: Transferring semantic knowledge from pre-trained vision-language models to a policy model paradigm for generating robot maneuvers.
[0024] The Selective State Space Model (SSM) is a sequence state propagation structure with constant-size hidden states, input-dependent gating, and linear time complexity, used for causal compressed memory of the entire history.
[0025] Full-history causal memory: Causally compress the entire history from the start of the task to the present with a constant-size hidden state, so that the strategy can distinguish non-Markov states that are "the same in appearance but different in history".
[0026] The Visual-Language Model (VLA) model, which transfers general semantic knowledge from a pre-trained Visual-Language Model (VLM) to robotic manipulation tasks, has become the mainstream paradigm in the field of embodied intelligence. Such models typically consist of three parts: a visual encoder, a language model backbone, and a motion head, and are fine-tuned through motion supervision on robot teaching data. When performing operations, robots are usually equipped with various sensors: cameras (vision), joint encoders (proprioception / proprioception state), force / torque sensors in the wrist or fingers, tactile arrays, etc.
[0027] Research has revealed that the existing VLA inference paradigm has the following two key technical flaws that have not yet been addressed simultaneously.
[0028] Question 1: Cross-modal frequency mismatch and action delay caused by synchronizing a single clock Existing VLAs generally inherit the synchronous single-clock structure of the vision-language pre-training paradigm, which aligns observations from all modalities on the same time grid, packages them at the same frequency, and then feeds them uniformly into the backbone network to generate actions. This structure has three major drawbacks: 1. Redundant computation: Slow modalities with low semantic change rates, such as vision (typical effective refresh rate requirement of about 3-10Hz), are forcibly processed with high-frequency repetition, and a lot of computing power is wasted on almost unchanging images. 2. Cross-modal frequency mismatch: Fast modes such as force / torque (typically 100–500Hz) and proprioception are undersampled, and high-frequency dynamic information in the contact phase is smoothed out or lost on the sampling grid; 3. Action delay and frequency capping: The strategy must wait for the complete observation package to arrive before inference can be performed. The action generation frequency is capped at the slowest mode frequency, making it impossible to achieve high-frequency closed-loop response, which limits the performance of contact / compliant tasks.
[0029] Question 2: Markov or short-window strategies fail on long-duration non-Markov tasks. Existing strategies are mostly Markov strategies, making decisions based solely on the current observation or a short window of the most recent frames. However, many real-world operational tasks are non-Markovian, and their fundamental difficulty lies in the fact that the same observation requires different actions due to different historical contexts. For example, in a task that requires swapping the positions of two objects: at some point in the middle, the image looks almost identical to the initial moment, but the correct action depends on which object was grabbed first; this cannot be distinguished based on the current image alone, and must rely on the memory of the entire task.
[0030] Existing short-time memory schemes (such as using GRU or rolling buffers to compress the most recent K frames) can only retain information from the most recent few frames. As the task duration increases, early critical events are squeezed out of the window and lost, making them ineffective for long-duration tasks. If standard attention is used to model all historical frames directly, the computational complexity increases by O(L^2) as the sequence length L grows. 2 The cache expands uncontrollably over time, making it impractical for high-frequency, long-term scenarios.
[0031] Limitations of existing solutions: Existing technologies often address these two issues separately: either focusing on multimodal fusion while using a synchronous single-frequency clock, or focusing on long-range memory while assuming a single sampling frequency. Simply combining the two also presents fundamental obstacles: 1. If the full historical memory mechanism is based on synchronous single-frequency observation sequences, its time step is still capped by the slowest mode, and high-frequency response capability is out of the question; 2. If asynchronous multi-frequency buffers are only connected to short-term memory, long-term causal memory will still be missing; 3. The time axes on which the two depend are inconsistent, and there is a lack of a unified state propagation structure that can simultaneously support asynchronous multi-frequency input and constant-size full-history compression.
[0032] The robot control method, device, electronic equipment, computer-readable storage medium, and computer program product provided in this application are intended to solve the above-mentioned technical problems of the prior art.
[0033] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0034] Figure 1 This is a schematic diagram of an application environment for implementing the robot control method provided in this application embodiment. This application environment may include at least a multimodal sensing unit 10, a semantic command interface unit 20, an intelligent decision-making and memory unit 30, and a robot execution terminal 40. In practical applications, the multimodal sensing unit 10, the semantic command interface unit 20, the intelligent decision-making and memory unit 30, and the robot execution terminal 40 can be directly or indirectly connected via industrial Ethernet (such as EtherCAT, Profinet), fieldbus (such as CANopen), or distributed real-time network (such as TSN) to achieve high-speed, deterministic transmission between sensor data, semantic commands, and action commands. This application does not impose any limitations on this.
[0035] In this embodiment, the multimodal sensing unit 10 can be a hardware module integrating a visual sensor (such as an RGB-D camera), a force / tactile sensor (such as a six-dimensional force sensor or a joint torque sensor), a proprioceptive sensor (such as an encoder or an IMU), and an environmental sensing sensor (such as a lidar or ultrasonic sensor). It can also include a multimodal fusion algorithm running on an edge computing device. Specifically, the visual sensor can periodically (e.g., every 100ms) acquire scene images and extract visual latent features, storing them in a visual buffer; the force / tactile sensor acquires contact force information at a high frequency (e.g., 1ms) and encodes it as force latent features, storing them in a force buffer; the proprioceptive sensor acquires information such as motor current, position, and speed at the highest control frequency (e.g., 1ms) and encodes it as proprioceptive latent features, storing them in a proprioceptive buffer. Each modal sensor achieves asynchronous data acquisition and feature retention through an independent sampling clock and buffering mechanism.
[0036] In this embodiment, the semantic instruction interface unit 20 can be a natural language processing module, a task planner, a teach pendant, or a host computer instruction system. It receives high-level task instructions (such as grabbing the red box or placing it in the blue tray) and encodes them into fixed-dimensional semantic feature vectors, storing them in a language buffer. The semantic instruction interface unit 20 can issue instructions all at once or in segments when the task starts, without needing to resend them in each control cycle, thus supporting long-term task planning and context maintenance.
[0037] In this embodiment, the intelligent decision-making and memory unit 30 is the core control module of the solution. It can adopt an embedded ARM+DSP+FPGA heterogeneous architecture or run on an edge AI server, possessing high real-time performance and parallel processing capabilities. The intelligent decision-making and memory unit 30 receives control commands; According to the control command, at least one control step is executed at a preset refresh frequency until a new control command is received. The preset refresh frequency is greater than or equal to the maximum value among the refresh frequencies of all sensor modes. Each control step includes: The current latent features stored in the latent feature buffers corresponding to each mode are fused to obtain the token for the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the sensor based on its own refresh frequency for the corresponding mode. The token input has a linear recursive model and is recursively processed step by step. The hidden state of a fixed size is updated to obtain the hidden state of the current control step. The hidden state is a vector obtained by compressing all tokens from the effective date of the control command to the current control step in a time sequence. Based on the hidden state of the current control step and the token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics.
[0038] In this embodiment, the robot execution terminal 40 may be an industrial robotic arm (such as a 6-axis collaborative robot), an AGV mobile chassis with a lifting mechanism, a CNC machining center feed system, a humanoid robot joint module, etc., and may also include software running in the physical equipment (such as a kinematic model library, a dynamic parameter table, a sensor signal processing module).
[0039] The robot execution terminal 40 is used to collect real-time status data such as motor current, speed, position, and electrical angle, and send them to the intelligent decision and memory unit 30. It also receives action commands (such as target joint angle and end pose) issued by the intelligent decision and memory unit 30, and drives the servo motor + reducer + linkage to complete high-precision trajectory tracking. In contact tasks (such as grinding, assembly, and handling), it provides real-time feedback of force / tactile information and supports force gating mechanism to trigger contact perception and response.
[0040] In addition, it should be noted that, Figure 1 The example shown is merely an application environment for an adaptive servo control system based on multimodal asynchronous perception, selective memory, and force-gated injection. This application environment can include more or fewer nodes (e.g., adding a host computer monitoring system, a cloud-based parameter optimization platform, or a digital twin model), and this application does not impose any limitations on this. For example, in high-end manufacturing scenarios, a cloud-based digital twin model can be introduced to offline train multimodal weight mapping tables, force-gated thresholds, and SSM recursive parameter adaptive rules under different working conditions, and then distributed to the edge intelligent decision and memory unit 30 to achieve integrated cloud-edge-device intelligent control. In human-machine collaboration scenarios, a safety monitoring module can be added to detect personnel approaching in real time and trigger protective actions. At the same time, it can receive voice commands such as pause and resume through the semantic command interface unit 20 to achieve natural human-machine interaction.
[0041] The inventive concept of this application lies in coupling multi-frequency asynchronous observation with a constant-size full-history causal memory through a single state propagation path, thereby completely decoupling the action generation frequency from the refresh time of each modality.
[0042] Key Point 1: The selective state-space model operates on an asynchronous multi-frequency buffer, advancing the entire historical causal memory with the highest control frequency. Existing VLA follows the synchronous single-clock structure of vision-language pre-training, aligning all modalities with the same time grid and processing them at the same frequency. This results in the action generation frequency being stuck at the slowest modality (vision, typically effective refresh rate 3–10Hz), making high-frequency loop closure impossible. If full-history attention is used to model long-range memory, the complexity increases to O(L). 2 Expansion and cache grow unbounded over time.
[0043] The key point 1 solution includes: maintaining an independent latent feature buffer for each modality m. Each component refreshes asynchronously at its own frequency, maintaining the most recent value between two refreshes; at the highest control frequency For each control step of the beat, merge the current modal buffers into a state representative token. The process proceeds step by step by feeding in a selective state-space model with hidden states of constant size:
[0044]
[0045] Among them, the discretization step size With projection From the current input via the selective parameterization module Generate, continuous matrix according to Discretization.
[0046] This recursion is Linear complexity ( L To control the number of steps, the hidden state dimension is constant and the video memory does not increase over time; in the example, when the control frequency is 100Hz and the task duration is 60 seconds (approximately 6000 control steps), the video memory usage does not increase with the number of steps.
[0047] The advantage of key point 1 lies in the coexistence of causal memory and high-frequency response throughout the process. SSM advances every step, and the slow mode is filled with the hold value when it is not refreshed. The action head can produce an action every step, which fundamentally removes the constraint that the action is capped by the slowest mode.
[0048] Key Point 2: Decoupling the triggering time of the action head from the refresh time of each modal buffer. Existing strategies require waiting for the complete observation package to arrive before inference can be performed, resulting in high action latency, inability to achieve high-frequency closed-loop responses, and limited performance in contact / compliant tasks.
[0049] The solution for key point 2 includes: the action head is based on the current full history state at each inference step. And perform a read-only query to generate the buffer. Its triggering time is completely decoupled from the refresh time of each modality. Even if the vision is not refreshed at this moment, the motion head still produces an action in each control step. The SSM single-step recursion and motion head are deployed on the high-frequency real-time controller or edge computing unit, while the vision encoding is left in the low-frequency loop.
[0050] In this embodiment, on a platform equipped with a high-frequency real-time controller, the measured effective control frequency can reach 200Hz, while the visual system still maintains a low-frequency update of about 1Hz.
[0051] The advantage of key point 2 is that the action latency is reduced from waiting for a complete observation packet to single-step recursion latency; the computing power is allocated according to the semantic change rate of each modality, with slow modalities updated at low frequency and fast modalities updated at high frequency, reducing redundant computation.
[0052] Key Point 3: The full historical memory solves the long-range non-Markov problem of the same observation requiring different actions due to different historical contexts. Existing short-term memory schemes (GRU, rolling window) tend to push early critical events out of the window and lose them as the task duration increases, thus failing on long-term non-Markov tasks.
[0053] Key point 3's solution includes: constant-size hidden states in SSM for full-history causal compression, and selective parameters. Adapting to input, the model autonomously decides what to remember, forget, and output at each step, thus compressing the entire task into a constant-size dataset. Early critical events can be selectively retained.
[0054] In the task of swapping the positions of two objects, the view at the intermediate time is approximately the same as that at the initial time. The full-history hidden state retains historical information about which one to grab first, enabling the strategy to output the correct differentiated action under the same observation, while the short-window baseline fails at that time.
[0055] The advantage of key point 3 lies in achieving full-process causal memory with constant video memory, balancing long-term memory and high-frequency response, thus improving the success rate of long-term tasks.
[0056] This application provides a robot control method. The robot includes multiple modal sensors, including but not limited to cameras (as visual sensors), joint encoders (for sensing proprioception / proprioceptive state), force / torque sensors at the wrist or fingertips, and tactile arrays. Figure 2 As shown, the method includes: S101, Receive control commands; S102. According to the control command, execute at least one control step at a preset refresh frequency, and control the robot to move according to the action command generated by each control step until a new control command is received.
[0057] This embodiment takes the receipt of a control command as the trigger point and uses the highest control frequency during the effective period of the control command as the beat to perform single-step reasoning in a loop until the control command is updated or the task ends.
[0058] In some embodiments, control instructions can be natural language task instructions. Natural language task instructions are semantic instructions given in free text form that describe what task needs to be performed. They are typically entered once at the beginning of a task episode and remain static throughout the episode. For example, placing the red square on the left next to the cup on the right is a control instruction described in natural language, i.e., a natural language task instruction.
[0059] In some embodiments, control instructions can also be high-level behavioral instructions. High-level behavioral instructions are behavioral switching signals issued by upper-level planners, state machines, or human-machine interfaces. They are more abstract than lower-level joint movements but more structured than natural language tasks, and are used to switch strategies between different sub-stages, but do not directly specify motor commands. Examples include reach, grasp, insert, and release.
[0060] Taking the scenario of controlling a kitchen robot to pour tea from a teapot into a cup as an example, the natural language task instruction could be to pour the tea from the teapot into the cup on the right, while the higher-level behavioral instructions would be issued sequentially by the state machine: align_tea_pot→lift_pot→tilt_pour→put_down.
[0061] It should be noted that, regardless of the type of control instruction used, the control instructions are updated at a low frequency and participate in the construction of state representative lexies through the instruction-side buffer during their effective period, but do not participate in the SRL single-step recursion blocking that proceeds at the highest control frequency.
[0062] It should be understood that each modal sensor has its own physical sampling / encoding refresh rate. For example, a camera captures images at a frequency of 30Hz, but the effective range of semantic encoding is 1~10Hz. The update frequency of proprioception (joint angle / velocity) is 100Hz, the update frequency of six-dimensional force / torque is 500Hz, and the update frequency of the tactile array is 200Hz. This application uses a frequency greater than or equal to the maximum value among the refresh frequencies of all modal sensors as the preset refresh frequency. In the above example, that is, the control step is executed at least at a beat of 500Hz.
[0063] In some embodiments, each control step includes: S1021. The current latent features stored in the latent feature buffer corresponding to each mode are fused to obtain the token of the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the corresponding mode sensor based on its own refresh frequency.
[0064] In this embodiment, an independent latent feature buffer Z is maintained for each modal sensor. m Each modal buffer is refreshed independently, with the refresh time determined solely by the hardware sampling frequency and software coding strategy of that modal sensor. They do not block each other. When a modal sensor does not generate new sensing data in the current control step, its corresponding latent feature buffer retains the latent features obtained in the most recent refresh. Through this mechanism, the coding delay of slow modalities (such as vision) will not block the advancement of the high-frequency control loop.
[0065] In this embodiment, the latent features are generated online by encoders of the corresponding modalities. For example, a visual encoder extracts deep features from images captured by a camera, a proprioceptive encoder encodes joint angles and velocities, and a force sensor encoder encodes six-dimensional forces or torques.
[0066] In each control step t with the highest control frequency as the beat, the currently stored latent features in the latent feature buffer corresponding to each modality are read and fused into a vector representation of a unified dimension, called the state representative word (also called the state representative token or token) of the current control step: x t .
[0067] In some embodiments, the latent features of multiple modalities are concatenated along the feature dimension, and then linearly projected and layer normalized to obtain the state representative token. This token is semantically equivalent to compressing the multimodal observations and ontology state at the current moment into a compact representation in one frame.
[0068] S1022. The token input has a linear recursive model and is recursively processed step by step to update the hidden state of a fixed size to obtain the hidden state of the current control step. The hidden state is a vector obtained by compressing all tokens from the effective date of the control instruction to the current control step in a time sequence.
[0069] This embodiment uses a linear recursive model with a constant hidden state dimension as the historical memory carrier. The model takes a state representative token as input and performs a single-step recursion as follows: Maintain a hidden state h of fixed size. t Its dimensions do not increase with task duration; In the t-th control step (i.e., the current control step), the token is represented by the current state: x. t The hidden state h of the previous control step t-1 The hidden state is updated through linear transformation and accumulation operation: This embodiment does not limit the model structure of the linear recursive model. For example, it can be a selective state-space model (SSM), linear attention recursion, gated linear recursive unit, etc.
[0070] S1023. Based on the hidden state of the current control step and the token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics.
[0071] In some embodiments, based on the hidden state h of the current control step t And the status token: x t The action command 'a' for the current control step is generated through an action head. t The embodiments of this application do not limit the content of the action command, such as indicating joint position, joint velocity, end-effector pose increment and control signal, etc.
[0072] In this embodiment, the triggering time of the action head is completely decoupled from the refresh time of the sensors of each modality: even if a certain modality (such as vision) has not been refreshed in the current control step, the action head can still output action commands based on the current hidden state and the latest hidden features in the buffer. In other words, this embodiment does not need to wait for the complete observation packets of all modalities to arrive before inference can be performed, thereby significantly reducing action latency.
[0073] As an optional implementation, the single-step recursion of the linear recursive model and the action head are deployed on the controller or edge computing unit, while computationally intensive modules such as the visual encoder run in a low-frequency loop, thereby achieving a reasonable allocation of computing power according to the modal semantic change rate.
[0074] The robot control method provided in this embodiment breaks the forced alignment constraint of traditional synchronous single-frequency clocks on cross-modal sampling by establishing independent latent feature buffers for sensors of each modality and maintaining the latest values with a zero-order hold mechanism. This allows slow modalities such as vision to refresh at low frequency according to semantic change rate, while fast modalities such as ontology / force can sample at high frequency according to physical needs. This avoids redundant calculations in slow modalities and prevents undersampling of information in fast modalities. Secondly, the system control beat is set to the maximum refresh frequency of all modalities. Combined with the mechanism of generating state representative words from the fusion buffer features at each step, this ensures that the linear recursive model can continuously advance at a frequency much higher than the visual refresh rate. This fundamentally removes the physical limitation that the action generation frequency is capped by the slowest modality. In actual tests, high-frequency closed-loop control at the 200Hz level can be achieved. Thirdly, a linear recursive model with a fixed-size latent state is used to perform causal compression on the representative words of the entire history of states. This achieves the memory capability of long-term non-Markov states with O(L) linear complexity, avoiding the O(L) complexity of full-history attention. 2 While mitigating the memory explosion problem caused by complexity, the system retains decision-making information from early critical events, enabling it to output differentiated and correct actions even in scenarios with identical observation scenes but different task histories. Finally, by completely decoupling the process of generating actions from action commands from the modal refresh time, the system can output actions at each control step without waiting for a complete observation packet. This reduces action latency from hundreds of milliseconds of waiting for multimodal data synchronization to microseconds of single-step recursion, significantly improving the execution stability and success rate of tasks with extremely high real-time requirements, such as contact operations and compliant control. Simultaneously, it achieves dynamic optimization allocation of computing power based on the modal semantic change rate.
[0075] Based on the above embodiments, as an optional embodiment, receiving control instructions described in natural language includes: The control command is encoded to obtain the corresponding hidden feature and stored in the command modality buffer. The command modality buffer is used to store the hidden feature corresponding to the latest control command. The process of fusing the current latent features stored in the latent feature buffer corresponding to each modality to obtain the token for the current control step includes: The latent feature buffers corresponding to each mode and the currently stored latent features in the instruction mode buffer are fused to obtain the token for the current control step.
[0076] In this embodiment, control commands can be input in natural language, such as placing the red square behind the blue cup. To further reduce the computational overhead of the language modality and adapt to the asynchronous buffer architecture, this embodiment employs a static encoding and long-term retention strategy for the control commands.
[0077] Specifically, at the start of a task episode, the language encoder is invoked only once to perform semantic encoding on a new natural language control instruction, generating fixed-dimensional latent language features. This embodiment does not limit the type of language encoder, such as a pre-trained large language model or a text Transformer.
[0078] The language latent feature does not participate in subsequent frame-by-frame high-frequency updates. Instead, it is written into a dedicated Instruction Modality Buffer and remains unchanged throughout the entire task round unless a new control instruction is received to overwrite or reset it.
[0079] During the construction of state representative lexical units at each control step, the text is no longer re-encoded. Instead, the system directly reads the currently stored latent linguistic features from the instruction modality buffer and treats them as an input source equivalent to ordinary sensory modalities (visual, proprioceptive, etc.). Subsequently, the system fuses the latent linguistic features in the instruction modality buffer with the current values in the latent feature buffers corresponding to each modality. For example, it concatenates them along the feature dimension and then performs linear projection and layer normalization to obtain the current control step state representative lexical unit containing semantic constraints.
[0080] This embodiment achieves the effect of encoding once and reusing throughout the entire process by incorporating language instructions into a unified asynchronous buffer framework. This not only completely eliminates the redundant computational burden caused by repeatedly encoding text at each step of reasoning in traditional VLA models, but also ensures that semantic intent remains absolutely stable during high-frequency control processes lasting thousands of steps, avoiding instruction drift caused by floating-point errors or repeated reasoning. Simultaneously, because the instruction buffer refresh frequency is extremely low, it aligns with multi-frequency asynchronous processing mechanisms, preventing the high-frequency control loop from being blocked while waiting for text input. This ensures that language semantics can be seamlessly integrated into high-frequency, full-history causal memory, guiding the robot to complete long-range, complex operational tasks.
[0081] Please see Figure 3 This example illustrates a flowchart of the robot control method based on multimodal asynchronous buffering and SSM in this embodiment. In this embodiment, the robot control system includes four core input modalities: language commands, visual sensors, proprioceptive sensors, and force / tactile sensors, each with different physical refresh frequencies and semantic importance. The system achieves a high-frequency, low-latency, long-term causal memory control closed loop through a pipeline architecture of encoder → buffer → state representation token construction → SSM hidden state update → decoupled action head → action command output.
[0082] The language instruction path (low-frequency, once per round) inputs natural language instructions, input only once at the start of the task or when triggered by a human operator. These natural language instructions are encoded into fixed-dimensional latent features via a language encoder (such as a pre-trained LLM or a lightweight text Transformer). The encoded results are stored in a language buffer and remain unchanged throughout the task rounds unless a new instruction is received. Silver leaves in the language buffer serve as global semantic constraints, participating in the state representation token construction at each control step, but are not updated with the control frequency.
[0083] The input to the visual sensor path (approximately 1Hz) is an image captured by a camera, with a raw sampling rate of approximately 1Hz. The image is then processed by a visual encoder (such as ResNet, ViT, or a lightweight CNN) to extract high-level semantic features. The encoded result is stored in a visual buffer, maintaining the latest value between adjacent control steps. It's important to note that although the visual sensor refreshes slowly, its semantic information (such as object positions and scene layout) is continuously injected into state representation tokens, guiding long-term planning.
[0084] The input to the proprioceptive sensor path (100Hz) consists of state signals collected by proprioceptive sensors such as joint encoders, IMUs, and motor current sensors, refreshed at a high frequency of 100Hz. These state signals are encoded into latent proprioceptive state features by a proprioceptive encoder (such as an MLP or a small Transformer) and updated to the proprioceptive buffer in real time. The proprioceptive sensor provides high-precision, high-frequency kinematic / dynamic feedback, supporting real-time closed-loop control.
[0085] The force / tactile sensor path (high frequency, greater than 100Hz) receives force or tactile signals from a six-dimensional force sensor, tactile array, etc., with a refresh rate higher than or equal to proprioceptive sensing (e.g., 200Hz~1kHz). The force or tactile signals are processed by a force encoder to extract latent force or tactile features, which are then updated to the force buffer in real time. The force / tactile sensor captures high-frequency physical interaction signals such as contact transients and friction changes, used for compliant control and obstacle avoidance.
[0086] In each control step, the currently stored latent features are read from the language buffer, vision buffer, ontology buffer, and force buffer. These latent features are concatenated along the feature dimension, and after linear projection and layer normalization, a state representation token with a unified dimension is generated: x. t It should be understood that this token is a compressed representation of the current multimodal observations, ontology state, and global semantics, and serves as the input to the SSM.
[0087] SSM maintains a hidden state h of fixed dimension. t Based on x t and , The hidden state h is updated using a linear recursive formula. t It carries compressed causal information from the start of the task to the current control step, and can distinguish non-Markov states that have the same observations but different histories.
[0088] Based on the hidden state h of the current control step t (Optional, combined with the current x) t The action command a for the current control step is generated through a decoupled action head. t The triggering time of the motion head is completely decoupled from the refresh times of each modal sensor. Even if the vision does not refresh in the current control step, the motion head still operates based on h. t With x t Output actions to ensure the control frequency remains at 100Hz. Output action commands can be sent directly to robot actuators, such as joint servos, gripper controllers, etc.
[0089] Based on the above embodiments, as an optional embodiment, the linear recursive model is a selective state-space model, and the single-step recursion includes: Based on the token of the current control step, recursive parameters for the current control step are generated. The recursive parameters include a recursive step size parameter, an input projection coefficient, and a history transfer matrix. The recursive step size parameter is used to control the forgetting rate of the historical hidden states. The input projection coefficient is a matrix that maps the token to the hidden state update space. The history transfer matrix is a fixed-dimensional continuous system matrix. The historical transfer matrix is discretized based on the recursive step size parameter to obtain the discretized transfer matrix of the current control step. The hidden state of the previous control step is linearly transformed based on the discretized transfer matrix to obtain the historical transfer components in the hidden state update. The historical transmission component is added to the input component obtained by mapping the token of the current control step through the input projection coefficient to obtain the updated hidden state (i.e. the hidden state of the current control step).
[0090] Unlike traditional recursive models that use fixed parameters, the recursive parameters of this SSM are dynamically generated by the current input at each control step. Its core is to achieve selective memorization and forgetting of historical information based on the idea of discretization of continuous systems.
[0091] Specifically, in the t-th control step, first obtain the current state representative token: x t Subsequently, a lightweight selective parameterization module was used. ψ t (For example, implemented by a small multilayer perceptron MLP), with x t As input, generate the set of input-dependent parameters for the current control step, including at least: Recursive step size parameter Δ t : A scalar or diagonal matrix that physically represents the time scale or discretization step size of the current control step and is used to control the forgetting rate of historical hidden states; Input projection coefficients A learnable matrix used to represent the current state as a lexical x. t Mapped to the hidden state update space; The history transfer matrix A is a fixed-dimensional continuous system matrix whose parameters are determined during the model training phase and remain unchanged during the inference phase, representing the intrinsic dynamic structure of the state space.
[0092] After obtaining the above parameters, this embodiment does not directly use matrix A for recursion, but first discretizes the system. Specifically, it uses the recursion step size parameter Δ t Discretize the fixed continuous system matrix A to obtain the discretized transfer matrix of the current control step. In a typical implementation, this discretization process follows the zeroth-order preservation assumption, and the calculation formula is as follows: =exp(Δ t A) Where, exp This indicates matrix exponentiation operations.
[0093] Based on this, SSM performs a single-step recursion to update the hidden state h. t Its calculation process is divided into two parts: Historical transmission component: the hidden state of the previous control step Discretized transfer matrix Obtained by linear transformation, i.e. This component determines the degree to which historical information is retained in the current step; Input component: The word x represented by the current state t After inputting projection coefficients The mapping yields, i.e. This component determines the intensity of the current observation's update of the hidden state.
[0094] Finally, the historical propagation components and the input components are added together to obtain the updated hidden state. :
[0095] Through the above mechanism, when Δ t When the value is large, the discretized transfer matrix A matrix closer to zero has its historical information quickly forgotten; when Δ t When smaller, It is closer to the identity matrix, allowing historical information to be preserved for a long time. This determines the current input x t The contribution weight to the hidden state update. Due to Δ t and All are composed of x t Dynamically generated by the selective parameterization module, the model can autonomously decide what to remember, forget, and update based on the current observation content (e.g., whether it is at the moment of contact or whether a visual mutation has occurred), thereby achieving causal compression of all historical information with a constant-size hidden state.
[0096] This recursive process involves only matrix multiplication and addition operations, with a computational complexity of O(L) (where L is the number of control steps), and the matrix... With fixed dimensions, it does not increase memory overhead as the task duration increases, making it very suitable for high-frequency, long-duration control scenarios for robots.
[0097] Based on the above embodiments, as an optional embodiment, the recursive parameters also include output projection coefficients, which are input dependency matrices generated by the token of the current control step and map the hidden state to the output feature space.
[0098] Based on the hidden state and token of the current control step, obtain the output features of the current control step, including: The hidden state of the current control step is mapped through the output projection coefficients to obtain the hidden state contribution component; The token of the current control step is mapped through the pre-learned jump connection parameters to obtain the observation direct transmission component; The hidden state contribution component is added to the observation direct transmission component to obtain the output feature of the current control step.
[0099] The hidden state h is completed through the selective parameterization mechanism described in the above embodiments. t After a single-step update, a feature representation that can be used by the action head needs to be generated based on the hidden state. To this end, this application introduces an input-dependent output projection mechanism and an observation direct connection path to balance the abstract representation of long-term historical information with the preservation of details of the current observation.
[0100] Specifically, in the t-th control step, the selective parameterization module... ψt ( In addition to generating the recursive step size parameter and the input projection coefficients, an output projection coefficient is also generated simultaneously. .and similar, It is also a word whose current state represents the word. x t The input-dependent learnable matrix, dynamically generated by the selective parameterization module, serves to transform the fixed-dimensional hidden states h... t Mapped to the output feature space suitable for action generation.
[0101] Based on the above parameters, the system executes the process of constructing output features, which includes two parallel branches: Hidden state contribution component: The hidden state h of the current control step t Output projection coefficients Linear mapping, to obtain h t This component carries compressed causal information from the start of the mission to the current control step, providing a long-term decision context related to mission objectives and historical trajectories. Because... It is determined by the current input x t Dynamically generated, the model can adaptively adjust which features to extract from historical memory based on the current observation context, such as the stage of the task and the reliability of the current perception. For example, in the fine motor skills stage, This can enhance the weights of latent state dimensions related to force / tactile sensation; while during free movement, Then we can focus more on the hidden state dimension related to spatial location.
[0102] Direct observation component: To compensate for potential detail loss during information transmission in the recursive model, the system also includes a skip connection path. The state representation word x for the current control step. t D is obtained by directly performing a linear mapping through a pre-learned fixed jump connection parameter D. x t This component does not undergo recursive compression of the hidden state, preserving the original details of the current observation, such as instantaneous changes in contact force and precise joint positions, thus providing high-fidelity real-time feedback for motion generation.
[0103] The two components mentioned above are fused to obtain the final output feature of the current control step. :
[0104] This fusion process can be carried out using an element-wise addition method to achieve complementary advantages between historical memory and current observations.
[0105] Through the above mechanism It contains both task-level memories formed through long-term evolution and precise observational details of the current moment. Subsequently, this output feature... The action head (e.g., a lightweight MLP regression head) is input to the downstream action head to predict the action command a for the current control step. t .
[0106] The output projection + jump connection structure design provided in this embodiment not only alleviates the gradient vanishing and information bottleneck problems common in deep recursive networks, but also enables action generation to benefit from both historical experience accumulation and current accurate perception. It is particularly suitable for robot operation tasks that require both long-term planning and high-frequency fine control.
[0107] Please see Figure 4 The example illustrates a flowchart of a selectively parameterized SSM calculation provided in an embodiment of this application, in which the calculation is performed once at each control step t. Figure 4 The forward computation flow shown enables high-frequency, constant-size hidden state updates and output generation for the robot's current state. This flow starts with the current time-stamped token as input: Initially, the process proceeds through four stages: parameterization, discretization, state recursion, and output mapping, ultimately generating output features for action decision-making. .
[0108] Specifically, it first receives the multimodal fusion input of the current control step. It is formed by fusing the current values of modal buffers such as language, vision, proprioception, and force / touch. Then, it is processed through a lightweight selective parameterization module (e.g., a small MLP or linear projection network) based on... Dynamically generate three parameters: Δ t , and ,in: Δ t It is a scalar or diagonal matrix representing the effective time step of the current step, used to control the rate of forgetting historical information; The input projection matrix is used to... Mapped to the hidden state space; The output projection matrix is used to represent the hidden states. Map back to the output space.
[0109] In the discretization stage, the generated Δ t Discretize the matrix A of the continuous system to obtain the discretized transfer matrix of the current step. (or denoted as Δ) t(It itself serves as a scaling factor). This embodiment can employ matrix exponential discretization under the zero-order preservation assumption: = exp( · A).
[0110] Where A∈R d×d This is the continuous system matrix learned and fixed during the model training phase, with dimension d being the size of the hidden state (constantly unchanged). This step transforms the continuous-time system into a form suitable for discrete control step updates.
[0111] During the hidden state recursion phase, the hidden state of the previous control step is... (Constant size, dimension d) and the current input Substitute into the recursive formula:
[0112] in, It represents the transmission and forgetting of historical information. This represents the injection and update of the current input. This recursive process involves only matrix multiplication and addition, with a computational complexity of O(d). 2 The frequency does not increase with the length of the task, making it ideal for high-frequency control requirements of robots above 100Hz. (Updated) This is the constant-size hidden state of the current step, which carries the compressed causal information of the entire history from the start of the task to the current step.
[0113] During the output mapping phase, based on the hidden state of the current control step... and current input Generate the final output features The process consists of two components: The weight of historical memory: , from the output projection matrix For hidden states Perform linear transformations to extract long-term memory information; Current observed components: D ,in, D These are pre-learned fixed jump parameters used to directly retain detailed information about the current input.
[0114] Adding the two together gives the final output: , It can be directly sent to the downstream decoupled action head (such as MLP) to generate the action command for the current step. .
[0115] Based on the above embodiments, as an optional embodiment, to address the problem of outdated observations caused by inconsistent refresh frequencies of various modal sensors, this application introduces an adaptive weighting mechanism based on staleness in the above multimodal fusion and state update process. This mechanism can effectively suppress the interference of outdated information introduced by slow modalities (such as vision) that have not been refreshed for a long time on the current decision, while ensuring the real-time advantage of fast modalities (such as proprioception and force perception). Specifically, in this embodiment, the staleness of the latent features corresponding to each modality is recorded. The staleness is set to zero when the sensor data of the corresponding modality is refreshed, and is incremented by 1 after each control step.
[0116] Latent feature buffer Z for each mode m m Maintain an obsolescence counter. The update rules for this counter are as follows: When mode m is successfully refreshed and generates new latent features in the current control step Reset to 0; When mode m is not refreshed in the current control step (i.e., retains the old value from the previous control step), The value automatically increments by 1 after each control step.
[0117] thus, The numerical value intuitively reflects the time span between the latent features of the modality and the last actual sampling. For example, if the vision sensor refreshes every 100 control steps, then in the control steps that have not been refreshed, its staleness... It will gradually accumulate from 0 to 99, until it is reset to zero at the next refresh.
[0118] The process of fusing the current latent features stored in the latent feature buffer corresponding to each modality to obtain the token for the current control step includes: For each mode, the corresponding aging weight is determined based on the age of the mode. Based on the aging weights corresponding to each modality, the current latent features of each modality are weighted and fused to obtain the token for the current control step; wherein, the aging weights are negatively correlated with the obsolescence.
[0119] In this embodiment, state representative terms are constructed in each control step. Previously, the system first determined the obsolescence of each modality. Calculate the corresponding aging weight In some embodiments, the aging weight is calculated using an exponential decay function:
[0120] in, >0 is a learnable or preset aging factor associated with mode m, used to control the rate at which the modal information decays over time.
[0121] This embodiment assigns greater weight to low-frequency, high-delay modalities such as vision. This causes its weight to decrease rapidly with age; while high-frequency, low-delay modes such as proprioception and force perception are assigned smaller weights. This allows it to maintain a high weight over a longer period of time.
[0122] After obtaining the aging weights for each modality, this embodiment no longer directly concatenates the original latent features, but first applies aging weights to the latent features of each modality: =
[0123] in, This refers to the current latent features stored in the buffer. Subsequently, the weighted modal features are... } m∈M The data is concatenated along the feature dimensions and then subjected to linear projection and layer normalization to obtain the current control step. : =LN( concat[ , ,…]+ ) in, It is a learnable bias vector.
[0124] This embodiment can dynamically adjust the level of trust in different modal information: when the vision has just been refreshed, its weight is close to 1, playing a dominant role in decision-making; when the vision has not been updated for a long time, its weight decays exponentially, and the system automatically reduces its dependence on it, instead relying more on high-frequency ontology and force information for short-term closed-loop control. This adaptive weighting strategy not only avoids illusory decisions caused by outdated visual observations, but also makes the model more robust to anomalies such as network latency and sensor failure.
[0125] It is worth noting that the staleness weighting mechanism provided in this embodiment is applied to the token construction stage and is located upstream of the SSM recursion. Therefore, it will not affect the constant size characteristic of the hidden state inside the SSM and the linear recursion efficiency. It is a lightweight and efficient multimodal fusion optimization method.
[0126] Please see Figure 5 The figure illustrates, as an example, a schematic diagram of the multi-frequency asynchronous buffer timing refresh process provided in an embodiment of this application. In this embodiment, at the beginning of each control step, the system first polls the data readiness status of each modal sensor. For low-frequency modalities such as vision, the system checks whether the current step has reached the preset encoding interval (e.g., once every 4 control steps); for high-frequency modalities such as proprioception and force / touch, the system checks for hardware interrupts or whether new data has arrived in the data buffer.
[0127] Based on the readiness status of each modality, an independent buffer update strategy is executed, and an obsolete counter corresponding to each modality is maintained synchronously: If the modal data is ready, for example, if vision has reached step 4, or if there is new data at each step of proprioception / force perception, the corresponding encoder is invoked to process the new data, and the generated latent features are stored in the independent buffer of that modality. At the same time, the staleness counter of that modality is reset to zero, marking the data as fresh.
[0128] If the modal data is not ready, for example, if vision is in steps 1, 2, or 3, the system does not re-encode it, but keeps the value in the modality buffer unchanged. At the same time, the staleness counter for that modality is automatically incremented by 1, recording the time span since the last actual sampling of the data.
[0129] In this way, the system ensures that the update delay of the slow modality (vision) does not block the high-frequency flow of the fast modality (proprioception / mechanics), and each buffer value is accompanied by a quantitative indicator reflecting its freshness.
[0130] Before constructing the state representative token for the current control step, the weights of each modality in the fusion process are dynamically adjusted based on its age. This is done by reading the current age value of each modality. Higher age indicates older data and lower reliability. Based on a preset aging strategy, an aging weight between 0 and 1 is calculated for each modality. Modalities like vision, which change slowly but are crucial for decision-making, have their weights decay rapidly with age; while modalities like ontology and force sensing, which require real-time responses, have their weights decay more slowly, maintaining high confidence even with minor delays.
[0131] Before feature stitching, this embodiment uses aging weights to weight the latent features in the buffer item by item. This means that when visual data is newly updated, it dominates the state token; while when the visual data becomes outdated, its influence automatically weakens, and the system relies more on high-frequency ontology and force data to maintain stable robot operation. This mechanism effectively avoids illusory decisions caused by using outdated visual information and improves the system's robustness to sensor latency or network jitter.
[0132] The weighted fusion of multimodal features is fed into a linear projection layer and normalized to generate a token for the current step. This token is then input into the SSM (Synchronous Module System), which dynamically generates recursive parameters based on the token's content to update the hidden state from the previous time step, resulting in a constant-size hidden state for the current time step, thus achieving causal compression of the entire historical information. Subsequently, combining the hidden state of the current control step and the current token, the final output features are generated through output projection and jump connections. Finally, a decoupled action head reads these output features and generates the action command for the current control step (such as joint torque or end-effector pose increment).
[0133] It's worth noting that motion commands are not generated while all modalities are refreshed simultaneously; instead, each control step is executed based on the latest token. This allows the robot to output smooth, continuous movements at the highest control frequency even if the vision is not updated temporarily.
[0134] Based on the above embodiments, as an optional embodiment, the recursive step size parameter is further adjusted according to the staleness of each mode: When the obsolescence of any mode exceeds a preset threshold, the recursive step size parameter is increased to improve the forgetting rate of the historical hidden states corresponding to that mode.
[0135] This embodiment, based on the obsolescence perception mechanism, further feeds the freshness of modal data into the internal dynamics of the SSM. Specifically, it not only uses obsolescence to adjust the weights of input features, but also uses obsolescence to dynamically control the rate at which the SSM forgets historical information, thereby achieving frequency-aware memory compression.
[0136] As mentioned earlier, the system maintains an age rating for each mode at each control step t. In this embodiment, staleness is incorporated as part of the input dependency parameter generation process, directly affecting the SSM recursive step size parameter Δ. t The value of .
[0137] Specifically, this embodiment first constructs a freshness vector reflecting the freshness of data for each modality. While oldness could be used directly... However, in a physical sense, it is more likely to be converted into a freshness score, for example, taking The inverse of the result is then normalized and mapped. Subsequently, this freshness information is combined with the current state representative lexical through a lightweight feature fusion layer (e.g., a fully connected layer or a simple linear transformation). Perform joint encoding.
[0138] The encoded result is used to generate the recursive step size parameter Δ for the current control step. tUnlike the standard SSM, the Δ in this embodiment... t No longer solely based on current observations The modulation logic is not determined by modal staleness, but rather by explicit modulation of modal age. It follows these intuitive physical constraints: When key modal data becomes outdated (e.g., vision hasn't been refreshed for a long time), it's determined that the reliability of the current observation has decreased or the environmental information hasn't been updated. In this case, the generated Δ... t It will be adjusted to a larger value. During the discretization process of SSM (i.e., the calculation...) =exp(Δ t When A), the larger Δ t This will lead to a discretized matrix The elements of the hidden state h approach zero. This means that the hidden state h... t During updates, historical information is forgotten more, or in other words, the decay rate of historical information is accelerated. This mechanism forces the model not to rely too much on outdated visual context for long-term inference, but rather to rely more on recent reliable ontological or force information for short-term responses.
[0139] When all modal data are fresh, the current observations are considered comprehensive and reliable. At this point, Δ t Keep it within the normal range or a relatively small range, so that By approximating the identity matrix, it ensures that historical hidden states can be effectively transmitted and preserved, maintaining the integrity of the entire historical causal memory.
[0140] In terms of specific algorithm implementation, this embodiment can achieve this through a linear superposition with learnable coefficients. For example, the freshness vector can be multiplied by a learnable weight matrix, and the output can be superimposed as a bias term onto the matrix x. t The generated benchmark Δ t Above. To prevent numerical instability, a Softplus or Sigmoid activation function is usually applied to the output to ensure Δ. t The positivity of.
[0141] Through the aforementioned mechanism, this application achieves adaptive adjustment of the memory time constant. Compared to traditional fixed forgetting rate models, this frequency-aware strategy allows the hidden state of the SSM to dynamically switch between trusting old memories and rapid forgetting based on the actual conditions of the sensing environment. This not only solves the problem of erroneous associations that may be introduced by memory when the visual delay is long, but also further reduces the computational interference of unnecessary historical information on current decision-making, making the model more robust and efficient in complex and changing physical environments.
[0142] To overcome the lag in instantaneous contact detection by pure vision or proprioception, this application introduces a gated injection mechanism based on high-frequency force / tactile features. This mechanism can directly embed millisecond-level physical interaction information into the token of each control step, enabling the SSM's full history memory to perceive and remember key contact events. Specifically, when the robot's multimodal sensor configuration includes force sensors (such as six-dimensional force / torque sensors) or high-density tactile arrays, after completing x in each control step... t After the initial construction (e.g., after age-weighted fusion), x t Before inputting SSM for recursion, the following force information injection process is executed: The first mean scalar is obtained by averaging the currently stored latent features in the latent feature buffer corresponding to the first sensor along the feature channel dimension. The first mean scalar is transformed by a learnable linear transformation and then activated by the Sigmoid activation function to obtain the force gating coefficients. Using the force gating coefficient as a weight, the token of the current control step is weighted and corrected by the latent features currently stored in the latent feature buffer corresponding to the first sensor, so as to obtain the token of fused force information.
[0143] This embodiment reads the latent feature buffer Z corresponding to the current control step force / tactile sensor. force The latent features are stored in the memory. Since force signals are usually high-dimensional, directly injecting the original features would introduce excessive computational overhead and may obscure key information.
[0144] Therefore, in this embodiment, Z is processed along the feature channel dimension. force Perform global average pooling to obtain the first mean. : =
[0145] in, The number of force feature channels encapsulates the overall force or tactile activation level of the current control step.
[0146] Further, this first mean Input a lightweight gated generator subnetwork, which can be a linear transform layer followed by a sigmoid activation function to dynamically generate force gating coefficients. g t : g t = (W g + bg ) Among them, W g and b g These are learnable parameters, where σ is the Sigmoid function used to constrain the gating coefficients within the (0,1) interval. t The physical significance lies in quantifying the information content or importance of the current force signal. During the free motion phase, the force signal is close to zero, and the gating coefficient g... t As the gating coefficient approaches zero, the contribution of force information to the state representation is suppressed; however, at the instant of contact or collision, the amplitude of the force signal surges, and the gating coefficient g... t The force feedback quickly approaches 1, forcibly activating the force information injection pathway. This design allows the model to learn autonomously when and how to utilize force feedback without relying on explicit contact detection algorithms.
[0147] To accurately inject force information while preserving the semantics of the current state, the system employs a cross-attention mechanism to calculate the force feature residual. Specifically, the initially constructed state represents the lexical unit x. t As a query, the force / tactile latent feature buffer Z force As keys and values, perform single-head or multi-head cross-attention operations: CrossAttn(x t Z force =Softmax( ) Among them W Q and W K d is a learnable projection matrix. k is the dimension of the key vector. The output of this cross-attention function represents the part of the force features most relevant to the current task from the perspective of the current state.
[0148] This embodiment further calculates the force characteristic residual using the following formula. : =CrossAttn(x t Z force )- x t The residual term measures the difference between the state after taking force information into account and the state based solely on the current multimodal fusion state. Essentially, it is a force-aware correction to the current state.
[0149] Finally, using the aforementioned force gate coefficient g t As a weight, the force characteristic residual Weighted summation to x t Above, the state representative word element of the fusion force information is obtained. : =xt +g t .
[0150] The step of inputting the token into a linear recursive model for single-step iteration specifically involves inputting the token containing the fusion force information into the linear recursive model for single-step iteration.
[0151] It should be understood that the state representative term of this fusion force information The event will be sent to the SSM for single-step recursion, allowing the contact event to be immediately written into the constant-size hidden state h of the SSM. t In the process, it participates in the subsequent whole-history causal memory and action generation process.
[0152] Based on the above embodiments, as an optional embodiment, the multi-modal sensors include a vision sensor, and the refresh rate of the vision sensor is lower than the maximum value of the refresh rates of all the modal sensors; In the control step between two consecutive refreshes of the visual sensor, the latent feature buffer corresponding to the visual sensor keeps the latent features obtained in the most recent refresh unchanged and does not block the single-step recursive process of the linear recursive model.
[0153] It should be noted that, due to their inherent physical characteristics and computational overhead, visual sensors (such as RGB-D cameras and stereo cameras) have an effective semantic refresh rate that is much lower than that of high-bandwidth sensors such as proprioceptive and force sensors. To address this cross-modal frequency mismatch problem, this application proposes a visual low-frequency buffer and zero-order hold mechanism, which enables visual information to be seamlessly integrated into the high-frequency control loop asynchronously without becoming a bottleneck for system performance.
[0154] Specifically, in this embodiment, an independent latent feature buffer is maintained for each modality, wherein the latent feature buffer Z corresponding to the visual sensor is... vision It has the following special refresh and access strategies: A preset visual encoding stride K is used (e.g., K=4 or K=100, the specific value depends on hardware performance and task requirements). The visual encoder is only triggered when control step t satisfies t mod K=0 and the camera has captured a new image frame. It then performs forward inference on the latest image, extracts high-level semantic features (such as object category, pose, scene layout, etc.), and writes the results to the visual buffer Z. vision .
[0155] In the vast majority of control steps (t mod K not equal to 0), the visual encoder is idle and does not perform any computation. This design significantly reduces the average computational load of the system, allowing expensive visual computing resources to be allocated to moments that require semantic understanding, rather than being consumed by high-frequency control loops idling.
[0156] In the control step where the visual encoder is not triggered, Z vision The feature values stored in the database remain unchanged from the most recent values written during the last refresh. This strategy of preserving old values is called zero-order preservation.
[0157] When the constructed state represents the lexical x t At that time, directly from Z vision The system reads the currently stored features. Regardless of whether the feature is a freshly encoded value from the current control step or an old value encoded dozens or even hundreds of control steps ago, it is treated as a currently valid visual observation and participates in the fusion along with other high-frequency modal features.
[0158] The key to this mechanism is that whether or not the visual buffer is refreshed does not constitute a blocking condition for SSM recursion and action generation. Even if the visual buffer remains at an old value for multiple consecutive control steps, SSM will still continuously advance its hidden state h based on current (potentially outdated) visual features, fresh ontology / force features, and language command features. t The update is performed, and the decoupled action head stably outputs the action command a at each control step. t .
[0159] For example, in a system with a control frequency of 100Hz and a visual coding interval K=100: In step 1, the visual encoder works to update Z. vision SSM updates h1 based on the new visual features and outputs a1; In steps 2 through 99, the visual encoder goes to sleep, Z vision Keeping the features from step 1 unchanged, SSM will still update h2,…,h at a frequency of 100Hz based on this retained value. 99 And output a2,…,a 99 ; In step 100, the visual encoder works again, updating Z. vision SSM updates h100 based on the new visual features and outputs a100.
[0160] Throughout the process, the frequency of motion generation remains at 100Hz, completely unaffected by the visual effective refresh rate of 1Hz.
[0161] Please see Figure 6 The figure exemplifies a flowchart illustrating a robot control method provided in another embodiment of this application, as shown in the figure, which includes: S201, Receive control commands in natural language The language encoder encodes the control command once to obtain the language latent features, and writes them into the language modality buffer. The language buffer remains unchanged throughout the entire task round and is only refreshed when the command changes, avoiding redundant calculations.
[0162] S202, Multimodal Asynchronous Buffer Refresh and Staleness Update In each control step, the system processes the modal data in parallel: Visual buffer: If the current step reaches the preset visual encoding interval and the camera has a new image, the visual encoder extracts high-level semantic features such as object category and pose, writes them into the visual buffer, and sets the visual staleness counter to zero; otherwise, the visual buffer retains the most recent encoding result, and the staleness counter is incremented by one.
[0163] Proprioceptive buffer: If the joint encoder has the latest data, it is written to the proprioceptive buffer after encoding, and the obsolescence counter is set to zero; otherwise, the old value is retained, and the obsolescence counter is incremented by one.
[0164] Force / Haptic Buffer: If the force sensor or haptic array has the latest data, it is smoothed and encoded into the force buffer, and the obsolescence counter is set to zero; otherwise, the old value is retained and the obsolescence counter is incremented by one.
[0165] Each modal buffer does not block the others, and the failure of any one modal to be refreshed will not affect the advancement of other modalities or the overall system.
[0166] S203, Multimodal Weighted Fusion Based on Aging Calculate the corresponding aging weight based on the current age of each mode: Low-frequency modalities such as vision have a larger aging coefficient, and their weight decreases rapidly with age; high-frequency modalities such as proprioception and force perception have a smaller aging coefficient, and their weight remains at a higher level for a longer period of time.
[0167] Subsequently, the latent features of each modality are weighted using aging weights, and then the weighted multimodal features are concatenated with the language buffer features. After linear transformation and normalization, the token of the current control step is obtained.
[0168] S204, Force-Gated Cross-Attention Injection Read the current hidden feature in the force buffer and perform the following operations: The average value along the feature channel direction is obtained as a scalar that reflects the overall force level. This scalar is then used to generate force gating coefficients with values ranging from 0 to 1 through a learnable linear transformation and a Sigmoid activation function. Using the current token as the query and the force buffer feature as the key, perform cross-attention operation to obtain the force perception feature related to the current state; The difference between the force-sensing features and the current token is calculated to obtain the force feature residual; Using the force gating coefficient as a weight, the force feature residual is superimposed onto the current token to obtain a token that integrates force information.
[0169] S205, Selective State-Space Model Parameter Generation Input the token containing the fusion force information into the selective parameterization module to dynamically generate the recursive parameters for the current control step: The recursion step size parameter is used to control the rate at which historical information is forgotten; Input projection coefficients are used to map the current input to the hidden state update space; Output projection coefficients, used to map hidden states to the action generation space.
[0170] S206, Hidden State Update and Full Historical Causal Memory The fixed continuous system matrix is discretized using the recursive step size parameter to obtain the discretized transfer matrix of the current step. Based on the hidden state of the previous control step and the current input, the current hidden state is updated by linear recursion.
[0171] S207, Constructing Output Features Based on the hidden state and current input of the current control step, generate the output features of the current control step: The hidden state is mapped by the output projection coefficients to obtain the historical memory component, which carries the long-term task context. The current input is mapped by fixed jump connection parameters to obtain the observation direct transmission component, which retains instantaneous detail information. The two components are added together to obtain the output feature.
[0172] S208, Decoupling Action Generation The motion head takes the output features as input and generates motion commands for the current control step, such as joint position and end-effector pose increment.
[0173] In this embodiment, the action generation is completely decoupled from the refresh time of each modality: even if the vision is not refreshed in this control step, the system still outputs actions based on the weighted vision hold value, fresh high-frequency body / force information, and the entire history of hidden states to achieve stable high-frequency closed-loop control.
[0174] This application provides a robot control device, such as... Figure 7 As shown, the robot's control device may include: a control command receiving module 701 and a control step execution module 702, wherein, The control command receiving module 701 is used to receive control commands; The control step execution module 702 is used to execute at least one control step according to the control command at a preset refresh frequency until a new control command is received. The preset refresh frequency is greater than or equal to the maximum value among the refresh frequencies of all modes of sensors. Each control step includes: The current latent features stored in the latent feature buffers corresponding to each mode are fused to obtain the token for the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the sensor based on its own refresh frequency for the corresponding mode. The token input has a linear recursive model and is recursively processed step by step to update the hidden state of a fixed size. The hidden state is a vector obtained by compressing all tokens from the effective date of the control command to the current control step in a time sequence. Based on the hidden state of the current control step and the token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics.
[0175] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0176] Based on the above embodiments, as an optional embodiment, the receiving of control instructions described in natural language includes: The control command is encoded to obtain the corresponding hidden feature and stored in the command modality buffer. The command modality buffer is used to store the hidden feature corresponding to the latest control command. The process of fusing the current latent features stored in the latent feature buffer corresponding to each modality to obtain the token for the current control step includes: The latent feature buffers corresponding to each mode and the currently stored latent features in the instruction mode buffer are fused to obtain the token for the current control step.
[0177] Based on the above embodiments, as an optional embodiment, the linear recursive model is a selective state-space model, and the single-step recursion includes: Based on the token of the current control step, recursive parameters for the current control step are generated. The recursive parameters include a recursive step size parameter, an input projection coefficient, and a history transfer matrix. The recursive step size parameter is used to control the forgetting rate of the historical hidden states. The input projection coefficient is a matrix that maps the token to the hidden state update space. The history transfer matrix is a fixed-dimensional continuous system matrix. The historical transfer matrix is discretized based on the recursive step size parameter to obtain the discretized transfer matrix of the current control step. The hidden state of the previous control step is linearly transformed based on the discretized transfer matrix to obtain the historical transfer components in the hidden state update. The hidden state of the current control step is obtained by adding the historical transmission component to the input component obtained by mapping the token of the current control step through the input projection coefficient.
[0178] Based on the above embodiments, as an optional embodiment, the recursive parameters further include output projection coefficients, which are input dependency matrices generated by the token of the current control step and map the hidden state to the output feature space; The step of obtaining the output features of the current control step based on the hidden state and the token of the current control step includes: The hidden state contribution component is obtained by mapping the hidden state of the current control step through the output projection coefficients. The token of the current control step is mapped through the pre-learned jump connection parameters to obtain the observation direct transmission component; The hidden state contribution component is added to the observation direct transmission component to obtain the output feature of the current control step.
[0179] Based on the above embodiments, as an optional embodiment, the device is further used for: The obsolescence of the latent feature association record for each modality is set to zero when the sensor data of the corresponding modality is refreshed, and incremented by 1 after each control step. The process of fusing the current latent features stored in the latent feature buffer corresponding to each modality to obtain the token for the current control step includes: For each mode, the corresponding aging weight is determined based on the age of the mode. Based on the aging weights corresponding to each modality, the current latent features of each modality are weighted and fused to obtain the token for the current control step; wherein, the aging weights are negatively correlated with the obsolescence.
[0180] Based on the above embodiments, as an optional embodiment, when the obsolescence of any mode exceeds a preset threshold, the recursive step size parameter is increased to improve the forgetting rate of the historical hidden states corresponding to that mode.
[0181] Based on the above embodiments, as an optional embodiment, the multiple modal sensors include a first sensor, which is a force or tactile sensor; The step of performing a single-step recursive iteration on the token input using a linear recursive model includes, prior to: The first mean value is obtained by averaging the currently stored latent features in the latent feature buffer corresponding to the first sensor along the feature channel dimension. The first mean is transformed by a learnable linear transformation and then activated by the Sigmoid activation function to obtain the force gating coefficients. Using the force gating coefficient as a weight, the token of the current control step is weighted and corrected by the hidden features currently stored in the hidden feature buffer corresponding to the first sensor, so as to obtain the token of fused force information. The step of inputting the token into a linear recursive model for single-step iteration specifically involves inputting the token containing the fusion force information into the linear recursive model for single-step iteration.
[0182] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of a robot control method. Compared with related technologies, this method achieves the following: By establishing independent latent feature buffers for sensors of each modality and maintaining the latest values with a zero-order hold mechanism, the forced alignment constraint of traditional synchronous single-frequency clocks on cross-modal sampling is broken. This allows slow modalities such as vision to be refreshed at low frequency according to semantic change rate, while fast modalities such as ontology / force can be sampled at high frequency according to physical requirements. This avoids redundant calculations in slow modalities and prevents undersampling of fast modal information. Secondly, the system control beat is set to the maximum value of the refresh frequency of all modalities. Combined with the mechanism of generating state representative words from the fusion buffer features at each step, this ensures that the linear recursive model can continuously advance at a frequency much higher than the visual refresh rate. This fundamentally removes the physical limitation that the action generation frequency is capped by the slowest modality. In actual tests, high-frequency closed-loop control at the 200Hz level can be achieved. Furthermore, a linear recursive model with a fixed-size hidden state is used to causally compress the representative lexical units of the entire historical state, achieving the memory capability of long-term non-Markov states with O(L) linear complexity. This avoids the memory explosion problem caused by the O(L²) complexity of full-history attention while retaining decision information from early key events, enabling differentiated and correct actions to be output even in scenarios with the same observation but different task histories. Finally, by completely decoupling the process of generating actions from action instructions from the modal refresh time, the system can output actions at each control step without waiting for a complete observation packet. This reduces the action latency from hundreds of milliseconds of waiting for multimodal data synchronization to microseconds of single-step recursion, significantly improving the execution stability and success rate of tasks with extremely high real-time requirements, such as contact operations and compliant control. At the same time, it achieves dynamic optimization allocation of computing power according to the modal semantic change rate.
[0183] In one alternative embodiment, an electronic device is provided, such as Figure 8 As shown, Figure 8 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0184] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0185] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, bus 4002 is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus.
[0186] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0187] The memory 4003 stores computer programs that execute embodiments of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0188] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.
[0189] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0190] The terms first, second, third, fourth, 1, 2, etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.
[0191] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0192] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A method for controlling a robot, characterized in that, The robot includes sensors with multiple modalities, and the method includes: Receive control commands; According to the control command, at least one control step is executed at a preset refresh frequency until a new control command is received. The preset refresh frequency is greater than or equal to the maximum value among the refresh frequencies of all modes of sensors. Each control step includes: The current latent features stored in the latent feature buffers corresponding to each mode are fused to obtain the token for the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the sensor based on its own refresh frequency for the corresponding mode. The token input has a linear recursive model and is recursively processed step by step. The hidden state of a fixed size is updated to obtain the hidden state of the current control step. The hidden state is a vector obtained by compressing all tokens from the effective date of the control command to the current control step in a time sequence. Based on the hidden state and token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics; The linear recursive model is a selective state-space model, and the single-step recursion includes: Based on the token of the current control step, recursive parameters for the current control step are generated. The recursive parameters include a recursive step size parameter, an input projection coefficient, and a historical transfer matrix. The recursive step size parameter is used to control the forgetting rate of the historical hidden states. The input projection coefficient is a matrix that maps the token to the hidden state update space. The historical transfer matrix is a fixed-dimensional continuous system matrix. The historical transfer matrix is discretized based on the recursive step size parameter to obtain the discretized transfer matrix of the current control step. The hidden state of the previous control step is linearly transformed based on the discretized transfer matrix to obtain the historical transfer components in the hidden state update. The hidden state of the current control step is obtained by adding the historical transmission component to the input component obtained by mapping the token of the current control step through the input projection coefficient.
2. The method according to claim 1, characterized in that, The received control command includes: The control command is encoded to obtain the corresponding hidden feature and stored in the command modality buffer. The command modality buffer is used to store the hidden feature corresponding to the latest control command. The process of fusing the current latent features stored in the latent feature buffer corresponding to each modality to obtain the token for the current control step includes: The latent feature buffers corresponding to each mode and the currently stored latent features in the instruction mode buffer are fused to obtain the token for the current control step.
3. The method according to claim 1, characterized in that, The recursive parameters also include output projection coefficients, which are input dependency matrices generated by the token of the current control step and map the hidden state to the output feature space. The step of obtaining the output features of the current control step based on the hidden state and the token includes: The hidden state of the current control step is mapped through the output projection coefficients to obtain the hidden state contribution component; The token of the current control step is mapped through the pre-learned jump connection parameters to obtain the observation direct transmission component; The hidden state contribution component is added to the observation direct transmission component to obtain the output feature of the current control step.
4. The method according to claim 2, characterized in that, Also includes: The obsolescence of the latent feature association record for each modality is set to zero when the sensor data of the corresponding modality is refreshed, and incremented by 1 after each control step. The process of fusing the current latent features stored in the latent feature buffer corresponding to each modality to obtain the token for the current control step includes: For each mode, the corresponding aging weight is determined based on the age of the mode. Based on the aging weights corresponding to each modality, the current latent features of each modality are weighted and fused to obtain the token for the current control step; wherein, the aging weights are negatively correlated with the obsolescence.
5. The method according to claim 4, characterized in that, When the obsolescence of any mode exceeds a preset threshold, the recursive step size parameter is increased to improve the forgetting rate of the historical hidden states corresponding to the mode.
6. The method according to claim 1, characterized in that, The multiple modal sensors include a first sensor, which is a force or tactile sensor; The step of performing a single-step recursive iteration on the token input using a linear recursive model, prior to which the following is also included: The first mean value is obtained by averaging the currently stored latent features in the latent feature buffer corresponding to the first sensor along the feature channel dimension. The first mean is transformed by a learnable linear transformation and then activated by the Sigmoid activation function to obtain the force gating coefficients. Using the force gating coefficient as a weight, the token of the current control step is weighted and corrected by the hidden features currently stored in the hidden feature buffer corresponding to the first sensor, so as to obtain the token of fused force information. The step of inputting the token into a linear recursive model for single-step iteration specifically involves inputting the token containing the fusion force information into the linear recursive model for single-step iteration.
7. A control device for a robot, characterized in that, The robot includes sensors with multiple modalities, and the device includes: The control command receiving module is used to receive control commands; The control step execution module is used to execute at least one control step according to the control command at a preset refresh frequency until a new control command is received. The preset refresh frequency is greater than or equal to the maximum value among the refresh frequencies of all modes of sensors. Each control step includes: The current latent features stored in the latent feature buffers corresponding to each mode are fused to obtain the token for the current control step; each latent feature buffer is used to store the latent features obtained by encoding the latest sensor data collected by the sensor based on its own refresh frequency for the corresponding mode. The token input has a linear recursive model and is recursively processed step by step. The hidden state of a fixed size is updated to obtain the hidden state of the current control step. The hidden state is a vector obtained by compressing all tokens from the effective date of the control command to the current control step in a time sequence. Based on the hidden state and token of the current control step, obtain the output characteristics of the current control step, and generate the action command of the current control step based on the output characteristics; The linear recursive model is a selective state-space model, and the single-step recursion includes: Based on the token of the current control step, recursive parameters for the current control step are generated. The recursive parameters include a recursive step size parameter, an input projection coefficient, and a historical transfer matrix. The recursive step size parameter is used to control the forgetting rate of the historical hidden states. The input projection coefficient is a matrix that maps the token to the hidden state update space. The historical transfer matrix is a fixed-dimensional continuous system matrix. The historical transfer matrix is discretized based on the recursive step size parameter to obtain the discretized transfer matrix of the current control step. The hidden state of the previous control step is linearly transformed based on the discretized transfer matrix to obtain the historical transfer components in the hidden state update. The hidden state of the current control step is obtained by adding the historical transmission component to the input component obtained by mapping the token of the current control step through the input projection coefficient.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the robot control method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the robot control method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the robot control method according to any one of claims 1-6.