Method and apparatus for generating robot action, electronic device, medium and product
Patent Information
- Application Number
- CN202611318067.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]由于视觉信息和本体感觉信息难以及时、准确地反映机器人执行操作时发生的接触变化,在插入、按压和擦拭等连续操作中容易出现动作切换错误或位置控制偏差,造成任务成功率下降和动作执行错误率上升,进而导致机器人动作生成准确率较低
[0010]根据本公开实施例的一个方面,提供了一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现上述任一种机器人动作的生成方法的步骤。
Smart Images

Figure CN122807961A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to a method, apparatus, electronic device, medium, and product for generating robot actions. Background Technology
[0002] Existing robot motion generation methods typically collect the robot's visual and proprioceptive information at the current moment, fuse the visual and proprioceptive information into operational features, use a state propagation model to update the current model state based on the operational features and the model state at the previous moment, and then generate robot motion based on the current model state.
[0003] Because visual and proprioceptive information cannot reflect contact changes that occur when the robot performs operations in a timely and accurate manner, errors in action switching or position control deviations are prone to occur in continuous operations such as insertion, pressing, and wiping, resulting in a decrease in task success rate and an increase in action execution error rate, which in turn leads to a low accuracy rate in robot action generation.
[0004] No effective solution has yet been proposed to address the technical issues, such as the low accuracy of robot motion generation in related technologies. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, medium, and product for generating robot motions, to at least address the technical problem of low accuracy in generating robot motions in related technologies.
[0006] According to one aspect of the present disclosure, a method for generating robot actions is provided, comprising: Obtain the robot's contact features at the current moment, wherein the contact features are used to characterize the contact state of the robot when performing an operation at the current moment; Based on the contact features, contact gating information for adjusting the injection amount of the contact features is determined, and the contact features are injected into the visual ontological features of the robot's state propagation model at the current moment according to the injection amount indicated by the contact gating information, so as to obtain the recursive input features of the state propagation model at the current moment. The visual ontological features are used to characterize the operation scene state and the robot's ontological motion state in the operation scene at the current moment. The recursive state at the previous time step is updated by the state propagation model based on the recursive input features to obtain the recursive state at the current time step, wherein the recursive state is the model state used by the state propagation model to retain the contact process experienced by the robot up to the corresponding time step; Generate the robot action for the current moment based on the recursive state at the current moment.
[0007] According to another aspect of the present disclosure, a robot motion generation apparatus is also provided, comprising: The first acquisition module is used to acquire the contact features of the robot at the current moment, wherein the contact features are used to characterize the contact state of the robot when performing an operation at the current moment; An injection module is used to determine contact gating information for adjusting the injection amount of the contact features based on the contact features, and to inject the contact features into the visual ontology features of the robot's state propagation model at the current moment according to the injection amount indicated by the contact gating information, so as to obtain the recursive input features of the state propagation model at the current moment. The visual ontology features are used to characterize the operation scene state and the robot's ontology motion state in the operation scene at the current moment. The update module is used to update the recursive state of the previous time step by the state propagation model according to the recursive input features to obtain the recursive state of the current time step, wherein the recursive state is the model state used by the state propagation model to retain the contact process experienced by the robot up to the corresponding time step; The generation module is used to generate the robot action at the current moment based on the recursive state at the current moment.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor is configured to execute the computer program to implement the steps of any of the above-described methods for generating robot actions.
[0009] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the above-described methods for generating robot actions.
[0010] According to one aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for generating robot actions.
[0011] The technical solution provided in this disclosure offers a method for generating robot actions. It involves acquiring contact features characterizing the contact state of the robot performing an operation at the current moment, determining contact gating information to adjust the injection amount of these features based on the contact features, and injecting the contact features into the visual ontology features of the state propagation model at the current moment according to the injection amount indicated by the contact gating information, thus obtaining recursive input features. The state propagation model updates the recursive state of the previous moment based on the recursive input features, obtaining a recursive state that retains the contact process experienced by the robot up to the current moment, and generating robot actions based on the recursive state at the current moment. Therefore, by combining the contact state at the current moment with the contact process experienced up to the current moment when generating robot actions, this method can solve the technical problem of low accuracy in robot action generation in related technologies, achieving the technical effect of improving the accuracy of robot action generation. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments of this disclosure will be briefly introduced below.
[0013] Figure 1 This is a schematic diagram of an application environment for an optional robot motion generation method according to an embodiment of the present disclosure; Figure 2 This is a flowchart illustrating an optional method for generating robot actions according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of an optional dual-path feature injection process according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of an optional contact feature-gated injection position according to an embodiment of the present disclosure; Figure 5 A schematic diagram illustrating the implementation steps of a method for generating robot actions according to an embodiment of this disclosure; Figure 6 A schematic diagram of a robot motion generation device provided in an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of an optional electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0014] The embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this disclosure, and do not constitute a limitation on the technical solutions of the embodiments of this disclosure.
[0015] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this disclosure mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element are connected through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term, for example, “A and / or B” or “A, B” indicates implementation as “A,” or implementation as “B,” or implementation as “A and B.”
[0016] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings.
[0017] First, the technical terms used in this disclosure will be introduced and explained: Vision-Language-Action Model (VLA): This refers to a model that generates robot actions based on visual observations and verbal commands.
[0018] Proprioception: refers to a robot's ability to perceive its own posture and motion state based on internal measurements such as joint position and joint velocity.
[0019] Selective State Space Model (SSM) is a sequential model that recursively updates the current hidden state based on the current input and the hidden state of the previous time step, allowing the model parameters to change with the current input.
[0020] Hidden State: refers to the internal state of a sequence model used to retain information from previous inputs that is relevant to subsequent calculations during the recursive process. It is usually represented by a fixed-dimensional vector.
[0021] Input token: refers to the feature representation input into the model and corresponding to the processing time or data unit.
[0022] To address at least one of the aforementioned technical problems or areas requiring improvement in related technologies, this disclosure proposes a method for generating robot actions. This method acquires contact features characterizing the contact state of the robot performing an operation at the current moment. Based on these contact features, it determines contact gating information to adjust the injection amount of the contact features. Following the injection amount indicated by the contact gating information, it injects the contact features into the visual ontology features of the state propagation model at the current moment, obtaining recursive input features. The state propagation model updates the recursive state of the previous moment based on the recursive input features, obtaining a recursive state that retains the contact process experienced by the robot up to the current moment. Finally, it generates robot actions based on the recursive state at the current moment. Therefore, by combining the contact state at the current moment with the contact process experienced up to the current moment when generating robot actions, this method solves the technical problem of low accuracy in robot action generation in related technologies, achieving the technical effect of improving the accuracy of robot action generation.
[0023] The following description of several exemplary embodiments illustrates the technical solutions of this disclosure and the technical effects produced by these solutions. It should be noted that the following embodiments can be referenced, learned from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0024] According to one aspect of the embodiments of this disclosure, a method for generating robot actions is provided. Optionally, in this embodiment, Figure 1 This is a schematic diagram illustrating an application environment of an optional robot motion generation method according to an embodiment of the present disclosure. The robot motion generation method described above can be applied to, but is not limited to, [examples of applications]. Figure 1 The hardware environment shown includes terminal device 102 and server 104. Server 104 can be connected to terminal device 102 via a network and can be used to provide services (e.g., application 107, etc.) to terminal device 102 or clients installed on terminal device 102. Database 105 can be set up on or independently of server 104 to provide data storage services for server 104.
[0025] The aforementioned network may include, but is not limited to, at least one of the following: wired network and wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network (WAN), metropolitan area network (MAN), and local area network (LAN). The aforementioned wireless network may include, but is not limited to, at least one of the following: Wireless Fidelity (WIFI) and Bluetooth. Terminal device 102 may be, but is not limited to, a personal computer (PC), mobile phone, tablet computer, etc. Server 104 may be, but is not limited to, a cloud server, server cluster, or other server types.
[0026] The robot motion generation method of this disclosure embodiment can be executed by server 104, terminal device 102, or jointly by server 104 and terminal device 102. Alternatively, the robot motion generation method of this disclosure embodiment can be executed by a client installed on the terminal device 102.
[0027] Taking the method for generating robot actions in this embodiment, executed by terminal device 102 (server 104), as an example, Figure 2 This is a flowchart illustrating an optional method for generating robot actions according to an embodiment of this disclosure, as shown below. Figure 2 As shown, the process of this method may include the following steps: Step S11: Obtain the contact features of the robot at the current moment, wherein the contact features are used to characterize the contact state of the robot when performing the operation at the current moment; Step S12: Determine contact gating information for adjusting the injection amount of the contact feature based on the contact feature, and inject the contact feature into the visual ontology feature of the robot's state propagation model at the current moment according to the injection amount indicated by the contact gating information, so as to obtain the recursive input feature of the state propagation model at the current moment. The visual ontology feature is used to characterize the operation scene state and the robot's ontology motion state in the operation scene at the current moment. Step S13: The state propagation model updates the recursive state of the previous time step according to the recursive input features to obtain the recursive state of the current time step, wherein the recursive state is the model state used by the state propagation model to retain the contact process experienced by the robot up to the corresponding time step. Step S14: Generate the robot action at the current moment based on the recursive state at the current moment.
[0028] Optionally, in this embodiment, Figure 3 This is a schematic diagram of an optional dual-path feature injection process according to an embodiment of this disclosure. Figure 3 As shown, in this embodiment, the control time corresponding to one recursive update of the state propagation model is taken as the current time. Recursive input features are formed at each control time, and the recursive state of the state propagation model is updated accordingly to generate the action to be performed by the robot at the current time.
[0029] In step S11, the robot's contact features at the current moment are acquired. Contact features can be generated based on one or more of the force, torque, or tactile information collected during robot operations, and their data form can be a feature vector or a sequence of feature vectors. The contact state corresponding to the contact features refers to the state of the contact action at the current moment when the robot comes into contact with the object being operated on. For example, the contact state can reflect whether the robot has formed contact with the object, the trend of the contact action, and the duration of the contact action. If the sampling frequency of contact information is higher than the robot's control frequency, the contact features at the current moment can be generated based on a segment of contact information corresponding to the current control moment, allowing contact changes formed at multiple contact sampling moments to participate in the action generation at the current control moment.
[0030] In step S12, contact gating information is determined based on the contact features, and the contact features are injected into the visual ontology features of the state propagation model at the current moment according to the injection amount indicated by the contact gating information. The visual ontology features can be formed by the robot's visual features and ontology features at the current moment, where the visual features reflect the current operating scene state of the robot, and the ontology features reflect the robot's posture or movement. The visual ontology features thus constitute the basis of the visual and ontology information used by the state propagation model at the current moment.
[0031] Contact gating information is determined based on the contact features at the current moment and is used to adjust the degree of influence of the contact features on the recursive input features. The larger the injection amount indicated by the contact gating information, the greater the proportion of the contact injection component formed by the contact features in the recursive input features; the smaller the injection amount, the greater the proportion of the visual ontological features in the recursive input features. Therefore, contact features do not participate in the state update at each control moment in a fixed proportion, but rather the corresponding injection degree is determined according to the current contact action state.
[0032] Figure 3 The current word in the model corresponds to the basic representation used when constructing the recursive input features at the current time step. Let the state propagation model be in the... The current word at each feature processing position is The current lexical unit can be a basic visual ontology feature formed based on the visual features and ontology features at the current moment, or it can be an intermediate representation formed after the basic visual ontology features have passed through one or more feature processing blocks in the state propagation model.
[0033] Figure 3Path A shown is the visual short-term memory path, used to inject visual short-term memory features acquired by the robot before the current moment into the current word. Visual short-term memory features can be used to characterize the state of the operational scene observed by the robot at multiple consecutive visual observation moments before the current moment, as well as the changes in the state of the operational scene over time.
[0034] Let the characteristics of visual short-term memory be... The current word element, after layer normalization, can be used as the source of query features, and visual short-term memory features can be used as the source of key and value features. Visual memory content related to the current word element can be extracted from the visual short-term memory features through cross-attention. The update of the current word element by path A can be represented as:
[0035] in, Presentation layer normalization processing; This indicates cross-attention processing, where the first input to cross-attention processing is the source of the query features, and the second input is the source of the key features and value features. This represents the learnable global gate parameter corresponding to path A; This indicates the amount of visual short-term memory injected as indicated by visual gating information; This represents the updated lexical unit obtained after injecting visual short-term memory into the current lexical unit. The updated lexical unit can serve as the visual ontological feature at the current moment, or as an intermediate representation in forming the visual ontological feature at the current moment.
[0036] Learnable global gate parameters It can be initialized to zero. At the start of model training, The value of is zero or close to zero, ensuring that the initial injection amount of visual short-term memory (STM) into the current word is zero or within a small range. As the model training progresses, the learnable global gate parameters are gradually updated, allowing path A to determine the amount of visual STM injection appropriate to the current operational state based on the training results. Therefore, visual STM related to the current operation can be gradually introduced while preserving the existing visual ontology representation of the current word.
[0037] Figure 3 Path B shown is the contact feature injection path. Figure 3 The "force / high frequency" in the text corresponds to the contact characteristics formed based on force, torque, or tactile information acquired at high frequencies. Let the contact characteristics at the current moment be... ,in, This is used to represent contact perception modalities such as force, torque, or touch. The current lexical unit after layer normalization can be used as the source of query features, and the contact features can be used as the source of key and value features. Contact fusion features that match the current lexical unit can be obtained through cross attention.
[0038] Figure 3 The purely additive orthogonal increment shown can be expressed as:
[0039] in, This represents the contact increment characteristic at the current control moment; The contact fusion feature represents the output of cross-attention; This indicates cross attention; Representation layer normalization; Indicates the first Visual ontology features at the injection locations corresponding to each state propagation block; This represents the contact feature at the current moment. The contact increment feature is the difference between the contact fusion feature and the visual ontology feature at the injection location, used to represent the incremental content provided by the contact feature relative to the current visual ontology feature.
[0040] Optionally, path A is the visual memory injection path, where visual memory is used to supplement the operational scene observed by the robot before the current moment, and the zero-initialization global gate is used to adjust the amount of visual memory injected; the visual memory component formed by path A, after being superimposed with the current word, can obtain or update the visual ontology features at the current moment. Path B is the contact feature injection path. Figure 3 The “Force / High Frequency” shown corresponds to the contact features formed based on high-frequency contact information such as force, torque, or tactile sensation; the “Pure Additive Orthogonal Increment” corresponds to the contact increment formed based on the contact features; and the “Input Related Gate” corresponds to the contact gating information determined based on the contact features.
[0041] Figure 3 The orthogonal superposition in this model is specifically manifested in that the visual memory injection path and the contact feature injection path each form a corresponding incremental component, which is then additively superimposed onto the current word unit. Specifically, the current word unit and the visual memory increment form the visual ontology feature, while the contact increment, after being adjusted by contact gating information, forms the contact injection component. The contact injection component is then superimposed onto the visual ontology feature, and the resulting superposition is the recursive input feature of the state propagation model at the current moment. Incremental injection of contact features allows for the addition of the current contact interaction state to the operational scenario state and ontology motion state reflected by the visual ontology feature.
[0042] In step S13, the state propagation model updates the recursive state from the previous time step based on the recursive input characteristics at the current time step. As an optional implementation, the state propagation model can employ a Selective State Space Model (SSM). The recursive state, in data form, can be an internal state of the model with a fixed dimension, and can be, but is not limited to, a hidden state in the selective state space model. Its stored content corresponds to the contact process experienced by the robot up to the corresponding time step, including contact events that occurred at previous control times and the changes in contact interaction as the operation progresses.
[0043] Specifically, the state propagation model propagates the contact processes preserved in the previous recursive state to the current time step, and writes the current recursive input features into the propagated model state, forming the current recursive state. Since the contact features of the current time step have already been added to the recursive input features, the current recursive state not only preserves the previous contact processes but also incorporates the current operation scenario state, the robot's motion state, and the contact interaction state. After continuously executing the above updates according to the control time step, the contact interactions at each control time step can be propagated sequentially along the recursive state, thereby preserving the contact processes experienced by the robot up to the current time step in a fixed-dimensional model state.
[0044] For example, during a robot's pressing operation, contact features can reflect the gradual increase in contact between the robot's end effector and the object being manipulated. The state propagation model writes the corresponding contact features into the recursive state at each control moment, ensuring that the current recursive state retains the changes in contact action from before. Even if current visual observation cannot directly reflect whether the target contact state has been reached after the robot's end effector has completed pressing, the state propagation model can still use the contact process retained in the recursive state to generate subsequent actions.
[0045] In step S14, the robot action for the current moment is generated based on the recursive state at the current moment. The state propagation model can extract historical operation features related to action generation from the recursive state at the current moment, and combine them with the recursive input features to form the model output features for the current moment. Then, the action head maps the model output features into robot control quantities used to drive the movement of the robot actuator. The robot action generated in this way is related to the current operation scene state, the body motion state, and the contact interaction state, and is also affected by the previous contact process.
[0046] By determining contact gating information based on contact characteristics and writing the contact characteristics, adjusted by the injection amount, into the recursive input characteristics of the state propagation model, the state propagation model can retain the contact processes already experienced by the robot during continuous updates of the recursive state. Subsequently, robot actions are generated based on the current recursive state, allowing previously occurring contact events to participate in the action determination at subsequent control moments, thereby improving the accuracy of robot action generation.
[0047] As an optional approach, obtaining the robot's contact features at the current moment includes: Within a sampling period ending at the current time, the contact effects experienced by the robot during its operation are sampled multiple times to obtain a contact information sequence. The contact information sequence includes multiple contact information items arranged according to the sampling time, and each contact information item is a measurement result of the contact effect at the corresponding sampling time. The contact information in the contact information sequence is smoothed and filtered according to the sampling order to obtain a smoothed contact information sequence. Extract the contact features at the current moment from the contact smoothing information sequence.
[0048] Optionally, in this embodiment, obtaining the robot's contact features at the current moment may include: sampling the contact effects experienced by the robot during operation multiple times within a sampling period ending at the current moment to obtain a contact information sequence; smoothing and filtering each contact information in the contact information sequence according to the sampling order to obtain a contact smoothing information sequence; and extracting the contact features at the current moment from the contact smoothing information sequence.
[0049] The current moment can be the control moment corresponding to a single robot action. The sampling period is a time window that ends at the current control moment and covers a period preceding it. The sampling frequency of contact information such as force, torque, or tactile sensation can be higher than the robot's action generation frequency; therefore, a single sampling period can include multiple sampling moments. Arranging the contact information obtained at each sampling moment in chronological order forms a contact information sequence corresponding to the current control moment. As an example, contact information can be acquired using a wrist six-dimensional force / torque sensor at a sampling frequency of 200Hz; contact information can also originate from tactile arrays, joint current sensors, wrist torque sensors, or fingertip pressure sensors.
[0050] Each contact information segment represents the measurement result of the contact action at the corresponding sampling time. Depending on the source of the contact information, each segment may include data from one or more measurement channels. For example, the contact information output by a six-dimensional force / torque sensor on the wrist may include force measurement results in three directions and torque measurement results in three directions, while the contact information output by a haptic array may include pressure measurement results from multiple haptic sensing locations. Each contact information segment in the contact information sequence uses the same data structure and corresponds to a sampling time, thereby preserving the temporal relationship between the contact measurement results.
[0051] Contact information may contain short-term fluctuations caused by sensor noise or sampling jitter. When smoothing the contact information sequence according to the sampling order, contact smoothing information corresponding to each contact piece of information can be generated sequentially, and these smoothing pieces of information are arranged according to the original sampling order to obtain a contact smoothing information sequence. Therefore, the number of contact smoothing pieces of information in the contact smoothing information sequence is consistent with the number of contact pieces of information in the contact information sequence, and the same correspondence between sampling times is maintained. Each contact smoothing piece of information represents the result obtained after smoothing and filtering the contact action measurement result at the corresponding sampling time.
[0052] Smoothing filtering can employ exponential moving average (EMA) filtering. EMA filtering generates contact smoothing information for the current sampling moment by combining contact information from the current sampling moment with contact smoothing information from the previous sampling moment, according to the sampling order, ensuring continuous variation in contact measurement results between adjacent sampling moments. Alternatively, smoothing filtering methods that reduce sensor noise and sampling jitter based on the temporal relationship between contact information can also be used.
[0053] The contact smoothing information sequence can be written to a short-time buffer. The short-time buffer is a storage area used to store contact smoothing information within the current sampling period in sampling order; the data stored in it constitutes the contact smoothing information sequence. The contact smoothing information sequence describes the data content arranged chronologically, and the short-time buffer stores this data content and provides input to the subsequent feature extraction process.
[0054] When extracting contact features from a contact smoothing information sequence, the temporal changes of contact action within the sampling period can be extracted based on each contact smoothing information and its sampling order. As an optional implementation, the contact smoothing information sequence can be input into a gated recurrent unit (GRU) according to the sampling order. The GRU then performs recursive calculations on each contact smoothing information, converting the contact smoothing information sequence into contact features of fixed dimensions. Alternatively, the GRU can be replaced with other temporal feature extraction models capable of extracting temporal changes from contact smoothing information arranged in time.
[0055] The resulting contact features can reflect the changes in contact action between adjacent sampling moments and the persistence of contact across multiple consecutive sampling moments. For example, contact features can reflect the impact changes that occur when contact occurs or the plateau changes that occur during sustained pressure. The extracted contact features are used as the contact features at the current moment, providing the input features for subsequent determination of contact gating information and construction of the recursive input features for the state propagation model at the current moment.
[0056] By collecting multiple contact information within a sampling period ending at the current time, the high-frequency contact measurement results can be correlated with the robot's current control time. By performing smoothing filtering according to the sampling order, the impact of sensor noise and sampling jitter on the contact measurement results can be reduced. By extracting the temporal changes of contact action from the contact smoothing information sequence, the contact characteristics at the current time can simultaneously reflect the instantaneous changes and duration of contact action, thereby improving the accuracy and stability of contact characteristics.
[0057] As an optional approach, the step of smoothing and filtering each contact information in the contact information sequence according to the sampling order to obtain a contact smoothing information sequence includes: For the t-th contact information sampled at the t-th sampling time among the N contact information information included in the contact information sequence, the t-th contact information is multiplied by the weight parameter to obtain the current weighted component at the t-th sampling time. The (t-1)-th contact smoothing information at the (t-1)-th sampling time is multiplied by the smoothing coefficient to obtain the historical weighted component at the t-th sampling time. Here, t is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When t is 1, the (t-1)-th contact smoothing information is the preset contact smoothing information. The weight parameter is obtained by subtracting 1 from the smoothing coefficient. The smoothing coefficient is greater than 0 and less than 1. The current weighted component at the t-th sampling time and the historical weighted component at the t-th sampling time are added together to obtain the t-th contact smoothing information of the t-th contact information collected at the t-th sampling time. The contact smoothing information sequence is obtained by arranging the contact smoothing information determined for each of the N contact information in the sampling order.
[0058] Optionally, in this embodiment, an exponential moving average filter can be used to recursively smooth each contact information in the contact information sequence according to the sampling order. Suppose the contact information sequence includes N contact information items arranged according to the sampling time. Starting from the first contact information item, the filter is applied sequentially based on the t-th contact information item and the t-th contact information item. One contact smoothing information determines the t-th contact smoothing information, which can be calculated using the following formula:
[0059] in, This represents the contact smoothing information at the t-th sampling time. This represents the t-th contact information obtained at the t-th sampling time; Indicates the t-th The t-th time of a sampling time One contact smoothing message; Represents the smoothing coefficient, and Greater than 0 and less than 1; This represents the weight parameter; t is an integer greater than or equal to 1 and less than or equal to N, where N is an integer greater than 1.
[0060] In the above calculation process, and Multiply to obtain the current weighted components. ;Will and Multiply by this to obtain the historical weighted component. Add the current weighted component to the historical weighted component; the result is... Therefore, the t-th contact smoothing information simultaneously retains the smoothing changes formed by the previous contact measurement results and incorporates the contact measurement results at the t-th sampling time.
[0061] When t takes the value 1, the formula is as follows: Corresponding to preset contact smoothing information The preset contact smoothing information is determined before the smoothing filtering of the contact information sequence begins. Its data structure and dimensions are consistent with the contact information and can be set according to the initial reference measurement results of the contact sensor. Therefore, the first contact smoothing information can be calculated according to the same recursive relationship, and then the first contact smoothing information can be used to calculate the second contact smoothing information, and so on, until the Nth contact smoothing information is obtained.
[0062] When each contact information includes multiple measurement channels, the aforementioned multiplication and addition operations can be performed on the measurement results in the same measurement channel. For example, for contact information including force measurement results in three directions and torque measurement results in three directions, the contact smoothing information corresponding to the six measurement channels can be calculated separately, and then combined according to the arrangement relationship of each measurement channel in the contact information to form the t-th contact smoothing information. Therefore, each contact smoothing information has the same data structure and dimension as the corresponding contact information.
[0063] After performing N recursive calculations according to the sampling order, the first to Nth contact smoothing information are arranged according to their corresponding sampling times to obtain a contact smoothing information sequence. The number of contact smoothing information in the contact smoothing information sequence is consistent with the number of contact information in the contact information sequence, and the correspondence between each contact smoothing information and its corresponding sampling time is maintained, so as to extract the changing characteristics of the contact action according to the chronological relationship.
[0064] Smoothing coefficient This is used to adjust the relative weights of historical and current contact measurement results in the smoothed result. Increasing the smoothing coefficient increases the weight of the historical weighted component in the t-th contact smoothing information, making the contact smoothing information change more smoothly with each sampling time. Conversely, decreasing the smoothing coefficient increases the weight of the current weighted component, making the contact smoothing information more sensitive to changes in the current contact measurement result. As an example, the smoothing coefficient can be set to 0.9, giving the historical weighted component a weight of 0.9 and the current weighted component a weight of 0.1.
[0065] By weighting and combining the contact smoothing information from the previous sampling moment with the contact information from the current sampling moment according to the sampling order, the continuity of the contact action between adjacent sampling moments can be utilized to reduce short-term fluctuations caused by sensor noise and sampling jitter, while preserving the changes in the contact action over time. The resulting contact smoothing information sequence can provide a stable temporal input for subsequent extraction of the instantaneous changes and duration of the contact action.
[0066] As an optional approach, extracting the contact features at the current moment from the contact smoothing information sequence includes: The contact timing states in the gated loop unit are recursively updated according to the contact smoothing information sequence to obtain the first to Nth contact timing states in the gated loop unit: According to the sampling order, for the i-th contact smoothing information among the N contact smoothing information included in the contact smoothing information sequence, the gated loop unit recursively updates the (i-1)-th contact timing state in the gated loop unit according to the i-th contact smoothing information to obtain the i-th contact timing state in the gated loop unit, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When i equals 1, the (i-1)-th contact timing state is the initial unit state of the gated loop unit, and the contact timing state is the unit state used by the gated loop unit to retain the changes in contact action received by the robot up to the corresponding sampling time during the sampling period; The first to Nth contact timing states in the gated loop unit are arranged into a contact timing state sequence according to the sampling order, and the contact timing state sequence is determined as the contact feature at the current time. The contact feature is used to characterize the changes in the contact action received by the robot during the sampling period between adjacent sampling times and the duration of the action during multiple consecutive sampling times.
[0067] Optionally, in this embodiment, a gated recurrent unit (GRU) can be used to extract the contact features at the current moment from the contact smoothing information sequence. The contact smoothing information sequence includes N contact smoothing information items arranged in the sampling order. The gated recurrent unit receives each contact smoothing information item sequentially according to the sampling order, and updates the contact timing state once after each reception of contact smoothing information item.
[0068] Specifically, for the i-th contact smoothing information, the gated loop unit updates the i-th contact timing state based on the i-th contact smoothing information and the (i-1)-th contact timing state, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When i equals 1, the (i-1)-th contact timing state is the initial unit state of the gated loop unit. The dimension of the initial unit state is consistent with the dimension of each subsequent contact timing state, enabling the first contact smoothing information to participate in the extraction of contact features according to the same state recursion relationship.
[0069] The contact timing state is a unit state formed by the gated loop unit when processing the contact smoothing information sequence. It is used to retain the changes in contact effects experienced by the robot up to the corresponding sampling time within the sampling period. The i-th contact timing state adds the contact effects reflected by the i-th contact smoothing information to the contact effect changes retained in the (i-1)-th contact timing state. Thus, the i-th contact timing state simultaneously associates the contact effect at the corresponding sampling time with the contact effect changes before the corresponding sampling time.
[0070] After obtaining the first to Nth contact timing states sequentially, these states can be arranged according to the sampling order to form a contact timing state sequence, which is then used as the contact feature at the current moment. Each contact timing state in the sequence corresponds to a sampling moment and maintains the same temporal arrangement as the contact smoothing information sequence. The overall relationship for the gated loop unit to generate the contact feature at the current moment based on the contact smoothing information sequence can be expressed as:
[0071] in, Indicates the contact characteristics at the current moment; This represents a sequence of contact smoothing information arranged in the order of sampling. This means that the contact smoothing information sequence is retained in a short-time buffer to form the input received by the gated loop unit in the sampling order; This indicates that the gated loop unit updates the contact timing state sequentially based on each contact smoothing information and the previous contact timing state, and outputs the operations performed by the first to Nth contact timing states according to the sampling order. The short-time buffer is a storage area used to store each contact smoothing information within the current sampling period. The data stored in the short-time buffer forms a contact smoothing information sequence according to the sampling order.
[0072] When the robot comes into contact with the object being manipulated, the rapidly changing contact action between adjacent sampling moments can create an impact change at the instant of contact; when the robot continuously presses down on the object, the contact action maintained over multiple consecutive sampling moments can create a plateau change due to continuous pressing. The gated loop unit recursively updates the contact timing state along the sampling sequence, ensuring that adjacent states in the contact timing state sequence reflect the changes in contact action between adjacent sampling moments, and that multiple consecutive states reflect the continuity of contact action over multiple consecutive sampling moments.
[0073] By updating the contact timing state sequentially according to the contact smoothing information sequence, and determining the multiple contact timing states arranged in the sampling order as the contact features at the current moment, the instantaneous changes, duration, and temporal relationships of the contact action within the sampling period can be retained in the contact features. This improves the accuracy of the contact features in describing the robot's contact action state and provides stable contact features for subsequently determining contact gating information and constructing recursive input features for the state propagation model.
[0074] As an optional approach, the contact feature is a sequence of feature vectors comprising multiple feature vectors, each feature vector in the sequence comprising the same number of feature values, the feature values being used to describe the contact action state of the robot. The step of determining contact gating information for adjusting the injection amount of the contact feature based on the contact feature includes: Calculate the average value of the multiple feature vectors included in the contact feature to obtain the average contact feature; A multiplication operation is performed on the average contact feature and the gating weight to obtain a gating operation result. The gating operation result is then mapped to an injection amount greater than 0 and less than 1 to obtain the contact gating information used to indicate the injection amount. The gating weight is a parameter used to adjust the degree of influence of the average contact feature on the injection amount.
[0075] Optionally, in this embodiment, the contact features may include multiple feature vectors arranged in chronological order, each feature vector including the same number of feature values. The same feature position in each feature vector corresponds to the same feature dimension; the same feature position refers to the same feature dimension formed by the contact temporal feature extraction process, not a spatial position on the robot or the manipulated object. The average contact feature is obtained by averaging the feature values at the same feature position in the multiple feature vectors. The average contact feature has the same number of feature values as any feature vector and is used to summarize the contact interaction state reflected by the multiple feature vectors.
[0076] The averaging process described above corresponds to one implementation of contact feature pooling. Pooling refers to aggregating multiple feature vectors along the direction of feature vector arrangement while retaining the feature dimension of each feature vector. Therefore, the average contact feature can comprehensively reflect the feature response formed by the contact action within the corresponding time range and serve as the basis for determining the amount of contact feature injection.
[0077] After obtaining the average contact characteristics, the injection amount indicated by the contact gating information can be determined based on the average contact characteristics and the gating weight. The injection amount can be determined according to the following formula:
[0078] in, Indicates the injection volume indicated by the contact gating information; The average contact characteristic represents the contact features. This represents the gating weights determined during model training. The gating weights are used to determine the injection amount based on the feature values in the average contact feature. The sigmoid function maps the result of the gating weights and average contact features to an injection amount greater than 0 and less than 1. A larger injection amount results in a larger proportion of the contact increment formed by the contact features in the recursive input features; conversely, a smaller injection amount results in a smaller proportion of the contact increment formed by the contact features in the recursive input features.
[0079] In other implementations, the sigmoid function can be replaced with other gate functions that can map the gating result to a bounded range of values. The visual short-term memory path can use a hyperbolic tangent function to adjust the injection amount, or it can use other gating functions whose output is zero or close to zero in the initial state.
[0080] As an optional implementation, contact gating can be set to a near-closed state at the start of model training, and the gating weights can be adjusted during training based on the impact of contact states on robot motion generation. Contact gating information is used to map current contact features to obtain continuously varying injection amounts, and the injection intensity under different contact states is determined by training data. This data-driven approach adjusts the timing and extent of contact information's participation in recursive state updates. By determining injection amounts within a bounded range based on contact features, the influence of contact features on recursive input features can be continuously adjusted, ensuring that contact information adapted to the current contact state participates in recursive state updates, thereby improving the accuracy of robot motion generation.
[0081] As an optional approach, injecting the contact features into the visual ontological features of the robot's state propagation model at the current moment to obtain the recursive input features of the state propagation model at the current moment includes: The query features are determined based on the visual ontology features, and the key features and value features are determined based on the contact features; Based on the correlation between the query feature and multiple key sub-features in the key feature, multiple value sub-features in the value feature are weighted and fused to obtain the contact fusion feature; Subtracting the visual ontology features from the contact fusion features yields the contact increment features; The contact increment feature is multiplied by the injection amount indicated by the contact gating information to obtain the contact injection component, and the contact injection component is added to the visual ontology feature to obtain the recursive input feature.
[0082] Optionally, in this embodiment, Figure 4 This is a schematic diagram of an optional contact feature-gated injection position according to an embodiment of the present disclosure, such as... Figure 4 As shown, the initial lexical units sequentially pass through multiple state propagation blocks, and gated injection is performed after a preset state propagation block. At each injection position, the lexical units output by the state propagation block correspond to the visual ontology features at the current time. Gated injection is used to form contact increment features based on the visual ontology features and contact features, and the contact increment features are superimposed on the visual ontology features according to the injection amount indicated by the contact gating information. The lexical units formed after gated injection continue to be passed to subsequent state propagation blocks. The lexical units formed after the last injection position and subsequent state propagation blocks serve as the recursive input features of the state propagation model at the current time.
[0083] Figure 4 Taking the example of performing gated injection every four state propagation blocks, the injection interval can be set according to the structure of the state propagation model; four state propagation blocks are an example value. Gated injection is located on the lexical propagation path used to construct recursive input features; therefore, the lexical units formed after injecting contact features can be used to update the recursive state of the state propagation model.
[0084] During gated injection, visual ontology features can be layer-normalized, and query features can be determined based on the layer-normalized visual ontology features. Key features and value features are determined based on contact features. Query features reflect the feature requirements corresponding to the current operation scene state and ontology motion state. Multiple key sub-features in the key features are used to establish correlations with the query features, and multiple value sub-features in the value features carry contact content associated with their respective key features. By weighting and fusing the corresponding value sub-features according to the correlation between the query features and each key feature, contact fusion features that are adapted to the dimension of the visual ontology features can be obtained. The higher the correlation, the greater the weight of the corresponding value sub-feature in the contact fusion features.
[0085] The above process can be implemented using cross-attention. The first parameter of cross-attention serves as the query source, and the second parameter serves as the source of the key and value. When the layer-normalized visual ontology features are used as the first parameter and the contact features as the second parameter, the output of cross-attention corresponds to the contact fusion feature obtained by fusing the current visual ontology features from the contact features. The contact increment feature can be determined according to the following formula:
[0086] in, Indicates contact increment characteristics; The contact fusion feature represents the output of cross-attention; This indicates cross attention; Representation layer normalization; Indicates the first Visual ontology features at the injection locations corresponding to each state propagation block; This represents the contact feature at the current moment. The contact increment feature is the difference between the contact fusion feature and the visual ontology feature at the injection location, used to represent the incremental content provided by the contact feature relative to the current visual ontology feature.
[0087] Based on the pooling summary and gating weights of the contact features, the injection amount indicated by the contact gating information can be obtained. The product of the injection amount and the contact increment feature is taken as the contact injection component, and this component is superimposed on the visual ontology features. The term at the injection location can then be updated according to the following formula:
[0088] in, This represents the word element formed after the current gated injection is completed; Indicates visual ontological features prior to injection; Indicates the injection volume indicated by the contact gating information; Indicates the gating weight; Pooling summaries representing contact features; This represents the sigmoid function. This represents the contact increment feature. The product of the injection amount and the contact increment feature corresponds to the contact injection component, which is added to the visual ontology feature. When multiple injection positions are set, the tokens formed after the current gated injection continue to pass through subsequent state propagation blocks and are updated in the same way at subsequent injection positions; after the update of each injection position is completed, the resulting tokens form the recursive input feature.
[0089] By using visual ontology features as the query source and contact features as the key and value source, the contact fusion features can centrally reflect the contact content related to the current operating state. Subtracting visual ontology features from the contact fusion features creates a contact increment relative to the current visual ontology features. By adding the contact increment to lexical constructs according to the injection amount indicated by the contact gating information, the influence of contact features on the recursive input features can be adjusted. The resulting recursive input features are used to update the recursive state of the state propagation model, enabling contact interactions to participate in robot motion generation, thereby improving the accuracy of robot motion generation.
[0090] As an optional approach, the contact features include multiple modal contact features corresponding to various sensing modes, and the contact gating information includes multiple independent modal contact gating information corresponding to the multiple modal contact features. Each modal contact gating information is determined based on the corresponding modal contact feature. The various sensing modes include at least two of force, torque, tactile sensation, and articular current. The step of determining the key features and value features based on the contact features includes: for each modal contact feature, determining the corresponding modal key features and modal value features based on the modal contact features; The step of weightedly fusing multiple value sub-features in the value features based on the correlation between the query feature and multiple key sub-features in the key features to obtain contact fusion features includes: for each modal contact feature, weightedly fusing multiple modal value sub-features in the modal value features based on the correlation between the query feature and multiple modal key sub-features in the modal key features to obtain the modal contact fusion feature corresponding to each modal contact feature; The step of subtracting the visual ontology feature from the contact fusion feature to obtain the contact increment feature includes: for each modal contact feature, subtracting the visual ontology feature from the modal contact fusion feature corresponding to each modal contact feature to obtain the modal contact increment feature corresponding to each modal contact feature; The step of performing a multiplication operation between the contact increment feature and the injection amount indicated by the contact gating information to obtain a contact injection component, and adding the contact injection component to the visual ontology feature to obtain the recursive input feature, includes: for each modal contact feature, performing a multiplication operation between the modal contact increment feature corresponding to each modal contact feature and the injection amount indicated by the modal contact gating information of each modal contact feature to obtain a modal contact injection component corresponding to each modal contact feature; and adding multiple modal contact injection components corresponding to multiple modal contact features to the visual ontology feature to obtain the recursive input feature.
[0091] Optionally, in this embodiment, when the robot is configured with multiple sensing modalities, corresponding modal contact features can be formed based on the contact information collected by each sensing modality. Multiple sensing modalities may include at least two of force, torque, tactile sensation, and joint current. A modal contact feature refers to a contact feature corresponding to one sensing modality; for example, a modal contact feature can be formed based on force information, and another modal contact feature can be formed based on tactile information. Different sensing modalities can update their modal contact features according to the sampling frequency of their corresponding sensors, and each can save the most recently updated modal contact feature for use by the state propagation model at the current control moment. The state propagation model and the motion head advance step-by-step according to the control frequency upon which the robot's motion output is based. Different sensing modalities asynchronously update their modal contact features according to the sampling frequency of their corresponding sensors; between two adjacent feature updates, the most recently updated modal contact feature is called at the current control moment, decoupling the state propagation process from the refresh time of each sensing modality.
[0092] Each modal contact feature corresponds to an independent contact injection path. Within a contact injection path, modal key features and modal value features are determined based on the corresponding modal contact features, and modal contact gating information is determined separately based on the corresponding modal contact features. The modal contact gating information corresponds one-to-one with the modal contact features and is used to adjust the injection amount of the modal contact increment features formed by the corresponding modal contact features. Thus, different sensing modalities can adjust the degree of influence of their respective contact information on the recursive input features according to the contact interaction state they reflect.
[0093] Building upon the aforementioned determination of query features based on visual ontology features, the same query feature can be applied to each contact injection path. For each modal contact feature, based on the correlation between the query feature and the multiple modal key sub-features included in the corresponding modal key feature, the multiple modal value sub-features included in the corresponding modal value feature are weighted and fused to obtain the corresponding modal contact fusion feature. The query feature reflects the robot's current operational scenario state and ontology motion state, while each modal key sub-feature is used to establish the correlation between the contact content of the corresponding modality and the current operational state. The higher the correlation, the greater the weight of the corresponding modal value sub-feature in the modal contact fusion feature.
[0094] For each modal contact feature, the visual ontology feature is subtracted from the corresponding modal contact fusion feature to obtain the corresponding modal contact increment feature. Each modal contact increment feature is formed in the contact injection path of the corresponding perceptual modality and is used to reflect the increment of the contact content provided by the corresponding perceptual modality relative to the current visual ontology feature. By forming each modal contact increment feature separately, the increment content provided by different perceptual modalities can be kept independent of each other before injection.
[0095] Furthermore, each modal contact increment feature is multiplied by the injection amount indicated by the corresponding modal contact gating information to obtain the corresponding modal contact injection component. Each modal contact injection component represents the incremental content and degree of writing of the corresponding perceptual modality into the visual ontology feature at the current time. Multiple modal contact injection components are additively superimposed onto the visual ontology feature to obtain the recursive input feature of the state propagation model at the current time. The recursive input feature is composed of the visual ontology feature and the modal contact injection components formed independently by each perceptual modality, and is used to update the recursive state of the state propagation model.
[0096] For example, when a robot simultaneously collects force, torque, and tactile information, force modal contact features, torque modal contact features, and tactile modal contact features can be formed separately, and force modal contact gating information, torque modal contact gating information, and tactile modal contact gating information can be determined separately. Each modality undergoes feature fusion, contact increment determination, and gating adjustment to form a corresponding modal contact injection component. The three modal contact injection components are then superimposed on the visual ontological features to form recursive input features.
[0097] By setting contact injection paths and modal contact gating information for different perception modalities, the injection amount of corresponding contact information can be adjusted according to the contact interaction state reflected by each perception modality at the current moment. By additively superimposing the contact injection components of each modality onto the visual ontological features, multiple contact information can be introduced while preserving the visual ontological features, and mutual interference between different perception modalities can be reduced. The resulting recursive input features can carry the contact interaction states reflected by multiple perception modalities, thereby improving the accuracy of robot motion generation.
[0098] As an optional approach, the step of weightedly fusing multiple value sub-features in the value features based on the correlation between the query features and multiple key sub-features in the key features to obtain contact fusion features includes: The degree of correlation between the query feature and each of the key features is determined respectively, and a weighting coefficient corresponding to each of the key features is determined according to the degree of correlation, wherein the higher the degree of correlation, the larger the weighting coefficient; Each of the value sub-features is multiplied by the weighting coefficient of the corresponding key sub-feature to obtain multiple weighted value sub-features corresponding to the multiple value sub-features. The multiple weighted value sub-features are then added to obtain the contact fusion feature with the same dimension as the visual ontology feature.
[0099] Optionally, in this embodiment, the query features are used to reflect the robot's current operational scenario state and body motion state. Both key features and value features are determined based on contact features. Key features include multiple key sub-features, and value features include multiple value sub-features corresponding to the key features. The correspondence between a key feature and a value sub-feature indicates that they originate from the same feature unit in the contact features. Key features are used to determine the degree of correlation between the corresponding feature unit and the query features, and the corresponding value sub-features are used to carry the contact content reflected by the corresponding feature unit.
[0100] The correlation between a query feature and a key feature represents the degree of matching between the contact content corresponding to the key feature and the robot's current operating state. The correlation between each query feature and each key feature is determined, and a corresponding weighting coefficient is determined based on each correlation level, so that key features with higher correlation levels correspond to larger weighting coefficients. Since each weighting coefficient corresponds to a key feature and its corresponding value sub-feature, the weighting coefficient can be used to adjust the contribution of the corresponding value sub-feature in the contact fusion feature.
[0101] After determining the weighting coefficients, each sub-feature is multiplied by its corresponding weighting coefficient to obtain multiple weighted sub-features. Each weighted sub-feature retains the contact content carried by its corresponding sub-feature, and the proportion of contact content in the subsequent fusion process is adjusted according to the corresponding weighting coefficient. Adding multiple weighted sub-features yields the contact fusion feature. When determining the value feature based on the contact feature, the dimension of each sub-feature can be made the same as the dimension of the visual ontology feature. This ensures that the contact fusion feature obtained by adding multiple weighted sub-features has the same dimension as the visual ontology feature and can be used to subsequently determine the contact increment feature.
[0102] The aforementioned correlation determination and weighted fusion process can be achieved through cross-attention. Here, the query feature serves as the query source for cross-attention, while the key and value features serve as the key and value sources, respectively. Based on the correlation between the query feature and each key sub-feature, cross-attention applies corresponding weighting coefficients to the corresponding value sub-features, and then fuses the resulting multiple weighted value sub-features to output a contact-fused feature.
[0103] By adjusting the weighting coefficients of corresponding sub-features according to the correlation between query features and each key feature, contact content that is highly matched with the current operation scene state and the body's motion state can occupy a larger proportion in the contact fusion features. This improves the matching degree between the contact fusion features and the robot's current operation state, providing more accurate contact information for the subsequent formation of contact increment features and the generation of robot actions.
[0104] As an optional approach, the step of updating the recursive state of the previous time step by the state propagation model based on the recursive input features to obtain the recursive state of the current time step includes: The state propagation parameters and input writing parameters at the current time are determined based on the recursive input features. The state propagation parameters are used to determine the propagation method from the previous recursive state to the current recursive state, and the input writing parameters are used to determine the method of writing the recursive input features into the current recursive state. The recursive state of the previous time step is multiplied with the state propagation parameters to obtain the historical state component. The recursive input feature is multiplied with the input writing parameter to obtain the recursive writing component. The historical state component and the recursive writing component are added to obtain the recursive state of the current time step.
[0105] Optionally, in this embodiment, the state propagation model can be a selective state space model. The selective state space model updates the recursive state sequentially according to the control moments upon which the robot's actions are generated, and retains the contact process experienced by the robot from the start of the operation to the corresponding control moment with a constant-size recursive state. At the current moment, the state propagation model receives the recursive input features of the current moment. The recursive input features are the input features formed after the contact feature gating injection is completed, and therefore simultaneously carry the current operation scene state, the robot's motion state, and the contact interaction state.
[0106] The state propagation model determines state propagation parameters and input writing parameters based on the recursive input features at the current moment. The state propagation parameters determine how each state component in the previous recursive state propagates to the current moment, while the input writing parameters determine how each feature component in the recursive input features is written into the current recursive state. Both the state propagation parameters and the input writing parameters are related to the recursive input features at the current moment, enabling the state propagation model to adjust the propagation of historical contact processes and the writing of current contact content based on the current operating state and the current contact interaction state.
[0107] Multiplying the previous recursive state with the current state propagation parameters yields the historical state component. This historical state component is the state component formed after the previous recursive state is processed by the current state propagation parameters, and is used to extend the operation process retained up to the previous time step to the current time step. Multiplying the current recursive input features with the input writing parameters yields the recursive write component. This recursive write component is the state component formed after the recursive input features are processed by the input writing parameters, and is used to write the current operation state and contact state into the recursive state.
[0108] The current recursive state can be determined using the following formula:
[0109] in, This indicates the current recursive state. This indicates the recursive state at the previous time step; Represents the recursive input features at the current time step; This represents the state propagation parameters determined based on the recursive input characteristics at the current moment; This represents the input writing parameters determined by the recursive input characteristics at the current moment; Represents historical state components; This represents the recursive write component. The multiplication operation described above can be a matrix multiplication between the parameters in the state propagation model and the recursive state or recursive input features. After performing an addition operation between the historical state component and the recursive write component, the resulting state retains the same state dimension as the recursive state at the previous time step.
[0110] For example, during a robot's pressing operation, the previous recursive state can retain the contact process that has already occurred since the pressing began, while the current recursive input features can carry the current operation state and the contact interaction state detected at the current moment. State propagation parameters extend the contact state already formed during the pressing process to the current moment, while input writing parameters write the current contact content into the recursive state. The current recursive state is formed by adding the historical state components and the recursive writing components, thus simultaneously retaining the contact process up to the previous moment and the contact content written at the current moment, and is used to generate the robot's action at the current moment.
[0111] By determining the state propagation parameters and input writing parameters based on the recursive input characteristics, the propagation method of historical contact processes and the writing method of current contact content can be adapted to the current operating state. By adding the historical state components to the recursive writing components, the contact processes already experienced by the robot can be continuously retained and the current contact content can be written within a constant-size recursive state. Therefore, the state propagation model can use the contact events that have already occurred to generate robot actions in subsequent control moments, thereby improving the accuracy of robot action generation.
[0112] As an optional approach, generating the robot action at the current moment based on the recursive state at the current moment includes: The state propagation model determines the output projection parameters at the current moment based on the recursive input features. The recursive state at the current moment is multiplied by the output projection parameters to obtain the state output component. The recursive input features are multiplied by the direct mapping parameters of the state propagation model to obtain the direct output component. The output projection parameters are used to extract historical operation features related to the robot's actions at the current moment from the recursive state at the current moment. The direct mapping parameters are used to directly transmit the operation state and contact interaction state at the current moment reflected by the recursive input features to the output of the state propagation model. The state output component and the direct output component are added together to obtain the model output feature of the state propagation model at the current time. The model output feature is used to characterize the contact process experienced by the robot up to the current time, the operation state at the current time, and the contact action state when performing the operation at the current time. The robot's motion head maps the model output features into robot control quantities for driving the robot's actuators, thus obtaining the robot's motion at the current moment.
[0113] Optionally, in this embodiment, the state propagation model can form the model output features for the current moment based on the recursive state and recursive input features, and the robot's motion head can generate the robot's motion for the current moment based on the model output features. The recursive state retains the contact process experienced by the robot up to the current moment, and the recursive input features reflect the operation state and contact interaction state at the current moment. The model output features are formed by the state output components corresponding to the recursive state and the direct output components corresponding to the recursive input features, enabling the motion head to utilize both historical operation features and current operation features simultaneously.
[0114] The state propagation model determines the output projection parameters for the current moment based on the recursive input features, and multiplies the current recursive state with the output projection parameters to obtain the state output components. The output projection parameters are determined by the recursive input features at the current moment and are used to extract historical operation features related to the robot's action at the current moment from the recursive state. Therefore, when the recursive state retains contact processes formed at multiple moments, the output projection parameters can select state content related to the generation of the current action from the recursive state according to the current operation state and contact interaction state.
[0115] The state propagation model also multiplies the recursive input features by the direct mapping parameters to obtain the direct output components. The direct mapping parameters are used to transmit the current operational state and contact interaction state, reflected in the recursive input features, to the output of the state propagation model. The state output components are added to the direct output components to obtain the model output features at the current time step. The calculation relationship of the model output features can be expressed as:
[0116] in, This represents the output feature of the state propagation model at the current moment; This represents the output projection parameters determined by the state propagation model based on the recursive input characteristics at the current moment; This indicates the current recursive state. Represents the direct mapping parameters of the state propagation model; This represents the recursive input features at the current time step. The product of the output projection parameters and the recursive state constitutes the state output component, and the product of the direct mapping parameters and the recursive input features constitutes the direct output component. The sum of the two components constitutes the model output features.
[0117] The state output component originates from the recursive state that preserves the contact process, thus allowing the robot to pass on the contact process experienced up to the current moment to the model output features. The direct output component originates from the recursive input features at the current moment, thus allowing the current operation state and contact action state to be passed to the model output features. The model output features are formed by both components, and can simultaneously include historical operation features and current operation features related to the current action, serving as input to the action head.
[0118] The action head is used to map the model output features from the state propagation model to robot control quantities that the robot actuators can execute. The relationship between the action head generating the robot's actions at the current moment based on the model output features can be represented as:
[0119] in, This indicates the robot's current action. This represents the action head that maps the model's output features to robot control variables; This represents the output characteristics of the state propagation model at the current moment. Robot control variables are used to drive the robot's actuators; the robot control variables output by the action head constitute the robot's action at the current moment.
[0120] As an optional implementation, the robot motion generation method can be applied to tasks that require subsequent actions to be generated based on the contact process, such as shaft hole insertion, pressing followed by wiping, and wiping coverage. During the execution of the task, the robot can acquire visual observations of the operating scene through a vision sensor, obtain the robot's current joint position, joint velocity, or end-effector pose through a body measurement component, and collect contact information between the robot and the object being operated on through force sensors, torque sensors, or tactile sensors located at the robot's end effector or joint positions.
[0121] Contact sensors can continuously collect contact information at a sampling frequency higher than the robot's motion output frequency. Multiple contact information corresponding to the current control moment are smoothed and temporal feature extracted to form the contact features at the current moment. The contact features are used to reflect the changes in contact action between adjacent sampling moments and the continuity of contact action over multiple consecutive sampling moments, such as the impact changes when contact occurs, the changes in the force plateau during continuous pressing, and the continuous tactile changes generated during wiping.
[0122] The state propagation model determines contact gating information based on the contact features at the current moment, and then, according to the injection amount indicated by the contact gating information, superimposes the contact increment formed by the contact features onto the visual ontology features at the current moment to obtain the recursive input features at the current moment. The recursive input features are used to update the recursive state of the state propagation model, enabling the contact event occurring at the current moment to enter the recursive state, and propagating it backward with the state propagation model in subsequent control moments. Thus, the recursive state can preserve the contact process experienced by the robot from the start of the operation to the current moment.
[0123] In the shaft-hole insertion scenario, the robot determines the relative position between the pin and the hole based on visual observations and controls the end effector to move the pin into the hole. After the pin contacts the hole wall, the contact information collected by the wrist force / torque sensor changes. The corresponding contact features are adjusted by contact gating information and then enter the recursive input features, writing the contact process of the pin having contacted the hole wall and having begun to enter the hole into the recursive state.
[0124] As the insertion process continues, the state propagation model updates the recursive state at each control moment, retaining the previously formed contact process. Even when the current visual observation cannot fully reflect the insertion state due to occlusion by the robot's end effector, pin, or the manipulated object, the state propagation model can still determine that the robot has entered the insertion phase based on the insertion contact process retained in the recursive state. It can then generate robot actions to maintain the insertion direction, adjust the insertion posture, or control the insertion depth based on the current contact state. This allows the robot to continue completing the insertion operation even when visual information is limited.
[0125] In a scenario where pressing is followed by wiping, the robot first controls the end effector to press the object. As the pressing depth increases, the contact characteristics reflect the gradual increase in pressing force and the process of reaching a stable platform. The state propagation model writes this contact process into a recursive state, so that the recursive state can characterize the contact state that the robot has reached corresponding to the pressing completion state.
[0126] The motion head generates robot actions based on the recursive state encompassing the aforementioned contact process, switching the robot from the pressing phase to the lateral wiping phase. Upon entering the wiping phase, the state propagation model continues to update the recursive state based on the contact features formed during wiping, and the motion head generates robot actions that move along the surface of the object being operated on. Thus, the robot can determine the timing of action switching based on the already occurred pressing process, improving the continuity between the pressing action and subsequent wiping actions.
[0127] In wiping and covering scenarios, the robot can collect contact information between the end effector and the surface being wiped through tactile arrays or other contact sensors. As the robot passes through different areas, the tactile changes corresponding to each area are written into the recursive state along with the recursive input features, so that the recursive state can retain the contact process that the robot has touched and wiped the corresponding area.
[0128] When the robot subsequently encounters a similar visual scene, the state propagation model can combine the wiping process preserved in the recursive state to distinguish whether the current area has been wiped. For areas that have been wiped, the action head can generate subsequent wiping actions for other areas based on the recursive state, thereby reducing repeated wiping of the same area and improving the coverage integrity of the wiping task.
[0129] Using the above method, contact events such as contact occurrence, insertion start, pressure reaching a stable state, and the corresponding area being wiped can be written into the recursive state of the state propagation model through recursive input features, and continuously participate in the generation of robot actions across multiple control moments. Even if the current visual observation is close or some visual information is occluded, the robot can still determine the current operation stage based on the contact process retained in the recursive state, thereby improving the accuracy, continuity, and stability of robot action generation in contact-intensive operations such as shaft hole insertion, pressure switching, and wiping coverage.
[0130] As an optional approach, before injecting the contact features into the visual ontological features of the robot's state propagation model at the current moment, the method further includes: The robot's visual ontology basic features and visual short-term memory features at the current moment are obtained. The visual ontology basic features include visual features generated based on the robot's visual observation results at the current moment and ontology features generated based on the robot's ontology measurement results at the current moment. The visual short-term memory features are used to characterize the state of the operation scene observed by the robot at multiple consecutive visual observation moments before the current moment and the change of the state of the operation scene over time. Visual query features are determined based on the basic features of the visual ontology, and visual key features and visual value features are determined based on the visual short-term memory features. Based on the degree of correlation between the visual query features and multiple visual key sub-features in the visual key features, and in a manner where the higher the degree of correlation, the greater the weight of the corresponding visual value sub-feature, multiple visual value sub-features in the visual value features are weighted and fused to obtain a visual memory fusion feature with the same dimension as the basic features of the visual ontology. Obtain visual gating information for adjusting the amount of visual memory fusion feature injection; The visual memory fusion feature is multiplied by the injection amount indicated by the visual gating information to obtain the visual memory injection component; The visual memory injection component is superimposed on the visual ontology basic feature to obtain the visual ontology feature.
[0131] Optionally, in this embodiment, before injecting the contact feature into the visual ontology features of the robot's state propagation model at the current moment, visual short-term memory features can be generated based on the robot's visual observation information before the current moment, and the visual short-term memory features can be injected into the visual ontology base features at the current moment to obtain the visual ontology features.
[0132] Specifically, the basic visual ontology features and visual short-term memory features of the robot at the current moment are obtained. The basic visual ontology features include visual features generated based on the robot's visual observations at the current moment, and ontology features generated based on the robot's ontology measurements at the current moment.
[0133] The visual features are used to characterize the state of the current operating scene of the robot, such as the state of the object being operated on, the spatial relationship between the object and the robot, and the visual information in the current operating environment; the ontological features are used to characterize the robot's own motion state, such as the state of the robot's joints, the state of the end effector, or other information that can reflect the robot's motion.
[0134] The visual short-term memory features are used to characterize the state of the operational scene observed by the robot at multiple consecutive visual observation moments prior to the current moment, as well as the changes in the operational scene state over time. By acquiring the visual short-term memory features corresponding to multiple historical visual observation moments, the robot's visual representation at the current moment can include not only the current operational scene information but also the scene change information during historical operations.
[0135] After obtaining the basic features of the visual ontology and the visual short-term memory features, visual query features are determined based on the basic features of the visual ontology, and visual key features and visual value features are determined based on the visual short-term memory features.
[0136] Optionally, the basic features of the visual ontology can be used as the query source, and the visual short-term memory features can be used as the source of keys and values. The visual short-term memory features can be fused through a cross-attention mechanism to obtain visual memory fusion features.
[0137] Specifically, the visual memory fusion features can be determined in the following way:
[0138] in, The state propagation model is represented by the first... The current lexical representation at each position; This indicates the characteristics of visual short-term memory; Presentation layer normalization processing; This refers to cross-attention processing, where the first input to the cross-attention processing is the query source, and the second input is the source of key features and value features.
[0139] Specifically, the correlation between the visual query feature and multiple visual key sub-features in the visual key feature is calculated, and the weight of the corresponding visual value sub-feature is determined based on each correlation level. A higher correlation level indicates a higher degree of matching between the corresponding historical visual information and the current operation state, and the corresponding visual value sub-feature has a greater weight in the fusion process.
[0140] Through the above processing, historical visual information related to the current operating state can be extracted from the visual short-term memory features corresponding to multiple consecutive visual observation moments, and multiple historical visual information are weighted and fused to obtain visual memory fusion features.
[0141] Furthermore, visual gating information is obtained for adjusting the amount of visual memory fusion feature injection.
[0142] The visual gating information is used to adjust the degree of influence of the visual memory fusion features on the current visual ontology base features. As an optional implementation, the visual gating information can be determined by learnable scalar gate parameters, which can be zero-initialized.
[0143] Specifically, the gating process in the visual short-term memory pathway can be represented as:
[0144] in, This represents a learnable scalar gate parameter, which is used to adjust the injection level of the visual short-term memory fusion result.
[0145] In the initial stage of model training, since the learnable scalar gate parameters are initialized to zero, therefore:
[0146] At this point, the contribution of the visual short-term memory path is zero. As the model training process progresses, the learnable scalar gate parameters are gradually adjusted based on the training data, enabling the visual short-term memory path to progressively learn the appropriate level of injection for the current task.
[0147] Subsequently, the visual memory fusion feature is multiplied by the injection amount indicated by the visual gating information to obtain the visual memory injection component, and the visual memory injection component is superimposed on the visual ontology basic feature to obtain the visual ontology feature.
[0148] The visual ontology features serve as the basic input features for the current state propagation model, and are used for the gating injection of subsequent contact features. In other words, the visual ontology features not only include the robot's current visual state and ontological motion state, but also historical operational scenario change information introduced through the visual short-term memory path.
[0149] In one optional implementation, the visual short-term memory features can be generated based on visual information collected from multiple historical visual observation moments during the robot's continuous operations. For example, when the robot performs an insertion operation, the visual short-term memory features may include information on changes in the robot's end-effector position, changes in the position of the object being operated on, and changes in the scene during the insertion process; when the robot performs a wiping operation, the visual short-term memory features may include historical visual information corresponding to the areas the robot has already traversed.
[0150] By employing the above method, before injecting contact features into the state propagation model, the basic visual ontology features are first enhanced using visual short-term memory features. This allows the resulting visual ontology features to simultaneously reflect the current operational scene state, the ontology's motion state, and historical operational scene changes. Furthermore, by adjusting the injection amount of visual short-term memory features through a zero-initialization global gate, the visual short-term memory path can gradually participate in feature updates during training. This improves the state propagation model's ability to understand the scene of continuous operational processes, providing a more accurate feature foundation for subsequent generation of recursive input features based on contact features and for generating robot actions based on recursive states.
[0151] To better understand and illustrate the solutions provided by the embodiments of this disclosure and their beneficial effects, the following description, in conjunction with a specific scenario embodiment, will illustrate the solutions provided by the embodiments of this disclosure.
[0152] Figure 5 This is a schematic diagram illustrating the implementation steps of a method for generating robot actions according to an embodiment of this disclosure, as shown below. Figure 5 As shown, this embodiment uses the process of a robot performing a contact-type operation task as an example for illustration. During the operation task, the robot acquires contact information between itself and the object being operated on through a high-frequency force / tactile sensor, and processes the acquired high-frequency contact information to form a force feature that can characterize the current contact state, i.e., a contact feature. Subsequently, the force feature and visual short-term memory information are jointly input into a dual-path gating injection module to form a fused information token (word) containing visual information, robot body state information, and contact information. The force information token (word) shown in the figure is the recursive input feature. The state propagation model continuously advances the hidden state based on the fused information token and generates robot actions based on the hidden state.
[0153] Specifically, when a robot performs an operation task, contact information between the robot and the object being operated can be collected by force sensors, torque sensors, tactile sensors, etc., installed at the robot's end effector, robotic arm joints, or other locations. Among these, high-frequency force / tactile information can reflect changes in contact during the operation, such as changes in the force generated when the robot comes into contact with the object, the duration of contact, and dynamic interaction information generated during the operation.
[0154] Since high-frequency force / tactile information typically has a high sampling frequency, while robot motion generation usually operates at a lower control frequency, high-frequency force / tactile information acquired at multiple consecutive sampling moments can be processed. First, the acquired contact information is smoothed to reduce the impact of sensor noise and instantaneous fluctuations on the contact feature extraction process. Then, the smoothed contact information is stored in a short-time buffer, enabling the robot to utilize continuous contact information within the time range corresponding to the current control moment. Further, the contact information in the short-time buffer is compressed to extract force features that reflect the contact change process.
[0155] Force features are used to characterize the contact state of the robot at the current moment when it performs an operation. Compared to contact information at a single sampling point, force features, after smoothing, buffering, and GRU compression, can comprehensively reflect the changes in contact over a period of time. For example, when the robot performs a pressing operation, force features can reflect the process of contact establishment, force increase, and force stabilization; when the robot performs an insertion operation, force features can reflect the change in force after the robot's end effector comes into contact with the object being operated on.
[0156] After obtaining the force features, these features are input into the dual-path gated injection module. The dual-path gated injection module includes a contact feature injection path and a visual short-term memory path. The contact feature injection path is used to introduce force features formed based on high-frequency force / tactile information into the input features of the state propagation model; the visual short-term memory path is used to introduce visual short-term memory information formed during continuous visual observations prior to the current moment into the current feature representation.
[0157] Visual short-term memory (STM) information is used to characterize the state of the operational scene at multiple historical visual observation moments during a robot's continuous operation, as well as the changes in the operational scene state over time. For example, during continuous operations such as grasping, inserting, pressing, or wiping objects, visual STM information can include historical visual information such as changes in the state of the manipulated object, changes in the relative position between the robot and the manipulated object, and changes in the operational area.
[0158] After processing by the dual-path gating injection module, the current visual information, robot body state information, and contact information can be fused to form a fused information token. The fused information token serves as the input feature for the state propagation model to update the state at the current moment, and is used to simultaneously represent the current operating scene state, the robot's own motion state, and the current contact state.
[0159] After receiving the fusion information token, the state propagation model recursively updates the model's hidden states based on the token. The state propagation model can employ a Selective State Space Model (SSM). The SSM recursively updates a fixed-dimensional model state according to the control time, compressing the historical input causality from the start of the task to the current control time into this model state. Thus, even when current visual observations are similar but previous contact processes differ, the model can distinguish the corresponding operational states based on the recursive state and generate robot actions adapted to the contact history. The state propagation model continuously updates the hidden states according to the corresponding control time generated by the robot's actions. The hidden states are used to store the operational and contact processes experienced by the robot up to the current time, enabling the robot to utilize previously occurred contact events in subsequent control processes.
[0160] For example, when a robot performs a shaft-hole insertion task, the contact change information generated after the robot's end effector contacts the hole wall is extracted by force features and entered into the fusion information token, which is then further written into the hidden state of the state propagation model. As the control timeline continues to advance, the hidden state can retain the contact changes during the insertion process, enabling the robot to utilize the information from the previous contact process during subsequent motion generation.
[0161] After the hidden state update is completed, the updated hidden state is input into the action head, which then generates robot actions based on the hidden state. The robot actions may include robot joint motion control quantities, end effector motion control quantities, or other control parameters used to drive the robot to complete the operation task.
[0162] In this embodiment, by employing high-frequency force / tactile information acquisition, smoothing, short-term caching, feature compression, and dual-path gating injection, the contact states during robot operation are incorporated into the state propagation model. This allows the model to simultaneously utilize current visual information, body motion information, and contact interaction information for state updates. Furthermore, by continuously recursively probing the hidden states through the state propagation model, the contact processes already experienced by the robot continue to play a role in subsequent action generation, thereby improving the accuracy and continuity of action generation in contact-based tasks such as pressing, inserting, and wiping.
[0163] This disclosure provides an apparatus for generating robot motions. Figure 6 This is a schematic diagram of the structure of a robot motion generation device provided in an embodiment of the present disclosure, as shown below. Figure 6 As shown, the robot motion generation device 60 may include: The first acquisition module 601 is used to acquire the contact features of the robot at the current moment, wherein the contact features are used to characterize the contact state of the robot when performing an operation at the current moment; The injection module 602 is used to determine contact gating information for adjusting the injection amount of the contact feature based on the contact feature, and inject the contact feature into the visual ontology feature of the robot's state propagation model at the current moment according to the injection amount indicated by the contact gating information, so as to obtain the recursive input feature of the state propagation model at the current moment, wherein the visual ontology feature is used to characterize the operation scene state of the robot in the operation scene at the current moment and the ontology motion state of the robot; The update module 603 is used to update the recursive state of the previous time step by the state propagation model according to the recursive input features to obtain the recursive state of the current time step, wherein the recursive state is the model state used by the state propagation model to retain the contact process experienced by the robot up to the corresponding time step. The generation module 604 is used to generate the robot action at the current moment based on the recursive state at the current moment.
[0164] In one exemplary embodiment, the injection module includes: The first determining unit is used to determine query features based on the visual ontology features, and to determine key features and value features based on the contact features; The fusion unit is used to perform weighted fusion of multiple value sub-features in the value features based on the correlation between the query features and multiple key sub-features in the key features, to obtain contact fusion features; An incremental unit is used to subtract the visual ontological feature from the contact fusion feature to obtain the contact incremental feature; The first recursive unit is used to perform a multiplication operation between the contact increment feature and the injection amount indicated by the contact gating information to obtain the contact injection component, and add the contact injection component to the visual ontology feature to obtain the recursive input feature.
[0165] In one exemplary embodiment, the contact features include multiple modal contact features corresponding to multiple sensing modalities, and the contact gating information includes multiple independent modal contact gating information corresponding to the multiple modal contact features. Each modal contact gating information is determined based on the corresponding modal contact feature. The multiple sensing modalities include at least two of force, torque, tactile sensation, and joint current. The first determining unit is further configured to: for each modal contact feature, determine the corresponding modal bond feature and modal value feature based on the modal contact feature; The fusion unit is further configured to: for each modal contact feature, perform weighted fusion of multiple modal value sub-features in the modal value feature according to the correlation between the query feature and multiple modal key sub-features in the modal key feature, to obtain a modal contact fusion feature corresponding to each modal contact feature; The incremental unit is further configured to: for each modal contact feature, subtract the visual ontology feature from the modal contact fusion feature corresponding to each modal contact feature to obtain the modal contact incremental feature corresponding to each modal contact feature; The first recursive unit is further configured to: for each modal contact feature, perform a multiplication operation on the modal contact increment feature corresponding to each modal contact feature and the injection amount indicated by the modal contact gating information of each modal contact feature to obtain the modal contact injection component corresponding to each modal contact feature, and add the multiple modal contact injection components corresponding to multiple modal contact features to the visual ontology feature to obtain the recursive input feature.
[0166] In one exemplary embodiment, the fusion unit is further configured to: The degree of correlation between the query feature and each of the key features is determined respectively, and a weighting coefficient corresponding to each of the key features is determined according to the degree of correlation, wherein the higher the degree of correlation, the larger the weighting coefficient; Each of the value sub-features is multiplied by the weighting coefficient of the corresponding key sub-feature to obtain multiple weighted value sub-features corresponding to the multiple value sub-features. The multiple weighted value sub-features are then added to obtain the contact fusion feature with the same dimension as the visual ontology feature.
[0167] In an exemplary embodiment, the first acquisition module includes: The sampling unit is used to sample the contact action experienced by the robot when it performs an operation multiple times during a sampling period with the current time as the end time, to obtain a contact information sequence, wherein the contact information sequence includes multiple contact information arranged according to the sampling time, and each contact information is the measurement result of the contact action at the corresponding sampling time; A smoothing unit is used to smooth and filter each contact information in the contact information sequence according to the sampling order to obtain a smoothed contact information sequence. An extraction unit is used to extract the contact features at the current moment from the contact smoothing information sequence.
[0168] In one exemplary embodiment, the smoothing unit is further configured to: For the t-th contact information sampled at the t-th sampling time among the N contact information information included in the contact information sequence, the t-th contact information is multiplied by the weight parameter to obtain the current weighted component at the t-th sampling time. The (t-1)-th contact smoothing information at the (t-1)-th sampling time is multiplied by the smoothing coefficient to obtain the historical weighted component at the t-th sampling time. Here, t is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When t is 1, the (t-1)-th contact smoothing information is the preset contact smoothing information. The weight parameter is obtained by subtracting 1 from the smoothing coefficient. The smoothing coefficient is greater than 0 and less than 1. The current weighted component at the t-th sampling time and the historical weighted component at the t-th sampling time are added together to obtain the t-th contact smoothing information of the t-th contact information collected at the t-th sampling time. The contact smoothing information sequence is obtained by arranging the contact smoothing information determined for each of the N contact information in the sampling order.
[0169] In one exemplary embodiment, the extraction unit is further configured to: The contact timing states in the gated loop unit are recursively updated according to the contact smoothing information sequence to obtain the first to Nth contact timing states in the gated loop unit: According to the sampling order, for the i-th contact smoothing information among the N contact smoothing information included in the contact smoothing information sequence, the gated loop unit recursively updates the (i-1)-th contact timing state in the gated loop unit according to the i-th contact smoothing information to obtain the i-th contact timing state in the gated loop unit, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When i equals 1, the (i-1)-th contact timing state is the initial unit state of the gated loop unit, and the contact timing state is the unit state used by the gated loop unit to retain the changes in contact action received by the robot up to the corresponding sampling time during the sampling period; The first to Nth contact timing states in the gated loop unit are arranged into a contact timing state sequence according to the sampling order, and the contact timing state sequence is determined as the contact feature at the current time. The contact feature is used to characterize the changes in the contact action received by the robot during the sampling period between adjacent sampling times and the duration of the action during multiple consecutive sampling times.
[0170] In an exemplary embodiment, the contact feature is a feature vector sequence comprising multiple feature vectors, each feature vector in the feature vector sequence comprising the same number of feature values, the feature values being used to describe the contact action state of the robot, and the injection module comprising: A calculation unit is used to calculate the average value of the multiple feature vectors included in the contact feature to obtain the average contact feature; A gating unit is used to perform a multiplication operation on the average contact feature and the gating weight to obtain a gating operation result, and to map the gating operation result to an injection amount greater than 0 and less than 1 to obtain contact gating information for indicating the injection amount, wherein the gating weight is a parameter used to adjust the degree of influence of the average contact feature on the injection amount.
[0171] In one exemplary embodiment, the update module includes: The second determining unit is used to determine the state propagation parameters and input writing parameters at the current time based on the recursive input features. The state propagation parameters are used to determine the propagation method from the recursive state at the previous time to the recursive state at the current time, and the input writing parameters are used to determine the method by which the recursive input features are written into the recursive state at the current time. The second recursive unit is used to perform a multiplication operation between the recursive state at the previous time step and the state propagation parameter to obtain a historical state component, perform a multiplication operation between the recursive input feature and the input writing parameter to obtain a recursive writing component, and perform an addition operation between the historical state component and the recursive writing component to obtain the recursive state at the current time step.
[0172] In one exemplary embodiment, the generation module includes: The first computation unit is used to determine the output projection parameters at the current moment by the state propagation model based on the recursive input features, perform a multiplication operation between the recursive state at the current moment and the output projection parameters to obtain a state output component, and perform a multiplication operation between the recursive input features and the direct mapping parameters of the state propagation model to obtain a direct output component. The output projection parameters are used to extract historical operation features related to the robot's action at the current moment from the recursive state at the current moment, and the direct mapping parameters are used to directly transmit the operation state and the contact interaction state at the current moment reflected by the recursive input features to the output of the state propagation model. The second computation unit is used to perform an addition operation on the state output component and the direct output component to obtain the model output feature of the state propagation model at the current time. The model output feature is used to characterize the contact process experienced by the robot up to the current time, the operation state at the current time, and the contact action state when performing the operation at the current time. The mapping unit is used to map the model output features from the robot's action head into robot control quantities for driving the movement of the robot's actuators, thereby obtaining the robot's action at the current moment.
[0173] In one exemplary embodiment, the apparatus further includes: The second acquisition module is used to acquire the basic visual ontology features and visual short-term memory features of the robot at the current moment before injecting the contact features into the visual ontology features of the robot's state propagation model at the current moment. The basic visual ontology features include visual features generated based on the robot's visual observation results at the current moment and ontology features generated based on the robot's ontology measurement results at the current moment. The visual short-term memory features are used to characterize the state of the operation scene observed by the robot at multiple consecutive visual observation moments before the current moment and the change of the state of the operation scene over time. The determination module is used to determine visual query features based on the basic features of the visual ontology, and to determine visual key features and visual value features based on the visual short-term memory features; The fusion module is used to perform weighted fusion of multiple visual value sub-features in the visual value features according to the degree of correlation between the visual query features and multiple visual key sub-features in the visual key features, with the higher the degree of correlation, the greater the weight of the corresponding visual value sub-feature, to obtain a visual memory fusion feature with the same dimension as the basic features of the visual ontology; The third acquisition module is used to acquire visual gating information for adjusting the amount of visual memory fusion feature injection. The computation module is used to perform a multiplication operation on the visual memory fusion feature and the injection amount indicated by the visual gating information to obtain the visual memory injection component; The overlay module is used to overlay the visual memory injection component onto the visual ontology basic feature to obtain the visual ontology feature.
[0174] The apparatus of this disclosure embodiment can execute the method provided in this disclosure embodiment, and its implementation principle is similar, and it has corresponding technical effects. The actions performed by each module in the apparatus of each embodiment of this disclosure correspond to the steps in the method of each embodiment of this disclosure. For a detailed functional description of each module of the apparatus, please refer to the description in the corresponding method shown above, and it will not be repeated here.
[0175] This disclosure provides an electronic device (computer apparatus / device / system) including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method provided in any optional embodiment of this disclosure.
[0176] In one alternative embodiment, an electronic device is provided. Figure 7 The following is a schematic diagram of the structure of an optional electronic device provided in an embodiment of this disclosure, such as... Figure 7 As shown, Figure 7 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this disclosure.
[0177] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0178] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0179] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0180] The memory 4003 is used to store computer programs that execute embodiments of the present disclosure, and is controlled by the processor 4001 to execute them. The processor 4001 is used to execute the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0181] This disclosure provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0182] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0183] It should be understood that although arrows indicate various operation steps in the flowcharts of the embodiments of this disclosure, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of this disclosure, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured as required, and the embodiments of this disclosure do not limit this.
[0184] The above description is only an optional implementation method for some implementation scenarios of this disclosure. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this disclosure without departing from the technical concept of this disclosure also fall within the protection scope of the embodiments of this disclosure.
Claims
1. A method for generating robot actions, characterized in that, include: Obtain the robot's contact features at the current moment, wherein the contact features are used to characterize the contact state of the robot when performing an operation at the current moment; Based on the contact features, contact gating information for adjusting the injection amount of the contact features is determined, and the contact features are injected into the visual ontological features of the robot's state propagation model at the current moment according to the injection amount indicated by the contact gating information, so as to obtain the recursive input features of the state propagation model at the current moment. The visual ontological features are used to characterize the operation scene state and the robot's ontological motion state in the operation scene at the current moment. The recursive state at the previous time step is updated by the state propagation model based on the recursive input features to obtain the recursive state at the current time step, wherein the recursive state is the model state used by the state propagation model to retain the contact process experienced by the robot up to the corresponding time step; Generate the robot action for the current moment based on the recursive state at the current moment.
2. The method for generating robot actions according to claim 1, characterized in that, The step of injecting the contact features into the visual ontology features of the robot's state propagation model at the current moment to obtain the recursive input features of the state propagation model at the current moment includes: The query features are determined based on the visual ontology features, and the key features and value features are determined based on the contact features; Based on the correlation between the query feature and multiple key sub-features in the key feature, multiple value sub-features in the value feature are weighted and fused to obtain the contact fusion feature; Subtracting the visual ontology features from the contact fusion features yields the contact increment features; The contact increment feature is multiplied by the injection amount indicated by the contact gating information to obtain the contact injection component, and the contact injection component is added to the visual ontology feature to obtain the recursive input feature.
3. The method for generating robot actions according to claim 2, characterized in that, The contact features include multiple modal contact features corresponding to multiple sensing modalities. The contact gating information includes multiple independent modal contact gating information corresponding to the multiple modal contact features. Each modal contact gating information is determined based on the corresponding modal contact feature. The multiple sensing modalities include at least two of force, torque, tactile sensation, and joint current. The step of determining the key features and value features based on the contact features includes: for each modal contact feature, determining the corresponding modal key features and modal value features based on the modal contact features; The step of weightedly fusing multiple value sub-features in the value features based on the correlation between the query feature and multiple key sub-features in the key features to obtain contact fusion features includes: for each modal contact feature, weightedly fusing multiple modal value sub-features in the modal value features based on the correlation between the query feature and multiple modal key sub-features in the modal key features to obtain the modal contact fusion feature corresponding to each modal contact feature; The step of subtracting the visual ontology feature from the contact fusion feature to obtain the contact increment feature includes: for each modal contact feature, subtracting the visual ontology feature from the modal contact fusion feature corresponding to each modal contact feature to obtain the modal contact increment feature corresponding to each modal contact feature; The step of performing a multiplication operation between the contact increment feature and the injection amount indicated by the contact gating information to obtain a contact injection component, and adding the contact injection component to the visual ontology feature to obtain the recursive input feature, includes: for each modal contact feature, performing a multiplication operation between the modal contact increment feature corresponding to each modal contact feature and the injection amount indicated by the modal contact gating information of each modal contact feature to obtain a modal contact injection component corresponding to each modal contact feature; and adding multiple modal contact injection components corresponding to multiple modal contact features to the visual ontology feature to obtain the recursive input feature.
4. The method for generating robot actions according to claim 2, characterized in that, The step of weightedly fusing multiple value sub-features in the value features based on the correlation between the query features and multiple key sub-features in the key features to obtain contact fusion features includes: The degree of correlation between the query feature and each of the key features is determined respectively, and a weighting coefficient corresponding to each of the key features is determined according to the degree of correlation, wherein the higher the degree of correlation, the larger the weighting coefficient; Each of the value sub-features is multiplied by the weighting coefficient of the corresponding key sub-feature to obtain multiple weighted value sub-features corresponding to the multiple value sub-features. The multiple weighted value sub-features are then added to obtain the contact fusion feature with the same dimension as the visual ontology feature.
5. The method for generating robot actions according to claim 1, characterized in that, The acquisition of the robot's current contact characteristics includes: Within a sampling period ending at the current time, the contact effects experienced by the robot during its operation are sampled multiple times to obtain a contact information sequence. The contact information sequence includes multiple contact information items arranged according to the sampling time, and each contact information item is a measurement result of the contact effect at the corresponding sampling time. The contact information in the contact information sequence is smoothed and filtered according to the sampling order to obtain a smoothed contact information sequence. Extract the contact features at the current moment from the contact smoothing information sequence.
6. The method for generating robot actions according to claim 5, characterized in that, The step of smoothing and filtering each contact information in the contact information sequence according to the sampling order to obtain a smoothed contact information sequence includes: For the t-th contact information sampled at the t-th sampling time among the N contact information information included in the contact information sequence, the t-th contact information is multiplied by the weight parameter to obtain the current weighted component at the t-th sampling time. The (t-1)-th contact smoothing information at the (t-1)-th sampling time is multiplied by the smoothing coefficient to obtain the historical weighted component at the t-th sampling time. Here, t is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When t is 1, the (t-1)-th contact smoothing information is the preset contact smoothing information. The weight parameter is obtained by subtracting 1 from the smoothing coefficient. The smoothing coefficient is greater than 0 and less than 1. The current weighted component at the t-th sampling time and the historical weighted component at the t-th sampling time are added together to obtain the t-th contact smoothing information of the t-th contact information collected at the t-th sampling time. The contact smoothing information sequence is obtained by arranging the contact smoothing information determined for each of the N contact information in the sampling order.
7. The method for generating robot actions according to claim 5, characterized in that, Extracting the contact features at the current moment from the contact smoothing information sequence includes: The contact timing states in the gated loop unit are recursively updated according to the contact smoothing information sequence to obtain the first to Nth contact timing states in the gated loop unit: According to the sampling order, for the i-th contact smoothing information among the N contact smoothing information included in the contact smoothing information sequence, the gated loop unit recursively updates the (i-1)-th contact timing state in the gated loop unit according to the i-th contact smoothing information to obtain the i-th contact timing state in the gated loop unit, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. When i equals 1, the (i-1)-th contact timing state is the initial unit state of the gated loop unit, and the contact timing state is the unit state used by the gated loop unit to retain the changes in contact action received by the robot up to the corresponding sampling time during the sampling period; The first to Nth contact timing states in the gated loop unit are arranged into a contact timing state sequence according to the sampling order, and the contact timing state sequence is determined as the contact feature at the current time. The contact feature is used to characterize the changes in the contact action received by the robot during the sampling period between adjacent sampling times and the duration of the action during multiple consecutive sampling times.
8. The method for generating robot actions according to claim 1, characterized in that, The contact feature is a sequence of feature vectors comprising multiple feature vectors, each feature vector in the sequence comprising the same number of feature values, the feature values being used to describe the contact action state of the robot. The step of determining contact gating information for adjusting the injection amount of the contact feature based on the contact feature includes: Calculate the average value of the multiple feature vectors included in the contact feature to obtain the average contact feature; A multiplication operation is performed on the average contact feature and the gating weight to obtain a gating operation result. The gating operation result is then mapped to an injection amount greater than 0 and less than 1 to obtain the contact gating information used to indicate the injection amount. The gating weight is a parameter used to adjust the degree of influence of the average contact feature on the injection amount.
9. The method for generating robot actions according to claim 1, characterized in that, The step of updating the recursive state of the previous time step by the state propagation model based on the recursive input features to obtain the recursive state of the current time step includes: The state propagation parameters and input writing parameters at the current time are determined based on the recursive input features. The state propagation parameters are used to determine the propagation method from the previous recursive state to the current recursive state, and the input writing parameters are used to determine the method of writing the recursive input features into the current recursive state. The recursive state of the previous time step is multiplied with the state propagation parameters to obtain the historical state component. The recursive input feature is multiplied with the input writing parameter to obtain the recursive writing component. The historical state component and the recursive writing component are added to obtain the recursive state of the current time step.
10. The method for generating robot actions according to claim 1, characterized in that, The step of generating the robot action at the current moment based on the recursive state at the current moment includes: The state propagation model determines the output projection parameters at the current moment based on the recursive input features. The recursive state at the current moment is multiplied by the output projection parameters to obtain the state output component. The recursive input features are multiplied by the direct mapping parameters of the state propagation model to obtain the direct output component. The output projection parameters are used to extract historical operation features related to the robot's actions at the current moment from the recursive state at the current moment. The direct mapping parameters are used to directly transmit the operation state and contact interaction state at the current moment reflected by the recursive input features to the output of the state propagation model. The state output component and the direct output component are added together to obtain the model output feature of the state propagation model at the current time. The model output feature is used to characterize the contact process experienced by the robot up to the current time, the operation state at the current time, and the contact action state when performing the operation at the current time. The robot's motion head maps the model output features into robot control quantities for driving the robot's actuators, thus obtaining the robot's motion at the current moment.
11. The method for generating robot actions according to claim 1, characterized in that, Before injecting the contact features into the visual ontological features of the robot's state propagation model at the current moment, the method further includes: The robot's visual ontology basic features and visual short-term memory features at the current moment are obtained. The visual ontology basic features include visual features generated based on the robot's visual observation results at the current moment and ontology features generated based on the robot's ontology measurement results at the current moment. The visual short-term memory features are used to characterize the state of the operation scene observed by the robot at multiple consecutive visual observation moments before the current moment and the change of the state of the operation scene over time. Visual query features are determined based on the basic features of the visual ontology, and visual key features and visual value features are determined based on the visual short-term memory features. Based on the degree of correlation between the visual query features and multiple visual key sub-features in the visual key features, and in a manner where the higher the degree of correlation, the greater the weight of the corresponding visual value sub-feature, multiple visual value sub-features in the visual value features are weighted and fused to obtain a visual memory fusion feature with the same dimension as the basic features of the visual ontology. Obtain visual gating information for adjusting the amount of visual memory fusion feature injection; The visual memory fusion feature is multiplied by the injection amount indicated by the visual gating information to obtain the visual memory injection component; The visual memory injection component is superimposed on the visual ontology basic feature to obtain the visual ontology feature.
12. A device for generating robot motion, characterized in that, include: The first acquisition module is used to acquire the contact features of the robot at the current moment, wherein the contact features are used to characterize the contact state of the robot when performing an operation at the current moment; An injection module is used to determine contact gating information for adjusting the injection amount of the contact features based on the contact features, and to inject the contact features into the visual ontology features of the robot's state propagation model at the current moment according to the injection amount indicated by the contact gating information, so as to obtain the recursive input features of the state propagation model at the current moment. The visual ontology features are used to characterize the operation scene state and the robot's ontology motion state in the operation scene at the current moment. The update module is used to update the recursive state of the previous time step by the state propagation model according to the recursive input features to obtain the recursive state of the current time step, wherein the recursive state is the model state used by the state propagation model to retain the contact process experienced by the robot up to the corresponding time step; The generation module is used to generate the robot action at the current moment based on the recursive state at the current moment.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 11.