Adaptive multi-granularity representation method and system of robot action space
Patent Information
- Application Number
- CN202611157901.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本申请提供一种机器人动作空间的自适应多粒度表示方法及系统,旨在解决机器人动作预测和控制过程中,现有动作空间表示方式难以适应不同任务的精度和效率要求、不同机器人平台之间的动作描述难以统一,以及复杂动作生成的实时性和准确性难以兼顾的技术问题
[0010]本申请实施例提供一种机器人动作空间的自适应多粒度表示方法及系统,该方法通过获取机器人执行目标任务对应的视觉观测信息、任务指令和机器人条件信息,从视觉与语言信息中确定多模态融合特征和用于表征动作需求的任务复杂度特征,并结合机器人条件信息针对不同动作维度自适应确定离散化粒度,使不同动作维度能够按照实际任务需求采用相适应的表示精度;进一步根据粒度配置匹配相应的目标动作码本,并通过条件化动作解码器生成动作标记序列,由此能够在统一的动作表示框架下适配不同任务需求和机器人平台;在将动作标记序列重建为连续动作序列后,再结合机器人条件信息进行运动约束校验和可行动作域投影,从而保证生成动作与机器人实际运动能力相适应。因此,本申请能够兼顾机器人动作表示的精度、生成效率和跨平台适用性,并提高复杂任务下机器人控制动作的准确性、可执行性和安全性。。
Smart Images

Figure CN122807906A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and in particular to an adaptive multi-granularity representation method and system for robot motion space. Background Technology
[0002] With the continuous development of artificial intelligence, machine vision, and robot control technologies, robots are increasingly being applied in industrial manufacturing, warehousing and logistics, medical assistance, and home services. To enable robots to autonomously perform operations such as movement, grasping, placement, and assembly based on environmental conditions and task requirements, it is typically necessary to process the environmental information perceived by the robot and the task instructions received, and convert the processing results into action information that the robot control system can execute. Since robot movements usually involve multiple joints, end effectors, or motion mechanisms, and each motion component has different value ranges and control requirements, how to reasonably represent the robot's motion space directly affects the accuracy and real-time performance of robot motion prediction and control execution.
[0003] Existing robot motion space representation methods typically describe robot movements according to pre-defined motion structures and data formats, and use this to perform motion prediction or generate control commands. However, the operational range, control precision, and response speed requirements of the tasks performed by the robot may vary with the task type and execution stage. Using relatively fixed motion representation methods makes it difficult to simultaneously achieve both motion description accuracy and data processing efficiency under different task conditions. In complex or delicate operations, insufficient motion representation precision may lead to deviations in the robot's end-effector position or posture; conversely, in situations with many motion dimensions or long task durations, complex motion representations may increase the computational burden and motion generation time of the model. Furthermore, different robots differ in mechanical structure, number of degrees of freedom, joint range of motion, and motion constraints, limiting the applicability of existing motion representation methods across different robot platforms and increasing the development, training, and migration costs of robot control models.
[0004] Therefore, in the process of robot motion prediction and control, the motion space representation is difficult to adapt to the accuracy and efficiency requirements of different tasks, the motion description between different robot platforms is difficult to unify, and the real-time performance and accuracy of complex motion generation are difficult to balance, which have become urgent problems to be solved. Summary of the Invention
[0005] This application provides an adaptive multi-granularity representation method and system for robot motion space, aiming to solve the technical problems in robot motion prediction and control, such as the difficulty of existing motion space representation methods to adapt to the accuracy and efficiency requirements of different tasks, the difficulty of unifying motion descriptions between different robot platforms, and the difficulty of balancing real-time performance and accuracy in complex motion generation.
[0006] In a first aspect, this application provides an adaptive multi-granularity representation method for robot motion space, the method comprising: In one possible design, Secondly, this application provides an adaptive multi-granularity representation system for robot motion space, the system comprising: The information acquisition module is used to acquire visual observation information, task instructions, and robot condition information corresponding to the robot's execution of the target task; The feature determination module is used to determine multimodal fusion features and task complexity features based on the visual observation information and the task instructions, wherein the task complexity features are used to characterize the action requirements of the target task; The granularity configuration module is used to determine the corresponding discretized granularity from multiple preset granularity levels for each action dimension of the robot based on the task complexity characteristics and the robot condition information, so as to obtain the granularity configuration. The codebook determination module is used to determine the target action codebook corresponding to each action dimension from the multi-granularity action codebook library according to the granularity configuration, wherein the target action codebook is used to represent the mapping relationship between the continuous action values and action tags of the corresponding action dimension; The action tag generation module is used to take the multimodal fusion features, the robot conditional information and the granular configuration input conditional action decoder to generate an action tag sequence that matches the target action codebook corresponding to each action dimension; The action reconstruction module is used to reconstruct the action tag sequence into a continuous action sequence based on the target action codebook corresponding to each action dimension. The constraint processing module is used to perform motion constraint verification on the continuous action sequence according to the robot condition information, and project the actions in the continuous action sequence that do not meet the motion constraints to the corresponding movable action domain of the robot to obtain a robot control action sequence for controlling the robot to perform the target task.
[0007] Thirdly, this application provides an electronic device, including: a memory and at least one processor; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method described in the first aspect or various possible designs of the first aspect.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed, implement the method described in the first aspect or various possible designs of the first aspect.
[0009] Fifthly, this application provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to implement the method described in the first aspect or various possible designs of the first aspect.
[0010] This application provides an adaptive multi-granularity representation method and system for robot motion space. The method acquires visual observation information, task instructions, and robot condition information corresponding to the robot's execution of a target task. It determines multimodal fusion features and task complexity features representing motion requirements from visual and linguistic information, and adaptively determines the discretization granularity for different motion dimensions based on the robot condition information. This allows different motion dimensions to adopt appropriate representation accuracy according to actual task requirements. Furthermore, it matches the corresponding target motion codebook according to the granularity configuration and generates motion marker sequences through a conditional motion decoder. This enables adaptation to different task requirements and robot platforms within a unified motion representation framework. After reconstructing the motion marker sequences into continuous motion sequences, it combines robot condition information for motion constraint verification and projection of the movable motion domain, ensuring that the generated motion is adapted to the robot's actual motion capabilities. Therefore, this application can balance the accuracy, generation efficiency, and cross-platform applicability of robot motion representation, and improve the accuracy, executability, and safety of robot control actions under complex tasks. Attached Figure Description
[0011] Figure 1 A flowchart illustrating an adaptive multi-granularity representation method for robot motion space provided in this application embodiment; Figure 2 A schematic diagram illustrating a unified training and migration deployment process for multiple robot platforms provided in an embodiment of this application; Figure 3 A schematic diagram of a hierarchical action prediction architecture provided in an embodiment of this application; Figure 4 A schematic diagram of the module architecture of an adaptive multi-granularity representation system for robot motion space provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion.
[0014] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] In this article, the term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B can exist simultaneously, and B exists. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0016] Furthermore, the terms "first," "second," etc., in the specification and claims of this application or in the aforementioned drawings are used to distinguish different objects rather than to describe a specific order, and may explicitly or implicitly include one or more of the features.
[0017] In the description of this application, unless otherwise stated, "multiple" and "at least two" mean two or more (including two), and similarly, "multiple groups" and "at least two groups" mean two or more (including two groups).
[0018] In the description of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, "connected" or "linked" can refer not only to a physical connection, but also to an electrical connection or a signal connection. For instance, it can be a direct connection, i.e., a physical connection, or an indirect connection through at least one intermediate component, as long as the circuit is connected. It can also refer to the internal connection between two components. A signal connection can refer not only to a signal connection through a circuit, but also to a signal connection through a medium, such as radio waves. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, different technical features in this application can be combined with each other.
[0020] Figure 1 This is a flowchart illustrating an adaptive multi-granularity representation method for robot motion space provided in an embodiment of this application. Figure 1 As shown, the method includes S101 to S107, and S101 to S107 are described in detail below.
[0021] It should be noted that the method provided in this application embodiment can be executed by a robot controller, task computer, edge computing device, cloud server, motion decision module in the robot control system, or processor in an electronic device. This method can be deployed on the robot body itself, or it can be deployed collaboratively with the cloud and the robot body. Visual observation information, task instructions, and robot condition information can be input into a pre-trained vision-language-action model, which performs task understanding, motion space representation, and robot control motion generation. The robot can be an industrial robotic arm, a mobile manipulation robot, a service robot, a medical robot, or other robots with one or more motion dimensions.
[0022] In one possible embodiment, before executing the reasoning process shown in S101 to S107, the visual feature extraction network, language feature extraction network, multimodal fusion network, granularity selection network, conditional action decoder, and multigranularity action codebook can be trained based on multi-robot demonstration data. Each piece of multi-robot demonstration data may include visual observation information, task instructions, robot condition information, continuous action sequences, and task execution results collected when the robot performs the demonstration task, and establish the correspondence between the various data according to the action sequence.
[0023] During training, a multi-granularity action codebook library can be constructed based on the range of continuous action values for different robots and multiple preset granularity levels. The continuous action sequences in the demonstration data are then converted into corresponding action label sequences by an action encoder, which serve as training supervision information for the conditional action decoder. Visual feature extraction and language feature extraction networks extract visual observation features and task instruction features, respectively. A multimodal fusion network fuses these features. A granularity selection network learns the granularity selection relationship for each action dimension based on task complexity features and robot conditional features. The conditional action decoder learns the generation relationship of action label sequences based on multimodal fusion features, robot conditional information, and granularity configuration.
[0024] After initial training, the network parameters of the visual feature extraction network, language feature extraction network, multimodal fusion network, granularity selection network, and conditional action decoder, as well as the codebook parameters of the multi-granularity action codebook library, can be saved. During the inference phase, the aforementioned network parameters and codebook parameters are loaded, the visual observation information, task instructions, and robot condition information corresponding to the current target task are obtained, and a sequence of robot control actions is generated according to steps S101 to S107. The action encoder is mainly used during the training phase to convert the demonstrated continuous actions into action labels required for training; it does not need to be called during the inference phase.
[0025] S101. Obtain visual observation information, task instructions, and robot condition information corresponding to the robot's execution of the target task.
[0026] The target task refers to the operational task that the robot needs to perform, such as object grasping, object placement, component assembly, insertion, welding, navigation, or a complex task consisting of multiple operational stages. Visual observation information can be acquired through visual sensors installed on the robot itself or in the working environment. Visual sensors can include color cameras, depth cameras, or color depth cameras. Visual observation information can include some or all of the following: images of the working environment, images of the target object, depth information, and spatial position data formed by the depth information, used to characterize the target object, obstacles, operating area, and the spatial relationship between the robot and the target object.
[0027] Task instructions describe the execution content of the target task. These can be user-input natural language text instructions or speech instructions obtained through speech recognition. For example, a task instruction could be "Place the part in the specified position" or "Insert the part into the corresponding slot." Robot condition information describes the current structural characteristics and motion capabilities of the robot. This information can be obtained from the robot's device configuration file, robot description model, controller parameters, or real-time status interface. Different robots have different degrees of freedom, motion dimensions, and motion constraints. By obtaining robot condition information, subsequent motion space representation processes can be adapted to the actual structure and motion capabilities of the current robot.
[0028] In one implementation, visual observation information, task instructions, and robot condition information are correlated in time and by task, so that the acquired information corresponds to the same target task and the same task execution stage, thereby avoiding the impact of asynchronous input information on the generation of subsequent actions.
[0029] S102. Based on visual observation information and task instructions, determine the multimodal fusion features and task complexity features.
[0030] Among them, task complexity features are used to characterize the action requirements of the target task.
[0031] Specifically, features can be extracted from visual observation information and task instructions separately, and then the extracted visual and linguistic features can be fused to obtain multimodal fusion features. Multimodal fusion features can simultaneously represent visual information such as objects, spatial locations, and obstacle distribution in the work environment, as well as semantic information such as operation objects, operation methods, and task objectives contained in task instructions, which can then be used by the conditional action decoder to generate action tag sequences that match the current work environment and task semantics.
[0032] Task complexity features can be determined based on multimodal fusion features and are used to characterize the target task's requirements in terms of action execution. Task complexity features can reflect the target task's requirements for action spatial accuracy, action execution speed, adaptability to the manipulated object, and adaptability to environmental constraints. For example, for large-scale movement or coarse positioning tasks, the spatial accuracy requirement is relatively low; for component assembly, insertion, or precision welding tasks, the spatial accuracy requirement is relatively high; for tasks involving fragile or flexible objects, it is also necessary to determine the corresponding action requirements based on the characteristics of the manipulated object.
[0033] S103. Based on the task complexity characteristics and robot condition information, for each action dimension of the robot, determine the corresponding discretization granularity from multiple preset granularity levels to obtain the granularity configuration.
[0034] Discretization granularity is used to characterize the fineness of discretizing the range of continuous action values for a given action dimension. A preset granularity level corresponds to a number of discrete intervals. The higher the discretization granularity, the more continuous action intervals are obtained from the range of continuous action values, and the higher the resolution of the action representation. The lower the discretization granularity, the fewer continuous action intervals are obtained from the range of continuous action values, and the lower the processing overhead required for action representation and action prediction.
[0035] The motion dimension of a robot can be determined based on its type. For example, for a robotic arm, the motion dimension can include the joint motion dimension of each joint and the opening and closing motion dimension of the gripper; for a mobile robot, the motion dimension can also include the translational speed dimension and rotational speed dimension of the mobile base. Based on the task complexity characteristics and robot condition information, the discretization granularity can be determined for each motion dimension, without requiring all motion dimensions to use the same discretization granularity.
[0036] For example, when performing precision insertion tasks, motion dimensions closely related to end-effector position and attitude control can use a higher discretization granularity, while motion dimensions related to coarse movement or gripper opening and closing can use a relatively lower discretization granularity. Conversely, when performing large-range movement tasks, multiple motion dimensions can use relatively lower discretization granularity. Thus, a granularity configuration including multiple motion dimensions and their corresponding discretization granularities can be formed within the same target task. Multiple preset granularity levels can include, for example, 16, 32, 64, 128, and 256 granularity levels, but these values are only examples; the specific number and values of preset granularity levels can be set according to the robot's motion range and control accuracy requirements.
[0037] S104. Based on the granularity configuration, determine the target action codebook corresponding to each action dimension from the multi-granularity action codebook library.
[0038] The target action codebook is used to represent the mapping relationship between the continuous action values and action tags of the corresponding action dimension.
[0039] The multi-granularity action codebook library can pre-store multiple action codebooks corresponding to different preset granularity levels. Each action codebook can divide the range of continuous action values into multiple continuous action intervals that match the corresponding preset granularity level, and configure corresponding action tags for each continuous action interval. Action tags can be discrete identifiers that can be predicted by the conditional action decoder, and different action tags are used to represent different continuous action intervals.
[0040] After determining the granularity configuration, the discretized granularity corresponding to each action dimension can be read, and the action codebook corresponding to that discretized granularity can be selected from the multi-granularity action codebook library as the target action codebook for that action dimension. For example, when an action dimension corresponds to 32 granularities, an action codebook containing 32 consecutive action intervals and corresponding action tags can be selected; when another action dimension corresponds to 128 granularities, an action codebook containing 128 consecutive action intervals and corresponding action tags can be selected.
[0041] For different motion dimensions, the target motion codebook can be parameter-adapted based on the range of continuous motion values for the corresponding motion dimension. For example, the joint angle dimension, gripper opening / closing dimension, and movement speed dimension can each have different ranges of continuous motion values. The target motion codebook serves two purposes: firstly, it limits the motion markers that can be used for the corresponding motion dimension; secondly, it is used to reconstruct continuous motion values from the generated motion markers in subsequent steps. The composition of the multi-granularity motion codebook library and the specific process for determining the target motion codebook will be further explained in subsequent embodiments.
[0042] S105. Input the multimodal fusion features, robot conditional information and granular configuration into the conditionalized action decoder to generate an action tag sequence that matches the target action codebook corresponding to each action dimension.
[0043] A conditional motion decoder is a decoding model with sequence prediction capabilities. It understands the current working environment and target task based on multimodal fusion features, determines the robot's motion dimensions and capabilities based on robot condition information, and determines the discretization granularity for each motion dimension based on granularity configuration. Robot condition information and granularity configuration can be embedded as conditional information into the motion prediction process of the conditional motion decoder, enabling it to generate corresponding motion labels for different robots, tasks, and motion dimensions.
[0044] In one implementation, the conditional action decoder predicts action tags progressively according to the generation order corresponding to the action temporal sequence and action dimensions. When predicting the action tag corresponding to a certain action dimension, the set of predictable action tags is limited according to the target action codebook corresponding to that action dimension, so that the generated action tags can match the discretization granularity used for that action dimension. The action tag sequence can represent robot actions within one control cycle, or robot actions within multiple consecutive control cycles or a preset prediction time domain.
[0045] By configuring the granularity to conditionally constrain the generation process of action tags, simple tasks or low-precision action dimensions can use a lower-granularity target action codebook, thereby reducing unnecessary action representation complexity and decoding computation overhead; fine tasks or high-precision action dimensions can use a higher-granularity target action codebook, thereby improving the action resolution corresponding to the action tag.
[0046] S106. Based on the target action codebook corresponding to each action dimension, reconstruct the action tag sequence into a continuous action sequence.
[0047] Specifically, based on the action dimension corresponding to each action marker in the action marker sequence, the target action codebook corresponding to that action dimension can be queried to determine the continuous action interval mapped to each action marker, and the corresponding continuous action value can be determined based on the continuous action interval. The continuous action value can be the midpoint value, center value, representative value determined by training, or other values calculated based on the boundary of the continuous action interval.
[0048] After obtaining the continuous motion values corresponding to each motion marker, these values can be combined according to the temporal and dimensional relationships of the motion marker sequence to form a continuous motion sequence. This continuous motion sequence can include joint position sequences, joint velocity sequences, end-effector pose sequences, gripper opening / closing sequences, mobile base velocity sequences, or other continuous control parameters that match the robot's motion interface. Reconstruction using the target motion codebook converts the discrete motion markers generated by the conditional motion decoder into continuous motion values that can be further processed by the robot controller.
[0049] S107. Perform motion constraint verification on the continuous action sequence based on the robot condition information, and project the actions in the continuous action sequence that do not meet the motion constraints to the corresponding movable action domain of the robot to obtain the robot control action sequence used to control the robot to perform the target task.
[0050] In this step, the motion constraints corresponding to the current robot can be determined based on the robot's condition information. Motion constraints may include some or all of the following: joint position constraints, joint velocity constraints, joint acceleration constraints, end effector range of motion constraints, pedestal motion constraints, and collision constraints between the robot body or the robot and its environment. Based on these motion constraints, each continuous action in the sequence of continuous actions is verified to determine whether each continuous action is within the current robot's movable action domain.
[0051] When all actions in a continuous sequence of actions satisfy motion constraints, the sequence can be directly used as the robot's control sequence. When there are actions in the continuous sequence that do not satisfy motion constraints, these actions are identified as actions to be corrected. Under the condition that the robot's motion constraints are satisfied, these corrected actions are projected onto the movable action domain to obtain corrected actions that satisfy the motion constraints. The movable action domain refers to the set of action values that the robot can safely execute, jointly defined by the robot's current structural parameters and motion constraints.
[0052] In one implementation, a constraint optimization method can be used to project the movable action domain, with the optimization objective being to minimize the difference between the corrected action and the action to be corrected, and robot motion constraints as the constraint conditions, to obtain the corrected action located within the movable action domain. For example, a quadratic programming method can be used to solve for the corrected action. Subsequently, the actions that satisfy the motion constraints in the continuous action sequence are combined with the corrected action according to the original action sequence and action dimension to obtain the robot control action sequence. The robot control action sequence is then sent to the robot controller, which drives the corresponding joints, grippers, mobile bases, or other actuators to perform the target task according to the robot control action sequence.
[0053] In this embodiment, multimodal fusion features and task complexity features representing action requirements are determined from visual observation information and task instructions. Discretization granularity is then determined for different action dimensions based on robot conditional information, enabling each action dimension to adopt a representation method adapted to its accuracy requirements according to the actual task needs. Furthermore, a target action codebook corresponding to each action dimension is selected based on the granularity configuration, and a conditional action decoder generates an action tag sequence matching the target action codebook. This allows the same action representation framework to adapt to different task requirements and different robot platforms. After reconstructing the action tag sequence into a continuous action sequence, motion constraint verification and feasible action domain projection are performed based on robot conditional information, thereby avoiding the generation of control actions that exceed the robot's actual motion capabilities. Therefore, this application can reduce unnecessary action representation and decoding overhead in simple tasks while ensuring the accuracy of action representation for fine tasks, balancing the accuracy, real-time performance, and cross-platform applicability of robot action generation, and improving the executability and safety of robot control actions.
[0054] In one possible embodiment, the method steps shown in S102 are implemented by S1021 to S1025, which are described in detail below.
[0055] S1021. Visual feature extraction is performed on the visual observation information to obtain visual observation features.
[0056] Specifically, visual observation information can be input into a visual feature extraction network. This network then extracts visual features corresponding to the target object, work area, obstacles, and spatial relationships between objects, thus obtaining visual observation features. The visual feature extraction network can employ convolutional neural networks, visual Transformer networks, or other network structures capable of encoding image information.
[0057] When visual observation information includes color images and depth information, features can be extracted from the color images and depth information separately, and the extracted results can be concatenated, weighted, or fused with attention to enable the visual observation features to simultaneously represent the appearance and spatial location attributes of the target object. Visual observation features can be represented in the form of feature vectors, feature matrices, or sequences of visual markers.
[0058] S1022. Extract language features from the task instructions to obtain task instruction features.
[0059] Specifically, the task instructions can first be segmented or tokenized to convert them into a sequence of linguistic tags. This sequence is then input into a language feature extraction network to obtain the task instruction features. The language feature extraction network can employ a Transformer encoder or other network structures with semantic encoding capabilities.
[0060] Task instruction features can be used to characterize the operation object, operation action, target position, execution order, accuracy requirements, and time requirements contained in the task instruction. For example, for the task instruction "insert part into slot", the task instruction features can characterize "part" as the operation object, "insert" as the operation action, and "slot" as the target position.
[0061] S1023. Perform feature fusion on visual observation features and task instruction features to obtain multimodal fusion features.
[0062] Specifically, visual observation features and task instruction features can be mapped to the same feature dimension first, and then feature fusion can be performed through feature concatenation, weighted fusion, cross-attention mechanism or multimodal Transformer network to establish a relationship between environmental objects in visual observation information and operational semantics in task instructions, thus obtaining multimodal fused features.
[0063] For example, attention weighting can be applied to visual observation features based on task instruction features to highlight target objects, target areas, and obstacles related to the target task from the visual observation information; semantic supplementation can also be applied to task instruction features based on visual observation features to determine the actual location and state of the operation object in the current working environment. The resulting multimodal fusion features can uniformly represent the semantic requirements of the target task and the state of the current working environment.
[0064] In one implementation, during the determination of multimodal fusion features, object detection and scene understanding can also be performed on the visual observation information. Object detection is used to determine the object category, object position, object size, object outline, or object pose corresponding to the manipulated object, target area, and obstacles from the visual observation information, and the object detection results are fused with the visual observation features. Scene understanding is used to combine the object detection results and task instruction features to determine the spatial relationship between the manipulated object and the target area, the robot's available operating space, and the operation process required to complete the target task.
[0065] Object detection results can be used to determine the spatial accuracy requirements and characteristics of the manipulated object for a target task. For example, spatial accuracy requirements can be determined based on the size differences, relative positional relationships, and fit clearances between the manipulated object and the target area, and the rigidity, flexibility, or fragility of the manipulated object can be determined based on the object category, appearance, or preset object attributes. Scene understanding results can be used to determine the time urgency and environmental constraints. For example, the time urgency can be determined based on the motion state of the manipulated object, and the environmental constraints can be determined based on the number of obstacles around the target area, the available operating space, and the passage range.
[0066] S1024. Based on the multimodal fusion characteristics, determine the spatial accuracy requirements, time urgency, characteristics of the operation object, and degree of environmental constraints corresponding to the target task.
[0067] Specifically, multimodal fusion features can be input into a task complexity evaluation network. Through multiple feature evaluation branches within this network, the network outputs spatial accuracy requirements, time urgency, characteristics of the operational object, and environmental constraints, respectively. Each feature evaluation branch can output a corresponding category using a classification method or a corresponding quantized value using a regression method.
[0068] Spatial accuracy requirements characterize the accuracy requirements of a robot's position, posture, or motion trajectory when performing a target task. For example, tasks involving large-scale movement or coarse positioning may require lower spatial accuracy, while assembly, insertion, or precision welding tasks may require higher spatial accuracy.
[0069] The time urgency level is used to characterize the requirements of the target task for the speed of action execution and response latency. It can be determined based on the time-related description in the task instruction, the motion state of the target object, and the allowable execution time of the operation process.
[0070] Object characteristics are used to characterize the physical properties of the object being manipulated in the target task, and may include some or all of rigidity, flexibility, fragility, size, or weight. For example, when performing a grasping task on a fragile object, it is necessary to reduce the range of motion and improve the precision of motion control.
[0071] Environmental constraints characterize the degree to which a robot's available operating space is limited. They can be determined based on the size of the work area, the distribution of obstacles, the available operating space around the target object, and the distance between the robot and the environment. Open work environments correspond to lower environmental constraints, while confined spaces or environments with dense obstacles correspond to higher environmental constraints.
[0072] S1025. Based on the spatial accuracy requirements, time urgency, characteristics of the operation object, and degree of environmental constraints, the task complexity characteristics are obtained.
[0073] Specifically, spatial accuracy requirements, time urgency, characteristics of the operation object, and environmental constraints can be numerically encoded or feature-embedded separately, and then concatenated according to a preset dimensional order to obtain task complexity features. Alternatively, each feature can be normalized and weighted to ensure that different features have comparable numerical ranges.
[0074] Task complexity features can be represented as a fixed-dimensional feature vector, where different feature dimensions characterize spatial accuracy requirements, time urgency, characteristics of the operand, and environmental constraints, respectively. Task complexity features can characterize the overall action requirements of the target task, or they can be determined separately for different execution stages of the target task. For example, in a part insertion task, the spatial accuracy requirement of the coarse positioning stage can be lower than that of the fine insertion stage, thus allowing different execution stages to acquire different task complexity features.
[0075] In one implementation, the process of determining task complexity characteristics based on visual observation information and task instructions can be represented as follows: ;in, This represents the characteristics of task complexity. Represents visual observation information; Indicates task instructions; This represents the task complexity evaluation function; It represents a real feature space consisting of K feature dimensions; K represents the number of feature dimensions of the task complexity feature.
[0076] The task complexity evaluation function may include visual feature extraction, linguistic feature extraction, feature fusion, and complexity evaluation processes. Specifically, the task complexity evaluation function extracts features from visual observation information and task instructions, fuses the obtained visual observation features and task instruction features, and determines spatial accuracy requirements, time urgency, characteristics of the operation object, and environmental constraints based on the formed multimodal fused features, thereby forming task complexity features. Different feature dimensions in the task complexity features are used to characterize the complexity of the target task in terms of corresponding action requirements.
[0077] In one example, task complexity features can be represented as a four-dimensional feature vector consisting of spatial precision requirements, time urgency, characteristics of the manipulated object, and environmental constraints, where K can be 4. Spatial precision requirements can characterize fine or coarse operational needs, time urgency can characterize fast or slow execution needs, characteristics of the manipulated object can characterize the rigidity, flexibility, or fragility of the manipulated object, and environmental constraints can characterize whether the robot's operating environment is open or restricted.
[0078] In this embodiment, by extracting visual observation features and task instruction features separately and fusing them, the task complexity assessment can be combined with the actual working environment and task semantics, avoiding the determination of action requirements based on only a single piece of information. Furthermore, spatial accuracy requirements, time urgency, characteristics of the operation object, and degree of environmental constraints are determined from the multimodal fusion features, forming task complexity features. This provides a feature basis that matches the actual needs of the target task for the subsequent selection of discretization granularity for each action dimension, thereby improving the accuracy of discretization granularity selection.
[0079] In one possible embodiment, the robot condition information processing procedure shown in the above embodiment and the application of robot condition information in S103, S105 and S107 are implemented through S301 to S305, and S301 to S305 are described in detail below.
[0080] S301. Obtain robot structural information and motion constraint information.
[0081] Robot condition information includes robot structural information and motion constraint information. Robot structural information characterizes the current structural composition and motion patterns of the robot and can be obtained from the robot description file, device configuration file, robot controller, or robot state interface. Robot structural information includes at least one of the following: number of degrees of freedom, joint type, joint range of motion, and kinematic type.
[0082] The number of degrees of freedom (DOFs) characterizes the number of motion dimensions a robot can independently move. For example, a robotic arm may include six or seven joint DDFs, and a mobile manipulator may also include translational and rotational DDFs corresponding to its mobile base. Joint types may include rotational joints, translational joints, and gripper opening / closing joints, etc. Joint range of motion characterizes the minimum and maximum positions that each joint can reach. Kinematic type characterizes the robot's motion structure and may include serial robotic arms, parallel robots, wheeled robots, legged robots, or mobile manipulators, etc.
[0083] Motion constraint information is used to characterize the constraints that a robot must satisfy when performing actions, including at least one of joint position constraints, joint velocity constraints, joint acceleration constraints, and collision constraints. Specifically, joint position constraints limit joint movements from exceeding their corresponding range of motion; joint velocity and joint acceleration constraints limit joint velocity and acceleration between adjacent control moments, respectively; and collision constraints prevent collisions between structural components of the robot body or between the robot and obstacles in the working environment.
[0084] S302. Perform conditional encoding on the robot's structural information to obtain the robot's conditional features.
[0085] Specifically, different types of data in the robot's structural information can be numerically processed first. For example, the number of degrees of freedom can be numerically encoded, different joint types and kinematic types can be categorically encoded or embedded encoded, and the range of motion of each joint can be normalized. Subsequently, the processed number of degrees of freedom, joint types, range of motion, and kinematic types are combined and input into a conditional coding network to obtain the robot's conditional features.
[0086] Conditional coding networks can employ multilayer perceptrons, fully connected neural networks, or other network structures capable of feature mapping of robot structural information. In one implementation, the number of robot degrees of freedom, joint range of motion, and kinematic type can be concatenated and input into a multilayer perceptron. The multilayer perceptron then maps the robot structural information corresponding to different robots to a preset feature space, thereby obtaining the robot's conditional features.
[0087] For robots with different degrees of freedom, robot structural information can be processed through padding, masking, or variable-length feature aggregation. This allows robot conditional features corresponding to different robots to be input into the same granularity determination model and conditional motion decoder. Robot conditional features are used to preserve information such as the current robot's motion dimensions, joint structure, and motion patterns, without requiring different robots to have completely identical structures.
[0088] S303. Based on the task complexity characteristics and robot condition characteristics, determine the discretization granularity corresponding to each action dimension to obtain the granularity configuration.
[0089] Specifically, task complexity features and robot condition features can be input into the granularity determination model, so that the granularity determination model can identify the target task action requirements and, in combination with the current robot action dimensions and structural characteristics, determine the corresponding discretized granularity for each action dimension.
[0090] For example, even if a six-DOF (degrees of freedom) robotic arm and a seven-DOF robotic arm perform the same target task, the discretization granularity corresponding to each motion dimension can differ due to differences in the number of degrees of freedom, joint types, and kinematic structures. For mobile manipulation robots, the corresponding discretization granularity can be determined separately for the robotic arm joint motion dimension, the gripper opening and closing motion dimension, and the mobile base motion dimension.
[0091] The resulting granularity configuration reflects both the action requirements of the target task and is compatible with the current robot structure. The specific process of determining the discretization granularity based on task complexity and robot condition characteristics will be further explained in subsequent implementations.
[0092] S304. Input multimodal fusion features, robot conditional features, and granular configuration into the conditionalized action decoder.
[0093] Specifically, robot conditional features are used as robot conditional inputs to the conditional motion decoder, enabling the decoder to identify the number of degrees of freedom, motion dimension composition, and kinematic type of the current robot; granularity configuration is used as granularity conditional input, enabling the decoder to determine the discretization granularity used for each motion dimension; and multimodal fusion features are used as task and environmental conditional inputs, enabling the decoder to generate motion label sequences by combining the current working environment, target task, and robot structure.
[0094] By conditionalizing the motion decoding process using robot conditional features, the same conditional motion decoder can generate motion labels that match the corresponding motion dimensions for different robot platforms, without having to set up completely independent motion decoding models for each robot structure.
[0095] S305. Perform motion constraint verification on the continuous action sequence based on the motion constraint information.
[0096] Specifically, after reconstructing the action marker sequence into a continuous action sequence, the joint position constraints, joint velocity constraints, and joint acceleration constraints corresponding to each action dimension can be read from the motion constraint information, and it can be determined one by one whether the actions in the continuous action sequence satisfy the corresponding constraints.
[0097] For joint position constraints, it can be determined whether each joint position is within the corresponding joint motion range; for joint velocity constraints, the joint velocity can be determined based on the change in joint position and the control time interval between adjacent control moments, and it can be determined whether the joint velocity exceeds the allowable velocity range; for joint acceleration constraints, it can be determined whether the joint acceleration exceeds the allowable acceleration range based on the change between adjacent joint velocities.
[0098] When motion constraint information includes collision constraints, the distances between various structural components of the robot body or between the robot and obstacles can be determined based on the robot pose corresponding to the continuous action sequence and the positions of obstacles in the working environment. This distance is then used to determine whether the corresponding action satisfies the collision constraints. After motion constraint verification, actions in the continuous action sequence that satisfy or do not meet the motion constraints can be identified, providing a basis for subsequent projection of the movable action domain.
[0099] In this embodiment, by encoding robot structural information into robot conditional features, robots with different numbers of degrees of freedom, joint types, and kinematic types can be mapped to a unified conditional representation space. This allows task complexity features to determine the discretization granularity corresponding to each action dimension based on the current robot's structural characteristics, and enables the conditional action decoder to generate action label sequences that match the current robot's action space. Simultaneously, by independently preserving motion constraint information and performing motion constraint verification on the reconstructed continuous action sequences, actions exceeding the robot's actual motion capabilities or potentially causing collisions can be identified in a timely manner. Therefore, this implementation improves the adaptability of the robot's action space representation and action generation process to different robot platforms, and enhances the executability and safety of the generated actions.
[0100] Figure 2 This is a schematic diagram illustrating a unified training and migration deployment process for multiple robot platforms, provided as an embodiment of this application. Figure 2 As shown, robot demonstration data corresponding to robotic arm A, robotic arm B, and mobile operation robot can be obtained separately, and the robot demonstration data of different robot platforms can be summarized to obtain a multi-robot hybrid dataset.
[0101] Robotic arm A may include six joint motion dimensions and one gripper motion dimension, robotic arm B may include seven joint motion dimensions and one gripper motion dimension, and the mobile manipulation robot may include six robotic arm joint motion dimensions, two mobile base motion dimensions, and one gripper motion dimension. Different robot platforms may have different numbers of degrees of freedom, joint types, joint range of motion, and kinematic types.
[0102] The robot structural information corresponding to each robot platform is conditionally encoded to obtain robot conditional features, which are then used as conditional inputs to a unified training model. This unified training model can be trained using a multi-robot hybrid dataset, enabling the model to learn the motion dimension composition, motion value range, and motion characteristics corresponding to different robot platforms.
[0103] When deploying a unified training model to a new robot platform, corresponding robot conditional features can be generated based on the number of degrees of freedom, joint types, joint range of motion, and kinematic type of the new robot platform. These features can then be adapted to the new robot platform using a small amount of demonstration data, thus enabling the migration and deployment of the new robot platform. Distinguishing between different robot platforms through robot conditional coding reduces the amount of data and training process required to train independent models for each platform.
[0104] In another possible embodiment, multiple robot platforms with different robot structure information and motion constraint information can be used to jointly construct a multi-robot hybrid dataset, and the same motion generation model can be trained using the multi-robot hybrid dataset.
[0105] For example, robot platform A can be a six-degree-of-freedom robotic arm, whose motion dimensions include six joint motion dimensions and one gripper motion dimension, and the motion constraint information includes joint position constraints; robot platform B can be a seven-degree-of-freedom robotic arm, whose motion dimensions include seven joint motion dimensions and one gripper motion dimension, and has redundant degree-of-freedom constraints; robot platform C can be a mobile manipulation robot, whose motion dimensions include six robotic arm joint motion dimensions, two mobile base motion dimensions and one gripper motion dimension, and the motion constraint information also includes mobile base motion constraints.
[0106] By conditionally encoding the robot structural information of each robot platform, different robot platforms can be mapped to corresponding robot conditional features. These robot conditional features are then used to conditionalize the granularity selection network and the conditional motion decoder, enabling the same motion generation model to distinguish the motion dimensions and movement capabilities of different robot platforms. When deployed to a new robot platform, the motion generation model can be provided with the robot structural and motion constraint information of the new platform, and adapted using a small amount of demonstration data, without needing to train an independent model from scratch for the new robot platform.
[0107] In one adaptation example, model adaptation can be completed using demonstration data from approximately 200 new robot platforms. Compared to performing full training for each robot platform separately, the amount of adaptation data required can be reduced by approximately 80% to 85%.
[0108] In one possible embodiment, the method steps shown in S303 are implemented by S3031 to S3034, which are described in detail below.
[0109] S3031. Input the task complexity features and robot condition features into the granularity selection network.
[0110] Specifically, the task complexity features and robot condition features can first be concatenated, weighted fused, or attention-based fused to form granularity-selection input features. These granularity-selection input features are then input into a granularity-selection network. The granularity-selection network can employ a multilayer perceptron, a fully connected neural network, a Transformer network, or other network structures with multidimensional classification capabilities.
[0111] Task complexity features characterize the spatial accuracy requirements, time constraints, characteristics of the manipulated object, and environmental constraints corresponding to the target task; robot condition features characterize the number of degrees of freedom, joint composition, joint range of motion, and kinematic type of the current robot. The granularity selection network, by simultaneously processing task complexity features and robot condition features, can combine the action requirements of the target task with the actual structure of the current robot to perform granularity selection for different action dimensions.
[0112] In one implementation, the effective motion dimensions of the current robot can be determined based on the robot's conditional characteristics, and motion dimensions that do not belong to the current robot can be masked, so that the granularity selection network outputs granularity selection results only for the motion dimensions that the current robot actually has.
[0113] In one implementation, the process by which the granularity selection network determines the discretization granularity for the iii-th action dimension can be represented as: ;in, This represents the discretization granularity corresponding to the i-th action dimension; i represents the index of the action dimension. R represents the granularity selection function corresponding to the granularity selection network; R represents the robot conditional features obtained by conditionally encoding the robot structure information; set This indicates multiple preset granularity levels.
[0114] The granularity selection network can perform the granularity selection process separately for each action dimension of the robot, outputting the granularity selection result for each action dimension. The granularity selection result can be a granularity level index corresponding to a preset granularity level, or a granularity selection score or probability corresponding to each preset granularity level. Based on the granularity selection result, the discretized granularity corresponding to the i-th action dimension is determined from multiple preset granularity levels. .
[0115] In one example, when the task complexity characteristics indicate that the target task has a high spatial accuracy requirement, a granularity of 128 or 256 can be determined for the action dimension related to fine position or attitude adjustment; when the spatial accuracy requirement of the target task is low, a granularity of 16 or 32 can be determined for the corresponding action dimension; when the target task includes both coarse movement and fine operation, different discretization granularities can be determined for different action dimensions to form a mixed granularity configuration.
[0116] For a robot with D effective action dimensions, the granularity configuration can be expressed as: ;in, Indicates granularity configuration; to These represent the discretization granularity corresponding to each action dimension, determined according to the order of action dimensions. Each discretization granularity in the granularity configuration can be further used as a codebook index for a multi-granularity action codebook library to determine the target action codebook corresponding to each action dimension.
[0117] S3032. Using the granularity selection network, output the granularity selection results corresponding to each action dimension.
[0118] Specifically, the granularity selection network can set corresponding granularity output units for each action dimension of the current robot. Each granularity output unit is used to output the correspondence between the action dimension and multiple preset granularity levels. The granularity selection result can be represented in the form of granularity level index, granularity selection score, or granularity selection probability.
[0119] When the granularity selection result adopts the granularity selection probability, the granularity output unit corresponding to each action dimension can output the probability of that action dimension adopting each preset granularity level. For example, when there are multiple preset granularity levels including 16, 32, 64, 128 and 256, the granularity selection result corresponding to an action dimension can include five granularity selection probabilities corresponding to the above five preset granularity levels.
[0120] The granularity selection results for each motion dimension are independent of each other. Therefore, different motion dimensions in the same target task can output different granularity selection results. For example, in a fine-grained insertion task, joint motion dimensions related to end-effector position and attitude adjustment can tend to have a higher preset granularity level, while gripper opening and closing motion dimensions or motion dimensions used for large-range movements can tend to have a lower preset granularity level.
[0121] S3033. Based on the granularity selection results corresponding to each action dimension, determine the discretization granularity corresponding to each action dimension from multiple preset granularity levels.
[0122] Specifically, for each action dimension, a preset granularity level can be selected from multiple preset granularity levels as the discretization granularity corresponding to that action dimension, based on the granularity selection result corresponding to that action dimension.
[0123] When the granularity selection result is a granularity level index, the preset granularity level indicated by the granularity level index can be determined as the discretized granularity of the corresponding action dimension; when the granularity selection result is a granularity selection score or a granularity selection probability, the preset granularity level with the highest granularity selection score or granularity selection probability can be determined as the discretized granularity of the corresponding action dimension.
[0124] In other implementations, the granularity selection result can be modified according to a preset selection threshold, task stage, or granularity configuration of adjacent control cycles to avoid frequent switching of discretized granularity between adjacent control cycles.
[0125] Multiple preset granularity levels can include 16, 32, 64, 128, and 256 granularity levels. Lower preset granularity levels can be used for large-scale movements, coarse positioning, or motion dimensions with lower precision requirements; higher preset granularity levels can be used for assembly, insertion, welding, or motion dimensions with higher precision requirements. For target tasks that simultaneously involve coarse movement and fine manipulation, different discretization granularities can be determined for different motion dimensions. The above granularity levels are merely examples; in practical applications, other preset granularity levels can be set according to the range of continuous motion values, robot control precision, and computational resources.
[0126] S3034. According to the order of the action dimensions, combine the discretized granularities corresponding to each action dimension to obtain the granularity configuration.
[0127] Specifically, the discretization granularity corresponding to each action dimension can be arranged and combined according to the predefined order of action dimensions in the robot control interface, robot description file, or action dimension set to form a granularity configuration. Each configuration position in the granularity configuration corresponds to an action dimension, and the discretization granularity used for that action dimension is recorded.
[0128] For example, for a robot with six joint motion dimensions and one gripper opening / closing motion dimension, the discretization granularity corresponding to the seven motion dimensions can be arranged according to the order of the first to sixth joints and the gripper opening / closing motion. If the first to third joints use 32 granularity, the fourth to sixth joints use 128 granularity, and the gripper opening / closing motion dimension uses 16 granularity, then the corresponding granularity configuration can be formed according to the above-mentioned motion dimension arrangement order.
[0129] Granularity configuration can be represented in the form of granularity level sequence, granularity index sequence, or granularity conditional features, and is passed to the subsequent target action codebook determination process and conditional action decoding process, so that each action dimension can adopt a target action codebook and action tag set that match its discretization granularity.
[0130] In this embodiment, by inputting both task complexity features and robot condition features into the granularity selection network, the granularity selection process can simultaneously consider the action requirements of the target task and the current structural characteristics of the robot. By outputting granularity selection results separately for each action dimension and independently determining the discretized granularity, fine-grained action dimensions can be represented with higher precision, while coarse-grained action dimensions can be represented with lower precision. Furthermore, by forming a granularity configuration according to the order of action dimensions, a clear granularity correspondence can be provided for subsequent target action codebook selection and action tag generation. Therefore, this implementation can avoid insufficient precision or wasted computational resources caused by using the same fixed granularity for all action dimensions, while balancing the accuracy and processing efficiency of robot action representation.
[0131] In one specific application embodiment, the robot can be a six-DOF robotic arm and gripper, with its motion representation set to eight motion dimensions, including six joint motion dimensions, one gripper opening and closing motion dimension, and one reserved motion dimension; the preset granularity levels corresponding to the multi-granularity motion codebook library include 16 granularity, 32 granularity, 64 granularity, 128 granularity, and 256 granularity. The target task is to insert a part into a slot, and the visual observation information is acquired by a color depth camera.
[0132] In the coarse localization phase of the target task, the discretization granularity corresponding to the first to third joints can be set to 32 granularities, the discretization granularity corresponding to the fourth to sixth joints can be set to 32 granularities, and the discretization granularity corresponding to the gripper opening and closing motion dimension can be set to 16 granularities. In the fine insertion phase, the discretization granularity corresponding to the first to third joints can be adjusted to 128 granularities, the discretization granularity corresponding to the fourth to sixth joints can be adjusted to 256 granularities, and the discretization granularity corresponding to the gripper opening and closing motion dimension can be adjusted to 64 granularities.
[0133] In the corresponding simulation tests, the coarse localization stage generates eight action tags with an inference time of approximately 45 milliseconds; the fine-grained socket stage generates twelve action tags with an inference time of approximately 65 milliseconds; the average control frequency is approximately 14 Hz. Compared to a fixed granularity of 256 for all action dimensions, the number of action tags is reduced by approximately 35%, and the inference time is reduced by approximately 42%; compared to a fixed granularity of 32 for all action dimensions, the socket success rate is increased from approximately 68% to approximately 87%.
[0134] In one possible embodiment, the method steps shown in S104 are implemented by S1041 and S1042. The multi-granularity action codebook library and S1041 and S1042 are described in detail below.
[0135] The multi-granularity motion codebook library includes multiple motion codebooks corresponding to different preset granularity levels. Each motion codebook is used to discretize the range of continuous motion values according to its corresponding preset granularity level. The preset granularity level characterizes the number of continuous motion intervals obtained by dividing the range of continuous motion values. For example, preset granularity levels may include 16, 32, 64, 128, and 256, and the corresponding motion codebooks may include 16, 32, 64, 128, and 256 motion markers, respectively. The above preset granularity levels are only examples; other preset granularity levels can be configured according to the robot's control precision, the range of continuous motion values, and computational resources.
[0136] In one example, multiple preset granularity levels can include 16 granularity, 32 granularity, 64 granularity, 128 granularity, and 256 granularity. Taking a joint angle motion dimension with a continuous motion value range of 0 to 180 degrees as an example, 16 granularity, 32 granularity, 64 granularity, 128 granularity, and 256 granularity correspond to continuous motion interval widths of approximately 11.25 degrees, 5.625 degrees, 2.8125 degrees, 1.40625 degrees, and 0.703125 degrees, respectively. Specifically, 16 granularity can be used for large-scale movement or navigation, 32 granularity for coarse positioning or obstacle avoidance, 64 granularity for general grasping or placement, 128 granularity for fine manipulation or assembly, and 256 granularity for precision insertion or welding. The above continuous motion interval widths and applicable tasks are only examples; in actual applications, settings can be made according to the continuous motion value range and control precision requirements of each motion dimension.
[0137] For each action codebook, the range of continuous action values can be divided into multiple continuous action intervals, consistent with the number of discrete intervals represented by the corresponding preset granularity level. A unique action tag is assigned to each continuous action interval, thus establishing a mapping relationship between the action tag and the continuous action interval. The range of continuous action values can be determined based on the actual action range of the corresponding action dimension; for example, it can be a joint angle range, a joint speed range, a gripper opening / closing range, or a moving base speed range.
[0138] In one implementation, the range of continuous action values can be evenly divided into multiple continuous action intervals using an equally spaced division method. For example, for an action dimension with a continuous action value range of 0 to 180 degrees, when using 16 granularity, the action range corresponding to each continuous action interval is larger than that corresponding to each continuous action interval when using 256 granularity. Therefore, the higher the preset granularity level, the smaller the continuous action interval corresponding to a single action marker, and the higher the resolution of the action representation. In other implementations, the range of continuous action values can also be non-uniformly divided according to the distribution of action values in the training data, so that areas with denser action value distribution correspond to smaller continuous action intervals.
[0139] To address situations where different motion dimensions have different ranges of continuous motion values, the continuous motion values can be normalized according to the range of continuous motion values for each motion dimension. Then, a mapping relationship between the normalized motion values and motion tags can be established using the corresponding motion codebook. Alternatively, motion codebooks with the same preset granularity level but different ranges of continuous motion values can be configured for different motion dimensions. This allows for the representation of different types of continuous motion values, such as joint angles, gripper opening / closing amounts, and movement speeds, within a unified preset granularity level system.
[0140] During the training phase or when it is necessary to convert known continuous actions into action labels, for the continuous action value of the i-th action dimension, the corresponding action label can be determined according to the following formula: ;in, This represents the index of the action tag corresponding to the i-th action dimension; Represents the continuous action value of the i-th action dimension; and These represent the lower bound and upper bound of the values for consecutive actions in the i-th action dimension, respectively. Indicates the discretization granularity corresponding to the i-th action dimension; symbol This indicates rounding down. Based on the calculated action tag index, the corresponding action tag can be determined from the target action codebook.
[0141] To ensure that the action tag index is within the set of action tags defined by the target action codebook, the action tag index can be restricted to 0 to... Within the range. When the value of continuous action equals the upper limit of the value of continuous action, the calculation result reaches... At that time, the action tag index can be corrected to .
[0142] S1041. Obtain the discretized granularity corresponding to each action dimension from the granularity configuration.
[0143] Specifically, the granularity configuration records the discretized granularity corresponding to each action dimension according to the order of their arrangement. The configuration positions in the granularity configuration can be read according to the order of action dimensions in the robot control interface, robot description file, or preset action dimension set, and the action dimension and its corresponding discretized granularity can be determined for each configuration position.
[0144] For example, for a robot that includes joints one through six and the gripper opening / closing motion dimension, the first to sixth configuration positions in the granularity configuration can correspond to joints one through six respectively, and the seventh configuration position can correspond to the gripper opening / closing motion dimension. Based on the preset granularity level recorded at each configuration position, the discretized granularity corresponding to each motion dimension can be obtained.
[0145] When obtaining the discretized granularity corresponding to each action dimension, it is also possible to verify whether the number of action dimensions in the granularity configuration is consistent with the number of effective action dimensions of the current robot, and to mask invalid action dimensions based on robot condition information, thereby ensuring that the granularity configuration corresponds to the action space of the current robot.
[0146] S1042. For each action dimension, determine the action codebook corresponding to the discretization granularity of the action dimension from multiple action codebooks, and use the determined action codebook as the target action codebook corresponding to the action dimension.
[0147] Specifically, each action codebook in the multi-granularity action codebook library can be assigned a granularity level identifier, which indicates the preset granularity level corresponding to the action codebook. For each action dimension, the discretized granularity corresponding to that action dimension can be matched with the granularity level identifier of each action codebook, and the action codebook whose granularity level identifier matches the discretized granularity is determined as the target action codebook corresponding to that action dimension.
[0148] For example, when the discretization granularity corresponding to the first action dimension is 32, the action codebook with granularity of 32 can be determined as the target action codebook corresponding to the first action dimension; when the discretization granularity corresponding to the second action dimension is 128, the action codebook with granularity of 128 can be determined as the target action codebook corresponding to the second action dimension. Thus, different action dimensions in the same robot can use target action codebooks with different preset granularity levels.
[0149] The target action codebook is used to define the set of action tags that can be used for a given action dimension, and records the mapping relationship between each action tag and the corresponding continuous action interval. During the subsequent action tag generation process, the conditional action decoder can generate action tags from the set of action tags defined by the target action codebook; during the action reconstruction process, it can query the continuous action interval corresponding to the action tag based on the target action codebook and convert the action tag into a continuous action value.
[0150] In this embodiment, by setting multiple action codebooks corresponding to different preset granularity levels in a multi-granularity action codebook library, and determining the target action codebook for each action dimension according to the granularity configuration, different action dimensions can adopt action representation resolutions that match their action requirements. Target action codebooks corresponding to higher discretization granularity can improve the representation accuracy of fine action dimensions, while target action codebooks corresponding to lower discretization granularity can reduce the representation complexity of coarse action dimensions. Simultaneously, the target action codebook provides a unified bidirectional mapping basis for action tag generation and continuous action reconstruction, thereby improving the accuracy of the conversion between discrete representation of action space and continuous action control.
[0151] In one possible embodiment, the method steps shown in S105 are implemented by S1051 to S1055, which are described in detail below.
[0152] S1051. Embedding and encoding the robot condition information and granular configuration respectively, to obtain robot condition embedding features and granular condition embedding features.
[0153] Specifically, the number of degrees of freedom, joint types, joint range of motion, kinematic type, and motion constraint information in the robot's conditional information can first be numerically processed. Then, the numerically processed robot conditional information is mapped to a preset feature space through a robot conditional embedding layer to obtain robot conditional embedding features. The robot conditional embedding features are used to indicate the current robot's motion dimension composition, structural type, and motion capabilities to the conditional motion decoder.
[0154] For granularity configuration, the discrete granularity corresponding to each action dimension can be read according to the order of the action dimensions, and different preset granularity levels can be converted into corresponding granularity identifiers. Subsequently, each granularity identifier is embedded and encoded through a granularity conditional embedding layer to obtain granularity conditional embedding features. Different feature positions in the granularity conditional embedding features correspond to different action dimensions, which are used to represent the discrete granularity adopted for each action dimension.
[0155] Robot conditional embedding layers and granular conditional embedding layers can be implemented using embedding matrices, fully connected networks, or multilayer perceptrons. Through embedding encoding, robot conditional information and granular configurations in different data formats can be converted into feature representations that can be uniformly processed by the conditional action decoder.
[0156] S1052. Use multimodal fusion features, robot conditional embedding features, and granular conditional embedding features as conditional inputs to the conditional action decoder.
[0157] Specifically, multimodal fusion features, robot conditional embedding features, and granular conditional embedding features can be concatenated, weighted, or attention-based fused to obtain conditional inputs; alternatively, the three types of features can be used as inputs to different attention layers in the conditional action decoder, so that the conditional action decoder focuses on the target task, the current robot, and the granular configuration respectively when generating action tags.
[0158] Among them, multimodal fusion features are used to characterize the semantic requirements of the current working environment and the target task; robot conditional embedding features are used to characterize the current robot's motion space and motion capabilities; and granular conditional embedding features are used to characterize the discretization granularity corresponding to each action dimension. Thus, the conditional action decoder can perform action label prediction by combining task semantics, robot structure, and the action representation accuracy of each action dimension during the same action generation process.
[0159] The conditional action decoder can be implemented using a Transformer decoder, a recurrent neural network, or other decoding networks with sequence generation capabilities. In one implementation, the aforementioned conditional inputs can be used as keys and values for a cross-attention layer, and the features corresponding to the generated action tag sequence can be used as inputs for a self-attention layer to achieve conditional action tag generation.
[0160] In one implementation, the conditional action decoder generates an action tag sequence in an autoregressive manner, and its generation relationship can be expressed as follows: ;in, This represents the sequence of action tags from the first action tag to the Tth action tag; T represents the length of the action tag sequence. This represents the t-th action tag to be generated; This represents the sequence of action tags generated before the t-th action tag; F represents the multimodal fusion feature; R represents the robot conditional embedding feature; and G represents the granular conditional embedding feature. This represents the conditional probability of generating the t-th action tag under the conditions of multimodal fusion features, robot conditional embedding features, granular conditional embedding features, and an already generated action tag sequence.
[0161] The conditional action decoder determines each action tag sequentially based on the above conditional probabilities, so that the generation of the current action tag can utilize the action timing information contained in the previously generated action tags, thereby maintaining the correlation between adjacent actions in the action tag sequence.
[0162] S1053. According to the preset action tag generation order, determine the action dimension corresponding to the current action tag to be generated, and determine the corresponding target action codebook based on the action dimension.
[0163] The preset motion marker generation order specifies the arrangement of motion markers corresponding to different motion times and different motion dimensions in the motion marker sequence. The preset motion marker generation order can be arranged first according to the motion time sequence, and then according to the motion dimension within each motion time; or it can be determined according to the motion data format specified by the robot control interface.
[0164] For example, for a robot with multiple joint motion dimensions and gripper opening / closing motion dimensions, motion markers corresponding to each joint motion dimension and gripper opening / closing motion dimension can be generated sequentially within a single motion moment, and then the motion markers for the next motion moment can be generated. Based on the current generation position within the preset motion marker generation order, the motion moment and motion dimension corresponding to the current motion marker to be generated can be determined.
[0165] After determining the action dimension corresponding to the action tag to be generated, the discretization granularity corresponding to the action dimension can be determined based on the correspondence between the action dimension and the granularity configuration, and the target action codebook corresponding to the action dimension can be further determined. The target action codebook is used to limit the set of action tags that the action tag to be generated can use, thereby avoiding the generation of action tags that do not match the action dimension or the corresponding discretization granularity.
[0166] S1054. Based on the conditional input and the generated action tag sequence, predict the current action tag to be generated from the action tag set defined by the target action codebook.
[0167] Specifically, the generated action tag sequence can be converted into an action tag embedding sequence and input into the conditional action decoder. The conditional action decoder determines the action tag prediction result corresponding to the current generation position based on the historical action information contained in the generated action tag sequence, as well as multimodal fusion features, robot conditional embedding features, and granular conditional embedding features.
[0168] The action tag prediction result can include the prediction score or prediction probability of each action tag defined by the target action codebook. The action tag with the highest prediction score or prediction probability can be determined as the current action tag to be generated, or the current action tag to be generated can be determined from multiple candidate action tags according to a preset sampling strategy.
[0169] During the prediction process, the output space of the conditional action decoder can be masked using the target action codebook, preventing action tags that do not belong to the target action codebook from participating in the prediction of the current action dimension. Therefore, even if different action dimensions use different discretization granularities, the conditional action decoder can still generate valid action tags from the target action codebook corresponding to each action dimension.
[0170] The generated action tag sequence can include all action tags generated in the current prediction time domain, as well as historical action tags generated or executed in the previous control cycle. By utilizing the generated action tag sequence, the conditional action decoder can maintain the action correlation between adjacent action moments and different action dimensions, reducing the possibility of discontinuous or conflicting actions in the action tag sequence.
[0171] S1055. Add the predicted action tags to the generated action tag sequence, and iteratively perform action tag prediction based on the updated generated action tag sequence until the preset generation termination condition is met, and obtain the action tag sequence.
[0172] Specifically, after obtaining the current action tag to be generated, the action tag can be added to the sequence of generated action tags according to the preset action tag generation order, and the action tag generation position can be moved to the next position. Subsequently, the action dimension and target action codebook corresponding to the next action tag to be generated are re-determined, and the next action tag is predicted based on the updated sequence of generated action tags.
[0173] The action label prediction is performed iteratively in the manner described above, so that the generation of subsequent action labels can be based on the action labels that have been generated previously, thus forming an autoregressive action label generation process.
[0174] The preset termination conditions may include at least one of the following: the number of generated action tags has reached the preset sequence length, the generation of action tags for all action dimensions in the preset prediction time domain has been completed, the generation of a preset termination tag has been completed, or the conditional action decoder has determined that the current action stage corresponding to the target task has been completed.
[0175] For example, when a prediction time domain includes multiple action moments, and each action moment includes multiple effective action dimensions, the preset generation termination condition can be determined after all action moments and action tags corresponding to effective action dimensions have been generated. The final action tag sequence records each action tag according to the action time sequence and action dimension, and matches them with the target action codebook corresponding to each action dimension.
[0176] In this embodiment, by embedding and encoding robot conditional information and granular configuration, and using them together with multimodal fusion features as conditional inputs to the conditional action decoder, the action tag generation process can simultaneously adapt to the target task, robot structure, and the discretized granularity of each action dimension. By determining the current action dimension according to a preset action tag generation order and predicting from the action tag set defined by the target action codebook corresponding to that action dimension, the mixing of action tags of different action dimensions or different granularity levels can be avoided. Furthermore, by iteratively predicting subsequent action tags based on the generated action tag sequence, the temporal correlation between action tags can be maintained. Therefore, this implementation can improve the matching degree between the action tag sequence and the target task and robot action space, while taking into account the continuity, accuracy, and cross-robot platform applicability of action generation.
[0177] In one possible embodiment, the method steps shown in S105 can also be implemented by S105A1 to S105A4, which are described in detail below.
[0178] In this embodiment, the conditional action decoder includes a high-level decoder and a low-level decoder. The high-level decoder decomposes the target task at the task semantic level and determines the task sub-objectives that the robot needs to achieve sequentially to complete the target task; the low-level decoder generates action tag sub-sequences that can be executed by the robot based on each task sub-objective. The high-level decoder and the low-level decoder can be implemented using a Transformer decoder, a recurrent neural network, or other neural networks with sequence generation capabilities, respectively.
[0179] Figure 3 This is a schematic diagram of a hierarchical action prediction architecture provided in an embodiment of this application. Figure 3 As shown, the conditional action decoder includes a high-level decoder and a low-level decoder, wherein the high-level decoder operates at a lower frequency than the low-level decoder.
[0180] Multimodal fusion features and robot conditional information are input into a high-level decoder. The high-level decoder decomposes the target task based on the task semantics, the current working environment, and the robot's motion capabilities, generating a sequence of task sub-targets arranged in the execution order. For example, for the target task of "go to the kitchen to get a glass of water," the sequence of task sub-targets could include, in sequence, navigating to the kitchen, locating the glass, grabbing the glass, navigating to the living room, and placing the glass.
[0181] During task execution, the current task sub-objective, multimodal fusion features, robot condition information, and granularity configuration are sequentially input into the low-level decoder according to the order of the task sub-objectives in the task sub-objective sequence. The low-level decoder determines the action dimensions required to complete the task sub-objective based on the current task sub-objective, and determines the target action codebook corresponding to each action dimension based on the granularity configuration.
[0182] For each task sub-objective, the low-level decoder generates a sequence of motion markers from the set of motion markers defined by the corresponding target action codebook. For example, the sequence of motion markers for navigating to the kitchen may include translation and rotation motion markers for the moving base, and the sequence of motion markers for grabbing a water cup may include joint motion markers for the robotic arm and opening and closing motion markers for the gripper.
[0183] After obtaining the action marker sub-sequences corresponding to each task sub-objective, the action marker sub-sequences are combined according to the order of each task sub-objective in the task sub-objective sequence to obtain the action marker sequence corresponding to the target task. The high-level decoder can update the task sub-objective sequence at an operating frequency of 2Hz, and the low-level decoder can generate the action marker sub-sequences at an operating frequency of 15Hz. Thus, by combining low-frequency task planning and high-frequency action generation, the efficiency of action generation and real-time control capability of multi-step target tasks are improved.
[0184] S105A1: Input the multimodal fusion features and robot condition information into the high-level decoder to generate the task sub-target sequence corresponding to the target task.
[0185] Specifically, multimodal fusion features are used to characterize the task semantics, operation objects, target location, and current working environment of the target task, while robot condition information is used to characterize the current robot's structural composition, action dimensions, and motion capabilities. Based on the multimodal fusion features and robot condition information, the high-level decoder performs semantic-level action planning for the target task, decomposing the target task into multiple task sub-objectives arranged according to their execution sequence, thus obtaining a sequence of task sub-objectives.
[0186] Task sub-objectives describe the phased results of a task during its execution, with a granularity higher than that of specific joint movements. For example, for the task "go to the kitchen to get a water glass and place it in the designated location," the sequence of task sub-objectives could sequentially include navigating to the kitchen, locating the water glass, grabbing the water glass, moving to the designated location, and placing the water glass. For a part insertion task, the sequence of task sub-objectives could include approaching the part, grabbing the part, moving to the slot, adjusting the part's orientation, and performing the insertion.
[0187] Each task sub-objective in the task sub-objective sequence can be represented in the form of semantic tags, task state vectors, or sub-objective features, and includes sub-objective position identifiers to characterize the execution order. The high-level decoder can also combine robot condition information to determine whether the current robot has the action capability required to complete the corresponding task sub-objective, avoiding the generation of task sub-objectives that are beyond the current robot's functional range.
[0188] S105A2. According to the order of each task sub-objective in the task sub-objective sequence, input each task sub-objective, multimodal fusion feature, robot condition information and granular configuration into the low-level decoder in sequence.
[0189] Specifically, task sub-objectives can be read sequentially from the task sub-objective sequence according to their arrangement, and the current task sub-objective can be converted into corresponding sub-objective conditional features. The sub-objective conditional features, multimodal fusion features, robot conditional information, and granularity configuration are input into the low-level decoder, so that the low-level decoder considers the current stage task results to be achieved, the working environment, the robot structure, and the discretized granularity corresponding to each action dimension when generating specific actions.
[0190] Among them, the current task sub-objective is used to limit the phased results that need to be achieved in this action generation; the multimodal fusion features are used to provide the target object state and environment state related to the current task sub-objective; the robot condition information is used to limit the effective action dimensions and motion capabilities of the current robot; and the granularity configuration is used to indicate the discretization granularity adopted for each action dimension.
[0191] In one implementation, before inputting into the low-level decoder, the motion dimensions required to complete the current task sub-objective can be determined based on the current task sub-objective and robot condition information, and these motion dimensions can be identified as the motion dimensions corresponding to the current task sub-objective. For example, the navigation task sub-objective can correspond to the translation and rotation motion dimensions of the mobile base, and the grasping task sub-objective can correspond to the joint motion dimensions of the robotic arm and the opening and closing motion dimensions of the gripper. This avoids the low-level decoder generating unnecessary motion labels for motion dimensions unrelated to the current task sub-objective.
[0192] S105A3: Through the low-level decoder, for each task sub-objective, generate an action tag sub-sequence that matches the target action codebook for each action dimension corresponding to the task sub-objective.
[0193] Specifically, the low-level decoder determines the action dimensions corresponding to the current task sub-objective based on the current task sub-objective, multimodal fusion features, robot condition information, and granularity configuration, and determines the corresponding target action codebook based on the discretization granularity corresponding to each action dimension. Subsequently, the low-level decoder generates action tags from the action tag set defined by each target action codebook to form an action tag subsequence for achieving the current task sub-objective.
[0194] The action tag subsequence may include action tags corresponding to one or more control moments, and each action tag is recorded according to the action sequence and action dimension. The low-level decoder may adopt a per-action tag generation method, or it may predict multiple action tags corresponding to the same control moment simultaneously while satisfying the target action codebook constraints for each action dimension. This embodiment does not limit this.
[0195] For example, under the task sub-objective of grasping a water cup, the low-level decoder can generate the motion marker sub-sequence required for the robotic arm to approach the water cup, adjust the end effector posture, and control the gripper closure; under the task sub-objective of navigating to a specified location, the low-level decoder can generate the motion marker sub-sequence required for the translation and rotation of the mobile base. Each motion marker is determined from the motion marker set defined by the target motion codebook of the corresponding motion dimension, thus ensuring that the motion marker sub-sequence matches the granularity configuration.
[0196] In one example, the higher-level decoder can generate or update task sub-objectives at an operating frequency of 1Hz to 5Hz, while the lower-level decoder can generate action marker sub-sequences at an operating frequency of 10Hz to 20Hz. For example, the higher-level decoder could operate at 2Hz, and the lower-level decoder at 15Hz. These frequencies are merely examples; the actual operating frequencies can be set according to the complexity of the target task, the robot control cycle, and computational resources.
[0197] S105A4. According to the order of the task sub-objectives, combine the action tag sub-sequences corresponding to each task sub-objective to obtain the action tag sequence.
[0198] Specifically, the action marker sub-sequences generated by the low-level decoder for each task sub-target can be sequentially concatenated according to the arrangement order of the task sub-targets in the task sub-target sequence to form the action marker sequence corresponding to the target task. During the combination process, the temporal order and dimensional arrangement relationship of each action marker in the corresponding action marker sub-sequence can be preserved so that the combined action marker sequence can reflect the complete execution process of the target task.
[0199] In one implementation, sub-target boundary markers can be set between adjacent action marker sub-sequences to distinguish action markers corresponding to different task sub-targets. When a task sub-target is completed, the system can switch to the action marker sub-sequence corresponding to the next task sub-target. When the execution result is inconsistent with the current task sub-target, the higher-level decoder can also redetermine the subsequent task sub-targets based on the updated multimodal fusion features, thereby adjusting the action marker sub-sequences that have not yet been executed.
[0200] The higher-level decoder operates at a lower frequency than the lower-level decoder, allowing the higher-level decoder to process task understanding and decomposition at a relatively lower frequency, while the lower-level decoder generates action tag subsequences adapted to the robot control cycle at a relatively higher frequency. This reduces the computational overhead of repeatedly performing task semantic planning at each control moment, while ensuring that specific robot actions are updated in a timely manner.
[0201] In one implementation, a closed-loop update relationship based on execution feedback can be formed between the high-level decoder and the low-level decoder. The low-level decoder generates a sequence of action markers based on the current task sub-objective and obtains the task execution result after the robot executes the corresponding robot control action sequence. The task execution result can be fed back to the high-level decoder, which determines whether the current task sub-objective has been completed based on the task execution result. When the current task sub-objective has been completed, the high-level decoder outputs the next task sub-objective; when the current task sub-objective has not been completed or the execution state has changed, the high-level decoder can regenerate the current task sub-objective or adjust the subsequent task sub-objective sequence based on the updated multimodal fusion features.
[0202] High-level and low-level decoders can also be deployed collaboratively in the cloud and at the edge. High-level decoders can be deployed on cloud servers or task computers with high computing power to perform task semantic understanding and task sub-objective generation at a lower frequency. Low-level decoders can be deployed on edge computing devices on the robot body to generate motion tag sub-sequences at a higher frequency. This reduces the computational burden on the robot body's high-level semantics while ensuring that low-level motion generation meets the robot's real-time control requirements.
[0203] In a simulation example of a service robot performing a multi-step task, the high-level decoder operates at 2Hz and the low-level decoder operates at 15Hz, generating 15 to 20 action markers for each task sub-objective. Compared to methods without hierarchical action prediction, the overall completion time of the target task is reduced by approximately 30%, and the cumulative action error in long sequence target tasks is reduced by approximately 45%.
[0204] In this embodiment, the target task is decomposed into sub-tasks arranged in execution order by a high-level decoder, and a low-level decoder generates action tag sub-sequences that match the target action codebook of the corresponding action dimension for each sub-task. This allows for hierarchical processing of task semantic planning and specific action generation. Since the high-level decoder operates at a lower frequency, redundant computations at the task semantic level are reduced, while the low-level decoder operates at a higher frequency, meeting the update requirements of real-time robot motion control. Furthermore, combining the action tag sub-sequences according to the arrangement order of the sub-tasks maintains the logical continuity between execution stages in complex target tasks. Therefore, this implementation improves the action generation efficiency of multi-stage and long-sequence target tasks, reduces error accumulation during long-sequence action generation, and balances the integrity of task planning with the real-time performance of robot motion control.
[0205] In one possible embodiment, the method steps shown in S106 and S107 are implemented by S1061, S1062 and S1071 to S1073, which are described in detail below.
[0206] S1061. For each action marker in the action marker sequence, determine the continuous action interval corresponding to the action marker based on the action dimension corresponding to the action marker and the target action codebook corresponding to the action dimension.
[0207] Specifically, the action tokens in the action token sequence are arranged according to a preset action token generation order. Therefore, the action timing and action dimension corresponding to an action token can be determined based on its position in the action token sequence. For each action token, the target action codebook corresponding to that action dimension can be obtained based on its corresponding action dimension, and the continuous action interval mapped by that action token can be queried in the target action codebook.
[0208] The target action codebook records the mapping relationships between multiple action tags and multiple consecutive action intervals. For a target action codebook with a higher discretization granularity, the range of consecutive action values is divided into a larger number of consecutive action intervals with smaller ranges; for a target action codebook with a lower discretization granularity, the range of consecutive action values is divided into a smaller number of consecutive action intervals with larger ranges.
[0209] When the target action codebook is established based on the normalized continuous action value range, the normalized continuous action interval corresponding to the action tag can be determined first. Then, based on the actual continuous action value range of the corresponding action dimension, the normalized continuous action interval can be de-normalized to obtain the continuous action interval under that action dimension. For example, for the joint angle action dimension, the normalized continuous action interval can be converted into a joint angle interval; for the gripper opening and closing action dimension, it can be converted into a gripper opening and closing amount interval.
[0210] S1062. Based on the representative action values corresponding to the continuous action intervals, determine the continuous action values corresponding to each action marker, and combine the continuous action values according to the action sequence and action dimension to obtain a continuous action sequence.
[0211] Specifically, a representative action value can be assigned to each consecutive action interval in the target action codebook. The representative action value can be the midpoint value, center value, typical action value determined based on training data, or the center value of the codebook obtained through codebook training.
[0212] For each action marker, the representative action value corresponding to the continuous action interval mapped by the action marker can be determined as the continuous action value corresponding to the action marker. If the representative action value is represented in a normalized form, it can be denormalized according to the actual continuous action value range of the corresponding action dimension to obtain the continuous action value that can be recognized by the robot controller.
[0213] After determining the consecutive action values corresponding to each action marker, these consecutive action values can be arranged and combined according to the action sequence and action dimension corresponding to each action marker. Specifically, consecutive action values corresponding to different action dimensions at the same action moment can first be combined into a consecutive action vector, and then multiple consecutive action vectors can be arranged according to the action sequence to obtain a consecutive action sequence.
[0214] For example, for a robot with six joint motion dimensions and one gripper opening / closing motion dimension, a continuous motion vector can sequentially include the continuous motion values of the first to sixth joints and the gripper opening / closing motion value. Multiple continuous motion vectors are arranged according to the control time sequence to form a continuous motion sequence within a preset prediction time domain.
[0215] When the midpoint value of a continuous action interval is used as the representative action value, the continuous action value corresponding to the i-th action dimension can be represented as: ;in This represents the continuous action value of the i-th action dimension reconstructed from the action markers.
[0216] S1071. Determine the robot's motion constraints based on the robot's condition information, and determine the actions to be corrected in the continuous action sequence that do not meet the motion constraints based on the motion constraints.
[0217] Specifically, at least one of the following constraints can be obtained from the robot's condition information: joint position constraints, joint velocity constraints, joint acceleration constraints, and collision constraints. Based on these motion constraints, a feasible action domain for the current robot can be constructed. The feasible action domain represents the set of action values that the robot can perform under the current robot structure and operating environment.
[0218] For joint position constraints, it can be determined whether the position of each joint in a continuous motion sequence is within the allowable position range of the corresponding joint; for joint velocity constraints, the corresponding joint velocity can be determined based on the continuous motion values of the same joint at adjacent motion times and the control time interval, and it can be determined whether the joint velocity is within the allowable velocity range; for joint acceleration constraints, the joint acceleration can be determined based on the changes in the velocities of adjacent joints, and it can be determined whether the joint acceleration exceeds the allowable acceleration range.
[0219] For collision constraints, the predicted pose of the robot at each moment of a continuous action sequence can be determined, and collision detection can be performed using the robot's kinematic model and the environmental space model. When the distance between different structural components of the robot or the distance between the robot and an obstacle is less than a preset safety distance, it can be determined that the corresponding action does not meet the collision constraints.
[0220] Actions that do not satisfy at least one of the above motion constraints are identified as actions to be corrected, and the action sequence, action dimension, and violated motion constraints corresponding to the actions to be corrected are recorded. When all actions in a continuous action sequence satisfy the motion constraints, the continuous action sequence can be directly used to form a robot control action sequence.
[0221] S1072. Based on the motion constraints, project the motion to be corrected onto the corresponding movable domain of the robot to obtain the corrected motion that satisfies the motion constraints.
[0222] Specifically, the action to be corrected can be taken as the action to be projected, and an executable action corresponding to the action to be corrected can be determined in the actionable domain defined by motion constraints. This executable action is then taken as the corrected action. The corrected action satisfies the corresponding constraints among the robot's joint position constraints, joint velocity constraints, joint acceleration constraints, and collision constraints.
[0223] In one implementation, boundary correction can be performed according to the motion constraints violated by the action to be corrected. For example, when the joint position exceeds the allowable position range, the corresponding joint position can be corrected to the nearest boundary value within the allowable position range; when the joint velocity or joint acceleration exceeds the corresponding range, the continuous motion values at adjacent motion moments can be smoothly adjusted so that the adjusted joint velocity and joint acceleration satisfy the corresponding constraints.
[0224] In another implementation, a constraint optimization method can be used to perform the action domain projection. Specifically, the optimization objective can be to minimize the difference between the corrected action and the action to be corrected, and the motion constraints corresponding to the robot can be used as constraints to construct an action projection optimization problem. The action projection optimization problem is then solved to obtain a corrected action that is located within the action domain and is close to the action to be corrected.
[0225] For example, quadratic programming can be used to solve the motion projection optimization problem. For multiple related motions to be corrected, multiple motions can be jointly projected to maintain the smoothness of the continuous motion sequence in terms of motion time while satisfying motion constraints, and to avoid abrupt changes between adjacent motions caused by correcting only a single motion independently.
[0226] The projection of the action domain to be corrected can be represented as: ;in, This indicates the corrective action that satisfies the motion constraints after projection; This represents the motion constraints determined based on the robot's condition information; This represents the projection operation that maps continuous motion values to a movable action domain defined by motion constraints.
[0227] When using quadratic programming to project the action domain, the optimization objective can be minimizing the distance between the corrected action and the preceding consecutive actions, with the constraint that the corrected action lies within the action domain. The corresponding action projection optimization relationship can be expressed as: , ;in, This represents a continuous action vector reconstructed from an action tag sequence; This represents the corrected action vector after projection; This represents the feasible action domain defined by the current robot's joint position constraints, joint velocity constraints, joint acceleration constraints, and collision constraints. By solving the above motion projection optimization relationship, we can minimize the impact of motion corrections on the original motion generation result while satisfying the motion constraints.
[0228] S1073. Combine the actions and correction actions that satisfy the motion constraints in the continuous action sequence according to the corresponding action timing and action dimension to obtain the robot control action sequence.
[0229] Specifically, actions that satisfy motion constraints in a continuous sequence of actions are retained, and the corresponding actions to be corrected are replaced with corrective actions based on the action sequence and action dimension. Subsequently, the actions that satisfy motion constraints and the corrective actions are combined according to the original action sequence and action dimension arrangement to obtain the robot control action sequence.
[0230] During the assembly process, it is possible to verify whether the number of actions, the order of action dimensions, and the timing of actions in the robot control action sequence are consistent with the continuous action sequence, and to perform motion constraint verification on the robot control action sequence again to confirm that each action in the robot control action sequence is within the movable action domain.
[0231] When there are no actions to be corrected in the continuous action sequence, the actions and their arrangement in the continuous action sequence can be kept unchanged, and the continuous action sequence can be determined as the robot control action sequence. When there are one or more actions to be corrected in the continuous action sequence, the corrected complete action sequence is determined as the robot control action sequence.
[0232] Once the robot control motion sequence is obtained, it can be sent to the robot controller. The robot controller can then read each consecutive motion vector sequentially according to the control timing and convert each consecutive motion vector into control commands for the corresponding actuators, thereby driving the robot's joints, grippers, mobile base, or other actuators to complete the target task.
[0233] In this embodiment, by restoring discrete action markers to continuous action values based on the target action codebook corresponding to each action dimension, and forming a continuous action sequence according to action timing and action dimensions, the conversion between discrete action space representation and robot continuous control interface can be realized. Furthermore, by performing motion constraint verification on the continuous action sequence based on robot condition information, actions exceeding the robot's structural capabilities or potentially causing collisions can be identified and corrected. By projecting the actions to be corrected onto the movable action domain and recombining the corrected actions with actions that satisfy motion constraints, the robot control action sequence can satisfy actual motion constraints while preserving the original action generation results as much as possible. Therefore, this implementation improves the accuracy of action marker reconstruction results, avoids directly issuing unexecutable actions to the robot, and improves the executability, continuity, and safety of the robot in performing the target task.
[0234] In one possible embodiment, the update process of the multi-granularity action codebook library is implemented through S901 to S904, which are described in detail below.
[0235] S901. Obtain the task execution result generated by the robot executing the target task according to the robot control action sequence.
[0236] Specifically, during the process of the robot executing the target task according to the robot control action sequence, task execution data generated by the robot executing the target task can be collected through the robot controller, joint sensors, vision sensors, force sensors or task status detection modules, and the task execution result can be determined based on the task execution data.
[0237] Task execution results can include some or all of the following: the robot's actual motion trajectory, the actual position and orientation of each joint, the actual position and orientation of the end effector, the actual state of the manipulated object, the completion status of the target task, and whether the task execution was successful. For example, for a part insertion task, the task execution results can include the actual insertion position of the part, the actual insertion depth, the orientation deviation between the part and the slot, and whether the insertion was successful; for an object grasping task, the task execution results can include the actual grasping position of the end effector, the opening and closing state of the gripper, whether the object is stably grasped, and the actual movement position of the object.
[0238] In one implementation, task execution data can be correlated according to the action timing and action dimension in the robot control action sequence, so that each actual action state in the task execution result can be mapped to the corresponding action time, action dimension and the discretization granularity used, thereby providing data basis for subsequent determination of codebook adjustment information.
[0239] S902. Based on the difference between the task execution result and the expected task state corresponding to the target task, determine the codebook adjustment information related to the corresponding action dimension and discretization granularity.
[0240] The expected task state characterizes the state a robot should reach after completing actions according to the target task requirements. It can be determined based on task instructions, multimodal fusion features, task planning results corresponding to the target task, or pre-set task completion conditions. The expected task state may include the expected motion trajectory, the expected end-effector position and posture, the expected position and posture of the manipulated object, and some or all of the expected task completion states.
[0241] Specifically, the actual state in the task execution result can be compared with the corresponding expected task state to determine the difference between the two. The difference can include at least one of the following: position difference, posture difference, trajectory difference, speed difference, task completion status difference, or operation success rate difference.
[0242] Subsequently, based on the timing and dimension of the actions that caused the differences, the discretization granularity and target action codebook used when generating the robot control action sequence for the corresponding action dimension are queried to determine the action dimension and discretization granularity that need to be adjusted. For example, if a joint action dimension consistently produces large positional differences when using a low discretization granularity, it can be determined that the action dimension and its corresponding discretization granularity need codebook adjustment; if the continuous action values corresponding to a certain action mark produce the same direction of action deviation in multiple task executions, it can be determined that the codebook parameters corresponding to that action mark need to be corrected.
[0243] The codebook adjustment information may include some or all of the following: the codebook identifier of the action codebook to be adjusted, the action marker to be adjusted, the corresponding action dimension, the corresponding discretization granularity, the adjustment direction, and the adjustment amount. The adjustment direction indicates whether the corresponding codebook parameter needs to be increased or decreased, and the adjustment amount can be determined based on the difference between the actual state and the expected task state.
[0244] S903. If the difference meets the preset codebook update conditions, update the codebook parameters of the corresponding action codebook in the multi-granularity action codebook library according to the codebook adjustment information to obtain the updated action codebook.
[0245] The preset codebook update conditions are used to determine whether the differences generated by the current task execution are sufficient to trigger an action codebook update. The preset codebook update conditions may include at least one of the following: the difference exceeds a preset difference threshold, the number of consecutive differences in the same action dimension reaches a preset number, the cumulative difference of action execution corresponding to the same action tag reaches a preset cumulative threshold, or the execution success rate of the corresponding target task is lower than a preset success rate.
[0246] In one implementation, the differences generated in multiple task execution cycles can be statistically analyzed, and the action codebook can be updated only when similar differences are continuously generated in the same action dimension, the same discretization granularity, or the same action tag, thereby reducing the impact of sensor noise, occasional disturbances, or single execution anomalies on the action codebook.
[0247] When the differences meet the preset codebook update conditions, the action codebooks and their parameters that need to be updated in the multi-granularity action codebook library can be determined based on the codebook adjustment information. Codebook parameters may include representative action values corresponding to action tags, interval boundaries of continuous action intervals, or some or all of the codebook features used to characterize action tags.
[0248] For example, when the actual action value corresponding to a certain action mark is consistently less than the expected action value, the representative action value corresponding to the action mark can be increased within the constraints of the corresponding continuous action value range and adjacent continuous action intervals; when the difference in action execution within a certain continuous action interval is significantly greater than that between adjacent continuous action intervals, the interval boundary of the continuous action interval can be adjusted so that the corresponding action value can be mapped to a more suitable action mark.
[0249] During the update process, the maximum adjustment amount of the codebook parameters can be set at one time, and the continuous action intervals should be arranged in the order of continuous action values to avoid overlapping, inversion, or exceeding the continuous action value range of the corresponding action dimension after the update. After the codebook parameters are updated, the updated action codebook is obtained.
[0250] When the difference does not meet the preset codebook update conditions, the current action codebook can be kept unchanged, and the difference generated by this task execution can be used as statistical data for subsequent judgment on whether to trigger a codebook update.
[0251] S904. Store the updated action codebook in a multi-granularity action codebook library so that the updated action codebook can be used for the action space representation of subsequent target tasks.
[0252] Specifically, the updated action codebook can replace the original action codebook with the same codebook identifier in the multi-granularity action codebook library, or the updated action codebook can be stored as a new codebook version in the multi-granularity action codebook library, while the original action codebook is retained for rollback or update effect comparison.
[0253] When storing updated action codebooks, the preset granularity level, applicable action dimension, update time, and codebook version corresponding to the action codebook can be associated with the records. When executing the target task subsequently, the updated action codebook can be read from the multi-granularity action codebook library according to the granularity configuration, and used as the target action codebook for the corresponding action dimension for action tag generation and continuous action reconstruction.
[0254] In one implementation, a preset verification task can also be used to verify the updated action codebook. When the updated action codebook can reduce the difference between the task execution result and the expected task state, the updated action codebook is determined to be a valid action codebook; when the updated action codebook fails to improve the task execution result, the action codebook before the update can be restored or the codebook adjustment information can be redefined.
[0255] In this embodiment, by acquiring the task execution result generated by the robot performing the target task and comparing the task execution result with the expected task state, the difference between the action mark mapping and the actual control effect of the robot can be identified from the actual action execution feedback. Furthermore, based on the difference, the codebook adjustment information of the corresponding action dimension and discretization granularity is determined, and the corresponding action codebook is updated when the preset codebook update conditions are met. This allows the mapping relationship between action marks and continuous action values to gradually adapt to the specific robot and task environment. Using the updated action codebook for the action space representation of subsequent target tasks can reduce the repeated action deviations under the same action dimension and discretization granularity, improve the accuracy of continuous action reconstruction, and increase the success rate of the robot performing the target task.
[0256] In an experimental verification example, under the same hardware environment, a vision-language-action model employing an adaptive multi-granularity representation was compared with a diffusion action model using fifty-step iterative denoising. With the adaptive multi-granularity representation, the single inference time was approximately 60 milliseconds, the control frequency was approximately 15 Hz, and the success rate for long-sequence target tasks was approximately 85%. With the diffusion action model, the single inference time was approximately 150 milliseconds, the control frequency was approximately 6 Hz, and the success rate for long-sequence target tasks was approximately 72%. Therefore, using the adaptive multi-granularity representation reduces inference time by approximately 60% and improves the update frequency of robot control actions and the success rate of long-sequence target tasks.
[0257] Regarding continuous motion accuracy, the continuous motion accuracy using the adaptive multi-granularity representation is approximately ±0.7 degrees, while the continuous motion accuracy using the diffusion motion model is approximately ±0.5 degrees. Although the former has slightly lower continuous motion accuracy, the quantization error caused by discrete motion representation can be reduced by selecting a higher discretization granularity for the fine motion dimension.
[0258] Under other testing conditions, compared with OpenVLA, the inference time was reduced by approximately 45% when inference tests were conducted on an RTX 4090 computing device; the grasping success rate was improved by approximately 12% when compared with RT-2 on the grasping task in the REALM benchmark; the amount of adaptation data required was reduced by approximately 80% when performing cross-robot adaptation among the three types of robotic arms; and the cumulative motion error was reduced by approximately 40% in long sequence target tasks with more than five task steps compared to a fixed-granularity vision-language-action model.
[0259] In one possible embodiment, this application provides an adaptive multi-granularity representation system for robot motion space, such as... Figure 4 As shown, the adaptive multi-granularity representation system for robot motion space includes an information acquisition module, a feature determination module, a granularity configuration module, a codebook determination module, an action tag generation module, an action reconstruction module, and a constraint processing module connected in sequence, and includes a multi-granularity action codebook library connected to the codebook determination module, the action tag generation module, and the action reconstruction module respectively.
[0260] The information acquisition module is used to acquire visual observation information, task instructions, and robot condition information corresponding to the robot's execution of the target task. The visual observation information can be collected by visual sensors on the robot itself or in the working environment; the task instructions can be text instructions or text information converted from voice instructions; and the robot condition information is used to characterize the robot's structural characteristics and motion capabilities.
[0261] The feature determination module is used to extract and fuse features from visual observation information and task instructions to obtain multimodal fusion features, and to determine task complexity features to characterize the action requirements of the target task based on the multimodal fusion features.
[0262] The granularity configuration module determines the corresponding discretization granularity from multiple preset granularity levels for each motion dimension of the robot based on task complexity characteristics and robot condition information, and forms the granularity configuration according to the order of the motion dimensions. Therefore, different motion dimensions can adopt different discretization granularities according to the actual accuracy requirements of the target task.
[0263] The codebook determination module is used to determine the target action codebook corresponding to each action dimension from the multi-granularity action codebook library based on the granularity configuration. The multi-granularity action codebook library stores action codebooks corresponding to different preset granularity levels. Each target action codebook is used to record the mapping relationship between the continuous action values and action tags of the corresponding action dimension.
[0264] The action tag generation module is used to input multimodal fusion features, robot condition information and granular configuration into the conditional action decoder, and generate action tag sequences according to the target action codebook corresponding to each action dimension, so that the generated action tags match the discretized granularity of the corresponding action dimension.
[0265] The action reconstruction module is used to convert each action tag in the action tag sequence into a continuous action value according to the target action codebook corresponding to each action dimension, and to combine each continuous action value according to the action sequence and action dimension to obtain a continuous action sequence.
[0266] The constraint processing module is used to determine the robot's motion constraints based on the robot's condition information, perform motion constraint verification on continuous motion sequences, and project actions that do not meet the motion constraints to the robot's corresponding movable action domain to obtain a robot control motion sequence that meets the robot's actual motion capabilities.
[0267] The modules described above can be implemented by a processor executing computer programs stored in memory, or they can be implemented by software modules, hardware circuits, or a combination of both. These modules can be located within the robot's control equipment, or they can be distributed across the robot itself, edge computing devices, and cloud servers.
[0268] Through the collaborative processing among the above modules, the appropriate discretization granularity can be configured for different action dimensions according to the action requirements of the target task and the structural characteristics of the robot. The corresponding target action codebook is used to complete the generation of action tags and the reconstruction of continuous actions. At the same time, motion constraint verification and movable action domain projection ensure that the generated actions can be actually executed by the robot, thus taking into account the accuracy of robot action representation, processing efficiency, cross-platform applicability and action execution safety.
[0269] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 500 provided in this embodiment includes a memory 501 and a processor 502.
[0270] The memory 501 can be a separate physical unit, connected to the processor 502 via a bus 503. Alternatively, the memory 501 and processor 502 can be integrated and implemented in hardware. The memory 501 stores program instructions, which the processor 502 calls to execute the operations performed by the adaptive multi-granularity representation system of the robot's motion space in any of the above method embodiments.
[0271] Optionally, when some or all of the methods in the above embodiments are implemented by software, the electronic device 500 may also include only the processor 502. A memory 501 for storing programs is located outside the electronic device 500, and the processor 502 is connected to the memory via circuits / wires to read and execute the programs stored in the memory. The processor 502 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. The processor 502 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0272] The memory 501 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory may also include a combination of the above types of memory.
[0273] For example, this application provides a chip including: an interface circuit and a logic circuit. The interface circuit is used to receive signals from other chips outside the chip and transmit them to the logic circuit, or to send signals from the logic circuit to other chips outside the chip. The logic circuit is used to perform the operations performed by the adaptive multi-granularity representation system of the robot motion space in the above method embodiments.
[0274] For example, this application provides a computer-readable storage medium storing computer program instructions thereon, which are executed by the processor of an electronic device to cause the electronic device to perform the operations performed by the adaptive multi-granularity representation system of the robot motion space in the above method embodiments.
[0275] For example, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform operations executed by the adaptive multi-granularity representation system of the robot motion space in the above method embodiments.
[0276] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An adaptive multi-granularity representation method for robot motion space, characterized in that, The method includes: Acquire visual observation information, task instructions, and robot condition information corresponding to the robot's execution of the target task; Based on the visual observation information and the task instructions, multimodal fusion features and task complexity features are determined; wherein, the task complexity features are used to characterize the action requirements of the target task; Based on the task complexity characteristics and the robot condition information, for each action dimension of the robot, the corresponding discretization granularity is determined from multiple preset granularity levels to obtain the granularity configuration; Based on the granularity configuration, the target action codebook corresponding to each action dimension is determined from the multi-granularity action codebook library, wherein the target action codebook is used to represent the mapping relationship between the continuous action values and action tags of the corresponding action dimension; The multimodal fusion features, the robot conditional information, and the granular configuration input conditional action decoder are used to generate an action tag sequence that matches the target action codebook corresponding to each action dimension. Based on the target action codebook corresponding to each action dimension, the action tag sequence is reconstructed into a continuous action sequence; Based on the robot condition information, the continuous action sequence is subjected to motion constraint verification, and the actions in the continuous action sequence that do not meet the motion constraints are projected to the corresponding movable action domain of the robot to obtain a robot control action sequence for controlling the robot to perform the target task.
2. The method according to claim 1, characterized in that, The step of determining multimodal fusion features and task complexity features based on the visual observation information and the task instructions includes: Visual features are extracted from the visual observation information to obtain visual observation features; Language feature extraction is performed on the task instructions to obtain task instruction features; The visual observation features and the task instruction features are fused to obtain the multimodal fusion features; Based on the multimodal fusion features, determine the spatial accuracy requirements, time urgency, characteristics of the operation object, and degree of environmental constraints corresponding to the target task; The task complexity characteristics are obtained based on the spatial accuracy requirements, the time urgency, the characteristics of the operation object, and the degree of environmental constraints.
3. The method according to claim 1, characterized in that, The robot condition information includes robot structural information and motion constraint information. The robot structural information includes at least one of the following: the number of degrees of freedom of the robot, joint type, joint range of motion, and kinematic type. The motion constraint information includes at least one of the following: joint position constraint, joint velocity constraint, joint acceleration constraint, and collision constraint. The method further includes: The robot's structural information is conditionally encoded to obtain the robot's conditional features; The step of determining the corresponding discretization granularity from multiple preset granularity levels for each action dimension of the robot based on the task complexity characteristics and the robot condition information, to obtain the granularity configuration, includes: Based on the task complexity characteristics and the robot condition characteristics, the discretization granularity corresponding to each action dimension is determined to obtain the granularity configuration; The step of fusing the multimodal features, the robot conditional information, and the granular configuration input conditional action decoder includes: The multimodal fusion features, the robot conditional features, and the granular configuration are input into the conditionalized action decoder; The step of performing motion constraint verification on the continuous action sequence based on the robot condition information includes: The motion constraint of the continuous action sequence is verified based on the motion constraint information.
4. The method according to claim 3, characterized in that, The step of determining the discretization granularity corresponding to each action dimension based on the task complexity characteristics and the robot condition characteristics, to obtain the granularity configuration, includes: The task complexity features and the robot condition features are input into the granularity selection network; The granularity selection network outputs the granularity selection results for each action dimension. Based on the granularity selection results corresponding to each action dimension, the discretization granularity corresponding to each action dimension is determined from the multiple preset granularity levels respectively; According to the order of the action dimensions, the discretized granularities corresponding to each action dimension are combined to obtain the granularity configuration.
5. The method according to claim 1, characterized in that, The multi-granularity action codebook library includes multiple action codebooks corresponding to different preset granularity levels. Each action codebook includes multiple action tags that are consistent with the number of discrete intervals represented by the corresponding preset granularity level. The multiple action tags are respectively mapped to multiple continuous action intervals in the range of continuous action values. The step of determining the target action codebook corresponding to each action dimension from the multi-granularity action codebook library according to the granularity configuration includes: Obtain the discretized granularity corresponding to each action dimension from the granularity configuration; For each action dimension, the action codebook corresponding to the discretization granularity of the action dimension is determined from the plurality of action codebooks, and the determined action codebook is used as the target action codebook corresponding to the action dimension.
6. The method according to claim 1, characterized in that, The step of generating an action tag sequence that matches the target action codebook corresponding to each action dimension by combining the multimodal fusion features, the robot conditional information, and the granular configuration input conditional action decoder includes: The robot condition information and the granularity configuration are embedded and encoded respectively to obtain robot condition embedding features and granularity condition embedding features; The multimodal fusion feature, the robot conditional embedding feature, and the granular conditional embedding feature are used as conditional inputs to the conditional action decoder; According to the preset action tag generation order, determine the action dimension corresponding to the current action tag to be generated, and determine the corresponding target action codebook based on the action dimension; Based on the conditional input and the generated action tag sequence, predict the current action tag to be generated from the action tag set defined by the target action codebook; The predicted action tags are added to the generated action tag sequence, and action tag prediction is iteratively performed based on the updated generated action tag sequence until the preset generation termination condition is met, thus obtaining the action tag sequence.
7. The method according to claim 1, characterized in that, The conditional action decoder includes a high-level decoder and a low-level decoder; The step of generating an action tag sequence that matches the target action codebook corresponding to each action dimension by combining the multimodal fusion features, the robot conditional information, and the granular configuration input conditional action decoder includes: The multimodal fusion features and the robot condition information are input into the high-level decoder to generate a task sub-target sequence corresponding to the target task; According to the order of the task sub-objectives in the task sub-objective sequence, each task sub-objective, the multimodal fusion feature, the robot condition information, and the granularity configuration are sequentially input into the low-level decoder; Through the low-level decoder, for each task sub-objective, an action tag sub-sequence is generated that matches the target action codebook for each action dimension corresponding to the task sub-objective; According to the order of the task sub-objectives, the action tag sub-sequences corresponding to each task sub-objective are combined to obtain the action tag sequence; wherein, the operating frequency of the higher layer decoder is lower than the operating frequency of the lower layer decoder.
8. The method according to claim 1, characterized in that, The step of reconstructing the action marker sequence into a continuous action sequence based on the target action codebook corresponding to each action dimension, performing motion constraint verification on the continuous action sequence based on the robot condition information, and projecting actions in the continuous action sequence that do not meet the motion constraints to the corresponding movable action domain of the robot includes: For each action marker in the action marker sequence, the continuous action interval corresponding to the action marker is determined according to the action dimension corresponding to the action marker and the target action codebook corresponding to the action dimension. Based on the representative action value corresponding to the continuous action interval, determine the continuous action value corresponding to each action mark, and combine the continuous action values according to the action sequence and action dimension to obtain the continuous action sequence. The motion constraints of the robot are determined based on the robot condition information, and the actions to be corrected in the continuous action sequence that do not meet the motion constraints are determined based on the motion constraints. Based on the motion constraints, the motion to be corrected is projected onto the corresponding movable action domain of the robot to obtain a corrected motion that satisfies the motion constraints. The robot control action sequence is obtained by combining the actions that satisfy the motion constraints and the corrective actions in the continuous action sequence according to the corresponding action timing and action dimension.
9. The method according to claim 1, characterized in that, The method further includes: Obtain the task execution result generated by the robot executing the target task according to the robot control action sequence; Based on the difference between the task execution result and the expected task state corresponding to the target task, codebook adjustment information related to the corresponding action dimension and discretization granularity is determined. If the difference meets the preset codebook update conditions, the codebook parameters of the corresponding action codebook in the multi-granularity action codebook library are updated according to the codebook adjustment information to obtain the updated action codebook. The updated action codebook is stored in the multi-granularity action codebook library so that the updated action codebook can be used for the action space representation of subsequent target tasks.
10. An adaptive multi-granularity representation system for robot motion space, characterized in that, The system includes: The information acquisition module is used to acquire visual observation information, task instructions, and robot condition information corresponding to the robot's execution of the target task; The feature determination module is used to determine multimodal fusion features and task complexity features based on the visual observation information and the task instructions, wherein the task complexity features are used to characterize the action requirements of the target task; The granularity configuration module is used to determine the corresponding discretized granularity from multiple preset granularity levels for each action dimension of the robot based on the task complexity characteristics and the robot condition information, so as to obtain the granularity configuration. The codebook determination module is used to determine the target action codebook corresponding to each action dimension from the multi-granularity action codebook library according to the granularity configuration, wherein the target action codebook is used to represent the mapping relationship between the continuous action values and action tags of the corresponding action dimension. The action tag generation module is used to generate an action tag sequence that matches the target action codebook corresponding to each action dimension by taking the multimodal fusion features, the robot conditional information and the granular configuration input conditional action decoder; The action reconstruction module is used to reconstruct the action tag sequence into a continuous action sequence based on the target action codebook corresponding to each action dimension. The constraint processing module is used to perform motion constraint verification on the continuous action sequence according to the robot condition information, and project the actions in the continuous action sequence that do not meet the motion constraints to the corresponding movable action domain of the robot to obtain a robot control action sequence for controlling the robot to perform the target task.