A wearable robotic arm and embodied intelligence method thereof
By acquiring visual language information and combining imitation learning and visual language action models, the robotic arm movement trajectory is generated, which solves the shortcomings of wearable robotic arms in intelligent interaction and environmental perception, and achieves efficient and adaptive task execution.
Patent Information
- Application Number
- CN202510813247.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing wearable robotic arms have weak performance in intelligent interaction, fail to achieve natural synchronization and efficient coordination with human body movements, limited environmental perception ability, lack adaptability, and difficulty in dealing with dynamic tasks.
By obtaining visual language information, combining imitation learning and visual language action models, the robotic arm movement trajectory is generated, and the robotic arm is driven to perform tasks using inverse kinematics methods to enhance environmental perception and adaptability.
It improves the intelligence level and environmental perception ability of the robotic arm, improves the task completion rate and response speed, and enhances the adaptability and flexibility to the environment.
Smart Images

Figure CN120347769B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a wearable robotic arm and an embodied intelligence method thereof. Background Art
[0002] Existing intelligent robotic arm devices are primarily used in industrial settings and are typically designed as fixed or semi-fixed equipment for tasks such as assembly line production, welding, and handling. These robotic arms rely on pre-set operating procedures, have fixed working positions, and are limited in their scope of tasks. They lack portability and flexibility, making them difficult to adapt to diverse and dynamically changing environments. Furthermore, traditional robotic arm designs focus primarily on precision and repeatability, typically offering only a single control mechanism, such as operating through preset programs or remote control. While suitable for industrial scenarios, this single control approach often exhibits significant limitations when rapid response or complex tasks are required.
[0003] Beyond industrial applications, with the rapid development of artificial intelligence and embodied intelligence, robotic arms are gradually evolving towards multifunctionality, portability, and intelligence. For example, wearable robotic arms, as an emerging technology, are beginning to be used in scenarios such as assisted rehabilitation, personal life, and rescue operations.
[0004] However, compared to industrial robotic arms, the development of wearable robotic arms is still in its early stages, with several key issues that need to be addressed. Existing wearable robotic arms are weak in intelligent interaction, failing to achieve natural synchronization and efficient coordination with human movements. When handling dynamic tasks, the robotic arms' reaction speed and decision-making capabilities are insufficient, making it impossible to balance real-time performance with complex task processing. Traditional wearable robotic arms have a relatively simple way of perceiving the external environment, typically used as feedback signals during task execution. Traditional robotic arm control methods rely on preset rules and procedures, lack adaptability, and are unable to continuously optimize behavior through human-machine interaction and feedback mechanisms.
[0005] Against this backdrop, society is increasingly demanding intelligent devices that can overcome the limitations of traditional robotic arms. In particular, in the field of personalized portable devices, there is an urgent need for more flexible and intelligent solutions that can adapt to diverse mission scenarios and user needs. Summary of the Invention
[0006] The purpose of the present invention is to provide a wearable robotic arm and its embodied intelligence method to address all or part of the above-mentioned problems, so as to solve the problems of insufficient intelligence and limited environmental perception capabilities of current wearable robotic arms.
[0007] The technical solution adopted in the present invention is as follows:
[0008] A wearable robotic arm embodied intelligence method, comprising:
[0009] Obtain visual language information for the target task;
[0010] Matching a first motion trajectory from a first database according to the visual language information, and / or inputting the visual language information into a visual language motion model to obtain a second motion trajectory; the first database is constructed by imitation learning;
[0011] generating a robot arm motion instruction according to the first motion trajectory and / or the second motion trajectory;
[0012] Execute the robot arm motion instruction to drive the robot arm to perform the target task.
[0013] On the other hand, the present invention also provides an embodied intelligent wearable robotic arm, which includes a robotic arm, a processor and a storage medium, wherein the storage medium stores computer instructions. When the processor runs the computer instructions, it can execute the above-mentioned wearable robotic arm embodied intelligence method to drive the robotic arm to perform the target task.
[0014] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0015] The wearable robotic arm and its embodied intelligence method designed in this application, through the perception and recognition of the environment at the visual level, combined with the description of the target task through language information, adaptively obtains the corresponding motion trajectory, and then solves the robotic arm motion instructions through inverse kinematics methods, thereby driving the robotic arm to perform the target task. It has the characteristics of high intelligence and strong adaptability. In addition, the visual language information obtained in this application includes spatial feature information and visual feature information at the visual level. The multimodal information improves the ability to perceive the environment, thereby improving the accuracy of motion trajectory planning and improving the task completion rate. In addition, this application designs two sets of motion trajectory planning modes. The robotic arm can quickly obtain the target motion trajectory based on imitation learning records, and can also generate rational motion trajectories with the help of visual language motion models, thereby improving the response speed of the robotic arm to the target task. As the frequency of use increases, the learned motion knowledge can be continuously enriched and the response speed can be continuously improved. The two modes can be used in combination as needed to further improve the adaptability to the environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0017] Figure 1 This is a flowchart of the execution of the wearable robotic arm embodied intelligence method provided in an embodiment of the present application.
[0018] Figure 2 、 Figure 3They are structural diagrams of the wearable robotic arm in the embodiment of the present application at two different perspectives.
[0019] Figure 4 This is a flow chart of the method for constructing the first database in an embodiment of the present application.
[0020] Figure 5 This is a state diagram demonstrating picking up a mobile phone in an embodiment of the present application.
[0021] Figure 6 It is a flow chart of the spatial structure feature and motion trajectory detection method in an embodiment of the present application.
[0022] Figure 7 This is a flowchart of a method for implementing embodied intelligence in a wearable robotic arm in an embodiment of the present application.
[0023] In the figure, 1 is the robotic arm camera, 2 is the wearable backpack, 3 is the robotic arm, 4 is the rotating disk, 5 is the power data line, 6 is the processor, 7 is the chest camera, 8 is the power supply, and 9 is the lidar. DETAILED DESCRIPTION
[0024] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.
[0025] Any feature disclosed in this specification (including any appended claims and abstract), unless otherwise stated, may be replaced by other equivalent or similar features. In other words, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.
[0026] In response to the shortcomings of traditional wearable robotic arms in terms of environmental perception and adaptability, the embodiments of the present application provide a wearable robotic arm and its embodied intelligence method, aiming to enhance the intelligence level of the robotic arm and improve its ability to perceive the environment.
[0027] The wearable robotic arm embodied intelligence method provided in the embodiments of the present application is as follows: Figure 1 As shown, the following steps are included:
[0028] Step S1: Obtain visual language information of the target task.
[0029] In the embodiment of the present application, the visual language information includes visual information of environmental perception and language information of task description, wherein the visual information includes spatial feature information and visual feature information to enhance spatial perception ability.
[0030] In some feasible implementations, the spatial feature information is point cloud features of the perceived environment, and the visual feature information is image features of the perceived environment. A laser radar is used to scan the target task's work area to obtain point cloud data, and thus point cloud features; a camera is used to capture the target task's work area to obtain image data, and thus image features.
[0031] like Figure 2 、 Figure 3 The figure shows the structure of a wearable robotic arm in some possible implementations. The wearable robotic arm includes a robotic arm 3, a robotic arm camera 1 fixedly mounted on the robotic arm (which can be used for real-time status feedback), a wearable backpack 2 for wear, a rotating disk 4 serving as a rotating base for the robotic arm, power and data lines 5, a processor 6 responsible for data processing, a chest camera 7 mounted on the chest, a power supply 8 for powering the wearable robotic arm, and a chest-mounted laser radar 9. A microphone (not shown) for voice collection is mounted on the wearable backpack 2 or elsewhere. When the wearable robotic arm is worn on the user, the chest camera 7 captures image data, such as RGB images, and the chest-mounted laser radar 9 captures point cloud data. This point cloud data contains the three-dimensional spatial features of objects and backgrounds in the environment.
[0032] Step S2: Match the first motion trajectory from the first database according to the visual language information, and / or input the visual language information into the visual language motion model to obtain the second motion trajectory.
[0033] The first database stores visual language information and motion trajectories in association with each other. The visual language information can be used to try to match the associated motion trajectory from the first database. If a match is found, the first motion trajectory is obtained. The first database is constructed through imitation learning.
[0034] Imitation learning is to record the actions performed by humans when performing target tasks, collect corresponding visual language information, and associate the visual language information with the recorded action trajectory.
[0035] As an optional implementation, Figure 4 As shown, the method for constructing the first database includes the following process:
[0036] Step S21: Acquire environment information for executing the current operation task and language information describing the current operation task.
[0037] Environmental information is information perceived from the environment, including spatial and visual features of the work environment. The work environment includes the target object being manipulated and the human arms manipulating the target object. According to the previous embodiment, this information includes image data of the work environment captured by the chest camera 7 and point cloud data of the work environment captured by the lidar 9.
[0038] Language information is a verbal description of the current task, indicating the current operational behavior. Language information can be collected in text or voice. Alternatively, language information can be extracted from the collected descriptive language, including keywords such as actions, goals, and constraints.
[0039] In some specific implementations, the method for acquiring language information includes:
[0040] a. Collect language instructions through voice or text form. For language instructions collected in voice form, use voice recognition tools to convert them into text form.
[0041] b. Segment the textual instructions, removing punctuation and stop words; extract the subject-verb-object structure of the instructions using part-of-speech tagging and syntactic analysis; and identify keywords in the instructions based on predefined task templates.
[0042] c. Use a pre-trained language model to convert text into word vectors in the form of high-dimensional semantic feature embeddings, and extract action information, target information, and constraints from the embeddings.
[0043] d. If the language instruction is part of a multi-round conversation, associate the current language instruction with the previous content and complete the missing keywords.
[0044] For example, suppose the current operation task is to operate the mobile phone, such as Figure 5 As shown, the chest camera 7 captures image data of the work environment to record the visual feature information of the mobile phone, including color, shape, and surface texture, as well as the movement process of the human arms operating the mobile phone. The laser radar 9 also performs a three-dimensional scan of the work environment to obtain point cloud data including the mobile phone and the human arms, to record the three-dimensional position and spatial size of the mobile phone, as well as the movement trajectory of the human arms. The integrated microphone synchronously collects the language description of the current operation task, such as "picking up the mobile phone", and transmits it to the processor 6 through the power data line 5 for post-processing. After the processor 6 converts the voice data into text data, it extracts keywords, such as the action keyword "pick up" and the target keyword "mobile phone", without constraint keywords. In addition, if the language description of the operation task is collected in text form, the processor 6 directly extracts the keywords. The storage form of the language information is a vector representation of the semantic features of the keywords, that is, the keywords are converted into corresponding word embeddings using a pre-trained language model, thereby extracting the semantic feature vector of the keywords.
[0045] S22: Detecting the spatial structural features of the target object and the motion trajectory for executing the current operation task from the environmental information.
[0046] The working environment contains different objects, and different objects have different spatial structural characteristics. The target objects can be identified by spatial structural characteristics.
[0047] In some possible implementations, see the attached Figure 6 ,The method of detecting the spatial structural features of the target object from the environmental information includes the following process:
[0048] Step S221: De-noise the point cloud data in the environment information and crop the first point cloud data in the workspace. Pre-process the image data in the environment information to obtain the first image data.
[0049] Denoising of point cloud data can be accomplished using formula (1):
[0050] Formula (1): ;
[0051] Where, represents the denoised point cloud data, It is the first data points, is the Gaussian filter kernel function, is the standard deviation, is the total number of data points.
[0052] By using formula (1), the point cloud data is smoothed to remove isolated points and noise points, making the remaining data points more continuous and smooth.
[0053] After denoising the point cloud data, the point cloud data is cropped to retain only the valid point cloud on the work desktop, which is the first point cloud data.
[0054] As for the preprocessing of the image data, in some feasible implementations, it may include denoising, color adjustment, etc. After the preprocessing, the first image data is obtained.
[0055] S222: spatially align the first point cloud data and the first image data.
[0056] The alignment of the data of the two modalities can be completed based on the calibration parameters of the laser radar 9 and the chest camera 7. Taking the alignment of the first image data to the first point cloud data as an example, the data alignment process can be completed using the global rigid transformation matrix T of formula (2).
[0057] Formula (2): ;
[0058] Where R is the rotation matrix and t is the translation vector. Using the calibration parameters of the chest camera 7 and the lidar 9, the overall rigid transformation matrix T can be obtained based on the coordinate points representing the same position, thereby completing the alignment between the first image data and the first point cloud data.
[0059] S223 : Filter out background information from the first point cloud data, cluster the remaining point cloud data, and separate different target objects.
[0060] For the first point cloud data, a point cloud segmentation algorithm can be used to extract static background information such as the ground and walls, filter out the background information, and then cluster the remaining point cloud data to separate different target objects based on the clustered clusters. As an optional implementation, the point cloud data clustering method uses the K-means clustering algorithm, and the formula is as follows:
[0061] Formula (3): ;
[0062] In the formula, k represents the cluster number, is the kth cluster, is the i-th data point, is the cluster center of the kth cluster, n is the number of data points involved in the cluster, Express request The norm of .
[0063] By clustering point cloud data, the point clouds of different objects can be separated, thereby isolating the point cloud data of the target object. Point cloud data carries the spatial characteristics of the object, from which information such as the 3D position, shape, and size of each object can be derived. For example, in the previous embodiment, if the target object is a mobile phone, this method can be used to isolate the point cloud data of the mobile phone from the collected point cloud data, thereby obtaining information such as the mobile phone's 3D position, shape, and size.
[0064] S224 : Detecting the categories and positions of different target objects from the first image data, and matching the three-dimensional positions and shapes and sizes of the different target objects in the first point cloud data according to the positions.
[0065] The target detection algorithm can detect the category and position of different target objects from the first image data. Image data can only reflect the pixel position of an object in the image, and it is difficult to reflect its specific position in three-dimensional space. Point cloud data can make up for this deficiency. Based on the position of the object detected from the first image data, the point cloud data of the object at the corresponding position is matched from the first point cloud data aligned with the first image data, thereby identifying the three-dimensional position and shape and size of the target object based on the point cloud data of the target object. For example, the position of a mobile phone (i.e., the category) is identified in the first image data, and then based on the position of the mobile phone in the first image data, the point cloud data of the mobile phone is matched from the first point cloud data to identify the three-dimensional position and shape and size of the mobile phone.
[0066] The category, three-dimensional position, shape and size of the target object constitute the spatial structural characteristics of the target object.
[0067] In addition, the above-mentioned motion trajectory can also be obtained by a method similar to that of obtaining the spatial structural features of the target object. Figure 6 , the method of obtaining the motion trajectory includes:
[0068] S225 , detecting the joint positions of the human arms from the first image data of each key frame respectively, and matching the three-dimensional positions of the joint points in the first point cloud data according to the joint point positions.
[0069] Keyframes are image data and point cloud data captured at key locations extracted from the continuously and synchronously acquired image data and point cloud data. Each keyframe corresponds to first image data and first point cloud data. The alignment operation between the first image data and the first point cloud data has been described previously. Using human pose estimation algorithms, the joints of the human arm, including key points such as the shoulder, elbow, and wrist, can be located from the first image data. The joints in two-dimensional space are then mapped to three-dimensional space using a calibration mapping relationship between the chest camera 7 and the lidar 9. This allows the three-dimensional positions of the joints to be isolated from the first point cloud data.
[0070] S226: Track the three-dimensional positions of the joint points in the key frames to calculate the three-dimensional motion trajectory of the human arms.
[0071] The joints are located at the pivotal positions of the arms, and the movement trajectory of the joints can reflect the three-dimensional movement trajectory of human arms.
[0072] S227. Use Kalman filtering or optical flow method to smooth the three-dimensional motion trajectory.
[0073] Taking the Kalman filter method as an example, its processing is shown in formula (4) and formula (5).
[0074] Formula (4): ;
[0075] Formula (5): ;
[0076] In the formula, k represents the current moment, k-1 represents the previous moment, is the state estimate at the current moment, A is the state transfer matrix, B is the control matrix, is the control input, is the process noise, and are the covariance matrices estimated at the current moment and the previous moment, respectively, and Q is the covariance of the process noise.
[0077] S228. Extract the time series features of the three-dimensional motion trajectory and represent the motion trajectory of performing the operation task in the form of a feature vector.
[0078] The extracted time series features include the speed, acceleration, and movement direction of human arms.
[0079] In another feasible implementation, the method for obtaining the motion trajectory may also include:
[0080] A human grips the robotic arm to perform an operation, and the robot's trajectory is recorded simultaneously during the operation. This trajectory is also in the form of a time series feature vector, including features such as the robot's velocity, acceleration, and direction of movement.
[0081] Step S23: The motion trajectory, the spatial structure characteristics of the target object, and the language information obtained by executing the current operation task are associated and stored in the first database.
[0082] The above method achieves this by imitating human arms, associating and storing the motion trajectory obtained from each operational task, the spatial structural characteristics of the target object, and language information into a single entry. Each piece of information is stored in the form of word vectors, thereby constructing a first database. This first database can quickly respond to subsequent target tasks, quickly matching feasible motion trajectories for the target task, and thus quickly completing the target task.
[0083] In addition, in some optional implementations, in order to improve the retrieval efficiency in the first database, based on the associated storage of the motion trajectory, the spatial structure characteristics of the target object, and the language information, the method of constructing the first data further includes:
[0084] Step S231: Hash index the spatial structure features and language information of the target object. This can effectively reduce the amount of character data to be calculated when traversing each entry during retrieval. While effectively reducing the amount of data storage, retrieval efficiency can also be improved.
[0085] Step S232: Cluster the motion trajectory and the spatial structure features of the target object to generate different task templates.
[0086] Representing the motion trajectory and the spatial features of the target object in the form of a template can further simplify the matching process and improve retrieval efficiency.
[0087] After constructing the first database, we search the first database based on the visual language information of the current target task, that is, the language information describing the target task, and the environmental information (including point cloud data and image data) obtained in the current scene. If the search is successful, the corresponding action trajectory is obtained. If the search fails, that is, the first database does not include the action trajectory that can be used to perform the current task, we need to obtain the second action trajectory through the visual language action model to perform the target task, such as Figure 7 shown.
[0088] The method for obtaining the second motion trajectory is obtained with the help of the reasoning ability of artificial intelligence. The visual language action model needs to be pre-trained to have reasoning ability. The input feature of the visual language action model is visual language information, and the output result is the target posture, that is, the target position and posture that the robot arm wants to achieve; the target position is a three-dimensional position, which represents the target coordinates of the robot arm end effector in three-dimensional space, and should be used to locate the specific position of the target object; the target posture is a three-dimensional posture, which is used to control the orientation and movement of the robot arm end effector. The orientation is represented by the roll, pitch, yaw or rotation vector parameters of the Euler angle to ensure that the robot arm end effector is aligned with the target object. The movement is represented by the gripper action parameters, which represent the opening and closing state of the robot arm end effector, and is used to indicate the grasping or releasing of the target object. The visual language action model adopts the Transformer structure to achieve the fusion understanding of multimodal features (spatial features, visual features and text features). Its model representation is as follows:
[0089] ,
[0090] Among them, Image represents image data, PointCloud represents point cloud data, Text represents language information, TargetAction represents target pose, and Transformer() represents the mapping function of the visual language action model to the input features.
[0091] Training the vision-language-action model includes the data preparation stage and the model training stage.
[0092] During the data preparation phase, we first collect a large amount of visual language information from different manipulation tasks and the corresponding target poses (the position and pose after the relevant action sequence is completed) as samples. This visual language information includes image data, point cloud data, and language information describing the manipulation tasks, and is multimodal data. For example, data samples for robotic arm training can be downloaded from the Open X-Embodiment dataset.
[0093] During the model training phase, the large amount of data samples obtained during the data preparation phase are divided into training and test sets for training and testing the visual-language-action model, respectively. The visual-language information of the samples (i.e., image data, point cloud data, and language information) is input into the model. The model parameters of the model are iteratively optimized using the corresponding target pose as the true value. By minimizing the loss between the predicted and true values of the target pose, a visual-language-action model is trained that uses the visual-language information as input and the target pose as output.
[0094] As an optional implementation, the loss function for visual language action model training is:
[0095] Formula (6): ;
[0096] In the formula, L represents the calculated loss, t refers to the sequence number of the sample, is the sample size, is the target pose predicted by the visual language action model for the t-th sample, is the actual target pose of the t-th sample, Express request The norm of .
[0097] After obtaining a trained visual language action model, the visual language information is input into the model to obtain the target pose when performing the target task. Based on this, the path of the robot arm from the initial pose to the target pose is optimized through trajectory planning algorithms, thereby obtaining the desired second action trajectory.
[0098] In addition, in the embodiment of the present application, the visual language action model is also used to continuously expand the candidate entries in the first database. In some feasible implementations, when the second action trajectory is obtained through the visual language action model, the visual language information and the second action trajectory are also associated and stored in the first database.
[0099] For this design, there are three modes of operation:
[0100] Method 1: First try to match the first action trajectory from the first database based on visual language information.
[0101] If the first motion trajectory is not matched in the first database, and the second motion trajectory is derived using the visual language motion model, the visual language information and the second motion trajectory are directly associated and stored in the first database as a new entry. If the first database has specific storage format requirements, the corresponding format is converted according to the format requirements of the first database, such as hash indexing the spatial structure characteristics and language information of the target object.
[0102] Alternatively, if the first motion trajectory is matched in the first database but ultimately fails to execute the target task, and the second motion trajectory is obtained using the visual language motion model, the visual language information and the second motion trajectory are directly associated and stored in the first database to form a new entry.
[0103] Method 2: Only visual language information is input into the visual language action model to obtain the second action trajectory.
[0104] If the second motion trajectory is directly derived using the visual language motion model without matching the first database, the visual language information is first matched against the first database. If a match is found with the first motion rule, the second motion trajectory replaces the first motion rule. If no match is found with the first motion rule, the visual language information and the second motion trajectory are directly associated and stored in the first database as a new entry. Similarly, if the first database has specific storage format requirements, the corresponding format conversion is performed according to the first database's format requirements.
[0105] Method three: a method of matching a first motion trajectory from a first database according to visual language information, and further inputting the visual language information into a visual language motion model to obtain a second motion trajectory.
[0106] A first motion trajectory is matched from the first database based on the visual language information. If the robotic arm fails to perform the target task according to the first motion trajectory, the current visual language information is input into the visual language motion model to obtain a second motion trajectory. If the robotic arm successfully performs the target task according to the second motion trajectory, the matched first motion trajectory is combined with the second motion trajectory to update the corresponding first motion trajectory in the first database.
[0107] As an optional embodiment, the method for optimizing the path of the robotic arm from the initial posture to the target posture includes:
[0108] a. Identify the spatial structural features of obstacles in the environment where the robotic arm performs the target task based on visual language information.
[0109] Obstacles are objects other than the target object in the environment. The spatial structural characteristics of obstacles are exactly the same as those of the objects described above, including the type, three-dimensional position, shape, and size of the obstacle. Therefore, the method for extracting the spatial structural characteristics of obstacles is the same as that in the previous embodiment and will not be repeated here.
[0110] Based on the spatial structural characteristics of the identified obstacle, the impact of traversing the obstacle on the robot arm can be determined, that is, whether the obstacle can be traversed. For example, rigid obstacles (such as racks, glass, etc.) can be traversed statically, while some flexible objects (such as cotton, leaves, etc.) can be traversed. The impact of the obstacle on the robot arm is called the potential penalty of the obstacle, and it is used to determine whether the obstacle can be traversed. In other words, the potential penalty of an obstacle is not equivalent to the path cost. In this embodiment, the potential penalty of different obstacles is also equivalent to the path cost dimension to facilitate the calculation and comparison of the path cost. This equivalent factor is called the obstacle penalty weight. The potential penalty and penalty weight of an obstacle are obtained by matching the spatial structure characteristics of the obstacle (the type in ) from the association table that stores the item type, potential penalty, and penalty weight.
[0111] Based on the spatial structural characteristics of each identified obstacle (its three-dimensional position and shape and size), it can be determined whether the planned path will pass through the obstacle. If so, the corresponding obstacle will be activated, increasing the path cost.
[0112] b. In the initial pose and target pose Search for the path with the minimum cost between. Among them, the cost of the path It consists of the length of the path and the path penalty caused by colliding with obstacles along the path. The path penalty caused by colliding with obstacles is calculated by the penalty weight and potential penalty of each obstacle collided with. Specifically, there are:
[0113] Formula (7): ;
[0114] Where, It represents the speed on the path, that is, the rate of change of the path length, and t represents time. The length of the path can be obtained by integrating the path speed.
[0115] The cost of the path from the initial position to the target position of the robot arm can be calculated by formula (7). On this basis, the path with the minimum cost is selected, which is the second action trajectory.
[0116] As for how to construct an alternative path between the initial pose and the target pose, in some optional implementations, the following methods are included:
[0117] a. Using the initial pose as the root pose node, iterate the following steps until the target pose is reached or the maximum number of iterations is reached:
[0118] b. Generate a random pose node within the robot's workspace and search for the pose node with the lowest jump cost among the generated pose nodes. The jump cost can include only the Euclidean distance between nodes, or it can also include the weighted norm of the joint angle change, with different weights for different joints, or other factors.
[0119] c. Extend a predetermined step length from the searched pose node toward the random pose node to obtain a candidate pose node. If there is no rigid obstacle in the path from the searched pose node to the candidate pose node (the three-dimensional position and shape of the obstacle can be used to determine whether it is traversed, and the type of obstacle can be used to determine whether it is a rigid obstacle), then add the candidate pose node to the generated pose node and use the searched pose node as the parent node.
[0120] d. Search all generated pose nodes within a predetermined radius (expressed by r, the value can be set by yourself) around the newly added generated pose node, and determine the pose node with the lowest jump cost to the root pose node (i.e., the initial pose) as the parent node.
[0121] The process of step d involves traversing each pose node found as a candidate parent node, calculating the total jump cost from the root pose node to the newly added pose node, and if the total jump cost is lower than the total jump cost from the root pose node to the root pose node via the current parent node, then the current candidate parent node is replaced with the currently traversed candidate parent node. The same process is repeated for subsequent candidate parent nodes.
[0122] It should be noted that in step a, the target pose will be reached within the set maximum number of iterations. The random pose node generated in each iteration has a certain probability (e.g., 5%) of being the target pose node, and within the set maximum number of iterations, the random pose node generated in one iteration will definitely be the target pose node.
[0123] The motion trajectory is a sequence of actions composed of a series of discrete posture nodes, which cannot directly guide the movement of the robotic arm. It is necessary to generate driving instructions that can directly drive the movement of the robotic arm based on the motion trajectory.
[0124] S3. Generate a robot arm motion instruction according to the first motion trajectory and / or the second motion trajectory.
[0125] As an optional implementation, after obtaining the first motion trajectory and / or the second motion trajectory, a kinematic planning algorithm is used to optimize the motion trajectory to ensure the smoothness and feasibility of the robot arm movement.
[0126] After obtaining the first motion trajectory and / or the second motion trajectory, the corresponding motion trajectory is converted into a robotic arm motion instruction executable by the robotic arm controller through the inverse kinematics (IK) algorithm, including a series of joint angle control instructions.
[0127] Assume that the target position of the end effector of the robot arm is expressed as , where x, y, z are the coordinates of the target object in three-dimensional space. The motion of the robot arm is determined by a set of joint angles The purpose of the control calculation is to calculate the angle q of each joint through the IK algorithm so that the end effector of the robot arm reaches the target position. The inverse kinematics is solved according to the following formula (8).
[0128] Formula (8): ;
[0129] Where f() represents the mapping function of the IK algorithm; is the target location; is the velocity or acceleration of the target position; These are control parameters (such as inertia and damping) determined by the specific application of the robot arm and are pre-defined. Given a target position and other relevant control information, an inverse kinematics solution can be used to determine a sequence of robot arm joint angles that meets the task requirements.
[0130] S4. Execute the robot arm motion instructions to drive the robot arm to perform the target task.
[0131] After receiving the robot arm motion instructions, the motion instructions are passed to the low-level controller through ROS (Robot Operating System). The low-level controller is responsible for controlling the specific motor rotation according to the robot arm motion instructions, thereby controlling the corresponding joint angles.
[0132] Based on the concepts of this application, an embodiment of this application further provides an embodied intelligent wearable robotic arm, comprising a robotic arm, a processor, and a storage medium. The storage medium stores computer instructions, and when the processor executes the computer instructions, it executes the aforementioned wearable robotic arm embodied intelligence method to drive the robotic arm to perform a target task.
[0133] The present invention is not limited to the aforementioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.
Claims
1. A wearable robotic arm embodied intelligence method, characterized in that: include: Obtain visual language information for the target task; Matching a first motion trajectory from a first database according to the visual language information, and inputting the visual language information into a visual language motion model to obtain a second motion trajectory; the first database is constructed by imitation learning; The method of matching a first motion trajectory from a first database according to the visual language information and inputting the visual language information into a visual language motion model to obtain a second motion trajectory comprises: matching a first motion trajectory from a first database according to the visual language information; if the robot arm (3) fails to perform a target task according to the first motion trajectory, inputting the current visual language information into the visual language motion model to obtain a second motion trajectory; if the robot arm (3) successfully performs the target task according to the second motion trajectory, combining the matched first motion trajectory with the second motion trajectory to update the corresponding first motion trajectory in the first database; generating a robot arm motion instruction according to the first motion trajectory and the second motion trajectory; The robot arm motion instruction is executed to drive the robot arm (3) to perform the target task.
2. The wearable robotic arm embodied intelligence method according to claim 1, characterized in that: The method for constructing the first database includes: Obtaining the environment information for executing the current operation task and the language information describing the current operation task; Detecting the spatial structural features of the target object and the motion trajectory of performing the current operation task from the environmental information; The motion trajectory, the spatial structure features of the target object, and the language information obtained by executing the current operation task are associated and stored in the first database.
3. The wearable robotic arm embodied intelligence method according to claim 2, wherein: The method for constructing the first database further includes: Performing hash indexing on the spatial structure features and language information of the target object; and, The motion trajectory and the spatial structural features of the target object are clustered to generate different task templates.
4. The wearable robotic arm embodied intelligence method according to claim 2 or 3, wherein: Detecting the spatial structural features of the target object from the environmental information includes: Denoising the point cloud data in the environmental information to crop first point cloud data in the workspace; preprocessing the image data in the environmental information to obtain first image data; spatially aligning the first point cloud data and the first image data; filtering out background information from the first point cloud data, clustering the remaining point cloud data, and separating different target objects; The categories and positions of different target objects are detected from the first image data, and the three-dimensional positions and shape sizes of the different target objects are matched in the first point cloud data according to the positions.
5. The wearable robotic arm embodied intelligence method according to claim 1, wherein: When obtaining the second action trajectory through the visual language action model, it also includes: The visual language information and the second motion trajectory are associated and stored in the first database.
6. The wearable robotic arm embodied intelligence method according to claim 1, wherein: Inputting the visual language information into a visual language action model to obtain a second action trajectory includes: Inputting the visual language information into a visual language action model to obtain a target pose; The path of the robot arm (3) from the initial posture to the target posture is optimized to obtain the second motion trajectory.
7. The wearable robotic arm embodied intelligence method according to claim 6, characterized in that: The optimization of the path of the robot arm (3) from the initial posture to the target posture includes: Identifying the spatial structural features of various obstacles in the environment where the robot arm (3) performs the target task based on the visual language information; The path with the minimum cost is searched between the initial and target positions. The cost of the path is composed of the length of the path and the path penalty caused by colliding with obstacles. The path penalty caused by colliding with obstacles is calculated by the penalty weight and potential penalty of each obstacle collided with. The penalty weight and potential penalty of each obstacle are matched according to the spatial structural characteristics of the obstacle.
8. The wearable robotic arm embodied intelligence method according to claim 7, wherein: Methods for constructing a path between an initial pose and a target pose include: With the initial pose as the root pose node, iteratively perform the following steps until the target pose is reached or the maximum number of iterations is reached: Generate a random pose node in the working area of the robot arm (3), and search for a pose node with the minimum jump cost among the generated pose nodes; Extend a predetermined step length from the searched pose node toward the random pose node to obtain a candidate pose node. If there is no rigid obstacle in the path from the searched pose node to the candidate pose node, then add the candidate pose node to the generated pose node and use the searched pose node as the parent node. All generated pose nodes are searched within a predetermined radius around the newly added generated pose node, and the pose node with the lowest jump cost to the root pose node is determined as the parent node.
9. An embodied intelligent wearable robotic arm, characterized in that: The wearable robotic arm comprises a robotic arm (3), a processor (6) and a storage medium, wherein the storage medium stores computer instructions. When the processor (6) runs the computer instructions, the wearable robotic arm embodied intelligence method according to any one of claims 1 to 8 can be executed to drive the robotic arm (3) to perform a target task.
Citation Information
Patent Citations
Mechanical arm control method, system and equipment for realizing multi-mode general operation task
CN119772905A