Wearable mechanical arm and intelligent method thereof
By obtaining visual language information and using imitation learning and visual language action models to generate action trajectories, the problem of insufficient intelligence of wearable robotic arms is solved, high intelligence and adaptive environment perception is achieved, and task completion rate and response speed are improved.
Patent Information
- Application Number
- CN202510813247.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing wearable robotic arms are not intelligent enough and have limited environmental perception capabilities, so they cannot adapt and complete tasks efficiently in a diverse and dynamically changing environment.
By obtaining visual language information, using imitation learning to build a first database and match the action trajectory, or inputting a visual language action model to generate the action trajectory, combining inverse kinematics methods to drive the robotic arm to perform tasks, enhance environmental perception and adaptability.
It improves the intelligence level and environmental perception ability of the robotic arm, improves the task completion rate and response speed, and enhances the adaptability to the environment.
Smart Images

Figure CN120347769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a wearable robotic arm and its embodied intelligence method. Background Art
[0002] Existing intelligent robotic arm devices mainly focus on industrial scenarios and are usually designed as fixed or semi-fixed devices for tasks such as assembly line production, welding, and handling. These robotic arms rely on pre-set operation processes, have fixed working positions, limited task scopes, lack portability and flexibility, and are difficult to adapt to diverse and dynamically changing environments. In addition, the design of traditional robotic arms focuses more on precision and repeatability and usually provides only a single control mechanism, such as operating through a pre-set program or remote control. Although this single control method is suitable for industrial scenarios, it often shows obvious limitations in situations that require quick response or complex task processing.
[0003] Outside of industrial scenarios, with the rapid development of artificial intelligence and embodied intelligence technologies, robotic arms are gradually evolving towards multi-functionality, portability, and intelligence. For example, wearable robotic arms, as an emerging technology, are beginning to be applied in scenarios such as assisted rehabilitation, personal life, and rescue operations.
[0004] However, compared with industrial robotic arms, the development of wearable robotic arms is still in its infancy, and there are multiple key problems that need to be solved urgently. Existing wearable robotic arms are weak in intelligent interaction and fail to achieve natural synchronization and efficient cooperation with human movements. When dealing with dynamic tasks, the reaction speed and decision-making ability of the robotic arm are insufficient, and it cannot balance real-time performance and complex task processing. The traditional wearable robotic arm has a relatively single way of perceiving the external environment and is usually used as a feedback signal when performing tasks. The control method of traditional robotic arms relies on pre-set rules and programs, lacks adaptability, and cannot continuously optimize behavior through human-machine interaction and feedback mechanisms.
[0005] Against the above background, society's demand for intelligent devices that can break through the limitations of traditional robotic arms is increasing. Especially in the field of personalized portable devices, there is an urgent need for a more flexible and intelligent solution to adapt to diverse task scenarios and user needs. Summary of the Invention
[0006] The object of the present invention is to: address all or part of the above problems and provide a wearable robotic arm and its embodied intelligence method to solve the problems of insufficient intelligence and limited environmental perception ability of current wearable robotic arms.
[0007] The technical solution adopted by the present invention is as follows: An embodied intelligence method for a wearable robotic arm, which includes: Obtain the visual language information of the target task; Match the first action trajectory from the first database according to the visual language information, and / or input the visual language information into the visual language action model to obtain the second action trajectory; the first database is constructed by imitation learning; Generate a robotic arm action instruction according to the first action trajectory and / or the second action trajectory; Execute the robotic arm action instruction to drive the robotic arm to perform the target task.
[0008] On the other hand, the present invention also provides an embodied intelligent wearable robotic arm, which includes a robotic arm, a processor, and a storage medium. Computer instructions are stored in the storage medium. When the processor runs the computer instructions, the above-mentioned wearable robotic arm embodied intelligent method can be executed to drive the robotic arm to perform the target task.
[0009] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are: The wearable robotic arm and its embodied intelligent method designed in this application adaptively obtain the corresponding action trajectory by perceiving and recognizing the environment at the visual level and combining the description of the target task in language information, and then calculate the robotic arm action instruction through the inverse kinematics method, etc., so as to drive the robotic arm to perform the target task. It has the characteristics of high intelligence and strong adaptability. In addition, in the visual language information obtained in this application, spatial feature information and visual feature information are included at the visual level. The multi-modal information improves the ability to perceive the environment, thereby improving the accuracy of action trajectory planning and the task completion rate. And, this application designs two sets of action trajectory planning modes. The robotic arm can quickly obtain the target action trajectory according to the imitation learning record, or generate a rational action trajectory with the help of the visual language action model, so as to improve the response speed of the robotic arm to the target task. As the usage frequency increases, the learned action knowledge can be continuously enriched, and the response speed can be continuously improved. The two modes can be used in combination according to needs, further improving the adaptability to the environment. Description of the Drawings
[0010] The present invention will be described by way of examples and with reference to the drawings, where: Figure 1 is the flowchart of the wearable robotic arm embodied intelligent method provided by the embodiment of the present application.
[0011] Figure 2 、 Figure 3 are the structural diagrams of the wearable robotic arm in two different perspectives in the embodiment of the present application.
[0012] Figure 4 is the flowchart of the method for constructing the first database in the embodiment of the present application.
[0013] Figure 5 It is a state diagram showing the state of picking up a mobile phone in an embodiment of the present application.
[0014] Figure 6 It is a flowchart of a method for detecting spatial structure features and action trajectories in an embodiment of the present application.
[0015] Figure 7 It is a flowchart of a method for implementing the embodied intelligence of a wearable robotic arm in an embodiment of the present application.
[0016] In the figure, 1 is a robotic arm camera, 2 is a wearable backpack, 3 is a robotic arm, 4 is a rotating disk, 5 is a power and data line, 6 is a processor, 7 is a chest camera, 8 is a power supply, and 9 is a lidar. Detailed implementation manners
[0017] All features disclosed in this specification, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
[0018] Any feature disclosed in this specification (including any additional claims, abstract) can be replaced by other equivalent or features with similar purposes, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only an example of a series of equivalent or similar features.
[0019] Aiming at the deficiencies of traditional wearable robotic arms in aspects such as environmental perception ability and adaptability, the embodiments of the present application provide a wearable robotic arm and its embodied intelligence method, aiming to improve the intelligence level of the robotic arm and the ability to perceive the environment.
[0020] The method for implementing the embodied intelligence of the wearable robotic arm provided by the embodiments of the present application, as Figure 1 shown, includes the following steps: Step S1, obtaining visual language information of the target task.
[0021] In the embodiments of the present application, the visual language information includes visual information for environmental perception and language information for task description. Among them, the visual information includes spatial feature information and visual feature information to improve the spatial perception ability.
[0022] In some feasible implementation manners, the spatial feature information is the point cloud feature for environmental perception, and the visual feature information is the image feature for environmental perception. By scanning the working area of the target task with a lidar, point cloud data is obtained, and then the point cloud feature is obtained; by taking pictures of the working area of the target task with a camera, image data is obtained, and then the image feature is obtained.
[0023] As Figure 2 ,Figure 3 Shown is a structural diagram of a wearable robotic arm in some feasible embodiments. The wearable robotic arm includes a robotic arm 3, a robotic arm camera 1 fixedly installed on the robotic arm (which can be used for real-time status feedback), a wearable backpack 2 for wearing, a rotating disk 4 serving as a rotating base of the robotic arm, a power and data line 5, a processor 6 responsible for data processing, a chest camera 7 installed on the chest, a power source 8 for powering the wearable robotic arm, and a lidar 9 installed on the chest. A microphone for collecting voice (not shown) is installed on the wearable backpack 2 or other positions. When the wearable robotic arm is worn on a user, image data, such as RGB images, is captured by the chest camera 7 on the chest, and point cloud data is collected by the lidar 9 on the chest. The point cloud data includes three-dimensional spatial features of various objects and the background in the environment.
[0024] Step S2: Match a first action trajectory from a first database according to the visual language information. And / or input the visual language information into a visual language action model to obtain a second action trajectory.
[0025] The first database is associated with and stores visual language information and action trajectories. Through the visual language information, an attempt can be made to match the associated action trajectory from the first database. If a match is found, the first action trajectory is obtained. The first database is constructed through imitation learning.
[0026] The so-called imitation learning means recording the actions performed by a human when executing a target task, collecting the corresponding visual language information, and associating the visual language information with the recorded action trajectory.
[0027] As an optional embodiment, as Figure 4 shown, the method for constructing the first database includes the following process: Step S21: Obtain the environmental information of the current operation task and the language information describing the current operation task.
[0028] The environmental information is the information perceived by the environment, including the spatial feature information and visual feature information of the working environment. The working environment includes the target object to be operated and the human arms operating the target object. According to the previous embodiments, it is the image data of the working environment collected by the chest camera 7 and the point cloud data of the working environment collected by the lidar 9.
[0029] The language information is the language description of the current task, indicating the current operation behavior. The collection method of the language information can be in the form of text collection or voice collection. In addition, the language information can be the keywords extracted from the collected descriptive language. The extracted keywords include actions, targets, and constraints.
[0030] In some specific embodiments, the method for obtaining language information includes: a. Collect language instructions in the form of speech or text. For the language instructions collected in the form of speech, use a speech recognition tool to convert them into text form.
[0031] b. Segment the language instructions in text form, remove punctuation marks and stop words; use part-of-speech tagging and syntactic analysis means to extract the subject-predicate-object structure of the language instructions; identify the keywords in the language instructions according to a predefined task template; c. Use a pre-trained language model to convert the text into word vectors in the form of high-dimensional semantic feature embeddings, and extract the action information, target information, and constraint conditions in the embeddings.
[0032] d. If the language instruction is part of a multi-turn conversation, associate the current language instruction with the previous content and complete the missing keywords.
[0033] For example, assume that the current operation task is to operate a mobile phone. As Figure 5 shown, the image data of the working environment is captured by the chest camera 7 to record the visual feature information of the mobile phone, including color, shape, and surface texture, etc., as well as the movement process of the human arms operating the mobile phone. The working environment is also scanned three-dimensionally by the lidar 9 to obtain the point cloud data including the mobile phone and the human arms, to record the three-dimensional position and spatial size of the mobile phone, as well as the action trajectory of the human arms. The language description of the current operation task, such as "picking up the mobile phone", is synchronously collected through the integrated microphone and transmitted to the processor 6 through the power data line 5 for post-processing. After the processor 6 converts the voice data into text data, the keywords are extracted, such as the action keyword "pick up" and the target keyword "mobile phone", and there is no constraint condition keyword. In addition, if the language description of the operation task is collected in text form, the processor 6 directly extracts the keywords. For the storage form of the language information, it is the vector representation of the semantic features of the keywords, that is, the keywords are converted into corresponding word embeddings by using a pre-trained language model, so as to extract the semantic feature vectors of the keywords.
[0034] S22. Detect the spatial structure features of the target object and the action trajectory of performing the current operation task from the environmental information.
[0035] The working environment contains different objects, and different objects have different spatial structure features. The target object can be identified through the spatial structure features.
[0036] In some feasible embodiments, see the appendix Figure 6 , the method for detecting the spatial structure features of the target object from the environmental information includes the following process: Step S221: Denoise the point cloud data in the environmental information and crop out the first point cloud data within the working space. Preprocess the image data in the environmental information to obtain the first image data.
[0037] The denoising of the point cloud data can be completed by the method of formula (1): Formula (1): ; In the formula, represents the denoised point cloud data, is the th data point of the point cloud data, is the Gaussian filtering kernel function, is the standard deviation, is the total number of data points.
[0038] Through formula (1), the point cloud data is smoothed, and after removing isolated points and noise points, the remaining data points become more continuous and smooth.
[0039] After denoising the point cloud data, crop the point cloud data to retain only the valid point cloud on the working desktop, which is the first point cloud data.
[0040] For the preprocessing of the image data, in some feasible implementation manners, it may include denoising, color adjustment, etc. After preprocessing, the first image data is obtained.
[0041] S222: Align the first point cloud data and the first image data in space.
[0042] The alignment of the data of the above two modalities can be completed according to the calibration parameters of the lidar 9 and the chest camera 7. Taking the alignment of the first image data to the first point cloud data as an example, the data alignment process can be completed through the overall rigid transformation matrix T of formula (2).
[0043] Formula (2): ; In the formula, R is the rotation matrix and t is the translation vector. Through the calibration parameters of the chest camera 7 and the lidar 9, based on the coordinate points representing the same position, the above overall rigid transformation matrix T can be obtained, and then the alignment between the first image data and the first point cloud data can be completed.
[0044] S223: Filter out the background information from the first point cloud data, cluster the remaining point cloud data, and separate different target objects.
[0045] For the first point cloud data, static background information such as the ground and walls can be extracted using point cloud segmentation algorithms, etc. After filtering out the background information, the remaining point cloud data is clustered, and different target objects are separated according to the clusters obtained from the clustering. As an alternative implementation, the K-means clustering algorithm is used for clustering the point cloud data, and the formula is as follows: Formula (3): ; In the formula, k represents the serial number of the cluster, is the k-th cluster, is the i-th data point, is the clustering center of the k-th cluster, n is the number of data points participating in the clustering, represents the calculation of norm.
[0046] By clustering the point cloud data, the point clouds of different objects can be separated, and thus the point cloud data of the target object can be isolated. The point cloud data carries the spatial feature information of the object, and based on this, information such as the three-dimensional position and shape size (dimensions) of different objects can be obtained respectively. For example, in the previous embodiment, if the target object is a mobile phone, then through this method, the point cloud data of the mobile phone can be separated from the collected point cloud data, and thus information such as the three-dimensional position and shape size of the mobile phone can be obtained.
[0047] S224. Detect the categories and positions of different target objects from the first image data, and respectively match the three-dimensional positions and shape sizes of different target objects in the first point cloud data according to the positions.
[0048] Through the target detection algorithm, the categories and positions of different target objects can be detected from the first image data. Image data can only reflect the pixel positions of objects in the image and is difficult to reflect their specific positions in the three-dimensional space, while the point cloud data can make up for this deficiency. According to the positions of the objects detected from the first image data, the point cloud data of the objects at the corresponding positions is matched from the first point cloud data aligned with the first image data, so as to identify the three-dimensional positions and shape sizes of the target objects based on the point cloud data of the target objects. For example, if the position of a mobile phone (i.e., the category) is identified in the first image data, then according to the position of the mobile phone in the first image data, the point cloud data of the mobile phone is matched from the first point cloud data, and the three-dimensional position and shape size of the mobile phone are identified.
[0049] The category, three-dimensional position and shape size of the target object constitute the spatial structure characteristics of the target object.
[0050] In addition, the above-mentioned motion trajectory can also be obtained by a method similar to that for obtaining the spatial structure characteristics of the target object. As a feasible implementation, refer to Appendix Figure 6 , and the method for obtaining the motion trajectory includes: S225. Detect the joint positions of the human arms from the first image data of each key frame, and match the three-dimensional joint positions in the first point cloud data according to the joint positions.
[0051] The so-called key frame refers to the image data and point cloud data collected at key positions extracted from the continuously and synchronously collected image data and point cloud data. Each key frame corresponds to first image data and first point cloud data. The alignment operation between the first image data and the first point cloud data has been introduced above. Through human pose estimation algorithms, etc., the joint points of the human arms can be located from the first image data, including key points such as shoulders, elbows, and wrists. Then, the joint points in the two-dimensional space are mapped to the three-dimensional space through the calibration mapping relationship between the chest camera 7 and the lidar 9, so that the three-dimensional positions of the joint points can be separated from the first point cloud data.
[0052] S226. Perform trajectory tracking on the three-dimensional positions of the joint points in the key frame, and calculate the three-dimensional motion trajectory of the human arms.
[0053] The joint points are at the hub positions of the arms, and the motion trajectory of the joint points can reflect the three-dimensional motion trajectory of the human arms.
[0054] S227. Use the Kalman filtering method or the optical flow method to smooth the three-dimensional motion trajectory.
[0055] Taking the Kalman filtering method as an example, its processing is shown in formulas (4) and (5).
[0056] Formula (4): ; Formula (5): ; In the formula, k represents the current moment, k - 1 represents the previous moment, is the state estimate at the current moment, A is the state transition matrix, B is the control matrix, is the control input, is the process noise, and are the covariance matrices of the estimates at the current moment and the previous moment respectively, and Q is the covariance of the process noise.
[0057] S228. Extract the time series features of the three-dimensional motion trajectory, and represent the action trajectory of performing the operation task in the form of a feature vector.
[0058] The extracted time series features include features such as the speed, acceleration, and motion direction of the human arms.
[0059] In another feasible implementation manner, the method for obtaining the action trajectory may also include: The operation task is performed by a human holding a robotic arm, and the motion trajectory of the robotic arm during the performance of the operation task is synchronously recorded. This motion trajectory is also in the form of a time series feature vector, including features such as the speed, acceleration, and motion direction of the robotic arm.
[0060] Step S23: Correlatively store the action trajectory obtained by performing the current operation task, the spatial structure features of the target object, and the language information in the first database.
[0061] The above method realizes, through imitative learning of the human arms, correlatively storing the action trajectory, the spatial structure features of the target object, and the language information obtained from each performed operation task as an entry, and each piece of information is stored in the form of a word vector, thereby constructing the first database. The first database can satisfy the quick response to subsequent target tasks, quickly match a feasible action trajectory for the target task, and thus quickly complete the target task.
[0062] In addition, in some alternative embodiments, in order to improve the retrieval efficiency in the first database, on the basis of correlatively storing the action trajectory, the spatial structure features of the target object, and the language information, the method for constructing the first data further includes: Step S231: Perform hash indexing on the spatial structure features of the target object and the language information. In this way, when retrieving, the amount of character data to be calculated for traversing each entry can be effectively reduced, and the retrieval efficiency can also be improved while effectively reducing the data storage amount.
[0063] Step S232: Cluster the action trajectory and the spatial structure features of the target object to generate different task templates.
[0064] Representing the action trajectory and the spatial features of the target object in the form of templates can further simplify the matching process and improve the retrieval efficiency.
[0065] After the first database is constructed, according to the visual language information of the current target task, that is, the language information describing the target task, and the environmental information (including point cloud data and image data) obtained in the current scene, a retrieval is performed in the first database. If the retrieval is successful, the corresponding action trajectory is obtained. If the retrieval fails, that is, the first database does not include an action trajectory available for performing the current task, then a second action trajectory needs to be obtained through the visual language action model to perform the target task, as Figure 7 shown.
[0066] The method for obtaining the second action trajectory is obtained by leveraging the inference ability of artificial intelligence. The vision-language-action model needs to be pre-trained to possess the inference ability. The input features of the vision-language-action model are vision-language information, and the output result is the target pose, that is, the target position and orientation that the robotic arm needs to reach. The target position is a three-dimensional position, representing the target coordinates of the end effector of the robotic arm in three-dimensional space, and should be used to locate the specific position of the target object. The target orientation is a three-dimensional orientation, used to control the orientation and movement of the end effector of the robotic arm. The orientation is represented by the roll, pitch, yaw of Euler angles or rotation vector parameters to ensure that the end effector of the robotic arm is aligned with the target object. The movement is represented by the gripper movement parameters, indicating the opening and closing state of the end effector of the robotic arm, used to represent grasping or releasing the target object. The vision-language-action model adopts a Transformer structure to achieve the fusion and understanding of multi-modal features (spatial features, visual features, and text features). Its model representation is: , where Image represents image data, PointCloud represents point cloud data, Text represents language information, TargetAction represents the target pose, and Transformer() represents the mapping function of the vision-language-action model for the input features.
[0067] Training the vision-language-action model includes a data preparation stage and a model training stage.
[0068] In the data preparation stage, a large amount of vision-language information and corresponding target poses (the positions and orientations after the execution of relevant action sequences) when performing different operation tasks are first obtained as samples. The vision-language information includes image data, point cloud data, and language information describing the operation tasks, which is multi-modal data. For example, data samples for robotic arm training can be downloaded from the Open X-Embodiment dataset.
[0069] In the model training stage, a large number of data samples obtained in the data preparation stage are divided into a training set and a test set, which are used for training and testing the vision-language-action model respectively. Among them, the vision-language information (i.e., image data, point cloud data, and language information) of the samples is input into the vision-language-action model, and the corresponding target pose is used as the ground truth to iteratively optimize the model parameters of the vision-language-action model. By minimizing the loss between the predicted value and the ground truth of the target pose, a vision-language-action model with vision-language information as the input and the target pose as the output is trained.
[0070] As an optional implementation manner, the loss function for training the vision-language-action model is: Formula (6): ; where \(L\) represents the calculated loss, \(t\) refers to the serial number of the sample, is the number of samples, is the target pose predicted by the vision-language-action model for the \(t\)-th sample, is the actual target pose of the \(t\)-th sample, denotes the operation of the norm of.
[0071] After obtaining the trained vision-language-action model, when performing the target task, the vision-language information is input into the vision-language-action model to obtain the target pose. On this basis, the path of the robotic arm from the initial pose to the target pose is optimized through a trajectory planning algorithm, etc., and the required second action trajectory can be obtained.
[0072] In addition, in the embodiments of the present application, the vision-language-action model is also used to continuously expand the alternative entries in the first database. In some feasible embodiments, when obtaining the second action trajectory through the vision-language-action model, the vision-language information and the second action trajectory are also associated and stored in the first database.
[0073] For this design, there are the following three operation modes: Mode 1: First, try to match the first action trajectory from the first database according to the vision-language information.
[0074] If the first action trajectory is not matched in the first database, and then the second action trajectory is obtained by using the vision-language-action model, the vision-language information and the second action trajectory are directly associated and stored in the first database to form a new entry. If there are specific requirements for the storage format in the first database, corresponding format conversion is performed according to the format requirements of the first database, such as performing hash indexing on the spatial structure features and language information of the target object.
[0075] Or, if the first action trajectory is matched in the first database, but the final execution of the target task fails, and then the second action trajectory is obtained by using the vision-language-action model, the vision-language information and the second action trajectory are directly associated and stored in the first database to form a new entry.
[0076] Mode 2: The mode of only inputting the vision-language information into the vision-language-action model to obtain the second action trajectory.
[0077] If the second action trajectory is directly obtained using the visual language action model without matching in the first database, first match the visual language information in the first database. If a first action rule is matched, use the second action trajectory to replace the first action rule. If no first action rule is matched, directly associate and store the visual language information and the second action trajectory in the first database to form a new entry. Similarly, if there are specific requirements for the storage format in the first database, perform corresponding format conversion according to the format requirements of the first database.
[0078] Method 3: A method of matching the first action trajectory from the first database according to the visual language information and also inputting the visual language information into the visual language action model to obtain the second action trajectory.
[0079] Match the first action trajectory from the first database according to the visual language information. If the manipulator fails to execute the target task according to the first action trajectory, input the current visual language information into the visual language action model to obtain the second action trajectory; if the manipulator successfully executes the target task according to the second action trajectory, combine the matched first action trajectory with the second action trajectory to update the corresponding first action trajectory in the first database.
[0080] As an optional implementation manner, the above method for optimizing the path of the manipulator from the initial pose to the target pose includes: a. Based on the visual language information, identify the spatial structure characteristics of each obstacle in the environment where the manipulator executes the target task.
[0081] The obstacle is other items in the environment except the target object. The spatial structure characteristics of the obstacle are exactly the same as the nature of the spatial structure characteristics of the items described above, that is, including the type, three-dimensional position, and shape and size of the obstacle. Therefore, the extraction method of the spatial structure characteristics of the obstacle is the same as that in the previous embodiment and will not be elaborated here.
[0082] According to the identified spatial structure characteristics of the obstacle, the degree of influence on the manipulator caused by crossing the obstacle can be judged, that is, whether the obstacle can be crossed is judged. For example, a rigid obstacle (such as a machine frame, glass, etc.) can be crossed statically, and some flexible items (such as cotton, leaves, etc.) can be crossed. The degree of influence on the manipulator caused by the obstacle is called the potential penalty of the obstacle, which is represented by where the subscript i is the identifier of the obstacle serial number. In addition, the potential penalty of the obstacle is not equivalent to the path cost. In the embodiments of the present application, the potential penalties of different obstacles are also equivalent to the path cost dimension to facilitate the calculation and comparison of the path cost. This equivalent factor is called the penalty weight of the obstacle, which is represented by Representation. The potential penalty and penalty weight of the obstacle are both obtained by matching from the association relation table storing item types, potential penalties, and penalty weights through the spatial structure features (types therein) of the obstacle.
[0083] According to the spatial structure features (3D position and shape size therein) of each identified obstacle, it can be determined whether the planned path will cross the obstacle. If so, the corresponding obstacle will be activated to increase the path cost.
[0084] b. At the initial pose and the target pose search for the path with the minimum cost. Among them, the cost of the path consists of the length of the path and the path penalty generated by colliding with obstacles in the path. The path penalty generated by colliding with obstacles is calculated from the penalty weight and potential penalty of each collided obstacle. Specifically, there is: Formula (7): ; In the formula, represents the velocity on the path, that is, the change rate of the path length, t represents time, and the length of the path can be obtained by integrating the path velocity.
[0085] Through Formula (7), the cost of the path for the robotic arm to reach the target pose from the initial pose can be calculated. On this basis, select the path with the minimum cost as the second motion trajectory.
[0086] As for how to construct alternative paths between the initial pose and the target pose, in some alternative embodiments, the following methods are included: a. Taking the initial pose as the root pose node, iteratively execute the following steps until reaching the target pose or reaching the maximum number of iterations: b. Generate a random pose node within the working area of the robotic arm, and search for the pose node with the minimum jump cost among the generated pose nodes. The so-called jump cost can only include the Euclidean distance between nodes, or can also include the weighted norm of the joint angle change amount, and the weights of different joints are different; or can also include other factors.
[0087] c. Extend a predetermined step length from the searched pose node in the direction of the random pose node to obtain a candidate pose node. If the path from the searched pose node to the candidate pose node has no rigid obstacles (it can be determined whether to cross from the 3D position and shape size of the obstacle, and it can be determined whether it belongs to a rigid obstacle from the type of the obstacle), then add the candidate pose node to the generated pose nodes and use the searched pose node as the parent node.
[0088] d. Search for all generated pose nodes within a predetermined radius (denoted as r, and the value can be set by yourself) around the newly added generated pose node, and determine the pose node with the lowest jump cost to the root pose node (i.e., the initial pose) as the parent node from them.
[0089] The process of step d includes traversing each searched pose node as an alternative parent node, calculating the total jump cost from the root pose node to the newly added pose node. If the total jump cost is lower than the total jump cost of reaching the root pose node through the current parent node, then replace the current parent node with the currently traversed alternative parent node. The same applies to traversing subsequent alternative parent nodes.
[0090] It should be noted that in step a, within the set maximum number of iterations, the target pose will surely be reached. Each randomly generated pose node in each iteration has a certain probability (such as 5%) of being the target pose node, and within the set maximum number of iterations, there will surely be an iteration in which the randomly generated pose node is the target pose node.
[0091] The motion trajectory belongs to a sequence of actions composed of a series of discrete pose nodes and cannot directly guide the movement of the robotic arm. It is necessary to generate drive instructions that can directly drive the movement of the robotic arm according to the motion trajectory.
[0092] S3. Generate robotic arm motion instructions according to the first motion trajectory and / or the second motion trajectory.
[0093] As an alternative implementation, after obtaining the first motion trajectory and / or the second motion trajectory, the kinematic planning algorithm is also used to optimize the motion trajectory to ensure the smoothness and feasibility of the robotic arm motion.
[0094] After obtaining the first motion trajectory and / or the second motion trajectory, the corresponding motion trajectory is converted into robotic arm motion instructions executable by the robotic arm controller through the inverse kinematics (IK) algorithm, including a series of joint angle control instructions.
[0095] Assume that the target position of the end effector of the robotic arm is represented as , where x, y, z are the coordinates of the target object in three-dimensional space. The motion of the robotic arm is controlled by a set of joint angles . The purpose of the calculation is to calculate the angle q of each joint through the IK algorithm so that the end effector of the robotic arm reaches the target position . The solution of the inverse kinematics is realized according to the following formula (8).
[0096] Formula (8): ; In the formula, f() represents the mapping function of the IK algorithm; is the target position; is the speed or acceleration of the target position; is a control parameter (such as inertia, damping, etc.), which is determined by the robotic arm in specific applications and is a parameter given in advance. Given the target position and other relevant control information, a sequence of robotic arm joint angles that meets the task requirements can be obtained through inverse kinematics solution.
[0097] S4. Execute the robotic arm action instruction to drive the robotic arm to perform the target task.
[0098] After obtaining the robotic arm action instruction, the action instruction is transmitted to the low-level controller through ROS (Robot Operating System), and the low-level controller is responsible for controlling the rotation of specific motors according to the robotic arm action instruction, thereby controlling the corresponding joint angles.
[0099] According to the idea of the present application, in an embodiment of the present application, an embodied intelligent wearable robotic arm is further provided. The wearable robotic arm includes a robotic arm, a processor, and a storage medium. Computer instructions are stored in the storage medium, and when the processor runs the computer instructions, the above-mentioned embodied intelligent method of the wearable robotic arm can be executed to drive the robotic arm to perform the target task.
[0100] The present invention is not limited to the foregoing specific embodiments. The present invention extends to any new feature or any new combination disclosed in this specification, as well as any new method or process step or any new combination disclosed.
Claims
1. A method for embodied intelligence of a wearable robotic arm, characterized in that, Including: Obtain the visual language information of the target task; Match the first action trajectory from the first database according to the visual language information, and / or input the visual language information into the visual language action model to obtain the second action trajectory; the first database is constructed by imitation learning; Generate a robotic arm action instruction according to the first action trajectory and / or the second action trajectory; Execute the robotic arm action instruction to drive the robotic arm (3) to execute the target task.
2. The embodied intelligence method of the wearable robotic arm according to claim 1, wherein The method for constructing the first database includes: Obtain the environmental information for executing the current operation task and the language information describing the current operation task; Detect the spatial structure features of the target object and the action trajectory for executing the current operation task from the environmental information; Associatively store the obtained action trajectory, the spatial structure features of the target object, and the language information for executing the current operation task into the first database.
3. The embodied intelligence method of the wearable robotic arm according to claim 2, characterized in that, The method for constructing the first database further includes: Perform hash indexing on the spatial structure features and the language information of the target object; and, Cluster the action trajectory and the spatial structure features of the target object to generate different task templates.
4. The wearable robotic arm embodied intelligence method according to claim 2 or 3, characterized in that, Detecting the spatial structure features of the target object from the environmental information includes: Denoise the point cloud data in the environmental information and crop the first point cloud data within the working space; preprocess the image data in the environmental information to obtain the first image data; Perform spatial alignment on the first point cloud data and the first image data; Filter out the background information from the first point cloud data, cluster the remaining point cloud data, and separate different target objects; Detect the categories and positions of different target objects from the first image data, and respectively match the three-dimensional positions and shape dimensions of different target objects in the first point cloud data according to the positions.
5. The wearable robotic arm embodied intelligence method according to claim 1, characterized in that, When obtaining the second action trajectory through the visual language action model, it further includes: Associatively store the visual language information and the second action trajectory into the first database.
6. The wearable robotic arm embodied intelligence method according to claim 1, wherein Inputting the visual language information into the visual language action model to obtain the second action trajectory includes: Input the visual language information into the visual language action model to obtain the target pose; Optimize the path for the robotic arm (3) to reach the target pose from the initial pose to obtain the second action trajectory.
7. The embodied intelligence method of the wearable robotic arm according to claim 6, wherein, The method for optimizing the path for the robotic arm (3) to reach the target pose from the initial pose includes: Based on the visual language information, identify the spatial structure features of each obstacle in the environment where the robotic arm (3) executes the target task; Search for the path with the minimum cost between the initial pose and the target pose; wherein, the cost of the path is composed of the length of the path and the path penalty generated by colliding with obstacles; the path penalty generated by colliding with obstacles is calculated from the penalty weight and the potential penalty of each collided obstacle; the penalty weights and potential penalties of each obstacle are obtained by matching according to the spatial structure features of the obstacles.
8. The embodied intelligence method of the wearable robotic arm according to claim 7, characterized in that, The method for constructing a path between the initial pose and the target pose includes: Taking the initial pose as the root pose node, iteratively execute the following steps until reaching the target pose or reaching the maximum number of iterations: Generate a random pose node within the working area of the robotic arm (3), and search for the pose node with the minimum jump cost among the generated pose nodes; Extend a predetermined step length from the searched pose node in the direction of the random pose node to obtain a candidate pose node. If there is no rigid obstacle on the path from the searched pose node to the candidate pose node, add the candidate pose node to the generated pose nodes, and use the searched pose node as the parent node; Search for all generated pose nodes within a predetermined radius around the newly added generated pose node, and determine the pose node with the lowest jump cost to the root pose node as the parent node.
9. The wearable robotic arm embodied intelligence method according to claim 1, wherein The matching of the first action trajectory from the first database according to the visual language information, and the input of the visual language information into the visual language action model to obtain the second action trajectory, includes: Match the first action trajectory from the first database according to the visual language information. If the robotic arm (3) fails to execute the target task according to the first action trajectory, input the current visual language information into the visual language action model to obtain the second action trajectory; if the robotic arm (3) successfully executes the target task according to the second action trajectory, combine the matched first action trajectory with the second action trajectory to update the corresponding first action trajectory in the first database.
10. An embodied intelligent wearable robotic arm, characterized in that, It includes a robotic arm (3), a processor (6) and a storage medium. Computer instructions are stored in the storage medium. When the processor (6) runs the computer instructions, it can execute the wearable robotic arm embodied intelligence method according to any one of claims 1-9 to drive the robotic arm (3) to execute the target task.
Citation Information
Patent Citations
Mechanical vision-based robot welding method
CN109465582A
Hotel service method and device based on robot
CN111222939A
Mechanical arm control method, device and equipment and storage medium
CN119319568A
Natural language control method for humanoid robot
CN119610090A
Mechanical arm control method, system and equipment for realizing multi-mode general operation task
CN119772905A
Cited By
Human-computer interaction method and system based on vision-language-action model
CN121245915A
A human-computer interaction method and system based on a vision-language-action model
CN121245915B