A Multimodal Demonstration Method and System for Robot Assembly Skill Learning
Through multimodal data acquisition and fusion model, the robot assembly operation is identified and segmented, which solves the problem of low efficiency of complex assembly operation trajectories in the prior art, and realizes high-precision and high-efficiency robot assembly motion planning.
Patent Information
- Application Number
- CN202211167429.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-23
AI Technical Summary
The prior art is difficult to effectively solve the problems of complex, high-precision and agile robot assembly operations, and the problem of low trajectory efficiency, and the data obtained by a single motion capture system is insufficient, so effective trajectory optimization cannot be carried out.
The multimodal data acquisition device is used to collect assembly trajectory and video data, and the assembly operations are identified and segmented through the multimodal fusion continuous action recognition and segmentation model to generate a robot assembly motion plan.
By integrating video and six-dimensional trajectory information, identifying and segmenting assembly action elements, the success rate and efficiency of robot assembly are improved, and are suitable for complex and high-precision assembly tasks.
Smart Images

Figure CN115565243B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal demonstration method and system for robot assembly skill learning, and particularly to a method for recognizing and segmenting manual assembly actions by multi-modal fusion combining video and precise pose, belonging to the technical field of action recognition and segmentation, and can be applied to robot assembly motion planning. Background Art
[0002] The assembly process is complex, with high precision requirements, and the working space is small, which brings high programming difficulty for robot technicians. For such problems, the traditional robot programming control methods have too high time cost and are not flexible enough, which not only limits the speed of robots to adapt to new tasks, but also restricts the application of robots in the assembly field. A technology that can quickly and flexibly transfer the assembly skills of workers on the assembly production line to robots is of great significance for expanding the application scope of robots and improving the automation level of precision assembly production lines.
[0003] However, the current research on demonstration learning in the assembly field lacks research on complex, high-precision, and dexterous robot assembly operations, and does not solve the problem of low efficiency of robot assembly trajectories. Moreover, the manual assembly demonstration data itself is not optimal, and the data information obtained only by a single motion capture system is insufficient, and it is impossible to clearly know the true meaning of each part of the trajectory to perform trajectory optimization that meets the constraint conditions and has a basis. When applied to industrial robots, the assembly efficiency and success rate are average. Summary of the Invention
[0004] The present invention provides a multi-modal demonstration method and system for robot assembly skill learning, aiming to solve at least one of the technical problems existing in the prior art.
[0005] The technical solution of the present invention on the one hand relates to a multi-modal demonstration method for robot assembly skill learning, including the following steps:
[0006] S100. Collect the assembly trajectory and assembly video of the target object and the human hand through a multi-modal data acquisition device, and associate the assembly trajectory and the assembly video;
[0007] S200. Based on the assembly action unit library, according to the assembly video and the assembly trajectory, recognize and segment the assembly operation of the target object through a continuous action recognition and segmentation model of multi-modal fusion to obtain a series of continuous single action elements;
[0008] S300. Based on the motion planning strategy library, according to the type of each single action element, obtain its corresponding motion planning strategy to generate a robot assembly motion planning for the entire assembly operation of the target object.
[0009] Further, in the step S200:
[0010] The continuous action recognition and segmentation model includes an action recognition model and a continuous action segmentation model based on a 3D convolutional neural network model; wherein, the action recognition module is trained by a single-action dataset, and the continuous action segmentation module is trained by a continuous-action dataset.
[0011] Further, the step S200 includes:
[0012] S210. Based on the assembly action unit library, the trajectory recognition module of the action recognition model based on the LSTM network model extracts features from the assembly trajectory to achieve trajectory action recognition, and the video action recognition module of the action recognition model based on the C3D network model performs video action recognition on the assembly video, and fuses the trajectory action recognition result and the video action recognition result to obtain the single-action recognition result of the target object;
[0013] S220. Based on the assembly action unit library, according to the single-action recognition result, and according to the assembly trajectory and the assembly video, the continuous action segmentation model segments the continuous assembly operation of the target object to obtain a series of continuous single action elements;
[0014] Wherein, the types of single action elements included in the assembly action unit library include carrying action elements, alignment action elements, plug-in action elements, fastening action elements, and tearing action elements.
[0015] Further, the network model of the action recognition module includes:
[0016] A first feature extraction sub-module for extracting assembly trajectory features; a first LSTM network for processing trajectory features;
[0017] Multiple single-action video processing sub-modules for processing the assembly video, wherein the single-action video processing sub-module consists of a first 3D convolutional layer and a first 3D pooling layer, and the zero-padding in the time dimension of the first 3D convolutional layer is replaced by partial information in the previous time step;
[0018] A first fully connected layer for processing the outputs of the multiple single-action video processing sub-modules;
[0019] Two second fully connected layers for fusing the output of the first LSTM network and the output of the first fully connected layer;
[0020] A softmax layer is arranged behind the two second fully connected layers.
[0021] Further, the network model of the continuous action segmentation model includes:
[0022] The second feature extraction sub-module for extracting assembly trajectory features; the second LSTM network for processing trajectory features;
[0023] Multiple continuous action video processing sub-modules for processing assembly videos, wherein the continuous action video processing sub-module consists of a second 3D convolutional layer and a third 3D convolutional layer, the third 3D convolutional layer is the bias term of the second 3D convolutional layer, and a ReLU activation function is set after the third 3D convolutional layer;
[0024] The third fully connected layer for processing the outputs of the multiple continuous action video processing sub-modules;
[0025] Two fourth fully connected layers for fusing the output of the second LSTM network and the output of the third fully connected layer;
[0026] A softmax layer is arranged behind the two fourth fully connected layers.
[0027] Furthermore, for the trajectory recognition module:
[0028] The data of the assembly trajectory includes the three-dimensional spatial position coordinates p and the unit quaternion q: , and the extracted trajectory features include velocity, acceleration, angular velocity, angular acceleration, curvature, and torsion;
[0029] wherein, the curvature and the torsion τ are calculated as follows:
[0030] ,
[0031] ,
[0032] In the formula, the dot above the vector represents the derivative with respect to time.
[0033] Furthermore, the training of the neural network model includes the following steps:
[0034] Optimize the recognition segmentation result through a weight function, and its calculation method is as follows:
[0035] ,
[0036] ,
[0037] wherein, the input of the neural network model is a total of 16 frames of video data from frame to frame, represents the frame in the model input data is recognized as Probability of an action Denote a weight function Denote the optimal weight
[0038] Furthermore, the training of the network model of the continuous action recognition and segmentation model includes the following steps:
[0039] Obtain the assembly trajectory and assembly video of a single assembly operation of the single action element, and gather them into a single-action dataset;
[0040] Obtain the assembly trajectory and assembly video of the continuous entire assembly operation of the target object, and gather them into a continuous-action dataset.
[0041] On the other hand, the technical solution of the present invention relates to a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above-mentioned method is implemented.
[0042] On the other hand, the technical solution of the present invention relates to a system applied to robot assembly skill learning, including: a computer device, and the computer device includes the above-mentioned computer-readable storage medium.
[0043] The beneficial effects of the present invention are as follows.
[0044] The multi-modal assembly demonstration method and system applied to robot assembly skill learning in the technical solution of the present invention are a multi-modal action acquisition, recognition, segmentation and planning method, system and device based on an artificial assembly action demonstration platform. By fusing the assembly video and the six-dimensional assembly trajectory, different assembly action units (single action elements) are recognized, so as to perform semantic segmentation on the continuous assembly trajectory, provide semantic support for assembly skill learning, and then formulate different robot assembly motion planning strategies to achieve the goal of improving the success rate and efficiency of robot assembly; propose to segment the assembly process into the starting state, the state to be aligned, the state to be assembled, and the completed state, and design an assembly action unit library to effectively decompose industrial assembly tasks; in the assembly skill learning, fuse the video dimension information and the accurate six-dimensional trajectory information for action recognition, which can not only make up for the disadvantage that it is difficult to obtain the accurate trajectory of the target in pure video action recognition, but also make up for the problem of low semantic segmentation accuracy in pure trajectory action recognition, and is applicable to most industrial assemblies; based on the action recognition and segmentation results, different motion planning strategies are adopted for different actions, which can improve the success rate and efficiency of robot assembly, and can be applied to the robot programming situation with complex assembly process and high precision requirements, and can quickly and flexibly transfer the assembly skills of workers to robot assembly, which is beneficial to expanding the application scope of robot assembly. Description of the Drawings
[0045] FIG. 1 is a schematic structural diagram and a physical photograph of a multi-modal data acquisition device according to an embodiment of the present invention.
[0046] Figure 2 is a schematic structural diagram and a physical photo of the human-machine interface of the multi-modal data acquisition device according to an embodiment of the present invention.
[0047] Figure 3 is a schematic structural diagram and a physical photo of the marking device of the multi-modal data acquisition device according to an embodiment of the present invention.
[0048] Figure 4 is the basic flowchart of the method according to the present invention.
[0049] Figure 5 is a schematic diagram of the network structure of the action recognition model of the method according to the present invention.
[0050] Figure 6 is a schematic diagram of the network structure of the continuous action segmentation model of the method according to the present invention.
[0051] Figure 7 is a schematic diagram of the recognition and segmentation result of the continuous action recognition and segmentation model of the method according to the present invention.
[0052] Figure 8 is the network model training flowchart of the continuous action recognition and segmentation model of the method according to the present invention.
[0053] Reference numerals:
[0054] 100, multi-modal data acquisition device; 110, optical motion capture platform; 111, capture camera; 112, human-machine interface; 113, connection part; 114, support part; 115, first reflective marking point; 116, connecting rod; 120, video acquisition camera; 130, data marking device; 131, actively emitting marking point; 132, second reflective marking point; 133, lamp bracket; 140, marking point switch.
[0055] It should be understood that the text content presented in the above-mentioned specification drawings has all been taken as part of the text content of this specification, and any combination with the embodiments in the specific implementation part of this specification, as long as it does not violate the technical solution principle of the present invention. Detailed implementation manners
[0056] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention in combination with the embodiments and drawings, so as to fully understand the purpose, solution and effects of the present invention.
[0057] It should be noted that, unless otherwise specified, when a certain feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. The singular forms of "a", "the", and "said" used in this article are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used in this article have the same meaning as those commonly understood by those skilled in the technical field of this technology. The terms used in the description of this article are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this article includes any combination of one or more of the related listed items.
[0058] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language (such as "for example", "such as", etc.) provided in this article is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.
[0059] Referring to FIGS. 1 to 3, a multi-modal data acquisition device 100 based on robot assembly skill learning according to the present invention, wherein the target object is disposed on an artificial assembly action demonstration platform. The multi-modal data acquisition device 100 includes an optical motion capture platform 110, a video acquisition camera 120, a data marking device 130, and a marking point switch 140. Among them, the optical motion capture platform 110 includes eight capture cameras 111 and a human-machine interface 112.
[0060] Eight capture cameras 111 are arranged in four pairs at two heights, and the eight capture cameras 111 are arranged in pairs on the four sides (four corners) of the learning platform. Referring to FIG. 1, the eight capture cameras 111 are mounted on four camera brackets. Two capture cameras 111 are mounted at different height positions on the same camera bracket, where the higher capture camera 111 is mounted at the top of the camera bracket, and the lower capture camera 111 is mounted in the middle of the camera bracket and below the higher capture camera 111. The four lower capture cameras 111 are on the same horizontal plane, that is, the four lower capture cameras 111 are all at the first height. At the same time, the four higher capture cameras 111 are on the same horizontal plane, that is, the four higher capture cameras 111 are all at the second height, where the first height is lower than the second height. The four camera brackets are evenly distributed at the four corners of the manual assembly operation demonstration platform, and the lenses of the eight capture cameras 111 all face the parts to be assembled (target objects). Further, in the embodiment of the present invention, an Optitrack optical motion capture platform 110 can be adopted, and the capture camera 111 adopts an Optitrack Prime X41. The above embodiments are only for illustrative purposes and are not specifically limited. By arranging multiple lenses at different heights and orientations, it helps to obtain diversified perspectives and can obtain the object motion trajectories of the target objects in the working space in all directions. At the same time, by adjusting the angles of the capture cameras 111, it is beneficial to make all the capture cameras 111 cover the entire working space and make the center of the field of view of the capture cameras 111 coincide with the center of the working space, so as to obtain more accurate assembly trajectory data.
[0061] The video acquisition camera 120 is arranged on one side of the manual assembly operation demonstration platform and is used to acquire the assembly operation video. By adjusting the angle of the video acquisition camera 120, it can cover the working area of the manual assembly operation demonstration platform. Further, the video acquisition camera 120 in the embodiment of the present invention can adopt an RGB camera. The above embodiment is only for illustration and is not specifically limited. The data marking device 130 is arranged between the video acquisition camera 120 and the target object. The data marking device 130 includes a lamp bracket 133, an active light-emitting marking point 131, and two second reflective marking points 132, and the active light-emitting marking point 131 is arranged within the shooting field of view of the video acquisition camera 120. As shown in Figure 2, the lamp bracket 133 is arranged on the learning platform, the active light-emitting marking point 131 is arranged at the top of the lamp bracket 133, and the two second reflective marking points 132 are respectively connected to the lamp bracket 133 through two connecting rods 116. The connecting rods 116 of the two second reflective marking points 132 are perpendicular to each other, and the two second reflective marking points 132 are arranged on the same horizontal plane. Further, the active light-emitting marking point 131 of the data marking device 130 in the embodiment of the present invention is set as an infrared lamp, and the wavelength of the infrared lamp is set to 850 mm, so as to facilitate ensuring that the optical movement after the active light-emitting marking point 131 lights up can be acquired. The turning on and off (lighting and extinguishing) of the active light-emitting marking point 131 is controlled by the marking point switch 140. Further, the marking point switch 140 can adopt a foot switch or an electric switch electrically connected to the control system.
[0062] Referring to Figure 3, the human-machine interface 112 includes a support part 114, a connecting part 113, and a plurality of first reflective marking points 115. The support part 114 is arranged on one side of the connecting part 113, and the other side of the connecting part 113 is connected to the target object. It should be noted that the connecting part 113 can also be used for connecting the target object and the robotic arm. The plurality of first reflective marking points 115 are arranged on the support part 114 and are used to obtain the multi-degree-of-freedom pose data of the target object. Further, the human-machine interface 112 in the embodiment of the present invention is provided with four first reflective marking points 115, and the four first reflective marking points 115 are not coplanar. Thus, by setting the four first reflective marking points 115, the first reflective marking points can form an asymmetric rigid body. By connecting the human-machine interface 112 with the target object, the six-degree-of-freedom pose data (assembly trajectory) of the target object can be obtained, and at the same time, occlusion can be reduced and the signal can be amplified, so as to facilitate obtaining more subtle and sensitive object movement data.
[0063] Further, in some specific embodiments of the present invention, the connecting portion 113 is a flat plate, the target object and the supporting portion 114 are respectively arranged on two opposite planes of the connecting portion 113, the supporting portion 114 is a long strip, one end of the supporting portion 114 is connected to the connecting portion 113, and the other end is connected to the first reflective marking point 115 through a connecting rod 116. One of the first reflective marking points 115 is arranged on the center line of the supporting portion 114 and on the side of the supporting portion 114 away from the connecting portion 113. The remaining three first reflective marking points 115 are all arranged on the first plane, and the first plane is perpendicular to the center line of the supporting portion 114. Two of the first reflective marking points 115 on the first plane are on the first straight line, and the connecting rod 116 of the remaining first reflective marking point 115 on the first plane is perpendicular to the first straight line. The above embodiments are only for illustration and are not specifically limited.
[0064] Referring to Figure 4 , the technical solution of the present invention is a multi-modal demonstration method (i.e., an assembly demonstration method) and system applied to the learning of robot assembly skills. Based on an artificial assembly action demonstration platform, a multi-modal assembly demonstration method applied to the movement trajectory of a robot assembly generally includes the following steps:
[0065] S100. Collect the assembly trajectory and assembly video of the target object and the human hand through the multi-modal data acquisition device 100, and associate the assembly trajectory and the assembly video;
[0066] S300. Based on the assembly action unit library, according to the assembly video and the assembly trajectory, identify and segment the assembly operation of the target object through a continuous action recognition and segmentation model of multi-modal fusion to obtain a series of continuous single action elements;
[0067] S400. Based on the motion planning strategy library, according to the type of each single action element, obtain its corresponding motion planning strategy to generate a robot assembly motion planning for the entire assembly operation of the target object.
[0068] Specific implementation manner of step S100
[0069] The data acquisition in the embodiments of the present invention is collected and processed by the multimodal data acquisition device 100 to obtain the assembly trajectory and assembly video of the target object, and associate the assembly trajectory and the assembly video. Specifically, the target object is placed on the artificial assembly action demonstration platform, the human-machine interface 112 is connected to the target object, and the capture cameras 111 and the video acquisition camera 120 are arranged around the learning platform. Among them, the six-dimensional pose data (i.e., the assembly trajectory) of the target object is obtained by arranging eight capture cameras 111 around the target object (at the four corners) and installing the first reflective marker point 115 on the target object. At the same time, the video data (i.e., the assembly video) of the assembly operations of the target object and the human hand is synchronously collected by the RGB camera (video acquisition camera), and the marker data of the assembly actions of the object to be assembled (target object) is generated by the data marking device 130. The data acquisition and processing steps include:
[0070] S110. Collect the assembly trajectory data, assembly video data and marker data of the target object during the assembly process through the multimodal data acquisition device 100. The data acquisition includes the following steps:
[0071] S111. Turn on the capture camera 111 and the video acquisition camera 120 to perform continuous data acquisition. Among them, the capture camera 111 is used to collect the assembly trajectory data and the marker data, and the video acquisition camera 120 is used to collect the assembly video data;
[0072] S112. At the moment when the assembly of the target object starts, obtain the action head marker through the active light-emitting marker point 131, the second reflective marker point 132 and the first reflective marker point 115. Specifically, when starting the assembly of the target object, turn on the active light-emitting marker point 131 through the marker point switch 140 to make it emit light. When the lighting time reaches about 0.2 seconds, turn off the active light-emitting marker point 131 through the marker point switch 140 to make it go out.
[0073] S113. During the assembly process of the target object, control the active light-emitting marker point 131 to emit light and go out to obtain the action middle segment marker. Specifically, the active light-emitting marker point 131 maintains a fixed frequency of emitting light and going out, and continuous data acquisition is performed by the capture camera 111 and the video acquisition camera 120.
[0074] S114. At the moment when the assembly of the target object ends, obtain the action end marker through the active light-emitting marker point 131, the second reflective marker point 132, and the first reflective marker point 115. Specifically, after the assembly action is completed, turn on the active light-emitting marker point 131, the second reflective marker point 132, and the first reflective marker point 115 through the marker point switch 140. When the lighting time reaches about 0.2 seconds, turn off the active light-emitting marker point 131, the second reflective marker point 132, and the first reflective marker point 115 through the marker point switch 140.
[0075] S120. According to the marker data, perform data processing on the assembly trajectory data and the assembly video data to align the assembly trajectory data and the assembly video data. The data processing operations include the following steps:
[0076] S121. According to the action middle segment marker, perform cropping processing on the assembly trajectory data to obtain multiple trajectory segments. Specifically, offline process the six-dimensional trajectory data (i.e., the assembly trajectory data) captured by the optical motion capture platform 110, and perform cropping trajectory processing on the assembly trajectory data of the entire assembly operation of the target object according to the action middle segment marker obtained in step S110, so as to obtain the desired multiple trajectory segments.
[0077] S122. According to the lighting and extinguishing of the active light-emitting marker point 131, perform frame-by-frame tagging processing on the assembly video data, and black out the area where the data marking device 130 is located in each frame of the video. Specifically, offline process the assembly video data collected by the video acquisition camera 120. First, add tags to the assembly video data frame by frame according to the brightness and darkness of the active light-emitting marker point 121 (i.e., the lighting and extinguishing of the active light-emitting marker point 131), and then black out the area where the data marking device 130 is located in each frame of the assembly video data.
[0078] S123. Perform trajectory compensation on the assembly trajectory data, and align the assembly trajectory data and the assembly video data according to the action start marker and the action end marker. Specifically, according to the action start marker, align the assembly trajectory data and the assembly video data at the moment when the assembly of the target object starts, and at the same time, according to the action end marker, align the assembly trajectory data and the assembly video data at the moment when the assembly of the target object is completed. Further, since the acquisition frequencies of the assembly trajectory data captured by the optical motion capture platform 110 and the assembly video data collected by the video acquisition camera 120 cannot be exactly the same or in multiples, after aligning the video frames at the start and end moments of the assembly with the assembly trajectory, it cannot be guaranteed that each frame of the assembly video data can correspond to the original motion capture trajectory data. To address the above problem, the method of this embodiment of the present invention uses a trajectory compensation strategy based on Gaussian process regression to compensate the assembly trajectory data.
[0079] Specific implementation manner of step S200
[0080] The continuous action recognition and segmentation model in the method of the present invention integrates assembly video data and accurate six-dimensional assembly trajectory data, and designs an effective tooling assembly task decomposition by designing an assembly action unit library. The continuous assembly action is semantically segmented by identifying the continuous action elements of different assembly actions. The multi-modal fusion continuous action recognition and segmentation model includes an action recognition model and a continuous action segmentation model based on a 3D convolutional neural network model. The recognition and segmentation of the assembly action include the following steps:
[0081] S210. Based on the assembly action unit library, the trajectory recognition module of the action recognition model based on the LSTM network model extracts features from the assembly trajectory to achieve trajectory action recognition. The video action recognition module of the action recognition model based on the C3D network model performs video action recognition on the assembly video. The trajectory action recognition result and the video action recognition result are fused to obtain the single action recognition result of the target object.
[0082] S220. Based on the assembly action unit library, according to the single action recognition result, as well as according to the assembly trajectory and the assembly video, the continuous assembly operation of the target object is segmented by the continuous action segmentation model to obtain a series of continuous single action elements (see Figure 7 ).
[0083] In some specific embodiments of the present invention, the types of single action elements included in the assembly action unit library include a moving action element, an aligning action element, a plugging action element, a fastening action element, and a tearing action element. Specifically, in the setting of the assembly action unit library, the entire process of the assembly operation is first divided into four states, namely, the starting state, the waiting-to-align state, the waiting-to-assemble state, and the completed state. The entire assembly process of the target item is segmented by the starting state, the waiting-to-align state, the waiting-to-assemble state, and the completed state, and the assembly actions are divided into different single action elements such as "moving action element, aligning action element, plugging action element, fastening action element, and tearing action element".
[0084] Among them, the action from the starting state to the waiting-to-align state is named the "moving action element", which is specifically defined as the action of moving two objects to be assembled from the completely separated starting state to the waiting-to-align state in free space. Next, the action from the waiting-to-align state to the waiting-to-assemble state is defined as the "aligning action element", and the "aligning action element" is specifically defined as the action of finely adjusting the relative pose between the parts to be assembled so that the connection and mating parts of the two parts to be assembled are in a relatively parallel and other assembly-friendly relative poses. Finally, the actions from the waiting-to-assemble state to the completed state are respectively defined as the "plugging action element", the "fastening action element", and the "tearing action element".
[0085] Further, the "plugin action element" can be further divided into "translation plugin action element", "rotation plugin action element" and "helical plugin action element". Among them, the "translation plugin action element" is specifically defined as the component in the to-be-assembled state being translated along the axis into a hole, key or groove to a certain depth. The "rotation plugin action element" is specifically defined as the component in the to-be-assembled state being rotated around the axis into a hole, key or groove to a certain depth. The "helical plugin action element" is specifically defined as the component in the to-be-assembled state being rotated around the central axis and translated along the central axis into a hole, key or groove to a certain depth. Further, the "fastening action element" can be further divided into "pressing action element" and "flat pressing action element". The "pressing action element" is specifically defined as applying an axial force perpendicular to the contact surface to make the to-be-assembled object move along the pressure direction, which includes two types: pressing for buckling and pressing for fitting. The "flat pressing action element" is specifically defined as applying a force diagonally downward perpendicular to the contact surface, and the applicator moves along the plane (perpendicular to the pressure direction). Further, the action of the "tearing action element" in the to-be-assembled state is specifically defined as the finger being aligned with the protruding part of the film, and in the completed state, the finger is completely separated from the attachment of the film.
[0086] In some specific embodiments of the present invention, the action recognition model includes a trajectory recognition module based on the LSTM network model for performing trajectory action recognition on the assembly trajectory and a video action recognition module based on the C3D network model for performing video action recognition on the assembly video. For the network model of the action recognition model, see Figure 5 . Among them, the trajectory action recognition includes a first feature extraction sub-module for extracting the assembly trajectory features and a first LSTM network for processing the trajectory features. The video action recognition module includes a plurality of single-action video processing sub-modules for processing the assembly video and a first fully connected layer for processing the outputs of the plurality of single-action video processing sub-modules. Among them, the single-action video processing sub-module is composed of a first 3D convolutional layer and a first 3D pooling layer, and the zero-padding in the time dimension of the first 3D convolutional layer is replaced with partial information in the previous time step. Two second fully connected layers fuse the output of the first LSTM network (trajectory action recognition result) and the output of the first fully connected layer (video action recognition result), and then enter the softmax layer.
[0087] Further, see Figure 8 . The action recognition module is trained by a single-action data set. Perform a single assembly operation of a single action element on the manual assembly action demonstration platform, and collect the operation trajectory data and operation video data of the single operation process through the multi-modal data acquisition module. The obtained data is collected into the single-action data set, and the action recognition module is trained for action recognition through the single-action data set.
[0088] In the specific embodiments of the present invention, for the trajectory recognition module, the original data of the obtained assembly trajectory includes the three-dimensional spatial position coordinates p and the unit quaternion q: , the trajectory features extracted by preprocessing include speed, acceleration, angular velocity, angular acceleration, curvature, and torsion. Among them, the calculation of speed, acceleration, angular velocity, and angular acceleration adopts the difference method, while the curvature and the calculation method of torsion τ are as follows:
[0089] ,
[0090] ,
[0091] In the formula, the dot above the vector represents the derivative with respect to time. And the vector cross product operation rule is as follows:
[0092] ,
[0093] After feature extraction, the trajectory recognition module uses an LSTM network. The three-dimensional spatial position coordinates, unit quaternion, speed, acceleration, angular velocity, angular acceleration, curvature, and torsion are used as the inputs of the trajectory recognition module for pre-training of the trajectory module, so that it can recognize single assembly actions.
[0094] In the specific embodiment of the present invention, for the video action recognition module, a publicly available video dataset is used to pre-train the video action recognition module in the network, and then the video data in the constructed single-action dataset is used to train the pre-trained video action recognition module. It should be noted that the video action recognition module can be based on the C3D network or replaced with other video action recognition networks.
[0095] In some specific embodiments of the present invention, the network model of the continuous action segmentation module is shown in Figure 6 , and its network architecture includes: a second feature extraction sub-module for extracting assembly trajectory features; a second LSTM network for processing trajectory features; multiple continuous action video processing sub-modules for processing assembly videos, where the continuous action video processing sub-module consists of a second 3D convolutional layer and a third 3D convolutional layer, the third 3D convolutional layer is the bias term of the second 3D convolutional layer, and a ReLU activation function is set after the third 3D convolutional layer; a third fully connected layer for processing the outputs of multiple continuous action video processing sub-modules; two fourth fully connected layers for fusing the output of the second LSTM network and the output of the third fully connected layer; a softmax layer is arranged behind the two fourth fully connected layers.
[0096] In an application scenario, the data size of the input 3D convolutional layer is set to , the number of channels is , the 3D convolutional kernel size is , then the output data size after convolution operation is:
[0097] ,
[0098] In the neural network of the action recognition model in step S210 (see the convolutional layer in Figure 5 ), the size of the 3D convolutional kernel in the video stream is , and the 3D convolutional layer is zero-padded. Remove the zero-padding in the time dimension of the 3D convolutional layer and replace it with partial information from the same layer in the previous time step. In this case, the number of removed and added axes is 2, which does not affect the network size and facilitates migrating the training parameters of step S210 to the continuous action segmentation module.
[0099] Furthermore, optimize the recognition segmentation result through a weight function, and its calculation method is as follows:
[0100] ,
[0101] ,
[0102] where the input of the neural network model is a total of 16 frames of video data from the th frame to the th frame, represents the probability that the th frame in the model input data is recognized as action, represents the weight function, represents the optimal weight. And multiple weight functions are proposed as follows, where x is i - k, and the optimal weight function is selected through experiments.
[0103] ,
[0104] ,
[0105] ,
[0106] ,
[0107] In some specific embodiments of the present invention, see Figure 8 , the continuous action segmentation module is trained by a continuous action dataset. Perform continuous assembly operations on the entire assembly process of the target object on the artificial assembly action demonstration platform, and collect the operation trajectory data and operation video data of the continuous operation process through the multi-modal data acquisition module. Pool the obtained data into the continuous action dataset, and perform action segmentation training on the continuous action segmentation module through the continuous action dataset.
[0108] Specific implementation manner of step S300
[0109] Based on the assembly action unit library, different motion planning strategies are designed for different types of assembly actions (single action elements) according to their accuracy requirements, motion characteristics, etc., to form a motion planning strategy library. Then, according to the series of consecutive single action elements obtained in step 200, and based on the type to which each single action element belongs, the corresponding motion planning strategy of its type is found in the motion planning measurement library. Then, a series of consecutive single action elements can obtain a series of motion planning strategies. Connecting the obtained different motion planning strategies can complete the robot assembly motion planning for the entire assembly process of the target object.
[0110] It should be recognized that the method steps in the embodiments of the present invention can be implemented or executed by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.
[0111] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes multiple instructions executable by one or more processors.
[0112] Further, the method may be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention may be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RS1M, ROM, etc., such that it is readable by a programmable computer and, when read by the storage medium or device, can be used to configure and operate the computer to perform the processes described herein. Additionally, the machine-readable code, or portions thereof, may be transmitted via wired or wireless networks. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention may also include the computer itself.
[0113] A computer program can be applied to input data to perform the functions described herein, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0114] As described above, these are only the preferred embodiments of the present invention, and the present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners may have various different modifications and variations.
Claims
1. A multimodal demonstration method applied to robot assembly skill learning, characterized in that The method includes the following steps: S100. Collect the assembly trajectory and assembly video of the target object and the human hand through the multimodal data acquisition device (100), and associate the assembly trajectory and the assembly video; S200. Based on the assembly action unit library, according to the assembly video and the assembly trajectory, identify and segment the assembly operation of the target object through the continuous action recognition and segmentation model of multimodal fusion to obtain a series of continuous single action elements; Based on the assembly action unit library, design different motion planning strategies for different types of assembly actions according to their accuracy requirements and motion characteristics to form a motion planning strategy library; S300. Based on the motion planning strategy library, obtain the corresponding motion planning strategy according to the type of each single action element to generate the robot assembly motion planning for the entire assembly operation of the target object; Among them, the step S200 includes: S210. Based on the assembly action unit library, the trajectory recognition module based on the LSTM network model of the action recognition module extracts features from the assembly trajectory to realize trajectory action recognition, and the video action recognition module based on the C3D network model of the action recognition module performs video action recognition on the assembly video, and fuses the trajectory action recognition result and the video action recognition result to obtain the single action recognition result of the target object; S220. Based on the assembly action unit library, according to the single action recognition result, and according to the assembly trajectory and the assembly video, segment the continuous assembly operation of the target object through the continuous action segmentation module to obtain a series of continuous single action elements; Among them, the types of single action elements included in the assembly action unit library include carrying action elements, alignment action elements, plug-in action elements, fastening action elements, and tearing action elements.
2. The method according to claim 1, wherein, In the step S200: The continuous action recognition and segmentation model includes an action recognition module based on a 3D convolutional neural network model and a continuous action segmentation module; among them, the action recognition module is trained by a single action data set, and the continuous action segmentation module is trained by a continuous action data set.
3. The method according to claim 1, characterized in that, The network model of the action recognition module includes: A first feature extraction sub-module for extracting assembly trajectory features; a first LSTM network for processing trajectory features; Multiple single action video processing sub-modules for processing the assembly video, where the single action video processing sub-module consists of a first 3D convolutional layer and a first 3D pooling layer, and the zero-padding in the time dimension of the first 3D convolutional layer is replaced by partial information in the previous time step; A first fully connected layer for processing the outputs of the multiple single action video processing sub-modules; Two second fully connected layers for fusing the output of the first LSTM network and the output of the first fully connected layer; A softmax layer is arranged behind the two second fully connected layers.
4. The method according to claim 1, characterized in that, The network model of the continuous action segmentation module includes: A second feature extraction sub-module for extracting assembly trajectory features; a second LSTM network for processing trajectory features; Multiple consecutive-action video processing sub-modules for processing assembly videos, wherein each of the consecutive-action video processing sub-modules consists of a second 3D convolutional layer and a third 3D convolutional layer, the third 3D convolutional layer is the bias term of the second 3D convolutional layer, and a ReLU activation function is provided after the third 3D convolutional layer; A third fully-connected layer for processing the outputs of the multiple consecutive-action video processing sub-modules; Two fourth fully-connected layers for fusing the output of the second LSTM network and the output of the third fully-connected layer; A softmax layer is provided after the two fourth fully-connected layers.
5. The method according to claim 1, characterized in that For the trajectory recognition module: The data of the assembly trajectory includes the three-dimensional spatial position coordinates p and the unit quaternion q: , and the extracted trajectory features include velocity, acceleration, angular velocity, angular acceleration, curvature, and torsion; Among them, the curvature and the torsion τ are calculated as follows: ; ; In the formula, the dot above the vector represents the derivative with respect to time.
6. The method according to claim 2, wherein The training of the neural network model includes the following steps: Optimize the recognition segmentation result through a weight function, and its calculation method is as follows: ; ; Among them, the input of the neural network model is from frame to frame, a total of 16 frames of video data. Indicates the probability that the th frame in the model input data is recognized as action. Indicates the weight function. Indicates the optimal weight.
7. The method according to claim 1, wherein The training of the network model of the consecutive-action recognition segmentation model includes the following steps: Obtain the assembly trajectories and assembly videos of the single assembly operations of the single action elements, and gather them into a single-action dataset; Obtain the assembly trajectories and assembly videos of the continuous entire assembly operations of the target object, and gather them into a consecutive-action dataset.
8. A computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
9. A multimodal demonstration system applied to robot assembly skill learning, characterized in that, Comprising: A computer device, which contains the computer-readable storage medium according to claim 8.
Citation Information
Patent Citations
Robot assembly trajectory optimization method and device for offline example learning
CN110561430A
Intelligent body element action learning method based on joint grouping strategy
CN114170454A