Assembly robot component assembly control methods, electronic equipment and storage media
By learning new component assembly tasks from designated training videos and generating assembly strategies using pre-trained assembly models, the problems of cumbersome operation and poor generalization ability of assembly robots are solved. This enables rapid adaptation to the assembly of new components and improves the robot's operational flexibility and environmental adaptability.
Patent Information
- Application Number
- CN202511280918.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing technologies for assembling robots involve cumbersome operation and poor generalization ability. Detailed assembly programs need to be written for each new model or specification of server component, which is time-consuming, labor-intensive, and difficult to learn from and apply from existing experience.
By acquiring a specific training video of a task to assemble a specified component, analyzing the specified reference trajectory information, and inputting it into a pre-trained assembly model, a predicted sequence of assembly actions is generated to control the robot to assemble the component, thus avoiding the need to write complex control programs for each new component.
It significantly reduces operational complexity and preparation time, improves the robot's adaptability and operational flexibility in complex and ever-changing production environments, and enables the rapid generation of reasonable assembly strategies.
Smart Images

Figure CN120755892B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot manufacturing technology, and in particular to component assembly control methods, electronic devices, and storage media for assembly robots. Background Technology
[0002] In the field of server manufacturing and maintenance, the application of robotics is gradually expanding, especially in component assembly. Related robotic component assembly methods primarily rely on precise pre-programmed trajectories and fixed rules. While this method offers high accuracy for specific tasks, it reveals significant limitations when dealing with diverse server components. For example, it suffers from cumbersome operation and poor generalization ability. For each new model or specification of server component, a detailed assembly program needs to be manually written. This process is not only time-consuming and labor-intensive, but also demands high skill from the operator, requiring extensive programming experience and a deep understanding of the mechanical system. Furthermore, pre-programmed robotic systems struggle to learn from existing experience and apply it to the assembly of new components. This means that reprogramming is required every time a new component is encountered, resulting in poor generalization ability.
[0003] Therefore, the component assembly control methods for assembly robots in related technologies suffer from technical problems such as cumbersome operation and poor generalization ability. Summary of the Invention
[0004] This application provides a component assembly control method, electronic equipment, and storage medium for assembly robots, in order to at least solve the problems of cumbersome operation and poor generalization ability of component assembly control methods for assembly robots in related technologies.
[0005] This application provides a component assembly control method for an assembly robot, including:
[0006] Obtain a specified component assembly task and a specified training video corresponding to the specified component assembly task. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not performed before. The specified training video is used to record the assembly scenario of the specified component assembly task.
[0007] The specified training video is analyzed to obtain specified reference trajectory information, wherein the specified reference trajectory information is used to describe the reference assembly trajectory corresponding to the specified component assembly task;
[0008] The specified reference trajectory information and the state information of the assembly robot are input into the pre-trained assembly model to obtain the predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot.
[0009] The assembly robot is controlled to assemble the specified component according to the predicted assembly action sequence, thereby performing the specified component assembly task.
[0010] This application also provides a component assembly control device for an assembly robot, comprising:
[0011] The acquisition module is used to acquire a specified component assembly task and a specified training video corresponding to the specified component assembly task. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not performed before. The specified training video is used to record the assembly scenario of the specified component assembly task.
[0012] The parsing module is used to parse the specified training video to obtain specified reference trajectory information, wherein the specified reference trajectory information is used to describe the reference assembly trajectory corresponding to the specified component assembly task;
[0013] An execution module is used to input the specified reference trajectory information and the state information of the assembly robot into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot;
[0014] An assembly module is used to control the assembly robot to assemble the specified component according to the predicted assembly actions in the predicted assembly action sequence, so as to perform the assembly task of the specified component.
[0015] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the component assembly control method of any of the above-described assembly robots.
[0016] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the component assembly control method of any of the above-described assembly robots.
[0017] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the component assembly control method for any of the above-described assembly robots.
[0018] This application obtains a specified component assembly task and a corresponding specified training video for the specified component assembly task. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not previously executed. The specified training video records the assembly scenario of the specified component assembly task. The specified training video is parsed to obtain specified reference trajectory information, which describes the reference assembly trajectory corresponding to the specified component assembly task. The specified reference trajectory information and the assembly robot's state information are input into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model. The assembly robot's state information indicates the state of the end effector of the assembly robot's robotic arm. Following the predicted assembly actions in the predicted assembly action sequence, the assembly robot is controlled to assemble the specified component to execute the specified component assembly task. Using this application, by learning new component assembly tasks from specified training videos, it is no longer necessary to write complex control programs for each new component, significantly reducing operational complexity and preparation time. By providing training videos and assembly models, the robot can quickly generate reasonable assembly strategies when faced with unseen parts, without the need for additional training. This greatly improves the robot's adaptability in complex and ever-changing production environments, thereby solving the technical problems of cumbersome operation and poor generalization ability in the component assembly control methods of assembly robots in related technologies. It also improves the operational flexibility and environmental adaptability of assembly robots. Attached Figure Description
[0019] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram illustrating an application scenario of a component assembly control method for an assembly robot according to an embodiment of this application.
[0021] Figure 2 This is a schematic flowchart of an optional component assembly control method for an assembly robot according to an embodiment of this application.
[0022] Figure 3 This is a schematic diagram of an optional task encoder according to an embodiment of this application.
[0023] Figure 4 This is a schematic diagram of an optional component assembly control method for an assembly robot according to an embodiment of this application.
[0024] Figure 5 This is a structural block diagram of an optional component assembly control device for an assembly robot according to an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0026] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0027] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] According to one aspect of the embodiments of this application, a component assembly control method for an assembly robot is provided. Optionally, in this embodiment, the above-described component assembly control method for an assembly robot may be applied, but is not limited to, to applications such as... Figure 1 The hardware environment shown includes terminal device 102 and server 104. Server 104 can be connected to terminal device 102 via a network and can be used to provide services (e.g., application services, etc.) to terminal device 102 or clients installed on terminal device 102. A database can be set up on server 104 or independently of server 104 to provide data storage services for server 104.
[0029] The aforementioned network may include, but is not limited to, at least one of the following: wired network and wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network (WAN), metropolitan area network (MAN), and local area network (LAN). The aforementioned wireless network may include, but is not limited to, at least one of the following: Wireless Fidelity (WIFI) and Bluetooth. Terminal device 102 may be, but is not limited to, a personal computer (PC), mobile phone, tablet computer, etc. Server 104 may be, but is not limited to, a cloud server, server cluster, or other server types.
[0030] The component assembly control method of the assembly robot in this application embodiment can be executed by server 104, terminal device 102, or jointly by server 104 and terminal device 102. Alternatively, the component assembly control method of the assembly robot in this application embodiment can be executed by a client installed on the terminal device 102.
[0031] Taking the component assembly control method of the assembly robot in this embodiment, executed by the terminal device 102, as an example, Figure 2 This is a flowchart illustrating an optional component assembly control method for an assembly robot according to an embodiment of this application, as shown below. Figure 2 As shown, the process of this method may include the following steps:
[0032] Step S202: Obtain the specified component assembly task and the specified training video corresponding to the specified component assembly task. The specified component assembly task is an assembly task performed on the specified component that the assembly robot has not performed before. The specified training video is used to record the assembly scenario of the specified component assembly task.
[0033] Step S204: Analyze the specified exercise video to obtain specified reference trajectory information, wherein the specified reference trajectory information is used to describe the reference assembly trajectory corresponding to the specified component assembly task;
[0034] Step S206: Input the specified reference trajectory information and the state information of the assembly robot into the pre-trained assembly model to obtain the predicted assembly action sequence output by the assembly model. The state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot.
[0035] Step S208: According to the predicted assembly action in the predicted assembly action sequence, control the assembly robot to assemble the specified parts to perform the specified parts assembly task.
[0036] The component assembly control method of the assembly robot in this embodiment can be applied to the field of robot manufacturing technology, specifically to server production and maintenance scenarios. The specific scenario is an automated server assembly line. On such a production line, the assembly robot is responsible for the precise and safe assembly of various server components, including but not limited to CPUs, memory modules, hard drives, and heat sinks. Automated assembly lines aim to improve production efficiency, reduce costs, and ensure the quality and reliability of server assembly. In related technologies, the component assembly control method of the assembly robot typically relies on precise programming and preset assembly paths; that is, engineers write detailed assembly steps and robotic arm movement trajectories for each server component. While this method can automate assembly tasks to a certain extent, it suffers from the following problems: cumbersome operation and poor generalization ability. For each new component or assembly task, complex reprogramming is required, which is time-consuming and prone to errors. As server technology iterates and the types and specifications of components increase, the efficiency and flexibility of this programming method are greatly challenged, resulting in cumbersome operation. Furthermore, it is difficult to quickly learn from existing assembly experience and apply it to new components or assembly tasks, limiting the assembly robot's adaptability to diverse production needs. Therefore, the component assembly control methods for assembly robots in related technologies suffer from technical problems such as cumbersome operation and poor generalization ability.
[0037] To at least partially address the aforementioned technical problems, this embodiment learns new component assembly tasks from designated training videos, eliminating the need to write complex control programs for each new component, significantly reducing operational complexity and preparation time. By using designated training videos and assembly models, a reasonable assembly strategy can be quickly generated when faced with unseen components, requiring no additional training. This greatly improves the robot's adaptability in complex and ever-changing production environments, thus solving the technical problems of cumbersome operation and poor generalization ability in component assembly control methods for assembly robots in related technologies, and improving the operational flexibility and environmental adaptability of assembly robots.
[0038] It should be noted that the specified component assembly task can refer to a new type of component assembly task that the assembly robot has never performed before, such as the installation of a new type of hard drive bracket encountered for the first time in the server assembly process.
[0039] Optionally, upon receiving a component assembly instruction, it can be identified whether the instruction indicates the execution of a new component assembly task, i.e., a specified component assembly task. If a specified component assembly task is identified, a corresponding training video can be retrieved from a preset database. The specified training video can be used to record the assembly scenario of the specified component assembly task, mainly recording the process of an expert completing the specified component assembly task, and can be regarded as a "teaching material" for the assembly robot to learn.
[0040] It should be noted that the specified reference trajectory information can be used to describe the reference assembly trajectory corresponding to the assembly task of a specified component. The reference assembly trajectory can be the reference trajectory of the assembly task corresponding to the reference training video generated based on the reference training video, and can include a series of state-action pairs, which can constitute the ideal path for experts to complete the assembly task.
[0041] Optionally, parsing the specified training video to obtain the specified reference trajectory information may include: decomposing the specified training video to obtain a series of video frames corresponding to the specified training video, identifying a series of key video frames from the series of video frames, and identifying the position, attitude, speed of the robotic arm of the assembly robot and the relevant state information of the parts to be assembled based on the series of key video frames, thereby obtaining the specified reference trajectory information.
[0042] Optionally, the reference trajectory information can be specified by using a preset encoder to convert the video frames corresponding to the specified training video into a low-dimensional embedding vector, which can represent the expert's assembly strategy.
[0043] It should be noted that the pre-trained assembly model can be a deep learning model that has been trained on a large dataset. It can learn specified reference trajectory information and robot state information, and predict a suitable sequence of assembly actions to guide the robot in completing the assembly task. By inputting the specified reference trajectory information and the robot's state information into the pre-trained assembly model, the model outputs a predicted sequence of assembly actions. The robot's state information can be used to indicate the state of the end effector of the robot's robotic arm. The predicted sequence of assembly actions can include a series of action commands with a specified execution order, used to instruct and guide the robot to assemble the robot and complete the assembly task. Optionally, the predicted sequence of assembly actions generated by the assembly model can include a series of robotic arm movements for actually assembling the specified components.
[0044] Optionally, the assembly robot controls the end effector of its robotic arm to perform precise assembly operations on designated parts by predicting the assembly action sequence. Optionally, the robot executes each action in the sequence one by one until the assembly of the entire part is completed.
[0045] This application provides an embodiment that obtains a specified component assembly task and a corresponding practice video. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not previously executed. The practice video records the assembly scenario of the specified component assembly task. The practice video is parsed to obtain specified reference trajectory information, which describes the reference assembly trajectory corresponding to the specified component assembly task. The specified reference trajectory information and the assembly robot's state information are input into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model. The assembly robot's state information indicates the state of the end effector of the robot's robotic arm. Following the predicted assembly actions in the predicted assembly action sequence, the assembly robot is controlled to assemble the specified component to execute the specified component assembly task. By learning new component assembly tasks from specified practice videos, this application eliminates the need to write complex control programs for each new component, significantly reducing operational complexity and preparation time. By providing training videos and assembly models, the robot can quickly generate reasonable assembly strategies when faced with unseen parts, without the need for additional training. This greatly improves the robot's adaptability in complex and ever-changing production environments, thereby solving the technical problems of cumbersome operation and poor generalization ability in the component assembly control methods of assembly robots in related technologies. It also improves the operational flexibility and environmental adaptability of assembly robots.
[0046] In one exemplary embodiment, parsing a specified training video to obtain specified reference trajectory information includes: performing target detection on the specified training video to extract at least two training images from the specified training video; and inputting the at least two training images into a pre-trained task encoder to obtain the specified reference trajectory information output by the task encoder.
[0047] It should be noted that the designated training videos refer to video materials demonstrating the assembly of server components by experienced operators (or experts). These videos contain specific actions of how the operator precisely aligns, grasps, and installs the components. Object detection can refer to a computer vision technique that can be used to identify and locate objects in images. In this embodiment, object detection is used to extract image frames (training images) containing key assembly actions from the training videos, such as the moment when the robotic arm end effector contacts and aligns with the server component.
[0048] By performing object detection on a specified training video, training images corresponding to assembly operations in a specified component assembly task are extracted from the video. At least two training images can be used to cover multiple key stages of the corresponding assembly operations in the component assembly task.
[0049] A pre-trained task encoder can be a deep learning model that has been trained on a large dataset and can transform an input image into a set of compact digital representations (i.e., task embedding vectors) that captures the key properties of the assembly task in the video.
[0050] Reference trajectory information can be obtained by analyzing training images using a pre-trained task encoder, and includes a quantitative description of expert operations during the assembly process. Reference trajectory information may include the robotic arm's position coordinates, attitude angles, movement speed, and the relative positions and angles of components.
[0051] In one example, object detection technology is used to locate the moment of the specified component assembly task in a specified training video to obtain at least two training images. At least two image frames are input into a pre-trained task encoder. The task encoder uses its internal convolutional neural network (CNN) and other layer structures to extract and compress features in the images and output a low-dimensional vector, i.e., the specified reference trajectory information.
[0052] Optionally, for each component assembly task, there can be one or more reference training videos consisting of different server component assembly scenarios, such as CPU installation, memory module insertion / removal, and hard drive bracket installation. Preprocessing of the reference training videos yields key state-action pairs, which may include the position, orientation, and speed of the robot's end effector, as well as the position and orientation of components (such as the component to be installed).
[0053] Through this embodiment, by accurately extracting specified reference trajectory information from a specified training video, assembly skills can be learned and imitated more accurately, thereby improving the assembly accuracy of the robotic arm of the assembly robot.
[0054] In one exemplary embodiment, target detection is performed on a specified training video to extract at least two training images from the specified training video, including: using a target detection model to sequentially perform target detection on video frames in the specified training video to obtain target detection results for the video frames in the specified training video; based on the target detection results for the video frames in the specified training video, identifying video frames containing target components, wherein at least two training images include video frames containing target components, the target components are related components for performing a specified component assembly task, and the target components include the specified components.
[0055] It should be noted that the object detection model can be a machine learning model, such as YOLOv4, which can be used to identify and locate specific objects in an image. In this embodiment, the object detection model can be an end effector of a robotic arm trained to recognize robot components and server parts.
[0056] Optionally, a set of video frames in a specified training video is identified, and a target detection model is used to sequentially detect targets in the set of video frames according to their order, obtaining the detection result for each video frame. Optionally, through frame-by-frame analysis, the specified training video can be decomposed into a series of static images, i.e., a set of video frames. A target detection model can be used to identify each frame to recognize the target components and the state of the robotic arm involved in the assembly process.
[0057] In one example, suppose there is a demonstration video showing the hard drive installation process. The object detection model would search for the hard drive and the robotic arm's end effector in each frame of the demonstration video. For instance, in the first few seconds of the video, the model might detect the robotic arm approaching the hard drive, and in the next frame, the hard drive would be clamped by the robotic arm and ready to be inserted into the tray.
[0058] Based on the target detection results of video frames in the reference training video, video frames containing target components, i.e., keyframes, are identified. These are image frames in the demonstration video that contain key assembly actions, such as the moment the robotic arm aligns with the hard drive slot. Optionally, the target detection results can be used to filter out image frames in which the robotic arm's end effector is interacting with the target component. Selecting keyframes from numerous frames allows for focusing on images that truly contain details of assembly skills, thus providing more accurate data for subsequent processing.
[0059] At least two training images include video frames containing target components, which are components related to performing the specified component assembly task. The target components include the specified components. One of the at least two training images is an image that identifies a specific assembly action. The training images may contain state-action information during the assembly process.
[0060] In this embodiment, a target detection model is used to identify video frames containing target components, which accurately identifies key information of the assembly process in a specified training video. This not only speeds up the model's learning speed but also greatly enhances the model's generalization ability and robustness, enabling the assembly robot to learn and perform complex server component assembly tasks with only a few demonstrations.
[0061] In an exemplary embodiment, after performing object detection on a specified training video to extract at least two training images from the specified training video, the method further includes: determining image distribution parameters of the at least two training images in the specified training video, wherein the video time of the specified training video is divided into multiple time periods, and the image distribution parameters are used to describe the distribution of the at least two training images in different time periods of the multiple time periods; if the image distribution parameters do not meet the specified distribution conditions, performing an image addition operation or an image subtraction operation on the at least two training images to update the at least two training images.
[0062] It should be noted that image distribution parameters can refer to the distribution characteristics of selected training images in a video at different time periods, including but not limited to frequency, density, and time intervals.
[0063] Optionally, after performing object detection and extracting at least two key training images, the distribution parameters of these at least two training images on the video timeline of a specified training video can be determined by analyzing the time of the frames in which they appear. This can help in understanding the rhythm of the demonstration and the time sequence of key actions. For example, taking hardware installation as an example, suppose the specified training video is 30 seconds long, and two key moments are identified. The first moment occurs at 10 seconds, when the robotic arm begins to grasp the hard drive; the second moment occurs at 20 seconds, when the hard drive is successfully inserted into the tray. Analyzing the time distribution of these two moments reveals that they occupy 1 / 3 and 2 / 3 of the video's time, respectively; these are the image distribution parameters.
[0064] Optionally, the specified distribution conditions may refer to the ideal temporal distribution standard of the training images in the video, based on domain knowledge or task requirements.
[0065] Optionally, the image distribution parameters are compared with preset specified distribution conditions, wherein the specified distribution conditions are determined based on the component assembly task. Optionally, the specified distribution conditions are determined according to the complexity of the assembly operations corresponding to the component assembly task. The number of training images in each time period under the specified distribution conditions is positively correlated with the number of assembly operations in the component assembly task.
[0066] If the image distribution parameters do not meet the specified distribution conditions, an image adjustment operation is performed, which may include: if a first specified time period exists, randomly extracting a specified number of training images from the first specified time period in the specified training video, or processing at least two training images to obtain new images, and supplementing the images to the first specified time period; the first specified time period is the time period in which the number of training images in multiple time periods is less than a first threshold.
[0067] In the presence of a second specified time period, the pre-trained image filtering model automatically filters and deletes duplicate or redundant training images from the second specified time period, or calculates the similarity between at least two training images to reduce training images with the highest similarity or a similarity greater than a specified similarity threshold. The second specified time period is the time period in which the number of training images in multiple time periods is greater than the first threshold.
[0068] Through the above embodiments, by finely adjusting the key images in the training video, sufficient learning samples are ensured for each assembly stage, while avoiding excessive data redundancy, thereby improving the learning efficiency and generalization ability of the assembly model.
[0069] Optionally, at least two training images have a certain temporal sequence, which enables the assembly model to better understand the temporal sequence of the assembly task and improve the coherence and success rate of the assembly.
[0070] In one embodiment, time series analysis techniques, such as Long Short-Term Memory (LSTM) networks, can be introduced to further optimize the model's ability to process time-series information and address the problem of strong time dependence in assembly tasks.
[0071] This embodiment optimizes image distribution by adding or removing images, enabling the model to learn more efficiently from key assembly actions, reducing redundant learning time, and focusing on learning key skill points, thus improving model learning efficiency. The image distribution parameters ensure the model receives practice images covering all stages of the assembly process, helping the model understand the transitions between different assembly stages and improving its adaptability to unfamiliar parts or assembly scenarios.
[0072] In one exemplary embodiment, the task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. Inputting at least two training images into the pre-trained task encoder to obtain specified reference trajectory information output by the task encoder includes: performing a convolution operation on the at least two training images through at least one convolutional layer to obtain at least two training images after the convolution operation; performing pooling processing on the at least two training images after the convolution operation through at least one pooling layer to obtain at least two training images after pooling processing; and performing integration processing on the at least two training images after pooling processing through at least one fully connected layer to obtain the specified reference trajectory information output by the task encoder.
[0073] It should be noted that the task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. The convolutional layer within the at least one convolutional layer can automatically detect and extract key features about the memory module, the robotic arm end effector, and their interaction by sliding the convolutional kernel across the input image (at least two training images). For example, the convolutional layer can identify edge features where the robotic arm end effector aligns with the memory slot, as well as texture changes as the memory module gradually penetrates deeper into the slot.
[0074] The feature maps output by convolutional layers may contain a large number of data points. These feature maps are reduced in dimensionality by pooling layers (such as max pooling or average pooling), which retain the most significant visual features, remove redundant information, and improve the computational efficiency of the algorithm.
[0075] At least one fully connected layer can be used to linearly combine the feature vectors output from the previous layer as inputs to the weights between nodes, outputting a fixed-size vector for classification or regression tasks; here, it's used to generate reference trajectory information. The pooled features are fed into the fully connected layer for intensive neural network computation, transforming the image features into a compact set of numerical representations—the reference trajectory information. This information contains core operational details, such as the initial position of the memory module, the robotic arm's trajectory, and the final insertion posture, guiding the model on how to perform similar assembly actions.
[0076] The convolutional layer is responsible for extracting local features from at least two training images, the pooling layer is used for feature dimensionality reduction and anti-interference processing, and the fully connected layer integrates the information to form compact specified reference trajectory information, which can be a low-dimensional embedding vector.
[0077] In one example, the task encoder can be as follows: Figure 3 As shown, the task encoder can include convolutional layers, pooling layers, and fully connected layers. The input image (which can be video frames from a specified training video, such as at least two training images) is fed into the convolutional layer. Multiple kernels in the convolutional layer perform convolutions to generate corresponding feature maps. The pooling layer performs max pooling or average pooling (subsampling) to obtain the input image after convolution and pooling operations. The fully connected layers then integrate the input image after convolution and pooling operations to output a low-dimensional embedding vector, which represents the specified reference trajectory information.
[0078] In this embodiment, a task encoder comprising convolutional layers, pooling layers, and fully connected layers is used to efficiently and accurately extract key features of the assembly task from training images and transform them into compact reference trajectory information.
[0079] In an exemplary embodiment, the method further includes: constructing a first training set based on a component assembly trajectory library, wherein the component assembly trajectory library is used to record the following trajectories corresponding to the component assembly task: a reference assembly trajectory and a trial assembly trajectory, wherein the trial assembly trajectory is an assembly trajectory obtained by assembling components in a real environment using a pre-trained assembly model, and a first sample in the first training set includes specified trajectory information, label information, and reference trajectory information, wherein in a first sample in the first training set, the label information is used to indicate whether the specified trajectory corresponding to the specified trajectory information is a reference assembly trajectory or a trial assembly trajectory, and the reference trajectory information is generated based on a reference exercise video corresponding to the component assembly task; and training a task evaluation network and a pre-trained assembly model according to the first training set until the task evaluation network and the assembly model converge, wherein the task evaluation network is an evaluation network that evaluates the trial assembly trajectory using the reference assembly trajectory as a reference value.
[0080] It should be noted that the component assembly trajectory library is a database that stores reference assembly trajectories and experimental assembly trajectories for different component assembly tasks. The reference assembly trajectories are derived from expert demonstration videos, while the experimental assembly trajectories are trajectory records generated by the initially trained assembly model when attempting to assemble components in a real environment.
[0081] Optionally, a first training set covering reference assembly trajectories and experimental assembly trajectories is selected from the component assembly trajectory library, ensuring that each trajectory is equipped with a set of label information to indicate the source and type of the trajectory.
[0082] Optionally, each first sample in the first training set may contain specified trajectory information, label information, and reference trajectory information, which constitute the data points required for training the network and assembling the model.
[0083] Optionally, the task evaluation network can learn how to assess the quality of assembly operations by comparing experimental assembly trajectories with reference assembly trajectories. The assembly model can continuously adjust its strategy by receiving feedback from the task evaluation network, making the experimental assembly trajectory increasingly closer to the reference assembly trajectory. After constructing the first training set, the task evaluation network and the assembly model can be sequentially trained so that the task evaluation network can accurately distinguish and evaluate different types of trajectories, while the assembly model can optimize and adjust itself based on this evaluation feedback, thereby gradually improving its actual assembly capabilities.
[0084] In one example, such as a memory module insertion task, the task evaluation network determines whether the trajectory is similar to a reference assembly trajectory by identifying the positional accuracy and speed control of the robotic arm's end effector. The assembly model then adjusts its algorithm parameters based on the task evaluation network's score to reduce insertion deviations and improve insertion speed and stability.
[0085] In one example, m assembly tasks can be randomly selected from the training tasks. Using the current assembly model and task encoder, at least one trajectory (i.e., a state-action pair) is generated for each task in the real world, and the trajectory is stored in the buffer corresponding to the assembly task. That is, for each task, the assembly model is used to obtain a predicted action sequence, and the predicted action sequence is used to generate a trajectory for recording.
[0086] In this embodiment, by constructing a first training set and fusing reference assembly trajectories and experimental assembly trajectories, the model can learn from a wider range of data, improving its adaptability to new tasks and environments. Furthermore, by introducing a task evaluation network, quantitative evaluation criteria are provided for the experimental assembly trajectories, enabling accurate identification and quantification of the success rate and quality of assembly operations, thus accelerating the model's iteration and optimization process.
[0087] In one exemplary embodiment, training a task evaluation network and a pre-trained assembly model based on a first training set includes: using a first sample in the first training set as the current first sample, and performing the following training operations on the task evaluation network and the pre-trained assembly model: inputting specified trajectory information and reference trajectory information in the current first sample into the task evaluation network, and outputting an evaluation value corresponding to the current first sample; adjusting the network parameters of the task evaluation network based on the difference between the evaluation value corresponding to the current first sample and the reference evaluation value indicated by the label information of the current first sample, to train the task evaluation network; and training the pre-trained assembly model based on the parameter-adjusted task evaluation network.
[0088] It should be noted that the specified trajectory information refers to a specific assembly operation trajectory directly output from the training video or assembly model, while the reference trajectory information is the optimal trajectory for assembly operations generated based on expert demonstrations. Specifically, the specified trajectory information (which can be a trial assembly trajectory or a reference assembly trajectory) and the reference trajectory information from the first sample are simultaneously fed into the task evaluation network. Through the network's calculation, an evaluation value is output, representing the degree of conformity between the trajectory and the optimal assembly trajectory. By inputting different types of trajectory information, the network can learn the ability to evaluate the quality of assembly operations, providing a basis for subsequent adjustments to network parameters.
[0089] The reference evaluation value, based on the label information of the first sample, can be a value given manually or according to pre-defined rules, and can be used to represent the ideal evaluation result. By calculating the difference between the evaluation value output by the task evaluation network and the reference evaluation value, and by adjusting the parameters of the task evaluation network through the backpropagation algorithm, the network's output is made closer to the reference evaluation value. Adjusting the parameters of the task evaluation network allows it to more accurately evaluate the quality of the assembly operation. This evaluation is then fed back to the assembly model to guide its training.
[0090] Optionally, the task evaluation network and the pre-trained assembly model can be trained by specifying a loss function. In one example, taking the IRL loss function as the specified loss function and the Q function as the task evaluation network, the calculation process of the difference between the evaluation value output by the corresponding task evaluation network and the reference evaluation value can be shown in the following formula (1):
[0091] (1)
[0092] in, Represents the loss function. Represents a state-action pair. Indicates use In the output of the task encoder, The corresponding classification results, i.e., label information, for example, In this case, the indicated state-action pair comes from the reference exercise video; in In this case, the indication state-action pair comes from the test assembly trajectory; This represents a Q-function whose inputs include state (s), action (a), and task embedding vector (z), and whose output is a probability value representing the probability that the state-action pair in a given task embedding vector comes from a reference exercise video (i.e., an expert). It refers to The expected loss function under a large number of samples, such as in and The expected loss function under these two strategies This represents the assembly model (such as a policy network). This refers to a strategy model derived from an expert, such as reference trajectory information in a reference exercise video, which can be understood as an expert strategy.
[0093] Specifically, the learning process of the Q-function is cleverly designed to resemble the principle of Generative Adversarial Networks (GANs). By introducing the concepts of discriminator and generator, the training of the Q-function is transformed into an adversarial learning framework. In this framework, the Q-function acts as the discriminator, while the robot policy (i.e., the trial assembly trajectory) acts as the generator. As a discriminator, the Q-function aims to distinguish whether state-action pairs originate from expert "real samples" (labeled y(s,a)=1) or from "generated samples" (labeled y(s,a)=0) generated by the robot policy. By minimizing the classification error—that is, accurately determining the source of the state-action pair—it can learn a discriminative ability to differentiate between the two. The robot policy, acting as a generator, aims to make its generated state-action pairs increasingly difficult for the Q-function to distinguish during the optimization process; in other words, it "deceives" the Q-function, making it judge the robot policy's generated state-action pairs as good as, or even better than, those of the expert. This process drives the robot policy to continuously improve, achieving expert-level assembly performance. By minimizing the loss function of inverse reinforcement learning (IRL), which calculates the error of the Q-function in distinguishing between expert and robot behavior, the Q-function can more accurately evaluate the reward of state-action pairs. This guides the generator part of the robot's policy to produce higher-quality action sequences in the real world.
[0094] It's important to note that the learning process of the Q-function is essentially the training of a fully connected network. The input includes state information, actions, and the task embedding vector output by the task encoder. The output is a value in the range [0,1], representing the probability that the state-action pair comes from the expert. By adjusting the network parameters to minimize the IRL loss, the Q-function can gradually learn the intrinsic reward structure of expert behavior, providing guidance for the robot's policy generator. This enables the robot to quickly and robustly complete assembly tasks when faced with new tasks without requiring extensive trial and error.
[0095] In another example, during training, the reference trajectory contains state-action pairs of "memory precisely installed into the slot," while the robot's trial-and-error trajectory may contain state-action pairs of "deviation from the slot." Training is performed using an IRL-based loss function: the discriminator classifies "precise installation into the slot" as expert behavior (high Q-value) and "deviation from the slot" as robot behavior (low Q-value). As the robot gradually learns and adjusts its position, it becomes harder for the discriminator to distinguish between these behaviors. This increases the Q-value of the Q-function for these types of actions, ultimately enabling the Q-function to not only reflect the optimal behavior of experts but also guide the robot to explore and return to the correct path when it deviates from the expert trajectory.
[0096] In this embodiment, the task evaluation network learns and compares reference assembly trajectories and experimental assembly trajectories to form a set of quantitative evaluation criteria, thereby more accurately evaluating the assembled model and improving the accuracy of model evaluation. Furthermore, based on the feedback from the task evaluation network, a joint training mechanism is implemented, which promotes the interaction between the task evaluation network and the assembled model, forming a closed-loop feedback, accelerating the convergence process of the entire model training, and saving training time.
[0097] In an exemplary embodiment, training the pre-trained assembly model based on the parameter-adjusted task evaluation network includes: randomly sampling a first sample from a first training set to form a second training set; using the second samples from the second training set as current second samples, performing the following training operations on the pre-trained assembly model: inputting specified trajectory information and reference trajectory information from the current second samples into the parameter-adjusted task evaluation network, and outputting the evaluation value corresponding to the current second sample; adjusting the parameters of the pre-trained assembly model based on the difference between the evaluation value corresponding to the current second sample and the reference evaluation value indicated by the label information of the current second sample, in order to train the pre-trained assembly model.
[0098] It should be noted that the second training set is a series of second samples randomly selected from the first training set, containing specified trajectory information, reference trajectory information, and label information, which can be used for subsequent model training. Optionally, multiple first samples can be randomly selected from the first training set to form the second training set.
[0099] The second training set may contain specified trajectory information and reference trajectory information for a specific assembly task, as well as label information indicating its true source (expert demonstration or experimental assembly). Optionally, the second samples in the second training set are used as the current second samples, and the following training operation is performed on the pre-trained assembly model: the specified trajectory information and reference trajectory information in the current second samples are fed into the parameter-adjusted task evaluation network to obtain an evaluation value representing the quality of the assembly operation. Subsequently, the assembly model adjusts its parameters and optimizes the assembly strategy based on the difference between this evaluation value and the reference evaluation value indicated by its label.
[0100] Optionally, the difference between the evaluated value and the reference evaluated value indicated by its label can be calculated using a specified loss function. The specified loss function can be the IRL loss function. The assembled model can be a fully connected network structure, such as a policy network.
[0101] In one example, taking the IRL loss function as the specified loss function, the updated task evaluation network is used to extract samples. Based on the difference value obtained from the specified loss function, the assembled model is updated. The above sample extraction training operation is repeated until the assembled model converges. Specifically, as shown in the following formula (2):
[0102]
[0103] in, This refers to the output of the assembly model. The corresponding input is and , This is the output of the task encoder, i.e., the embedding vector; This is the conditional policy loss function, i.e., the loss function for the assembly model.
[0104] Through this embodiment, the assembly model can quickly identify its own operational shortcomings and accurately adjust parameters by receiving feedback from the task evaluation network, thereby improving assembly efficiency and success rate in real-world environments.
[0105] In an exemplary embodiment, before constructing a first training set based on a component assembly trajectory library, the method includes: taking a reference training video from a plurality of reference training videos as a first reference training video, and sequentially performing the following first construction operation on the first reference training video to obtain a plurality of reference assembly trajectories corresponding to the plurality of reference training videos; extracting features from the first reference training videos to obtain first reference state-action pair information corresponding to the first reference training videos, so as to construct the reference assembly trajectory corresponding to the first reference training videos based on the first reference state-action pair information; obtaining a set of experimental assembly tasks, wherein the experimental assembly tasks in the set of experimental assembly tasks belong to one of the component assembly tasks corresponding to the plurality of reference training videos; executing a set of experimental assembly tasks according to the pre-trained task encoder and the pre-trained assembly model to construct the experimental assembly trajectory corresponding to the experimental assembly tasks in the set of experimental assembly tasks; and constructing a component assembly trajectory library based on the plurality of reference assembly trajectories and the experimental assembly trajectories corresponding to the set of experimental assembly tasks.
[0106] It should be noted that the reference training video can refer to a demonstration video obtained from experienced engineers. By analyzing and processing the reference training video, a series of state-action pairs can be extracted, thereby constructing a reference assembly trajectory. Optionally, individual reference training videos can be used sequentially as analysis objects. By extracting features from the keyframes in an individual reference training video, the position, attitude, and speed information of the robotic arm's end effector, as well as the position and attitude information of components, can be identified to form state-action pairs, thus constituting the reference assembly trajectory corresponding to a single reference training video.
[0107] Experimental assembly tasks can be assembly operations attempted by a pre-trained assembly model in a real-world environment, thereby collecting actual trajectory data of the model during execution. Optionally, based on a pre-trained task encoder, the assembly model can assemble multiple parts in a real-world environment, collecting their operational trajectories, i.e., experimental assembly trajectories.
[0108] In this embodiment, by constructing a component assembly trajectory library, data from multiple sources, including reference demonstrations and robot experiments, can be integrated to provide comprehensive training materials for the assembly model, accelerating training. Furthermore, the component assembly trajectory library can enrich the model's training data, enabling the initially trained assembly model to better understand and execute various component assembly tasks, thereby improving the model's performance.
[0109] In an exemplary embodiment, a set of experimental assembly tasks is executed based on a pre-trained task encoder and a pre-trained assembly model to construct experimental assembly trajectories corresponding to the experimental assembly tasks in the set of experimental assembly tasks. This includes: taking a set of experimental assembly tasks as the current experimental assembly task in sequence, and performing the following second construction operation respectively to construct the experimental assembly trajectory corresponding to the current experimental assembly task: inputting the reference trajectory information corresponding to the current experimental assembly task and the current state information of the assembly robot into the pre-trained assembly model, and outputting the current experimental action sequence corresponding to the current experimental assembly task; constructing the experimental assembly trajectory corresponding to the current experimental assembly task based on the operation information of the assembly robot executing the current experimental assembly task through the current experimental action sequence.
[0110] It should be noted that the reference trajectory information is the optimal assembly path extracted from expert demonstrations. The current state information of the assembly robot can include the position, orientation, and speed of the robotic arm's end effector, reflecting the real-time state of the assembly robot while performing the current task. The experimental assembly trajectory is a sequence of state changes of the assembly robot during the execution of the experimental assembly task, including changes in the position, orientation, and speed of the robot arm, as well as changes in the position and orientation of the components.
[0111] The reference trajectory information and the current state information of the assembly robot corresponding to the current test assembly task are input into the assembly model after preliminary training to obtain the operation information of executing the current test assembly task through the current test action sequence, so as to form the test assembly trajectory corresponding to the current test assembly task.
[0112] Optionally, the operational information for executing the current test assembly task through the current test action sequence can be obtained through test rehearsal videos. The test rehearsal videos can be videos of the test process captured during the execution of the current test action sequence.
[0113] In one example, if there are n component assembly tasks that need to be trained, a trajectory cache can be created for each component assembly task to store the experimental assembly trajectories corresponding to the assembly task. H can be the reference assembly trajectory for assembly task i, where H is the time length corresponding to the reference assembly trajectory. To represent a state-action pair, you can use... It can be represented based on assembly models (such as policy networks). The output of assembly task i is the experimental assembly trajectory.
[0114] Through this embodiment, by performing experimental assembly tasks in a real environment, valuable information on environmental factors (such as component materials, external interference, etc.) can be collected, thereby improving operational capability and robustness in actual production environments.
[0115] In one exemplary embodiment, constructing the test assembly trajectory corresponding to the current test assembly task based on the operation information of the assembly robot executing the current test assembly task through the current test action sequence includes: acquiring operation images of the assembly robot executing the current test assembly task when the assembly robot executes the current test assembly task based on the current test action sequence; and constructing the test assembly trajectory corresponding to the current test assembly task based on the operation images.
[0116] It should be noted that the operation image can refer to the image corresponding to each step of the operation captured during the assembly task performed by the assembly robot. It can include the status information of the assembly robot and the parts at various points in time. Optionally, it can be captured by a high frame rate camera.
[0117] Optionally, while the assembly robot is executing the current experimental action sequence given by the assembly model, the camera will track and record every operational detail of the assembly process in real time, forming a series of operational images.
[0118] Image processing can be used to identify and extract the state information of the assembly robot and its components from the operation images, and this information can be reconstructed in chronological order to form the experimental assembly trajectory.
[0119] Through this embodiment, by acquiring and recording operation images of the assembly robot's action sequence in real time, it is possible to intuitively understand the dynamic operation status of the robot when handling various component assembly tasks, thus forming a more complete and high-quality test assembly trajectory.
[0120] In an exemplary embodiment, the method further includes: using each of the multiple reference training videos as a second reference training video, and sequentially training the task encoder and assembly model to obtain a pre-trained task encoder and a pre-trained assembly model; inputting a set of training images corresponding to the second reference training video into the task encoder, and outputting reference trajectory information corresponding to the second reference training video, wherein the reference trajectory information corresponding to the second reference training video is used to indicate the reference assembly action sequence corresponding to the second reference training video; inputting the reference trajectory information and initial state information corresponding to the second reference training video into the assembly model to obtain the assembly action sequence corresponding to the second reference training video, wherein the initial state information is used to indicate the initial state of the assembly robot in the second reference training video; adjusting the parameters of the task encoder and the assembly model according to the difference between the assembly action sequence and the reference assembly action sequence corresponding to the second reference training video, until the task encoder and the assembly model converge.
[0121] It should be noted that a task encoder can be a neural network model that transforms image sequences from a reference training video into task feature vectors (also known as task embedding vectors). These vectors contain key information about the assembly actions, such as action type, target components, and operation sequence.
[0122] Optionally, a set of corresponding training images from the reference training video is input into the task encoder. The task encoder uses its corresponding convolutional layer, pooling layer, and fully connected layer to convert the set of corresponding training images from the reference training video into reference trajectory information representing the action sequence of the video.
[0123] The assembly module can be a deep learning-based control strategy model that can predict the sequence of actions of the robot's arm during the assembly process based on the input reference trajectory information and initial state information. The initial state information can include the position and orientation of the robot arm.
[0124] The assembly motion sequence is the flow of movements of the robotic arm during the assembly process output by the assembly model, while the reference assembly motion sequence is the optimal motion flow recorded in the expert demonstration. By comparing the motion sequence predicted by the assembly model with the reference assembly motion sequence in the expert demonstration, the differences between the two are identified. Based on these differences, the parameters of the task encoder and the assembly model are adjusted to narrow the gap between the model prediction and the expert standard.
[0125] Specifically, the action sequence predicted by the assembly model is compared with the reference assembly action sequence in the expert demonstration. This can be determined by a preset loss function, such as the behavior cloning loss function. The assembly model and task encoder are trained using the behavior cloning loss function. Specifically, as shown in the following formula (3):
[0126] (3)
[0127] in, and These are state-action pairs extracted from reference exercise videos, and the task encoder. This indicates that the output corresponding to the i-th assembly task is ,in The parameters of the CNN network are as follows: The network structure of the task encoder is as follows Figure 3 As shown, To assemble the model, for example, a fully connected network structure, the inputs are the state information s of the robotic arm end effector (including position, orientation, and velocity) and the output of the task encoder. This constitutes the task condition strategy. .
[0128] In this embodiment, multiple reference training videos are used for training, exposing the assembly model to diverse assembly tasks and scenarios, thereby enhancing the generalization ability of the assembly model and task encoder, enabling it to make reasonable assembly operations more quickly when faced with new parts and new assembly environments.
[0129] Figure 4 This is a schematic diagram of the component assembly control method of the assembly robot in this optional example, such as... Figure 4 As shown, this can specifically include: pre-training of the task encoder and assembly model, value function and model optimization, and testing on new tasks.
[0130] The pre-training process for the task encoder and assembly model can be as follows: Initial learning of the assembly model and task encoder is performed using Behavior Cloning (BC), enabling the robot to initially mimic the assembly behavior of an expert. Specifically, demonstration data of the expert during the server component assembly process is collected, and state-action pair sequences are extracted. These state-action pairs are used for supervised learning of the policy and task encoder, optimizing the parameters by minimizing the difference (e.g., mean squared error) between the actions predicted by the model and the actions actually performed by the expert.
[0131] The process of value function and model optimization can be as follows: In a real physical environment, data on robot operation is collected using a pre-trained assembly model and task encoder. Specifically, multiple tasks are selected (these tasks have corresponding reference demonstration videos and task embedding vectors for each task), and the robot is controlled to attempt assembly actions in these tasks, collecting state-action pair sequences to obtain the real experience of the robotic arm for multiple tasks. A hybrid experience pool is formed by combining the data from expert demonstrations and robot collection. A task-conditional Q function (i.e., evaluation function) that can evaluate the value of states and actions under given task conditions is learned through the hybrid experience pool. The Q function is constructed, with inputs including environmental state (5), executed action (a), and task embedding vector (z). Using the data in the hybrid experience pool, the Q function (i.e., the improved evaluation function) is trained through inverse learning (IRL) to distinguish the value of expert behavior from robot behavior, guiding the optimization of the assembly model. The cumulative experience of the robot in each task is evaluated according to the Q function, and the parameters of the assembly model are adjusted to maximize the expected cumulative reward. The experiment and model parameter optimization are repeated until the assembly model converges, that is, the robot's assembly performance in each task is stable and reaches the preset standard.
[0132] The new task testing process may include: processing expert demonstration examples (i.e., expert demonstration videos) of the new task, extracting key information, and using a task encoder to generate a new task embedding vector (z). The new task embedding vector and the current environment state are then input into the optimized assembly model to directly generate action sequences, thus assembling the new task without additional trial-and-error learning.
[0133] This optional example addresses the problems of traditional robot control methods in server component assembly tasks, such as programming burden, low experience efficiency, and weak ability to handle complex scenarios, by employing behavioral cloning and iterative learning in real-world environments. Through pre-training, Q-function learning, assembly model optimization, and new task testing, it achieves rapid learning, high robustness, and flexibility of robot control strategies, significantly improving assembly efficiency and quality. It is suitable for automated and intelligent server production lines.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0135] According to another aspect of the embodiments of this application, a component assembly control device for an assembly robot is also provided. This component assembly control device can be used to implement the component assembly control method for the assembly robot provided in the above embodiments, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0136] Embodiments of this application also provide a component assembly control device for an assembly robot, such as... Figure 5 As shown, the device includes:
[0137] The acquisition module 502 is used to acquire a specified component assembly task and a specified training video corresponding to the specified component assembly task. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not performed before, and the specified training video is used to record the assembly scene of the specified component assembly task.
[0138] The parsing module 504 is used to parse the specified exercise video to obtain the specified reference trajectory information, wherein the specified reference trajectory information is used to describe the reference assembly trajectory corresponding to the specified component assembly task;
[0139] The execution module 506 is used to input the specified reference trajectory information and the state information of the assembly robot into the pre-trained assembly model to obtain the predicted assembly action sequence output by the assembly model. The state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot.
[0140] Assembly module 508 is used to control the assembly robot to assemble specified parts according to the predicted assembly actions in the predicted assembly action sequence, so as to perform the specified part assembly task.
[0141] It should be noted that the acquisition module 502 in this embodiment can be used to execute the above step S202, the parsing module 504 in this embodiment can be used to execute the above step S204, the execution module 506 in this embodiment can be used to execute the above step S206, and the assembly module 508 in this embodiment can be used to execute the above step S208.
[0142] The embodiments provided in this application obtain a specified component assembly task and a corresponding specified training video. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not previously executed. The specified training video records the assembly scenario of the specified component assembly task. The specified training video is parsed to obtain specified reference trajectory information, which describes the reference assembly trajectory corresponding to the specified component assembly task. The specified reference trajectory information and the assembly robot's state information are input into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model. The assembly robot's state information indicates the state of the end effector of the assembly robot's robotic arm. Following the predicted assembly actions in the predicted assembly action sequence, the assembly robot is controlled to assemble the specified component to execute the specified component assembly task. By using this application, new component assembly tasks are learned from specified training videos, eliminating the need to write complex control programs for each new component, significantly reducing operational complexity and preparation time. By providing training videos and assembly models, the robot can quickly generate reasonable assembly strategies when faced with unseen parts, without the need for additional training. This greatly improves the robot's adaptability in complex and ever-changing production environments, thereby solving the technical problems of cumbersome operation and poor generalization ability in the component assembly control methods of assembly robots in related technologies. It also improves the operational flexibility and environmental adaptability of assembly robots.
[0143] In an exemplary embodiment, the parsing module 504 is further configured to: perform target detection on the specified training video to extract at least two training images from the specified training video; and input the at least two training images into a pre-trained task encoder to obtain specified reference trajectory information output by the task encoder.
[0144] In an exemplary embodiment, the parsing module 504 is further configured to: use a target detection model to sequentially perform target detection on video frames in a specified training video to obtain target detection results for the video frames in the specified training video; and identify video frames containing target components based on the target detection results for the video frames in the specified training video, wherein at least two training images include video frames containing target components, the target components are related components for performing the specified component assembly task, and the target components include the specified components.
[0145] In an exemplary embodiment, the component assembly control device for assembling the robot further includes: a parameter determination module, configured to determine image distribution parameters of at least two training images in a specified training video, wherein the video time of the specified training video is divided into multiple time periods, and the image distribution parameters are used to describe the distribution of at least two training images in different time periods of the multiple time periods;
[0146] The operation execution module is used to perform an image addition operation on at least two training images or an image reduction operation on at least two training images to update at least two training images when the image distribution parameters do not meet the specified distribution conditions.
[0147] In one exemplary embodiment, the task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. The parsing module 504 is further configured to: perform a convolution operation on at least two training images using at least one convolutional layer to obtain at least two training images after the convolution operation; perform pooling processing on the at least two training images after the convolution operation using at least one pooling layer to obtain at least two training images after pooling processing; and perform integration processing on the at least two training images after pooling processing using at least one fully connected layer to obtain specified reference trajectory information output by the task encoder.
[0148] In an exemplary embodiment, the component assembly control device for the assembly robot further includes: a first construction module, configured to construct a first training set based on a component assembly trajectory library, wherein the component assembly trajectory library is used to record the following trajectories corresponding to the component assembly task: a reference assembly trajectory and a trial assembly trajectory, wherein the trial assembly trajectory is an assembly trajectory obtained by assembling components in a real environment using a pre-trained assembly model, and a first sample in the first training set includes specified trajectory information, label information, and reference trajectory information, wherein in a first sample in the first training set, the label information is used to indicate whether the specified trajectory corresponding to the specified trajectory information is a reference assembly trajectory or a trial assembly trajectory, and the reference trajectory information is generated based on a reference practice video corresponding to the component assembly task; and a first training module, configured to train a task evaluation network and a pre-trained assembly model according to the first training set until the task evaluation network and the assembly model converge, wherein the task evaluation network is an evaluation network that evaluates the trial assembly trajectory using the reference assembly trajectory as a reference value.
[0149] In an exemplary embodiment, the first training module is further configured to use the first samples in the first training set as the current first samples, and perform the following training operations on the task evaluation network and the pre-trained assembly model: inputting the specified trajectory information in the current first sample and the reference trajectory information in the current first sample into the task evaluation network, and outputting the evaluation value corresponding to the current first sample; adjusting the network parameters of the task evaluation network based on the difference between the evaluation value corresponding to the current first sample and the reference evaluation value indicated by the label information of the current first sample, so as to train the task evaluation network; and training the pre-trained assembly model according to the task evaluation network with adjusted parameters.
[0150] In an exemplary embodiment, the first training module is further configured to: randomly sample a first sample from the first training set to form a second training set; and use the second samples from the second training set as current second samples to perform the following training operations on the pre-trained assembly model: inputting the specified trajectory information in the current second sample and the reference trajectory information in the current second sample into the parameter-adjusted task evaluation network, and outputting the evaluation value corresponding to the current second sample; and adjusting the parameters of the pre-trained assembly model based on the difference between the evaluation value corresponding to the current second sample and the reference evaluation value indicated by the label information of the current second sample, so as to train the pre-trained assembly model.
[0151] In one exemplary embodiment, the component assembly control device for the assembly robot further includes:
[0152] The second construction module is used to, before constructing the first training set based on the component assembly trajectory library, take the reference training videos from multiple reference training videos as first reference training videos, and sequentially perform the following first construction operations on the first reference training videos to obtain multiple reference assembly trajectories corresponding to the multiple reference training videos: extract features from the first reference training videos to obtain first reference state-action pair information corresponding to the first reference training videos, so as to construct the reference assembly trajectory corresponding to the first reference training videos based on the first reference state-action pair information; obtain a set of experimental assembly tasks, wherein the experimental assembly tasks in the set of experimental assembly tasks belong to one of the component assembly tasks corresponding to the multiple reference training videos; execute a set of experimental assembly tasks according to the pre-trained task encoder and the preliminarily trained assembly model to construct the experimental assembly trajectory corresponding to the experimental assembly tasks in the set of experimental assembly tasks; and construct a component assembly trajectory library based on the multiple reference assembly trajectories and the experimental assembly trajectories corresponding to the set of experimental assembly tasks.
[0153] In an exemplary embodiment, the second construction module is further configured to: sequentially take a group of experimental assembly tasks as the current experimental assembly tasks, and perform the following second construction operations respectively to construct the experimental assembly trajectory corresponding to the current experimental assembly task: input the reference trajectory information corresponding to the current experimental assembly task and the current state information of the assembly robot into the pre-trained assembly model, and output the current experimental action sequence corresponding to the current experimental assembly task; construct the experimental assembly trajectory corresponding to the current experimental assembly task based on the operation information of the assembly robot executing the current experimental assembly task through the current experimental action sequence.
[0154] In one exemplary embodiment, the second construction module is further configured to: acquire operation images of the assembly robot performing the current test assembly task based on the current test action sequence; and construct the test assembly trajectory corresponding to the current test assembly task based on the operation images.
[0155] In an exemplary embodiment, the component assembly control device for the assembly robot further includes: a second training module, configured to: use reference training videos from a plurality of reference training videos as second reference training videos, and sequentially train the task encoder and the assembly model to obtain a pre-trained task encoder and a pre-trained assembly model; input a set of training images corresponding to the second reference training videos into the task encoder, and output reference trajectory information corresponding to the second reference training videos, wherein the reference trajectory information corresponding to the second reference training videos is used to indicate the reference assembly action sequence corresponding to the second reference training videos; input the reference trajectory information and initial state information corresponding to the second reference training videos into the assembly model to obtain the assembly action sequence corresponding to the second reference training videos, wherein the initial state information is used to indicate the initial state of the assembly robot in the second reference training videos; and adjust the parameters of the task encoder and the assembly model according to the difference between the assembly action sequence and the reference assembly action sequence corresponding to the second reference training videos, until the task encoder and the assembly model converge.
[0156] For a description of the features in the embodiment of the component assembly control device for the assembly robot, please refer to the relevant description of the embodiment of the component assembly control method for the assembly robot, which will not be repeated here.
[0157] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the component assembly control method for assembly robots.
[0158] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the component assembly control method for assembly robots when it is run.
[0159] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0160] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the component assembly control method for assembly robots.
[0161] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the component assembly control method for assembly robots.
[0162] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0163] The foregoing has provided a detailed description of the component assembly control method, electronic device, and storage medium for an assembly robot provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A component assembly control method for an assembly robot, characterized in that, include: Obtain a specified component assembly task and a specified training video corresponding to the specified component assembly task. The specified component assembly task is an assembly task performed on a specified component that the assembly robot has not performed before. The specified training video is used to record the assembly scenario of the specified component assembly task. The specified training video is analyzed to obtain specified reference trajectory information, wherein the specified reference trajectory information is used to describe the reference assembly trajectory corresponding to the specified component assembly task; The specified reference trajectory information and the state information of the assembly robot are input into the pre-trained assembly model to obtain the predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot; According to the predicted assembly actions in the predicted assembly action sequence, the assembly robot is controlled to assemble the specified component to perform the specified component assembly task; The step of parsing the specified training video to obtain specified reference trajectory information includes: Target detection is performed on the specified training video to extract at least two training images from the specified training video; The at least two training images are input into a pre-trained task encoder to obtain the specified reference trajectory information output by the task encoder; After performing object detection on the specified training video to extract at least two training images from the specified training video, the method further includes: Determine the image distribution parameters of the at least two training images in the specified training video, wherein the video time of the specified training video is divided into multiple time periods, and the image distribution parameters are used to describe the distribution of the at least two training images in different time periods of the multiple time periods; If the image distribution parameters do not meet the specified distribution conditions, perform an image addition operation or an image reduction operation on the at least two training images to update the at least two training images.
2. The method according to claim 1, characterized in that, The step of performing target detection on the specified training video to extract at least two training images from the specified training video includes: Using an object detection model, object detection is performed sequentially on the video frames in the specified training video to obtain the object detection results of the video frames in the specified training video; Based on the target detection results of the video frames in the specified training video, video frames containing target components are identified, wherein the at least two training images include video frames containing the target components, the target components are related components for performing the specified component assembly task, and the target components include the specified components.
3. The method according to claim 1, characterized in that, The task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. The step of inputting the at least two training images into a pre-trained task encoder to obtain the specified reference trajectory information output by the task encoder includes: The at least two training images are convolved by the at least one convolutional layer to obtain the at least two training images after the convolution operation. The at least two training images after the convolution operation are pooled through the at least one pooling layer to obtain the at least two training images after pooling. The at least two training images after pooling are integrated through the at least one fully connected layer to obtain the specified reference trajectory information output by the task encoder.
4. The method according to claim 1, characterized in that, The method further includes: Based on the component assembly trajectory library, a first training set is constructed. The component assembly trajectory library is used to record the following trajectories corresponding to the component assembly task: the reference assembly trajectory and the experimental assembly trajectory. The experimental assembly trajectory is the assembly trajectory obtained by assembling components in a real environment using the pre-trained assembly model. The first sample in the first training set includes specified trajectory information, label information, and reference trajectory information. In a first sample in the first training set, the label information is used to indicate that the specified trajectory corresponding to the specified trajectory information is the reference assembly trajectory or the experimental assembly trajectory. The reference trajectory information is generated based on the reference training video corresponding to the component assembly task. Based on the first training set, the task evaluation network and the pre-trained assembly model are trained until the task evaluation network and the assembly model converge, wherein the task evaluation network is an evaluation network that evaluates the experimental assembly trajectory using the reference assembly trajectory as a reference value.
5. The method according to claim 4, characterized in that, The step of training the task evaluation network and the pre-trained assembly model based on the first training set includes: Using the first sample in the first training set as the current first sample, perform the following training operations on the task evaluation network and the pre-trained assembled model: The specified trajectory information and the reference trajectory information in the current first sample are input into the task evaluation network, and the evaluation value corresponding to the current first sample is output. Based on the difference between the evaluation value corresponding to the current first sample and the reference evaluation value indicated by the label information of the current first sample, the network parameters of the task evaluation network are adjusted to train the task evaluation network. The assembled model, after initial training, is trained based on the task evaluation network with adjusted parameters.
6. The method according to claim 5, characterized in that, The step of training the pre-trained assembly model using the task evaluation network adjusted according to the parameters includes: The first sample in the first training set is randomly selected to form the second training set; Using the second samples from the second training set as the current second samples, perform the following training operations on the pre-trained assembled model: The specified trajectory information and the reference trajectory information in the current second sample are input into the task evaluation network after parameter adjustment, and the evaluation value corresponding to the current second sample is output. Based on the difference between the evaluation value corresponding to the current second sample and the reference evaluation value indicated by the label information of the current second sample, the parameters of the assembled model after preliminary training are adjusted to train the assembled model after preliminary training.
7. The method according to claim 4, characterized in that, Before constructing the first training set based on the component assembly trajectory library, the method includes: Each of the multiple reference training videos is used as a first reference training video. The following first construction operation is then performed on each of the first reference training videos to obtain multiple reference assembly trajectories corresponding to the multiple reference training videos: Feature extraction is performed on the first reference training video to obtain the first reference state action pair information corresponding to the first reference training video, so as to construct the reference assembly trajectory corresponding to the first reference training video based on the first reference state action pair information. Obtain a set of test assembly tasks, wherein the test assembly tasks in the set of test assembly tasks belong to one of the component assembly tasks corresponding to the multiple reference exercise videos; Based on the pre-trained task encoder and the pre-trained assembly model, the set of experimental assembly tasks is executed to construct the experimental assembly trajectory corresponding to the experimental assembly task in the set of experimental assembly tasks; The component assembly trajectory library is constructed based on the multiple reference assembly trajectories and the test assembly trajectories corresponding to the set of test assembly tasks.
8. The method according to claim 7, characterized in that, The step of executing the set of experimental assembly tasks based on the pre-trained task encoder and the initially trained assembly model to construct the experimental assembly trajectory corresponding to the experimental assembly tasks in the set of experimental assembly tasks includes: The set of test assembly tasks are sequentially used as the current test assembly task, and the following second construction operation is performed respectively to construct the test assembly trajectory corresponding to the current test assembly task: The reference trajectory information corresponding to the current test assembly task and the current state information of the assembly robot are input into the assembly model after preliminary training, and the current test action sequence corresponding to the current test assembly task is output. Based on the operation information of the assembly robot executing the current test assembly task through the current test action sequence, the test assembly trajectory corresponding to the current test assembly task is constructed.
9. The method according to claim 8, characterized in that, The step of constructing the test assembly trajectory corresponding to the current test assembly task based on the operation information of the assembly robot executing the current test assembly task through the current test action sequence includes: When the assembly robot performs the current test assembly task based on the current test action sequence, an operation image of the assembly robot performing the current test assembly task is acquired; Based on the operation image, construct the test assembly trajectory corresponding to the current test assembly task.
10. The method according to claim 7, characterized in that, The method further includes: Using each of the multiple reference training videos as a second reference training video, the task encoder and the assembly model are trained sequentially to obtain the pre-trained task encoder and the pre-trained assembly model. A set of training images corresponding to the second reference training video is input to the task encoder, and reference trajectory information corresponding to the second reference training video is output. The reference trajectory information corresponding to the second reference training video is used to indicate the reference assembly action sequence corresponding to the second reference training video. The reference trajectory information and initial state information corresponding to the second reference training video are input into the assembly model to obtain the assembly action sequence corresponding to the second reference training video, wherein the initial state information is used to indicate the initial state of the assembly robot in the second reference training video; Based on the difference between the assembly action sequence and the reference assembly action sequence corresponding to the second reference training video, adjust the parameters of the task encoder and the assembly model until the task encoder and the assembly model converge.
11. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the component assembly control method for the assembly robot as described in any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the component assembly control method of the assembly robot as described in any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the component assembly control method for the assembly robot as described in any one of claims 1 to 10.
Citation Information
Patent Citations
3C assembly programming-free method based on large model action analysis and automatic programming
CN119027618A