Component assembly control method of assembly robot, electronic equipment and storage medium
By learning new component assembly tasks from specified rehearsal videos and generating predicted assembly action sequences, the problems of cumbersome assembly robot operations and poor generalization ability are solved, and the robot's efficient adaptation and flexible assembly are achieved in diverse production environments.
Patent Information
- Application Number
- CN202511280918.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
The component assembly control method of the assembly robot in the existing technology has the problems of cumbersome operation and poor generalization ability. Especially when faced with a variety of server components, detailed assembly programs need to be manually written, and it is difficult to quickly learn from existing experience and apply it to the assembly of new components.
By obtaining the specified rehearsal video corresponding to the specified component assembly task, parsing the video to obtain reference trajectory information, and inputting it into the pre-trained assembly model, a predicted assembly action sequence is generated and the robot is controlled to perform assembly, avoiding the need to write complex control programs for each new component.
It greatly reduces the operational complexity and preliminary preparation time, improves the robot's adaptability and operational flexibility in complex and changing production environments, and can quickly generate reasonable assembly strategies without additional training.
Smart Images

Figure CN120755892A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot manufacturing technology, and in particular to a component assembly control method, electronic equipment, and storage medium for an assembly robot. Background Art
[0002] In the field of server production and maintenance, the application of robotic technology is gradually expanding, especially in the component assembly link. The robot component assembly method in the related art mainly relies on precise pre-programmed trajectories and fixed rules. Although this method has high accuracy in specific tasks, it exposes obvious limitations when faced with a variety of server components. For example, the operation is cumbersome and the generalization ability is poor. For each new model or specification of server component, a detailed assembly program needs to be written manually. This process is not only time-consuming and labor-intensive, but also has the problem of cumbersome operation. It also has high requirements for the operator, requiring rich programming experience and a deep understanding of mechanical systems. Pre-programmed robotic systems are difficult to learn from existing experience and apply it to the assembly of new components. This means that every time a new type of component is encountered, it needs to be reprogrammed, and there is a problem of poor generalization ability.
[0003] Therefore, the component assembly control method of the assembly robot in the related art has technical problems of complicated operation and poor generalization ability. Summary of the Invention
[0004] The present application provides a component assembly control method, electronic equipment and storage medium for an assembly robot, so as to at least solve the problems of cumbersome operation and poor generalization ability of the component assembly control method for an assembly robot in the related art.
[0005] The present application provides a component assembly control method for an assembly robot, comprising:
[0006] Obtaining a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task;
[0007] Parsing the designated drill video to obtain designated reference trajectory information, wherein the designated reference trajectory information is used to describe a reference assembly trajectory corresponding to the designated component assembly task;
[0008] Inputting the specified reference trajectory information and state information of the assembly robot into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot;
[0009] According to the predicted assembly action in the predicted assembly action sequence, the assembly robot is controlled to assemble the designated component to perform the designated component assembly task.
[0010] The present application also provides a component assembly control device for an assembly robot, comprising:
[0011] an acquisition module, configured to acquire a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task;
[0012] a parsing module, configured to parse the designated drill video to obtain designated reference trajectory information, wherein the designated reference trajectory information is used to describe a reference assembly trajectory corresponding to the designated component assembly task;
[0013] an execution module, configured to input the specified reference trajectory information and state information of the assembly robot into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate a state of an end effector of a robotic arm of the assembly robot;
[0014] An assembly module is used to control the assembly robot to assemble the designated components according to the predicted assembly actions in the predicted assembly action sequence, so as to perform the designated component assembly task.
[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned component assembly control methods for an assembly robot when executing the computer program.
[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned component assembly control methods of the assembly robot are implemented.
[0017] The present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned component assembly control methods of the assembly robot when the computer program is executed by a processor.
[0018] Through this application, a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task are obtained, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task; the designated rehearsal video is parsed to obtain designated reference trajectory information, wherein the designated reference trajectory information is used to describe the reference assembly trajectory corresponding to the designated component assembly task; the designated reference trajectory information and the state information of the assembly robot are input into the pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot; according to the predicted assembly action in the predicted assembly action sequence, the assembly robot is controlled to assemble the designated component to perform the designated component assembly task. By adopting this application, by learning new component assembly tasks from designated rehearsal videos, it is no longer necessary to write complex control programs for each new component, which greatly reduces the complexity of the operation and the pre-preparation time. By specifying rehearsal videos and assembly models, it is possible to quickly generate reasonable assembly strategies when faced with unprecedented parts without the need for additional training, greatly improving the robot's adaptability in complex and changing production environments. This solves the technical problems of cumbersome operation and poor generalization ability of component assembly control methods of assembly robots in related technologies, and improves the operational flexibility and environmental adaptability of assembly robots. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 It is a schematic diagram of an application scenario of a component assembly control method of an assembly robot according to an embodiment of the present application.
[0021] Figure 2 It is a flow chart of an optional component assembly control method of an assembly robot according to an embodiment of the present application.
[0022] Figure 3 is a schematic diagram of an optional task encoder according to an embodiment of the present application.
[0023] Figure 4 It is a schematic diagram of an optional component assembly control method of an assembly robot according to an embodiment of the present application.
[0024] Figure 5 This is a structural block diagram of an optional component assembly control device of an assembly robot according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0026] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0027] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0028] According to an aspect of the embodiments of the present application, a component assembly control method for assembling a robot is provided. Optionally, in the present embodiment, the component assembly control method for assembling a robot can be applied to, but is not limited to, a hardware environment as shown in the figure comprising a terminal device 102 and a server 104. The server 104 can be connected with the terminal device 102 through a network, and can be used to provide services (for example, application services, etc.) for the terminal device 102 or a client installed on the terminal device 102, and a database can be set on the server 104 or independently of the server 104, used to provide data storage services for the server 104. Figure 1 The network can include, but is not limited to, at least one of the following: a wired network, a wireless network. The wired network can include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The wireless network can include, but is not limited to, at least one of the following: Wireless Fidelity (WIFI), Bluetooth. The terminal device 102 can be, but is not limited to, a personal computer (PC), a mobile phone, a tablet computer, etc. The server 104 can be, but is not limited to, a cloud server, a server cluster or other server types.
[0029]
[0030] The component assembly control method for the assembly robot according to the embodiment of the present application may be executed by the server 104, or by the terminal device 102, or jointly by the server 104 and the terminal device 102. The component assembly control method for the assembly robot according to the embodiment of the present application may be executed by the terminal device 102 or by a client installed thereon.
[0031] Taking the terminal device 102 as an example, the component assembly control method of the assembly robot in this embodiment is executed. Figure 2 is a flow chart of an optional component assembly control method of an assembly robot according to an embodiment of the present application, such as Figure 2 As shown, the process of the method may include the following steps:
[0032] Step S202: obtaining a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task;
[0033] Step S204: parsing the designated drill video to obtain designated reference trajectory information, wherein the designated reference trajectory information is used to describe a reference assembly trajectory corresponding to the designated component assembly task;
[0034] Step S206: Input the specified reference trajectory information and the state information of the assembly robot into the pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the assembly robot's robotic arm;
[0035] Step S208 : controlling the assembly robot to assemble the designated components according to the predicted assembly actions in the predicted assembly action sequence, so as to perform the designated component assembly task.
[0036] The component assembly control method for an assembly robot in this embodiment can be applied to the field of robotic manufacturing technology, specifically in server production and maintenance scenarios. A specific scenario is an automated server assembly line. On such a production line, assembly robots are responsible for accurately and safely assembling various server components, including but not limited to CPUs, memory modules, hard drives, heat sinks, and more. Automated assembly lines aim to improve production efficiency and reduce costs while ensuring the quality and reliability of server assembly. In related art, component assembly control methods for assembly robots typically rely on precise programming and pre-set assembly paths. Engineers typically program detailed assembly steps and robotic arm motion trajectories for each server component. While this approach can achieve a certain degree of automation for assembly tasks, it suffers from cumbersome operations and poor generalization. Complex reprogramming is required for each new component or assembly task, which is time-consuming and error-prone. As server technology evolves and the number of component types and specifications continues to increase, the efficiency and flexibility of this programming approach are significantly challenged, and the operation is cumbersome. Furthermore, it is difficult to quickly learn from existing assembly experience and apply it to new components or assembly tasks, limiting the assembly robot's ability to adapt to diverse production needs. Therefore, the component assembly control method of the assembly robot in the related art has technical problems of complicated operation and poor generalization ability.
[0037] To at least partially address the aforementioned technical issues, this embodiment learns new component assembly tasks from designated practice videos, eliminating the need to write complex control programs for each new component, significantly reducing operational complexity and pre-production time. By specifying practice videos and assembly models, the robot can quickly generate reasonable assembly strategies for unseen components without requiring additional training, significantly improving its adaptability in complex and changing production environments. This addresses the technical issues of cumbersome operation and poor generalization capabilities of component assembly control methods used by assembly robots in related technologies, improving the robot's operational flexibility and environmental adaptability.
[0038] It should be noted that the designated component assembly task may refer to a component assembly task of a new type that has never been performed by the assembly robot, such as the installation of a new model hard disk bracket encountered for the first time during the server assembly process.
[0039] Optionally, upon receiving a component assembly instruction, the system can identify whether the instruction indicates a new component assembly task to be performed, i.e., a designated component assembly task. Upon identifying the designated component assembly task, a designated practice video corresponding to the designated component assembly task can be retrieved from a pre-set database. The designated practice video can be used to record the assembly scenario of the designated component assembly task, primarily recording the process of an expert completing the designated component assembly task, and can be considered a "teaching material" for the assembly robot to learn.
[0040] It should be noted that the specified reference trajectory information can be used to describe a reference assembly trajectory corresponding to the specified component assembly task. The reference assembly trajectory can be a reference trajectory of an assembly task corresponding to a reference demonstration video generated based on the reference demonstration video, can include a series of state-action pair information, and can constitute an ideal path of an expert completing the assembly task.
[0041] Optionally, the specified demonstration video is parsed to obtain the specified reference trajectory information can include: the specified demonstration video is decomposed to obtain a series of video frames corresponding to the specified demonstration video, a series of key video frames are identified from the series of video frames, the position, posture, speed of the mechanical arm of the assembly robot and the related state information of the component to be assembled are identified according to the series of key video frames, and thus the specified reference trajectory information is obtained.
[0042] Optionally, the specified reference trajectory information is a low-dimensional embedding vector converted from the video frames corresponding to the specified demonstration video using a preset encoder, which can represent the assembly strategy of the expert.
[0043] It should be noted that the pre-trained assembly model can be a deep learning model that has been trained on a large amount of data set in advance, can learn the specified reference trajectory information and the robot state information, and can predict a set of suitable assembly action sequences for guiding the robot to complete the assembly task. The specified reference trajectory information and the state information of the assembly robot are input into the pre-trained assembly model to obtain the predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot can be used to indicate the state of the end effector of the mechanical arm of the assembly robot. The predicted assembly action sequence can include a series of action instructions, and the series of action instructions have a specified execution order for indicating and guiding the robot to assemble to complete the assembly task. Optionally, the predicted assembly action sequence generated by the assembly model can include a series of mechanical arm actions for actual assembly of the specified component.
[0044] Optionally, the assembly robot controls the end effector of the mechanical arm of the assembly robot to perform precise assembly operations on the specified component through the predicted assembly action sequence. Optionally, the robot will execute each action in the sequence one by one until the entire component assembly task is completed.
[0045] By the embodiments of the present application, a specified component assembly task and a specified demonstration video corresponding to the specified component assembly task are obtained, wherein the specified component assembly task is an assembly task performed on a specified component and not performed by the assembly robot, and the specified demonstration video is used to record an assembly scene of the specified component assembly task; the specified demonstration video is analyzed to obtain specified reference trajectory information, wherein the specified reference trajectory information is used to describe a reference assembly trajectory corresponding to the specified component assembly task; the specified reference trajectory information and state information of the assembly robot are input into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate a state of an end effector of a mechanical arm of the assembly robot; and the assembly robot is controlled to assemble the specified component according to a predicted assembly action in the predicted assembly action sequence to perform the specified component assembly task. By the present application, a new component assembly task is learned from the specified demonstration video, and a complex control program does not need to be written for each new component, thereby greatly reducing the complexity of operation and the preparation time in advance. Through the specified demonstration video and the assembly model, a reasonable assembly strategy can be quickly generated when facing an unknown component without additional training, thereby greatly improving the adaptability of the robot in a complex and variable production environment, so as to solve the technical problems of complicated operation and poor generalization ability of the component assembly control method of the assembly robot in the related art, and improve the operation flexibility and environmental adaptability of the assembly robot.
[0046] In one example embodiment, the specified demonstration video is analyzed to obtain the specified reference trajectory information, including: target detection is performed on the specified demonstration video to extract at least two demonstration images in the specified demonstration video; and the at least two demonstration images are input into a pre-trained task encoder to obtain specified reference trajectory information output by the task encoder.
[0047] It should be noted that the specified demonstration video refers to video materials in which an experienced operator (or expert) demonstrates assembly of a server component. These videos contain specific actions of how the operator accurately aligns, grabs, and installs the components. Target detection can refer to a computer vision technology that can be used to identify and locate objects in an image. In the present embodiment, target detection is used to extract image frames (demonstration images) containing key assembly actions from the demonstration video, such as the moment when the end effector of the mechanical arm contacts and aligns with the server component.
[0048] By performing target detection on the specified demonstration video, demonstration images corresponding to assembly operations in the specified component assembly task are extracted from the specified demonstration video. The at least two demonstration images can be used to cover multiple key stages of the corresponding assembly operations in the component assembly task.
[0049] The pre-trained task encoder can be a deep learning model that has been trained on a large dataset and can convert the input image into a set of compact numerical representations (i.e., task embedding vectors) that capture the key properties of the assembly task in the video.
[0050] The reference trajectory information, derived from the analysis of the rehearsal images by a pre-trained task encoder, contains a quantitative description of the expert's actions during the assembly process. This information can include the robot's position coordinates, posture angle, movement speed, and the relative positions and angles of the components.
[0051] In one example, target detection technology is used to locate the moment of a specified component assembly task in a specified rehearsal video to obtain at least two rehearsal images, and the at least two image frames are input into a pre-trained task encoder. The task encoder uses its internal convolutional neural network (CNN) and other layer structures to extract and compress features in the image and output a low-dimensional vector, namely the specified reference trajectory information.
[0052] Optionally, each component assembly task can correspond to one or more reference drill videos consisting of different server component assembly scenarios, such as CPU installation, memory module insertion and removal, and hard drive tray installation. Preprocessing is performed on the reference drill videos to obtain key state-action pairs, which can include information such as the position, posture, and speed of the robot end effector, as well as the position and posture of the component (e.g., the component to be installed).
[0053] Through this embodiment, by accurately extracting the specified reference trajectory information from the specified rehearsal video, it is possible to more accurately learn and imitate assembly skills, thereby improving the assembly accuracy of the robotic arm of the assembly robot.
[0054] In an exemplary embodiment, target detection is performed on a designated rehearsal video to extract at least two rehearsal images from the designated rehearsal video, including: using a target detection model to sequentially perform target detection on video frames in the designated rehearsal video to obtain target detection results for the video frames in the designated rehearsal video; and identifying video frames containing target components based on the target detection results for the video frames in the designated rehearsal video, wherein at least two rehearsal images include video frames containing target components, the target components are related components for performing designated component assembly tasks, and the target components include designated components.
[0055] It should be noted that the target detection model can be a machine learning model, such as YOLOv4, which can be used to identify and locate specific objects in an image. In an embodiment, the target detection model can be trained to identify the end effector and server components of a robot arm.
[0056] Optionally, a set of video frames in a specified rehearsal video is identified, and target detection is performed sequentially on the set of video frames in the specified rehearsal video using an object detection model in the order of the frames, thereby obtaining a detection result for each video frame. Optionally, the specified rehearsal video can be decomposed into a series of static images, i.e., a set of video frames, through frame-by-frame analysis. The target detection model can be used to identify each frame to identify the target component and the state of the robotic arm involved in the assembly process.
[0057] In one example, consider a walkthrough video showing a hard drive installation. The object detection model would look for the hard drive and the robotic arm's end effector in each frame. For example, in the first few seconds of the video, the model would identify the robotic arm approaching the hard drive, and in the next frame, the hard drive would be grasped by the robotic arm and ready to be inserted into the carrier.
[0058] Based on the object detection results from the reference walkthrough video, video frames containing the target component are identified. These are known as key frames, which are those in the demonstration video that contain key assembly actions, such as the moment when the robotic arm aligns with the hard drive slot. Optionally, the object detection results can be used to filter out frames where the robotic arm's end effector interacts with the target component. Selecting key frames from the numerous frames allows for focusing on images that truly capture assembly skill details, providing more accurate data for subsequent processing.
[0059] At least two of the training images include video frames containing a target component, wherein the target component is a component associated with performing a specified component assembly task, and the target component includes a specified component. The training image in the at least two training images is an image identified as containing a specific assembly action. The training image may include state-action information during the assembly process.
[0060] Through this embodiment, the target detection model is used to identify video frames containing target components, thereby accurately identifying key information of the assembly process in a specified rehearsal video. This not only speeds up the learning speed of the model, but also greatly enhances the generalization ability and robustness of the model, enabling the assembly robot to learn and perform complex server component assembly tasks with only a small amount of demonstration.
[0061] In an exemplary embodiment, after performing target detection on a specified rehearsal video to extract at least two rehearsal images in the specified rehearsal video, the above method further includes: determining image distribution parameters of the at least two rehearsal images in the specified rehearsal video, wherein the video time of the specified rehearsal video is divided into multiple time periods, and the image distribution parameters are used to describe the distribution of the at least two rehearsal images in different time periods among the multiple time periods; if the image distribution parameters do not meet the specified distribution conditions, performing an image addition operation on the at least two rehearsal images, or performing an image reduction operation on the at least two rehearsal images, to update the at least two rehearsal images.
[0062] It should be noted that the image distribution parameters may refer to the distribution characteristics of the selected drill images in different time periods in a video, which may include but are not limited to frequency, density, time interval, etc.
[0063] Optionally, after performing target detection and extracting at least two key rehearsal images, the distribution parameters of at least two rehearsal images on the video timeline of a specified rehearsal video can be determined by the time of the frames in which the at least two rehearsal images are located, which can help understand the rhythm of the demonstration and the time sequence of key actions. For example, taking hardware installation as an example, assuming that the specified rehearsal video is 30 seconds in total, two key moments are identified. The first moment occurs at the 10th second, when the robotic arm begins to grab the hard drive; the second moment occurs at the 20th second, when the hard drive is successfully inserted into the bracket. At this point, the time distribution of these two moments is analyzed, and it is found that they occupy 1 / 3 and 2 / 3 of the time points of the video respectively. These are the image distribution parameters.
[0064] Optionally, the specified distribution condition may refer to an ideal time distribution standard of preset rehearsal images in the video based on domain knowledge or task requirements.
[0065] Optionally, the image distribution parameter is compared with a preset specified distribution condition, wherein the specified distribution condition is determined based on the component assembly task, and optionally, the specified distribution condition is determined based on the complexity of the assembly operations corresponding to the component assembly task. The number of practice images in each time period in the specified distribution condition is positively correlated with the number of assembly operations in the component assembly task.
[0066] When the image distribution parameters do not meet the specified distribution conditions, an image adjustment operation is performed, which may specifically include: when there is a first specified time period, randomly extracting a specified number of rehearsal images from the first specified time period in the specified rehearsal video, or processing at least two rehearsal images to obtain new images, and supplementing the images to the first specified time period; the first specified time period is a time period in which the number of rehearsal images in multiple time periods is less than a first threshold value.
[0067] In the presence of a second specified time period, a pre-trained image screening model is used to automatically screen and delete repeated or redundant rehearsal images from the second specified time period, or the similarity between at least two rehearsal images is calculated to reduce the rehearsal images with the greatest similarity or a similarity greater than a specified similarity threshold. The second specified time period is a time period in which the number of rehearsal images in multiple time periods is greater than the first threshold.
[0068] Through the above embodiment, by fine-tuning the key images in the rehearsal video, it is ensured that there are sufficient learning samples in each assembly stage, and excessive redundancy of data is avoided, thereby improving the learning efficiency and generalization ability of the assembly model.
[0069] Optionally, there is a certain temporal sequence between at least two training images, which can enable the assembly model to better understand the temporal sequence of the assembly task and improve the consistency and success rate of the assembly.
[0070] In one embodiment, time series analysis techniques, such as long short-term memory (LSTM) networks, can be introduced to further optimize the model's ability to process time series information and solve the problem of strong time dependence in assembly tasks.
[0071] This embodiment optimizes image distribution by adding or subtracting images, enabling the model to more efficiently learn key assembly actions, reducing redundant learning time and focusing on key skill points, thereby improving model learning efficiency. Image distribution parameters ensure that the model receives rehearsal images covering all stages of the assembly process, helping the model understand the transitions between different assembly stages and improving its adaptability to unseen parts or assembly scenarios.
[0072] In an exemplary embodiment, the task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer; at least two training images are input into the pre-trained task encoder to obtain specified reference trajectory information output by the task encoder, including: performing a convolution operation on the at least two training images through at least one convolutional layer to obtain at least two training images after the convolution operation; performing pooling processing on the at least two training images after the convolution operation through at least one pooling layer to obtain at least two pooled training images; and performing integration processing on the at least two pooled training images through at least one fully connected layer to obtain the specified reference trajectory information output by the task encoder.
[0073] It should be noted that the task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. The convolutional layer in at least one of the convolutional layers can automatically detect and extract key features of the memory stick, the robotic arm end effector, and their interactions by sliding the convolution kernel over the input image (at least two training images). For example, the convolutional layer can identify the edge features where the robotic arm end effector aligns with the memory slot, as well as the texture changes as the memory stick gradually moves deeper into the slot.
[0074] The feature maps output by the convolutional layer may contain a large number of data points. Pooling layers (such as maximum pooling or average pooling) are used to reduce the dimensionality of these feature maps, retain the most significant visual features, remove redundant information, and improve the computational efficiency of the algorithm.
[0075] At least one fully connected layer in the fully connected layer can be used to use the feature vector output by the previous layer as the input of the weights between nodes, perform linear combination, and output a fixed-size vector for classification or regression tasks, in this case, to generate reference trajectory information. The pooled features are fed into the fully connected layer for intensive neural network calculations, converting the image features into a compact set of digital representations, namely the reference trajectory information. This set of information contains the core details of the operation, such as the initial position of the memory stick, the motion trajectory of the robot arm, and the final insertion posture, which is used to guide the model how to complete similar assembly actions.
[0076] The convolution layer is responsible for extracting local features of at least two training images, the pooling layer is used to perform feature dimensionality reduction and anti-interference processing, and the fully connected layer is used to integrate information to form compact specified reference trajectory information, which can be a low-dimensional embedding vector.
[0077] In one example, the task encoder can be Figure 3 As shown, the task encoder may include convolutional layers, pooling layers, and fully connected layers. The input image (input) is fed into the convolutional layer. The input image can be a video frame from a specified rehearsal video, such as at least two rehearsal images. Multiple convolution kernels in the convolutional layer are used, each of which performs a convolution operation to generate corresponding feature maps. The pooling layer performs maximum pooling or average pooling, i.e., downsampling, to obtain the input image after convolution and pooling operations. The fully connected layer integrates the input image after convolution and pooling operations to output a low-dimensional embedding vector, i.e., the specified reference trajectory information.
[0078] Through this embodiment, through the task encoder including convolutional layers, pooling layers and fully connected layers, it is possible to efficiently and accurately extract key features of the assembly task from the practice images and convert them into compact reference trajectory information.
[0079] In an exemplary embodiment, the above method also includes: constructing a first training set based on a component assembly trajectory library, wherein the component assembly trajectory library is used to record the following trajectories corresponding to the component assembly task: a reference assembly trajectory and a test assembly trajectory, the test assembly trajectory is an assembly trajectory obtained by assembling components in a real environment using a preliminarily trained assembly model, the first sample in the first training set includes specified trajectory information, label information and reference trajectory information, in a first sample in the first training set, the label information is used to indicate that the specified trajectory corresponding to the specified trajectory information is a reference assembly trajectory or a test assembly trajectory, and the reference trajectory information is generated based on a reference rehearsal video corresponding to the component assembly task; according to the first training set, the task evaluation network and the preliminarily trained assembly model are trained until the task evaluation network and the assembly model converge, wherein the task evaluation network is an evaluation network that uses the reference assembly trajectory as a reference value to evaluate the test assembly trajectory.
[0080] It should be noted that the component assembly trajectory library is a database that stores reference assembly trajectories and experimental assembly trajectories for different component assembly tasks. The reference assembly trajectories are derived from expert demonstration videos, while the experimental assembly trajectories are trajectory records generated when the preliminarily trained assembly model attempts to assemble components in a real environment.
[0081] Optionally, a first training set covering reference assembly trajectories and experimental assembly trajectories is selected from the component assembly trajectory library, ensuring that each trajectory is equipped with a set of label information for indicating the source and type of the trajectory.
[0082] Optionally, each first sample in the first training set may include specified trajectory information, label information, and reference trajectory information, constituting data points required for training the task evaluation network and assembling the model.
[0083] Optionally, a task evaluation network can learn how to evaluate the quality of assembly operations by comparing trial assembly trajectories with reference assembly trajectories. The assembly model, receiving feedback from the task evaluation network, can continuously adjust its strategy to bring the trial assembly trajectories closer and closer to the reference assembly trajectories. After constructing the first training set, the task evaluation network and assembly model can be sequenced to enable the task evaluation network to accurately distinguish and evaluate different types of trajectories. The assembly model, in turn, optimizes and adjusts itself based on this evaluation feedback, gradually improving its actual assembly capabilities.
[0084] In one example, for a memory stick insertion task, the task evaluation network identifies the position accuracy and velocity control of the robot arm's end effector during the trajectory and determines whether the trajectory is similar to a reference assembly trajectory. Based on the task evaluation network's score, the assembly model adjusts its algorithm parameters to reduce insertion deviation and improve insertion speed and stability.
[0085] In one example, m assembly tasks can be randomly selected from the training tasks. The current assembly model and task encoder are used to generate at least one trajectory (i.e., state-action pair) for each task in the real world. The trajectory is then stored in a buffer corresponding to the assembly task. Specifically, for each task, the assembly model is used to obtain a predicted action sequence, which is then used to generate a trajectory for recording.
[0086] This embodiment constructs a first training set and integrates reference and experimental assembly trajectories, enabling the model to learn from a wider range of data and improve its adaptability to new tasks and environments. Furthermore, by introducing a task evaluation network, a quantitative evaluation metric is provided for the experimental assembly trajectories, enabling accurate identification and quantification of the success rate and quality of assembly operations, accelerating the model's iteration and optimization process.
[0087] In an exemplary embodiment, a task evaluation network and a preliminarily trained assembly model are trained based on a first training set, including: taking the first samples in the first training set as the current first samples, respectively, and performing the following training operations on the task evaluation network and the preliminarily trained assembly model: inputting the specified trajectory information in the current first sample and the reference trajectory information in the current first sample into the task evaluation network, and outputting the evaluation value corresponding to the current first sample; adjusting the network parameters of the task evaluation network based on the difference between the evaluation value corresponding to the current first sample and the reference evaluation value indicated by the label information of the current first sample to train the task evaluation network; and training the preliminarily trained assembly model based on the task evaluation network after parameter adjustment.
[0088] It should be noted that designated trajectory information refers to a specific assembly operation trajectory directly output from a rehearsal video or assembly model, while reference trajectory information is the optimal assembly operation trajectory generated based on expert demonstrations. Specifically, the designated trajectory information (which can be a test assembly trajectory or a reference assembly trajectory) and the reference trajectory information in the first sample are simultaneously fed into the task evaluation network. The network calculates and outputs an evaluation value indicating the trajectory's conformance to the optimal assembly trajectory. By inputting different types of trajectory information, the network learns to evaluate the quality of assembly operations, providing a basis for subsequent network parameter adjustments.
[0089] The reference evaluation value can be an evaluation value given manually or by preset rules based on the label information of the first sample, and can be used to represent the ideal evaluation result. By calculating the difference between the evaluation value output by the task evaluation network and the reference evaluation value, and adjusting the parameters of the task evaluation network through the backpropagation algorithm, the network output is made closer to the reference evaluation value. The adjustment of the task evaluation network parameters is to enable the network to more accurately evaluate the quality of the assembly operation. This quality evaluation will be fed back to the assembly model to guide the training of the assembly model.
[0090] Optionally, the task evaluation network and the preliminarily trained assembly model can be trained by specifying a loss function. In one example, taking the IRL loss function as the specified loss function and the Q function as the task evaluation network as an example, the calculation process of the difference between the evaluation value output by the corresponding task evaluation network and the reference evaluation value can be shown as the following formula (1):
[0091] (1)
[0092] in, represents the loss function, represents a state-action pair, Indicates use At the output of the task encoder, The corresponding classification results, that is, label information, for example, In the case of , the indicated state-action pair comes from the reference walkthrough video; In the case of , the state-action pairs are indicated to test the assembly trajectory; It represents a Q function whose input includes state (s), action (a) and task embedding vector (z), and its output is a probability value, which is used to represent the probability that the state-action pair comes from the reference rehearsal video (i.e., the expert) given the task embedding vector. refers to The expected loss function under a large number of samples, such as and The expected loss function under these two strategies is, represents an assembly model (such as a policy network), Represents a policy model from an expert, such as the reference trajectory information in the reference rehearsal video, which can be understood as an expert policy.
[0093] Specifically, the Q-function learning process is cleverly designed to resemble the principles of generative adversarial networks (GANs). By introducing the concepts of a discriminator and a generator, Q-function training is transformed into an adversarial learning framework. In this framework, the Q-function plays the role of the discriminator, while the robot policy (i.e., the trial assembly trajectory) plays the role of the generator. As the discriminator, the Q-function aims to distinguish between state-action pairs as "real samples" from an expert (labeled y(s, a) = 1) and "generated samples" generated by the robot policy (labeled y(s, a) = 0). By minimizing classification error—accurately determining the source of state-action pairs—the robot learns to distinguish between the two. The robot policy, as the generator, aims to make its generated state-action pairs increasingly difficult for the Q-function to distinguish during the optimization process, essentially "deceiving" the Q-function into judging the policy's state-action pairs as good as or even better than those generated by the expert. This process encourages the robot policy to continuously improve, ultimately achieving expert-level assembly performance. By minimizing an inverse reinforcement learning (IRL) loss function, which calculates the error of the Q function in distinguishing between expert actions and robot actions, optimizing this loss allows the Q function to more accurately estimate the reward of a state-action pair, thereby guiding the generator part of the robot's policy to produce higher-quality action sequences in the real world.
[0094] It's important to note that the Q-function learning process is essentially the training of a fully connected network. Its inputs include state information, actions, and the task embedding vector output by the task encoder. Its output is a value in the range [0, 1], representing the probability that the state-action pair originated from an expert. By adjusting network parameters to minimize the IRL loss, the Q-function gradually learns the intrinsic reward structure of expert behavior, providing guidance for the robot's policy generator. This enables the robot to quickly and robustly complete assembly tasks when faced with new tasks without extensive trial and error.
[0095] In another example, during training, the reference trajectory includes the state-action pair "memory fits precisely into the slot," while the robot's trial-and-error trajectory may include the state-action pair "deviating from the slot." Training is performed using an IRL-based loss function: the discriminator classifies "fitting precisely into the slot" as expert behavior (high Q-value) and "deviating from the slot" as robot behavior (low Q-value). As the robot gradually learns and adjusts its position, making it more difficult for the discriminator to distinguish, the Q function increases the Q-value of such behaviors. Ultimately, the Q function not only reflects the expert's optimal behavior but also guides the robot back to the correct path through exploration when it deviates from the expert's trajectory.
[0096] Through the embodiment, the task evaluation network forms a set of quantitative evaluation standards by learning and comparing the reference assembly trajectory and the test assembly trajectory, so as to more accurately evaluate the assembly model and improve the accuracy of model evaluation. Moreover, based on the feedback of the task evaluation network, a joint training mechanism is realized, which can promote the interaction between the task evaluation network and the assembly model, form a closed-loop feedback, accelerate the convergence process of the whole model training, and save the training time.
[0097] In one example embodiment, the assembly model preliminarily trained is trained according to the task evaluation network adjusted in parameters, including: performing a random extraction operation on the first samples in the first training set to form a second training set; taking the second samples in the second training set as current second samples respectively, and performing the following training operation on the assembly model preliminarily trained: inputting the specified trajectory information in the current second sample and the reference trajectory information in the current second sample into the task evaluation network adjusted in parameters to output an evaluation value corresponding to the current second sample; and adjusting the parameters of the assembly model preliminarily trained based on the difference between the evaluation value corresponding to the current second sample and the reference evaluation value indicated by the label information of the current second sample, to train the assembly model preliminarily trained.
[0098] It should be noted that the second training set is a series of second samples randomly extracted from the first training set, contains the specified trajectory information, the reference trajectory information and the label information, and can be used for subsequent model training. Alternatively, a plurality of first samples are randomly selected from the first training set by random extraction to form the second training set.
[0099] The second training set can contain the specified trajectory information and the reference trajectory information of a specific assembly task, and the label information indicating the true source (expert demonstration or test assembly). Alternatively, the second samples in the second training set are taken as current second samples respectively, and the following training operation is performed on the assembly model preliminarily trained: the specified trajectory information and the reference trajectory information in the current second sample are input into the task evaluation network adjusted in parameters to obtain an evaluation value, which represents the pros and cons of the assembly operation. Subsequently, the assembly model adjusts the parameters of the assembly model according to the difference between the evaluation value and the reference evaluation value indicated by the label, and optimizes the assembly strategy.
[0100] Alternatively, the difference between the evaluation value and the reference evaluation value indicated by the label can be calculated using a specified loss function. The specified loss function can be an IRL loss function. The assembly model can be a fully connected network structure, such as a policy network.
[0101] In one example, taking the IRL loss function as an example, the updated task evaluation network is used to extract samples, and the assembly model is updated based on the difference value obtained by the specified loss function. The above sample extraction training operation is repeated until the assembly model converges. Specifically, as shown in the following formula (2):
[0102]
[0103] in, Refers to the output of the assembly model , and its corresponding input is and , is the output of the task encoder, i.e., the embedding vector; is the conditional policy loss function, that is, the loss function for the assembly model.
[0104] Through this embodiment, through the feedback of the task evaluation network, the assembly model can quickly identify deficiencies in its own operations, accurately adjust parameters, and improve the assembly efficiency and success rate in a real environment.
[0105] In an exemplary embodiment, before constructing a first training set based on a component assembly trajectory library, the above method includes: taking reference rehearsal videos from multiple reference rehearsal videos as first reference rehearsal videos, and performing the following first construction operations on the first reference rehearsal videos in turn to obtain multiple reference assembly trajectories corresponding to the multiple reference rehearsal videos: performing feature extraction on the first reference rehearsal video to obtain first reference state-action pair information corresponding to the first reference rehearsal video, and constructing a reference assembly trajectory corresponding to the first reference rehearsal video based on the first reference state-action pair information; obtaining a set of experimental assembly tasks, wherein the experimental assembly task in a set of experimental assembly tasks belongs to one of the component assembly tasks corresponding to multiple reference rehearsal videos; executing a set of experimental assembly tasks according to a pre-trained task encoder and a preliminarily trained assembly model to construct experimental assembly trajectories corresponding to the experimental assembly tasks in a set of experimental assembly tasks; and constructing a component assembly trajectory library based on multiple reference assembly trajectories and the experimental assembly trajectories corresponding to a set of experimental assembly tasks.
[0106] It should be noted that the reference rehearsal video can refer to a demonstration video obtained from an experienced engineer. By analyzing and processing the reference rehearsal video, a series of state-action pairs can be extracted, thereby constructing a reference assembly trajectory. Optionally, a single reference rehearsal video can be used as the analysis object in turn, and feature extraction is performed on the key frames in the single reference rehearsal video to identify the position, posture, and velocity information of the robot end effector, as well as the position and posture information of the components, to form state-action pairs, thereby forming a reference assembly trajectory corresponding to the single reference rehearsal video.
[0107] A trial assembly task can be an assembly operation attempted by a pre-trained assembly model in a real-world environment, thereby collecting actual trajectory data during the model's execution. Alternatively, based on the pre-trained task encoder, the assembly model can assemble various components in a real-world environment and collect their operation trajectories, known as trial assembly trajectories.
[0108] Through this embodiment, by constructing a component assembly trajectory library, it is possible to integrate data from multiple sources, including reference demonstrations and robot experiments, to provide comprehensive training materials for the assembly model and accelerate training. Moreover, through the component assembly trajectory library, the model training data can be enriched, and the initially trained assembly model can better understand and perform various component assembly tasks, thereby improving the performance of the model.
[0109] In an exemplary embodiment, a group of trial assembly tasks are performed based on a pre-trained task encoder and a preliminarily trained assembly model to construct a trial assembly trajectory corresponding to the trial assembly tasks in a group of trial assembly tasks, including: taking a group of trial assembly tasks as current trial assembly tasks in sequence, and performing the following second construction operations respectively to construct the trial assembly trajectory corresponding to the current trial assembly task: inputting the reference trajectory information corresponding to the current trial assembly task and the current state information of the assembly robot into the preliminarily trained assembly model, and outputting the current trial action sequence corresponding to the current trial assembly task; based on the operation information of the assembly robot performing the current trial assembly task through the current trial action sequence, constructing the trial assembly trajectory corresponding to the current trial assembly task.
[0110] It should be noted that the reference trajectory information is the optimal assembly path extracted from expert demonstrations. The assembly robot's current state information, including the position, posture, and speed of the robot's end effector, reflects the robot's real-time status while performing the current task. The test assembly trajectory is the sequence of state changes of the assembly robot during the test assembly task, including changes in the position, posture, and speed of the robot's arm, as well as changes in the position and posture of the components.
[0111] The reference trajectory information corresponding to the current experimental assembly task and the current state information of the assembly robot are input into the assembly model after preliminary training to obtain the operation information for executing the current experimental assembly task through the current experimental action sequence to form the experimental assembly trajectory corresponding to the current experimental assembly task.
[0112] Optionally, the operation information of executing the current test assembly task through the current test action sequence can be obtained through a test rehearsal video. The test rehearsal video can be a video of the test process shot when executing the current test action sequence.
[0113] In an example, if there are n component assembly tasks that need to be trained, a trajectory buffer can be created for each component assembly task, which can be used to store the test assembly trajectory corresponding to the assembly task. Can be the reference assembly trajectory of assembly task i, H is the time length corresponding to the reference assembly trajectory, To represent a state-action pair, you can use It can represent the combination of models (such as policy networks ) is the experimental assembly trajectory of assembly task i output by .
[0114] Through this embodiment, by performing a test assembly task in a real environment, valuable information on environmental factors (such as component materials, external interference, etc.) can be collected, thereby improving the operational capability and robustness in an actual production environment.
[0115] In an exemplary embodiment, based on the operation information of the assembly robot performing the current test assembly task through the current test action sequence, a test assembly trajectory corresponding to the current test assembly task is constructed, including: when the assembly robot performs the current test assembly task based on the current test action sequence, collecting the operation image of the assembly robot performing the current test assembly task; based on the operation image, constructing the test assembly trajectory corresponding to the current test assembly task.
[0116] It should be noted that the operation image may refer to the image corresponding to each operation step taken during the assembly robot performing the assembly task, which may include the status information of the assembly robot and components at each time point, and optionally, may be captured by a high frame rate camera.
[0117] Optionally, when the assembly robot is executing the current trial action sequence given by the assembly model, the camera will track and record every operation detail in the assembly process in real time, forming a series of operation images.
[0118] Image processing can be used to identify and extract the status information of the assembly robot and components from the operation image, and reconstruct this information in chronological order to form a test assembly trajectory.
[0119] Through this embodiment, by real-time acquisition and recording of the operation images of the assembly robot executing the action sequence, it is possible to intuitively understand the dynamic operation state of the robot when processing various component assembly tasks, forming a more complete and high-quality test assembly trajectory.
[0120] In an example embodiment, the method further comprises: training the task encoder and the assembly model in sequence by taking each of the plurality of reference demonstration videos as a second reference demonstration video, to obtain a pre-trained task encoder and a preliminarily trained assembly model; inputting a set of demonstration images corresponding to the second reference demonstration video into the task encoder, to output reference trajectory information corresponding to the second reference demonstration video, wherein the reference trajectory information corresponding to the second reference demonstration video is used to indicate a reference assembly action sequence corresponding to the second reference demonstration video; inputting the reference trajectory information corresponding to the second reference demonstration video and initial state information into the assembly model, to obtain an assembly action sequence corresponding to the second reference demonstration video, wherein the initial state information is used to indicate an initial state of the assembly robot in the second reference demonstration video; and adjusting parameters of the task encoder and parameters of the assembly model according to a difference between the assembly action sequence and the reference assembly action sequence corresponding to the second reference demonstration video, until the task encoder and the assembly model converge.
[0121] It should be noted that the task encoder can be a neural network model that converts an image sequence in a reference demonstration video into a task feature vector (also referred to as a task embedding vector), which contains key information of an assembly action, such as an action type, a target component, an operation sequence, and the like.
[0122] Optionally, a set of demonstration images corresponding to the reference demonstration video is input into the task encoder, and the task encoder converts the set of demonstration images corresponding to the reference demonstration video into reference trajectory information representing an action sequence of the video through its corresponding convolutional layer, pooling layer, and fully connected layer.
[0123] The assembly module can be a deep learning-based control strategy model that can predict an action sequence of a robot arm in an assembly process according to input reference trajectory information and initial state information. The initial state information can include information such as a position and an attitude of the robot arm.
[0124] The assembly action sequence is an action flow of the robot arm in the assembly process output by the assembly model, and the reference assembly action sequence is an optimal action flow recorded in expert demonstration. By comparing the action sequence predicted by the assembly model with the reference assembly action sequence in expert demonstration, differences between the two are found, and parameters of the task encoder and the assembly model are adjusted based on the differences, so as to narrow the gap between the model prediction and the expert standard.
[0125] Specifically, comparing the action sequence predicted by the assembly model with the reference assembly action sequence in expert demonstration can be determined by a preset loss function, such as a behavior cloning loss function, by which the assembly model and the task encoder are trained. Specifically, as shown in the following formula (3):
[0126] (3)
[0127] in, and The state-action pairs extracted from the reference rehearsal video and the task encoder are The output corresponding to the i-th assembly task is ,in is the parameter of CNN network, and the network structure of task encoder is as follows Figure 3 As shown, To assemble a model, for example, a fully connected network structure, whose input is the state information s of the end effector of the robot arm (including position, posture, speed), the output of the task encoder , which constitutes the task condition strategy .
[0128] Through this embodiment, multiple reference rehearsal videos are used for training, so that the assembly model is exposed to a variety of assembly tasks and scenarios, thereby enhancing the generalization ability of the assembly model and task encoder, enabling it to perform reasonable assembly operations more quickly when faced with new parts and new assembly environments.
[0129] Figure 4 is a schematic diagram of a component assembly control method of an assembly robot in this optional example, such as Figure 4 As shown, it can specifically include: pre-training of task encoders and assembly models, value function and model optimization, and new task testing.
[0130] The pre-training process for the task encoder and assembly model can be as follows: Initializing the assembly model and task encoder through behavioral cloning (BC) allows the robot to initially mimic the expert's assembly behavior. Specifically, data demonstrating the expert's assembly of server components is collected to extract sequences of state-action pairs. These state-action pairs are used for supervised learning of the policy and task encoder, optimizing parameters by minimizing the difference (e.g., mean squared error) between the actions predicted by the model and the actions actually performed by the expert.
[0131] The process of value function and model optimization can be as follows: in a real physical environment, using a pre-trained assembly model and task encoder to collect data on robot operations. Specifically, multiple tasks are selected (these tasks have corresponding reference rehearsal videos and task embedding vectors corresponding to each task), and the robot is controlled to try assembly actions in these tasks, and state-action pair sequences are collected to obtain real experience of the robotic arm corresponding to multiple tasks. Combining expert demonstrations and data collected by the robot, a mixed experience pool is formed. Through the mixed experience pool, a task condition Q function (i.e., evaluation function) that can evaluate the value of state and action under given task conditions is learned. The Q function is constructed, and the input includes the environment state (5), the executed action (a), and the task embedding vector (z). Using the data in the mixed experience pool, the Q function (i.e., improved evaluation function) is trained through inverse learning (IRL) so that it can distinguish the value of expert behavior and robot behavior, guiding the optimization of the assembly model. According to the Q function, the robot's accumulated experience in each task is evaluated, and the parameters of the assembly model are adjusted to maximize the expected cumulative reward. The experiment and model parameter optimization are repeated until the assembly model converges, that is, the robot's assembly performance in each task is stable and meets the preset standards.
[0132] The new task testing process involves processing expert demonstration examples (i.e., expert demonstration videos) of the new task, extracting key information, and using the task encoder to generate a new task embedding vector (z). The new task embedding vector and the current environment state are then fed into the optimized assembly model to directly generate action sequences, enabling assembly of the new task without the need for additional trial-and-error learning.
[0133] This optional example uses behavioral cloning and iterative learning in real environments to solve the problems encountered by traditional robot control methods in server component assembly tasks, such as programming burden, low experience efficiency, and weak ability to cope with complex scenarios. Through pre-training, Q-function learning, assembly model optimization, and new task testing, the robot control strategy achieves rapid learning, high robustness, and flexibility, significantly improving the efficiency and quality of assembly, making it suitable for automated and intelligent server production lines.
[0134] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM) / random access memory (RAM), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0135] According to another aspect of the embodiments of the present application, a component assembly control device of an assembly robot is also provided, and the component assembly control device of the assembly robot can be used to implement the component assembly control method of the assembly robot provided in the above-mentioned embodiment, and will not be repeated hereafter. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0136] The embodiment of the present application also provides a component assembly control device for an assembly robot, such as Figure 5 As shown, the device includes:
[0137] An acquisition module 502 is configured to acquire a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task;
[0138] The parsing module 504 is used to parse the specified drill video to obtain specified reference trajectory information, wherein the specified reference trajectory information is used to describe the reference assembly trajectory corresponding to the specified component assembly task;
[0139] An execution module 506 is configured to input the specified reference trajectory information and the state information of the assembly robot into the pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the assembly robot's robotic arm;
[0140] The assembly module 508 is used to control the assembly robot to assemble the specified components according to the predicted assembly actions in the predicted assembly action sequence, so as to perform the specified component assembly task.
[0141] It should be noted that the acquisition module 502 in this embodiment can be used to execute the above step S202, the parsing module 504 in this embodiment can be used to execute the above step S204, the execution module 506 in this embodiment can be used to execute the above step S206, and the assembly module 508 in this embodiment can be used to execute the above step S208.
[0142] Through the embodiments provided by the present application, a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task are obtained, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task; the designated rehearsal video is parsed to obtain designated reference trajectory information, wherein the designated reference trajectory information is used to describe the reference assembly trajectory corresponding to the designated component assembly task; the designated reference trajectory information and the state information of the assembly robot are input into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the assembly robot's robotic arm; according to the predicted assembly action in the predicted assembly action sequence, the assembly robot is controlled to assemble the designated component to perform the designated component assembly task. By adopting the present application, by learning new component assembly tasks from designated rehearsal videos, it is no longer necessary to write complex control programs for each new component, which greatly reduces the complexity of the operation and the pre-preparation time. By specifying rehearsal videos and assembly models, it is possible to quickly generate reasonable assembly strategies when faced with unprecedented parts without the need for additional training, greatly improving the robot's adaptability in complex and changing production environments. This solves the technical problems of cumbersome operation and poor generalization ability of component assembly control methods of assembly robots in related technologies, and improves the operational flexibility and environmental adaptability of assembly robots.
[0143] In an exemplary embodiment, the parsing module 504 is further used to: perform target detection on a specified rehearsal video to extract at least two rehearsal images in the specified rehearsal video; input the at least two rehearsal images into a pre-trained task encoder to obtain specified reference trajectory information output by the task encoder.
[0144] In an exemplary embodiment, the parsing module 504 is further used to: use a target detection model to perform target detection on the video frames in the specified rehearsal video in sequence to obtain target detection results for the video frames in the specified rehearsal video; identify video frames containing target components based on the target detection results of the video frames in the specified rehearsal video, wherein at least two rehearsal images include video frames containing target components, the target components are related components for performing the specified component assembly task, and the target components include specified components.
[0145] In an exemplary embodiment, the component assembly control device of the assembly robot further includes: a parameter determination module for determining image distribution parameters of at least two drill images in a specified drill video, wherein the video time of the specified drill video is divided into a plurality of time periods, and the image distribution parameters are used to describe the distribution of the at least two drill images in different time periods of the plurality of time periods;
[0146] The operation execution module is used to execute an image adding operation on at least two training images or an image reducing operation on at least two training images to update the at least two training images when the image distribution parameters do not meet the specified distribution conditions.
[0147] In an exemplary embodiment, the task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer. The parsing module 504 is further configured to: perform a convolution operation on at least two training images through at least one convolutional layer to obtain at least two training images after the convolution operation; perform pooling processing on the at least two training images after the convolution operation through at least one pooling layer to obtain at least two training images after the pooling processing; and perform integration processing on the at least two training images after the pooling processing through at least one fully connected layer to obtain the specified reference trajectory information output by the task encoder.
[0148] In an exemplary embodiment, the component assembly control device of the assembly robot further includes: a first construction module for constructing a first training set based on a component assembly trajectory library, wherein the component assembly trajectory library is used to record the following trajectories corresponding to the component assembly task: a reference assembly trajectory and a trial assembly trajectory, the trial assembly trajectory is an assembly trajectory obtained by assembling components in a real environment using a preliminarily trained assembly model, the first sample in the first training set includes designated trajectory information, label information and reference trajectory information, in a first sample in the first training set, the label information is used to indicate that the designated trajectory corresponding to the designated trajectory information is a reference assembly trajectory or a trial assembly trajectory, and the reference trajectory information is generated based on a reference rehearsal video corresponding to the component assembly task; a first training module for training the task evaluation network and the preliminarily trained assembly model according to the first training set until the task evaluation network and the assembly model converge, wherein the task evaluation network is an evaluation network that uses the reference assembly trajectory as a reference value to evaluate the trial assembly trajectory.
[0149] In an exemplary embodiment, the first training module is further used to use the first samples in the first training set as the current first samples, and perform the following training operations on the task evaluation network and the assembly model after preliminary training: input the specified trajectory information in the current first sample and the reference trajectory information in the current first sample into the task evaluation network, and output the evaluation value corresponding to the current first sample; based on the difference between the evaluation value corresponding to the current first sample and the reference evaluation value indicated by the label information of the current first sample, adjust the network parameters of the task evaluation network to train the task evaluation network; and train the assembly model after preliminary training based on the task evaluation network with adjusted parameters.
[0150] In an exemplary embodiment, the first training module is further used to: randomly extract the first samples in the first training set to form a second training set; use the second samples in the second training set as the current second samples, and perform the following training operations on the assembly model after preliminary training: input the specified trajectory information in the current second sample and the reference trajectory information in the current second sample into the task evaluation network after parameter adjustment, and output the evaluation value corresponding to the current second sample; based on the difference between the evaluation value corresponding to the current second sample and the reference evaluation value indicated by the label information of the current second sample, adjust the parameters of the assembly model after preliminary training to train the assembly model after preliminary training.
[0151] In an exemplary embodiment, the component assembly control device of the assembly robot further includes:
[0152] The second construction module is used to, before constructing the first training set based on the component assembly trajectory library, use the reference rehearsal videos in multiple reference rehearsal videos as the first reference rehearsal videos, and perform the following first construction operations on the first reference rehearsal videos in turn to obtain multiple reference assembly trajectories corresponding to the multiple reference rehearsal videos: perform feature extraction on the first reference rehearsal video to obtain first reference state-action pair information corresponding to the first reference rehearsal video, so as to construct a reference assembly trajectory corresponding to the first reference rehearsal video based on the first reference state-action pair information; obtain a group of experimental assembly tasks, wherein the experimental assembly task in the group of experimental assembly tasks belongs to one of the component assembly tasks corresponding to the multiple reference rehearsal videos; execute a group of experimental assembly tasks according to the pre-trained task encoder and the preliminarily trained assembly model to construct the experimental assembly trajectory corresponding to the experimental assembly task in the group of experimental assembly tasks; and construct a component assembly trajectory library based on the multiple reference assembly trajectories and the experimental assembly trajectory corresponding to the group of experimental assembly tasks.
[0153] In an example embodiment, the second constructing module is further configured to: in a case where the assembly robot performs the current trial assembly task based on the current trial action sequence, collect an operation image of the assembly robot performing the current trial assembly task; and construct the trial assembly trajectory corresponding to the current trial assembly task based on the operation image.
[0154] In an example embodiment, the second constructing module is further configured to: in a case where the assembly robot performs the current trial assembly task based on the current trial action sequence, collect an operation image of the assembly robot performing the current trial assembly task; and construct the trial assembly trajectory corresponding to the current trial assembly task based on the operation image.
[0155] In an example embodiment, the component assembly control device of the assembly robot further comprises a second training module configured to: sequentially train the task encoder and the assembly model by taking each of a plurality of reference rehearsal videos as a second reference rehearsal video, to obtain a pre-trained task encoder and a preliminarily trained assembly model; input a set of rehearsal images corresponding to the second reference rehearsal video to the task encoder, to output reference trajectory information corresponding to the second reference rehearsal video, wherein the reference trajectory information corresponding to the second reference rehearsal video is used to indicate a reference assembly action sequence corresponding to the second reference rehearsal video; input the reference trajectory information corresponding to the second reference rehearsal video and initial state information to the assembly model, to obtain an assembly action sequence corresponding to the second reference rehearsal video, wherein the initial state information is used to indicate an initial state of the assembly robot in the second reference rehearsal video; and adjust parameters of the task encoder and parameters of the assembly model according to a difference between the assembly action sequence and the reference assembly action sequence corresponding to the second reference rehearsal video, until the task encoder and the assembly model converge.
[0156] The description of the features in the embodiments of the component assembly control device of the assembly robot can be referred to the related description of the embodiments of the component assembly control method of the assembly robot, which will not be repeated here.
[0157] Embodiments of the present application also provide an electronic device comprising a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to perform the steps in any of the above-described embodiments of the component assembly control method of the assembly robot.
[0158] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the component assembly control method of the assembly robot when running.
[0159] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0160] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned embodiments of the component assembly control method of the assembly robot are implemented.
[0161] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the component assembly control method of the assembly robot.
[0162] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0163] The above is a detailed introduction to the component assembly control method, electronic device and storage medium of an assembly robot provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A component assembly control method for an assembly robot, characterized in that: include: Obtaining a designated component assembly task and a designated rehearsal video corresponding to the designated component assembly task, wherein the designated component assembly task is an assembly task performed on a designated component that has not been performed by the assembly robot, and the designated rehearsal video is used to record the assembly scene of the designated component assembly task; Parsing the designated drill video to obtain designated reference trajectory information, wherein the designated reference trajectory information is used to describe a reference assembly trajectory corresponding to the designated component assembly task; Inputting the specified reference trajectory information and state information of the assembly robot into a pre-trained assembly model to obtain a predicted assembly action sequence output by the assembly model, wherein the state information of the assembly robot is used to indicate the state of the end effector of the robotic arm of the assembly robot; According to the predicted assembly action in the predicted assembly action sequence, the assembly robot is controlled to assemble the designated component to perform the designated component assembly task.
2. The method according to claim 1, characterized in that The step of parsing the designated drill video to obtain designated reference trajectory information includes: Performing object detection on the designated drill video to extract at least two drill images from the designated drill video; The at least two training images are input into a pre-trained task encoder to obtain the specified reference trajectory information output by the task encoder.
3. The method according to claim 2, characterized in that The performing target detection on the designated drill video to extract at least two drill images from the designated drill video includes: Using the target detection model, sequentially performing target detection on the video frames in the specified drill video to obtain target detection results for the video frames in the specified drill video; Based on the target detection results of the video frames in the specified rehearsal video, the video frames containing the target components are identified, wherein the at least two rehearsal images include video frames containing the target components, the target components are related components for performing the specified component assembly task, and the target components include the specified components.
4. The method according to claim 2, characterized in that After performing target detection on the designated drill video to extract at least two drill images from the designated drill video, the method further includes: determining image distribution parameters of the at least two drill images in the designated drill video, wherein the video time of the designated drill video is divided into a plurality of time periods, and the image distribution parameters are used to describe distribution of the at least two drill images in different time periods among the plurality of time periods; In a case where the image distribution parameter does not satisfy a specified distribution condition, an image adding operation is performed on the at least two training images, or an image reducing operation is performed on the at least two training images to update the at least two training images.
5. The method according to claim 2, characterized in that The task encoder includes at least one convolutional layer, at least one pooling layer, and at least one fully connected layer; Inputting the at least two training images into a pre-trained task encoder to obtain the designated reference trajectory information output by the task encoder includes: Performing a convolution operation on the at least two training images through the at least one convolution layer to obtain the at least two training images after the convolution operation; performing pooling processing on the at least two training images after the convolution operation through the at least one pooling layer to obtain the at least two training images after the pooling processing; The at least two training images after pooling are integrated through the at least one fully connected layer to obtain the specified reference trajectory information output by the task encoder.
6. The method according to claim 2, characterized in that The method further comprises: A first training set is constructed based on a component assembly trajectory library, wherein the component assembly trajectory library is used to record the following trajectories corresponding to the component assembly task: the reference assembly trajectory and the experimental assembly trajectory, the experimental assembly trajectory being an assembly trajectory obtained by assembling components in a real environment using the preliminarily trained assembly model, the first sample in the first training set including designated trajectory information, label information, and reference trajectory information, in a first sample in the first training set, the label information is used to indicate that the designated trajectory corresponding to the designated trajectory information is the reference assembly trajectory or the experimental assembly trajectory, and the reference trajectory information is generated based on a reference rehearsal video corresponding to the component assembly task; According to the first training set, the task evaluation network and the preliminarily trained assembly model are trained until the task evaluation network and the assembly model converge, wherein the task evaluation network is an evaluation network that evaluates the test assembly trajectory using the reference assembly trajectory as a reference value.
7. The method according to claim 6, characterized in that The step of training the task evaluation network and the preliminarily trained assembly model according to the first training set includes: The first samples in the first training set are respectively used as current first samples, and the following training operations are performed on the task evaluation network and the preliminarily trained assembly model: Inputting the designated trajectory information in the current first sample and the reference trajectory information in the current first sample into the task evaluation network, and outputting an evaluation value corresponding to the current first sample; Adjusting network parameters of the task evaluation network based on a difference between an evaluation value corresponding to the current first sample and a reference evaluation value indicated by the label information of the current first sample to train the task evaluation network; The assembly model after preliminary training is trained based on the task evaluation network after parameter adjustment.
8. The method according to claim 7, characterized in that The step of training the preliminarily trained assembly model based on the task evaluation network after parameter adjustment includes: Performing a random sampling operation on the first sample in the first training set to form a second training set; The second samples in the second training set are respectively used as the current second samples, and the following training operations are performed on the assembled model after preliminary training: Inputting the designated trajectory information in the current second sample and the reference trajectory information in the current second sample into the task evaluation network after parameter adjustment, and outputting an evaluation value corresponding to the current second sample; Based on the difference between the evaluation value corresponding to the current second sample and the reference evaluation value indicated by the label information of the current second sample, the parameters of the assembly model after preliminary training are adjusted to train the assembly model after preliminary training.
9. The method according to claim 6, characterized in that Before constructing the first training set based on the component assembly trajectory library, the method includes: A reference rehearsal video from the plurality of reference rehearsal videos is respectively used as a first reference rehearsal video, and the following first construction operation is performed on the first reference rehearsal videos in sequence to obtain a plurality of reference assembly trajectories corresponding to the plurality of reference rehearsal videos: performing feature extraction on the first reference rehearsal video to obtain first reference state-action pair information corresponding to the first reference rehearsal video, and constructing a reference assembly trajectory corresponding to the first reference rehearsal video based on the first reference state-action pair information; Acquire a set of test assembly tasks, wherein the test assembly tasks in the set of test assembly tasks belong to one of the component assembly tasks corresponding to the plurality of reference drill videos; executing the set of trial assembly tasks according to the pre-trained task encoder and the preliminarily trained assembly model to construct trial assembly trajectories corresponding to the trial assembly tasks in the set of trial assembly tasks; The component assembly trajectory library is constructed according to the multiple reference assembly trajectories and the test assembly trajectories corresponding to the set of test assembly tasks.
10. The method according to claim 9, characterized in that The step of executing the set of trial assembly tasks according to the pre-trained task encoder and the preliminarily trained assembly model to construct a trial assembly trajectory corresponding to the trial assembly task in the set of trial assembly tasks includes: The set of test assembly tasks is sequentially used as the current test assembly task, and the following second construction operations are respectively performed to construct the test assembly trajectory corresponding to the current test assembly task: Inputting the reference trajectory information corresponding to the current test assembly task and the current state information of the assembly robot into the preliminarily trained assembly model, and outputting the current test action sequence corresponding to the current test assembly task; Based on the operation information of the assembly robot performing the current test assembly task through the current test action sequence, a test assembly trajectory corresponding to the current test assembly task is constructed.
11. The method according to claim 10, characterized in that The constructing a test assembly trajectory corresponding to the current test assembly task based on the operation information of the assembly robot performing the current test assembly task through the current test action sequence includes: When the assembly robot performs the current test assembly task based on the current test action sequence, collecting an operation image of the assembly robot performing the current test assembly task; Based on the operation image, a test assembly trajectory corresponding to the current test assembly task is constructed.
12. The method according to claim 9, characterized in that The method further comprises: Using reference rehearsal videos from a plurality of reference rehearsal videos as second reference rehearsal videos, the task encoder and the assembly model are trained in sequence to obtain the pre-trained task encoder and the preliminarily trained assembly model: Inputting a set of rehearsal images corresponding to the second reference rehearsal video into the task encoder, and outputting reference trajectory information corresponding to the second reference rehearsal video, wherein the reference trajectory information corresponding to the second reference rehearsal video is used to indicate a reference assembly action sequence corresponding to the second reference rehearsal video; Inputting reference trajectory information and initial state information corresponding to the second reference rehearsal video into the assembly model to obtain an assembly action sequence corresponding to the second reference rehearsal video, wherein the initial state information is used to indicate the initial state of the assembly robot in the second reference rehearsal video; According to the difference between the assembly action sequence and the reference assembly action sequence corresponding to the second reference rehearsal video, the parameters of the task encoder and the parameters of the assembly model are adjusted until the task encoder and the assembly model converge.
13. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the component assembly control method of the assembly robot as claimed in any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the component assembly control method of the assembly robot according to any one of claims 1 to 12 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the component assembly control method of the assembly robot according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Moving object rapid detection method based on video sequence
CN106846359A
Micro-nano robot assembly track learning method based on dynamic motion primitives
CN114571458A
Video action scoring method, computer readable storage medium and system
CN116311536A
Teaching-based electronic assembly control method, system and equipment and storage medium
CN118617402A
3C assembly programming-free method based on large model action analysis and automatic programming
CN119027618A