An assembly method, system, electronic device, and storage medium
Patent Information
- Application Number
- CN202610502147.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-16
- Publication Date
- 2026-08-18
AI Technical Summary
然而,当装配对象规格、装配工位位置或作业环境发生变化时,预编的控制程序难以适配新的装配工况,需技术人员重新调整参数并编写程序,不仅增加了工程维护成本,还大幅降低了装配效率及装配设备对不同装配任务的自适应能力
[0015]The assembly method, system, electronic device, and storage medium of this application receive an assembly instruction, which includes an assembly task, comprising an assembly station and an assembly method for the component to be assembled; in response to the assembly instruction, the assembly equipment, whose end effector is holding the component to be assembled, moves to an initial pose corresponding to the assembly station, and establishes a local task coordinate system with the initial pose as the origin; acquires the end effector's end-effector mechanical data and end-effector pose data in the local task coordinate system; based on the end effector mechanical data and end-effector pose data, uses a pre-trained learning strategy model to determine the motion increment data of the end effector in the local task coordinate system; and generates motion control instructions based on the motion increment data, which control the end effector to assemble the component to be assembled to the assembly station according to the assembly method. This eliminates the need to pre-write control logic or state switching programs for specific assembly objects and working conditions, but instead, through an end-to-end learning strategy model, directly outputs motion increment instructions based on the current mechanical and pose states, autonomously deciding on the next movement. When the specifications, installation location, or working environment of the assembled object change, the model can adaptively adjust the end effector based on the current perception information without manual reprogramming or debugging. This effectively solves the problems of low assembly efficiency and insufficient adaptability to different tasks caused by repeated program adjustments due to changes in working conditions.
Smart Images

Figure CN122593155A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control and intelligent assembly technology, and in particular to an assembly method, system, electronic device and storage medium. Background Technology
[0002] In the field of industrial automation assembly, for assembly tasks involving plugging and mating, such as plugs and sockets, pins and sockets, and pluggable connectors, technicians typically need to pre-write control programs based on the specific assembly object and working conditions to guide the assembly equipment to automatically complete assembly actions such as hole alignment, insertion, and mounting. However, when the specifications of the assembly object, the location of the assembly station, or the working environment changes, the pre-written control programs become difficult to adapt to the new assembly conditions. Technicians need to readjust the parameters and rewrite the program, which not only increases engineering maintenance costs but also significantly reduces assembly efficiency and the assembly equipment's adaptability to different assembly tasks. Summary of the Invention
[0003] This application provides an assembly method, system, electronic device, and storage medium to at least solve the above-mentioned technical problems existing in the prior art.
[0004] According to a first aspect of this application, an assembly method is provided, the method comprising:
[0005] Receive assembly instructions, the assembly instructions including assembly tasks, the assembly tasks including the assembly station and assembly method of the parts to be assembled; In response to the assembly command, the control end of the assembly equipment, which has clamped the part to be assembled, moves to an initial pose corresponding to the assembly station, and establishes a local task coordinate system with the initial pose as the origin. Acquire the end-effector mechanical data and the end-effector pose data in the local task coordinate system; Based on the end-effector mechanical data and the end-effector pose data, a pre-trained learning-based policy model is used to determine the motion increment data of the end-effector in the local task coordinate system. Based on the incremental motion data, motion control commands are generated. These commands are used to control the end effector to assemble the parts to be assembled to the assembly station according to the assembly method.
[0006] In one possible implementation, determining the motion increment data of the end effector in the local task coordinate system based on the end effector mechanical data and the end effector pose data using a pre-trained learning-based policy model includes: The end-effector mechanical data and the end-effector pose data are combined into a state vector; The state vector is input into the learning policy model to obtain motion increment data output by the learning policy model. The motion increment data includes end-effector translation increment or end-effector translation increment and attitude increment in the local task coordinate system.
[0007] In one possible implementation, the learned policy model is trained in the following manner: Acquire multiple assembly demonstration trajectory data collected through assembly demonstrations, wherein the assembly demonstration trajectory data is a normal assembly trajectory or a specific assembly trajectory that includes contact loss and recovery processes; Training samples are extracted from the assembly demonstration trajectory data. Each training sample includes a sample state vector consisting of sample end pose data and corresponding sample end mechanical data, and sample motion increment data corresponding to the sample state vector. Using the sample state vector in the training samples as input and the sample motion increment data in the training samples as output, the model used to output the motion increment data is trained by imitation learning and / or reinforcement learning to obtain the learning policy model.
[0008] In one possible implementation, the end-effector pose data includes the relative pose of the end-effector with respect to the local task coordinate system at the current moment; the end-effector mechanical data includes multiple forces and moments arranged in time sequence; correspondingly, the method further includes: Detect whether the relative pose exceeds the preset search range based on the initial pose, and / or detect whether the force or torque in the force and torque exceeds their respective safety thresholds; An error message is generated in response to the relative pose exceeding the search range, and / or the force or torque in the force and torque exceeding the corresponding safety threshold.
[0009] In one possible implementation, the assembly equipment holding the part to be assembled by the control terminal moves to an initial position corresponding to the assembly station, including: Acquire an image of the assembly station; Feature extraction is performed on the image to obtain the visual features of the assembly station; Based on the visual features, calculate the pose error between the current end pose and the initial pose; Based on the pose error, the end effector is controlled to move to the initial pose.
[0010] In one possible implementation, the assembly task is a plug-in and mating assembly task with contact constraints and guiding features, including at least one of the following: plug and socket insertion, pin and socket mating, connector insertion, and snap-fit engagement.
[0011] According to a second aspect of this application, an assembly system is provided, the system comprising: An assembly device, the assembly device including a robotic arm; An end effector, mounted at the end of the robotic arm, is used to grip the part to be assembled; An end-effector force and torque sensor is installed between the robotic arm flange and the end effector to collect end-effector force and torque data. The controller, connected to the assembly equipment and the end force and torque sensor, is used to execute the method described in this application.
[0012] In one possible implementation, the system further includes: A teleoperation device is used to allow users to control the robotic arm to perform assembly demonstrations and generate assembly demonstration trajectory data during offline teaching and online training phases. The teleoperation device is at least one of the following: a teaching robot arm, a three-dimensional force feedback handle, a three-dimensional mouse, a six-degree-of-freedom spatial input device, a joystick, or a virtual reality / augmented reality controller, which may be isomorphic or heteromorphic to the robotic arm.
[0013] According to a third aspect of this application, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in this application.
[0014] According to a fourth aspect of this application, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described in this application.
[0015] The assembly method, system, electronic device, and storage medium of this application receive an assembly instruction, which includes an assembly task, comprising an assembly station and an assembly method for the component to be assembled; in response to the assembly instruction, the assembly equipment, whose end effector is holding the component to be assembled, moves to an initial pose corresponding to the assembly station, and establishes a local task coordinate system with the initial pose as the origin; acquires the end effector's end-effector mechanical data and end-effector pose data in the local task coordinate system; based on the end effector mechanical data and end-effector pose data, uses a pre-trained learning strategy model to determine the motion increment data of the end effector in the local task coordinate system; and generates motion control instructions based on the motion increment data, which control the end effector to assemble the component to be assembled to the assembly station according to the assembly method. This eliminates the need to pre-write control logic or state switching programs for specific assembly objects and working conditions, but instead, through an end-to-end learning strategy model, directly outputs motion increment instructions based on the current mechanical and pose states, autonomously deciding on the next movement. When the specifications, installation location, or working environment of the assembled object change, the model can adaptively adjust the end effector based on the current perception information without manual reprogramming or debugging. This effectively solves the problems of low assembly efficiency and insufficient adaptability to different tasks caused by repeated program adjustments due to changes in working conditions.
[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0017] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of this application are illustrated in the drawings by way of example and not limitation, in which: In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0018] Figure 1 A schematic diagram illustrating the implementation flow of the assembly method provided in an embodiment of this application is shown; Figure 2 A schematic diagram illustrating the implementation flow of the motion increment determination operation of the assembly method provided in this application embodiment is shown; Figure 3 A structural block diagram of the assembly system provided in an embodiment of this application is shown; Figure 4 A schematic diagram of the composition structure of an electronic device according to an embodiment of this application is shown. Detailed Implementation
[0019] To make the objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] First, the application scenario of this application will be described. Currently, for plug-in and mating assembly tasks such as plugs and sockets, pins and sockets, and pluggable connectors, the most widely used traditional plug-in and mating assembly control method is the rigid-flexible switching / impedance control method based on force-position control. This method requires installing force / torque sensors at the end of the robotic arm, first establishing the geometric and contact model between the end and the environment, and then designing control laws such as impedance control, rigid-flexible switching, and force / position composite control. During the actual plugging process, the position and attitude of the end are adjusted in real time according to the force applied to the end, thereby realizing the introduction of the parts to be assembled, hole alignment, and anti-jamming operation. However, this method has obvious shortcomings. On the one hand, it requires manual establishment of the contact model and completion of parameter debugging. When facing different socket structures and different robot installation postures, a lot of engineering adjustment work is required, resulting in poor adaptability. On the other hand, its control logic usually relies on manually written multi-stage state machines. Once the assembly conditions change or abnormal contact occurs, the maintenance of the state machine becomes very complex, and the overall robustness of the control is limited.
[0021] Besides the force-position control method mentioned above, visual servo-based plug-in mating control is also a common approach in traditional assembly control. This method acquires image information of the plug-in hole using an end-effector camera or an external camera, extracts feature points or contour features from the image, and then uses visual servo technology to reduce pose errors in the image plane, thereby guiding the robotic arm to gradually approach and complete the insertion action. The shortcomings of this method are mainly twofold: first, it has stringent requirements for the visual conditions of the external environment, heavily relying on stable lighting, color, and texture features. When the color or material of the plug-in hole changes, or the assembly background changes, the control effect of this method will significantly decrease; second, its ability to handle complex contact processes is limited. It is difficult to achieve precise control when faced with slight jamming, friction, or self-alignment during assembly. In practical applications, it still needs to be combined with a manual state machine or force threshold logic to complete the assembly, resulting in insufficient autonomy and adaptability of control. Overall, traditional plug-in and mating assembly control methods are based on analytical control laws combined with manual state machines as their core architecture, relying excessively on the experience of technicians for modeling and parameter tuning. When assembly conditions change, the generalization and adaptability of these methods are significantly insufficient.
[0022] In recent years, some research has attempted to use learning-based methods to solve precision assembly tasks involving insertion and mating. Typical implementations of these methods include two approaches: one involves inputting the robot's body state (including joint angles, joint velocities, end-effector poses, etc.) and image information from multiple cameras into a single policy network, using reinforcement learning or imitation learning to allow the network to directly learn the mapping relationship from visual / body state to assembly actions; the other involves introducing a human-in-the-loop policy, accelerating the training process of the policy network and improving its control accuracy through real-time human intervention and reward-based learning. Compared to traditional control methods, these learning-based methods possess stronger expressive power in complex nonlinear contact scenarios and can better adapt to complex force and pose changes during assembly. However, they still have many significant shortcomings in practical industrial applications.
[0023] First, learning-based assembly control methods are strongly coupled with the robot's body state, making cross-platform migration difficult. The policy network of such methods usually takes the robot's joint angles, joint velocities, and other body states as input. During training, the network is prone to overfitting, and what is ultimately learned is only the insertion force control behavior under a specific robot and a specific installation posture. When the robot platform is changed or the robot's initial joint posture is adjusted, the distribution of the input data of the policy network will change significantly. The trained policy often cannot be directly reused and needs to be retrained or undergo large-scale parameter fine-tuning, which increases the cost and difficulty of industrial applications.
[0024] Secondly, this type of method also suffers from overfitting to the appearance of objects, resulting in poor generalization. During the training phase of the policy network, only a certain color, material, and background of the socket are usually used as training samples. The network will implicitly memorize the visual distribution characteristics of this type of socket during the training process. However, in the actual inference application phase, when the socket with a significantly different appearance is replaced, or when the lighting and background of the assembly scene change, the distribution of the collected visual features will change significantly, directly leading to a significant decrease in the control performance of the policy network, making it unable to adapt to assembly objects with different appearances and different assembly environments.
[0025] Finally, existing learning-based assembly control methods are insufficient in their ability to recover from lost contacts. Most of the research and training of these methods focus on the insertion process of "continuous contact". When abnormal situations such as slippage or complete loss of contact occur during the insertion process, there is a lack of a unified and robust learning-based recovery mechanism for finding the socket again and restoring the assembly in a local area. In practical applications, traditional state machines or hard-coded recovery logic are usually still required in high-level control to handle abnormal situations, thus failing to achieve end-to-end autonomous assembly control.
[0026] In summary, while existing learning-based plugging and mating control methods have theoretical advantages, they still have significant shortcomings in cross-robot platform compatibility, scenario generalization, and contact loss recovery, limiting practical industrial applications. Therefore, to address these technical problems, this application provides an assembly method, system, electronic device, and storage medium.
[0027] Figure 1 A schematic diagram illustrating the implementation flow of the assembly method provided in the embodiments of this application is shown.
[0028] refer to Figure 1 This application provides an assembly method, the method comprising: Operation 101: Receive assembly instructions. Assembly instructions include assembly tasks, which include the assembly station and assembly method of the parts to be assembled.
[0029] The assembly method described in this application can be applied to various assembly tasks requiring precise alignment and contact assembly. Before starting the assembly operation, an assembly instruction is first received. This assembly instruction can be triggered by an upper-level control system (such as a production management computer, programmable logic controller, or manual control console) based on the production task. The assembly instruction carries key information for this assembly task, specifically including the assembly station of the parts to be assembled and the assembly method. The parts to be assembled can be various workpieces that require insertion or mating, such as plugs, pins, connectors, or clips; correspondingly, the assembly station can be a socket, socket, interface, or slot; the assembly method can be insertion, pressing, or snap-fit.
[0030] In one embodiment of this application, the assembly task is a plug-in and mating assembly task with contact constraints and guiding features, including at least one of the following: plug and socket insertion, pin and socket mating, connector insertion, and snap-fit engagement. It should be noted that the assembly task in this application is not limited to these; the assembly task in this application can be a variety of plug-in and mating assembly tasks with contact constraints and guiding features.
[0031] Operation 102, in response to the assembly command, controls the assembly equipment with the parts to be assembled already clamped at the end to move to the initial pose corresponding to the assembly station, and establishes a local task coordinate system with the initial pose as the origin.
[0032] Upon receiving an assembly command, the assembly equipment must first move its end effector (the part holding the component to be assembled) to an initial pose corresponding to the assembly station. This process is typically called coarse alignment, and its purpose is to roughly align the end effector with the assembly station in space, establishing a starting position for subsequent fine insertion. Once the end effector reaches this initial pose, a local task coordinate system is established with this pose as the origin. This coordinate system describes the local movement of the end effector relative to the assembly starting point, constraining subsequent assembly actions within a local range centered on this origin, facilitating fine control of the assembly process.
[0033] Operation 103: Obtain the end-effector mechanical data and the end-effector pose data in the local task coordinate system.
[0034] During the assembly process initiated by the end effector from its initial pose, it is necessary to perceive the current state in real time. This requires acquiring two types of data in real time: first, end effector mechanical data, reflecting the forces and torques acting on the end effector during assembly, used to perceive the contact state with the assembly station; and second, end effector pose data in the local task coordinate system, reflecting the position and attitude changes of the end effector relative to the assembly start point. These two types of data together form the basis for subsequent decision-making.
[0035] Operation 104, based on end-effector mechanical data and end-effector pose data, uses a pre-trained learning-based policy model to determine the incremental motion data of the end-effector in the local task coordinate system.
[0036] Based on real-time acquired end-effector mechanical and pose data, these are input into a pre-trained learning-based policy model. After training, this model can output incremental motion data that the end-effector should execute in the local task coordinate system, based on the current end-effector mechanical and pose data. This process achieves a mapping from perception information to motion decisions without requiring manual pre-setting of specific contact models or state switching logic.
[0037] Motion increment data can be understood as minute adjustments used to guide the next movement of the end effector. This increment data reflects the fine movements that should be performed in the current contact state, such as a small translation along a certain direction or a small rotation around a certain axis, so that the end effector can gradually adapt to changes in the contact state and complete actions such as insertion, alignment, and mounting. By continuously outputting and accumulating these motion increments, continuous and compliant control of the assembly process is achieved.
[0038] Operation 105: Based on the incremental motion data, generate motion control commands. The motion control commands are used to control the end effector to assemble the parts to be assembled to the assembly station according to the assembly method.
[0039] Based on the motion increment data output by the learning strategy model, corresponding motion control commands are generated. These commands can be viewed as Cartesian spatial position control commands or velocity control commands. These commands drive the end effector of the assembly equipment to perform precise position and attitude adjustments within the local task coordinate system according to the specified assembly method, gradually importing and assembling the parts to be assembled at the assembly station until the entire assembly task is completed. Through this cycle, end-to-end control of the plug-in and mating assembly processes is achieved.
[0040] Thus, this embodiment acquires end-effector mechanical data and end-effector pose data in the local coordinate system in real time, outputs incremental motion data using a pre-trained learning strategy model, and generates control commands accordingly to complete assembly. There is no need to pre-write control logic or state switching programs for specific assembly objects and working conditions. Instead, through an end-to-end learning strategy model, it directly outputs incremental motion commands based on the current mechanical and pose states, autonomously deciding on the next movement. When the specifications of the assembly object, its installation location, or the working environment changes, the model can adaptively adjust the end-effector action based on the current perception information, without requiring manual reprogramming or debugging. This effectively solves the problems of low assembly efficiency and insufficient adaptability to different tasks caused by repeated program adjustments due to changes in working conditions.
[0041] Figure 2 The diagram illustrates the implementation flow of the motion increment determination operation of the assembly method provided in this application embodiment.
[0042] refer to Figure 2 In one embodiment of this application, the above-described operation 104, based on the end-effector mechanical data and end-effector pose data, utilizes a pre-trained learning-based policy model to determine the motion increment data of the end-effector in the local task coordinate system, including: Operation 201 combines the end-effector mechanical data and end-effector pose data into a state vector; Operation 202: Input the state vector into the learning policy model to obtain the motion increment data output by the learning policy model. The motion increment data includes the end translation increment or the end translation increment and attitude increment in the local task coordinate system.
[0043] Specifically, the end-effector's mechanical data and pose data are combined to form the state vector of the learning-based policy model. This state vector integrates the current force state and spatial position information of the end-effector, providing the policy model with multi-dimensional perceptual inputs required for decision-making. Furthermore, this state vector does not contain ontological state information such as robot joint angles and velocities, nor does it include scene image information. This allows the model to focus on the interaction between the end-effector and its environment during the learning process, rather than the specific robot configuration or the appearance features of the scene. This enables it to adapt to different robot platforms and operational scenarios during deployment, solving the problem of policy performance degradation due to changes in the robot's physical characteristics or environment.
[0044] Furthermore, the state vector is input into the learning policy model to obtain motion increment data output by the learning policy model. The motion increment data includes: end-effector translation increments in the local task coordinate system (e.g., fine-tuning displacements along the X, Y, and Z directions), used to control the end-effector's approach to or insertion into the assembly station; and optional attitude increments (e.g., small angular adjustments around the insertion axis or other coordinate axes), used to adjust the angular deviation between the end-effector and the assembly station. By continuously outputting such increments, the end-effector can gradually approach the assembly station based on its current perception state.
[0045] In one embodiment of this application, the learning strategy model is trained in the following manner: acquiring multiple assembly demonstration trajectory data collected through assembly demonstrations, wherein the assembly demonstration trajectory data are normal assembly trajectories or specific assembly trajectories containing contact loss and recovery processes; extracting training samples from the assembly demonstration trajectory data, wherein each training sample includes a sample state vector composed of sample end pose data and corresponding sample end mechanical data, and sample motion increment data corresponding to the sample state vector; using the sample state vector in the training samples as input and the sample motion increment data in the training samples as output, performing imitation learning and / or reinforcement learning training on the model used to output motion increment data to obtain the learning strategy model.
[0046] To achieve autonomous control of the assembly process, this application embodiment pre-constructs a learning strategy model as an end-to-end control strategy. This learning strategy model is trained to take a state vector as input and the end-effector motion increment (i.e., motion increment data) in the local task coordinate system as output, thereby directly deciding the next action based on the real-time perceived mechanical and pose states during the assembly process.
[0047] In one embodiment of this application, the learning policy model is a neural network policy model, also known as a policy network, which may include one or more temporal feature extraction layers, such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or transformer networks, and may further include one or more fully connected layers. This application does not limit the specific network structure type or parameter configuration, as long as it can achieve the mapping from state vectors to motion increment data.
[0048] The training process of this learning-based strategy model may include: First, multiple assembly demonstration trajectory data are collected through assembly demonstrations. In one embodiment of this application, a remote operation method can be used, allowing a human operator to control the assembly equipment to complete multiple insertion task demonstrations. During the demonstration, the operator not only demonstrates normal, slip-free insertion trajectories but also deliberately demonstrates trajectories where slippage or loss of contact occurs during insertion, followed by re-insertion through local search. In this way, the collected trajectory data can cover normal operating conditions and abnormal recovery conditions that may occur during the assembly process. During the demonstration, relevant data are recorded simultaneously, including end-effector mechanical data, end-effector pose data, and corresponding motion increment data.
[0049] After data acquisition, training samples are extracted from the assembly demonstration trajectory data. For each time step in the trajectory, the current end-effector pose data and end-effector mechanical data are combined to form a sample state vector. Simultaneously, the next motion action corresponding to that time step, i.e., the sample motion increment data, is used as the label or target output for that sample. In this way, a large number of "state-action" pairs are formed, serving as the basis for subsequent training.
[0050] After obtaining training samples, the model is trained using the sample state vector as input and the sample motion increment data as output. This allows the model to gradually learn which motion increments should be executed under different end-effector poses and mechanical states. The training continues until the model's success rate reaches a predetermined threshold across various socket appearances, different initial deviations, and multiple robot platforms. At this point, the model parameters are frozen, resulting in the final learning-based policy model. The model training can employ one or more of the following methods in combination: The model is initially trained using imitation learning (such as behavior cloning). That is, through supervised learning, the model learns to imitate the action mapping relationships in human demonstrations, enabling it to initially possess basic insertion and simple recovery capabilities.
[0051] Building upon imitation learning, reinforcement learning is further employed to optimize the model. For example, offline-collected teaching data can be used for policy initialization and / or training data buffering. Then, in a simulated or real robotic environment, a policy optimization-based reinforcement learning algorithm (such as PPO, SAC, etc.) can be used to train the policy network. During training, a sparse reward function can be designed, such as providing a positive reward for successful insertion and a negative reward for insertion force or torque exceeding a safety threshold. Simultaneously, initial pose deviations, contact disturbances, or simulated slippage conditions can be randomly introduced during training, enabling the policy network to learn reasonable behaviors in both contact and non-contact states. Ultimately, this achieves a continuous control process from contact to non-contact and then to re-establishing contact.
[0052] In one embodiment of this application, the end-effector pose data includes the relative pose of the end-effector relative to the local task coordinate system at the current moment; the end-effector mechanical data includes multiple forces and torques arranged in time sequence; correspondingly, the method further includes: detecting whether the relative pose exceeds a search range preset based on the initial pose, and / or detecting whether the force or torque in the force and torque exceeds its corresponding safety threshold; in response to the relative pose exceeding the search range, and / or the force or torque in the force and torque exceeding the corresponding safety threshold, generating error information.
[0053] Specifically, during the assembly process, to ensure the safety of equipment and workpieces and avoid damage or jamming due to abnormal contact, this application embodiment includes a safety constraint mechanism. During assembly, boundary constraints are applied to the relative pose of the end effector in the local task coordinate system. That is, an allowable search range is preset based on the initial pose, such as limiting the maximum range of relative displacement of the end effector. At the same time, corresponding safety thresholds are set for the forces and torques experienced by the end effector, namely the maximum allowable values of insertion force and torque.
[0054] Furthermore, based on the aforementioned safety constraints, during the assembly process, the system continuously monitors whether the current end effector's relative pose exceeds the preset search range, and whether the current force and torque exceed their respective safety thresholds. If the relative pose exceeds the search range, or the force / torque exceeds the safety threshold, it indicates that the current assembly process is in an abnormal state, and an error message is generated. This error message can be used to trigger subsequent processing, such as pausing the current action, terminating task execution, or controlling the end effector to perform a safety rollback operation to prevent equipment damage or workpiece scrapping, thereby improving the safety and reliability of the assembly process.
[0055] In one embodiment of this application, the above operation 102, which controls the assembly equipment with the end effector holding the part to be assembled to move to the initial pose corresponding to the assembly station, includes: acquiring an image of the assembly station; extracting features from the image to obtain the visual features of the assembly station; calculating the pose error between the current end effector pose and the initial pose based on the visual features; and controlling the end effector to move to the initial pose according to the pose error.
[0056] This application embodiment achieves coarse alignment via visual servoing. Specifically, images of the assembly station are acquired using an end-effector camera or an external global camera. Feature extraction is performed on the images to obtain visual features reflecting the position and orientation of the assembly station, such as the outline of the socket, feature points, or marker points. Based on the extracted visual features and the current end-effector pose, the pose error between the current end-effector pose and the target initial pose is calculated. According to this pose error, corresponding motion control commands are generated to drive the end-effector to gradually move to the initial pose, completing the coarse alignment.
[0057] It should be noted that the above-described visual servoing coarse alignment method is only one implementation of this application, and not the only one. In other embodiments of this application, coarse alignment can also be implemented in other ways, such as pre-determining the initial pose through manual teaching, or directly learning the coarse alignment strategy from the image using a learning-based method.
[0058] Figure 3 A structural block diagram of the assembly system provided in an embodiment of this application is shown.
[0059] refer to Figure 3 Based on the above assembly method, this application also provides an assembly system, also known as a learning-type plug-in and mating assembly control system, which includes: Assembly equipment (shown as the robot body), which includes a robotic arm.
[0060] An end effector, mounted at the end of a robotic arm, is used to grip parts to be assembled.
[0061] The end effector force and torque sensor (shown as a six-dimensional end effector force / torque sensor) is installed between the robotic arm flange and the end effector to collect end effector mechanical data.
[0062] The controller, connected to the assembly equipment and the end effector force and torque sensor, is used to execute the assembly method of this application. The functional modules implemented by the controller include, but are not limited to: a "force / torque data preprocessing" module, a "local task coordinate system establishment and relative pose calculation" module, a "learning policy network (i.e., learning policy model)" module, and a "motion command generation and safety constraint" module.
[0063] In one embodiment of this application, the system further includes a teleoperation device for allowing a user to control the robotic arm to perform assembly demonstrations and generate assembly demonstration trajectory data during offline teaching and online training phases; wherein, the teleoperation device is at least one of the following: a teaching robotic arm that is isomorphic or heteromorphic to the robotic arm, a three-dimensional force feedback handle, a three-dimensional mouse, a six-degree-of-freedom spatial input device, a joystick, or a virtual reality / augmented reality controller.
[0064] Specifically, the data flow and workflow of the assembly system are as follows: During the demonstration phase ( Figure 3(Above the dashed line) The user inputs motion commands via a teleoperation device. These commands are processed by the "Motion Command Generation and Safety Constraints" module, generating Cartesian space control commands which are then sent to the assembly equipment to drive the robot body to complete the assembly demonstration actions. Simultaneously, the end-effector pose data fed back by the assembly equipment is input to the "Local Task Coordinate System Establishment and Relative Pose Calculation" module to establish the local task coordinate system and calculate the relative pose. Raw data collected by the end-effector's six-dimensional force / torque sensor is input to the "Force / Torque Data Preprocessing" module for processing. The data generated in the above process is recorded and used for subsequent strategy training.
[0065] During the strategy execution phase ( Figure 3 (Below the dashed line) When the assembly equipment performs an assembly task, it feeds back the end-effector pose data in real time to the "Local Task Coordinate System Establishment and Relative Pose Calculation" module to obtain the relative pose of the end-effector relative to the local task coordinate system (i.e., end-effector pose data). The end-effector's six-dimensional force / torque sensor collects force / torque data in real time. After processing by the "Force / Torque Data Preprocessing" module, the data, together with the relative pose data, forms a state vector, which is then input to the "Learning Policy Network" module. The "Learning Policy Network" outputs motion increment data based on the state vector. This data is input to the "Motion Command Generation and Safety Constraints" module to generate specific Cartesian space control commands and send them to the assembly equipment to drive the end-effector to complete the next moment's motion. This cycle repeats, achieving end-to-end control of the entire assembly process.
[0066] It should be noted that the description of the system in this application's embodiments is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments; therefore, it will not be repeated. For technical details not covered in the assembly system provided in this application's embodiments, please refer to... Figures 1 to 2 The meaning is understood in accordance with the description of any of the accompanying drawings.
[0067] According to embodiments of this application, this application also provides an electronic device and a readable storage medium.
[0068] Figure 4 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0069] like Figure 4As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0070] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0071] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as assembly methods. For example, in some embodiments, the assembly method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the assembly method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the assembly method by any other suitable means (e.g., by means of firmware).
[0072] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0073] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0074] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0075] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0076] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0077] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0078] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0079] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0080] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An assembly method, characterized in that, The method includes: Receive assembly instructions, the assembly instructions including assembly tasks, the assembly tasks including the assembly station and assembly method of the parts to be assembled; In response to the assembly command, the control end of the assembly equipment, which has clamped the part to be assembled, moves to an initial pose corresponding to the assembly station, and establishes a local task coordinate system with the initial pose as the origin. Acquire the end-effector mechanical data and the end-effector pose data in the local task coordinate system; Based on the end-effector mechanical data and the end-effector pose data, a pre-trained learning-based policy model is used to determine the motion increment data of the end-effector in the local task coordinate system. Based on the incremental motion data, motion control commands are generated. These commands are used to control the end effector to assemble the parts to be assembled to the assembly station according to the assembly method.
2. The method according to claim 1, characterized in that, The step of determining the motion increment data of the end effector in the local task coordinate system based on the end effector mechanical data and the end effector pose data, using a pre-trained learned policy model, includes: The end-effector mechanical data and the end-effector pose data are combined into a state vector; The state vector is input into the learning policy model to obtain motion increment data output by the learning policy model. The motion increment data includes end-effector translation increment or end-effector translation increment and attitude increment in the local task coordinate system.
3. The method according to claim 2, characterized in that, The learning-based policy model is trained in the following way: Acquire multiple assembly demonstration trajectory data collected through assembly demonstrations, wherein the assembly demonstration trajectory data is a normal assembly trajectory or a specific assembly trajectory that includes contact loss and recovery processes; Training samples are extracted from the assembly demonstration trajectory data. Each training sample includes a sample state vector consisting of sample end pose data and corresponding sample end mechanical data, and sample motion increment data corresponding to the sample state vector. Using the sample state vector in the training samples as input and the sample motion increment data in the training samples as output, the model used to output the motion increment data is trained by imitation learning and / or reinforcement learning to obtain the learning policy model.
4. The method according to claim 1, characterized in that, The end-effector pose data includes the relative pose of the end-effector with respect to the local task coordinate system at the current moment; the end-effector mechanical data includes multiple forces and moments arranged in time sequence; correspondingly, the method further includes: Detect whether the relative pose exceeds the preset search range based on the initial pose, and / or detect whether the force or torque in the force and torque exceeds their respective safety thresholds; An error message is generated in response to the relative pose exceeding the search range, and / or the force or torque in the force and torque exceeding the corresponding safety threshold.
5. The method according to claim 1, characterized in that, The control terminal has moved the assembly equipment holding the part to be assembled to the initial position corresponding to the assembly station, including: Acquire an image of the assembly station; Feature extraction is performed on the image to obtain the visual features of the assembly station; Based on the visual features, calculate the pose error between the current end pose and the initial pose; Based on the pose error, the end effector is controlled to move to the initial pose.
6. The method according to claim 1, characterized in that, The assembly task is a plug-in and mating assembly task with contact constraints and guiding characteristics, including at least one of the following: plug and socket plugging, pin and socket mating, connector plugging, and snap-fit engagement.
7. An assembly system, characterized in that, The system includes: An assembly device, the assembly device including a robotic arm; An end effector, mounted at the end of the robotic arm, is used to grip the part to be assembled; An end-effector force and torque sensor is installed between the robotic arm flange and the end effector to collect end-effector force and torque data. A controller, connected to the assembly equipment and the end force and torque sensor, is used to perform the method according to any one of claims 1-6.
8. The assembly system according to claim 7, characterized in that, The system also includes: A teleoperation device is used to allow users to control the robotic arm to perform assembly demonstrations and generate assembly demonstration trajectory data during offline teaching and online training phases. The teleoperation device is at least one of the following: a teaching robot arm, a three-dimensional force feedback handle, a three-dimensional mouse, a six-degree-of-freedom spatial input device, a joystick, or a virtual reality / augmented reality controller, which may be isomorphic or heteromorphic to the robotic arm.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.