Robotic precision assembly method, apparatus, device, and medium

CN122807999APending Publication Date: 2026-09-25BINZHOU WEIQIAO NATIONAL SCIENCE & TECHNOLOGY ADVANCED TECHNOLOGY RESEARCH INSTITUTE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611037484.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该装配过程存在接触动力学复杂、静摩擦易引发卡滞、环境扰动不确定和装配容错率极低等特性,对机器人控制的稳定性、自适应能力与感知精度提出了极高要求

Benefits of technology

[0008]根据本发明的另一方面,提供了一种计算机可读存储介质,所述计算机可读存储介质存储有计算机指令,所述计算机指令用于使处理器执行时实现本发明任一实施例所述的机器人精密装配方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807999A_ABST
    Figure CN122807999A_ABST
Patent Text Reader

Abstract

The application discloses a robot precision assembly method, device, equipment and medium, and particularly relates to the technical field of industrial manufacturing. The method comprises the following steps: in the scene of robot bolt hole assembly, a time-aligned first observation vector is obtained; a strategy network is used for feature extraction and action reasoning on the first observation vector to generate a strategy action; anti-stuck intervention is performed on the strategy action to determine a joint torque instruction; the robot is driven to execute the joint torque instruction, and the first observation vector is updated to obtain a second observation vector, and a reward value of action execution result is obtained; single-step transition data is generated and stored according to the first observation vector, the strategy action, the reward value and the second observation vector; and the strategy network is updated according to the historically stored single-step transition data, so that the updated strategy network is applied to robot control in the next round. The embodiment of the application can improve assembly stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial manufacturing technology, and in particular to a robot precision assembly method, apparatus, equipment and medium. Background Technology

[0002] In the field of precision industrial manufacturing, the assembly of robot pin holes with sub-millimeter tolerances is a core and critical process in component assembly, widely used in high-precision production scenarios such as high-end equipment and precision electronics. This assembly process is characterized by complex contact dynamics, static friction that can easily cause jamming, uncertain environmental disturbances, and extremely low assembly fault tolerance, which places extremely high demands on the stability, adaptability, and perception accuracy of robot control.

[0003] Existing robot precision assembly reinforcement learning control technology lacks an active intervention and release mechanism for assembly stall and jamming conditions. It cannot effectively break the local optimal jamming state in the insertion process and has inherent technical bottlenecks such as easy assembly failure, easy training stagnation and poor operation stability. Summary of the Invention

[0004] This invention provides a robot precision assembly method, apparatus, equipment, and medium that can intervene in assembly jamming problems and improve assembly stability.

[0005] According to one aspect of the present invention, a robot precision assembly method is provided, the method comprising: In the scenario of assembling a robot pin hole, the multimodal perception data of the robot is processed in time synchronization to obtain a time-aligned first observation vector; The policy network is used to extract features and infer actions from the first observation vector to generate policy actions. Anti-jamming intervention is performed on the aforementioned strategy actions to determine joint torque commands; The robot is driven to execute the joint torque command, and the first observation vector is updated to obtain the second observation vector, as well as the reward value of the action execution result is obtained; Based on the first observation vector, the policy action, the reward value, and the second observation vector, generate and store single-step transition data; The policy network is updated based on historically stored single-step transition data so that the updated policy network can be applied to the next round of robot control.

[0006] According to one aspect of the present invention, a robot precision assembly apparatus is provided, the apparatus comprising: The time synchronization module is used to perform time synchronization processing on the multimodal perception data of the robot in the scenario of robot pin hole assembly, so as to obtain the first observation vector with time alignment. The policy reasoning module is used to perform feature extraction and action reasoning on the first observation vector through a policy network to generate policy actions. The anti-jamming intervention module is used to perform anti-jamming intervention on the strategy action and determine the joint torque command; The reward calculation module is used to drive the robot to execute the joint torque command, update the first observation vector, obtain the second observation vector, and obtain the reward value of the action execution result; The data storage module is used to generate and store single-step transition data based on the first observation vector, the policy action, the reward value, and the second observation vector; The network update module is used to update the policy network based on historically stored single-step transition data, so that the updated policy network can be applied to the next round of robot control.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the robot precision assembly method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the robot precision assembly method according to any embodiment of the present invention.

[0009] The technical solution of this invention addresses the precision assembly scenario of robot pin holes. It obtains a first observation vector that accurately matches the working condition through multimodal perception data time-series alignment processing, providing a reliable input foundation for policy network feature extraction and action inference. This effectively avoids observation bias caused by asynchronous acquisition from multiple sensors, ensuring the accuracy of assembly policy decisions. Based on this, the policy network adaptively infers and outputs policy actions adapted to the current assembly condition, meeting the requirements of refined assembly control. Furthermore, an anti-jamming intervention step is added after the policy action output to adaptively determine reasonable joint torque commands. This proactively identifies and eliminates motion stagnation and jamming caused by static friction and contact constraints during assembly, solving the problems of traditional reinforcement learning assembly control lacking proactive working condition intervention, easily getting trapped in local optima, and assembly stability issues. To address the issue of poor performance, the robot executes torque commands to update the assembly status in real time and obtain corresponding reward values. This provides accurate feedback on the single-step assembly operation effect, thereby constructing and storing complete single-step transfer data containing observation vectors, policy actions, reward values, and updated observation vectors. This continuously accumulates real and effective assembly interaction samples, providing sufficient data support for the iterative optimization of the policy network. The policy network can continuously adapt and upgrade based on historical interaction data, continuously adapting to changes in assembly environment disturbances and pose deviations. This effectively improves the policy's generalization ability under working conditions and long-term assembly control accuracy. Overall, it achieves a closed-loop control effect of precise perception, fine action, strong anti-jamming ability, and adaptive iterative optimization of the policy in precision pin hole assembly, significantly improving the operational stability and product success rate of industrial precision pin assembly.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of a robot precision assembly method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of another robot precision assembly method provided according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a robot precision assembly device according to an embodiment of the present invention; Figure 4This is a schematic diagram of the structure of an electronic device that implements the robot precision assembly method of the present invention. Detailed Implementation

[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0015] Figure 1 This is a flowchart illustrating a robot precision assembly method according to an embodiment of the present invention. This embodiment is applicable to situations where a robot is used to assemble pin holes. The method can be executed by a robot precision assembly device, which can be implemented in hardware and / or software. The robot precision assembly device can be configured in a robot controller.

[0016] See Figure 1 The robot precision assembly method shown includes: S101. In the scenario of assembling the robot's pin holes, the multimodal perception data of the robot is processed in time synchronization to obtain a time-aligned first observation vector.

[0017] In this context, the robot pin-hole assembly scenario refers to the precision assembly process of accurately inserting cylindrical pins into corresponding circular holes. Multimodal perception data can include image data and body data. Body data can include pose data of the body, such as position, velocity, and acceleration. The body can include joints and end effectors. In practice, the sampling frequencies of perception data from different modalities differ, requiring time alignment. Analyzing time-aligned data reduces interference caused by time delays. The first observation vector can be a vector formed by concatenating time-aligned perception data from multiple modalities. The first observation vector describes the environmental state at the current moment.

[0018] In some embodiments, the mode with the lowest sampling frequency is determined from the multimodal sensing data. Using the sensing data of this mode as a reference, sensing data from other modes whose sensing data is closest in time to this mode are acquired and used as a set of time-aligned multimodal data. This time-aligned multimodal data is then used to generate a first observation vector. By continuously acquiring multimodal sensing data during the assembly process, multiple sets of time-aligned multimodal data can be generated, resulting in multiple first observation vectors.

[0019] In an optional embodiment, the step of performing time synchronization processing on the multimodal perception data of the robot to obtain an aligned first observation vector includes: when a visual image acquired by the robot is received, obtaining the acquisition timestamp of the visual image; the visual image includes: a local image of the assembly area and a global image; for at least one robot modality, based on the difference between the perception timestamp of the robot modality's perception data and the acquisition timestamp of the visual image, selecting the perception data closest to the visual image as the perception data for which the robot modality is aligned with the visual image; and generating an aligned first observation vector based on the visual image and the aligned perception data of each robot modality.

[0020] In this context, a local image of the assembly area can refer to a close-up image of the parts to be assembled taken by an end-effector camera; these parts can be pins or sockets. A global image can refer to a video image of the entire assembly workbench taken by a fixed camera. Robot modality refers to the type of robot body data. Robot modality can include joint positions, joint velocities, and end-effector poses, etc. Alignment refers to the temporal pairing and binding of images with body data.

[0021] Specifically, the wrist camera and the still camera are activated to asynchronously acquire visual images of the assembly area, including both local and global images, at a low frequency, and the corresponding image timestamps t are recorded. imgSimultaneously, the robot's proprioceptive internal state, i.e., perception data of multiple robot modalities, is acquired at high frequency in real time and bound to the corresponding robot timestamp t. robot Specifically, 1) Visual input: Raw RGB (Red, Green, Blue) image frames asynchronously acquired by the static camera and wrist camera at 30Hz, and the corresponding image timestamps t. img 2) Body input: Body perception data (including joint position q) fed back by the robot's underlying controller at a high frequency of 1kHz. t Joint velocity End effector pose x ee ), and the corresponding ontology timestamp t robot The processing procedure is as follows: A high-frequency updated circular buffer is established to temporarily store the ontology-aware data. Last-Hold (the last-hold synchronization mechanism) is triggered: when a new video frame is received, the absolute value of the time difference |t| is retrieved in reverse order within the circular buffer. robot -t img The smallest ontology data is extracted and physically aligned and concatenated with the video image to obtain the aligned reinforcement learning multimodal environment state observation vector, i.e., the first observation vector s. t : Among them, I wrist Video images captured by a wrist camera, specifically local images of the assembly area; I static The video image captured by the static camera, i.e., the global image; T represents vector transpose.

[0022] It is evident that by aligning the images with the ontology data in a timely manner, the problem of asynchrony caused by inconsistent sampling frequencies and time misalignment of multimodal sensing data can be avoided, ensuring that the input observation vector conforms to the real physical working conditions, reducing decision-making errors from the data source, and lowering the probability of assembly misjudgment and jamming; by using local and global dual images in combination with multimodal ontology data, the integrity of observation information is increased, and the accuracy of judging the relative position of pin holes is improved.

[0023] S102. The first observation vector is subjected to feature extraction and action reasoning through a policy network to generate policy actions.

[0024] Feature extraction is used to extract image features from the first observation vector. Action reasoning is used to infer policy actions based on image features and ontology data. A policy action can refer to the predicted incremental action of the current end-effector pose, i.e., the desired fine-tuning action. The action directly output by the policy network is the output action, which is not executed by the robot. The policy action is the final action obtained by introducing re-parameter sampling and is used for the robot's actual motion control. Correspondingly, the output action is the deterministic mean directly calculated by the policy network, representing the optimal intention or the original output. The policy action is the action actually executed after adding random noise to the mean, representing the actual physical action performed.

[0025] In an optional embodiment, the step of performing feature extraction and action reasoning on the first observation vector through a policy network to generate policy actions includes: extracting features from the visual image in the first observation vector through the policy network to obtain a feature map; reducing the dimensionality of the feature map through the policy network to obtain key point coordinates; fusing the key point coordinates with the robot body data in the first observation vector to obtain a fusion result; predicting the action probability distribution based on the fusion result; and reparameterizing the action probability distribution according to a preset standard distribution noise to obtain policy actions.

[0026] Feature extraction refers to extracting the spatial texture information of pins and holes from video images and outputting a multidimensional feature map. The feature map can be a multidimensional feature tensor obtained from the video image after feature extraction, containing spatial location information of the target. For example, a ResNet (Residual Network) backbone network can be used to extract features from video images. Dimensionality reduction of the feature map to obtain keypoint coordinates can be achieved by compressing the feature map and calculating the two-dimensional pixel coordinates of the pin and hole centers, i.e., quantizing the position of key components using coordinates. The fusion result can refer to concatenating the keypoint coordinates and the body data. The action probability distribution can refer to the robot's end effector following a specific distribution, such as a Gaussian distribution. The standard distribution noise can be standard normal distribution noise. Based on the preset standard distribution noise, the action probability distribution is reparameterized and sampled to obtain the policy action. This can be achieved by multiplying the mean μ, standard deviation σ of the action probability distribution with the preset standard normal noise ξ, then adding it back to the mean μ, and compressing the result to the [-1,1] interval using the tanh function to obtain the policy action.

[0027] For example, in S21, a ResNet-10 residual backbone network is used to perform convolutional feature extraction on the visual image within the first observation vector, outputting a multi-channel feature map, i.e., feature map Z∈R. C×H×WS22. Spatial Softmax is used to replace global average pooling for dimensionality reduction of the feature map, and the expected 2D (two-dimensional) keypoint coordinates of each channel are calculated: Where C is the number of channels, H is the height of the video image, and W is the width of the video image. c , v c ) represents the coordinates of the keypoint output in the c-th channel. h represents the row coordinates of the pixels in the feature image, and w represents the column coordinates of the pixels in the feature image. To iterate through the row coordinates of all pixels in the c-th channel, the values ​​range from 1 to H. The column coordinates of all pixels in the c-th channel are traversed, with values ​​ranging from 1 to W.

[0028] S23. The keypoint coordinates are concatenated and fused with the robot body data such as joint position, joint velocity, and end-effector pose in the first observation vector, and then fed into a multi-layer perceptron (MLP) to output the mean of the Gaussian distribution of the motion. With log standard deviation S24. Introduce random noise following a standard normal distribution ξ~N(0,I), and use the reparameterization formula to calculate the policy action. : The sampling generates an incremental pose strategy action with values ​​in the range [-1, 1]. Using a hyperbolic tangent function, the action is compressed into a fixed interval to avoid abrupt changes in output limits, preventing sudden robot stops and workpiece collisions. This approach is suitable for sub-millimeter precision pin assembly scenarios. The subsequent linear mapping can be converted to a Cartesian displacement command Δx. cmd Cartesian displacement can refer to the physical displacement of the robot's end effector in a three-dimensional Cartesian coordinate system, involving translation and rotation.

[0029] As can be seen, by extracting features from images and reducing dimensionality to obtain the coordinates of key points, the relative spatial position of pin holes can be accurately represented, geometric information can be quantified, and the assembly positioning accuracy can be improved. By fusing visual key points with robot body data, the spatial information of the image and the robot's own motion state can be integrated, making the observation dimensions more comprehensive. The distribution of actions is reparameterized, taking into account both the determinism of assembly actions and reasonable exploration space, which ensures smooth insertion and allows for small-scale probing and searching of pin holes, reducing the probability of assembly jamming.

[0030] S103. Perform anti-jamming intervention on the strategy action and determine the joint torque command.

[0031] Anti-jamming intervention refers to determining whether the robot is jammed, actively intervening after jamming occurs, and restoring normal control after the jammed state is resolved. Jamming can occur when the robot's end effector pins and sockets physically interfere with each other, causing static friction lock-up. The robot's end effector cannot produce displacement as instructed, and although the joints output electrical signals, the end effector does not actually move; there is no hardware fault in the joints themselves. Joint torque commands can include commands to release jamming and normal drive control commands when jamming does not occur. Anti-jamming intervention can be used to break out of the original control strategy when jamming occurs, execute a control strategy to release jamming, and then resume the original control strategy after the jamming is resolved. Joint torque commands can refer to commands sent to the robot's underlying actuators, controlling joint rotation and end effector movement by setting the output torque values ​​of each joint.

[0032] S104. Drive the robot to execute the joint torque command, update the first observation vector, obtain the second observation vector, and obtain the reward value of the action execution result.

[0033] In this process, the robot executing joint torque commands is equivalent to completing one assembly step. The second observation vector can refer to the observation vector formed after re-acquiring multimodal perception data and temporally aligning it after the action is completed. The reward value can refer to the assembly effect after executing a single joint torque command (single-step assembly).

[0034] S105. Generate and store single-step transition data based on the first observation vector, the policy action, the reward value, and the second observation vector.

[0035] Specifically, the first observation vector, policy action, reward value, and second observation vector are concatenated to obtain the single-step transition data. Storing the single-step transition data for each step yields a large amount of single-step transition data.

[0036] S106. Update the policy network based on the historically stored single-step transfer data so that the updated policy network can be applied to the next round of robot control.

[0037] Among them, the historically stored single-step transfer data represents historical experience data, and the historically stored single-step transfer data update strategy network can be used for the next round of assembly decisions so as to drive the robot to continue to perform assembly actions.

[0038] The technical solution of this invention addresses the precision assembly scenario of robot pin holes. It obtains a first observation vector that accurately matches the working condition through multimodal perception data time-series alignment processing, providing a reliable input foundation for policy network feature extraction and action inference. This effectively avoids observation bias caused by asynchronous acquisition from multiple sensors, ensuring the accuracy of assembly policy decisions. Based on this, the policy network adaptively infers and outputs policy actions adapted to the current assembly condition, meeting the requirements of refined assembly control. Furthermore, an anti-jamming intervention step is added after the policy action output to adaptively determine reasonable joint torque commands. This proactively identifies and eliminates motion stagnation and jamming caused by static friction and contact constraints during assembly, solving the problems of traditional reinforcement learning assembly control lacking proactive working condition intervention, easily getting trapped in local optima, and assembly stability issues. To address the issue of poor performance, the robot executes torque commands to update the assembly status in real time and obtain corresponding reward values. This provides accurate feedback on the single-step assembly operation effect, thereby constructing and storing complete single-step transfer data containing observation vectors, policy actions, reward values, and updated observation vectors. This continuously accumulates real and effective assembly interaction samples, providing sufficient data support for the iterative optimization of the policy network. The policy network can continuously adapt and upgrade based on historical interaction data, continuously adapting to changes in assembly environment disturbances and pose deviations. This effectively improves the policy's generalization ability under working conditions and long-term assembly control accuracy. Overall, it achieves a closed-loop control effect of precise perception, fine action, strong anti-jamming ability, and adaptive iterative optimization of the policy in precision pin hole assembly, significantly improving the operational stability and product success rate of industrial precision pin assembly.

[0039] Figure 2 This is a flowchart illustrating a robot precision assembly method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention provides anti-jamming intervention for the strategy action and determines joint torque commands. Specifically, it involves: detecting whether the robot is in a motion-stagnant state based on the strategy action and historical robot end-effector velocities; when the robot is detected to be in a motion-stagnant state, sequentially generating and outputting corresponding joint torque commands according to a jamming release process; the jamming release process includes: freezing strategy, upward retraction, axial stiffness zeroing, lateral random search, and recovery strategy; when the robot is detected not to be in a motion-stagnant state, mapping the strategy action to an anisotropic impedance controller to determine the end-effector target pose; calculating the target velocity and the actual velocity; calculating the feedback torque based on the pose difference between the end-effector target pose and the current end-effector pose, and the velocity difference between the target velocity and the actual velocity; and determining the joint torque commands based on the feedback torque. It should be noted that parts not detailed in this embodiment of the present invention can be found in the descriptions of other embodiments.

[0040] See Figure 2 The robot precision assembly method shown includes: S201. In the scenario of assembling the robot's pin holes, the multimodal perception data of the robot is processed in time synchronization to obtain a time-aligned first observation vector.

[0041] S202. The first observation vector is subjected to feature extraction and action reasoning through a policy network to generate policy actions.

[0042] S203. Based on the strategy action and the historical robot end-effector velocity, detect whether the robot is in a motion-stagnant state.

[0043] The motion stagnation state refers to a state where there is an intention to move, but the measured velocity of the end effector is 0. The motion stagnation state characterizes the robot's end effector having a continuously zero velocity while executing the intended action, and is used to determine if the robot is in a pin-hole assembly jamming condition. The strategy action is used to determine whether there is an intention to move. Historical robot end effector velocities are used to determine whether the measured velocity of the robot's end effector is 0, in order to determine whether static friction has caused a lock-up between the end effector pin and the socket.

[0044] In one example, the existence of a state of motion stagnation is determined based on the following formula: Among them, the robot end effector speed collected within the historical W sliding time windows is a t This represents the desired action without noise. This is the threshold for stagnant speed. Ψt is the threshold for the intended action. Ψt=1: The robot is determined to be in a motion-stagnant state, that is, the end effector is stopped and the pin hole is stuck; Ψt=0: The end effector is moving normally and there is no sticking, and conventional impedance control can be performed.

[0045] It should be noted that, It is the strategy action after sampling with superimposed normal noise, a t For the desired action that does not contain noise, a t This actually represents an ideal, undisturbed strategy action. In this embodiment of the invention, the strategy action is... a t It is the output action directly output by the policy network.

[0046] S204. When the robot is detected to be in a motion-stagnant state, the corresponding joint torque commands are generated and output sequentially according to the jam release procedure. The jam release procedure includes: freeze strategy, upward retraction, axial stiffness zeroing, lateral random search, and recovery strategy.

[0047] The jamming release process refers to a pre-defined procedure for resolving jamming. Each step in the jamming release process can generate a corresponding joint torque command. Specifically, the freeze strategy temporarily disables the strategy actions output by the original strategy network; the upward retraction is used to output a reverse torque to lift the pin away from the jammed position at the orifice; the axial stiffness is zeroed to lower the Z-axis control stiffness, giving the end effector passive compliant floating capability to avoid hard jamming; the lateral random search applies small random torques to the X and Y axes (XYZ is a pre-constructed three-dimensional coordinate system) to explore left and right to find the center of the orifice; the recovery strategy is used to reactivate the original strategy actions for normal insertion after the jamming is resolved. In one example, the following steps are executed in sequence: freeze the reinforcement learning strategy to prevent accumulated noise; apply a retraction command to the end effector to break the contact; temporarily set the axial stiffness of the anisotropic impedance controller to 0 to unload stress; apply lateral random noise for exploration; and then reinstate the reinforcement learning strategy network. For example, if the pin is stuck at the edge of the hole, the freezing strategy is as follows: first stop the feed, then retract upwards: raise by 0.02mm, reduce the axial stiffness to zero: soften the Z-axis, apply lateral random noise to search: slightly shake left and right to find the hole and then press down normally for assembly.

[0048] S205. When it is detected that the robot is not in a motion-stagnant state, the strategy action is mapped to the anisotropic impedance controller to determine the end-effector pose.

[0049] The fact that the robot is not in a state of motion stagnation indicates that the robot can be controlled normally. Mapping to an anisotropic impedance controller can refer to mapping policy actions to real physical displacements, such as Cartesian displacement Δx. cmd An anisotropic impedance controller can be a compliant controller with different stiffness and damping parameters along the X, Y, and Z axes. The end-effector target pose can be the desired spatial position and orientation of the robot's end effector. By mapping the strategy actions to the anisotropic impedance controller, higher axial compliance is provided while maintaining lateral accuracy. The feedback control target torque is calculated to control the robot to perform assembly. For example, if no jamming occurs, the strategy actions are mapped to the X-axis 0.03 mm to the right as the target point of the impedance controller.

[0050] S206. Calculate the target speed and the actual speed.

[0051] The target velocity can be the desired motion velocity obtained from the end-effector pose difference, while the actual velocity can be the real motion velocity of the end-effector acquired in real time.

[0052] S207. Calculate the feedback torque based on the pose difference between the end target pose and the current end pose, and the velocity difference between the target velocity and the actual velocity.

[0053] The current pose of the end effector can refer to the actual pose of the robot's end effector acquired in real time. Pose difference describes the positional deviation between the desired and actual positions. Velocity difference describes the velocity deviation between the desired and actual velocities. Feedback torque can refer to a correction torque to bring the current pose of the end effector closer to the target pose, and to bring the actual velocity closer to the target velocity. In some embodiments, the feedback torque can be calculated based on the following formula: Feedback torque = Position deviation × Stiffness coefficient + Velocity deviation × Damping coefficient. Specifically: Among them, F cmd For feedback torque, K p This is the position stiffness coefficient matrix. d Let be the end-point target pose, the desired end-point position obtained by linear mapping of the policy action. Let x be the current end-point pose. Let D be the damping coefficient matrix. It could be the target speed. The actual speed can be calculated. The feedback torque of each axis is obtained, and the results are summarized to generate the final joint torque command sent to the driver.

[0054] S208. Determine the joint torque command based on the feedback torque.

[0055] Among them, the feedback torque is the end torque in Cartesian space, which can be mapped to the output torque of each robot joint through the Jacobian transpose matrix, thus obtaining the joint torque command of the final driver.

[0056] S209. Drive the robot to execute the joint torque command, update the first observation vector, obtain the second observation vector, and obtain the reward value of the action execution result.

[0057] S210. Generate and store single-step transition data based on the first observation vector, the policy action, the reward value, and the second observation vector.

[0058] S211. Update the policy network based on the historically stored single-step transfer data so that the updated policy network can be applied to the next round of robot control.

[0059] The technical solution of this invention accurately distinguishes between normal feed and physical jamming of the pin hole by relying on the strategy action and the actual speed of the end effector to make a stop judgment, avoiding misjudgment. For jamming conditions, segmented unblocking torque logic is activated. Through lifting, low stiffness and lateral search, the pin is actively unblocked, which greatly reduces the probability of workpiece collision, jamming and scrap. When there is no jamming, anisotropic impedance control is adopted, which generates a compliant torque by relying on the dual deviation of position and speed. The end effector automatically buffers when encountering slight contact, which is suitable for precision pin assembly with micron-level gaps, improving the assembly fault tolerance and yield rate.

[0060] In an optional embodiment, obtaining the reward value of the action execution result includes: calculating a distance reward based on the assembly pose error in the action execution result; calculating an insertion indication reward based on the insertion depth in the hole in the action execution result; calculating an action smoothing penalty based on the current output action and the adjacent previous step output action in the action execution result; and weighting the distance reward, the insertion indication reward, and the action smoothing penalty to obtain the reward value.

[0061] The assembly pose error can be the spatial distance between the center of the pin and the center of the socket. Distance reward guides the reduction of assembly pose error, characterizing the alignment quality; the smaller the assembly pose error, the higher the distance reward, and vice versa. Insertion depth within the socket refers to the real-time feed depth of the pin extending into the socket, characterizing the degree of assembly advancement. Insertion indication reward guides a greater insertion depth; the greater the insertion depth within the socket, the higher the insertion indication reward, and vice versa. Adjacent previous output action refers to the previous step's current output action. Since this is actually updating the policy network, noisy policy actions are unnecessary, as they would introduce unnecessary interference. Action smoothing penalty guides smoother continuous actions, reducing the intensity of actions. The greater the action difference (intense), the higher the action smoothing penalty; the smaller the action difference (smooth), the lower the action smoothing penalty.

[0062] For example, after the robot performs a policy action, the distance reward R is calculated. dist Insertion Instruction Reward I depth and motion smoothing penalty R smooth .

[0063] Distance Reward R dist : ;x err This refers to the assembly pose error.

[0064] Insertion Instruction Reward I depth When z insertion -z hole > d thresh At that time, I depth = 1, otherwise 0; where z insertion Let z be the actual coordinate of the pin in the Z direction. hole The Z-coordinate of the socket reference is z. insertion -z hole d is the insertion depth inside the hole. thresh The depth threshold for determining whether a pin has successfully entered the hole.

[0065] 3) Motion smoothing penalty R smooth : a tThis is the current output action, a t-1 It is the output action immediately preceding the current output action. Motion smoothing penalty R smooth Used to penalize aggressive control commands to prevent triggering the robot's emergency stop protection.

[0066] Reward value r t The calculation formula is: r t = w d ×R dist +w i ×I depth -w a ×R smooth Among them, w d w is the weight of the distance reward. i To insert the weight of the indicator reward, w a The weight of the action smoothing penalty.

[0067] It is evident that by constraining alignment accuracy through distance rewards, the pins are continuously pulled closer to the center of the insertion hole, improving the coaxial positioning capability of the micro-hole; by guiding the feed trend through insertion depth rewards, the assembly progress is quantified by the insertion amount, incentivizing steady feeding at the end and avoiding hovering; by suppressing impacts and collisions through motion smoothing penalties, the constraint strategy network outputs smooth motions, reducing the impact of sudden large movements on the hole wall, and lowering the probability of jamming and workpiece scratches; finally, through multi-dimensional weighted composite dense rewards, feedback signals can be obtained at every step of the process, accelerating the convergence of the strategy network and improving the success rate and control stability of small-gap precision insertion.

[0068] In an optional embodiment, updating the policy network based on historically stored single-step transfer data includes: sampling mixed training data from historically stored single-step transfer data and human teaching experience data; for each training sample in the mixed training data, generating a current action value using a value network based on the earlier observation vector and policy action in the training sample; generating a target action value using the value network based on the later observation vector and reward value in the training sample; updating the value network based on the difference between the current action value and the target action value of each training sample; for each training sample in the mixed training data, generating a current predicted action using the policy network based on the earlier observation vector in the training sample; calculating the loss value of the policy network based on the difference between the teaching action corresponding to each training sample and the current predicted action; and updating the policy network based on the loss value of the policy network.

[0069] In this context, the single-step transition data is used to update the policy network, and thus the output action in the single-step transition data is the unprocessed output action 'a' of the policy network. tHuman instruction experience data can refer to data from manually performing pin hole assembly. Hybrid training data can be training data obtained by separately extracting data from historically stored single-step transfer data and human instruction experience data, and then mixing them together. Each training sample in the hybrid training data includes the observation vector, output action, reward value, and observation vector for the next adjacent step. For example, single-step transfer data can be stored in an online experience replay buffer D. rl ;D rl Teaching buffer D with pre-stored expert teaching samples demo (Data containing human teaching experience is stored independently and in parallel. It is retrieved from the online buffer D according to a preset ratio (e.g., 50% each). rl and teaching buffer D demo Samples are extracted and concatenated to form a mixed training batch, which serves as the input data for this round of training.

[0070] The value network processes the input observation vectors and outputs the long-term total reward that can be obtained by performing the action in the current state, i.e., the current action value. The earlier observation vectors correspond to the current observation vectors. The later observation vectors correspond to the observation vectors updated after executing the policy action corresponding to the output action. The value network also processes the later input observation vectors, traversing all candidate actions in the current state and outputting the long-term reward corresponding to each candidate action, selecting the maximum reward as the target action value. The current action value is the predicted value of the value network, which is the predicted reward corresponding to the measured action. The target action value is the constructed ground truth value of the current action value, obtained by summing the actual reward value obtained in the current step with the future optimal reward. The difference between the current action value and the target action value is the value obtained by subtracting the two. Updating the value network based on the difference between the current action value and the target action value allows the value network to learn in real time the ability to accurately predict the quality of actions. Backward gradient optimization of the value network weights can be used to adjust the parameters of the value network with the goal of reducing the difference between the current action value and the target action value. For example, the value network can be updated by gradient descent that minimizes the Bellman error between the actual reward and the network's predicted Q-value. The value network can employ a critic network.

[0071] The current predicted action can be the output action generated by the current policy network processing the earlier observation vectors. Since the output actions in the historically stored single-step transition data are based on the historical policy network processing the earlier observation vectors, but the current policy network may differ from the historical ones, this requires updating the latest policy network. Therefore, the output action needs to be re-predicted based on the latest policy network. The corresponding taught action can refer to the standard action of manual assembly corresponding to the current predicted action. Both the current predicted action and the corresponding taught action are the same action, but the current predicted action is the robot's action, while the taught action is the action performed manually. Updating the policy network based on the difference between the taught action corresponding to the training samples and the current predicted action allows the policy network to learn the correct assembly action in real time. Backward gradient optimization of the policy network weights can be used to adjust the policy network parameters with the goal of reducing the difference between the taught action corresponding to the training samples and the current predicted action. The policy network can be an actor network.

[0072] It is evident that by combining online training with teaching and mixed sampling to obtain training data, high-quality teaching samples provide correct assembly priors, while random interactive samples supplement boundary fault conditions, avoiding the shortcomings of poor generalization from simple teaching and high trial-and-error costs from simple autonomous exploration. The value network accurately quantifies the future benefits of each action, providing a reliable evaluation basis for strategy optimization. Based on teaching experience data, it quickly converges and obtains usable assembly strategies. At the same time, based on continuous optimization of the value network, it autonomously optimizes alignment and anti-jamming on the basis of demonstration, achieving stable training, effectively improving the success rate of small-gap insertion, and reducing the jamming failure rate.

[0073] In an optional embodiment, calculating the loss value of the policy network based on the difference between the taught action corresponding to each training sample and the current predicted action includes: for each training sample, generating a current probability distribution through the policy network based on the earlier observation vector in the training sample; calculating the current entropy regularization term of the training sample based on the current probability distribution; calculating a first loss value based on the difference between the current entropy regularization term of each training sample and the current action value of the training sample; calculating a second loss value based on the difference between the taught action corresponding to each training sample and the current predicted action; and calculating the loss value of the policy network based on the first loss value and the second loss value.

[0074] The current probability distribution refers to the action probability distribution obtained by the latest policy network based on earlier observation vectors. For example, the latest policy network extracts features from the visual images in earlier observation vectors to obtain a feature map; the feature map is then dimensionality-reduced to obtain keypoint coordinates; these keypoint coordinates are fused with the robot body data in earlier observation vectors to obtain a fusion result; and the current probability distribution is predicted based on this fusion result. The current entropy regularization term encourages action diversity and a wider action exploration range. The difference between the current entropy regularization term and the current action value of the training samples is used to balance action value and entropy regularization term, ensuring both excellent and diverse actions. The current action value is used to converge towards better actions, while entropy regularization prevents actions from being completely locked down; the difference reflects their mutual restraint. This approach continuously optimizes insertion accuracy based on rewards while preserving action tolerance space, improving the device's adaptability to workpiece position deviations and reducing jamming failure rates. The first loss value describes the degree of matching between the policy output action and the long-term environmental benefits, as well as the divergence of the action probability distribution, and is used for optimal optimization and anti-rigidity exploration. The second loss value is used to enable the robot to learn assembly capabilities by imitating the action specifications of human assembly, thereby reducing the deviation between the predicted action and the human-taught action.

[0075] In one example, the SAC (Soft Actor-Critic) algorithm framework is employed to minimize the Bellman error based on mixed training batch data, completing the iterative optimization of the value network parameters θ. An improved SAC algorithm with behavioral cloning regularization is used, introducing a behavioral cloning (BC) term to suppress covariate bias and constructing a hybrid loss function. : in, The first loss value is the SAC maximum entropy enhancement loss, and α is the adaptive entropy temperature coefficient, which balances reward maximization and action exploration. The second loss value, λ, is the penalty term for cloning the taught behavior. BC The weights of the second loss value constrain the output action of the policy network. Closely follow the expert's demonstration movements. exp The loss value of the policy network calculated by minimizing the above-mentioned mixed loss function through gradient descent is used to realize the policy network parameters. Updated to overcome the covariate shift problem caused by pure imitation learning.

[0076] It is evident that by calculating the first loss value to optimize the strategy network based on the difference between the current entropy regularization term and the current action value of the training sample, the strategy can be prevented from rapidly converging to a single fixed action, preserving the action exploration space, adapting to non-standard working conditions such as part deviations and assembly disturbances, and reducing the probability of jamming. By calculating the second loss value to optimize the strategy network based on the difference between the taught action corresponding to the training sample and the current predicted action, the basic insertion logic can be quickly learned, significantly shortening the early training and exploration costs. The strategy network optimized based on the combination of the first and second loss values ​​ensures that the actions are close to mature processes, and can autonomously optimize based on demonstrations, taking into account both assembly accuracy and environmental adaptability, and improving the success rate of small-gap insertion.

[0077] Figure 3 This is a schematic diagram of a robot precision assembly device provided in an embodiment of the present invention. The present invention is applicable to situations where a robot is used to assemble pin holes. This device can execute a robot precision assembly method and can be implemented in hardware and / or software.

[0078] See Figure 3 The robot precision assembly device shown includes: The time synchronization module 301 is used to perform time synchronization processing on the multimodal perception data of the robot in the scenario of robot pin hole assembly to obtain a time-aligned first observation vector. The strategy reasoning module 302 is used to perform feature extraction and action reasoning on the first observation vector through a policy network to generate policy actions; Anti-jamming intervention module 303 is used to perform anti-jamming intervention on the strategy action and determine the joint torque command; The reward calculation module 304 is used to drive the robot to execute the joint torque command, update the first observation vector, obtain the second observation vector, and obtain the reward value of the action execution result; The data storage module 305 is used to generate and store single-step transfer data based on the first observation vector, the strategy action, the reward value and the second observation vector; The network update module 306 is used to update the policy network based on the historically stored single-step transfer data, so that the updated policy network can be applied to the next round of robot control.

[0079] The technical solution of this invention addresses the precision assembly scenario of robot pin holes. It obtains a first observation vector that accurately matches the working condition through multimodal perception data time-series alignment processing, providing a reliable input foundation for policy network feature extraction and action inference. This effectively avoids observation bias caused by asynchronous acquisition from multiple sensors, ensuring the accuracy of assembly policy decisions. Based on this, the policy network adaptively infers and outputs policy actions adapted to the current assembly condition, meeting the requirements of refined assembly control. Furthermore, an anti-jamming intervention step is added after the policy action output to adaptively determine reasonable joint torque commands. This proactively identifies and eliminates motion stagnation and jamming caused by static friction and contact constraints during assembly, solving the problems of traditional reinforcement learning assembly control lacking proactive working condition intervention, easily getting trapped in local optima, and assembly stability issues. To address the issue of poor performance, the robot executes torque commands to update the assembly status in real time and obtain corresponding reward values. This provides accurate feedback on the single-step assembly operation effect, thereby constructing and storing complete single-step transfer data containing observation vectors, policy actions, reward values, and updated observation vectors. This continuously accumulates real and effective assembly interaction samples, providing sufficient data support for the iterative optimization of the policy network. The policy network can continuously adapt and upgrade based on historical interaction data, continuously adapting to changes in assembly environment disturbances and pose deviations. This effectively improves the policy's generalization ability under working conditions and long-term assembly control accuracy. Overall, it achieves a closed-loop control effect of precise perception, fine action, strong anti-jamming ability, and adaptive iterative optimization of the policy in precision pin hole assembly, significantly improving the operational stability and product success rate of industrial precision pin assembly.

[0080] Optional, the anti-jamming intervention module 303 is specifically used for: Based on the strategy action and the historical robot end-effector velocity, detect whether the robot is in a state of motion stagnation; When the robot is detected to be in a motion stagnation state, the corresponding joint torque commands are generated and output sequentially according to the jam release process; the jam release process includes: freeze strategy, upward retraction, axial stiffness to zero, lateral random search and recovery strategy; When it is detected that the robot is not in a state of motion stagnation, the strategy action is mapped to the anisotropic impedance controller to determine the end-effector pose; Calculate the target speed and the actual speed; The feedback torque is calculated based on the pose difference between the end target pose and the current end pose, and the velocity difference between the target velocity and the actual velocity. The joint torque command is determined based on the feedback torque.

[0081] Optional, the strategy reasoning module 302 is specifically used for: The visual image in the first observation vector is feature extracted using a policy network to obtain a feature map. The feature map is reduced in dimensionality using the policy network to obtain the coordinates of key points; The key point coordinates are fused with the robot body data in the first observation vector to obtain the fusion result; Predict the action probability distribution based on the fusion results; Based on a preset standard noise distribution, the action probability distribution is reparameterized and sampled to obtain the strategy action.

[0082] Optional, the reward calculation module 304 is specifically used for: Based on the assembly pose error in the action execution result, calculate the distance reward; Calculate the insertion indication reward based on the insertion depth within the hole in the action execution result; Based on the current output action and the adjacent previous step output action in the action execution result, calculate the action smoothing penalty; The distance reward, the insertion indication reward, and the motion smoothing penalty are weighted to obtain a reward value.

[0083] Optional, network update module 306, specifically used for: Hybrid training data is obtained by sampling from historically stored single-step transfer data and human teaching experience data; For each training sample in the mixed training data, the current action value is generated by the value network based on the earlier observation vector and policy action in the training sample. The target action value is generated through the value network based on the observation vector and reward value that are later in time in the training samples; The value network is updated based on the difference between the current action value and the target action value of each training sample; For each training sample in the mixed training data, the policy network generates the current predicted action based on the earlier observation vector in the training sample. The loss value of the policy network is calculated based on the difference between the taught action corresponding to each training sample and the current predicted action. The policy network is updated based on its loss value.

[0084] Optional, network update module 306, specifically used for: For each of the training samples, the policy network generates the current probability distribution based on the earlier observation vectors in the training samples; Calculate the current entropy regularization term of the training samples based on the current probability distribution; The first loss value is calculated based on the difference between the current entropy regularization term of each training sample and the current action value of the training sample. The second loss value is calculated based on the difference between the taught action corresponding to each training sample and the current predicted action; The loss value of the policy network is calculated based on the first loss value and the second loss value.

[0085] Optional, time synchronization module 301, specifically used for: When the visual image acquired by the robot is received, the acquisition timestamp of the visual image is obtained; the visual image includes: a local image of the assembly area and a global image; For at least one robot modality, based on the difference between the perception timestamp of the robot modality's perception data and the acquisition timestamp of the visual image, the perception data closest to the visual image is selected as the perception data for aligning the robot modality with the visual image; Based on the visual image and the perception data of each of the aligned robot modalities, an aligned first observation vector is generated.

[0086] The acquisition, storage, and application of data involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0087] The robot precision assembly device provided in the embodiments of the present invention can execute the robot precision assembly method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0088] Figure 4 A schematic diagram of the structure of an electronic device 400 that can be used to implement an embodiment of the present invention is shown.

[0089]

[01] As Figure 4 As shown, the electronic device 400 includes at least one processor 401 and a memory, such as a read-only memory 402 or a random access memory 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 402 or loaded from storage unit 408 into the random access memory 403. The random access memory 403 may also store various programs and data required for the operation of the electronic device 400. The processor 401, read-only memory 402, and random access memory 403 are interconnected via a bus 404. An input / output interface 405 is also connected to the bus 404.

[0090]

[02] Multiple components in the electronic device 400 are connected to the input / output interface 405, including: an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0091]

[03] The processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 401 include, but are not limited to, a central processing unit, a graphics processing unit, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. The processor 401 performs the various methods and processes described above, such as robotic precision assembly methods.

[0092]

[04] In some embodiments, the robot precision assembly method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 400 via read-only memory 402 and / or communication unit 409. When the computer program is loaded into random access memory 403 and executed by processor 401, one or more steps of the robot precision assembly method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform the robot precision assembly method by any other suitable means (e.g., by means of firmware).

[0093]

[05] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits, application-specific standard products, systems-on-a-chip systems, complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implementation in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0094]

[06] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. Such computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer program may be executed entirely on the machine, partially on the machine, or as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0095]

[07] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0096]

[08] To provide interaction with the user, the systems and techniques described herein can be implemented on an operational detection device. The robot has: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the robot. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0097]

[09] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0098]

[10] A computing system may include a target user terminal and a server. The target user terminal and the server are generally far apart and usually interact through a communication network. The relationship between the target user terminal and the server is generated by computer programs running on the respective computers and having a target user terminal-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of high management difficulty and weak business scalability in traditional physical host and virtual private server services.

[0099] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0100] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A robot precision assembly method, characterized in that, include: In the scenario of assembling a robot pin hole, the multimodal perception data of the robot is processed in time synchronization to obtain a time-aligned first observation vector; The policy network is used to extract features and infer actions from the first observation vector to generate policy actions. Anti-jamming intervention is performed on the aforementioned strategy actions to determine joint torque commands; The robot is driven to execute the joint torque command, and the first observation vector is updated to obtain the second observation vector, as well as the reward value of the action execution result is obtained; Based on the first observation vector, the policy action, the reward value, and the second observation vector, generate and store single-step transition data; The policy network is updated based on historically stored single-step transition data so that the updated policy network can be applied to the next round of robot control.

2. The method according to claim 1, characterized in that, The anti-jamming intervention for the strategy action and the determination of joint torque commands include: Based on the strategy action and the historical robot end-effector velocity, detect whether the robot is in a state of motion stagnation; When the robot is detected to be in a motion stagnation state, the corresponding joint torque commands are generated and output sequentially according to the jam release process; the jam release process includes: freeze strategy, upward retraction, axial stiffness to zero, lateral random search and recovery strategy; When it is detected that the robot is not in a state of motion stagnation, the strategy action is mapped to the anisotropic impedance controller to determine the end-effector pose; Calculate the target speed and the actual speed; The feedback torque is calculated based on the pose difference between the end target pose and the current end pose, and the velocity difference between the target velocity and the actual velocity. The joint torque command is determined based on the feedback torque.

3. The method according to claim 1, characterized in that, The step of extracting features and inferring actions from the first observation vector through a policy network to generate policy actions includes: The visual image in the first observation vector is feature extracted using a policy network to obtain a feature map. The feature map is reduced in dimensionality using the policy network to obtain the coordinates of key points; The key point coordinates are fused with the robot body data in the first observation vector to obtain the fusion result; Predict the action probability distribution based on the fusion results; Based on a preset standard noise distribution, the action probability distribution is reparameterized and sampled to obtain the strategy action.

4. The method according to claim 1, characterized in that, The reward value for obtaining the result of the action includes: Based on the assembly pose error in the action execution result, calculate the distance reward; Calculate the insertion indication reward based on the insertion depth within the hole in the action execution result; Based on the current output action and the adjacent previous step output action in the action execution result, calculate the action smoothing penalty; The distance reward, the insertion indication reward, and the motion smoothing penalty are weighted to obtain a reward value.

5. The method according to claim 1, characterized in that, The step of updating the policy network based on historically stored single-step transfer data includes: Hybrid training data is obtained by sampling from historically stored single-step transfer data and human teaching experience data; For each training sample in the mixed training data, the current action value is generated by the value network based on the earlier observation vector and policy action in the training sample. The target action value is generated through the value network based on the observation vector and reward value that are later in time in the training samples; The value network is updated based on the difference between the current action value and the target action value of each training sample; For each training sample in the mixed training data, the policy network generates the current predicted action based on the earlier observation vector in the training sample. The loss value of the policy network is calculated based on the difference between the taught action corresponding to each training sample and the current predicted action. The policy network is updated based on its loss value.

6. The method according to claim 5, characterized in that, The step of calculating the loss value of the policy network based on the difference between the taught action corresponding to each training sample and the current predicted action includes: For each of the training samples, the policy network generates the current probability distribution based on the earlier observation vectors in the training samples; Calculate the current entropy regularization term of the training samples based on the current probability distribution; The first loss value is calculated based on the difference between the current entropy regularization term of each training sample and the current action value of the training sample. The second loss value is calculated based on the difference between the taught action corresponding to each training sample and the current predicted action; The loss value of the policy network is calculated based on the first loss value and the second loss value.

7. The method according to claim 1, characterized in that, The step of performing time synchronization processing on the multimodal perception data of the robot to obtain the aligned first observation vector includes: When the visual image acquired by the robot is received, the acquisition timestamp of the visual image is obtained; the visual image includes: a local image of the assembly area and a global image; For at least one robot modality, based on the difference between the perception timestamp of the robot modality's perception data and the acquisition timestamp of the visual image, the perception data closest to the visual image is selected as the perception data for aligning the robot modality with the visual image; Based on the visual image and the perception data of each of the aligned robot modalities, an aligned first observation vector is generated.

8. A robot precision assembly device, characterized in that, The device includes: The time synchronization module is used to perform time synchronization processing on the multimodal perception data of the robot in the scenario of robot pin hole assembly, so as to obtain the first observation vector with time alignment. The policy reasoning module is used to perform feature extraction and action reasoning on the first observation vector through a policy network to generate policy actions. The anti-jamming intervention module is used to perform anti-jamming intervention on the strategy action and determine the joint torque command; The reward calculation module is used to drive the robot to execute the joint torque command, update the first observation vector, obtain the second observation vector, and obtain the reward value of the action execution result; The data storage module is used to generate and store single-step transition data based on the first observation vector, the policy action, the reward value, and the second observation vector; The network update module is used to update the policy network based on historically stored single-step transition data, so that the updated policy network can be applied to the next round of robot control.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the robot precision assembly method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the robot precision assembly method according to any one of claims 1-7.