Robot shaft hole assembly rotating axial decoupling and six-degree-of-freedom pose correction method, system and equipment and medium

By acquiring multimodal representations of visual and non-visual data through a three-view camera and decoupling constraints of rotational axis, the stability and accuracy problems of traditional robot shaft hole assembly under complex working conditions are solved. This achieves stability and generalization ability of six-degree-of-freedom pose correction, improving assembly success rate and compliance.

CN121374569APending Publication Date: 2026-01-23BAIHE POWER SUPPLY BUREAU OF GUANGXI POWER GRID CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511557050.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Traditional robot shaft and hole assembly suffers from poor stability under complex working conditions, and it is difficult to balance pose correction accuracy and assembly generalization ability. Especially when dealing with workpieces with weak texture, occluded or transparent materials, existing solutions cannot effectively decouple rotational degree of freedom information, leading to assembly failure or instability.

Method used

Visual data is acquired using a three-view camera and combined with non-visual data. Through multimodal representation and rotation axis decoupling constraints, the view-axial component is fused using gated weights. The network is optimized by combining cross-view consistency loss and prediction loss, and the robot end-effector pose fine-tuning is output to achieve six-degree-of-freedom pose correction.

Benefits of technology

It improves the stability and generalization ability of the assembly process, avoids centering jitter and jamming problems caused by rotational degree of freedom coupling, enhances compliance and safety, and improves the success rate and adaptability of robot shaft and hole assembly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374569A_ABST
    Figure CN121374569A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a system, equipment and a medium for axial decoupling and six-degree-of-freedom pose correction of a shaft hole assembly rotation shaft of a robot. The method comprises the following steps: collecting visual data and non-visual data in a shaft hole assembly process; inputting the visual data and the non-visual data into an image branch encoder and a non-image branch encoder respectively, and performing deterministic fusion through a fusion network to obtain multi-modal representation; setting a rotary axial decoupling constraint in the multi-modal representation to obtain a sub-representation channel; on the basis of multi-modal characterization, the robot tail end pose fine adjustment amount is obtained through a strategy learning model; and the pose fine adjustment amount is input into an admittance controller to generate a compliance contact control instruction, and the robot executes six-degree-of-freedom pose correction and assembly operation of shaft hole assembly based on the instruction. According to the method, the shielding or weak texture working condition adaptability, the six-degree-of-freedom pose correction precision and the assembly generalization ability are effectively improved, collision damage of parts is reduced, and the robot shaft hole assembly success rate, the convergence speed and the industrial scene adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot operation control, and in particular to a robot shaft hole assembly rotating shaft decoupling and six-degree-of-freedom pose correction method, system, device and medium. BACKGROUND

[0002] The shaft hole automatic assembly requires high accuracy of pose perception and contact compliance control capability of the robot. The traditional geometric vision scheme often relies on single-view images or RGB-D devices to complete positioning by identifying edge, circular and other features and combining consistency algorithms. In the scene with regular working conditions and no interference, stable operation can be realized. However, in the face of weak texture workpieces, part occlusion in the assembly process, or workpieces with strong glare and transparent materials, the scheme is extremely sensitive to feature thresholds and quality, and it is difficult to fully capture the rotational degree of freedom information around the x-axis (rx), y-axis (ry) and z-axis (rz), resulting in a significant decrease in pose observation accuracy. Especially for transparent material workpieces, the depth sensor is prone to systematic distortion problems, further limiting the robustness of the "only rely on RGB-D" scheme in real industrial assembly environment, increasing the cost of engineering parameter setting, and being difficult to adapt to multiple types of workpieces and complex working conditions, and lacking generalization ability.

[0003] To improve the assembly robustness, the industry gradually adopts learning-driven multi-modal perception technology to optimize the anti-interference effect of workpiece grasping and registration by fusing visual and tactile data. In the control strategy level, reinforcement learning is also applied to fine assembly tasks with high contact, showing certain operation potential. However, the existing scheme still has obvious limitations: if the upstream pose observation link cannot effectively decouple the rx, ry and rz three-axis rotation information, or the multi-view data lacks consistency constraints, it is easy to cause the downstream control strategy to have centering jitter, assembly jamming and even task failure problems, which is difficult to guarantee the success rate and stability of actual assembly. Although the multi-view learning idea can improve the representation stability to a certain extent, how to efficiently fuse three-view RGB data and tactile information without relying on depth cameras, realize rx, ry and rz axis decoupling in the observation stage, adaptively adjust the weight of the occluded and quality degraded view, and at the same time stably drive the compliance control, force control and other compliance execution strategies, a unified and engineering feasible solution has not yet been formed. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a robot shaft hole assembly rotating shaft decoupling and six-degree-of-freedom pose correction method to solve the problems of poor stability in traditional robot shaft hole assembly under complex working conditions, and the difficulty in balancing pose correction accuracy and assembly generalization ability.

[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a robot shaft-hole assembly rotation-axis decoupling and six-degree-of-freedom pose correction method, comprising: Collecting visual data and non-visual data of the shaft-hole assembly process; Inputting the visual data and the non-visual data into an image branch encoder and a non-image branch encoder respectively to obtain a visual feature vector, a pose feature vector and a motion feature vector, and determining the feature vectors through a fusion network for deterministic fusion to obtain a multi-modal representation; Setting a rotation-axis decoupling constraint in the multi-modal representation to obtain a sub-representation channel, combining the coupling relationship between the view and each rotation axis, and weighting and fusing the view-axis component through a gating weight and applying an orthogonal constraint; Combining the multi-modal representation and the non-visual data in the shaft-hole assembly process as an environmental state, inputting it into a policy learning model, and outputting a robot end pose fine-tuning amount; Inputting the pose fine-tuning amount into an admittance controller to generate a compliant contact control instruction, and the robot performs six-degree-of-freedom pose correction and assembly operation of the shaft-hole assembly based on the compliant contact control instruction.

[0007] As a preferred scheme of the robot shaft-hole assembly rotation-axis decoupling and six-degree-of-freedom pose correction method of the present application, wherein the visual data and the non-visual data of the shaft-hole assembly process are collected, comprising: The visual data includes three-view RGB image data, and the non-visual data includes robot pose data and robot motion amount data.

[0008] As a preferred scheme of the robot shaft-hole assembly rotation-axis decoupling and six-degree-of-freedom pose correction method of the present application, wherein the visual feature vector, the pose feature vector and the motion feature vector are obtained, comprising: Inputting the three-view RGB image data into a neural network model for feature extraction, the neural network model adopts a 6-layer convolutional neural network for data encoding, and processes the size change of the feature map through the convolutional layer, and finally converts the feature vector obtained by encoding into a visual feature vector through a fully connected layer; Based on the robot pose data, a 4-layer multilayer perceptron is used to encode and extract features of the position and attitude of the assembly shaft end center at the current time to generate a pose feature vector; Based on the robot motion amount data, a 2-layer multilayer perceptron is used to encode and extract features of the pose adjustment amount information at the current time to generate a motion feature vector.

[0009] As a preferred scheme of the robot shaft hole assembly rotation-axial decoupling and six-degree-of-freedom pose correction method of the application, wherein: in the multi-modal representation, a rotation-axial decoupling constraint acquisition sub-representation channel is arranged, including: The multi-modal representation is divided into three sub-representation channels corresponding to rotation around the x-axis, rotation around the y-axis and rotation around the z-axis respectively, and the dimensions of each channel are equal or approximately equal; The view-axis components corresponding to the axial direction are obtained by linear mapping of the top view, left view and right view visual coding features respectively, and the axial sub-representation is obtained by weighted fusion according to the gating weight; If any view is blocked or has a quality score lower than a threshold, the gating weight corresponding to the view is adaptively reduced.

[0010] As a preferred scheme of the robot shaft hole assembly rotation-axial decoupling and six-degree-of-freedom pose correction method of the application, wherein: the strategy learning model is based on a continuous control algorithm of maximum entropy thought, and the multi-modal representation and tactile are observed to output an end pose fine-tuning or stage switching strategy.

[0011] As a preferred scheme of the robot shaft hole assembly rotation-axial decoupling and six-degree-of-freedom pose correction method of the application, wherein: further comprising: A loss function is arranged to train and optimize the encoder and the fusion network; The loss function is used to constrain the consistency of the representations of the same physical state encoded from different image shooting angles, that is, the encoding vectors of the top view channel and the two side view channels are aligned, and the difference in the feature space is minimized; Let the three-view encoding results at the same time be Then the cross-view Figure One consistency loss can be expressed as: + , In the supervised learning part, the multi-modal feature representation and the action feature vector are spliced and input into the decoder to predict the environmental state data at the next moment. The RGB image decoding predictor adopts a four-layer deconvolutional neural network and four skip connections, and combines the up-sampling result of the action feature vector to finally decode the RGB image at the next moment; The pose decoder is composed of four layers of multilayer perceptron, which is used to predict the pose of the assembly shaft end at the next moment. According to the difference of the prediction task, the decoder output is optimized by using the endpoint error loss and the mean square error respectively, which is expressed as: + , Wherein, , are the predicted and real RGB images, , predicted and real end poses, respectively; The overall optimization objective of the fusion network is composed of the consistency loss and the prediction task loss, and is expressed as: Figure One wherein, , is a trade-off coefficient.

[0012] As a preferred scheme of the robot shaft hole assembly rotation axis decoupling and six-degree-of-freedom pose correction method of the application, wherein: The top-view camera and the two side-view cameras are arranged at the assembly station, the optical axes of the three cameras approximately correspond to the -z, -x and -y directions of the station coordinate system, and a unified reference system is aligned for the multi-view through calibration alignment; The included angle tolerance of the optical axes of the three cameras is 90°±3°; When any view is blocked or the quality score is lower than a threshold, a gating mechanism is used to automatically reduce the contribution of the view to the representation.

[0013] In a second aspect, the application provides a robot shaft hole assembly rotation axis decoupling and six-degree-of-freedom pose correction system, comprising: A data acquisition module for acquiring visual data and non-visual data of the shaft hole assembly process; A feature extraction module for inputting the visual data and the non-visual data into an image branch encoder and a non-image branch encoder, respectively, to obtain visual feature vectors, pose feature vectors and action feature vectors, and performing deterministic fusion on the feature vectors through a fusion network to obtain a multi-modal representation; A constraint processing module for setting a rotation axis decoupling constraint in the multi-modal representation to obtain a sub-representation channel, combining the coupling relationship between the view and each rotation axis, and performing gating weight weighted fusion of the view-axis component and applying an orthogonal constraint; An instruction generation module for combining the multi-modal representation and the non-visual data in the shaft hole assembly process as an environmental state, inputting the environmental state into a policy learning model, and outputting a robot end pose fine adjustment amount; inputting the pose fine adjustment amount into a steering controller to generate a compliant contact control instruction, and executing six-degree-of-freedom pose correction and assembly operation of the shaft hole assembly based on the compliant contact control instruction.

[0014] In a third aspect, the application provides an electronic device, comprising: A memory and a processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the robot shaft hole assembly rotation axis decoupling and six-degree-of-freedom pose correction method.​

[0015] In a fourth aspect, the present application provides a computer readable storage medium storing computer executable instructions, which, when executed by a processor, implement the steps of the robot shaft hole assembly rotation axis decoupling and six-degree-of-freedom pose correction method.

[0016] Compared with the prior art, the present application has the following beneficial effects: the present application deploys a three-view camera at an assembly station and synchronously collects visual data and non-visual data, obtains multi-modal representation through multi-branch coding and deterministic fusion, realizes rotation error decoupling through rotation axis decoupling constraint, combines cross-view consistency loss and prediction loss to optimize the network and improve the robustness of the representation. Figure One The present application does not need to rely on an RGB-D depth sensor, effectively adapts to complex working conditions such as occlusion and weak texture; the rotation axis decoupling design greatly improves the six-degree-of-freedom pose correction accuracy, avoids the problems of centering jitter and sticking caused by the coupling of the rotation degree of freedom; the cross-view consistency and gating mechanism enhance the feature generalization ability, which can cope with larger initial pose deviation and multi-axis hole type scenes; the maximum entropy strategy combined with the mobility control ensures the flexibility and safety of the assembly process, reduces the risk of part collision and damage, and improves the success rate, convergence speed and industrial scene adaptability of the robot shaft hole assembly. Figure One BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 The figure is a schematic diagram of the robot shaft hole assembly rotation axis decoupling and six-degree-of-freedom pose correction method of an embodiment of the present application.

[0019] Figure 2 The figure is a schematic diagram of the robot shaft hole assembly rotation axis decoupling and six-degree-of-freedom pose correction method of an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0021] Embodiment 1, refer to​Figures 1-2 For an embodiment of the present application, a robot shaft hole assembly rotation-axial decoupling and six-degree-of-freedom pose correction method is provided, comprising: S1: collecting visual data and non-visual data of the shaft hole assembly process; Preferably, the visual data includes three-view RGB image data, and the non-visual data includes robot pose data and robot action amount data.

[0022] Preferably, an overhead camera and two side-view cameras are arranged at the assembly station, the optical axes of the three cameras approximately correspond to the -z, -x, and -y directions of the station coordinate system, and a calibration alignment is used to align the unified reference system of the multi-view; the included angle tolerance of the optical axes of the three cameras is 90°±3°; when any view is blocked or the quality score is lower than a threshold, a gating mechanism is used to automatically reduce the contribution of the view to the representation.

[0023] In the embodiments of the present application, for the collected original visual data, image preprocessing operations are first performed, including cutting, normalization, smoothing processing, and noise removal operations, so as to obtain 128×128×3 RGB image data of each view of the three views.

[0024] In the embodiments of the present application, occlusion is defined as that the visible proportion of the target region in a view is insufficient or the cross-view geometric registration is significantly mismatched. The following determination is used: if the following conditions are met, it is considered as occlusion: In each view, a fixed threshold or a lightweight segmentation is used to binarize the target visible region to obtain a mask (the region is a calibrated rectangle in which the hole / insert is located determined by the camera extrinsic parameters + 90°±3° installation priori); The visible proportion is calculated as: When ⇒ occlusion, record ; otherwise .

[0025] In the embodiments of the present application, three fast indicators are taken and normalized to [0, 1] to calculate the quality score, which is represented as: wherein, represents the sharpness image; represents the exposure, 1−saturation pixel proportion (the closer to 1, the more overexposure / underexposure), represents the occlusion penalty.

[0026] S2: input the visual data and the non-visual data into the image branch encoder and the non-image branch encoder respectively to obtain a visual feature vector, a pose feature vector and a motion feature vector, and the feature vectors are deterministically fused through a fusion network to obtain a multi-modal representation; Preferably, the three-view RGB image data is input into a neural network model for feature extraction, the neural network model adopts a 6-layer convolutional neural network to encode the data, and the size change of the feature map is processed through a convolutional layer, and finally the feature vector obtained by encoding is converted into a visual feature vector through a fully connected layer; based on the robot pose data, a 4-layer multilayer perceptron is used to encode and extract features of the position and attitude of the center of the end of the assembly shaft at the current time, to generate a pose feature vector; based on the robot motion amount data, a 2-layer multilayer perceptron is used to encode and extract features of the pose adjustment amount information at the current time, to generate a motion feature vector.

[0027] In the embodiment of the application, the image data obtained by preprocessing is input into a corresponding neural network for feature extraction, a 6-layer convolutional neural network is used to encode the data, and the size change of the feature map is processed through a convolutional layer, instead of directly using a pooling layer for processing. In addition, in order to further extract potential features, a fully connected layer is added at the end of the feature extraction channel of the RGB image, which converts the feature vector obtained by encoding into a 2x128-dimensional RGB feature vector. For the pose information of the robot body, a 4-layer multilayer perceptron (MLP) is used to encode and extract features of the position and attitude (attitude represented by Euler angles) of the center of the end of the assembly shaft at the current time, and finally generate a 128-dimensional pose feature vector. For the pose adjustment amount information of the robot, a 2-layer MLP is used to encode and extract features of the current pose adjustment amount information, and finally generate a 32-dimensional motion feature vector.

[0028] In the embodiment of the application, the image branch is a three-way convolutional encoder, and the non-image branch encodes the pose / motion using a multilayer perceptron; the deterministic fusion model uses a 2-layer multilayer perceptron as a feature fusion module to complete feature extraction and fusion of the RGB and pose feature vectors, thereby learning to obtain a deterministic multi-modal feature representation.

[0029] S3: setting a rotation axis decoupling constraint in the multi-modal representation to obtain a sub-representation channel, combining the coupling relationship between the view and each rotation axis, and weighting and fusing the view-axis component through a gating weight and applying an orthogonal constraint; Preferably, the multi-modal representation is divided into three sub-representation channels corresponding to rotation around the x-axis, rotation around the y-axis, and rotation around the z-axis, respectively, each channel having equal or approximately equal dimensions; the view-axis components corresponding to the axial directions are obtained through linear mapping of the top view, left view, and right view visual encoding features, respectively, and then weighted fusion is performed according to the gating weights to obtain the axial sub-representation; if any view is blocked or has a quality score lower than a threshold, the gating weight corresponding to the view is adaptively reduced.

[0030] In the embodiment of the present application, the multi-modal representation vector output by the fusion network is , and a part of the vector is divided into three sub-representation channels corresponding to rotation around the x-axis, rotation around the y-axis, and rotation around the z-axis, respectively , wherein each channel has equal or approximately equal dimensions; and the view-axis components corresponding to the axial directions are obtained through linear mapping of the top view, left view, and right view visual encoding features, respectively (wherein , ), and then weighted fusion is performed according to the gating weights to obtain the axial sub-representation, that is: The gating weights satisfy and , and are determined by the content-related item and the prior bias item together, so that the weight of the top view channel on the rz sub-representation is higher than the weight of the side view channel, the weight of the left view channel on the ry sub-representation is higher than the weight of the remaining channels, and the weight of the right view channel on the rx sub-representation is higher than the weight of the remaining channels; and when any view is blocked or has a quality score lower than a threshold, the gating weight corresponding to the view is adaptively reduced, thereby realizing decoupling of the rx / ry / rz rotation error and improving the stability and generalization ability of the assembly strategy. The three views Figure One consistency refers to the consistency of the multi-view re-projection / regulation results of the same target under a unified reference system; the fusion regression refers to a learnable regression mapping based on multi-modal features to estimate the translation amount, without limiting the specific model structure.

[0031] S4: The multi-modal representation and the non-visual data during the axial hole assembly process are combined as the environmental state, input into the policy learning model, and the output is the robot end position fine adjustment amount; Preferably, the policy learning model is based on a continuous control algorithm based on the maximum entropy idea, and the multi-modal representation and the tactile sensation are used as observations to output the end position fine adjustment or stage switching strategy; the admittance controller realizes compliant contact and safety constraints.

[0032] In this embodiment, three-view RGB images and robot pose / motion quantities are acquired simultaneously. The data are input into the image branch encoder and the non-image branch encoder respectively. The low-dimensional multimodal representation z (128-dimensional multimodal feature vector) is obtained by deterministic fusion. The 128-dimensional multimodal feature vector is combined with the contact force information and used as the environmental state description of the shaft hole assembly process. Contact force information refers to the wrench vector output by the six-axis force / torque sensor (F / T) at the end effector, after zero drift and gravity compensation, and unified to the end effector coordinate system. , denoted as: Where F is in units of N and T is in units of Nm.

[0033] Combining multimodal characterization with non-visual data from the shaft-hole assembly process, it can be represented as: S5: Input the pose fine-tuning amount into the admittance controller to generate compliant contact control instructions. The robot performs six-degree-of-freedom pose correction and assembly operations for shaft and hole assembly based on the compliant contact control instructions.

[0034] In this embodiment of the application, cross-view representation is introduced during the training phase to ensure the consistency and robustness of multi-view representations. Figure One Consistency Loss. This loss is used to ensure that the representation of the same physical state remains consistent after encoding from different camera viewpoints. Specifically, it aligns the encoding vectors of the top-view channel with those of the side view channels, minimizing their differences in the feature space. In particular, let the encoding results of the three views at the same time be... Then cross-view Figure One Loss of inertia can be expressed as: + , This constraint enables the network to extract stable and consistent low-dimensional representations of the same assembly state under multi-view conditions. In the supervised learning part, the 128-dimensional multimodal feature representation is concatenated with the action feature vector and input into the decoder to predict the environmental state data at the next moment. The RGB image decoding predictor uses a four-layer deconvolutional neural network and four skip connections, and combines the upsampling results of the action feature vector to finally decode the RGB image at the next moment; the pose decoder consists of a four-layer multilayer perceptron (MLP) to predict the pose of the assembly axis end at the next moment. Depending on the prediction task, the decoder output is optimized using endpoint error loss (EPE) and mean squared error (MSE), respectively: + , in , are the predicted and real RGB images respectively, , are the predicted and real end-effector poses respectively.

[0035] Finally, the overall optimization objective of the deterministic fusion model is composed of the cross-view consistency loss and the prediction task loss: Figure One wherein , is the trade-off coefficient. This design not only ensures the uniformity of feature representation across different views, but also guides the model to learn effective representations related to the assembly physical process through the prediction task.

[0036] In the training stage, the next time image and end-effector pose prediction can be introduced as auxiliary tasks to improve the stability of the representation; in the inference stage, this prediction is turned off, and only the encoding and fusion are retained.

[0037] The embodiment also provides a robot shaft-hole assembly rotation axis decoupling and six-degree-of-freedom pose correction system, comprising: a data acquisition module configured to acquire visual data and non-visual data of a shaft-hole assembly process; a feature extraction module configured to input the visual data and the non-visual data into an image branch encoder and a non-image branch encoder respectively, to obtain a visual feature vector, a pose feature vector and a motion feature vector, and to perform deterministic fusion on the feature vectors through a fusion network to obtain a multi-modal representation; a constraint processing module configured to set a rotation axis decoupling constraint in the multi-modal representation to obtain a sub-representation channel, to combine a view and a coupling relationship of each rotation axis, to fuse a view-axis component through a gating weight and to apply an orthogonal constraint; an instruction generation module configured to combine the multi-modal representation and the non-visual data in the shaft-hole assembly process as an environment state, to input the environment state into a policy learning model, and to output a robot end-effector pose fine-tuning amount; and configured to input the pose fine-tuning amount into a steering controller to generate a compliant contact control instruction, and to execute a six-degree-of-freedom pose correction and assembly operation of the shaft-hole assembly based on the compliant contact control instruction.

[0038] The embodiment also provides an electronic device suitable for robot shaft-hole assembly rotation axis decoupling and six-degree-of-freedom pose correction, comprising a memory and a processor; the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to implement the robot shaft-hole assembly rotation axis decoupling and six-degree-of-freedom pose correction method proposed in the above embodiment.

[0039] ​The embodiment also provides a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method for decoupling a rotating axis and correcting a six-degree-of-freedom pose of a robot shaft hole assembly.

[0040] The storage medium provided by the embodiment belongs to the same inventive concept as the method for decoupling a rotating axis and correcting a six-degree-of-freedom pose of a robot shaft hole assembly, and the technical details not described in the embodiment can be referred to the above embodiment, and the embodiment has the same beneficial effects as the above embodiment.

[0041] Embodiment 2, refer to Figures 1-2 Based on the above embodiment, a method for decoupling a rotating axis and correcting a six-degree-of-freedom pose of a robot shaft hole assembly is provided, comprising: The assembly station is configured with three industrial cameras, which are used as a top view camera (the optical axis is approximately along the -z of the station coordinate system) and two side view cameras (the optical axes are approximately along -x and -y), and the included angle tolerance of the optical axes of the three cameras is controlled to be 90°±3°; a six-dimensional force / torque sensor is integrated at the end of the mechanical arm. A calibration plate is used to complete the internal and external parameter calibration of the three cameras, and the three view coordinates are unified to the station reference system; the quality score of each view is calculated online (based on indicators such as clarity and occlusion), and when occlusion occurs, the view is automatically down-weighted in the subsequent fusion stage. The system control cycle frequency is 100 Hz, and the three view RGB, end Cartesian pose , motion increment and six-dimensional force / torque are synchronously acquired.

[0042] For visual data, each image is first geometrically cropped and scaled to 128x128, and then normalized, lightly denoised, and randomly brightness / contrast jittered (±10%). The dataset covers multi-hole type and multi-initial bias, and about 60,000 time series samples are collected, which are divided into training / validation / testing sets at a ratio of 8:1:1; small SE(3) perturbations (translation ±1.5mm, rotation ±3°) are added in the training stage to enhance the geometric robustness. The pose is expressed by Cartesian translation and Euler angle, and the motion is the difference between the end poses at adjacent time points. The tactile data, image / pose are strictly aligned by timestamp.

[0043] The perception network adopts a multi-branch coding and deterministic fusion structure. Each of the three image branches is composed of a 6-layer convolutional network, and the end of each branch is connected to a fully connected layer to obtain a 128-dimensional vector. In the non-image branch, the pose is mapped to 128 dimensions by a 4-layer MLP, and the motion is obtained by a 2-layer MLP to obtain a 32-dimensional vector. The three image features, pose features, and motion features are concatenated and input into a 2-layer MLP to obtain a 128-dimensional multi-modal latent representation .

[0044] To achieve observable decoupling of six degrees of freedom, it is preferable to... Explicitly divided into ,in 16 dimensions each (48 dimensions in total) A total of 48 dimensions, with the remaining 32 dimensions used for task-related residuals. For views In the axial direction The amount Soft gating fusion is employed: in, Encoding features for the view, For axial query vectors, For prior bias (viewed from above) Translation in the plane ( The left eye has a higher weight; the right eye has a higher weight. and( The right side has a higher weight; the right side has a higher weight. and( (Higher weighting) This is the mass gain coefficient.

[0045] During the training phase, "predicting the next moment of the action condition" is used as explicit supervision. This involves a 128-dimensional... The image decoder, concatenated with action features, is fed into the decoder. The image decoder consists of 4 deconvolutional layers with 4 skip connections to predict the RGB values ​​at the next time step. The pose decoder consists of 4 MLP layers to predict the end-effector pose. The loss function is derived from cross-view... Figure One Consistency term (constraining different views at the same time) The target is composed of distance), image endpoint error (EPE), and pose mean square error (MSE), and is superimposed with gating regularization and axial orthogonality regularization. The total target can be written as , The optimizer uses Adam with a learning rate of 1×10⁻ 4 Batch size 32, training 100k steps, first 5k steps linear warm-up; during the inference phase, the decoder is turned off, and only the 128-dimensional representation obtained from encoding and fusion is retained. Use online.

[0046] Strategy learning A 134-dimensional vector is used as the state input, and a continuous control algorithm (SAC) based on the maximum entropy concept is employed. The Actor / Critic network width is 256-256, with a discount factor... , soft update coefficient , target entropy is set to -6 (for six-dimensional motion), replay buffer capacity 5x10 5 , update once per step. The policy output is the end pose fine-tuning amount, , single-step limit ±1.5mm, ±1°. The policy output is input to the admittance controller after being limited by the trajectory generator and limited acceleration smoothing. Thus, stable and generalizable six-degree-of-freedom error decoupling perception and compliant control can be achieved in complex assembly scenarios with large initial pose deviations and multi-hole type changes.

[0047] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and necessary general hardware, and of course can also be realized by hardware. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a FLASH, a hard disk, or an optical disc, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various embodiments of the present application.

[0048] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A robot shaft hole assembly rotation-axial decoupling and six-degree-of-freedom pose correction method, characterized in that, The method comprises the following steps: Collecting visual data and non-visual data of the shaft-hole assembly process; Inputting the visual data and non-visual data into an image branch encoder and a non-image branch encoder respectively to obtain visual feature vectors, pose feature vectors and action feature vectors, and determining the feature vectors by a fusion network to obtain a multi-modal representation; Setting a rotational axis decoupling constraint in the multi-modal representation to obtain a sub-representation channel, combining the coupling relationship between the view and each rotational axis, and weighting and fusing the view-axis components by a gating weight and applying an orthogonal constraint; Combining the multi-modal representation and the non-visual data in the shaft-hole assembly process as an environmental state, inputting the environmental state into a policy learning model, and outputting a robot end pose fine adjustment amount; Inputting the pose fine adjustment amount into an admittance controller to generate a compliant contact control instruction, and executing six-degree-of-freedom pose correction and assembly operation of the shaft-hole assembly based on the compliant contact control instruction.

2. The robotic shaft-hole assembly rotation-to-axial decoupling and six-DOF pose correction method of claim 1, wherein, The method comprises the following steps: The visual data includes three-view RGB image data, and the non-visual data includes robot pose data and robot action amount data.

3. The robotic shaft-hole assembly spin-to-axial decoupling and six-DOF pose correction method of claim 2, wherein, The method comprises the following steps: Inputting the three-view RGB image data into a neural network model for feature extraction, using a 6-layer convolutional neural network to encode the data, processing the size change of the feature map through the convolutional layer, and finally converting the feature vector obtained by encoding into a visual feature vector through the fully connected layer; Based on the robot pose data, using a 4-layer multilayer perceptron to encode and extract features of the position and attitude of the assembly shaft end center at the current time to generate a pose feature vector; Based on the robot action amount data, using a 2-layer multilayer perceptron to encode and extract features of the pose adjustment amount information at the current time to generate an action feature vector.

4. The robotic shaft bore fitment spin-to-axial decoupling and six-DOF pose correction method of claim 3, wherein, The method comprises the following steps: The method comprises the following steps: The multi-modal representation is divided into three sub-representation channels corresponding to rotation around the x-axis, rotation around the y-axis and rotation around the z-axis, and the dimensions of each channel are equal or approximately equal; The three-view visual encoding features are linearly mapped to obtain view-axis components corresponding to the axis, and then the axis sub-representation is obtained by weighting and fusing the view-axis components according to the gating weight; 5. The robotic shaft bore fitment spin-to-axial decoupling and six-DOF pose correction method of claim 1, wherein, If any view is blocked or has a quality score lower than a threshold, the gating weight corresponding to the view is adaptively reduced.

6. The robotic shaft bore fitment spin-to-axial decoupling and six-DOF pose correction method of claim 1, wherein, The policy learning model is based on a continuous control algorithm of the maximum entropy idea, and outputs an end pose fine adjustment or stage switching strategy based on the multi-modal representation and the tactile sensation. The method further comprises the following steps: Setting a loss function to train and optimize the encoder and the fusion network; The loss function is used to constrain the representations of the same physical state encoded from different image shooting angles to be consistent, that is, the encoding vectors of the top view channel and the two side view channels are aligned, and the difference between them in the feature space is minimized. Let the three-view encoding results at the same time be Then the cross-view consistency loss can be expressed as: + , In the supervised learning part, the multi-modal feature representation and the action feature vector are concatenated and input into the decoder to predict the next time step of the environment state data. The RGB image decoding predictor adopts a four-layer de-convolutional neural network and four skip connections, and combines the up-sampling results of the action feature vector to finally decode the next time step of the RGB image. The pose decoder is composed of four layers of multilayer perceptron, which is used to predict the next time step of the end of the assembly axis. According to the different prediction tasks, the decoder output is optimized by the endpoint error loss and the mean square error, respectively, which is represented as: + , wherein, , are predicted and real RGB images, respectively, , are predicted and real end pose, respectively. The overall optimization objective of the fusion network is composed of the cross-view consistency loss and the prediction task loss, which is represented as: wherein , is a trade-off coefficient.

7. The robotic shaft-hole assembly spin-to-axial decoupling and six-DOF pose correction method of claim 1 or 2, wherein, Also includes: The overhead camera and the two side-view cameras are arranged at the assembly station, and the optical axes of the three cameras approximately correspond to the -z, -x and -y directions of the station coordinate system. The unified reference system of the multi-view is aligned by calibration alignment. The included angle tolerance of the optical axes of the three cameras is 90°±3°. When any view is blocked or the quality score is lower than the threshold, a gating mechanism is used to automatically reduce the contribution of the view to the representation.

8. A robotic shaft-hole assembly rotation-axial decoupling and six-degree-of-freedom pose correction system applying the method of any one of claims 1-7, characterized in that, It includes: A data acquisition module for acquiring visual data and non-visual data of the shaft hole assembly process; A feature extraction module for inputting visual data and non-visual data into an image branch encoder and a non-image branch encoder respectively to obtain visual feature vectors, pose feature vectors and action feature vectors. The feature vectors are deterministically fused by a fusion network to obtain a multi-modal representation. A constraint processing module for setting a rotational axis decoupling constraint in the multi-modal representation to obtain a sub-representation channel, combining the coupling relationship between the view and each rotational axis, and weighting and fusing the view-axis component through a gating weight and applying an orthogonal constraint. An instruction generation module for combining the multi-modal representation and the non-visual data in the shaft hole assembly process as an environment state, inputting it into a policy learning model, and outputting a robot end pose fine adjustment amount; inputting the pose fine adjustment amount into a steering controller to generate a compliant contact control instruction, and the robot performs six-degree-of-freedom pose correction and assembly operation based on the compliant contact control instruction.

9. An electronic device, comprising: It includes: A memory and a processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the robot shaft hole assembly rotational axis decoupling and six-degree-of-freedom pose correction method of any one of claims 1 to 7 are realized.

10. A computer-readable storage medium, characterized in that, It stores computer executable instructions, which are executed by the processor to realize the steps of the robot shaft hole assembly rotational axis decoupling and six-degree-of-freedom pose correction method of any one of claims 1 to 7.

Citation Information

Cited By

  • Mechanical arm self-adaptive grading grabbing control method and system based on visual feedback

    CN121989258A