A vision-guided robotic arm assembly method and system, and a computer-readable medium

By using 2D cameras and deep learning technology, the problem of error accumulation in the assembly of vision-guided robotic arms was solved, achieving high-precision and robust assembly results, and reducing equipment costs and deployment difficulty.

CN120620242BActive Publication Date: 2025-12-02PUDA DITAI (CHENGDU) INTELLIGENT MFG RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511140760.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-02
Estimated Expiration
2045-08-15

Smart Images

  • Figure CN120620242B_ABST
    Figure CN120620242B_ABST
Patent Text Reader

Abstract

This application provides a vision-guided robotic arm assembly method and system, and a computer-readable medium, which can solve the technical problem of poor accuracy and stability in guided assembly in related technologies. The method includes: determining an initial detection point; obtaining a first center position based on the position of the target to be assembled in a first camera image; using the first center position to guide the collaborative arm to move so that the position of the target to be assembled in the camera image is a first preset center of the camera image; obtaining a feature localization result based on the geometric features of the target to be assembled in a second camera image; obtaining a second center position based on the feature localization result and a second preset center; using the second center position to make the position of the target to be assembled in the camera image a second preset center of the camera image; and obtaining a guided assembly path based on the initial detection point, the first center position, and the second center position. Using a 2D camera for object recognition and feature detection improves the accuracy and stability of the guided assembly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of equipment assembly technology, and in particular to a vision-guided robotic arm assembly method and system, and a computer-readable medium. Background Technology

[0002] In the field of industrial automation, especially in high-precision assembly, using vision systems to guide robotic arms to grasp, position, and assemble workpieces has become a key technology for improving production efficiency and ensuring product quality. This vision-guided assembly technology senses the spatial position and posture of the target workpiece, calculates in real time, and guides the robotic arm to move to the predetermined target pose, effectively replacing traditional manual operation or reliance on high-precision tooling fixtures, and significantly improving production flexibility and automation levels.

[0003] Currently, such as Figure 1 As shown, vision-based robotic arm guidance methods typically include the following core steps: First, "eye-in-hand" or "eye-to-hand" calibration is performed to determine the precise transformation relationship between the camera coordinate system installed at the end effector or fixed position of the robotic arm and the robotic arm's base coordinate system. Second, in the offline phase, through a "pre-teaching" process, the robotic arm is manually guided to carry a camera (or a fixed camera observes the workpiece being taught) to acquire multi-angle views of a standard workpiece or assembly position, and a high-precision "point cloud template" is established based on this. This template represents the ideal three-dimensional geometric model of the target workpiece and its reference pose in the calibration coordinate system. During actual operation, the system acquires point cloud data of the workpiece to be operated in real time through the camera, and uses "point cloud registration" algorithms (such as ICP, feature matching, etc.) to match and align the real-time point cloud with the pre-established "point cloud template," calculating the pose transformation of the real-time workpiece relative to the reference template. Finally, combined with the eye-in-hand calibration relationship, this pose transformation is converted into a target pose command in the robotic arm's base coordinate system, thereby "guiding the robotic arm" to perform precise grasping or assembly actions.

[0004] However, the aforementioned existing technical solutions involve multiple stages during implementation (hand-eye calibration, teaching process, point cloud acquisition and processing, registration algorithm, etc.), and each stage inevitably introduces certain error variables. These errors risk accumulating and amplifying in subsequent steps. For example, the accuracy of hand-eye calibration is affected by the accuracy of the calibration board, the calibration algorithm, and the operation. The quality of the point cloud template established during the initial teaching depends on the workpiece's state, lighting conditions, and the stability of the point cloud reconstruction algorithm during teaching. Real-time point cloud acquisition is easily affected by workpiece surface reflection, texture, occlusion, and changes in ambient lighting. The convergence, robustness, and accuracy of the point cloud registration algorithm are also constrained by point cloud quality, initial pose estimation, and algorithm parameter selection. These interrelated variables work together to cause the final positioning accuracy of the entire guidance system to be unstable. Especially when dealing with workpieces with complex geometries, low texture features, or reflective surfaces, as well as the changing lighting and vibration environments in the production site, existing technical solutions often fail to consistently guarantee high-precision and high-robust assembly results, limiting their reliable application in scenarios with higher precision requirements. There is an urgent need for a new method that can effectively suppress error accumulation and improve the overall precision and stability of the system. Summary of the Invention

[0005] This application provides a vision-guided robotic arm assembly method and system, which can solve the technical problem of poor accuracy and stability in guided assembly in related technologies.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, embodiments of this application provide a vision-guided robotic arm assembly method, which includes: determining a detection initial point; the detection initial point being the position information of a collaborative arm in an initial posture; the initial posture being that the first surface of the collaborative arm is parallel to a preset plane of the target to be assembled; obtaining a first center position based on the position of the target to be assembled in a first camera image; the first center position being used to guide the collaborative arm to move so that the position of the target to be assembled in the camera image is a first preset center of the camera image, denoted as a second camera image; obtaining a feature localization result based on the geometric features of the target to be assembled in the second camera image; obtaining a second center position based on the feature localization result and the second preset center; the second center position being used to guide the collaborative arm to move so that the position of the target to be assembled in the camera image is a second preset center of the camera image; and obtaining a guided assembly path based on the detection initial point, the first center position, and the second center position.

[0008] Based on the above description of the vision-guided robotic arm assembly method provided in this application embodiment, it can be seen that this method includes reducing the guidance of the collaborative arm from six degrees of freedom to three degrees of freedom during teaching by limiting the initial posture of the collaborative arm (i.e., the initial point for subsequent detection). This eliminates the need for re-detection of the posture and, consequently, the need for point cloud acquisition and hand-eye guidance using a 3D camera, thus avoiding processes such as hand-eye calibration, point cloud template establishment, and point cloud registration. Secondly, a 2D camera is used for object recognition and feature detection, avoiding errors introduced during 3D point cloud imaging and registration, achieving an accuracy of three pixels, far exceeding that of a 3D camera. Finally, by avoiding the point cloud processing and hand-eye calibration processes, on-site deployment by professional vision engineers is unnecessary, significantly reducing deployment difficulty and greatly increasing deployment speed. This, in turn, improves the accuracy and stability of the guided assembly.

[0009] Furthermore, replacing 3D cameras with 2D cameras significantly reduces equipment investment costs from a hardware perspective. Simultaneously, the elimination of the need for high-performance GPUs or dedicated point cloud processing chips significantly reduces the computing power requirements of the robotic arm control system, allowing stable operation with a standard industrial control motherboard. This optimization not only reduces initial equipment procurement expenditures but also lowers long-term energy consumption costs.

[0010] In a feasible implementation of the first aspect, when performing the step of obtaining the first center position based on the position of the target to be assembled in the first camera image, the vision-guided robotic arm assembly method further includes: obtaining the position of the target to be assembled in the first camera image through deep learning.

[0011] In the feasible implementation of the first aspect, the vision-guided robotic arm assembly method further includes: determining an assembly object dataset; the assembly object dataset includes images of normal assembly objects, tilted assembly objects, damaged assembly objects, and missing assembly objects; expanding the assembly object dataset through symmetry transformation, scale transformation, and rotation transformation to obtain an expanded dataset; training a first learning model based on the expanded dataset; inputting a first camera image into the first learning model to obtain category results; filtering the category results to obtain target results; calculating the center position of the bounding box corresponding to the result with the highest confidence among the target results; and calculating a first center position based on the center position of the bounding box.

[0012] In one feasible implementation of the first aspect, the vision-guided robotic arm assembly method further includes: converting the first camera image to grayscale to obtain a black and white image; equalizing the black and white image according to adaptive histogram equalization to obtain an equalized image; normalizing the scale of the equalized image to obtain a normalized image; and obtaining the category result based on the normalized image.

[0013] In the feasible implementation of the first aspect, the vision-guided robotic arm assembly method also includes: obtaining feature localization results through feature detection algorithms.

[0014] In the feasible implementation of the first aspect, the vision-guided robotic arm assembly method further includes: preprocessing the second camera image through bilateral nonlinear filtering and adaptive thresholding to obtain a smooth image; and using edge drawing to detect image edges to achieve high-precision feature detection and obtain feature localization results.

[0015] In the feasible implementation of the first aspect, the vision-guided robotic arm assembly method further includes: filtering the image using a Gaussian function based on the image size, and obtaining the gradient extreme points in the X and Y directions by difference, which are used as drawing anchor points; filtering anchor points by introducing gradient values ​​and gradient angles, while limiting the minimum line length, to obtain target anchor points; and connecting the target anchor points using an addition strategy.

[0016] In the feasible implementation of the first aspect, the vision-guided robotic arm assembly method further includes: based on the symmetry of the kernel function, calculating 6 sets of kernel functions for a 5×5 window size, and simultaneously performing normalization; based on the discrete value characteristic of grayscale values, calculating the product of 0-255 and the normalized kernel function to obtain a 6×256 lookup table; using the sum of absolute position values ​​as the row number and the pixel grayscale as the column number, querying the value in the Gaussian filter function lookup table; and for changes in the size of the target to be assembled, guiding the robotic arm to adjust the depth direction, converting the pixel ratio into actual three-dimensional spatial depth changes.

[0017] Secondly, embodiments of this application provide a vision-guided robotic arm assembly system, which includes: at least one processor; a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method provided in the first aspect.

[0018] Thirdly, embodiments of this application provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the method provided in the first aspect. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the relevant technologies;

[0020] Figure 2 This is a schematic diagram illustrating an application scenario of a vision-guided robotic arm assembly system provided in an embodiment of this application.

[0021] Figure 3 A flowchart illustrating a vision-guided robotic arm assembly method provided in this application embodiment;

[0022] Figure 4 A schematic diagram of the first center position in a vision-guided robotic arm assembly method provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of the structure of the first learning model in a vision-guided robotic arm assembly method provided in an embodiment of this application;

[0024] Figure 6 A schematic diagram of block division in a vision-guided robotic arm assembly method provided in an embodiment of this application;

[0025] Figure 7 A flowchart illustrating a vision-guided robotic arm assembly method provided in this application embodiment;

[0026] Figure 8 This is a structural schematic diagram of a vision-guided robotic arm assembly system provided in an embodiment of this application. Detailed Implementation

[0027] The principles and features of this application are described below. The examples given are only for explaining this application and are not intended to limit the scope of this application.

[0028] This application provides a vision-guided robotic arm assembly method applicable to the field of guided assembly. For example, the assembly target is a standard product such as aluminum alloy or cast iron, which can have the same shape, color, and surface treatment process.

[0029] Figure 2 This is a schematic diagram illustrating an application scenario of a vision-guided robotic arm assembly system provided in an embodiment of this application. For example... Figure 2 As shown, in some embodiments, the vision-guided robotic arm assembly system includes: an assembly mechanism, a camera positioned directly above the assembly mechanism, and a six-degree-of-freedom robotic arm that controls the movement of the assembly mechanism. One end of the six-degree-of-freedom robotic arm is fixed to the ground, and the other end is connected to the assembly mechanism and the camera via connecting rods.

[0030] The six-degree-of-freedom robotic arm comprises a first link, a second link, a third link, and a fourth link connected in sequence. The first link is fixed to the ground, and the fourth link is connected to a connecting link. Variable angles are formed between the first and second links, between the second and third links, and between the third and fourth links. The fourth link serves as a cooperating arm.

[0031] Using deep learning methods, a camera mounted on the end effector of a robotic arm enables vision-guided assembly tasks. The arm is moved to ensure the device to be assembled is within the camera's field of view. A visual detection method is used to obtain the offset between the device and the center of the camera's field of view. Based on this offset, the robotic arm is automatically moved to position the device to be assembled at the center of the camera's field of view.

[0032] like Figure 2 As shown, in one implementation, the vision-guided robotic arm assembly system includes an assembly execution unit and a robotic arm motion unit.

[0033] The assembly execution unit includes an assembly mechanism 101 and an industrial camera 102.

[0034] Assembly mechanism 101 uses pneumatic grippers with an adjustable clamping force range of 0-20N.

[0035] The industrial camera 102 is rigidly fixed above the assembly mechanism by an aluminum alloy bracket 102a. The parallelism error between the camera optical axis and the central axis of the assembly mechanism is ≤0.05°, and the working distance is set to 150±5mm.

[0036] The assembly execution unit also includes a ring-shaped LED supplementary light source. This ring-shaped LED supplementary light source is integrated around the camera lens and has a programmable brightness control of 0-3000 lux.

[0037] The robotic arm motion unit includes a six-degree-of-freedom robotic arm 200 and connecting rods 300.

[0038] A six-degree-of-freedom robotic arm 200 comprises four links connected in series: the first link 201 is fixed to the ground base 200a via a flange, and its length is... =350mm; the second member 202 is connected to the first member via the first rotary joint J1, the joint axis is perpendicular to the ground, and the length is... =420mm, forming the first included angle ∈[0°,180°]; The third member 203 is connected to the second member via the second rotary joint J2, with the joint axis arranged horizontally and the length... =380mm, forming the second included angle ∈[-90°,+90°]; The fourth member 204 is connected to the third member through the third rotary joint J3, the joint axis is perpendicular to the second joint axis, and the length is... =250mm, forming the third included angle ∈[-180°,+180°].

[0039] The connecting rod 300 is a carbon fiber square tube with a cross section of 20×20mm. One end is connected to the end of the fourth rod through a universal joint, and the other end is bolted to the assembly execution unit integrated body.

[0040] In this way, with the assistance of the industrial camera 102, the equipment to be assembled 400 adjusts the position of each link in the six-degree-of-freedom robotic arm 200, and then completes the guided assembly through the assembly mechanism 101.

[0041] Figure 3 This is a flowchart illustrating a vision-guided robotic arm assembly method provided in an embodiment of this application. Figure 3 As shown, in some embodiments, the vision-guided robotic arm assembly method includes the following steps:

[0042] S1, determine the initial detection point.

[0043] The initial detection point is the position information of the collaborative arm in its initial posture.

[0044] like Figure 2 As shown, the initial orientation is that the first surface 204a of the collaborative arm is parallel to the preset plane of the target 400 to be assembled.

[0045] In some embodiments, the preset plane of the target 400 to be assembled may be the geometric center plane of the target 400 to be assembled.

[0046] S2, based on the position of the target to be assembled in the first camera image, obtain the first center position.

[0047] The first camera image is as follows: Figure 2 The image shown is within the camera's field of view.

[0048] like Figure 4 As shown, the first center position (e.g.) or This is used to guide the movement of the collaborative arm so that the position of the target to be assembled in the camera image is the first preset center of the camera image, denoted as the second camera image.

[0049] In some embodiments, when performing step S2, the vision-guided robotic arm assembly method further includes:

[0050] S21. Through deep learning, the position of the target to be assembled in the first camera image is obtained.

[0051] Deep learning centering mainly uses deep learning algorithms to identify assembly targets, especially for detecting and centering defective targets that are not completely in the field of view.

[0052] In some embodiments, when performing step S21, deep learning includes a training process and a detection process, and the vision-guided robotic arm assembly method further includes:

[0053] S211, Determine the assembly object dataset.

[0054] The assembly object dataset includes images of normal assembly objects, tilted assembly objects, damaged assembly objects, and missing assembly objects. This allows for the use of deep learning algorithms to identify assembly targets, particularly for detecting and centering missing targets that are not fully within the field of view.

[0055] The assembly object dataset can be an image dataset of common assembly target objects (such as various threaded holes), containing images of various conditions such as normal, tilted, damaged, and missing, which are then labeled to form the dataset. For example, ... Figure 4 As shown, the assembly target object can be a missing image located in the upper left corner of the image's field of view, meaning that only a portion of common assembly target objects appear in the image. Alternatively, the assembly target object can be a normal image located within the image's field of view, meaning that all common assembly target objects appear in the image.

[0056] S212 expands the assembly object dataset by performing symmetry transformation, scaling transformation, and rotation transformation to obtain an expanded dataset.

[0057] To reduce the difficulty of data collection, the data was expanded.

[0058] S213, Train the first learning model based on the expanded dataset.

[0059] The first learning model can be the YOLO model.

[0060] The dataset is divided into a training set and a validation set. The YOLO model is trained using images from the training set, and then tested using the validation set after training.

[0061] like Figure 5 As shown, the first learning model, namely the YOLO model, includes a backbone feature extraction network, a classifier, a regressor, and a reinforcement feature extraction network.

[0062] Starting with the inputs, a 640×640×3 input layer receives RGB three-channel color image data, serving as the starting point for the entire process. This is followed by the backbone feature extraction network (CSPDarknet), where the Focus layer (320×320×12, Focus) acts as a special downsampling and feature reconstruction structure. After receiving the input, it outputs a 320×320×12 feature map, reducing computational load while preserving image details, laying the foundation for subsequent operations.

[0063] In the backbone network, the 2-Dimensional Convolutional-Batch Normalization-Sigmoid-Linear Unit Layer (Conv2D_BN_SiLU) appears repeatedly. The 2-Dimensional Convolutional Layer (Conv2D) extracts local features by sliding the convolution kernel. The size and number of channels of the output feature map are represented by parameters (e.g., 20×20×64), with the first two digits representing the size and the third digit representing the number of channels. Batch Normalization (BN) normalizes the convolution results, accelerating training and improving stability. The Sigmoid-Linear Unit (SiLU) serves as the activation function, achieving a smooth nonlinear transformation through SiLU(x) = x·sigmoid(x), which is beneficial for gradient propagation.

[0064] The Cross-Stage Partial Layer (CspLayer) is based on CSPNet. After splitting the feature map, some parts are convolutional and some are directly connected, which enhances feature fusion while reducing computation. For example, the CspLayer layer that outputs a 160×160×128 feature map plays this role. The Spatial Pyramid Pooling Bottleneck (SPPBottleneck) combines Spatial Pyramid Pooling (SPP) and the Bottleneck layer. SPP performs multi-scale pooling and concatenates the results to extract multi-scale contextual features. The Bottleneck uses a "narrow-wide-narrow" channel design to control computation. The combination of the two enhances feature representation and outputs a 20×20×1024 feature map. Through these components, the backbone network gradually extracts multi-scale, deep features from the original image and feeds them into subsequent networks.

[0065] The enhanced feature extraction network PN (PN) first merges the channel dimension information of feature maps at different scales through concatenation (Concatenate, Cross-Stage Partial Layer, Concat+CSPLayer) in the Concatenate-Cross-Stage Partial Layer, and then strengthens the association through the CSP structure. The UpSampling layer enlarges the feature map size through interpolation, restoring low-resolution, high-semantic feature maps to a larger size and supplementing details. The DownSampling layer shrinks the feature map and increases the number of channels by using convolution stride or pooling, allowing high-resolution, low-semantic feature maps to incorporate more semantics. The PN network repeatedly merges multi-scale features through these operations, strengthening semantic and detailed information, resulting in more representative features for subsequent outputs.

[0066] Finally, the classifier and regressor (YoloHead) are used, which is further divided into classification (CIs), regression (Reg), and object confidence (Obj) branches. The classification branch processes and outputs the object class probability, the regression branch predicts the target bounding box coordinate offset for precise localization, and the object confidence branch determines the probability of the object's presence at the feature map location, distinguishing between foreground and background. YoloHead relies on the features output from the Neck to predict the three classes of information in parallel, completing the output of the object detection result. The backbone feature extraction network includes a Focus network structure. In the Focus network structure, a value is obtained for every single pixel in an image, resulting in four independent feature layers. These four independent feature layers are then stacked, concentrating the width and height information into channel information, expanding the input channels fourfold. The concatenated feature layer becomes twelve channels instead of the original three.

[0067] BaseConv is a term used in YOLOX networks. It includes convolution (Conv), batch normalization layer (BN), and SiLu activation function. Convolution operation is mainly responsible for feature extraction in the network and is one of the most important operations of the model.

[0068] Batch Normalization (BN) ensures that the output of each layer is as consistent as possible with the input data distribution of the next layer, making the model more stable during training. Activation functions provide the network with the ability to undergo non-linear changes, enabling hierarchical, progressively abstracting features in deep models.

[0069] Using the SiLU activation function, which is unbounded at the upper limit but bounded at the lower limit, smooth, and non-monotonic, SiLU outperforms ReLU in deep models. It can be viewed as a smooth ReLU activation function; the activation function is continuous and differentiable, and its goal is to make the neural network nonlinear. The activation function is bounded at the lower limit but unbounded at the upper limit. The lower bound avoids slow convergence caused by zero gradients during network training and also facilitates the regularization of network parameters. Since the activation function itself is nonlinear, introducing it into a neural network allows the network to arbitrarily approximate nonlinear functions, thereby enhancing the expressive power of deep neural networks.

[0070] UpSample refers to upsampling, which enlarges the image.

[0071] DownSample is for downsampling, which reduces the size of the image; Concat is for stitching together.

[0072] YoloHead is used to detect feature pyramids (generated from the middle section), which consists of a combination of convolutional layers, pooling layers, and fully connected layers. YoloHead, through CSPDarknet and FPN, can obtain three enhanced effective feature layers. Each feature layer has width, height, and number of channels. We can then view the feature map as a collection of feature points, each with a number of channels.

[0073] The function of YoloHead is to determine whether a feature point corresponds to an object. Using the FPN feature pyramid, we can obtain three enhanced features with shapes of (20, 20, 1024), (40, 40, 512), and (80, 80, 256), respectively. We then use these three feature layers to obtain the prediction results from YoloHead.

[0074] CBS describes the Convolution (Conv), Batch Normalization (BN), and SiLU activation functions.

[0075] CSPLayer is the core module of CSPDarknet, typically containing multiple convolutional layers (such as 3x3 convolutions) and residual connections. In this way, the model can learn complex features more effectively.

[0076] The SPP structure extracts features through max pooling with different kernel sizes, thereby increasing the network's receptive field. The receptive field refers to the region of the input image that a point on the feature map can see; that is, a point on the feature map is calculated from the receptive field size region in the input image. A larger neuron's receptive field value indicates a larger range of the original image it can access. This solves the problem of inconsistent input image sizes by fusing multiple receptive fields through three different pooling operations.

[0077] S214, input the first camera image into the first learning model to obtain the category result.

[0078] S215, filter the product category results to obtain the target results.

[0079] S216, calculate the center position of the bounding box corresponding to the result with the highest confidence level in the target results.

[0080] S217, Calculate the first center position based on the center position of the bounding box.

[0081] Guide the collaborative arm to move so that the assembly object returns to approximately the center of the field of view.

[0082] To improve training quality and speed, in some embodiments, the vision-guided robotic arm assembly method further includes the following steps before performing step S21:

[0083] S221, convert the first camera image to grayscale to obtain a black and white image.

[0084] Using grayscale images for training and detection reduces the training dimension, dropping from the three-dimensional RGB space to a black and white color space, thus minimizing the influence of color. Black and white images are binary images.

[0085] S222, equalize the black and white image using adaptive histogram equalization to obtain an equalized image.

[0086] Adaptive histogram equalization is used to equalize the illumination distribution density function, avoiding excessively dark or bright areas and reducing the impact of illumination.

[0087] S223, scale normalize the equalized image to obtain a normalized image.

[0088] The camera and images were selected and cropped using a 4:3 aspect ratio in the model to avoid the automatic stretching during training, which would increase the aliasing between features in the image.

[0089] like Figure 6As shown, the adaptive histogram equalization algorithm using bilinear interpolation improves the overall and local edge quality of the image while also considering processing speed. The adaptive histogram equalization algorithm first divides the image into blocks as shown in the figure below: edge regions, edge corner regions, and center regions. It then calculates the gray-level histogram and gray-level distribution function for each sub-region.

[0090] like Figure 6 As shown, in one implementation, pixels in the red region (edge ​​corner region) are grayscale mapped according to the transformation function of their respective subimages. Pixels in the green region (edge ​​region) are obtained by linear interpolation after being transformed by the transformation functions of their two adjacent subimages. Pixels in the purple region (center region) are obtained by bilinear interpolation after being transformed by the transformation functions of their four adjacent subimages.

[0091] Arrows represent direction and change trends, indicating the direction of movement or mapping of pixels or regions.

[0092] for example:

[0093] Blue double-headed arrows: represent "correspondence", indicating that the area or pixel in the large grid on the left will be mapped to the corresponding position in the small grid on the right according to the direction of the arrow, reflecting the coordinate mapping logic when the image is scaled or transformed (such as how the pixels in the large grid "move" to the small grid for rearrangement).

[0094] Green double-headed arrows: Emphasize the "direction of transmission", such as extracting features from the original region (pink or green block) and transmitting them to other regions for processing according to the direction of the arrow (such as extracting features first and then mapping them when scaling).

[0095] The solid black squares are "anchor points" or "key reference points," which can mark fixed positions or serve as references for feature extraction. In image transformations, the coordinates of these points are known and fixed, acting as "reference points" for scaling or mapping to ensure alignment of key positions before and after the transformation (e.g., a solid square in a large grid corresponds to specific coordinates in a small grid, ensuring structural integrity during image scaling). The arrows around these points represent pixels that need to be preserved or processed; for example, their grayscale and color are key features during image transformation. The arrows convey information around them, ensuring that features are not lost after the transformation.

[0096] S224. Based on the normalized image, obtain the category results.

[0097] Targeted image preprocessing reduces the difficulty of model training, allowing the model to converge in fewer iterations and reducing training time. Preprocessed images also improve the accuracy of subsequent feature detection.

[0098] In this way, by executing steps S221 to S224, a unified preprocessing process is performed on the training set and the detection images, so that the images to be trained and detected focus on their structural differences, thereby reducing the weight of influence from ambient lighting, color, etc.

[0099] S3. Based on the geometric features of the target to be assembled in the second camera image, obtain the feature localization result.

[0100] Feature detection centering uses a 2D camera to detect the geometric features of the target to be assembled, thereby achieving the localization of the target to be assembled.

[0101] In some embodiments, when performing step S3, the vision-guided robotic arm assembly method further includes:

[0102] S31, the feature localization result is obtained through the feature detection algorithm.

[0103] In addition to the image preprocessing performed in step S2, further processing of the two-dimensional (2D) image is required, including bilateral nonlinear filtering and adaptive thresholding.

[0104] In some embodiments, when performing step S31, the vision-guided robotic arm assembly method further includes:

[0105] S311 preprocesses the second camera image using bilateral nonlinear filtering and adaptive thresholding to obtain a smooth image.

[0106] Bilateral nonlinear filtering works by smoothing the image while preserving edges, reducing the impact of Gaussian noise, thermal noise, and other contaminants.

[0107] S312 uses edge rendering to detect image edges to achieve high-precision feature detection and obtain feature localization results.

[0108] High-precision feature detection is achieved by using edge rendering to detect image edges. The edge rendering accuracy is ±1 pixel. Considering the accuracy deviation in both the horizontal and vertical directions of the feature center, the maximum error is approximately 2.82 pixels, thus the accuracy is better than 3 pixels. It can be understood that the deviation of the feature center is essentially a deviation in the two-dimensional plane (the xy coordinate system of the image), which can be decomposed into deviations in two independent dimensions: the horizontal direction (x-axis) and the vertical direction (y-axis).

[0109] In some embodiments, when performing step S31, the vision-guided robotic arm assembly method further includes:

[0110] S321, based on the image size, a Gaussian function is used for filtering, and its weighting function is shown below:

[0111] ;

[0112] Where G(x,y) represents the function value of the two-dimensional Gaussian function at coordinates (x,y), used to describe the relative weight or probability density of that position in the Gaussian distribution; x represents the horizontal coordinate variable of a point in the two-dimensional plane, usually referring to the horizontal offset of that point relative to the center of the Gaussian distribution; y represents the vertical coordinate variable of a point in the two-dimensional plane, usually referring to the vertical offset of that point relative to the center of the Gaussian distribution; π represents the mathematical constant pi, with a value of approximately 3.14159, used in the formula to ensure the normalization property of the Gaussian distribution; σ represents the standard deviation of the Gaussian distribution, which is the core parameter controlling the shape of the Gaussian function, and its value determines the steepness or flatness of the function curve. , representing the variance of the Gaussian distribution, is the square of the standard deviation σ, and together with σ, it affects the diffusion range of the function; e, representing the natural constant, has a value of approximately 2.71828, which serves as the base of the exponential function, causing the function to exhibit the decaying characteristics of a bell curve. + , represents the square of the Euclidean distance from a point (x,y) in a two-dimensional plane to the center of the Gaussian distribution (usually the origin), and is used to measure the spatial distance between the point and the center.

[0113] S322 directly performs subtraction on the image to obtain the gradient extrema in the X and Y directions, which are then used as drawing anchor points.

[0114] S323 uses gradient values ​​and gradient angles to filter anchor points while limiting the minimum line length to obtain the target anchor point.

[0115] S324 uses an additive strategy to connect target anchor points.

[0116] In some embodiments, when performing step S321, the vision-guided robotic arm assembly method further includes:

[0117] According to S3211, based on the symmetry of the kernel function, for a 5×5 window size, only 6 sets of kernel functions need to be calculated, and normalization is completed simultaneously.

[0118] When dealing with large image sizes, using Gaussian filtering will significantly increase computational complexity due to the excessive exponential functions and multiplication and division operations. Therefore, a lookup table-based method is used to calculate the Gaussian kernel function.

[0119] S3212, based on the discrete nature of grayscale values, calculate the product of 0-255 and the normalized kernel function mentioned above to obtain a lookup table of 6×256.

[0120] S3213: Based on the sum of the absolute values ​​of the positions as the row number and the pixel gray level as the column number, look up the value in the lookup table of the Gaussian filter function.

[0121] S3214 guides the robotic arm to adjust the depth direction in response to changes in the size of the target to be assembled, converting the pixel ratio into actual three-dimensional depth changes.

[0122] In this way, when guiding the robotic arm, the depth direction can be adjusted according to the changes in the size of the assembly object (such as the radius of a circle), and the pixel ratio can be converted into the actual three-dimensional space depth change, so as to achieve greater flexibility in the depth direction to adapt to the changes in the depth of the assembly target.

[0123] S4. Based on the feature localization results and the second preset center, obtain the position of the second center.

[0124] The second center position is used to guide the movement of the collaborative arm so that the position of the target to be assembled in the camera image is the second preset center of the camera image.

[0125] S5. Based on the initial detection point, the first center position, and the second center position, the guided assembly path is obtained.

[0126] like Figure 7 As shown, in some embodiments, a deep learning method is used, employing a camera mounted on the end effector of a robotic arm to achieve vision-guided robotic arm assembly tasks. The assembly process involves moving the robotic arm so that the device to be assembled is within the camera's field of view. A visual detection method is used to obtain the first offset between the device to be assembled and the center of the camera's field of view. Based on the first offset, the robotic arm is automatically moved so that the device to be assembled is located at the center of the camera's field of view. This process is repeated until the updated first offset is less than three pixels, and the current position is recorded as P0. The robotic arm is then manually moved until the end effector reaches the assembly position, and the amount of movement is recorded as Move0, completing the teaching process. The robotic arm is moved to P0, and the deep learning YOLOX algorithm is used to detect the second offset between the device to be assembled and the center of the camera's field of view. Based on the second offset, the robotic arm is automatically moved so that the device to be assembled is located around the center of the camera's field of view. A visual detection method is used to detect the third offset between the device to be assembled and the center of the camera's field of view. Based on the third offset, the robotic arm is automatically moved so that the device to be assembled is located at the center of the camera's field of view. This process is repeated until the third offset calculated by the visual detection method is less than three pixels. Finally, the robotic arm is automatically moved according to Move0, completing the assembly.

[0127] In some embodiments, the YOLOX algorithm is used to identify the equipment to be assembled, which is flexible to changes in the assembly position;

[0128] In some embodiments, the Egerwing algorithm is used to achieve high-precision detection of circular targets, reaching a distance of three pixels.

[0129] In some embodiments, a visually guided calibration method is used, which eliminates the need for hand-eye calibration and enables rapid deployment.

[0130] In some embodiments, a host computer platform is used to integrate robotic arm control algorithms, deep learning algorithms, camera acquisition and image processing algorithms to achieve full automation of the assembly process.

[0131] Based on the same concept, this application also provides a vision-guided robotic arm assembly system. The method corresponding to the vision-guided robotic arm assembly system can be the vision-guided robotic arm assembly method in the aforementioned embodiments, and its problem-solving principle is similar to that method. Figure 8 As shown, the vision-guided robotic arm assembly system 001 includes at least one processor 011 and a memory 012 communicatively connected to the at least one processor; wherein, the memory 012 stores instructions that can be executed by the at least one processor 011, and the instructions are executed by the at least one processor 011 to enable the at least one processor 011 to execute the vision-guided robotic arm assembly method provided in the embodiments of this application.

[0132] Another embodiment of this application provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.

[0133] Specifically, this embodiment may employ any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0134] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0135] The program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0136] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

Claims

1. A method for assembling a vision-guided robotic arm, characterized in that, include: Determine the initial detection point; The detection initial point is the position information of the collaborative arm in its initial posture; The initial posture is that the first surface of the collaborative arm is parallel to the preset plane of the target to be assembled; Based on the position of the target to be assembled in the first camera image, a first center position is obtained; the first center position is used to guide the movement of the cooperative arm so that the position of the target to be assembled in the camera image is the first preset center of the camera image, denoted as the second camera image. Based on the geometric features of the target to be assembled in the second camera image, the feature localization result is obtained; Based on the feature localization result and the second preset center, the second center position is obtained; the second center position is used to guide the movement of the collaborative arm so that the position of the target to be assembled in the camera image is the second preset center of the camera image. Based on the initial detection point, the first center position, and the second center position, a guided assembly path is obtained; The vision-guided robotic arm assembly method further includes: Determine the assembly object dataset; the assembly object dataset includes images of normal assembly objects, tilted assembly objects, damaged assembly objects, and missing assembly objects. The assembly object dataset is expanded by symmetry transformation, scaling transformation, and rotation transformation to obtain an expanded dataset; Train the first learning model based on the expanded dataset; Input the first camera image into the first learning model to obtain the category result; Filter the results by category to obtain the target results; Calculate the center position of the bounding box corresponding to the result with the highest confidence among the target results; The first center position is calculated based on the center position of the bounding box.

2. The visually guided robotic arm assembly method according to claim 1, characterized in that, When performing the step of obtaining the first center position based on the position of the target to be assembled in the first camera image, the vision-guided robotic arm assembly method further includes: The position of the target to be assembled in the first camera image is obtained through deep learning.

3. The assembly method for a vision-guided robotic arm according to claim 1 or 2, characterized in that, The vision-guided robotic arm assembly method further includes: The first camera image is converted to grayscale to obtain a black and white image; The black and white image is equalized using adaptive histogram equalization to obtain an equalized image. The equalized image is scale-normalized to obtain a normalized image; Based on the normalized image, the category results are obtained.

4. The assembly method for a vision-guided robotic arm according to claim 1 or 2, characterized in that, The vision-guided robotic arm assembly method further includes: The feature localization result is obtained through a feature detection algorithm.

5. The visually guided robotic arm assembly method according to claim 4, characterized in that, The vision-guided robotic arm assembly method further includes: The second camera image is preprocessed using bilateral nonlinear filtering and adaptive thresholding to obtain a smooth image; High-precision feature detection is achieved by using edge rendering to detect image edges, and the feature localization result is obtained.

6. The visually guided robotic arm assembly method according to claim 5, characterized in that, The vision-guided robotic arm assembly method further includes: Based on the image size, a Gaussian function is used for filtering, and the image is subtracted to obtain the gradient extreme points in the X and Y directions, which are used as drawing anchor points; By introducing gradient values ​​and gradient angles to filter anchor points, and limiting the minimum line length, the target anchor points are obtained. An additive strategy is used to connect the target anchor points.

7. The vision-guided robotic arm assembly method according to claim 6, characterized in that, The vision-guided robotic arm assembly method further includes: Based on the symmetry of the kernel function, six sets of kernel functions are calculated for a 5×5 window size, and normalization is performed simultaneously. Based on the discrete nature of grayscale values, the product of 0-255 and the normalized kernel function is calculated to obtain a 6×256 lookup table. The value of the Gaussian filter function is retrieved from the lookup table using the sum of the absolute values ​​of the positions as the row number and the pixel gray level as the column number. In response to the size changes of the target to be assembled, the robotic arm is guided to adjust in the depth direction, converting the pixel ratio into actual depth changes in three-dimensional space.

8. A vision-guided robotic arm assembly system, characterized in that, include: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

9. A computer-readable medium having computer program instructions stored thereon, characterized in that, The computer program instructions can be executed by a processor to implement the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Spacecraft assembly part identifying and positioning method based on target detection and composite target code

    CN116091401A