Robot arm 6D pose estimation method and system based on fuzzy visual angle prior

By using a 6D pose estimation method for a robotic arm based on fuzzy perspective priors, combined with RGB and depth image processing, the determination of candidate pose imaging points for rotation is simplified. This solves the problems of high labor intensity and slow calculation speed in existing technologies for rebar tying, and achieves accurate and safe automation of rebar tying.

CN121391997AActive Publication Date: 2026-01-23HUNAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511966590.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-01-23
Estimated Expiration
2045-12-24

AI Technical Summary

Technical Problem

In existing technologies, rebar tying relies on manual operation, which is labor-intensive, takes place in harsh environments, and is prone to quality defects. Traditional pose estimation methods have poor robustness under occlusion and noise conditions, while zero-shot deep learning 6D pose estimation networks are heavily dependent on computing hardware and have slow running speeds, making it difficult to meet the requirements of real-time performance and low-cost deployment on site.

Method used

A 6D pose estimation method for a robotic arm based on fuzzy perspective priors is adopted. By acquiring RGB and depth images and combining them with camera intrinsic parameters to obtain rebar point cloud data, plane fitting segmentation and target detection and recognition are performed. A zero-shot pose estimation model is used for translation and rotation processing, which simplifies the determination of rotational candidate pose imaging points, reduces hardware dependence, and enables localized deployment.

Benefits of technology

It improves computing speed, reduces hardware dependence, and achieves precise and safe automation of rebar tying, making it suitable for real-time field applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121391997A_ABST
    Figure CN121391997A_ABST
Patent Text Reader

Abstract

The invention discloses a robot arm 6D pose estimation method and system based on fuzzy visual angle prior, and relates to the technical field of robot intelligent manufacturing, and the method comprises the steps: obtaining steel bar point cloud data based on a color image and a depth image in combination with camera internal parameters; performing plane fitting segmentation processing on the steel bar point cloud data to obtain a current steel bar layer point cloud plane; mapping the current steel bar layer point cloud plane into a current layer steel bar image; performing target detection identification and pixel binarization processing on the current layer steel bar image to obtain a frame mask image; and taking the frame mask image, the color image and the depth image as input, performing translation part processing and rotation part processing based on a zero sample pose estimation model, and coupling processing results to obtain a 6D estimation pose corresponding to the reinforcing steel bar node frame image. According to the method provided by the invention, the processing process is simplified when the 6D estimation pose corresponding to each reinforcing steel bar node is accurately obtained, and localized deployment is favorably realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent work of industrial robots, and in particular to a robot arm 6D pose estimation method and system based on fuzzy visual angle prior. BACKGROUND

[0002] In the construction process of reinforced concrete structures, steel binding and steel welding are one of the most basic and key processes, and the quality thereof is directly related to the overall stability of the steel reinforcement and the stress performance after concrete pouring. In existing construction, steel binding mainly relies on manual operation, and workers need to complete the winding and fixing of the binding wire at complex steel intersection nodes. This method not only has high labor intensity and a poor working environment, but also is prone to cause quality problems such as node loosening and steel displacement due to non-standard operation or omission. With the development of intelligent construction and industrial robot technology, steel binding automation has gradually become a research hotspot. One important problem is how to realize six-degree-of-freedom (6D) pose estimation of the binding robot arm during the execution of the binding action. 6D pose estimation refers to determining the position and orientation of an object in three-dimensional space. 6D represents six degrees of freedom, including translation along X, Y, and Z directions, and rotation around X, Y, and Z directions.

[0003] Traditional pose estimation methods (such as point cloud registration) rely on point cloud geometric features (surface points, normal vectors, feature descriptors, etc.), and obtain the relative pose through iterative alignment or feature matching. The method has high requirements for camera quality and poor robustness under occlusion and noise conditions. Although the existing zero-shot deep learning 6D pose estimation network has strong adaptability, it has a large dependence on computing hardware (GPU) and relatively slow running speed, which is difficult to meet the real-time and low-cost deployment requirements on site.

[0004] Therefore, it is necessary to provide a robot arm 6D pose estimation method and system based on fuzzy visual angle prior. When performing 6D pose estimation based on a zero-shot deep learning 6D pose estimation network, multi-source information is fused to realize processing based on fuzzy visual angle prior, improve the calculation speed, and reduce the dependence on hardware, thereby facilitating the implementation of local deployment on construction robots. SUMMARY

[0005] The main purpose of the present application is to provide a robot arm 6D pose estimation method and system based on fuzzy visual angle prior, which aims to solve the technical problems of complex processing and difficulty in realizing local deployment when performing 6D pose estimation using a zero-shot deep learning 6D pose estimation network in the prior art.

[0006] To achieve the above object, the application provides a robot arm 6D pose estimation method based on a fuzzy perspective prior, comprising the steps of: S10, obtaining steel bar point cloud data based on a target workpiece corresponding RGB image and depth image and combining camera internal parameters; S20, performing plane fitting segmentation processing on the steel bar point cloud data to obtain a current steel bar layer point cloud plane; S30, mapping the current steel bar layer point cloud plane into a current layer steel bar image; S40, performing target detection and recognition and pixel binarization processing on the current layer steel bar image to obtain a frame mask image corresponding to each steel bar node in the current layer steel bar image; S50, taking the frame mask image, the RGB image and the depth image as inputs, performing translation part processing and rotation part processing based on a zero sample pose estimation model, and coupling the processing results to obtain a 6D estimated pose corresponding to the steel bar node frame image; The rotation part processing comprises rotation pose initialization, rotation network optimization and rotation network scoring processing, and the rotation pose initialization is used to generate a rotation candidate pose photographing point and a candidate photographing pose corresponding to each rotation candidate pose photographing point. The rotation pose initialization specifically comprises: selecting a point on a target circle on a regular polyhedral spherical subdivision grid as a rotation candidate pose photographing point, and generating a candidate photographing pose based on each rotation candidate pose photographing point. The angle between the camera optical axis corresponding to the rotation candidate pose photographing point and the axis of the object coordinate system is , and the difference between and is not greater than a preset angle. The vector expression of the camera optical axis is and the expression of the axis in the object coordinate system is

[0007] Further, the step of generating a candidate photographing pose based on each rotation candidate pose photographing point specifically comprises: Within a rotation range of -β to β, a preset rotation step is adopted to rotate around the camera optical axis to obtain a plurality of candidate photographing poses based on each rotation candidate pose photographing point, and β is not greater than 180 degrees.

[0008] Further, the preset rotation step is 60 degrees, and β is 60 degrees or 100 degrees. One rotation candidate pose photographing point corresponds to three candidate photographing poses.

[0009] Further, the steel bar point cloud data is subjected to plane fitting segmentation processing to obtain a plurality of target steel bar layer point cloud planes, the optimal target steel bar layer point cloud plane is determined as the current steel bar layer point cloud plane, and the normal vector of the target steel bar layer point cloud plane is determined as the axis of the object coordinate system.

[0010] Further, in step S30, the current steel bar layer point cloud plane is mapped into a current layer steel bar image by using an inflation kernel operation.

[0011] Further, in step S40, a steel bar node frame image in the current layer steel bar image is obtained by performing target detection and recognition on the current layer steel bar image based on a steel bar node target detection model; wherein the establishment of the steel bar node target detection model comprises the following steps: Collecting training data, collecting a steel bar node data set by using a 3D camera; Processing training data, segmenting and filtering a plurality of layer steel bar images in the steel bar node data set to obtain a single layer steel bar image; Labeling training data, manually labeling the single layer steel bar image, and the labeling of one steel bar intersection object consists of one bounding box and one key point; Converging and training the initial steel bar node target detection model to obtain an updated steel bar node target detection model.

[0012] Further, pixel binary processing is performed on each steel bar node frame image, non-white pixels in the steel bar node frame image are set to white, and a corresponding frame mask image is generated.

[0013] Further, in step S40, the steel bar node frame images are sequentially sorted based on an S-shaped path working strategy, which specifically comprises: Clustering and analyzing the steel bar node frame images based on the size of the axis of the image coordinate system of the current layer steel bar image to obtain a plurality of rows of steel bar node frame image rows arranged along the axis, wherein each row of steel bar node frame image rows has at least one steel bar node frame image arranged along the horizontal direction; Judging whether the current row of steel bar node frame images is an odd row; If the current row of steel bar node frame images is an odd row, then the steel bar node frame images in the row of steel bar node frame images are sequentially sorted along the first direction of the axis based on the size of the axis coordinate of the image coordinate system; If the current row of steel bar node frame images is an even row, then the steel bar node frame images in the row of steel bar node frame images are sequentially sorted along the second direction of the axis based on the size of the The second direction of the shaft sequentially sorts the steel bar node frame images in the steel bar node frame image row; wherein the second direction is opposite to the first direction. The sorted steel bar node frame images are sequentially connected to form an S shape.

[0014] The application also provides a robot arm 6D pose estimation system based on a fuzzy perspective prior, comprising a welding robot and a visual perception device, the welding robot having a base and a robot hand, and the visual perception device being arranged on the robot hand. A processing device is arranged on the welding robot, and the processing device is used to realize the steps of the above-mentioned robot arm 6D pose estimation method based on a fuzzy perspective prior.

[0015] The application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to realize the steps of the above-mentioned robot arm 6D pose estimation method based on a fuzzy perspective prior.

[0016] Compared with the prior art, the robot arm 6D pose estimation method based on a fuzzy perspective prior has the following beneficial effects: The robot arm 6D pose estimation method based on a fuzzy perspective prior provided by the application firstly acquires an RGB image and a depth image of a target workpiece, combines and processes the RGB image, the depth image and camera intrinsic parameters to acquire steel bar point cloud data; then performs plane fitting segmentation processing on the steel bar point cloud data to obtain a plurality of target steel bar layer point cloud planes, and determines an optimal target steel bar layer point cloud plane as a current steel bar layer point cloud plane; next, after mapping the current steel bar layer point cloud plane into a current layer steel bar image, target detection and recognition and pixel binarization processing are performed to acquire a frame mask image corresponding to the current layer steel bar image; finally, the frame mask image, the RGB image and the depth image are taken as inputs, and a zero-shot pose estimation model is used for recognition to obtain a 6D estimated pose corresponding to each steel bar node; in the scheme of the application, when determining a rotation candidate pose shooting point and a candidate shooting pose, instead of using each vertex of a special icosphere as a shooting point, only the included angle between the optical axis of the camera and the object coordinate system is used as the rotation candidate pose shooting point, and several points on a corresponding circle are used as the rotation candidate pose shooting point, so that the processing process is simplified when accurately acquiring the 6D estimated pose corresponding to each steel bar node, and local deployment is facilitated. The included angle between the optical axis of the camera and the object coordinate system is The several points on the corresponding circle are the rotation candidate pose shooting points, and the processing process is simplified when accurately acquiring the 6D estimated pose corresponding to each steel bar node, which is conducive to realizing local deployment. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the structures shown in these drawings without creative labor.

[0018] Figure 1 is a flowchart of a robot arm 6D pose estimation method based on a fuzzy view angle prior in an embodiment of the present application. Figure 2 is an icosphere (special sphere) diagram under different subdivision values in the prior art, wherein a is an icosphere diagram when subdivision = 1, and b is an icosphere diagram when subdivision = 2. Figure 3 is a pose initialization processing principle diagram of a zero-sample pose estimation model in an embodiment of the present application. Figure 4 is an icosphere pose initialization diagram of a FoundationPose model in the prior art. Figure 5 is an icosphere pose initialization diagram of a FoundationPose model in an embodiment of the present application, wherein a is an icosphere pose initialization diagram when the angle between the camera optical axis and the axis of the object coordinate system is 166.40°, b is an icosphere pose initialization diagram when the angle between the camera optical axis and the axis of the object coordinate system is 93.69°, and c is an icosphere pose initialization diagram when the angle between the camera optical axis and the axis of the object coordinate system is 12.01°. Figure 6 is a current layer reinforcement image diagram obtained after an inflation operation in an embodiment of the present application.

[0019] The purposes, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0021] ​​​With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0022] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative positional relationship, movement condition, etc. between components in a certain posture (as shown in the drawings), and if the certain posture changes, the directional indications will also change accordingly.

[0023] In addition, the descriptions involving “first”, “second” and the like in the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined as “first” and “second” can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize it, and when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the scope of protection claimed by the present application.

[0024] Please refer to the drawings Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 , the present application provides a robot arm 6D pose estimation method based on fuzzy view prior, comprising the following steps: S10, based on the target workpiece corresponding RGB image and depth image and combining camera internal parameter to obtain steel bar point cloud data; S20, the steel bar point cloud data is subjected to plane fitting segmentation processing to obtain the current steel bar layer point cloud plane; S30, the current steel bar layer point cloud plane is mapped to the current layer steel bar image; S40, target detection and recognition and pixel binarization processing are performed on the current layer steel bar image to obtain the frame mask image corresponding to each steel bar node in the current layer steel bar image; S50, taking the frame mask image, the RGB image and the depth image as input, performing translation part processing and rotation part processing based on the zero sample pose estimation model, and coupling the processing results to obtain the 6D estimated pose corresponding to the steel bar node frame image; The rotation part processing includes rotation pose initialization, rotation network optimization, and rotation network scoring processing, and the rotation pose initialization is used to generate a rotation candidate pose photographing point and a candidate photographing pose corresponding to each rotation candidate pose photographing point. The rotation pose initialization specifically includes: selecting a point on a target circle on a regular polyhedral spherical subdivision grid as a rotation candidate pose photographing point, and generating a candidate photographing pose based on each rotation candidate pose photographing point. The angle between the camera optical axis corresponding to the rotation candidate pose photographing point and the axis of the object coordinate system is The difference between , and is not greater than a preset angle, wherein , is a vector expression of the camera optical axis, is an axis expression of the object coordinate system, , , are respectively a component of the vector in the x-axis direction, a component of the vector in the y-axis direction, and a component of the vector in the z-axis direction.

[0025] Through research, it is found that the existing 6D estimated pose method during robot binding and robot welding mainly includes the following: a method based on traditional geometry, mainly based on point cloud registration, such as using the ICP algorithm, iteratively aligning the actual collected reinforcement node point cloud with the standard model point cloud, so that the two sets of point clouds coincide as much as possible, solving the spatial position and attitude of the node, the principle is simple, and it does not depend on a large amount of data, but the initial position requirement is high, and the point cloud is prone to failure when there is noise or occlusion, and the stability is poor; a method based on template matching, the first is image template matching, rendering the RGB image template of the reinforcement node at different viewing angles in advance, which can also combine gradient, normal vector and other features, and in the image taken on the construction site, these templates are used to match pixel by pixel, find the closest viewing angle, and estimate the pose; the second is to generate a point cloud subset template of different attitudes of the object offline in advance (such as rendering multiple angle point clouds from a CAD model), and then compare the point cloud collected on site with these templates one by one, find the most similar one, and directly use it as the estimation result, in the actual construction site environment with many reinforcement and high appearance similarity, it is easy to confuse, and the template quantity is large, and the expansibility is poor.

[0026] The application provides a robot arm 6D pose estimation method based on a fuzzy perspective prior, first, an RGB image and a depth image of a target workpiece are acquired, and the RGB image, the depth image and camera intrinsic parameters are combined and processed to acquire rebar point cloud data; then, the rebar point cloud data is subjected to plane fitting segmentation processing to obtain a plurality of target rebar layer point cloud planes, and an optimal target rebar layer point cloud plane is determined as a current rebar layer point cloud plane; next, after the current rebar layer point cloud plane is mapped into a current layer rebar image, target detection and recognition and pixel binarization processing are performed to acquire a frame mask image corresponding to the current layer rebar image; finally, the frame mask image, the RGB image and the depth image are taken as inputs, and a zero sample pose estimation model is used for recognition to obtain a 6D estimated pose corresponding to each rebar node; in the scheme of the application, when determining a rotation candidate pose shooting point and a candidate shooting pose, instead of using each vertex of a special icosphere as a shooting point, an included angle between an optical axis of a camera and an axis of an object coordinate system is used as a rotation candidate pose shooting point on a corresponding circle, and the processing process is simplified when accurately acquiring a 6D estimated pose corresponding to each rebar node, which is beneficial to local deployment. . . .

[0027] . . .

[0028] . .

[0029] . .

[0030] . . . . . .

[0031] It can be understood that in the scheme of the present application, Icosphere (icosahedron) is mainly studied, which is a geometric primitive in computer graphics, which approximates a sphere by icosahedron, and a parameter of Icosphere is subdivision, when subdivision=1, each edge of the equilateral triangle is bisected, that is, each triangle is cut into four small equilateral triangles, as shown by a and b in the following formula: Figure 2

[0032] Further, the step of generating a candidate photographing pose based on each rotation candidate pose photographing point specifically comprises: rotating around the camera optical axis by a preset rotation step within a rotation range of -β to β to obtain a plurality of candidate photographing poses corresponding to each rotation candidate pose photographing point, and β is not greater than 180 degrees.

[0033] Further, the preset rotation step is 60 degrees, β is 60 degrees or 100 degrees, and one rotation candidate pose photographing point corresponds to three candidate photographing poses.

[0034] Further, the steel bar point cloud data is subjected to plane fitting segmentation processing to obtain a plurality of target steel bar layer point cloud planes, the optimal target steel bar layer point cloud plane is determined as the current steel bar layer point cloud plane, and the normal vector of the target steel bar layer point cloud plane is determined as the axis of the object coordinate system.

[0035] In the scheme of the present application, after determining the rotation candidate pose photographing point, the candidate photographing pose corresponding to each rotation candidate pose photographing point is further determined, which not only reduces the number of rotation candidate pose photographing points, but also solves the technical problem that mirror images exist when obtaining pictures within a 360-degree circumferential range in the prior art by determining the candidate photographing pose within a rotation range of -β to β, and further identification processing of the mirror image is required to determine the state of the robot arm.

[0036] It can be understood that the zero-shot network is a commonly used end-to-end identification network in the current 6D pose estimation network, which can be trained based on objects such as water cups and pen containers and used for pose estimation of steel bar nodes, without the need for training with steel bar nodes as training objects to still complete the pose estimation task of the steel bar nodes, which shows strong generalization. The representative network with good performance and high accuracy in the zero-shot network is the FoundationPose network, and the main structure of the FoundationPose network is divided into three parts of pose initialization, optimization network and scoring network.

[0037] Please refer to Figure 3 ​In a preferred embodiment of the present invention, the Foundationpose pose estimation network decouples the rotation part and the translation part and performs them in two parts. The Foundationpose pose estimation network includes translation part estimation and rotation part estimation. In the solution of the present invention, the translation part pose estimation includes translation initialization and fine alignment. It mainly improves the rotation pose initialization of the rotation part. The rotation part initialization includes rotation pose initialization to generate candidate poses, candidate pose optimization (rotation network optimization), and optimized pose (rotation network scoring) to select the final rotation part estimate.

[0038] In practical application, one of the improvements of this invention lies in the initialization strategy for the rotation part. The ultimate result is improved speed and accuracy by reducing the generation of irrelevant rotational candidate pose capture points. In existing technologies, the translation part initialization of the Foundationpose pose estimation network mainly refers to initializing the translation using 3D points located at the mid-depth of the detected 2D bounding box. The rotation part initialization refers to uniformly sampling from the Icosphere on the object facing the camera towards the center. Viewpoint, camera pose will be through The discrete in-plane rotations are further enhanced, resulting in A global pose initialization is sent as input to the pose optimization network. Please refer to [reference needed]. Figure 4 Before its improvement, the Foundationpose pose estimation network initializes the object by placing it at the center of a special sphere (icosphere) with a radius of 1. It then assumes the camera will take pictures at each vertex of the icosphere. At each vertex, the camera rotates 360 degrees around the z-axis of the camera coordinate system, taking a picture every 60 degrees. In practice, 42 vertices are captured, each rotating 6 times, resulting in 6 × 42 = 252 possible poses. Please refer to [reference needed]. Figure 5 In Figures a, b, and c, the solution of this invention, based on a rough perspective, uses a target circle corresponding to the shooting angle. Rotating the candidate pose shooting point onto this target circle helps reduce the camera's search range. For actual binding tasks, when segmenting the rebar layers to obtain the target rebar layer point cloud plane, all rebar nodes corresponding to the binding layer are on the current rebar layer point cloud plane. For each rebar node, its object coordinate system... The plane is approximately parallel to the dividing surface; the normal vector of the current reinforcement layer point cloud plane and the coordinate system of the object are... The axes are represented the same in the camera coordinate system; assuming the normal vector of the target rebar layer point cloud plane is (A, B, C), the condition that can be obtained is that the coordinates of the object system are the same in the camera coordinate system. The axis expression is In the actual algorithm, since it is known that the camera is facing the reinforcement cage, since the plane normal vector has two directions, it will be judged to take The normal vector with an axis angle greater than 90 degrees, the camera optical axis is the Axis of the camera coordinate system, the expression is From which it can be deduced that when the camera shoots the object, the angle between the camera optical axis And the Axis of the object coordinate system , .

[0039] Understandably, after obtaining the angle between the camera optical axis And the Axis of the object coordinate system , return to the initialization in the Foundationpose pose estimation network, it is known that the reference coordinate system becomes the object coordinate system, and in the object coordinate system, the camera is assumed to be at a shooting position , (assuming the icosphere radius is 1, the distance from the vertex to the origin is 1, that is =1, the optical axis vector expression is , the object coordinate system Z-axis expression ), at this time the camera optical axis And the object coordinate system Axis expression is: .

[0040] Understandably, since the selection of the coordinate system does not change the angle between the two vectors, under the correct pose , thus in the initialization process, the vector of the camera shooting point pointing to the origin of the object coordinate system and the object coordinate system Axis is Angle, at this time the candidate pose camera shooting point is a circle on the icosphere rather than the entire spherical surface, and then set the point selected on the target circle on the subdivision grid of the regular polyhedral spherical surface as the rotation candidate pose shooting point.

[0041] By adopting the scheme of the present application, even if the icosphere sphere takes subdivision=2, the maximum number of shooting point positions searched at around 90° is 24, which is much lower than the 42 shooting point positions before the FoundationPose is not improved subdivision=1.

[0042] Further, the scheme of the present application is another improvement based on the rough view angle initialization, which solves the geometric symmetry problem when the candidate pose photographing point is rotated. The existing scheme adopts a 360-degree rotation around the camera optical axis at the candidate pose photographing point. Since the steel bar nodes are symmetrical up and down, the deep learning network cannot distinguish the accurate direction of an object (cannot distinguish whether the head of an object is up or down). This will have an impact on accurately determining the pose in actual engineering. For example, if a mechanical arm is required to hold a chopstick to touch the steel bar node in the axial direction, without improvement, the three nodes in the middle row estimated to be upside down require the mechanical arm to move from bottom to top to complete the task, and the steel bar nodes estimated to be right side up (top left and bottom left) require the mechanical arm to move from top to bottom to complete the task, which will cause the mechanical arm to collide during the movement from bottom to top. The two pose estimation results have different impacts on the actual engineering application.

[0043] Through analysis, it is known that the existing upside-down pose is generated because the camera at each photographing position is rotated 360 degrees around the optical axis during initialization, thereby generating an upside-down pose. Since the prior condition is known, it is impossible to take a photograph upside down during photographing. Therefore, when rotating around the optical axis, 360° rotation is no longer used, but only 100° or 60° rotation is used.

[0044] Please refer to Figure 6 Further, in step S30, an inflation kernel operation is used to map the current layer steel bar image.

[0045] ​In another optional embodiment of the present application, due to the feature similarity of the intersection points of the multi-layer reinforcement structure in the image, the points of different layers are difficult to distinguish, and there are false points, in the identification of the intersection points of the current layer reinforcement, due to the lack of depth information, the existing image recognition algorithm is easily disturbed by the intersection points of the background layer reinforcement, and it is difficult to determine which layer the recognized target belongs to, therefore, excluding the interference of the background layer reinforcement is the premise of detecting the intersection points of the current layer reinforcement. Specifically, the present application proposes a single-layer reinforcement segmentation technology to filter the background and background layer reinforcement in the image, realize effective extraction of only the pixels of the current layer reinforcement, first, a 3D camera is used to capture and obtain a color image and a depth image; the built-in function geometry.RGBDImage.create_from_color_and_depth of the open source library open3d is used to combine the camera intrinsic parameters to convert them into a reinforcement point cloud model, the reinforcement point cloud model not only includes the reinforcement point cloud data, but also includes the point cloud data of the surrounding environment or other objects and a certain degree of noise; then, the straight-through filter is used to preliminarily filter and denoise the reinforcement point cloud data, by specifying the range on one or more axes, the region of interest containing the reinforcement point cloud is retained, and the unnecessary data and noise are deleted, to obtain the processed reinforcement point cloud data (denoised point cloud) . In the process of RANSAC fitting the point cloud plane, first, three points are randomly selected from the denoised point cloud set M, respectively 、 、 , the plane model parameters fitted by the three points are calculated 、 、 、 , then the distance of the remaining point set to the fitted plane is calculated , if is less than a predetermined threshold, the inlier set is added, otherwise iteration is performed. When the number of inliers meets the requirement, the model parameters of the point cloud fitting plane are output, if the number of inliers does not meet the requirement, the iteration is continued until the model parameters of the point cloud fitting plane with the largest number of inliers are found, and the multiple initial target reinforcement point cloud planes are obtained after plane fitting segmentation; further, the current reinforcement layer point cloud plane is determined from the multiple initial target reinforcement point cloud planes, considering that the coordinate reference system origin of the point cloud generated by the color image and the depth image is the position of the camera, the distance dis between the origin (0, 0, 0) and each initial target reinforcement point cloud plane is calculated, and the initial target reinforcement point cloud plane with the shortest distance is the current reinforcement layer point cloud plane.

[0046] Further, since there is a discontinuous phenomenon when the point cloud is mapped into an image, the method adopts an inflation kernel operation to expand the display range of the steel bar pixels, After the image processed by the inflation kernel, the identification of the steel bar intersection point can effectively avoid the interference of other layer steel bars.

[0047] Further, in step S40, the steel bar node target detection model is used to detect and identify the current layer steel bar image to obtain a steel bar node framework image in the current layer steel bar image; wherein the establishment of the steel bar node target detection model comprises the following steps: Collecting training data, using a 3D camera to collect a steel bar node data set; Processing the training data, segmenting and filtering the multi-layer steel bar images in the steel bar node data set to obtain single-layer steel bar images; Labeling the training data, manually labeling the single-layer steel bar images, and the labeling of one steel bar intersection point object consists of one bounding box and one key point; Convergent training of the initial steel bar node target detection model, and updating to obtain the steel bar node target detection model.

[0048] Further, pixel binarization processing is performed on each steel bar node framework image, non-white pixels in the steel bar node framework image are set to white, and a corresponding framework mask image is generated.

[0049] In step S40, the steel bar node framework image is sequentially sorted based on the S-shaped path working strategy, which specifically includes: Based on the size of the axis of the image coordinate system of the current layer steel bar image, clustering analysis is performed on the steel bar node framework image to obtain a plurality of rows of steel bar node framework images arranged along the axis, wherein each row of steel bar node framework images has at least one steel bar node framework image arranged along the horizontal direction; If the current row of steel bar node framework images is an odd row, then based on the size of the axis coordinate of the image coordinate system, the steel bar node framework images in the row of steel bar node framework images are sequentially sorted along the first direction of the axis; If the current row of steel bar node framework images is an even row, then based on the size of the axis coordinate of the image coordinate system, the steel bar node framework images in the row of steel bar node framework images are sequentially sorted along the second direction of the axis; wherein the second direction is opposite to the first direction; The sorted steel bar node frame images are sequentially connected to form an S shape.

[0050] In another optional embodiment of the present application, the training of the steel bar node target detection model specifically comprises: data collection, collecting a steel bar node data set using a 3D camera, obtaining a color image and a corresponding depth map, considering various backgrounds, light, different shooting angles and distances, different diameters, spacings, layers and rusted reinforcement cages when collecting the data set, collecting data using different multiple 3D cameras, considering the good and bad of single-layer steel bar segmentation effect by using the difference in depth information acquisition accuracy of them, so as to increase the diversity of the data set; data processing, processing the collected data set, for multi-layer steel bar images, using the above single-layer steel bar segmentation technology, filtering the background layer steel bar, obtaining a white background image containing only the current layer steel bar; for single-layer steel bar images, no filtering operation is performed, if the steel bar in the image is not substantially parallel to the x and y axes, the minimum angle rotation technology is used to automatically rotate the image into a steel bar image substantially parallel to the x and y axes, and both the pre-rotation and post-rotation images are used as data sets, so as to increase the diversity of the data set; data labeling: using the labeling tool Labelme to manually label the processed images, the labeling of a steel bar intersection object consists of a bounding box and a key point, when labeling the steel bar intersection point with a bounding box, since the steel bar is continuous, the steel bar intersection point is not an independent individual with clear boundaries, therefore, the part where two steel bars intersect and the part where the four ends of the steel bar extend a small amount are regarded as a steel bar intersection point and labeled with a bounding box. The key point is labeled as much as possible at the center of the two steel bar intersection parts. Finally, a.json format labeling file is obtained. Convert the labeling file into a.txt format labeling file required for training the key point detection model. Steel bar node target detection model training: using the labeled steel bar node data set, dividing it into a training set and a validation set in a ratio of 8:2, and then training the target detection model in deep learning to learn the features of the steel bar nodes in the image. During the training process, the performance of the model on the training set and the validation set is monitored, and adjustments and optimizations are made to ensure that the model has good generalization ability and accuracy.

[0051] In an optional embodiment of the present application, the frame mask image is obtained and the frame mask image on the current layer steel bar image is sorted based on the binding path planning. Specifically, pixel binary processing is performed on each node frame, and non-white pixels in the detection frame are set to white. Each node generates a mask image. If there are three nodes in one picture, three mask images will be generated. It is found through research that the S-shaped binding path is more economical than the Z-shaped binding path during the binding of steel bars. The S-shaped binding path is shorter and the mechanical arm runs for a shorter time. The pixel of the color image detection frame (steel bar node frame image) is used for binding path planning. The geometric center of each detection frame can know the approximate pixel position of the steel bar intersection point. The color image pixel coordinates are defined as the origin at the top left corner of the image, the positive direction of the axis is to the right, the positive direction of the axis is downward; specifically, first, the values of all pixel intersection point coordinates are clustered to determine which steel bar intersection points are in the same row; then the pixel points in each row are sorted again. According to the coordinate system definition, the class with the smallest value is the uppermost row of steel bar nodes in the image. At this time, the binding sequence is from left to right, so the smaller the x, the further to the front of this class. However, the second class is in the second row. At this time, it should be bound from right to left, so the larger the x, the further to the front of this class. The summary of the entire process is as follows: first, all pixel coordinates are roughly divided into several classes according to the value, the class with the smallest value is the first class, the second smallest is the second class, and so on. In the odd class (also representing the odd row, 1, 3, 5 rows), from small to large, and the even class x from large to small. As a specific example: if the pixel point information obtained is [(345, 435), (50, 100), (280, 230), (70, 240), (260, 120), (100, 430)] a total of 6 binding points, first, the axis is clustered, and the values that differ by no more than 50 are clustered into one class. The clustering result is [(345, 435), (100, 430)], [(50, 100), (260, 120)], [(280, 230), (70, 240)], and each class is sorted according to the y value from small to large. The class sorting is: [(50, 100), (260, 120)], [(280, 230), (70, 240)], [(345, 435), (100, 430)], where the i-th class represents the i-th row from top to bottom of the image. Then, sorting is performed in each class. When it is an odd row, it is bound from left to right, so it should be sorted according to Sort from small to large, so the first and third internal sorting of the class is [(50,100), (260,120)], [(100,430), (345,435)]; when it is an even row, it is sorted from right to left, and it should be sorted according to Sort from large to small, so the second class sorting is [(280,230), (70,240)]. The final sorting result is: [(50,100), (260,120), (280,230), (70,240), (100,430), (345,435)], these pixel points are corresponding to the bounding box, and the bounding box is corresponding to the mask, and the mask order is corresponding to each specific steel bar node order, so the sorting of the pixel points is the sorting of the pose of the steel bar node, and the robot will be bound according to the pose order.

[0052] The application also provides a robot arm 6D pose estimation system based on fuzzy view prior, comprising a welding robot and a visual perception device, the welding robot has a base and a robot hand, and the visual perception device is arranged on the robot hand. The welding robot is provided with a processing device, and the processing device is used to realize the steps of the robot arm 6D pose estimation method based on fuzzy view prior.

[0053] Optionally, in the scheme of the application, in order to accurately obtain the two-dimensional image and three-dimensional information of the target workpiece, an industrial-grade 3D structured light camera (Mech-Eye NANO) installed at the end of the mechanical arm is selected as the visual perception device.

[0054] The application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the robot arm 6D pose estimation method based on fuzzy view prior.

[0055] The robot arm 6D pose estimation method and system based on fuzzy view prior have the following beneficial effects: Color image information (RGB image) and depth information (depth image) are used for multi-modal fusion to realize 6D pose estimation of the steel bar node; the zero sample pose estimation model is used for rotation part processing based on fuzzy view prior, angle The prior condition reduces the search range of the icosphere virtual camera shooting position, realizes the end-to-end 6D pose estimation of the steel bar joint with high efficiency and low computing power, generates a candidate shooting pose by using the shooting point of the rotation candidate pose in the preset rotation range (-beta to beta), which is beneficial to avoid generating a mirror image; the current layer steel bar image is sorted in an S-shaped path to realize light computing of the steel bar binding path planning; the single-layer steel bar segmentation technology is used to obtain the current layer steel bar image of the single layer, which effectively realizes the segmentation and rotation of the current layer steel bar image, and is not limited by the distance and angle of the camera, and has good robustness; the current layer steel bar image is obtained in advance, and the steel bar node target detection model is used to identify the steel bar node framework image, to realize the generation of the rough steel bar node mask in the multi-layer steel bar overlapping scene, without affecting the accuracy of the final 6D pose estimation.

[0056] The above embodiments only express several implementation manners of the application, and the description is relatively specific and detailed, but it cannot be understood as a limitation on the patent scope of the application. It should be noted that, for ordinary skilled persons in the art, without departing from the concept of the application, several modifications and improvements can be made, which all belong to the protection scope of the application. Therefore, the protection scope of the application should be subject to the appended claims.

Claims

1. A method for estimating the 6D pose of a robotic arm based on fuzzy view priors, characterized in that, Includes the following steps: S10: Obtain steel bar point cloud data based on the RGB image and depth image corresponding to the target workpiece and combined with camera intrinsic parameters; S20, Perform plane fitting and segmentation processing on the rebar point cloud data to obtain the current rebar layer point cloud plane; S30, map the current reinforcement layer point cloud plane to the current layer reinforcement image; S40, Perform target detection and recognition and pixel binarization processing on the current layer steel reinforcement image to obtain the frame mask image corresponding to each steel reinforcement node in the current layer steel reinforcement image; S50, taking the frame mask image, the RGB image and the depth image as input, performing translation and rotation processing based on the zero-sample pose estimation model, and coupling the processing results to obtain the 6D estimated pose corresponding to the steel bar node frame image; The rotation process includes rotation pose initialization, rotation network optimization, and rotation network scoring. The rotation pose initialization is used to generate rotation candidate pose imaging points and candidate imaging poses corresponding to each rotation candidate pose imaging point. Specifically, the rotation pose initialization includes: selecting a point on a target circle on a regular polyhedral spherical subdivision mesh as the rotation candidate pose imaging point, and generating the candidate imaging pose based on each of the rotation candidate pose imaging points; The camera optical axis and object coordinate system corresponding to the rotated candidate pose imaging point The included angle between the axes is , and The difference between them is no greater than a preset angle, where, , Let be the vector expression for the optical axis of the camera. In the coordinate system of the object Axis expression.

2. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to claim 1, characterized in that, The step of generating the candidate image pose based on each of the rotated candidate pose image points specifically includes: exist to Within the rotation range, a preset rotation step size is used to rotate around the camera's optical axis to obtain multiple candidate image poses corresponding to each of the rotation candidate pose image points. No more than 180 degrees.

3. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to claim 2, characterized in that, The preset rotation step size is 60 degrees. The angle is 60 degrees or 100 degrees, and one of the rotating candidate poses for taking pictures corresponds to three candidate poses for taking pictures.

4. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to claim 1, characterized in that, The rebar point cloud data is subjected to planar fitting and segmentation to obtain multiple target rebar layer point cloud planes. The optimal target rebar layer point cloud plane is determined as the current rebar layer point cloud plane, and the normal vector of the target rebar layer point cloud plane is determined as the coordinate of the object coordinate system. axis.

5. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to any one of claims 1 to 4, characterized in that, In step S30, the point cloud plane of the current rebar layer is mapped to the image of the current rebar layer using an expansion kernel operation.

6. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to any one of claims 1 to 4, characterized in that, In step S40, target detection and recognition are performed on the current layer of rebar image based on the rebar node target detection model to obtain the rebar node frame image in the current layer of rebar image; wherein, the establishment of the rebar node target detection model includes the following steps: Training data collection: A 3D camera was used to collect a dataset of rebar nodes; Training data processing involves segmenting and filtering multi-layer rebar images in the rebar node dataset to obtain single-layer rebar images. Training data annotation involves manually annotating the single-layer rebar image. The annotation of a rebar intersection object consists of a bounding box and a key point. The initial rebar node target detection model is trained to convergence, and then updated to obtain the rebar node target detection model.

7. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to claim 6, characterized in that, Each of the rebar node frame images is subjected to pixel binarization processing, and non-white pixels in the rebar node frame image are set to white to generate a corresponding frame mask image.

8. The 6D pose estimation method for a robotic arm based on fuzzy view priors according to claim 6, characterized in that, In step S40, the rebar node frame images are sequentially sorted based on the S-shaped path working strategy, specifically including: Based on the image coordinate system of the current layer of reinforcement image Cluster analysis was performed on the image of the steel reinforcement node frame based on the size of the axis to obtain the result along the axis. A series of images of multiple rows of reinforcing steel node frames arranged along an axial direction, wherein each row of the reinforcing steel node frame images has at least one reinforcing steel node frame image arranged laterally. Determine whether the current row of the rebar node frame image is an odd number of rows; If the current row of the rebar node frame image is an odd number of rows, then based on the image coordinate system axial coordinates along The first direction of the axis sorts the rebar node frame images in the row of rebar node frame images sequentially; If the current row of the rebar node frame image is an even number of rows, then based on the image coordinate system... axial coordinates along The second direction of the axis is used to sequentially sort the rebar node frame images in the row of rebar node frame images; wherein, the second direction is opposite to the first direction; The sorted rebar node frame images are connected sequentially to form an S-shape.

9. A 6D pose estimation system for a robotic arm based on fuzzy view priors, characterized in that, It includes a welding robot and a vision sensing device. The welding robot has a base and a robotic arm, and the vision sensing device is mounted on the robotic arm. The welding robot is equipped with a processing device, which is used to implement the steps of the 6D pose estimation method for a robot arm based on fuzzy perspective prior as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Object 6D posture prediction method based on RGB image and coordinate system transformation

    CN110660101A

  • Disorderly stacked workpiece grabbing system based on 3D vision and interaction method

    CN111508066A

  • Registration method and readable storage medium

    CN116912294A

  • Steel bar binding method, steel bar binding system, storage medium and electronic equipment

    CN119180865A

  • Robot visual identification, positioning and grabbing system based on RGBD point cloud

    CN119625052A