A highly robust visual recognition and posture detection method for the unloading hole of large tank tooling

By using auxiliary laser points and triangular holes on the robot, and combining deep learning technology for coarse and precise positioning, the problem of inaccurate unloading positioning of the robot in complex industrial environments is solved, achieving higher detection accuracy and stability.

CN114399500BActive Publication Date: 2025-05-13CHONGQING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210089387.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-05-13
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

In complex industrial environments, when the robot unloads the material from a large tank, the space positioning of the unloading hole is inaccurate due to interference factors such as environmental noise and dust, which affects the accuracy of the unloading.

Method used

The auxiliary laser point and auxiliary triangular hole are used, combined with coarse positioning and precision positioning technology based on deep learning, helping the robot accurately locate the position of the tooling and unloading hole. The specific steps include building a binocular camera system, performing binocular camera calibration, using the improved Yolov3 network for coarse positioning, and fine positioning through the Mask R-CNN network, and finally obtaining the height and offset through laser ranging and parameter calculation.

Benefits of technology

It improves the accuracy, stability and robustness of the inspection results in complex industrial environments, ensures that the robot can accurately enter the hole and hook the tooling out, and improves the accuracy of the unloading process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399500B_ABST
    Figure CN114399500B_ABST
Patent Text Reader

Abstract

The present invention discloses a highly robust visual recognition and posture detection method for the unloading hole of a large tank tooling. A static binocular image acquisition system is built, distortion correction and polar line correction are performed on the binocular camera system, and the data set is expanded using a conventional method for expanding the data set; the ROI area is extracted through an improved Yolov3 network to complete the first step of coarse positioning; the detection results obtained in the coarse positioning stage are clipped with a bounding box and sent to the Mask R‑CNN network for the second step of fine positioning to obtain accurate segmentation results. The centroid of the operating hole and the auxiliary hole are then extracted and the recognition success rate is verified; finally, parameter calculation is performed to obtain the height, offset angle, and x and y direction offsets. Auxiliary laser points and auxiliary triangular holes are used, and a technical route combining coarse positioning + fine positioning based on deep learning is used to help the manipulator accurately locate the position of the tooling unloading hole, and the tooling is horizontally hooked out from the inside of the tank device, thereby improving the accuracy and stability of the detection results in complex industrial environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a visual recognition and posture detection method, and in particular to a highly robust visual recognition and posture detection method for a discharge hole of a large tank tooling. Background Art

[0002] Removing tooling from large tanks traditionally relies primarily on manual labor, but with the advancement of automation technology, robots are being used. A robot is an automatic machine that simulates human hand operations. It can repeatedly grasp and move products or handle tools according to a program to complete specific processes. Robots can replace humans in monotonous, repetitive, or strenuous physical labor, achieving mechanization and automation of production. In an environment with continuously rising labor costs, robots are being chosen to automate production lines, effectively replacing manual labor while significantly improving operational accuracy. In actual operation, the robot's hook extends into the tank through the discharge port to remove the tooling. However, in complex industrial environments, various uncertainties in ambient noise, such as dust and smoke, can interfere with target detection and easily lead to inaccurate spatial positioning of the discharge port. Summary of the Invention

[0003] In response to the shortcomings of the above-mentioned existing technologies, the present invention provides a highly robust visual recognition and posture detection method for the unloading hole of large tank tooling. This method uses auxiliary laser points and auxiliary triangular holes, and through a technical route combining coarse positioning + fine positioning based on deep learning, helps the robot to accurately locate the position of the tooling unloading hole, and horizontally hook out the tooling from the inside of the tank device, thereby improving the accuracy and stability of the detection results in complex industrial environments.

[0004] In order to achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0005] A highly robust visual recognition and posture detection method for the discharge hole of a large tank tooling is proposed. An auxiliary triangular hole is added below the rectangular discharge hole. An auxiliary laser is mounted on a manipulator, with the relative positions of the laser and the mechanism fixed. The laser spot is projected onto the surface of the tooling's circular disc. The manipulator's posture is adjusted by calculating the offset of the laser spot relative to the line connecting the center of mass of the rectangular discharge hole and the auxiliary triangular hole, ensuring that the hook can accurately enter the hole and remove the tooling.

[0006] At the same time, a technical route combining coarse and fine positioning based on deep learning is adopted to improve the accuracy, stability, and robustness of detection results in complex industrial environments. The specific steps are as follows:

[0007] S1: Build a binocular camera system, perform binocular camera calibration, and collect images;

[0008] S2: Use conventional dataset expansion methods to expand the dataset;

[0009] S3: The images captured by the binocular camera are fed into the improved Yolov3 network for training, ROI area extraction is performed, and the first step of coarse positioning is completed;

[0010] S4: The rectangular discharge hole and auxiliary triangular hole region ROI extracted by rough positioning are sent to the Mask R-CNN network to obtain accurate segmentation results and complete the second step of fine positioning;

[0011] S5: Extract the centroid of the rectangular discharge hole and the auxiliary triangle hole and verify the recognition success rate;

[0012] S6: Perform parameter calculation to obtain the height, offset angle, and x and y direction offsets.

[0013] As a preferred solution of the present invention, in step S1, the main purpose of binocular camera calibration is to obtain the camera intrinsic parameter matrix A, extrinsic parameter matrix [R|T], distortion coefficients [k1, k2, k3, ~, p1, p2, ~], adjust the position of the distortion point on the imager, and then perform extreme correction to make the same object the same size in the left and right images and on the same horizontal line; specifically:

[0014] S-A1: Collect chessboard images. The photos for binocular camera calibration must be taken simultaneously by both cameras. The chessboard should occupy as much of the image as possible to obtain more information about lens distortion. Shoot from multiple angles and use at least five pairs of photos.

[0015] S-A2: Use the stereoCalibrate function provided by Opencv to obtain the camera's intrinsic parameter matrix A, extrinsic parameter matrix [R|T], and distortion coefficients [k1, k2, k3, ~, p1, p2, ~];

[0016] S-A3: Decompose the extrinsic parameter matrix [R|T] solved by OpenCV into the matrices R1, T1 and R2, T2 of the rotation and translation halves of the left and right cameras respectively;

[0017] S-A4: The source image pixel coordinate system is converted into the camera coordinate system through the intrinsic parameter matrix A, parallel epipolar correction is performed through R1 and R2, and the camera coordinates of the image are corrected by the distortion coefficient;

[0018] S-A5: Convert the calibrated camera coordinate system into the image pixel coordinate system and assign the new image coordinates according to the pixel values ​​of the source image coordinates.

[0019] As a preferred solution of the present invention, in step S3, because the background in the industrial environment is complex and there are many interference factors, in order to prevent the detection of rectangular targets other than the rectangular unloading hole in the coarse positioning stage, a fault-tolerant mechanism is added to this stage. Based on the Euclidean distance method, it is judged whether the target is in the background of the anchor. When the Euclidean distance between the predicted rectangular unloading hole boundary box and the auxiliary triangular hole boundary box is greater than the diameter of the tooling disc, it is determined to be a recognition error, thereby ensuring reliable recognition.

[0020] As a preferred solution of the present invention, in step S4, pixel-by-pixel target detection is continued through a mask algorithm; in the fine positioning stage, the detection results obtained in the coarse positioning stage are first clipped with a bounding box and then sent to the Mask R-CNN network to obtain accurate pixel-by-pixel detection results.

[0021] As an improvement of the present invention, in step S5, after the precise positioning is completed, in order to ensure the recognition rate, it is necessary to introduce the epipolar constraint in the binocular system calibration.

[0022] As an improved solution of the present invention, step S6 performs parameter calculation to obtain the height, offset angle, and x and y direction offsets, specifically:

[0023] S-B1: Use a laser rangefinder to measure the distance from the tooling to the tank opening;

[0024] S-B2: Establish a coordinate system and measure vector p r p t The angle between the vector (0, -1), where p r is the coordinate of the center point of the rectangular discharge hole, p t is the center coordinate of the auxiliary triangle hole;

[0025] S-B3: vector p r p t The actual length is Lmm, and the coordinates of the center point of the rectangular discharge hole are p r (x r ,y r ), the center coordinate of the auxiliary triangle hole is p t (x t ,y t ), the laser point coordinate is p l (x l ,y l ); First calculate the vector p r p t The equation of the straight line is:

[0026] Ax+By+C=0;

[0027] The above formula is the general form of the equation of a straight line. First, let the vector p r pt The equation of the straight line is y=kx+b (slope-intercept form of the straight line equation), k is the slope of the straight line, b is the intercept of the straight line, and the coordinates of the two known center points are: p r (x r ,y r ), p t (x t ,y t ), substitute the slope-intercept form of the straight line equation to solve for k and b; then reorganize y=kx+b into the general form of the straight line equation, which is Ax+By+C=0. At this time, the slope of the straight line is expressed as The intercept of the line is expressed as

[0028] S-B4: Calculate point p l Shortest distance to a straight line:

[0029]

[0030] S-B5: Calculating Vectors The length l:

[0031]

[0032] S-B6: Convert pixel distance d to actual size

[0033]

[0034] S-B7: Calculate altitude

[0035]

[0036] The technical effect of the present invention is as follows: the present invention collects images of tooling by building a static binocular image acquisition system, performs distortion correction and epipolar line correction on the binocular camera system, and uses conventional data set expansion methods such as random cropping to expand the data set; performs ROI area extraction through an improved Yolov3 network to complete the first step of coarse positioning; performs bounding box cropping on the detection results obtained in the coarse positioning stage and sends them to the Mask R-CNN network for the second step of fine positioning to obtain accurate segmentation results; then extracts the centroid of the rectangular discharge hole and the auxiliary triangular hole and verifies the recognition success rate; finally, performs parameter calculation to obtain the height, offset angle, and x and y direction offset. Using auxiliary laser points and auxiliary triangular holes, through a technical route based on the combination of coarse positioning + fine positioning based on deep learning, the robot is helped to accurately locate the position of the tooling discharge hole, and the tooling is horizontally hooked out from the inside of the tank device, thereby improving the accuracy and stability of the detection results in complex industrial environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1This is a flowchart of a highly robust visual recognition and posture detection method for the unloading hole of a large tank tooling;

[0038] Figure 2 To identify the effect diagram;

[0039] Figure 3 There are 9 anchor boxes in YOLOv3;

[0040] Figure 4 Schematic diagram of the prediction box center;

[0041] Figure 5 Schematic diagram of the fine positioning stage;

[0042] Figure 6 This is the Mask R-CNN network diagram;

[0043] Figure 7 is a bilinear interpolation graph;

[0044] Figure 8 shows the centroid extraction results. Figure 8(a) shows the centroid of the rectangular discharge hole, and Figure 8(b) shows the centroid of the auxiliary triangle hole.

[0045] Figure 9 A schematic diagram of the coordinate system. DETAILED DESCRIPTION

[0046] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0047] A highly robust visual recognition and posture detection method for the unloading hole of large tank tooling uses auxiliary laser points and auxiliary triangular holes. It adopts a technical route based on coarse positioning + fine positioning based on deep learning to help the robot accurately locate the position of the tooling unloading hole and horizontally hook out the tooling from the inside of the tank device. The process flow of this method is as follows: Figure 1 As shown,

[0048] In order to better adjust the posture of the manipulator relative to the discharge hole, an auxiliary triangular hole 2 is added below the rectangular discharge hole 1. The auxiliary laser is installed on the manipulator. The relative position of the laser and the machine is fixed, and the laser point is projected onto the surface of the disc of the tooling. The manipulator posture is adjusted by calculating the offset of the laser point relative to the center of mass of the two holes (rectangular discharge hole and auxiliary triangular hole) to ensure that the hook can accurately enter the hole and hook out the tooling. At the same time, the technical route of combining coarse positioning and fine positioning based on deep learning improves the accuracy, stability, and robustness of the detection results in complex industrial environments. The steps are as follows:

[0049] S1: Build a binocular camera system, perform binocular camera calibration, and collect images, such as Figure 2 As stated.

[0050] Step S1 calibrates the built binocular camera system. The main purpose of binocular camera calibration is to obtain the camera's intrinsic parameter matrix A, extrinsic parameter matrix [R|T], and distortion coefficients [k1, k2, k3, ~, p1, p2, ~], adjust the position of the distortion point on the imager, and then use extreme correction to make the same object the same size in the left and right images and on the same horizontal line. Specifically:

[0051] S-A1: Collect chessboard images. The photos used for binocular camera calibration must be taken simultaneously with both cameras. The chessboard should occupy as much of the image as possible to obtain more information about lens distortion. Shoot from multiple angles, preferably with at least five pairs of photos.

[0052] S-A2: Use the library functions such as stereoCalibrate() provided by Opencv to obtain the camera intrinsic parameter matrix A, extrinsic parameter matrix [R|T], and distortion coefficients [k1, k2, k3, ~, p1, p2, ~].

[0053] S-A3: Decompose the extrinsic parameter matrix [R|T] solved by OpenCV into matrices R1, T1 and R2, T2 for half of the rotation and translation of the left and right cameras respectively.

[0054] S-A4: The source image pixel coordinate system is converted into the camera coordinate system through the intrinsic parameter matrix A, parallel epipolar correction is performed through R1 and R2, and the camera coordinates of the image are corrected by the distortion coefficient.

[0055] S-A5: Convert the calibrated camera coordinate system into the image pixel coordinate system and assign the new image coordinates according to the pixel values ​​of the source image coordinates.

[0056] S2: Use conventional methods to expand the dataset (such as random cropping, flipping or mirroring, adjusting image brightness or contrast, Gaussian blurring, adding noise, etc.) to expand the dataset.

[0057] S3: Images captured by the binocular camera are fed into an improved Yolov3 network for training and ROI extraction, completing the first step of coarse positioning. Due to the complex backgrounds and numerous interference factors in industrial environments, a fault-tolerant mechanism is implemented in this stage to prevent the detection of rectangular targets outside the rectangular discharge hole during coarse positioning. This mechanism uses the Euclidean distance method to determine whether the target is within the background of the anchor. If the Euclidean distance between the predicted rectangular discharge hole bounding box and the auxiliary triangular hole bounding box is greater than the tooling disk diameter, a recognition error is detected, ensuring reliable recognition.

[0058] Deep learning-based object detection algorithms can be divided into two categories: one is based on the Region Proposal + CNN framework (a two-stage approach), represented by the RCNN series; the other is based on the Regression framework (a one-stage approach), represented by the YOLO series. The YOLO algorithm achieves predictions directly from images in one step, eliminating the need for region extraction. This simplifies the object detection task into a regression problem.

[0059] YOLOv3 combines the advantages of YOLOv1 and YOLOv2, improving its ability to detect small objects while also increasing detection accuracy. YOLOv3 uses Darknet53 as its backbone feature extraction network. Darknet53 uses residual units. This skip connection structure alleviates the vanishing gradient problem caused by increased network depth.

[0060] The network also considers the problem of detecting multi-scale targets, draws on the idea of ​​pyramid feature maps, and fuses feature maps of various scales through a series of convolution and upsampling operations, ultimately obtaining three feature layers of different sizes. Small-size feature layers are used to detect large-size targets, and large-size feature layers are used to detect small-size targets. YoloV3 has three Anchor boxes for each feature point in each feature layer. Each Anchor box has a 4-dimensional prediction box value (tx, ty, tw, th) and a 1-dimensional prediction box confidence pc (whether it contains the target). Taking the coco training set (80 classes) as an example, the number of feature map channels should be 255=3*(80+1+4), and the shapes of the three feature layers are (13,13,255), (26,26,255), and (52,52,255). Figure 3 shown.

[0061] After obtaining the output feature map, decoding is required to obtain the detection information. The IOU of the three anchor boxes and the actual box is calculated separately, and the prediction generated by the anchor box with the largest IOU is selected for correction. The decoding process of YoloV3 is divided into two steps:

[0062] ①The coordinates of the upper left point of the grid where the center point of the bounding box is located (c x ,c y ), and the offset of the center point relative to the grid (σ(t x ),σ(t y )) is added together to get the center of the prediction box (b x ,b y ),like Figure 4 shown.

[0063] ② Using the preset prior frame relative to the width and height of the feature map (p w ,p h ) and scale scaling (t w ,t h ) Calculate the width and height of the prediction box (b w ,b h ).

[0064]

[0065] right Figure 4 And the calculation formula is declared: (t x ,t y ,t w ,t h ) is the offsets learned by the deep network relative to the prior box, (b x ,b y ,b w ,b h ) is the position and size of the bounding box relative to the feature map. Use the sigmoid function to convert (t x ,t y ) is compressed into the range [0,1], which can effectively ensure that the center of the bounding box is in the grid cell where the prediction is performed, preventing excessive deviation. (t w ,t h ) to ensure that the scaling factor is greater than 0.

[0066] In the YOLOv3 algorithm, the bounding box error, confidence error, and classification error are integrated to form the overall loss function. The bounding box loss function includes the center coordinate error and the width and height coordinate errors, and a weight coefficient is added to reduce the contribution of the absence of objects in the grid to the overall loss function. This part of the loss function is:

[0067]

[0068] The confidence error is expressed using cross entropy. Most grids in an image do not contain the target to be measured, which results in the contribution weight of the target-free calculation part being greater than the target-containing part. Therefore, the weight coefficient of the target-free loss function needs to be increased.

[0069]

[0070] The classification error loss function also uses the cross entropy loss function. When the j-th anchor box of the i-th grid can accurately hit the target, the prediction box generated by this anchor box will be used to calculate the classification loss function.

[0071]

[0072] S4: The rectangular unloading hole (operating hole) and auxiliary triangular hole area ROI extracted by rough positioning are sent to the Mask R-CNN network to obtain accurate segmentation results and complete the second step of fine positioning.

[0073] Using YOLOv3 alone can only obtain the bounding boxes of the rectangular discharge hole (operating hole) and the auxiliary triangular holes, but cannot perform accurate visual recognition and pose detection. Therefore, a masking algorithm is used to continue pixel-by-pixel object detection. In digital image processing, image masks are commonly used to extract regions of interest or shield certain areas of the image from processing.

[0074] Precision positioning stage, such as Figure 5 As shown in the figure, the detection results obtained in the coarse positioning stage are first cropped by the bounding box, and then sent to the Mask R-CNN network to obtain accurate detection results at the pixel level.

[0075] The mask algorithm used is Mask R-CNN, which goes a step further based on Faster R-CNN: it can obtain pixel-level detection results, add a mask prediction branch in parallel with the Faster R-CNN prediction box, and use the mask branch to predict a binary mask for each ROI. The mask branch used to predict the mask is a small fully convolutional network that predicts a semantic mask for each ROI at the pixel level. This network structure decouples the prediction of category and mask, and predicts a mask independently for each class, relying on the ROI classification branch of the network to make category predictions, such as Figure 6 shown.

[0076] Another important improvement in Mask R-CNN is ROIAlign. A problem with Faster R-CNN is that its feature maps are misaligned with the original image, affecting detection accuracy. Mask R-CNN proposes RoIAlign to replace ROI pooling, which preserves the approximate spatial position. ROI Align incorporates two key concepts: bilinear interpolation and ROI pooling. The following explains these two aspects.

[0077] Bilinear interpolation is essentially linear interpolation in two directions. Given the data (x0, y0) and (x1, y1), to calculate the y value of a position x on the line within the interval [x0, x1], the relationship between y and x can be constructed by making the slope equal, as follows:

[0078]

[0079] Suppose we want to get the interpolation value of point P in the figure below, such as Figure 7 As shown, we can first adjust Q in the x direction. 11 and Q 21 Linear interpolation is performed between them to obtain R1, and R2 can be obtained similarly. Then, linear interpolation of R1 and R2 in the y direction can be performed to obtain the final P.

[0080] First, perform linear interpolation in the x direction and obtain:

[0081]

[0082] Then perform linear interpolation in the y direction to obtain:

[0083]

[0084] This will give you the desired result

[0085]

[0086] The ROI pooling layer can significantly accelerate training and testing and improve detection accuracy. This layer has two inputs: fixed-size feature maps obtained from a deep network with multiple convolution kernels; an N*5 matrix representing all ROIs, where N represents the number of ROIs. The first column represents the image index, and the remaining four columns represent the coordinates of the remaining upper left and lower right corners.

[0087] The specific operations of ROI pooling are as follows:

[0088] (1) According to the input image, map the ROI to the corresponding position of the feature map;

[0089] (2) Divide the mapped area into sections of the same size;

[0090] (3) Perform max pooling operation on each section;

[0091] This allows us to generate feature maps of fixed size from boxes of different sizes. It's worth noting that the size of the output feature maps is independent of the size of the ROI or the convolutional feature maps. The biggest benefit of ROI pooling is that it significantly improves processing speed.

[0092] S5: Extract the centroid of the operating hole (rectangular unloading hole) and the auxiliary hole (auxiliary triangular hole) and verify the recognition success rate.

[0093] The center of mass of an image is also called the center of gravity of an image. The pixel value of each point in the image can be understood as the mass at that point. As shown in Figure 8, Figure 8(a) is the center of mass of the rectangular discharge hole, and Figure 8(b) is the center of mass of the auxiliary triangle hole. For the center of mass in the x (y) direction, the sum of the pixels on the left and right (top and bottom) sides of the image are equal. The center of mass formula is as follows, where x i (y i ) is the coordinate of each pixel in the x(y) direction, p i is the corresponding pixel value, and (x, y) is the obtained centroid coordinate.

[0094]

[0095] After the precise positioning is completed, in order to ensure the recognition rate, it is necessary to introduce the epipolar constraint in the binocular system calibration. At time t, the centroid coordinates of the rectangular hole in the left camera image are The coordinates of the center of mass of the rectangular hole in the right camera image are If satisfied It is believed that the recognition of rectangular holes is effective, and the same is true for triangular holes.

[0096] S6: Perform parameter calculations to obtain the height, offset angle, and x and y direction offsets. Specifically:

[0097] S-B1: Use a laser rangefinder to measure the distance from the tooling to the tank opening.

[0098] S-B2: Establish Figure 9 The coordinate system shown, the measurement vector p r p t The angle between the vector (0, -1).

[0099] S-B3: vector p r p t The actual length Lmm. Assume Figure 9 The coordinates of the center point of the rectangular hole are p r (x r ,y r ), the coordinate of the center of the triangular hole is p t (x t ,y t ), the laser point coordinate is p l (x l ,y l ). First calculate the vector p r p t The equation of the straight line is:

[0100] Ax+By+C=0 (10)

[0101] S-B4: Calculate point p l Shortest distance to a straight line:

[0102]

[0103] S-B5: Calculating Vectors The length l:

[0104]

[0105] S-B6: Convert pixel distance d to actual size

[0106]

[0107] S-B7: Calculate altitude

[0108]

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A highly robust visual recognition and posture detection method for a large tank tooling unloading hole, characterized in that: Add an auxiliary triangular hole below the rectangular discharge hole; install the auxiliary laser on the manipulator, the relative position of the laser and the machine is fixed, and the laser point is projected on the disc surface of the tooling; adjust the manipulator posture by calculating the offset of the laser point relative to the centroid line of the rectangular discharge hole and the auxiliary triangular hole to ensure that the hook can accurately enter the hole to hook out the tooling; At the same time, a technical route combining coarse positioning and fine positioning based on deep learning is adopted to improve the accuracy, stability and robustness of detection results in complex industrial environments; The specific steps are as follows: S1: Build a binocular camera system, calibrate the binocular camera, and collect images; S2: Use conventional dataset expansion methods to expand the dataset; S3: Send the pictures collected by the binocular camera to the Yolov3 network for training, extract the ROI area, and complete the first step of rough positioning; S4: Send the rectangular discharge hole and auxiliary triangle hole area ROI extracted by rough positioning into the Mask R-CNN network to obtain accurate segmentation results and complete the second step of fine positioning; S5: Extract the centroid of the rectangular discharge hole and the auxiliary triangle hole and verify the recognition success rate; S6: Perform parameter calculation to obtain the height, offset angle, and x and y direction offsets.

2. According to claim 1, a highly robust visual recognition and posture detection method for a large tank tooling discharge hole is characterized in that: In step S1, the main purpose of binocular camera calibration is to obtain the camera intrinsic parameter matrix A, extrinsic parameter matrix [R|T], distortion coefficients [k1, k2, k3, ~, p1, p2, ~], adjust the position of the distortion point on the imager, and then make the same object the same size in the left and right images and on the same horizontal line through limit correction; specifically: S-A1: Collect chessboard images. The photos taken during binocular camera calibration must be taken by the left and right cameras at the same time. The chessboard should occupy as much of the picture as possible to obtain more information about lens distortion. The photos should be taken from multiple angles and the number of photos should be more than 5 pairs. S-A2: Use the stereoCalibrate function provided by Opencv to obtain the camera's intrinsic parameter matrix A, extrinsic parameter matrix [R|T], and distortion coefficients [k1, k2, k3, ~, p1, p2, ~]; S-A3: Decompose the external parameter matrix [R|T] solved by Opencv into the matrices R1, T1 and R2, T2 of the rotation and translation of the left and right cameras respectively; S-A4: The source image pixel coordinate system is transformed into the camera coordinate system through the intrinsic parameter matrix A, parallel epipolar correction is performed through R1 and R2, and the camera coordinates of the image are corrected through the distortion coefficient; S-A5: Convert the calibrated camera coordinate system into the image pixel coordinate system and assign the new image coordinates according to the pixel values ​​of the source image coordinates.

3. According to the highly robust large tank tooling unloading hole visual recognition and posture detection method of claim 1, its characteristics are The feature is that in step S3, because the background is complex and there are many interference factors in the industrial environment, in order to prevent the detection of rectangular targets other than the rectangular unloading hole in the rough positioning stage, a fault-tolerant mechanism is added in this stage. Based on the Euclidean distance method, it is judged whether it is a target in the background of the anchor. When the Euclidean distance between the predicted rectangular unloading hole bounding box and the auxiliary triangular hole bounding box is greater than the diameter of the tooling disc, it is judged as a recognition error, thereby ensuring reliable recognition.

4. According to claim 1, a highly robust visual recognition and posture detection method for a large tank tooling discharge hole is characterized by: In step S4, a mask algorithm is used to continue pixel-by-pixel target detection. In the fine positioning stage, the detection results obtained in the coarse positioning stage are first cropped with bounding boxes, and then sent to the Mask R-CNN network to obtain accurate pixel-by-pixel detection results.

5. The highly robust visual recognition and posture detection method for a large tank tooling discharge hole according to claim 1 is characterized in that: In step S5, after the precise positioning is completed, in order to ensure the recognition rate, it is necessary to introduce the epipolar constraint in the binocular system calibration.

6. A highly robust visual recognition and posture detection method for a large tank tooling discharge hole according to claim 1, characterized in that: Step S6 performs parameter calculation to obtain the height, offset angle, and x and y direction offsets, specifically: S-B1: Use a laser rangefinder to measure the distance from the tooling to the tank opening; S-B2: Establish a coordinate system and measure the angle between the vector prpt and the vector (0, -1), where pr is the coordinate of the center point of the rectangular discharge hole and pt is the coordinate of the center point of the auxiliary triangle hole; S-B3: The actual length of the vector prpt is Lmm, the coordinates of the center point of the rectangular discharge hole are pr(xr,yr), the coordinates of the center point of the auxiliary triangle hole are pt(xt,yt), and the coordinates of the laser point are pl(xl,yl); first calculate the equation of the line where the vector prpt is located: Ax+By+C=0; The above formula is the general form of the straight line equation. First, assume that the straight line equation where the vector prpt is located is y=kx+b, k is the slope of the straight line, and b is the intercept of the straight line. Substitute the known coordinates of the two center points: pr(xr, yr) and pt(xt, yt) into the slope-intercept straight line equation to solve for k and b; then y=kx+b is organized into the general form of the straight line equation, that is, Ax+By+C=0. At this time, the slope of the straight line is expressed as The intercept of the line is expressed as S-B4: Calculate the shortest distance from point pl to the straight line: S-B5: Calculate vectors The length l: S-B6: Convert pixel distance d to actual size S-B7: Calculate height

Citation Information

Patent Citations

  • CNN-based fruit and obstacle synchronous identification method and system and robot

    CN109948444A

  • Intelligent piled brick loading system based on computer vision and loading method thereof

    CN113192058A