A method of occluded target data acquisition and posture recognition based on image fusion

By automatically generating the occlusion target data set based on image fusion and using point cloud registration to obtain the target pose, the problems of high manual data acquisition cost, unbalanced data and low pose recognition efficiency in the prior art are solved, and efficient and balanced occlusion target detection and pose recognition are achieved.

CN116030316BActive Publication Date: 2025-05-20FUZHOU UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211680663.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-05-20
Estimated Expiration
2042-12-26

Smart Images

  • Figure CN116030316B_ABST
    Figure CN116030316B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for data collection and posture recognition of occluded targets based on image fusion. It includes randomly fusing images to generate a specific data set and using template point cloud information to perform grid downsampling and registration to obtain the target posture. The occluded data set of a specific target is generated by image fusion, reducing the cost of manual image collection; by statistically analyzing the amount of data of different targets in the image fusion process, the image fusion balance is ensured, and dynamic data set supplementation is performed to ensure the balance of the amount of data of different targets and reduce the overfitting degree of the detector; by grid downsampling, setting thresholds and other methods, the calculation amount of the point cloud registration process is reduced, and finally the lateral deviation angle is calculated to obtain the target posture; thus, a method for detecting and recognizing occluded targets and postures is proposed. The present invention can identify severely occluded targets and obtain the posture of the target, thereby realizing efficient occluded target image fusion and posture recognition, and has a very high engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for occluded target data acquisition and pose recognition based on image fusion. Background Art

[0002] In the industrial manufacturing industry, target detection, recognition, and grasping based on robot vision can greatly reduce the labor intensity of workers and improve production efficiency. Especially on assembly lines with stacked targets, accurately detecting and recognizing occluded targets is crucial for improving work efficiency and grasping success rate.

[0003] In real-world scenarios, due to the existence of occlusion, target detection remains a challenging task. Humans can continue and infer through the contours present in the scene even when partial information of an object is occluded or lost, so as to judge the attributes of the object. However, when an artificial neural network learns an occluded target, it cannot well extract feature information, thus making it difficult to accurately classify the object. Currently, there are three difficulties in the problem of occluded target detection: 1) Manually collecting occluded data takes more time and has inconsistent quality; 2) The quantity between different targets in the model training data is unbalanced, which may lead to overfitting of the model or a single training direction, and manual image collection cannot ensure the balance of data collection; 3) Existing methods have inaccurate target pose recognition or excessive computational complexity, only considering two-dimensional images without considering three-dimensional coordinates. In view of this, the present invention aims to automatically generate an occluded data set with reduced labor costs and obtain the target pose through efficient point cloud registration. By analyzing image fusion and occluded target detection algorithms, a method for occluded target data acquisition and pose recognition based on image fusion is proposed.

[0004] The Chinese patent application number is: CN 201911014564.X, and the name is: A target detection method based on multi-source sensor fusion. This method extracts the regions of interest of infrared images and visible light images; performs image fusion based on rolling guidance filtering and a weighted least squares optimization function to obtain the fused image F; inputs the fused image of the region of interest, and completes the detection of small and slow targets through a background modeling method. This method has a relatively large computational complexity for image fusion, and the detection process is not efficient, and it cannot ensure the rapid and efficient completion of tasks in an industrial environment.

[0005] The Chinese patent application number is: CN202210143660.X, and the title is: A method and device for detecting occluded objects. Training images and images to be recognized in a dense scene are obtained and preprocessed; the categories and position coordinates of the objects to be recognized in the preprocessed training images are labeled; an improved neural network model is established and the improved neural network model is trained using the labeled training images to obtain a detection model. The problem with this method is that a large amount of manpower is required to collect data during the process of obtaining training images in a dense scene, and the data quality cannot be guaranteed. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for collecting occluded target data and recognizing poses based on image fusion, including randomly fusing images to generate a specific data set and using template point cloud information for grid downsampling and registration to obtain the target pose. By fusing images in a certain way to generate an occluded data set for specific targets, the cost of manually collecting images is reduced; by statistically analyzing the data volume of different targets during the image fusion process, the balance of image fusion is ensured, dynamic data set supplementation is carried out to ensure the balance of the data volume of different targets, and the overfitting degree of the detector is reduced; by methods such as grid downsampling and setting thresholds, the computational amount in the point cloud registration process is reduced, and finally the horizontal deviation angle is calculated to obtain the target pose; thus, a method for detecting occluded targets and recognizing poses is proposed. The present invention can automatically generate an occluded target data set, reduce the cost of manually collecting images, ensure the balance of the data set, reduce the overfitting degree of the detector, recognize severely occluded targets, and obtain the pose of the target, realizing high-efficiency image fusion and pose recognition of occluded targets, and having very high engineering application value.

[0007] To achieve the above object, the technical solution of the present invention is: A method for collecting occluded target data and recognizing poses based on image fusion, including the following steps:

[0008] S1. Image preprocessing: Collect K RGB images of K types of target objects directly below the robotic arm and with the pose maintained at 90° with the end of the robotic arm; segment the images, and extract each type of target object in an image with pixels of (g rd , g gd , g bd ) to obtain K foreground images of different target objects, and collect a background image of the workbench without any target objects, that is, obtain K + 1 images;

[0009] S2. Image fusion: Randomly select i images from the K foreground images, and randomly rotate a predetermined angle with the center point coordinates of the target object as the rotation center; for the foreground images after random rotation, traverse all pixel values, and set the pixel values that are not (g rd , g gd , g bd) The pixel coordinates are stored in the set M = [(x 1 , y 1 ), (x 2 , y 2 ), … (x m , y m )]. To ensure that the rotated target object can be fully displayed in the picture, the center point of all coordinates in the set M is used as the expansion or cropping center, and the size of the rotated picture is expanded or cropped to a unified size (W, H);

[0010] S3. Traverse the pixel matrix again and update the coordinate information in M; determine whether all the point coordinates in M fall within the workbench range x ∈ (x s , x e ), y ∈ (y s , y e ). Take the center of the workbench If there is a point not within the workbench range, make a case-by-case judgment and perform a translation operation;

[0011] S4. After completing the translation operation, replace the pixel values in the workbench background image that are not (g rd , g gd , g bd ) with the pixel values at the same position in the foreground image, generate n pictures in batches, and count the number of times C i that K foreground images are randomly used. Store Ci in the set N = [C 1 , C 2 , … C K ;

[0012] S5. To ensure the balance of data collection, determine whether C 1 , C 2 , … C K in the set N are equal; if not, sort the set N from largest to smallest to get N = [C m1 , C m2 , … C mK . Using the largest C m1 value in the set N as the standard, randomly rotate the foreground images except C m1 and fuse them to generate C m1 - C m2 pictures; update the set N and determine again whether the other C i-1 in N are equal to C m1 . If one or more of the remaining C i-1 are equal to C m1 , then only perform the same operation on the remaining C i that are not equal to C m1 until all C i = Cm1 ;

[0013] S6. Divide the generated data set into a training set and a validation set, and use a convolutional neural network to train to obtain a detector; use the detector to detect the target object, locate the coordinates of the target object detection box in the depth map, obtain the RGB image and the depth map of the target object through the coordinates of the detection box, and obtain the corresponding point cloud map from the RGB-D image;

[0014] S7. Perform grid downsampling on the obtained point cloud map. The object point cloud map that is 90° to the end of the robotic arm is used as the template point, denoted as the set Q = {q 1 , q 2 ,....q n}; The detected target object point cloud map is used as the target point, denoted as the set P = {p 1 , p 2 ,....p n}; Estimate the coordinates of the corresponding points in the set P according to the points in the set Q using the nearest neighbor method. To simplify the calculation amount, set a distance threshold θ. When the distance is less than the threshold θ, it is used as the corresponding point; Calculate the Euclidean transformation: p i = Rq i + t; where R is the rotation matrix and t is the translation matrix; Perform iterative calculation based on the least squares method to make the sum of squared errors reach the minimum value: Output the final rotation matrix R and translation matrix t as the pose transformation;

[0015] S8. Online detection: When the model trained with the image fusion data set detects the target, the coordinate information of the detection box of the target object is located in the depth map, and the corresponding target point cloud map is obtained from the RGB-D image; Then perform point cloud registration on the target to obtain the rotation matrix R and the translation matrix t. The end pose of the robotic arm meets the requirements of the rotation matrix and the translation matrix, and reaches the target point to grasp the target object.

[0016] In an embodiment of the present invention, step S2 is specifically implemented as follows:

[0017] S21. Before random rotation, traverse the pixel matrix of the foreground image, and store the pixel point coordinates whose pixel values are not (g rd , g gd , g bd ) in the set M = [(x 1 , y 1 ), (x 2 , y 2 ), …(x m , y m )]; Take the th or th coordinate in the set M th or Randomly rotate the foreground image around the rotation center (x c , y c );

[0018] S22. Obtain the size of the foreground image after rotation. At the same time, for the foreground image after random rotation, traverse all pixel values, and store the coordinates of the pixel points whose pixel values are not (g rd , g gd , g bd ) in the set M = [(x 1 , y 1 ), (x 2 , y 2 ), … (x m , y m )]. To ensure that the sizes of the training set images are all (W, H), it is necessary to pad or crop the foreground image after rotation; judge the difference between the size (W R , H R ) of the rotated image and (W, H), calculate the differences w R , h R = W 1 - W and h R = H 1 - H R , and select the correction form according to different situations; the specific correction methods are as follows:

[0019] Case 1: When h 1 < 0 and w 1 > 0, expand the height using pixel points with pixel values of (g rd , g gd , g bd ) to H; if supplement columns of pixel points with pixel values of (g rd , g gd , g bd ) at the right end of the pixel matrix; if then supplement columns of pixel points with pixel values of (g rd , g gd , g bd ) at the left end of the pixel matrix; if do not make any supplement; finally, use as the width and (0, H) as the height;

[0020] Case 2: When h 1 < 0 and w 1 < 0, expand the height using pixel points with pixel values of (g rd , g gd , g bd ​) Expand the pixel points to H; for the width, use pixel values of (g rd , g gd , g bd ) to expand the pixel points to W; finally, use (0, W) as the width and (0, H) as the height;

[0021] Case 3: When h 1 > 0 and w 1 < 0, for the width, use pixel values of (g rd , g gd , g bd ) to expand the pixel points to W; if Add at the end of the pixel matrix rows of pixel values of (g rd , g gd , g bd ) of pixel points; if Add at the start of the pixel matrix rows of pixel values of (g rd , g gd , g bd ) of pixel points; if Do not make any addition; finally, use (0, W) as the width and as the height;

[0022] Case 4: When h 1 > 0 and w 1 > 0, if Add at the right end of the pixel matrix columns of pixel values of (g rd , g gd , g bd ) of pixel points; if Add at the left end of the pixel matrix columns of pixel values of (g rd , g gd , g bd ) of pixel points; if Then do not make any addition; if Add at the end of the pixel matrix rows of pixel values of (g rd , g gd , g bd ) of pixel points; if Add at the start of the pixel matrix rows of pixel values of (g rd , g gd , g bd ) of pixel points; if Do not make any processing; finally, use as the width, as the height;

[0023] Case 5: 1) When h 1 = 0 and w 1 = 0, no cropping is required, and the original image can be retained; 2) When h 1 = 0 and w 1 > 0, if add columns of pixels with pixel values (g rd , g gd , g bd ) at the right end of the pixel matrix; if add columns of pixels with pixel values (g rd , g gd , g bd ) at the left end of the pixel matrix; finally, both use as the width and (0, H) as the width; 3) When h 1 = 0 and w 1 < 0, expand the width to W using pixels with pixel values (g rd , g gd , g bd ); finally, both use (0, W) as the width and (0, H) as the height; 4) When w 1 = 0 and h 1 > 0, if add rows of pixels with pixel values (g rd , g gd , g bd ) at the end of the pixel matrix; if add rows of pixels with pixel values (g rd , g gd , g bd ) at the start of the pixel matrix; finally, both use (0, W) as the width and as the height; 5) When w 1 = 0 and h 1 < 0, expand the height to H using pixels with pixel values (g rd , g gd , g bd ); finally, both use (0, W) as the width and (0, H) as the height.

[0024] In an embodiment of the present invention, in step S3, it is determined whether all the point coordinates in M fall within the workbench range x ∈ (x s , x e ), y ∈ (y s , y e ), and the center of the workbench is taken If there is a point not within the workbench range, it is judged and translated in different cases:

[0025] Case 1: If there is an abscissa outside the workbench range, the difference between the first point outside the range and the abscissa of the center point is obtained

[0026] Update the abscissa of all points in set M to x + Δw; at the same time, traverse all points in set M again to determine whether there are points outside the range. If so, repeat the above operation until the abscissas of all points are within the workbench range;

[0027] Case 2: If there is an ordinate outside the workbench range, the difference between the first point outside the range and the ordinate of the center point is obtained

[0028] Update the ordinate of all points in set M to y + Δw; at the same time, traverse all points in set M again to determine whether there are points outside the range. If so, repeat the above operation until the ordinates of all points are within the workbench range.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) The present invention effectively improves the efficiency of collecting images. There is no need for manual spending a lot of time to collect occluded target data. The image fusion method can be used to generate images online, greatly reducing the labor cost.

[0031] (2) Count the different target data amounts during the image fusion process, ensure the balance of image fusion, perform dynamic data set supplementation, ensure the balance of different target data amounts, reduce the overfitting degree of the detector, and greatly reduce the error of the manual image collection that cannot ensure the balance of the data set.

[0032] (3) Obtain the detection frame coordinate information through the detected target RGB image and locate it in the depth image, and obtain the point cloud image of the corresponding target from RGB-D. Finally, register it with the template point cloud to obtain the rotation matrix R and the translation matrix t, ensuring the success rate of the robotic arm grasping. Description of the Drawings

[0033] Figure 1 It is the flowchart of the present invention for completing image fusion and pose recognition.

[0034] Figure 2 It is the RGB image of the pliers collected by the present invention directly below the robotic arm and with the pose maintained at 90° with the end of the robotic arm.

[0035] Figure 3 It is the randomly rotated pliers image after image segmentation of the present invention.

[0036] Figure 4 It is the flowchart of the present invention for completing the correction of the image size after rotation.

[0037] Figure 5 This is the flowchart for optimizing the translation of the target point coordinates in the present invention.

[0038] Figure 6 This is the flowchart for optimizing the balance of the occlusion dataset in the present invention.

[0039] Figure 7 This is the occluded image after fusing the images of scissors and pliers in the present invention.

[0040] Figure 8 This is the image after point cloud registration of scissors in the present invention. Detailed implementation manners

[0041] The technical solution of the present invention will be specifically described below with reference to the accompanying drawings.

[0042] As Figure 1 shown, the flowchart of the method for collecting and posture recognizing occluded target data based on image fusion is as follows:

[0043] 1) Use Kinect 2.0 to take 5 RGB and depth images of K = 5 kinds of target objects directly below the robotic arm and with the posture maintained at 90° to the end of the robotic arm, as Figure 2 shown. Segment the images, and extract each type of target in a colorless background image with pixels of (g rd , g gd , g bd ) = (255, 255, 255) to obtain 5 foreground images of different targets, and collect a background image of the workbench without any target objects, that is, 6 images are obtained.

[0044] 2) Randomly select 2 - 4 of the K foreground images and randomly rotate them by a certain angle, as Figure 3 shown. For the foreground images after random rotation, judge all pixel values, and store the coordinates of the pixel points that are not (255, 255, 255) in the set M = [(x 1 , y 1 ), (x 2 , y 2 ), … (x m , y m )]. To ensure that the rotated target can be completely displayed in the picture, use the center point of all coordinates in the set M as the center for expansion or cropping, and expand or crop the size of the rotated picture to a unified size (1920, 1080), specifically as follows:

[0045] 2.1) Before random rotation, traverse the pixel matrix of the foreground image, and store the coordinates of the pixel points with pixel values not (255, 255, 255) in the set M = [(x 1,y 1 ),(x 2 ,y 2 ),…(x m ,y m )], take the or Coordinates or As the center of rotation (x c ,y c ).

[0046] 2.2) Get the size of the foreground image after rotation. At the same time, for the foreground image after random rotation, traverse all pixel values ​​and store the pixel coordinates of the pixel points whose pixel values ​​are not (255,255,255) in the set M = [(x 1 ,y 1 ),(x 2 ,y 2 ),…(x m ,y m )], in order to ensure that the training set image size is (1920, 1080), it is necessary to fill or crop the foreground image after rotation. Determine the image size after rotation (W R ,H R ) and (1920, 1080), calculate (W R ,H R The difference between (1920, 1080) and (1920, 1080) is w 1 =W R -1920 and h 1 =H R -1080. Choose the correct form according to different situations, such as Figure 4 , the specific correction method is as follows:

[0047] Case 1: When h 1 <0 and w 1 >0, the high-utilization pixel value is (255,255,255) expanded to 1080. If x c >960, add 960 columns of pixels with pixel values ​​of (255,255,255) at the right end of the pixel matrix; if x c <960, then add 960 columns of pixels with pixel values ​​of (255,255,255) at the left end of the pixel matrix. If x c =960, no additional information is added. Finally, (x c -960,x c +960) as the width and (0,1080) as the height.

[0048] Case 2: When h 1 <0 and w 1 ​​​​​​< 0, expand the height using the pixel value (255, 255, 255) to 1080. Expand the width using the pixel value (255, 255, 255) to 1920. Finally, use (0, 1920) as the width and (0, 1080) as the height.

[0049] Case 3: When h 1 > 0 and w 1 < 0, expand the width using the pixel value (255, 255, 255) to 1920. If y c > 540, supplement 540 rows of pixel points with the pixel value (255, 255, 255) at the end of the pixel matrix; if y c < 540, supplement 540 rows of pixel points with the pixel value (255, 255, 255) at the beginning of the pixel matrix. If y c = 540, do not make any supplement. Finally, use (0, 1920) as the width and (y c - 540, y c + 540) as the height.

[0050] Case 4: When h 1 > 0 and w 1 > 0, if x c > 960, supplement 960 columns of pixel points with the pixel value (255, 255, 255) at the right end of the pixel matrix; if x c < 960, supplement 960 columns of pixel points with the pixel value (255, 255, 255) at the left end of the pixel matrix; if x c = 960, then do not make any supplement. Use (x c - 960, x c + 960) as the width. If y c > 540, supplement 540 rows of pixel points with the pixel value (255, 255, 255) at the end of the pixel matrix; if y c < 540, supplement 540 rows of pixel points with the pixel value (255, 255, 255) at the beginning of the pixel matrix. If y c = 540, do not make any processing. Use (x c - 960, x c + 960) as the width and (y c - 540, y c + 540) as the height.

[0051] Case 5: When h 1 = 0 and w 1 = 0, no cropping is required, just keep the original picture. When h 1 = 0 and w 1 > 0, if x c> 960, supplement 960 columns of pixels with pixel values (255, 255, 255) at the right end of the pixel matrix; if x c <960, supplement 960 columns of pixels with pixel values (255, 255, 255) at the left end of the pixel matrix. Both use (x c -960, x c +960) as the width and (0, 1080) as the height. When h 1 =0 and w 1 <0, expand the width to 1920 using pixels with pixel value (255, 255, 255). Finally, both use (0, 1920) as the width and (0, 1080) as the height. When w 1 =0 and h 1 >0, if y c >540, supplement 540 rows of pixels with pixel values (255, 255, 255) at the end of the pixel matrix; if y c <540, supplement 540 rows of pixels with pixel values (255, 255, 255) at the start of the pixel matrix. Finally, both use (0, 1920) as the width and (y c -540, y c +540) as the height. When w 1 =0 and h 1 <0, expand the height to H using pixels with pixel value (255, 255, 255). Finally, both use (0, 1920) as the width and (0, 1080) as the height.

[0052] 3) Traverse the pixel matrix again and update the coordinate information in M. Determine whether all the point coordinates in M fall within the workbench range x ∈ (500, 970), y ∈ (200, 690), and take the workbench center (735, 445). If there is a point not within the workbench range, make a case-by-case judgment and perform a translation operation, such as Figure 5 , the specific method is as follows:

[0053] Case 1: If the abscissa is not within the workbench range, calculate the difference between the abscissa of the first point outside the range and the abscissa of the center point to get Δw = 735 - x m , update the abscissa of all points in set M to x + Δw. At the same time, traverse all points in set M again to determine whether there are points outside the range. If there are, repeat the above operation until all the abscissas of the points are within the workbench range.

[0054] Case 2: If the ordinate is not within the workbench range, calculate the difference between the ordinate of the first point outside the range and the ordinate of the center point to get Δh = 445 - y m, update the vertical coordinates of all points in set M to y + Δw. At the same time, traverse all points in set M again to determine if there are any points outside the range. If there are, repeat the above operations until the vertical coordinates of all points are within the workbench range.

[0055] 4) After completing the translation operation, replace the pixel values in the background image that are not (255, 255, 255) with the pixel values at the same position in the foreground image, and batch generate 2000 images, such as Figure 6 , count the number of times C that 5 foreground images are randomly used i , store Ci in set N = [C 1 , C 2 , … C 5 .

[0056] 5) Determine if C 1 , C 2 , … C 5 in set N are equal. If N = [731, 780, 822, 844, 872]. Balance the dataset, such as Figure 7 , the specific method is as follows: sort set N from largest to smallest to get N = [872, 844, 822, 780, 731]. Use the largest C 5 = 872 value in set N as the standard, and use the foreground images except C 5 to perform random rotation and fusion to generate C 5 - C 4 = 28 images. Update set N and determine if other C i-1 in N are equal to C 5 . If one or more of the remaining C i-1 are equal to C 5 , then only perform the above operations on the remaining C i that are not equal to C 5 until all C i = C 5 .

[0057] 6) Divide the generated dataset into a training set and a validation set, and use a convolutional neural network for training to generate a detector. Use the detector to detect the target object, locate the coordinates of the target object detection box in the depth map as well, obtain the RGB image and depth map of the target object through the coordinates of the detection box, and obtain the corresponding point cloud map from the RGB - D map.

[0058] 7) Perform grid downsampling on the obtained point cloud map. The object point cloud map that is 90° to the end of the robotic arm is used as the template point, denoted as set Q = {q 1 , q 2 ,....q n}, and the detected point cloud map of the target object is used as the target point, denoted as the set P = {p 1 , p 2 ,.... p n}. According to the points in the set Q, the coordinates of the corresponding points in the set P are estimated using the nearest neighbor method. To simplify the calculation, a distance threshold θ is set. When the distance is less than the threshold θ, it is regarded as the corresponding point. Calculate the Euclidean transformation: p i = Rq i + t. Where R is the rotation matrix and t is the translation matrix. Based on the least squares method, iterative calculations are performed to make the sum of squared errors reach the minimum value: to obtain the point cloud registration map, as shown in Figure 8 , and the final rotation matrix R and translation matrix t are output as the pose transformation.

[0059] 8) Online detection: When the model trained with the image fusion dataset is used to detect the target, the coordinate information of the detection box of the target object is located in the depth map, and the corresponding target point cloud map is obtained from the RGB-D map. Then, point cloud registration is performed on the target to obtain the rotation matrix R and the translation matrix t. The pose of the end of the robotic arm meets the requirements of the rotation matrix and the translation matrix, and it reaches the target point to grasp the target object.

[0060] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects produced do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.

Claims

1. A method for occluded target data collection and posture recognition based on image fusion, characterized in that: The steps include: S1. Image preprocessing: Collect K RGB images of K kinds of target objects directly under the robot arm and with the posture maintained at 90° to the end of the robot arm; segment the image and extract each type of target object into a pixel (g rd ,g gd ,g bd ), obtain K foreground images of different target objects, and collect a workbench background image that does not contain any target object, that is, obtain K+1 images; S2, image fusion: randomly select i images from K foreground images, and randomly rotate them by a predetermined angle with the coordinates of the center point of the target object as the rotation center; for the foreground image after random rotation, traverse all pixel values ​​and convert pixel values ​​other than (g rd ,g gd ,g bd )’s pixel coordinates are stored in the set M = [(x1,y1),(x2,y2),…(x m ,y m )], in order to ensure that the rotated target object can be fully displayed in the picture, the center point of all coordinates in the set M is used as the expansion or cropping center, and the size of the rotated picture is expanded or cropped to a uniform size (W, H); S3, re-traverse the pixel matrix and update the coordinate information in M; determine whether the coordinates of all points in M ​​fall within the workbench range x∈(x s ,x e ),y∈(y s ,y e ), take the center of the workbench If there is a point that is not within the range of the workbench, judge and perform translation operation according to the situation; S4. After the translation operation is completed, the pixel value in the background image of the workbench is non-(g rd ,g gd ,g bd ) with the pixel value of the same position in the foreground image, generate n images in batches, and count the number of times C the K foreground images are randomly used i , store Ci into the set N = [C1, C2, ... C K ]; S5. To ensure the balance of data collection, determine C1, C2, ...C in the set N. K Are they equal? ​​If not, sort the set N from large to small to get N = [C m1 ,C m2 ,…C mK ], with the largest C in the set N m1 Values ​​are used as standard, except for C m1 The foreground image outside is randomly rotated and fused to generate C m1 -C m2 pictures; update the set N, and judge the other C in N again i-1 Is it equal to C? m1 , if the rest of C i-1 There are one or more equal to C m1 , then only for the remaining C i Not equal to C m1 The same operation is performed on the foreground image until all C i =C m1 ; S6, dividing the generated data set into a training set and a validation set, and using a convolutional neural network to train a detector; The detector is used to detect the target object, and the coordinates of the target object detection frame are located in the depth map. The RGB image and depth map of the target object are obtained through the coordinates of the detection frame, and the corresponding point cloud image is obtained from the RGB-D image. S7. Grid downsampling is performed on the obtained point cloud image, and the object point cloud image at 90° to the end of the robot arm is used as the template point, which is recorded as a set Q = {q1, q2, ....q n }; The point cloud of the detected target object is taken as the target point, which is recorded as the set P = {p1, p2, .... p n }; Based on the points in set Q, the nearest neighbor method is used to estimate the coordinates of the corresponding points in set P. To simplify the calculation, a distance threshold θ is set. When the distance is less than the threshold θ, it is regarded as the corresponding point; Calculate the Euclidean transformation: p i =Rq i +t; where R is the rotation matrix and t is the translation matrix; iterative calculation is performed based on the least squares method to minimize the sum of squared errors: Output the final rotation matrix R and translation matrix t as the pose transformation; S8, online detection: After the model trained with the image fusion dataset detects the target, the detection frame coordinate information of the target object is located in the depth map, and the corresponding target point cloud map is obtained from the RGB-D map; then the point cloud of the target is registered to obtain the rotation matrix R and the translation matrix t. The posture of the end of the robot arm meets the requirements of the rotation matrix and the translation matrix, and reaches the target point to grab the target object; Step S2 is specifically implemented as follows: S21, before random rotation, traverse the pixel matrix of the foreground image and convert the pixel values ​​non-(g rd ,g gd ,g bd )’s pixel coordinates are stored in the set M = [(x1,y1),(x2,y2),…(x m ,y m )], take the first or Coordinates or As the rotation center (x c ,y c ) Randomly rotate the foreground image; S22, obtain the size of the foreground image after rotation, and at the same time, traverse all pixel values ​​of the foreground image after random rotation, and convert the pixel values ​​non-(g rd ,g gd ,g bd )’s pixel coordinates are stored in the set M = [(x1,y1),(x2,y2),…(x m ,y m )], in order to ensure that the image size of the training set is (W, H), it is necessary to fill or crop the foreground image after rotation; determine the image size after rotation (W R ,H R ) and (W, H), calculate (W R ,H R ) and W, H) w1 = W R -W and h1 = H R -H, select the correction form according to different situations; the specific correction methods are as follows: Case 1: When h1<0 and w1>0, the high utilization pixel value is (g rd ,g gd ,g bd ) is expanded to H; if Add at the right end of the pixel matrix The column pixel value is (g rd ,g gd ,g bd ) pixels; if Then add at the left end of the pixel matrix The column pixel value is (g rd ,g gd ,g bd ) pixels; if No additional information is provided; as width and (0,H) as height; Case 2: When h1<0 and w1<0, the high utilization pixel value is (g rd ,g gd ,g bd ) pixels are expanded to H; the pixel value for width is (g rd ,g gd ,g bd ) is expanded to W; finally, (0,W) is used as the width and (0,H) is used as the height; Case 3: When h1>0 and w1<0, the pixel value for width is (g rd ,g gd ,g bd ) is expanded to W; if Add at the end of the pixel matrix The row pixel value is (g rd ,g gd ,g bd ) pixels; if Add at the beginning of the pixel matrix The row pixel value is (g rd ,g gd ,g bd ) pixels; if No additional information is added; finally, (0, W) is used as the width. as height; Case 4: When h1>0 and w1>0, if Add at the right end of the pixel matrix The column pixel value is (g rd ,g gd ,g bd ) pixels; if Supplement at the left end of the pixel matrix The column pixel value is (g rd ,g gd ,g bd ) pixels; if No supplement will be made; if Add at the end of the pixel matrix The row pixel value is (g rd ,g gd ,g bd ) pixels; if Add at the beginning of the pixel matrix The row pixel value is (g rd ,g gd ,g bd ) pixels; if No processing is done; all are finally As the width, as height; Case 5: 1) When h1 = 0 and w1 = 0, no need to crop, just keep the original image; 2) When h1 = 0 and w1>0, if Add at the right end of the pixel matrix The column pixel value is (g rd ,g gd ,g bd ) pixels; if Supplement at the left end of the pixel matrix The column pixel value is (g rd ,g gd ,g bd ) pixels; finally As the width, (0,H) is used as the width; 3) When h1=0 and w1<0, the pixel value for the width is (g rd ,g gd ,g bd ) pixels are expanded to W, and finally (0,W) is used as the width and (0,H) is used as the height; 4) When w1=0 and h1>0, if Add at the end of the pixel matrix The row pixel value is (g rd ,g gd ,g bd ) pixels; if Add at the beginning of the pixel matrix The row pixel value is (g rd ,g gd ,g bd ) pixels; finally, (0,W) is used as the width. As height; 5) When w1 = 0 and h1 < 0, the pixel value for height is (g rd ,g gd ,g bd ) is expanded to H; finally, (0,W) is used as the width and (0,H) is used as the height.

2. The method for occluded target data collection and posture recognition based on image fusion according to claim 1, characterized in that: In step S3, it is determined whether the coordinates of all points in M ​​fall within the workbench range x∈(x s ,x e ),y∈(y s ,y e ), take the center of the workbench If a point is not within the range of the workbench, judge and perform translation operations according to the situation: Case 1: If the horizontal coordinate is not within the range of the workbench, the first point outside the range is obtained by subtracting the horizontal coordinate of the center point. Update the horizontal coordinates of all points in the set M to x+Δw; at the same time, traverse all points in the set M again to determine whether there are points outside the range. If so, repeat the above operation until the horizontal coordinates of all points are within the range of the workbench; Case 2: If there is a vertical coordinate that is not within the range of the workbench, the vertical coordinate of the first point outside the range is obtained by subtracting the vertical coordinate of the center point. Update the ordinates of all points in the set M to y+Δw; at the same time, traverse all points in the set M again to determine whether there are points outside the range. If so, repeat the above operation until the ordinates of all points are within the range of the workbench.

Citation Information

Patent Citations

  • Target detection method based on multi-source sensor fusion

    CN110766676B

  • Shielding object detection method and device

    CN114187491A

  • Three-dimensional target sensing method in vehicle-mounted edge scene

    CN113506318A

  • Method of constructing indoor two-dimensional semantic map with wall corner as critical feature based on robot platform

    US20220244740A1