Object pose parameter acquisition method, object disordered grabbing method and system and medium

Through deep learning and image processing technology, a convolutional neural network combining threshold segmentation and attention mechanisms solves the problem of accurate grabbing of material boxes under complex stacking in turnover boxes, and realizes an efficient and intelligent box grabbing process, improving production efficiency and grabbing accuracy.

CN120388068APending Publication Date: 2025-07-29JINSHENG INTELLIGENT EQUIPMENT (SUZHOU) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510404316.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the stacking of material boxes in the turnover box is complex and changeable, the lighting conditions are different, the number of material boxes is uncertain, and there are multiple stacking methods, which leads to the inefficiency of traditional loading methods and is difficult to achieve accurate grasping. The calculation of point cloud map matching method of 3D camera material grabbing technology is time-consuming and complex, making it difficult to meet efficient production needs.

Method used

Deep learning and image processing technology are used to preprocess and threshold segment the depth image, and the convolutional neural network with attention mechanism is used to extract the box features, combine the clustering algorithm to dynamically calculate the threshold, accurately segment the box area, and generate the driving parameters of the robot for grabbing.

Benefits of technology

It improves the accuracy and speed of material box grabbing, reduces calculation time, adapts to different lighting conditions and stacking methods, realizes intelligent and automated material box grabbing, reduces labor intensity, and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388068A_ABST
    Figure CN120388068A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an object pose parameter acquisition method, an object disordered grabbing method and system and a medium, and the object pose parameter acquisition method comprises the steps: carrying out the preprocessing of a depth image containing an object, and obtaining a preprocessed image; dividing the preprocessed image by using a threshold segmentation algorithm to obtain an object region only containing the object; wherein a threshold value used by the threshold value segmentation algorithm is calculated according to a gray value of each pixel in the preprocessed image; and inputting the object area into a preset pose parameter prediction model for pose prediction, and outputting to obtain pose parameters of the object. According to the object pose parameter obtaining method, the pose parameters of the material boxes can be rapidly and accurately obtained under different illumination conditions and in the disordered stacking state of the material boxes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of material grasping, and specifically relates to a method for obtaining object pose parameters, a method for disorderly grasping objects, a system and a medium. Background Art

[0002] In the semiconductor industry, chip products are usually packaged in cassettes and stacked in turnover boxes for circulation. The traditional feeding method relies on manual workers to take out the cassettes from the turnover boxes. This method not only has a large labor intensity but also low efficiency, and it is difficult to meet the high requirements for efficiency and automation in modern production. With the continuous growth of the demand for automated production, it has become an inevitable trend in the industry to introduce manipulators to replace manual workers for feeding.

[0003] In the actual production environment, the stacking conditions of the cassettes in the turnover boxes are complex and variable. For example, the lighting conditions are different, or the number of cassettes is uncertain and there are various stacking methods, including non-full layer stacking and rotation of the cassette tubes, etc. This makes it impossible to adopt a fixed-point strategy to achieve accurate grasping in the feeding process. In addition, in the existing 3D camera material grasping technology, the commonly used point cloud map matching method has many problems, such as a large amount of data, complex algorithm writing and time-consuming calculation, etc., and it is difficult to meet the requirements of high-efficiency production. Summary of the Invention

[0004] Aiming at the technical defects in the prior art, the purpose of the embodiments of the present invention is to provide a method for obtaining object pose parameters, a method for disorderly grasping objects, a system and a medium, which can accurately obtain the pose information of the cassettes in the turnover box, use a manipulator to achieve accurate grasping and feeding, and overcome the shortcomings of the point cloud map method in the prior art, improving production efficiency and grasping accuracy.

[0005] To achieve the above purpose, in the first aspect, the embodiments of the present invention provide a method for obtaining object pose parameters, including:

[0006] Preprocess a depth image containing an object to obtain a preprocessed image;

[0007] Use a threshold segmentation algorithm to divide the preprocessed image to obtain an object region containing only the object; wherein, the threshold used in the threshold segmentation algorithm is calculated according to the gray value of each pixel in the preprocessed image; and,

[0008] Input the object region into a preset pose parameter prediction model for pose prediction, and output the pose parameters of the object.

[0009] Optionally, the method for calculating the threshold according to the gray value of each pixel in the preprocessed image includes:

[0010] The pixel gray values in the preprocessed image are divided into two different clusters by a clustering algorithm, and the central gray values of the two clusters are found;

[0011] The average value of the central gray values of the two clusters is calculated as the threshold; where,

[0012] In the two clusters, one cluster represents the cartridge area and the other cluster represents the background area.

[0013] Optionally, the pose parameter prediction model includes a residual network, and a channel attention module, a multi-layer perceptron, and a spatial attention module in sequential combination are inserted between the last convolutional layer and the fully connected layer of the residual network; where, the multi-layer perceptron includes two fully connected layers;

[0014] The loss function of the pose parameter prediction model is the mean square error loss function.

[0015] Optionally, preprocessing the depth image containing the object includes:

[0016] Performing Gaussian filtering on the depth image; where,

[0017] The convolution kernel radius and the standard deviation of the Gaussian distribution used for the Gaussian filtering are obtained by predicting by inputting the depth image into a preset filtering parameter acquisition model.

[0018] Optionally, before inputting the depth image into the preset filtering parameter acquisition model, the method further includes:

[0019] Dividing the depth image into several local blocks; where, each local block is used as the input of the filtering parameter acquisition model.

[0020] Optionally, the filtering parameter acquisition model includes an input layer, a first convolutional layer, an activation function, a second convolutional layer, a max pooling layer, and an output layer; where,

[0021] The output layer includes a fully connected layer;

[0022] The first convolutional layer uses n1 convolutional kernels of size f1×f1 with a stride of s1;

[0023] The activation function uses the ReLU activation function;

[0024] The second convolutional layer uses n2 convolutional kernels of size f2×f2 with a stride of s2;

[0025] The pooling window size of the max pooling layer is p×p with a stride of p.

[0026] Second aspect, an embodiment of the present invention provides an object pose parameter acquisition system, including:

[0027] A preprocessing module, configured to preprocess a depth image including an object to obtain a preprocessed image;

[0028] An image partitioning module, configured to partition the preprocessed image by using a threshold segmentation algorithm to obtain an object region that only includes the object; wherein, the threshold used by the threshold segmentation algorithm is calculated according to the gray value of each pixel in the preprocessed image; and,

[0029] A pose parameter acquisition module, configured to input the object region into a preset pose parameter prediction model for pose prediction, and output the pose parameters of the object.

[0030] Third aspect, an embodiment of the present invention provides an object disordered grasping method, including:

[0031] Receiving a depth image including an object captured by a 3D camera;

[0032] Obtaining the pose parameters of the object according to the depth image;

[0033] Generating driving parameters of the manipulator according to the pose parameters, in combination with the kinematic model and workspace parameters of the manipulator; wherein, the driving parameters include grasping coordinates and a motion path;

[0034] Driving the manipulator to grasp the object according to the driving parameters; wherein,

[0035] The method for obtaining the pose parameters of the object according to the depth image is as described in the first aspect.

[0036] Fourth aspect, an embodiment of the present invention provides an object disordered grasping system, including:

[0037] A depth image receiving module, configured to receive a depth image including an object captured by a 3D camera;

[0038] A pose parameter acquisition module, configured to obtain the pose parameters of the object according to the depth image;

[0039] A driving parameter generation module, configured to generate driving parameters of the manipulator according to the pose parameters, in combination with the kinematic model and workspace parameters of the manipulator; wherein, the driving parameters include grasping coordinates and a motion path;

[0040] A driving control module, configured to drive the manipulator to grasp the object according to the driving parameters; wherein,

[0041] The method for the pose parameter acquisition module to acquire the pose parameters of the object according to the depth image is as described in the first aspect.

[0042] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method as described in the third aspect.

[0043] When implementing the object pose parameter acquisition method provided by the embodiments of the present invention, since the threshold T used in the threshold segmentation algorithm is dynamically calculated according to actual image data, it can adapt to different lighting conditions and depth map characteristics, thereby improving the accuracy of object region segmentation in the depth image.

[0044] Moreover, in the pose parameter prediction model used in this object pose parameter acquisition method, an attention module is inserted after the last convolutional layer of the basic residual network, enabling the pose parameter prediction model to weight features according to the importance of the channel and spatial dimensions, thus paying more attention to the key information related to the object pose, which helps improve the accuracy and robustness of pose estimation. At the same time, by abandoning the traditional point cloud map matching method and adopting deep learning and image processing technologies, the data processing volume and calculation time are reduced, greatly improving the speed of pose parameter acquisition.

[0045] Finally, in the depth image preprocessing stage, by inputting the depth image into a preset filtering parameter acquisition model, the convolution kernel radius and the standard deviation of the Gaussian distribution used for Gaussian filtering processing are obtained, thereby realizing adaptive filtering processing to obtain the best image quality, greatly improving the efficiency and accuracy of image processing.

[0046] Similarly, when the object disordered grasping method provided by the embodiments of the present invention is applied in the field of material grasping, it can effectively handle various stacking methods of the material box in the turnover box, including the situation of not being full-layer and the rotation of the material pipe, etc. Without relying on the fixed-point strategy, it has strong adaptability and can be widely applied to the material box grasping tasks under different scenarios.

[0047] Since the adopted object pose parameter acquisition method abandons the traditional point cloud map matching method and adopts deep learning and image processing technologies, reducing the data processing volume and calculation time, it meets the requirements of high-efficiency production and significantly improves the material box grasping speed.

[0048] Accurately predict the Gaussian filtering parameters through a deep learning network, optimize the depth map preprocessing effect, dynamically calculate the threshold in combination with the clustering algorithm, accurately segment the material box area, and improve the grasping accuracy.

[0049] Using a convolutional neural network with an attention mechanism to deeply extract the features of the bin, accurately predict the pose parameters, enabling the manipulator to adjust the grasping posture in real time according to the actual pose of the bin, and ensuring the grasping success rate.

[0050] Finally, the entire grasping process is fully automated from depth map acquisition to manipulator control, without manual intervention, reducing labor intensity, improving production efficiency, and realizing the intelligence and automation of material grasping.

[0051] Based on the method for obtaining the object pose parameters, the present application also provides a system for obtaining object pose parameters. This system for obtaining object pose parameters has all the beneficial effects of the above method for obtaining object pose parameters and will not be elaborated specifically.

[0052] Based on the method for unordered grasping of objects, the present application also provides a system for unordered grasping of objects. This system for unordered grasping of objects has all the beneficial effects of the above method for unordered grasping of objects and will not be elaborated specifically. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art.

[0054] Figure 1 is a schematic flow chart of the method for unordered grasping of objects provided by an embodiment of the present invention;

[0055] Figure 2 is a schematic flow chart of the method for obtaining object pose parameters provided by an embodiment of the present invention;

[0056] Figure 3 is a schematic structural diagram of the system for unordered grasping of objects provided by an embodiment of the present invention;

[0057] Figure 4 is a schematic structural diagram of the system for obtaining object pose parameters provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0060] It should also be understood that the terminology used in this specification of the present invention is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0061] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0062] As used in this specification and the appended claims, the term "if" can be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0063] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meanings understood by those skilled in the art to which the present invention pertains.

[0064] In a first aspect, as Figure 1 shown, an object disordered grasping method provided by an embodiment of the present application can be used to control a manipulator to accurately grasp an object. The object can be a cassette for packaging chip products in the semiconductor industry, and the cassettes are stacked and palletized in a turnover box for circulation. By implementing the object disordered grasping method provided by the embodiment of the present application, the manipulator can be controlled to accurately grasp and load the cassettes with complex and variable stacking conditions in the turnover box, and overcome the shortcomings of the point cloud map method in the prior art, improving production efficiency and grasping accuracy.

[0065] Specifically, continuing to refer to Figure 1 , the object disordered grasping method may include the following steps:

[0066] S110: Receive a depth image containing the object captured by a 3D camera.

[0067] Specifically, a depth image containing an object is captured by using a 3D camera, and the gray value of each pixel point in the depth image represents the distance information from this point to the camera.

[0068] Specifically in the semiconductor industry, a depth image of a cartridge in a turnover bin can be obtained by using a 3D camera, and the gray value of each pixel point in the depth image of the cartridge represents the distance information from this point to the camera.

[0069] S120: Obtain the pose parameters of the object according to the depth image.

[0070] As Figure 2 shown, it shows a method for obtaining the pose parameters of the object according to the depth image, and this method specifically includes the following steps:

[0071] S121: Preprocess the depth image containing the object to obtain a preprocessed image.

[0072] In this embodiment, preprocessing the depth image containing the object includes: performing Gaussian filtering on the depth image. Among them, the convolution kernel radius and the standard deviation of the Gaussian distribution used in the Gaussian filtering are obtained by inputting the depth image into a preset filtering parameter acquisition model for prediction.

[0073] Specifically, the filtering parameter acquisition model includes an input layer, a first convolutional layer, an activation function, a second convolutional layer, a max pooling layer, and an output layer; among them, the output layer includes a fully connected layer; the first convolutional layer uses n1 convolution kernels of size f1×f1, with a stride of s1; the activation function uses the ReLU activation function; the second convolutional layer uses n2 convolution kernels of size f2×f2, with a stride of s2; the pooling window size of the max pooling layer is p×p, and the stride is p.

[0074] Next, a detailed description of the process of obtaining the convolution kernel radius and the standard deviation of the Gaussian distribution used in the Gaussian filtering corresponding to the depth image through the filtering parameter acquisition model in this embodiment is as follows:

[0075] Let the original depth map be I(x, y), the filtered depth map be I f (x, y), and the Gaussian kernel function be G(x, y, σ), then the filtering process can be expressed as:

[0076]

[0077] Among them, k is the convolution kernel radius, σ is the standard deviation of the Gaussian distribution, and the filtering effect can be controlled by adjusting the values of k and σ.

[0078] The specific steps of deep learning for adaptively adjusting the σ value and k value of Gaussian filtering are as follows:

[0079] a. Divide the depth image into several local blocks, and each local block serves as the input to the filtering parameter acquisition model. Assume the size of the depth map is H×W, and select the local block size as h×w. Normalize the local blocks of the depth map so that their pixel values are in the range of [0,1] to improve the stability and convergence speed of network training.

[0080] b. First convolutional layer: Use n1 convolutional kernels of size f1×f1 with a stride of s1. After the convolutional operation on the input local block, the size of the output feature map is ((h - f1) / s1 + 1)×((w - f1) / s1 + 1)×n1.

[0081] c. Activation function: Adopt the ReLU activation function. For the input x, its output is y = max(0, x).

[0082] d. Second convolutional layer: Use n2 convolutional kernels of size f2×f2 with a stride of s2. The size of the output feature map changes accordingly.

[0083] Among them, for the first convolutional layer, assume the input local block is X, the convolutional kernel is K1, and the mathematical representation of the convolutional operation is:

[0084] The calculation formula for the output feature map F1(i,j) is:

[0085]

[0086] where i and j are the coordinates of the output feature map, and m and n are the coordinates of the convolutional kernel.

[0087] The output after passing through the ReLU activation function is:

[0088] A1(i,j) = max(0, F1(i,j))

[0089] For the second convolutional layer, similarly, assume the convolutional kernel is K2, and the calculation formula for the output feature map F2 is:

[0090]

[0091] After passing through the ReLU activation function, we get A2(i,j) = max(0, F2(i,j)).

[0092] e. Max pooling layer: Add a max pooling layer after the convolutional layer. The size of the pooling window is p×p with a stride of p. Its function is to downsample the feature map, reduce the amount of data, and retain the main features at the same time. After pooling, the size of the feature map shrinks.

[0093] Taking max pooling as an example, assume the input feature map is A, the pooling window size is p×p, and the stride is p. For the element at coordinates (i, j) in the output feature map P after pooling, its calculation formula is:

[0094]

[0095] f. Fully connected layer: After several convolutional and pooling operations, the feature map is flattened into a one-dimensional vector. Connected to the fully connected layer, the number of output nodes in the fully connected layer is 2, corresponding to the σ value and k value of Gaussian filtering respectively.

[0096] Assume the flattened feature vector after the pooling layer is x, the weight matrix of the fully connected layer is W, and the bias vector is b. Then the output y of the fully connected layer (i.e., the predicted σ and k values) is:

[0097] y = σ, k = Wx + b

[0098] The final output of the filtering parameter acquisition model is two values, namely the predicted Gaussian filtering parameters σ and k. These two values will be used for Gaussian filtering operations on the local blocks of the input depth map.

[0099] It can be understood that the filtering parameter acquisition model provided in this embodiment can be obtained by training a convolutional neural network (CNN) model. The specific training process can refer to the existing neural network model training process, which will not be elaborated in this embodiment.

[0100] By inputting the depth image into a preset filtering parameter acquisition model to obtain the convolution kernel radius and the standard deviation of the Gaussian distribution used for Gaussian filtering processing, adaptive filtering processing is realized to obtain the best image quality, greatly improving the efficiency and accuracy of image processing.

[0101] S122: Use a threshold segmentation algorithm to divide the preprocessed image to obtain an object region that only contains the object; wherein, the threshold used in the threshold segmentation algorithm is calculated based on the gray value of each pixel in the preprocessed image.

[0102] Specifically, the method for calculating the threshold based on the gray value of each pixel in the preprocessed image includes:

[0103] Divide the pixel gray values in the preprocessed image into two different clusters through a clustering algorithm, and find the central gray values of the two clusters; calculate the average value of the central gray values of the two clusters as the threshold; wherein, in the two clusters, one cluster represents the cartridge area and the other cluster represents the background area. The specific process is as follows:

[0104] Assume the pixel set of the depth map is P, the number of pixels is N, and the gray value of each pixel is g p , p ∈ P;

[0105] Step A: Initialize two cluster centers of K-means clustering and (Two different grayscale values can be randomly selected).

[0106] Step B: Iterate the following steps until the cluster centers converge:

[0107] a. For each pixel p, calculate the distance d from it to the two cluster centers p1 =|g p -C1| and d p2 =

[0108] |g p -C2|, assign pixel p to the cluster with a closer distance;

[0109] b. Update cluster centers: Cluster1 and Cluster2 are the pixel sets of two clusters respectively, and |Cluster1| and |Cluster2| are the corresponding pixel numbers.

[0110] Step C: After convergence, obtain the final cluster centers C1 and C2 and calculate the threshold

[0111] Because threshold T is dynamically calculated based on actual image data, it can adapt to varying lighting conditions and depth map characteristics, thereby improving segmentation accuracy. Threshold T acts as a dividing point, dividing pixels in the depth map into two parts. For example, pixels with grayscale values above T can be considered part of the container area, while pixels with grayscale values below T can be considered the background area.

[0112] S123: Inputting the object area into a preset pose parameter prediction model to perform pose prediction, and outputting pose parameters of the object.

[0113] In this embodiment, the object region is a box used to package chip products in the semiconductor industry. A preset pose parameter prediction model can be used to extract and learn features from the box region image to predict the pose parameters of the box.

[0114] In this embodiment, the posture parameter prediction model includes a residual network, and a sequentially combined channel attention module, a multi-layer perceptron and a spatial attention module are inserted between the last convolutional layer and the fully connected layer of the residual network; wherein the multi-layer perceptron includes two fully connected layers; the loss function of the posture parameter prediction model is a mean square error loss function.

[0115] Specifically, the training process and working principle of the pose parameter prediction model are as follows:

[0116] Data Preparation: Based on the above image segmentation method, the dataset is increased by changing the orientation of the cartridge, thereby obtaining the data augmentation effect of the deep learning training model. A large number of image data of cartridges with different poses are collected, and the central position coordinates (x c , y c ) and the rotation angle θ of the cartridge in each image are labeled. The specific method is as follows:

[0117] Collection of Central Position: For the segmented cartridge area, the coordinates (x i , y i ) of all pixel points in this area are statistically analyzed. The central point coordinates (x c , y c ) of the cartridge can be calculated by the following formula:

[0118] where N is the total number of pixel points in the cartridge area, ∑ i x i is the sum of the x coordinates of all pixel points, Similarly.

[0119] Collection of Rotation Angle: Use the edge detection algorithm to extract the edge of the cartridge area. Then, perform linear fitting on the edge pixels to find the two main edge lines of the cartridge (for example, the lines corresponding to the long sides). Let the slopes of these two lines be k1 and k2 respectively, then the rotation angle θ can be calculated by the following formula for the rotation angle:

[0120]

[0121] Design of Pose Parameter Prediction Model with Attention Mechanism Introduction:

[0122] The network architecture selects the Residual Network (ResNet). By adding an attention mechanism to the last convolutional layer of ResNet, the feature map extracted by the basic network is first processed by the attention module and then input into the subsequent fully connected layer for pose parameter prediction. Specifically:

[0123] ① Feature Map Acquisition (Output of Basic Network)

[0124] Assume that the feature map output by the last convolutional layer of the ResNet network is F ∈ R C×H×W , where C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map.

[0125] ② Channel Attention Module

[0126] Global Average Pooling: Perform global average pooling on the feature map F to compress the feature map of each channel into a value, obtaining a C-dimensional vector z ∈ R C , and its calculation method is:

[0127] Among them, c represents the channel index, and i, j respectively represent the row and column indices of the feature map.

[0128] ③ Multilayer perceptron (MLP): Input the vector z into a multilayer perceptron containing two fully connected layers. The number of neurons in the middle layer is set to (r is a reduction ratio), and the number of neurons in the output layer is C.

[0129] The first fully connected layer: The output is δ(W1z + b1), where δ is the ReLU activation function.

[0130] The second fully connected layer: b2 ∈ R C , and the output is S c = σ(W2δ(W1z + b1) + b2), where σ is the Sigmoid activation function, to obtain the channel attention weight S c , S ∈ R C .

[0131] Channel weighting: Multiply the channel attention weight σ by the original feature map F to obtain the channel-weighted feature map F ′ :

[0132] F′ c,i,j = S c × F c,i,j

[0133] ④ Spatial attention module

[0134] Spatial information extraction: Perform global average pooling on the channel-weighted feature map F ′ to obtain a 1×H×W feature map The calculation method is:

[0135]

[0136] Perform global max pooling on F ′ to obtain a 1×H×W feature map That is:

[0137]

[0138] Convolution operation: Concatenate and in the channel dimension to obtain a 2×H×W feature map, and then pass it through a convolutional layer with a kernel size of 7×7 and an output channel number of 1, and then through the Sigmoid activation function to obtain the spatial attention weight M ∈ R 1×H×W .

[0139] Spatial weighting: Multiply the spatial attention weight M by the feature map F after channel weighting to obtain the final feature map F″ processed by the attention module: ′ F″

[0140] F″ c,i,j = M i,j × F′ c,i,j

[0141] ⑤ Fully connected layer and pose parameter prediction

[0142] Flatten the feature map F″ processed by the attention module into a vector and input it into the subsequent fully connected layer. Assume the weight of the fully connected layer is W fc ∈ R D×(C×H×W) and the bias is b fc ∈ R D (D is the number of neurons in the fully connected layer), then the output of the fully connected layer is:

[0143] y fc = W fc × flatten(F″)+ b fc

[0144] After passing through several fully connected layers, the final output vector where and are the pose parameters of the bin predicted by the network.

[0145] ⑥ Loss function and training

[0146] Adopt the mean squared error (MSE) loss function where N is the number of training samples, is the true pose annotation of the i-th sample.

[0147] Through the backpropagation algorithm, update all the parameters of the network (including the basic network parameters, attention module parameters, and fully connected layer parameters) according to the loss function, so that the pose parameters predicted by the network gradually approach the true values.

[0148] By inserting an attention module after the last convolutional layer of the basic network, the network can weight the features according to the importance of the channel and spatial dimensions, thus paying more attention to the key information related to the bin pose, which helps to improve the accuracy and robustness of pose estimation. At the same time, abandoning the traditional point cloud map matching method and adopting deep learning and image processing technologies reduce the amount of data processing and calculation time, greatly improving the speed of obtaining pose parameters.

[0149] S130: Generate the driving parameters of the manipulator based on the pose parameters, in combination with the kinematic model and workspace parameters of the manipulator; wherein, the driving parameters include the grasping coordinates and the motion path.

[0150] S140: Drive the manipulator to grasp the object according to the driving parameters.

[0151] Based on the calculated pose information of the magazine, in combination with the kinematic model and workspace parameters of the manipulator, generate the grasping coordinates and motion path of the manipulator. The corresponding control module drives the manipulator to move to the magazine position along a predetermined path according to the grasping coordinates and motion path of the manipulator, and performs the grasping operation.

[0152] In the semiconductor chip packaging production line workshop, the object disordered grasping method provided in this embodiment can be applied to the production line. The turnover box contains chip magazines with different stacking methods. The 3D camera acquires images and generates depth maps according to predetermined shooting angles and parameters. After steps such as depth map preprocessing, magazine recognition and segmentation, pose calculation, and manipulator control, the manipulator successfully grasps the magazine and places it at a designated position of the production equipment for subsequent processes.

[0153] In continuous running tests, this method can achieve the grasping of the magazine under different lighting conditions and the disordered stacking state of the magazine, and can significantly improve the production efficiency and automation level, proving the effectiveness and practicality of this method.

[0154] In a second aspect, based on the same inventive concept as the object disordered grasping method provided in this embodiment, this embodiment also provides an object disordered grasping system. As Figure 3 shown, the object disordered grasping system includes:

[0155] A depth image receiving module 201, configured to receive a depth image of an object captured by a 3D camera;

[0156] A pose parameter obtaining module 202, configured to obtain the pose parameters of the object according to the depth image;

[0157] A driving parameter generating module 203, configured to generate the driving parameters of the manipulator based on the pose parameters, in combination with the kinematic model and workspace parameters of the manipulator; wherein, the driving parameters include the grasping coordinates and the motion path;

[0158] A driving control module 204, configured to drive the manipulator to grasp the object according to the driving parameters; wherein, the method by which the pose parameter obtaining module obtains the pose parameters of the object according to the depth image is as described in steps S121 - S123 of the first aspect.

[0159] Thirdly, based on the same inventive concept as the object pose parameter acquisition method in the object disordered grasping method provided in this embodiment, that is, step S120, an embodiment of the present invention provides an object pose parameter acquisition system. As shown in FIG. 4, the object pose parameter acquisition system includes:

[0160] A preprocessing module 301, configured to preprocess a depth image including an object to obtain a preprocessed image;

[0161] An image division module 302, configured to divide the preprocessed image by using a threshold segmentation algorithm to obtain an object region that only includes the object; wherein, the threshold used in the threshold segmentation algorithm is calculated according to the gray value of each pixel in the preprocessed image; and,

[0162] A pose parameter acquisition module 303, configured to input the object region into a preset pose parameter prediction model for pose prediction, and output the pose parameters of the object.

[0163] Further, fifthly, an embodiment of the present invention further provides a readable storage medium, storing a computer program, where the computer program includes program instructions, and when the program instructions are executed by a processor, the following is implemented: the above object disordered grasping method.

[0164] The computer-readable storage medium may be an internal storage unit of the background server described in the foregoing embodiment, such as the hard disk or memory of the system. The computer-readable storage medium may also be an external storage device of the system, such as a plug-in hard disk equipped on the system, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the system. The computer-readable storage medium is used to store the computer program and other programs and data required by the system. The computer-readable storage medium may also be used to temporarily store data that has been output or will be output.

[0165] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0166] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0167] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0168] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for obtaining object pose parameters, characterized in that, Comprising: Preprocessing a depth image containing an object to obtain a preprocessed image; Using a threshold segmentation algorithm to divide the preprocessed image to obtain an object region containing only the object; wherein, the threshold used in the threshold segmentation algorithm is calculated based on the gray value of each pixel in the preprocessed image; and, Inputting the object region into a preset pose parameter prediction model for pose prediction, and outputting the pose parameters of the object.

2. The method for obtaining object pose parameters according to claim 1, characterized in that, The method for calculating the threshold based on the gray value of each pixel in the preprocessed image includes: Dividing the pixel gray values in the preprocessed image into two different clusters through a clustering algorithm, and finding the central gray values of the two clusters; Calculating the average value of the central gray values of the two clusters as the threshold; wherein, In the two clusters, one cluster represents the cartridge region and the other cluster represents the background region.

3. The method for obtaining object pose parameters according to claim 1, characterized in that, The pose parameter prediction model includes a residual network, and a channel attention module, a multi-layer perceptron, and a spatial attention module in sequential combination are inserted between the last convolutional layer and the fully connected layer of the residual network; wherein, the multi-layer perceptron includes two fully connected layers; The loss function of the pose parameter prediction model is the mean square error loss function.

4. The method for obtaining object pose parameters according to claim 1, characterized in that Preprocessing a depth image containing an object includes: Performing Gaussian filtering on the depth image; wherein, The convolution kernel radius and the standard deviation of the Gaussian distribution used in the Gaussian filtering are obtained by inputting the depth image into a preset filtering parameter acquisition model for prediction.

5. The method for obtaining object pose parameters according to claim 4, wherein, Before inputting the depth image into the preset filtering parameter acquisition model, the method further includes: Dividing the depth image into several local blocks; wherein, each local block is used as the input of the filtering parameter acquisition model.

6. The method for obtaining the pose parameters of an object according to claim 4, characterized in that, The filtering parameter acquisition model includes an input layer, a first convolutional layer, an activation function, a second convolutional layer, a max pooling layer, and an output layer; wherein, The output layer includes a fully connected layer; The first convolutional layer uses n1 convolutional kernels of size f1×f1 with a stride of s1; The activation function uses the ReLU activation function; The second convolutional layer uses n2 convolutional kernels of size f2×f2 with a stride of s2; The pooling window size of the max pooling layer is p×p with a stride of p.

7. An object pose parameter acquisition system, characterized in that Comprising: A preprocessing module for preprocessing a depth image containing an object to obtain a preprocessed image; An image division module for dividing the preprocessed image using a threshold segmentation algorithm to obtain an object region containing only the object; wherein, the threshold used in the threshold segmentation algorithm is calculated based on the gray value of each pixel in the preprocessed image; and, A pose parameter acquisition module for inputting the object region into a preset pose parameter prediction model for pose prediction, and outputting the pose parameters of the object.

8. A method for randomly grasping an object, characterized in that, Comprising: Receiving a depth image containing an object captured by a 3D camera; Obtaining the pose parameters of the object based on the depth image; Generating drive parameters of the manipulator based on the pose parameters in combination with the kinematic model and the workspace parameters of the manipulator; wherein, the drive parameters include grasping coordinates and motion paths; Drive the manipulator to grasp the object according to the driving parameters; wherein, The method for obtaining the pose parameters of the object according to the depth image is as described in any one of claims 1-6.

9. An object disordered grasping system, characterized in that, Comprising: A depth image receiving module, configured to receive a depth image of an object captured by a 3D camera; A pose parameter obtaining module, configured to obtain the pose parameters of the object according to the depth image; A driving parameter generating module, configured to generate driving parameters of the manipulator according to the pose parameters, in combination with the kinematic model and workspace parameters of the manipulator; wherein, the driving parameters include a grasping coordinate and a motion path; A driving control module, configured to drive the manipulator to grasp the object according to the driving parameters; wherein, The method for obtaining the pose parameters of the object by the pose parameter obtaining module according to the depth image is as described in any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and the program instructions, when executed by a processor, cause the processor to execute the method as described in claim 8.

Citation Information

Cited By

  • Robot disordered grabbing method and system

    CN122008259A