Grabbing control method and system of mechanical arm and electronic equipment
By detecting and segmenting the target camera image, combined with multi-level coordinate transformation and angle correction, the problem of insufficient grasping and positioning accuracy of the robotic arm in complex scenes is solved, and high-precision grasping control is achieved.
Patent Information
- Application Number
- CN202510886312.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-26
AI Technical Summary
The existing technology has poor accuracy in the positioning data captured by the robotic arm in complex scenarios, resulting in reduced grasping accuracy.
By performing target detection and segmentation on the image captured by the target camera, an initial mask area is generated. The target pixels are screened out by combining the position and color information of non-zero pixels. A multi-level coordinate transformation chain is established and the image rotation angle correction is performed to calculate high-precision grasping coordinates and angles.
It significantly improves the grasping precision and accuracy of the robotic arm in complex scenarios, reduces the probability of misgrasping and missed grasping, and improves the stability and robustness of the grasping operation.
Smart Images

Figure CN120697012A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision recognition, and in particular to a grasping control method, system, and electronic equipment for a robotic arm. Background Art
[0002] With the rapid development of mechanical automation technology, robotic arm grasping technology has been widely used in industrial manufacturing, logistics sorting, medical surgery, agricultural harvesting and other fields. However, in complex scenarios, the accuracy of grasping positioning data measured by existing technologies is poor, resulting in reduced robotic arm grasping accuracy. Summary of the Invention
[0003] The embodiments of the present application provide a grasping control method, system and electronic device for a robotic arm to solve the problem that the grasping positioning data measured in complex scenarios in the prior art has poor accuracy, resulting in reduced grasping accuracy of the robotic arm.
[0004] A first aspect of an embodiment of the present application provides a grasping control method of a robotic arm, comprising: Performing target detection and segmentation processing on the image captured by the target camera to obtain an initial mask area corresponding to the detected target; the initial mask area contains multiple non-zero pixels; Based on the first position information and color information of each of the non-zero pixels, target pixels constituting the detection target are screened out from the initial mask area to obtain a target mask area containing the target pixels; Calculate the target three-dimensional coordinates of the detection target in the robot arm base coordinate system according to the depth value of the target pixel point and the second position information of the central pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the end of the robot arm, and the second transformation matrix from the end of the robot arm to the base of the robot arm; Calculating a target grasping angle based on an image rotation angle of the detection target and a deviation angle between a grasping tool at the end of the robotic arm and the detection target; According to the target three-dimensional coordinates and the target grasping angle, the grasping tool is controlled to grasp the detection target.
[0005] A second aspect of an embodiment of the present application provides a grasping control system for a robotic arm, comprising: A detection and segmentation module is used to perform target detection and segmentation processing on the image captured by the target camera to obtain an initial mask area corresponding to the detected target; the initial mask area contains multiple non-zero pixels; a screening module, configured to screen target pixels constituting the detection target from the initial mask area based on the first position information and color information of each of the non-zero pixels, to obtain a target mask area containing the target pixels; A first calculation module is configured to calculate the target three-dimensional coordinates of the detection target in the manipulator base coordinate system according to the depth value of the target pixel point and the second position information of the center pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the manipulator end, and the second transformation matrix from the manipulator end to the manipulator base; A second calculation module is configured to calculate a target grasping angle based on an image rotation angle of the detection target and a deviation angle between a grasping tool at an end of the robotic arm and the detection target; The grabbing module is used to control the grabbing tool to grab the detection target according to the target three-dimensional coordinates and the target grabbing angle.
[0006] A third aspect of an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the computer program.
[0007] A fourth aspect of an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0008] A fifth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0009] As can be seen from the above, the present application first generates an initial mask area through target detection and segmentation processing to effectively isolate environmental noise interference. On this basis, the dual-modal screening mechanism of pixel position and color information is integrated to accurately separate the target pixel points in the initial mask area to form a target mask area. Next, a coordinate transformation chain is established based on the target mask area, that is, mask pixel → target camera → end of the robotic arm → robotic arm base, and the cumulative error of mechanical assembly is eliminated through the spatial mapping relationship. At the same time, the dual-angle closed-loop correction of the image rotation angle and the tool deviation angle is combined to calculate high-precision grasping coordinates and grasping angles to realize the grasping of the robotic arm. That is, the purpose of improving the accuracy of grasping positioning data is achieved through a multi-stage collaborative mechanism, thereby significantly improving the grasping accuracy of the robotic arm in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 This is a flow chart of a grasping control method of a robotic arm provided in an embodiment of the present application; Figure 2 This is a schematic diagram of the conversion relationship between three-dimensional coordinates in different coordinate systems provided by an embodiment of the present application; Figure 3 This is a structural diagram of a gripping control system for a robotic arm provided in an embodiment of the present application; Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0013] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0014] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0015] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0016] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0017] In specific implementations, the terminals described in the embodiments of the present application include, but are not limited to, other portable devices such as mobile phones, laptop computers, or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or touch pads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., touch screen displays and / or touch pads).
[0018] In the following discussion, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and / or joystick.
[0019] The terminal supports various applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk burning application, a spreadsheet application, a game application, a phone application, a video conferencing application, an email application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0020] Various applications that can be executed on the terminal can use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and corresponding information displayed on the terminal can be adjusted and / or changed between applications and / or within a corresponding application. In this way, the common physical architecture of the terminal (e.g., the touch-sensitive surface) can support a variety of applications with user interfaces that are intuitive and transparent to the user.
[0021] It should be understood that the size of the serial numbers of each step in this embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0022] In order to illustrate the technical solution described in this application, specific embodiments are provided below.
[0023] See also Figure 1 , Figure 1 This is a flow chart of a method for controlling the gripping of a robotic arm provided in an embodiment of the present application. Figure 1 As shown, a grasping control method of a robotic arm comprises the following steps: Step 101 : performing target detection and segmentation processing on an image captured by a target camera to obtain an initial mask area corresponding to the detected target; the initial mask area includes a plurality of non-zero pixels.
[0024] In some embodiments, the target camera is a camera mounted on the gripper at the end of the robotic arm. The target camera can be a red, green, blue (RGB) industrial camera, a red, green, blue depth (RGBD) camera, a multispectral camera, or a binocular vision camera.
[0025] The image captured by the target camera is an RGB image or an RGBD image.
[0026] The detection target is the target object to be grasped, such as fruit, electronic devices, express boxes, medicine bottles, etc.
[0027] The target camera captures images within the camera's visible area, performs target detection on the image, and filters out some of the noise in the image to achieve the detection of the detection target. By performing segmentation processing, an initial mask area corresponding to the detection target is obtained. This initial mask area is a sub-image area of the image and contains multiple non-zero pixels, which can be used to accurately locate the detection target.
[0028] In some embodiments, the target detection and segmentation processing is performed on the image captured by the target camera to obtain an initial mask area corresponding to the detected target, including: using the target camera to perform image capture to obtain a first image; performing target detection on the first image, and filtering out an initial detection frame with a first confidence level greater than a first set confidence level from the detection frame included in the first image; using the detection frame center point of the initial detection frame as a camera adjustment reference point, adjusting the posture of the target camera, and using the target camera with the adjusted posture to perform image capture to obtain a second image; performing target detection on the second image, and filtering out a target detection frame with a second confidence level greater than a second set confidence level from the detection frame included in the second image; and performing pixel-level segmentation on the target detection frame to obtain the initial mask area.
[0029] The target camera is used to capture images multiple times and perform multiple target detections, improving the reliability of the target detection results. Pixel-level segmentation is then performed on the target detection box containing the detected target. This pixel-level segmentation transforms the recognition carrier from a geometric bounding box to non-zero pixels, achieving a precise transition from coarse-grained regional representation to pixel-level geometric elements.
[0030] In some embodiments, the target camera is used to capture an image of its visible area according to its current position to obtain a first image, which is an RGB image or an RGBD image. Object detection can be performed on the first image using a single-stage object detection method, such as the You Only Look Once (YOLO) method.
[0031] During target detection, objects in the first image are marked as detection frames. Each detection frame is associated with a confidence level, namely a first confidence level, which measures the similarity between the area within the detection frame and the detected target. The first confidence levels of the detection frames contained in the first image are compared with a preset first confidence level. Detection frames with a first confidence level greater than the first confidence level are determined as initial detection frames, and detection frames with insufficient similarity are discarded to reduce the false detection rate. The first confidence level is set as needed, for example, 0.7.
[0032] The center of the initial detection frame serves as the camera adjustment reference point, i.e., the reference point for the target camera's optical center point for subsequent image acquisition. The target camera's pose is adjusted so that its optical center point is at a set height directly above the real-space point corresponding to the center of the detection frame.
[0033] In some embodiments, the set height may be the same as the initial height of the target camera during the previous acquisition, or may be smaller than the initial height of the target camera during the previous acquisition.
[0034] The pose-adjusted target camera is used to capture an image, generating a second image, which is an RGB image or an RGBD image. After pose adjustment, the resulting second image exhibits reduced visual deviation compared to the first image, resulting in higher accuracy in subsequent target detection.
[0035] After obtaining the second image, target detection is performed on the second image. During target detection, objects in the second image are marked as detection frames. Each detection frame is associated with a confidence level, i.e., a second confidence level, which measures the similarity between the area within the detection frame and the detected target. The second confidence levels of the detection frames contained in the second image are compared with a preset second confidence level. Detection frames with a second confidence level greater than the second confidence level are determined to be target detection frames, thereby obtaining target detection frames with higher confidence levels.
[0036] In some embodiments, the second set confidence level is set as needed. The second set confidence level can be set equal to the first set confidence level to avoid missed detections. The second set confidence level can also be set greater than the first set confidence level to reduce the probability of false detections.
[0037] Perform pixel-level fine segmentation on the target detection frame to obtain an initial mask area. This initial mask area is a binary mask area, that is, each pixel in the area is marked with 0 or 1. Among them, 1 is used to mark the pixel points that constitute the detection target, and 0 is used to mark the pixel points that do not belong to the detection target.
[0038] The target detection frame contains the detection target, and accordingly, the initial mask area also contains the detection target. In this way, pixel-level recognition and positioning can be achieved based on the initial mask area, and the positioning accuracy is higher.
[0039] In some embodiments, a general segmentation model or other method may be used to implement pixel-level segmentation processing of an image, such as using the Segment Anything Model (SAM) 2.1 to perform pixel-level segmentation processing.
[0040] Step 102 : Based on the first position information and color information of each of the non-zero pixels, target pixels constituting the detection target are screened out from the initial mask area to obtain a target mask area containing the target pixels.
[0041] In some embodiments, an image coordinate system is constructed based on the second image. For example, the lower left corner of the second image is used as the coordinate origin, the width edge where the coordinate origin is located is the horizontal axis, and the length edge where the coordinate origin is located is the vertical axis.
[0042] Based on the image coordinate system of the second image, the two-dimensional coordinates of the non-zero pixel points included in the initial mask area in the second image, that is, the first position information of the non-zero pixel points, are correspondingly determined.
[0043] In some embodiments, each non-zero pixel corresponds to an RGB color value. For each non-zero pixel, the RGB color value is mapped to the LAB color space defined by the Commission Internationale de l'Éclairage (CIE) through a nonlinear transformation to obtain a LAB color value, where L represents lightness and A and B represent chromaticity. The lightness component of the LAB color value is extracted as the color information of the non-zero pixel.
[0044] The LAB color space solves the illumination sensitivity problem that the RGB color space cannot overcome, and the brightness component used can improve calculation and recognition speed.
[0045] Through the first position information and color information of non-zero pixels, we can further screen out target pixels that are more likely to constitute the detection target from the initial mask area, and then obtain the target mask area containing the target pixels, that is, the mask area that is closer to the detection target, filter out irrelevant pixels, and achieve noise removal.
[0046] In some embodiments, based on the first position information and color information of each of the non-zero pixel points, the target pixel points constituting the detection target are screened out from the initial mask area to obtain the target mask area containing the target pixel points, including: constructing a multidimensional feature vector of the non-zero pixel point according to the first position information and the color information of each of the non-zero pixel points; performing iterative calculation based on the multidimensional feature vectors and Gaussian mixture model of the multiple non-zero pixel points to obtain a first probability value that the non-zero pixel point belongs to the detection target; based on the numerical size characteristics and spatial distribution characteristics of the first probability values of the multiple non-zero pixel points, the target pixel points constituting the detection target are screened out from the multiple non-zero pixel points to obtain the target mask area containing the target pixel points.
[0047] A multidimensional feature vector is constructed based on the first position and color information of non-zero pixels. It is then iterated using a Gaussian mixture model to accurately quantify the probability that each non-zero pixel belongs to the target being detected. This, combined with the numerical and spatial distribution characteristics of the first probability value, enables dual verification and screening of target pixels, effectively overcoming the positioning drift problem of traditional segmentation methods in complex scenes. This process significantly improves the accuracy of the resulting target mask area's fit to the target's outline, facilitating the determination of highly accurate capture and positioning data.
[0048] In some embodiments, the width and height of the second image and the two-dimensional coordinates of non-zero pixels in the second image are obtained, and combined with the brightness component to construct a multi-dimensional feature vector.
[0049] In some embodiments, for each non-zero pixel, the expression of the multi-dimensional feature vector is specifically: .in, is a multidimensional feature vector, ( ) is the two-dimensional coordinate of the non-zero pixel point, is the width of the second image, is the height of the second image, is the brightness component, Is a positive integer.
[0050] The multidimensional feature vector The scale-normalized position features and image brightness features are integrated to significantly enhance the robustness of subsequent calculations to image scale changes and lighting interference, while eliminating the edge positioning drift problem caused by lens distortion.
[0051] In some embodiments, the iterative calculation based on the multidimensional feature vector and Gaussian mixture model of the multiple non-zero pixel points to obtain the first probability value of the non-zero pixel point belonging to the detection target includes: inputting the multidimensional feature vector of each non-zero pixel point into the Gaussian mixture model to obtain the second probability value of the non-zero pixel point output by the Gaussian mixture model belonging to the detection target; based on the multiple second probability values, adjusting the model parameters of the Gaussian mixture model, and returning to execute the step of inputting the multidimensional feature vector of each non-zero pixel point into the Gaussian mixture model to obtain the second probability value of the non-zero pixel point output by the Gaussian mixture model belonging to the detection target until a convergence condition is reached; the convergence condition is that the change in the second probability values of the multiple non-zero pixel points is less than or equal to the set change amount, or the number of probability value calculation rounds reaches a set number of rounds; the second probability value when the convergence condition is reached is determined as the first probability value.
[0052] This iterative calculation process can be implemented in conjunction with the Expectation-Maximization Algorithm (EM algorithm). This assumes that the multidimensional feature vector is constructed using a Gaussian Mixture Model (GMM). However, the specific set of model parameters used to construct the GMM is unknown, and whether a non-zero pixel belongs to the detection target cannot be accurately determined. The EM algorithm is now used to solve for the model parameters and the first probability value.
[0053] First, set the initial model parameters for the Gaussian mixture model. The model parameters are the mixing coefficient, mean vector, and covariance matrix.
[0054] In step E, the multidimensional feature vector of each non-zero pixel point is input into the Gaussian mixture model to obtain a second probability value of each non-zero pixel point belonging to the detection target calculated and output by the Gaussian mixture model.
[0055] In step M, based on the multiple second probability values output by the calculation, the model parameters of the Gaussian mixture model are reversed, and the model parameters of the Gaussian mixture model are replaced with the latest calculated parameter values.
[0056] Return to execute step E. The Gaussian mixture models used in step E are all Gaussian mixture models with the latest model parameters at that time.
[0057] Determine whether the change in the second probability value of multiple non-zero pixel points is less than or equal to the set change, or whether the probability value calculation round has reached the set round. If any of the above convergence conditions is met, the second probability value when the convergence condition is met is determined as the first probability value. If any of the convergence conditions is not met, continue to execute the M step and the E step, and perform the convergence condition judgment until the convergence condition is met, and stop the iterative calculation. Among them, the change can be the change in the likelihood function, and the Gaussian mixture model under any set of model parameters performs a complete second probability value calculation on all non-zero pixel points, which is called a probability value calculation round.
[0058] In some embodiments, the setting change amount may be , set the round to Second, set according to actual needs.
[0059] The EM algorithm dynamically optimizes the model parameters of the Gaussian mixture model to achieve adaptive solution of pixel attribution probability. That is, in the E step, the pixel attribution probability is calculated based on the current optimal parameters, and in the M step, the model parameters are updated in reverse, the model fitting is strengthened, and the two-step iteration is carried out until convergence. Finally, a more reliable pixel-level target probability in a statistical sense is obtained, which effectively eliminates the noise in the complex background and ensures that the recognition result only retains the foreground pixels of the detection target, thereby improving the accuracy of depth calculation and pose estimation. In addition, the parameter optimization has a high level of automation, and the computational efficiency is greatly improved, which helps to improve the grasping efficiency.
[0060] By analyzing the numerical size characteristics and spatial distribution characteristics of the first probability value, highly robust extraction of target pixels is achieved.
[0061] Numerical size feature screening can lock candidate pixels with high attribution based on a preset probability threshold, effectively eliminating low-probability noise points, that is, pixels that do not belong to the detection target.
[0062] Spatial distribution characteristics screening, performing connected domain analysis on candidate pixels with high attribution. This application requires that the target pixels must form a spatially continuous cluster, which can eliminate island-type misjudgments caused by discrete noise points and probability pseudo-peaks.
[0063] In this way, the target pixel points that constitute the detection target are screened out from multiple non-zero pixel points, the fragmented, marginalized, irregular low-probability areas are eliminated, the interference of irrelevant factors is reduced, and the target mask area containing the target pixel points is obtained, which is highly consistent with the outline of the detection target.
[0064] The target mask area ensures pixel-level positioning accuracy, and its spatial distribution strictly meets the continuity requirements of actual application scenarios, providing a reliable spatial model for high-precision grasping.
[0065] In some embodiments, a size of The mask smoothing technology such as the elliptical morphological kernel is used to smooth the edges of the target mask area, eliminate boundary breaks, and ensure the geometric closure and physical rationality of the target mask area.
[0066] Step 103, calculate the target three-dimensional coordinates of the detection target in the robotic arm base coordinate system based on the depth value of the target pixel point and the second position information of the center pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the end of the robotic arm, and the second transformation matrix from the end of the robotic arm to the robotic arm base.
[0067] Based on the depth value of the target pixel point in the target mask area and the second position information of the central pixel point, the camera parameters of the target camera, the first transformation matrix from the target camera to the end of the robotic arm, and the second transformation matrix from the end of the robotic arm to the robotic arm base, calculations are performed from the image coordinate system to the camera coordinate system, then to the robotic arm end coordinate system, and finally to the robotic arm base coordinate system. Multi-level transformation eliminates error accumulation and avoids single-step error amplification to obtain more accurate three-dimensional coordinates suitable for drive control in the robotic arm base coordinate system.
[0068] The center pixel is the target pixel at the center of the target mask area.
[0069] In some embodiments, the target three-dimensional coordinates of the detection target in the robotic arm base coordinate system are calculated based on the depth value of the target pixel point in the target mask area and the second position information of the center pixel point, the camera parameters of the target camera, the first transformation matrix from the target camera to the end of the robotic arm, and the second transformation matrix from the end of the robotic arm to the base of the robotic arm, including: calculating the initial three-dimensional coordinates of the detection target in the camera coordinate system according to the depth value of the target pixel point in the target mask area and the second position information of the center pixel point and the camera parameters; constructing an initial coordinate matrix corresponding to the initial three-dimensional coordinates; performing matrix multiplication operations on the initial coordinate matrix, the first transformation matrix, and the second transformation matrix to obtain a target coordinate matrix; determining the first row of values in the target coordinate matrix as the target horizontal coordinate component, the second row of values as the target longitudinal coordinate component, and the third row of values as the target vertical coordinate component to obtain the target three-dimensional coordinates composed of the target horizontal coordinate component, the target longitudinal coordinate component, and the target vertical coordinate component.
[0070] The coordinate transformation from the image coordinate system to the camera coordinate system is realized according to the depth values of multiple target pixels in the target mask area, the coordinate value of the central pixel, and the camera parameters of the target camera.
[0071] In some embodiments, the initial three-dimensional coordinates of the detection target in the camera coordinate system are calculated based on the depth value of the target pixel point in the target mask area and the second position information of the center pixel point and the camera parameters, including: obtaining the depth values of multiple target pixel points in the target mask area; calculating the depth mean based on the depth value, and determining the depth mean as the vertical coordinate component of the initial three-dimensional coordinate; calculating the horizontal coordinate component of the initial three-dimensional coordinate based on the horizontal coordinate in the second position information, the vertical coordinate component, and the horizontal focal length and the horizontal coordinate of the optical center in the camera parameters; calculating the longitudinal coordinate component of the initial three-dimensional coordinate based on the vertical coordinate in the second position information, the vertical coordinate component, and the longitudinal focal length and the longitudinal coordinate of the optical center in the camera parameters; and obtaining the initial three-dimensional coordinate consisting of the horizontal coordinate component, the longitudinal coordinate component, and the vertical coordinate component.
[0072] In some embodiments, if the second image is an RGB image, the regional image corresponding to the target mask area is also an RGB image. In this case, a pre-trained depth estimation neural network can be used to determine the depth value of the target pixel in the target mask area. Specifically, the regional image corresponding to the target mask area is input into the depth estimation neural network, and the depth prediction value of each target pixel is obtained as the depth value of the target pixel.
[0073] In some embodiments, if the second image is an RGBD image, the regional image corresponding to the target mask area is also an RGBD image. In this case, the depth value of each target pixel can be directly extracted from the regional image.
[0074] The depth values of the plurality of target pixels are sorted in ascending order, and a depth value for calculating the depth mean is selected from the sorted result.
[0075] In some embodiments, depth values are obtained starting from the smallest depth value in ascending order until the ratio of the number of depth values obtained reaches a set ratio, at which point the number of depth values obtained stops. For example, if there are 6,000 depth values in total and the set ratio is 20%, the number of depth values to be obtained is 1,200. The average of the obtained depth values is then calculated to obtain the depth mean.
[0076] In some embodiments, the depth mean is calculated as: .in, is the depth mean, is the number of values, After sorting Depth value.
[0077] The obtained depth mean is determined as the vertical coordinate component of the initial three-dimensional coordinate of the detection target in the camera coordinate system.
[0078] At the same time, the horizontal coordinate component of the initial three-dimensional coordinates of the detected object in the camera coordinate system is calculated using the horizontal coordinate of the center pixel, the aforementioned vertical coordinate component, and the horizontal focal length and horizontal coordinate of the optical center in the camera parameters. The vertical coordinate component of the initial three-dimensional coordinates of the detected object in the camera coordinate system is calculated using the vertical coordinate of the center pixel, the aforementioned vertical coordinate component, and the longitudinal focal length and longitudinal coordinate of the optical center in the camera parameters. Ultimately, the initial three-dimensional coordinates of the detected object in the camera coordinate system, consisting of the horizontal coordinate component, the longitudinal coordinate component, and the vertical coordinate component, are obtained, and the two-dimensional coordinates in the image coordinate system are converted into three-dimensional coordinates in the camera coordinate system.
[0079] Here, the optical center transverse coordinate and the optical center longitudinal coordinate are the transverse coordinate and the longitudinal coordinate of the optical center of the target camera in the camera coordinate system when the second image is captured.
[0080] In some embodiments, the calculation formula for the initial three-dimensional coordinates is: .
[0081] in, is the horizontal coordinate of the center pixel, is the vertical coordinate of the center pixel, is the lateral coordinate of the optical center of the target camera, is the longitudinal coordinate of the optical center of the target camera, is the horizontal focal length of the target camera, is the longitudinal focal length of the target camera, is the depth mean, 、 and are the horizontal coordinate component, vertical coordinate component and vertical coordinate component of the initial three-dimensional coordinate respectively, ( , , ) is the initial three-dimensional coordinate of the detection target in the camera coordinate system.
[0082] In some embodiments, the initial three-dimensional coordinates are used to construct an initial coordinate matrix of four rows and one column, which is .use Refers to the initial coordinate matrix 、 and , get the initial coordinate matrix of four rows and one column .
[0083] Obtain the first transformation matrix from the target camera to the end of the manipulator and the second transformation matrix from the end of the manipulator to the base of the manipulator. Combine these with the initial coordinate matrix and calculate the target coordinate matrix through matrix multiplication. Both the first transformation matrix and the second transformation matrix are homogeneous transformation matrices.
[0084] In some embodiments, the target coordinate matrix is calculated as follows: .in, is the first transformation matrix, is the second transformation matrix, and are all known quantities, is the target coordinate matrix.
[0085] In some embodiments, Refers to the target coordinate matrix 、 and , accordingly, the target coordinate matrix is expressed as . The first row of values in the target coordinate matrix Determined as the target horizontal coordinate component, the second row value Determined as the target longitudinal coordinate component, the third row value Determine the target vertical coordinate component, and obtain the target three-dimensional coordinate composed of the target horizontal coordinate component, the target longitudinal coordinate component and the target vertical coordinate component ( , , ).
[0086] Figure 2 This is a schematic diagram of the conversion relationship between three-dimensional coordinates in different coordinate systems provided by an embodiment of the present application. The three-dimensional coordinates in the camera coordinate system ( , , ) Based on the first transformation matrix from the target camera to the end of the robotic arm By performing coordinate transformation, the three-dimensional coordinates in the end coordinate system of the robot arm can be obtained ( , , ), which is based on the second transformation matrix from the end of the manipulator to the base of the manipulator By performing coordinate transformation, the three-dimensional coordinates in the robot base coordinate system can be obtained ( , , ).
[0087] The first transformation matrix eliminates the target camera installation error, and the second transformation matrix compensates for the robot arm tolerance. The combination of the two transformations reduces the accuracy loss of the target three-dimensional coordinates.
[0088] The base of the robot arm is the more stable part of the robot arm. Determining it as the reference coordinate system for motion control helps to ensure grasping stability.
[0089] Step 104 : Calculate the target grasping angle based on the image rotation angle of the detection target and the deviation angle between the grasping tool at the end of the robotic arm and the detection target.
[0090] When grasping, coordinates are important, and the grasping angle is equally important. The target grasping angle required for grasping needs to be calculated based on the image rotation angle of the detection target and the deviation angle between the grasping tool at the end of the robot arm and the detection target.
[0091] In some embodiments, the rotation angle of the detection target in the regional image extracted from the regional image corresponding to the target mask area is used as the image rotation angle.
[0092] In some embodiments, the axis passing through the center of the detection target and intersecting the placement plane is defined as the principal axis. The angle between the principal axis and the placement plane is defined as the first angle. The angle between the placement plane and the line formed by the end of the robotic arm and the gripping tool located at the end of the robotic arm is defined as the second angle. The difference between the second angle and the first angle yields the deviation angle between the gripping tool at the end of the robotic arm and the detection target. The placement plane can be a real placement plane or a virtual placement plane.
[0093] The target grasping angle can be obtained by subtracting the image rotation angle and the deviation angle.
[0094] In some embodiments, the target grasping angle is calculated as follows: .in, is the image rotation angle, is the deviation angle, The target grab angle.
[0095] It should be noted that when calculating, it is necessary to ensure that the angle quantities in the same operation use the same angle unit system, that is, all are radians or all are degrees.
[0096] The target grasping angle is used to adjust the grasping direction of the grasping tool and achieve angle correction to improve grasping accuracy.
[0097] Step 105 : Control the grasping tool to grasp the detection target according to the target three-dimensional coordinates and the target grasping angle.
[0098] After obtaining the target three-dimensional coordinates and the target grasping angle, the grasping tool can be controlled to perform a grasping action based on the target three-dimensional coordinates and the target grasping angle to grasp the detection target.
[0099] In some embodiments, the grasping tool is a claw clamp type tool, a suction cup type tool, a needle puncture clamp or an electromagnet array tool, etc., and has various forms.
[0100] In some embodiments, controlling the grasping tool to grasp the detection target according to the target three-dimensional coordinates and the target grasping angle includes: generating a joint control signal of the robotic arm based on the target three-dimensional coordinates, and the joint control signal is used to drive the end of the robotic arm to move to the target three-dimensional coordinates; adjusting the grasping direction of the grasping tool according to the target grasping angle, and driving the grasping tool to grasp the detection target.
[0101] Based on the target 3D coordinates, the robot generates joint control signals for the robotic arm, driving the end of the robotic arm to move to the target 3D coordinates along an optimal path within the robot base coordinate system. Upon detecting that the end of the robotic arm has reached the target 3D coordinates, the gripping direction of the gripper at the end of the robotic arm is adjusted based on the target gripping angle.
[0102] In some embodiments, the gripping orientation is adjusted to align the opening and closing planes of the gripper tool perpendicularly with the main axis of the detection target, ensuring that the symmetry plane of the gripper tool is parallel to the normal vector of the detection target's gripping surface, resulting in evenly distributed contact force. The gripping orientation is also adjusted to align the suction cup tool's axis of attraction perpendicular to the detection target's gripping surface, ensuring that the center of negative pressure coincides with the geometric center of the gripped surface.
[0103] In some embodiments, the gripping direction of the gripping tool may be adjusted based on the target gripping angle before the end of the robotic arm is moved to the target three-dimensional coordinates. Alternatively, both adjustment operations may be performed simultaneously.
[0104] After the device is adjusted according to the target's three-dimensional coordinates and target grasping angle, the grasping operation is triggered to grasp the detection target.
[0105] In some embodiments, the claw-type tools control the clamping force through a servo motor, and the suction cup-type tools adjust the vacuum negative pressure to control the grasping force to implement grasping and ensure grasping reliability.
[0106] In some embodiments, the grasping of the robotic arm is controlled by joint space trajectory motion and Cartesian space motion, where the joint space trajectory motion is used for large-scale position adjustment and the Cartesian space motion is used for fine grasping.
[0107] In some embodiments, a target camera can be used to monitor the motion state of the robotic arm in real time, and closed-loop feedback can be used to ensure that the motion process is safe and stable.
[0108] Using the method described in this application, the robot arm was controlled to grasp objects and a grasping experiment was carried out, with a grasping success rate of over 95%. In the experiment of grasping a reagent bottle, the number of successful grasps was 96 out of 100.
[0109] The method described in this application involves two visual inspections, Gaussian mixture model modeling optimization, dual-modal information deep screening and precise homogeneous coordinate transformation, which improves the accuracy of target detection and segmentation. The identified detection targets are more accurate, and the accuracy of the obtained target three-dimensional coordinates and target grasping angles is improved, which reduces the probability of false grasping and missed grasping, improves the stability and robustness of the grasping operation, reduces the probability of grasping failure, and greatly improves the grasping accuracy and success rate.
[0110] The method described in this application is applicable to a variety of application scenarios, such as agriculture, product manufacturing, transportation, logistics, medical care, animal husbandry, etc. It provides an efficient and reliable solution for the field of industrial automatic grasping and helps promote the development of industrial automation.
[0111] In the embodiment of the present application, an initial mask area is first generated through target detection and segmentation processing to effectively isolate environmental noise interference. On this basis, a dual-modal screening mechanism is integrated with pixel position and color information to accurately separate target pixels in the initial mask area to form a target mask area. Next, a coordinate transformation chain is established based on the target mask area, namely, mask pixel → target camera → end of the robotic arm → robotic arm base. The accumulated error of mechanical assembly is eliminated through the spatial mapping relationship. At the same time, the dual-angle closed-loop correction of the image rotation angle and the tool deviation angle is combined to calculate high-precision grasping coordinates and grasping angles to realize the grasping of the robotic arm. That is, the purpose of improving the accuracy of grasping positioning data is achieved through a multi-stage collaborative mechanism, thereby significantly improving the grasping accuracy of the robotic arm in complex scenes.
[0112] See also Figure 3 , Figure 3 This is a structural diagram of a grasping control system of a robotic arm provided in an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0113] The robotic arm grasping control system 300 includes: a detection and segmentation module 301 , a screening module 302 , a first calculation module 303 , a second calculation module 304 , and a grasping module 305 .
[0114] The detection and segmentation module 301 is used to perform target detection and segmentation processing on the image captured by the target camera to obtain an initial mask area corresponding to the detected target; the initial mask area contains multiple non-zero pixels.
[0115] The screening module 302 is configured to screen out target pixels constituting the detection target from the initial mask area based on the first position information and color information of each non-zero pixel point, and obtain a target mask area containing the target pixels.
[0116] The first calculation module 303 is used to calculate the target three-dimensional coordinates of the detection target in the robotic arm base coordinate system based on the depth value of the target pixel point and the second position information of the center pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the end of the robotic arm, and the second transformation matrix from the end of the robotic arm to the robotic arm base.
[0117] The second calculation module 304 is configured to calculate a target grasping angle based on the image rotation angle of the detection target and the deviation angle between the grasping tool at the end of the robotic arm and the detection target.
[0118] The grabbing module 305 is used to control the grabbing tool to grab the detection target according to the target three-dimensional coordinates and the target grabbing angle.
[0119] In some embodiments, the detection and segmentation module is specifically configured to: Acquiring an image using the target camera to obtain a first image; Performing target detection on the first image, and selecting initial detection frames having a first confidence level greater than a first set confidence level from detection frames included in the first image; Using the center point of the initial detection frame as a camera adjustment reference point, adjusting the position of the target camera, and using the target camera after the position adjustment to perform image acquisition to obtain a second image; Performing target detection on the second image, and screening target detection frames having a second confidence level greater than a second set confidence level from detection frames included in the second image; Perform pixel-level segmentation on the target detection frame to obtain the initial mask area.
[0120] In some embodiments, the screening module is specifically configured to: constructing a multidimensional feature vector of each non-zero pixel point according to the first position information and the color information of the non-zero pixel point; Performing iterative calculation based on the multidimensional feature vectors and Gaussian mixture model of the plurality of non-zero pixels to obtain a first probability value of the non-zero pixel belonging to the detection target; Based on the numerical size characteristics and spatial distribution characteristics of the first probability values of the multiple non-zero pixel points, the target pixel points constituting the detection target are screened out from the multiple non-zero pixel points to obtain the target mask area containing the target pixel points.
[0121] In some embodiments, the screening module is further configured to: Inputting the multidimensional feature vector of each non-zero pixel into the Gaussian mixture model to obtain a second probability value output by the Gaussian mixture model that the non-zero pixel belongs to the detection target; Based on the plurality of second probability values, adjusting the model parameters of the Gaussian mixture model, and returning to the step of inputting the multidimensional feature vector of each non-zero pixel point into the Gaussian mixture model to obtain a second probability value output by the Gaussian mixture model that the non-zero pixel point belongs to the detection target, until a convergence condition is reached; the convergence condition is that the change in the second probability values of the plurality of non-zero pixels is less than or equal to a set change amount, or the number of probability value calculation rounds reaches a set number of rounds; The second probability value when the convergence condition is met is determined as the first probability value.
[0122] In some embodiments, the first computing module is specifically configured to: Calculating the initial three-dimensional coordinates of the detection target in the camera coordinate system according to the depth value of the target pixel point in the target mask area, the second position information of the central pixel point, and the camera parameters; Constructing an initial coordinate matrix corresponding to the initial three-dimensional coordinates; Performing a matrix multiplication operation on the initial coordinate matrix, the first transformation matrix, and the second transformation matrix to obtain a target coordinate matrix; The first row of values in the target coordinate matrix is determined as the target horizontal coordinate component, the second row of values is determined as the target vertical coordinate component, and the third row of values is determined as the target vertical coordinate component, to obtain the target three-dimensional coordinate composed of the target horizontal coordinate component, the target longitudinal coordinate component and the target vertical coordinate component.
[0123] In some embodiments, the first computing module is further configured to: Acquire the depth values of a plurality of target pixels in the target mask area; Calculating a depth mean based on the depth values, and determining the depth mean as a vertical coordinate component of the initial three-dimensional coordinate; Calculate the transverse coordinate component of the initial three-dimensional coordinate based on the transverse coordinate in the second position information, the vertical coordinate component, and the transverse focal length and the transverse coordinate of the optical center in the camera parameters; Calculate the longitudinal coordinate component of the initial three-dimensional coordinate based on the longitudinal coordinate in the second position information, the vertical coordinate component, and the longitudinal focal length and the longitudinal coordinate of the optical center in the camera parameters; The initial three-dimensional coordinates consisting of the horizontal coordinate component, the vertical coordinate component and the vertical coordinate component are obtained.
[0124] In some embodiments, the capture module is specifically configured to: generating a joint control signal for the robotic arm based on the target three-dimensional coordinates, wherein the joint control signal is used to drive the end of the robotic arm to move to the target three-dimensional coordinates; According to the target grasping angle, the grasping direction of the grasping tool is adjusted, and the grasping tool is driven to grasp the detection target.
[0125] The grasping control system of the robotic arm provided in the embodiment of the present application can implement each process of the embodiment of the grasping control method of the above-mentioned robotic arm and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0126] Figure 4 : is a structural diagram of an electronic device provided in an embodiment of the present application. As shown in the figure, the electronic device 4 of this embodiment includes: at least one processor 40 ( Figure 4 Only one is shown in the figure), a memory 41 and a computer program 42 stored in the memory 41 and executable on the at least one processor 40, wherein the processor 40 implements the steps of any of the above-mentioned method embodiments when executing the computer program 42.
[0127] The electronic device 4 can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device 4 can include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art will understand that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation of the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0128] The processor 40 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0129] The memory 41 may be an internal storage unit of the electronic device 4, such as a hard drive or memory of the electronic device 4. The memory 41 may also be an external storage device of the electronic device 4, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 41 may include both an internal storage unit of the electronic device 4 and an external storage device. The memory 41 is used to store the computer program and other programs and data required by the electronic device. The memory 41 may also be used to temporarily store data that has been output or is about to be output.
[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0131] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0132] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0133] In the embodiments provided in this application, it should be understood that the disclosed systems / electronic devices and methods can be implemented in other ways. For example, the system / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the system or unit, which can be electrical, mechanical or other forms.
[0134] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0135] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0136] If the integrated module / unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0137] The present application implements all or part of the processes in the above-mentioned embodiment methods, and may also be implemented through a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0138] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A grasping control method for a robotic arm, characterized in that: include: Perform target detection and segmentation on the image captured by the target camera to obtain the initial mask area corresponding to the detected target; The initial mask area contains a plurality of non-zero pixels; Based on the first position information and color information of each of the non-zero pixels, target pixels constituting the detection target are screened out from the initial mask area to obtain a target mask area containing the target pixels; Calculate the target three-dimensional coordinates of the detection target in the robot arm base coordinate system according to the depth value of the target pixel point and the second position information of the central pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the end of the robot arm, and the second transformation matrix from the end of the robot arm to the base of the robot arm; Calculating a target grasping angle based on an image rotation angle of the detection target and a deviation angle between a grasping tool at the end of the robotic arm and the detection target; According to the target three-dimensional coordinates and the target grasping angle, the grasping tool is controlled to grasp the detection target.
2. The method according to claim 1, characterized in that The target detection and segmentation processing is performed on the image captured by the target camera to obtain an initial mask area corresponding to the detected target, including: Acquiring an image using the target camera to obtain a first image; Performing target detection on the first image, and selecting initial detection frames having a first confidence level greater than a first set confidence level from detection frames included in the first image; Using the center point of the initial detection frame as a camera adjustment reference point, adjusting the position of the target camera, and using the target camera after the position adjustment to perform image acquisition to obtain a second image; Performing target detection on the second image, and screening target detection frames having a second confidence level greater than a second set confidence level from detection frames included in the second image; Perform pixel-level segmentation on the target detection frame to obtain the initial mask area.
3. The method according to claim 1, characterized in that The step of screening target pixels constituting the detection target from the initial mask area based on the first position information and color information of each of the non-zero pixels to obtain a target mask area containing the target pixels comprises: constructing a multidimensional feature vector of each non-zero pixel point according to the first position information and the color information of the non-zero pixel point; Performing iterative calculation based on the multidimensional feature vectors and Gaussian mixture model of the plurality of non-zero pixels to obtain a first probability value of the non-zero pixel belonging to the detection target; Based on the numerical size characteristics and spatial distribution characteristics of the first probability values of the multiple non-zero pixel points, the target pixel points constituting the detection target are screened out from the multiple non-zero pixel points to obtain the target mask area containing the target pixel points.
4. The method according to claim 3, characterized in that The iterative calculation based on the multidimensional feature vector and the Gaussian mixture model of the plurality of non-zero pixels to obtain a first probability value of the non-zero pixel belonging to the detection target includes: Inputting the multidimensional feature vector of each non-zero pixel into the Gaussian mixture model to obtain a second probability value output by the Gaussian mixture model that the non-zero pixel belongs to the detection target; Based on the plurality of second probability values, adjusting the model parameters of the Gaussian mixture model, and returning to the step of inputting the multidimensional feature vector of each non-zero pixel point into the Gaussian mixture model to obtain a second probability value output by the Gaussian mixture model that the non-zero pixel point belongs to the detection target, until a convergence condition is reached; the convergence condition is that the change in the second probability values of the plurality of non-zero pixels is less than or equal to a set change amount, or the number of probability value calculation rounds reaches a set number of rounds; The second probability value when the convergence condition is met is determined as the first probability value.
5. The method according to claim 1, wherein The method further comprises calculating the target three-dimensional coordinates of the detection target in the manipulator base coordinate system according to the depth value of the target pixel point and the second position information of the center pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the manipulator end, and the second transformation matrix from the manipulator end to the manipulator base, including: Calculating the initial three-dimensional coordinates of the detection target in the camera coordinate system according to the depth value of the target pixel point in the target mask area, the second position information of the central pixel point, and the camera parameters; Constructing an initial coordinate matrix corresponding to the initial three-dimensional coordinates; Performing a matrix multiplication operation on the initial coordinate matrix, the first transformation matrix, and the second transformation matrix to obtain a target coordinate matrix; The first row of values in the target coordinate matrix is determined as the target horizontal coordinate component, the second row of values is determined as the target vertical coordinate component, and the third row of values is determined as the target vertical coordinate component, to obtain the target three-dimensional coordinate composed of the target horizontal coordinate component, the target longitudinal coordinate component and the target vertical coordinate component.
6. The method according to claim 5, characterized in that The calculating, based on the depth value of the target pixel point in the target mask area, the second position information of the center pixel point, and the camera parameters, the initial three-dimensional coordinates of the detection target in the camera coordinate system includes: Acquire the depth values of a plurality of target pixels in the target mask area; Calculating a depth mean based on the depth values, and determining the depth mean as a vertical coordinate component of the initial three-dimensional coordinate; Calculate the transverse coordinate component of the initial three-dimensional coordinate based on the transverse coordinate in the second position information, the vertical coordinate component, and the transverse focal length and the transverse coordinate of the optical center in the camera parameters; Calculate the longitudinal coordinate component of the initial three-dimensional coordinate based on the longitudinal coordinate in the second position information, the vertical coordinate component, and the longitudinal focal length and the longitudinal coordinate of the optical center in the camera parameters; The initial three-dimensional coordinates consisting of the horizontal coordinate component, the vertical coordinate component and the vertical coordinate component are obtained.
7. The method according to claim 1, characterized in that The step of controlling the grasping tool to grasp the detection target according to the target three-dimensional coordinates and the target grasping angle includes: generating a joint control signal for the robotic arm based on the target three-dimensional coordinates, wherein the joint control signal is used to drive the end of the robotic arm to move to the target three-dimensional coordinates; According to the target grasping angle, the grasping direction of the grasping tool is adjusted, and the grasping tool is driven to grasp the detection target.
8. A gripping control system for a robotic arm, characterized in that: include: The detection and segmentation module is used to perform target detection and segmentation processing on the image captured by the target camera to obtain the initial mask area corresponding to the detection target; The initial mask area contains a plurality of non-zero pixels; a screening module, configured to screen target pixels constituting the detection target from the initial mask area based on the first position information and color information of each of the non-zero pixels, to obtain a target mask area containing the target pixels; A first calculation module is configured to calculate the target three-dimensional coordinates of the detection target in the manipulator base coordinate system according to the depth value of the target pixel point and the second position information of the center pixel point in the target mask area, the camera parameters of the target camera, the first transformation matrix from the target camera to the manipulator end, and the second transformation matrix from the manipulator end to the manipulator base; A second calculation module is configured to calculate a target grasping angle based on an image rotation angle of the detection target and a deviation angle between a grasping tool at an end of the robotic arm and the detection target; The grasping module is used to control the grasping tool to grasp the detection target according to the target three-dimensional coordinates and the target grasping angle.
9. An electronic device, characterized in that: The electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The invention comprises a computer program which, when executed, causes the method according to any one of claims 1 to 7 to be performed.
Citation Information
Cited By
Grabbing method and equipment of side wall outer plate, storage medium and program product
CN121083653A
Method, apparatus, storage medium and program product for gripping a side wall outer panel
CN121083653B