Grabbing guiding method and system based on binocular images under mobile robot platform

By using deep learning algorithms based on binocular images and optical flow tracking technology, the accuracy and adaptability issues of visual grasping guidance under mobile robot platforms were solved, achieving an efficient and stable grasping process.

CN117237414BActive Publication Date: 2026-03-24SHANDONG INST OF ADVANCED TECH CHINESE ACAD OF SCI CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies for visual grasping guidance on mobile robot platforms suffer from problems such as insufficient accuracy, strong environmental dependence, discontinuous grasping point detection, and difficulty in online calibration.

Method used

A binocular image-based grasping guidance method is adopted, which uses deep learning algorithms for feature extraction and stereo matching, and combines optical flow tracking for real-time calibration to achieve 3D reconstruction and grasping point detection.

Benefits of technology

It improves grasping accuracy and adaptability, enhances the robustness and real-time calibration capability of the grasping system, and adapts to scene changes and sudden interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237414B_ABST
    Figure CN117237414B_ABST
Patent Text Reader

Abstract

The application provides a kind of mobile robot platform based on binocular image grasping guidance method and system, respectively carry out feature extraction to the left eye image and right eye image of mobile robot at current time, obtain corresponding left eye feature map and right eye feature map;Stereo matching is carried out to left eye feature map and right eye feature map, and disparity map is obtained;Left eye feature map and right eye feature map carry out grasping point detection, and obtain grasping position;According to disparity map, grasping position and binocular camera external parameter matrix, three-dimensional reconstruction of grasping target is carried out, the relative pose of mobile robot and target object and the distance of current grasping point from mobile robot are obtained in combination with three-dimensional reconstruction result;In the process that mobile robot approaches grasping object, key feature point tracking is carried out using optical flow, according to key feature point tracking result, the relative pose of mobile robot and target object and the distance of current grasping point from mobile robot are corrected, until grasping is completed;The application realizes more accurate visual grasping guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot grasping guidance and control technology, and in particular to a grasping guidance method and system based on binocular images for a mobile robot platform. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Visual grasping guidance technology is a cross-disciplinary solution integrating visual perception, trajectory planning, and motion control. It is primarily used in intelligent manufacturing and smart logistics scenarios such as robotic arm grasping and mobile robot grasping. Mobile robots with intelligent grasping capabilities are the main operational platform for realizing intelligent manufacturing or smart logistics. Based on preset requirements, the robot begins its work after arriving at the initial working position, grasping the target object, and performing subsequent actions according to the scenario's needs after grasping. During the target object grasping process, the robot's vision system's accurate positioning and distance calculation of the target object are prerequisites for successful grasping. Continuously providing accurate visual perception signals and grasping reference information to the motion control system online is essential to ensure the successful completion of the grasping task.

[0004] Patent No. CN201810063064.4 discloses a visual recognition and localization method for intelligent grasping applications of robots. It uses RGB-D images as input, which are fed into a convolutional neural network and calculated to recover the overall point cloud structure of the target object. Then, the relative posture and distance information of the target object are obtained from the overall object point cloud map. This method requires three-dimensional modeling of the target object as a whole, and the depth camera used is relatively more expensive and more likely to be interfered with.

[0005] Patent No. CN202010932018.0 discloses a real-time tracking method for a robotic arm based on binocular vision guidance. This method employs binocular stereo vision calibration, correction, and matching, combined with hand-eye calibration for camera-robot coordinate transformation. Visual tracking control is then performed during the subsequent grasping process. However, this method relies solely on a fixed calibration matrix for stereo matching, making it highly dependent on the environment. On mobile robot platforms, if the robot has undergone prolonged movement or experiences scene changes or variations in working distance, this method cannot achieve high-precision matching and reconstruction results. Furthermore, this patent does not include the grasping point detection process, requiring manual calibration or placement of the target object in a fixed pose during grasping, which increases workload or reduces adaptability. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a grasping guidance method and system based on binocular images for a mobile robot platform. The mobile robot uses a binocular camera to acquire RGB images of the target object at its initial working position, and uses deep learning algorithms to infer information such as the relative distance, relative pose, and optimal grasping position between the mobile robot and the target object. Combined with optical flow tracking, online grasping information calibration is performed during the grasping process, thereby achieving more accurate visual grasping guidance.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a grasping and guidance method based on binocular images for a mobile robot platform.

[0009] A grasping and guidance method based on binocular images for a mobile robot platform includes the following steps:

[0010] Feature extraction is performed on the left and right eye images of the mobile robot at the current moment to obtain the corresponding left eye feature map and right eye feature map;

[0011] Stereo matching is performed on the left and right eye feature maps to obtain a disparity map; grasping point detection is performed on the left and right eye feature maps to obtain the grasping position;

[0012] Based on the disparity map, the grasping position, and the extrinsic matrix of the binocular camera, a 3D reconstruction of the grasping target is performed. The relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot, are obtained by combining the 3D reconstruction results.

[0013] During the process of the mobile robot approaching the object to be grasped, optical flow is used to track key feature points. Based on the tracking results, the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot, are corrected until the grasping is completed.

[0014] As a further limitation of the first aspect of the present invention, feature extraction is performed on the left and right eye images of the mobile robot at the current moment, including:

[0015] A deep neural network based on dilated convolution is used to extract semantic features from the left and right eye images; the deep neural networks used for feature extraction in the left and right eye images have the same structure and parameters.

[0016] As a further limitation of the first aspect of the present invention, the left and right eye feature maps are used to detect grasping points to obtain the grasping position, including:

[0017] The graspability of each position is scored on the left and right eye feature maps, and the position with the highest score is regarded as the grasping position.

[0018] As a further limitation of the first aspect of the present invention, stereo matching is performed on the left-eye feature map and the right-eye feature map to obtain a disparity map, including:

[0019] The disparity map is obtained by performing feature matching based on a 3D neural network on the left and right eye feature maps.

[0020] As a further limitation of the first aspect of the invention, the left-eye feature map and the right-eye feature map are connected at each disparity level to form a cost capacity matrix, wherein the matrix dimension of the cost capacity is 4, namely height, width, disparity and feature size.

[0021] The cost capacity matrix is ​​processed by 3D convolution with a basic residual structure. The heat map after 3D convolution is upsampled to restore the original image size by bilinear interpolation. Regression calculation is then applied to obtain a disparity map of a set size.

[0022] As a further definition of the first aspect of the invention, the key feature points include: gripping points and ORB feature points.

[0023] Secondly, the present invention provides a grasping and guidance system based on binocular images for a mobile robot platform.

[0024] A grasping and guidance system based on binocular images for a mobile robot platform, comprising:

[0025] The feature extraction module is configured to extract features from the left and right eye images of the mobile robot at the current moment, respectively, to obtain the corresponding left and right eye feature maps.

[0026] The stereo matching module is configured to perform stereo matching on the left and right eye feature maps to obtain a disparity map.

[0027] The grasping point detection module is configured to detect grasping points using the left and right eye feature maps to obtain the grasping position.

[0028] The grasping information calculation module is configured to: perform 3D reconstruction of the grasping target based on the disparity map, grasping position, and binocular camera extrinsic matrix; combine the 3D reconstruction results to obtain the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot; during the process of the mobile robot approaching the grasping object, optical flow is used to track key feature points; based on the key feature point tracking results, the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot, are corrected until the grasping is completed.

[0029] As a further limitation of the second aspect of the present invention, in the stereo matching module, feature matching based on a three-dimensional neural network is performed on the left-eye feature map and the right-eye feature map to obtain a disparity map, including:

[0030] At each disparity level, the left and right feature maps are concatenated to form a cost capacity matrix, which has a dimension of 4, namely height, width, disparity and feature size.

[0031] The cost capacity matrix is ​​processed by 3D convolution with a basic residual structure. The heat map after 3D convolution is upsampled to restore the original image size by bilinear interpolation. Regression calculation is then applied to obtain a disparity map of a set size.

[0032] Thirdly, the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the binocular image-based grasping and guiding method for a mobile robot platform as described in the first aspect of the present invention.

[0033] Fourthly, the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the binocular image-based grasping and guidance method for a mobile robot platform as described in the first aspect of the present invention.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] 1. Existing grasping guidance algorithms are generally designed for grasping with robotic arms on fixed platforms. The data is then transmitted to the control terminal. The robotic arm base is fixed in position, and only its grasping part is activated. Such grasping systems cannot be adapted to mobile robot platforms. This invention uses a binocular camera to complete visual detection and 3D reconstruction, outputting the grasping point position, relative posture, and relative distance information of the object. This improves the accuracy and efficiency of grasping work and enhances the adaptability of the object grasping system.

[0036] 2. Existing capture point detection algorithms can only output single capture point information, that is, determine the capture score of each pixel and auxiliary information such as width and angle. After outputting the capture point on the image, there is no further processing. Moreover, the newly acquired image sequence is processed independently frame by frame during the capture process, which fails to effectively utilize the spatial continuity of the video. To address this issue, this invention uses optical flow points to track the capture points, ensuring the continuity of the video processing process and exhibiting good robustness to sudden detection errors or abnormal states caused by occlusion or target object movement.

[0037] 3. Existing deep learning-based binocular depth estimation is generally used in autonomous driving or 3D mapping scenarios. In grasping scenarios, it is not fully utilized because the scene scale is small and grasping does not require global disparity information. This invention adopts a method that uses only the information at the grasping point position to participate in 3D reconstruction after outputting the global disparity map, which efficiently filters the effective information in the disparity map. In addition, structured light data is used for supervision and RGB data is used for inference during the training process, making the training process simpler and more efficient, and the inference process more accurate.

[0038] 4. Current grasping structures mostly adopt linear grasping behavior operations, that is, given initial grasping information, the robot grasping behavior is planned, and then the planned grasping action is executed. Due to the lack of intermediate state information, online grasping behavior calibration cannot be achieved. In this invention, by tracking, matching and updating the reconstruction results processing chain, the continuous output of state information during the grasping process is realized, and real-time calibration of grasping behavior can be achieved by receiving the current grasping information.

[0039] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0040] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0041] Figure 1 This is a flowchart illustrating the grasping and guidance method based on binocular images on a mobile robot platform provided in Embodiment 1 of the present invention.

[0042] Figure 2 This is a schematic diagram of the process for calculating the crawled information provided in Embodiment 1 of the present invention;

[0043] Figure 3 This is a schematic diagram of a binocular image-based grasping and guidance system for a mobile robot platform provided in Embodiment 2 of the present invention. Detailed Implementation

[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0045] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0046] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0047] Example 1:

[0048] like Figure 1 As shown, Embodiment 1 of the present invention provides a grasping guidance method based on binocular images for a mobile robot platform. The mobile robot starts the grasping guidance system at the initial working position. It uses a pre-calibrated binocular camera and a trained target detection algorithm to define the visual region of the current target, acquires the RGB image of the target object, and obtains the optimal grasping position through image processing. At the same time, it outputs the distance of the grasping position relative to the mobile robot and the grasping posture information. After completing this process, the mobile robot gradually approaches the target object and continues to acquire grasping posture information. The control system uses information such as the grasping point position and the relative distance of the grasping point to continuously adjust the mobile robot's own posture as it gradually approaches the target object until the grasping process is completed.

[0049] In this embodiment, the preferred method is to obtain the RGB image information of the target area captured by the binocular camera. The binocular camera has been visually calibrated and the extrinsic parameter matrix has been obtained before use. It is understood that in some other implementations, a light camera or a TOF camera can be used instead of a binocular RGB camera, as long as a binocular camera-like structure can be formed and the final generated image is an RGB image (which can be converted). This will not be elaborated here.

[0050] Specifically, such as Figure 1 As shown, the process includes the following:

[0051] S1: Extract semantic features of the image using a deep neural network based on dilated convolution, and output feature maps from the left and right cameras respectively.

[0052] S2: Perform feature matching based on a 3D neural network on the left and right eye feature maps to obtain the disparity maps of the left and right eye cameras;

[0053] S3: Score the graspability of each position on the left and right eye feature maps, and the position with the highest score is regarded as the grasping position;

[0054] S4: Use the disparity map in the video sequence, the extrinsic matrix of the original RGB image, and the position of the gripping point to perform 3D pose reconstruction, and obtain the final relative pose of the mobile robot and the target object, as well as the distance between the current gripping point and the mobile robot.

[0055] This invention completes binocular matching reconstruction and active selection of grasping points through deep learning image features, and combines the two to achieve real-time online grasping information feedback, which has strong adaptability and stability for target grasping operations.

[0056] In S1, more specifically, it includes:

[0057] When the mobile robot performs real-time perception, it acquires data streams of the surrounding environment through binocular RGB images, and uses a trained target detection algorithm to perform image region filtering on the target in the data stream.

[0058] In the image region filtering process, the detection algorithm will extract the two-dimensional images of the target object from the left and right eyes of the RGB image of the stereo camera. After the image is extracted, it will be processed by the deep learning algorithm module.

[0059] The image feature encoder in the feature extraction process consists of a spatial pyramid pooling module and dilated convolution. The encoder uses open-source web crawling datasets (such as YCB, LINEMOD, etc.) for model pre-training, and also uses some proprietary data in the pre-training process. Since the original dataset is rich enough, the features learned by the pre-trained model can be effectively used as a general model for crawling tasks. The image feature encoder performs the same feature extraction process on both left and right eye images, that is, the algorithm model used in the two-eye image processing is the same model. After the feature maps are extracted, they are respectively sent to the stereo matching module and the crawling point detection module for feature map processing.

[0060] In S2, more specifically, it includes:

[0061] After obtaining the feature maps of the left eye and the right eye, they are first scaled to 300×300. Then, the same deep neural network is used to perform point detection processing on the two feature maps respectively. The deep neural network here adopts a structure similar to GG-CNN, which consists of three convolutional layers and three transposed convolutional layers.

[0062] In the inference phase of grasp point detection, after the feature map is fed into the deep neural network, it outputs the grasp score corresponding to each pixel. Additionally, the grasp angle and grasp width can be output according to actual needs. In its training phase, it uses self-collected manually labeled data and open-source network data as data sources. The PyTorch framework is used to build the algorithm model and perform supervised training. The final output of grasp point detection is:

[0063]

[0064] Where q represents the capture score corresponding to the capture point, s is the image pixel coordinate, and the latter two are the corresponding capture angle and width. If the point with the largest capture score has a score greater than the capture threshold, and the capture angle and width meet the system range, then the point is finally selected as the capture point. The image is then transformed from the 300×300 feature map to the original image as the final capture point output. After the left and right feature maps obtain this output, the capture information is calculated in combination with the stereo matching results.

[0065] Where q represents the current pixel capture quality, which can only be represented by a single floating-point number, such as 0.9 or 0.8; Represents pixel coordinates, with two quantities: X and Y axes, for example (123, 456); the remaining two are angular physical quantities. same Similarity can contain multiple numbers or be expressed by a single floating-point number, depending on the crawling system.

[0066] More specifically, regarding the score capture, this includes:

[0067] The capture detection algorithm is an algorithm composed of a deep neural network. Its output is a matrix of the same size as the original image. Each matrix element corresponds to a pixel in the original image. Each matrix element has five values, corresponding to the five parameters in formula (1).

[0068] When training the neural network, grasping points are selected on the image by manual annotation. The selected grasping points are assigned a grasping score of 100, while the grasping scores of other points are set to 0. At the same time, the grasping width and grasping angle of the grasping points also need to be manually set, while the two values ​​of non-grasping points are automatically assigned to -1.

[0069] Understandably, in some other implementations, grab point detection can be performed using Harris features or other visual features instead of deep learning methods, which will not be elaborated here.

[0070] Specifically, in S3, this includes:

[0071] After obtaining the left and right feature maps, the disparity map of the target object in the current state is output through a deep learning algorithm. The network structure of the deep learning algorithm adopts the PSMnet network structure. After simple convolution, the left and right feature maps are connected to their corresponding right feature maps at each disparity level to form a cost volume matrix with a dimension of 4, which is height, width, disparity and feature size respectively.

[0072] After obtaining the cost capacity matrix, it is processed using a three-dimensional convolution with a basic residual structure. This three-dimensional convolution is simply constructed from residual blocks and contains 12 3×3×3 convolutional layers. Then, the heat map after the three-dimensional convolution is upsampled to restore the original image size through bilinear interpolation. Finally, regression is applied to calculate a disparity map of size H×W. The above process corresponds to the inference process of stereo matching.

[0073] During training, the stereo matching module was pre-trained using open-source datasets such as ScenceFlow and KITTI. Then, the network was fine-tuned using self-collected data. This self-collected data came from non-pure RGB sensors such as structured light and Time-of-Flight (TOF). Image data served as the network input, and disparity data was used as supervision during training. The same loss function as PSMnet was used for training.

[0074]

[0075] in:

[0076]

[0077] Where N represents the total number of pixels in the image, and i represents the pixel index. Let d represent the disparity calculated by the neural network algorithm model during training, and d represent the current true disparity; when i is used as Alternatively, when the subscript 'd' appears, it represents the calculated disparity and the actual disparity of the pixel position corresponding to 'i'.

[0078] Specifically, in S4, this includes:

[0079] After obtaining the disparity map and the position of the binocular capture point, the 3D reconstruction of the visual target can be completed by combining the binocular camera extrinsic matrix, and the triangulated depth information can be obtained.

[0080] More specifically, in the previous operation, the disparity map and the grab point map were scaled to the same size. Then, the disparity map and the grab point map can be image overlay aligned. The pixel position of the grab point can correspond to the same pixel position in the disparity map. Therefore, the disparity information of this pixel in the disparity map is the disparity information d of the grab point.

[0081] In addition to parallax information, the triangulation process also requires the optical center positions of the left and right eyes, the optical center distance T, and the camera focal length f (since the left and right eyes use the same camera, f is the same).

[0082] Taking left-eye camera triangulation as an example, let's assume the pixel position of the left-eye camera's capture point is x. l y l The optical center position of the left eye camera is c. x cy Then the triangulated expression is:

[0083]

[0084]

[0085]

[0086] Where Xw, Yw, and Zw represent the three-dimensional world coordinates with the center of the grabbing system as the origin and the position of the grabbing point, and X, Y, and Z represent the quantities of its three Cartesian axes, respectively.

[0087] In this embodiment, after the grasping point is detected, the best grasping position and the corresponding score are output in the binocular images respectively. When the score exceeds the threshold and the left and right images are matched correctly, the pixel coordinates of the grasping point are determined, and the depth information and reference coordinates relative to the mobile robot are obtained by combining the results of the three-dimensional reconstruction.

[0088] After completing the initialization of depth and coordinate values, the mobile robot begins the grasping operation and uses optical flow to track key feature points in the subsequently acquired binocular video sequence (the key feature points include two parts: grasping points and ORB feature points. The number of graspable points extracted is generally small and sparse in the image, so we use ORB feature points to supplement the graspable points and improve tracking stability). At the same time, the distance information and relative coordinates of the grasping points are continuously obtained for grasping calibration until the grasping is completed.

[0089] The key feature points described in this invention consist of two parts: one part is the graspable points calculated above, and the other part is ORB feature points. Using these points as the base point set, the updated feature point index and pixel coordinates are obtained in real time through inter-frame image matching, and the pixel coordinates are used to update the three-dimensional world coordinates Xw, Yw and Zw in real time to complete the online calibration of the grasp points.

[0090] like Figure 3 The diagram shown illustrates the specific process of information retrieval and calculation. The nonlinear optimization employs the bundle adjustment method, and its specific formula is as follows:

[0091]

[0092] Where i is the symbol of the key feature point, n is the total number of key feature points, K is the camera intrinsic parameter, T is the matrix representing the pose, ui is the pixel coordinate of the i-th key feature point, and Si is the distance of the i-th key feature point relative to the origin, obtained from Xw, Yw, and Zw (where Xw, Yw, and Zw are obtained through disparity), the formula is as follows: The subscript i represents the world coordinates of the key feature point, and the subscript 0 represents the world coordinates of the system origin.

[0093] Figure 3 In the process, after updating the depth and pose, if there is no new information input, the currently calculated pose is the final pose; if there is new information input, the latest pose is returned to participate in the calculation; as for whether new information is entered, it depends on whether the grasping robot system closes the pose estimation process, and closing the control part is not the subject of this invention, but depends on the specific use case.

[0094] Example 2:

[0095] like Figure 3 As shown, Embodiment 2 of the present invention provides a grasping and guidance system based on binocular images for a mobile robot platform, comprising:

[0096] The feature extraction module is configured to extract features from the left and right eye images of the mobile robot at the current moment, respectively, to obtain the corresponding left and right eye feature maps.

[0097] The stereo matching module is configured to perform stereo matching on the left and right eye feature maps to obtain a disparity map.

[0098] The grasping point detection module is configured to detect grasping points using the left and right eye feature maps to obtain the grasping position.

[0099] The grasping information calculation module is configured to: perform 3D reconstruction of the grasping target based on the disparity map, grasping position, and binocular camera extrinsic matrix; combine the 3D reconstruction results to obtain the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot; during the process of the mobile robot approaching the grasping object, optical flow is used to track key feature points; based on the key feature point tracking results, the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot, are corrected until the grasping is completed.

[0100] The working method of the system is the same as the grasping and guidance method based on binocular images under the mobile robot platform provided in Embodiment 1, and will not be described again here.

[0101] Example 3:

[0102] Embodiment 3 of the present invention provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, it implements the steps in the grasping and guiding method based on binocular images under the mobile robot platform as described in Embodiment 1 of the present invention.

[0103] Example 4:

[0104] Embodiment 4 of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the grasping and guiding method based on binocular images under the mobile robot platform as described in Embodiment 1 of the present invention.

[0105] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A grasping and guiding method based on binocular images for a mobile robot platform, characterized in that, The process includes the following: Feature extraction is performed on the left and right eye images of the mobile robot at the current moment to obtain the corresponding left eye feature map and right eye feature map; Stereo matching is performed on the left and right eye feature maps to obtain a disparity map; Grasping points are detected using the left and right eye feature maps to determine the grasping location. Based on the disparity map, the grasping position, and the extrinsic matrix of the binocular camera, a 3D reconstruction of the grasping target is performed. The relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot, are obtained by combining the 3D reconstruction results. The disparity map and the grab point map were scaled to the same size, and then the disparity map and the grab point map were image overlay aligned, with the pixel position of the grab point corresponding to the same pixel position in the disparity map. During the process of the mobile robot approaching the object to be grasped, optical flow is used to track key feature points. Based on the tracking results of key feature points, the relative pose of the mobile robot and the target object and the distance between the current grasping point and the mobile robot are corrected until the grasping is completed. The key feature points consist of two parts: grasp points and ORB feature points. Using the grasp points and ORB feature points as the basic point set, the updated feature point index and pixel coordinates are obtained in real time through inter-frame image matching. The pixel coordinates are then used to update the three-dimensional world coordinates Xw, Yw, and Zw in real time to complete the online calibration of the grasp points. The specific formula is as follows: in, i The code name for the key feature point. n This represents the total number of key feature points. K For camera internal parameters, T The matrix representing the pose, For the first i Pixel coordinates of key feature points; For the first i The distance of each key feature point relative to the origin is calculated using the formula: .

2. The grasping and guidance method based on binocular images for a mobile robot platform as described in claim 1, characterized in that, Feature extraction is performed on the left and right eye images of the mobile robot at the current moment, including: A deep neural network based on dilated convolution is used to extract semantic features from the left and right eye images; the deep neural networks used for feature extraction in the left and right eye images have the same structure and parameters.

3. The grasping and guidance method based on binocular images on a mobile robot platform as described in claim 1, characterized in that, Grasping point detection is performed on the left and right eye feature maps to obtain the grasping location, including: The graspability of each position is scored on the left and right eye feature maps, and the position with the highest score is regarded as the grasping position.

4. The grasping and guidance method based on binocular images on a mobile robot platform as described in claim 1, characterized in that, Stereo matching is performed on the left and right eye feature maps to obtain a disparity map, including: The disparity map is obtained by performing feature matching based on a 3D neural network on the left and right eye feature maps.

5. The grasping and guidance method based on binocular images on a mobile robot platform as described in claim 4, characterized in that, At each disparity level, the left and right feature maps are concatenated to form a cost capacity matrix, which has a dimension of 4, namely height, width, disparity and feature size. The cost capacity matrix is ​​processed by 3D convolution with a basic residual structure. The heat map after 3D convolution is upsampled to restore the original image size by bilinear interpolation. Regression calculation is then applied to obtain a disparity map of a set size.

6. A grasping and guidance system based on binocular images for a mobile robot platform, employing the grasping and guidance method based on binocular images for a mobile robot platform as described in any one of claims 1-5, characterized in that, include: The feature extraction module is configured to extract features from the left and right eye images of the mobile robot at the current moment, respectively, to obtain the corresponding left and right eye feature maps. The stereo matching module is configured to perform stereo matching on the left and right eye feature maps to obtain a disparity map. The grasping point detection module is configured to detect grasping points using the left and right eye feature maps to obtain the grasping position. The grasping information calculation module is configured to: perform 3D reconstruction of the grasping target based on the disparity map, grasping position, and binocular camera extrinsic matrix; combine the 3D reconstruction results to obtain the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot; during the process of the mobile robot approaching the grasping object, optical flow is used to track key feature points; based on the key feature point tracking results, the relative pose of the mobile robot and the target object, as well as the distance between the current grasping point and the mobile robot, are corrected until the grasping is completed.

7. The grasping and guiding system based on binocular images on a mobile robot platform as described in claim 6, characterized in that, The stereo matching module performs feature matching between the left and right eye feature maps based on a 3D neural network to obtain a disparity map, including: At each disparity level, the left and right feature maps are concatenated to form a cost capacity matrix, which has a dimension of 4, namely height, width, disparity and feature size. The cost capacity matrix is ​​processed by 3D convolution with a basic residual structure. The heat map after 3D convolution is upsampled to restore the original image size by bilinear interpolation. Regression calculation is then applied to obtain a disparity map of a set size.

8. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the binocular image-based grasping and guidance method for a mobile robot platform as described in any one of claims 1-5.

9. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the binocular image-based grasping and guidance method for a mobile robot platform as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Visual recognition and positioning method for robot intelligent capture application

    CN108171748A

  • Mechanical arm real-time tracking method based on binocular vision guidance

    CN112132894A

  • Binocular vision positioning method for target grabbing of underwater robot

    CN111062990A

  • Vision-based mobile robot positioning method

    CN112308917A

  • Specified object grabbing method based on target cutting area

    CN113888631A