Manipulator self-adaptive grabbing and releasing system based on multi-mode visual positioning

By combining the multimodal visual positioning module and the OpenCV fitting algorithm, the robot arm can adaptively grasp and place different goods, solving the problems of low grasping accuracy and insufficient adaptability in the existing technology, and improving the working efficiency and safety of the loading vehicle.

CN120791735APending Publication Date: 2025-10-17GUIZHOU AEROSPACE TIANMA ELECTRICAL TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510800847.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing manipulators have low grasping accuracy in cargo handling and are unable to adapt to the grasping needs of different cargoes, resulting in long loading time, low efficiency and insufficient safety.

Method used

A robot adaptive grasping and placing system based on multimodal visual positioning is adopted, which combines the display and control terminal, the robot main control module, the multimodal visual positioning module and the handheld remote control. The target detection model is loaded through the multimodal visual positioning module, and the approxPolyDP polygon fitting algorithm of OpenCV is used to identify the grasping and placing positions, realizing adaptive recognition and precise positioning.

Benefits of technology

It improves the grasping and releasing accuracy of the robot in different scenarios and environments, enhances its adaptive ability, improves the work efficiency and operation safety of cargo grasping, and improves the level of automation and unmanned operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120791735A_ABST
    Figure CN120791735A_ABST
Patent Text Reader

Abstract

A manipulator self-adaptive grabbing and releasing system based on multi-mode visual positioning is applied to a filling vehicle and comprises a manipulator master control module, a multi-mode visual positioning module, a display control terminal and a handheld remote controller. When the display control terminal responds to a user instruction and issues a grabbing and releasing control instruction to the mechanical arm main control module, a user inputs cargo parameters through the display control terminal, and the mechanical arm main control module is unfolded into different postures according to the cargo parameters. And meanwhile, a grabbing and releasing position measurement instruction is sent to the multi-mode visual positioning module, the multi-mode visual positioning module loads a target detection model after receiving the grabbing and releasing position measurement instruction, the grabbing and releasing position is measured, grabbing and releasing position coordinate information is fed back to the mechanical arm main control module, and the mechanical arm main control module controls the mechanical arm to work. And after receiving the grabbing and releasing position coordinate information, the manipulator master control module controls the manipulator to execute grabbing and releasing actions. The device can adapt to the states of different cargoes, and can still achieve precise positioning and precise grabbing and placing in different scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of cargo grabbing devices, and particularly relates to a mechanical hand self-adaptive grabbing and placing system based on multi-modal visual positioning. BACKGROUND

[0002] With the rapid development of intelligent manufacturing and robot technology, mechanical arms have been widely applied in industrial production, logistics and warehousing, cargo handling, medical treatment and other fields. As one of the core functions of mechanical arms, mechanical hand grabbing technology is of great significance for realizing automated, unmanned and intelligent grabbing and improving work efficiency. Therefore, the mechanical hand grabbing technology is expected to realize high integration of reliability and flexibility, and promote its application in more fields. Mechanical arms are also widely used in cargo loading. Traditional cargo handling requires a large amount of manpower and time, and has low work efficiency. Meanwhile, it is accompanied by human factors affecting cargo quality and personnel safety problems. Therefore, it is very important to study the application of mechanical hands in cargo handling.

[0003] At present, most of the mechanical hands in cargo handling still need to be manually controlled by personnel to grab the position, which has a distance error of visual observation and operation of the mechanical hand, and poor grabbing precision, so that the research on the mechanical hand in cargo handling gradually moves towards mechanical and system calculation of grabbing position. However, different calculation methods lead to different calculation accuracies, and the calculation accuracy of the grabbing position still needs to be improved. At the same time, in the existing grabbing technology, different cargo grabbing has different grabbing postures and positions, so it is often necessary to design and use mechanical hands and systems with strong pertinence.

[0004] For example, the patent document with publication number CN118478352A specifically discloses an industrial robot 3D visual positioning method and system. The three-dimensional coordinate system corresponding to the grabbing object is analyzed, and the coordinate points obtained from the image are judged in combination with the obtained image information. For abnormal situations, the coordinate points are corrected, and the grabbing points are generated according to the corrected coordinate points. Then, the grabbing object is judged and analyzed according to the generated grabbing points, and the most suitable stress point is generated by comprehensively considering the parameters of the grabbing object. The most suitable stress point is analyzed and used as the subsequent grabbing target point, which reduces the error in the subsequent grabbing process, improves the accuracy of 3D visual positioning, solves the technical problems of hand-eye calibration error, workpiece grabbing and placing position setting error and other system errors, and improves the positioning accuracy. However, it cannot adapt to the grabbing of different cargos.

[0005] Therefore, a technical scheme with high grabbing and positioning accuracy and capable of adapting to different cargo grabbing requirements needs to be researched to solve the problems existing in the prior art, shorten the loading time, improve the work efficiency and operation safety of cargo grabbing, and improve the automation, intelligence and unmanned level of cargo grabbing. SUMMARY

[0006] To solve the above technical problems, the application provides a mechanical hand adaptive picking and placing system based on multi-modal visual positioning, applied to a loading vehicle, comprising a mechanical hand master control module, a multi-modal visual positioning module, a display control terminal and a handheld remote controller.

[0007] The display control terminal is used to realize user interaction, and the user interaction content includes cargo parameter setting, position information display of each joint of the mechanical hand, and issuing picking and placing control instructions to the mechanical hand master control module in response to user instructions, wherein the cargo parameters include the type, length, width, height of the cargo and the position information of the cargo relative to the loading vehicle.

[0008] The mechanical hand master control module is used to receive the picking and placing control instructions issued by the display control terminal, and send picking and placing position measurement instructions to the multi-modal visual positioning module, and control the mechanical hand to perform picking and placing actions through the picking and placing position information fed back by the multi-modal visual positioning module.

[0009] The multi-modal visual positioning module is used to receive the picking and placing position recognition instructions sent by the mechanical hand master control module, perform picking and placing position recognition, and feed back picking and placing position coordinate information to the mechanical hand master control module.

[0010] The handheld remote controller is used to issue picking and placing control instructions to the mechanical hand master control module.

[0011] When the display control terminal issues picking and placing control instructions to the mechanical hand master control module in response to user instructions, the user inputs cargo parameters through the display control terminal, the mechanical hand master control module expands into different postures according to the cargo parameters, and sends picking and placing position measurement instructions to the multi-modal visual positioning module, the multi-modal visual positioning module loads a target detection model after receiving the picking and placing position measurement instructions, measures the picking and placing position, feeds back picking and placing position coordinate information to the mechanical hand master control module, and the mechanical hand master control module receives the picking and placing position coordinate information, controls the mechanical hand to perform picking and placing actions.

[0012] Further, the mechanical hand master control module comprises a master control unit, a driving unit and a first communication unit, the master control unit is used to control the mechanical hand to perform picking and placing actions, the driving unit drives the mechanical hand to complete picking and placing actions through a servo motor and a driver, and the first communication unit is used to establish communication between the master control unit and the driving unit, and between the master control unit and the multi-modal visual positioning module.

[0013] Further, the first communication unit comprises a CAN communication unit and an Ethernet communication unit, the CAN communication unit is used to establish communication between the master control unit and the driving unit, and the Ethernet communication unit is used to establish communication between the master control unit and the multi-modal visual positioning module.

[0014] Further, the multi-modal visual positioning module comprises a visual positioning control unit, a second communication unit, and a positioning recognition unit, the visual positioning control unit issues a position recognition instruction to the positioning recognition unit through the second communication unit, the positioning recognition unit performs position recognition after receiving the position recognition instruction, and feeds back to the visual positioning control unit through the second communication unit, and the visual positioning control unit transmits the recognized position coordinate information to the master control module of the manipulator in real time through the second communication unit.

[0015] Further, the visual positioning control unit adopts a GPU graphics card to realize the operation function, the second communication unit adopts an Ethernet communication unit to realize the communication connection between the multi-modal visual positioning module and the master control module of the manipulator, and the positioning recognition unit adopts a ToF RGB camera to realize the measurement of the pick-and-place position.

[0016] Further, the positioning recognition unit comprises a measurement camera and a calibration camera, the measurement camera and the calibration camera measure from different angles at the same time and have an overlapping part in the measurement area, so as to ensure the measurement of the same goods.

[0017] Further, the pick-and-place position recognition comprises pick-up position recognition and placement position recognition.

[0018] Further, when the multi-modal visual positioning module performs pick-up position recognition, the following contents are included:

[0019] When the multi-modal visual positioning module receives the pick-up position recognition instruction sent by the master control module of the manipulator, the visual positioning control unit loads a target detection model;

[0020] Based on the target detection model, the visual positioning control unit issues a recognition instruction to the positioning recognition unit, the positioning recognition unit performs target recognition on the pick-up hole contour structure of the goods, and obtains pick-up hole contour structure information;

[0021] According to the pick-up hole contour structure information, the visual positioning control unit adopts an approxPolyDP polygon fitting algorithm in OpenCV to fit the pick-up hole contour structure information into a quadrilateral model, calculates the four vertex coordinates of the quadrilateral model, obtains pick-up hole coordinate information, and feeds back to the master control module of the manipulator.

[0022] Further, when the multi-modal visual positioning module performs placement position recognition, the following is included:

[0023] When the mechanical hand runs to the placement area after grabbing the goods, the multi-modal visual positioning module automatically recognizes the placement position, and the visual positioning control unit loads the target detection model;

[0024] Based on the target detection model, the visual positioning control unit sends a recognition instruction to the positioning recognition unit, the positioning recognition unit performs target recognition on the placement positioning block contour to obtain placement positioning block contour information;

[0025] According to the placement positioning block contour information, the visual positioning control unit uses the approxPolyDP polygon fitting algorithm in OpenCV to fit the placement positioning block contour information into a quadrilateral model, calculates the coordinates of the four vertices of the quadrilateral model, obtains the placement positioning block coordinate information, and feeds back to the mechanical hand master control module.

[0026] Further, the method for the mechanical hand master control module to control the mechanical hand to perform the grab-and-place action includes:

[0027] When the mechanical hand master control module receives the grab hole coordinate information fed back by the multi-modal visual positioning module, the master control unit fits the four grab hole coordinate information into a quadrilateral through an algorithm, and controls the mechanical hand to move to the grab hole to realize the grab action of the goods according to the coordinate information of the quadrilateral;

[0028] When the mechanical hand runs to the placement area after grabbing the goods, the mechanical hand master control module fits the four placement positioning block coordinate information into a quadrilateral through an algorithm according to the placement positioning block coordinate information fed back by the multi-modal visual positioning module, and controls the mechanical hand to complete the drop action of the goods according to the coordinate information of the quadrilateral.

[0029] The beneficial effects of the present application are that: by loading a target detection model in the multi-modal visual positioning module, the contours of different goods are intelligently identified, so as to adaptively identify the grabbing position; by using the positioning method of fitting a quadrilateral model when the multi-modal visual positioning module performs grabbing position identification, the visual positioning accuracy of the manipulator adaptive grabbing and placing system is improved; by also using the method of fitting a quadrilateral in the execution of the grabbing and placing action of the manipulator master control module, the accuracy of the grabbing and placing action of the manipulator is improved, so that the manipulator adaptive grabbing and placing system of the present application can still accurately position and accurately perform the grabbing and placing action under complex conditions such as different scenes, different environmental conditions, and different types of goods, adapt to the state of different goods, adapt to grabbing and placing, adapt to unfolding, intelligently identify, improve the work efficiency and safety of goods grabbing, improve the automation, intelligence and unmanned level of goods grabbing, and to some extent, fill the blank in the field of domestic filling vehicle manipulator control systems, and have a certain reference significance for the development and application of computer vision in this industry. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is a structural block diagram of a manipulator adaptive grabbing and placing system based on multi-modal visual positioning provided by the present application;

[0031] Figure 2 is a schematic diagram of the installation position of a ToF RGB camera provided by the present application;

[0032] Figure 3 is a schematic diagram of the bottom of a manipulator end lock pin structure provided by the present application;

[0033] Figure 4 is a communication connection relationship diagram of a manipulator adaptive grabbing and placing system based on multi-modal visual positioning provided by the present application;

[0034] Figure 5 is a positioning and identification flowchart of a multi-modal visual positioning module provided by the present application;

[0035] Figure 6 is a training target retrieval model flowchart of a manipulator adaptive grabbing and placing system based on multi-modal visual positioning provided by the present application;

[0036] Figure 7 is a connection coordinate system diagram of a manipulator end lock pin structure, a ToF RGB camera and a grabbing hole provided by the present application;

[0037] Figure 8 is a ToF RGB camera and lock pin structure center point coordinate system conversion diagram provided by the present application;

[0038] Figure 9 is a multi-modal visual positioning module identification fitting quadrilateral model flowchart of a grabbing hole provided by the present application;

[0039] Figure 10 is a quadrilateral diagram fitted according to the grab hole coordinates fed back by the multi-modal visual positioning module, provided by the application;

[0040] Figure 11 is a device diagram for placing the positioning block, provided by the application;

[0041] Figure 12 is a flowchart of the adaptive grabbing algorithm of the manipulator, provided by the application;

[0042] Figure 13 is a flowchart of the cargo placing algorithm of the manipulator, provided by the application;

[0043] Figure 14 is the working principle of the ToF RGB depth camera in the multi-modal visual positioning module, provided by the application. DETAILED DESCRIPTION

[0044] The technical solutions of the application are further described below, but the scope of protection is not limited to the description.

[0045] The application embodiment provides a manipulator adaptive grabbing and placing system based on multi-modal visual positioning, which solves the problems of low grabbing precision, inability to adaptively grab different cargos, small application range and low intelligent level in the existing filling vehicle manipulator grabbing technology.

[0046] The manipulator adaptive grabbing and placing system based on multi-modal visual positioning is applied to a filling vehicle, as shown in Figure 1 , comprising a manipulator master control module, a multi-modal visual positioning module, a display control terminal and a handheld remote controller.

[0047] The display control terminal is used to realize user interaction, and the user interaction content includes cargo parameter setting, position information display of each joint of the manipulator, and issuing grabbing and placing control instructions to the manipulator master control module in response to user instructions, wherein the cargo parameters include the type, length, width, height of the cargo and the position information of the cargo relative to the filling vehicle.

[0048] The handheld remote controller is used to issue grabbing and placing control instructions to the manipulator master control module.

[0049] The grabbing and placing instructions issued by the handheld remote controller include manipulator unfolding, one-key grabbing, one-key placing and one-key grabbing and placing, etc., the display control terminal includes display control software and a display screen, the display screen is used to display the position information of each joint of the manipulator, and the display control software realizes control of the manipulator master control module and the multi-modal visual positioning module in response to user instructions to complete the grabbing and placing instructions.

[0050] The manipulator main control module is used to receive the grasping and releasing control instructions issued by the display and control terminal, and send grasping and releasing position measurement instructions to the multimodal visual positioning module, and control the manipulator to perform grasping and releasing actions through the grasping and releasing position information fed back by the multimodal visual positioning module;

[0051] The manipulator main control module includes a main control unit, a drive unit and a first communication unit. The main control unit is used to control the manipulator to perform grasping and releasing actions. The drive unit drives the manipulator to complete the grasping and releasing actions through a servo motor and a driver. The first communication unit is used to establish communication between the main control unit and the drive unit, and between the main control unit and the multimodal visual positioning module.

[0052] The first communication unit includes a CAN communication unit and an Ethernet communication unit. The CAN communication unit is used to establish communication between the main control unit and the drive unit, and the Ethernet communication unit is used to establish communication between the main control unit and the multimodal visual positioning module to transmit the position information of the grasping point and the placement point.

[0053] The multimodal visual positioning module is used to receive the grasping and placing position recognition instruction sent by the manipulator main control module, perform grasping and placing position recognition, and feedback grasping and placing position coordinate information to the manipulator main control module; the positioning recognition process of the multimodal visual positioning module is as follows Figure 5 shown.

[0054] The multimodal visual positioning module includes a visual positioning control unit, a second communication unit and a positioning identification unit. The visual positioning control unit sends a position identification instruction to the positioning identification unit through the second communication unit. After receiving the position identification instruction, the positioning identification unit performs position identification and feeds back to the visual positioning control unit through the second communication unit. The visual positioning control unit transmits the identified position coordinate information to the manipulator main control module through the second communication unit in real time.

[0055] The communication connection relationship of the manipulator adaptive grasping and placing system based on multimodal visual positioning is as follows: Figure 4 shown.

[0056] The visual positioning control unit uses a GPU graphics card to realize the computing function, the second communication unit uses an Ethernet communication unit to realize the communication connection between the multimodal visual positioning module and the manipulator main control module, and the positioning recognition unit uses a ToF RGB camera to realize the measurement of the grasping and releasing position. The installation position of the ToF RGB camera is as follows: Figure 2 shown.

[0057] The positioning recognition unit includes a measurement camera and a calibration camera, which measure simultaneously from different angles and have overlapping measurement areas to ensure measurement of the same goods. The installation position diagram of the measurement camera and the calibration camera is shown in Figure 8

[0058] The positioning recognition unit also includes a goods grabbing and placing positioning structure. The grabbing structure is an adapter for the lock pin structure of the mechanical hand when grabbing goods. The bottom diagram of the lock pin structure is shown in Figure 3 The placing positioning structure is a placing positioning block, which is a rectangular mechanical structure. The device diagram of the placing positioning block is shown in Figure 11 The surface of the placing positioning block is painted in bright color to assist the mechanical hand in accurately positioning the placing point when placing goods.

[0059] The grabbing and placing position recognition includes grabbing position recognition and placing position recognition.

[0060] When the multi-modal visual positioning module performs grabbing position recognition, the following contents are included:

[0061] When the multi-modal visual positioning module receives the grabbing position recognition instruction sent by the mechanical hand master control module, the visual positioning control unit loads the target detection model. The training process of the target detection model is shown in Figure 6

[0062] Based on the target detection model, the visual positioning control unit sends an identification instruction to the positioning recognition unit. The positioning recognition unit performs target recognition on the grabbing hole contour structure of the goods to obtain grabbing hole contour structure information.

[0063] According to the grabbing hole contour structure information, the visual positioning control unit uses the approxPolyDP polygon fitting algorithm in OpenCV to fit the grabbing hole contour structure information into a quadrilateral model, calculates the coordinates of the four vertices of the quadrilateral model, obtains the grabbing hole coordinate information, and feeds back to the mechanical hand master control module.

[0064] When the multi-modal visual positioning module performs placing position recognition, the following contents are included:

[0065] When the mechanical hand grabs goods and runs to the placing area, the multi-modal visual positioning module automatically recognizes the placing position. At the same time, the visual positioning control unit loads the target detection model.

[0066] ​​Based on the target detection model, the visual positioning control unit issues an identification instruction to the positioning identification unit, the positioning identification unit performs target identification on the placement positioning block contour, and obtains placement positioning block contour information;

[0067] According to the placement positioning block contour information, the visual positioning control unit uses the approxPolyDP polygon fitting algorithm in OpenCV to fit the placement positioning block contour information into a quadrilateral model, calculates the coordinates of the four vertices of the quadrilateral model, obtains the placement positioning block coordinate information, and feeds back to the manipulator master control module.

[0068] The method for the manipulator master control module to control the manipulator to perform a pick-and-place action includes:

[0069] When the manipulator master control module receives the pick hole coordinate information fed back by the multi-modal visual positioning module, the master control unit fits the four pick hole coordinate information into a quadrilateral through an algorithm, and controls the manipulator to move to the pick hole to realize the pick action of the goods according to the coordinate information of the quadrilateral;

[0070] When the manipulator picks up the goods and moves to the placement area, the manipulator master control module fits the four placement positioning block coordinate information into a quadrilateral through an algorithm according to the placement positioning block coordinate information fed back by the multi-modal visual positioning module, and controls the manipulator to complete the drop action of the goods according to the coordinate information of the quadrilateral.

[0071] The master control unit fits the quadrilateral of the pick hole coordinate fed back by the multi-modal visual positioning module as Figure 10 shown.

[0072] When the display control terminal responds to the user instruction and issues a pick-and-place control instruction to the manipulator master control module, the user inputs the goods parameters through the display control terminal, the manipulator master control module expands into different postures according to the goods parameters, and simultaneously sends a pick-and-place position measurement instruction to the multi-modal visual positioning module, the multi-modal visual positioning module loads a target detection model after receiving the pick-and-place position measurement instruction, measures the pick-and-place position, and feeds back pick-and-place position coordinate information to the manipulator master control module, the manipulator master control module receives the pick-and-place position coordinate information, and controls the manipulator to perform a pick-and-place action.

[0073] The relative position information between the grabbing hole and the measuring camera is calculated by the multi-modal visual positioning system during the grabbing process, and then transmitted to the manipulator master control module through the second communication unit. The master control unit inputs the position parameters fed back by the multi-modal visual positioning module, controls the driving unit to move the manipulator above the grabbing hole structure, and according to the coordinate points obtained by the positioning method, fits a quadrilateral formed by the center points of the cargo grabbing hole, such as a rectangle, to realize accurate docking of the cargo grabbing hole and ensure that the next grabbing action of the manipulator can be implemented.

[0074] During the placing process, the manipulator needs to move to the target position point of the goods, which is composed of four goods placing positioning blocks. When the manipulator grabs the goods and moves above the positioning blocks, the multi-modal visual positioning module will obtain the coordinate information of the goods placing positioning blocks, fit a quadrilateral area for placing the goods, calculate the relevant parameters of the manipulator moving to the target position, and complete the placing operation of the goods by the manipulator master control module.

[0075] The filling vehicle manipulator includes four supporting arms, each of which is composed of a shoulder joint, an elbow joint, a wrist joint and a lock pin. Each joint controls the unfolding and retracting posture of the manipulator, and the four supporting arms have the function of self-adapting unfolding and retracting. The self-adapting function in the multi-modal visual positioning-based manipulator self-adapting grabbing and placing system is realized by each joint of the manipulator, which has the same flexibility as human arms.

[0076] The manipulator master control module adaptively unfolds into different postures to realize grabbing and placing actions according to the cargo-related parameters input by the display control terminal.

[0077] The manipulator servo motor and driver include four shoulder joint motors, four elbow joint motors, four wrist joint motors and four lock pin motors, which are respectively connected with the driver in combination and connected with the manipulator master control module through cables.

[0078] The positioning method of the multi-modal visual positioning module when performing grabbing and placing position recognition is realized by the following steps:

[0079] Step 1: selection, installation and calibration of manipulator visual camera equipment.

[0080] The camera used by the positioning recognition unit in the multi-modal visual positioning method is a ToF RGB (Time of Flight) depth camera. Its ranging principle is to send light pulses to the target continuously, then use a sensor to receive the pulse signal reflected from the object, calculate the round-trip flight time of the light pulse to obtain the target distance, and combine the internal data processing to return the two-dimensional information and depth information of the object in the image. Its working principle is as follows: Figure 14The ToF RGB depth camera measures depth by light flight time, is not affected by the complex texture features of the object surface, and the depth information obtained at close range is more accurate and stable, meeting the requirements of close-range positioning. In view of the poor performance of the long-distance positioning method based on the binocular camera in close-range positioning, a ToF RGB depth camera is used to realize close-range three-dimensional target positioning.

[0081] The measuring camera is installed on the side wall of the end of the truck loading robot for fixing the mechanical mechanism, and the installation position is as shown in Figure 2 Since the measuring camera is installed on the side wall, translation and rotation calibration parameters need to be calculated before use, so as to obtain the coordinate and pose information of the end effector of the robot when grabbing the goods in real time.

[0082] The calculation process of the measuring camera calibration parameters is as follows:

[0083] a. Calibrate the intrinsic parameters: First, the intrinsic parameters of the measuring camera and the calibration camera need to be calibrated, that is, the focal length, principal point position and other parameters of the camera are determined. The intrinsic parameter matrix of the camera is obtained by using a camera calibration board or other calibration methods, and then the relevant parameters are calculated by shooting images at different angles and using a camera calibration algorithm.

[0084] b. Observe the common measured object: Ensure that the measuring camera and the calibration camera can observe the same measured object at the same time, which can be realized by placing the measured object in the overlapping area of the field of view of the two cameras.

[0085] c. Feature point extraction: Select feature points on the measured object, such as corner points, edges, etc. Ensure that these feature points can be accurately extracted in the images of the two cameras.

[0086] d. Feature point matching: Use feature point ORB matching algorithm, such as SIFT, SURF, ORB, to match the feature points of the two cameras, and establish the corresponding relationship between the two images.

[0087] e. Three-dimensional reconstruction: Use the corresponding relationship of the feature points to calculate the position of the feature points in the three-dimensional space by using triangulation or other three-dimensional reconstruction methods.

[0088] f. Extrinsic parameter calculation: According to the results of three-dimensional reconstruction, use camera pose estimation algorithm, such as ICP algorithm, to calculate the extrinsic parameters of the other camera, that is, the position and attitude information of the camera.

[0089] The feature point extraction and matching in the above steps may be affected by factors such as light and occlusion, so preprocessing or more complex algorithms can be used to improve the accuracy of matching in actual operation.

[0090] Step 2: Image data acquisition and preprocessing of grabbing hole and placing positioning block.

[0091] The acquisition and preprocessing of the grabbing hole image data is an important basis for identifying the grabbing hole, and is an important factor for improving the accuracy of identifying the grabbing hole in the cargo grabbing process. First, a ToF RGB depth camera is used to collect images of the structure of the grabbing hole, and the grabbing hole images are taken under multi-angle, natural light, and night environment conditions. The diversity of data and the robustness of the model are increased, and data enhancement techniques such as random cropping, rotation, and flipping are used to preprocess the images. Then, the quadrilateral contour in the collected grabbing hole image is labeled, and the LabelImg tool is used for labeling. Finally, the labeled data set is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1 for model training. The premise of image acquisition and preprocessing of the placing positioning block during the placing process is that the robot lifts the cargo and moves to the top of the placing positioning block to take images. The remaining processing process is consistent with the operation of the grabbing process.

[0092] Step 3: Select the grabbing hole and placing block target detection model.

[0093] Considering the accuracy requirements of the robot grabbing and placing cargo, YoloV7 is selected as the model for identifying the grabbing hole and placing positioning block. By introducing an attention mechanism module, the shortest and longest path of the gradient is controlled, allowing the deeper network to learn and converge more efficiently. YoloV7 mainly optimizes the re-parameterization module and dynamic label assignment, further reducing the complexity of the model and obtaining a relatively lightweight detection model. The present invention embeds a coordinate attention mechanism in the multi-modal visual positioning module backbone network structure, fully considers the relationship between channels and position information, and captures cross-channel information while considering direction-related position information, which helps the model to better locate and identify targets. Some traditional convolution modules in the multi-modal visual positioning module backbone network are replaced by deformable convolution modules to improve the model's fitting ability for different shapes of targets in the scene, thereby helping to extract refined features and reduce irrelevant feature interference.

[0094] Target detection in real-world scenarios faces numerous challenges, including complex backgrounds, occlusions between different objects, varying lighting conditions, scale variations of targets, and diversity of target features. These factors make target recognition and localization more difficult. For target detection methods in real-world scenarios, the introduction of attention mechanisms can significantly improve the model's ability to handle complex backgrounds, focusing on key areas and strengthening the learning of target features, thus more effectively extracting the most critical feature information for the current task. This mechanism not only helps to cope with challenges such as occlusion, lighting changes, and scale differences, but also adaptively handles targets of different scales, improving the model's robustness in variable environments.

[0095] Step 4: Train the target detection model for the grab hole.

[0096] Using the YoloV7 deep learning model to identify the contour information of the grab hole, a large number of grab hole shape image information under different lighting conditions, different directions, and different visual angles is needed during training. The model training process is as shown in Figure 6

[0097] The hardware environment for training the model is two work computers, with an intel Core i7 processor, 32G memory, and an Rtx3060Ti GPU card each. The main software tools and development languages involved are: Python3.7, Pytorch1.8, Pycharm2020, Labelme, etc.

[0098] Parameter configuration mainly includes: batch_size, learning rate, input image resolution.

[0099] Step 5: Load the trained target detection model to identify the grab hole and the placement positioning block to obtain the contour points of the target.

[0100] During the cargo grabbing process, when the multi-modal visual positioning module receives the grab position recognition instruction from the main control unit, the visual positioning control unit loads the target detection model to perform target recognition on the grab hole structure. During the cargo placement process, the visual positioning control unit loads the target detection model to perform target recognition on the placement positioning block, obtaining the contour points of the grab hole or the placement positioning block.

[0101] Step 6: Use the contour point position information obtained in step 5 to use the approxPolyDP polygon fitting algorithm in OpenCV to fit the contour of the grab hole into a quadrilateral.

[0102] ​The image is pre-processed by using grayscale and edge detection algorithms. The edge detection algorithm Canny finds the edges in the image. For the detected edge points, some edge points are randomly selected in each iteration of the approxPolyDP algorithm. A quadrilateral model is fitted according to these points of the grabbing hole. Distance measurement is used to calculate the fitting error of each fitted quadrilateral model with all edge points. According to a predefined threshold, the edge points with a fitting error less than the threshold are regarded as inner points, and the remaining edge points are regarded as outer points.

[0103] Then the local coordinate system is defined, as shown in Figure 9 The four vertex information of the grabbing hole contour is calculated. The four contour points after quadrilateral fitting are arranged in a clockwise direction, and the local coordinate system of the target is specified. The right-hand coordinate system is used, and the Z-axis direction is perpendicular to the XY plane and points to the inside of the grabbing hole. The length of the quadrilateral and the local coordinate are obtained by Euclidean distance.

[0104] Step 7: Calculate the coordinate information of the grabbing hole center point and the placement positioning block.

[0105] Under the condition that the camera intrinsic parameter is known, the position information and attitude information of the working camera relative to the grabbing hole are obtained by using the PnP algorithm according to the obtained four three-dimensional contour points and the corresponding two-dimensional plane coordinate points in the image, including the rotation variable R1 and the translation vector T1. Then the calibration camera mounted on the side wall of the manipulator is used to collect images of the grabbing hole, and point cloud information is generated. According to the point cloud information, the camera extrinsic parameter is calibrated, and the relative position and attitude information of the calibration camera relative to the measurement camera is calculated, including the rotation vector R2 and the translation vector T2. Through R1, T1 and R2, T2, the pose information of the grabbing point relative to the grabbing hole can be calculated, including the rotation vector R3 and the translation vector T3, and the three-dimensional position information of the grabbing structure in the local coordinate system of the grabbing hole is obtained, that is, the position deviation of the grabbing hole relative to the center point of the manipulator lock pin structure. The coordinate conversion process of the calibration camera is shown in Figure 8 The end of the manipulator lock pin structure, the ToF RGB camera and the grabbing hole connection coordinate system are shown in Figure 7 During the placement of the goods, the vision positioning control unit obtains the contour points of the four placement positioning blocks. The main control control unit fits the placement point position information fed back by the vision positioning control unit into a quadrilateral, and calculates the position deviation value of the placement positioning block relative to the manipulator lock pin structure. The main control control unit then issues a control instruction to make the manipulator move to the target position to complete the goods lowering action.

[0106] Step 8: Fit the quadrilateral composed of four manipulator grabbing point coordinates.

[0107] When the mechanical hand master control module receives the grasping point coordinate information (x, y, z) fed back by the multi-modal visual positioning module, four points are fitted into a rectangle through an algorithm, and the mechanical hand master control module controls the motor to move the mechanical arm to the grasping point according to the coordinate information of the rectangle, so as to realize the grasping of the goods.

[0108] Due to the influence of various factors such as illumination, calibration error and communication quality, the coordinate information fed back by the multi-modal visual positioning module in the actual application process may be missing, and 2-4 groups of coordinate information need to be fed back to fit the rectangle composed of the grasping points. A multi-point translation and rotation transformation method is used to iteratively optimize and solve the position and attitude angle deviation of the mechanical hand, and the parameter solving method is the Gauss gradient descent method.

[0109] The local position deviation fed back by the visual positioning control unit is as follows:

[0110]

[0111] Each row in the formula represents the position deviation information of a grasping point. In order to shorten the time of iterative optimization, the height error is averaged separately, and the calculation process is as follows:

[0112]

[0113] The theoretical translation and rotation matrix of the two-dimensional plane is as follows:

[0114]

[0115] That is:

[0116]

[0117] The actual coordinates fed back by the visual positioning control unit are:

[0118]

[0119] The design optimization index function is the error square of the translation and rotation coordinates and the measured coordinates:

[0120]

[0121] J according to the chain rule to solve the gradient of the variable (dx, dy, φ):

[0122]

[0123] The iterative optimization formula of the rotation and translation parameters by using the gradient descent method is as follows:

[0124]

[0125] Through the above calculation process, 4 sets of coordinate information are obtained, which correspond to the 4 corners of the quadrilateral, i.e. the target position controlled by the locking pin structure when the manipulator grabs.

[0126] Step 9: fitting the quadrilateral composed of the center point coordinates of the manipulator goods placement positioning block

[0127] According to the work flow of goods grabbing and placing, when the goods are grabbed by the manipulator and moved above the placement area, the multi-modal visual positioning module automatically obtains the position information of the placement positioning block, i.e. loads the target detection model to identify the goods placement positioning block, calculates the center point coordinate information of the placement point, and feeds back to the main control unit, which calls the manipulator control algorithm to complete the goods lowering operation. At the same time, the visual positioning control unit is affected by various factors, such as light, calibration error, and communication quality. The coordinate information fed back by the visual positioning control unit in the actual application process may be missing, and 2-4 sets of visual coordinate information need to be fed back to fit the rectangle composed of the center points of the placement positioning block, as shown in Figure 10 .

[0128] The manipulator grabbing and placing action control algorithm is realized by the main control unit, and the goods grabbing algorithm flow is as shown in Figure 12 , and the goods placing algorithm flow is as shown in Figure 13 .

[0129] The detailed steps of the manipulator grabbing and placing algorithm flow are as follows:

[0130] a) Manipulator goods grabbing process

[0131] The grabbing process includes manipulator unfolding, positioning measurement, pin insertion, rigid-flexible conversion, and lifting function. The grabbing algorithm flow is as shown in Figure 12 , the manipulator unfolding function is to move each joint to the target position according to the size of the goods, i.e. the locking pin structure of the 4 arms. After the manipulator is unfolded by one key, the user issues a grabbing and placing position recognition instruction through the display control terminal. The manipulator main control module receives the grabbing and placing position recognition instruction and transmits it to the multi-modal visual positioning module. The multi-modal visual positioning module obtains the target position information of the grabbing point and returns it to the main control unit. The main control unit controls the movement of the manipulator to the center point of the goods grabbing hole according to the feedback position information, and the locking pin structure can perform locking action, and then the manipulator performs one-key rigid-flexible conversion function.

[0132] b) Manipulator goods placing process

[0133] When the mechanical arm moves to the placement area above the goods, the multi-modal visual positioning module loads the target detection model to identify the placement positioning block, and feeds back the identified region center point coordinates to the master control unit. The master control unit calculates control parameters according to the target region center point coordinates, and controls the mechanical arm to complete the goods lowering action through the driving unit. After the goods are successfully lowered, the mechanical arm performs the pin pulling action to complete the entire goods hoisting process.

[0134] The above disclosed is only a specific embodiment of the present application, but the present application is not limited thereto, and any changes that can be thought of by those skilled in the art shall fall within the protection scope of the present application.

Claims

1. A manipulator adaptive pick-and-place system based on multimodal visual positioning, characterized in that: Applied to loading vehicles, including manipulator main control module, multi-modal visual positioning module, display and control terminal and handheld remote control; The display and control terminal is used to implement user interaction; the user interaction content includes: setting cargo parameters, displaying the position information of each joint of the manipulator, and issuing grab and release control instructions to the manipulator main control module in response to user instructions. The cargo parameters include the type, length, width, height, and position information of the cargo relative to the loading vehicle. The manipulator main control module is used to receive the grasping and releasing control instructions issued by the display and control terminal, and send grasping and releasing position measurement instructions to the multimodal visual positioning module, fit the grasping and releasing position information fed back by the multimodal visual positioning module into a quadrilateral, and control the manipulator to perform grasping and releasing actions according to the coordinates of the quadrilateral; The multimodal visual positioning module is used to receive the grasping and placing position recognition instruction sent by the manipulator main control module, perform grasping and placing position recognition, and feed back grasping and placing position coordinate information to the manipulator main control module; The handheld remote controller is used to send a grasping and releasing control instruction to the manipulator main control module; When the display and control terminal responds to user instructions and sends a grabbing and releasing control instruction to the manipulator main control module, the user inputs cargo parameters through the display and control terminal, and the manipulator main control module expands into different postures according to the cargo parameters, and at the same time sends a grabbing and releasing position measurement instruction to the multimodal visual positioning module. After receiving the grabbing and releasing position measurement instruction, the multimodal visual positioning module loads the target detection model, measures the grabbing and releasing position, and feeds back the grabbing and releasing position coordinate information to the manipulator main control module. After receiving the grabbing and releasing position coordinate information, the manipulator main control module controls the manipulator to perform the grabbing and releasing action.

2. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 1, wherein: The manipulator main control module includes a main control unit, a drive unit and a first communication unit. The main control unit is used to control the manipulator to perform grasping and releasing actions. The drive unit drives the manipulator to complete the grasping and releasing actions through a servo motor and a driver. The first communication unit is used to establish communication between the main control unit and the drive unit, and between the main control unit and the multimodal visual positioning module.

3. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 2, wherein: The first communication unit includes a CAN communication unit and an Ethernet communication unit. The CAN communication unit is used to establish communication between the main control unit and the drive unit, and the Ethernet communication unit is used to establish communication between the main control unit and the multimodal visual positioning module.

4. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 3, wherein: The multimodal visual positioning module includes a visual positioning control unit, a second communication unit and a positioning identification unit. The visual positioning control unit sends a position identification instruction to the positioning identification unit through the second communication unit. After receiving the position identification instruction, the positioning identification unit performs position identification and feeds back to the visual positioning control unit through the second communication unit. The visual positioning control unit transmits the identified position coordinate information to the manipulator main control module through the second communication unit in real time.

5. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 4, characterized in that: The visual positioning control unit uses a GPU graphics card to realize computing functions, the second communication unit uses an Ethernet communication unit to realize communication connection within the multimodal visual positioning module and with the manipulator main control module, and the positioning identification unit uses a ToF RGB camera to realize the measurement of the grasping and releasing position.

6. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 5, wherein: The positioning and identification unit includes a measuring camera and a calibration camera. The measuring camera and the calibration camera measure from different angles simultaneously and the measurement areas have overlapping parts, so as to ensure that the same goods are measured.

7. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 1, wherein: The grab and place position recognition includes grab position recognition and placement position recognition.

8. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 7, wherein: The grasping position identification includes the following contents: When the multimodal visual positioning module receives the grasping position recognition instruction sent by the manipulator main control module, the visual positioning control unit loads the target detection model; Based on the target detection model, the visual positioning control unit sends a recognition instruction to the positioning recognition unit, and the positioning recognition unit performs target recognition on the contour structure of the grabbing hole of the goods to obtain the contour structure information of the grabbing hole; According to the grasping hole contour structure information, the visual positioning control unit uses the approxPolyDP polygon fitting algorithm in OpenCV to fit the grasping hole contour structure information into a quadrilateral model, calculates the four vertex coordinates of the quadrilateral model, obtains the grasping hole coordinate information, and feeds it back to the manipulator main control module.

9. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 8, wherein: The placement position identification includes the following: When the robot grabs the goods and moves to the placement area, the multimodal visual positioning module automatically identifies the placement location, and at the same time, the visual positioning control unit loads the target detection model; Based on the target detection model, the visual positioning control unit sends a recognition instruction to the positioning recognition unit, and the positioning recognition unit performs target recognition on the placement positioning block outline to obtain the placement positioning block outline information; According to the placement positioning block contour information, the visual positioning control unit uses the approxPolyDP polygon fitting algorithm in OpenCV to fit the placement positioning block contour information into a quadrilateral model, calculates the four vertex coordinates of the quadrilateral model, obtains the placement positioning block coordinate information, and feeds it back to the manipulator main control module.

10. The multimodal vision-based adaptive pick-and-place system for manipulators according to claim 9, wherein: The method of controlling the manipulator to perform a pick-and-place action by the manipulator main control module includes: When the manipulator main control module receives the grabbing hole coordinate information fed back by the multimodal visual positioning module, the main control unit fits the four grabbing hole coordinate information into a quadrilateral through an algorithm, and controls the manipulator to move to the grabbing hole according to the coordinate information of the quadrilateral to realize the grabbing action of the goods; When the robot grabs the goods and runs to the placement area, the robot main control module fits the four placement positioning block coordinate information into a quadrilateral through an algorithm based on the placement positioning block coordinate information fed back by the multimodal visual positioning module, and controls the robot to complete the lowering of the goods according to the coordinate information of the quadrilateral.

Citation Information

Patent Citations

  • Industrial robot 3D visual positioning method and system

    CN118478352A