Mechanical arm grabbing control system and method based on visual perception

By building a highly coordinated system of visual perception, strategy planning and motion control, the perception and execution bottlenecks of existing robotic arm systems in complex environments have been resolved, achieving high-precision, high-efficiency recognition and precise placement of multiple categories of target objects, and improving the operational accuracy and robustness of the robotic arm.

CN120791798APending Publication Date: 2025-10-17HANGZHOU DIANZI UNIVERSTIY INFORMATION ENG SCHOOL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511282539.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing robotic arm systems face perception and execution bottlenecks when faced with high-mix, small-batch, and high-precision production and sorting requirements. This is especially true in scenarios such as 3C electronic precision assembly, automatic splicing of creative toys, and sorting of special-shaped parts in logistics. The visual recognition, strategy generation, and motion control modules are separated, and there is a lack of a unified adaptive optimization framework, resulting in rigid operations and poor fault tolerance.

Method used

A robotic arm grasping control system based on visual perception is constructed, including a visual perception module, a strategy planning module, and a motion control module. The HSV color space segmentation and k-means clustering algorithm are used to realize target object recognition and positioning. The nine-point calibration method is used to establish the mapping between the image and the robotic arm coordinate system. The depth-first search algorithm is combined to generate the optimal task sequence, and the motion trajectory is optimized through arc transition.

Benefits of technology

It achieves high-precision, high-efficiency recognition, grasping and precise placement of multiple categories of target objects in complex environments, improves the accuracy, efficiency and robustness of robotic arm operations, and can adapt to operational requirements in complex unstructured environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120791798A_ABST
    Figure CN120791798A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of mechanical arm intelligent control, and discloses a mechanical arm grabbing control system and method based on visual perception, and the system comprises a visual perception module, a strategy planning module and a motion control module; the visual perception module is used for collecting image data of a working area and identifying color, shape and pose information of a target object through image processing as an identification result; the strategy planning module is used for modeling the working area into a two-dimensional grid map, and generating an optimal task sequence including a grabbing sequence, a placement coordinate and a placement posture for a target object according to an identification result and a current working area state; and the motion control module is used for planning the motion track of the mechanical arm according to the optimal task sequence and controlling the mechanical arm to execute grabbing, transferring and placing operations. According to the invention, high-precision and high-efficiency recognition, grabbing and accurate placement of multiple types of target objects in a complex environment can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent control of mechanical arms, and particularly relates to a mechanical arm grasping control system and method based on visual perception. BACKGROUND

[0002] In the evolution of industrial automation and intelligent manufacturing, mechanical arm systems have achieved efficient operation in many structured scenarios, but there are still significant perception and execution bottlenecks when facing high-mix, small-batch, and high-precision production and sorting demands. This challenge is particularly prominent in 3C electronic precision assembly, creative toy automatic assembly, and logistics irregular part sorting applications.

[0003] Existing technical solutions usually rely on a single sensing modality or isolated functional modules. For example, in 3C assembly, traditional visual methods are easily disturbed by reflections, small structures, and changes in lighting, making it difficult to simultaneously achieve high-robust recognition of part types and sub-millimeter level pose estimation. In toy block assembly, due to the lack of joint analysis capabilities for color, shape, and multi-degree-of-freedom pose, existing systems cannot stably complete the insertion task requiring precise fitting of hole shafts. In the logistics sorting field, the diversity of irregular part appearances and the randomness of placement make it difficult for automated systems based on fixed grasping points and static path planning to balance efficiency and success rate. These limitations can essentially be attributed to a common core problem: most existing systems fail to achieve deep closed-loop coordination among perception, decision-making, and execution. Their visual recognition, strategy generation, and motion control modules are often fragmented, lacking a unified adaptive optimization framework, resulting in a rigid and poor fault tolerance when the overall system faces uncertainty.

[0004] From the current technical development status, vision-based mechanical arm guidance has become a mainstream research direction, but most achievements still focus on isolated optimization for specific tasks, and a general-purpose, cross-scenario reusable technical system has not yet been formed. For example, in the perception layer, although deep learning object detection methods have made significant progress, their dependence on labeled data and high computational cost restrict their deployment in real-time control systems. At the same time, high-precision pose estimation algorithms under monocular vision often rely on texture features on object surfaces, and for typical industrial parts with low texture, high reflectivity, or repeated structures (such as electronic components and block pieces), their stability and accuracy in solving are still difficult to meet the requirements of physical interaction. Issues such as high-frequency data acquisition and processing, low-latency communication and control also pose high requirements on the system's infrastructure and computing power. SUMMARY

[0005] The application aims to provide a mechanical arm grasping control system and method based on visual perception, which can solve the operation bottleneck problem caused by inaccurate perception, rigid planning and control lag in traditional mechanical arm systems in 3C assembly, toy splicing and logistics sorting scenes, and can realize high-precision and high-efficiency identification, grasping and accurate placement of multi-category target objects in complex environments.

[0006] To achieve the above-mentioned purpose, the application provides the following basic scheme.

[0007] Scheme one A mechanical arm grasping control system based on visual perception, comprising a visual perception module, a strategy planning module and a motion control module; The visual perception module is used to collect image data of a working area, and identify color, shape and pose information of a target object through image processing as an identification result; The strategy planning module is used to model the working area as a two-dimensional grid map, and generate an optimal task sequence for the target object containing grasping sequence, placement coordinates and placement posture according to the identification result and the current working area state; The motion control module is used to plan the motion trajectory of the mechanical arm according to the optimal task sequence, and control the mechanical arm to perform grasping, transferring and placing operations; the motion trajectory includes a pick-up safety point, a pick-up point, a placement safety point and a placement point; The visual perception module adopts a nine-point calibration method to establish the mapping relationship between the image coordinate system and the mechanical arm coordinate system, and realizes the identification and positioning of the target object through HSV color space segmentation and k-means clustering algorithm.

[0008] Scheme two A mechanical arm grasping control method based on visual perception, which applies a mechanical arm grasping control system based on visual perception as described in scheme one to control the motion of the mechanical arm, comprising the following steps: Image acquisition step: collecting the working area image containing the target object through the visual perception module; Visual identification step: pre-processing the working area image, converting to HSV color space and performing color segmentation, generating a binary image using k-means clustering algorithm, and then identifying the color, shape and pixel coordinates of the target object; Coordinate conversion step: obtaining the affine transformation matrix through the nine-point calibration method, and converting the pixel coordinates of the identified target object into three-dimensional grasping coordinates in the mechanical arm base coordinate system; Strategy planning step: based on the current working area state, using a depth-first search algorithm to generate an optimal task sequence of the target object; Trajectory planning step: according to the optimal task sequence, the motion trajectory of the mechanical arm is planned, and the path is optimized by adopting circular arc transition; Motion control step: control the mechanical arm to perform the grabbing and placing operation according to the motion trajectory, and perform real-time vacuum degree detection and pose calibration in the process.

[0009] The working principle and advantages of the present application are: The mechanical arm grabbing control system and method based on visual perception can realize high-precision and high-efficiency identification, grabbing and accurate placing of multi-category target objects in a complex environment. The focus is on: The present application constructs a highly coordinated visual-planning-control closed-loop system, rather than simply connecting independent modules in series, which can improve the precision, efficiency and robustness of the mechanical arm operation in a complex unstructured environment.

[0010] Among them, in the visual perception aspect, the present application adopts the strategy of combining HSV color space segmentation and k-means clustering algorithm, which not only effectively overcomes the sensitivity of traditional threshold segmentation to light changes, but also realizes excellent robustness to low contrast, reflection and different color backgrounds through the data-driven adaptive clustering process, ensuring that the object contour with complete topological structure is extracted from the image. In the aspect of coordinate mapping, a nine-point calibration method is adopted in stages to establish two independent affine transformation relationships between the camera-mechanical arm and the working area-mechanical arm, respectively, separating the coordinate conversion problems of grabbing and placing two links, and avoiding the inherent defect of single mapping model that the error increases in the edge area of the working space.

[0011] More importantly, the present application takes the "color-shape-pose" collaborative recognition as the core, seamlessly integrates color recognition, shape classification and pose calculation in the same processing pipeline, accurately maps two-dimensional pixel coordinates to three-dimensional working space of the mechanical arm through affine transformation and dynamic height compensation model, and the pose estimation accuracy can reach sub-millimeter level to meet the requirements of precision assembly, which can lay a reliable data foundation for subsequent decision and execution. In the strategy planning layer, the lower supporting rules are converted into a physical constraint and embedded into the depth-first search algorithm, so that the generated placing sequence can meet the physical requirements of stable placing of the mechanical arm. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 It is a system hardware device structure schematic diagram of the embodiment of the mechanical arm grabbing control system and method based on visual perception of the present application; Figure 2 It is a system structure schematic diagram of the embodiment of the mechanical arm grabbing control system and method based on visual perception of the present application; Figure 3A schematic diagram of a visual calibration process for an embodiment of a robotic arm grasping control system and method based on visual perception according to the present invention; Figure 4 A schematic diagram of the visual processing and block recognition flow of a robotic arm grasping control system and method embodiment based on visual perception of the present invention; Figure 5 This is a schematic diagram of the structure of a robot arm trajectory planning process according to an embodiment of a robot arm grasping control system and method based on visual perception of the present invention; Figure 6 The present invention is a schematic diagram of the method operation flow of a robotic arm grasping control system and method embodiment based on visual perception. DETAILED DESCRIPTION

[0013] The following is a further detailed description through specific implementation methods: The embodiment is basically as shown in the attached Figure 2 Shown: A robotic arm grasping control system based on visual perception, including a visual perception module, a strategy planning module, a motion control module, a safety protection module and a visualization module.

[0014] The visual perception module is used to collect image data of the working area and identify the color, shape and posture information of the target object through image processing as the recognition result.

[0015] The visual perception module also includes a data processing unit for smoothing the recognition results through Kalman filtering.

[0016] The visual perception module uses a nine-point calibration method to establish a mapping relationship between the image coordinate system and the robotic arm coordinate system, and realizes the recognition and positioning of the target object through HSV color space segmentation and k-means clustering algorithm.

[0017] In this embodiment, Figure 1 As shown, the robotic arm uses an xArm. A camera and suction cup tool are integrated at the end of the arm. The visual perception module uses the camera to collect image data of the work area. The suction cup tool is used to pick up target objects, and the suction force is controlled by an air valve. Each system module can be executed in software form on a computer, and the motion control module establishes a communication connection with the xArm controller of the robotic arm.

[0018] Specifically, if Figure 3 As shown, the nine-point calibration method includes the following steps: Step 1, within the workspace of the robot arm, select a region as uniform as possible within the target working plane (i.e. the working area plane) and arrange 9 calibration points. In this embodiment, physical markers with high contrast (e.g. black and white concentric circles with a diameter of 8 mm) are used as calibration points to ensure that they are stably and accurately positioned in the image by a sub-pixel edge detection algorithm.

[0019] Step 2, control the end of the robot arm to move vertically above each calibration point to ensure that the physical position of the end effector of the robot arm is aligned with the center of the marker point in three-dimensional space. At each point, two data are recorded synchronously and accurately: one is the three-dimensional coordinates in the base coordinate system of the robot arm controller feedback, i.e. the robot arm coordinates (X, Y, Z), and the other is the pixel coordinates of the center of the marker point extracted from the camera image, i.e. the image pixel coordinates (U, V). To reduce random errors, this process is repeated 3 times for each point and the average value is taken.

[0020] Step 3, define a coordinate system on the working area plane and mark 9 feature points, and then move the end of the robot arm to the feature points one by one and record the coordinates to achieve the coordinate conversion from the working plane to the robot arm.

[0021] Specifically, in this embodiment, two sets of core parameters are used to achieve coordinate conversion. The first set is the conversion parameters from the image pixel coordinate system to the robot arm coordinate system, including the element values of the rotation matrix and the element values of the translation vector , where the conversion relationship between the camera and the robot arm is: ; These parameters are used to convert the recognized block pixel position to the robot arm grasping coordinates.

[0022] The second set is the conversion parameters from the working area coordinate system to the robot arm coordinate system, including the element values of the rotation matrix and the element values of the translation vector , where the conversion relationship between the robot arm and the working area coordinates is: ; These parameters are used to convert the grid position generated by the strategy to the robot arm placement coordinates.

[0023] In addition, the calibration environment requires that the light intensity be stable at 500 ± 100 lux, the camera parameters be fixed at aperture f / 4 and exposure time 15 ms, and the calibration markers be black and white concentric circles with a diameter of 8 mm to ensure that the contrast exceeds 80%.

[0024] The calibration is based on a two-dimensional affine transformation mathematical model, the core of which is to solve the optimal solution of the rotation matrix R and the translation vector M, and the mathematical expression is that the mechanical arm coordinates are equal to the rotation matrix multiplied by the pixel coordinates plus the translation vector. The principle is realized by constructing an overdetermined equation set, substitiating 9 groups of corresponding coordinates into the transformation formula to form 18 equations containing 6 unknown parameters, and then solving the normal equation by the least square method: ; That is, the parameter vector is equal to the design matrix transpose multiplied by the inverse matrix of the design matrix, and then multiplied by the design matrix transpose and the observation vector, so as to obtain the optimal estimation of the transformation matrix. This principle effectively overcomes three major error sources, the camera lens radial distortion is controlled within 3%, the mechanical arm repeatability error is limited to ±0.3mm, and the plane installation inclination is less than 2°, so as to finally realize the sub-millimeter level positioning accuracy of the center area position error ≤0.8mm and the edge area ≤1.5mm.

[0025] As shown in Figure 4 , the HSV color space segmentation specifically includes the following contents: First, the RGB image collected by the camera is converted to the HSV color space. The HSV (hue, saturation, value) space separates the color information (H, S) from the brightness information (V), and this feature makes it more robust to changes in lighting (mainly represented by changes in the V component) than the RGB space. After conversion, a Gaussian filtering operation is also performed to suppress image noise.

[0026] According to the known color (such as red, green, blue) of the target object, the threshold range of its corresponding hue (H), saturation (S) and value (V) in the HSV space is set, for example, the green channel is set to H∈[60, 88], S∈[100, 255], V∈[120, 255]; the yellow channel is set to H∈[20, 30], S∈[80, 255], V∈[150, 255]; the blue channel is set to H∈[92, 105], S∈[160, 255], V∈[120, 255]; and the red channel uses a double-interval strategy to capture the H∈[0, 18] and [160, 180] ranges to solve the hue ring breakage problem.

[0027] After segmentation, differential morphological processing is implemented, the green channel uses a 5×5 kernel to perform three times of expansion to connect the broken areas and twice of corrosion to smooth the boundaries; the blue channel uses a 3×3 kernel for single corrosion to eliminate noise points; and the purple channel implements a corrosion-expansion combination operation of a 3×3 kernel to enhance the target integrity. The processing process is realized by inRange and morphologyEx functions, and each channel is independently and parallelly calculated.

[0028] The k-means clustering algorithm specifically includes the following contents: Firstly, the BGR values of all pixels in the image are extracted as an N x 3 matrix, and each pixel is regarded as a three-dimensional vector. Secondly, after randomly initializing two cluster centers, an iterative optimization loop is entered, the Euclidean distance of each pixel to the center is calculated, the color difference in three-dimensional space is measured, and it is classified into the nearest center; then, the mean vector of each cluster is recalculated as the new center.

[0029] The iteration continues until the triple convergence condition is met: no vector is reassigned, the center displacement is less than the threshold of 0.1, and the total distance of all objects to the center to which they belong is minimized; then, the foreground cluster pixels are set to white (255) and the background cluster is set to black (0) to generate a binary image; finally, a 5 x 5 closing operation is performed to fill the internal holes, a 3 x 3 opening operation is performed to eliminate discrete noise points, and a boundary smoothing process is performed to form a complete and connected target shape.

[0030] The entire process breaks through the limitations of traditional threshold segmentation through data-driven adaptive segmentation, improves the processing efficiency by 40%, eliminates light interference, and realizes one-time complete extraction of all target object shapes.

[0031] The visual perception module is also used to perform square type recognition, including a feature matching method based on aspect ratio and a template matching method based on grid binary coding; The feature matching method based on aspect ratio includes the following: For targets with an aspect ratio greater than 3:1, directly determine as long strip-shaped squares; For targets with an aspect ratio between 0.9 and 1.1, directly determine as square squares; The template matching method based on grid binary coding includes the following: For other shaped targets, after normalization to a standard size, divide them into 2x3 grids, calculate the proportion of white pixels in each grid, and generate a 6-bit binary feature code; By calculating the Hamming distance between the feature code and the pre-stored standard template library, accurate matching of the shape type is realized, and the matching threshold is Hamming distance ≤ 1.

[0032] Through the above settings, the system can realize industrial-level performance indicators without relying on expensive high-precision sensors or a large amount of labeled data. In the pose type recognition, by combining the feature matching based on aspect ratio and the binary coding template matching, the rotation state of the target object (such as a square) can be stably calculated using only a monocular camera, avoiding the complex and time-consuming three-dimensional point cloud registration process.

[0033] The visual perception module also includes a coordinate conversion unit for converting image coordinates to three-dimensional coordinates in the mechanical arm coordinate system through a pre-calibrated affine transformation matrix. Specifically, the following steps are included: First, the target center point pixel coordinates are obtained The mechanical arm X / Y coordinates are calculated through a pre-calibrated 2x3 transformation matrix, and the specific formula is: ; ; The matrix parameters are obtained by a nine-point calibration method, and nine calibration points are uniformly arranged on the workbench surface, and are fitted by the least squares method; Second, the height dimension adopts a dynamic compensation model: ; The linear coefficient is obtained by regression analysis of 25-point plane measurement data, and -2mm is a safety margin compensation.

[0034] For attitude calculation, first obtain the original angle θ from the minimum circumscribed rectangle, and perform 90° compensation when the rectangle height is greater than the width: ; Then superimpose the rotation state parameter s output by shape recognition (0 represents the standard state and 1 represents 180° rotation); The final attitude angle calculation formula is: ; Optimization processing for special scenarios includes: adding an X-axis 1mm surface characteristic compensation for green squares; when trigger the edge protection mechanism, calculate ; when the robot is operating, apply rotation to make the square return to the predefined standard orientation.

[0035] The entire calculation process introduces Kalman filter smoothing processing to effectively suppress measurement fluctuations caused by camera shaking.

[0036] Preferably, the visual perception module is provided with an adaptive image processing unit, a multi-scale shadow detection unit and a morphological processing unit.

[0037] The adaptive image processing unit is used to automatically correct color deviation caused by environmental light changes using the gray world algorithm, and to enhance image contrast using limited contrast adaptive histogram equalization; the multi-scale shadow detection unit is used to distinguish real object contours from shadow interference by calculating the gradient amplitude of the image and combining color invariance features; the morphological processing unit is used to dynamically select the size of the structure element according to the edge sharpness information of the image.

[0038] The strategy planning module is used to model the working area as a two-dimensional grid map, and according to the recognition results and the current working area state, generate an optimal task sequence for the target object containing the grabbing sequence, placement coordinates and placement attitude.

[0039] The strategy planning module adopts a depth-first search algorithm to traverse all possible placement positions and rotation states of the current target object, and combines a heuristic evaluation function to optimize the search priority.

[0040] The motion control module is configured to plan a motion trajectory of the robotic arm according to the optimal task sequence, and control the robotic arm to perform the picking, transferring and placing operations; the motion trajectory includes a pick safety point, a pick point, a place safety point and a place point.

[0041] The motion control module adopts a segmented trajectory planning, including a multi-segment path of a pick safety height, a pick height, a place safety height and a place height, and optimizes the path smoothness through a circular arc transition.

[0042] In this embodiment, the motion control relies on the moveit! framework of ROS, and the robotic arm execution process strictly follows a five-segment trajectory sequence, which is moving to a pick safety height, descending to a pick height, opening a valve, lifting to a safety height, moving to a place safety height, descending to a place height, closing the valve, and returning to the safety height, as shown in FIG. 8. Figure 5 Each action node is sent to the robotic arm controller through a service call, and a fixed delay of 300ms or 500ms is inserted between key steps to ensure that the action is executed. In terms of parameter configuration, the acceleration of all linear motions is uniformly set to , the speed is 600mm / s, and the transition radius is 20mm. The valve control is realized through a digital IO service, and the port is set to 1 to open the suction, and to 0 to release the block.

[0043] In addition, the motion control module subscribes to the robotic arm pose topic in real time, continuously monitors the XYZ axis position deviation, and before executing the key action, the difference between the target pose and the actual pose is detected in a loop. When the sum of the absolute differences of the three axes between the target pose and the actual pose is not greater than 0.1mm, it is determined to be in place, and the detection frequency reaches 10kHz. At the same time, the action synchronization is controlled by setting the algorithm parameters, the key node is set to true to ensure sequential execution, and the non-continuous action is set to false to improve efficiency.

[0044] In addition, a protection mechanism is established in the motion control module to ensure operation safety, including performing a zero reset operation at startup to ensure accurate reference position, reserving a safety margin in height calculation, lowering by 2mm during picking and lifting by 3mm during placing, actively avoiding singular point angles, setting a detection upper limit to prevent dead loops, and clearing error states through ROS services in emergency situations.

[0045] The safety protection module includes a self-collision detection unit, a safety boundary limiting unit, a vacuum degree detection unit and an emergency stop unit. The self-collision detection unit is used to detect the collision risk between the mechanical arm links and the environment in real time during motion planning. The safety boundary limiting unit is used to define the motion soft limit of each joint of the mechanical arm and the cubic safety region in the Cartesian space. The vacuum degree detection unit is used to monitor the grabbing state of the mechanical arm through a vacuum pressure sensor and trigger an automatic retry process when the suction fails. The emergency stop unit is used to provide a hardware emergency stop signal interface with the highest priority to directly cut off the driving power supply of the mechanical arm.

[0046] The visualization module is used to visually display the recognition results of the visual perception module.

[0047] The embodiment also provides a mechanical arm grabbing control method based on visual perception, which applies the mechanical arm grabbing control system based on visual perception to control the motion of the mechanical arm, and includes the following steps: An image acquisition step: acquiring a work area image containing a target object through the visual perception module; A visual recognition step: pre-processing the work area image, converting to an HSV color space and performing color segmentation, generating a binary image by using a k-means clustering algorithm, and then recognizing the color, shape and pixel coordinates of the target object; A coordinate conversion step: converting the pixel coordinates of the recognized target object into three-dimensional grabbing coordinates in the base coordinate system of the mechanical arm through an affine transformation matrix obtained by nine-point calibration; A strategy planning step: generating an optimal task sequence of the target object based on the current work area state by using a depth-first search algorithm; A trajectory planning step: planning the motion trajectory of the mechanical arm according to the optimal task sequence, and performing path optimization by using circular arc transition; A motion control step: controlling the mechanical arm to perform grabbing and placing operations according to the motion trajectory, and performing vacuum degree detection and pose calibration in real time during the process.

[0048] As Figure 6 shown, the mechanical arm grabbing control system and method based on visual perception can realize high-precision and high-efficiency recognition, grabbing and precise placing of multiple types of target objects in a complex environment.

[0049] The system can realize rapid fusion and robust recognition of multi-modal features (such as color and shape) under complex lighting and background interference, and simultaneously output high-precision spatial poses to meet the operation requirements of precise assembly and splicing. Moreover, the system can dynamically adjust the grabbing point, placing pose and motion trajectory based on real-time visual feedback, so as to adapt to the high uncertainty of objects in a logistics sorting scene.

[0050] The above is only an embodiment of the present application, and the common knowledge of the specific structure and characteristics in the scheme is not described in detail here. The ordinary skilled person in the art knows all the ordinary technical knowledge in the field of the application before the application date or the priority date, can know all the prior art in the field, and has the ability to apply conventional experimental means before that date. The ordinary skilled person in the art can perfect and implement the present scheme based on the disclosure given in the present application and in combination with their own ability. Some typical known structures or known methods should not be an obstacle for the ordinary skilled person in the art to implement the present application. It should be noted that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the present application. These will not affect the effect and practicality of the patent.

Claims

1. A robotic arm grasping control system based on visual perception, characterized in that: Includes visual perception module, strategy planning module and motion control module; The visual perception module is used to collect image data of the working area and identify the color, shape and posture information of the target object through image processing as the recognition result; The strategy planning module is used to model the work area as a two-dimensional grid map and generate an optimal task sequence for the target object, including a grasping order, placement coordinates, and placement posture, based on the recognition results and the current state of the work area; The motion control module is used to plan the motion trajectory of the robot arm according to the optimal task sequence and control the robot arm to perform grasping, transfer and placement operations; the motion trajectory includes a pick-up safety point, a pick-up point, a placement safety point and a placement point; Among them, the visual perception module uses the nine-point calibration method to establish the mapping relationship between the image coordinate system and the robotic arm coordinate system, and realizes the recognition and positioning of the target object through HSV color space segmentation and k-means clustering algorithm.

2. A robotic arm grasping control system based on visual perception according to claim 1, characterized in that: The visual perception module is provided with an adaptive image processing unit, a multi-scale shadow detection unit and a morphological processing unit; The adaptive image processing unit is used to correct color casts caused by changes in ambient lighting and enhance image contrast using contrast-limited adaptive histogram equalization. The multi-scale shadow detection unit is used to distinguish between real object contours and shadow interference by calculating the gradient amplitude of the image and combining color invariance features. The morphological processing unit is used to dynamically select the size of the structural element based on the edge sharpness information of the image.

3. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: The strategy planning module adopts a depth-first search algorithm to traverse all possible placement positions and rotation states of the current target object, and optimizes the search priority in combination with a heuristic evaluation function.

4. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: The motion control module adopts segmented trajectory planning, including multiple paths of grabbing safety height, grabbing height, placing safety height and placing height, and optimizes path smoothness through arc transition.

5. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: The visual perception module also includes a coordinate conversion unit for converting image coordinates into three-dimensional coordinates in a robotic arm coordinate system through a pre-calibrated affine transformation matrix.

6. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: The visual perception module is also used to perform block type recognition, including a feature matching method based on aspect ratio and a template matching method based on grid binary coding; The feature matching method based on aspect ratio includes the following contents: For targets with an aspect ratio greater than 3:1, they are directly judged as long rectangles; For targets with aspect ratios between 0.9 and 1.1, they are directly judged as squares; The template matching method based on grid binary coding includes the following contents: For targets of other shapes, they are normalized to a standard size and divided into 2x3 grids. The proportion of white pixels in each grid is calculated to generate a 6-bit binary feature code. By calculating the Hamming distance between the feature code and a pre-stored standard template library, accurate matching of shape types is achieved, and the matching threshold is a Hamming distance ≤ 1.

7. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: The visual perception module also includes a data processing unit for smoothing the recognition results through Kalman filtering.

8. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: It also includes a safety protection module, including a self-collision detection unit, a safety boundary limitation unit, a vacuum detection unit and an emergency stop unit; The self-collision detection unit is used to detect the collision risk between the robot arm links and the environment in real time during motion planning; The safety boundary limitation unit is used to define the soft limit of motion of each joint of the robot arm and the cube safety area in Cartesian space; The vacuum detection unit is used to monitor the gripping state of the robot arm through a vacuum pressure sensor and trigger an automatic retry process when the gripping fails; The emergency stop unit is used to provide a hardware emergency stop signal interface with the highest priority to directly cut off the driving power supply of the robotic arm.

9. The visual perception-based robotic arm grasping control system according to claim 1, characterized in that: It also includes a visualization module; the visualization module is used to visually display the recognition results of the visual perception module.

10. A robotic arm grasping control method based on visual perception, characterized in that: Applying a visual perception-based robotic arm grasping control system as described in any one of claims 1 to 9 to control the movement of a robotic arm comprises the following steps: Image acquisition step: using the visual perception module to acquire an image of the working area containing the target object; Visual recognition step: pre-processing the working area image, converting it into HSV color space and performing color segmentation, generating a binary image using the k-means clustering algorithm, and then identifying the color, shape and pixel coordinates of the target object; Coordinate transformation step: The affine transformation matrix obtained by the nine-point calibration method is used to convert the pixel coordinates of the identified target object into the three-dimensional grasping coordinates in the robot arm base coordinate system; Strategy planning step: Based on the current state of the work area, a depth-first search algorithm is used to generate the optimal task sequence for the target object; Trajectory planning step: planning the motion trajectory of the robot arm according to the optimal task sequence, and optimizing the path using arc transition; Motion control step: Control the robotic arm to perform grabbing and placing operations according to the motion trajectory, and perform vacuum detection and posture calibration in real time during the process.

Citation Information

Cited By

  • Panoramic vision and dynamic vision combined mobile manipulator system and guide grabbing method

    CN121608139A

  • Detection device and detection method

    CN121656138A