A Vision-Based Intelligent Positioning and Grasping Method and System for Industrial Sorting Robots

By integrating the spatial geometric and semantic features of multi-view images, the problem of unstable grasping in existing technologies has been solved, enabling reliable grasping of fragile and deformable objects and improving the intelligence level of sorting robots.

CN122125688APending Publication Date: 2026-06-02BEIJING BOYAN SHENGKE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BOYAN SHENGKE TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing industrial sorting robot positioning and grasping methods are unstable when dealing with fragile, deformable, or smooth objects, which can easily lead to objects slipping or being damaged. Furthermore, there is a lack of assessment of the physical feasibility and stability of the grasping process.

Method used

By acquiring multi-view images of the objects to be sorted, spatial geometric features and semantic features are fused to establish spatial correspondence, geometric transformation parameters are calculated and unified features are reconstructed, candidate grasping positions are determined, the spatial distribution and contact area of ​​the contact region are predicted, mechanical stability values ​​are calculated, feasible grasping points are selected, and grasping motion trajectories are planned.

Benefits of technology

It improved the sorting success rate, enabled reliable and adaptive grasping of flexible and irregularly shaped objects, enhanced the system's intelligence level, and ensured the physical stability and task adaptability of the grasping process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122125688A_ABST
    Figure CN122125688A_ABST
Patent Text Reader

Abstract

This invention provides a vision-based intelligent positioning and grasping method and system for industrial sorting robots, relating to the field of industrial robot technology. The method includes: acquiring multi-view images and fusing spatial geometric and semantic features; obtaining unified features through coordinate alignment; establishing a coordinate system mapping relationship to achieve precise object positioning and attitude estimation; predicting contact force distribution based on geometric features and constraints; calculating stability and adaptability values ​​and selecting the optimal grasping point based on reachability verification; and planning a motion trajectory to complete the grasping process. This invention can improve the accuracy of sorting robots in recognizing and positioning complex objects and enhance their grasping stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial robot technology, and in particular to an intelligent positioning and grasping method and system for industrial sorting robots based on vision recognition. Background Technology

[0002] In the field of industrial automation, vision-based sorting robots are key equipment for realizing flexible manufacturing and intelligent logistics. Existing conventional industrial sorting robot positioning and grasping methods usually rely on visual perception and pose estimation of target objects. Existing methods mostly rely on simple geometric heuristic rules, such as selecting the centroid, convex hull center of the object, or specifying the approach point and direction of the gripper according to a preset grasping template. They rarely consider the physical interaction characteristics and stability quantitative evaluation during the grasping process.

[0003] Existing technologies still suffer from incomplete and ambiguous feature extraction due to object occlusion, changes in ambient lighting, or arbitrary placement, which severely reduces the accuracy of subsequent pose estimation. Furthermore, they lack the ability to finely model the physical feasibility and stability of the grasping action itself, and cannot fully consider the real mechanical distribution, friction conditions, and stability of the grasping force closure area of ​​the contact area between the gripper and the object surface. This leads to problems such as unstable grasping, object slippage, or even damage to the workpiece when dealing with fragile, deformable, or smooth objects. Summary of the Invention

[0004] This invention provides a method and system for intelligent positioning and grasping of industrial sorting robots based on visual recognition, which can at least solve some of the problems existing in the prior art.

[0005] A first aspect of this invention provides a method for intelligent positioning and grasping of an industrial sorting robot based on visual recognition, comprising:

[0006] Multi-view images of the object to be sorted are acquired and the corresponding spatial geometric features and semantic features are extracted and fused to obtain fused features. Surface key points of the object to be sorted under different views are extracted and spatial correspondences are established. Geometric transformation parameters are calculated based on the spatial correspondences, and coordinate alignment and feature reconstruction are performed on the fused features to obtain unified features.

[0007] Based on the unified features, a transformation relationship between the image coordinate system and the robot coordinate system is established, and the coordinates of the object to be sorted are mapped from the image coordinate system to the robot coordinate system to obtain the positioning coordinates. The posture parameters corresponding to the object to be sorted are calculated by combining the unified features.

[0008] Based on the spatial geometric features, candidate gripping positions on the surface of the object to be sorted are determined. Based on the candidate gripping positions and preset geometric constraints, the spatial distribution and contact area of ​​the contact region are predicted, and the composite distribution of the contact force on the surface of the object to be sorted is determined. Based on the composite distribution, the mechanical stability value is calculated and combined with the semantic features to obtain the comprehensive adaptability value.

[0009] Based on the positioning coordinates and the attitude parameters, the reachability of the candidate grasping positions is verified, feasible candidate grasping positions are determined, and the candidate grasping position with the largest comprehensive adaptability value is selected as the target grasping point. Based on the target grasping point, the grasping direction vector is determined, and the motion trajectory is calculated in combination with the attitude parameters to complete the grasping.

[0010] In one alternative implementation,

[0011] The process involves acquiring multi-view images of the objects to be sorted, extracting corresponding spatial geometric features, and fusing them with semantic features to obtain fused features. It also involves extracting key surface points of the objects from different viewpoints and establishing spatial correspondences. Based on these spatial correspondences, geometric transformation parameters are calculated, and the fused features are aligned and reconstructed to obtain unified features, including:

[0012] Multi-view images are obtained by acquiring multiple perspectives of the objects to be sorted. Depth is estimated from the multi-view images to obtain depth maps, and the corresponding 3D point cloud data is extracted as spatial geometric features. Semantic features are obtained by feature extraction from the multi-view images. The spatial geometric features and the semantic features are spliced ​​together according to spatial position to obtain fused features.

[0013] Key point detection is performed on the surface of the object to be sorted in the multi-view image to obtain a set of surface key points. The similarity of the feature descriptors corresponding to the surface key points under different views is calculated, and feature matching is performed based on the similarity to obtain key point matching pairs. Spatial correspondence is established based on the image coordinates and depth information of different key points in the key point matching pairs under different views.

[0014] Based on the 3D coordinate differences of the key point matching pairs in the spatial correspondence, the rotation matrix and translation vector are calculated as geometric transformation parameters. Based on the geometric transformation parameters, the spatial geometric features in the fused features are transformed to a unified reference coordinate system to complete coordinate alignment. The spatial geometric features after coordinate alignment are re-associated with the corresponding semantic features in spatial position, and the spatial index relationship of the feature vector is updated to obtain unified features.

[0015] In one alternative implementation,

[0016] Based on the unified features, a transformation relationship between the image coordinate system and the robot coordinate system is established, and the coordinates of the object to be sorted are mapped from the image coordinate system to the robot coordinate system to obtain the positioning coordinates. Combined with the unified features, the posture parameters corresponding to the object to be sorted are calculated, including:

[0017] Extract three-dimensional feature points from the surface of the object to be sorted from the unified features and obtain the coordinates of the three-dimensional feature points in the image coordinate system as the original coordinates. Set a calibration reference point in the robot coordinate system. Determine the origin position and coordinate axis direction of the robot coordinate system based on the three-dimensional coordinates of the calibration reference point. Solve the correspondence between the original coordinates and the coordinates in the robot coordinate system to determine the rotation transformation matrix and translation transformation vector. Construct a coordinate transformation matrix based on the rotation transformation matrix and the translation transformation vector as the transformation relationship.

[0018] Substitute the coordinates of the centroid of the object to be sorted in the image coordinate system into the coordinate transformation matrix to perform matrix operations and obtain the positioning coordinates;

[0019] Extract the normal vector distribution of the surface of the object to be sorted from the unified features and calculate the principal direction vector of the object surface. Substitute the direction component of the principal direction vector in the image coordinate system into the rotation transformation matrix to perform coordinate system transformation to obtain the direction component of the principal direction vector in the robot coordinate system. Calculate the rotation angle of the object to be sorted relative to each coordinate axis of the robot coordinate system based on the direction component and combine them to form the attitude parameters.

[0020] In one alternative implementation,

[0021] Based on the spatial geometric features, candidate gripping positions on the surface of the object to be sorted are determined. Based on the candidate gripping positions and preset geometric constraints, the spatial distribution and contact area of ​​the contact region are predicted, and the composite distribution of contact forces on the surface of the object to be sorted is determined, including:

[0022] The curvature distribution of the surface of the object to be sorted is extracted from the spatial geometric features in the unified features. Based on the curvature distribution, flat areas and raised areas on the surface of the object to be sorted are identified. The surface normal vectors of the flat areas and the raised areas are calculated and the surface areas that meet the grasping stability conditions are selected as candidate grasping positions based on the angle between the surface normal vectors and the horizontal plane.

[0023] Obtain the contact surface geometry of the end effector under preset geometric constraints, project the contact surface geometry onto the surface area corresponding to each candidate gripping position, calculate the degree of contact between the contact surface and the surface of the object to be sorted based on the curvature distribution of the surface area corresponding to the contact surface geometry, determine the boundary contour of the actual contact area as the spatial distribution of the contact area based on the degree of contact, and perform area integration to obtain the contact area.

[0024] The gripping force applied by the end effector under the preset geometric constraints is decomposed into local forces at each contact point within the contact area according to the contact area. Based on the local forces and the corresponding surface normal vectors, the normal and tangential components of the local forces on the surface of the object to be sorted are calculated. The normal and tangential components of all contact points within the contact area are vector-superimposed to obtain the composite distribution.

[0025] In one alternative implementation,

[0026] The mechanical stability value is calculated based on the synthetic distribution, and the comprehensive fitness value is obtained by combining the semantic features, including:

[0027] The spatial distribution of contact forces on the surface of the object to be sorted is obtained from the synthetic distribution, and the force field is reconstructed to obtain the force distribution field. Based on the force distribution field, the stress tensor distribution on the surface of the object to be sorted is calculated, and the resultant force and the point of application of the resultant force in the contact area are obtained by integration. The position vector of the point of application of the resultant force relative to the center of mass of the object to be sorted is calculated, and the resultant moment of the object to be sorted relative to the center of mass is obtained by combining the resultant force. The gravitational moment is obtained based on the mass and center of mass of the object to be sorted. The vector sum of the resultant moment and the gravitational moment is calculated, and it is determined whether the vector sum satisfies the moment equilibrium condition. If so, the vector sum is used as the mechanical stability value.

[0028] Based on the semantic features in the unified features, extract the material properties and surface texture features of the object to be sorted. Based on the material properties, determine the elastic modulus and Poisson's ratio of the object to be sorted and calculate the surface deformation of the object under contact force. Based on the surface texture features, calculate the micro-protrusion distribution on the surface of the object to be sorted and determine the ratio of the actual contact area to the nominal contact area and record it as the first coefficient.

[0029] The deformation adaptability value is obtained by correlating the mechanical stability value and the surface deformation value, and the contact effectiveness value is obtained by correlating the mechanical stability value and the first coefficient. The comprehensive adaptability value is obtained by weighted fusion of the deformation adaptability value and the contact effectiveness value.

[0030] In one alternative implementation,

[0031] Based on the positioning coordinates and the attitude parameters, the reachability of candidate grasping positions is verified to determine feasible candidate grasping positions and select the candidate grasping position with the largest comprehensive adaptability value as the target grasping point, including:

[0032] Obtain the link length parameters and joint rotation axis direction of the robot end effector and establish a mathematical mapping relationship between the robot joint angle and the spatial position of the end effector. Based on the mathematical mapping relationship, derive the inverse solution expression of the spatial position and attitude of the end effector to the joint angle.

[0033] For each candidate gripping position, the positioning coordinates and attitude parameters corresponding to the current candidate gripping position are substituted into the inverse solving expression to calculate the target joint angle of each joint of the robot. The angle limit range of each joint of the robot is obtained and it is determined whether the target joint angle exceeds the angle limit range to obtain the first judgment result. The reachable distance range of the robot end effector is calculated based on the link length parameter and it is determined whether the positioning coordinates exceed the reachable distance range to obtain the second judgment result. Based on the first judgment result and the second judgment result, a feasible candidate gripping position is determined.

[0034] Extract the comprehensive adaptability value corresponding to the feasible candidate crawling position, sort the feasible candidate crawling positions in descending order based on the comprehensive adaptability value to obtain a priority sequence, and take the first feasible candidate crawling position in the priority sequence as the target crawling point.

[0035] In one alternative implementation,

[0036] Determining the grasping direction vector based on the target grasping point, calculating the motion trajectory based on the posture parameters, and completing the grasping process includes:

[0037] The system acquires the position and attitude parameters of the target gripping point. Based on the attitude parameters, it determines the normal approach angle and tangential deflection angle of the robot end effector relative to the surface of the object to be sorted. Based on the normal approach angle, it calculates the unit vector along the normal of the surface of the object to be sorted. Based on the tangential deflection angle, it calculates the rotation correction vector in the tangential plane of the surface of the object to be sorted. Based on the unit vector and the rotation correction vector, it determines the initial gripping vector and, combined with the synthetic distribution, determines the cosine value of the angle between the direction of the contact force and the initial gripping vector. Based on the cosine value of the angle, it performs mechanical optimization adjustment on the initial gripping vector to obtain the gripping direction vector.

[0038] Obtain the current position coordinates and current attitude angle of the robot's end effector, and calculate the spatial displacement vector of the position parameters from the current position coordinates to the target grasping point and the angular change of the attitude parameters from the current attitude angle to the target grasping point;

[0039] A pose transformation matrix is ​​constructed based on the spatial displacement vector and the grasping direction vector, and singular value decomposition is performed to obtain translation and rotation components. A translation path is generated based on the translation components, and the path is corrected along the grasping direction vector to obtain a corrected translation path. A rotation path is generated based on the rotation components and the angle change, ensuring the orthogonality between the rotation axis and the grasping direction vector. The corrected translation path and the rotation path are interpolated by a fifth-order polynomial to obtain the motion trajectory, and the object to be sorted is moved along the motion trajectory to the target grasping point to complete the grasping of the object to be sorted.

[0040] A second aspect of this invention provides a vision-based intelligent positioning and grasping system for industrial sorting robots, comprising:

[0041] The feature fusion unit is used to acquire multi-view images of the object to be sorted and extract the corresponding spatial geometric features and semantic features to fuse them into fused features. It also extracts the surface key points of the object to be sorted under different views and establishes spatial correspondence. Based on the spatial correspondence, it calculates geometric transformation parameters and performs coordinate alignment and feature reconstruction on the fused features to obtain unified features.

[0042] The coordinate transformation unit is used to establish the transformation relationship between the image coordinate system and the robot coordinate system based on the unified feature, and to map the coordinates of the object to be sorted from the image coordinate system to the robot coordinate system to obtain the positioning coordinates and calculate the posture parameters corresponding to the object to be sorted in combination with the unified feature.

[0043] The grasping evaluation unit is used to determine candidate grasping positions on the surface of the object to be sorted based on the spatial geometric features, predict the spatial distribution and contact area of ​​the contact region based on the candidate grasping positions and preset geometric constraints, determine the composite distribution of contact forces on the surface of the object to be sorted, calculate the mechanical stability value based on the composite distribution, and solve for the comprehensive adaptability value by combining the semantic features.

[0044] The grasping execution unit is used to verify the reachability of candidate grasping positions based on the positioning coordinates and the attitude parameters, determine feasible candidate grasping positions and select the candidate grasping position with the largest comprehensive adaptability value as the target grasping point, determine the grasping direction vector based on the target grasping point, calculate the motion trajectory in combination with the attitude parameters and complete the grasping.

[0045] A third aspect of the present invention provides an electronic device, comprising:

[0046] A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0047] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0048] This invention integrates spatial geometric and semantic features from multi-view images, performs coordinate alignment and feature reconstruction based on the spatial correspondence of key points, and generates unified features. This effectively overcomes interference caused by incomplete information from a single viewpoint, partial object occlusion, or changes in lighting, ensuring the integrity and consistency of object representation and laying a reliable foundation for subsequent precise positioning. The image-robot coordinate system transformation relationship established based on the unified features enables high-precision mapping of object coordinates from image space to robot operating space. Combined with the unified features, accurate posture parameters are calculated, solving the problem of grasping failure caused by coordinate transformation errors or inaccurate posture estimation. This allows the robot to accurately understand the position and orientation of the object in three-dimensional space. Candidate grasping positions are determined through spatial geometric features, and geometric constraints are introduced to predict the composite distribution of contact area and contact force, thereby calculating mechanical stability values. This achieves a quantitative assessment of the physical stability and task adaptability of the grasping point. After determining the positioning coordinates and posture parameters, the reachability of candidate grasping positions is verified, and feasible grasping points are selected. This significantly improves the sorting success rate and the overall intelligence level of the system, making it suitable for reliable and adaptive grasping operations on flexible and irregularly shaped objects. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating the intelligent positioning and grasping method for industrial sorting robots based on visual recognition, as described in an embodiment of the present invention.

[0050] Figure 2 This is a flowchart illustrating the object grasping mechanical adaptability evaluation of the vision recognition-based intelligent positioning and grasping method for industrial sorting robots according to an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0053] Figure 1 This is a flowchart illustrating the intelligent positioning and grasping method for industrial sorting robots based on visual recognition, as described in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0054] Multi-view images of the object to be sorted are acquired and the corresponding spatial geometric features and semantic features are extracted and fused to obtain fused features. Surface key points of the object to be sorted under different views are extracted and spatial correspondences are established. Geometric transformation parameters are calculated based on the spatial correspondences, and coordinate alignment and feature reconstruction are performed on the fused features to obtain unified features.

[0055] Based on the unified features, a transformation relationship between the image coordinate system and the robot coordinate system is established, and the coordinates of the object to be sorted are mapped from the image coordinate system to the robot coordinate system to obtain the positioning coordinates. The posture parameters corresponding to the object to be sorted are calculated by combining the unified features.

[0056] Based on the spatial geometric features, candidate gripping positions on the surface of the object to be sorted are determined. Based on the candidate gripping positions and preset geometric constraints, the spatial distribution and contact area of ​​the contact region are predicted, and the composite distribution of the contact force on the surface of the object to be sorted is determined. Based on the composite distribution, the mechanical stability value is calculated and combined with the semantic features to obtain the comprehensive adaptability value.

[0057] Based on the positioning coordinates and the attitude parameters, the reachability of the candidate grasping positions is verified, feasible candidate grasping positions are determined, and the candidate grasping position with the largest comprehensive adaptability value is selected as the target grasping point. Based on the target grasping point, the grasping direction vector is determined, and the motion trajectory is calculated in combination with the attitude parameters to complete the grasping.

[0058] In one alternative implementation,

[0059] The process involves acquiring multi-view images of the objects to be sorted, extracting corresponding spatial geometric features, and fusing them with semantic features to obtain fused features. It also involves extracting key surface points of the objects from different viewpoints and establishing spatial correspondences. Based on these spatial correspondences, geometric transformation parameters are calculated, and the fused features are aligned and reconstructed to obtain unified features, including:

[0060] Multi-view images are obtained by acquiring multiple perspectives of the objects to be sorted. Depth is estimated from the multi-view images to obtain depth maps, and the corresponding 3D point cloud data is extracted as spatial geometric features. Semantic features are obtained by feature extraction from the multi-view images. The spatial geometric features and the semantic features are spliced ​​together according to spatial position to obtain fused features.

[0061] Key point detection is performed on the surface of the object to be sorted in the multi-view image to obtain a set of surface key points. The similarity of the feature descriptors corresponding to the surface key points under different views is calculated, and feature matching is performed based on the similarity to obtain key point matching pairs. Spatial correspondence is established based on the image coordinates and depth information of different key points in the key point matching pairs under different views.

[0062] Based on the 3D coordinate differences of the key point matching pairs in the spatial correspondence, the rotation matrix and translation vector are calculated as geometric transformation parameters. Based on the geometric transformation parameters, the spatial geometric features in the fused features are transformed to a unified reference coordinate system to complete coordinate alignment. The spatial geometric features after coordinate alignment are re-associated with the corresponding semantic features in spatial position, and the spatial index relationship of the feature vector is updated to obtain unified features.

[0063] Multi-view cameras are used to acquire images of the objects to be sorted from multiple angles. Typically, 3 to 5 cameras at different angles are set up to obtain a comprehensive view of the objects to be sorted. Taking an industrial part as an example, images are acquired from the top, side, and a 45-degree tilt angle to obtain a color image with a resolution of 1920×1080 pixels.

[0064] After acquiring multi-view images, a depth estimation network is used to process these images and generate corresponding depth maps. The depth estimation network employs an encoder-decoder architecture. The encoder contains five convolutional layers, each using a 3×3 convolutional kernel with a stride of 2, and the number of channels is 64, 128, 256, 512, and 512 respectively. The decoder uses deconvolutional layers to progressively restore the spatial resolution, outputting a depth map of the same size as the input image. For acquired images of industrial parts, the depth map output by the depth estimation network achieves millimeter-level accuracy, with depth values ​​typically ranging from 0.1 meters to 2 meters.

[0065] Based on depth map information and combined with the camera intrinsic parameter matrix, the pixels of the 2D image are back-projected into 3D space to generate point cloud data as spatial geometric features. The point cloud density is typically set to 50 to 100 points per square centimeter to ensure accurate description of the object's surface shape. Point cloud data for industrial parts typically contains 5,000 to 20,000 points, each containing spatial coordinates and corresponding color information.

[0066] Feature extraction is performed on multi-view images to obtain semantic features. The feature extraction network adopts an improved residual network structure, containing four residual blocks, each containing three convolutional layers, with an output feature dimension of 2048. The output features can represent the material, texture, and semantic information of an object. For industrial parts, the extracted semantic features can distinguish different materials such as metal and plastic, as well as attributes such as surface smoothness and texture features.

[0067] Spatial geometric features and semantic features are concatenated according to spatial location to construct fused features. Specifically, for each 3D point, the corresponding semantic feature vector is connected to the spatial coordinates to form a high-dimensional feature vector. For industrial parts, the feature vector of each point has a dimension of 2051, including 3D spatial coordinates and 2048-dimensional semantic features.

[0068] To achieve a unified representation of features from different viewpoints, keypoint detection is performed on the surface of the objects to be sorted. Feature point detection algorithms are used to extract features from the object surface in multi-view images, typically selecting areas with significant curvature changes, edges, and corners as keypoints. For example, for industrial parts, 50 to 200 keypoints are typically detected at each viewpoint. These keypoints are evenly distributed on the object surface, effectively describing the object's geometric features.

[0069] Feature descriptors are calculated for the detected keypoints. These descriptors typically have a dimension of 128 or 256 and describe the gradient information and texture features of the pixels surrounding the keypoint. Euclidean distance or cosine similarity is calculated between the feature descriptors of the keypoints from different viewpoints. A threshold of 0.7 is usually set; two keypoints are considered to match when the similarity is higher than the threshold. For example, for industrial parts, 30 to 100 valid keypoint matching pairs can typically be established between different viewpoints.

[0070] Based on the image coordinates and depth information of different keypoints in keypoint matching under different viewpoints, the 3D spatial coordinates of the keypoints are calculated, and a spatial correspondence is established. The spatial correspondence includes the 3D coordinate mapping relationship of the keypoints under different viewpoints, and the coordinates are usually accurate to the millimeter level.

[0071] Based on the 3D coordinate differences of keypoint matching pairs in spatial correspondence, the least squares method is used to calculate the rotation matrix and translation vector as geometric transformation parameters. During the calculation, keypoint matching pairs with the highest confidence are selected for parameter solving; typically, the top 80% of matching pairs with the highest confidence are chosen. For example, for industrial parts, the obtained rotation matrix can usually align point clouds from different viewpoints, with an average alignment error of less than 2 millimeters.

[0072] Using the calculated geometric transformation parameters, the spatial geometric features in the fused features are transformed to a unified reference coordinate system, thus completing coordinate alignment. The reference coordinate system is typically chosen from the first viewpoint; point cloud data from other viewpoints are mapped to the reference coordinate system through rotation and translation transformations. For example, for industrial parts, the accuracy of the transformed point cloud can typically reach the millimeter level, and the overlap rate between point clouds is usually above 85%.

[0073] After coordinate alignment, the spatial geometric features and corresponding semantic features are re-associated spatially to establish a spatial index structure. This spatial index structure is implemented using an octree or KD-tree, with a partitioning precision typically set to 5 millimeters. For overlapping regions, features from different perspectives are fused, and the feature vector is updated using a weighted average or by taking the maximum value. For example, for industrial parts, the fused feature point cloud typically contains 8,000 to 30,000 points, each containing unified spatial coordinates and semantic features, constituting a unified feature set.

[0074] In this embodiment, a significant improvement in the accuracy and stability of the object representation to be sorted is achieved through feature fusion methods involving multi-view information collaborative modeling and spatial consistency constraints. Spatial geometric features and high-level semantic features are simultaneously introduced on the basis of multi-view images and fused through spatial position constraints. This results in features that not only contain the three-dimensional structural information of the object but also possess the ability to effectively express object categories, morphological differences, and local semantics, enhancing the ability to distinguish similar objects in complex scenes. By matching key points on the object surface under different viewpoints and establishing precise spatial correspondences, geometric transformation parameters between different viewpoints can be reliably estimated, enabling high-precision alignment of multi-source geometric information in a unified reference coordinate system. This effectively reduces spatial offset errors caused by viewpoint changes, occlusion, or pose differences, avoiding feature misalignment problems caused by simple feature superposition or coarse registration in traditional methods. The spatial geometric features and semantic features are re-spatialized and the index relationship is updated, ensuring that the fused unified features simultaneously maintain spatial and semantic consistency.

[0075] In one alternative implementation,

[0076] Based on the unified features, a transformation relationship between the image coordinate system and the robot coordinate system is established, and the coordinates of the object to be sorted are mapped from the image coordinate system to the robot coordinate system to obtain the positioning coordinates. Combined with the unified features, the posture parameters corresponding to the object to be sorted are calculated, including:

[0077] Extract three-dimensional feature points from the surface of the object to be sorted from the unified features and obtain the coordinates of the three-dimensional feature points in the image coordinate system as the original coordinates. Set a calibration reference point in the robot coordinate system. Determine the origin position and coordinate axis direction of the robot coordinate system based on the three-dimensional coordinates of the calibration reference point. Solve the correspondence between the original coordinates and the coordinates in the robot coordinate system to determine the rotation transformation matrix and translation transformation vector. Construct a coordinate transformation matrix based on the rotation transformation matrix and the translation transformation vector as the transformation relationship.

[0078] Substitute the coordinates of the centroid of the object to be sorted in the image coordinate system into the coordinate transformation matrix to perform matrix operations and obtain the positioning coordinates;

[0079] Extract the normal vector distribution of the surface of the object to be sorted from the unified features and calculate the principal direction vector of the object surface. Substitute the direction component of the principal direction vector in the image coordinate system into the rotation transformation matrix to perform coordinate system transformation to obtain the direction component of the principal direction vector in the robot coordinate system. Calculate the rotation angle of the object to be sorted relative to each coordinate axis of the robot coordinate system based on the direction component and combine them to form the attitude parameters.

[0080] Three-dimensional feature points are extracted from the surface of the object to be sorted based on uniform features. The extraction process uses a density clustering algorithm to analyze the point cloud data in the uniform features, with a cluster radius of 10 mm and a minimum number of points of 50. For example, for industrial parts, typically 100 to 300 representative three-dimensional feature points can be extracted. These three-dimensional feature points are located at the edges, corners, and areas with significant curvature changes on the object's surface, effectively representing the object's geometry. Each extracted three-dimensional feature point contains its coordinate values ​​in the image coordinate system, which serve as the original coordinates for coordinate transformation.

[0081] When setting calibration reference points in the robot coordinate system, 4 to 6 pre-marked fixed points on the worktable are typically selected as reference points. These reference points are arranged in a regular geometric shape, such as a square or rectangle, with side lengths typically ranging from 300 to 500 millimeters. The position of each reference point is precisely calibrated, and its coordinate values ​​in the robot coordinate system are known, achieving an accuracy of 0.1 millimeters. For example, in a typical industrial sorting scenario, four reference points can be marked on the worktable, with coordinates in the robot coordinate system of (0, 0, 0), (400, 0, 0), (0, 400, 0), and (400, 400, 0), in millimeters.

[0082] The origin and axis directions of the robot coordinate system are determined based on the three-dimensional coordinates of the calibration reference points. The origin of the robot coordinate system is typically set at the location of the first reference point. The x-axis points from the first reference point to the second reference point, the y-axis is perpendicular to the x-axis and points to the third reference point, and the z-axis is determined according to the right-hand coordinate system rule. In the example of the four reference points mentioned above, the origin is located at the reference point with coordinates (0, 0, 0), the x-axis points from the reference point with coordinates (400, 0, 0), the y-axis points from the reference point with coordinates (0, 400, 0), and the z-axis is perpendicular to the worktable surface and upwards.

[0083] To determine the correspondence between the original coordinates and the coordinates in the robot coordinate system, camera calibration technology is required. By analyzing the correspondence between the coordinates of reference points in the image coordinate system and the robot coordinate system, the least squares method is used to solve for the rotation transformation matrix and translation transformation vector. In practical applications, at least three non-collinear reference points are needed to solve for the complete transformation relationship; using more reference points can improve accuracy. For industrial parts sorting scenarios, six reference points are typically used for calibration, achieving a calibration accuracy of 0.5 mm. The resulting rotation transformation matrix is ​​a 3×3 matrix, and the translation transformation vector is a 3D vector.

[0084] Based on the rotation transformation matrix and translation transformation vector, a coordinate transformation matrix is ​​constructed to represent the transformation relationship between two coordinate systems. The coordinate transformation matrix is ​​a 4×4 matrix, where the top-left 3×3 part is the rotation transformation matrix, the top-right 3×1 part is the translation transformation vector, the bottom-left corner is a 0 vector, and the bottom-right corner is a 1 vector. The transformation matrix is ​​used to transform points in the image coordinate system to the robot coordinate system.

[0085] The centroid position and positioning coordinates of the objects to be sorted are calculated. Point cloud data of the objects to be sorted are extracted from uniform features, and the average coordinates of all points are calculated as the centroid position. For complex-shaped industrial parts, the centroid position can be calculated based on point cloud density weighting to improve positioning accuracy. Taking an industrial part as an example, its point cloud data contains 10,000 points, and the calculated centroid position in the image coordinate system is (150, 200, 50), in millimeters.

[0086] The coordinates of the centroid in the image coordinate system are extended to homogeneous coordinates (150, 200, 50, 1). Substituting these coordinates into the coordinate transformation matrix and performing matrix multiplication yields the coordinates of the centroid in the robot coordinate system, i.e., the positioning coordinates. For the aforementioned industrial part, the transformed positioning coordinates might be (250, 350, 30), in millimeters, representing the position of the object's centroid in the robot coordinate system.

[0087] Calculating the pose parameters of objects to be sorted requires extracting the normal vector distribution of the object's surface from a unified feature set. For each point in the point cloud data, a local plane is fitted using the corresponding neighborhood points, and the normal vector of the local plane is calculated. For industrial parts, neighborhood points within a 5 mm radius are typically selected for normal vector calculation, resulting in the normal vector for each point. The normal vector distribution of the object's surface reflects the object's orientation and pose information.

[0088] Based on the surface normal vector distribution, the principal direction vector of the object's surface is calculated. Covariance analysis is used to construct the covariance matrix of the surface point normal vectors. Eigenvalue decomposition is performed on the covariance matrix, and the eigenvector with the largest eigenvalue is the principal direction vector. For regularly shaped industrial parts, such as cylinders, the principal direction vector is usually along the cylinder's axis; for irregularly shaped industrial parts, the principal direction vector reflects the object's main extension direction.

[0089] Substituting the direction component of the principal direction vector in the image coordinate system into the rotation transformation matrix, we obtain the direction component of the principal direction vector in the robot coordinate system. The direction vector is not affected by translation transformation; only rotation transformation is required. For the aforementioned industrial part, the direction component of the principal direction vector in the image coordinate system might be (0.707, 0.707, 0), while the direction component after transformation to the robot coordinate system might become (0.866, 0.5, 0).

[0090] Based on the direction components of the principal direction vector in the robot coordinate system, the rotation angles of the object to be sorted relative to each coordinate axis of the robot coordinate system are calculated. Euler angles are typically used to represent the object's posture, and the rotation angles around the x, y, and z axes are calculated. For the aforementioned industrial part, the calculated Euler angles might be (0, 0, 30), in degrees, representing a 30-degree rotation around the z-axis. These three angles are combined to form the posture parameters.

[0091] In this embodiment, by establishing a precise spatial mapping relationship between the image coordinate system and the robot coordinate system, a high-precision unified expression of the position and orientation of the object to be sorted is achieved. The rotation transformation matrix and translation transformation vector are solved by utilizing the correspondence between the three-dimensional feature points and the robot calibration reference points, so that the image perception results can be accurately converted to the robot coordinate system. This significantly reduces the positioning deviation introduced by the inconsistency of the coordinate system, improves the accuracy and repeatability of the object's spatial positioning, and achieves a stable estimation of the spatial position of the object to be sorted by uniformly mapping the position of the object's centroid to the robot coordinate system. This avoids the positioning drift problem caused by changes in viewpoint or incomplete local features of the object, which is beneficial for the robot to perform precise path planning during actual grasping and handling.

[0092] In one alternative implementation,

[0093] Based on the spatial geometric features, candidate gripping positions on the surface of the object to be sorted are determined. Based on the candidate gripping positions and preset geometric constraints, the spatial distribution and contact area of ​​the contact region are predicted, and the composite distribution of contact forces on the surface of the object to be sorted is determined, including:

[0094] The curvature distribution of the surface of the object to be sorted is extracted from the spatial geometric features in the unified features. Based on the curvature distribution, flat areas and raised areas on the surface of the object to be sorted are identified. The surface normal vectors of the flat areas and the raised areas are calculated and the surface areas that meet the grasping stability conditions are selected as candidate grasping positions based on the angle between the surface normal vectors and the horizontal plane.

[0095] Obtain the contact surface geometry of the end effector under preset geometric constraints, project the contact surface geometry onto the surface area corresponding to each candidate gripping position, calculate the degree of contact between the contact surface and the surface of the object to be sorted based on the curvature distribution of the surface area corresponding to the contact surface geometry, determine the boundary contour of the actual contact area as the spatial distribution of the contact area based on the degree of contact, and perform area integration to obtain the contact area.

[0096] The gripping force applied by the end effector under the preset geometric constraints is decomposed into local forces at each contact point within the contact area according to the contact area. Based on the local forces and the corresponding surface normal vectors, the normal and tangential components of the local forces on the surface of the object to be sorted are calculated. The normal and tangential components of all contact points within the contact area are vector-superimposed to obtain the composite distribution.

[0097] The curvature distribution of the surface of the object to be sorted is extracted from uniform features. The spatial geometric features within the uniform features are processed, and the curvature value of each point is calculated based on point cloud data. Curvature calculation employs a local neighborhood analysis method, selecting neighborhood points within a 5 mm radius for each point, fitting a local surface, and calculating the principal curvature and mean curvature. The principal curvature represents the maximum and minimum degree of curvature of the surface at that point, and the mean curvature is the arithmetic mean of the two. Taking an industrial part as an example, its surface point cloud contains 15,000 points, and the curvature calculation precision is set to three decimal places. The curvature value distribution range is typically between 0 and 0.5, with units per millimeter.

[0098] Based on the calculated curvature distribution, a threshold segmentation method was used to identify flat and raised areas on the surface of the object to be sorted. The curvature threshold for flat areas was set to 0.05 per millimeter, meaning points with curvature values ​​less than the threshold were classified as flat areas; the curvature threshold for raised areas was set to greater than 0.15 per millimeter. For the aforementioned industrial part, curvature analysis identified 5 flat areas with a total area of ​​approximately 8000 square millimeters, accounting for 40% of the total surface area; and 3 raised areas with a total area of ​​approximately 3000 square millimeters, accounting for 15% of the total surface area.

[0099] Surface normal vectors are calculated for both identified flat and raised regions. Principal component analysis is used to calculate the surface normal vectors. A covariance matrix is ​​constructed for the point cloud data within each region, and the eigenvector corresponding to the smallest eigenvalue in the covariance matrix is ​​the normal vector for that region. To improve calculation accuracy, a weighted average is applied to the normal vectors of all points within the region, with the weights inversely proportional to the curvature value of the point. For flat regions of industrial parts, the consistency of normal vectors is usually high, with directional differences less than 5 degrees; however, the directional differences of normal vectors in raised regions can reach 30 degrees.

[0100] Surface areas meeting the gripping stability criteria were selected based on the angle between the surface normal vector and the horizontal plane. The stable gripping condition was set as areas where the angle between the normal vector and the horizontal plane was greater than 60 degrees and less than 120 degrees, ensuring that the gripping force could effectively act on the object's surface. For areas perpendicular to the horizontal plane, the angle between the normal vector and the horizontal plane was close to 90 degrees, making it the most suitable for gripping. Through angle selection, four areas meeting the stability criteria were selected from the eight areas of the aforementioned industrial part as candidate gripping positions, with a total area of ​​approximately 6000 square millimeters.

[0101] Obtain the contact surface geometry of the end effector under preset geometric constraints. In industrial sorting scenarios, commonly used robot end effectors include parallel grippers, suction cups, and multi-fingered dexterous hands. Taking a parallel gripper as an example, its contact surface is usually rectangular or curved, with a typical size of 20 mm × 30 mm. The contact surface material is rubber or silicone, and the coefficient of friction is approximately 0.8.

[0102] The geometry of the contact surface is projected onto the surface area corresponding to each candidate gripping position. The projection process considers the relative orientation of the contact surface and the object surface, aligning the center of the contact surface with the candidate gripping position, and ensuring that the normal vector of the contact surface is opposite in direction to the normal vector of the object surface. For parallel grippers, the distance between the two contact surfaces is dynamically adjusted according to the object size, typically set to the width of the object in the gripping direction plus a preload distance of 10 mm. For the aforementioned industrial part, contact surface projections are performed at the four candidate gripping positions, resulting in projected areas of 600 mm², 580 mm², 550 mm², and 520 mm², respectively.

[0103] The degree of fit between the contact surface and the surface of the object to be sorted is calculated based on the surface area corresponding to the geometry of the contact surface and the curvature distribution of the surface area. The fit calculation adopts the curvature matching evaluation method, comparing the curvature of the contact surface with the curvature of the object surface. The contact surface is usually designed as a plane or a curved surface with a specific curvature, and the corresponding curvature value is close to zero or a fixed value. The curvature difference of each point in the projection area is calculated, and the smaller the difference, the higher the fit. For the four candidate gripping positions of the industrial parts, the calculated fit degrees are 0.85, 0.92, 0.78 and 0.65, respectively, with values ​​ranging from 0 to 1. The larger the value, the higher the fit.

[0104] The boundary contour of the actual contact area is determined based on the degree of contact. Areas with a contact degree below a threshold of 0.5 are discarded, and the remaining areas form the actual contact area. A contour extraction algorithm is used to obtain the closed boundary of the contact area, with a contour accuracy set to 1 mm. For the second candidate gripping position with the highest contact degree, the contour of the actual contact area is approximately elliptical, with a major axis of approximately 28 mm and a minor axis of approximately 18 mm.

[0105] The contact area is calculated using area integration. The discrete integration method is employed, dividing the contact area into a grid of 0.5 mm × 0.5 mm cells, and the sum of the areas of all grid cells is calculated. For the aforementioned elliptical contact area, the calculated actual contact area is 500 square millimeters, less than the theoretical projected area of ​​580 square millimeters. This difference primarily stems from contact loss caused by surface curvature mismatch.

[0106] The gripping force applied by the end effector under preset geometric constraints is decomposed into a force across the contact area. The magnitude of the gripping force is typically determined based on the object's weight and coefficient of friction; for an industrial part weighing 500 grams, a gripping force of 20 Newtons is set. The gripping force is evenly distributed across the contact area, with the force distributed to each small area proportional to its size. For a contact area of ​​500 square millimeters, an average gripping force of 0.04 Newtons is applied per square millimeter.

[0107] The normal and tangential components of the local force on the surface of the object to be sorted are calculated based on the local force and the corresponding surface normal vector. The direction of the local force is usually consistent with the contact surface normal vector of the end effector, but at a certain angle to the surface normal vector of the object. Through vector decomposition, the local force is decomposed into a normal component along the direction of the surface normal vector and a tangential component perpendicular to the normal vector. For the case where the angle between the surface normal vector of the industrial part and the contact surface normal vector is 10 degrees, the normal component of the local force is approximately 0.0394 N / mm², and the tangential component is approximately 0.0069 N / mm².

[0108] The composite force is obtained by vector superposition of the normal and tangential components at all contact points within the contact area. Vector superposition considers both the magnitude and direction of the force; the normal components are added along the direction of the normal vector to the object's surface, and the tangential components are added within the tangential plane of the object's surface. For a contact area of ​​500 square millimeters, the resultant force of the normal components is approximately 19.7 Newtons, the resultant force of the tangential components is approximately 3.45 Newtons, and the composite force is 20 Newtons, with its direction essentially consistent with the original gripping force.

[0109] In this embodiment, flat and convex regions are distinguished based on the three-dimensional curvature distribution, and the selection is performed by combining the angle between the surface normal vector and the horizontal plane. This allows the candidate gripping positions to more realistically reflect the spatial morphological characteristics of the object's surface, thereby effectively improving the stability of the gripping area under spatial posture changes. By introducing an evaluation of the fit between the geometry of the end effector contact surface and the object's surface, gripping is no longer simplified to point contact or ideal surface contact. Instead, the actual contact area is spatially modeled and the contact area is calculated, avoiding gripping instability caused by ignoring curvature changes or insufficient contact. This makes the gripping planning results more consistent with real physical contact conditions. The gripping force is decomposed according to the contact area and the normal and tangential components are calculated by combining the surface normal vector. Then, the force in the entire contact area is synthesized and analyzed, realizing a quantitative description of gripping stability and potential slippage risk. This can more comprehensively reflect the force distribution characteristics during the gripping process and improve the adaptability to complex-shaped objects and non-uniform contact conditions.

[0110] In one alternative implementation,

[0111] The mechanical stability value is calculated based on the synthetic distribution, and the comprehensive fitness value is obtained by combining the semantic features, including:

[0112] The spatial distribution of contact forces on the surface of the object to be sorted is obtained from the synthetic distribution, and the force field is reconstructed to obtain the force distribution field. Based on the force distribution field, the stress tensor distribution on the surface of the object to be sorted is calculated, and the resultant force and the point of application of the resultant force in the contact area are obtained by integration. The position vector of the point of application of the resultant force relative to the center of mass of the object to be sorted is calculated, and the resultant moment of the object to be sorted relative to the center of mass is obtained by combining the resultant force. The gravitational moment is obtained based on the mass and center of mass of the object to be sorted. The vector sum of the resultant moment and the gravitational moment is calculated, and it is determined whether the vector sum satisfies the moment equilibrium condition. If so, the vector sum is used as the mechanical stability value.

[0113] Based on the semantic features in the unified features, extract the material properties and surface texture features of the object to be sorted. Based on the material properties, determine the elastic modulus and Poisson's ratio of the object to be sorted and calculate the surface deformation of the object under contact force. Based on the surface texture features, calculate the micro-protrusion distribution on the surface of the object to be sorted and determine the ratio of the actual contact area to the nominal contact area and record it as the first coefficient.

[0114] The deformation adaptability value is obtained by correlating the mechanical stability value and the surface deformation value, and the contact effectiveness value is obtained by correlating the mechanical stability value and the first coefficient. The comprehensive adaptability value is obtained by weighted fusion of the deformation adaptability value and the contact effectiveness value.

[0115] The spatial distribution of contact forces on the surface of the object to be sorted is obtained from the synthetic distribution, and the force field is reconstructed. The force field reconstruction uses discrete point sampling and interpolation algorithms to mesh the contact area into small grids of 0.5 mm × 0.5 mm. Each grid node records the magnitude and direction of the corresponding force. 2000 sampling points are set within the contact area, and the normal and tangential force components are calculated for each sampling point. Taking a metal part as an example, the normal force in the central region of the force distribution field within the contact area is approximately 0.06 N / mm², while that in the edge region is approximately 0.02 N / mm², exhibiting a distribution characteristic of high at the center and low at the edges.

[0116] The stress tensor distribution on the surface of the object to be sorted is calculated based on the force distribution field. The stress tensor calculation considers two components: normal stress and tangential stress. Normal stress equals normal force divided by the area of ​​application, and tangential stress equals tangential force divided by the area of ​​application. For the aforementioned metal parts, the normal stress at the center of the contact area is approximately 0.06 MPa, and the tangential stress is approximately 0.012 MPa; the normal stress at the edge area is approximately 0.02 MPa, and the tangential stress is approximately 0.004 MPa. The stress tensor comprehensively describes the stress state of the object's surface in all directions, including the normal component and two tangential components.

[0117] The stress tensor is integrated within the contact area to obtain the resultant force and its point of application. The integration process employs a numerical integration method, dividing the contact area into minute elements, calculating and summing the force contribution of each element. The resultant force is calculated by adding the force vectors of all minute elements. The point of application is calculated using the torque balance principle, ensuring that the torque generated by the resultant force at the point of application equals the sum of the torques generated by all distributed forces. For a contact area of ​​500 square millimeters, the resultant force is calculated to be 20 Newtons, with the point of application located near the geometric center of the contact area, approximately 1.5 millimeters off-center, at coordinates (125.5, 98.2, 45.3), in millimeters.

[0118] Calculate the position vector of the point of application of the resultant force relative to the center of mass of the object to be sorted. The center of mass of the object to be sorted is calculated from point cloud data and is located at coordinates (120, 100, 50), in millimeters. The position vector of the point of application of the resultant force relative to the center of mass is (5.5, -1.8, -4.7), in millimeters, indicating that the point of application of the resultant force is located 5.5 mm to the right, 1.8 mm in front, and 4.7 mm below the center of mass.

[0119] The resultant moment of the object to be sorted relative to its center of mass is obtained by combining the resultant force and the position vector. The resultant moment is calculated using the cross product principle, where the position vector and the resultant force vector are cross-multiplied to obtain the resultant moment. The direction of the resultant force is (0.1, 0.2, -0.98), indicating that the resultant force is slightly inclined, mainly along the negative z-axis. The calculated resultant moment is (0.56, 0.47, 1.28) Newton-meters, indicating that the rotational tendency of the object is mainly around the z-axis.

[0120] The gravitational torque is calculated based on the mass and center of mass of the object to be sorted. The metal part has a mass of 0.5 kg and a gravity of 4.9 N, with a direction of (0, 0, -1). The point of application of gravity is the object's center of mass; therefore, the position vector of gravity relative to the center of mass is zero, and the calculated gravitational torque is zero. When considering changes in the gripping posture, such as gripping the object at a 45-degree angle, the point of application of gravity remains the center of mass, but the position vector relative to the gripping point of the robotic arm will change. In this case, the gravitational torque needs to be recalculated.

[0121] Calculate the vector sum of the resultant torque and the gravitational torque, and determine if the torque balance condition is met. The vector sum calculation adds the resultant torque to the gravitational torque component, yielding a total torque of (0.56, 0.47, 1.28) N·m. The torque balance condition requires the magnitude of the total torque to be less than a preset threshold, typically set to 1.5 N·m. The calculated magnitude of the total torque is 1.45 N·m, which is less than the threshold, thus satisfying the torque balance condition. The magnitude of the total torque, 1.45 N·m, is used as the mechanical stability value; a smaller mechanical stability value indicates a more stable grip.

[0122] Material properties and surface texture features of the objects to be sorted are extracted based on semantic features from unified features. Material properties are identified from visual images using a deep learning model. The deep learning model employs a convolutional neural network structure, containing 16 convolutional layers and 3 fully connected layers. The input is a 256×256 pixel image of the object's surface, and the output is the material category and material parameters. Surface texture features are captured using a high-resolution camera with a resolution of 20 micrometers per pixel, resulting in an image size of 2048×2048 pixels, covering a 10 mm × 10 mm area of ​​the object's surface.

[0123] The elastic modulus and Poisson's ratio of the objects to be sorted were determined based on their material properties. For the identified metal material, the elastic modulus was found to be 210 GPa and the Poisson's ratio to be 0.3, according to the material database. For the plastic material, the elastic modulus was approximately 3 GPa and the Poisson's ratio was approximately 0.4. For the rubber material, the elastic modulus was approximately 0.01 GPa and the Poisson's ratio was approximately 0.48.

[0124] The surface deformation of the objects to be sorted under contact force was calculated. The deformation was calculated using Hertzian contact theory, considering factors such as contact area, contact pressure, material elastic modulus, and Poisson's ratio. For the aforementioned metal parts, the average pressure in the contact area was 0.04 MPa, and the calculated maximum surface deformation was 0.005 mm, with an average deformation of 0.003 mm. The deformation was mainly concentrated in the center of the contact area. For plastic parts under the same pressure, the maximum deformation reached 0.1 mm, and the average deformation was approximately 0.06 mm.

[0125] The distribution of micro-protrusions on the surface of objects to be sorted was calculated based on surface texture features. Surface micro-protrusion analysis employed Gaussian filtering and threshold segmentation to extract micro-protrusion features from high-resolution surface images. With a height threshold of 5 micrometers, approximately 200 micro-protrusions per unit area were identified. The average height of the micro-protrusions was 8 micrometers, with a standard deviation of 3 micrometers, and their spatial distribution exhibited randomness.

[0126] The ratio of the actual contact area to the nominal contact area is determined and denoted as the first coefficient. The actual contact area refers to the area that actually contacts after considering microscopic protrusions, while the nominal contact area refers to the contact area measured macroscopically. For a metal part with a surface roughness of 3.2 micrometers, the actual contact area is approximately 15% of the nominal contact area, and the first coefficient is 0.15. For a metal part with a surface roughness of 1.6 micrometers, the first coefficient can be increased to 0.25. For a metal part with a surface roughness of 6.4 micrometers, the first coefficient is decreased to 0.08.

[0127] The deformation fitness value is obtained by correlating the mechanical stability value and the surface deformation. The correlation calculation uses a nonlinear mapping function to consider the synergistic effect between deformation and mechanical stability. For a mechanical stability value of 1.45 N·m and a surface deformation of 0.003 mm, the calculated deformation fitness value is 0.82, ranging from 0 to 1. A higher deformation fitness value indicates a higher fitness.

[0128] The contact effectiveness value is obtained by correlating the mechanical stability value with the first coefficient. The correlation calculation uses a linear weighting method, considering the combined effect of mechanical stability and contact area ratio. For a mechanical stability value of 1.45 N·m and a first coefficient of 0.15, the calculated contact effectiveness value is 0.68, ranging from 0 to 1. A higher value indicates more effective contact.

[0129] The deformation fitness score and the contact effectiveness score are weighted and fused to obtain a comprehensive fitness score. The weighting fusion uses a geometric mean method, assigning a weight of 0.6 to deformation fitness and a weight of 0.4 to contact effectiveness. For a deformation fitness score of 0.82 and a contact effectiveness score of 0.68, the calculated comprehensive fitness score is 0.76, ranging from 0 to 1. A higher value indicates a higher grasping fitness.

[0130] In this embodiment, by reconstructing the force field and calculating the stress tensor distribution, the spatial distribution of forces within the contact area can be accurately reflected. Furthermore, the resultant force, the point of application of the resultant force, and the torque generated relative to the object's center of mass are solved, thereby systematically analyzing the force equilibrium state of the object during the grasping process. This effectively improves the ability to identify potential rollover or instability risks. By vector synthesis of the contact resultant torque and the gravitational torque generated by the object's own weight, and determining the torque equilibrium condition, the analysis is no longer limited to local contact stability. Material properties and surface texture information based on semantic features are introduced to model the surface deformation behavior and actual contact state of the object under contact forces. The distribution of microscopic protrusions is used to characterize the difference between the actual contact area and the nominal contact area, enabling the contact assessment results to adapt to objects with different materials and surface conditions. This overcomes the insufficient adaptability caused by ignoring material differences or uniformly assuming contact conditions. By correlating the mechanical stability values ​​with the surface deformation and the actual contact ratio, and then weighting and fusing them to form a comprehensive adaptability value, a unified quantitative evaluation of the grasping scheme under multiple dimensions such as stability, deformation adaptability, and contact effectiveness is achieved.

[0131] Figure 2 This is a flowchart illustrating the object grasping mechanical adaptability evaluation of the vision recognition-based intelligent positioning and grasping method for industrial sorting robots according to an embodiment of the present invention.

[0132] In one alternative implementation,

[0133] Based on the positioning coordinates and the attitude parameters, the reachability of candidate grasping positions is verified to determine feasible candidate grasping positions and select the candidate grasping position with the largest comprehensive adaptability value as the target grasping point, including:

[0134] Obtain the link length parameters and joint rotation axis direction of the robot end effector and establish a mathematical mapping relationship between the robot joint angle and the spatial position of the end effector. Based on the mathematical mapping relationship, derive the inverse solution expression of the spatial position and attitude of the end effector to the joint angle.

[0135] For each candidate gripping position, the positioning coordinates and attitude parameters corresponding to the current candidate gripping position are substituted into the inverse solving expression to calculate the target joint angle of each joint of the robot. The angle limit range of each joint of the robot is obtained and it is determined whether the target joint angle exceeds the angle limit range to obtain the first judgment result. The reachable distance range of the robot end effector is calculated based on the link length parameter and it is determined whether the positioning coordinates exceed the reachable distance range to obtain the second judgment result. Based on the first judgment result and the second judgment result, a feasible candidate gripping position is determined.

[0136] Extract the comprehensive adaptability value corresponding to the feasible candidate crawling position, sort the feasible candidate crawling positions in descending order based on the comprehensive adaptability value to obtain a priority sequence, and take the first feasible candidate crawling position in the priority sequence as the target crawling point.

[0137] To obtain the link length parameters and joint rotation axis directions of the robot's end effector, for a six-axis robot used in industrial sorting, the link length parameters are usually obtained from the robot's technical manual. Taking a commonly used industrial sorting robot as an example, its link lengths are as follows: height from the base to the first joint axis: 345 mm; distance from the first joint to the second joint: 320 mm; distance from the second joint to the third joint: 1150 mm; distance from the third joint to the fourth joint: 200 mm; distance from the fourth joint to the fifth joint: 200 mm; distance from the fifth joint to the sixth joint: 126 mm; distance from the sixth joint to the center point of the end effector: 167 mm. The joint rotation axis directions are defined as follows: first joint rotation axis along the z-axis; second joint rotation axis along the y-axis; third joint rotation axis along the y-axis; fourth joint rotation axis along the z-axis; fifth joint rotation axis along the y-axis; and sixth joint rotation axis along the z-axis.

[0138] Establishing a mathematical mapping relationship between robot joint angles and the spatial position of the end effector requires the use of a standard robot forward kinematics model. Based on the link coordinate system transformation, the transformation relationship between each joint coordinate system is recursively calculated to obtain the position and orientation of the end effector relative to the base coordinate system. During the forward kinematics calculation, the rotation or translation transformation of each joint is considered separately, and the transformation matrices are multiplied sequentially to obtain the total transformation matrix from the base coordinate system to the end effector coordinate system. The first three columns of the total transformation matrix describe the orientation of the end effector, and the first three elements of the last column describe the spatial position coordinates of the end effector. For the aforementioned six-axis robot, when the joint angles are 0 degrees, 45 degrees, -30 degrees, 0 degrees, 60 degrees, and 90 degrees, the spatial position coordinates of the end effector are (1358.5, 0, 765.3) mm, and the orientation is a rotation of 30 degrees around the x-axis, a rotation of 0 degrees around the y-axis, and a rotation of 90 degrees around the z-axis.

[0139] The inverse kinematics (IK) of the robot is derived by deriving the inverse kinematics expressions from the spatial position and orientation of the end effector to the joint angles based on mathematical mapping relationships. For a six-DOF robot, the IK is solved analytically, decomposing the position and orientation of the end effector into multiple geometric subproblems. The calculation process first solves the first three joint angles to determine the position of the robot's wrist center point, and then solves the last three joint angles to determine the orientation of the end effector. For the first three joints, the joint angles are inversely derived from the end effector position coordinates, involving trigonometric functions and vector operations; for the last three joints, the joint angles are inversely derived from the end effector orientation matrix, involving matrix decomposition and Euler angle calculation. The solution process usually yields multiple solutions, and the optimal solution must be selected based on the actual situation. For example, for an end effector with a position of (1358.5, 0, 765.3) mm and an orientation of 30 degrees rotation around the x-axis, 0 degrees rotation around the y-axis, and 90 degrees rotation around the z-axis, the inverse kinematics results in joint angles of 0 degrees, 45 degrees, -30 degrees, 0 degrees, 60 degrees, and 90 degrees, verifying the consistency between the forward and inverse kinematics models.

[0140] For each candidate grasping position, the positioning coordinates and attitude parameters corresponding to the current candidate grasping position are substituted into the inverse kinematics expression to calculate the target joint angles of each joint of the robot. Taking the first candidate grasping position as an example, the positioning coordinates are (950.2, 325.8, 420.5) mm, and the attitude is a rotation of -15 degrees around the x-axis, a rotation of 5 degrees around the y-axis, and a rotation of 78 degrees around the z-axis. Substituting into the inverse kinematics expression, the target angles of each joint are calculated to be 18.9 degrees, 32.6 degrees, -26.8 degrees, -12.4 degrees, 52.3 degrees, and 96.7 degrees, respectively.

[0141] The first judgment result is obtained by acquiring the angle limit range of each joint of the robot and determining whether the angle of the target joint exceeds the angle limit range. For the aforementioned six-axis robot, the angle limit range of each joint is as follows: first joint ±180 degrees, second joint -60 degrees to 120 degrees, third joint -120 degrees to 156 degrees, fourth joint ±180 degrees, fifth joint ±130 degrees, and sixth joint ±360 degrees. Comparing the target joint angle with the limit range, it is determined that the target joint angle of the first candidate grasping position is within the limit range, and the first judgment result is that the limit condition is met. For the second candidate grasping position, the positioning coordinates are (1050.5, -450.6, 320.8) mm, and the posture is a rotation of 25 degrees around the x-axis, a rotation of -8 degrees around the y-axis, and a rotation of -65 degrees around the z-axis. The calculated target angle of the second joint is 125.4 degrees, which exceeds the maximum limit of 120 degrees for the second joint, and the first judgment result is that the limit condition is not met.

[0142] The reachable distance range of the robot's end effector is calculated based on the link length parameters, and a second judgment result is obtained by determining whether the positioning coordinates exceed the reachable distance range. The robot's maximum reachable distance is calculated by the vector sum of the lengths of each link. Considering the range of motion of each joint, the theoretical maximum reachable distance of the aforementioned six-axis robot is 2163 mm, the actual effective reachable distance is 1850 mm, and the minimum reachable distance is 500 mm. For the first candidate gripping position, its positioning coordinates are 1060.3 mm away from the robot base, which is within the reachable distance range, and the second judgment result is that the reachability condition is met. For the third candidate gripping position, the positioning coordinates are (1920.8, 150.2, 680.5) mm, and the distance from the robot base is 2041.5 mm, which exceeds the actual effective reachable distance, and the second judgment result is that the reachability condition is not met.

[0143] Feasible candidate gripping positions are determined based on the first and second judgment results. A candidate gripping position is only deemed feasible if both the first and second judgment results meet the conditions. Of the three candidate gripping positions mentioned above, only the first candidate gripping position simultaneously meets the joint angle limitation and reachability range requirements, and is therefore determined to be a feasible candidate gripping position. In practical applications, multiple candidate gripping positions typically meet the conditions, forming a set of feasible candidate gripping positions. For a gripping task of a certain industrial part, four feasible candidate gripping positions are obtained after judgment: the first, fourth, sixth, and seventh candidate gripping positions.

[0144] Extract the comprehensive suitability values ​​corresponding to feasible candidate crawling positions, and sort the feasible candidate crawling positions in descending order based on the comprehensive suitability values ​​to obtain a priority sequence. For the aforementioned four feasible candidate crawling positions, the comprehensive suitability values ​​are 0.76, 0.82, 0.65, and 0.73, respectively. Sort these values ​​in descending order to obtain the ranking results as the fourth position (0.82), the first position (0.76), the seventh position (0.73), and the sixth position (0.65). Based on the ranking results, a priority sequence is formed, with the priority from high to low as the fourth, first, seventh, and sixth candidate crawling positions.

[0145] The first feasible candidate gripping position in the priority sequence is selected as the target gripping point. Based on the priority sequence, the fourth candidate gripping position has the highest overall fit value of 0.82, and is therefore chosen as the target gripping point. The location coordinates of this target gripping point are (865.3, 286.5, 495.2) mm, with an orientation of 5 degrees rotation around the x-axis, -3 degrees rotation around the y-axis, and 82 degrees rotation around the z-axis. The corresponding robot joint angles are 18.3 degrees, 40.5 degrees, -22.7 degrees, -5.8 degrees, 48.2 degrees, and 88.5 degrees.

[0146] In this embodiment, by establishing a precise mathematical mapping between joint angles and the spatial position of the end effector and deriving the inverse solution expression, each candidate grasping position can correspond to a clear joint target state. This allows for accurate determination of whether the joint angle exceeds physical limitations and whether the end effector is within the reachable space during the planning stage, significantly improving the accuracy and reliability of kinematic feasibility judgment and reducing the deviation between planning results and actual execution. Candidate grasping positions are retained only under the premise of satisfying joint angle constraints and reachability constraints, so that the subsequent decision space is focused on the truly executable grasping scheme, effectively reducing planning complexity and the proportion of invalid calculations. By introducing the aforementioned comprehensive adaptability value to sort feasible candidate grasping positions and selecting the position with the highest priority as the target grasping point, the finally selected grasping point not only satisfies the robot's kinematic constraints but also achieves overall optimization in terms of mechanical stability, deformation adaptability, and contact effectiveness.

[0147] In one alternative implementation,

[0148] Determining the grasping direction vector based on the target grasping point, calculating the motion trajectory based on the posture parameters, and completing the grasping process includes:

[0149] The system acquires the position and attitude parameters of the target gripping point. Based on the attitude parameters, it determines the normal approach angle and tangential deflection angle of the robot end effector relative to the surface of the object to be sorted. Based on the normal approach angle, it calculates the unit vector along the normal of the surface of the object to be sorted. Based on the tangential deflection angle, it calculates the rotation correction vector in the tangential plane of the surface of the object to be sorted. Based on the unit vector and the rotation correction vector, it determines the initial gripping vector and, combined with the synthetic distribution, determines the cosine value of the angle between the direction of the contact force and the initial gripping vector. Based on the cosine value of the angle, it performs mechanical optimization adjustment on the initial gripping vector to obtain the gripping direction vector.

[0150] Obtain the current position coordinates and current attitude angle of the robot's end effector, and calculate the spatial displacement vector of the position parameters from the current position coordinates to the target grasping point and the angular change of the attitude parameters from the current attitude angle to the target grasping point;

[0151] A pose transformation matrix is ​​constructed based on the spatial displacement vector and the grasping direction vector, and singular value decomposition is performed to obtain translation and rotation components. A translation path is generated based on the translation components, and the path is corrected along the grasping direction vector to obtain a corrected translation path. A rotation path is generated based on the rotation components and the angle change, ensuring the orthogonality between the rotation axis and the grasping direction vector. The corrected translation path and the rotation path are interpolated by a fifth-order polynomial to obtain the motion trajectory, and the object to be sorted is moved along the motion trajectory to the target grasping point to complete the grasping of the object to be sorted.

[0152] The target grasping point's position parameters are represented by three-dimensional spatial coordinates (865.3, 286.5, 495.2) mm, and the attitude parameters are represented by Euler angles (5, -3, 82) degrees, corresponding to the rotation angles around the three axes of the coordinate system. The position parameters determine the precise location of the grasping point in space, while the attitude parameters determine the direction and orientation of the end effector during grasping. Together, they constitute a complete pose description of the grasping point.

[0153] The normal approach angle and tangential deflection angle of the robot's end effector relative to the surface of the object to be sorted are determined based on the attitude parameters. The normal approach angle is defined as the angle between the end effector's spindle and the normal vector of the object's surface, while the tangential deflection angle is defined as the rotation angle of the end effector within the tangential plane of the object's surface. For the attitude parameters (5, -3, 82) degrees of the target gripping point, the normal approach angle is calculated to be 8 degrees and the tangential deflection angle to be 82 degrees through coordinate transformation. A smaller normal approach angle indicates that the end effector is almost perpendicular to the object's surface, which is beneficial for stable gripping; the tangential deflection angle determines the orientation of the gripper's opening and closing direction relative to the object's surface.

[0154] The unit vector along the surface normal of the object to be sorted is calculated based on the approach angle of the normal. The surface normal of the object to be sorted is calculated by the vision system from point cloud data. The surface normal vector at the target gripping point is (0.12, 0.05, 0.99), which is normalized to obtain the unit normal vector. Considering the 8-degree offset of the approach angle of the normal, the unit vector along the surface normal is calculated to be (0.14, 0.04, 0.99). The unit vector represents the ideal direction for the end effector to approach the object, ensuring that the gripping force can be effectively applied to the object surface.

[0155] The rotation correction vector within the tangential plane of the object to be sorted is calculated based on the tangential deflection angle. The tangential plane is determined by the surface normal vector, and the rotation within the tangential plane is calculated based on a tangential deflection angle of 82 degrees. By constructing a local coordinate system with the surface normal vector as the z-axis, the direction vector corresponding to the 82-degree rotation within this coordinate system is calculated, resulting in the rotation correction vector (-0.97, 0.23, 0.05). This vector represents the orientation of the end effector within the tangential plane and determines the alignment of the gripper relative to the object's features.

[0156] The initial grasping vector is determined based on the unit vector and the rotation correction vector. The initial grasping vector is obtained by combining the unit vector and the rotation correction vector, representing the initial direction in which the end effector approaches the object. During the calculation, the dominant role of the unit vector and the auxiliary adjustment of the rotation correction vector are considered, resulting in an initial grasping vector of (0.16, 0.04, 0.98).

[0157] The cosine of the angle between the applied contact force direction and the initial grasping vector is determined by combining the composite distribution. The applied contact force direction is determined by the aforementioned force distribution analysis, and its direction at the target grasping point is (0.12, 0.06, -0.99). The calculated cosine of the angle between the applied contact force direction and the initial grasping vector is -0.97, which is close to -1, indicating that the two vectors are almost opposite in direction. This is consistent with the physical characteristic that the grasping force and grasping direction are opposite.

[0158] The grasping direction vector is obtained by mechanically optimizing the initial grasping vector based on the cosine of the included angle. This optimization aims to better align the grasping direction with the contact force direction, thereby improving grasping stability. The adjustment process employs vector projection and orthogonal decomposition methods, fine-tuning the direction of the grasping vector while ensuring the cosine of the included angle remains close to -1. Optimizing the initial grasping vector (0.16, 0.04, 0.98) yields a grasping direction vector of (0.13, 0.06, 0.99).

[0159] Obtain the current position coordinates and current attitude angle of the robot's end effector. The real-time state of the end effector is obtained through feedback from the robot control system. The current position coordinates are (650.8, 420.3, 780.5) mm, and the current attitude angle is (15, 10, 45) degrees. The current state is the starting point for planning the grasping trajectory, and together with the target grasping point, it defines the spatial motion that the end effector needs to complete.

[0160] Calculate the spatial displacement vector from the current position coordinates to the corresponding position parameters of the target grasping point. The spatial displacement vector represents the distance and direction the end effector needs to move in three-dimensional space, calculated by subtracting the current position from the target position. For the current position (650.8, 420.3, 780.5) mm and the target position (865.3, 286.5, 495.2) mm, the calculated spatial displacement vector is (214.5, -133.8, -285.3) mm. This spatial displacement vector indicates that the end effector needs to move 214.5 mm to the right, 133.8 mm backward, and 285.3 mm downward.

[0161] Calculate the angular change in attitude parameters from the current attitude angle to the target grasping point. The angular change represents the angle the end effector needs to adjust along the three rotation axes, calculated by subtracting the current attitude from the target attitude. For the current attitude (15, 10, 45 degrees) and the target attitude (5, -3, 82 degrees), the calculated angular change is (-10, -13, 37 degrees). This means the end effector needs to rotate 10 degrees counterclockwise around the x-axis, 13 degrees counterclockwise around the y-axis, and 37 degrees clockwise around the z-axis.

[0162] A pose transformation matrix is ​​constructed based on the spatial displacement vector and the grasping direction vector, and singular value decomposition (SVD) is performed to obtain translational and rotational components. The pose transformation matrix comprehensively describes the position and attitude changes of the end effector, and is constructed using spatial displacement vectors and attitude angle changes. Singular value decomposition decomposes the transformation matrix into two basic motions: translation and rotation. Singular value decomposition is performed on the constructed pose transformation matrix to obtain the translational components, represented as a translational distance of 385.3 mm and a translational direction (0.56, -0.35, -0.74), and the rotational components, represented as a rotational angle of 40.6 degrees and a rotational axis direction (-0.24, -0.32, 0.92).

[0163] A corrected translation path is obtained by generating a translation path based on translation components and correcting the path along the grasping direction vector. The initial translation path is planned as a straight line, moving at a constant speed from the current position to near the target position along the translation direction. Considering the special nature of the grasping operation, the path needs to be corrected so that the end effector moves along the grasping direction vector when approaching the object, ensuring stable contact. The correction method adopts a segmented planning strategy, dividing the translation path into an approach segment and an alignment segment. The approach segment accounts for 80% of the total path and moves along the original translation direction; the alignment segment accounts for 20% of the total path and moves along the grasping direction vector (0.13, 0.06, 0.99). The corrected translation path is described by 100 discrete points: the first 80 points are evenly distributed along the original translation direction, and the last 20 points are evenly distributed along the grasping direction vector, forming a smoothly transitioned corrected translation path.

[0164] A rotation path is generated based on the rotation components and angle changes, ensuring the orthogonality between the rotation axis and the grasping direction vector. The rotation path describes the attitude change process of the end effector from its current posture to the target posture. To ensure grasping stability, the rotation axis must remain orthogonal to the grasping direction vector to avoid unnecessary attitude fluctuations when approaching the object. The dot product of the rotation axis (-0.24, -0.32, 0.92) and the grasping direction vector (0.13, 0.06, 0.99) is 0.89, which does not satisfy the orthogonality condition. Through vector orthogonalization, the rotation axis is adjusted to (-0.82, -0.57, 0.05) to make it orthogonal to the grasping direction vector. Based on the adjusted rotation axis and rotation angle of 40.6 degrees, a rotation path consisting of 50 discrete postures is generated to describe the continuous attitude change process of the end effector.

[0165] The motion trajectory is obtained by performing fifth-order polynomial interpolation on the corrected translation and rotation paths. Fifth-order polynomial interpolation generates smooth trajectories with continuous velocity and acceleration, avoiding abrupt acceleration and deceleration during robot motion. The interpolation calculation considers boundary conditions, including the position, velocity, and acceleration of the start and end points, ensuring a smooth transition of the motion state at the start and end points. Fifth-order polynomial interpolation is performed on the corrected translation path with 100 discrete points and the rotation path with 50 discrete poses, respectively, to obtain parameterized continuous trajectory functions. The total duration of the interpolated motion trajectory is set to 5 seconds, with a resolution of 10 milliseconds, containing 500 trajectory points, each point containing 6 parameters: position coordinates and pose angle.

[0166] The robot moves along its trajectory to the target gripping point to grasp the object to be sorted. The robot control system controls the motors of each joint according to the planned trajectory, driving the end effector to precisely track the trajectory points, achieving smooth movement from the current position to the target gripping point. In the final stage of approaching the object, the end effector moves strictly along the gripping direction vector to ensure that the direction of the contact force is consistent with the expectation. Upon reaching the target gripping point, the end effector closes its gripper with a preset gripping force, firmly clamping the object to be sorted, completing the gripping operation.

[0167] In this embodiment, by uniformly modeling the spatial displacement and attitude changes between the current position and the target grasping point, and decomposing the translation and rotation components using the pose transformation matrix, the translational motion and attitude adjustment process are decoupled, improving the predictability and stability of motion planning. When generating the translational path, the grasping direction vector is introduced for path correction, ensuring that the approach process of the end effector always remains consistent with the optimal grasping direction, reducing the probability of lateral contact or undesirable collisions. In the rotational path planning, the orthogonality between the rotation axis and the grasping direction vector is guaranteed, making the attitude adjustment more in line with the mechanical requirements of the grasping process, effectively improving the alignment accuracy between the end effector and the object surface. By performing fifth-order polynomial interpolation on the corrected translational and rotational paths, a smooth motion trajectory with continuous velocity and acceleration characteristics is generated, significantly reducing the impact and vibration during the motion process, and improving the compliance and control accuracy of the grasping process.

[0168] A second aspect of this invention provides a vision-based intelligent positioning and grasping system for industrial sorting robots, comprising:

[0169] The feature fusion unit is used to acquire multi-view images of the object to be sorted and extract the corresponding spatial geometric features and semantic features to fuse them into fused features. It also extracts the surface key points of the object to be sorted under different views and establishes spatial correspondence. Based on the spatial correspondence, it calculates geometric transformation parameters and performs coordinate alignment and feature reconstruction on the fused features to obtain unified features.

[0170] The coordinate transformation unit is used to establish the transformation relationship between the image coordinate system and the robot coordinate system based on the unified feature, and to map the coordinates of the object to be sorted from the image coordinate system to the robot coordinate system to obtain the positioning coordinates and calculate the posture parameters corresponding to the object to be sorted in combination with the unified feature.

[0171] The grasping evaluation unit is used to determine candidate grasping positions on the surface of the object to be sorted based on the spatial geometric features, predict the spatial distribution and contact area of ​​the contact region based on the candidate grasping positions and preset geometric constraints, determine the composite distribution of contact forces on the surface of the object to be sorted, calculate the mechanical stability value based on the composite distribution, and solve for the comprehensive adaptability value by combining the semantic features.

[0172] The grasping execution unit is used to verify the reachability of candidate grasping positions based on the positioning coordinates and the attitude parameters, determine feasible candidate grasping positions and select the candidate grasping position with the largest comprehensive adaptability value as the target grasping point, determine the grasping direction vector based on the target grasping point, calculate the motion trajectory in combination with the attitude parameters and complete the grasping.

[0173] A third aspect of the present invention provides an electronic device, comprising:

[0174] A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0175] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0176] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent positioning and grasping of industrial sorting robots based on visual recognition, characterized in that, include: Multi-view images of the object to be sorted are acquired and the corresponding spatial geometric features and semantic features are extracted and fused to obtain fused features. Surface key points of the object to be sorted under different views are extracted and spatial correspondences are established. Geometric transformation parameters are calculated based on the spatial correspondences, and coordinate alignment and feature reconstruction are performed on the fused features to obtain unified features. Based on the unified features, a transformation relationship between the image coordinate system and the robot coordinate system is established, and the coordinates of the object to be sorted are mapped from the image coordinate system to the robot coordinate system to obtain the positioning coordinates. The posture parameters corresponding to the object to be sorted are calculated by combining the unified features. Based on the spatial geometric features, candidate gripping positions on the surface of the object to be sorted are determined. Based on the candidate gripping positions and preset geometric constraints, the spatial distribution and contact area of ​​the contact region are predicted, and the composite distribution of the contact force on the surface of the object to be sorted is determined. Based on the composite distribution, the mechanical stability value is calculated and combined with the semantic features to obtain the comprehensive adaptability value. Based on the positioning coordinates and the attitude parameters, the reachability of the candidate grasping positions is verified, feasible candidate grasping positions are determined, and the candidate grasping position with the largest comprehensive adaptability value is selected as the target grasping point. Based on the target grasping point, the grasping direction vector is determined, and the motion trajectory is calculated in combination with the attitude parameters to complete the grasping.

2. The method according to claim 1, characterized in that, The process involves acquiring multi-view images of the objects to be sorted, extracting corresponding spatial geometric features, and fusing them with semantic features to obtain fused features. It also involves extracting key surface points of the objects from different viewpoints and establishing spatial correspondences. Based on these spatial correspondences, geometric transformation parameters are calculated, and the fused features are aligned and reconstructed to obtain unified features, including: Multi-view images are obtained by acquiring multiple perspectives of the objects to be sorted. Depth is estimated from the multi-view images to obtain depth maps, and the corresponding 3D point cloud data is extracted as spatial geometric features. Semantic features are obtained by feature extraction from the multi-view images. The spatial geometric features and the semantic features are spliced ​​together according to spatial position to obtain fused features. Key point detection is performed on the surface of the object to be sorted in the multi-view image to obtain a set of surface key points. The similarity of the feature descriptors corresponding to the surface key points under different views is calculated, and feature matching is performed based on the similarity to obtain key point matching pairs. Spatial correspondence is established based on the image coordinates and depth information of different key points in the key point matching pairs under different views. Based on the 3D coordinate differences of the key point matching pairs in the spatial correspondence, the rotation matrix and translation vector are calculated as geometric transformation parameters. Based on the geometric transformation parameters, the spatial geometric features in the fused features are transformed to a unified reference coordinate system to complete coordinate alignment. The spatial geometric features after coordinate alignment are re-associated with the corresponding semantic features in spatial position, and the spatial index relationship of the feature vector is updated to obtain unified features.

3. The method according to claim 1, characterized in that, Based on the unified features, a transformation relationship between the image coordinate system and the robot coordinate system is established, and the coordinates of the object to be sorted are mapped from the image coordinate system to the robot coordinate system to obtain the positioning coordinates. Combined with the unified features, the posture parameters corresponding to the object to be sorted are calculated, including: Extract three-dimensional feature points from the surface of the object to be sorted from the unified features and obtain the coordinates of the three-dimensional feature points in the image coordinate system as the original coordinates. Set a calibration reference point in the robot coordinate system. Determine the origin position and coordinate axis direction of the robot coordinate system based on the three-dimensional coordinates of the calibration reference point. Solve the correspondence between the original coordinates and the coordinates in the robot coordinate system to determine the rotation transformation matrix and translation transformation vector. Construct a coordinate transformation matrix based on the rotation transformation matrix and the translation transformation vector as the transformation relationship. Substitute the coordinates of the centroid of the object to be sorted in the image coordinate system into the coordinate transformation matrix to perform matrix operations and obtain the positioning coordinates; Extract the normal vector distribution of the surface of the object to be sorted from the unified features and calculate the principal direction vector of the object surface. Substitute the direction component of the principal direction vector in the image coordinate system into the rotation transformation matrix to perform coordinate system transformation to obtain the direction component of the principal direction vector in the robot coordinate system. Calculate the rotation angle of the object to be sorted relative to each coordinate axis of the robot coordinate system based on the direction component and combine them to form the attitude parameters.

4. The method according to claim 1, characterized in that, Based on the spatial geometric features, candidate gripping positions on the surface of the object to be sorted are determined. Based on the candidate gripping positions and preset geometric constraints, the spatial distribution and contact area of ​​the contact region are predicted, and the composite distribution of contact forces on the surface of the object to be sorted is determined, including: The curvature distribution of the surface of the object to be sorted is extracted from the spatial geometric features in the unified features. Based on the curvature distribution, flat areas and raised areas on the surface of the object to be sorted are identified. The surface normal vectors of the flat areas and the raised areas are calculated and the surface areas that meet the grasping stability conditions are selected as candidate grasping positions based on the angle between the surface normal vectors and the horizontal plane. Obtain the contact surface geometry of the end effector under preset geometric constraints, project the contact surface geometry onto the surface area corresponding to each candidate gripping position, calculate the degree of contact between the contact surface and the surface of the object to be sorted based on the curvature distribution of the surface area corresponding to the contact surface geometry, determine the boundary contour of the actual contact area as the spatial distribution of the contact area based on the degree of contact, and perform area integration to obtain the contact area. The gripping force applied by the end effector under the preset geometric constraints is decomposed into local forces at each contact point within the contact area according to the contact area. Based on the local forces and the corresponding surface normal vectors, the normal and tangential components of the local forces on the surface of the object to be sorted are calculated. The normal and tangential components of all contact points within the contact area are vector-superimposed to obtain the composite distribution.

5. The method according to claim 1, characterized in that, The mechanical stability value is calculated based on the synthetic distribution, and the comprehensive fitness value is obtained by combining the semantic features, including: The spatial distribution of contact forces on the surface of the object to be sorted is obtained from the synthetic distribution, and the force field is reconstructed to obtain the force distribution field. Based on the force distribution field, the stress tensor distribution on the surface of the object to be sorted is calculated, and the resultant force and the point of application of the resultant force in the contact area are obtained by integration. The position vector of the point of application of the resultant force relative to the center of mass of the object to be sorted is calculated, and the resultant moment of the object to be sorted relative to the center of mass is obtained by combining the resultant force. The gravitational moment is obtained based on the mass and center of mass of the object to be sorted. The vector sum of the resultant moment and the gravitational moment is calculated, and it is determined whether the vector sum satisfies the moment equilibrium condition. If so, the vector sum is used as the mechanical stability value. Based on the semantic features in the unified features, extract the material properties and surface texture features of the object to be sorted. Based on the material properties, determine the elastic modulus and Poisson's ratio of the object to be sorted and calculate the surface deformation of the object under contact force. Based on the surface texture features, calculate the micro-protrusion distribution on the surface of the object to be sorted and determine the ratio of the actual contact area to the nominal contact area and record it as the first coefficient. The deformation adaptability value is obtained by correlating the mechanical stability value and the surface deformation value, and the contact effectiveness value is obtained by correlating the mechanical stability value and the first coefficient. The comprehensive adaptability value is obtained by weighted fusion of the deformation adaptability value and the contact effectiveness value.

6. The method according to claim 1, characterized in that, Based on the positioning coordinates and the attitude parameters, the reachability of candidate grasping positions is verified to determine feasible candidate grasping positions and select the candidate grasping position with the largest comprehensive adaptability value as the target grasping point, including: Obtain the link length parameters and joint rotation axis direction of the robot end effector and establish a mathematical mapping relationship between the robot joint angle and the spatial position of the end effector. Based on the mathematical mapping relationship, derive the inverse solution expression of the spatial position and attitude of the end effector to the joint angle. For each candidate gripping position, the positioning coordinates and attitude parameters corresponding to the current candidate gripping position are substituted into the inverse solving expression to calculate the target joint angle of each joint of the robot. The angle limit range of each joint of the robot is obtained and it is determined whether the target joint angle exceeds the angle limit range to obtain the first judgment result. The reachable distance range of the robot end effector is calculated based on the link length parameter and it is determined whether the positioning coordinates exceed the reachable distance range to obtain the second judgment result. Based on the first judgment result and the second judgment result, a feasible candidate gripping position is determined. Extract the comprehensive adaptability value corresponding to the feasible candidate crawling position, sort the feasible candidate crawling positions in descending order based on the comprehensive adaptability value to obtain a priority sequence, and take the first feasible candidate crawling position in the priority sequence as the target crawling point.

7. The method according to claim 1, characterized in that, Determining the grasping direction vector based on the target grasping point, calculating the motion trajectory based on the posture parameters, and completing the grasping process includes: The system acquires the position and attitude parameters of the target gripping point. Based on the attitude parameters, it determines the normal approach angle and tangential deflection angle of the robot end effector relative to the surface of the object to be sorted. Based on the normal approach angle, it calculates the unit vector along the normal of the surface of the object to be sorted. Based on the tangential deflection angle, it calculates the rotation correction vector in the tangential plane of the surface of the object to be sorted. Based on the unit vector and the rotation correction vector, it determines the initial gripping vector and, combined with the synthetic distribution, determines the cosine value of the angle between the direction of the contact force and the initial gripping vector. Based on the cosine value of the angle, it performs mechanical optimization adjustment on the initial gripping vector to obtain the gripping direction vector. Obtain the current position coordinates and current attitude angle of the robot's end effector, and calculate the spatial displacement vector of the position parameters from the current position coordinates to the target grasping point and the angular change of the attitude parameters from the current attitude angle to the target grasping point; A pose transformation matrix is ​​constructed based on the spatial displacement vector and the grasping direction vector, and singular value decomposition is performed to obtain translation and rotation components. A translation path is generated based on the translation components, and the path is corrected along the grasping direction vector to obtain a corrected translation path. A rotation path is generated based on the rotation components and the angle change, ensuring the orthogonality between the rotation axis and the grasping direction vector. The corrected translation path and the rotation path are interpolated by a fifth-order polynomial to obtain the motion trajectory, and the object to be sorted is moved along the motion trajectory to the target grasping point to complete the grasping of the object to be sorted.

8. A vision-based intelligent positioning and grasping system for industrial sorting robots, used to implement the method of any one of claims 1-7, characterized in that, include: The feature fusion unit is used to acquire multi-view images of the object to be sorted and extract the corresponding spatial geometric features and semantic features to fuse them into fused features. It also extracts the surface key points of the object to be sorted under different views and establishes spatial correspondence. Based on the spatial correspondence, it calculates geometric transformation parameters and performs coordinate alignment and feature reconstruction on the fused features to obtain unified features. The coordinate transformation unit is used to establish the transformation relationship between the image coordinate system and the robot coordinate system based on the unified feature, and to map the coordinates of the object to be sorted from the image coordinate system to the robot coordinate system to obtain the positioning coordinates and calculate the posture parameters corresponding to the object to be sorted in combination with the unified feature. The grasping evaluation unit is used to determine candidate grasping positions on the surface of the object to be sorted based on the spatial geometric features, predict the spatial distribution and contact area of ​​the contact region based on the candidate grasping positions and preset geometric constraints, determine the composite distribution of contact forces on the surface of the object to be sorted, calculate the mechanical stability value based on the composite distribution, and solve for the comprehensive adaptability value by combining the semantic features. The grasping execution unit is used to verify the reachability of candidate grasping positions based on the positioning coordinates and the attitude parameters, determine feasible candidate grasping positions and select the candidate grasping position with the largest comprehensive adaptability value as the target grasping point, determine the grasping direction vector based on the target grasping point, calculate the motion trajectory in combination with the attitude parameters and complete the grasping.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.