Robot production line article grabbing method and system based on visual positioning
Through multimodal data fusion and dynamic adaptive grid mapping, the problem of insufficient coordinate offset and material recognition accuracy of conveyor belts in the visual positioning method is solved, and high accuracy and stability of item grabbing on industrial production lines is achieved.
Patent Information
- Application Number
- CN202510629058.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing visual positioning methods ignore the dynamic coordinate offset of the conveyor belt on the industrial production line, resulting in insufficient grasping accuracy and the single visual mode is not very accurate in identifying the material of the item.
Multimodal data fusion and dynamic adaptive grid mapping are used to establish a dynamic mapping relationship between production lines, robots and vision sensors through a non-rigid coordinate system alignment algorithm, and combine hierarchical feature matching and geometric constraint optimization algorithm to achieve accurate coordinate compensation and item positioning.
It improves the accuracy and stability of item grabbing, can adapt to complex materials and dynamic environments, reduces the risk of calculation load and slippage, and improves the success rate of grabbing.
Smart Images

Figure CN120543631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision positioning, and in particular to a method and system for grasping items in a robot production line based on vision positioning. Background Art
[0002] With the advancement of industrial automation, industrial robots are increasingly used for grasping objects on production lines. Current visual positioning methods are mostly based on fixed visual positioning or single-modal image analysis, achieving grasping operations through static coordinate alignment. However, these methods ignore the dynamic coordinate offsets caused by the continuous motion of production line conveyor belts, and the accuracy of object material recognition using a single visual modality is insufficient. Therefore, it is crucial to design a method and system for robotic production line object grasping based on visual positioning. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for robot production line item grasping based on visual positioning, so as to achieve precise coordinate compensation through multimodal data fusion and dynamic adaptive grid mapping, and improve the accuracy of item positioning in combination with hierarchical feature matching.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] A method for grasping items on a robot production line based on visual positioning includes the following steps:
[0006] Collect multimodal image data of the production line and perform preprocessing operations to obtain raw image data; the raw image data includes: RGB images, near-infrared band images and polarized light images;
[0007] Perform dynamic adaptive grid mapping on the original image data to obtain a multi-level feature descriptor; the multi-level feature descriptor includes: object material, surface curvature distribution and spatial pose;
[0008] Based on multi-level feature descriptors, a non-rigid coordinate system alignment algorithm is used to establish a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system, and the visual sensor coordinate system, and to compensate for the coordinate offset of the conveyor belt motion in real time.
[0009] Perform hierarchical feature matching between multi-level feature descriptors and the preset object template library to identify the object category and extract the object's contour geometric features;
[0010] Based on the contour geometric features and surface curvature distribution, the three-dimensional pose and candidate grasping points of the object are obtained through the geometric constraint optimization algorithm;
[0011] Generate the robot's obstacle avoidance trajectory based on the candidate grasping points and the robot motion model to grasp the object.
[0012] Optionally, multimodal image data of the production line is collected and preprocessed to obtain raw image data, including:
[0013] The visible spectrum of three bands, 450-500nm, 520-580nm, and 600-680nm, was captured on a single camera.
[0014] The three visible spectra are adaptively fused through the constructed spectral response weight function to obtain the RGB image;
[0015] The near-infrared reflectivity is obtained through a high-frequency modulated 940nm pulsed laser matrix, and the near-infrared reflectivity is dynamically adjusted according to the reflectivity of the object material;
[0016] By formula The near-infrared reflectivity is subjected to roughness compensation to obtain the near-infrared band image; where ρ r is the near-infrared reflectivity before compensation, σ s is the surface roughness coefficient, θ is the incident angle;
[0017] A polarization singularity is generated by a vortex phase plate, and the topological charge of the polarization singularity is calculated by phase gradient integration.
[0018] Based on the topological charge number, the light intensity and Stokes parameters of the three phases are mapped to the same coordinate system to obtain the polarization map; the polarization map includes: linear polarization contrast map, circular polarization response map and topological singular point distribution map;
[0019] The polarization images are weightedly fused according to the polarization fusion index calculated from the Stokes parameters to obtain a polarized light image.
[0020] Optionally, the polarization images are weightedly fused according to the polarization fusion index calculated from the Stokes parameters to obtain a polarized light image, including:
[0021] By formula Calculate the polarization fusion index, where I max is the maximum light intensity, I min is the minimum light intensity, S0 and S k All are Stokes parameters;
[0022] Both the linear polarization contrast map and the circular polarization response map were normalized by the Stokes parameter;
[0023] Add exponential weights to the topological singular point regions in the topological singular point distribution map; the exponential weights are determined by the topological charge number;
[0024] According to the polarization fusion index, the normalized linear polarization contrast map, the normalized circular polarization response map, and the topological singularity distribution map with added exponential weights are fused at the pixel level to obtain a polarized light image; the weight coefficient in the pixel-level weighted fusion is determined by the material of the object.
[0025] Optionally, dynamic adaptive grid mapping is performed on the original image data to obtain a multi-level feature descriptor, including:
[0026] The original image data is divided into a triangular mesh with dynamically variable resolution; the resolution is determined by the conveyor belt speed, and the expression for the resolution is: Where V is the conveyor belt speed, V th is the speed threshold, V max is the maximum speed of the conveyor belt;
[0027] Based on the Stokes parameters of the triangular mesh, the formula and Calculate the linear polarization degree η and circular polarization degree ξ, and determine the material of the object based on the linear polarization degree and circular polarization degree;
[0028] Taking the vertex of the triangular mesh as the center, select the area points within the radius of 3 times the resolution, fit the quadratic surface by the least squares method, and calculate the principal curvature and mean curvature to obtain the surface curvature distribution;
[0029] Perform eigenvalue decomposition on the vertex and object centroid positions, and take the eigenvector corresponding to the maximum eigenvalue as the principal axis direction;
[0030] Align the principal axis direction with the template coordinate system, obtain the rotation matrix, and then solve it using the Kabsch algorithm to obtain the spatial pose.
[0031] Optionally, determining the material of the object according to the linear polarization degree and the circular polarization degree includes:
[0032] When η>0.5 and ξ<0.1, the material of the object is determined to be metal;
[0033] When η<0.3 and ξ>0.4, the material of the item is determined to be transparent;
[0034] When 0.3≤η≤0.5 and ξ<0.2, the material of the object is determined to be a rough non-metallic material.
[0035] Optionally, based on multi-level feature descriptors, a non-rigid coordinate system alignment algorithm is used to establish a dynamic mapping relationship between the production line conveyor coordinate system, the robot base coordinate system, and the vision sensor coordinate system, and to compensate for the coordinate offset of the conveyor motion in real time, including:
[0036] Embed RFID tags at equal intervals on the conveyor belt surface as dynamic origins, and use the belt motion direction as the coordinate axis to establish the production line conveyor belt coordinate system;
[0037] The robot base coordinate system is established based on the vibration compensation matrix generated by the robot's built-in three-axis accelerometer;
[0038] Establish the visual sensor coordinate system based on the temperature sensor array and the single camera intrinsic parameter matrix;
[0039] The three-dimensional coordinates are obtained by setting the laser tracking control points in the robot base coordinate system, and a cubic spline interpolation function is constructed according to the three-dimensional coordinates;
[0040] The cubic spline interpolation function is extended to the four-dimensional spline space according to the conveyor belt speed, and the registration objective function is obtained by iterative optimization using the Manhattan update strategy.
[0041] The conveyor belt is discretized into tetrahedral unit grids using the finite element method, and the linear elastic constitutive equation is constructed;
[0042] Based on the displacement field obtained by solving the linear elastic constitutive equation, the registration objective function is inversely compensated to the visual sensor coordinate system to obtain a dynamic mapping relationship.
[0043] Optionally, hierarchical feature matching is performed between the multi-level feature descriptors and a preset object template library to identify the object category and extract the object's contour geometric features, including:
[0044] Use a multispectral 3D scanner to scan standard objects and build an object template library;
[0045] By formula Calculate the Mahalanobis distance between the object to be identified and the material fingerprint in the object template library, where M l is the multimodal material fingerprint of the object to be identified, μ M is the mean vector of template materials in the item template library, Σ M is the template material covariance matrix of the item template library;
[0046] By formula Calculate the curvature distribution similarity between the object to be identified and the object template library; where γ is the scale factor, is the average curvature of the object, is the curvature of the i-th point of the object to be identified, is the curvature of the i-th point in the item template library, and N is the number of curvature matching points;
[0047] Construct physical constraints between the object to be identified and the object template library, and solve them using the LM algorithm to obtain the pose similarity;
[0048] The Mahalanobis distance, curvature distribution similarity and pose similarity are weightedly fused to obtain the matching confidence;
[0049] When the matching confidence is greater than 0.8, the category of the object to be identified is determined;
[0050] The initial contour of the object to be identified is generated based on the curvature topology map in the object template library, and the curvature-guided completion operation is performed on the occluded area of the object to be identified to obtain the contour geometric features.
[0051] Optionally, based on the contour geometry and surface curvature distribution, a geometric constraint optimization algorithm is used to obtain the object's 3D pose and candidate grasping points, including:
[0052] Perform Gaussian convolution on the mean curvature in the contour geometric features to generate a multi-scale curvature map and select the curvature extreme points;
[0053] The pose constraint energy function is constructed based on the extreme points of curvature and solved by manifold differentiation to obtain the three-dimensional pose.
[0054] Monte Carlo sampling is performed on the extreme points of curvature, and effective grasping points are selected through convex hull screening;
[0055] By formula Calculate the stability index of the effective grasping point, and select the effective grasping point whose stability index is greater than the preset stability threshold as the candidate grasping point, where F n is the normal force at the effective grasping point, F t is the tangential force at the effective grasping point, A c is the contact area of the effective grasping point, is the curvature of the effective grasping point.
[0056] Optionally, a robot obstacle avoidance trajectory is generated based on the candidate grasping points and the robot motion model to grasp the object, including:
[0057] Predict the trajectory of moving obstacles through the speed obstacle model;
[0058] Based on the motion trajectory, a B-spline initial path is generated in Cartesian space;
[0059] The joint torque constraints of the robot motion model are mapped to the B-spline initial path through the Jacobian matrix to generate the robot motion optimization function;
[0060] The robot motion optimization function is solved by the sequential quadratic programming algorithm to obtain the robot's obstacle avoidance trajectory.
[0061] A robot production line object grasping system based on visual positioning, comprising:
[0062] The image acquisition module is used to collect multimodal image data of the production line and perform preprocessing operations to obtain raw image data; the raw image data includes: RGB images, near-infrared band images and polarized light images;
[0063] The feature extraction module is used to perform dynamic adaptive grid mapping on the original image data to obtain multi-level feature descriptors; the multi-level feature descriptors include: object material, surface curvature distribution and spatial pose;
[0064] The coordinate compensation module is used to establish a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system, and the vision sensor coordinate system through a non-rigid coordinate system alignment algorithm based on multi-level feature descriptors, and to compensate for the coordinate offset of the conveyor belt movement in real time;
[0065] The object recognition module is used to perform hierarchical feature matching between multi-level feature descriptors and a preset object template library, identify the object category, and extract the object's contour geometric features;
[0066] The grasping point positioning module is used to obtain the three-dimensional pose of the object and candidate grasping points based on the contour geometric features and surface curvature distribution through a geometric constraint optimization algorithm;
[0067] The control grasping module is used to generate the robot's obstacle avoidance trajectory based on the candidate grasping points and the robot motion model to grasp the object.
[0068] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: the present invention provides a method for robotic production line object grasping based on visual positioning, the method comprising: collecting multimodal image data of the production line and performing preprocessing operations to obtain raw image data; performing dynamic adaptive grid mapping on the raw image data to obtain a multi-level feature descriptor; based on the multi-level feature descriptor, establishing a dynamic mapping relationship between the production line conveyor coordinate system, the robot base coordinate system, and the visual sensor coordinate system through a non-rigid coordinate system alignment algorithm, and compensating for the coordinate offset of the conveyor motion in real time; performing hierarchical feature matching on the multi-level feature descriptor with a preset object template library to identify the object category and extract the object's contour geometric features; based on the contour geometric features and surface curvature distribution, obtaining the object's three-dimensional pose and candidate grasping points through a geometric constraint optimization algorithm; generating a robot obstacle avoidance trajectory based on the candidate grasping points and the robot motion model to grasp the object. This method achieves precise coordinate compensation through multimodal data fusion and dynamic adaptive grid mapping, and improves the stability of the grasping point by combining hierarchical feature matching with a geometric constraint optimization algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0070] Figure 1 This is a flow chart of the method for grabbing items on a robot production line of the present invention;
[0071] Figure 2 This is a flow chart of multimodal image data acquisition of the present invention;
[0072] Figure 3 This is a flow chart of the coordinate system dynamic mapping of the present invention. DETAILED DESCRIPTION
[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0074] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0075] like Figure 1 As shown, the present invention provides a method for grasping items in a robot production line based on visual positioning, comprising the following steps:
[0076] Step 100: Collect multimodal image data of the production line and perform preprocessing operations to obtain raw image data; the raw image data includes: RGB image, near-infrared band image and polarized light image;
[0077] Step 200: Perform dynamic adaptive grid mapping on the original image data to obtain a multi-level feature descriptor; the multi-level feature descriptor includes: object material, surface curvature distribution and spatial pose;
[0078] Step 300: Based on the multi-level feature descriptor, a non-rigid coordinate system alignment algorithm is used to establish a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system, and the vision sensor coordinate system, and to compensate for the coordinate offset of the conveyor belt motion in real time;
[0079] Step 400: Perform hierarchical feature matching between the multi-level feature descriptors and the preset object template library to identify the object category and extract the object's contour geometric features;
[0080] Step 500: Based on the contour geometric features and surface curvature distribution, the three-dimensional pose of the object and candidate grasping points are obtained through a geometric constraint optimization algorithm;
[0081] Step 600: Generate a robot obstacle avoidance trajectory based on the candidate grasping points and the robot motion model to grasp the object.
[0082] like Figure 2 As shown, the multimodal image data of the production line is collected and preprocessed to obtain the original image data, including:
[0083] Step 101: Capture the visible spectrum of three wavelength bands, 450-500 nm, 520-580 nm, and 600-680 nm, respectively, using a single camera;
[0084] Specifically, the system captures three wavelength bands of the visible spectrum: 450-500nm (blue-cyan), 520-580nm (green-yellow), and 600-680nm (red-orange). The filter switching frequency is synchronized with the production line conveyor speed to ensure that the exposure time of each wavelength band adapts to dynamic scenes. These three wavelength bands cover the main reflective characteristics of the object's surface color while avoiding the peak wavelengths of common industrial light sources (such as LEDs), reducing light interference.
[0085] Step 102: Adaptively fuse the three visible spectra using the constructed spectral response weight function to obtain an RGB image;
[0086] Specifically, the spectral response weight function predetermines the relative sensitivity of each band to different materials through calibration experiments, and introduces the ambient light intensity to dynamically adjust the weight coefficient. In some embodiments, the metal surface has a strong reflection in the blue light band, so the blue light band weight coefficient is increased accordingly, while the rough material scatters more significantly in the red light band, so the red light band weight coefficient is increased accordingly. The adaptive fusion process uses a pixel-by-pixel weighted superposition algorithm, and the formula is: where ω i is the dynamic weight, and λ is the visible light band.
[0087] Step 103: Obtain near-infrared reflectivity through a high-frequency modulated 940nm pulsed laser matrix, and dynamically adjust the near-infrared reflectivity according to the reflectivity of the object material;
[0088] Specifically, a high-frequency modulated (>1MHz) 940nm pulsed laser matrix is used to project a structured light spot onto the surface of the object, and the reflected light intensity is captured by synchronously triggering the camera. IR The calculation formula is: Among them I fis the reflected light intensity, I b is the background light intensity, I r is the incident light intensity, η c is the camera response coefficient.
[0089] Step 104: By formula The near-infrared reflectivity is subjected to roughness compensation to obtain the near-infrared band image; where ρ r is the near-infrared reflectivity before compensation, σ s is the surface roughness coefficient, θ is the angle of incidence; the formula is obtained by tan 2 The (θ) term quantifies the scattering effect, making the compensated reflectivity closer to the ideal smooth surface value, thereby improving the material classification accuracy.
[0090] Step 105: Generate a polarization singular point by using a vortex phase plate, and obtain the topological charge number of the polarization singular point by phase gradient integral calculation;
[0091] Specifically, a vortex phase plate is used to introduce a spiral phase delay into the laser beam, generating a light field with a circular polarization singularity. The polarization state at the singularity is discontinuous, and its topological charge, representing the number of rotations of the phase around the singularity, is calculated using a phase gradient integration algorithm.
[0092] Step 106: Based on the topological charge number, the light intensities and Stokes parameters of the three phases are mapped to the same coordinate system to obtain a polarization map; the polarization map includes: a linear polarization contrast map, a circular polarization response map, and a topological singular point distribution map;
[0093] Specifically, based on the polarized light intensity measurements at three phases (0°, 120°, and 240°), the Stokes parameters S0 (total light intensity), S1 (linear polarization component), S2 (45° linear polarization component), and S3 (45° linear polarization component) are calculated. A linear polarization contrast image is obtained based on S0, S1, and S2, highlighting edges and textures. The formula is: The circular polarization response diagram is obtained based on S0 and S3, which can identify optically active materials. The formula is: The topological singularity distribution map marks the areas where the topological charge number is not equal to 0, which can locate surface defects or special structures.
[0094] Step 107: Perform weighted fusion on the polarization maps according to the polarization fusion index calculated from the Stokes parameters to obtain a polarized light image.
[0095] Specifically, first use the formula Calculate the polarization fusion index, where I max and I min is the maximum and minimum light intensity values in the scene, determined by global histogram statistics, The integrated intensity of linear polarization and circular polarization was quantified, and P was achieved through S0. F Normalization of P F The range of the value is [0,1]. Then the linear polarization contrast map and the circular polarization response map are normalized by the Stokes parameter. Then, an exponential weight is added to the topological singular point region in the topological singular point distribution map. The exponential weight is dynamically adjusted by the topological charge number, and the expression is: e α|q| , where α is an empirical coefficient (usually 0.5 to 1.0), and q is the topological charge number. This exponential weight can make the high topological charge number area obtain a higher weight during fusion, thereby enhancing the defect detection capability. Finally, according to the polarization fusion index, the normalized linear polarization contrast image, the normalized circular polarization response image, and the topological singular point distribution map after adding the exponential weight are pixel-wise weighted fused to obtain the polarized light image. The formula is: I p =ω1·LPC+ω2·CPR; wherein ω1 and ω2 are both weight coefficients determined by the material of the object. In some embodiments, when the object is made of metal, ω1=0.6 and ω2=0.4; when the object is made of transparent material, ω1=0.3 and ω2=0.7; when the object is made of rough non-metallic material, ω1=0.4 and ω2=0.6.
[0096] It should be noted that through the collaborative processing of multimodal data, comprehensive capture and optimization of spectral, near-infrared, and polarization information is achieved. Multi-band fusion enhances color and material differentiation; dynamic near-infrared adjustment improves reflectivity measurement accuracy; and polarization singularity analysis and Stokes parameter mapping provide in-depth characterization of surface microscopic properties. The resulting RGB, near-infrared, and polarized images form complementary information sources, providing high-precision, interference-resistant input data for subsequent feature extraction and coordinate compensation, significantly addressing the limitations of traditional methods in dynamic environments, complex materials, and occluded scenes.
[0097] Specifically, the steps for dynamically adaptively mapping the original image data to obtain a multi-level feature descriptor are as follows: the original image data is divided into a triangular mesh with dynamically variable resolution, and the mesh density is reduced to reduce the amount of computation while ensuring that key features can still be captured under the high-speed movement of the conveyor belt; the triangular mesh is generated using the Delaunay triangulation algorithm, and the side length of the triangular mesh is constrained by the resolution, which is determined by the conveyor belt speed. The expression for the resolution is: Where V is the conveyor belt speed, V th is the speed threshold, which in some embodiments is 1 m / s, V max is the maximum speed of the conveyor belt. Then based on the Stokes parameters of the triangular mesh, the formula and The linear polarization degree η and the circular polarization degree ξ are calculated, and the material of the object is determined based on the linear polarization degree and the circular polarization degree. In some embodiments, when η>0.5 and ξ<0.1, the material of the object is determined to be a metal material, when η<0.3 and ξ>0.4, the material of the object is determined to be a transparent material, and when 0.3≤η≤0.5 and ξ<0.2, the material of the object is determined to be a rough non-metallic material. Then, with the vertex of the triangular mesh as the center, the area points within a radius of 3 times the resolution (usually containing 10 to 20 points) are selected, and the quadratic surface is fitted by the least squares method to obtain the expression z=ax 2 +by 2 +cxy+dx+ey+f, where a, b, c, d, and f are coefficients, to construct an overdetermined system of equations Ax=b1, where A is the design matrix and x=[a,b,c,d,e,f] T , b1 is the vertex height value. The coefficients are solved through singular value decomposition, and then the principal curvature and mean curvature are calculated to obtain the surface curvature distribution. The covariance matrix of the vertex coordinate matrix (including the mesh vertices and the object centroid positions) is calculated, and then the covariance matrix is decomposed by eigenvalue. The eigenvector corresponding to the maximum eigenvalue is used as the principal axis direction. The principal axis direction is aligned with the template coordinate system. The rotation matrix is obtained by calculating the minimum directional error between the principal axis direction and the template principal axis direction. The rotation matrix is then solved using the Kabsch algorithm to obtain the spatial pose.
[0098] like Figure 3 As shown in the figure, based on multi-level feature descriptors, a non-rigid coordinate system alignment algorithm is used to establish a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system, and the visual sensor coordinate system, and to compensate for the coordinate offset of the conveyor belt motion in real time, including:
[0099] Step 301: Embed RFID tags at equal intervals on the surface of the conveyor belt as dynamic origins, and establish a production line conveyor belt coordinate system using the direction of conveyor belt movement as the coordinate axis;
[0100] Specifically, radio frequency identification (RFID) tags are embedded every 200 mm on the surface of the conveyor belt. The tags have a built-in unique ID and position code. When the tag moves with the conveyor belt into the field of view of the visual sensor, the absolute position of the tag is acquired in real time through a UHF reader / writer (operating frequency 860-960 MHz). The coordinate system of the production line conveyor belt is constructed with the direction of conveyor belt movement as the X-axis, the vertical direction as the Y-axis, and the normal direction as the Z-axis. The tag spacing is dynamically adjusted according to the maximum conveying speed and sampling frequency to ensure that the displacement of adjacent tags within the sampling period does not exceed the field of view.
[0101] Step 302: Establishing a robot base coordinate system based on a vibration compensation matrix generated by a three-axis accelerometer built into the robot;
[0102] Specifically, the tiny displacement of the robot base caused by mechanical vibration is monitored in real time by the built-in three-axis MEMS accelerometer. The acceleration data is filtered by Kalman filter to remove high-frequency noise, and then integrated to generate displacement increments to construct the vibration compensation matrix. Where Δx, Δy, and Δz are the vibration displacements. The origin is calibrated using a laser tracker.
[0103] Step 303: Establishing a visual sensor coordinate system based on the temperature sensor array and the single-camera intrinsic parameter matrix;
[0104] Step 304: Acquire three-dimensional coordinates by setting laser tracking control points in the robot base coordinate system, and construct a cubic spline interpolation function based on the three-dimensional coordinates;
[0105] Step 305: Expand the cubic spline interpolation function to a four-dimensional spline space according to the conveyor belt speed, and perform iterative optimization using the Manhattan update strategy to obtain the registration objective function;
[0106] Specifically, the conveyor belt speed is introduced as the fourth dimension of the cubic spline interpolation function. The Manhattan update strategy adopts the L1 norm constraint to optimize the objective function.
[0107] Step 306: discretize the conveyor belt into a tetrahedral unit grid using the finite element method, and construct a linear elastic constitutive equation;
[0108] Specifically, the conveyor belt is divided into a tetrahedral grid with a node number set based on geometric complexity, ranging from 1,000 to 5,000 nodes in some embodiments. The equation is: Ku = F, where K is the stiffness matrix, calculated based on the material elastic modulus and Poisson's ratio, u is the node displacement vector, and F is the external load vector, calculated based on the conveyor belt tension and the material weight.
[0109] Step 307: Based on the displacement field obtained by solving the linear elastic constitutive equation, the registration objective function is reversely compensated to the visual sensor coordinate system to obtain a dynamic mapping relationship.
[0110] It should be noted that the RFID tag provides an absolute position reference, the vibration compensation matrix suppresses mechanical jitter, and the four-dimensional spline and finite element method accurately describe the nonlinear motion of the conveyor belt, significantly enhancing the robustness and real-time performance of grasping in complex industrial environments.
[0111] Specifically, the steps for performing hierarchical feature matching between the multi-level feature descriptor and the preset object template library, identifying the object category and extracting the object's contour geometric features are as follows:
[0112] First, a multispectral 3D scanner is used to scan the standard object to obtain multispectral point cloud data (including RGB, near-infrared reflectivity, polarization information) and high-precision geometric models to build an object template library. The scanner captures the surface details of the standard object with a resolution of 0.1mm, and aligns the multi-view data through the calibration plate to generate a curvature topology map (storing the main curvature, mean curvature and curvature derivative) and material fingerprints (including Stokes parameters, near-infrared reflectivity, polarization fusion index). The template library is stored by item category, and each category contains at least 10 instances to cover morphological variations. Then, the formula Calculate the Mahalanobis distance D between the object to be identified and the material fingerprint in the object template library M , where M l is the multimodal material fingerprint of the object to be identified, μ M is the mean vector of template materials in the item template library, Σ M is the template material covariance matrix of the item template library. At the same time, through the formula Calculate the curvature distribution similarity S between the object to be identified and the object template library C ; where γ is the scale factor, which is 0.1, is the average curvature of the object, is the curvature of the i-th point of the object to be identified, is the curvature of the i-th point in the object template library, and N is the number of curvature matching points. Next, physical constraints are established between the object to be identified and the object template library, including mass conservation (volume error <2%), moment of inertia matching (principal axis direction deviation <5°), and surface curvature continuity. An energy function is generated, and the LM algorithm is used to adjust the damping factor for iterative optimization. After the residual error drops to the convergence threshold, the pose similarity is output. The Mahalanobis distance, curvature distribution similarity, and pose similarity are weighted and fused with weights of 0.6, 0.2, and 0.2 to obtain the matching confidence. When the matching confidence is greater than 0.8, the object to be identified is classified as the corresponding item category. Simultaneously, an initial contour of the object to be identified is generated based on the curvature topology map in the object template library. Using the Poisson reconstruction algorithm, the occluded area is filled using the curvature gradient as a guiding field to ensure the curvature continuity of the occluded area, thereby obtaining the contour geometric features.
[0113] It is important to note that multi-dimensional feature matching, combining material fingerprints, curvature distribution, and pose constraints, effectively overcomes the limitations of single features in noisy, reflective, or occluded environments. The dynamic weight allocation strategy flexibly adjusts to the needs of different scenarios, balancing material sensitivity and geometric accuracy. Physical constraint optimization and curvature-guided contour completion ensure high-quality reconstruction and accurate pose estimation in occluded areas, providing reliable spatial information support for robotic grasping. This improves recognition robustness while enabling efficient and accurate processing of a wide range of industrial objects.
[0114] Specifically, based on the contour geometric features and surface curvature distribution, the specific steps for obtaining the object's 3D pose and candidate grasping points through the geometric constraint optimization algorithm are as follows:
[0115] First, the average curvature in the contour geometric features is subjected to Gaussian convolution operation, and a multi-scale curvature map is generated using convolution kernels of different scales. The curvature extreme points are extracted through scale space extreme value detection (such as non-maximum suppression) and noise interference is eliminated to retain significant feature points. Then, the pose constraint energy function is constructed based on the curvature extreme points. Where R and t are the rotation matrix and translation vector, p i and q i are the corresponding points of the object and the scene point cloud, κ 1j and κ 2j The principal curvature difference term is solved iteratively through manifold differentiation to minimize the energy function to obtain the 3D pose. Monte Carlo sampling is performed on the extreme points of curvature to generate a large number of random grasping directions. The points located on the concave surface of the object or in the geometrically inaccessible area are eliminated through convex hull screening, and the potential grasping points in the convex area are retained as valid grasping points. Calculate the stability index of the effective grasping point, and select the effective grasping point whose stability index is greater than the preset stability threshold as the candidate grasping point, where F n is the normal force at the effective grasping point, F t is the tangential force at the effective grasping point, A c is the contact area of the effective grasping point, is the curvature of the effective grasping point.
[0116] It should be noted that the multi-scale curvature map based on Gaussian convolution can effectively suppress noise interference while preserving key geometric features. Combining manifold differential optimization with physical constraints improves the robustness of pose estimation in dynamic environments. Monte Carlo sampling and convex hull screening work together to quickly locate the reachable grasping area. A comprehensive assessment of normal force, tangential force, and curvature significantly improves grasping reliability and significantly reduces the risk of slippage or instability. While improving the robot's grasping accuracy and success rate, it also enhances its adaptability to objects with diverse shapes and surface characteristics.
[0117] Specifically, the steps for generating the robot's obstacle avoidance trajectory based on the candidate grasping points and the robot motion model are as follows: predicting the trajectory of the moving obstacle using the velocity obstacle model, assuming that the obstacle moves at a constant speed at the current velocity, constructing the obstacle's velocity cone using the space-time dilation method, and determining the space-time region that the robot should avoid. Then, in Cartesian space, with the candidate grasping points as the target, an initial B-spline path is generated. Its control points consist of the starting point, the obstacle avoidance key point, and the end point. The path smoothness is ensured by cubic B-splines. The robot's kinematics and joint motion constraints are then mapped to Cartesian space using the Jacobian matrix to generate the robot's motion optimization function. Finally, the sequential quadratic programming (SQP) algorithm is used to solve the function and obtain the robot's obstacle avoidance trajectory.
[0118] It should be noted that the velocity obstacle model can predict the movement trends of moving obstacles in real time. Combined with B-spline path generation technology, it enhances the smoothness and adjustability of the initial path. Using the Jacobian matrix to map joint torque constraints effectively balances motion efficiency and mechanical load safety. The sequential quadratic programming algorithm integrates obstacle avoidance distance, path smoothness, and dynamic constraints to quickly solve for the optimal trajectory under complex constraints.
[0119] The present invention also provides a robot production line object grasping system based on visual positioning, comprising:
[0120] The image acquisition module is used to collect multimodal image data of the production line and perform preprocessing operations to obtain raw image data; the raw image data includes: RGB images, near-infrared band images and polarized light images;
[0121] The feature extraction module is used to perform dynamic adaptive grid mapping on the original image data to obtain multi-level feature descriptors; the multi-level feature descriptors include: object material, surface curvature distribution and spatial pose;
[0122] The coordinate compensation module is used to establish a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system, and the vision sensor coordinate system through a non-rigid coordinate system alignment algorithm based on multi-level feature descriptors, and to compensate for the coordinate offset of the conveyor belt movement in real time;
[0123] The object recognition module is used to perform hierarchical feature matching between multi-level feature descriptors and a preset object template library, identify the object category, and extract the object's contour geometric features;
[0124] The grasping point positioning module is used to obtain the three-dimensional pose of the object and candidate grasping points based on the contour geometric features and surface curvature distribution through a geometric constraint optimization algorithm;
[0125] The control grasping module is used to generate the robot's obstacle avoidance trajectory based on the candidate grasping points and the robot motion model to grasp the object.
[0126] The beneficial effects of the present invention are as follows:
[0127] 1) By fusing RGB, near-infrared, and polarized light images, the system overcomes the limitations of a single modality in identifying reflective, transparent, or rough materials, significantly enhances the ability to analyze complex surface characteristics, and improves the robustness of material classification and feature extraction.
[0128] 2) Real-time adjustment of the triangle mesh resolution based on the conveyor belt speed ensures the capture of key features while effectively reducing the computational load, adapting to the efficient processing requirements of high-speed production lines.
[0129] 3) By discretizing the conveyor belt model and constructing a linear elastic constitutive equation, combined with four-dimensional spline interpolation and Manhattan optimization strategy, the effect of conveyor belt deformation on visual positioning is inversely compensated, improving the coordinate mapping accuracy in dynamic environments.
[0130] 4) By integrating material fingerprints (Mahalanobis distance), curvature distribution similarity, and pose constraints, and using weighted confidence and Poisson reconstruction techniques, we effectively solve the contour completion problem in occluded scenes and improve the accuracy of 3D pose estimation.
[0131] 5) Based on multi-scale curvature analysis, manifold differential optimization and Monte Carlo sampling, combined with the mechanical stability model (normal force / tangential force ratio, contact area and curvature), the slip risk is reduced and the grasping success rate is improved.
[0132] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0133] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A method for grasping items in a robot production line based on visual positioning, characterized in that: The steps include: Collecting multimodal image data of the production line and performing preprocessing operations to obtain original image data; the original image data includes: RGB image, near-infrared band image and polarized light image; Performing dynamic adaptive grid mapping on the original image data to obtain a multi-level feature descriptor; the multi-level feature descriptor includes: object material, surface curvature distribution and spatial pose; Based on the multi-level feature descriptors, a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system and the visual sensor coordinate system is established through a non-rigid coordinate system alignment algorithm, and the coordinate offset of the conveyor belt motion is compensated in real time; Performing hierarchical feature matching between the multi-level feature descriptor and a preset object template library to identify the object category and extract the contour geometric features of the object; Based on the contour geometric features and the surface curvature distribution, a geometric constraint optimization algorithm is used to obtain the three-dimensional pose and candidate grasping points of the object; The robot's obstacle avoidance trajectory is generated according to the candidate grasping points and the robot motion model to grasp the object.
2. The method for grabbing items on a robot production line based on visual positioning according to claim 1 is characterized in that: Collect multimodal image data of the production line and perform preprocessing operations to obtain raw image data, including: The visible spectrum of three bands, 450-500nm, 520-580nm, and 600-680nm, was captured on a single camera. Adaptively fusing the three visible spectra using a constructed spectral response weight function to obtain the RGB image; The near-infrared reflectivity is obtained by a high-frequency modulated 940nm pulsed laser matrix, and the near-infrared reflectivity is dynamically adjusted according to the reflectivity of the object material; By formula Perform roughness compensation on the near-infrared reflectivity to obtain the near-infrared band image; wherein, ρ r is the near-infrared reflectivity before compensation, σ s is the surface roughness coefficient, θ is the incident angle; Generate a polarization singular point by a vortex phase plate, and obtain the topological charge number of the polarization singular point by phase gradient integral calculation; Based on the topological charge number, the light intensity and Stokes parameters of the three phases are mapped to the same coordinate system to obtain a polarization map; the polarization map includes: a linear polarization contrast map, a circular polarization response map and a topological singular point distribution map; The polarization images are weightedly fused according to the polarization fusion index calculated from the Stokes parameters to obtain the polarized light image.
3. The method for grabbing items on a robot production line based on visual positioning according to claim 2, characterized in that: The polarization images are weightedly fused according to the polarization fusion index calculated from the Stokes parameters to obtain the polarized light image, including: By formula Calculate the polarization fusion index, where I max is the maximum light intensity, I min is the minimum light intensity, S0 and S k All are Stokes parameters; Normalizing both the linear polarization contrast image and the circular polarization response image by using Stokes parameters; adding an exponential weight to a topological singular point region in the topological singular point distribution map; the exponential weight is determined by the topological charge number; According to the polarization fusion index, the normalized linear polarization contrast map, the normalized circular polarization response map, and the topological singular point distribution map after adding the index weight are subjected to pixel-level weighted fusion to obtain the polarized light image; the weight coefficient in the pixel-level weighted fusion is determined by the material of the object.
4. The method for grabbing items on a robot production line based on visual positioning according to claim 1, characterized in that: Performing dynamic adaptive grid mapping on the original image data to obtain a multi-level feature descriptor, including: The original image data is divided into a triangular grid with dynamically variable resolution; the resolution is determined by the conveyor belt speed, and the expression for the resolution is: Where V is the conveyor belt speed, V th is the speed threshold, V max is the maximum speed of the conveyor belt; Based on the Stokes parameters of the triangular mesh, the formula and Calculating the linear polarization degree η and the circular polarization degree ξ, and determining the material of the object based on the linear polarization degree and the circular polarization degree; Taking the vertex of the triangular mesh as the center, selecting a region point within a radius of 3 times the resolution, fitting a quadratic surface by the least squares method, and calculating the principal curvature and the mean curvature to obtain the surface curvature distribution; Performing eigenvalue decomposition on the vertex and object centroid positions, and taking the eigenvector corresponding to the maximum eigenvalue as the principal axis direction; The principal axis direction is aligned with the template coordinate system, and the rotation matrix is obtained and then solved by the Kabsch algorithm to obtain the spatial pose.
5. The method for grabbing items on a robot production line based on visual positioning according to claim 4 is characterized in that: Determining the material of the object according to the linear polarization degree and the circular polarization degree includes: When η>0.5 and ξ<0.1, the material of the object is determined to be metal; When η<0.3 and ξ>0.4, the material of the object is determined to be transparent; When 0.3≤η≤0.5 and ξ<0.2, the material of the object is determined to be a rough non-metallic material.
6. The method for grabbing items on a robot production line based on visual positioning according to claim 1, characterized in that: Based on the multi-level feature descriptors, a dynamic mapping relationship between the production line conveyor coordinate system, the robot base coordinate system, and the visual sensor coordinate system is established through a non-rigid coordinate system alignment algorithm, and the coordinate offset of the conveyor motion is compensated in real time, including: Embed RFID tags at equal intervals on the surface of the conveyor belt as dynamic origins, and establish the conveyor belt coordinate system of the production line using the direction of movement of the conveyor belt as the coordinate axis; Establishing the robot base coordinate system according to the vibration compensation matrix generated by the three-axis accelerometer built into the robot; Establishing the visual sensor coordinate system according to the temperature sensor array and the single camera intrinsic parameter matrix; Acquire three-dimensional coordinates by setting laser tracking control points in the robot base coordinate system, and construct a cubic spline interpolation function according to the three-dimensional coordinates; The cubic spline interpolation function is extended to a four-dimensional spline space according to the conveyor belt speed, and iteratively optimized using a Manhattan update strategy to obtain a registration objective function; The conveyor belt is discretized into a tetrahedral unit grid by a finite element method, and a linear elastic constitutive equation is constructed; Based on the displacement field obtained by solving the linear elastic constitutive equation, the registration objective function is reversely compensated to the visual sensor coordinate system to obtain the dynamic mapping relationship.
7. The method for grabbing items on a robot production line based on visual positioning according to claim 1, characterized in that: Performing hierarchical feature matching on the multi-level feature descriptor and a preset object template library to identify the object category and extract the contour geometric features of the object, including: Scanning standard objects using a multispectral 3D scanner to build the object template library; By formula Calculate the Mahalanobis distance between the object to be identified and the material fingerprint in the object template library, where M l is the multimodal material fingerprint of the object to be identified, μ M is the mean vector of template materials in the item template library, Σ M is the template material covariance matrix of the item template library; By formula Calculate the similarity of the curvature distribution of the object to be identified and the object template library; where γ is the scale factor, is the average curvature of the object, is the curvature of the i-th point of the object to be identified, is the curvature of the i-th point in the item template library, and N is the number of curvature matching points; Constructing physical constraints between the object to be identified and the object template library, and solving them using the LM algorithm to obtain posture similarity; Performing weighted fusion on the Mahalanobis distance, the curvature distribution similarity, and the posture similarity to obtain a matching confidence; When the matching confidence is greater than 0.8, determining the category of the object to be identified; An initial outline of the object to be identified is generated based on the curvature topology map in the object template library, and a curvature-guided completion operation is performed on the occluded area of the object to be identified to obtain the outline geometric features.
8. The method for grabbing items on a robot production line based on visual positioning according to claim 1, characterized in that: Based on the contour geometric features and the surface curvature distribution, a geometric constraint optimization algorithm is used to obtain the three-dimensional pose and candidate grasping points of the object, including: Performing a Gaussian convolution operation on the average curvature in the contour geometric features to generate a multi-scale curvature map and select curvature extreme points; Constructing a posture constraint energy function according to the curvature extreme point and solving it through manifold differentiation to obtain the three-dimensional posture; Performing Monte Carlo sampling on the curvature extreme points and selecting effective grasping points through convex hull screening; By formula Calculate the stability index of the effective grasping point, and take the effective grasping point whose stability index is greater than the preset stability threshold as the candidate grasping point, where F n is the normal force at the effective grasping point, F t is the tangential force at the effective grasping point, A c is the contact area of the effective grasping point, is the curvature of the effective grasping point.
9. The method for grabbing items on a robot production line based on visual positioning according to claim 1, characterized in that: Generating a robot obstacle avoidance trajectory based on the candidate grasping points and the robot motion model to grasp the object includes: Predict the trajectory of moving obstacles through the speed obstacle model; Based on the motion trajectory, generating a B-spline initial path in Cartesian space; Mapping the joint torque constraints of the robot motion model to the B-spline initial path through the Jacobian matrix to generate a robot motion optimization function; The robot motion optimization function is solved by a sequential quadratic programming algorithm to obtain the robot's obstacle avoidance trajectory.
10. A robot production line object grasping system based on visual positioning, characterized in that: include: Image acquisition module, used to collect multimodal image data of the production line and perform preprocessing operations to obtain original image data; The original image data includes: RGB image, near infrared band image and polarized light image; A feature extraction module is used to perform dynamic adaptive grid mapping on the original image data to obtain a multi-level feature descriptor; the multi-level feature descriptor includes: object material, surface curvature distribution and spatial pose; A coordinate compensation module is used to establish a dynamic mapping relationship between the production line conveyor belt coordinate system, the robot base coordinate system and the visual sensor coordinate system based on the multi-level feature descriptor through a non-rigid coordinate system alignment algorithm, and to compensate for the coordinate offset of the conveyor belt movement in real time; An object recognition module is used to perform hierarchical feature matching between the multi-level feature descriptors and a preset object template library, identify the object category, and extract the outline geometric features of the object; A grasping point positioning module, configured to obtain the three-dimensional pose and candidate grasping points of the object through a geometric constraint optimization algorithm based on the contour geometric features and the surface curvature distribution; The control grasping module is used to generate the robot's obstacle avoidance trajectory according to the candidate grasping points and the robot motion model to grasp the object.
Citation Information
Cited By
Laser auxiliary positioning system and method based on intelligent control
CN121163378A
Method and system for recognizing and positioning target grabbed by robot based on knowledge graph
CN121361102A
Autonomous navigation and grabbing control method and system for intelligent robot with body
CN121733592A
Autonomous navigation and grasping control method and system for body-aware robot
CN121733592B