Target collaborative control method and system based on deep learning
Through real-time multi-view image data processing and deep learning model, the problem of feature missing and misidentification of traditional visual guidance systems in complex scenarios is solved, accurate three-dimensional operation and stable object capture are achieved, and operation success rate and efficiency are improved.
Patent Information
- Application Number
- CN202510864673.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Traditional visual guidance operating systems have problems with missing features or misidentification in object occlusion, lighting changes, or unstructured environments, and the end effector control strategy fails to adapt to dynamically changing operation scenarios, resulting in clamping failure or surface damage.
Real-time multi-view image data is obtained through the image acquisition device, the improved YOLOv8 algorithm is used to identify objects, generate three-dimensional coordinate data, and calibrate and pose solution in combination with deep learning models, and generate control instructions to drive the operation execution device.
It realizes end-to-end optimization from multi-source heterogeneous visual data to precise mechanical operation, significantly improving object recognition integrity and operation success rate, reducing spatial positioning errors, ensuring operation stability and avoiding physical damage.
Smart Images

Figure CN120370719B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a target collaborative control method and system based on deep learning. Background Art
[0002] In the field of industrial automation, traditional technical solutions have significant limitations in meeting the demand for precise manipulation of objects in complex scenarios. Existing vision-guided operating systems mostly use monocular cameras or fixed-view sensor arrays, which leads to feature loss or misidentification problems when objects are occluded, lighting changes, or in unstructured environments. Although typical deep learning-based target detection methods can extract two-dimensional image features, they lack a spatial dimension association mechanism and are difficult to directly map to a three-dimensional operating space. For grasping tasks that require precise posture control, traditional solutions usually rely on pre-modeled object template libraries or offline calibration parameters, which cannot adapt to dynamically changing work scenarios. In addition, the end-effector control strategy generally adopts a fixed parameter drive mode, and has not established a correlation model between the physical properties of the object and the operating force, resulting in easy clamping failure or surface damage when handling objects of different materials, shapes or weights. Summary of the Invention
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a target collaborative control method based on deep learning, the method comprising:
[0004] Acquire real-time multi-view image data of the target area through an image acquisition device, wherein the real-time multi-view image data includes visual feature information of the object to be operated;
[0005] Invoking a real-time recognition model to perform object feature recognition on the real-time multi-view image data to obtain feature distribution data of the object to be operated, the feature distribution data including category identification information and boundary area coordinate information of the object to be operated;
[0006] Calibrate the characteristic distribution data to generate coordinate data of the object to be operated in a three-dimensional space coordinate system;
[0007] Performing a three-dimensional posture solution process based on the coordinate data to obtain a posture parameter set of the object to be operated, wherein the posture parameter set includes spatial position parameters and posture angle parameters of the object to be operated;
[0008] A control instruction is generated based on the posture parameter set and the object physical property parameters corresponding to the category identification information, and the operation execution device is driven to complete the object grasping action of the object to be operated. The control instruction includes a path control signal, an end effector posture control signal and a clamping force control signal based on the category identification information.
[0009] On the other hand, an embodiment of the present invention also provides a target collaborative control system based on deep learning, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0010] Based on the above aspects, the embodiments of the present invention achieve end-to-end optimization from multi-source heterogeneous visual data to precise mechanical operation. The dynamic multi-view image fusion mechanism is adopted to effectively overcome the occlusion defects of traditional monocular vision. The cross-view feature correlation extracted by the deep learning model significantly improves the completeness of object recognition in complex scenes. By combining the camera array calibration parameters and the deep learning error compensation network, the two-dimensional detection results are accurately mapped to the three-dimensional coordinate system, reducing the spatial positioning error to the sub-centimeter level. The closed-loop control system based on physical property parameters establishes a mapping model between object category and operation force, enabling the end effector to adaptively adjust the clamping strategy, while ensuring operational stability and avoiding physical damage to fragile objects. Through the real-time data interaction between the posture solution module and the motion planner, the dynamic optimization of the operation path and the synchronous response of the actuator are realized, which significantly improves the operation success rate and work efficiency under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic diagram of the execution flow of the target collaborative control method based on deep learning provided in an embodiment of the present invention.
[0012] Figure 2 Schematic diagram of exemplary hardware and software components of a deep learning-based target collaborative control system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a deep learning-based target collaborative control method provided by an embodiment of the present invention. The deep learning-based target collaborative control method is introduced in detail below.
[0014] Step S110: acquiring real-time multi-view image data of the target area through an image acquisition device, wherein the real-time multi-view image data includes visual feature information of the object to be operated.
[0015] In printed circuit board assembly (PCBA) production scenarios, the removal of auxiliary components from certain components requires accurate information about the components within the target area. This can be achieved using high-precision industrial cameras as image acquisition devices. These industrial cameras are selected based on parameters such as resolution, frame rate, and sensitivity to ensure clear and fast image capture.
[0016] For example, cameras are positioned at different locations and angles based on the actual PCBA production line. For example, a camera can be installed directly above the production line to capture the top visual features of components, while multiple cameras can be installed at different angles to the side to capture features such as the component's side profile. This way, capturing images from multiple perspectives allows for more comprehensive visual information about the components being processed.
[0017] The camera's image acquisition frequency needs to be dynamically adjusted based on the production line's speed and component movement. If the line is running fast, the camera acquisition frequency should be increased accordingly to ensure that the captured images reflect the component's status in real time. Conversely, if the line is running slowly, the acquisition frequency can be appropriately reduced to reduce data volume and processing burden. The captured real-time multi-view image data contains visual features such as component color, shape, and texture.
[0018] Step S120: calling a real-time recognition model to perform object feature recognition on the real-time multi-view image data to obtain feature distribution data of the object to be operated, wherein the feature distribution data includes category identification information and boundary area coordinate information of the object to be operated.
[0019] After acquiring real-time multi-view image data, it is necessary to call a real-time recognition model to process the data to identify the feature distribution of the object to be operated. In this embodiment, an improved YOLOv8 algorithm is used as the real-time recognition model.
[0020] Step S121: Perform parallel multi-scale feature extraction on the real-time multi-view image data through the multi-level dilated convolutional network in the real-time recognition model to generate a multi-level feature set containing feature maps of different resolutions, wherein the dilation rate of the dilated convolution kernel of each level is dynamically adjusted according to the inverse of the corresponding input image resolution to match the change in object size.
[0021] The multi-level dilated convolutional network in the real-time recognition model consists of multiple levels of dilated convolutional layers, each with a different dilation rate. The dilation rate is dynamically adjusted based on the inverse of the input image resolution.
[0022] Assuming the input real-time multi-view image data has different resolutions, for higher-resolution images, the reciprocal is smaller, and the corresponding dilation rate of the dilated convolution kernel is also smaller. This is because high-resolution images contain more detailed information, and a smaller dilation rate allows for more precise extraction of these detailed features. For example, for a high-resolution image block A with a resolution of R_A, the corresponding dilation rate is calculated as E_A according to the formula. Using the dilated convolution kernel with a dilation rate of E_A, the image block A is convolved to extract its detailed features.
[0023] For lower-resolution images, the reciprocal is larger, and the dilation rate of the dilated convolution kernel increases accordingly. This is to expand the receptive field of the convolution kernel and capture more macroscopic features. For example, for a low-resolution image block B with a resolution of R_B and a calculated dilation rate of E_B (E_B > E_A), convolution of image block B with a dilated convolution kernel with a dilation rate of E_B can capture more macroscopic feature information.
[0024] During parallel processing, dilated convolutional layers at different levels simultaneously perform convolution operations on the input real-time multi-view image data. Each convolution operation is performed independently, which improves processing efficiency. These convolution operations generate a multi-level feature set containing feature maps of varying resolutions, reflecting the feature information of the component being manipulated at various scales and resolutions.
[0025] Step S122: Input the multi-level feature set into the cross-level feature pyramid in the real-time recognition model, and enhance the resolution of the low-level large receptive field feature map through bilinear interpolation so that its size is aligned with the high-level small receptive field feature map, and perform channel-by-channel weighted fusion on the aligned feature map to generate an enhanced feature map that fuses global semantics and local details.
[0026] After the multi-level feature set is generated, it is input into the cross-level feature pyramid in the real-time recognition model. The role of the cross-level feature pyramid is to fuse the information of feature maps at different levels to obtain a more comprehensive and accurate feature representation.
[0027] First, the low-level, large receptive field feature maps are processed. Since the low-level feature maps have low resolution, bilinear interpolation is used to increase their resolution to enable effective fusion with the high-level, small receptive field feature maps. Bilinear interpolation calculates new pixel values based on a weighted average of surrounding pixels.
[0028] Assume that the low-level feature map with a large receptive field is F_L, and the high-level feature map with a small receptive field is F_H. Using bilinear interpolation, we increase the resolution of F_L to the same size as F_H, resulting in a new feature map F_L'. During interpolation, the value of each pixel in F_L' is calculated based on the grayscale values and positional relationships of the surrounding pixels in F_L.
[0029] Then, a channel-by-channel weighted fusion is performed on the aligned feature maps F_L' and F_H. For each channel, a weight is assigned. The weight is determined based on the importance of the channel features. These weights can be determined by training a small neural network or using statistical methods. Assuming that the weight of channel i is w_i, the value of the fused feature map in channel i is F_fused_i=w_i*F_L'_i+(1-w_i)*F_H_i, where F_L'_i and F_H_i are the values of F_L' and F_H in channel i, respectively. Through this channel-by-channel weighted fusion method, an enhanced feature map that fuses global semantics and local details is generated. The enhanced feature map integrates the information of feature maps at different levels, which is more conducive to subsequent object recognition and positioning.
[0030] Step S123: Apply channel dimension attention weighting and spatial dimension attention weighting to the enhanced feature map to obtain a weighted feature map, wherein the channel attention weight is generated by a two-layer fully connected network after global average pooling, and the spatial attention weight is generated by the spatial compression feature of a single-layer convolution kernel.
[0031] In order to further highlight the important feature information in the enhanced feature map, it is necessary to apply channel dimension attention weighting and spatial dimension attention weighting to the enhanced feature map.
[0032] To weight attention in the channel dimension, we first perform global average pooling on the enhanced feature map. Global average pooling averages the feature map across the spatial dimensions to obtain the average value for each channel. Assuming the enhanced feature map is F_enhanced, its dimensions are C×H×W (C is the number of channels, H is the height, and W is the width). After global average pooling, we obtain a vector V of length C, where each element V_c of V is the average value of F_enhanced for all pixels in channel c.
[0033] Next, vector V is input into a two-layer fully connected network. This network consists of an input layer, a hidden layer, and an output layer. The input layer receives vector V, and after a nonlinear transformation in the hidden layer, the output layer outputs a channel attention weight vector W_c. Each element W_c_i in the channel attention weight vector W_c represents the importance of channel i.
[0034] To weight spatial attention, a single-layer convolution kernel is used to generate spatially compressed features on the enhanced feature map. Parameters such as the kernel size and stride are set based on the actual situation. Assuming the kernel size is K, a convolution operation is performed on the enhanced feature map F_enhanced to generate a spatially compressed feature map F_spatial. The spatial attention weight matrix W_s is then calculated based on the value of F_spatial. Each element W_s_ij of the spatial attention weight matrix W_s represents the importance of position (i, j) in the feature map.
[0035] Finally, the channel attention weight vector W_c and the spatial attention weight matrix W_s are applied to the enhanced feature map F_enhanced. For each element F_enhanced_ijk in the enhanced feature map F_enhanced (where i represents the row, j represents the column, and k represents the channel), the weighted feature map element F_weighted_ijk = W_c_k * W_s_ij * F_enhanced_ijk. This yields a weighted feature map that emphasizes important features and suppresses unimportant ones.
[0036] Step S124: Divide the weighted feature map into a global classification branch and a local segmentation branch along the channel dimension, wherein the global classification branch is input into the multi-layer perceptron classifier after dimensionality reduction through global maximum pooling, outputs the category identification information and corresponding confidence of the object to be operated, and filters the category identification information with confidence lower than the dynamic threshold.
[0037] The weighted feature map is divided into a global classification branch and a local segmentation branch along the channel dimension. The purpose of the global classification branch is to classify and identify the object to be operated.
[0038] First, perform global max pooling on the feature map of the global classification branch to reduce its dimensionality. Global max pooling takes the maximum value of the feature map in the spatial dimension, thereby reducing the dimensionality of the feature map. Assume that the feature map of the global classification branch is F_global, and its size is C_g × H_g × W_g. After global max pooling, a vector V_global of length C_g is obtained.
[0039] The vector V_global is then input into a multilayer perceptron classifier. The multilayer perceptron classifier consists of multiple neural layers, including an input layer, a hidden layer, and an output layer. The input layer receives the vector V_global, and after a nonlinear transformation in the hidden layer, the output layer outputs the class identification information and the corresponding confidence level of the object being operated on. The class identification information can be represented by a class vector, where each element corresponds to a class. A value of 1 indicates that the class belongs to that class, and a value of 0 indicates that the class does not belong to that class. The confidence level is a probability value corresponding to the class identification information, indicating the reliability of the classification result.
[0040] To improve classification accuracy, a dynamic threshold is set. This threshold is adjusted based on actual conditions, such as production environments and component types. The output class identifiers and confidence levels are filtered out if their confidence level falls below the dynamic threshold. This eliminates unreliable classification results and improves classification accuracy.
[0041] Step S125: The local segmentation branch is restored to the original input image resolution through the transposed convolution layer and then input into the boundary coordinate regressor to generate the normalized coordinate parameters of the initial bounding box, and a non-maximum suppression algorithm based on the intersection-over-union ratio threshold is used to eliminate redundant bounding boxes with spatial overlap.
[0042] The main task of the local segmentation branch is to determine the boundary region of the object to be operated on. First, the feature map of the local segmentation branch is processed through a transposed convolutional layer. The transposed convolutional layer is a convolutional layer that can perform upsampling, which restores the resolution of the feature map to the resolution of the original input image.
[0043] Assume that the feature map of the local segmentation branch is F_local, which is smaller than the resolution of the original input image. Using a transposed convolutional layer, the resolution of F_local is increased to the same size as the original input image, resulting in the feature map F_local'. The parameters of the transposed convolutional layer, such as kernel size, stride, and padding, should be set as needed to ensure accurate resolution restoration.
[0044] The restored feature map F_local' is then fed into the bounding coordinate regressor. This model is used to predict the coordinates of an object's bounding box. It analyzes and calculates the feature map F_local' to generate the normalized coordinate parameters of the initial bounding box. Normalized coordinate parameters normalize the bounding box coordinates to within a specified range, facilitating subsequent processing and comparison.
[0045] To eliminate redundant, spatially overlapping bounding boxes, a non-maximum suppression algorithm based on an IoU threshold is employed. The IoU refers to the ratio of the intersection area to the union area of two bounding boxes. First, all generated initial bounding boxes are sorted by confidence. Then, the bounding box with the highest confidence score is selected as the baseline bounding box, and the IoU of the remaining bounding boxes with the baseline bounding box is calculated. If the IoU of a bounding box with the baseline bounding box exceeds the set IoU threshold, the bounding box is considered redundant and removed. This process is repeated until all bounding boxes have been processed, ultimately retaining only the highest-scoring, non-overlapping bounding boxes.
[0046] Step S126: Match and associate the filtered category identification information with the suppressed normalized bounding box coordinates according to the spatial grid area, and output the feature distribution data.
[0047] After the previous processing, we get the filtered category identification information and the suppressed normalized bounding box coordinates. In order to effectively integrate these two pieces of information, we match and associate them according to the spatial grid area.
[0048] First, the target area is divided into multiple spatial grid regions. The division method can be selected according to the actual situation, for example, it can be divided into equal intervals or adaptively divided according to the distribution of image features. Each spatial grid region has a certain range and coordinates.
[0049] Then, the spatial grid region to which the normalized bounding box coordinates belong is determined based on the normalized bounding box coordinates. For each normalized bounding box, the coordinates of its center point are calculated, and the spatial grid region to which it belongs is determined based on the coordinates of the center point. The corresponding category identification information is associated with the grid region. For example, if the center point of a normalized bounding box is located within a spatial grid region, and the category identification information corresponding to the bounding box is a certain component category, the category identification information is associated with the spatial grid region.
[0050] Through this matching and association method, the category identification information and normalized bounding box coordinates are integrated together to output feature distribution data containing the category identification information of the object to be operated and the boundary area coordinate information. This feature distribution data will be used for subsequent calibration and positioning processing.
[0051] Step S130: performing calibration processing on the feature distribution data to generate coordinate data of the object to be operated in a three-dimensional space coordinate system.
[0052] After obtaining the feature distribution data, it needs to be calibrated to generate accurate coordinate data of the object to be operated in the three-dimensional space coordinate system.
[0053] Step S131: Compensating the position offset in the feature distribution data according to a preset benchmark reference feature to obtain compensated feature position data. The compensation process includes pixel-level error correction based on a calibration template and reconstruction of a coordinate system mapping relationship.
[0054] First, you need to obtain the preset reference features. These can be specific markers or known feature points pre-set on the production line. These reference features have accurate position information and are used for subsequent calibration.
[0055] After acquiring the feature distribution data, we discovered that there was positional offset. This offset could be due to factors such as camera installation errors and lens distortion. To eliminate this offset, we performed compensation.
[0056] The first step in the compensation process is pixel-level error correction based on a calibration template. A calibration template is a pattern with known features, such as a checkerboard pattern. By comparing the calibration template image captured by the camera with the known calibration template information, the camera's distortion parameters, such as radial and tangential distortion coefficients, are calculated. Based on these distortion parameters, distortion correction is performed on the boundary region coordinates in the feature distribution data to eliminate pixel-level errors caused by lens distortion.
[0057] Next, coordinate system mapping relationships are reconstructed. The image acquisition device's calibration template intrinsic and extrinsic matrix are obtained. The intrinsic matrix contains the focal length and principal point coordinates, while the extrinsic matrix contains the camera pose parameters. Based on these two matrices, a coordinate system transformation relationship can be constructed to map image pixel coordinates to 3D physical coordinates.
[0058] Assuming that the calibration template intrinsic parameter matrix is K, the calibration template extrinsic parameter matrix is [R|t] (R is the rotation matrix, t is the translation vector), and the image pixel coordinate is p, the image pixel coordinate p can be converted to the three-dimensional physical coordinate P through the formula P=K^(-1)*[R|t]^(-1)*p. According to this coordinate system conversion relationship, the boundary area coordinate information in the feature distribution data is converted to obtain the corrected boundary coordinate data.
[0059] The compensation coefficient for the position offset is calculated based on the difference between the reference feature and the corrected boundary coordinate data. This difference can be determined using the Euclidean distance metric and least squares optimization. Assuming the reference feature coordinates are P_ref and the corrected boundary coordinate data are P_corrected, a least squares optimization method is used to find a compensation coefficient α that minimizes ∑(P_corrected + α - P_ref)^2.
[0060] Finally, the coordinate point data in the feature distribution data is updated based on the compensation coefficient. The update process includes coordinate translation and scaling. For each coordinate point p in the feature distribution data, the updated coordinate point p' = p + α is calculated, completing the compensation process and obtaining the compensated feature position data.
[0061] Step S132: Map the compensated feature position data to the target grid area in the image coordinate system, and extract the feature point set in the target grid area. The target grid area is dynamically divided according to the image resolution and object distribution density. The feature point set includes the edge feature points and internal key points of the object to be operated.
[0062] In order to further accurately extract the feature information of the object to be operated, the compensated feature position data is mapped to the target grid area in the image coordinate system.
[0063] First, the image coordinate system is dynamically divided into multiple grid cells based on the resolution parameters of the real-time multi-view image data and a preset object distribution density threshold. The object distribution density threshold is determined by statistically averaging the pixel spacing of objects in historical data. If the image resolution is high and the object distribution density is high, the grid cell size is smaller to more precisely capture object features. Conversely, if the image resolution is low and the object distribution density is low, the grid cell size is larger.
[0064] The coverage of the target grid area is then determined based on the distribution density of the coordinate points in the compensated feature location data. Density clustering algorithms and region growing methods can be used to determine coverage. Density clustering algorithms, such as DBSCAN, divide coordinate points into clusters based on their density. Region growing methods start from a seed point and continuously expand the region according to a set growth rule until a stopping condition is met.
[0065] Cluster the coordinate point data in the compensated feature location data according to grid cells. You can use algorithms such as Density-Based Spatial Clustering of Noise Applications (DBSCAN) and mean-shift algorithms. These clustering algorithms divide the coordinate point data into clusters, generating a set of coordinate point clusters within the target grid area.
[0066] Extract the target coordinate point cluster with the highest density from the set of coordinate point clusters. The target coordinate point cluster can be determined by counting the number of points within the cluster and evaluating the compactness of its spatial distribution. For example, select the cluster with the largest number of points within the cluster and the smallest distances between points as the target coordinate point cluster.
[0067] Calculate the center point of the target coordinate point cluster as the mapping reference point. The center point can be determined by geometric median calculation or centroid coordinates. Assume that the coordinate points in the target coordinate point cluster are p_1, p_2, ..., p_n, and the centroid coordinates are C = (∑p_i) / n.
[0068] Based on the mapping reference points, the compensated feature position data is mapped to the target grid area in the image coordinate system. Each coordinate point is translated and scaled relative to the mapping reference points so that it falls within the target grid area. Next, a set of feature points within the target grid area is extracted. These feature points include edge feature points and internal key points of the object to be operated. Edge feature points can be extracted using edge detection algorithms such as the Canny edge detection algorithm; internal key points can be extracted using corner detection algorithms such as the Harris corner detection algorithm.
[0069] Step S133: performing geometric transformation processing on the feature point set to generate center point coordinate data and boundary contour coordinate data of the object to be operated, wherein the geometric transformation processing includes affine transformation and perspective correction.
[0070] After obtaining the feature point set, in order to more accurately determine the position and posture of the object to be operated in three-dimensional space, the feature point set needs to be geometrically transformed. The geometric transformation here mainly includes affine transformation and perspective correction.
[0071] First, perform an affine transformation. Affine transformation is a linear transformation between two-dimensional coordinates and two-dimensional coordinates. It can maintain the parallelism of straight lines and can realize operations such as translation, rotation, scaling and shearing. In this process, the affine transformation matrix must be determined. This can be done by selecting at least three non-collinear feature points in the feature point set as reference points, and assuming that the coordinates of these three reference points in the original coordinate system are A, B, and C, and the expected corresponding coordinates in the target coordinate system are A', B', and C'. The affine transformation matrix M can be obtained by solving a set of linear equations. This linear equation set is constructed based on the properties of the affine transformation, that is, for each reference point, its coordinates in the original coordinate system can be obtained after the affine transformation matrix M is applied to it. By solving this set of linear equations, the elements in the affine transformation matrix M can be determined.
[0072] After obtaining the affine transformation matrix M, perform an affine transformation on each feature point in the feature point set. Let the coordinates of a feature point in the feature point set be P, and the coordinates after affine transformation be P'. P' is obtained by multiplying P by the affine transformation matrix M. This operation is performed on all feature points in the feature point set to obtain the feature point set after affine transformation.
[0073] Next, we perform perspective correction. The purpose of perspective correction is to eliminate the perspective distortion caused by the camera's viewing angle so that the objects in the image appear in normal proportions and shapes. To perform perspective correction, it is also necessary to select four non-collinear feature points from the feature point set as reference points. Let the coordinates of these four reference points in the original image be Q1, Q2, Q3, and Q4, and the corresponding coordinates in the desired corrected image be Q1', Q2', Q3', and Q4'. Through the coordinates of these four sets of corresponding points, the perspective transformation matrix H can be calculated. The calculation process of the perspective transformation matrix H is also based on a set of linear equations, which describe the relationship between the points in the original image and the corresponding points in the corrected image after the perspective transformation.
[0074] After obtaining the perspective transformation matrix H, the feature point set that has undergone affine transformation is further perspective transformed. Let the coordinates of a feature point after affine transformation be R, and the coordinates after perspective transformation be R'. R' is obtained by performing the corresponding operation on R and the perspective transformation matrix H. Perspective transformation is performed on all feature points in the feature point set that has undergone affine transformation, thus obtaining a feature point set that has undergone a complete geometric transformation.
[0075] After completing the geometric transformation processing, the center point coordinate data and boundary contour coordinate data of the object to be operated are generated based on the processed feature point set. For the calculation of the center point coordinate data, a minimum circumscribed geometric figure is constructed based on the distribution of coordinate points in the feature point set. The minimum circumscribed geometric figure can be a rectangle, polygon, or ellipse, etc. The specific choice of figure depends on the distribution of the feature points. If the distribution of the feature points is approximately rectangular, a minimum circumscribed rectangle is constructed; if the distribution is more complex, a minimum circumscribed polygon or ellipse is constructed. Taking the construction of the minimum circumscribed rectangle as an example, the minimum circumscribed rectangle is generated by the rotating caliper algorithm or the convex hull algorithm. The rotating caliper algorithm finds the circumscribed rectangle of the minimum area by continuously rotating a convex hull surrounding the feature point set; the convex hull algorithm first finds the convex hull of the feature point set, and then determines the minimum circumscribed rectangle based on the convex hull.
[0076] Calculate the geometric center coordinates of the smallest circumscribed geometric figure as the center coordinate data for the object being operated. For rectangles, the geometric center can be determined by weighted averaging of vertex coordinates or area centroid calculation. Assume the four vertices of the rectangle are V1, V2, V3, and V4. The weighted average vertex coordinates are calculated by adding and averaging the x-coordinates and y-coordinates of the four vertices. The resulting coordinates are the geometric center coordinates. The area centroid calculation takes into account the area distribution of the rectangle and determines the center coordinates through a more complex calculation.
[0077] For the generation of boundary contour coordinate data, the vertex coordinate data of the minimum circumscribed geometric figure is extracted. The vertex coordinate data can be obtained through edge detection and corner point recognition. The edge detection algorithm can detect the edge of the figure composed of a set of feature points, and the corner point recognition algorithm can accurately locate the intersection of the edges, that is, the vertex. For example, the Canny edge detection algorithm is used to detect the edge, and the Harris corner detection algorithm is used to identify the corner point. The boundary contour coordinate data is generated based on the connection relationship between the vertex coordinates. The connection relationship is realized through polygon fitting and spline curve interpolation processing. Polygon fitting is to connect the vertices in a set order to form a polygon; spline curve interpolation is to generate a smoother curve through the interpolation algorithm based on the polygon to more accurately describe the boundary contour of the object to be operated.
[0078] Step S134: constructing a spatial coordinate mapping relationship of the object to be operated according to the center point coordinate data and the boundary contour coordinate data, and converting the coordinate data in the image coordinate system into coordinate data in a three-dimensional space coordinate system through the spatial coordinate mapping relationship.
[0079] After obtaining the center point coordinate data and boundary contour coordinate data of the object to be operated, it is necessary to construct a spatial coordinate mapping relationship of the object to be operated so as to convert the coordinate data in the image coordinate system into the coordinate data in the three-dimensional space coordinate system.
[0080] First, let's analyze the relationship between the image coordinate system and the three-dimensional space coordinate system. The image coordinate system is two-dimensional and is used by cameras to capture images, while the three-dimensional space coordinate system describes the position and orientation of objects in real space. To establish the mapping relationship between these two coordinate systems, we need to consider the camera's intrinsic and extrinsic parameters. The camera's intrinsic parameters include parameters such as focal length and principal point coordinates, which describe the camera's internal optical properties. The camera's extrinsic parameters include the camera's position and orientation, that is, its coordinates and orientation in three-dimensional space.
[0081] Based on the center point coordinate data and the boundary contour coordinate data, combined with the camera's intrinsic and extrinsic parameters, a spatial coordinate mapping relationship is constructed. This mapping can be achieved through a series of transformations. Suppose the coordinates of a point in the image coordinate system are I, and the coordinates of the corresponding point in the three-dimensional space coordinate system are S. First, point I in the image coordinate system is converted to the camera coordinate system. This requires the use of the camera's intrinsic parameter matrix K. By multiplying I with the inverse matrix of K, we can obtain point I_c in the camera coordinate system.
[0082] Next, we transform point I_c in the camera coordinate system into a 3D coordinate system. This requires the use of the camera's extrinsic matrix [R|t], where R is the rotation matrix and t is the translation vector. By multiplying I_c by the rotation matrix R and adding the translation vector t, we obtain point S in the 3D coordinate system.
[0083] The above steps are used to convert each point in the center point coordinate data and the boundary outline coordinate data. This allows the center point coordinate data and boundary outline coordinate data in the image coordinate system to be converted into coordinate data in the three-dimensional space coordinate system. By constructing and converting this spatial coordinate mapping relationship, the position of the object to be operated in three-dimensional space can be accurately determined.
[0084] Step S140: performing three-dimensional posture calculation processing based on the coordinate data to obtain a posture parameter set of the object to be operated, wherein the posture parameter set includes spatial position parameters and posture angle parameters of the object to be operated.
[0085] After obtaining the coordinate data of the object to be operated in the three-dimensional space coordinate system, it is necessary to perform three-dimensional posture solution processing based on these coordinate data to obtain a posture parameter set of the object to be operated, which includes spatial position parameters and posture angle parameters.
[0086] Step S141: Perform spatial projection processing on the coordinate data according to the real-time multi-perspective image data to generate multi-perspective projection feature data of the object to be operated, wherein the multi-perspective image data includes synchronous observation information of the object to be operated at different perspectives, the spatial projection processing includes multi-perspective projection matrix construction and depth information fusion, and the multi-perspective projection feature data includes spatial distribution information of object surface features at different perspectives.
[0087] The coordinate data is spatially projected using real-time multi-view image data. This data contains synchronous observation information of the object to be operated from different perspectives, providing more comprehensive object features.
[0088] First, construct the multi-view projection matrix. For each camera perspective, there is a corresponding projection matrix. The projection matrix describes how points in three-dimensional space are projected onto a two-dimensional image plane. To construct a multi-view projection matrix, you need to know the intrinsic and extrinsic parameters of each camera. The intrinsic parameters include focal length, principal point coordinates, etc., and the extrinsic parameters include the position and posture of the camera. Based on these parameters, the projection matrix of each camera can be calculated. Let the intrinsic parameter matrix of the i-th camera be K_i and the extrinsic parameter matrix be [R_i|t_i]. Then the projection matrix P_i of the i-th camera can be obtained by multiplying K_i with [R_i|t_i].
[0089] Then, the coordinate data in the 3D space coordinate system is projected through the projection matrix of each camera to obtain the 2D projection coordinates at each viewing angle. Let the coordinates of a point in 3D space be X, and the 2D projection coordinates obtained by projecting it through the projection matrix P_i of the i-th camera be x_i.
[0090] While projection is being performed, depth information fusion is also required. Depth information can be acquired through methods such as binocular vision and structured light. Depth information from different perspectives may differ, requiring fusion to obtain more accurate depth information. Depth information fusion can be performed using a weighted average method. Let the depth information from the i-th perspective be d_i, and the corresponding weight be w_i. The fused depth information D can be obtained by taking the weighted sum of the depth information from all perspectives, i.e., D = ∑(w_i * d_i).
[0091] By constructing a multi-view projection matrix and fusing depth information, we generate multi-view projection feature data of the object to be operated. This multi-view projection feature data contains the spatial distribution information of the object's surface features at different viewing angles.
[0092] Step S142: Construct a pose solution matrix based on the multi-view projection feature data, and update the rotation parameters and translation parameters in the pose solution matrix through an iterative optimization algorithm until the reprojection error converges to a preset threshold, thereby obtaining optimized rotation parameters and translation parameters. The pose solution matrix includes initial estimated values of the rotation parameters and translation parameters, and the iterative optimization algorithm includes error minimization based on gradient descent and parameter convergence verification.
[0093] The pose solution matrix is constructed based on the multi-view projection feature data. The pose solution matrix contains initial estimates of the rotation parameters and translation parameters. The rotation parameters describe the object's rotation in 3D space, while the translation parameters describe the object's translation in 3D space.
[0094] First, obtain the observed position data of the operation actuator at different viewing angles. This data can be obtained through encoder feedback and kinematic forward solution. The encoder provides real-time feedback on the joint angle information of the operation actuator. Using the forward kinematic solution algorithm, the position and posture of the operation actuator in three-dimensional space can be calculated based on this joint angle information. Let the observed position data of the operation actuator at the jth viewing angle be O_j.
[0095] Construct an observation matrix based on the observation position data. The observation matrix contains the camera perspective transformation parameters and the object projection relationship. Let the observation matrix at the j-th perspective be M_j, which is related to the camera's intrinsic and extrinsic parameters and the observation position data of the operation execution device.
[0096] The initial parameters of the pose solution matrix are constructed based on the correspondence between the multi-view projection feature data and the observation matrix. This correspondence can be determined through feature point matching and projection error calculation. Feature point matching involves finding corresponding feature points in the multi-view projection feature data and the observation matrix, while projection error refers to the distance error between these corresponding feature points. By minimizing the projection error, the initial parameters of the pose solution matrix can be determined.
[0097] After obtaining the initial parameters of the pose solution matrix, the rotation parameters and translation parameters are updated using an iterative optimization algorithm. The iterative optimization algorithm uses error minimization and parameter convergence verification based on gradient descent. The gradient descent algorithm is a commonly used optimization algorithm that continuously updates parameters along the negative gradient direction of the error function to reduce the error. Suppose the rotation parameter in the pose solution matrix is R, the translation parameter is t, and the error function is E(R, t). In each iteration, according to the gradient of the error function, Update the rotation and translation parameters.
[0098] After each iteration, parameter convergence verification is performed to determine whether the reprojection error has converged to a preset threshold. The reprojection error is the difference between the projection result obtained by applying the updated pose parameters to the multi-view projection feature data and the observation matrix. If the reprojection error is less than the preset threshold, the parameters are considered converged and the iteration ends; otherwise, the iterative update continues.
[0099] After multiple iterations, until the reprojection error converges to the preset threshold, the optimized rotation parameters and translation parameters are obtained.
[0100] Step S143: Generate the pose parameter set based on the optimized rotation parameters and translation parameters, the pose parameter set including attitude angle parameters and spatial position parameters, the attitude angle parameters are used to describe the rotation state of the object in the three-dimensional space, and the spatial position parameters are used to describe the translation state of the object in the three-dimensional space.
[0101] After obtaining the optimized rotation parameters and translation parameters, a pose parameter set is generated based on these parameters.
[0102] Attitude angle parameters can be derived from the optimized rotation parameters. Rotation parameters are typically represented by rotation matrices or quaternions. If the rotation parameters are represented by a rotation matrix R, a predefined conversion method can be used to convert the rotation matrix into attitude angle parameters, such as Euler angles (including rotation angles around the X-axis, Y-axis, and Z-axis). The conversion process involves the mathematical relationship between the elements of the rotation matrix and the attitude angles. The attitude angle parameters can be obtained by solving a system of equations.
[0103] The spatial position parameters can be directly obtained from the optimized translation parameters. The translation parameters describe the translation state of the object in three-dimensional space, where the three components correspond to the translation distance of the object in the X, Y, and Z axis directions respectively.
[0104] The attitude angle parameters and spatial position parameters are combined to form a pose parameter set, which accurately describes the position and attitude of the object to be operated in three-dimensional space.
[0105] Step S150: Generate a control instruction based on the posture parameter set and the object physical property parameters corresponding to the category identification information, and drive the operation execution device to complete the object grasping action of the object to be operated. The control instruction includes a path control signal, an end effector posture control signal, and a clamping force control signal based on the category identification information.
[0106] After obtaining the posture parameter set and category identification information, a control instruction is generated according to the physical property parameters of the object corresponding to the category identification information to drive the operation execution device to complete the grasping action of the object to be operated.
[0107] Step S151: calling a preset grasping strategy library according to the category identification information, and selecting a clamping force parameter and motion trajectory constraint condition that matches the category identification information.
[0108] The preset grasping strategy library stores grasping strategies corresponding to different category identification information, including gripping force parameters and motion trajectory constraints. According to the category identification information of the object to be operated, matching information is searched from the grasping strategy library.
[0109] Different types of objects require different gripping forces, depending on their physical properties. For example, a harder, heavier object requires a stronger gripping force, while a softer, more fragile object requires a weaker gripping force. Based on the category identification information, the corresponding gripping force parameters can be determined.
[0110] Trajectory constraints are the trajectory rules that the actuator must follow when grasping an object. Different object categories may have different trajectory requirements. For example, for unusually shaped objects, the actuator may need to approach the object along a specific path to avoid collision or damage. Based on the category identification information, the corresponding trajectory constraint is selected from the grasping strategy library.
[0111] Step S152: Input the posture parameter set into the motion planning model, and generate the motion trajectory data and end effector posture data of the operation execution device based on inverse kinematics and path optimization algorithm through the motion planning model. The motion trajectory data includes joint motion sequence and speed planning parameters, and meets the motion trajectory constraint conditions.
[0112] For example, step S1521: constructing a target homogeneous transformation matrix of the end of the robotic arm based on the spatial position parameters and the posture angle parameters in the posture parameter set, wherein the target homogeneous transformation matrix includes a translation component and a rotation component.
[0113] The spatial position parameters in the pose parameter set describe the position of the end arm in three-dimensional space, and the attitude angle parameters describe the attitude of the end arm. The spatial position parameters are represented as a three-dimensional vector, and the attitude angle parameters are represented by a rotation matrix. The target homogeneous transformation matrix is a 4×4 matrix. The 3×3 submatrix in the upper left corner is the rotation matrix, representing the rotation information corresponding to the attitude angle parameters. The first three elements of the fourth column are the three components of the spatial position parameters, representing the translation information. The elements in the last row are [0, 0, 0, 1]. In this way, the spatial position parameters and attitude angle parameters are integrated into the target homogeneous transformation matrix for subsequent processing.
[0114] Step S1522: performing feasible space verification on the target homogeneous transformation matrix according to the motion trajectory constraint conditions, wherein the feasible space verification includes joint motion range limitation detection and end effector interference area avoidance.
[0115] Motion trajectory constraints are rules defined based on the type of object being manipulated and the actual working environment. Joint range constraint detection verifies that the robot arm joint angles corresponding to the target homogeneous transformation matrix are within the permitted range of motion for each joint. This is accomplished by presetting the minimum and maximum motion angles for each joint and comparing the joint angles calculated using inverse kinematics with these ranges. If a joint angle exceeds the permitted range, the target homogeneous transformation matrix is considered infeasible.
[0116] End-effector interference zone avoidance checks whether the end-effector will interfere with surrounding objects in the pose corresponding to the target homogeneous transformation matrix. This is done by building a 3D model of the work environment and comparing the end-effector model with this 3D model to determine if interference exists. If interference does occur, the target homogeneous transformation matrix needs to be adjusted or replanned.
[0117] Step S1523: Based on inverse kinematics, multiple solutions are screened for the verified target homogeneous transformation matrix to generate a set of candidate joint angles that meet the motion trajectory constraints. The screening criteria give priority to candidate solutions with the shortest joint motion path and the lowest energy consumption. The multiple solution screening includes Jacobian matrix singularity analysis and joint angular velocity threshold traversal.
[0118] Inverse kinematics is the process of solving for the angles of each joint based on the position and posture of the end effector. For a given target homogeneous transformation matrix, inverse kinematics may have multiple solutions. To identify a suitable solution, a Jacobian matrix singularity analysis is first performed. The Jacobian matrix describes the relationship between the joint velocities of the robot and the end effector. If the Jacobian matrix is singular, it indicates that the robot is in a singular configuration. In this case, small changes in the joints may result in large movements of the end effector, which is not conducive to control. Therefore, singular solutions need to be eliminated.
[0119] Next, the joint angular velocity thresholds are traversed, checking whether the corresponding joint angular velocity of each candidate solution is within the permitted range. Candidate solutions with excessive joint angular velocity may cause unstable arm motion or exceed the motor's capabilities, and therefore need to be excluded. Finally, based on the screening criteria, candidate solutions with the shortest joint motion path and the lowest energy consumption are prioritized to generate a set of candidate joint angles that meet the motion trajectory constraints.
[0120] Step S1524: A path optimization algorithm is used to perform smooth trajectory planning on the candidate joint angle set to generate an optimized joint motion sequence and velocity planning parameters. The path optimization algorithm includes dynamic time warping processing and joint acceleration continuity constraints.
[0121] The goal of the path optimization algorithm is to make the robot's motion smoother and more stable, reducing vibration and shock. Dynamic Time Warping (DTW) adjusts the temporal sequence of joint motions to achieve more coordinated movement of each joint. By performing DTW on each joint angle sequence in the candidate joint angle set, the joint motion rhythm can be made more reasonable.
[0122] Joint acceleration continuity constraints ensure that joint accelerations change continuously during motion, avoiding sudden changes. Interpolation and fitting of the joint angle sequence can be used to ensure smooth changes in joint acceleration between adjacent time points. These path optimization algorithms process the candidate joint angle sets to generate optimized joint motion sequences and velocity planning parameters. The joint motion sequence describes the angles of each joint of the manipulator at different time points, while the velocity planning parameters specify the velocity of each joint during motion.
[0123] Step S1525: Performing posture synchronization interpolation on the optimized joint motion sequence according to the end effector posture adjustment parameters to generate an end effector posture angle change sequence corresponding to each joint motion node, wherein the posture synchronization interpolation includes rotation matrix linear interpolation and posture angular rate consistency matching.
[0124] The end-effector posture adjustment parameters are determined based on the grasping requirements and actual conditions of the object being manipulated. They specify the posture the end-effector must achieve during the grasping process. The purpose of posture synchronization interpolation is to synchronize the end-effector posture changes with the motion of the manipulator joints.
[0125] Rotation matrix linear interpolation linearly interpolates the end-effector's rotation matrix between different joint motion nodes to ensure smooth transitions in posture. Attitude angular rate consistency matching ensures that the end-effector's attitude angular rate matches the speed of the joint motion. By combining the end-effector attitude adjustment parameters with the optimized joint motion sequence, attitude synchronization interpolation is performed to generate a sequence of end-effector attitude angle changes corresponding to each joint motion node. This sequence of end-effector attitude angle changes describes the end-effector's attitude angle at each moment during the manipulator's motion.
[0126] Step S1526: Associating the joint motion sequence, velocity planning parameters, and posture angle change sequence into motion trajectory data and end effector posture data that satisfy the motion trajectory constraint conditions.
[0127] The optimized joint motion sequence, velocity planning parameters, and the end-effector attitude angle change sequence obtained through attitude synchronization interpolation are associated to form complete motion trajectory data and end-effector attitude data. The motion trajectory data details the motion sequence, speed, and time of each joint of the operating actuator when performing a grasping task, while the end-effector attitude data specifies the attitude angle of the end-effector at different joint motion nodes. In this way, through the motion planning model, comprehensive consideration of inverse kinematics and path optimization algorithms, combined with motion trajectory constraints, motion trajectory data and end-effector attitude data suitable for the operating actuator to complete the grasping task are generated.
[0128] Step S153: generating a path control signal according to the motion trajectory data, and the path control signal drives the operation execution device to move to the target grasping position through pulse width modulation and a servo drive interface.
[0129] After obtaining the motion trajectory data, it needs to be converted into a path control signal that can drive the operation actuator. The motion trajectory data contains the joint motion sequence and speed planning parameters, which determine the position and movement speed of each joint of the operation actuator at different times.
[0130] The path control signal is generated based on the joint motion sequence and velocity planning parameters in the motion trajectory data. First, the position and velocity information of each joint in the joint motion sequence is encoded and converted into digital signals. These digital signals represent the position and velocity that each joint of the actuator should achieve at different times.
[0131] These digital signals are then processed using pulse-width modulation (PWM). PWM controls motor speed and torque by varying the duty cycle of the pulse signal. For each joint motor in the actuator, the duty cycle of the pulse signal is adjusted based on its corresponding speed planning parameters. If the joint requires rapid movement, the duty cycle of the pulse signal is increased; if the joint requires slow movement, the duty cycle is decreased.
[0132] The pulse-width modulated signal is transmitted to the joint motors of the actuator via the servo drive interface. The servo drive interface bridges the control system and the joint motors, converting the received control signal into an electrical signal that the motors can recognize and execute. Through the servo drive interface, the path control signal is accurately transmitted to each joint motor, driving the actuator to move to the target grasping position along the path specified by the motion trajectory data.
[0133] Step S154: Generate an end effector posture control signal and a clamping force control signal based on the end effector posture data and the clamping force parameter, and respectively perform end effector rotation axis control and clamping force feedback adjustment based on the end effector posture control signal and the clamping force control signal, triggering the clamping action to complete the grasping of the object to be operated.
[0134] Based on the end-effector posture data and gripping force parameters, the end-effector posture control signal and gripping force control signal are generated, respectively. The end-effector posture data describes the posture angle that the end-effector should achieve when grasping an object. Based on this posture angle information, the corresponding posture control signal is generated. The posture control signal is used to control the end-effector's rotation axis, enabling it to accurately adjust to the appropriate posture.
[0135] The gripping force parameter is selected from the grasping strategy library based on the object's category identification information. It determines the force the end effector must apply to grasp the object. Based on the gripping force parameter, a gripping force control signal is generated. This gripping force control signal is used to control the end effector's gripping mechanism, ensuring that it applies the appropriate force to grasp the object.
[0136] During execution, the end effector's rotation axis is controlled based on the end effector's posture control signal. The posture control signal is transmitted to the end effector's rotation axis motor via the corresponding drive circuit. The motor rotates according to the signal's instructions, thereby adjusting the end effector's posture. At the same time, to ensure that the end effector can accurately achieve the desired posture, posture feedback adjustment is also required. A posture sensor can be installed on the end effector to monitor the end effector's actual posture angle in real time and feed this information back to the control system. The control system compares the actual posture angle with the target posture angle. If there is a deviation, the posture control signal is adjusted to gradually bring the end effector's posture closer to the target posture.
[0137] Feedback adjustment is also required for the clamping force control signal. A force sensor is installed on the clamping device of the end effector to monitor the actual force applied by the clamping device in real time. The actual force information is fed back to the control system, which compares the actual force with the clamping force parameter. If the actual force is less than the clamping force parameter, it means that the clamping force is insufficient. The control system will increase the clamping force control signal to make the clamping device apply more force. If the actual force is greater than the clamping force parameter, it means that the clamping force is too great. The control system will reduce the clamping force control signal to avoid damage to the object to be operated.
[0138] When the end effector's posture is adjusted to the appropriate position and the gripping force reaches the set parameters, the gripping action is triggered. The gripping device, in accordance with the gripping force control signal, firmly grasps the object to be operated, completing the grasping task.
[0139] In the entire deep learning-based target collaborative control method, starting from the image acquisition device acquiring real-time multi-view image data, through the real-time recognition model's object feature recognition, feature distribution data calibration processing, three-dimensional posture solution processing, and finally generating control instructions based on the posture parameter set and the object's physical property parameters to drive the operation execution device to complete the object grasping action, each step is closely linked to form a complete control process.
[0140] During the data acquisition stage, high-precision industrial cameras are used as image acquisition devices. The selection of performance parameters and layout of their layout directly affect the quality of the collected real-time multi-view image data. Cameras at different positions and angles can capture visual feature information of the object to be operated from multiple aspects.
[0141] The real-time recognition model uses an improved YOLOv8 algorithm. Through a series of operations, including a multi-level dilated convolutional network, a cross-level feature pyramid, and channel- and spatial-dimensional attention weighting, it accurately identifies the category identification and boundary coordinates of the object being operated on. The multi-level dilated convolutional network dynamically adjusts the dilation rate based on the input image resolution to extract features at different scales. The cross-level feature pyramid integrates feature maps from different levels through bilinear interpolation and channel-by-channel weighted fusion. The channel- and spatial-dimensional attention weighting further highlights important feature information, improving recognition accuracy.
[0142] Calibrating feature distribution data is a critical step in ensuring the accuracy of subsequent operations. Pixel-level error correction and coordinate system mapping relationship reconstruction based on the calibration template eliminate position offsets caused by factors such as camera installation errors and lens distortion. The compensated feature position data is mapped to the target grid area in the image coordinate system, and a set of feature points is extracted. A geometric transformation is then performed to generate the center point coordinate data and boundary contour coordinate data for the object to be operated. Finally, a spatial coordinate mapping relationship is constructed, converting the image coordinate data into a three-dimensional spatial coordinate system.
[0143] The three-dimensional pose solution processing generates multi-view projection feature data through spatial projection processing, constructs a pose solution matrix based on this data, and updates the rotation parameters and translation parameters through an iterative optimization algorithm to obtain an optimized pose parameter set. These pose parameter sets accurately describe the position and posture of the object to be operated in three-dimensional space.
[0144] Finally, control instructions are generated based on the pose parameter set and the object's physical property parameters, driving the operator to complete the object grasping action. By calling a preset grasping strategy library, appropriate gripping force parameters and motion trajectory constraints are selected. The motion planning model generates motion trajectory data and end-effector posture data based on inverse kinematics and path optimization algorithms. The path control signal drives the operator to the target grasping position through pulse width modulation and a servo drive interface. The end-effector posture control signal and gripping force control signal respectively control the end-effector's posture and gripping force. Feedback adjustment ensures operational accuracy, ultimately completing the grasping task of the object to be manipulated.
[0145] During the above implementation process, attention must also be paid to data privacy protection and leakage prevention. The real-time multi-view image data collected may contain sensitive information about the object to be operated, such as product design details. To protect this privacy-sensitive data, encryption technology is used to encrypt the data. During data transmission, secure communication protocols, such as SSL / TLS, are used to ensure that the data is not stolen or tampered with during transmission. In terms of data storage, the data is stored on a secure server with strict access rights set so that only authorized personnel can access the data.
[0146] Furthermore, training real-time recognition models requires a large amount of labeled data. This labeled data should cover a wide variety of objects to be manipulated, as well as images of them from different viewing angles and lighting conditions. During training, optimization algorithms such as stochastic gradient descent are used to continuously adjust model parameters to continuously improve recognition accuracy. Parameters such as the learning rate and batch size also need to be appropriately set during training to ensure rapid convergence and stability.
[0147] The performance of the motion planning model directly impacts the motion performance of the actuator. When performing inverse kinematics calculations and path optimization, the dynamic characteristics of the manipulator, such as joint torque limits, acceleration, and deceleration, must be considered. The choice of path optimization algorithm is also crucial. Different algorithms may produce different motion trajectories, and the most appropriate algorithm must be selected based on the actual situation to ensure that the actuator can efficiently and stably complete the grasping task.
[0148] In practical applications, the entire system also requires real-time monitoring and debugging. By monitoring parameters such as the motion state of the actuator, the posture of the end effector, and the gripping force, system problems can be promptly identified and adjusted accordingly. Furthermore, continuous data collection from practical applications optimizes and improves the real-time recognition and motion planning models, enhancing system performance and adaptability.
[0149] Regarding the system's hardware, the performance of high-precision industrial cameras and the accuracy and stability of the actuators all impact the overall system's effectiveness. Therefore, when selecting hardware, factors such as performance, price, and reliability should be comprehensively considered. For example, parameters such as the resolution, frame rate, and sensitivity of a high-precision industrial camera will affect the quality of captured images; parameters such as the repeatability and load capacity of the actuators will affect the accuracy and stability of their grasping.
[0150] Regarding the system's software, the interfaces and communication protocols between various modules also require careful design. Data transmission and interaction are required between the real-time recognition model, calibration processing module, 3D pose calculation module, and control instruction generation module. Therefore, unified interfaces and communication protocols must be defined to ensure accurate data transmission and processing. Software stability and reliability are also crucial, requiring thorough testing and optimization to avoid system crashes or errors.
[0151] In terms of system maintenance, hardware devices require regular inspection and maintenance to ensure proper operation. Software systems require timely updates and upgrades to fix discovered vulnerabilities and improve system performance. Furthermore, a comprehensive fault diagnosis and resolution mechanism should be established to quickly locate and repair system failures, minimizing downtime.
[0152] Figure 2 A schematic diagram illustrating exemplary hardware and software components of a deep learning-based target collaborative control system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application, is shown. For example, the processor 120 can be used in the deep learning-based target collaborative control system 100 and used to perform the functions of the present application.
[0153] The deep learning-based target collaborative control system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the deep learning-based target collaborative control method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0154] For example, the target collaborative control system 100 based on deep learning may include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the target collaborative control system 100 based on deep learning may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The target collaborative control system 100 based on deep learning also includes an I / O interface 150 between the computer and other input and output devices.
[0155] For ease of explanation, only one processor is described in the deep learning-based target collaborative control system 100. However, it should be noted that the deep learning-based target collaborative control system 100 in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the deep learning-based target collaborative control system 100 executes step A and step B, it should be understood that step A and step B may also be performed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0156] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned target collaborative control method based on deep learning is implemented.
[0157] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A target collaborative control method based on deep learning, characterized in that: The method comprises: Acquire real-time multi-view image data of the target area through an image acquisition device, wherein the real-time multi-view image data includes visual feature information of the object to be operated; Invoking a real-time recognition model to perform object feature recognition on the real-time multi-view image data to obtain feature distribution data of the object to be operated, wherein the feature distribution data includes category identification information and boundary area coordinate information of the object to be operated; Calibrate the characteristic distribution data to generate coordinate data of the object to be operated in a three-dimensional space coordinate system; Performing a three-dimensional posture solution process based on the coordinate data to obtain a posture parameter set of the object to be operated, wherein the posture parameter set includes spatial position parameters and posture angle parameters of the object to be operated; Generate a control instruction based on the pose parameter set and the object physical property parameters corresponding to the category identification information, and drive the operation execution device to complete the object grasping action of the object to be operated, wherein the control instruction includes a path control signal, an end effector posture control signal, and a clamping force control signal based on the category identification information; The calling of the real-time recognition model to perform object feature recognition on the real-time multi-view image data to obtain feature distribution data of the object to be operated includes: Performing parallel multi-scale feature extraction on the real-time multi-view image data using a multi-level dilated convolutional network in the real-time recognition model to generate a multi-level feature set containing feature maps of different resolutions, wherein the dilation rate of the dilated convolution kernel at each level is dynamically adjusted according to the inverse of the corresponding input image resolution to match the change in object size; Inputting the multi-level feature set into the cross-level feature pyramid in the real-time recognition model, performing resolution enhancement on the low-level large receptive field feature map through bilinear interpolation to align its size with the high-level small receptive field feature map, and performing channel-by-channel weighted fusion on the aligned feature maps to generate an enhanced feature map that combines global semantics with local details; Applying channel-dimensional attention weighting and spatial-dimensional attention weighting to the enhanced feature map to obtain a weighted feature map, wherein the channel attention weight is generated by a two-layer fully connected network after global average pooling, and the spatial attention weight is generated by the spatial compression feature of a single-layer convolution kernel; The weighted feature map is divided into a global classification branch and a local segmentation branch along the channel dimension. The global classification branch is input into a multi-layer perceptron classifier after dimensionality reduction through global maximum pooling, outputs the category identification information and corresponding confidence of the object to be operated, and filters the category identification information with confidence lower than the dynamic threshold; The local segmentation branch is restored to the original input image resolution through the transposed convolution layer and then input into the boundary coordinate regressor to generate the normalized coordinate parameters of the initial bounding box. The non-maximum suppression algorithm based on the intersection-over-union ratio threshold is used to eliminate redundant bounding boxes with spatial overlap. The filtered category identification information is matched and associated with the suppressed normalized bounding box coordinates according to the spatial grid area, and the feature distribution data is output.
2. The target collaborative control method based on deep learning according to claim 1 is characterized in that: The calibrating the characteristic distribution data to generate coordinate data of the object to be operated in a three-dimensional space coordinate system includes: Compensating the position offset in the feature distribution data according to a preset benchmark reference feature to obtain compensated feature position data, wherein the compensation process includes pixel-level error correction based on a calibration template and reconstruction of a coordinate system mapping relationship; Mapping the compensated feature position data to a target grid area in an image coordinate system, extracting a feature point set within the target grid area, wherein the target grid area is dynamically divided according to image resolution and object distribution density, and the feature point set includes edge feature points and internal key points of the object to be operated; Performing geometric transformation processing on the feature point set to generate center point coordinate data and boundary contour coordinate data of the object to be operated, wherein the geometric transformation processing includes affine transformation and perspective correction; A spatial coordinate mapping relationship of the object to be operated is constructed according to the center point coordinate data and the boundary contour coordinate data, and the coordinate data in the image coordinate system is converted into coordinate data in a three-dimensional space coordinate system through the spatial coordinate mapping relationship.
3. The target collaborative control method based on deep learning according to claim 2 is characterized in that: The compensating the position offset in the feature distribution data according to the preset benchmark reference feature to obtain compensated feature position data includes: Obtaining a calibration template intrinsic parameter matrix and a calibration template extrinsic parameter matrix of the image acquisition device, wherein the calibration template intrinsic parameter matrix includes focal length and principal point coordinate parameters, and the calibration template extrinsic parameter matrix includes camera pose parameters; Constructing a coordinate system transformation relationship according to the calibration template intrinsic parameter matrix and the calibration template extrinsic parameter matrix, wherein the coordinate system transformation relationship is used to map image pixel coordinates to three-dimensional physical coordinates; Performing distortion correction processing on the coordinate information of the boundary area according to the coordinate system conversion relationship to obtain corrected boundary coordinate data, wherein the distortion correction processing includes radial distortion coefficient compensation and tangential distortion correction; Calculating a compensation coefficient for the position offset based on a difference between the benchmark reference feature and the corrected boundary coordinate data, the difference being determined by a Euclidean distance metric and least squares optimization; The coordinate point data in the boundary area coordinate information is updated according to the compensation coefficient to obtain compensated feature position data, and the updating process includes coordinate translation transformation and scale adjustment.
4. The target collaborative control method based on deep learning according to claim 2, characterized in that: Mapping the compensated feature position data to a target grid area in an image coordinate system and extracting a set of feature points in the target grid area includes: Dynamically dividing the image coordinate system into a plurality of grid cells according to the resolution parameters of the real-time multi-view image data and a preset object distribution density threshold, wherein the density threshold is determined by statistically analyzing the average pixel spacing of object distribution in historical data; Determining the coverage of the target grid area according to the distribution density of the coordinate points in the compensated feature position data, wherein the coverage is determined by a density clustering algorithm and a region growing method; Clustering the coordinate point data in the compensated feature position data according to grid units to generate a set of coordinate point clusters within the target grid area, wherein the clustering includes density-based noise application spatial clustering and mean shift algorithm; Extracting the target coordinate point cluster with the largest density from the set of coordinate point clusters, wherein the target coordinate point cluster is determined by counting the number of points within the cluster and evaluating the compactness of spatial distribution; Calculating the center point of the target coordinate point cluster as a mapping reference point, wherein the center point is determined by geometric median calculation or centroid coordinates; The compensated feature position data is mapped to a target grid area in an image coordinate system based on the mapping reference point, and a feature point set in the target grid area is extracted.
5. The target collaborative control method based on deep learning according to claim 2, characterized in that: The step of performing geometric transformation on the feature point set to generate center point coordinate data and boundary contour coordinate data of the object to be operated includes: Constructing a minimum circumscribed geometric figure based on the distribution of coordinate points in the feature point set, wherein the minimum circumscribed geometric figure includes a rectangle, a polygon, or an ellipse, and is generated by a rotating calculus algorithm or a convex hull algorithm; Calculating the coordinates of the geometric center point of the minimum circumscribed geometric figure as the coordinate data of the center point of the object to be operated, wherein the geometric center point is determined by weighted average of vertex coordinates or area centroid calculation; Extracting vertex coordinate data of the minimum circumscribed geometric figure, wherein the vertex coordinate data is obtained through edge detection and corner point recognition; The boundary contour coordinate data is generated according to the connection relationship between the vertex coordinates, and the connection relationship is realized by polygon fitting and spline curve interpolation processing.
6. The target collaborative control method based on deep learning according to claim 1, characterized in that: The performing of three-dimensional posture calculation processing based on the coordinate data to obtain a posture parameter set of the object to be operated includes: Performing spatial projection processing on the coordinate data according to the real-time multi-view image data to generate multi-view projection feature data of the object to be operated, wherein the multi-view image data includes synchronous observation information of the object to be operated at different viewing angles, the spatial projection processing includes multi-view projection matrix construction and depth information fusion, and the multi-view projection feature data includes spatial distribution information of surface features of the object at different viewing angles; Constructing a pose solution matrix based on the multi-view projection feature data, and updating the rotation parameters and translation parameters in the pose solution matrix through an iterative optimization algorithm until the reprojection error converges to a preset threshold, thereby obtaining optimized rotation parameters and translation parameters, wherein the pose solution matrix includes initial estimated values of the rotation parameters and translation parameters, and the iterative optimization algorithm includes error minimization based on gradient descent and parameter convergence verification; The pose parameter set is generated based on the optimized rotation parameters and translation parameters, and the pose parameter set includes attitude angle parameters and spatial position parameters. The attitude angle parameters are used to describe the rotation state of the object in the three-dimensional space, and the spatial position parameters are used to describe the translation state of the object in the three-dimensional space.
7. The target collaborative control method based on deep learning according to claim 6, characterized in that: The constructing of a pose solution matrix based on the multi-view projection feature data includes: Obtaining observation position data of the operation execution device at different viewing angles, wherein the observation position data is obtained through encoder feedback and kinematic forward solution calculation; Constructing an observation matrix based on the observation position data, wherein the observation matrix includes camera perspective transformation parameters and object projection relationships; Constructing initial parameters of a pose solution matrix based on the correspondence between the multi-view projection feature data and the observation matrix, wherein the correspondence is determined by feature point matching and projection error calculation; The initial parameters in the pose solution matrix are adjusted by a nonlinear optimization algorithm to minimize the error between the multi-view projection feature data and the observation matrix, thereby obtaining the final pose solution matrix. The error value is optimized by reprojection error metric and covariance matrix weighting.
8. The target collaborative control method based on deep learning according to claim 1, characterized in that: The step of generating a control instruction based on the pose parameter set and the object physical property parameters corresponding to the category identification information, and driving the operation execution device to complete the object grasping action of the object to be operated, includes: Calling a preset grasping strategy library according to the category identification information, and selecting a gripping force parameter and motion trajectory constraint condition that matches the category identification information; Inputting the pose parameter set into a motion planning model, and generating motion trajectory data and end effector posture data of the operation execution device through the motion planning model based on inverse kinematics and path optimization algorithm, wherein the motion trajectory data includes a joint motion sequence and velocity planning parameters and satisfies the motion trajectory constraint conditions; Generate a path control signal according to the motion trajectory data, wherein the path control signal drives the operation execution device to move to the target grasping position through pulse width modulation and a servo drive interface; Based on the end effector posture data and the clamping force parameters, an end effector posture control signal and a clamping force control signal are generated, and based on the end effector posture control signal and the clamping force control signal, the end effector's rotation axis control and the clamping force feedback adjustment are respectively performed to trigger the clamping action to complete the grasping of the object to be operated.
9. A target collaborative control system based on deep learning, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the deep learning-based target collaborative control method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Visual guidance system for movable grabbing device
CN115082926A
Wheat field wheat scab state evaluation method and system and electronic equipment
CN117392668A