A computer vision-based robot virtual commissioning path optimization method

By using multi-view visual acquisition and calibration technology, combined with Canni edge detection and an improved Poisson flow generation model, the automatic acquisition and consistency of workstation geometry and pose are achieved, solving the inconsistency problem of existing virtual debugging methods and improving trajectory generation efficiency and deployment reliability.

CN122425680APending Publication Date: 2026-07-21YIGONGCHANG DIGITAL TECHNOLOGY (CHONGQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing robot virtual debugging methods rely heavily on prior knowledge of workstation geometry, clamping posture, and environmental layout, and require high calibration accuracy. Furthermore, when faced with clamping deviations, part batch differences, and fixture wear, inconsistencies can easily arise between the virtual scene and the real workstation, leading to deviations between virtual verification results and actual execution. This makes it difficult to balance safety margins, process constraints, and execution feasibility.

Method used

By employing multi-view visual acquisition and calibration, Cannibal edge detection, improved Poisson flow generation model and voxel registration technology, and through rolling guided filtering pre-smoothing, sub-pixel spline extreme value scanning and annular shell prior module, the automatic acquisition and consistency of station geometry and pose are achieved, and uncertainty adaptive path optimization is performed. Combined with virtual simulation verification, a closed-loop iteration is formed.

Benefits of technology

It reduces errors in manual modeling and calibration, improves consistency between virtual debugging and on-site testing, enhances trajectory generation efficiency and deployment reliability, and ensures that collision safety and process constraints are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122425680A_ABST
    Figure CN122425680A_ABST
Patent Text Reader

Abstract

The application discloses a kind of robot virtual debugging path optimization methods based on computer vision, it is related to computer vision technical field, including: acquisition station multi-view image data, obtain robot base coordinate system conversion relationship;Perform Canny edge detection, output subpixel edge point set;Dominant sparse voxel field is constructed to edge, and pose parameter and geometric boundary parameter are obtained;Improved Poisson flow generation model is constructed, and three-dimensional geometry and uncertainty information are obtained;Boundary consistent deformable geometry fusion is executed, and final pose parameter is obtained;Adaptive two-domain path optimization is carried out, and candidate robot trajectory is output;Virtual simulation verification is carried out, and deployable robot trajectory is output.The application realizes the deployable robot optimal trajectory of virtual debugging stage by fusing Canny edge detection and improved Poisson flow generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method for optimizing virtual debugging paths for robots based on computer vision. Background Technology

[0002] Current virtual robot debugging typically relies on offline modeling and manual calibration to construct the workstation scene. Before the robot goes online, engineers manually create 3D models of the workpiece, fixtures, tooling, equipment body, and surrounding obstacles based on process documents, equipment layout diagrams, or on-site measurement data. They then manually set the coordinate system, calibrate reference points, and input clamping positions and posture parameters to ensure that the geometry and pose relationships in the virtual environment are as consistent as possible with the actual site. Subsequently, trajectory planning, interference checks, reachability analysis, posture verification, and cycle time evaluation are performed in the virtual environment, and the trajectory or program parameters are then exported and deployed to the actual robot controller. The current method is highly dependent on the prior knowledge of workstation geometry, clamping posture, and environmental layout, and has high requirements for calibration accuracy and model maintenance. When there are clamping deviations, part batch differences, fixture wear, equipment relocation, workstation modifications, or spatial obstructions on-site, inconsistencies between the virtual scene and the real workstation are likely to occur, leading to deviations between virtual verification results and actual execution, or even the risk of interference.

[0003] To reduce human intervention, existing technologies have attempted to introduce computer vision to acquire workstation information for path generation and optimization. However, industrial workstations are generally plagued by glare, low texture, occlusion, and noise interference. Traditional edge detection outputs are prone to edge breakage, positioning misalignment, and repetitive responses, thus affecting the stability of pose estimation and geometric boundary extraction. When multi-view observations are insufficient or depth data is sparse, 3D reconstruction results are prone to holes or deformations, making collision detection and accessibility assessment unreliable. This makes it difficult to balance safety margins, process constraints, and feasibility in path optimization. Existing virtual commissioning processes mostly rely on offline one-time modeling and planning, lacking adaptive modeling and simulation feedback closed-loop update mechanisms for uncertain geometry. This makes it difficult to achieve an integrated and stable process that automatically aligns with real workstations, solves multi-constraint trajectories, performs simulation closed-loop iteration, and provides deployable output.

[0004] Therefore, how to provide a computer vision-based method for optimizing robot virtual debugging paths is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a computer vision-based method for optimizing virtual debugging paths in robots. This invention comprehensively utilizes multi-view visual acquisition and calibration, Cannibal edge detection, an improved Poisson flow generation model, voxel registration, and multi-constraint trajectory solving techniques. It forms a complete process from multi-view image data acquisition at the workstation, sub-pixel edge extraction, edge-dominated sparse voxel field construction and coarse pose localization, 3D geometric completion and uncertainty estimation, boundary-consistent deformable geometric fusion, dual-domain path optimization, to virtual simulation verification and deployable program output. The edge detection structure incorporates rolling guided filtering pre-smoothing and sub-pixel spline extreme value scanning; the generation model structure incorporates toroidal shell priors, biphasic radius flipping, and local subspace convergent projection. This achieves automatic acquisition and consistency of workstation geometry and pose, and outputs candidate trajectories that satisfy collision safety, accessibility, and process constraints under the condition of considering geometric uncertainties, which are then verified through simulation to form a closed-loop iteration. Compared with existing technologies, this invention reduces manual modeling and calibration errors, improves the consistency between virtual debugging and on-site operations, and enhances trajectory generation efficiency and deployment reliability.

[0006] A robot virtual debugging path optimization method based on computer vision according to an embodiment of the present invention includes: Collect multi-view image data of the workstation and complete camera calibration and coordinate system establishment to obtain the transformation relationship between camera parameters and robot base coordinate system; Canni edge detection is performed on multi-view image data of the workstation. Noise suppression and main contour preservation are achieved by pre-smoothing with rolling guided filtering. Subpixel spline extreme value scanning is used for fine-grained extreme value localization, and subpixel edge point set is output. An edge-dominated sparse voxel field is constructed based on sub-pixel edge point sets, and voxel gradient consistent registration is used for coarse pose localization to obtain initial pose parameters and initial geometric boundary parameters. An improved Poisson flow generation model is constructed. The initial pose parameters and initial geometric boundary parameters are initialized by annular shell spatial sampling based on the annular shell prior module. The radial evolution process of the generated flow is reversed by the two-phase radius reversal module. The neighborhood subspace projection is introduced by the local subspace convergence projection module to obtain the three-dimensional geometric representation and geometric uncertainty information. Based on the 3D geometric representation and the sub-pixel edge point set, perform boundary-consistent deformable geometry fusion, iteratively transform the initial geometric boundary parameters into the final geometric boundary parameters, and obtain the final pose parameters; Based on the final pose parameters and final geometric boundary parameters, an uncertainty-adaptive dual-domain path optimization is performed to output candidate robot trajectories; The candidate robot trajectory is verified by virtual simulation, and the deployable robot trajectory and corresponding robot program parameters are output.

[0007] Optionally, the multi-view image data of the workstation includes a multi-view color image sequence, a grayscale image sequence, a depth map sequence, a camera pose sequence, a timestamp sequence, and a lighting status parameter sequence.

[0008] Optionally, completing camera calibration and coordinate system establishment, and obtaining the transformation relationship between camera parameters and the robot base coordinate system, includes: Imaging equipment is set up in the workstation area to collect multi-view image data of the workstation, and the data from each viewpoint are synchronized and associated according to the timestamp; Camera calibration is performed on the imaging device to obtain the camera intrinsic parameter matrix and distortion parameters. The camera intrinsic parameter matrix is ​​a three-row, three-column matrix containing focal length parameters and principal point coordinate parameters. The distortion parameters include radial distortion coefficients and tangential distortion coefficients. Distortion correction is performed on multi-view color images, grayscale images, and depth maps based on the camera intrinsic parameter matrix and distortion parameters. Establish the camera coordinate system, the workstation coordinate system, and the robot base coordinate system, and obtain the rigid body transformation relationship from the camera coordinate system to the workstation coordinate system and from the workstation coordinate system to the robot base coordinate system. Based on the matrix multiplication of the two rigid body transformation relationships, obtain the transformation relationship between the camera parameters and the robot base coordinate system.

[0009] Optionally, the output sub-pixel edge point set includes: Rolling guided filtering pre-smoothing is performed on the grayscale image in the multi-view image data of the workstation. The rolling guided filtering pre-smoothing is performed according to the number of iterations. The initial value of the iteration is the original grayscale image. In each iteration, the pre-smoothed grayscale image output by the previous iteration is used as the guide image of the current iteration and the pre-smoothed grayscale image of the current iteration is output to obtain the final pre-smoothed grayscale image. The horizontal and vertical gradient maps are calculated based on the final pre-smoothed grayscale image. The magnitude of each pixel in the gradient magnitude map is the square root of the sum of the squares of the horizontal and vertical gradient values. The direction of each pixel in the gradient direction map is the arctangent of the vertical and horizontal gradient values, thus obtaining the gradient magnitude map and gradient direction map. Subpixel spline extreme value scanning and positioning is performed on the gradient magnitude map. The scanning direction is determined by the gradient direction map. Three gradient magnitudes are selected with the current pixel as the center: the forward sampling point, the center sampling point, and the backward sampling point. Quadratic spline interpolation is used to fit the function of gradient magnitude changing with position, and the position where the first derivative of the function is zero is obtained to get the subpixel extreme value position. The subpixel extreme value position is used as the edge point position. The set of strong edge points and the set of weak edge points are obtained based on the dual threshold determination. The dual threshold determination uses a high threshold and a low threshold. The high threshold is the amplitude corresponding to the preset percentile of the gradient amplitude map, and the low threshold is the high threshold multiplied by a scaling factor. Perform a hysteresis connection on the set of strong edge points, search for edge points connected to weak edge points in the eight neighborhoods of the strong edge points and add the connected weak edge points to the edge point set, delete the weak edge points that are not connected to any strong edge point, and obtain the sub-pixel edge point set.

[0010] Optionally, obtaining the initial pose parameters and initial geometric boundary parameters includes: The sub-pixel edge point set is divided into a three-dimensional voxel grid according to the spatial range of the work station coordinate system. Each voxel corresponds to a spatial cubic region. Sub-pixel edge points falling into the same voxel are aggregated into the edge occupancy marker of the current voxel. An edge-dominated sparse voxel field is constructed in a three-dimensional voxel mesh. The edge-dominated sparse voxel field consists of all voxels whose edges are occupied and marked as true. Each occupied voxel records the coordinates of the voxel center point and the average gradient direction vector of the edge points within the voxel. Voxel gradient consistency registration is performed between the edge-dominated sparse voxel field and the preset target geometric template. The target geometric template is represented as a set of template voxels in the work station coordinate system, and each template voxel contains the coordinates of the template voxel center point and the template gradient direction vector. By searching for rigid body transformations in the work station coordinate system to minimize the registration cost, coarse pose localization results are obtained. The coarse localization result is output as the initial pose parameters. Under the initial pose parameters, the outer surface voxels of the target geometric template are projected onto the station coordinate system to obtain the initial geometric boundary parameters.

[0011] Optionally, obtaining the three-dimensional geometric representation and geometric uncertainty information includes: An improved Poisson flow generation model is constructed. The improved Poisson flow generation model sets a ring-shell prior module at the initial sampling position of the original Poisson flow generation model, sets a two-phase radius flipping module at the middle position of the radial evolution process, and sets a local subspace convergence projection module at the output position of the final generation stage. The toroidal shell prior module is executed to perform initial state sampling within the toroidal shell space enclosed by the inner and outer spherical surfaces at the same center point. The radii of the inner and outer spherical surfaces are preset real numbers, and the radius of the outer spherical surface is greater than the radius of the inner spherical surface. The initial state sampling yields the initial state vector, which is then input into the Poisson flow generation process. The two-phase radius flipping module is executed. When the intermediate state sequence output by the Poisson flow generation process satisfies the radial interval, the radial component is flipped and mapped. The flipping and mapping transforms the radial component from the first radial value to the second radial value. The flipped and mapped radial component and the corresponding spatial variable component form the flipped intermediate state and continue to be input into the Poisson flow generation process until the final generation result is obtained. The local subspace convergence projection module is executed. Neighborhood sample points are retrieved in the spatial neighborhood corresponding to the generated final result. The covariance matrix of the neighborhood sample points is calculated and the eigenvectors of the covariance matrix are obtained to obtain the local subspace basis vector set. The spatial variable components of the generated final result are orthogonally projected according to the local subspace basis vector set to obtain the projected spatial variable components. The projected spatial variable components are mapped to the radial components of the generated final result to form a three-dimensional geometric representation. The improved Poisson flow generation model is trained, training sample pairs are constructed, and the improved Poisson flow generation model is executed on each training sample pair to obtain the predicted 3D geometric representation. The loss function between the predicted 3D geometric representation and the real 3D geometric label is calculated and the model parameters are updated. The loss function includes the mean Euclidean distance term of the corresponding points in the point cloud and the mean absolute error term of the sampling points in the distance field, until the training termination condition is met.

[0012] Optionally, the step of iteratively transforming the initial geometric boundary parameters into final geometric boundary parameters and obtaining the final pose parameters includes: The three-dimensional geometric representation is transformed into a set of boundary surfaces of the target geometry in the work station coordinate system. The set of boundary surfaces of the target geometry is a set of surface points obtained by performing neighborhood connectivity clustering on the point cloud set. Normal vectors are calculated for the surface elements in the set of boundary surfaces of the target geometry. The sub-pixel edge point set is back-projected onto the work station coordinate system according to the camera pose sequence to obtain the edge three-dimensional point set. Each point in the edge three-dimensional point set is determined by the intersection of the light rays emitted from the camera center and the target geometry boundary surface set. The intersection points are obtained by solving the intersection equation of the camera ray parameter equation and the boundary surface element. Within the boundary range defined by the edge 3D point set, perform boundary-consistent deformable geometry fusion, establish the deformable boundary model determined by the initial geometric boundary parameters and perform iterative updates, and update the rotation matrix and translation vector by minimizing the sum of the Euclidean distances from the boundary vertices after rigid body transformation to the corresponding target geometry boundary surface elements. When the iteration terminates, output the final geometric boundary parameters and the final pose parameters.

[0013] Optionally, the output candidate robot trajectory includes: Based on the final pose parameters and final geometric boundary parameters, construct the environmental geometry set and the task target set, and generate the uncertainty region set based on the geometric uncertainty information; A discrete feasible region is constructed and a safety-critical pose sequence is selected. The discrete feasible region consists of a set of states obtained by discrete sampling of the robot joint space. The safety-critical pose sequence is obtained by connecting the discrete states in the discrete feasible region that satisfy collision safety constraints, accessibility constraints and process constraints in the order of the task objective set. Candidate robot trajectories are generated and output in a continuous space. Collision safety constraints, accessibility constraints, and process constraints are calculated for each time sampling point on the candidate robot trajectory, and trajectory segment reconstruction is performed for sampling points that do not meet the constraints.

[0014] Optionally, the output can deploy robot trajectories and corresponding robot program parameters, including: Virtual simulation verification is performed on the candidate robot trajectory. Kinematic and collision simulations are performed at the trajectory time sampling points, and the minimum collision distance sequence, joint limit margin sequence, velocity margin sequence, and process parameter deviation sequence are output. Simulation results are generated based on the minimum collision distance sequence, joint limit margin sequence, velocity margin sequence, and process parameter deviation sequence. When the simulation result is "fail", the candidate robot trajectory is updated repeatedly. When the simulation result is "pass", the deployable robot trajectory and the corresponding robot program parameters are output. The deployable robot trajectory is discretized into a sequence of control instructions.

[0015] The beneficial effects of this invention are: This invention combines multi-view visual acquisition and calibration at the workstation, Canny edge detection, edge-dominated sparse voxel field registration, and an improved Poisson flow generation model to achieve automatic acquisition and consistency of the pose and geometric boundaries of workpieces, fixtures, and obstacles. Compared with existing virtual debugging methods that rely on manual modeling, manual alignment, and experience-based parameter tuning, this invention utilizes rolling guided filtering pre-smoothing and sub-pixel spline extreme value scanning to obtain a highly stable and highly accurate edge point set. Furthermore, it generates a three-dimensional geometric representation and geometric uncertainty information through annular shell priors, biphasic radius flipping, and local subspace convergent projection. This reduces modeling deviations and calibration errors caused by complex working conditions such as occlusion, reflection, and low texture, thereby improving the consistency and robustness between the virtual scene and the real workstation.

[0016] This invention constructs an uncertainty-adaptive dual-domain path optimization process based on final pose parameters and final geometric boundary parameters, and forms a closed-loop iteration by combining virtual simulation verification. It can output deployable robot trajectories and program parameters that meet collision safety, accessibility, and process constraints. Compared with existing path generation methods that involve one-time planning or are based solely on static models, this invention selects safety-critical pose sequences in the discrete feasible domain and reconstructs the trajectory in continuous space. Simultaneously, it establishes a more reliable safety margin setting based on geometric uncertainty information, improving trajectory executability and safety margin. Furthermore, it reduces on-site trial-and-error costs through simulation judgment and iterative updates, improving debugging efficiency and deployment reliability. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a robot virtual debugging path optimization method based on computer vision proposed in this invention; Figure 2 This is a structural block diagram of the Canney edge detection method for a robot virtual debugging path optimization method based on computer vision proposed in this invention. Figure 3 This is a functional diagram of the improved Poisson flow generation model for a robot virtual debugging path optimization method based on computer vision proposed in this invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figure 1 , Figure 2 and Figure 3 A computer vision-based method for optimizing virtual debugging paths in robots, comprising: Collect multi-view image data of the workstation and complete camera calibration and coordinate system establishment to obtain the transformation relationship between camera parameters and robot base coordinate system; Canni edge detection is performed on multi-view image data of the workstation. Noise suppression and main contour preservation are achieved by pre-smoothing with rolling guided filtering. Subpixel spline extreme value scanning is used for fine-grained extreme value localization, and subpixel edge point set is output. An edge-dominated sparse voxel field is constructed based on sub-pixel edge point sets, and voxel gradient consistent registration is used for coarse pose localization to obtain initial pose parameters and initial geometric boundary parameters. An improved Poisson flow generation model is constructed. The initial pose parameters and initial geometric boundary parameters are initialized by annular shell spatial sampling based on the annular shell prior module. The radial evolution process of the generated flow is reversed by the two-phase radius reversal module. The neighborhood subspace projection is introduced by the local subspace convergence projection module to obtain the three-dimensional geometric representation and geometric uncertainty information. Based on the 3D geometric representation and the sub-pixel edge point set, perform boundary-consistent deformable geometry fusion, iteratively transform the initial geometric boundary parameters into the final geometric boundary parameters, and obtain the final pose parameters; Based on the final pose parameters and final geometric boundary parameters, an uncertainty-adaptive dual-domain path optimization is performed to output candidate robot trajectories; The candidate robot trajectory is verified by virtual simulation, and the deployable robot trajectory and corresponding robot program parameters are output.

[0020] In this embodiment, the multi-view image data of the workstation includes a multi-view color image sequence, a grayscale image sequence, a depth map sequence, a camera pose sequence, a timestamp sequence, and a lighting state parameter sequence.

[0021] In this embodiment, completing camera calibration and coordinate system establishment, and obtaining the transformation relationship between camera parameters and the robot base coordinate system includes: Imaging equipment is set up in the workstation area to collect multi-view image data of the workstation, and the data from each viewpoint are synchronized and associated according to the timestamp; Camera calibration is performed on the imaging device to obtain the camera intrinsic parameter matrix and distortion parameters. The camera intrinsic parameter matrix is ​​a three-row, three-column matrix containing focal length parameters and principal point coordinate parameters. The distortion parameters include radial distortion coefficients and tangential distortion coefficients. Distortion correction is performed on multi-view color images, grayscale images, and depth maps based on the camera intrinsic parameter matrix and distortion parameters. Establish the camera coordinate system, workstation coordinate system, and robot base coordinate system, and obtain the rigid body transformation relationships from the camera coordinate system to the workstation coordinate system and from the workstation coordinate system to the robot base coordinate system. Based on the matrix multiplication of the two rigid body transformation relationships, obtain the transformation relationship between the camera parameters and the robot base coordinate system. Specifically, establishing the camera coordinate system, workstation coordinate system, and robot base coordinate system involves: The camera coordinate system is based on the camera's optical center as the origin, with the optical axis direction defined as the Z-axis, the horizontal direction of the imaging plane defined as the X-axis, and the Y-axis determined by the right-hand rule. The workstation coordinate system is established using a calibration plate fixed on the workstation. Three non-collinear reference points are selected, and the origin, X-axis direction, and Y-axis direction are defined respectively. The Z-axis is determined by the X-axis and Y-axis. The robot base coordinate system adopts the base coordinate system of the robot controller. The rigid body transformation from the camera coordinate system to the workstation coordinate system is achieved by detecting the corner points of the calibration plate and combining them with the depth map to obtain the three-dimensional coordinates of the corner points. After corresponding with the workstation reference points, the rotation matrix and translation vector are solved using the least squares method to minimize the mean square error between corresponding points. The iteration terminates when the root mean square error change is less than 0.1 pixels or the number of iterations reaches 50. The rigid body transformation from the workstation coordinate system to the robot base coordinate system is achieved by the robot end effector touching each reference point in sequence and recording the three-dimensional coordinates of the end effector in the base coordinate system. The rotation matrix and translation vector are solved using the same least squares method. The iteration terminates when the root mean square error change is less than 0.2 millimeters or the number of iterations reaches 50. The two rigid body transformations are multiplied in sequence to obtain the transformation relationship between the camera parameters and the robot base coordinate system.

[0022] In this embodiment, the output sub-pixel edge point set includes: A rolling guided filtering pre-smoothing process is performed on the grayscale images in the multi-view image data of the workstation. The rolling guided filtering pre-smoothing is performed according to the number of iterations. The initial value of each iteration is the original grayscale image. In each iteration, the pre-smoothed grayscale image output from the previous iteration is used as the guiding image for the current iteration, and the pre-smoothed grayscale image of the current iteration is output to obtain the final pre-smoothed grayscale image. Specifically, the rolling guided filtering pre-smoothing process on the grayscale images in the multi-view image data of the workstation is as follows: For each frame of grayscale image data from the multi-view image data of the workstation, the pixel value is first normalized and used as the initial input image. Then, rolling guided filtering pre-smoothing is performed. In each iteration, the output image of the previous round is used as the guide image and the input image. Within a local window centered on the current pixel, the spatial distance from the neighboring pixels to the center pixel and the grayscale difference of the guide image are calculated to generate a two-factor weight. The neighboring pixels are weighted and summed according to the weights and divided by the weight sum to obtain the smoothing result of the pixels, thereby obtaining the output image of the current round. The local window is 7×7, the spatial distance control parameter is 2.0 pixels, the guide grayscale difference control parameter is 0.1, and the number of iterations is 3 rounds. If the average absolute difference between the output images of the 2nd and 3rd rounds is less than 0.002, the process is terminated early and the final pre-smoothed grayscale image is output. The horizontal and vertical gradient maps are calculated based on the final pre-smoothed grayscale image. In the gradient magnitude map, the magnitude of each pixel is the square root of the sum of the squares of the horizontal and vertical gradient values. In the gradient direction map, the direction of each pixel is the arctangent of the vertical and horizontal gradient values. This yields the gradient magnitude map and gradient direction map. Specifically, the calculation of the horizontal and vertical gradient maps based on the final pre-smoothed grayscale image is as follows: Based on the final pre-smoothed grayscale image, a three-row, three-column differential convolution kernel is used to calculate the horizontal gradient map and the vertical gradient map respectively. For each pixel, a 3×3 neighborhood is taken as the center. The grayscale of the neighborhood is multiplied element-wise by the horizontal differential kernel and summed to obtain the horizontal gradient value of the pixel. The grayscale of the neighborhood is multiplied element-wise by the vertical differential kernel and summed to obtain the vertical gradient value of the pixel. The gradient magnitude is calculated using the horizontal gradient value and the vertical gradient value of each pixel. The arctangent is obtained by the ratio of the vertical gradient value to the horizontal gradient value to obtain the gradient direction. Thus, the horizontal gradient map, the vertical gradient map, the gradient magnitude map, and the gradient direction map are obtained. Subpixel spline extreme value scanning and localization are performed on the gradient magnitude map. The scanning direction is determined by the gradient direction map. Along the scanning direction, three gradient magnitude points—the forward sampling point, the center sampling point, and the backward sampling point—are selected with the current pixel as the center. Quadratic spline interpolation is used to fit the function of gradient magnitude as a function of position, and the location where the first derivative of the function is zero is obtained. This subpixel extreme value location is then used as the edge point location. Specifically, the quadratic spline interpolation is used to fit the function of gradient magnitude as a function of position and the location where the first derivative of the function is zero. For each candidate pixel, the scanning direction is first determined by the gradient direction. The gradient magnitudes of the three sampling positions (backward, center, and forward) are taken along the scanning direction. The distances of the three sampling positions relative to the center pixel are defined as negative one, zero, and positive one pixels, respectively. The corresponding magnitudes are recorded as backward magnitude, center magnitude, and forward magnitude, respectively. A quadratic polynomial is constructed, and the magnitude is equal to the quadratic term coefficient multiplied by the square of the distance, plus the linear term coefficient multiplied by the distance, plus the constant term. The polynomial is set to equal the backward magnitude, center magnitude, and forward magnitude when the distance is negative one, zero, and positive one, respectively. A set of three equations is obtained, and the quadratic term coefficient, linear term coefficient, and constant term are solved. Take the first derivative of the obtained quadratic polynomial. The first derivative is twice the coefficient of the quadratic term multiplied by the distance plus the coefficient of the linear term. Set the first derivative to zero to get the distance of the extreme value position, which is equal to the coefficient of the negative linear term divided by twice the coefficient of the quadratic term. Limit the extreme value distance to the range of negative one to positive one pixels. Translate the center pixel coordinates along the scanning direction by the extreme value distance to obtain the sub-pixel extreme value position. Use the sub-pixel extreme value position as the edge point position. Strong edge point sets and weak edge point sets are obtained based on a dual-threshold determination. The dual-threshold determination uses a high threshold and a low threshold. The high threshold is the amplitude corresponding to a preset percentile of the gradient magnitude map, and the low threshold is the high threshold multiplied by a scaling factor. Specifically, the strong edge point sets and weak edge point sets are obtained based on the dual-threshold determination as follows: The magnitude values ​​of all non-zero pixels in the gradient magnitude map are statistically analyzed and sorted from smallest to largest. The magnitude values ​​corresponding to 90% are taken as the high threshold, and the high threshold is multiplied by 0.4 to obtain the low threshold. The gradient magnitude map is traversed pixel by pixel. When the pixel magnitude value is greater than or equal to the high threshold, the pixel is marked as a strong edge point and added to the strong edge point set. When the pixel magnitude value is less than the high threshold but greater than or equal to the low threshold, the pixel is marked as a weak edge point and added to the weak edge point set. When the pixel magnitude value is less than the low threshold, the pixel is marked as a non-edge point and discarded. This results in the strong edge point set and the weak edge point set. Perform a hysteresis connection on the set of strong edge points, search for edge points connected to weak edge points in the eight neighborhoods of the strong edge points and add the connected weak edge points to the edge point set, delete the weak edge points that are not connected to any strong edge point, and obtain the sub-pixel edge point set.

[0023] In this embodiment, obtaining the initial pose parameters and initial geometric boundary parameters includes: The sub-pixel edge point set is divided into a three-dimensional voxel mesh according to the spatial range of the workstation coordinate system. Each voxel corresponds to a spatial cubic region. Sub-pixel edge points falling into the same voxel are aggregated into the edge occupancy marker of the current voxel. Specifically, the sub-pixel edge point set is divided into a three-dimensional voxel mesh according to the spatial range of the workstation coordinate system as follows: In the workstation coordinate system, the coverage area of ​​the three-dimensional voxel mesh is determined. The minimum and maximum values ​​of the three-dimensional coordinates of all sub-pixel edge points in the X, Y, and Z directions are taken respectively. The mesh boundary is extended by 10 mm in each direction. The voxel size is set to 5 mm. The number of voxels in each direction is calculated. For each sub-pixel edge point, the three-dimensional coordinates in the workstation coordinate system are read. The minimum boundary of the mesh in the corresponding direction is subtracted from the coordinates and then divided by the voxel size and rounded down to obtain the voxel index. Edge points with the same voxel index are grouped into the same voxel to obtain the three-dimensional voxel mesh and edge occupancy markers divided according to the workstation space range. An edge-dominated sparse voxel field is constructed in a 3D voxel mesh. This field consists of all voxels whose edges are occupied and marked as true. Each occupied voxel records the coordinates of its center point and the average gradient direction vector of its internal edge points. Specifically, the construction of the edge-dominated sparse voxel field in the 3D voxel mesh is as follows: Traverse all voxels and filter voxels whose edge occupancy is marked as true. Add the voxels to the edge-dominated sparse voxel field. For each added occupancy voxel, the coordinates of the voxel center point are obtained by adding the voxel index multiplied by the voxel size and half the voxel size to the minimum boundary values ​​of the voxel in the X, Y, and Z directions, respectively. Collect the gradient directions corresponding to all sub-pixel edge points within the voxel. Convert the gradient direction of each edge point into a unit direction vector and then perform vector summation. Divide the summation result by the number of edge points within the voxel to obtain the average gradient direction vector. The edge-dominated sparse voxel field is obtained by the coordinates of the occupancy voxel center point and the average gradient direction vector. Voxel gradient consistency registration is performed between the edge-dominant sparse voxel field and the preset target geometric template. The target geometric template is represented as a set of template voxels in the workstation coordinate system, and each template voxel contains the coordinates of the template voxel center point and the template gradient direction vector. By searching for rigid body transformations in the workstation coordinate system to minimize the registration cost, a coarse pose localization result is obtained. Specifically, the voxel gradient consistency registration between the edge-dominant sparse voxel field and the preset target geometric template is performed as follows: The target geometric template is pre-discretized into a set of template voxels, and the center point coordinates and template gradient direction vector are stored for each template voxel. The center point coordinates and average gradient direction vector of the voxels dominated by the edge sparse voxel field are used as the observation set to perform registration search. For rotation, a hierarchical enumeration is adopted. The rotation angles around the X-axis, Y-axis, and Z-axis are generated in a step of 5 degrees within the range of -30 degrees to +30 degrees and evaluated one by one. For each candidate rotation, the center point coordinates and direction vector of the observation set are rotated. In the work position coordinate system, the translation vector is enumerated in the X, Y, and Z directions in a step of 10 mm within the range of -100 mm to +100 mm to form candidate rigid body transformations. For each candidate rigid body transformation, the rotated observation center point is translated again and the nearest template voxel is searched in the template voxel set. The Euclidean distance between the two center points is calculated and the cosine difference of the angle between the two direction vectors is calculated. The registration cost is the sum of the distances of all observation voxels and the sum of the cosine differences. The candidate rigid body transformation with the minimum registration cost is taken as the coarse pose localization result. The coarse localization result is output as the initial pose parameters. Under the initial pose parameters, the outer surface voxels of the target geometric template are projected onto the station coordinate system to obtain the initial geometric boundary parameters.

[0024] In this embodiment, obtaining the three-dimensional geometric representation and geometric uncertainty information includes: An improved Poisson flow generation model is constructed. The improved Poisson flow generation model sets a ring-shell prior module at the initial sampling position of the original Poisson flow generation model, sets a two-phase radius flipping module at the middle position of the radial evolution process, and sets a local subspace convergence projection module at the output position of the final generation stage. The toroidal shell prior module is executed, and initial state sampling is performed within the toroidal shell space enclosed by the inner and outer spherical surfaces at the same center point. The radii of the inner and outer spherical surfaces are preset real numbers, with the outer spherical radius being larger than the inner spherical radius. The initial state sampling yields an initial state vector, which is then input into the Poisson flow generation process. Specifically, the initial state sampling within the toroidal shell space enclosed by the inner and outer spherical surfaces at the same center point is performed as follows: In the workstation coordinate system, the target center point determined by the initial pose parameters is used as the sampling center. An annular shell space is formed by setting the inner sphere radius to 0.6 meters and the outer sphere radius to 1.0 meters. For each sampling, a three-dimensional direction vector is generated. The three components are taken from the uniform distribution of the interval from negative one to positive one and then normalized to obtain the unit direction vector. A radial length is taken from the uniform distribution of the interval from 0.6 meters to 1.0 meters as the sampling radius. The unit direction vector is scaled according to the sampling radius and superimposed on the sampling center point to obtain the initial spatial coordinates. The two-phase radius reversal module is executed. When the intermediate state sequence output by the Poisson flow generation process satisfies the radial interval, the radial component is reversed and mapped. The reversal mapping transforms the radial component from a first radial value to a second radial value. The reversed radial component and the corresponding spatial variable component form the reversed intermediate state and continue to be input into the Poisson flow generation process until the final generation result is obtained. Specifically, the reversal mapping transforms the radial component from the first radial value to the second radial value as follows: During the generation of Poisson flow, the radial component of the intermediate state obtained at each time step is read as the first radial value. When the first radial value falls into the trigger range of 0.75 m to 0.85 m, the flip mapping is performed. The flip mapping uses a preset constant of 1.0 m as the flip reference. The first radial value is truncated at the lower limit to avoid numerical instability. The truncation threshold is 0.05 m. The second radial value is calculated as the square of the flip reference divided by the first radial value. The second radial value is limited to the range of 0.6 m to 1.0 m. If it exceeds the range, it is truncated to 0.6 m. The local subspace convergence projection module is executed. Neighborhood sample points are retrieved in the spatial neighborhood corresponding to the generated final result. The covariance matrix of the neighborhood sample points is calculated and the eigenvectors of the covariance matrix are obtained to obtain the local subspace basis vector set. The spatial variable components of the generated final result are orthogonally projected according to the local subspace basis vector set to obtain the projected spatial variable components. The projected spatial variable components are mapped to the radial components of the generated final result to form a three-dimensional geometric representation. The improved Poisson flow generation model is trained by constructing training sample pairs. For each training sample pair, the improved Poisson flow generation model is executed to obtain a predicted 3D geometric representation. The loss function between the predicted 3D geometric representation and the true 3D geometric label is calculated, and the model parameters are updated. The loss function includes the mean Euclidean distance term of the corresponding points in the point cloud and the mean absolute error term of the sampling points in the distance field. This process continues until the training termination condition is met. Specifically, training the improved Poisson flow generation model involves: During training, training sample pairs are constructed for each workstation sample. The input includes multi-view image data of the workstation, sub-pixel edge point set, initial pose parameters, and initial geometric boundary parameters. The ground truth label is the ground truth distance field corresponding to the workstation. In each iteration, sample pairs with a batch size of 8 are randomly selected from the training set. The input is subjected to toroidal shell prior initialization, biphasic radius flipping, and local subspace convergence projection to obtain the predicted 3D geometric representation. Then, the loss function is calculated and the model parameters are updated in reverse. The loss function is obtained by adding two parts. The first part is the mean Euclidean distance between corresponding points in the predicted point cloud and the ground truth point cloud, and the corresponding points are established by nearest neighbor matching. The second part is the mean absolute error between the predicted distance value and the ground truth distance value in the voxel grid. Training is terminated when the loss function does not decrease by more than 0.0005 in 10 consecutive evaluations or when the number of training rounds reaches 200. The parameters of the improved Poisson flow generation model after training are output.

[0025] In this embodiment, the step of iteratively transforming the initial geometric boundary parameters into the final geometric boundary parameters and obtaining the final pose parameters includes: The three-dimensional geometric representation is transformed into a set of boundary surfaces of the target geometry in the work station coordinate system. The set of boundary surfaces of the target geometry is a set of surface points obtained by performing neighborhood connectivity clustering on the point cloud set. Normal vectors are calculated for the surface elements in the set of boundary surfaces of the target geometry. The sub-pixel edge point set is back-projected onto the work station coordinate system according to the camera pose sequence to obtain the edge three-dimensional point set. Each point in the edge three-dimensional point set is determined by the intersection of the light rays emitted from the camera center and the target geometry boundary surface set. The intersection points are obtained by solving the intersection equation of the camera ray parameter equation and the boundary surface element. Within the boundary defined by the edge 3D point set, a boundary-consistent deformable geometric fusion is performed. A deformable boundary model determined by initial geometric boundary parameters is established and iteratively updated. Simultaneously, the rotation matrix and translation vector are updated by minimizing the sum of Euclidean distances from the boundary vertices after rigid body transformation to the corresponding target geometry boundary surface elements. Upon termination of the iteration, the final geometric boundary parameters and final pose parameters are output. Specifically, the boundary-consistent deformable geometric fusion within the boundary defined by the edge 3D point set is performed as follows: The fusion range is formed by taking the minimum and maximum values ​​of the edge 3D point set in the X, Y, and Z directions and expanding it outward by 5 mm. Only the vertices of the boundary model to be deformed within the current range are updated. The boundary model to be deformed consists of vertices and faces with initial geometric boundary parameters. In each iteration, the nearest point is searched for in the target boundary surface set and the edge 3D point set for each vertex. The two distances are calculated and the larger one is taken as the update distance. The vertex moves along the normal vector direction of the adjacent face, and the step size is the update distance multiplied by 0.5 with a single upper limit of 2 mm. After the vertex update is completed, the rotation matrix and translation vector are solved to minimize the sum of the Euclidean distances from the vertex after rigid body transformation to the corresponding nearest surface point.

[0026] In this embodiment, the output candidate robot trajectory includes: Based on the final pose parameters and final geometric boundary parameters, construct the environmental geometry set and the task target set, and generate the uncertainty region set based on the geometric uncertainty information; A discrete feasible region is constructed and a safety-critical pose sequence is selected. The discrete feasible region consists of a set of states obtained by discrete sampling of the robot joint space. The safety-critical pose sequence is obtained by connecting the discrete states in the discrete feasible region that satisfy collision safety constraints, accessibility constraints and process constraints in the order of the task objective set. Candidate robot trajectories are generated and output within a continuous space. For each time sampling point on the candidate robot trajectory, collision safety constraints, reachability constraints, and process constraints are calculated. Trajectory segment reconstruction is performed on sampling points that do not meet the constraints. Specifically, the calculation of collision safety constraints, reachability constraints, and process constraints, and the reconstructing of trajectory segments for sampling points that do not meet the constraints, are as follows: Sampling is performed on the continuous trajectory with a step size of 10 milliseconds. For each sampling point, positive kinematics is performed to obtain the pose of the link and the end effector. Collision safety constraints are determined by calculating the minimum distance between the link envelope and the geometric set of the environment. If the minimum distance is less than 5 mm, it is considered unsatisfactory. The threshold in the uncertainty area is 10 mm. The accessibility constraints are determined by checking whether the joint angle is within the upper and lower limits, whether the joint angular velocity does not exceed 1.5 radians per second, and whether the joint angular acceleration does not exceed 3 radians per second. If any limit is exceeded, it is considered unsatisfactory. The process constraints are determined by the deviation between the end effector and the target. If the position deviation is greater than 1 mm or the attitude deviation is greater than 1 degree, it is considered unsatisfactory. For sampling points that do not meet the requirements, determine the trajectory segment to which they belong and extend it forward and backward by 0.2 seconds as the reconstruction interval. Refit the joint trajectory of the interval into a cubic polynomial so that the position and velocity at the start and end of the interval are continuous and meet the threshold within the interval. After reconstruction, resample and verify.

[0027] In this embodiment, the output of deployable robot trajectory and corresponding robot program parameters includes: Virtual simulation verification is performed on the candidate robot trajectory. Kinematic and collision simulations are performed at the trajectory time sampling points, and the minimum collision distance sequence, joint limit margin sequence, velocity margin sequence, and process parameter deviation sequence are output. Simulation results are generated based on the minimum collision distance sequence, joint limit margin sequence, velocity margin sequence, and process parameter deviation sequence. When the simulation result is "fail", the candidate robot trajectory is updated repeatedly. When the simulation result is "pass", the deployable robot trajectory and the corresponding robot program parameters are output. The deployable robot trajectory is discretized into a sequence of control instructions.

[0028] Example 1: To verify the feasibility of this invention in practice, it was applied to a virtual debugging cycle of a robot assembly station. The system received a batch of multi-view image data from the station, including 6 views, with 240 frames continuously acquired from each view, totaling 1440 color images and 1440 grayscale images, with a resolution of 1920×1080 pixels; 1440 depth maps were acquired simultaneously, with an effective depth range of 0.35–1.50 meters and a depth hole pixel ratio of 14.8%; the camera pose sequence was obtained by extrinsic parameter calculation, with an angle standard deviation of 0.12 degrees for inter-frame pose jitter; the illumination parameters were recorded as exposure time and gain, with an exposure fluctuation range of ±18%. Due to the reflectivity of the workpiece surface and the low texture of the fixture, the average contrast of the original grayscale image was 0.19, and the pixel ratio of gradient spikes caused by noise was approximately 6.3%. Additionally, occlusion existed, and approximately 21.5% of the edges of the workpiece's outer contour showed breakage in the original image.

[0029] After the data enters the process of this invention, the grayscale image is first normalized, scaling the pixel values ​​to the 0–1 range, and then a rolling guided filter pre-smoothing is performed. The window is 7×7, the spatial control parameter is 2.0 pixels, the guided grayscale difference control parameter is 0.1, and the process is repeated 3 times. After pre-smoothing, the proportion of background noise gradient peak pixels decreases from 6.3% to 2.1%, and the grayscale contrast increases from 0.19 to 0.24. A 3×3 difference convolution kernel is used to calculate the horizontal and vertical gradients, and the boundaries are filled with mirror images. The mean of the gradient magnitude map is 0.037, and the variance is 0.012. In the subpixel spline extreme value scanning stage, the magnitudes at three points (backward, center, and forward) along the gradient direction are taken. A quadratic polynomial is constructed, and the function values ​​at these three points are equal to the magnitudes at the three points. The derivative of the quadratic polynomial is calculated and set to zero to obtain the extreme value displacement, which is limited to ±1 pixel. Statistical analysis shows that the absolute value of the subpixel displacement is 0.31 pixels, and the 95th percentile is 0.86 pixels. When using dual thresholds, the high threshold is the 90th percentile of the gradient magnitude, and the low threshold is the high threshold multiplied by 0.4; hysteresis connections use an eight-neighborhood approach. The final sub-pixel edge point set is obtained, with an average of 46,200 edge points per frame. The edge breakage rate decreases from 21.5% to 6.8%, the repetition response rate decreases from 8.9% to 2.7%, and the edge localization error, using the straight edge of the calibration board as a reference, decreases from 0.62 pixels to 0.18 pixels.

[0030] In the coarse pose localization stage, edge points are back-projected onto the workstation coordinate system and a 3D voxel mesh is created. The voxel size is 5 mm, and the boundary is expanded 10 mm outward from the point cloud extrema. Each voxel contains at least three edge points, occupying approximately 18,100 voxels. After constructing the edge-dominated sparse voxel field, voxel gradient consistency registration is performed with the target geometry template. Rotation candidates are scanned at 5-degree steps ±30 degrees, and translation candidates at 10-mm steps ±100 mm. The closest distance greater than 20 mm is not considered in the cost. The mesh is then refined again at 1-degree and 2-mm steps to obtain the initial pose and initial boundary. At this stage, the workpiece translation error is 3.1 mm with an angular error of 1.0 degree, the fixture translation error is 2.4 mm with an angular error of 0.8 degrees, and the obstacle boundary deviation is 5.3 mm.

[0031] An improved Poisson flow generation model was implemented. For the toroidal shell prior, the target center point was used as the sampling center, with an inner radius of 0.6 meters and an outer radius of 1.0 meter. 1024 initial state samples were sampled. Samples falling inside the initial envelope were discarded and resampled, with a resampling rate of 7.0%. The biphasic radius flip was triggered radially within 0.75–0.85 meters, with 1.0 meter as the flip reference and a radial lower limit truncation of 0.05 meters, resulting in a range of 0.6–1.0 meters. For the local subspace convergence projection, 64 neighborhood sample points were taken, consisting of edge 3D points and depth 3D points. The covariance matrix was calculated, and the eigenvectors were obtained to obtain the local subspace basis vector set. An orthogonal projection was performed on the output points, outputting the point cloud, occupied voxels, and uncertainty scalar field. The completed region occupied 11.0% of the point cloud, with a mean uncertainty of 0.72 in the completed region and 0.29 in the stable region. Compared with traditional deep direct fusion point clouds, the proportion of occluded side hole area decreased from 17.2% to 6.4%, and the proportion of maximum connected surface increased from 0.63 to 0.87.

[0032] In the geometry unification stage, a boundary-consistent deformable geometry fusion is performed, with the fusion range extending 5 mm beyond the bounding box of the edge 3D points. Vertices move along the normal direction with a step size of 0.5 times the update distance and a single maximum of 2 mm. Simultaneously, rigid body transformations are solved to minimize the sum of distances from vertices to the nearest surface point. Termination occurs when the average displacement between two adjacent iterations is less than 0.05 mm or the number of iterations reaches 30. This batch of data averaged 14 iterations. After output, the workpiece translation error decreased from 3.1 mm to 0.9 mm, and the angle error decreased from 1.0 degree to 0.23 degrees; the fixture translation error decreased from 2.4 mm to 0.7 mm, and the angle error decreased from 0.8 degrees to 0.18 degrees; the obstacle boundary deviation decreased from 5.3 mm to 1.6 mm.

[0033] The path optimization phase employs an uncertainty-adaptive dual-domain strategy. Constraints are calculated for continuous trajectories using 10-millisecond sampling. The collision threshold is set at 5 mm for the stable region and 10 mm for the uncertain region; the joint velocity threshold is 1.5 radians per second, and the acceleration threshold is 3 radians per second; the process thresholds are 1 mm for position and 1 degree for attitude. If the constraints are not met, the trajectory segment is reconstructed by extending the interval by 0.2 seconds before and after it, and a cubic polynomial is refitted. On average, each trajectory triggers 1.2 reconstructions. The final minimum collision distance is 14.0 mm, the average process position deviation is 0.63 mm, and the average attitude deviation is 0.55 degrees.

[0034] In comparative experiments, the traditional method employs deep direct fusion modeling, single-domain trajectory generation, and one-time simulation verification; the present invention employs Canney edge detection, an improved Poisson flow generation model, boundary consistency fusion, and uncertainty dual-domain optimization. Statistical analysis of the two methods on the same batch of 600 test data sets showed the following differences: the average workpiece translation error was 2.5 mm for the traditional method and 0.9 mm for the present invention; the average fixture translation error was 2.0 mm for the traditional method and 0.7 mm for the present invention; the average obstacle boundary deviation was 4.8 mm for the traditional method and 1.6 mm for the present invention; the average hole area ratio was 16.5% for the traditional method and 6.4% for the present invention; the average minimum collision distance was 7.5 mm for the traditional method and 14.0 mm for the present invention; the average first-pass rate of virtual simulation was 63.0% for the traditional method and 90.8% for the present invention; the average number of manually adjusted path points before deployment was 12.1 for the traditional method and 1.8 for the present invention; and the average total time from virtual debugging to output program parameters was 126 minutes for the traditional method and 30 minutes for the present invention. The data demonstrate that the present invention exhibits quantifiable process changes and comparative advantages in each key step, proving its engineering feasibility and effectiveness.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for optimizing virtual debugging paths for robots based on computer vision, characterized in that, include: Collect multi-view image data of the workstation and complete camera calibration and coordinate system establishment to obtain the transformation relationship between camera parameters and robot base coordinate system; Canni edge detection is performed on multi-view image data of the workstation. Noise suppression and main contour preservation are achieved by pre-smoothing with rolling guided filtering. Subpixel spline extreme value scanning is used for fine-grained extreme value localization, and subpixel edge point set is output. An edge-dominated sparse voxel field is constructed based on sub-pixel edge point sets, and voxel gradient consistent registration is used for coarse pose localization to obtain initial pose parameters and initial geometric boundary parameters. An improved Poisson flow generation model is constructed. The initial pose parameters and initial geometric boundary parameters are initialized by annular shell spatial sampling based on the annular shell prior module. The radial evolution process of the generated flow is reversed by the two-phase radius reversal module. The neighborhood subspace projection is introduced by the local subspace convergence projection module to obtain the three-dimensional geometric representation and geometric uncertainty information. Based on the 3D geometric representation and the sub-pixel edge point set, perform boundary-consistent deformable geometry fusion, iteratively transform the initial geometric boundary parameters into the final geometric boundary parameters, and obtain the final pose parameters; Based on the final pose parameters and final geometric boundary parameters, perform uncertainty adaptive dual-domain path optimization to output candidate robot trajectories; The candidate robot trajectory is verified by virtual simulation, and the deployable robot trajectory and corresponding robot program parameters are output.

2. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The multi-view image data of the workstation includes multi-view color image sequences, grayscale image sequences, depth map sequences, camera pose sequences, timestamp sequences, and lighting status parameter sequences.

3. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The process of completing camera calibration and coordinate system establishment, and obtaining the transformation relationship between camera parameters and the robot's base coordinate system, includes: Imaging equipment is set up in the workstation area to collect multi-view image data of the workstation, and the data from each viewpoint are synchronized and associated according to the timestamp; Camera calibration is performed on the imaging device to obtain the camera intrinsic parameter matrix and distortion parameters. The camera intrinsic parameter matrix is ​​a three-row, three-column matrix containing focal length parameters and principal point coordinate parameters. The distortion parameters include radial distortion coefficients and tangential distortion coefficients. Distortion correction is performed on multi-view color images, grayscale images, and depth maps based on the camera intrinsic parameter matrix and distortion parameters. Establish the camera coordinate system, the workstation coordinate system, and the robot base coordinate system, and obtain the rigid body transformation relationship from the camera coordinate system to the workstation coordinate system and from the workstation coordinate system to the robot base coordinate system. Based on the matrix multiplication of the two rigid body transformation relationships, obtain the transformation relationship between the camera parameters and the robot base coordinate system.

4. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The output sub-pixel edge point set includes: Rolling guided filtering pre-smoothing is performed on the grayscale image in the multi-view image data of the workstation. The rolling guided filtering pre-smoothing is performed according to the number of iterations. The initial value of the iteration is the original grayscale image. In each iteration, the pre-smoothed grayscale image output by the previous iteration is used as the guide image of the current iteration and the pre-smoothed grayscale image of the current iteration is output to obtain the final pre-smoothed grayscale image. The horizontal and vertical gradient maps are calculated based on the final pre-smoothed grayscale image. The magnitude of each pixel in the gradient magnitude map is the square root of the sum of the squares of the horizontal and vertical gradient values. The direction of each pixel in the gradient direction map is the arctangent of the vertical and horizontal gradient values, thus obtaining the gradient magnitude map and gradient direction map. Subpixel spline extreme value scanning and positioning is performed on the gradient magnitude map. The scanning direction is determined by the gradient direction map. Three gradient magnitudes are selected with the current pixel as the center: the forward sampling point, the center sampling point, and the backward sampling point. Quadratic spline interpolation is used to fit the function of gradient magnitude changing with position, and the position where the first derivative of the function is zero is obtained to get the subpixel extreme value position. The subpixel extreme value position is used as the edge point position. The set of strong edge points and the set of weak edge points are obtained based on the dual threshold determination. The dual threshold determination uses a high threshold and a low threshold. The high threshold is the amplitude corresponding to the preset percentile of the gradient amplitude map, and the low threshold is the high threshold multiplied by a scaling factor. Perform a hysteresis connection on the set of strong edge points, search for edge points connected to weak edge points in the eight neighborhoods of the strong edge points and add the connected weak edge points to the edge point set, delete the weak edge points that are not connected to any strong edge point, and obtain the sub-pixel edge point set.

5. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The process of obtaining the initial pose parameters and initial geometric boundary parameters includes: The sub-pixel edge point set is divided into a three-dimensional voxel grid according to the spatial range of the work station coordinate system. Each voxel corresponds to a spatial cubic region. Sub-pixel edge points falling into the same voxel are aggregated into the edge occupancy marker of the current voxel. An edge-dominated sparse voxel field is constructed in a three-dimensional voxel mesh. The edge-dominated sparse voxel field consists of all voxels whose edges are occupied and marked as true. Each occupied voxel records the coordinates of the voxel center point and the average gradient direction vector of the edge points within the voxel. Voxel gradient consistency registration is performed between the edge-dominated sparse voxel field and the preset target geometric template. The target geometric template is represented as a set of template voxels in the work station coordinate system, and each template voxel contains the coordinates of the template voxel center point and the template gradient direction vector. By searching for rigid body transformations in the work station coordinate system to minimize the registration cost, coarse pose localization results are obtained. The coarse localization result is output as the initial pose parameters. Under the initial pose parameters, the outer surface voxels of the target geometric template are projected onto the station coordinate system to obtain the initial geometric boundary parameters.

6. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The obtained three-dimensional geometric representation and geometric uncertainty information include: An improved Poisson flow generation model is constructed. The improved Poisson flow generation model sets a ring-shell prior module at the initial sampling position of the original Poisson flow generation model, sets a two-phase radius flipping module at the middle position of the radial evolution process, and sets a local subspace convergence projection module at the output position of the final generation stage. The toroidal shell prior module is executed to perform initial state sampling within the toroidal shell space enclosed by the inner and outer spherical surfaces at the same center point. The radii of the inner and outer spherical surfaces are preset real numbers, and the radius of the outer spherical surface is greater than the radius of the inner spherical surface. The initial state sampling yields the initial state vector, which is then input into the Poisson flow generation process. The two-phase radius flipping module is executed. When the intermediate state sequence output by the Poisson flow generation process satisfies the radial interval, the radial component is flipped and mapped. The flipping and mapping transforms the radial component from the first radial value to the second radial value. The flipped and mapped radial component and the corresponding spatial variable component form the flipped intermediate state and continue to be input into the Poisson flow generation process until the final generation result is obtained. The local subspace convergence projection module is executed. Neighborhood sample points are retrieved in the spatial neighborhood corresponding to the generated final result. The covariance matrix of the neighborhood sample points is calculated and the eigenvectors of the covariance matrix are obtained to obtain the local subspace basis vector set. The spatial variable components of the generated final result are orthogonally projected according to the local subspace basis vector set to obtain the projected spatial variable components. The projected spatial variable components are mapped to the radial components of the generated final result to form a three-dimensional geometric representation. The improved Poisson flow generation model is trained, training sample pairs are constructed, and the improved Poisson flow generation model is executed on each training sample pair to obtain the predicted 3D geometric representation. The loss function between the predicted 3D geometric representation and the real 3D geometric label is calculated and the model parameters are updated. The loss function includes the mean Euclidean distance term of the corresponding points in the point cloud and the mean absolute error term of the sampling points in the distance field, until the training termination condition is met.

7. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The step of iteratively transforming the initial geometric boundary parameters into the final geometric boundary parameters and obtaining the final pose parameters includes: The three-dimensional geometric representation is transformed into a set of boundary surfaces of the target geometry in the work station coordinate system. The set of boundary surfaces of the target geometry is a set of surface points obtained by performing neighborhood connectivity clustering on the point cloud set. Normal vectors are calculated for the surface elements in the set of boundary surfaces of the target geometry. The sub-pixel edge point set is back-projected onto the work station coordinate system according to the camera pose sequence to obtain the edge three-dimensional point set. Each point in the edge three-dimensional point set is determined by the intersection of the light rays emitted from the camera center and the target geometry boundary surface set. The intersection points are obtained by solving the intersection equation of the camera ray parameter equation and the boundary surface element. Within the boundary range defined by the edge 3D point set, perform boundary-consistent deformable geometry fusion, establish the deformable boundary model determined by the initial geometric boundary parameters and perform iterative updates, and update the rotation matrix and translation vector by minimizing the sum of the Euclidean distances from the boundary vertices after rigid body transformation to the corresponding target geometry boundary surface elements. When the iteration terminates, output the final geometric boundary parameters and the final pose parameters.

8. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The output candidate robot trajectory includes: Based on the final pose parameters and final geometric boundary parameters, construct the environmental geometry set and the task target set, and generate the uncertainty region set based on the geometric uncertainty information; A discrete feasible region is constructed and a safety-critical pose sequence is selected. The discrete feasible region consists of a set of states obtained by discrete sampling of the robot joint space. The safety-critical pose sequence is obtained by connecting the discrete states in the discrete feasible region that satisfy collision safety constraints, accessibility constraints and process constraints in the order of the task objective set. Candidate robot trajectories are generated and output in a continuous space. Collision safety constraints, accessibility constraints, and process constraints are calculated for each time sampling point on the candidate robot trajectory, and trajectory segment reconstruction is performed for sampling points that do not meet the constraints.

9. The robot virtual debugging path optimization method based on computer vision according to claim 1, characterized in that, The output can deploy robot trajectories and corresponding robot program parameters, including: Virtual simulation verification is performed on the candidate robot trajectory. Kinematic and collision simulations are performed at the trajectory time sampling points, and the minimum collision distance sequence, joint limit margin sequence, velocity margin sequence, and process parameter deviation sequence are output. Simulation results are generated based on the minimum collision distance sequence, joint limit margin sequence, velocity margin sequence, and process parameter deviation sequence. When the simulation result is "fail", the candidate robot trajectory is updated repeatedly. When the simulation result is "pass", the deployable robot trajectory and the corresponding robot program parameters are output. The deployable robot trajectory is discretized into a sequence of control instructions.