Automatic docking method and system based on monocular vision
By constructing a combination of a concentric double-ring structure and a near-infrared light source filter, the problem of low accuracy in oil gun docking caused by imaging reflection interference was solved, achieving high-precision oil nozzle pose estimation, meeting the technical requirements of automatic vehicle refueling, reducing manpower and material resources and improving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-08
AI Technical Summary
In existing automated refueling solutions, the accuracy of the fuel nozzle docking is not high due to interference from imaging reflections, and a large amount of manpower and resources are required, posing a safety threat.
An automatic docking method based on monocular vision is adopted. By constructing a double-ring structure with concentric opposite surfaces, combining a near-infrared light source and a narrow-band filter to suppress reflective interference, and using arc segment adaptive segmentation and multi-constraint aggregation strategy, the pose is directly calculated to achieve high-precision docking.
High-precision refueling nozzle pose estimation was achieved in complex reflective environments, meeting the technical requirements of automatic vehicle refueling, reducing manpower and material resources, and improving safety.
Smart Images

Figure CN121998949A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision, and more particularly to an automatic docking method and system based on monocular vision. Background Technology
[0002] Currently, gas stations primarily rely on hired workers to manually refuel vehicles. However, factors such as the time required for worker training and their grasp of theoretical knowledge introduce uncertainty into the workers' operations, potentially posing safety threats. Furthermore, manual refueling requires significant manpower and resources.
[0003] With the rapid development of machine vision, the application prospects of automated refueling at gas stations are very broad. In automated refueling systems for automobiles, the precise automatic docking of the fuel filler neck and the connector is a core component, and accurate perception of the spatial pose of the fuel filler neck is a prerequisite for reliable docking. Currently, most automotive fuel filler necks have a circular, symmetrical structure and are made of high-reflectivity materials. In actual refueling scenarios, the reflection from the fuel filler neck surface overlaps with the reflection from the equipment's optical components, leading to broken image edges and increased stray edges, severely affecting the accuracy of feature extraction.
[0004] Therefore, developing a monocular vision pose estimation method that does not rely on manual control points, has strong anti-glare interference capabilities, and high pose estimation accuracy is of great significance for promoting the engineering application of automatic refueling technology for automobiles. Summary of the Invention
[0005] In view of this, in order to solve the technical problem of low nozzle docking accuracy caused by imaging reflection interference in existing automatic refueling solutions, the present invention proposes an automatic docking method based on monocular vision, which includes the following steps: Geometric Model Establishment: First, a concentric double-ring structure is constructed as the spatial geometric representation of the fuel filler nozzle. Image Preprocessing: Images of the fuel filler nozzle are acquired and filtered and enhanced based on brightness characteristics to obtain clear feature images. Edge and Arc Extraction: Edge chains are extracted from the preprocessed images, and arc segmentation is performed to form a set of candidate arc segments. Arc Reassembly: Candidate arc segments are screened and aggregated through multiple constraints to generate a structured set of arc segments. Ellipse Fitting and Parameter Calculation: Projected ellipses are fitted based on the reassembled arc segments, and their geometric parameters are calculated. Initial Pose Estimation: The possible poses of the spatial circle are inferred from the ellipse parameters, and the optimal initial solution is selected by combining geometric structure error evaluation. Pose Optimization: An optimization objective function is constructed, and iterative optimization is performed starting from the initial solution to determine the precise spatial pose of the fuel filler nozzle. Docking Execution: The final pose guides the collaborative robotic arm to adjust the docking joint position and complete the precise docking operation.
[0006] In addition to the above method, the present invention also proposes an automatic docking system based on monocular vision, which includes a modeling unit, an image acquisition unit, a pose calculation unit, and a control unit.
[0007] Based on the above scheme, this invention provides an automatic docking method and system based on monocular vision. It proposes a non-surface concentric double-ring oil port model, which eliminates pose ambiguity by utilizing natural geometric constraints, eliminates the need for manual control points, and can be adapted to different types of oil ports. Furthermore, it suppresses ambient light interference by using a near-infrared light source and a narrow-band filter, and solves the edge breakage problem caused by reflection by combining arc segment adaptive segmentation and multi-constraint aggregation strategies. In addition, it designs a control point-free direct solution model, which solves the pose based on the correspondence between the two-dimensional elliptical projection of the double ring and the three-dimensional geometric structure, and has both high accuracy and strong physical interpretability. Attached Figure Description
[0008] Figure 1 This is a flowchart of the steps of an automatic docking method based on monocular vision according to the present invention; Figure 2 This is a line graph comparing the position error of the method of this invention with that of the unoptimized circular ring method; Figure 3 Based on Figure 2 Box plot of location error in data statistics; Figure 4 This is a line graph comparing the angle error of the method of this invention with that of the unoptimized double-ring method; Figure 5 Based on Figure 4 Angle error box plot of data statistics; Figure 6 This is a line graph comparing the relative position error between the method of this invention and the unoptimized circular ring method; Figure 7 Based on Figure 6 The relative error box plot of the data statistics; Figure 8 This is a comparison chart of the target measurement position distribution results of the method of this invention and the unoptimized circular ring method. Detailed Implementation
[0009] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0010] It should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0011] It should be understood that the terms "system," "apparatus," "unit," and / or "module" used in this application are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0012] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "a," and / or "the" are not specifically singular and may include the plural. Generally, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. An element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.
[0013] In the description of the embodiments of this application, "a plurality of" refers to two or more. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0014] Furthermore, flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, the steps can be processed in reverse order or simultaneously. Additionally, other operations can be added to these processes, or one or more steps can be removed from them.
[0015] Reference Figure 1 This is a flowchart illustrating an optional example of the automatic docking method based on monocular vision proposed in this invention. This method can be applied to computer devices, and the docking method proposed in this embodiment may include, but is not limited to, the following steps: Step S1: Model the filler port and construct a model of a non-circular concentric double-ring filler port. Step S2: Obtain the image of the fuel filler neck and filter it based on brightness information; Step S3: Extract edge chains and segment arcs from the image filtered in S2; Step S4: Reorganize the candidate arc segment set after S3 segmentation using a multi-constraint arc segment aggregation strategy; Step S5: Based on the candidate arc segments recombined in S4, filter candidate ellipses and calculate the parameters of the projected ellipse; Step S6: Solve the initial pose of the spatial circle based on the projection ellipse parameters of S5, and select the optimal initial solution through the geometric structure error function; Step S7: Construct the objective function and perform pose optimization using the optimal initial solution from S6; Step S8: Based on the optimized oil port position and orientation from S7, guide the collaborating arm to drive the connector to complete the docking.
[0016] In some feasible embodiments, step S1 specifically includes: The fuel filler neck is modeled as two coaxial parallel rings, with the inner radius of the upper smaller ring being... outer radius The inner radius of the large lower ring outer radius ,satisfy The distance between the two rings is With the center of the small ring as the pose coordinate point of the oil port, and the direction of the normal vector perpendicular to the plane of the small ring downwards; In this embodiment, the filler port is abstracted as a non-plane concentric double ring structure with fixed geometric constraints. The collinearity of the centers of the two rings, the consistency of their normal vectors, and the fixed axial distance provide structural priors for eliminating pose ambiguity.
[0017] In some feasible embodiments, step S2 specifically includes: A preprocessing scheme combining hardware and software is adopted to suppress reflective interference. Near-infrared LED light source configured in a monocular vision module and narrow-band filter are used to filter visible light interference. Combined with Sobel gradient filtering and brightness threshold screening, low-gradient light clusters formed by reflections from the oil port surface and optical components are effectively suppressed, resulting in a more prominent reflective ring image.
[0018] In some feasible embodiments, step S3 specifically includes: The Edge Drawing algorithm is used to extract the initial edge chain, which alleviates the threshold sensitivity problem of the traditional Canny algorithm to reflective edges. The RDP algorithm is used to adaptively sample the dominator point, and the edge chain is accurately segmented based on the vector angle and inflection point determination, thus eliminating stray short arc segments.
[0019] In some feasible embodiments, step S4 specifically includes: Coarse screening is performed using distance constraints and directional consistency constraints. Then, accurate aggregation of arc segments from the same origin is achieved through local angle constraints, global directional consistency verification, and a secondary correction mechanism. Finally, the arc segment combination is expanded using a disjoint-set data structure algorithm to restore the complete elliptical contour.
[0020] The distance constraint is determined as follows: the Euclidean distance between the midpoints of two arc segments is no greater than 3 times the sum of the lengths of the two arc segments; the direction consistency constraint determines the bidirectional direction consistency by calculating the cross product of the key vectors of the arc segments; the local angle constraint requires that the angle between adjacent vectors does not exceed 90°, and the length ratio of the connecting vector to the local direction vector is within the range of [0.5, 2]; the secondary correction mechanism is triggered when the endpoint spacing is < 5 pixels or > 20 pixels, recalculating the connecting endpoint vectors and performing angle and direction verification.
[0021] In some feasible embodiments, step S5 specifically includes: By combining geometric features and gradient confidence scores, a two-layer clustering strategy is used to separate the inner and outer ring ellipses, remove redundant interference ellipses, and obtain the parameters of the four projected ellipses corresponding to the double rings.
[0022] The gradient confidence score is obtained by calculating the cosine similarity between the local gradient direction of the ellipse contour sampling point and the tangent direction of the ellipse. The two-layer clustering strategy includes: coarse clustering is grouped by constraining the difference between the center distance and the axis length, and fine clustering selects the optimal ellipse pair based on the principle of the highest center concentricity and the smallest axis ratio difference.
[0023] In some feasible embodiments, steps S6 and S7 specifically include: The initial pose of the spatial circle is solved based on the characteristics of elliptical projection. The optimal initial solution is selected by the geometric structure error function. The objective function is constructed by the multiple geometric constraints of the double ring. The high-precision oil port pose is obtained by LM optimization.
[0024] The formula for the geometric structure error function is as follows: In the above formula, This represents the coordinates of the center of the inner edge of the large annulus. This represents the coordinates of the center of the outer edge of the large annulus. This represents the coordinates of the center of the inner edge of the small annulus. This represents the coordinates of the center of the outer edge of the small ring. , , , These are the corresponding normal vectors.
[0025] The objective function is expressed as follows: in, This represents the consistency constraint term for the centers of concentric circles. This represents the normal vector consistency constraint term. This represents the normalization constraint term for the normal vector. This indicates the collinearity constraint term. This indicates the axial distance constraint term. This represents the corresponding weighting coefficient. It represents the distance between the center points of two coaxial rings in space.
[0026] In some feasible embodiments, step S8 specifically includes: converting the pose in the camera coordinate system to the base coordinate system of the collaborating arm through a preset coordinate transformation relationship, and guiding the collaborating arm to drive the connector to complete the precise docking with the oil port.
[0027] Based on the overall process of the above method, this invention also provides relevant data examples: Set the inner radius of the small ring outer radius The inner radius of the large ring outer radius The distance between the two rings With the center of the small ring With the origin as the point, Establish the oil port coordinate system downwards along the normal direction.
[0028] The monocular vision module turns on the near-infrared LED light source (wavelength 850nm), the camera image acquisition frame rate is set to 30fps, the exposure time is adjusted to 10ms, the image is acquired through an 850nm narrowband filter (bandwidth 10nm), the Sobel gradient matrix is calculated, the gradient threshold is set to 80, and low gradient reflective areas are filtered out.
[0029] The Edge Drawing algorithm is used to extract the initial edge chain, with the edge seed point threshold set to 50 and the direction tracking step size set to 1 pixel. The RDP algorithm is used for sampling, with a sampling interval of 1 pixel for regions of abrupt curvature change and a sampling interval of 3 pixels for regions with smooth grayscale, and an angle threshold is set. The length threshold is 10 pixels, and short, stray arc segments are removed. The distance constraint is set to the Euclidean distance between the midpoints of the arc segments ≤ 3 times the sum of the arc lengths; the direction consistency constraint is determined by the cross product of the vectors, and if the signs of the cross product results are consistent, the directions are considered consistent; the local angle constraint is set with an included angle threshold of 90°, the ratio of the length of the connecting vector to the length of the local direction vector is in the range of [0.5, 2], the secondary correction trigger condition is the endpoint spacing < 5 pixels or > 20 pixels, and the included angle threshold of the new connecting vector is 15°; In the ellipse selection and optimization stage, the initial screening and gradient feature extraction of candidate ellipses are first completed: ellipses with multi-arc aggregation and satisfactory fit are collected, and single-arc fitting ellipses with a length ≥ 4 pixels are added to form a candidate set. The image gradient is calculated using the Sobel operator, and the mean cosine similarity between the gradient direction and the tangent direction of the ellipse contour sampling points is used as the gradient confidence score. Subsequently, coarse clustering of ellipses is carried out, and ellipses with a center Euclidean distance < 20 pixels, a minor axis relative deviation < 6%, and a major axis relative deviation < 5% are grouped into one class, retaining only effective clusters containing at least 2 ellipses; if there are fewer than 2 effective clusters, the top 4 ellipses are selected in descending order of gradient confidence, otherwise, the optimal ellipse pair is selected for each effective cluster. During the screening process, the ellipse boundary type and axis ratio parameter are marked first. The combination of "outer boundary + inner boundary" is selected first. The ellipse pair corresponding to the minimum value is selected by calculating the comprehensive difference value (center distance + axis ratio difference × 1000). If there is no such combination, the ellipse pair with the axis ratio parameter closest to the outer boundary feature value (0) and the inner boundary feature value (1) is selected. Finally, the optimal ellipse pair of the two effective clusters is extracted to obtain the four ellipses corresponding to the inner and outer boundary projections of the double ring of the fuel filler, which provide high-quality feature input for subsequent pose calculation. Dual-constraint pose calculation: The SAF AEE-RAD method is used to solve the initial pose of the spatial circle, traversing all combinations of the four elliptical solutions (a total of 24 combinations), and selecting the optimal initial solution through the geometric structure error function; constraint weights are set. , , , , The final pose is obtained by iterating through the LM optimization algorithm until convergence. Automatic docking guidance: Based on the preset camera intrinsic parameter calibration, hand-eye calibration and tool coordinate system calibration results, a coordinate transformation relationship is established to transform the oil port pose in the camera coordinate system to the base coordinate system of the collaborating arm. The collaborating arm adopts a speed control mode with an approximation speed of 5mm / s. During the docking process, the docking action is controlled by the force feedback threshold to complete the precise docking.
[0030] An experimental platform was built to simulate an automated refueling scenario. The monocular vision module was equipped with explosion-proof adapters to match actual working conditions. The camera was positioned with its vertex 30cm above the fuel filler neck, and 300 observation positions were selected within an inverted cone-shaped area 40cm high and 25cm in diameter at the base, ranging from near to far, covering different positions and orientation changes. Using the ArUco code marker method as the ground truth reference, comparative experiments were conducted with the EPnP method and the unoptimized double-ring method. The results are as follows: Figure 2To reflect changes in position from near to far during observation, the position calculated using the ArUco code is taken as the true value. The absolute position errors of the proposed method are compared with those of the EPnP pose estimation method, the unoptimized position, and the optimized position using the double-ring geometry. The horizontal axis represents the sampling position, and the vertical axis represents the absolute error (mm). The average absolute position error of the optimized method is 4.59 mm, a 22.0% reduction compared to the EPnP method (5.88 mm) and a 42.3% reduction compared to the unoptimized double-ring method (7.96 mm).
[0031] Figure 3 For statistics Figure 2 The box plot of absolute positional error of the data shows that the absolute positional accuracy distribution after the double-ring geometric optimization is the best, with a better error concentration trend, fewer outliers, and a significant improvement compared to before optimization.
[0032] Figure 4 To reflect changes in attitude as observation progresses from near to far, the attitude calculated using ArUco codes is used as the true value. The proposed method compares the attitude calculated using the method described in this paper with that calculated using the EPnP method. The attitude error is represented by the angle between the calculated normal vector and the true normal vector calculated using ArUco codes. The horizontal axis represents the number of iterations (from 1 to 300), and the vertical axis represents the angle error (°). The average angle error of the proposed method is 1.37°, the average angle error of the EPnP method is 1.89°, and the average angle error of the unoptimized double-ring method is 2.15°.
[0033] Figure 5 For statistics Figure 4 The angle error box plot of the data shows that the method presented in this paper has fewer outliers and lower overall error compared to the EPnP method. Figure 6 To reflect changes in position from near to far, the position calculated using the ArUco code is taken as the true value. The relative position errors of the proposed method are compared with those of the EPnP pose estimation method, the unoptimized position, and the position after optimization using the double-ring geometry. The horizontal axis represents the sampling position, and the vertical axis represents the relative error (%). The mean relative position error is 0.98%, a 28.0% reduction compared to the EPnP method (1.36%), and a 42.7% reduction compared to the unoptimized double-ring method (1.71%).
[0034] Figure 7 For statistics Figure 6 The relative error box plot of the data shows that, with the change of observation distance, the double-ring geometric optimization method used in this paper has the lowest overall error and effectively reduces outliers compared to the EPnP method. Compared to the estimated relative position before optimization, although the error distribution is larger, the error is smaller.
[0035] Figure 8The estimated positions in the camera coordinate system are represented among 300 samples. It can be seen intuitively that the solution positions and true values are most concentrated after the double-ring geometric constraints, which is a significant improvement compared to before optimization.
[0036] In summary, the method of the present invention outperforms the prior art.
[0037] Position accuracy: The average absolute position error of the method of this invention is 4.59 mm, which is 22.0% lower than that of the EPnP method (5.88 mm) and 42.3% lower than that of the unoptimized double-ring method (7.96 mm); the average relative position error is 0.98%, which is 28.0% lower than that of the EPnP method (1.36%) and 42.7% lower than that of the unoptimized double-ring method (1.71%). Angle accuracy: The average angle error of the method of this invention is 1.37°, the average angle error of the EPnP method is 1.89°, and the average angle error of the unoptimized double ring method is 2.15°; Experimental results show that the method of the present invention can still maintain high pose estimation accuracy in complex reflective environments, fully meeting the technical requirements for docking of automatic refueling ports for automobiles.
[0038] An automated docking system based on monocular vision includes: A modeling unit is used to perform step S1; The image acquisition unit is used to perform step S2; The pose calculation unit is used to execute steps S3-S7; The control unit is used to execute step S8.
[0039] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0040] A storage medium storing processor-executable instructions, which, when executed by a processor, are used to implement an automatic docking method based on monocular vision as described above.
[0041] The content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0042] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. An automatic docking method based on monocular vision, characterized in that, Includes the following steps: Model the filler port and construct a non-surface concentric double-ring filler port model; The image of the fuel filler neck is acquired and filtered based on brightness information to obtain a preprocessed image; Edge chain extraction and arc segmentation are performed on the preprocessed image to obtain a set of candidate arc segments; The candidate arc segment set is reorganized using a multi-constraint arc segment aggregation strategy to obtain a reorganized arc segment set. Based on the recombined set of arc segments, candidate ellipses are selected, and the parameters of the projected ellipse are calculated. The initial pose of the spatial circle is solved based on the projection ellipse parameters, and the optimal initial solution is selected through the geometric structure error function. Construct an objective function and combine it with the optimal initial solution to optimize the pose, thereby obtaining the final pose of the oil port; Based on the final position of the oil port, the guiding arm drives the docking joint to complete the docking.
2. The automatic docking method based on monocular vision according to claim 1, characterized in that, The step of acquiring the fuel filler neck image and filtering it based on brightness information to obtain a preprocessed image specifically includes: A filter is introduced, and an image of the fuel filler neck is acquired; The Sobel gradient matrix of the fuel filler image is calculated and filtered in conjunction with brightness information to obtain the preprocessed image.
3. The automatic docking method based on monocular vision according to claim 1, characterized in that, The step of extracting edge chains and segmenting arcs from the preprocessed image to obtain a set of candidate arc segments specifically includes: An edge chain set is obtained by extracting edge chains from the preprocessed image using an edge detection algorithm. By sampling the dominant points and calculating the vector angle and determining the inflection point, the initial edge chain set is segmented into arc segments to obtain a candidate arc segment set.
4. The automatic docking method based on monocular vision according to claim 1, characterized in that, The multi-constraint arc segment aggregation strategy specifically includes coarse screening and fine geometric constraints, wherein: The coarse screening includes distance constraints and direction consistency constraints; The determination condition for the distance constraint is: the Euclidean distance between the midpoints of the two arc segments is not greater than 3 times the sum of the lengths of the two arc segments; The fine geometric screening includes local angle constraints, global direction consistency verification, and a secondary correction mechanism. The global directional consistency constraint is determined by the cross product of the arc segment key vectors.
5. The automatic docking method based on monocular vision according to claim 1, characterized in that, The step of filtering candidate ellipses based on the recombined arc segment set and calculating the projected ellipse parameters specifically includes: In the recombined set of arc segments, ellipses are selected and optimized based on geometric features and gradient information to obtain candidate ellipses; The gradient confidence scores of the candidate ellipses are calculated, and the inner and outer ring ellipses are separated by a two-layer clustering strategy to obtain the parameters of the four projected ellipses corresponding to the double rings.
6. The automatic docking method based on monocular vision according to claim 5, characterized in that, The gradient confidence score is obtained by calculating the cosine similarity between the local gradient direction of the elliptical contour sampling point and the tangent direction of the ellipse.
7. The automatic docking method based on monocular vision according to claim 5, characterized in that, The two-layer clustering strategy includes: Coarse clustering is constrained by the difference between center distance and axis length; Fine clustering selects the optimal ellipse pair based on the principle of the highest concentricity and the smallest difference in axial ratio.
8. The automatic docking method based on monocular vision according to claim 1, characterized in that, The formula for the geometric structure error function is as follows: in, This represents the coordinates of the center of the inner edge of the large annulus. This represents the coordinates of the center of the outer edge of the large annulus. This represents the coordinates of the center of the inner edge of the small annulus. This represents the coordinates of the center of the outer edge of the small ring. , , , These are the corresponding normal vectors.
9. The automatic docking method based on monocular vision according to claim 8, characterized in that, The objective function is expressed as follows: in, This represents the consistency constraint term for the centers of concentric circles. This represents the normal vector consistency constraint term. This represents the normalization constraint term for the normal vector. This indicates the collinearity constraint term. This indicates the axial distance constraint term. This represents the corresponding weighting coefficient. It represents the distance between the center points of two coaxial rings in space.
10. An automatic docking system based on monocular vision, characterized in that, include: The modeling unit is used to model the filler port and construct a non-surface concentric double-ring filler port model; The image acquisition unit is used to acquire the image of the fuel filler neck and filter it based on brightness information to obtain a preprocessed image; The pose calculation unit is used to extract edge chains and segment arcs in the preprocessed image to obtain a set of candidate arcs; and to reorganize the set of candidate arcs using a multi-constraint arc aggregation strategy to obtain a reorganized set of arcs. Candidate ellipses are selected based on the recombined arc segment set, and the parameters of the projected ellipse are calculated. The initial pose of the spatial circle is solved based on the parameters of the projected ellipse, and the optimal initial solution is selected through the geometric structure error function. Construct an objective function and combine it with the optimal initial solution to optimize the pose, thereby obtaining the final pose of the oil port; The control unit guides the cooperating arm to drive the connector to complete the docking based on the final position of the oil port.