Unmanned aerial vehicle vision autonomous docking and locking control system based on deep learning

The UAV vision-based autonomous docking system, which utilizes deep learning and an improved VMamba model, combined with light field enhancement and texture phase code matching, solves the problems of accurate perception and path correction of locking structures in UAV autonomous docking. It achieves high-precision and reliable docking and locking control, making it suitable for UAV missions in complex environments.

CN121545086BActive Publication Date: 2026-04-28HUNAN ZHONGDIAN JINJUN TECH GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN ZHONGDIAN JINJUN TECH GRP CO LTD
Filing Date
2026-01-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing autonomous docking systems for unmanned aerial vehicles (UAVs) lack the ability to accurately perceive the microscopic features of the locking structure in scenarios with limited space, complex lighting, or extremely small docking tolerances. Their path planning methods also lack a dynamic path correction mechanism, and their locking status judgment is not accurate enough, resulting in large docking errors and easy jamming or misalignment.

Method used

A deep learning-based UAV vision-based autonomous docking and locking control system is adopted, which integrates an improved VMamba model and multimodal texture localization technology. Through image acquisition, structural edge extraction, path correction, locking slot localization, insertion path construction, and locking state detection modules, the system realizes autonomous docking and locking control of the UAV. Edge perception and path correction are performed using light field enhancement maps, improved VMamba models, and Lucas-Kanade optical flow algorithms. The locking state is determined by combining texture phase code matching and artifact motion trajectory maps.

Benefits of technology

It improves the accuracy and stability of UAV autonomous docking, enhances the ability to identify locking structure features, ensures docking success rate and connection reliability, and is suitable for precise docking tasks in a variety of complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545086B_ABST
    Figure CN121545086B_ABST
Patent Text Reader

Abstract

The application discloses a deep learning-based unmanned aerial vehicle vision autonomous docking and locking control system, comprising the following modules: an image acquisition and light field enhancement module, which is used for constructing a light field interference enhancement graph and generating an interference saliency mask graph; a structure edge extraction module, which is used for outputting a structure dynamic edge response graph based on an improved VMamba model; a path correction module, which is used for extracting a set of conical profile tracking points, calculating time sequence displacement vectors by adopting a Lucas-Kanade optical flow algorithm, and generating a landing path correction amount; a locking groove positioning module, which is used for determining a locking groove center position; an insertion path construction module, which is used for outputting an insertion path selection result; a locking state detection module, which is used for determining a locking state; and a docking judgment module, which is used for outputting a docking judgment result. The application fuses the improved VMamba model and multi-modal texture positioning, and realizes unmanned aerial vehicle autonomous docking and locking control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous control and computer vision recognition technology for unmanned aerial vehicles (UAVs), and particularly to a deep learning-based UAV vision-based autonomous docking and locking control system. Background Technology

[0002] With the increasing demand for autonomous operation of intelligent unmanned systems in complex environments, visual docking and mechanical locking technologies for UAVs, oriented towards vertical takeoff and landing, multi-task collaboration, and precise control, have become a research hotspot. Existing autonomous docking systems mostly rely on macroscopic positioning based on GPS, inertial navigation, or lidar, combined with structured markers, for alignment. However, in scenarios with limited space, complex lighting, or extremely small docking tolerances, the following problems are commonly encountered:

[0003] During the docking process, UAVs lack the ability to accurately perceive the microscopic features of the locking structure. Traditional image recognition methods based on template matching or geometric reconstruction are highly dependent on target texture and are easily affected by image blurring, interference noise, and changes in viewing angle, resulting in low accuracy and poor robustness in structure recognition. Most path planning methods are based on preset alignment trajectories and lack a dynamic insertion path correction mechanism that combines the relative pose of the actual locking components, resulting in large insertion errors and easy jamming or misalignment during docking. In addition, existing locking status judgments mostly rely on motor current changes or structural travel feedback, which cannot detect minute sliding misalignments or non-standard contact situations. They also lack a dynamic evaluation mechanism based on image texture motion trajectory, making it impossible to achieve closed-loop verification and accurate judgment of the docking process.

[0004] Therefore, how to provide a deep learning-based UAV vision-based autonomous docking and locking control system is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a deep learning-based visual autonomous docking and locking control system for unmanned aerial vehicles (UAVs). This invention integrates an improved VMamba model and multimodal texture localization technology to establish a full-process visual guidance system, from structural dynamic edge perception, cone path correction, texture feature phase encoding, insertion path construction to locking state judgment. This enables autonomous docking and locking control of UAVs, and has the advantages of high edge extraction accuracy, strong path planning reliability, and high docking and locking success rate. It is suitable for precise docking and stable connection tasks in various complex environments.

[0006] The deep learning-based UAV visual autonomous docking and locking control system according to an embodiment of the present invention includes the following modules:

[0007] The image acquisition and light field enhancement module is used to continuously acquire multiple frames of images from the UAV's downward-looking direction using an industrial camera, construct a light field interferometric enhancement map, and generate an interferometric saliency mask map.

[0008] The structural edge extraction module is used to input the light field interference enhancement map and the interference saliency mask map into the improved VMamba model. The improved VMamba model introduces an interference-guided context folding mechanism and outputs a structural dynamic edge response map.

[0009] The path correction module is used to extract the set of cone contour tracking points based on the dynamic edge response map of the structure, and to calculate the time series displacement vector of the cone contour tracking points using the Lucas-Kanade optical flow algorithm to generate the landing path correction amount.

[0010] The locking groove positioning module is used to adjust the attitude of the UAV based on the landing path correction amount, perform texture block and frequency domain analysis on the image of the locking groove area, generate texture phase code and match it with the standard texture phase code library to determine the center position of the locking groove.

[0011] The insertion path construction module is used to construct the insertion path between the locking toe and the locking groove based on the center position of the locking groove and the current attitude of the UAV, calculate the insertion angle and the insertion offset vector, and output the insertion path selection result.

[0012] The locking state detection module is used to control the rotation of the lead screw based on the insertion path selection result, acquire the locking structure image sequence, extract texture artifact information, construct the image artifact motion trajectory map, and determine the locking state.

[0013] The docking determination module is used to output the docking determination result based on the insertion path selection result and the locking status.

[0014] The deep learning-based UAV visual autonomous docking and locking control method according to an embodiment of the present invention includes the following steps:

[0015] Step 1: Continuously acquire multiple frames of images from the UAV's downward-looking direction using an industrial camera, construct an optical field interferometry enhancement map, and generate an interferometry saliency mask map based on the optical field interferometry enhancement map;

[0016] Step 2: Input the light field interference enhancement map and the interference saliency mask map into the improved VMamba model. The improved VMamba model introduces an interference-guided context folding mechanism to obtain the dynamic edge response map of the structure.

[0017] Step 3: Based on the dynamic edge response map of the structure, determine the set of cone contour tracking points, and use the Lucas-Kanade optical flow algorithm to calculate the time series displacement vector of the cone contour tracking points to generate the landing path correction amount;

[0018] Step 4: Adjust the UAV attitude based on the landing path correction amount, perform texture block processing on the image of the locking groove area, generate texture phase code and match it with the standard texture phase code library to determine the center position of the locking groove.

[0019] Step 5: Based on the center position of the locking groove and the current attitude of the UAV, construct the insertion path between the locking toe and the locking groove, obtain the insertion angle and insertion offset vector through the geometric calculation method of the insertion path, and output the insertion path selection result;

[0020] Step 6: Based on the insertion path selection result, control the lead screw rotation, acquire the locking structure image sequence, extract texture artifact information, construct the image artifact motion trajectory map, and determine the locking state;

[0021] Step 7: Based on the insertion path selection result and locking status, output the docking determination result.

[0022] Optionally, the construction of the optical field interference enhancement map specifically involves:

[0023] The image pairs are formed by combining any two frames from multiple frames continuously acquired by an industrial camera. Pixel registration is performed on each pair of images, and the brightness difference, color channel difference, and Sobel gradient difference between corresponding pixels are calculated to construct an interferometric difference map group.

[0024] For each interferometric difference image in the interferometric difference image group, a weighted enhancement process is performed pixel by pixel, including: taking the currently selected pixel as the center pixel, selecting several neighboring pixels within a preset range, calculating the Euclidean distance between each neighboring pixel and the center pixel in the image coordinate system, taking the square inverse of the Euclidean distance as the spatial weight of the neighboring pixels, weighting and summing the brightness difference, color difference, and Sobel gradient difference of the neighboring pixels according to the spatial weight and taking the average, and updating the value of the center pixel;

[0025] All the weighted and enhanced interferometric differential images are superimposed at the corresponding pixel positions, and the superimposed results are normalized and averaged according to the number of images to generate an optical field interferometric enhancement image.

[0026] Optionally, the generation of the interference saliency mask map based on the optical field interference enhancement map specifically involves:

[0027] In the light field interference enhancement image, with each pixel as the center pixel, several adjacent pixels within a preset range are selected to form a local neighborhood. The average gray value of all adjacent pixels in the local neighborhood is calculated, and the difference between the gray value of each adjacent pixel and the average value is calculated. The difference is squared and accumulated, and then divided by the total number of adjacent pixels in the local neighborhood to obtain the local gray variance value of the pixel.

[0028] For the light field interference enhancement image, calculate the gray level change rate of each pixel in the horizontal and vertical directions respectively. Square the change rate of each pixel in the two directions and add them together to obtain the sum of squared directional gradients of the pixel. Taking each pixel as the center pixel, average the sum of squared directional gradients of all adjacent pixels in the local neighborhood to obtain the directional gradient energy value of the pixel.

[0029] The local grayscale variance of each pixel is averaged and weighted with the directional gradient energy value to obtain the fusion response value of each pixel and generate a fusion response map.

[0030] A pixel-by-pixel judgment is performed on the fused response map. When the fused response value is greater than the set response threshold, the pixel is assigned a value of 1, representing a significant region pixel; when the fused response value is less than or equal to the set response threshold, the pixel is assigned a value of 0, representing a non-significant region pixel.

[0031] All the 0 and 1 binary saliency markers corresponding to the pixels are combined to form an interference saliency mask with the same size as the light field interference enhancement map.

[0032] Optionally, the improved VMamba model is composed of a feature encoding module, a context folding and fusion module, a phase position embedding module, and a dynamic structure decoding module connected in sequence;

[0033] The feature encoding module is used to receive the light field interference enhancement map and the interference saliency mask map respectively, and take the pixel points corresponding to the saliency region positions with a value of 1 in the interference saliency mask map as the enhancement positions, multiply the pixel values ​​of each channel at the corresponding positions in the light field interference enhancement map by a preset enhancement coefficient, perform a numerical amplification operation, and generate an enhancement encoding feature tensor.

[0034] The context folding fusion module is used to receive the enhanced coded feature tensor and introduce an interference-guided context folding mechanism. The interference-guided context folding mechanism constructs a time series vector centered on each spatial location of the enhanced coded feature tensor, stacks feature segments at the same location in consecutive frames in frame order to form a time segment group, and calculates the channel numerical difference sequence between adjacent frame segments within each time segment group. The root mean square result of the numerical difference sequence is used as the interference phase stability factor. The interference phase stability factor is used as a weighting coefficient to perform a weighted summation of each frame segment within the time segment group to obtain the time-folded feature tensor.

[0035] The phase position embedding module is used to receive the time-folded feature tensor and obtain the set of boundary pixels of the salient region in the interference saliency mask image. The image coordinates are extracted by edge scanning. Based on the image center point, the angle between the line connecting each boundary pixel and the image center point in the image coordinate system is calculated. The angle value ranges from 0 to 360 degrees. All angles are re-encoded into a two-dimensional angle distribution map according to their position in the image space. The distribution map is then mapped into a phase position vector through sine and cosine functions. The phase position vector is embedded into the time-folded feature tensor by channel splicing to obtain the phase enhancement feature tensor.

[0036] The dynamic structure decoding module receives the phase-enhanced feature tensor, extracts temporal edge change features by setting a one-dimensional convolutional layer with a kernel size of 1×3 along the time dimension, and performs channel compression by using a convolutional layer with a kernel size of 1×1 in the channel dimension. The channel-compressed feature map is then upsampled in space using bilinear interpolation, and the upsampled result is concatenated with the feature map before channel compression in the channel dimension. The concatenated result is then convolved by a two-dimensional convolutional layer with a kernel size of 3×3 to output the dynamic edge response map of the structure.

[0037] Optionally, step three specifically includes:

[0038] Obtain a dynamic edge response map of the structure, wherein the value of each pixel in the dynamic edge response map is the edge response value;

[0039] A pixel-by-pixel threshold judgment is performed on the dynamic edge response map of the structure. Pixels with edge response values ​​greater than the set edge response threshold are marked as candidate edge points. The candidate edge points are sorted from largest to smallest edge response value. Pixels with edge response values ​​in the first set proportion range are selected from the sorting results as a set of high-confidence edge points.

[0040] Based on the set of high-confidence edge points, non-maximum suppression processing is performed in the local neighborhood of each high-confidence edge point. The edge response values ​​in the local neighborhood are compared pixel by pixel. Only local maximum pixels are retained and the rest are removed. All the retained local maximum pixels are combined into a cone contour tracking point set.

[0041] In the continuous multi-frame images corresponding to the light field interferometry enhancement map, an optical flow window of a set size is constructed with each cone contour tracking point as the center. The gray values ​​of all pixels in the optical flow window are extracted in the current frame and the next frame respectively. Based on the Lucas-Kanade optical flow algorithm, a set of linear constraint equations between pixel gray-level differences and pixel gradients in the optical flow window is constructed. The set of linear constraint equations is solved by the least squares method to obtain the displacement vector of each cone contour tracking point between adjacent frames.

[0042] For each cone contour tracking point, the displacement vector is recorded in frame order to form a time series displacement vector of the cone contour tracking point;

[0043] The global average displacement vector of the time series corresponding to all cone contour tracking points at the same time step is obtained by averaging the displacement vectors in space.

[0044] The global average displacement vectors of consecutive time steps are connected in chronological order to form the landing path correction.

[0045] Optionally, step four specifically includes:

[0046] The spatial attitude of the UAV is adjusted based on the landing path correction, so that the UAV maintains a stable alignment with the preset locking groove area. The current frame image after attitude adjustment is obtained, and the locking groove area image is extracted from the current frame image.

[0047] The image of the locking groove region is divided into several texture blocks. A two-dimensional fast Fourier transform is performed on each texture block to extract the frequency domain amplitude spectrum and phase spectrum of the texture block.

[0048] Extract the main frequency component from the frequency domain amplitude spectrum, extract the phase center and phase spread from the phase spectrum, and concatenate the main frequency component, phase center and phase spread to form the texture phase code of the current texture block;

[0049] The texture phase codes corresponding to all texture blocks in the image of the locking groove region are combined according to their positions in the image space to form the combined texture phase code of the entire locking groove region.

[0050] The combined texture phase code is matched one by one with the pre-established standard texture phase code library, the cosine similarity is calculated, and the center position corresponding to the standard texture phase code with the highest cosine similarity is selected as the center position of the locking groove.

[0051] Optionally, step five specifically includes:

[0052] Obtain the center position of the locking groove and the spatial attitude parameters of the current UAV, the spatial attitude parameters including the UAV's pose vector and attitude direction vector in three-dimensional space;

[0053] Establish a relative positional relationship between the center position of the locking groove and the current pose vector of the UAV in a spatial rectangular coordinate system, and construct an insertion path vector from the tip of the UAV locking toe to the center position of the locking groove. The insertion path vector starts from the UAV body coordinate system and ends at the center of the locking groove.

[0054] The insertion angle is calculated using the three-dimensional vector dot product formula based on the spatial angle between the insertion path vector and the current attitude direction vector of the UAV. The insertion angle is the inverse cosine of the angle between the insertion path vector and the attitude direction vector.

[0055] Based on the horizontal and vertical components of the insertion path vector, combined with the normal vector of the locking groove opening direction and the mating surface, the lateral and longitudinal offsets of the UAV's current position relative to the starting direction of the insertion path are calculated by geometric projection, thus forming the insertion offset vector.

[0056] The insertion angle and insertion offset vector are used as the decision parameters for insertion path selection, and the insertion path selection result is output.

[0057] Optionally, step six specifically includes:

[0058] Based on the insertion path selection result, the screw structure in the UAV locking mechanism is controlled to perform a clockwise rotation insertion operation along the selected insertion path.

[0059] During the rotation of the lead screw, a sequence of images of the locking structure is continuously acquired at set time intervals, and the sequence of images of the locking structure covers the entire screwing process.

[0060] Texture enhancement processing is performed on each frame of the locked structure image sequence, and the Sobel operator is applied to calculate the horizontal and vertical gradients of the image to extract the texture artifact regions generated by the locked structure components.

[0061] In the locked structure image sequence, the spatial position change of the same texture artifact region in the image coordinate system is detected frame by frame, a texture artifact displacement vector sequence is constructed, and the texture artifact displacement vectors are combined according to the frame order to generate an image artifact motion trajectory map.

[0062] Based on the image artifact motion trajectory map, calculate the cumulative trajectory length of the texture artifact, the consistency of the motion direction, the closure state of the edge contour of the artifact region in the termination frame, and the variance of the gray-level distribution within the region.

[0063] When the cumulative trajectory length exceeds the set length threshold, the consistency of the motion direction is higher than the set standard, the edge of the artifact region of the termination frame is closed and the gray-scale variance is lower than the set variance threshold, the locking state is determined to be successful; otherwise, the locking state is determined to be unsuccessful.

[0064] Optionally, step seven specifically includes:

[0065] If the insertion angle in the insertion path selection result is less than the set safety angle threshold, the insertion offset vector is less than the set safety offset threshold, and the locking status is judged as successful, then the insertion path is considered to be consistent with the actual locking process, and the docking is judged to be successful; otherwise, the docking is judged to be failed.

[0066] If docking fails, the UAV will be controlled to perform position reversal or fine-tuning operations, and the center position of the locking groove and the current spatial attitude parameters will be reacquired. The insertion path will be recalculated and a new round of docking attempts will be initiated until docking is successful or the maximum number of attempts is exceeded.

[0067] The beneficial effects of this invention are:

[0068] This invention constructs a high-contrast optical field interferometric enhancement map and introduces an interferometric saliency mask map as spatial attention guidance. Utilizing an improved VMamba model's interferometric guidance context folding mechanism, it significantly enhances the temporal consistency and edge saliency of target boundaries in the dynamic edge response map of the structure, solving the problem of recognizing locking structure features affected by background interference and weak texture occlusion. The path correction module combines the Lucas-Kanade optical flow algorithm to extract the time-series displacement vector of the cone contour tracking points, generating a landing path correction amount to achieve dynamic correction of the UAV's attitude and improve docking accuracy. The texture phase code constructed in the locking slot positioning module introduces a joint expression mechanism of frequency and phase features, and is compared with a standard texture phase code library. String similarity matching ensures the accuracy and robustness of locking groove center position identification. In the insertion path construction module, a three-dimensional angle calculation method between the insertion path vector and the attitude direction vector is proposed. Combined with the insertion offset vector, the insertion path selection result is quantitatively judged. The locking state detection module introduces the cumulative trajectory length of texture artifacts, motion direction consistency, and gray-scale distribution variance as judgment indicators based on the image artifact motion trajectory map. It can accurately evaluate the dynamic execution state of the locking process. Finally, through the docking judgment module, the insertion path selection result and locking state are jointly verified to achieve intelligent closed-loop judgment of docking success and failure, which improves the execution stability and task completion rate of the UAV autonomous docking and locking control system in complex environments. Attached Figure Description

[0069] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0070] Figure 1 This is a schematic diagram of the structure of the UAV visual autonomous docking and locking control system based on deep learning proposed in this invention;

[0071] Figure 2 This is an overall flowchart of the deep learning-based UAV visual autonomous docking and locking control method proposed in this invention;

[0072] Figure 3 This is a schematic diagram of the docking and locking process of the locking structure in the deep learning-based UAV visual autonomous docking and locking control method proposed in this invention. Detailed Implementation

[0073] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0074] refer to Figure 1 The deep learning-based UAV vision-based autonomous docking and locking control system includes the following modules:

[0075] The image acquisition and light field enhancement module is used to continuously acquire multiple frames of images from the UAV's downward-looking direction using an industrial camera, construct a light field interferometric enhancement map, and generate an interferometric saliency mask map.

[0076] The structural edge extraction module is used to input the light field interference enhancement map and the interference saliency mask map into the improved VMamba model. The improved VMamba model introduces an interference-guided context folding mechanism and outputs a structural dynamic edge response map.

[0077] The path correction module is used to extract the set of cone contour tracking points based on the dynamic edge response map of the structure, and to calculate the time series displacement vector of the cone contour tracking points using the Lucas-Kanade optical flow algorithm to generate the landing path correction amount.

[0078] The locking groove positioning module is used to adjust the attitude of the UAV based on the landing path correction amount, perform texture block and frequency domain analysis on the image of the locking groove area, generate texture phase code and match it with the standard texture phase code library to determine the center position of the locking groove.

[0079] The insertion path construction module is used to construct the insertion path between the locking toe and the locking groove based on the center position of the locking groove and the current attitude of the UAV, calculate the insertion angle and the insertion offset vector, and output the insertion path selection result.

[0080] The locking state detection module is used to control the rotation of the lead screw based on the insertion path selection result, acquire the locking structure image sequence, extract texture artifact information, construct the image artifact motion trajectory map, and determine the locking state.

[0081] The docking determination module is used to output the docking determination result based on the insertion path selection result and the locking status.

[0082] refer to Figure 2 A deep learning-based method for autonomous docking and locking control of unmanned aerial vehicles (UAVs) includes the following steps:

[0083] Step 1: Continuously acquire multiple frames of images from the UAV's downward-looking direction using an industrial camera, construct an optical field interferometry enhancement map, and generate an interferometry saliency mask map based on the optical field interferometry enhancement map;

[0084] Step 2: Input the light field interference enhancement map and the interference saliency mask map into the improved VMamba model. The improved VMamba model introduces an interference-guided context folding mechanism to obtain the dynamic edge response map of the structure.

[0085] Step 3: Based on the dynamic edge response map of the structure, determine the set of cone contour tracking points, and use the Lucas-Kanade optical flow algorithm to calculate the time series displacement vector of the cone contour tracking points to generate the landing path correction amount;

[0086] Step 4: Adjust the UAV attitude based on the landing path correction amount, perform texture block processing on the image of the locking groove area, generate texture phase code and match it with the standard texture phase code library to determine the center position of the locking groove.

[0087] Step 5: Based on the center position of the locking groove and the current attitude of the UAV, construct the insertion path between the locking toe and the locking groove, obtain the insertion angle and insertion offset vector through the geometric calculation method of the insertion path, and output the insertion path selection result;

[0088] Step 6: Based on the insertion path selection result, control the lead screw rotation, acquire the locking structure image sequence, extract texture artifact information, construct the image artifact motion trajectory map, and determine the locking state;

[0089] Step 7: Based on the insertion path selection result and locking status, output the docking determination result.

[0090] In this embodiment, the construction of the optical field interference enhancement map specifically involves:

[0091] An image pair is formed by taking any two frames from multiple frames continuously captured by an industrial camera. Pixel registration is performed on each image pair, and the brightness difference, color channel difference, and Sobel gradient difference between corresponding pixels are calculated to construct an interferometric difference map group. The brightness difference, color difference, and Sobel gradient difference reflect the differences of the image at the gray level, color level, and structural contour level, respectively, providing a comprehensive description of the multidimensional changes of the image.

[0092] For each interferometric difference image in the interferometric difference image group, a weighted enhancement process is performed pixel by pixel, including: taking the currently selected pixel as the center pixel, selecting several neighboring pixels within a preset range, calculating the Euclidean distance between each neighboring pixel and the center pixel in the image coordinate system, taking the square inverse of the Euclidean distance as the spatial weight of the neighboring pixels, weighting and summing the brightness difference, color difference, and Sobel gradient difference of the neighboring pixels according to the spatial weight and taking the average, and updating the value of the center pixel;

[0093] All the weighted and enhanced interferometric differential images are superimposed at the corresponding pixel positions, and the superimposed results are normalized and averaged according to the number of images to generate an optical field interferometric enhancement image.

[0094] In this embodiment, the generation of the interference saliency mask based on the optical field interference enhancement map specifically involves:

[0095] In the light field interference enhancement image, with each pixel as the center pixel, several adjacent pixels within a preset range are selected to form a local neighborhood. The average gray value of all adjacent pixels in the local neighborhood is calculated, and the difference between the gray value of each adjacent pixel and the average value is calculated. The difference is squared and accumulated, and then divided by the total number of adjacent pixels in the local neighborhood to obtain the local gray variance value of the pixel.

[0096] For the light field interference enhancement image, calculate the gray level change rate of each pixel in the horizontal and vertical directions respectively. Square the change rate of each pixel in the two directions and add them together to obtain the sum of squared directional gradients of the pixel. Taking each pixel as the center pixel, average the sum of squared directional gradients of all adjacent pixels in the local neighborhood to obtain the directional gradient energy value of the pixel.

[0097] The local grayscale variance of each pixel is averaged and weighted with the directional gradient energy value to obtain the fusion response value of each pixel and generate a fusion response map.

[0098] A pixel-by-pixel judgment is performed on the fused response map. When the fused response value is greater than the set response threshold, the pixel is assigned a value of 1, representing a significant region pixel; when the fused response value is less than or equal to the set response threshold, the pixel is assigned a value of 0, representing a non-significant region pixel.

[0099] All the 0 and 1 binary saliency markers corresponding to the pixels are combined to form an interference saliency mask with the same size as the light field interference enhancement map;

[0100] Compared to traditional edge extraction and region segmentation methods, this invention, based on the fusion of multi-frame image information, effectively highlights dynamic structural regions and weakens low-frequency background interference by integrating a multi-dimensional interference difference map group formed by brightness difference, color channel difference, and gradient direction response. This avoids the problem that single-frame edge information is easily affected by noise, occlusion, and changes in illumination.

[0101] Furthermore, the method of constructing a response map by integrating local gray-level variance and directional gradient energy can accurately extract features of high interference regions such as locking grooves and cone edges while maintaining the integrity of the image structure. It has stronger adaptability and robustness to complex environments such as non-uniform illumination and weak expression of structural texture. It is especially suitable for the visual docking preprocessing of UAVs under complex ground conditions such as natural light and multi-source shadows, significantly improving the docking recognition accuracy and structural stability.

[0102] In this embodiment, the improved VMamba model is composed of a feature encoding module, a context folding and fusion module, a phase position embedding module, and a dynamic structure decoding module connected in sequence;

[0103] The feature encoding module is used to receive the light field interference enhancement map and the interference saliency mask map respectively, and take the pixel points corresponding to the saliency region positions with a value of 1 in the interference saliency mask map as the enhancement positions, multiply the pixel values ​​of each channel at the corresponding positions in the light field interference enhancement map by a preset enhancement coefficient, perform a numerical amplification operation, and generate an enhancement encoding feature tensor.

[0104] The context folding fusion module is used to receive the enhanced coded feature tensor and introduce an interference-guided context folding mechanism. The interference-guided context folding mechanism constructs a time series vector centered on each spatial location of the enhanced coded feature tensor, stacks feature segments at the same location in consecutive frames in frame order to form a time segment group, and calculates the channel numerical difference sequence between adjacent frame segments within each time segment group. The root mean square result of the numerical difference sequence is used as the interference phase stability factor. The interference phase stability factor is used as a weighting coefficient to perform a weighted summation of each frame segment within the time segment group to obtain the time-folded feature tensor.

[0105] The phase position embedding module is used to receive the time-folded feature tensor and obtain the set of boundary pixels of the salient region in the interference saliency mask image. The image coordinates are extracted by edge scanning. Based on the image center point, the angle between the line connecting each boundary pixel and the image center point in the image coordinate system is calculated. The angle value ranges from 0 to 360 degrees. All angles are re-encoded into a two-dimensional angle distribution map according to their position in the image space. The distribution map is then mapped into a phase position vector through sine and cosine functions. The phase position vector is embedded into the time-folded feature tensor by channel splicing to obtain the phase enhancement feature tensor.

[0106] The dynamic structure decoding module receives the phase-enhanced feature tensor, extracts temporal edge change features by setting a one-dimensional convolutional layer with a kernel size of 1×3 along the time dimension, and performs channel compression by using a convolutional layer with a kernel size of 1×1 in the channel dimension. The channel-compressed feature map is then upsampled in space using bilinear interpolation, and the upsampled result is concatenated with the feature map before channel compression in the channel dimension. The concatenated result is then convolved by a two-dimensional convolutional layer with a kernel size of 3×3 to output a dynamic edge response map of the structure.

[0107] The improved VMamba model described in this invention integrates interference saliency guidance, temporal stability measurement, and spatial phase awareness on the basis of traditional temporal modeling structure, thereby more effectively improving the edge modeling capability of key structures for UAV docking. The model is composed of a feature encoding module, a context folding fusion module, a phase position embedding module, and a dynamic structure decoding module in sequence, and has strong local guidance, adaptive modeling, and structural analysis capabilities.

[0108] In the feature encoding stage, an interferometric saliency mask is introduced as a guiding source. Based on the binary labels of salient regions, the corresponding channel values ​​of the optical field interferometric enhancement map are numerically enhanced, thereby making the downstream model more focused on salient regions with edge structure features. The context folding fusion module extracts the interferometric phase stability features of each spatial location in the temporal dimension by constructing time-slice groups and calculating the root mean square of the inter-frame differences, strengthening the model's ability to discriminate the structure of temporally stable regions. The phase position embedding module constructs a two-dimensional angle distribution map based on the spatial angle information of the salient region boundaries, uses sine and cosine functions for phase encoding, and embeds it into the backbone features, providing the model with a priori perception capability of spatial structure.

[0109] The final dynamic structure decoding module combines one-dimensional convolution, channel compression, and spatial reconstruction to construct a multi-scale fusion decoding path that can perceive temporal changes, spatial phase, and channel differences. It outputs stable edge position information with physical structure orientation. Through structured collaboration between modules, the VMamba model exhibits superior spatiotemporal structure perception and modeling performance in docking scenarios, effectively improving the accuracy of edge tracking and position prediction.

[0110] In this embodiment, step three specifically includes:

[0111] Obtain a structural dynamic edge response map. The value of each pixel in the structural dynamic edge response map is the edge response value, which represents the degree of edge change at that spatial location in the time series. The larger the edge response value, the more significant the structural contour change of that pixel in multiple frames of images.

[0112] A pixel-by-pixel threshold judgment is performed on the dynamic edge response map of the structure. Pixels with edge response values ​​greater than the set edge response threshold are marked as candidate edge points. The candidate edge points are sorted from largest to smallest edge response value. Pixels with edge response values ​​in the first set proportion range are selected from the sorting results as a set of high-confidence edge points.

[0113] Based on the set of high-confidence edge points, non-maximum suppression processing is performed in the local neighborhood of each high-confidence edge point. The edge response values ​​in the local neighborhood are compared pixel by pixel. Only local maximum pixels are retained and the rest are removed. All the retained local maximum pixels are combined into a cone contour tracking point set.

[0114] In the continuous multi-frame images corresponding to the light field interferometry enhancement map, an optical flow window of a set size is constructed with each cone contour tracking point as the center. The gray values ​​of all pixels in the optical flow window are extracted in the current frame and the next frame respectively. Based on the Lucas-Kanade optical flow algorithm, a set of linear constraint equations between pixel gray-level differences and pixel gradients in the optical flow window is constructed. The set of linear constraint equations is solved by the least squares method to obtain the displacement vector of each cone contour tracking point between adjacent frames.

[0115] The system of linear constraint equations is as follows:

[0116] ;

[0117] in, Indicates the first The image gradient of each pixel in the horizontal direction (x direction). Indicates the first The image gradient of each pixel in the vertical direction (y-direction). Indicates the first The grayscale difference of a pixel between two frames and For the pixel point to be solved in And displacement in the y-direction (optical flow). Indicates the total number of pixels;

[0118] For each cone contour tracking point, the displacement vector is recorded in frame order to form a time series displacement vector of the cone contour tracking point;

[0119] The global average displacement vector of the time series corresponding to all cone contour tracking points at the same time step is obtained by averaging the displacement vectors in space.

[0120] The global average displacement vectors of consecutive time steps are connected in chronological order to form the landing path correction.

[0121] In this embodiment, step four specifically includes:

[0122] The spatial attitude of the UAV is adjusted based on the landing path correction, so that the UAV maintains a stable alignment with the preset locking groove area. The current frame image after attitude adjustment is obtained, and the locking groove area image is extracted from the current frame image.

[0123] The image of the locking groove region is divided into several texture blocks. A two-dimensional fast Fourier transform is performed on each texture block to extract the frequency domain amplitude spectrum and phase spectrum of the texture block.

[0124] Extract the main frequency component from the frequency domain amplitude spectrum, extract the phase center and phase spread from the phase spectrum, and concatenate the main frequency component, phase center and phase spread to form the texture phase code of the current texture block;

[0125] The texture phase codes corresponding to all texture blocks in the image of the locking groove region are combined according to their positions in the image space to form the combined texture phase code of the entire locking groove region.

[0126] The combined texture phase code is matched one by one with the pre-established standard texture phase code library, the cosine similarity is calculated, and the center position corresponding to the standard texture phase code with the highest cosine similarity is selected as the center position of the locking groove.

[0127] In this invention, to extract discriminative texture phase features from the locking groove region image, a two-dimensional fast Fourier transform is first performed on each divided texture block to obtain its frequency domain amplitude spectrum and phase spectrum. In the frequency domain amplitude spectrum, the frequency components are sorted from high to low amplitude, and the highest amplitude components are selected as the set of principal frequency components to reflect the main periodic structure of the texture block in the frequency domain. For the phase spectrum, based on the phase angle distribution in polar coordinates, the phase angle values ​​of all frequency points are statistically analyzed, and their circular mean is calculated as the phase center. Simultaneously, by calculating the variance of all phase angle values, the phase diffusion degree, reflecting the degree of phase diffusion, is obtained to characterize the concentration and variability of the texture direction distribution. The texture phase code, composed of these three components, can simultaneously reflect the principal frequency characteristics and phase organization morphology of the texture structure, exhibiting strong robustness and discriminative power.

[0128] To achieve rapid identification and localization of unknown locking slot regions, a standard texture phase code library is pre-constructed. This library consists of multiple sets of manually annotated and accurately located locking slot region images. Each set of images undergoes texture segmentation, Fourier transform, and texture phase code generation operations, ultimately forming multiple standard texture block code images. Each code image is then bound and stored to its corresponding real physical center position. During actual matching, the input combined texture phase code is compared with the standard code images in the code library using a cosine similarity calculation. This ensures that the center position of the locking slot can be accurately reconstructed based on the highest matching result, improving not only positioning accuracy but also enhancing the system's anti-interference capability under complex texture interference.

[0129] In this embodiment, step five specifically includes:

[0130] The center position of the locking groove and the current spatial attitude parameters of the UAV are obtained. The spatial attitude parameters include the UAV's pose vector and attitude direction vector in three-dimensional space. The pose vector is a six-dimensional parameter combination of the UAV's position coordinates and attitude angles in three-dimensional space. The attitude angles include pitch angle, roll angle and yaw angle. The attitude direction vector is a unit direction vector calculated based on the attitude angles, representing the spatial direction in which the UAV is currently facing.

[0131] Establish a relative positional relationship between the center position of the locking groove and the current pose vector of the UAV in a spatial rectangular coordinate system, and construct an insertion path vector from the tip of the UAV locking toe to the center position of the locking groove. The insertion path vector starts from the UAV body coordinate system and ends at the center of the locking groove.

[0132] The insertion angle is calculated using the three-dimensional vector dot product formula based on the spatial angle between the insertion path vector and the current attitude direction vector of the UAV. The insertion angle is the inverse cosine of the angle between the insertion path vector and the attitude direction vector.

[0133] Based on the horizontal and vertical components of the insertion path vector, combined with the normal vector of the locking groove opening direction and the mating surface, the lateral and longitudinal offsets of the UAV's current position relative to the starting direction of the insertion path are calculated by geometric projection, thus forming the insertion offset vector.

[0134] The insertion angle and insertion offset vector are used as the decision parameters for insertion path selection, and the insertion path selection result is output.

[0135] In this embodiment, step six specifically includes:

[0136] Based on the insertion path selection result, the screw structure in the UAV locking mechanism is controlled to perform a clockwise rotation insertion operation along the selected insertion path.

[0137] During the rotation of the lead screw, a sequence of images of the locking structure is continuously acquired at set time intervals, and the sequence of images of the locking structure covers the entire screwing process.

[0138] Texture enhancement processing is performed on each frame of the locked structure image sequence, and the Sobel operator is applied to calculate the horizontal and vertical gradients of the image to extract the texture artifact regions generated by the locked structure components.

[0139] This invention extracts texture artifact regions generated by locking structure components from a sequence of locked structure images, primarily based on the periodic interference texture features caused by the rotation of the metal structure. First, the image undergoes contrast stretching and sharpening enhancement to strengthen the texture characteristics of the edges and details of each structure. Then, the Sobel operator is applied to calculate the gradients in the horizontal and vertical directions, obtaining gradient intensity images. The overall edge intensity distribution is obtained by calculating the gradient magnitude map (i.e., the square root of the sum of the squares of the horizontal and vertical gradients). Since the locking structure components form strip-shaped, ring-shaped, or radial repeating texture changes in the image during rotation, these appear as high-frequency and directionally consistent texture response regions in the gradient magnitude map. Based on these characteristics, a gradient intensity threshold is set to filter out significant edge regions. Combined with local direction consistency analysis, regions with consistent texture direction and continuous edge intensity are identified as texture artifact regions, thus distinguishing typical texture interference image regions caused by rotating components and providing a stable reference object for displacement trajectory calculation.

[0140] In the locked structure image sequence, the spatial position change of the same texture artifact region in the image coordinate system is detected frame by frame, a texture artifact displacement vector sequence is constructed, and the texture artifact displacement vectors are combined according to the frame order to generate an image artifact motion trajectory map.

[0141] Based on the image artifact motion trajectory map, calculate the cumulative trajectory length of the texture artifact, the consistency of the motion direction, the closure state of the edge contour of the artifact region in the termination frame, and the variance of the gray-level distribution within the region.

[0142] When the cumulative trajectory length exceeds the set length threshold, the consistency of the motion direction is higher than the set standard, the edge of the artifact region of the termination frame is closed and the gray-scale variance is lower than the set variance threshold, the locking state is determined to be successful; otherwise, the locking state is determined to be unsuccessful.

[0143] In this embodiment, step seven specifically includes:

[0144] If the insertion angle in the insertion path selection result is less than the set safety angle threshold, the insertion offset vector is less than the set safety offset threshold, and the locking status is judged as successful, then the insertion path is considered to be consistent with the actual locking process, and the docking is judged to be successful; otherwise, the docking is judged to be failed.

[0145] If docking fails, the UAV will be controlled to perform position reversal or fine-tuning operations, and the center position of the locking groove and the current spatial attitude parameters will be reacquired. The insertion path will be recalculated and a new round of docking attempts will be initiated until docking is successful or the maximum number of attempts is exceeded.

[0146] Example 1:

[0147] To verify the feasibility of this invention in practice, it was applied to the docking and locking control process of a certain type of vertical take-off and landing UAV in a platform-based autonomous charging and docking mission. In this scenario, the UAV needs to return to the ground base station after the mission, and complete the fixation and charging connection by visually recognizing, adjusting its attitude, and inserting the locking mechanism. Traditional methods mainly rely on external high-precision GNSS positioning and manual monitoring, which suffer from high insertion failure rate and unstable locking in complex environments.

[0148] In this embodiment, the UAV uses a bottom-mounted industrial camera to acquire multiple frames of images of the ground locking structure, constructs an optical field interference enhancement map, and suppresses background interference through a saliency mask to generate an interference saliency mask map. The optical field interference enhancement map and the interference saliency mask map are input into an improved VMamba model to output a structural dynamic edge response map. Based on the extracted structural dynamic edge response map, the Lucas-Kanade optical flow algorithm is used to track the contour points of the locking groove cone, calculate the time-series displacement vector, and obtain a precise landing path correction amount, thereby adjusting the UAV attitude.

[0149] After attitude adjustment, the system acquires an image of the locking groove area, performs texture segmentation and frequency domain analysis, extracts the main frequency component, phase center and diffusion to form a texture phase code, and matches it with the standard texture phase code library to locate the center position of the locking groove. Then, it constructs an insertion path vector and calculates the insertion angle and insertion offset vector by combining the UAV's attitude direction vector.

[0150] During the lead screw insertion process, the system continuously acquires images and extracts texture artifacts, generating an image artifact motion trajectory map. By statistically analyzing the cumulative trajectory length of the texture artifacts, the consistency of their motion direction, the closure state of the edge contour of the artifact region in the termination frame, and the variance of the gray-level distribution within the region, the system determines whether the locking is successful. Finally, combining the insertion path selection result and the locking status, the system outputs the final docking determination result.

[0151] In 120 consecutive platform-based autonomous locking task tests, the method of this invention showed significantly better stability and accuracy than traditional methods, as shown in Table 1 below.

[0152] Table 1. Comparison of self-connection data between the present invention and traditional methods.

[0153]

[0154] As can be seen from the data in Table 1, the UAV visual autonomous docking and locking control method based on the improved VMamba model and light field interference enhancement proposed in this invention is significantly better than the traditional method in all key performance indicators.

[0155] Regarding the docking success rate, the method of this invention achieved 97.5%, an improvement of 18.9% compared to the traditional method's 82.0%, indicating that the system has stronger robustness and positioning accuracy in complex environments. The average insertion time was reduced from 4.8 seconds to 2.3 seconds, with an insertion efficiency improvement of over 52%, significantly reducing the time consumed in the docking process and improving overall operational efficiency in high-frequency tasks. The number of attitude corrections was reduced from the traditional 3.4 times to 1.2 times, a decrease of 64.7%, reflecting the higher accuracy of this invention in cone contour tracking and path correction, avoiding the accumulation of errors caused by multiple adjustments.

[0156] The consistency score of locking artifact motion direction improved to 92.6, indicating a significant enhancement in the dynamic tracking stability of the artifact region. The variance of the locking grayscale distribution in the terminating frame decreased from 13.4 to 7.8, a reduction of 41.8%, indicating that the locking final state is more stable and reliable, and the texture changes are more convergent, which is beneficial for accurately judging the locking completion state.

[0157] This embodiment improves the image perception accuracy and structural edge recognition capability of UAVs during autonomous docking by introducing an enhanced optical field interference map and an improved VMamba model. By combining the structural dynamic edge response map and optical flow algorithm, it accurately generates the landing path correction amount, realizes the precise positioning of the locking groove center position, enhances the reliability of the insertion path, and improves the overall system stability and intelligence level.

[0158] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A deep learning-based UAV vision-based autonomous docking and locking control system, characterized in that, Includes the following modules: The image acquisition and light field enhancement module is used to continuously acquire multiple frames of images from the UAV's downward-looking direction using an industrial camera, construct a light field interferometric enhancement map, and generate an interferometric saliency mask map. The construction of the optical field interference enhancement map specifically involves: The image pairs are formed by combining any two frames from multiple frames continuously acquired by an industrial camera. Pixel registration is performed on each pair of images, and the brightness difference, color channel difference, and Sobel gradient difference between corresponding pixels are calculated to construct an interferometric difference map group. For each interferometric difference image in the interferometric difference image group, a weighted enhancement process is performed pixel by pixel, including: taking the currently selected pixel as the center pixel, selecting several neighboring pixels within a preset range, calculating the Euclidean distance between each neighboring pixel and the center pixel in the image coordinate system, taking the square inverse of the Euclidean distance as the spatial weight of the neighboring pixels, weighting and summing the brightness difference, color difference, and Sobel gradient difference of the neighboring pixels according to the spatial weight and taking the average, and updating the value of the center pixel; All the weighted and enhanced interferometric differential images are superimposed at the corresponding pixel positions, and the superimposed results are normalized and averaged according to the number of images to generate an optical field interferometric enhancement image. The generation of the interference saliency mask image is specifically as follows: In the light field interference enhancement image, with each pixel as the center pixel, several adjacent pixels within a preset range are selected to form a local neighborhood. The average gray value of all adjacent pixels in the local neighborhood is calculated, and the difference between the gray value of each adjacent pixel and the average value is calculated. The difference is squared and accumulated, and then divided by the total number of adjacent pixels in the local neighborhood to obtain the local gray variance value of the pixel. For the light field interference enhancement image, calculate the gray level change rate of each pixel in the horizontal and vertical directions respectively. Square the change rate of each pixel in the two directions and add them together to obtain the sum of squared directional gradients of the pixel. Taking each pixel as the center pixel, average the sum of squared directional gradients of all adjacent pixels in the local neighborhood to obtain the directional gradient energy value of the pixel. The local grayscale variance of each pixel is averaged and weighted with the directional gradient energy value to obtain the fusion response value of each pixel and generate a fusion response map. A pixel-by-pixel judgment is performed on the fused response map. When the fused response value is greater than the set response threshold, the pixel is assigned a value of 1, representing a significant region pixel; when the fused response value is less than or equal to the set response threshold, the pixel is assigned a value of 0, representing a non-significant region pixel. All the 0 and 1 binary saliency markers corresponding to the pixels are combined to form an interference saliency mask with the same size as the light field interference enhancement map; The structural edge extraction module is used to input the light field interference enhancement map and the interference saliency mask map into the improved VMamba model. The improved VMamba model introduces an interference-guided context folding mechanism and outputs a structural dynamic edge response map. The improved VMamba model is composed of a feature encoding module, a context folding and fusion module, a phase position embedding module, and a dynamic structure decoding module connected in sequence. The feature encoding module is used to receive the light field interference enhancement map and the interference saliency mask map respectively, and take the pixel points corresponding to the saliency region positions with a value of 1 in the interference saliency mask map as the enhancement positions, multiply the pixel values ​​of each channel at the corresponding positions in the light field interference enhancement map by a preset enhancement coefficient, perform a numerical amplification operation, and generate an enhancement encoding feature tensor. The context folding fusion module is used to receive the enhanced coded feature tensor and introduce an interference-guided context folding mechanism. The interference-guided context folding mechanism constructs a time series vector centered on each spatial location of the enhanced coded feature tensor, stacks feature segments at the same location in consecutive frames in frame order to form a time segment group, and calculates the channel numerical difference sequence between adjacent frame segments within each time segment group. The root mean square result of the numerical difference sequence is used as the interference phase stability factor. The interference phase stability factor is used as a weighting coefficient to perform a weighted summation of each frame segment within the time segment group to obtain the time-folded feature tensor. The phase position embedding module is used to receive the time-folded feature tensor and obtain the set of boundary pixels of the salient region in the interference saliency mask image. It extracts the image coordinates by edge scanning and calculates the angle between the line connecting each boundary pixel and the image center point in the image coordinate system with the image center point as the reference. All angles are re-encoded into a two-dimensional angle distribution map according to their position in the image space and mapped into a phase position vector through sine and cosine functions. The phase position vector is embedded into the time-folded feature tensor by channel splicing to obtain the phase enhancement feature tensor. The dynamic structure decoding module receives the phase-enhanced feature tensor, extracts temporal edge change features by setting a one-dimensional convolutional layer with a kernel size of 1×3 along the time dimension, and performs channel compression by using a convolutional layer with a kernel size of 1×1 in the channel dimension. The channel-compressed feature map is then upsampled in space using bilinear interpolation, and the upsampled result is concatenated with the feature map before channel compression in the channel dimension. The concatenated result is then convolved by a two-dimensional convolutional layer with a kernel size of 3×3 to output a dynamic edge response map of the structure. The path correction module is used to extract a set of cone contour tracking points based on the dynamic edge response map of the structure, and to calculate the time-series displacement vector of the cone contour tracking points using the Lucas-Kanade optical flow algorithm to generate the landing path correction amount, specifically: Obtain a dynamic edge response map of the structure, wherein the value of each pixel in the dynamic edge response map is the edge response value; A pixel-by-pixel threshold judgment is performed on the dynamic edge response map of the structure. Pixels with edge response values ​​greater than the set edge response threshold are marked as candidate edge points. The candidate edge points are sorted from largest to smallest edge response value. Pixels with edge response values ​​in the first set proportion range are selected from the sorting results as a set of high-confidence edge points. Based on the set of high-confidence edge points, non-maximum suppression processing is performed in the local neighborhood of each high-confidence edge point. The edge response values ​​in the local neighborhood are compared pixel by pixel. Only local maximum pixels are retained and the rest are removed. All the retained local maximum pixels are combined into a cone contour tracking point set. In the continuous multi-frame images corresponding to the light field interferometry enhancement map, an optical flow window of a set size is constructed with each cone contour tracking point as the center. The gray values ​​of all pixels in the optical flow window are extracted in the current frame and the next frame respectively. Based on the Lucas-Kanade optical flow algorithm, a set of linear constraint equations between pixel gray-level differences and pixel gradients in the optical flow window is constructed. The set of linear constraint equations is solved by the least squares method to obtain the displacement vector of each cone contour tracking point between adjacent frames. For each cone contour tracking point, the displacement vector is recorded in frame order to form a time series displacement vector of the cone contour tracking point; The global average displacement vector of the time series corresponding to all cone contour tracking points at the same time step is obtained by averaging the displacement vectors in space. Connect the global average displacement vectors of consecutive time steps in time order to form the landing path correction. The locking slot positioning module is used to adjust the UAV attitude based on the landing path correction amount. It performs texture segmentation and frequency domain analysis on the locking slot region image, generates a texture phase code, and matches it with a standard texture phase code library to determine the center position of the locking slot. Specifically: The spatial attitude of the UAV is adjusted based on the landing path correction, so that the UAV maintains a stable alignment with the preset locking groove area. The current frame image after attitude adjustment is obtained, and the locking groove area image is extracted from the current frame image. The image of the locking groove region is divided into several texture blocks. A two-dimensional fast Fourier transform is performed on each texture block to extract the frequency domain amplitude spectrum and phase spectrum of the texture block. Extract the main frequency component from the frequency domain amplitude spectrum, extract the phase center and phase spread from the phase spectrum, and concatenate the main frequency component, phase center and phase spread to form the texture phase code of the current texture block; The texture phase codes corresponding to all texture blocks in the image of the locking groove region are combined according to their positions in the image space to form the combined texture phase code of the entire locking groove region. The combined texture phase code is matched one by one with the pre-established standard texture phase code library, the cosine similarity is calculated, and the center position corresponding to the standard texture phase code with the highest cosine similarity is selected as the center position of the locking groove. The insertion path construction module is used to construct the insertion path between the locking toe and the locking groove based on the center position of the locking groove and the current attitude of the UAV, calculate the insertion angle and the insertion offset vector, and output the insertion path selection result. The locking state detection module is used to control the rotation of the lead screw based on the insertion path selection result, acquire the locking structure image sequence, extract texture artifact information, construct the image artifact motion trajectory map, and determine the locking state. The docking determination module is used to output the docking determination result based on the insertion path selection result and the locking status.

2. The deep learning-based UAV visual autonomous docking and locking control system according to claim 1, characterized in that, The modules are connected in the following way: Step 1: Continuously acquire multiple frames of images from the UAV's downward-looking direction using an industrial camera, construct an optical field interferometry enhancement map, and generate an interferometry saliency mask map based on the optical field interferometry enhancement map; Step 2: Input the light field interference enhancement map and the interference saliency mask map into the improved VMamba model. The improved VMamba model introduces an interference-guided context folding mechanism to obtain the dynamic edge response map of the structure. Step 3: Based on the dynamic edge response map of the structure, determine the set of cone contour tracking points, and use the Lucas-Kanade optical flow algorithm to calculate the time series displacement vector of the cone contour tracking points to generate the landing path correction amount; Step 4: Adjust the UAV attitude based on the landing path correction amount, perform texture block processing on the image of the locking groove area, generate texture phase code and match it with the standard texture phase code library to determine the center position of the locking groove. Step 5: Based on the center position of the locking groove and the current attitude of the UAV, construct the insertion path between the locking toe and the locking groove, obtain the insertion angle and insertion offset vector through the geometric calculation method of the insertion path, and output the insertion path selection result; Step 6: Based on the insertion path selection result, control the lead screw rotation, acquire the locking structure image sequence, extract texture artifact information, construct the image artifact motion trajectory map, and determine the locking state; Step 7: Based on the insertion path selection result and locking status, output the docking determination result.

3. The deep learning-based UAV visual autonomous docking and locking control system according to claim 2, characterized in that, Step five specifically involves: Obtain the center position of the locking groove and the spatial attitude parameters of the current UAV, the spatial attitude parameters including the UAV's pose vector and attitude direction vector in three-dimensional space; Establish a relative positional relationship between the center position of the locking groove and the current pose vector of the UAV in a spatial rectangular coordinate system, and construct an insertion path vector from the tip of the UAV locking toe to the center position of the locking groove. The insertion path vector starts from the UAV body coordinate system and ends at the center of the locking groove. The insertion angle is calculated using the three-dimensional vector dot product formula based on the spatial angle between the insertion path vector and the current attitude direction vector of the UAV. The insertion angle is the inverse cosine of the angle between the insertion path vector and the attitude direction vector. Based on the horizontal and vertical components of the insertion path vector, combined with the normal vector of the locking groove opening direction and the mating surface, the lateral and longitudinal offsets of the UAV's current position relative to the starting direction of the insertion path are calculated by geometric projection, thus forming the insertion offset vector. The insertion angle and insertion offset vector are used as the decision parameters for insertion path selection, and the insertion path selection result is output.

4. The deep learning-based UAV visual autonomous docking and locking control system according to claim 2, characterized in that, Step six specifically involves: Based on the insertion path selection result, the screw structure in the UAV locking mechanism is controlled to perform a clockwise rotation insertion operation along the selected insertion path. During the rotation of the lead screw, a sequence of images of the locking structure is continuously acquired at set time intervals, and the sequence of images of the locking structure covers the entire screwing process. Texture enhancement processing is performed on each frame of the locked structure image sequence, and the Sobel operator is applied to calculate the horizontal and vertical gradients of the image to extract the texture artifact regions generated by the locked structure components. In the locked structure image sequence, the spatial position change of the same texture artifact region in the image coordinate system is detected frame by frame, a texture artifact displacement vector sequence is constructed, and the texture artifact displacement vectors are combined according to the frame order to generate an image artifact motion trajectory map. Based on the image artifact motion trajectory map, calculate the cumulative trajectory length of the texture artifact, the consistency of the motion direction, the closure state of the edge contour of the artifact region in the termination frame, and the variance of the gray-level distribution within the region. When the cumulative trajectory length exceeds the set length threshold, the consistency of the motion direction is higher than the set standard, the edge of the artifact region of the termination frame is closed and the gray-scale variance is lower than the set variance threshold, the locking state is determined to be successful; otherwise, the locking state is determined to be unsuccessful.

5. The deep learning-based UAV visual autonomous docking and locking control system according to claim 2, characterized in that, Step seven specifically involves: If the insertion angle in the insertion path selection result is less than the set safety angle threshold, the insertion offset vector is less than the set safety offset threshold, and the locking status is judged as successful, then the insertion path is considered to be consistent with the actual locking process, and the docking is judged to be successful; otherwise, the docking is judged to be failed.

Citation Information

Patent Citations

  • Hybrid vision backbone architecture combining selective state space model blocks and transformer blocks

    US20250371326A1