An In-air Refueling Drogue Detection Method Based on Improved YOLOv5

By improving the detection method of YOLOv5, combining binocular image registration and adaptive pyramid feature fusion, the problem of insufficient adaptation to the scale changes of the cone sleeve during aerial refueling is solved, and efficient and accurate cone sleeve detection and positioning is achieved.

CN116843732BActive Publication Date: 2025-06-17AERONAUTICS RES INST OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310542755.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-15
Publication Date
2025-06-17
Estimated Expiration
2043-05-15

AI Technical Summary

Technical Problem

The existing target detection algorithm cannot effectively adapt to the scale changes of the cone sleeve from the remote small target to the nearest large target during aerial refueling, and the error detection rate is high and the real-time performance is insufficient.

Method used

Using the detection method of improved YOLOv5, through binocular image registration and adaptive pyramid feature fusion, YOLOv5's backbone network and feature fusion module are optimized, combined with artificial operators to extract optical marking features, and three-dimensional coordinate information is inferred.

Benefits of technology

It realizes fast and accurate detection and positioning of the cone sleeve, improves the real-time and adaptability of the detection algorithm, and reduces the false detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116843732B_ABST
    Figure CN116843732B_ABST
Patent Text Reader

Abstract

The present invention discloses an in-air refueling drogue detection method based on improved YOLOv5, belonging to the technical field of computer vision recognition. The present invention uses a feature-based binocular image registration and fusion algorithm to register the drogue images detected by binocular cameras, and then further designs and optimizes the YOLOv5 lightweight object detection algorithm using adaptive pyramid feature fusion, which can achieve the rotational invariance and size invariance of the drogue image recognition algorithm, enabling the network to directly learn how to perform spatial filtering on features and adaptively fuse features at different levels. Finally, the optimized object detection method is used to solve the dual-view images, and then the pose of the drogue relative to the observation camera is obtained. This method ensures the stability of the results by using geometric matching based on binocular vision, and at the same time greatly improves the detection accuracy and speed of the target by using the adaptive neural network algorithm based on deep learning, realizing the rapid detection and positioning of the refueling drogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method for detecting an in-air refueling drogue based on improved YOLOv5, belonging to the technical field of computer vision recognition. Background Art

[0002] The in-air refueling technology can increase the range, endurance, operating space, payload, etc. of an aircraft without changing the original aircraft structure, thereby enhancing the combat effectiveness. Autonomous in-air refueling requires that the sensor error during the docking process is not less than the centimeter level. Visual navigation, with the characteristics of fast real-time and high precision, has become the main method to solve this problem, and achieving accurate detection and tracking of the drogue has become the key to completing this technology.

[0003] Existing drogue detection technologies often combine deep learning networks for method research. For example, the Unmanned System Technology Research Institute of Northwestern Polytechnical University combines the optical flow method with the YOLOv3 (You Only Look Once, YOLO) network, aligns the feature maps extracted from adjacent frames to the current frame for feature aggregation, to solve the problem of degraded frames due to motion blur, defocus, occlusion, etc. at the camera imaging end (see Tao Chengyang, Yuan Jie, Hui Tian, "An Algorithm for Detecting In-air Refueling Drogue Based on Feature Aggregation", "Modern Electronics Technique", 2022, 45(13): 152-158). In addition, Wang Gangzhi et al. from the China National Aeronautical Radio Electronics Research Institute use an improved Single Shot MultiBox Detector (SSD) object detection algorithm to design feature fusion of multi-scale information, improving the original test speed and detection accuracy (see Wang Gangzhi, Wang Xinhua, Chen Guanyu, "Research on Vision-based UAV In-air Refueling Target Recognition Technology", "Electronic Measurement Technology", 2020, 43(13): 89-94). Since only deep learning networks are used to detect the refueling drogue, the real-time performance of the above detection methods cannot meet the requirements of actual applications, and at the same time, they cannot well adapt to the scale change characteristics of the detection object size target.

[0004] With the increasing urgency of the demand for autonomous in-air refueling, it is particularly necessary to develop an algorithm that can have good real-time performance while occupying and consuming less on-board computing power, especially to meet the requirements of quickly identifying and accurately positioning the position information of the refueling drogue. This solution is generated based on the above ideas. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the problem that the existing object detection algorithm is insufficient in adapting to the scale change of the drogue from a small target at a far end to a large target at a near end during the in-air refueling process, as well as the problems of relatively high false detection rate of the current algorithm and insufficient real-time performance of the airborne algorithm. Therefore, a method for detecting the in-air refueling drogue based on improved YOLOv5 is proposed. Drawing on the advanced achievement of the YOLOv5 object detection framework in the field of object detection, the binocular images are first encoded and then decoded using the image registration method, and the lightweight object detection algorithm of YOLOv5 is further designed and optimized using adaptively spatial feature fusion (ASFF). The artificial operator is used to extract and match the features of the optical markers, and then the three-dimensional coordinate information of the optical markers is inferred.

[0006] The present invention adopts the following technical solutions to achieve the above invention purpose. A method for detecting the in-air refueling drogue based on improved YOLOv5 includes the following steps:

[0007] A. Use a binocular camera to collect the images of the refueling drogue in the refueling scene;

[0008] B. Use a feature-based binocular image registration and fusion algorithm to register the two sets of drogue images collected by the binocular camera, ensure that each row of pixels of the drogue target in the dual-view images is aligned, obtain the transformation matrix of the two sets of images, and the images can be successfully matched through inverse operation;

[0009] C. Optimize the backbone network and the feature fusion part of the YOLOv5 detection model, and replace the backbone network with the MobileNetV3 structure;

[0010] D. Introduce an efficient channel attention (ECA) module for deep convolutional neural networks into MobileNetV3 to reduce the network parameter quantity;

[0011] E. Introduce an adaptive feature pyramid fusion module into the YOLOv5 convolutional network;

[0012] F. Use the optimized detection model to detect the two sets of drogue images to obtain the pixel range of the drogue target;

[0013] G. Use the traditional artificial operator to solve the drogue target to obtain the final pose of the drogue.

[0014] Step A specifically is as follows: The camera acquires binocular images camera_image_1 and camera_image_2 of the refueling scene and saves them into Image[1, 2], where 1 and 2 represent the image serial numbers, and the expression Image[1, 2] = [camera_image_1, camera_image_2] is adopted;

[0015] Among them, camera_image_1 and camera_image_2 represent two original RGB images acquired by the binocular camera. The RGB image refers to a color image composed of three primary color components, namely R (Red), G (Green), and B (blue).

[0016] Step B adopts a feature-based binocular image registration and fusion algorithm, which uses the accelerated feature point extraction and description ORB algorithm to extract and match two groups of cone sleeve images collected by the binocular camera. Feature points are detected in the two groups of images, and then specific descriptors are generated. The transformation matrix of the two groups of images is calculated using the corresponding multiple groups of feature points, and the inverse operation can successfully match the images.

[0017] The extraction process of the cone sleeve image is as follows: The defects at the feature points on the cone sleeve are improved by constructing an image pyramid and the gray centroid method to obtain image information with scale and rotation invariance ORB features. The image pyramid is to sample and process the image into different sizes and stack them, extract feature points on images of different sizes, and perform feature matching on adjacent two-layer graphics at the same time. The gray centroid method is to take the center of the gray value of the image block as the weight.

[0018] The image information is processed, and at the same time, the corner response of each feature point is calculated, and all the calculation results are arranged, and the required number of feature points are selected in order as the final result;

[0019] The formula for calculating the corner response value is:

[0020] R = detM - k(traceM) 2 (1)

[0021] In the formula, R is the corner response function, k is an empirical constant usually taking a value of 0.04 - 0.06, M is a real symmetric matrix of eigenvalues, detM is the determinant of matrix M, and traceM is the matrix trace.

[0022] The matching process of the conical sleeve image is as follows: Describe the corner points according to the information in the adjacent areas of the extracted corner points to form a binary coding BRIEF descriptor. Retain the feature points with a higher matching degree of the conical sleeve feature points in the two groups of images as matching points, which serve as the basic input information for image registration. Establish a homography transformation to complete the registration of the two groups of images.

[0023] There is a homography between the images of the same planar object taken by a lensless distorted camera from different positions, which can be represented by the homography transformation:

[0024]

[0025] In the formula, x l , y l represent the coordinates of the feature points in the left-eye image, and x r , y r represent the coordinates of the feature points in the right-eye image that match the left-eye image. H 3×3 represents the homography matrix. Expanding the homography matrix gives:

[0026]

[0027] It can be converted to:

[0028]

[0029] One set of matching points can obtain 2 sets of equations, and at least five sets of matching points can be used to solve the homography matrix H 3×3 , thereby obtaining the conversion relationship between the left-eye and right-eye images. By using the registration of binocular images, the coordinate positions of the conical sleeves in both binocular images can be obtained simultaneously with only one detection.

[0030] Specifically, step D is as follows: In the efficient channel attention module, the number of channels is C, and each channel attention module has k×C parameters. The interaction between different channels is as follows:

[0031]

[0032] In the formula, σ is the Sigmoid activation function, represents the set of k adjacent channels of y i ;

[0033] The efficient channel attention module is quickly implemented using a one-dimensional convolutional kernel of size k, that is:

[0034] ω = σ(C1D k (y)) (6)

[0035] In the formula, C1D k represents a one-dimensional convolutional kernel of size k;

[0036] For the size of the one-dimensional convolutional kernel k, through an adaptive selection method, local cross-channel interaction information is obtained, and there is a mapping relationship:

[0037]

[0038] In addition, since the number of channels C is usually a power of 2, use 2 (γ*k-b) to replace exp(γ*k - b); given the number of channels C, the size of the kernel k is adaptively determined as:

[0039]

[0040] where |t| odd represents the nearest odd number to t, and both γ and b are hyperparameters.

[0041] The function of the adaptive pyramid feature fusion ASFF module described in step E is: enabling the network to directly learn how to spatially filter features at other levels, so as to only retain useful information for combination, thereby enhancing the network's feature fusion ability for targets of different scales, and can solve the inconsistency inside the feature pyramid in the first-order detector;

[0042] The method for the ASFF module to fuse is: after identity scaling, the vectors at the position (i, j) of the three feature maps are weighted and fused to obtain the vector at the new spatial position (i, j), and the coefficient weights are learned by the network adaptively. The process is expressed as:

[0043]

[0044]

[0045] where the respective weight matrices and represents the feature vector at the position (i, j) after adjusting the features from level n to level l.

[0046] Step F is specifically: using the expression object_location[1,2] =

[0047] F_network(I_match[1,2]), the pixel range of the cone target is obtained. By inputting the matched binocular image match I_ into the optimized detection network F_network, we can obtain the pixel coordinates object_location of the detected cone target, and its format is

[0048] [x,y,w,h], where x and y represent the center point of the target, and w and h represent the width and height of the detection box.

[0049] Step G is specifically as follows: The expression object_pose = pose_calculate(I_match[1,2], object_location[1,2]) is adopted, and the traditional artificial operator is used to calculate the pose of the target feature. Here, pose_calculate represents the artificial feature pose calculation method. By extracting and calculating the artificial features within the target range object_location of the binocular image I_match after matching, the six-degree-of-freedom pose [X, Y, Z, yaw, pitch, roll] of the cone sleeve target can be finally obtained. Among them, X, Y, and Z represent the center coordinates of the cone sleeve in the world coordinates, and yaw, pitch, and roll represent the rotation angles of the three axes.

[0050] The present invention adopts the above technical solutions and has the following beneficial effects: The present invention uses a feature-based binocular image registration and fusion algorithm to register the cone sleeve images detected by the binocular camera, and then further designs and optimizes the YOLOv5 lightweight target detection algorithm using adaptive pyramid feature fusion, which can achieve the rotation invariance and size invariance of the cone sleeve image recognition algorithm, enabling the network to directly learn how to perform spatial filtering on features and adaptively fuse features at different levels. Finally, the optimized target detection method is used to calculate the dual-view images, and thus the pose of the cone sleeve relative to the observation camera is obtained. This method ensures the stability of the results by using geometric matching based on binocular vision, and at the same time greatly improves the detection accuracy and speed of the target by using the adaptive neural network algorithm based on deep learning, realizing the rapid detection and positioning of the refueling cone sleeve. Description of the Drawings

[0051] Figure 1 It is the calculation flow chart of the method of the present invention;

[0052] Figure 2 It is the cone sleeve with artificial marked feature points added;

[0053] Figure 3 It is the performance comparison chart of different algorithm models and the improved YOLOv5-l and improved YOLOv5-s of the present invention. Specific Embodiments

[0054] The technical solutions of the invention will be described in detail below with reference to the drawings.

[0055] An in-air refueling drogue detection method based on improved YOLOv5 uses a binocular camera to collect drogue images in the refueling scenario, registers the drogue images collected by the binocular camera using a feature-based binocular image registration and fusion algorithm, and then further designs and optimizes the backbone network and feature fusion module of the YOLOv5 object detection algorithm. The optimized object detection method is used to detect and fuse information from the dual-view images, and then the pose of the drogue relative to the observation camera is obtained. The specific flowchart is as shown in Figure 1 shown, and includes the following steps:

[0056] Step 1, the camera acquires the binocular images camera_image_1 and camera_image_2 of the refueling scenario and saves them in Image[1, 2], where 1 and 2 represent the image numbers. The expression Image[1, 2] = [camera_image_1, camera_image_2] is used;

[0057] where camera_image_1 and camera_image_2 are two original RGB images acquired by the binocular camera; the RGB image refers to a color image composed of three primary color components, namely R (Red), G (Green), and B (blue).

[0058] Step 2, the feature-based binocular image registration and fusion algorithm is used. The accelerated feature point extraction and description (Oriented FAST and Rotated BRIEF, i.e., ORB) algorithm is used to extract and match two sets of drogue images collected by the binocular camera. Feature points are detected in the two sets of images, and then specific descriptors are generated. The transformation matrix of the two sets of images is calculated using the corresponding multiple sets of feature points, and the images can be successfully matched through inverse operation. The drogue with manually marked feature points is as shown in Figure 2 shown.

[0059] Step 201, the specific feature extraction process is as follows:

[0060] The defects at the feature points (the image features do not have rotation invariance, and the features between different scales are inconsistent when the image size changes) are improved by constructing an image pyramid and the gray centroid method to obtain an ORB feature extraction method with scale and rotation invariance. The image pyramid is to sample and process the image into different sizes and stack them, extract feature points on images of different sizes, and perform feature matching on adjacent two-layer images at the same time. The gray centroid method is to use the gray value of the image block as the center of weight to achieve the scale invariance of the feature points. The image information is processed, and at the same time, the corner response of each feature point is calculated, and all the calculation results are arranged, and the required number of feature points are selected in order as the final result.

[0061] The formula for calculating the corner response value is:

[0062] R = detM - k(traceM) 2 (1)

[0063] Where R is the corner response function, k is an empirical constant usually taking values between 0.04 - 0.06, M is a real symmetric matrix of eigenvalues, detM is the determinant of matrix M, and traceM is the matrix trace.

[0064] Step 202, the specific feature matching process is as follows:

[0065] If there are enough pixels around a pixel with a large difference in value from it, then this point is very likely to be a corner point. Describe the corner points according to the extracted image information to form a binary coding descriptor to accelerate the feature matching rate. Retain some feature points with a higher matching degree as the basic input information for image registration, establish a homography transformation, and then the registration of two sets of images can be completed. There is a homography between the images of the same planar object taken by a camera without lens distortion from different positions, which can be represented by a homography transformation:

[0066]

[0067] Where x l , y l represent the coordinates of the feature points of the left-eye image, x r , y r represent the coordinates of the feature points of the right-eye image that match the left-eye image, H 3×3 represents the homography matrix. Expanding the homography matrix gives:

[0068]

[0069] It can be converted to:

[0070]

[0071] One set of matching points can obtain 2 sets of equations, and at least five sets of matching points can be used to solve the homography matrix H 3×3 , so as to obtain the conversion relationship between the left-eye and right-eye images, and the coordinate positions of the cone sleeves in the binocular images can be achieved simultaneously or only by detecting once using the registration of binocular images.

[0072] Step 3, optimize the backbone network of the YOLOv5 detection model, and replace the backbone network with the MobileNetV3 structure. For MobileNetV3, by introducing an ECA module for deep convolutional neural networks, dimensionality reduction is avoided, and local cross-channel interaction is efficiently achieved using one-dimensional convolution, effectively reducing the number of network parameters.

[0073] In the efficient channel attention module, the number of channels is C, and each channel attention module has k×C parameters. The interaction between different channels is as follows:

[0074]

[0075] where is the Sigmoid activation function, represents the set of k adjacent channels of y i .

[0076] The efficient channel attention module can be quickly implemented using a one-dimensional convolutional kernel of size k, that is:

[0077] ω = σ(C1D k (y)) (6)

[0078] where C1D k represents a one-dimensional convolutional kernel of size k.

[0079] For the size of the one-dimensional convolutional kernel k, through an adaptive selection method, manual adjustment of the k value is avoided, which consumes a large amount of computing resources. Obtaining local cross-channel interaction information determines the coverage range of the interaction, and this coverage range may vary with different numbers of channels and different convolutional neural network structures, and there is a mapping relationship:

[0080]

[0081] In addition, since the number of channels C is usually a power of 2, use 2 (γ*k-b) to replace exp(γ*k - b). Given the number of channels C, the size of the kernel k is adaptively determined as:

[0082]

[0083] where |t| odd represents the nearest odd number of t, and the hyperparameters γ = 2 and b = 1 are set.

[0084] Step 4, introduce adaptive pyramid feature fusion, enabling the network to directly learn how to perform spatial filtering on features at other levels, so as to only retain useful information for combination, thereby enhancing the network's feature fusion ability for targets of different scales and solving the inconsistency inside the feature pyramid in the first-order detector.

[0085] After identity scaling, the vectors at the position (i, j) of the three feature maps are weighted and fused to obtain the vector at the new spatial position (i, j), and the coefficient weights are learned by the network adaptively. The process is expressed as:

[0086]

[0087]

[0088] wherein each weight matrix and represents the feature vector at (i, j) after the feature adjustment from level n to level l.

[0089] Step 5: Use the optimized detection model to detect the dual-view image to obtain the pixel range of the cone sleeve target. object_location[1,2] =

[0090] F_network(I_match[1,2]). Specifically, input the registered binocular image I_match[1,2] into the modified and optimized detection network F_network, and the result output by the network is saved in object_location in the format of [x, y, w, h], where x and y represent the center point of the target, and w and h represent the width and height of the detection box.

[0091] Step 6: Use traditional artificial operators to calculate the pose of the cone sleeve to obtain the final pose of the cone sleeve. Specifically, use the target area coordinates object_location obtained by the detection algorithm in the previous step and the registered image I_match to extract the local pixels of the target in the two sets of images, and perform matching calculation through the pre-set feature information to obtain the final pose of the cone sleeve [X, Y, Z, yaw, pitch, roll], where X, Y, and Z represent the cone sleeve center coordinates in the world coordinate system, and yaw, pitch, and roll represent the rotation angles. Finally, the performance comparison diagrams of different algorithm models with the improved YOLOv5-l and improved YOLOv5-s of the present invention are as Figure 3 shown.

Claims

1. An in-air refueling drogue detection method based on improved YOLOv5, characterized in that It includes the following steps: A. Use a binocular camera to collect the images of the fueling cone sleeve in the fueling scenario; B. Use a feature-based binocular image registration and fusion algorithm to register the two sets of cone sleeve images collected by the binocular camera, ensure that the pixels of the cone sleeve target in the dual-view images are aligned row by row, obtain the transformation matrix of the two sets of images, and the inverse operation can successfully match the images; C. Optimize the backbone network and the feature fusion part of the YOLOv5 detection model, and replace the backbone network with the MobileNetV3 structure; D. Introduce an efficient channel attention ECA module for deep convolutional neural networks into MobileNetV3 to reduce the number of network parameters; E. Introduce an adaptive feature pyramid fusion ASFF module into the YOLOv5 convolutional network; F. Use the optimized detection model to detect the two sets of registered cone sleeve images to obtain the pixel range of the cone sleeve target; adopt the expression object_location[1,2] = F_network(I_match[1,2]), obtain the pixel range of the cone sleeve target, and by inputting the matched binocular image match I_ into the optimized detection network F_network, obtain the pixel coordinates object_location of the detected cone sleeve target, and its format is [x,y,w,h], where x and y represent the center point of the target, and w and h represent the width and height of the detection box; G. Use traditional artificial operators to solve the cone sleeve target to obtain the final pose of the cone sleeve; adopt the expression object_pose = pose_calculate(I_match[1,2],object_location[1,2]), use traditional artificial operators to solve the target features to obtain the pose, where pose_calculate represents the artificial feature pose calculation method, and by extracting and calculating the artificial features within the target range object_location of the matched binocular image I_match, the six-degree-of-freedom pose of the cone sleeve target can be finally obtained [X,Y,Z,yaw,pitch,roll], where X, Y, and Z represent the cone sleeve center coordinates in the world coordinates, and yaw, pitch, and roll represent the three-axis rotation angles.

2. The in-air refueling drogue detection method based on improved YOLOv5 according to claim 1, characterized in that Step A is specifically as follows: The camera obtains the binocular images camera_image_1 and camera_image_2 of the fueling scenario and saves them in Image[1,2]. 1 and 2 represent the image numbers, and the expression Image[1,2] = [camera_image_1, camera_image_2] is adopted; Among them, camera_image_1 and camera_image_2 represent the two original RGB images obtained by the binocular camera, and the RGB image refers to a color image composed of three primary color components of Red, Green, and blue.

3. The in-air refueling drogue detection method based on improved YOLOv5 according to claim 1, characterized in that Step B adopts a feature-based binocular image registration and fusion algorithm, which uses the ORB algorithm for accelerated feature point extraction and description to extract and match two sets of cone sleeve images collected by the binocular camera. Feature points are detected in the two sets of images, and then specific descriptors are generated. The transformation matrix of the two sets of images is calculated using the corresponding multiple sets of feature points, and the images can be successfully matched through inverse operation.

4. The in-air refueling drogue detection method based on improved YOLOv5 according to claim 3, characterized in that The extraction process of the cone sleeve images is as follows: It is improved by constructing an image pyramid and the gray centroid method to extract the defects at the feature points on the cone sleeve, and the image information with ORB features with scale and rotation invariance is obtained. The image pyramid is to sample the image into different sizes and stack them, extract feature points on the images of different sizes, and perform feature matching on adjacent two-layer graphics at the same time. The gray centroid method is to use the gray value of the image block as the center of weight. Process the image information, calculate the corner response of each feature point at the same time, and arrange all the calculation results. Select the required number of feature points in order as the final result. The formula for calculating the corner response value is: R = det M - k(trace M) 2 (1) In the formula, R is the corner response function, k is an empirical constant usually taking a value of 0.04 - 0.06, M is the real symmetric matrix of eigenvalues, detM is the determinant of matrix M, and traceM is the matrix trace.

5. The in-air refueling drogue detection method based on improved YOLOv5 according to claim 4, characterized in that The matching process of the cone sleeve images is as follows: Describe the corner according to the information in the adjacent area of the extracted corner to form a binary encoding BRIEF descriptor. Retain the feature points with higher matching degrees of the cone sleeve feature points in the two sets of images as the matching points, which are used as the basic input information for image registration. Establish a homography transformation to complete the registration of the two sets of images.

6. The in-air refueling drogue detection method based on improved YOLOv5 according to claim 5, characterized in that The homography transformation means that there is a homography between the images of the same planar object taken by a lensless-distortion camera from different positions, which can be represented by the homography transformation: where x l , y l represent the coordinates of the feature points of the left-eye image, and x r , y r represent the coordinates of the feature points of the right-eye image that match the left-eye image. H 3×3 represents the homography matrix. Expanding the homography matrix gives: It can be converted to: A set of matching points can yield two sets of equations, and at least five sets of matching points are required to solve for the homography matrix H 3×3 , thereby obtaining the transformation relationship between the left and right eye images. By using the registration of binocular images, it is possible to simultaneously obtain the coordinate positions of the cone sleeves in both binocular images with only one detection.

7. The in-air refueling drogue detection method based on improved YOLOv5 according to claim 1, characterized in that The specific content of Step D is: In the efficient channel attention module, the number of channels is C, and each channel attention module has k×C parameters. The interaction between different channels is as follows: where σ is the Sigmoid activation function, represents the set of i k adjacent channels of y; The efficient channel attention module is quickly implemented using a one-dimensional convolutional kernel of size k, that is: ω=σ(C1D k (y)) (6) where C1D k represents a one-dimensional convolutional kernel of size k; Through an adaptive selection method for the size of the one-dimensional convolutional kernel k, local cross-channel interaction information is obtained, and there is a mapping relationship: In addition, since the number of channels C is usually a power of 2, use 2 (γ*k-b) to replace exp(γ*k - b); given the number of channels C, adaptively determine the kernel k size as: where |t| odd represents the nearest odd number of t, and both γ and b are hyperparameters.

8. The method for detecting the aerial refueling drogue based on the improved YOLOv5 according to claim 1, wherein, The function of the adaptive pyramid feature fusion ASFF module in Step E is: enabling the network to directly learn how to perform spatial filtering on features at other levels, so as to only retain useful information for combination, thereby enhancing the network's feature fusion ability for targets of different scales, and can solve the inconsistency inside the feature pyramid in the first-order detector. The fusion method of the ASFF module is: After identity scaling, the vectors at the position (i, j) of the three feature maps are weighted and fused to obtain the vector at the new spatial position (i, j). The coefficient weights are learned by the network adaptively. The process is expressed as: Wherein, each weight matrix and represents the feature vector at (i, j) after the feature adjustment from level n to level l.

Citation Information

Patent Citations

  • Rapid pest detection method based on improved YOLO V4

    CN114220035A

  • Agaricus bisporus detection method based on shielding condition

    CN115294036A