Unmanned aerial vehicle logistics distribution method and device, computer equipment and storage medium

By combining the improved YOLOv8 model with FMCW radar, high-precision target detection and positioning of drones in foggy conditions was achieved, solving the positioning accuracy and obstacle avoidance problems in harsh environments, and improving the safety and efficiency of drone delivery.

CN120595820APending Publication Date: 2025-09-05SHENZHEN HUANGUOYUN LOGISTICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510709964.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The target detection accuracy and positioning efficiency of drones in harsh environments are low, especially in low visibility conditions such as fog. Existing vision and radar technologies are difficult to meet the needs of high-precision positioning and obstacle avoidance.

Method used

The improved YOLOv8 model with integrated feature attention fusion network is combined with FMCW radar for foggy image dehazing and enhancement. Targets are identified through channel and pixel attention, and combined with dynamic correction of three-dimensional coordinates and secondary identity verification to achieve high-precision target detection and positioning of UAVs.

Benefits of technology

In foggy scenarios, it significantly improves the drone’s target detection accuracy and positioning efficiency, reduces hardware costs and power consumption, and improves delivery safety and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595820A_ABST
    Figure CN120595820A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle navigation, and discloses an unmanned aerial vehicle logistics distribution method and device, computer equipment and a storage medium, and the method comprises the steps: building a feature attention fusion network based on a YOLOv8 model, enhancing the foggy day image feature extraction capability through a channel and pixel attention module, and improving the target recognition precision; resolving three-dimensional coordinates of a target and dynamically correcting a flight path by combining a frequency-modulated continuous wave radar and array antenna phase difference technology; the unmanned aerial vehicle is controlled to move through an ROS framework, hovers when approaching a target and performs secondary visual verification, and safe landing is ensured. According to the method, the improved model and the multi-sensor technology are fused, the detection reliability, the positioning precision and the distribution efficiency of the unmanned aerial vehicle in foggy days, rainy days and other scenes are remarkably improved through a low-cost scheme, and the method is suitable for urban logistics and remote region transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drone navigation technology, and in particular to a drone logistics distribution method, device, computer equipment and storage medium. Background Art

[0002] With the rapid development of the logistics industry, drone delivery has gradually become an important supplement to traditional logistics due to its high efficiency and flexibility. Compared with ground transportation, drones can significantly shorten delivery times by using aerial routes, especially in urban areas with traffic congestion or remote areas. In addition, the automated operation of drones significantly reduces labor costs, alleviates the burden of manual sorting and delivery, and further improves logistics efficiency. However, existing drone technology still faces severe challenges in practical application. In particular, its target detection and positioning performance is significantly reduced in complex environments (such as fog, rain, heavy fog and other adverse weather conditions), which seriously restricts the reliability and safety of drones.

[0003] Traditional drones rely heavily on vision-based target detection algorithms (such as YOLOv8) for obstacle identification and navigation. However, in low-visibility environments, issues such as image blur and noise interference significantly reduce the accuracy of the algorithms. For example, in foggy scenes, light scattering and occlusion make it difficult to extract target features. Traditional algorithms are prone to false detection or missed detection, which directly affects the drone's obstacle avoidance and positioning capabilities. In addition, while existing radar positioning technology can calculate target distance and speed using frequency-modulated continuous wave (FMCW) signals, it still has shortcomings in the high-precision calculation of three-dimensional spatial coordinates. In particular, the dynamic tracking accuracy of the target's azimuth and elevation is limited, making it difficult to meet the real-time positioning requirements in complex environments.

[0004] Current research attempts to improve performance by refining algorithms or fusing multi-sensor data, but significant limitations remain. For example, some approaches enhance feature extraction capabilities by increasing the depth of convolutional neural networks, but this increases computational complexity, making efficient deployment in embedded drones difficult. Other approaches, while combining lidar with vision fusion technology, improve positioning robustness but significantly increase hardware cost and system power consumption. Therefore, improving target detection accuracy and positioning efficiency in harsh environments while ensuring real-time performance and low cost remains a pressing technical challenge. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: target detection accuracy and positioning efficiency of UAVs in harsh environments.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a drone logistics distribution method, comprising the following steps:

[0007] S10, obtaining the three-dimensional spatial coordinates of the target through the FMCW radar carried by the UAV;

[0008] S20, using an improved YOLOv8 model with integrated feature attention fusion network to perform dehazing and enhancement processing on foggy images collected by drones, and identifying targets in foggy images through a joint weighted mechanism of channel attention and pixel attention;

[0009] S30, the flight control system of the UAV drives the UAV to move toward the target according to the three-dimensional coordinates of the target, and dynamically corrects the coordinates during the movement;

[0010] S40. When the distance threshold between the drone and the target is within a preset range, the drone hovers and performs secondary identity verification on the target;

[0011] S50. After the secondary identity verification is passed, the drone is controlled to land at the target location to complete the delivery mission.

[0012] Furthermore, step S10 specifically includes:

[0013] S11, transmitting a linear frequency modulation signal through the FMCW radar carried by the UAV and receiving the target reflection signal;

[0014] S12. Generate a beat signal based on the frequency mixing process, calculate the three-dimensional distance, speed and azimuth of the target through two-dimensional Fourier transform, and obtain the three-dimensional spatial coordinates of the target.

[0015] Furthermore, step S12 specifically includes:

[0016] S121, calculating the three-dimensional distance of the target by extracting the range frequency based on a fast-time two-dimensional Fourier transform;

[0017] S122, calculating the target speed by extracting the Doppler frequency based on a slow-time two-dimensional Fourier transform;

[0018] S123. Estimate the target horizontal azimuth and vertical elevation angle using the array antenna phase difference, and decompose the three-dimensional coordinates into X, Y, and Z axes.

[0019] Furthermore, step S20 specifically includes:

[0020] S21, performing shallow feature extraction on the input foggy image to generate a multi-channel feature map retaining the original resolution;

[0021] S22, deepening the multi-channel feature map through multiple groups of serially connected group structures, each group of serially connected group structures includes multiple basic blocks, each basic block adaptively weights the fog area features through a local residual connection and a feature attention module to obtain multiple groups of intermediate feature maps, wherein the feature attention module includes a channel attention submodule and a pixel attention submodule;

[0022] S23, perform attention feature fusion on multiple sets of intermediate feature maps;

[0023] S24. The feature map fused with the attention features is reconstructed through global residual to generate a defogging target feature map.

[0024] Furthermore, step S23 specifically includes:

[0025] S231, perform global average pooling of the channel dimension on the intermediate feature map to generate a channel weight matrix;

[0026] S232, performing convolution processing on the channel-weighted feature map in the spatial dimension to generate a pixel weight matrix;

[0027] S233. Multiply the channel weight matrix and the pixel weight matrix element by element to obtain a feature map of attention feature fusion.

[0028] Furthermore, in step S30, the coordinates are dynamically corrected by continuously receiving the three-dimensional spatial coordinates of the target updated by the FMCW radar during the movement of the UAV, and adjusting the flight path in real time through the ROS topic subscription mechanism.

[0029] Furthermore, in step S40, secondary identity verification includes face recognition and communication with the drone flight control system through the MAVROS toolkit to verify the legitimacy of the target identity.

[0030] The present invention also provides a UAV logistics distribution device, comprising:

[0031] The target position acquisition module is used to obtain the three-dimensional spatial coordinates of the target through the FMCW radar carried by the UAV;

[0032] The target image recognition module uses an improved YOLOv8 model with an integrated feature attention fusion network to perform dehazing and enhancement processing on foggy images collected by drones, and identifies targets in foggy images through a joint weighted mechanism of channel attention and pixel attention;

[0033] The flight control module is used in the UAV's flight control system to drive the UAV towards the target according to the target's three-dimensional coordinates and dynamically correct the coordinates during the movement;

[0034] The target verification module is used to hover the drone and perform secondary identity verification on the target when the distance threshold between the drone and the target is within a preset range;

[0035] The drone landing module is used to control the drone to land at the target location after the secondary identity verification is passed to complete the delivery mission.

[0036] The present invention also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the drone logistics distribution method described above.

[0037] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, can implement the drone logistics distribution method described above.

[0038] The beneficial effects of the present invention are: by improving the YOLOv8 model and integrating the feature attention fusion network (FFA-Net), combined with the three-dimensional positioning technology of the frequency-modulated continuous wave radar, high-precision target detection and dynamic path planning in foggy scenes are achieved, significantly improving the safety and efficiency of drone delivery in harsh environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The specific structure of the present invention is described in detail below with reference to the accompanying drawings.

[0040] Figure 1 This is a flow chart of the drone logistics distribution method according to an embodiment of the present invention;

[0041] Figure 2 A schematic diagram of radar angle positioning according to an embodiment of the present invention;

[0042] Figure 3 Schematic diagram of the FFA-NET basic blocks of an embodiment of the present invention;

[0043] Figure 4 A schematic diagram of an FFA-NET network according to an embodiment of the present invention;

[0044] Figure 5 This is a schematic diagram of the FFA-NET network embedded in YOLOv8 according to an embodiment of the present invention;

[0045] Figure 6 This is a schematic diagram of communication between an edge device and a drone according to an embodiment of the present invention;

[0046] Figure 7 This is a block diagram of a UAV logistics distribution device according to an embodiment of the present invention;

[0047] Figure 8 A schematic block diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0049] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0050] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0051] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0052] like Figure 1 As shown, an embodiment of the present invention is: a drone logistics distribution method, comprising the following steps:

[0053] S10. Obtain the three-dimensional spatial coordinates of the target through the FMCW radar carried by the UAV.

[0054] In a specific embodiment, step S10 specifically includes:

[0055] S11. The FMCW radar carried by the UAV transmits a linear frequency modulation signal and receives the target reflection signal.

[0056] In this step, the transmission and reception of FMCW radar signals:

[0057] Transmitted signal: The radar transmits a linear frequency modulated signal whose phase varies quadratically with time:

[0058]

[0059] Where f0 is the initial frequency, μ is the frequency modulation slope, and φ0 is the initial phase.

[0060] Received signal: target reflected signal after time delay After being received,

[0061] Its expression is:

[0062] Where A1 is the amplitude gain of the received signal, T is the time width (time window length) of the radar signal, and φ0 is the initial phase.

[0063] S12. Generate a beat signal based on the frequency mixing process, calculate the three-dimensional distance, speed and azimuth of the target through two-dimensional Fourier transform, and obtain the three-dimensional spatial coordinates of the target.

[0064] In this step, the mixing process is to mix the received signal with the local oscillator signal to generate a beat signal S b (t), whose frequency contains target distance and speed information:

[0065]

[0066] Where A2 is the amplitude gain of the radar local oscillator signal, and R0 is the initial distance of the target (the initial distance from the radar).

[0067] Distance-dependent frequencies:

[0068] Doppler frequency:

[0069] In a specific embodiment, step S12 specifically includes:

[0070] S121, calculating the three-dimensional distance of the target by extracting the range frequency based on a fast-time two-dimensional Fourier transform;

[0071] S122, calculating the target speed by extracting the Doppler frequency based on a slow-time two-dimensional Fourier transform;

[0072] In this step, the distance and speed are solved and two-dimensional FFT processing is performed on the beat signal: a two-dimensional Fourier transform is performed on the fast time (distance dimension) and slow time (speed dimension):

[0073] Fast time FFT: extract the range frequency f r , calculate the target distance:

[0074] Slow-time FFT: Extract the Doppler frequency f a , calculate the target speed:

[0075] S123. Estimate the target horizontal azimuth and vertical elevation angle using the array antenna phase difference, and decompose the three-dimensional coordinates into X, Y, and Z axes.

[0076] In this step, the azimuth angle (angle of arrival DOA) is estimated:

[0077] Array antenna phase difference: Utilizes the phase difference of the signals received by multiple antennas to determine the target azimuth through the angle of arrival estimation algorithm.

[0078] In interferometric direction-finding systems, based on the principle of spatial phase coherence, the target azimuth angle θ can be accurately calculated using the phase difference between the dual-channel received signals. Assuming the distance between the two receiving antennas along the baseline is d, when a plane wavefront reaches the antenna array at an incident angle θ, the electromagnetic wave propagation path will produce a geometric path difference ΔR = d·sinθ. Based on the phase propagation characteristics of the wave equation, this path difference is converted into a spatial phase difference between the received signals:

[0079]

[0080] Where λ is the radar operating wavelength. This relationship shows that the spatial phase difference φ and the incident azimuth θ exhibit a strict sinusoidal mapping relationship. After using a high-precision phase detection circuit to measure the instantaneous phase difference φ between the two channels, the inversion operation is performed:

[0081]

[0082] The arrival angle information of the target can be reconstructed. Figure 2 shown.

[0083]

[0084] Similarly, with the support of a two-dimensional or three-dimensional array, the horizontal azimuth and vertical elevation angle of the target can be measured simultaneously.

[0085] After obtaining the horizontal azimuth and vertical elevation, decompose the straight-line distance D into three coordinate axes:

[0086] Where D is the straight-line distance to the target (i.e. the distance R measured by the radar).

[0087] 1. Horizontal projection (XY plane)

[0088] First calculate the projected length L of the target in the horizontal plane (ignoring the height):

[0089] L=D·cosθ v (11)

[0090] Then decompose L into X and Y axes:

[0091] x=L·cosθ h =D·cosθ v cosθ h (12)

[0092] Y=L·sinθ h =D·cosθ v sinθ h(13)

[0093] Horizontal azimuth angle θ h Determines the ratio of X and Y.

[0094] 2. Vertical projection (Z axis)

[0095] Directly calculate the height Z:

[0096] Z=D·sinθ v (14)

[0097] The final three-dimensional coordinates are:

[0098] The three-dimensional coordinates are transmitted to the flight control, which will control the drone to identify the target at each coordinate. At this time, the aircraft will judge each target in turn until the target user is found.

[0099] This embodiment uses frequency-modulated continuous wave radar combined with array antenna phase difference technology to calculate target range and velocity using fast / slow time two-dimensional FFTs, and uses antenna phase difference to estimate azimuth and elevation angles. This overcomes the limitations of traditional monocular vision ranging in foggy conditions, enabling dynamic updating of three-dimensional coordinates and minimizing positioning errors, thus resolving the positioning challenge in complex weather conditions where GPS is denied.

[0100] S20. An improved YOLOv8 model with an integrated feature attention fusion network is used to perform dehazing and enhancement processing on foggy images collected by drones, and targets in foggy images are identified through a joint weighted mechanism of channel attention and pixel attention.

[0101] In this step, the YOLOv8 of FFA-NET is introduced to recognize the collected images and identify the objects with specific postures in the images. Improve the YOLOv8 model training execution process:

[0102] Input data: paired fog image-clear image dataset

[0103] Data augmentation: random 90°, 180°, 270° rotation, horizontal flip (50% probability), and random cropping of 240×240 pixel blocks.

[0104] Forward propagation, perform the complete calculation according to the above model structure to obtain the predicted output J pred .

[0105] Optimization strategy: loss calculation, using L1 norm loss function:

[0106]

[0107] Learning rate adjustment: cosine annealing strategy:

[0108] Where T is the total number of training cycles (epochs).

[0109] During the posture detection model deployment training process, the model is trained using the RTTS dataset. The trained model is deployed on the designated chip of the drone. After training, the model is evaluated to determine its performance on the test dataset. Common evaluation metrics include accuracy, recall, and F1 score. Based on the evaluation results, the model is further optimized by adjusting the network structure and increasing the training data.

[0110] In a specific embodiment, step S20 specifically includes:

[0111] S21, performing shallow feature extraction on the input foggy image to generate a multi-channel feature map retaining the original resolution;

[0112] In this step, the shallow feature extraction input (RGB three-channel fog image I haze ∈R H×W×3 ) This operation extracts shallow features F0∈R through a 3×3 convolutional layer H×W×64 To achieve:

[0113] Formula: F0=Conv 3x3 (I haze ) (18)

[0114] Finally, a 64-channel feature map is output, retaining the original resolution.

[0115] S22, deepening the multi-channel feature map through multiple groups of serially connected group structures, each group of serially connected group structures includes multiple basic blocks, each basic block adaptively weights the fog area features through a local residual connection and a feature attention module to obtain multiple groups of intermediate feature maps, wherein the feature attention module includes a channel attention submodule and a pixel attention submodule;

[0116] In this step, three group architectures are set up in series, and each group contains 19 basic blocks.

[0117] like Figure 3 As shown in Figure 3, the basic block structure consists of local residual learning and feature attention (FA) modules. Local residual learning allows bypassing less important information (such as haze or low-frequency areas) through multiple local residual connections, allowing the main network to focus on valid information.

[0118] like Figure 4As shown in Figure 1, the group structure combines 19 basic block structures and skip connection modules. The 19 consecutive blocks increase the depth and expressiveness of FFA-Net. The skip connection enables FFA-Net to bypass training difficulties.

[0119] At the end of FFA-Net, a restoration part is added using a two-layer convolutional network implementation and a long shortcut global residual learning module. Finally, the desired dehazed image is restored.

[0120] The group architecture and global residual learning combine multiple basic block structures and skip connection modules. The successive basic blocks increase the depth and expressiveness of FFA-Net. The skip connections enable FFA-Net to circumvent training difficulties. At the end of FFA-Net, a restoration component is added using a two-layer convolutional network implementation and a long-shortcut global residual learning module. Finally, the desired dehazed image is restored. The group architecture increases network depth and utilizes multiple basic blocks to gradually extract and optimize features.

[0121] S23, perform attention feature fusion on multiple sets of intermediate feature maps;

[0122] In a specific embodiment, step S23 specifically includes:

[0123] S231, perform global average pooling of the channel dimension on the intermediate feature map to generate a channel weight matrix;

[0124] In this step, the basic blocks are iteratively processed (double residual connections are used in the basic blocks of FFA-Net to allow the hazy areas (low-frequency information) to bypass the main path directly).

[0125] Each group contains B basic blocks (Basic Block, B = 19), and the i-th basic block is processed as follows:

[0126] Local residual connection: input feature F in Generate the intermediate feature F through two 3×3 convolutional layers (in Block, conv1 and conv2 perform feature transformation and are the core of actual processing) mid :

[0127] F mid =Conv 3x3 (ReLU(Conv 3x3 (F in ))) (19)

[0128] Feature attention weighting: for F mid Perform channel attention and pixel attention calculations.

[0129] Channel Attention (CA) obtains channel-level statistical information through global average pooling and generates channel weights to adjust the importance of different channels.

[0130] Channel attention weight calculation: Generate channel weights through global average pooling and two layers of 1×1 convolution:

[0131] CA=σ(Conv 1x1 (δ(Conv 1x1 (GAP(F mid )))) (20)

[0132] Where σ is the Sigmoid activation function, δ is the ReLU activation function, and GAP is the global average pooling

[0133] S232, performing convolution processing on the channel-weighted feature map in the spatial dimension to generate a pixel weight matrix;

[0134] In this step, pixel attention (PA) generates weights for each pixel in the spatial dimension, emphasizing important areas such as thick fog. Pixel attention weight calculation: Spatial weight calculation is performed on the weighted features:

[0135] PA=σ(Conv 3x3 (δ(Conv 3x3 (F mid ⊙CA)))) (21)

[0136] Final feature output:

[0137] F out =(F mid ⊙CA)⊙PA (22)

[0138] Local residual connection: F final =F in +F out (twenty three)

[0139] Inter-group feature transfer, each group outputs feature G i ∈R HxWx64 Passed to subsequent processing via skip connections.

[0140] S233. Multiply the channel weight matrix and the pixel weight matrix element by element to obtain a feature map of attention feature fusion.

[0141] In this step, attention feature fusion (FFA) is performed, including:

[0142] Feature splicing: splicing the output features of 3 group structures:

[0143] F concat=Concat(G1,G2,G3)∈R HxWx192 (twenty four)

[0144] Adaptive weighted fusion: Generate channel-pixel joint weight matrix W through feature attention module FFA .

[0145] Weighted fusion formula: Generate fusion weight WFFA through the feature attention module and perform adaptive weighting:

[0146] F fused =F concat ⊙W FFA (25)

[0147] S24. The feature map fused with the attention features is reconstructed through global residual to generate a defogging target feature map.

[0148] In this step, global residual reconstruction: Feature reconstruction: Generate a residual map through two 3×3 convolutional layers;

[0149] R=Conv 3x3 (ReLU(Conv 3x3 (F fused ))) (26)

[0150] Final output: Superimpose the global residual to get the dehazed image:

[0151] J=I haze +R (27)

[0152] Among them I haze is the input fog image, R is the residual image, and J is the final defogging image

[0153] like Figure 5 The figure shows the schematic diagram of embedding the FFA-NET network into yolov8.

[0154] In this embodiment, channel attention (CA) and pixel attention (PA) modules are introduced after shallow feature extraction, achieving adaptive feature weighting through global average pooling and spatial convolution. Group structure concatenation and global residual connections are used to preserve multi-scale feature information. This significantly enhances structured objects. The computational complexity is maintained at a high level, taking into account the real-time requirements of embedded device deployment.

[0155] S30. The flight control system of the UAV drives the UAV to move toward the target according to the three-dimensional coordinates of the target, and dynamically corrects the coordinates during the movement.

[0156] In a specific embodiment, in step S30, the coordinates are dynamically corrected by continuously receiving the three-dimensional spatial coordinates of the target updated by the FMCW radar during the movement of the UAV, and adjusting the flight path in real time through the ROS topic subscription mechanism.

[0157] In this step, the three-dimensional coordinates of the target are input into the UAV flight control system, and the UAV is driven to move toward the target through the control instructions encapsulated by ROS, and the coordinates are updated in real time during the movement.

[0158] After locating the target, the drone moves towards it and encapsulates the deployed algorithm model in the ROS system. In actual scenarios, when the drone successfully locates the target using pinhole imaging and yolov8 visual recognition, it transmits the target's three-dimensional coordinates (x, y, z) to the flight control board through the MAVROS toolkit in the form of a topic to control the drone's next action. Figure 6 The figure shows a schematic diagram of communication between the edge end and the drone.

[0159] In this implementation, radar coordinate updates are acquired in real time through a ROS topic subscription mechanism, and a PID control algorithm is used to generate flight control commands. This shortens the response time for path corrections, significantly improving timeliness and avoiding coordinate offsets caused by target movement.

[0160] S40: When the distance threshold between the drone and the target is within a preset range, the drone hovers and performs secondary identity authentication on the target.

[0161] In a specific embodiment, in step S40, the secondary identity verification includes face recognition and communication with the drone flight control system through the MAVROS toolkit to verify the legitimacy of the target identity.

[0162] In this step, when the distance threshold between the drone and the target is ≤10 meters, the drone hovers and starts the FFA-Net-based visual recognition module to perform secondary target verification.

[0163] When the distance between the drone and the target is x≤10m, y≤10m, and z≤10m, the drone hovers and performs target face recognition. When the target is confirmed to be the recipient, the drone begins to land and locks onto the target.

[0164] In this example, the FFA-Net-based facial recognition module is activated during the hover phase and verified via the MAVROS toolkit in conjunction with the flight control system. This significantly reduces the false recognition rate, and combined with radar and visual verification, ensures the uniqueness of the delivery target and prevents the risk of misdelivery.

[0165] S50. After the secondary identity verification is passed, the drone is controlled to land at the target location to complete the delivery mission.

[0166] In this embodiment, FMCW radar complements an improved visual model, significantly improving target detection accuracy in foggy scenarios. FFA-Net significantly improves accuracy while increasing the number of parameters by only 1.7%. Combined with the lightweight communication architecture of ROS, the system can run on edge devices such as the Jetson Nano, significantly reducing hardware costs. This complete technology chain, from coarse positioning using radar to precise recognition, dynamic tracking, and identity verification using visual models, significantly improves delivery success rates in complex weather conditions.

[0167] The training experimental results of the improved YOLOv8 model are as follows:

[0168] Comparative ablation experiments: To validate the model's defogging capabilities across multiple dimensions, the dataset selection included several additional target categories, including but not limited to buses, cars, and motorcycles, in addition to pedestrians. This allowed for a more multi-dimensional analysis of the model's performance. This experiment used the publicly available foggy RTTS dataset, selecting 1,000 images for training, 100 for testing, and 100 for validation, for a total of 1,200 images. The dataset was trained using the Yolov8 model.

[0169] Table 1: Comparison of ablation parameters and computational effort

[0170]

[0171] As shown in Table 1 above, the parameters and computational complexity of the improved model are as follows: the improved model slightly increases the parameters by adding CA (Channel Attention) and PA (possibly Physical Augmentation or Position Attention) modules, but maintains a similar GFLOPs (8.1) to the original version, indicating better computational efficiency.

[0172] Inference speed: The inference time of the improved model (0.9-1.1ms) is basically the same as the original version (1.0ms), and the module design does not significantly increase the computational burden.

[0173] Table 2: Comparison of global ablation metrics

[0174]

[0175] As shown in Table 2 above, the four parameters are important indicators for measuring the quality of the model;

[0176] Precision: measures the accuracy of the detection results (avoiding false detections).

[0177] Recall: measures the coverage of detection (avoiding missed detections).

[0178] mAP50: The mean average precision calculated with an IoU threshold of 0.5, reflecting the comprehensive detection capability of the model under relaxed bounding box matching requirements.

[0179] mAP50-95: The mean average precision with IoU thresholds ranging from 0.5 to 0.95 (step size 0.05), which measures the robustness and generalization of the model at different levels of strictness.

[0180] The CA+PA combination is optimal: it outperforms the original version (0.503 / 0.47) in both mAP50 (0.53) and recall (0.48). Its precision (P = 0.673) is significantly higher than the original version (0.649), indicating that the model reduces false positives. Therefore, it offers the best overall performance, with a low false positive rate, making it suitable for object detection in foggy conditions.

[0181] The CA module improves precision: its mAP50 (0.51) is close to the original version, but its mAP50 for the "bus" class reaches 0.615, indicating that channel attention significantly improves structured objects (such as buses). The CA group alone achieves the highest P-value (0.74), but its recall decreases (0.447), likely due to the attention mechanism filtering out some ambiguous objects.

[0182] The PA module achieves stable generalization: its mAP50 (0.507) is comparable to the original version, but its mAP50 for the bicycle class improves to 0.306, demonstrating pixel attention's increased sensitivity to small objects. Only the PA group achieves a slight improvement in mAP50-95 (0.331), likely due to improved multi-scale detection robustness through physical augmentation.

[0183] Table 3: Comparative ablation performance analysis by category

[0184]

[0185] As shown in Table 3 above, the improved model performs best in four of the five target categories, fully demonstrating the effectiveness of the improved module. Pedestrians saw the most significant improvement across these five categories. The two modules also differ significantly in their applicability. Channel Attention (CA) is more suitable for processing objects with high color contrast and clear structure (such as buses). Pixel Attention (PA) is more sensitive to small objects (such as bicycles) and local details.

[0186] like Figure 7 As shown, an embodiment of the present invention further provides a drone logistics distribution device, comprising:

[0187] The target position acquisition module 10 is used to obtain the three-dimensional spatial coordinates of the target through the FMCW radar carried by the UAV;

[0188] The target image recognition module 20 is used to perform defogging and enhancement processing on foggy images collected by the UAV using an improved YOLOv8 model with an integrated feature attention fusion network, and to identify targets in foggy images through a joint weighted mechanism of channel attention and pixel attention;

[0189] The flight control module 30 is used for the flight control system of the UAV to drive the UAV to move towards the target according to the three-dimensional coordinates of the target and dynamically correct the coordinates during the movement;

[0190] The target verification module 40 is used to cause the drone to hover and perform secondary identity verification on the target when the distance threshold between the drone and the target is within a preset range;

[0191] The drone landing module 50 is used to control the drone to land at the target location after the secondary identity verification is passed to complete the delivery task.

[0192] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned drone logistics distribution device can refer to the corresponding description in the aforementioned method embodiment. For the convenience and conciseness of the description, it will not be repeated here.

[0193] The above-mentioned UAV logistics distribution device can be implemented in the form of a computer program, which can be used in Figure 8 Runs on the computer device shown.

[0194] See also Figure 8 , Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 can be a terminal or a server. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.

[0195] See Figure 8 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0196] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can cause the processor 502 to execute a drone logistics delivery method.

[0197] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0198] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a drone logistics distribution method.

[0199] The network interface 505 is used to communicate with other devices through the network. Figure 8 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0200] The processor 502 is configured to execute a computer program 5032 stored in a memory to implement the drone logistics distribution method described above.

[0201] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0202] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0203] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to execute the drone logistics delivery method described above.

[0204] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0205] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0206] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0207] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0208] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0209] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A drone logistics distribution method, characterized in that: The following steps are involved: S10, obtaining the three-dimensional spatial coordinates of the target through the FMCW radar carried by the UAV; S20, using an improved YOLOv8 model with integrated feature attention fusion network to perform dehazing and enhancement processing on foggy images collected by drones, and identifying targets in foggy images through a joint weighted mechanism of channel attention and pixel attention; S30, the flight control system of the UAV drives the UAV to move toward the target according to the three-dimensional coordinates of the target, and dynamically corrects the coordinates during the movement; S40. When the distance threshold between the drone and the target is within a preset range, the drone hovers and performs secondary identity verification on the target; S50. After the secondary identity verification is passed, the drone is controlled to land at the target location to complete the delivery mission.

2. The drone logistics distribution method according to claim 1, characterized in that: Step S10 specifically includes: S11, transmitting a linear frequency modulation signal through the FMCW radar carried by the UAV and receiving the target reflection signal; S12. Generate a beat signal based on the frequency mixing process, calculate the three-dimensional distance, speed and azimuth of the target through two-dimensional Fourier transform, and obtain the three-dimensional spatial coordinates of the target.

3. The drone logistics distribution method according to claim 2, characterized in that: Step S12 specifically includes: S121, calculating the three-dimensional distance of the target by extracting the range frequency based on a fast-time two-dimensional Fourier transform; S122, calculating the target speed by extracting the Doppler frequency based on a slow-time two-dimensional Fourier transform; S123. Estimate the target horizontal azimuth and vertical elevation angle using the array antenna phase difference, and decompose the three-dimensional coordinates into X, Y, and Z axes.

4. The drone logistics distribution method according to claim 1, characterized in that: Step S20 specifically includes: S21, performing shallow feature extraction on the input foggy image to generate a multi-channel feature map retaining the original resolution; S22, deepening the multi-channel feature map through multiple groups of serially connected group structures, each group of serially connected group structures includes multiple basic blocks, each basic block adaptively weights the fog area features through a local residual connection and a feature attention module to obtain multiple groups of intermediate feature maps, wherein the feature attention module includes a channel attention submodule and a pixel attention submodule; S23, perform attention feature fusion on multiple sets of intermediate feature maps; S24. The feature map fused with the attention features is reconstructed through global residual to generate a defogging target feature map.

5. The drone logistics distribution method according to claim 4, characterized in that: Step S23 specifically includes: S231, perform global average pooling of the channel dimension on the intermediate feature map to generate a channel weight matrix; S232, performing convolution processing on the channel-weighted feature map in the spatial dimension to generate a pixel weight matrix; S233. Multiply the channel weight matrix and the pixel weight matrix element by element to obtain a feature map of attention feature fusion.

6. The drone logistics distribution method according to claim 1, characterized in that: In step S30, the coordinates are dynamically corrected by continuously receiving the three-dimensional spatial coordinates of the target updated by the FMCW radar during the movement of the UAV, and adjusting the flight path in real time through the ROS topic subscription mechanism.

7. The drone logistics distribution method according to claim 1, characterized in that: In step S40, the secondary identity verification includes face recognition and communication with the drone flight control system through the MAVROS toolkit to verify the legitimacy of the target identity.

8. A UAV logistics distribution device, characterized in that: include: The target position acquisition module is used to obtain the three-dimensional spatial coordinates of the target through the FMCW radar carried by the UAV; The target image recognition module uses an improved YOLOv8 model with an integrated feature attention fusion network to perform dehazing and enhancement processing on foggy images collected by drones, and identifies targets in foggy images through a joint weighted mechanism of channel attention and pixel attention; The flight control module is used in the UAV's flight control system to drive the UAV towards the target according to the target's three-dimensional coordinates and dynamically correct the coordinates during the movement; The target verification module is used to hover the drone and perform secondary identity verification on the target when the distance threshold between the drone and the target is within a preset range; The drone landing module is used to control the drone to land at the target location after the secondary identity verification is passed to complete the delivery mission.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the drone logistics distribution method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, can implement the drone logistics distribution method according to any one of claims 1 to 7.