An unmanned aerial vehicle logistics distribution method based on multi-modal visual perception
Patent Information
- Application Number
- CN202611079776.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-21
AI Technical Summary
然而,单一传感器易受光照变化、遮挡干扰,难以全面捕捉环境信息,CNN架构难以有效融合多模态数据,语义理解能力受限,且简单的特征拼接或加权平均忽略了模态间的非线性关联,静态路径规划无法适应动态障碍物,主流DRL方法需要大量仿真训练,迁移到真实场景时易出现安全性问题
本申请提供了一种基于多模态视觉感知的无人机物流配送方法,首先获取目标区域的环境数据,环境数据包括RGB图像、深度数据、热成像图像以及激光雷达点云数据,将热成像图像引入无人机物流配送方法中,通过温度特征增强了对金属货架以及电子设备的识别能力;基于环境感知与路径规划模型,实时生成避障搬运路径;环境感知与路径规划模型包括时空对齐模块、Transformer特征编码器、轻量化蒸馏模块、多模态特征融合模块、智能分析模块和航路规划模块;时空对齐模块实现了多源的环境数据的硬件级同步;Transformer特征编码器分别对对齐后的环境特征进行特征提取,提高了特征多样性的同时提高了Transformer特征编码器的鲁棒性;轻量化蒸馏模块将预训练的大语言模型通过知识蒸馏压缩为轻量化的学生模型,保留语义理解能力的同时减少了参数计算量;多模态特征融合模块基于跨模态注意力机制进行加权融合,增强了关键信息对统一图谱的影响权重,抑制了噪声影响;智能分析模块基于图神经网络和时空Transformer进行多尺度特征聚合,图神经网络的设置使得本申请能够实时反映动态障碍物的位置变化情况,提高了障碍物的检测精度;航路规划模块基于改进的RRT*算法和深度强化学习方法,实时确定无人机的避障搬运路径,改进的RRT*算法和深度强化学习方法相结合,能够在动态环境中实现毫秒级响应,提高无人机的避障能力。综上,本申请能够实时确定无人机的避障搬运路径并进行物流配送,提高无人机在复杂低空场景下的感知能力,同时提高目标检测精度和航路规划效率。
Smart Images

Figure CN122596811B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal perception technology, and in particular to a drone logistics delivery method based on multimodal visual perception. Background Technology
[0002] Urban logistics drones (such as early models of Amazon Prime Air) rely on GPS signals for navigation in the "urban canyons" formed by dense high-rise buildings. Multipath reflections and blockages of GPS signals result in positioning accuracy errors exceeding 5 meters, making it difficult for drones to land safely and accurately on designated helipads or balconies, increasing the mission failure rate in complex urban areas. Furthermore, the onboard visual obstacle avoidance systems of power line inspection or long-distance delivery drones struggle to identify small, low-contrast obstacles. At dusk, dawn, or against complex backgrounds (such as forests or buildings), the rate of missed detection of critical hazards such as power lines and guy wires is high, leading to mission failures and aircraft damage.
[0003] In related methods, a single sensor is typically used to acquire environmental data of the target area. A CNN architecture is then employed to extract features from the data, followed by data concatenation or weighted averaging. Static path planning algorithms such as A* are then used to plan the delivery path for the drone. However, single sensors are susceptible to changes in lighting and occlusion interference, making it difficult to comprehensively capture environmental information. CNN architectures struggle to effectively integrate multimodal data, limiting their semantic understanding capabilities. Furthermore, simple feature concatenation or weighted averaging ignores the nonlinear relationships between modalities. Static path planning cannot adapt to dynamic obstacles, and mainstream DRL methods require extensive simulation training, which can lead to safety issues when transferred to real-world scenarios.
[0004] Therefore, based on the above problems, there is an urgent need to provide a drone logistics delivery method based on multimodal visual perception, which can improve the perception capability of drones in complex low-altitude scenarios, improve target detection accuracy, and improve route planning efficiency. Summary of the Invention
[0005] The purpose of this application is to provide a drone logistics delivery method based on multimodal visual perception. This application can improve the perception capability of drones in complex low-altitude scenarios, improve target detection accuracy, and improve route planning efficiency.
[0006] To achieve the above objectives, this application provides the following solution: This application provides a drone logistics delivery method based on multimodal visual perception, including: Acquire environmental data of the target area; the target area includes multiple dynamic and static obstacles; the environmental data includes: RGB images, depth data, thermal imaging images, and LiDAR point cloud data; Based on the environmental data, an obstacle avoidance and transport path is generated in real time using an environmental perception and path planning model. The environmental perception and path planning model includes: a spatiotemporal alignment module, a Transformer feature encoder, a lightweight distillation module, a multimodal feature fusion module, an intelligent analysis module, and a route planning module. The spatiotemporal alignment module performs sequential temporal and spatiotemporal alignment on the environmental data to obtain aligned environmental data. The Transformer feature encoder extracts features from the aligned environmental data to obtain multimodal features. The multimodal features include: RGB features, depth map features, and thermodynamic features. The lightweight distillation module utilizes pre-trained... The large language model is used as the teacher model, and a lightweight student model is trained through knowledge distillation. Semantic features are obtained based on multimodal features. The multimodal feature fusion module is used to perform weighted fusion based on the semantic features and the multimodal features using a cross-modal attention mechanism to obtain a unified graph. The intelligent analysis module is used to perform multi-scale feature aggregation based on graph neural networks and spatiotemporal Transformers based on the unified graph to obtain an aggregated result including object detection results and dynamic obstacle future state prediction results. The route planning module is used to determine the obstacle avoidance and transport path of the UAV in real time based on the aggregated result, using an improved RRT* algorithm and deep reinforcement learning methods. The drone is controlled to perform logistics delivery based on real-time obstacle avoidance and handling paths.
[0007] Optionally, environmental data of the target area can be acquired using a sensor array; the sensor array includes: an RGB camera, a time-of-flight depth sensor, a thermal imager, and a lidar. The RGB camera is used to acquire RGB images of the target area; The time-of-flight depth sensor is used to acquire depth data of the target area; The thermal imager is used to acquire thermal images of the target area; The lidar is used to acquire lidar point cloud data of the target area.
[0008] Optionally, the processing procedure of the spatiotemporal alignment module includes: Timing alignment of RGB images, depth data, thermal imaging images, and LiDAR point cloud data is performed using pulse synchronization signals and field-programmable gate array logic circuits. The efficient perspective n-point algorithm is used to spatially align the temporally aligned RGB image, temporally aligned depth data, temporally aligned thermal imaging image, and temporally aligned LiDAR point cloud data, resulting in aligned RGB image, aligned depth data, aligned thermal imaging image, and aligned LiDAR point cloud data.
[0009] Optionally, the Transformer feature encoder includes independent RGB branches, depth and LiDAR branches, and thermal imaging branches; The RGB branch is used to extract global semantic features of the aligned RGB image using a visual Transformer to obtain RGB features; The depth and lidar branches are used to extract multi-scale spatial features of aligned depth data and aligned lidar point cloud data using PointNet++ to obtain depth map features; The thermal imaging branch is used to extract the temperature distribution and temporal changes of the aligned thermal imaging image using a three-dimensional convolutional neural network, thereby obtaining thermodynamic features.
[0010] Optionally, the lightweight distillation module specifically includes: Using formula Knowledge distillation is performed on pre-trained large language models; in, Distillation of knowledge For feature channel index, The Kullback-Leibler divergence function, To reduce the feature distribution output by the lightweight student model, This refers to semantic prior knowledge.
[0011] Optionally, the processing procedure of the multimodal feature fusion module includes: Determine the internal interaction weights of semantic features, RGB features, depth map features, and thermodynamic features respectively; Based on the internal interaction weights and using the cross-attention mechanism, the formula is applied. Determine the cross-modal interaction weights; where, For cross-modal interaction weights, for function, This is a query vector generated based on RGB features. A key vector generated based on semantic features. The key vector is generated based on depth map features. The bond vector is generated based on thermodynamic characteristics. The dimension of the key vector; Based on cross-modal interaction weights, using the formula A unified map is obtained by weighted fusion of semantic features, RGB features, depth map features, and thermodynamic features; among them, To unify the map, The spatial location index of the feature represents the first... Okay, number Feature points of the column A value vector generated based on semantic features. A value vector generated based on depth map features. This is a value vector generated based on thermodynamic characteristics.
[0012] Optionally, the processing procedure of the intelligent analysis module includes: Based on the unified graph, an environmental topology graph is constructed using a graph neural network to obtain object detection results; Based on the unified map and environmental topology map, using the formula The positional changes of dynamic obstacles in the target area are predicted to obtain the predicted future state of the dynamic obstacles; among which, This is the prediction result for the future state of dynamic obstacles. For spacetime Transformer, To unify the map, This is an environmental topology diagram.
[0013] Optionally, the processing procedure of the route planning module includes: The search radius is adjusted using the improved RRT* algorithm to determine the search range; Based on the prediction results and the search range, a deep reinforcement learning method is used to determine the transport path of the UAV.
[0014] Optionally, adjusting the search radius using the improved RRT* algorithm to determine the search range specifically includes: Using formula Define the search scope; in, For the search radius, are constant parameters. This is the third hyperparameter. Let cross-entropy be the loss function. This is the prediction result for the future state of dynamic obstacles. To unify the map.
[0015] Optionally, determining the transport path of the UAV using deep reinforcement learning methods based on the prediction results and the search range specifically includes: Using formula Determine the transport route for the drone; in, Let be the objective function. For the search radius, As the first hyperparameter, This is the second hyperparameter. This is the penalty value for a drone colliding with an obstacle.
[0016] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a drone logistics delivery method based on multimodal visual perception. First, environmental data of the target area is acquired, including RGB images, depth data, thermal imaging images, and LiDAR point cloud data. Thermal imaging images are introduced into the drone logistics delivery method, enhancing the recognition capability of metal shelves and electronic devices through temperature features. Based on an environmental perception and path planning model, obstacle avoidance and handling paths are generated in real time. The environmental perception and path planning model includes a spatiotemporal alignment module, a Transformer feature encoder, a lightweight distillation module, a multimodal feature fusion module, an intelligent analysis module, and a route planning module. The spatiotemporal alignment module achieves hardware-level synchronization of multi-source environmental data. The Transformer feature encoder extracts features from the aligned environmental features, improving feature diversity and enhancing the performance of the Transformer. The robustness of the RMER feature encoder is enhanced; the lightweight distillation module compresses the pre-trained large language model into a lightweight student model through knowledge distillation, preserving semantic understanding capabilities while reducing parameter computation; the multimodal feature fusion module performs weighted fusion based on a cross-modal attention mechanism, enhancing the influence weight of key information on the unified graph and suppressing noise; the intelligent analysis module performs multi-scale feature aggregation based on graph neural networks and spatiotemporal Transformers. The graph neural network configuration enables the application to reflect the positional changes of dynamic obstacles in real time, improving obstacle detection accuracy; the route planning module, based on an improved RRT* algorithm and deep reinforcement learning methods, determines the obstacle avoidance and transport path of the UAV in real time. The combination of the improved RRT* algorithm and deep reinforcement learning methods enables millisecond-level response in dynamic environments, improving the obstacle avoidance capability of the UAV. In summary, this application can determine the obstacle avoidance and transport path of the UAV in real time and perform logistics delivery, improving the perception capability of the UAV in complex low-altitude scenarios, while improving target detection accuracy and route planning efficiency. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a drone logistics delivery method based on multimodal visual perception in one embodiment of this application. Figure 2 This is a flowchart illustrating a drone logistics delivery method based on multimodal visual perception in one embodiment of this application. Figure 3 This is a schematic diagram of the processing procedure of the spatiotemporal alignment module in one embodiment of this application; Figure 4 This is a schematic diagram of the processing procedure of the Transformer feature encoder in one embodiment of this application; Figure 5 This is a schematic diagram of the processing procedure of the lightweight distillation module in one embodiment of this application; Figure 6 This is a schematic diagram of the processing procedure of the multimodal feature fusion module in one embodiment of this application; Figure 7 This is a schematic diagram of the processing procedure of the intelligent analysis module in one embodiment of this application; Figure 8 This is a schematic diagram of the processing procedure of the route planning module in one embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] In one exemplary embodiment, such as Figure 1 and Figure 2 As shown, a drone logistics delivery method based on multimodal visual perception is provided, including the following S1 to S3. Wherein: S1: Obtain environmental data for the target area.
[0022] The target area includes multiple dynamic and static obstacles; the environmental data includes RGB images, depth data, thermal imaging images, and LiDAR point cloud data.
[0023] Specifically, this application utilizes a sensor array to acquire environmental data of the target area. The sensor array includes an RGB camera, a time-of-flight depth sensor (ToF sensor), a thermal imager, and a LiDAR (Light Detection and Ranging) system. The time-of-flight depth sensor is also known as a ToF depth sensor. The RGB camera is used to acquire RGB images of the target area; the time-of-flight depth sensor is used to acquire depth data of the target area; the thermal imager is used to acquire thermal images of the target area; and the LiDAR is used to acquire LiDAR point cloud data of the target area.
[0024] In one exemplary embodiment, the RGB camera employs a Sony IMX585 sensor with a resolution of 3840×2160@60fps, supporting High Dynamic Range Imaging (HDR) with a dynamic range of 120dB, ensuring the capture of detailed images and the generation of RGB images in both bright and low-light environments. The ToF depth sensor has a measurement accuracy of ±1mm and can generate millimeter-level precision depth data, i.e., depth maps, in real time for accurate spatial positioning. The thermal imager integrates a FLIR Boson 640 with a thermal sensitivity of 50mK, capable of identifying targets with significant thermal characteristics, such as metal shelves and electronic devices, by detecting infrared radiation from object surfaces, and generating thermal images to compensate for the shortcomings of optical sensors in smoke and low-light environments. The LiDAR uses a Velodyne VLP-16 capable of 16-line scanning at a scanning frequency of 20Hz, thereby generating 360° horizontal field-of-view LiDAR point cloud data, i.e., 3D point clouds, for constructing high-precision geometric models of the environment.
[0025] In addition, after acquiring environmental data, the data was preprocessed. Specifically, the RGB images underwent Contrast Limited Adaptive Histogram Equalization (CLAHE) to eliminate illumination interference; bilateral filtering and motion prediction were used to remove noise from the depth data, generating a dense point cloud (resolution 1024×768), which suppressed random noise while preserving edge details, improving the signal-to-noise ratio by 40%; the thermal imaging images were temperature normalized (range 0~100℃) and converted into pseudo-color images; and the LiDAR point cloud data was voxelized (voxel size 0.1m). 3 ), and extract surface normal features.
[0026] S2: Based on environmental data and an environmental perception and path planning model, generate obstacle avoidance and transport paths in real time.
[0027] The environmental perception and path planning model includes a spatiotemporal alignment module, a Transformer feature encoder, a lightweight distillation module, a multimodal feature fusion module, an intelligent analysis module, and a route planning module.
[0028] The spatiotemporal alignment module is used to sequentially align environmental data and then perform spatiotemporal alignment to obtain aligned environmental data, such as... Figure 3 As shown.
[0029] The spatiotemporal alignment module's processing includes hardware alignment and software alignment. Hardware alignment is timing alignment, while software alignment is spatiotemporal alignment.
[0030] Specifically, pulse synchronization signals (Pulse Per Second, PPS) and field-programmable gate array (FPGA) logic circuits are used to perform time-series alignment of RGB images, depth data, thermal imaging images, and LiDAR point cloud data; thereby achieving microsecond-level time-series alignment (time deviation <2ms) between environmental data, avoiding perception errors caused by sampling time differences, and realizing hardware-level synchronization.
[0031] The Efficient Perspective-n-Point (EPnP) algorithm is used to spatially align the temporally aligned RGB image, temporally aligned depth data, temporally aligned thermal image, and temporally aligned LiDAR point cloud data, resulting in aligned RGB images, aligned depth data, aligned thermal images, and aligned LiDAR point cloud data. Spatial alignment involves aligning the coordinate systems of the four data sets, ensuring that the spatial coordinates of the temporally aligned RGB image, temporally aligned depth data, temporally aligned thermal image, and temporally aligned LiDAR point cloud data are unified, with a reprojection error ≤ 0.5 pixels.
[0032] The introduction of thermal imaging and LiDAR point cloud data has solved the problem of traditional vision systems failing in extreme lighting and smoky environments. Thermal imagers can penetrate smoke to detect high-temperature targets (such as faulty motor equipment), while LiDAR can accurately model complex shelving structures.
[0033] The Transformer feature encoder is used to extract features from aligned environmental data to obtain multimodal features.
[0034] Specifically, such as Figure 4As shown, the Transformer feature encoder includes independent RGB, depth and LiDAR, and thermal imaging branches. The RGB branch uses a Vision Transformer (ViT) to extract global semantic features from the aligned RGB image, obtaining RGB features. The depth and LiDAR branch uses PointNet++ to extract multi-scale spatial features from the aligned depth data and aligned LiDAR point cloud data, obtaining depth map features. The thermal imaging branch uses a 3D Convolutional Neural Network (3D-CNN) to extract the temperature distribution and temporal changes of the aligned thermal imaging image, obtaining thermodynamic features. RGB features, depth map features, and thermodynamic features are collectively referred to as multimodal features. .
[0035] In the RGB branch, the input RGB image (size 1024×768) is segmented into 16×16 non-overlapping blocks (4096 patches in total). Each patch is transformed into a 768-dimensional vector through a linear projection matrix, and RGB features are generated through a multi-head self-attention (MHSA) mechanism. The MHSA mechanism uses a 12-head attention mechanism to calculate the global context association weights. The calculation formula for the MHSA mechanism is as follows: ; in, This is a multi-head self-attention mechanism. For querying the matrix, The key matrix, For value matrices, for function, For transpose, is the dimension of the key vector, used as a normalization factor to scale the gradient.
[0036] The input RGB image is independently encoded through the RGB branch to generate RGB features. RGB features can capture global information such as texture and shape of the target. The calculation formula for RGB features is as follows: ; in, It is an RGB feature. For visual Transformer, For the aligned RGB image, The height of the aligned RGB image. For the set of real numbers, 3 represents the width of the aligned RGB image, and 3 represents the number of channels in the aligned RGB image.
[0037] In the depth and LiDAR branches, the Farthest Point Sampling (FPS) method is used to select 1024 key points from the aligned LiDAR point cloud data (approximately 100,000 points). This reduces computational cost while preserving geometric structure. Subsequently, multi-scale spatial feature extraction is performed on the aligned depth data and aligned LiDAR point cloud data based on the PointNet++ method. Detail features are gradually recovered through deconvolution layers to obtain depth map features. The calculation formula for the depth map features is as follows: ; in, For depth map features, For PointNet++ methods, This is a joint dataset that combines aligned LiDAR point cloud data and aligned depth data. This represents the number of key points in the point cloud.
[0038] In the thermal imaging branch, a 3D convolutional neural network with a 7×7×3 kernel is used to capture temperature distribution and temporal changes, obtaining thermodynamic features. The calculation formula for the thermodynamic features is as follows: ; in, Thermodynamic characteristics, It is a three-dimensional convolutional neural network. For the aligned thermal imaging image, The time series length, i.e., the number of frames in the input video stream, is used to capture temporal changes. The height of the aligned thermal imaging image. The width of the aligned thermal image.
[0039] This application independently encodes aligned environmental data using a Transformer feature encoder to obtain learnable multimodal-specific vector information (multimodal features). The hierarchical encoding architecture allows for the design of dedicated encoders for different modal characteristics. For example, ViT excels at capturing global semantics, PointNet++ optimizes geometric feature extraction, and 3D-CNN adapts to temporal hot data, improving the specificity of feature representation. Through a multi-head self-attention mechanism, the ViT in this application reduces the computational cost by 60% compared to the traditional standard ViT. The lightweight design enables this application to be adapted for embedded deployment.
[0040] The lightweight distillation module is used to utilize a pre-trained large language model as a teacher model, train a lightweight student model through knowledge distillation, and obtain semantic features based on multimodal features.
[0041] Specifically, such as Figure 5 As shown, a pre-trained large language model (32 layers, 32 attention heads) is used as the teacher model to generate semantic prior knowledge. This paper utilizes knowledge distillation to compress the teacher model, transferring its semantic prior knowledge to a lightweight network with four layers and four attention heads—a lightweight student model. Different layers of this lightweight network share a partial projection matrix. Knowledge distillation is performed using KL divergence loss, calculated as follows: ; in, Distillation of knowledge For feature channel index, The Kullback-Leibler divergence function is used to measure... and The differences between them The feature distribution output by the lightweight student model, i.e., the initial semantic features. This refers to semantic prior knowledge, specifically the feature distribution output by the pre-trained large language model.
[0042] Specifically, the aligned multimodal features are input into the teacher model to obtain semantic prior knowledge; then, the aligned multimodal features are input into the lightweight student model to obtain initial semantic features; knowledge distillation is performed using KL divergence loss, and the lightweight student model is iteratively trained to make the initial semantic features as close as possible to the semantic prior knowledge; in each training epoch, the temperature is dynamically adjusted, and the attention alignment loss is calculated to force the lightweight student model to mimic the teacher's attention distribution. The formula for calculating the temperature parameter decay with training epochs is as follows: ; in, For temperature parameters, This is the initial maximum temperature value. This is the temperature decay factor, also known as the decay rate. , This serves as an index for training rounds.
[0043] Attention alignment loss The calculation formula is as follows: ; in, For indexing the number of network layers in a lightweight student model, For the network layer index of the teacher model, For lightweight student models in the first Attention matrix of layer, For the teacher model in the first Attention matrix of layer, for Norm, used to determine and The differences.
[0044] Multimodal features are input into the lightweight student model to obtain semantic features. An attention alignment loss is used to force the lightweight student model to mimic the teacher's feature attention pattern, resulting in a semantic understanding retention rate of ≥85%. Furthermore, while preserving semantic understanding, the lightweight design reduces the number of parameters in the lightweight student model to 10% of that in the teacher model (from approximately 700M parameters to 70M), a reduction of 90%. Sharing partial projection matrices across different layers further reduces memory usage. This enables the lightweight student model to achieve an inference latency of ≤15ms on the Jetson AGX platform and supports real-time semantic interaction, achieving real-time inference optimization.
[0045] The multimodal feature fusion module is used to perform weighted fusion based on semantic features and multimodal features using a cross-modal attention mechanism to obtain a unified map.
[0046] like Figure 6 As shown, the internal interaction weights of semantic features, RGB features, depth map features, and thermodynamic features are determined respectively; based on the internal interaction weights, cross-modal interaction weights are determined based on the cross-attention mechanism; based on the cross-modal interaction weights, the semantic features, RGB features, depth map features, and thermodynamic features are weighted and fused to obtain a unified map.
[0047] Specifically, determining the internal interaction weights can enhance the representation of salient regions, including shelf edges or high-temperature targets. For example, the RGB branch focuses on moving targets through a multi-head self-attention mechanism. Cross-modal interaction of semantic features, RGB features, depth map features, and thermodynamic features is achieved using a cross-attention mechanism. Specifically, the cross-modal interaction weights are determined using RGB features as the query vector and depth map features, thermodynamic features, and semantic features as the key and value vectors, respectively. The calculation formula is as follows: ; in, for function, This is a query vector generated based on RGB features. A key vector generated based on semantic features. The key vector is generated based on depth map features. The bond vector is generated based on thermodynamic characteristics. The dimension of the key vector is used to scale the dot product and prevent gradient vanishing.
[0048] Based on modal interaction weights, semantic features, RGB features, depth map features, and thermodynamic features are weighted and fused to obtain a unified map covering geometric, semantic, and thermodynamic information. The calculation process is as follows: ; in, To unify the map, The spatial location index of the feature represents the first... Okay, number Feature points of the column A value vector generated based on semantic features. A value vector generated based on depth map features. This is a value vector generated based on thermodynamic characteristics.
[0049] This application first determines the internal interaction weights to extract local salient features, and then determines the cross-modal interaction weights to achieve global feature alignment, thus avoiding the feature overload problem.
[0050] like Figure 7 As shown, the intelligent analysis module is used to perform multi-scale feature aggregation based on the unified graph, using graph neural networks (GNN) and spatiotemporal transformers, to obtain aggregated results including object detection results and prediction results of the future state of dynamic obstacles.
[0051] Based on the unified graph, an environmental topology graph is constructed using a graph neural network. Each detected object (dynamic and static obstacles) is treated as a node, and the corresponding feature vector (i.e., the object detection result) is determined. The feature vector includes category, position, velocity, color, shape, orientation, size, and trajectory. Spatial distances between nodes less than a preset distance threshold (e.g., 2m) are used as edges. The edge weights reflect the interaction strength between objects, and the edge weights are the reciprocal of the spatial distance.
[0052] The environment topology graph G=(V,E) updates the feature vectors of nodes through message passing. In this application, the pre-trained large language model, i.e., the teacher model, plays a guiding role in updating the environment topology graph. Therefore, the hierarchical feature information in the teacher model is used to update the environment topology graph. The update formula is as follows: ; in, For nodes In the teacher model The feature vector of the layer, also known as the updated feature vector. For activation function, For the teacher model The learnable weight matrix of the layer, This is a vector concatenation operation. For nodes In the teacher model The feature vector of the layer, also known as the feature vector before the update. Indexing neighboring nodes, For nodes The set of neighboring nodes, For neighboring nodes In the teacher model The feature vector of the layer.
[0053] Based on the unified map and environmental topology map, the positional changes of dynamic obstacles in the target area are predicted using the spatiotemporal Transformer, and the future state prediction results of the dynamic obstacles are obtained.
[0054] Dilated convolution is used to capture long-term temporal dependencies in the unified graph and the environment topology graph. A decoder is designed to generate multi-step prediction results, i.e., the prediction results of the future state of dynamic obstacles, through a masked self-attention mechanism. The calculation formula is as follows: ; in, For spacetime Transformer, To unify the map, This is an environmental topology diagram.
[0055] This application updates the environmental topology map every 100ms, reflects changes in object position in real time, and supports continuous perception in dynamic scenes.
[0056] The route planning module is used to determine the obstacle avoidance and transport path of the UAV in real time based on the aggregation results, using the improved RRT* algorithm and deep reinforcement learning (DRL) method.
[0057] like Figure 8 As shown, the improved RRT* algorithm is used to adjust the search radius and determine the search range; based on the prediction results and the search range, a deep reinforcement learning method is used to determine the transport path of the UAV.
[0058] Specifically, the search radius is adaptively adjusted based on environmental complexity, automatically reducing the search radius in complex environments to improve safety. The calculation formula is shown below: ; in, For the search radius, are constant parameters. The third hyperparameter controls the rate (sensitivity) at which the search radius decays with environmental complexity (entropy). The cross-entropy loss function is used to reflect the complexity of the environment. To unify the map.
[0059] After determining the search radius, the state space is further determined. The state space includes the UAV pose, obstacle distribution, target location, and unified map. The UAV pose includes the UAV's position and attitude, which includes forward, backward, left, and right. The obstacle distribution is determined by the pixel density distribution. When the pixel density distribution is greater than a preset threshold, it indicates that there is an obstacle at that location.
[0060] Based on the search radius and state space, a reward function is constructed using deep reinforcement learning to optimize the length and safety of the obstacle avoidance and transport path. The reward function calculation formula is as follows: ; in, Let be the objective function. For the search radius, The first hyperparameter is used to adjust the weight of the search radius in the reward function; in this application, it is set to 0.7. The second hyperparameter, which is set to 0.3 in this application, is used to adjust the weight of the collision penalty in the reward function. This is the penalty value for a drone colliding with an obstacle, also known as the collision risk cost.
[0061] This application combines an improved RRT* algorithm and deep reinforcement learning methods to generate safe and efficient obstacle-avoidance transport paths in real time. Dynamic adjustment of the search radius based on actual conditions improves computational efficiency by 50% and path length optimization by ≥30%. By setting hyperparameters for biased sampling, candidate nodes are generated preferentially in the target direction, improving path convergence speed by 40%. The improved RRT* algorithm handles global path exploration, while deep reinforcement learning optimizes local obstacle avoidance, achieving millisecond-level response in dynamic environments. A unified graph identifies passable areas (such as gaps between shelves), reducing invalid search nodes by 50%.
[0062] S3: Control the drone for logistics delivery based on the real-time obstacle avoidance and handling path.
[0063] This application addresses the problems of insufficient perception capabilities, low target detection accuracy, and inefficient route planning in traditional UAVs in complex low-altitude scenarios. Through the deep integration of multimodal perception, knowledge distillation, dynamic fusion of multimodal features, and route planning technologies, a highly efficient and robust environmental perception and path planning model is constructed. Each module is designed to address the shortcomings of traditional methods, significantly improving perception accuracy, decision-making speed, and system energy efficiency in complex scenarios through technological innovation. This provides reliable technical support for practical applications in intelligent logistics and promotes the large-scale application of UAVs in complex low-altitude logistics environments.
[0064] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0065] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A drone logistics delivery method based on multimodal visual perception, characterized in that, The drone logistics delivery method based on multimodal visual perception includes: Acquire environmental data of the target area; the target area includes multiple dynamic and static obstacles; the environmental data includes: RGB images, depth data, thermal imaging images, and LiDAR point cloud data; Based on the environmental data, an obstacle avoidance and transport path is generated in real time using an environmental perception and path planning model. The environmental perception and path planning model includes: a spatiotemporal alignment module, a Transformer feature encoder, a lightweight distillation module, a multimodal feature fusion module, an intelligent analysis module, and a route planning module. The spatiotemporal alignment module performs sequential temporal and spatiotemporal alignment on the environmental data to obtain aligned environmental data. The Transformer feature encoder extracts features from the aligned environmental data to obtain multimodal features. The multimodal features include: RGB features, depth map features, and thermodynamic features. The lightweight distillation module utilizes pre-trained... The large language model is used as the teacher model, and a lightweight student model is trained through knowledge distillation. Semantic features are obtained based on multimodal features. The multimodal feature fusion module is used to perform weighted fusion based on the semantic features and the multimodal features using a cross-modal attention mechanism to obtain a unified graph. The intelligent analysis module is used to perform multi-scale feature aggregation based on graph neural networks and spatiotemporal Transformers based on the unified graph to obtain an aggregated result including object detection results and dynamic obstacle future state prediction results. The route planning module is used to determine the obstacle avoidance and transport path of the UAV in real time based on the aggregated result, using an improved RRT* algorithm and deep reinforcement learning methods. The drone is controlled to perform logistics delivery based on real-time obstacle avoidance and handling paths. The processing steps of the route planning module include: The search radius is adjusted using the improved RRT* algorithm to determine the search range; Based on the prediction results and the search range, the transport path of the UAV is determined using deep reinforcement learning methods; The process of adjusting the search radius and determining the search range using the improved RRT* algorithm specifically includes: Using formula Define the search scope; in, For the search radius, are constant parameters. This is the third hyperparameter. Let cross-entropy be the loss function. This is the prediction result for the future state of dynamic obstacles. To unify the map.
2. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The system utilizes a sensor array to acquire environmental data of the target area; the sensor array includes: an RGB camera, a time-of-flight depth sensor, a thermal imager, and a lidar. The RGB camera is used to acquire RGB images of the target area; The time-of-flight depth sensor is used to acquire depth data of the target area; The thermal imager is used to acquire thermal imaging images of the target area; The lidar is used to acquire lidar point cloud data of the target area.
3. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The processing procedure of the spatiotemporal alignment module includes: Timing alignment of RGB images, depth data, thermal imaging images, and LiDAR point cloud data is performed using pulse synchronization signals and field-programmable gate array logic circuits. The efficient perspective n-point algorithm is used to spatially align the temporally aligned RGB image, temporally aligned depth data, temporally aligned thermal imaging image, and temporally aligned LiDAR point cloud data, resulting in aligned RGB image, aligned depth data, aligned thermal imaging image, and aligned LiDAR point cloud data.
4. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The Transformer feature encoder includes independent RGB branches, depth and LiDAR branches, and a thermal imaging branch; The RGB branch is used to extract global semantic features of the aligned RGB image using a visual Transformer to obtain RGB features; The depth and lidar branches are used to extract multi-scale spatial features of aligned depth data and aligned lidar point cloud data using PointNet++ to obtain depth map features; The thermal imaging branch is used to extract the temperature distribution and temporal changes of the aligned thermal imaging image using a three-dimensional convolutional neural network, thereby obtaining thermodynamic features.
5. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The lightweight distillation module specifically includes: Using formula Knowledge distillation is performed on pre-trained large language models; in, Distillation of knowledge For feature channel index, The Kullback-Leibler divergence function, To reduce the feature distribution output by the lightweight student model, This refers to semantic prior knowledge.
6. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The processing procedure of the multimodal feature fusion module includes: Determine the internal interaction weights of semantic features, RGB features, depth map features, and thermodynamic features respectively; Based on the internal interaction weights and using the cross-attention mechanism, the formula is applied. Determine the cross-modal interaction weights; where, For cross-modal interaction weights, for function, This is a query vector generated based on RGB features. A key vector generated based on semantic features. The key vector is generated based on depth map features. The bond vector is generated based on thermodynamic characteristics. The dimension of the key vector; Based on cross-modal interaction weights, using the formula A unified map is obtained by weighted fusion of semantic features, RGB features, depth map features, and thermodynamic features; among them, To unify the map, The spatial location index of the feature represents the first... Okay, number Feature points of the column A value vector generated based on semantic features. A value vector generated based on depth map features. This is a value vector generated based on thermodynamic characteristics.
7. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The processing steps of the intelligent analysis module include: Based on the unified graph, an environmental topology graph is constructed using a graph neural network to obtain object detection results; Based on the unified map and environmental topology map, using the formula The positional changes of dynamic obstacles in the target area are predicted to obtain the predicted future state of the dynamic obstacles; among which, This is the prediction result for the future state of dynamic obstacles. For spacetime Transformer, To unify the map, This is an environmental topology diagram.
8. The UAV logistics delivery method based on multimodal visual perception according to claim 1, characterized in that, The step of determining the transport path of the UAV using deep reinforcement learning methods based on the prediction results and the search range specifically includes: Using formula Determine the transport route for the drone; in, Let be the objective function. For the search radius, As the first hyperparameter, This is the second hyperparameter. This is the penalty value for a drone colliding with an obstacle.
Citation Information
Patent Citations
Multi-mode perception adaptive unmanned aerial vehicle obstacle avoidance method, device, equipment and medium
CN120722926A
Three-dimensional semantic segmentation method based on multi-modal alignment and knowledge distillation in rail transit scene
CN121213931A