Unmanned aerial vehicle low-altitude remote sensing data ai intelligent analysis method based on multi-modal fusion

By employing a multimodal fusion AI method in the intelligent analysis of UAV low-altitude remote sensing data, and utilizing shared feature components to load fusion analysis branches in parallel, the problem of poor multimodal data adaptability in existing technologies is solved, achieving efficient and accurate remote sensing data analysis.

CN122223591APending Publication Date: 2026-06-16CHANGZHOU XINYI SPACE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGZHOU XINYI SPACE INFORMATION TECH CO LTD
Filing Date
2026-03-17
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing intelligent analysis technologies for low-altitude remote sensing data from unmanned aerial vehicles (UAVs) often rely on single-modal data processing models or fixed-structure multi-branch analysis frameworks, failing to fully leverage the information advantages of multimodal fusion data and lacking a unified feature extraction and sharing mechanism, resulting in inaccurate analysis results and poor adaptability.

Method used

An AI-based intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion is adopted. By sharing feature components and attaching several fusion analysis branches in parallel, the multimodal feature sets are extracted using the shared feature components. The parallel fusion analysis branches can be dynamically adjusted to achieve automated analysis of remote sensing fusion data.

Benefits of technology

It achieves integrated collaborative reasoning across multiple tasks, improving analysis efficiency and accuracy, enhancing the reliability and anti-interference capabilities of analysis results, and adapting to the high-precision, automated analysis needs of multiple industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223591A_ABST
    Figure CN122223591A_ABST
Patent Text Reader

Abstract

The application discloses a UAV low-altitude remote sensing data AI intelligent analysis method based on multi-modal fusion, and relates to the technical field of low-altitude remote sensing data intelligent analysis; the application matches and fuses analysis branches through data analysis requirements or in combination with fusion feature data, and is hung on the output layer of a shared feature component in parallel through a standardized interface and an adaptive interface; multi-modal feature sets are extracted through the shared feature component, the multi-modal feature sets are analyzed by each fusion analysis branch, branch output results are stored in a branch result shared library, and mutual reuse is realized according to a reuse mapping rule; the application effectively solves the problems of incomplete feature extraction, poor branch adaptability and the like in the prior art, realizes multi-task integrated collaborative reasoning, avoids repeated feature extraction, improves analysis efficiency, and results can be cross-verified; branches can dynamically adapt to data characteristics and scene requirements, the universality is significantly enhanced, and meanwhile, the analysis precision and anti-interference capability are improved by relying on multi-modal features and a reuse mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent analysis technology for low-altitude remote sensing data, specifically an AI-based intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion. Background Technology

[0002] In existing technologies, intelligent analysis of UAV low-altitude remote sensing data often relies on single-modal data processing models or fixed-structure multi-branch analysis frameworks. Typically, independent models perform single tasks such as ground cover classification, target detection, and quantitative inversion, or fixed-configuration branch combinations are used to analyze the fused data. Some solutions attempt to achieve joint processing of multimodal data through simple feature stitching, but lack a unified feature extraction and sharing mechanism. In practical applications, these technologies are often driven by pre-set analysis objectives, directly calling fixed analysis modules without accurately adapting to the characteristics of remote sensing fusion data and specific analysis needs, thus failing to fully leverage the informational advantages of multimodal fusion data.

[0003] This invention provides an AI-based intelligent analysis method for low-altitude remote sensing data from unmanned aerial vehicles (UAVs) based on multimodal fusion to solve the aforementioned technical problems. Summary of the Invention

[0004] The present invention aims to at least solve one of the technical problems existing in the prior art; to this end, the present invention proposes an AI intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion.

[0005] To achieve the above objectives, a first aspect of the present invention provides an AI-based intelligent analysis method for low-altitude remote sensing data from unmanned aerial vehicles (UAVs) based on multimodal fusion, comprising: Several fusion analysis branches are obtained based on data analysis requirements; these fusion analysis branches are then mounted in parallel on the output layer of the shared feature component. Multimodal feature sets of remote sensing fusion data are extracted by sharing feature components; the multimodal feature sets are analyzed using several fusion analysis branches to achieve automated analysis of remote sensing fusion data.

[0006] In one possible implementation, several fusion analysis branches are mounted in parallel on the output layer of a shared feature component, including: The output layer of the shared feature component is equipped with a standardized output interface, and the feature input layer of the fusion analysis branch is equipped with an adapter interface that matches the standardized output interface. The parallel mounting of fusion analysis branches is achieved based on standardized output interfaces and adaptation interfaces; among them, the fusion analysis branches mounted in parallel on the shared feature component output layer can be dynamically adjusted.

[0007] In one possible implementation, the shared feature components include an input layer, a single-modal feature extraction layer, a cross-modal feature fusion layer, a feature optimization layer, and an output layer.

[0008] In one possible implementation, the multimodal feature set includes basic general features and branch-specific features; wherein, the basic general features include spatial geometric features, global texture features, modality fusion weight features, and feature quality features; the branch-specific features correspond to each fusion analysis branch and are used to meet the analysis requirements of the fusion analysis branch.

[0009] In one possible implementation, the branch outputs of several parallel-mounted fusion analysis branches are reused, including: The branch output results of the fusion analysis are stored in the branch result shared library; Each fusion analysis branch calls the branch output results from the branch result sharing library according to the pre-built reuse mapping rules, so as to realize the mutual reuse of branch output results.

[0010] In one possible implementation, a fusion analysis branch is matched based on the fusion feature data and data analysis requirements, including: Identify the fusion feature data of remote sensing fusion data; among which, the fusion feature data is used to determine whether the fusion analysis branch is compatible with the remote sensing fusion data; Several fusion analysis branches were obtained by matching data analysis requirements; Based on the data matching of several fusion analysis branches, the fusion analysis branches are selected according to the data matching analysis results.

[0011] In one possible implementation, the fusion feature data of the remote sensing fusion data is identified, including: Call a pre-trained lightweight CNN model; the lightweight CNN model includes an input layer, lightweight convolutional layers, pooling layers, fully connected layers, and an output layer; A lightweight CNN model is used to identify attribute features of remote sensing fusion data to obtain fused feature data. The fused feature data includes modality composition features, data quality features, dominant feature types, and scene adaptation features.

[0012] In one possible implementation, several fusion analysis branches mounted in parallel are jointly fine-tuned, including: Select a fine-tuning dataset from the remote sensing fusion data to be processed; The parameters of several fusion analysis branches are fine-tuned using a fine-tuning dataset; the fine-tuning layers include feature input layers, attention modules, or output layers.

[0013] In one possible implementation, the parameters of the fine-tuning layers for several fusion analysis branches are fine-tuned using a fine-tuning dataset, including: Retrieve and reuse mapping rules; Based on the reuse mapping rules, fine-tuning priorities are set for several fusion analysis branches, and the corresponding fusion analysis branches are fine-tuned in sequence according to the fine-tuning priorities.

[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention, by individually matching fusion analysis branches based on data analysis needs or in combination with fusion feature data, is mounted in parallel on the output layer of a shared feature component via standardized and adaptive interfaces. A multimodal feature set is extracted through this shared feature component, and each fusion analysis branch analyzes this multimodal feature set, with the branch output results stored in a shared branch result library. These results are reused according to reuse mapping rules. This invention effectively solves the problems of incomplete feature extraction and poor branch adaptability in existing technologies, achieving integrated collaborative reasoning across multiple tasks, avoiding redundant feature extraction, improving analysis efficiency, and enabling cross-validation of results. Branches can dynamically adapt to data characteristics and scenario requirements, significantly enhancing versatility. Simultaneously, relying on multimodal features and reuse mechanisms, it improves analysis accuracy and anti-interference capabilities, meeting the high-precision, automated, and intelligent analysis needs of multiple industries.

[0015] 2. This invention selects a fine-tuning dataset from the remote sensing fusion data to be processed, performs parameter fine-tuning on the fine-tuning layer of the fusion analysis branch, retrieves the reuse mapping rules to set the fine-tuning priority, and performs parameter fine-tuning sequentially according to the fine-tuning priority without changing the core structure of the branch. This invention quickly corrects the adaptation deviation between the branch and the current data while ensuring the original performance of the pre-trained basic model, and enhances the synergy between branches. The fine-tuning process is standardized and quantifiable, taking into account both efficiency and effectiveness, further improving the accuracy and reliability of the analysis results, making the model more flexibly adaptable to diverse data scenarios, and providing more stable support for remote sensing data decision-making in various industries. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the method flow for AI intelligent analysis of low-altitude remote sensing data from unmanned aerial vehicles in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the method for joint fine-tuning of parallel mounted fusion analysis branches in Embodiment 2 of the present invention. Detailed Implementation

[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example:

[0020] Please see Figure 1 The first aspect of the present invention provides an AI-based intelligent analysis method for low-altitude remote sensing data from unmanned aerial vehicles (UAVs) based on multimodal fusion, comprising: matching several fusion analysis branches according to data analysis requirements; mounting the several fusion analysis branches in parallel on the output layer of a shared feature component; extracting a multimodal feature set of the remote sensing fusion data through the shared feature component; and analyzing the multimodal feature set using the several fusion analysis branches to achieve automated analysis of the remote sensing fusion data.

[0021] The data analysis requirements are used to match the fusion analysis branch of the remote sensing fusion data. Therefore, the data analysis requirements need to match the core functions of the fusion analysis branch, and the rationality of the matching can be verified. The remote sensing fusion data is obtained by fusing multimodal low-altitude remote sensing data collected by UAVs.

[0022] In one example, if the data analysis requirement is "land cover classification map, land cover distribution statistics table", then the semantic segmentation branch needs to be matched; if the data analysis requirement is "target coordinate list, target outline map, target count report", then the target detection branch needs to be matched; if the data analysis requirement is "parameter raster map, quantitative statistics report, parameter spatial distribution map", then the quantitative inversion branch needs to be matched; if the data analysis requirement is "abnormal point coordinates, abnormal risk list, abnormal distribution raster map", then the abnormal identification branch needs to be matched; if the data analysis requirement is multiple types of requirements, such as "classification map, parameter report, abnormal list", then multiple corresponding branches need to be matched.

[0023] The fusion analysis branches include semantic segmentation branches, object detection branches, quantitative inversion branches, anomaly recognition branches, etc. The fusion analysis branches can be set and called according to specific task requirements and are not limited to the branch types listed above.

[0024] The semantic segmentation branch is used to perform semantic classification of all land features in remote sensing fusion data, accurately distinguishing different types of land features and completing pixel-level annotation. The workflow of the semantic segmentation branch is as follows: it receives a multimodal feature set with shared feature components, improves the U-Net model to achieve pixel-level segmentation of various land features, and generates a standardized semantic segmentation map. Furthermore, the semantic segmentation map can also serve as auxiliary input for other fusion analysis branches, improving the analysis accuracy of the corresponding fusion analysis branches. It should be noted that the semantic segmentation branch can also dynamically adjust the weights of the attention gating module based on the dominant feature type to adapt to the semantic classification requirements of different scenarios.

[0025] The basic structure of semantic segmentation includes a feature input layer, an encoding layer, a decoding layer, and an output layer.

[0026] The feature input layer is used to receive the multimodal feature set output by the shared feature component and adjust the feature dimension of the multimodal feature set.

[0027] The encoding layer consists of 4-5 convolutional blocks, each containing two 3×3 standard convolutional layers, one BN layer, and one ReLU activation function. Max pooling layers (with a stride of 2) are interspersed to achieve feature dimensionality reduction, focusing on mining land cover category association features in the multimodal feature set. The encoding layer can also embed a cross-modal attention gating module to enhance the discriminative power between land cover boundary features and category features and suppress noise interference.

[0028] The decoding layer and the encoding layer are symmetrically set with 4-5 deconvolution blocks. The transposed convolution is used to achieve feature dimensionality enhancement and gradually restore the spatial details of ground features. Each deconvolution block is connected to the feature map of the corresponding layer of the encoding layer to supplement shallow details such as ground feature edges and textures, thereby improving the accuracy of segmentation boundaries. The output layer uses a 1×1 convolutional layer to map the feature map into a segmentation map with the same size as the input remote sensing fusion data. The softmax activation function is used to output the land cover category probability of each pixel, and finally outputs a single-channel semantic segmentation map.

[0029] The target detection branch is used to accurately locate, count, and extract the contours of specific targets in remote sensing fusion data, focusing on individual-level target analysis. The workflow of the target detection branch includes: receiving a multimodal feature set of shared feature components; detecting and statistically analyzing typical targets (such as buildings, utility poles, disaster sites, etc.) in the remote sensing fusion data based on the multimodal feature set; and outputting target detection information such as target type, target quantity, and bounding box coordinates. It should be noted that the target detection branch can also dynamically adjust the detection threshold and feature weights according to data quality characteristics to adapt to the target detection needs of different scenarios.

[0030] The object detection branch includes a feature input layer, an object detection head, an instance segmentation head, and an output layer.

[0031] The feature input layer receives the multimodal feature set output by the shared feature component and constructs a multi-scale feature pyramid through the Feature Pyramid Network (FPN) to extract features of small-scale (such as single trees and utility poles), medium-scale (such as small buildings), and large-scale (such as landslides and contiguous orchards) targets respectively, thus solving the problem of large differences in target scale in remote sensing data.

[0032] The target detection head is based on the YOLOv9 detection head structure and has three detection branches at different scales, corresponding to different levels of the multi-scale feature pyramid. Each detection branch outputs the bounding box coordinates, class probability, and confidence of the target. At the same time, a coordinate attention module is embedded to enhance the distinction between target features and background features and reduce the false negative rate of small and dense targets.

[0033] The instance segmentation head integrates the instance segmentation module of Mask R-CNN, which works in parallel with the object detection head to extract pixel-level contours for each detected object and output the object instance mask.

[0034] Output layer: Non-maximum suppression (NMS) algorithm is used to remove redundant bounding boxes and filter candidate targets with confidence scores below the threshold (which can be dynamically adjusted); combined with the elevation and spectral features of shared feature components, the position of the target bounding box is corrected to improve the localization accuracy; finally, the target detection information is output.

[0035] The quantitative inversion branch is used to obtain continuous geographic parameters with practical business significance. The workflow of the quantitative inversion branch is as follows: it receives the multimodal feature set output from the shared output layer, analyzes the multimodal feature set to obtain continuous geographic parameters such as vegetation cover (FVC), leaf area index (LAI), tree height, crown width, land surface temperature (LST), earthwork volume, and land slope. It should be noted that the quantitative inversion branch can dynamically adjust the regression model weights according to the dominant feature type to control the relative error of the inversion.

[0036] The quantitative inversion branch includes a feature input layer, a feature dimensionality reduction layer, a regression prediction layer, and an output layer.

[0037] The feature input layer is used to receive the multimodal feature set output by the shared feature component. Through a lightweight encoding module composed of three convolutional blocks, it further mines the deep correlation features related to quantitative parameters in the feature set, such as the correlation between spectral features and vegetation growth, and the correlation between elevation features and terrain parameters. Each convolutional block contains one 3×3 convolutional layer, one BN layer and one ReLU activation function. No pooling layer is set to avoid loss of feature information.

[0038] The feature dimensionality reduction layer uses a global average pooling layer to convert the encoded feature map into a fixed-dimensional feature vector. It then uses a fully connected layer to further reduce the dimensionality, remove redundant features, and retain the core features that are relevant to quantitative inversion.

[0039] The regression prediction layer consists of 3-4 fully connected layers (MLP), using the ReLU activation function (the last layer uses a linear activation function) to map the feature vectors to specific quantitative parameter values; the embedded residual connection module alleviates the gradient vanishing problem in the regression process and improves the inversion accuracy.

[0040] The output layer, combined with a pre-defined inversion error correction model, corrects the parameter values ​​of the regression prediction output, removes outliers, and ensures the stability and accuracy of the inversion results. The final output is a quantitative parameter raster map with the same size as the input remote sensing fusion data. The inversion error correction model is trained based on a large amount of measured data.

[0041] The anomaly identification branch is used to accurately identify anomalies in remote sensing fusion data and perform risk classification. The workflow of the anomaly identification branch is as follows: it receives the multimodal feature set output by the shared feature component and the output results of the other fusion analysis branches mentioned above, identifies anomalous targets in the remote sensing fusion data, and outputs the type and region of the anomalous targets.

[0042] The anomaly detection branch includes a feature input layer, an anomaly feature extraction layer, a fusion decision layer, and an output layer. It is also equipped with a business rule base, which allows for the customization of business rules, such as threshold rules and topology rules. Moreover, the business rules can be dynamically updated according to data analysis needs.

[0043] The feature input layer is used to receive multimodal feature sets of shared feature components, semantic segmentation maps of semantic segmentation branches, target detection information of target detection branches, and parameter inversion results of quantitative inversion branches, so as to realize the collaborative utilization of multi-source parameters.

[0044] The anomaly feature extraction layer uses a lightweight CNN network (5-7 layers) to extract anomaly association features of multi-source parameters, such as spectral anomalies of diseased and pest-infested vegetation, elevation deformation anomalies of landslide bodies, and land cover category anomalies of illegal buildings. The anomaly feature extraction layer also embeds an attention mechanism to enhance the distinction between anomaly features and normal features.

[0045] The fusion decision layer combines the output of the anomaly feature layer and the business rule base, and uses a Bayesian inference algorithm to make fusion decisions to determine whether the target or region is abnormal, and at the same time classify the anomaly risk level (such as high, medium and low).

[0046] The output layer is used to output anomaly information, including the coordinates of the anomaly points, the anomaly type, the risk level, etc., and to generate an anomaly warning list and an anomaly distribution raster map.

[0047] It should be noted that the types of branches included in the fusion analysis branch are not unique. Multiple basic models based on different models but with the same function can be pre-trained, such as the object detection branch for object detection built based on multiple artificial intelligence models.

[0048] It is worth noting that in some cases, although the data analysis requirement may be singular, if the analytical precision of this requirement needs to be verified / supplemented by other branches, multiple branches still need to be matched to ensure the accuracy and reliability of the results. The additional matched fusion analysis branches provide auxiliary data to achieve the data analysis requirement. In one example, when the target detection branch performs target detection on remote sensing fusion data, utilizing land cover classification results can improve target detection efficiency and accuracy. In this case, the semantic segmentation branch can be matched simultaneously, and its output semantic segmentation map can be used as one of the input features for the target detection branch. When identifying anomalies, the anomaly detection branch can utilize the target detection information from the target detection branch to assist in focusing on the target area and avoiding false alarms in area-level anomaly detection. Therefore, the target detection branch can be matched simultaneously for assistance. Thus, if the data analysis requirement is anomaly detection, it is necessary to match the anomaly detection branch, the target detection branch, and the semantic segmentation branch.

[0049] In a preferred embodiment, several fusion analysis branches are mounted in parallel on the output layer of a shared feature component, including: the output layer of the shared feature component is provided with a standardized output interface, and the feature input layer of the fusion analysis branches is preset with an adapter interface that matches the standardized output interface; the parallel mounting of the fusion analysis branches is realized based on the standardized output interface and the adapter interface; wherein, the fusion analysis branches mounted in parallel on the output layer of the shared feature component can be dynamically adjusted.

[0050] The shared feature components include an input layer, a single-modal feature extraction layer, a cross-modal feature fusion layer, a feature optimization layer, and an output layer.

[0051] The input layer receives preprocessed remote sensing fusion data and performs unified adaptation and normalization of multimodal data. The input layer has a built-in normalization module that quantizes different modal data to the range [0,1].

[0052] The single-modal feature extraction layer is used to extract modal features from single-modal data. 1. Optical modality (RGB) feature extraction module: Employs a 3-layer lightweight CNN structure to extract modal features such as texture, edge, and color features from images, adapting to the needs of feature / object recognition in semantic segmentation and object detection branches; 2. Elevation modality (LiDAR) feature extraction module: Employs a 2-layer convolutional + global pooling structure to extract modal features such as terrain elevation, slope, and texture roughness, adapting to the needs of terrain parameter inversion and landslide detection in quantitative inversion and object detection branches; 3. Spectral modality (multispectral) feature extraction module: Employs a 4-layer deep separateable... 4. Temperature modality (thermal infrared) feature extraction module: Adopting a lightweight MLP + convolutional structure, it extracts modal features such as surface temperature features and temperature gradient features, adapting to the needs of vegetation classification and pest identification in semantic segmentation branch, quantitative inversion branch, and anomaly identification branch;

[0053] The cross-modal fusion layer adaptively fuses the modal features output from the single-modal feature extraction layer to generate fused features that combine the advantages of each modality. This layer employs an attention mechanism and feature concatenation fusion strategy, with a built-in cross-modal attention module that automatically identifies the importance of each modal feature (e.g., in disaster scenarios, the elevation and temperature modalities have higher weights than the optical modal), dynamically assigns fusion weights, and avoids interference from invalid modal features. The fusion process automatically adapts to the input modal composition; if a certain modality is missing (e.g., the non-thermal infrared modality), the fusion logic of the corresponding modality is automatically skipped without affecting the overall fusion effect, adapting to the diversity of UAV remote sensing fusion data. After fusion, multi-scale fused features are output. Shallow features are adapted to the target / ground feature localization requirements of target detection and semantic segmentation branches, while deep features are adapted to the parameter inversion and anomaly feature extraction requirements of quantitative inversion and anomaly identification branches.

[0054] The feature optimization layer is used to denoise and enhance the multi-scale fusion features output by the cross-modal fusion layer, improving feature discriminability and preventing noisy features from interfering with the analysis accuracy of each fusion analysis branch. Simultaneously, it standardizes the feature dimensions to ensure that the output features can directly adapt to the feature input requirements of each fusion analysis branch. This layer employs lightweight Gaussian filtering and residual connection structures to remove noise interference in the fusion features (such as flight noise and pixel interference in the fusion data), improving feature purity. It uses an attention gating mechanism to strengthen core features (such as ground feature edges and target contours) and suppress redundant features, supporting accurate analysis and result reuse for each branch. It uniformly adjusts the multi-scale fusion features to a fixed dimension to avoid branch adaptation problems caused by incompatible feature dimensions, while providing standardized input for the feature input layer of each fusion analysis branch.

[0055] The output layer outputs the optimized multi-scale fused features as a multimodal feature set. This layer is the core interface connecting the shared feature components to various fusion analysis branches. The standardized output interface of the output layer supports simultaneous access from multiple fusion analysis branches, automatically allocates feature output channels, and ensures that each fusion analysis branch receives the multimodal feature set in parallel. The multimodal feature set includes basic general features and branch-specific features.

[0056] The basic general features include: 1) Spatial geometric features: covering the coordinate features of image pixels, spatial neighborhood relationship features, and geometric features of land cover contours (such as contour perimeter, area, and shape factor), adapting to the land cover classification and target localization requirements of semantic segmentation and target detection branches, while supporting the unification of coordinate systems for the results of each branch; 2) Global texture features: including the gray-level co-occurrence matrix features of the image, texture roughness, and texture uniformity, adapting to the land cover / target differentiation requirements of each branch (such as the texture difference between bare soil and forest land, and the texture difference between normal vegetation and diseased vegetation); 3) Modal fusion weight features: recording the fusion weight of each single modality during cross-modal fusion, providing a basis for cross-validation of the analysis results of each branch (such as determining whether an anomaly in a certain area is related to the excessively high weight of the elevation modality feature); 4) Feature quality features: including feature purity, noise content, and completeness, used by each branch to verify the accuracy of its own analysis results.

[0057] Branch-specific features include: 1) Semantic segmentation branch-specific features: mainly optical + spectral modal fusion features, including land cover category association features, pixel-level spectral response difference features, and land cover boundary enhancement features, adapting to the pixel-level land cover classification requirements of the semantic segmentation branch, and supporting the reuse of its results by the target detection and anomaly recognition branches; 2) Target detection branch-specific features: mainly optical + elevation modal fusion features, including small-scale target enhancement features, target contour edge features, and elevation gradient features of terrain targets, adapting to the target localization, counting, and contour extraction requirements of the target detection branch, and supporting the reuse of its results by the anomaly recognition branch; 3) Quantitative inversion branch-specific features: The main features are the fusion of elevation, spectral and temperature modes, including vegetation growth parameter correlation features (such as the correlation between spectrum and tree height), topographic parameter features (fusion features of elevation, slope and aspect), and surface temperature gradient features. These features are adapted to the continuous parameter inversion requirements of the quantitative inversion branch and support the reuse of its results by the anomaly identification branch. 4) Anomaly identification branch-specific features: These features cover anomaly correlation features of all modes, including spectral anomaly features, elevation deformation anomaly features, temperature anomaly gradient features, and land cover category anomaly correlation features. These features are adapted to the anomaly screening and risk classification requirements of the anomaly identification branch and support the feedback of its results to the other three major branches, thereby optimizing the feature extraction accuracy of each branch.

[0058] After matching and obtaining the fusion analysis branch, a unified interface and parallel mounting method are used to connect with the shared feature components to ensure stable connection, efficient information exchange, and adapt to the dynamic adjustment characteristics of the branch, achieving fully automated connection without manual configuration.

[0059] In the design of the connection interface between the shared feature component and the fusion analysis branch, the output layer of the shared feature component is equipped with a standardized output interface. This interface is compatible with the feature input layers of all fusion analysis branches, and it can automatically identify the number of matched fusion analysis branches and feature requirements, dynamically adjusting the feature output dimensions to adapt to the feature input requirements of different fusion analysis branches. The feature input layer of each fusion analysis branch has a pre-set unified adaptation interface, through which a quick and instant connection can be achieved with the standardized output interface of the shared feature component.

[0060] After connection, the shared feature component completes feature extraction of the remote sensing fusion data and then sends the extracted multimodal feature set to the attached fusion analysis branch through a standardized output interface. Each fusion analysis branch receives the required multimodal features through the adaptation interface of the feature input layer, processes the multimodal features, and uses them for subsequent analysis.

[0061] In this invention, the shared feature components and each fusion analysis branch are mounted in parallel using a unified interface, and the dynamic addition or removal of fusion analysis branches is supported. Each fusion analysis branch shares the same multimodal feature set, avoiding redundant feature extraction by each branch, which can significantly reduce computing power consumption and improve inference efficiency; the connection process requires no manual intervention, adapting to the automated analysis needs of batch remote sensing fusion data.

[0062] In a preferred embodiment, the branch output results of several parallel fusion analysis branches are reused, including: the branch output results of the fusion analysis branches are stored in a branch result shared library; each fusion analysis branch calls the branch output results from the branch result shared library according to a pre-built reuse mapping rule, thereby realizing the mutual reuse of branch output results.

[0063] There is no need to set direct hard links between the matched fusion analysis branches, that is, there is no need to use fixed network layer connections, so as to avoid excessive coupling between the fusion analysis branches. This also ensures that the adjustment of each fusion analysis branch does not affect the normal operation of other branches.

[0064] The branch outputs of each fusion analysis branch are processed according to a preset standardized format and stored in a branch result shared library. During data analysis, each fusion analysis branch retrieves the corresponding branch output from the branch result shared library using pre-built reuse mapping rules.

[0065] The reuse mapping rules include which other fusion analysis branches each branch can utilize to improve analysis accuracy. For example, the object detection branch can reuse the semantic segmentation map from the semantic segmentation branch to locate the analysis range and remove background interference; the anomaly detection analysis branch can reuse the branch outputs from the semantic segmentation branch, the object detection branch, and the quantitative inversion branch to enhance the feature extraction accuracy of anomaly regions.

[0066] It should be noted that the reuse mapping rules are pre-stored and can be read and called at any time. The rules included in the reuse mapping rules are not limited to the examples above; different reuse rules can be preset according to different scenarios, and the reuse mapping rules can be dynamically updated.

[0067] In a preferred embodiment, matching fusion analysis branches based on fusion feature data and data analysis requirements includes: identifying fusion feature data of remote sensing fusion data; wherein, the fusion feature data is used to determine whether the fusion analysis branch is compatible with the remote sensing fusion data; obtaining several fusion analysis branches through data analysis requirement matching; analyzing the data matching of several fusion analysis branches based on the fusion feature data; and selecting several fusion analysis branches based on the data matching analysis results.

[0068] Identifying the fusion feature data of remote sensing fusion data includes: calling a pre-trained lightweight CNN model; using the lightweight CNN model to identify attribute features of the remote sensing fusion data to obtain fusion feature data; wherein, the fusion feature data includes modality composition features, data quality features, dominant feature types and scene adaptation features.

[0069] The goal of lightweight CNN models is to quickly and efficiently identify attribute features from remote sensing fusion data. This requires balancing recognition accuracy and inference speed while avoiding excessive computational resource consumption that could negatively impact overall analysis efficiency. A lightweight CNN model consists of an input layer, lightweight convolutional layers, pooling layers, fully connected layers, and an output layer.

[0070] The input layer receives the multimodal fused remote sensing data. The number of input channels corresponds to the modal composition of the fused remote sensing data; for example, RGB and LiDAR fused remote sensing data has 4 channels, while RGB, multispectral, and thermal infrared fused remote sensing data has 6-8 channels. The input size of the input layer is adapted to the standard resolution of the fused remote sensing data, thus eliminating the need for scaling and reducing preprocessing time.

[0071] The lightweight convolutional layer uses depthwise separable convolution instead of traditional standard convolution to reduce the number of parameters and computational cost. It sets 3-4 convolutional blocks, each containing a depthwise separable convolutional layer, a BN (Batch Normalization) layer and a ReLU activation function. The kernel size is 3×3 with a stride of 1 and zero padding at the edges to ensure that the feature map size is not reduced and to focus on extracting shallow attribute features from the fused data.

[0072] Pooling layer: Each convolutional block is followed by a global average pooling layer to replace the traditional max pooling. This can further simplify parameters and reduce dimensionality, while preserving global attribute information and avoiding the loss of attribute features caused by local pooling.

[0073] The fully connected layer maps the feature vectors output by the pooling layer to fixed-dimensional attribute recognition results, which are then used for subsequent fusion analysis and branch matching.

[0074] The output layer uses the Softmax activation function to output fused feature data of the remote sensing fusion data. The fused feature data can be directly used for matching branches in the fusion analysis.

[0075] The training objective of lightweight CNN models is to accurately identify various attribute features of remote sensing fusion data, providing a reliable basis for the accurate matching of fusion analysis branches.

[0076] Training samples are pre-constructed by collecting multimodal remote sensing fusion data from multiple scenarios and types. Specifically, this needs to cover different application scenarios such as agriculture, forestry, disaster relief, and urban areas, including remote sensing fusion data with different modalities, resolutions, noise intensities (e.g., low, medium, and high noise), and modal proportions. Attribute feature annotations are performed on the training samples, and the annotations are matched with subsequent fusion analysis branches. The training samples and their corresponding annotations are then integrated into a training dataset.

[0077] The training parameters were set to use the cross-entropy loss function as the training loss function to optimize the attribute recognition accuracy of the model; the Adam optimizer was used, the initial learning rate was set to 0.001, and a learning rate decay strategy was adopted (decaying to 0.5 every 20 rounds); the number of iterations was set to 80-100 rounds, the batch size was set to 32, and an early stopping strategy was adopted (training was stopped if the accuracy on the validation set did not improve for 10 consecutive rounds) to avoid model overfitting.

[0078] The training dataset was divided into training, validation, and test sets in a 7:2:1 ratio, ensuring that each set included remote sensing fusion data from different scenarios and modalities to guarantee sample balance. The training set was input into a lightweight CNN model, and the attribute recognition results were output through forward propagation. The loss value was calculated by comparing the output with the labeled attribute features, and the model parameters (convolutional kernel parameters, fully connected layer weights) were updated through backpropagation. The model hyperparameters were adjusted in real time using the validation set to optimize the model structure. After training, the model performance was validated using the test set to ensure that the attribute feature recognition accuracy and inference time met the requirements. The trained lightweight CNN model was stored in a model library and associated with the pre-trained base models for each fusion analysis branch.

[0079] Fusion feature data is the output of a lightweight CNN model. It provides accurate basis for matching fusion analysis branches, ensuring that the matched fusion analysis branches are adapted to the characteristics of the current remote sensing fusion data. Fusion feature data specifically includes modality composition features, data quality features, dominant feature types, and scene adaptation features.

[0080] Modal composition features are used to quantitatively characterize the modal types in remote sensing fusion data, such as "RGB+LiDAR", "multispectral+thermal infrared", "RGB+multispectral+LiDAR+thermal infrared", etc. They can be identified using binary encoding, for example, RGB is 0001, LiDAR is 0010, and RGB+LiDAR is 0011. Modal composition features are used to match fusion analysis branches with compatible modal fit.

[0081] Data quality features include resolution and noise intensity. Resolution is quantified into specific numerical values, such as 0.1m, 0.5m, 1m, etc., while noise intensity can be quantified between 0 and 1, where 0 represents low noise and 1 represents high noise. Data quality features are used to match fusion analysis branches that meet accuracy requirements.

[0082] Dominant feature types are used to quantitatively characterize the most prevalent and core feature types in remote sensing fusion data, such as optical feature dominance, spectral feature dominance, elevation feature dominance, and temperature feature dominance. Dominant feature types are used to match fusion analysis branches that emphasize functional compatibility.

[0083] Scene adaptation features are used to quantitatively characterize the application scenario tendency of remote sensing fusion data, such as agricultural and forestry scenarios, urban management scenarios, etc. Based on the shallow texture and feature distribution of remote sensing fusion data, quantitative judgment is made to further accurately match the fusion analysis branches. Example

[0084] In a preferred embodiment, please refer to Figure 2 The process involves joint fine-tuning of several parallel fusion analysis branches, including: selecting a fine-tuning dataset from the remote sensing fusion data to be processed; and fine-tuning the parameters of the fine-tuning layers of several fusion analysis branches using the fine-tuning dataset. The fine-tuning layers include feature input layers, attention modules, or output layers.

[0085] The matching fusion analysis branches are pre-trained basic analysis models, which are then fine-tuned using the remote sensing fusion data to be processed after being mounted onto the output layer of the shared feature components.

[0086] The training process for each fusion analysis branch is as follows: Standard training sets are constructed for each fusion analysis branch, such as semantic segmentation, object detection, quantitative inversion, and anomaly recognition. Following the existing data preprocessing and training processes for artificial intelligence models, the corresponding fusion analysis branches are trained using the standard training sets. The trained fusion analysis branches are then saved to the model library for later matching and retrieval.

[0087] It should be noted that the standard dataset includes not only remote sensing fusion data, but may also include the branch outputs of other fusion analysis branches. For a detailed understanding of the inclusion relationship, please refer to the reuse mapping relationship.

[0088] In a preferred embodiment, the fine-tuning of parameters of the fine-tuning layer of several fusion analysis branches is performed using the fine-tuning dataset, including: retrieving the reuse mapping rules; setting the fine-tuning priority for several fusion analysis branches according to the reuse mapping rules; and fine-tuning the corresponding fusion analysis branches in sequence according to the fine-tuning priority.

[0089] The core objective of joint fine-tuning of fusion analysis branches is to enhance the synergy between pre-trained fusion analysis branches. Fine-tuning is performed only at specific levels of the branches without altering the core structure. The specific implementation is as follows: Non-core layers of each fusion analysis branch are designated as fine-tuning layers, such as the feature input layer, attention module, and output layer in the fusion analysis branch. Core layers of the fusion analysis branch, such as the encoding layer of the semantic segmentation branch and the feature pyramid network of the object detection branch, are not adjusted to ensure fine-tuning efficiency and to avoid affecting the performance of the fusion model.

[0090] When selecting fine-tuning samples to form the fine-tuning sample set, a small number of samples (≤5%) are randomly selected from the remote sensing fusion data input in the current batch as fine-tuning samples. These samples can be labeled using a lightweight automatic labeling algorithm (labeling accuracy must be ≥95%, samples that do not meet the standard should be manually corrected). The fine-tuning samples and their corresponding labeling are then integrated into the fine-tuning sample set. During the joint fine-tuning process, the fine-tuning parameters are automatically set without manual intervention. The fine-tuning parameters can be set as follows: 1) Optimizer: Use the Adam optimizer from the pre-training stage, and lower the learning rate to 0.0001 (to avoid over-fine-tuning and damaging the basic model); 2) Number of iterations: 5-10 rounds (only to correct the bias, no need for a large number of iterations); 3) Loss function: consistent with the pre-training stage; 4) Early stopping strategy: if the accuracy of the validation set of the fine-tuned samples does not improve for two consecutive rounds, stop fine-tuning immediately to avoid overfitting.

[0091] When fine-tuning parameters of the fine-tuning layers of several fusion analysis branches using the fine-tuning dataset, if there is mutual reuse of branch output results among the fusion analysis branches, it is necessary to determine the fine-tuning priority of each fusion analysis branch according to the reuse mapping rules, and then fine-tune the fusion analysis branches in sequence according to the fine-tuning priority.

[0092] In one example, the parallel fusion analysis branches include anomaly detection, object detection, and semantic segmentation. According to the reuse mapping rules, the branch outputs of the object detection and semantic segmentation branches can assist the anomaly detection branch in data analysis, while the branch outputs of the semantic segmentation branch can assist the object detection branch in data analysis. Therefore, the fine-tuning priority of the three fusion analysis branches from high to low is the semantic segmentation branch, the object detection branch, and the anomaly detection branch.

[0093] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments.

[0094] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any other combination thereof. When implemented using a software program, it can be implemented entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0095] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. An AI-based intelligent analysis method for low-altitude remote sensing data from unmanned aerial vehicles (UAVs) based on multimodal fusion, characterized in that: include: Several fusion analysis branches are obtained based on the data analysis requirements; Several of the aforementioned fusion analysis branches are mounted in parallel on the output layer of a shared feature component; The shared feature component is used to extract a multimodal feature set from the remote sensing fusion data; the multimodal feature set is analyzed using several fusion analysis branches to achieve automated analysis of the remote sensing fusion data.

2. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 1, characterized in that, Several of the aforementioned fusion analysis branches are mounted in parallel on the output layer of a shared feature component, including: The output layer of the shared feature component is provided with a standardized output interface, and the feature input layer of the fusion analysis branch is preset with an adaptation interface that matches the standardized output interface. The fusion analysis branches are mounted in parallel based on the standardized output interface and the adapter interface; wherein, the fusion analysis branches mounted in parallel on the shared feature component output layer can be dynamically adjusted.

3. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 1, characterized in that, The shared feature component includes an input layer, a single-modal feature extraction layer, a cross-modal feature fusion layer, a feature optimization layer, and an output layer.

4. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 1, characterized in that, The multimodal feature set includes basic general features and branch-specific features; the basic general features include spatial geometric features, global texture features, modality fusion weight features, and feature quality features; the branch-specific features correspond to each fusion analysis branch and are used to meet the analysis requirements of the fusion analysis branch.

5. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 1, characterized in that, The branch outputs of the several fusion analysis branches mounted in parallel are reused, including: The branch output results of the fusion analysis branch are stored in the branch result shared library; Each of the fusion analysis branches calls the branch output results from the branch result sharing library according to the pre-built reuse mapping rules, so as to realize the mutual reuse of branch output results.

6. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 1, characterized in that, Based on the fusion feature data and the data analysis requirements, a fusion analysis branch is matched, including: Identify the fusion feature data of the remote sensing fusion data; wherein, the fusion feature data is used to determine whether the fusion analysis branch is compatible with the remote sensing fusion data; Several fusion analysis branches are obtained by matching the data analysis requirements; Based on the data matching of several fusion analysis branches, the data matching results of the fusion feature data analysis are used to select several fusion analysis branches.

7. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 6, characterized in that, The fusion feature data for identifying the remote sensing fusion data includes: Call a pre-trained lightweight CNN model; the lightweight CNN model includes an input layer, lightweight convolutional layers, pooling layers, fully connected layers, and an output layer; The lightweight CNN model is used to identify attribute features of the remote sensing fusion data to obtain fusion feature data; wherein, the fusion feature data includes modality composition features, data quality features, dominant feature types and scene adaptation features.

8. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion according to claim 1, characterized in that, Joint fine-tuning of several fusion analysis branches mounted in parallel includes: Select a fine-tuning dataset from the remote sensing fusion data to be processed; The fine-tuning dataset is used to fine-tune the parameters of the fine-tuning layers of several fusion analysis branches; wherein the fine-tuning layer includes a feature input layer, an attention module, or an output layer.

9. The AI-powered intelligent analysis method for UAV low-altitude remote sensing data based on multimodal fusion as described in claim 8, characterized in that, Using the fine-tuning dataset, the parameters of the fine-tuning layers of several fusion analysis branches are fine-tuned, including: Retrieve and reuse mapping rules; According to the reuse mapping rules, fine-tuning priorities are set for several fusion analysis branches, and the corresponding fusion analysis branches are fine-tuned sequentially according to the fine-tuning priorities.