Unmanned aerial vehicle multi-modal visual identification method based on cross-scene transfer learning

Through real-time environment perception and scene prediction, select the subset of source domains, dynamically adjust the fusion of multimodal features, and combine time series feature accumulation and phased transfer learning, the problem of decreasing recognition accuracy in drone cross-scene migration is solved, and efficient and adaptive visual recognition is achieved.

CN120339887AActive Publication Date: 2025-07-18ZHEJIANG FULIN TECH CO LTD

Patent Information

Application Number
CN202510791756.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-18
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing multimodal visual recognition methods of drone have problems such as decreasing recognition accuracy, insufficient modal fusion strategies and relying on labeled data in cross-scene migration, making it difficult to achieve efficient adaptive recognition in complex environments.

Method used

The most matching subset of source domains is selected through real-time environment perception and scene prediction, the target domain subdomain is divided based on environmental features, the multimodal feature fusion strategy is dynamically adjusted, and the visual recognition model is trained using staged transfer learning, combining time series feature accumulation and historical state correction to achieve cross-scene migration.

Benefits of technology

It improves the recognition accuracy and stability of the drone in complex environments, enhances its adaptability to conditions such as drastic lighting changes and severe weather, reduces negative migration risks, and improves the robustness and generalization capabilities of the identification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339887A_ABST
    Figure CN120339887A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to an unmanned aerial vehicle multi-modal visual recognition method for cross-scene transfer learning, which comprises the following steps: synchronously acquiring environmental parameter information and multi-modal image data in real time, and predicting the scene category of a current flight area by using a scene prediction model; selecting the source domain subset with the highest matching degree by calculating the matching degree between the scene category and the source domain subset in the preset source domain data set; dividing the target domain sample into a plurality of sub-domains to obtain target sub-domain data; extracting single-mode features of each image in the multi-mode image data, and performing fusion based on a quality evaluation result of each image and a scene category dynamic switching fusion strategy to generate a first fusion feature; a sliding window mechanism is adopted to accumulate the first fusion features, and correction is carried out according to the historical feature state at the previous moment to obtain second fusion features; and performing visual identification on the second fusion feature of the current target domain by using the visual identification model after the completion of the staged transfer learning training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and specifically to a multi-modal visual recognition method for unmanned aerial vehicles with cross-scene transfer learning. Background Art

[0002] In recent years, with the wide application of deep learning technologies, especially convolutional neural networks (CNNs), Transformer networks, etc. in the field of image recognition, vision-based unmanned aerial vehicle (UAV) perception systems have developed rapidly. By integrating high-definition visible light cameras, infrared sensors, and other imaging devices, UAVs can detect, classify, and identify various objects (such as people, vehicles, infrastructure, disaster traces, etc.) in ground scenes during flight missions. In the prior art, recognition models pre-trained with large-scale datasets (such as ImageNet, COCO, DOTA), as well as object detection (such as YOLO, Faster R-CNN) and multi-modal fusion methods (such as visible light and infrared fusion recognition), have been applied to various scenarios such as disaster rescue, urban patrol, and agricultural monitoring. Especially in aspects of image feature extraction, feature fusion, and discriminant classification, existing methods can achieve high recognition accuracy under standard environmental conditions, providing important support for the intelligent task execution of UAVs.

[0003] However, existing image recognition-based methods still face significant challenges in actual UAV applications. First, most recognition models rely on training samples in specific source domain environments (such as urban roads, regular buildings, etc.), lacking the cross-scene transfer ability for different flight scenarios (such as mountains, farmlands, post-disaster ruins), resulting in a significant decrease in the accuracy of target detection and recognition in new environments. Second, due to the variability of the flight environment (such as direct sunlight, haze occlusion, low illuminance at night, sandstorm interference, etc.), there are significant differences in the quality and effective information content of different modal imaging data. Existing multi-modal image fusion methods usually adopt a fixed weighting strategy and are difficult to dynamically adjust the modal contribution ratio according to environmental conditions, resulting in insufficient recognition robustness in extreme environments. In addition, traditional transfer learning methods mostly rely on supervised learning, requiring a large number of labeled samples in the new scene, which restricts the ability of UAV vision recognition systems to be quickly deployed and adaptively optimized in practical applications. Therefore, there is an urgent need for a multi-modal visual recognition method for UAVs that can dynamically adjust the feature fusion strategy in combination with environmental perception under the condition of lacking labeled data and achieve efficient cross-scene transfer.

[0004] Therefore, a multi-modal visual recognition method for UAVs with cross-scene transfer learning is proposed. Summary of the Invention

[0005] The object of the present invention is to provide a multi-modal visual recognition method for drones with cross-scenario transfer learning, which can dynamically adjust the feature fusion strategy in combination with environmental perception under the condition of lack of labeled data and achieve efficient cross-scenario transfer.

[0006] To achieve the above object, the present invention provides the following technical solutions: A multi-modal visual recognition method for drones with cross-scenario transfer learning, comprising: Real-time synchronously collecting environmental parameter information and multi-modal image data through sensors carried by the drone, and predicting the scene category of the current flight area by using a scene prediction model; By calculating the matching degree between the scene category and the source domain subsets in the preset source domain dataset, selecting the source domain subset with the highest matching degree; dividing the target domain samples into multiple sub-domains based on the environmental feature information of the target domain samples to obtain target sub-domain data; Respectively extracting the single-modal features of each modal image in the multi-modal image data, and dynamically switching the fusion strategy of the single-modal features based on the quality evaluation results and scene categories of each modal image, and fusing the single-modal features to generate a first fusion feature; For multiple frames of the first fusion feature continuously collected during flight, adopting a sliding window mechanism to accumulate time series features and correcting according to the historical feature state at the previous moment to obtain a second fusion feature; Based on the source domain subset and the target sub-domain data, performing staged transfer learning to train a visual recognition model and using the trained visual recognition model to perform visual recognition on the second fusion feature of the current target domain.

[0007] Preferably, the environmental parameter information includes at least one of light intensity, visibility, temperature, humidity and wind speed; the multi-modal image data includes visible light images, infrared images and lidar point clouds; The scene prediction model includes: an environmental parameter processing unit, a preliminary image feature extraction unit, a multi-source information fusion unit and a scene classification prediction unit; The environmental parameter processing unit normalizes and / or standardizes the environmental parameter information, encodes categorical parameters, and outputs an environmental feature vector; The preliminary image feature extraction unit extracts visual feature vectors by using a lightweight feature extraction network based on the multi-modal image data; the visual feature vectors contain scene content information; The multi-source information fusion unit fuses the environmental feature vector and the visual feature vector to obtain a scene representation feature; The scene classification prediction unit calculates the probabilities of the scene representation features corresponding to each predefined scene category through M fully connected layers and a Softmax activation function, and outputs the scene category with the highest probability as the prediction result.

[0008] Preferably, the process of obtaining the source domain subset includes: Pre-annotate the scene category to which each source domain sample belongs in the source domain dataset; According to the current scene category predicted by the scene prediction model, calculate the matching degree between the current scene category and each scene category in the source domain dataset using the Mahalanobis distance; Sort in descending order according to the matching degree, and select the source domain sample set with a matching degree exceeding the set threshold as the source domain subset, while excluding the source domain samples with a matching degree lower than the set threshold.

[0009] Preferably, the process of obtaining the target sub-domain data includes: Collect the unlabeled sample data of the current target domain; according to the environmental feature vector synchronously obtained with the unlabeled sample data, use a clustering algorithm to perform clustering analysis on the target domain samples to obtain N clustering clusters; the number of clustering clusters is adaptively adjusted according to the scene category; Divide the target domain samples belonging to the same clustering cluster into one sub-domain to obtain N pieces of the target sub-domain data.

[0010] Preferably, the process of generating the first fusion feature includes: Use a feature extraction network to extract the depth features of the visible light image, infrared image, and lidar point cloud in the multi-modal image data as single-modal features; By calculating the clarity score, signal-to-noise ratio, contrast, and point cloud density index of the visible light image, infrared image, and lidar point cloud in the multi-modal image data, obtain the quality evaluation results of each modal image; Based on the real-time environmental parameters, the predicted scene category, and the modal image quality evaluation results, set a fusion strategy switching rule table, dynamically select and execute one and / or more feature fusion strategies to fuse the single-modal features to generate the first fusion feature.

[0011] Preferably, the process of obtaining the second fusion feature includes: Set a time sliding window to accumulate the features of K consecutive frames of the first fusion feature; Introduce the historical feature state to correct the first fusion feature within the current time window to obtain the second fusion feature; the correction is based on an error compensation mechanism between the first fusion feature at the previous moment and the first fusion feature at the current moment.

[0012] Preferably, the visual recognition model is formed during the phased transfer learning training process and is ultimately used to perform specific recognition tasks on the second fusion feature. Specifically, it includes: a feature processing and local adaptation unit, a task recognition head unit, and a visual recognition output unit; The feature processing and local adaptation unit includes a feature transformation layer and a local adaptation layer; the feature transformation layer contains a series of neural network layers for non-linearly processing the input second fusion feature; the local adaptation layer aligns the source domain subset with each target sub-domain based on the least mean square error loss; The task recognition head unit is connected after the feature processing and local adaptation unit, and its specific structure is determined according to the final visual recognition task; the visual recognition task includes image classification tasks, object detection tasks, and other recognition tasks; The visual recognition output unit is connected after the task recognition head unit and is used to output the visual recognition result.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By introducing an environmental perception and scene prediction mechanism, the present invention enables the drone to understand the specific environment and macro scene type it is currently in in real time. This information is critically applied in multiple subsequent links: First, when selecting the source domain for transfer learning, instead of blindly using the entire source data set, the most matching source domain subset is preferentially selected according to the predicted scene category, ensuring the relevance of knowledge transfer; Second, when processing target domain data, the target domain is divided into sub-domains based on environmental features, achieving more refined local environmental adaptation; Finally, when fusing multi-modal features, instead of adopting a fixed fusion strategy, the fusion method is dynamically switched or adjusted according to real-time environmental parameters, scene categories, and the quality evaluation results of each modal data itself. This full-process adaptive adjustment driven by environmental and scene perception enables the visual recognition model to better adapt to complex conditions such as drastic changes in lighting, bad weather, and different landforms. Compared with traditional methods that use fixed models and fixed fusion strategies, the method of the present invention shows higher recognition accuracy and stronger stability in various non-ideal or dynamically changing actual flight environments.

[0014] 2. On the one hand, through scene prediction and matching degree calculation, the present invention only screens out the source domain subset most relevant to the current target scene for migration, avoiding the interference of irrelevant source domain data. On the other hand, the present invention recognizes that there may also be significant environmental differences within the target domain. Innovatively, the target domain is divided into multiple more homogeneous sub-domains based on environmental characteristics, and staged and localized adaptation training is carried out. This strategy of "accurate source selection" and "fine target division + local adaptation" significantly reduces the distribution difference between the source domain and the target sub-domains, making the transfer learning more focused and effective, maximizing the utilization of the effective knowledge of the source domain, and avoiding the negative impacts that may be brought by overall adaptation, thereby improving the recognition accuracy and overall performance of the subsequent visual recognition model on each sub-domain of the target domain.

[0015] 3. Based on the real-time quality assessment and scene context of each modality, the present invention dynamically adjusts the fusion strategy, which ensures that the information of high-quality modalities is fully utilized and the influence of low-quality or unreliable modalities is suppressed, thereby generating a more robust "first fusion feature". In view of the continuity characteristics of UAV flight, the present invention also introduces a time series feature accumulation and optimization mechanism. By sliding window to accumulate the first fusion features of multiple frames and using the historical feature states for correction, the mutation and noise of single-frame features are smoothed, effectively reducing the influence of instantaneous interferences such as motion blur and short-term occlusion on the recognition results, and generating a "second fusion feature" that is more stable in time and can better reflect the scene dynamics. This feature representation method combining modality quality-aware fusion and temporal optimization significantly improves the accuracy and reliability of the final visual recognition model in the case of continuous flight and interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of a cross-scene transfer learning-based UAV multi-modal visual recognition method provided by an embodiment of the present invention; Figure 2 It is a schematic flowchart of generating the second fusion feature provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a visual recognition model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] With the wide application of drones in fields such as security patrol, disaster monitoring, and environmental perception, the complex and variable environment poses higher requirements for their visual recognition capabilities. Traditional visual recognition methods are prone to problems such as a decline in recognition accuracy when facing scene changes such as different lighting, climate, or terrain, especially when the data distribution changes and the model generalization ability is limited. The cross-scene transfer learning-based multi-modal visual recognition method for drones can break through the bottleneck of the sharp drop in recognition accuracy of traditional models in new environments, make full use of the complementary advantages of multi-modal data such as visible light, infrared, and laser point cloud, and achieve robust recognition of targets in different geographical regions, weather conditions, and lighting environments by dynamically adapting the source domain knowledge and target domain features, which is of great significance for improving the intelligent perception ability of drones in complex real-world scenarios such as disaster monitoring, security patrol, and traffic supervision.

[0019] The present invention proposes a cross-scene transfer learning-based multi-modal visual recognition method for drones, which can dynamically adjust the feature fusion strategy in combination with environmental perception and achieve efficient cross-scene transfer under the condition of lack of labeled data. In order to illustrate that the method of the present invention can play a role in realizing efficient cross-scene transfer, the effectiveness of the present invention will be described from two embodiments below.

[0020] Embodiment 1 In the embodiment of the present application, the method proposed by the present invention is used to elaborate in detail the process of dynamically adjusting the feature fusion strategy in combination with environmental perception and achieving efficient cross-scene transfer under the condition of lack of labeled data. The embodiment of the present application is applicable to the multi-modal visual recognition of drones in a continuous flight environment with dynamic scene changes or interference. The following will elaborate on the multi-modal visual recognition process of drones according to Figure 1 the content; among them, Figure 1The specific flowchart of the method proposed by the present invention includes: real-time synchronously collecting environmental parameter information and multi-modal image data through sensors carried by a drone, and predicting the scene category of the current flight area using a scene prediction model; calculating the matching degree between the scene category and source domain subsets in a preset source domain dataset, and selecting the source domain subset with the highest matching degree; dividing target domain samples into multiple sub-domains based on the environmental feature information of the target domain samples to obtain target sub-domain data; respectively extracting the single-modal features of each modal image in the multi-modal image data, and dynamically switching the fusion strategy of the single-modal features based on the quality evaluation results and scene category of each modal image, and fusing the single-modal features to generate a first fusion feature; for multiple frames of the first fusion feature continuously collected during flight, adopting a sliding window mechanism to accumulate time series features and correcting according to the historical feature state at the previous moment to obtain a second fusion feature; based on the source domain subset and the target sub-domain data, performing staged transfer learning to train a visual recognition model and using the trained visual recognition model to perform visual recognition on the second fusion feature of the current target domain. In combination with Figure 1 the content in Real-time synchronously collecting environmental parameter information and multi-modal image data through sensors carried by a drone; The environmental parameter information includes at least one of light intensity, visibility, temperature, humidity, and wind speed; The multi-modal image data includes visible light images, infrared images, and lidar point clouds; Specifically, a visible light camera (resolution 1920×1080, frame rate 30fps), a thermal infrared camera (resolution 640×512, frame rate 30fps), a lidar (16 lines, scanning frequency 10Hz), and an environmental sensor kit are carried. The kit includes a digital illuminance meter (measurement range 0-200kLux), a visibility sensor, an integrated temperature and humidity sensor, and a three-dimensional wind speed and direction meter. The controller synchronously collects the data streams of all sensors in real time. The collected environmental parameter information and preliminary multi-modal image data are sent into a pre-trained lightweight scene prediction model.

[0021] Preferably, predicting the scene category of the current flight area using a scene prediction model; The scene prediction model includes: an environmental parameter processing unit, a preliminary image feature extraction unit, a multi-source information fusion unit, and a scene classification prediction unit; the environmental parameter processing unit performs normalization and / or standardization processing on the environmental parameter information, encodes categorical parameters, and outputs an environmental feature vector; The preliminary image feature extraction unit extracts a visual feature vector using a lightweight feature extraction network based on the multi-modal image data; the visual feature vector contains scene content information; The multi-source information fusion unit fuses the environmental feature vector and the visual feature vector to obtain a scene representation feature; The scene classification prediction unit calculates the probability that the scene representation feature corresponds to each predefined scene category through M fully connected layers and a Softmax activation function, and outputs the scene category with the highest probability as a prediction result.

[0022] Specifically, the preliminary image feature extraction unit uses MobileNetV3-Small as a backbone network to process the input multimodal image data respectively and extract visual feature vectors containing scene information; The multi-source information fusion unit concatenates the environmental feature vector and the visual feature vector, and performs fusion and dimensionality reduction through one or two fully connected layers to obtain scene representation features; The scene classification prediction unit consists of three fully connected layers, with the number of neurons being 128, 64 and 9 respectively (corresponding to 9 predefined scene categories). The last layer uses the Softmax activation function to calculate the probability that the current scene belongs to each predefined scene category, and outputs the scene category with the highest probability as the prediction result; the predefined categories include but are not limited to: city-clear, city-rainy, city-night, suburb-clear, suburb-rainy, suburb-night, mountain-fog, mountain-rainy and post-disaster-smoke.

[0023] By integrating multiple environmental sensors and multimodal imaging devices, the physical environment and visual scene of the drone are fully perceived. The lightweight scene prediction model containing a specific processing unit can efficiently integrate environmental parameters and preliminary visual information in real time, and accurately and quickly predict the current complex flight scene category. It provides a key basis for the subsequent adaptive source domain selection, target domain division and feature fusion according to the scene, improves the environmental adaptability and response speed in rapidly changing scenes, and ensures the feasibility of airborne computing.

[0024] Preferably, by calculating the matching degree between the scene category and the source domain subset in the preset source domain data set, the source domain subset with the highest matching degree is selected; the process of obtaining the source domain subset includes: Pre-label the scene category to which each source domain sample belongs in the source domain dataset; According to the current scene category predicted by the scene prediction model, the Mahalanobis distance is used to calculate the matching degree between the current scene category and each scene category in the source domain dataset; The matching degree is sorted from high to low, and a set of source domain samples whose matching degree exceeds a set threshold is selected as the source domain subset, while source domain samples whose matching degree is lower than the set threshold are eliminated.

[0025] Specifically, a source domain dataset containing multiple scene categories is pre-constructed. In this embodiment, the source domain dataset contains 100,000 multi-modal images, covering 9 scene categories. Each source domain sample is labeled with the belonging scene category and recognition task label; the scene category labeling is completed by combining manual labeling and automatic matching to ensure the labeling accuracy; When the drone enters a new environment, the current scene category is output through the scene prediction model; assuming that the current predicted scene category is "city - rainy day", then it is necessary to select a sample subset with a high matching degree with the "city - rainy day" scene from the source domain dataset as the knowledge source for transfer learning; The Mahalanobis distance is used to calculate the matching degree between the current scene category and each scene category in the source domain dataset; Sort the calculated matching degrees from high to low, set the matching degree threshold to 0.7; select the source domain sample set with a matching degree exceeding the matching degree threshold as the source domain subset; for example, for the "city - rainy day" scene, the matching degrees from high to low are: city - rainy day (0.92), suburban - rainy day (0.78), city - sunny (0.71), suburban - night (0.65), etc.; select the samples of the top three scene categories with a matching degree ≥ 0.7 to form the source domain subset; at the same time, eliminate the source domain samples with a matching degree lower than the threshold.

[0026] By pre-labeling the source domain scene categories and using the Mahalanobis distance to calculate the matching degree between the predicted target scene and the source domain scenes, the matching degree between scenes can be measured more accurately. Setting the threshold further ensures that only highly relevant source domain samples are selected. This source domain subset selection strategy based on scene matching effectively avoids introducing irrelevant or conflicting source domain knowledge, significantly reduces the risk of negative transfer, improves the pertinence and efficiency of transfer learning, especially when the scene differences are large, the recognition accuracy is significantly improved. Table 1 is a comparison table of recognition accuracies of different methods in different scenes.

[0027] Table 1 Comparison table of recognition accuracies in different scenes Test scenario Accuracy of the comparison method (%) (fixed model / fusion) Accuracy of the method of the present invention (%) Accuracy improvement (%) City - sunny 92.5 94.8 2.3 City - night 75.3 91.2 15.9 Suburb - rainy day 78.1 89.5 11.4 Mountain - foggy day 65.7 85.1 19.4 Post - disaster - smoky dust 70.1 89.6 19.5 Preferably, based on the environmental feature information of the target domain samples, the target domain samples are divided into multiple sub-domains to obtain target sub-domain data; the process of obtaining the target sub-domain data includes: collecting unlabeled sample data of the current target domain; according to the environmental feature vectors synchronously obtained with the unlabeled sample data, using a clustering algorithm to perform clustering analysis on the target domain samples to obtain N clustering clusters; the number of the clustering clusters is adaptively adjusted according to the scene category; The target domain samples belonging to the same clustering cluster are divided into one sub-domain to obtain N pieces of the target sub-domain data.

[0028] Specifically, after the drone enters the target domain environment, continuously collect unlabeled sample data of the current target domain; According to the flight speed and coverage range of the unmanned aerial vehicle, 5000 - 10000 multi-modal images are set to be collected as unlabeled samples in the target domain environment. Meanwhile, the environmental feature vectors obtained synchronously with each image are recorded; The K-Means++ clustering algorithm is used to perform clustering analysis on the target domain samples: First, calculate the similarity matrix between samples based on the environmental feature vectors; then, identify the density peak points as the clustering centers; finally, assign the remaining points to the nearest clustering center to form clustering clusters; the number N of clustering clusters is adaptively adjusted according to the scene category. For example, for the "city - sunny" scene, there may be various lighting changes, and N = 5 is set; for the "suburb - night" scene, the environment is relatively simple, and N = 2 is set. The specific adjustment rules are as follows: Urban scene: N = 5 + amplitude of lighting change / 20; Suburban scene: N = 2 + amplitude of lighting change / 30; Mountain scene: N = 2 + visibility change / 5; Rainy day scene: N = 2 + rainfall level; Among them, the unit of the amplitude of lighting change is klux, the unit of visibility change is km, and the rainfall level is an integer from 0 to 3.

[0029] The target domain samples belonging to the same clustering cluster are divided into a sub-domain, and finally N target sub-domain data are obtained.

[0030] Table 2 is a comparison table of the recognition accuracy of the target domain in transfer learning for different methods.

[0031] Table 2 Comparison Table of Recognition Accuracy of Target Domain in Transfer Learning Target domain scenario / sub - domain Accuracy of the comparison method (%) (global migration) Accuracy of the method of the present invention (%) Target domain: Mixed lighting in the city 80.5 - Sub - domain 1: Strong light - 91.3 Sub - domain 2: Weak light and shadow - 86.8 Target domain: Rural foggy day 72.3 - Sub - domain 1: Light fog - 89.5 Sub - domain 2: Dense fog - 92.3 In this embodiment, through the adaptive clustering method driven by environmental features, fine-grained division of the target domain samples is achieved, and the data distribution differences under different environmental conditions within the same scene can be captured. At the same time, the number N of clustering clusters is adaptively adjusted according to the scene category, so that the division granularity can match the complexity of the scene. Compared with the traditional method that regards the target domain as a whole, this method fully considers the influence of environmental factors on visual features, finely depicts the environmental diversity within the target domain, makes the subsequent localized domain adaptation process more accurate, and helps to reduce the interference caused by intra-domain differences.

[0032] Preferably, the single-modal features of each modal image in the multi-modal image data are respectively extracted, and the fusion strategy of the single-modal features is dynamically switched based on the quality evaluation results of each modal image and the scene category, and the single-modal features are fused to generate a first fusion feature; The generation process of the first fusion feature includes: The feature extraction network is used to extract the depth features of visible light images, infrared images, and lidar point clouds in the multimodal image data as single-modal features respectively; By calculating the clarity score, signal-to-noise ratio, contrast, and point cloud density index of the visible light image, infrared image, and lidar point cloud in the multimodal image data, the quality evaluation results of each modal image are obtained; Based on the real-time environmental parameters, predicted scene categories, and modal image quality evaluation results, a fusion strategy switching rule table is set, and one and / or more feature fusion strategies are dynamically selected and executed to fuse the single-modal features to generate the first fusion feature.

[0033] Specifically, the pre-trained EfficientNet-B3 is used to extract the features of the visible light image (output dimension is 1536), the improved ResNet-34 is used to extract the features of the infrared image (output dimension is 512), and PointNet++ is used to extract the features of the lidar point cloud (output dimension is 1024); each modal feature passes through an adaptive pooling layer and is uniformly adjusted to a 512-dimensional feature vector; Calculate the quality evaluation indicators of each modal image: for the visible light image, calculate the clarity score of the Laplacian gradient method, the signal-to-noise ratio based on wavelet transform, and the RMS contrast to obtain the quality evaluation result of the visible light image; for the infrared image, calculate the clarity index and contrast index based on the thermal gradient to obtain the quality evaluation result of the infrared image; for the lidar point cloud, calculate the point cloud density and point cloud integrity index to obtain the quality evaluation result of the lidar point cloud. Normalize the quality evaluation results of each modality to form a modal quality score vector; Based on the real-time environmental parameters, predicted scene categories, and modal image quality evaluation results, a fusion strategy switching rule table is set, and a feature fusion strategy is dynamically selected.

[0034] This embodiment realizes a dynamic feature fusion strategy switching mechanism based on environmental parameters, scene categories, and modal quality, and solves the problem that traditional fixed fusion strategies are difficult to cope with changing environments. The system can adaptively select the optimal fusion strategy according to the current environmental conditions and the quality of each modal data, give full play to the advantages of each modality, maximize the utilization of effective information, suppress noise interference, and generate a first fusion feature with higher quality and stronger environmental adaptability.

[0035] Preferably, for multiple frames of the first fusion feature continuously collected during flight, a sliding window mechanism is adopted for time series feature accumulation, and correction is performed according to the historical feature state at the previous moment to obtain a second fusion feature; The process of obtaining the second fusion feature includes: Set a time sliding window to accumulate features of the continuously acquired K frames of the first fused feature; Introduce the historical feature state to correct the first fused feature within the current time window to obtain the second fused feature; the correction is based on an error compensation mechanism between the first fused feature at the previous moment and the first fused feature at the current moment.

[0036] Specifically, set the sliding window size K = 5, corresponding to approximately 0.25 seconds of data continuously acquired by the drone at a sampling frequency of 20 Hz. The sliding window moves forward by 1 frame each time, keeping the window size unchanged.

[0037] Store the first fused feature sequence of the most recent K frames, accumulate the features within the window, and calculate the weighted average, where the weight decays exponentially with time; Calculate the difference between the current frame and the first fused feature of the previous frame, and correct the first fused feature within the current time window through a correction mechanism based on error compensation to obtain the second fused feature.

[0038] Table 3 gives a comparison table of the recognition accuracies of different methods under different interferences.

[0039] Table 3 Comparison table of recognition accuracies under different interferences Test conditions (continuous flight) Interference type Accuracy of the comparison method (%) (simple fusion + single - frame recognition) Accuracy of the method of the present invention (%) Normal flight No obvious interference 93.1 97.5 Quick turn Motion blur (partial frames) 76.5 91.3 Low - altitude flight through branches Short - term occlusion (partial frames) 72.8 92.7 Weak infrared signal at night Single - modality quality degradation 79.4 89.1 Sparse lidar point cloud on rainy days Single - modality quality degradation 75.6 88.2 In this embodiment, multi-frame features are accumulated through a time sliding window, integrating the temporal information within a short period and enhancing the stability of the features. Introducing a correction mechanism based on the feature error compensation between consecutive frames can effectively identify and suppress the drastic feature changes caused by single-frame anomalies (such as jitter blur and short-term occlusion), while retaining the true dynamic changes of the scene. This makes the generated second fused feature smoother in time, significantly improving the reliability and accuracy of visual recognition under continuous flight and instantaneous interference conditions.

[0040] Preferably, based on the source domain subset and the target sub-domain data, perform staged transfer learning to train a visual recognition model and use the trained visual recognition model to perform visual recognition on the second fused feature of the current target domain; refer to Figure 3 ; The visual recognition model is formed during the staged transfer learning training process and is ultimately used to perform specific recognition tasks on the second fused feature, specifically including: a feature processing and local adaptation unit, a task recognition head unit, and a visual recognition output unit; The feature processing and local adaptation unit includes a feature transformation layer and a local adaptation layer; The feature transformation layer contains a series of neural network layers for non-linearly processing the input second fused feature; The local adaptation layer aligns the source domain subset with each target sub - domain based on the least mean square error loss; The task recognition head unit is connected after the feature processing and local adaptation unit, and its specific structure is determined according to the final visual recognition task; the visual recognition tasks include image classification tasks, object detection tasks, and other recognition tasks; The visual recognition output unit is connected after the task recognition head unit and is used to output the visual recognition result.

[0041] Specifically, the specific structure of the feature processing and local adaptation unit includes: Feature transformation layer: It contains three residual blocks, and each residual block consists of two convolutional layers and a short - circuit connection. The convolutional kernel size is 3×3, and the number of channels is 512, 512, and 512 in sequence. After each convolutional layer, a batch normalization layer and a ReLU activation function are connected to perform a non - linear transformation on the input second - fused feature to enhance the feature expression ability.

[0042] Local adaptation layer: It adopts a multi - kernel maximum average pooling structure, uses kernels of different scales (1×1, 3×3, 5×5) to extract multi - scale features, and then calculates the weighted sum of the maximum pooling result and the average pooling result. A fully - connected layer is connected after this layer, and the output dimension is 256, which is used for feature alignment between the source domain subset and each target sub - domain.

[0043] The structure of the task recognition head unit is determined according to the specific visual recognition task: Image classification task: It contains two fully - connected layers, with the number of neurons being 128 and the number of classes C (C = 9) respectively. The last layer uses the Softmax activation function to output the class probability distribution; Object detection task: It adopts an improved FPN structure, which contains a feature pyramid network and two parallel branches, including a classification branch and a bounding box regression branch; Other recognition tasks (such as scene segmentation, anomaly detection, etc.): Customize the corresponding network structure according to the specific task requirements.

[0044] Visual recognition output unit: Connected after the task recognition head unit, it outputs the corresponding recognition result according to the task type. For the classification task, it outputs the class label and confidence; for the detection task, it outputs the target position, class, and confidence.

[0045] The staged transfer learning training process includes three stages: source domain pre - training, local domain adaptation, and target domain fine - tuning.

[0046] The visual recognition model architecture designed in this embodiment fully considers the requirements of feature processing, domain adaptation, and task recognition, and realizes knowledge transfer from the source domain to the target domain through a phased transfer learning strategy. The feature processing and local adaptation unit enhances the feature expression ability and realizes feature alignment between different sub-domains; the task recognition head unit is customized according to specific tasks, improving the task specificity of the model; the phased training and local adaptation loss for each target sub-domain ensure that the model can accurately adapt to the unique data distributions of each sub-domain within the target domain caused by environmental changes. This overcomes the limitations of traditional global domain adaptation methods, significantly improves the comprehensive recognition performance of the model on the entire heterogeneous target domain, and makes the final visual recognition results more accurate and reliable.

[0047] A multi-modal visual recognition method for unmanned aerial vehicles based on cross-scene transfer learning provided by the present invention significantly improves the visual recognition robustness and generalization ability of unmanned aerial vehicles in complex and changing environments by integrating mechanisms such as environmental perception, scene prediction, modal dynamic weighing, and time series modeling. The method first uses the unmanned aerial vehicle sensor to synchronously collect environmental parameters and multi-modal image data in real time, and combines the environmental parameters and preliminary image features for scene prediction, so as to dynamically select the source domain subset that best matches the current scene, improving the pertinence and effectiveness of the initial transfer starting point. At the same time, a target domain sub-domain division mechanism based on environmental features is adopted to construct a finer-grained sub-domain structure, and the transfer process is optimized by combining local adaptation, effectively suppressing the risk of negative transfer. In terms of feature processing, aiming at the differences and environmental adaptability of multi-modal images, a fusion strategy dynamic switching mechanism that combines quality assessment and scene categories is proposed to achieve the optimal combination of multi-modal features in different scenes. In addition, aiming at the temporal disturbance problems such as image blurring and occlusion during the flight of the unmanned aerial vehicle, a sliding window accumulation and historical state correction mechanism is designed to significantly improve the consistency and recognition stability of feature expression in consecutive frames. Finally, the phased transfer learning is performed by combining the source domain subset and the target sub-domain data to train the visual recognition model, effectively improving the recognition accuracy and generalization ability of the recognition model in new scenes. The overall solution has the significant advantages of reasonable structure, strong adaptability, and high robustness, and is applicable to the intelligent unmanned aerial vehicle visual perception system in dynamic task scenarios.

[0048] Embodiment 2 In Embodiment 1, the method proposed by the present invention successfully realizes the dynamic adjustment of the feature fusion strategy in the absence of labeled data and achieves efficient cross-scene transfer. To further verify the effectiveness of the present invention, in the embodiments of the present application, multi-modal visual recognition is also performed on another unmanned aerial vehicle in a continuous flight environment with dynamic scene changes or interference.

[0049] A multi-modal visual recognition method for unmanned aerial vehicles based on cross-scene transfer learning includes: The environmental parameter information and multi-modal image data are collected in real-time and synchronously by sensors carried on the UAV. The environmental parameter information includes at least one of light intensity, visibility, temperature, humidity, and wind speed; the multi-modal image data includes visible light images, infrared images, and lidar point clouds. Preferably, a scene prediction model is used to predict the scene category of the current flight area. The scene prediction model includes an environmental parameter processing unit, a preliminary image feature extraction unit, a multi-source information fusion unit, and a scene classification prediction unit. The environmental parameter processing unit performs normalization and / or standardization processing on the environmental parameter information, encodes categorical parameters, and outputs an environmental feature vector; the preliminary image feature extraction unit extracts visual feature vectors based on the multi-modal image data using a lightweight feature extraction network; the visual feature vector contains scene content information; the multi-source information fusion unit fuses the environmental feature vector and the visual feature vector to obtain a scene representation feature; the scene classification prediction unit calculates the probabilities of the scene representation feature corresponding to each predefined scene category through M fully connected layers and a Softmax activation function, and outputs the scene category with the highest probability as the prediction result.

[0050] Preferably, by calculating the matching degree between the scene category and the source domain subsets in the preset source domain dataset, the source domain subset with the highest matching degree is selected. The acquisition process of the source domain subset includes: Each source domain sample in the source domain dataset is pre-annotated with the scene category it belongs to. According to the current scene category predicted by the scene prediction model, the Mahalanobis distance is used to calculate the matching degree between the current scene category and each scene category in the source domain dataset. Sorted from high to low according to the matching degree, the source domain sample set with a matching degree exceeding the set threshold is selected as the source domain subset, and at the same time, the source domain samples with a matching degree lower than the set threshold are excluded.

[0051] Preferably, based on the environmental feature information of the target domain samples, the target domain samples are divided into multiple sub-domains to obtain target sub-domain data; the acquisition process of the target sub-domain data includes: Collect the unlabeled sample data of the current target domain; according to the environmental feature vector synchronously obtained with the unlabeled sample data, the target domain samples are clustered and analyzed using a clustering algorithm to obtain N clustering clusters; the number of clustering clusters is adaptively adjusted according to the scene category. The target domain samples belonging to the same clustering cluster are divided into one sub-domain to obtain N pieces of the target sub-domain data.

[0052] Preferably, the unimodal features of each modality image in the multimodal image data are extracted respectively, and the fusion strategy of the unimodal features is dynamically switched based on the quality evaluation results and scene categories of each modality image, and the unimodal features are fused to generate a first fusion feature; The generation process of the first fusion feature includes: The depth features of the visible light image, infrared image and lidar point cloud in the multimodal image data are extracted respectively by using a feature extraction network as unimodal features; By calculating the clarity score, signal-to-noise ratio, contrast and point cloud density indexes of the visible light image, infrared image and lidar point cloud in the multimodal image data, the quality evaluation results of each modality image are obtained; Based on the real-time environmental parameters, predicted scene categories and modality image quality evaluation results, a fusion strategy switching rule table is set, and one and / or more feature fusion strategies are dynamically selected and executed to fuse the unimodal features to generate the first fusion feature.

[0053] Part of the fusion strategy switching rule table is shown in Table 4 for example.

[0054] Table 4 Partial example of fusion strategy switching rule table Scene category Environmental conditions Modality quality conditions Fusion strategy City - sunny Illumination > 80 klux Visible light clarity > 0.8 Weighted fusion (visible light 0.7, infrared 0.1, point cloud 0.2) City - sunny Illumination 20 - 80 klux Visible light clarity > 0.6 Attention mechanism Suburb - night Illumination < 0.1 klux Infrared clarity > 0.7 Weighted fusion Mountain - foggy day Visibility < 0.1 km <![CDATA[Point cloud density > 10pts / m 2 > Feature complementary fusion Any scene Any condition All modal masses are < 0.4 Decision - level fusion For multiple frames of the first fusion feature continuously collected during flight, a sliding window mechanism is used for time series feature accumulation, and correction is performed according to the historical feature state at the previous moment to obtain a second fusion feature; The acquisition process of the second fusion feature includes: A time sliding window is set to perform feature accumulation on K frames of the first fusion feature continuously collected; The historical feature state is introduced to correct the first fusion feature within the current time window to obtain the second fusion feature; the correction is based on an error compensation mechanism between the first fusion feature at the previous moment and the first fusion feature at the current moment.

[0055] Based on the source domain subset and the target subdomain data, phased transfer learning is performed to train a visual recognition model, and the trained visual recognition model is used to perform visual recognition on the second fusion feature of the current target domain.

[0056] The visual recognition model is formed during the staged transfer learning training process and is ultimately used to perform specific recognition tasks on the second fusion feature. Specifically, it includes: a feature processing and local adaptation unit, a task recognition head unit, and a visual recognition output unit; the feature processing and local adaptation unit includes a feature transformation layer and a local adaptation layer; the feature transformation layer includes a series of neural network layers for non-linearly processing the input second fusion feature; the local adaptation layer aligns the source domain subset with each target subdomain based on the least mean square error loss; the task recognition head unit is connected after the feature processing and local adaptation unit, and its specific structure is determined according to the final visual recognition task; the visual recognition tasks include image classification tasks, target detection tasks, and other recognition tasks; the visual recognition output unit is connected after the task recognition head unit and is used to output visual recognition results.

[0057] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-modal visual recognition method for drones based on cross-scenario transfer learning, characterized in that, include: The sensors carried by the drone collect environmental parameter information and multimodal image data in real time and synchronously, and use the scene prediction model to predict the scene category of the current flight area; By calculating the matching degree between the scene category and the source domain subset in the preset source domain data set, selecting the source domain subset with the highest matching degree; Dividing the target domain sample into multiple subdomains based on the environmental feature information of the target domain sample to obtain target subdomain data; Extracting the single-modal features of each modality image in the multimodal image data respectively, and dynamically switching the fusion strategy of the single-modal features based on the quality assessment results of each modality image and the scene category, fusing the single-modal features, and generating a first fusion feature; For the first fusion features of multiple frames continuously collected during the flight, a sliding window mechanism is used to accumulate time series features, and the features are corrected according to the historical feature state at the previous moment to obtain the second fusion features; Based on the source domain subset and the target subdomain data, staged transfer learning is performed to train a visual recognition model and the trained visual recognition model is used to perform visual recognition on the second fused features of the current target domain.

2. The method for multi-modal visual recognition of an unmanned aerial vehicle based on cross-scenario transfer learning according to claim 1, wherein The environmental parameter information includes: at least one of light intensity, visibility, temperature, humidity and wind speed; the multimodal image data includes: visible light image, infrared image and lidar point cloud; The scene prediction model includes: an environmental parameter processing unit, a preliminary image feature extraction unit, a multi-source information fusion unit and a scene classification prediction unit; The environmental parameter processing unit normalizes and / or standardizes the environmental parameter information, encodes the category parameters, and outputs an environmental feature vector; the preliminary image feature extraction unit extracts a visual feature vector based on the multimodal image data using a lightweight feature extraction network; the visual feature vector contains scene content information; the multi-source information fusion unit fuses the environmental feature vector and the visual feature vector to obtain a scene representation feature; the scene classification prediction unit calculates the probability that the scene representation feature corresponds to each predefined scene category through M fully connected layers and a Softmax activation function, and outputs the scene category with the highest probability as a prediction result.

3. A multi-modal visual recognition method for drones with cross-scenario transfer learning according to claim 1, characterized in that, The process of obtaining the source domain subset includes: Pre-label the scene category to which each source domain sample belongs in the source domain dataset; According to the current scene category predicted by the scene prediction model, the Mahalanobis distance is used to calculate the matching degree between the current scene category and each scene category in the source domain dataset; The matching degree is sorted from high to low, and a set of source domain samples whose matching degree exceeds a set threshold is selected as the source domain subset, while source domain samples whose matching degree is lower than the set threshold are eliminated.

4. A multi-modal visual recognition method for drones with cross-scenario transfer learning according to claim 1, characterized in that, The process of acquiring the target subdomain data includes: Collect unlabeled sample data of the current target domain; perform cluster analysis on the target domain samples using a clustering algorithm according to the environmental feature vectors obtained synchronously with the unlabeled sample data to obtain N clusters; the number of the clusters is adaptively adjusted according to the scene category; Divide the target domain samples belonging to the same clustering cluster into a sub-domain to obtain N pieces of the target sub-domain data.

5. A multi-modal visual recognition method for drones with cross-scenario transfer learning according to claim 1, characterized in that The generation process of the first fusion feature includes: Use a feature extraction network to extract the depth features of the visible light image, infrared image, and lidar point cloud in the multi-modal image data as single-modal features respectively; Calculate the sharpness score, signal-to-noise ratio, contrast, and point cloud density index of the visible light image, infrared image, and lidar point cloud in the multi-modal image data to obtain the quality evaluation results of each modal image; Based on the real-time environmental parameters, predicted scene categories, and modal image quality evaluation results, set a fusion strategy switching rule table, dynamically select and execute one and / or more feature fusion strategies to fuse the single-modal features and generate the first fusion feature.

6. The method for multi-modal visual recognition of drones with cross-scenario transfer learning according to claim 1, wherein, The acquisition process of the second fusion feature includes: Set a time sliding window to accumulate the features of K consecutive frames of the first fusion feature; Introduce the historical feature state to correct the first fusion feature within the current time window to obtain the second fusion feature; the correction is based on the error compensation mechanism between the first fusion feature at the previous moment and the first fusion feature at the current moment.

7. A multi-modal visual recognition method for drones with cross-scenario transfer learning according to claim 1, characterized in that, The visual recognition model is formed during the staged transfer learning training process and is finally used to perform specific recognition tasks on the second fusion feature, specifically including: a feature processing and local adaptation unit, a task recognition head unit, and a visual recognition output unit; The feature processing and local adaptation unit includes a feature transformation layer and a local adaptation layer; the feature transformation layer contains a series of neural network layers for non-linearly processing the input second fusion feature; the local adaptation layer aligns the source domain subset with each target sub-domain based on the least mean square error loss; the task recognition head unit is connected after the feature processing and local adaptation unit, and its specific structure is determined according to the final visual recognition task; the visual recognition task includes an image classification task, an object detection task, and other recognition tasks; the visual recognition output unit is connected after the task recognition head unit for outputting the visual recognition result.

Citation Information

Patent Citations

  • Anti-interference target detection method based on automatic driving scene multi-mode fusion

    CN115393684A

  • Complex scene target detection method based on multi-modal fusion and storage medium

    CN115713481A

  • Three-dimensional scene segmentation domain migration method and device based on multi-source heterogeneous data fusion

    CN116246070A

  • Multi-scale feature fusion unmanned aerial vehicle target detection method for bad weather

    CN119131622A

  • Multi-target detection method for complex traffic scene

    CN119442145A

Cited By

  • Automatic scene calibration method based on machine vision

    CN120635605A

  • Target identification method and system based on deep learning

    CN120876834A

  • Efficient intelligent delivery parcel detection method based on transfer learning

    CN121121316A

  • Efficient intelligent detection method for delivery parcels based on transfer learning

    CN121121316B

  • Unmanned aerial vehicle load control method, system and equipment integrating visible light and laser radar

    CN121541662A