Unmanned aerial vehicle inspection intrusion detection method and system based on deep learning

By fusing real-time UAV status data into a multimodal intrusion detection model, a multi-label risk probability vector is generated, which solves the problems of single detection dimension and coarse threat identification granularity in UAV inspection. It enables accurate identification and immediate response to complex attacks, and improves the security and traceability of UAV data links.

CN121842686APending Publication Date: 2026-04-10CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
Filing Date
2025-11-20
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing drone-based intrusion detection technologies have limited detection dimensions and fail to deeply integrate network and physical conditions. This results in weak ability to identify complex attacks, coarse threat identification granularity, inability to effectively characterize concurrent attacks, and a lack of intelligent reporting mechanisms, which can easily lead to channel congestion or delays in critical alarms.

Method used

By acquiring real-time status data of drones, a multimodal intrusion detection model is used for fusion analysis to generate multi-label risk probability vectors. Adaptive risk assessment and hierarchical uplink strategies are then implemented to identify and classify attack types.

Benefits of technology

It enables accurate identification of single and compound attacks, reduces false alarms and missed alarms, provides rich threat situation information, ensures the timeliness of critical alarms, reduces the load on communication links, and improves the anti-intrusion capability and event traceability of UAV data links.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842686A_ABST
    Figure CN121842686A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle inspection intrusion detection method and system based on deep learning, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining the real-time state data of an unmanned aerial vehicle; performing fusion analysis on the real-time state data by using a pre-constructed multi-mode intrusion detection model, and generating a real-time multi-label risk probability vector representing an attack type; and performing adaptive risk assessment based on the real-time multi-label risk probability vector, and executing a hierarchical uplink strategy to obtain an unmanned aerial vehicle inspection intrusion detection result. According to the invention, through multi-modal information, complex cross-domain composite attacks are accurately identified, and false alarms and missing alarms are reduced; through an output form of a multi-label risk index, the system can identify and report various attacks at the same time, and richer threat situation information is provided for a ground station; and the anti-intrusion capability of the unmanned aerial vehicle data link and the traceability of the event are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a UAV inspection intrusion detection method and system based on deep learning. BACKGROUND

[0002] Unmanned aerial vehicles (UAVs) have been widely used in power inspection due to their high mobility, low cost and easy deployment. When performing tasks, UAVs are not only simple flight platforms, but also aerial intelligent nodes equipped with high-definition cameras, sensors and communication modules. Therefore, the flight safety and data security of UAVs are of great importance. However, the openness of UAV systems and their dependence on wireless communication expose them to an increasing number of complex security threats. Intrusion detection technology has become an essential barrier to ensuring the safe inspection of UAVs.

[0003] However, the existing technology has single detection dimension when detecting UAV inspection intrusion, and focuses on network side traffic analysis. It fails to deeply integrate with the physical flight state of the UAV, resulting in weak recognition ability for complex attacks that require cross-domain collaboration. The threat recognition granularity is coarse, and the detection output is mostly mutually exclusive single labels or single alarms, which cannot effectively represent and distinguish multiple concurrent attacks occurring simultaneously, making it difficult to support fine-grained threat situation assessment and response. There is a lack of intelligent reporting mechanism matching the limited communication resources of UAVs, which cannot adaptively adjust the reporting content and priority according to the risk level, easily causing channel congestion or delaying or losing critical alarms. SUMMARY

[0004] To overcome the shortcomings of the prior art, the present application provides a UAV inspection intrusion detection method based on deep learning, comprising: obtaining real-time state data of the UAV, including real-time network traffic data, real-time dynamic network topology data, real-time geographic positioning data and real-time inertial measurement unit data of the UAV; using a pre-constructed multi-modal intrusion detection model to analyze the real-time state data, generating a real-time multi-label risk probability vector representing the type of attack suffered; based on the real-time multi-label risk probability vector, performing adaptive risk assessment and executing a hierarchical uplink strategy to obtain the UAV inspection intrusion detection result; the hierarchical uplink strategy includes dividing the evaluated current risk into different risk levels and determining the data content and priority reported to the ground end according to the risk level.

[0005] Preferably, the construction of the multi-modal intrusion detection model comprises: The unmanned aerial vehicle collects original data under normal inspection and preset attack scenarios, and the original data is under four modalities of network traffic, dynamic network topology, geographical spatial positioning and inertial measurement unit, and time synchronization preprocessing is performed to form a unified sample slice with time alignment; Based on the credibility weight, the quality of the unified sample slice is detected and the credibility is evaluated and optimized to obtain an optimized unified sample slice; The coupling gate perception mechanism is used to dynamically weight and fuse each modal feature in the optimized unified sample slice to generate a unified high-dimensional fusion feature; The heterogeneous special collaborative array is used to process the high-dimensional fusion feature to output a multi-label risk probability vector representing the attack type; the heterogeneous special collaborative array includes multiple parallel special sub-networks, and each special sub-network is used to identify a specific attack type; Through multi-label oriented end-to-end joint training, the coupling gate perception mechanism and the heterogeneous special collaborative array are jointly optimized to obtain a mapping relationship from the original data to the multi-label risk probability vector, and a multi-modal intrusion detection model is constructed based on the mapping relationship.

[0006] Preferably, the unmanned aerial vehicle collects original data under normal inspection and preset attack scenarios, and the original data is under four modalities of network traffic, dynamic network topology, geographical spatial positioning and inertial measurement unit, and time synchronization preprocessing is performed to form a unified sample slice with time alignment, including: The unmanned aerial vehicle collects original data under normal inspection and preset attack scenarios, and the original data is under four modalities of network traffic, dynamic network topology, geographical spatial positioning and inertial measurement unit; According to the original time stamps corresponding to each modal data in the original data, the neighboring observations are aligned to the same time window to form a preliminary correspondence relationship, and the preliminary corresponding time stamps of each modality are obtained; Taking the time data of geographical spatial positioning as a reference time base, the preliminary corresponding time stamps of each modality are corrected by an estimation linear clock mapping algorithm to obtain corrected data of each modality; The corrected data of each modality is resampled and aligned on an equal-interval time grid point, and the corrected data of each modality is spliced in an adaptive sliding time window centered on the current time to form a unified sample slice with time alignment.

[0007] Preferably, taking the time data of geographical spatial positioning as a reference time base, the preliminary corresponding time stamps of each modality are corrected by an estimation linear clock mapping algorithm to obtain corrected data of each modality, including: Taking the time data of geographical spatial positioning as a reference time base, the clock stretching coefficients and offsets of each modality are estimated by least squares method at selected anchor points; The clock stretching coefficient and the offset of the preliminary corresponding time stamp of each modality are corrected by estimating a linear clock mapping algorithm, and the corrected data of each modality is obtained according to the corrected time stamp; The original time stamp of each modality is , The original time stamp of each modality is corrected by estimating a linear clock mapping algorithm, and the corrected data of each modality is obtained according to the corrected time stamp; The linear clock mapping algorithm is as follows: , wherein, The preliminary corresponding time stamp of modality m is The corrected time stamp after mapping to the GPS reference time base is ; The formula of the least square method is as follows:

[0008] is fitted by the least square method on the anchor point pair , is the time index of the anchor point in modality m, is the time index corresponding to the anchor point of the GPS sequence of modality m, is the clock stretching coefficient of modality m, is the clock offset of modality m, is the index of the anchor point in modality m, is the index of the corresponding anchor point in the GPS sequence, is the anchor point pair number, is the clock stretching coefficient, is the clock offset coefficient.

[0009] Preferably, based on the credibility weight, the quality detection and credibility evaluation optimization of the unified sample slice are performed to obtain the optimized unified sample slice, including: Based on the statistical sufficiency and the time sequence consistency, the independent quality indicators of each modality are calculated; The modality credibility of each modality in the current sample slice and the comprehensive credibility of the current sample slice are calculated by fusing the independent quality indicators of each modality and the cross-modality alignment residual; Based on the comprehensive credibility and the preset quality threshold, the quality gating is performed to obtain the optimized unified sample slice; The calculation formula of the modality credibility is as follows: ; Let m be the confidence level of the m-th mode at time n. Let m be the coverage of the m-th mode. Let m be the stability of the m-th mode. To align residuals for cross-modal temporal consistency, To adjust the power exponent of quality discrimination, The preset lower limit of credibility weight, To truncate the result to an interval ; The formula for calculating overall credibility is as follows:

[0010] The overall confidence level of all modes at time n is given by: Let m be the confidence level of the m-th mode at time n. Let be the confidence weight of the m-th mode.

[0011] Preferably, a coupled gated sensing mechanism is used to dynamically weight and fuse the modal features of each segment in the optimized unified sample slice to generate unified high-dimensional fusion features, including: Encode each modal data in the optimized unified sample slice to obtain the encoding features of each modality; The baseline importance score of each modality is calculated based on the modality coding features, and the confidence of each modality is injected into the baseline importance score to obtain the corrected score result. Calculate the similarity of each modality's encoded features in the common projection space to form a cross-modal support matrix; The corrected scoring results are coupled and updated based on the cross-modal support matrix to obtain the final gating vector; The modality-coded features are weighted and summed based on the final gating vector to generate a unified high-dimensional fusion feature. The formulas for each modality coding feature are as follows:

[0012] For the encoding features of the m-th modality, For the preprocessed features of the m-th mode, For the encoder of the m-th mode; The formula for baseline importance scoring is as follows:

[0013] Assign a baseline importance score to the m-th mode on sample n. Encode features for each modality Projected onto common dimensions linear mapping of the m-th modality, is the encoding feature of the m-th modality, is the context vector, is the linear readout parameter of the m-th modality, is the transpose of is the bias of the m-th modality; The modified scoring result is calculated as follows:

[0014] is the modified scoring result, is the baseline importance score of the m-th modality on sample n, is the injection intensity coefficient, is the credibility of the m-th modality at time n; The calculation formula of the cross-modal support matrix is as follows:

[0015]

[0016] is the time window index, is the support of the m-th modality in the n-th time window, is the support of the m-th modality in the n-th time window, is the state encoding of the m-th modality in the n-th time window, is the state encoding of the m-th modality in the n-th time window, , is the modality index, is the self-loop of the m-th modality in the n-th time window, is the self-loop of the m-th modality in the n-th time window, The calculation process of the final gating vector includes the following: The initial gating vector is as follows:

[0017]

[0018]

[0019] The cross-modal support amount is as follows:

[0020]

[0021] The final gating vector is as follows: ​​​​​

[0022]

[0023]

[0024] wherein, is a time window index; , and m, l are modal indices; is a total number of modalities; is a category index; is a total number of categories; is an iteration step index; is an initial gating vector, obtained by , is an initial gating vector of modality m, is a scoring result after modal injection of credibility; is a cross-modal support quantity vector, is a cross-modal support quantity vector of modality m, m, l are modal indices, is a final gating vector of modality l, wherein is a scoring vector after joining cross-modal support, is a weight update coefficient; is a final gating / fusion weight, satisfying ; is a scoring vector after joining cross-modal support; is a vector composed of each , indicates a scoring result after modal injection of credibility, the first component is , is a final gating vector of the (t+1)th iteration, is a final gating vector of the tth iteration.

[0025] Preferably, a heterogeneous special direction collaborative array is adopted to process the high-dimensional fusion feature, and a multi-label risk probability vector representing the type of attack suffered is output, including: The high-dimensional fusion feature is obtained through a shared transformation function to obtain a unified public feature; The unified public feature is input in parallel to k structurally different special direction sub-networks to obtain an independent intermediate representation and a logarithmic risk score of each special direction sub-network; the k structurally different special direction sub-networks are constructed using at least two different neural network operator families of different computing paradigms, for capturing different time scales and related patterns; The logarithmic risk score output by each special direction sub-network is collaboratively corrected through a non-diagonal label correlation matrix to obtain a logarithmic risk vector, and the co-occurrence or mutual exclusion relationship between different attack labels is obtained. Convert the logarithmic risk vector into a multi-label risk probability vector; The shared transformation function is as follows:

[0026] To share the transformation function, As a high-dimensional fusion feature, To unify public characteristics; The intermediate formula is as follows:

[0027] This represents the intermediate feature representation output by the k-th specialized sub-network. To unify common characteristics, It is a linear discriminant layer; The formula for collaborative correction is as follows:

[0028]

[0029] To obtain the basic logarithmic risk in the dedicated subnetwork, To correct for logarithmic risk, To learn in conjunction with sub-network parameters, This indicates the co-occurrence tendency of i and j. This indicates the mutual exclusion tendency of i and j. As the first fundamental logarithmic risk component, This is the fourth fundamental logarithmic risk component; The formula for calculating the multi-label risk probability vector is as follows:

[0030] As a high-dimensional fusion feature, The parameters of the first classifier for the k-th class are... The parameters of the second classifier for the k-th class are... The Sigmoid function outputs multi-label probabilities for each category. Let be the probability of the k-th type of risk occurring at time n. This represents the total number of tags.

[0031] Preferably, before performing adaptive risk assessment based on real-time multi-label risk probability vectors and executing a tiered uplink strategy to obtain the UAV inspection intrusion detection results, the method further includes: The real-time multi-label risk probability vector is weighted according to preset risk weights to obtain the comprehensive risk index under the current situation; Based on the comprehensive risk index and adaptively set high and low thresholds, the risk level is divided into three levels: high, medium and low, and an adaptive risk assessment algorithm for adaptive risk assessment is constructed. The formula for calculating the comprehensive risk index is as follows:

[0032] Let n be the comprehensive risk index at time n. Let n be the weight coefficient of the k-th type of risk at time n. Let be the probability of the k-th type of risk occurring at time n.

[0033] Preferably, before performing adaptive risk assessment based on real-time multi-label risk probability vectors and executing the tiered uplink strategy, the method further includes: When the current risk level is high, a high-priority uplink strategy is generated. The high-priority uplink strategy includes: a high-priority alarm data packet, and uplinking to the ground terminal by preempting resources through the communication link. The high-priority alarm data packet encapsulates at least the following data: risk token, multi-label attack vector, comprehensive risk index, and key forensic summary data for post-event analysis. When the current risk level is medium, a medium-priority uplink strategy is generated; the medium-priority uplink strategy includes: a lightweight alarm data packet and sending it to the ground terminal; the lightweight alarm data packet encapsulates a risk token and a multi-tag attack vector; When the current risk level is low, a low-priority uplink strategy is generated. The low-priority uplink strategy includes: not triggering real-time uplink alarms, recording the current risk index, status slice and timestamp in the drone's local memory, and transmitting them back to the ground in batch compression when the communication link is idle or bandwidth is sufficient.

[0034] Preferably, after performing adaptive risk assessment based on real-time multi-label risk probability vectors and implementing a tiered uplink strategy, the method further includes: A buffer zone is set up on the drone. The buffer zone adopts a circular queue and is equipped with metadata index to store alarm and evidence data to be reported when the communication link is interrupted. When the communication link is restored, the priority of the backhaul is calculated based on the freshness and risk level of the alarm and evidence data to be reported, and the compensation backhaul is performed in order of priority.

[0035] Based on the same inventive concept, this invention also provides a deep learning-based drone inspection intrusion detection system, which includes: The status data acquisition module is used to acquire real-time status data of the UAV, including real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data. The risk probability vector generation module is used to perform fusion analysis on real-time status data using a pre-built multimodal intrusion detection model to generate a real-time multi-label risk probability vector representing the type of attack suffered. The intrusion detection result acquisition module is used to perform adaptive risk assessment based on real-time multi-label risk probability vectors and execute a hierarchical uplink strategy to obtain intrusion detection results of UAV inspections. The hierarchical uplink strategy includes: classifying the assessed current risk into different risk levels, and determining the data content and priority to be reported to the ground end based on the risk level.

[0036] Based on the same inventive concept, the present invention also provides an electronic device, comprising: at least one processor and a memory; wherein the memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a deep learning-based drone inspection intrusion detection method as described above is implemented.

[0037] Based on the same inventive concept, the present invention also provides a readable storage medium having an executable program stored thereon, wherein when the executable program is executed, it implements the deep learning-based UAV inspection intrusion detection method described above.

[0038] Compared with the closest existing technology, the present invention has the following beneficial effects: This invention provides a deep learning-based method for UAV inspection intrusion detection, comprising: acquiring real-time status data of the UAV, including real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data; fusing and analyzing the real-time status data using a pre-built multimodal intrusion detection model to generate a real-time multi-label risk probability vector representing the type of attack suffered; performing adaptive risk assessment based on the real-time multi-label risk probability vector and executing a hierarchical uplink strategy to obtain UAV inspection intrusion detection results; the hierarchical uplink strategy includes: classifying the assessed current risk into different risk levels, and determining the content and priority of data reported to the ground terminal according to the risk level. This invention, by integrating multimodal information from the network and physical space, can more accurately identify single attacks and complex cross-domain composite attacks, reducing false positives and false negatives. Through the output format of a multi-label risk index, the system can simultaneously identify and report multiple attacks, providing ground stations with richer threat situation information. A tiered uplink strategy transforms abstract risk indexes into specific risk levels, intelligently allocating communication resources based on these levels to ensure the timeliness of critical alarms and reduce the average load on the link. Overall, it improves the anti-intrusion capability and event traceability of the UAV data link. Attached Figure Description

[0039] Figure 1 A flowchart illustrating a deep learning-based drone inspection intrusion detection method provided by this invention; Figure 2 This is a diagram of the multimodal intrusion detection model architecture provided by the present invention; Figure 3 The overall flowchart of the deep learning-based drone intrusion detection method provided by this invention; Figure 4 The present invention provides a structural diagram of a deep learning-based unmanned aerial vehicle (UAV) inspection and intrusion detection system. Figure 5 A schematic diagram of the electronic device provided by the present invention. Detailed Implementation

[0040] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0041] Example 1: This invention provides a deep learning-based method for drone inspection and intrusion detection. Specifically, Figure 1 The flowchart of the deep learning-based UAV inspection intrusion detection method provided in this embodiment of the invention is shown in the figure, and includes the following steps: S101: Acquire real-time status data of the UAV, including real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data. S102: Utilize a pre-built multimodal intrusion detection model to fuse and analyze real-time state data, generating a real-time multi-label risk probability vector representing the type of attack suffered; S103: Adaptive risk assessment is performed based on real-time multi-label risk probability vectors, and a hierarchical uplink strategy is executed to obtain drone inspection intrusion detection results; the hierarchical uplink strategy includes: classifying the assessed current risk into different risk levels, and determining the data content and priority to be reported to the ground end according to the risk level.

[0042] This invention, by integrating multimodal information from the network and physical space, can more accurately identify single attacks and complex cross-domain composite attacks, reducing false positives and false negatives. Through the output format of a multi-label risk index, the system can simultaneously identify and report multiple attacks, providing ground stations with richer threat situation information. A tiered uplink strategy transforms abstract risk indexes into specific risk levels, intelligently allocating communication resources based on these levels to ensure the timeliness of critical alarms and reduce the average load on the link. Overall, it improves the anti-intrusion capability and event traceability of the UAV data link.

[0043] In this invention, the real-time status data of the UAV is first acquired. The real-time status data includes the UAV's real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data.

[0044] Specifically, it collects real-time network traffic data (NetFlow), real-time dynamic network topology data (Topo), real-time geospatial positioning data (GPS, Global Positioning System), and real-time inertial measurement unit (IMU) data from the UAV.

[0045] Furthermore, a pre-built multimodal intrusion detection model is used to fuse and analyze real-time state data, generating a real-time multi-label risk probability vector representing the type of attack suffered.

[0046] Before performing deep learning-based drone inspection intrusion detection, this invention requires the pre-construction of a multimodal intrusion detection model. This multimodal intrusion detection model is based on a deep learning model and is ultimately obtained through offline training. The core task of the multimodal intrusion detection model is to use relevant datasets to train a deep learning model capable of understanding and recognizing complex attack patterns.

[0047] The process of building a multimodal intrusion detection model includes: S201: Collect raw data from UAVs under normal inspection and preset attack scenarios, targeting four modes: network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit, and perform time-synchronized preprocessing to form time-aligned unified sample slices. S202: Based on the credibility weight, perform quality detection and credibility assessment on the unified sample slice to obtain the optimized unified sample slice; S203: A coupled gated sensing mechanism is used to dynamically weight and fuse the modal features in the optimized unified sample slice to generate unified high-dimensional fusion features; S204: A heterogeneous specialized cooperative array is used to process high-dimensional fusion features and output a multi-label risk probability vector representing the type of attack suffered. The heterogeneous specialized cooperative array contains multiple parallel specialized sub-networks, each of which is used to identify a specific type of attack. S205: Through multi-label-guided end-to-end joint training, the coupled gating sensing mechanism and the heterogeneous specialized cooperative array are jointly optimized to obtain the mapping relationship from the original data to the multi-label risk probability vector, and a multimodal intrusion detection model is constructed based on the mapping relationship.

[0048] Specifically, raw data from UAVs under normal inspection and pre-set attack scenarios are collected for four modes: network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit (IMU). Time-synchronized preprocessing is then performed to form a time-aligned unified sample slice. This process includes: collecting raw data from UAVs under normal inspection and pre-set attack scenarios for these four modes; aligning neighboring observations to the same time window based on the raw timestamps corresponding to each mode in the raw data to establish a preliminary correspondence and obtain preliminary corresponding timestamps for each mode; using geospatial positioning time data as a reference time base, correcting the preliminary corresponding timestamps for each mode using an estimation linear clock mapping algorithm to obtain corrected data for each mode; resampling and aligning the corrected data for each mode at equally spaced time grid points; and stitching the corrected data for each mode together within an adaptive sliding time window centered on the current time to ensure that at any given moment, the data from the four modes can form a time-aligned unified sample slice describing the complete state of the UAV.

[0049] Specifically, using geospatial positioning time data as a reference time base, the preliminary corresponding timestamps of each mode are corrected using an estimation linear clock mapping algorithm to obtain corrected modal data. This includes: using geospatial positioning time data as a reference time base, estimating the clock scaling factor and offset of each mode at selected anchor points using the least squares method; correcting the clock scaling factor and offset of the preliminary corresponding timestamps of each mode using an estimation linear clock mapping algorithm; and obtaining corrected modal data based on the corrected timestamps. The original timestamps for each modality are as follows: , Various modes NetFlow is network traffic data, Topo is dynamic network topology data, GPS is geospatial positioning data, and IMU is inertial measurement unit data. The algorithm for estimating linear clock mapping is as follows: , in, This represents the initial corresponding timestamp for mode m. Indicates to Corrected timestamp mapped to GPS reference time base; The formula for the least squares method is as follows:

[0050] For least squares at anchor point pairs Upper fitting, This is the time index of the anchor point in mode m. This is the time index corresponding to the GPS sequence anchor point of mode m. Let m be the clock scaling factor for mode m. Let m be the clock offset of mode m. This is the index of the anchor point in mode m. This is the index of the corresponding anchor point in the GPS sequence. Number the anchor points. This is the clock scaling factor. This is the clock offset coefficient.

[0051] Multimodal data is collected and preprocessed synchronously to prepare a high-quality, time-aligned dataset for model training, targeting both normal UAV inspections and scenarios subjected to pre-defined attacks. Network traffic (NetFlow), dynamic network topology (Topo), geospatial positioning (GPS), and inertial measurement unit (IMU) data are collected. The raw data is preprocessed to ensure that at any given time, the data from the four modalities can form a unified sample slice describing the complete state of the UAV.

[0052] In one specific implementation, multimodal data synchronization preprocessing includes: collecting network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit data from UAVs under normal inspection and pre-defined attack scenarios; performing time base unification and drift correction on each modality's data; and using a "three-step alignment" process for coarse alignment in the data synchronization preprocessing stage. First, based on the original timestamps / frame sequences of each modality, aligning neighboring observations to the same time window to form a preliminary correspondence; second, performing linear time synchronization, using GPS time as a reference, and estimating the clock scaling factor and offset of each modality using the least squares method on selected anchor point pairs. And based on this, map the modal time axis as This simultaneously corrects clock scaling and overall offset; finally, fine alignment is performed, aligning the calibrated multimodal data according to a unified grid. Perform interpolation / filtering resampling, in length of Slices of a uniform state are formed by splicing within a sliding time window. Among them, grid points The timeline, unified to the GPS reference time axis, is discretized to obtain a series of equally spaced sampling times, which serve as the anchor points for each sample slice. A starting point is then selected. and step length Grid points are The grid points are anchor points that appear on a uniform, equally spaced time grid at each moment, rather than the original, unequal-interval timestamps of each sensor.

[0053] At each grid point, the four modalities are spliced ​​to obtain a uniform sample slice. ,for: ,in This is the encoding of the network topology mode at time n. Let n be the feature vector in the NetFlow mode at time n. Let be the feature vector of the GPS mode at time n. Let be the eigenvector of the IMU mode at time n.

[0054] After forming a time-aligned unified sample slice, the unified sample slice is further optimized by quality detection and credibility assessment based on credibility weights to obtain an optimized unified sample slice.

[0055] The specific optimization process includes: calculating the independent quality index of each modality based on statistical sufficiency and temporal consistency; fusing the independent quality index of each modality with the cross-modal alignment residual to calculate the modal credibility of each modality in the current sample slice, and calculating the comprehensive credibility of the current sample slice; performing quality gating based on the comprehensive credibility and a preset quality threshold to obtain the optimized unified sample slice.

[0056] The formula for calculating modal confidence is as follows: ; Let m be the confidence level of the m-th mode at time n. Let m be the coverage of the m-th mode. Let m be the stability of the m-th mode. To align residuals based on cross-modal temporal consistency, To adjust the power exponent of quality discrimination, The preset lower limit of credibility weight, To truncate the result to an interval ; The formula for calculating overall credibility is as follows:

[0057] The overall confidence level of all modes at time n is given by: Let m be the confidence level of the m-th mode at time n. Let be the confidence weight of the m-th mode.

[0058] In one specific implementation, an upper limit is set for the cross-modal alignment residuals of a uniform sample slice, quality checks are performed to obtain modal confidence weights, and a comprehensive confidence score is calculated. When the comprehensive confidence score falls below a threshold, gating or rejection is implemented. The central window Within the sample slice, a quality control process of "statistical sufficiency + temporal consistency" is first performed, and the results are mapped to modal confidence.

[0059] Taking network traffic data as an example, packet coverage is used. and arrival rate stability Calculate the credibility weight of network traffic. For packet coverage, This represents the number of network packets actually observed within the current sliding window. To determine the desired number of packets to observe; For network-side arrival rate stability, For network-side packet arrival rate, The minimum threshold to be reached to ensure stable estimation.

[0060] Taking inertial measurement unit data as an example, frame coverage is used. And combined with the mean angular velocity Saturation penalty Calculate the confidence weight of the inertial measurement unit to avoid noise amplification caused by violent maneuvers. For frame coverage, This represents the number of valid frames actually observed within the current sliding window. The desired number of frames to be captured; This is the saturation penalty value. This is the average of the L2 norms of the angular velocity vectors of all frames within the current sliding window. This is the angular velocity saturation threshold.

[0061] All modalities share a time-series consistency metric: cross-modal alignment residuals. and mapped to ; For the mapped cross-modal alignment residual, The scaling factor for aligning the residuals; the residuals are less than a preset upper limit to satisfy resampling and synchronous alignment.

[0062] Modal reliability is: The overall credibility is .when During the training phase, the slice is directly removed to prevent noise from being included in the database; among which... This is the overall credibility threshold.

[0063] The optimized unified sample slices are labeled to form a high-quality, time-aligned dataset suitable for training multimodal intrusion detection models.

[0064] After obtaining the optimized unified sample slice, an end-to-end neural network architecture consisting of a cascaded fusion module and a detection module is defined. A gated perception mechanism is used as the fusion module. It receives four modal feature vectors through a gating unit, dynamically evaluates the information contribution of each modality in the current context through learnable parameters, and generates a unified high-dimensional feature representation that has undergone intelligent weighting and deep fusion. A heterogeneous specialized cooperative array (Heco-Array) is used for backend detection. This module consists of four parallel, structurally simplified specialized sub-networks. Each sub-network is designed as a "specialized identifier" that specifically identifies a certain type of attack. They collaboratively receive the unified feature representation from the coupled gated perception mechanism fusion module and perform inference independently.

[0065] Specifically, a coupled gated sensing mechanism is used to dynamically weight and fuse the modal features in the optimized unified sample slice to generate a unified high-dimensional fusion feature. This includes: encoding the modal data in the optimized unified sample slice to obtain modal encoded features; calculating the baseline importance score for each modality based on the modal encoded features, injecting the confidence level of each modality into the baseline importance score to obtain a corrected score; calculating the similarity of the modal encoded features in the common projection space to form a cross-modal support matrix; coupling and updating the corrected score based on the cross-modal support matrix to obtain the final gating vector; and weighting and summing the modal encoded features according to the final gating vector to generate a unified high-dimensional fusion feature.

[0066] The formulas for each modality coding feature are as follows:

[0067] For the encoding features of the m-th modality, For the preprocessed features of the m-th mode, For the encoder of the m-th mode; The formula for baseline importance scoring is as follows:

[0068] Assign a baseline importance score to the m-th mode on sample n. Encode features for each modality Projected onto common dimensions linear mapping, For the encoding features of the m-th modality, For context vectors, Let be the linear readout parameters for the m-th mode. for transpose, This is the bias for the m-th mode; The revised scoring formula is as follows:

[0069] The revised scoring results Assign a baseline importance score to the m-th mode on sample n. The injection strength coefficient, Let m be the confidence level of the m-th mode at time n; The formula for calculating the cross-modal support matrix is ​​as follows:

[0070]

[0071] For the mode on sample n For modes Support For the mode on sample n Status coding, For the mode on sample n Status coding, , For modal indexing, For the mode on sample n With mode Self-loop; The calculation process for the final gating vector includes the following: The initial gating vector is as follows:

[0072]

[0073]

[0074] The cross-modal support is as follows:

[0075]

[0076] The final gating vector is as follows:

[0077]

[0078]

[0079] in, Index for time windows; And m and l are modal indices; The total number of modes; Indexed by category; Total number of categories; For iteration step index; For the final gating / fusion weights, satisfy ; Let be the initial gating vector, given by get, Let m be the initial gating vector for mode m. Scoring results after injecting credibility into the modality; For cross-modal support vectors, Let l be the cross-modal support vector of mode m, where m and l are mode indices. Let be the final gating vector of mode l; This is the scoring vector after adding cross-modal support; For each The vector formed This represents the scoring result after the confidence level of each modality injection, the th Each component is , Let be the final gating vector of the (t+1)th iteration. Let be the final gating vector for the t-th iteration.

[0080] In one specific implementation, at time Establish a length of centered Adaptive sliding time window ,in The selection is adaptive based on the link packet arrival rate, and the specific formula is as follows:

[0081]

[0082]

[0083] in, For adaptive sliding time, The recommended window length to ensure network-side statistical stability. The minimum average number of network packets observed within a time window. The package arrival rate is estimated using EWMA (Exponentially Weighted Moving Average). Recommended window duration to ensure inertial navigation information coverage. For IMU frame count target, This refers to the IMU sampling frequency.

[0084] Within the adaptive sliding time window, the modal features of the optimized unified sample slice are adaptively weighted and fused. The coupled gated perception (C-GaPe) mechanism includes: First, modal encoding is performed using their respective encoders: , For the encoding features of the m-th modality, For the preprocessed features of the m-th mode, For the encoder of the m-th mode; Next, basic gate calculations are performed, based on the content and context vectors. Find:

[0085] Assign a baseline importance score to the m-th mode on sample n. Encode features for each modality Projected onto common dimensions linear mapping, For the encoding features of the m-th modality, For context vectors, Let be the linear readout parameters for the m-th mode. for transpose, This is the bias for the m-th mode; Then, credibility is injected, combined with modal credibility obtained from quality control. Revised to: , Let m be the confidence level of the m-th mode at time n. The injection strength coefficient, This is the revised scoring result.

[0086] Next, cross-modal coupling is performed to form the support matrix using projection similarity:

[0087]

[0088] in, For time indexing, as well as For modal indexing, NetFlow is network traffic data, Topo is dynamic network topology data, GPS is geospatial positioning data, and IMU is inertial measurement unit data. This is the state encoding for the m-th mode within the n-th time window. Encoding the state of the l-th mode within the n-th time window. For modality For modes The support at time n. Let... This will remove the self-loop.

[0089] Then, the threshold value is calculated. Let the number of modes be M, and define a vector.

[0090]

[0091] First, obtain the initial uncoupled gate. ,

[0092] Next, the cross-modal support quantity is calculated:

[0093]

[0094] Then update the scoring and calculate the final gate vector:

[0095]

[0096]

[0097] The above formulas simultaneously satisfy: .

[0098] in, Index for time windows; And m and l are modal indices; The total number of modes; Indexed by category; Total number of categories; For iteration step index; Represents the scoring vector after injecting confidence for each modality, where the i-th Each component is Let be the initial gating vector, given by get, Let m be the initial gating vector for mode m. Scoring results after injecting credibility into the modality; For cross-modal support vectors, Let l be the cross-modal support vector of mode m, where m and l are mode indices. Let be the final gating vector of mode l, where This is the scoring vector after adding cross-modal support. Update the coefficients for the weights; For the final gating / fusion weights; ,satisfy For the first Cross-modal support matrix for each window.

[0099] After generating unified high-dimensional fusion features, a heterogeneous specialized cooperative array is used to process these features, outputting a multi-label risk probability vector representing the type of attack suffered. The heterogeneous specialized cooperative array contains multiple parallel specialized sub-networks, each dedicated to identifying a specific attack type. In other words, a multi-label-oriented end-to-end joint training is defined, aiming to teach the model an accurate mapping from input data to attack labels. The goal is to train the model to accurately identify single attacks as well as the complex ability to simultaneously identify multiple concurrent attacks.

[0100] Specifically, this includes: obtaining a unified common feature by sharing a transformation function to integrate high-dimensional fusion features; inputting the unified common feature in parallel into k structurally dissimilar targeted subnetworks to obtain an independent intermediate representation and log risk score for each targeted subnetwork; constructing the k structurally dissimilar targeted subnetworks using at least two families of neural network operators with different computational paradigms to capture different time scales and related patterns; co-correcting the log risk score output by each targeted subnetwork using an off-diagonal label correlation matrix to obtain a log risk vector, thereby obtaining the co-occurrence or mutual exclusion relationship between different attack labels; and converting the log risk vector into a multi-label risk probability vector. The shared transformation function is as follows:

[0101] To share the transformation function, As a high-dimensional fusion feature, To unify public characteristics; The intermediate formula is as follows:

[0102] This represents the intermediate feature representation output by the k-th specialized sub-network. To unify common characteristics, It is a linear discriminant layer; The formula for collaborative correction is as follows:

[0103]

[0104] To obtain the basic logarithmic risk in the dedicated subnetwork, To correct for logarithmic risk, To learn in conjunction with sub-network parameters, This indicates the co-occurrence tendency of i and j. This indicates the mutual exclusion tendency of i and j. As the first fundamental logarithmic risk component, This is the fourth fundamental logarithmic risk component; The formula for calculating the multi-label risk probability vector is as follows:

[0105] As a high-dimensional fusion feature, The parameters of the first classifier for the k-th class are... The parameters of the second classifier for the k-th class are... The Sigmoid function outputs multi-label probabilities for each category. Let be the probability of the k-th type of risk occurring at time n. This represents the total number of tags.

[0106] Furthermore, the multi-label risk probability vector can be calibrated. Specifically, airborne confidence calibration employs temperature and bias calibration, outputting... ,in, For the k-th class logarithmic risk, the linear discriminant output comes from the directional subnet. For bias correction, For temperature parameters, Let be the probability of the k-th logarithmic risk label occurring at time n after calibration.

[0107] and according to the threshold Obtain a four-bit binary multi-label attack vector ; During the training phase, weighted multi-label binary cross-entropy and array diversity regularization are used:

[0108] in, This is a diversity regularization term used to encourage each specialized subnetwork to learn complementary representations; For the k-th targeted subnetwork, the intermediate feature representation output for the n-th input sample is... This is the intermediate feature representation output by the first specialized subnetwork for the nth input sample; in subsequent outputs... It serves as a multi-label risk index, and is used for evidence collection with evidence scores attached to each specialized sub-network.

[0109] Then, through multi-label-guided end-to-end joint training, the coupled gating sensing mechanism and the heterogeneous specialized collaborative array are jointly optimized to obtain the mapping relationship from the original data to the multi-label risk probability vector, and a multimodal intrusion detection model is constructed based on the mapping relationship.

[0110] like Figure 2 The diagram shown illustrates the architecture of the multimodal intrusion detection model provided by this invention. In one specific implementation, it consists of four parallel dedicated sub-networks. The system simultaneously acquires raw network packet / session-level data flow graphs, GPS and sensor data, performs lightweight preprocessing such as token transformation and segmentation, graph tiling, and I / Q in-phase orthogonal differential analysis, and organizes the data into sample slices on a unified time grid. The outputs are then fed into a four-branch encoder for parallel modeling. The four branches together constitute an engineering example of a heterogeneous specialized cooperative array (Heco-Array): one branch employs a lightweight Transformer with local attention; another branch uses Edge-GraphSAGE, an edge-graph sampling and aggregation algorithm, superimposed with dilated temporal convolutions to capture topology and intermediate dependencies; and the remaining two branches first extract short-window dynamics using 1-D one-dimensional convolutions, then connect to a lightweight long-dependency modeler based on state space. The four outputs are aligned to a common dimension via linear projection and then fed into a coupled gated sensing mechanism (C-GaPe) and the heterogeneous specialized cooperative array (HeCo-Array), where each modality is encoded. A basic score is assigned and quality control credibility is injected, then a cross-modal support matrix is ​​formed using projection similarity. Then, the gate vector is obtained through maximum gating. Therefore, according to The fused representation is then fed into the detection head to output the multi-label probability. ; For high-dimensional fusion feature representation, All of them are the first Classifier parameters, , for the first The probability of a class For the real number field. Finally, through a multi-task detection head, the intrusion probability, attack type, and interpretable heatmap are obtained.

[0111] In this invention, "heterogeneous" in heterogeneous dedicated cooperative arrays does not refer to a specific set of networks, but rather emphasizes the differences and complementarity of operator families: the array covers at least two different computational paradigms of subheadings, which can be a combination of convolutional and attention families, or any combination of graph neural networks, recurrent / state space, and frequency domain operators; the illustrated implementation precisely provides a lightweight combination of four parallel families. The array adopts a "shared stem + dedicated head" organization: shared transformation Unified public characteristics, each specialized head Logarithmic risk arises independently on it. With probability and with diversity regularization This encourages each head to learn complementary representations, satisfying heterogeneous requirements while adhering to onboard computing power and latency budget constraints.

[0112] Set up a tag relevance collaboration layer (Diagonal elements are 0), perform coordinated sag correction on the outputs of each sub-network:

[0113]

[0114] To obtain the basic logarithmic risk in the dedicated subnetwork, To correct for logarithmic risk, To learn in conjunction with sub-network parameters, This indicates the co-occurrence tendency of i and j. This indicates the mutual exclusion tendency of i and j. As the first fundamental logarithmic risk component, This is the fourth basic logarithmic risk component.

[0115] Through steps S201-S205 described above, a multimodal intrusion detection model is constructed. This constructed model is then used to fuse and analyze real-time state data, thereby generating a real-time multi-label risk probability vector representing the type of attack suffered. This step, involving online inference and the generation of a multi-label risk index, is executed in real-time during UAV inspection missions. This is the core online application of the UAV inspection intrusion detection method of this invention. The core task is to continuously analyze the real-time state of the UAV using the deployed multimodal intrusion detection model and output structured risk assessment results.

[0116] Tag relevance collaboration layer, with dimensions set as follows Furthermore, the correlation matrix R, which has zero diagonal, represents the basic logarithmic risk given by the directional subnet. The previous linear cooperative operation was performed once. ,in Indicates label The co-occurrence tendency Indicates label To mitigate the mutual exclusion tendency, matrix parameters and subnet parameters are jointly learned end-to-end. During the training phase, weighted multi-label binary cross-entropy is used as the main loss, and array diversity regularization is further applied. If necessary, sparsity and symmetry constraints are added to R as a preferred implementation method, thereby suppressing noise co-occurrence and obtaining interpretable correlation structures even with a small sample size. To ensure that risk outputs of different categories are comparable at the threshold, temperature-bias calibration is used for each category. Then, based on the configured threshold, a multi-label binary vector is obtained; multi-label risk index. Further mapped to a comprehensive risk index according to weights ,in It can be preset according to task attributes or updated online to drive the hierarchical uplink strategy.

[0117] After generating a real-time multi-label risk probability vector, adaptive risk assessment is performed based on the real-time multi-label risk probability vector, and a hierarchical uplink strategy is executed to obtain the UAV inspection intrusion detection results. Among these, the adaptive risk assessment and the execution of the hierarchical uplink strategy are key to achieving intelligent communication and efficient response. The core task is to transform the abstract risk index into a specific risk level and execute differentiated, resource-optimized communication strategies accordingly.

[0118] Prior to this, an adaptive risk assessment algorithm and a tiered uplink strategy can be constructed. Specifically, constructing the adaptive risk assessment algorithm includes: weighting the real-time multi-label risk probability vector according to preset risk weights to obtain a comprehensive risk index under the current situation; and dividing the risk level into three levels—high, medium, and low—based on the comprehensive risk index and adaptively set high and low thresholds, thus constructing an adaptive risk assessment algorithm for adaptive risk assessment. The formula for calculating the comprehensive risk index is as follows:

[0119] Let n be the comprehensive risk index at time n. Let n be the weight coefficient of the k-th type of risk at time n. Let be the probability of the k-th type of risk occurring at time n.

[0120] The process of constructing a tiered uplink strategy includes: When the current risk level is high, a high-priority uplink strategy is generated. The high-priority uplink strategy includes: a high-priority alarm data packet, and uplinking to the ground terminal by preempting resources through the communication link. The high-priority alarm data packet encapsulates at least the following data: risk token, multi-label attack vector, comprehensive risk index, and key forensic summary data for post-event analysis. When the current risk level is medium, a medium-priority uplink strategy is generated; the medium-priority uplink strategy includes: a lightweight alarm data packet and sending it to the ground terminal; the lightweight alarm data packet encapsulates a risk token and a multi-tag attack vector; When the current risk level is low, a low-priority uplink strategy is generated. The low-priority uplink strategy includes: not triggering real-time uplink alarms, recording the current risk index, status slice and timestamp in the drone's local memory, and transmitting them back to the ground in batch compression when the communication link is idle or bandwidth is sufficient.

[0121] In one specific implementation, adaptive risk assessment and tiered uplink strategy execution include: The tag risk index is mapped to a composite risk index (CRI) based on weights. Then, based on the task scenario, high / low thresholds are adaptively set to divide the risk into three levels: high, medium, and low. Then, a response is made for different risk levels. When the risk level is high, it is immediately reported with high priority and accompanied by key evidence. When it is medium, only a lightweight token (risk index) is reported. When it is low, it is recorded locally and sent back in batches when necessary.

[0122] In some optional implementations, after performing adaptive risk assessment based on real-time multi-label risk probability vectors and executing a tiered uplink strategy, the method further includes: A buffer zone is set up on the drone. The buffer zone adopts a circular queue and is equipped with metadata index to store alarm and evidence data to be reported when the communication link is interrupted. When the communication link is restored, the priority of the backhaul is calculated based on the freshness and risk level of the alarm and evidence data to be reported, and the compensation backhaul is performed in order of priority.

[0123] Specifically, to ensure that alarms are not lost during link anomalies and that efficient compensation is achieved after recovery, this invention further sets up a link failure buffer and redundant backhaul mechanism, with the buffer size based on the "maximum expected link failure duration". With the highest reporting rate The system is configured based on the product of risk and risk levels, with a 20% to 30% security margin. The cache uses a circular queue and is equipped with metadata indexes (including timestamps, Comprehensive Risk Index (CRI), tag vectors, and evidence fingerprints). When space is limited, low-risk heartbeat frames are selectively discarded while high-risk summaries and evidence elements are retained. During recovery and transmission, a priority queue scheduling based on "CRI × Freshness Decay" is initiated (freshness decays exponentially with arrival time), achieving the sequence of "token first, evidence following, batch compensation": first, risk tokens and key summaries are quickly uploaded; then, hash chain summaries and evidence scores from each dedicated subnet are transmitted back; finally, low-priority records are compensated in batches. If bandwidth remains limited, low-risk heartbeats are discarded or downsampled and merged for transmission only, provided the integrity of high / medium risk alarms is maintained.

[0124] like Figure 3 The diagram shows the overall flowchart of the deep learning-based UAV inspection intrusion detection method provided by this invention. Specifically: a multimodal intrusion detection model is pre-trained, a four-source monitoring parameter table is generated and loaded with C-GaPe+HeCo-Array model weights, the UAV simultaneously collects four-source data from NetFlow, ToPo, GPS and IMU, and performs cross-modal fusion of C-GaPe + HeCo-Array subnet inference to generate a four-bit binary attack vector. The system explicitly identifies one or more attacks currently being attacked, followed by an ALT (Alternate Threshold) threshold assessment. At the severe (high) level, token encapsulation and forensics are performed; at the medium level, only token encapsulation is performed; and at the low level, only heartbeat recording is performed. Based on the execution result, the system further determines whether the primary uplink is available. If available, encrypted information is sent to the remote control; if unavailable, it is written to the circular buffer for multipath backhaul.

[0125] This invention provides a deep learning-based UAV intrusion detection method. On the data side, it unifies and aligns NetFlow, network topology, GPS, and IMU modalities, performing time base unification and drift correction to ensure a complete state slice can be formed at any given time. On the model side, it proposes a coupled gating perception mechanism (C-GaPe), dynamically evaluating and fusing the contributions of each modality through credibility injection and cross-modal gating. The backend uses a heterogeneous dedicated cooperative array (HeCo-Array) for parallel discrimination, natively outputting K-dimensional multi-label risk and binary attack vectors to adapt to concurrent threats. On the communication decision side, it uses a comprehensive risk index (CRI) to drive adaptive hierarchical threshold (ALT) and hysteresis to achieve strict / medium / wide-range uplink tiers. This invention forms a practical closed loop through network-air cross-domain fusion, concurrent threat characterization, and communication perception collaboration, exhibiting stronger robustness against complex attacks, higher alarm efficiency under constrained links, and more complete evidence collection protection. It is particularly suitable for task environments such as power line inspection, which require both real-time performance and traceability.

[0126] This invention provides a deep learning-based intrusion detection method for UAV inspection, targeting power grid UAV inspection scenarios to improve the intrusion resistance and event traceability of UAV data links. In the offline stage, a multi-label model is trained based on network traffic, network topology, GPS coordinates, and inertial measurement unit data. The trained detection model is then loaded onto the UAV for online inference. During UAV flight, the aforementioned data streams are collected synchronously. To accurately characterize the threat situation, this invention proposes a Coupled Gated Perception (C-GaPe) mechanism to dynamically fuse cross-modal features and utilizes a Heterogeneous Cooperative Array (HeCo-Array) to assign directional sub-networks to output a multi-label risk probability vector of length K. and its binary attack vector (Current Implementation) =4) Used to indicate one or more attacks suffered; the system uses Adaptive Layered Thresholding (ALT) and hysteresis to implement hierarchical uplink risk assessment. High-risk attacks send a fast response token and forensic digest, medium-risk attacks send only the token, and low-risk attacks maintain heartbeat and local auditing; when the main chain is abnormal, a chain break buffer-redundant backhaul is triggered to ensure that alarm delays are not lost; when a suspected intrusion event is detected, the drone encapsulates the encrypted alarm data into an identity token frame and transmits it to the remote control; the remote control parses out a multi-label attack vector of length K. In addition, a summary of the Comprehensive Risk Index (CRI) is provided to support fine-grained assessment and rapid forensics of concurrent threats.

[0127] Example 2: Based on the same inventive concept, this invention also provides a deep learning-based drone inspection intrusion detection system, the structure of which is as follows: Figure 4 As shown, the system includes: The status data acquisition module 401 is used to acquire the real-time status data of the UAV, which includes the UAV's real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data. The risk probability vector generation module 402 is used to perform fusion analysis on real-time status data using a pre-built multimodal intrusion detection model to generate a real-time multi-label risk probability vector representing the type of attack suffered. The intrusion detection result acquisition module 403 is used to perform adaptive risk assessment based on real-time multi-label risk probability vector and execute a hierarchical uplink strategy to obtain the intrusion detection result of UAV inspection. The hierarchical uplink strategy includes: classifying the assessed current risk into different risk levels, and determining the data content and priority to be reported to the ground end according to the risk level.

[0128] Preferably, the system also includes a multimodal intrusion detection model building module, used for: The system collects raw data from drones under normal inspection and pre-set attack scenarios, covering four modes: network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit. It then performs time-synchronized preprocessing to form time-aligned unified sample slices. Based on the credibility weight, the unified sample slice is optimized by quality detection and credibility assessment to obtain the optimized unified sample slice; A coupled gated sensing mechanism is used to dynamically weight and fuse the modal features of each modality in the optimized unified sample slice to generate a unified high-dimensional fusion feature. A heterogeneous targeted cooperative array is used to process high-dimensional fusion features and output a multi-label risk probability vector representing the type of attack suffered. The heterogeneous targeted cooperative array contains multiple parallel targeted sub-networks, each of which is used to identify a specific attack type. By conducting multi-label-guided end-to-end joint training, the coupled gating sensing mechanism and the heterogeneous specialized cooperative array are jointly optimized to obtain the mapping relationship from the original data to the multi-label risk probability vector. Based on the mapping relationship, a multimodal intrusion detection model is constructed.

[0129] Preferably, the multimodal intrusion detection model building module is specifically used for: The system collects raw data from drones under four modes: network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit, both during normal inspections and under pre-set attack scenarios. Based on the original timestamps corresponding to each modality in the original data, the nearest observations are aligned to the same time window to form a preliminary correspondence and obtain the preliminary corresponding timestamps for each modality. Using geospatial positioning time data as a reference time base, the preliminary corresponding timestamps of each mode are corrected by estimating a linear clock mapping algorithm to obtain the corrected data of each mode. The corrected modal data are resampled and aligned at equally spaced time grid points. Within an adaptive sliding time window centered on the current time, the corrected modal data are stitched together to form a time-aligned unified sample slice.

[0130] Preferably, the multimodal intrusion detection model building module is also specifically used for: Using geospatial positioning time data as a reference time base, the least squares method is used to estimate the clock scaling factor and offset of each mode at the selected anchor point. By estimating the linear clock mapping algorithm, the clock scaling factor and offset on the initial corresponding timestamp of each mode are corrected, and the corrected mode data are obtained based on the corrected timestamp. The original timestamps for each modality are as follows: , Various modes NetFlow is network traffic data, Topo is dynamic network topology data, GPS is geospatial positioning data, and IMU is inertial measurement unit data. The algorithm for estimating linear clock mapping is as follows: , in, This represents the initial corresponding timestamp for mode m. Indicates to Corrected timestamp mapped to GPS reference time base; The formula for the least squares method is as follows:

[0131] For least squares at anchor point pairs Upper fitting, This is the time index of the anchor point in mode m. This is the time index corresponding to the GPS sequence anchor point of mode m. Let m be the clock scaling factor for mode m. Let m be the clock offset of mode m. This is the index of the anchor point in mode m. This is the index of the corresponding anchor point in the GPS sequence. Number the anchor points. This is the clock scaling factor. This is the clock offset coefficient.

[0132] Preferably, the multimodal intrusion detection model building module is also specifically used for: Based on statistical sufficiency and temporal consistency, independent quality indices for each modality are calculated. By fusing the independent quality metrics of each modality with the cross-modal alignment residuals, the modal credibility of each modality in the current sample slice is calculated, and the overall credibility of the current sample slice is also calculated. Based on the comprehensive credibility and the preset quality threshold, quality gating is performed to obtain an optimized unified sample slice. The formula for calculating modal confidence is as follows: ; Let m be the confidence level of the m-th mode at time n. Let m be the coverage of the m-th mode. Let m be the stability of the m-th mode. To align residuals based on cross-modal temporal consistency, To adjust the power exponent of quality discrimination, The preset lower limit of credibility weight, To truncate the result to an interval ; The formula for calculating overall credibility is as follows:

[0133] The overall confidence level of all modes at time n is given by: Let m be the confidence level of the m-th mode at time n. Let be the confidence weight of the m-th mode.

[0134] Preferably, the multimodal intrusion detection model building module is also specifically used for: Encode each modal data in the optimized unified sample slice to obtain the encoding features of each modality; The baseline importance score of each modality is calculated based on the modality coding features, and the confidence of each modality is injected into the baseline importance score to obtain the corrected score result. Calculate the similarity of each modality's encoded features in the common projection space to form a cross-modal support matrix; The corrected scoring results are coupled and updated based on the cross-modal support matrix to obtain the final gating vector; The modality-coded features are weighted and summed based on the final gating vector to generate a unified high-dimensional fusion feature. The formulas for each modality coding feature are as follows:

[0135] For the encoding features of the m-th modality, For the preprocessed features of the m-th mode, For the encoder of the m-th mode; The formula for baseline importance scoring is as follows:

[0136] Assign a baseline importance score to the m-th mode on sample n. Encode features for each modality Projected onto common dimensions linear mapping, For the encoding features of the m-th modality, For context vectors, Let be the linear readout parameters for the m-th mode. for transpose, This is the bias for the m-th mode; The revised scoring formula is as follows:

[0137] The revised scoring results Assign a baseline importance score to the m-th mode on sample n. The injection strength coefficient, Let m be the confidence level of the m-th mode at time n; The formula for calculating the cross-modal support matrix is ​​as follows:

[0138]

[0139] For time window indexing, For the mode within the nth time window For modes Support For the mode within the nth time window Status coding, For the mode within the nth time window Status coding, , For modal indexing, For the mode within the nth time window With mode Self-loop; The calculation process for the final gating vector includes the following: The initial gating vector is as follows:

[0140]

[0141]

[0142] The cross-modal support is as follows:

[0143]

[0144] The final gating vector is as follows:

[0145]

[0146]

[0147] in, Index for time windows; And m and l are modal indices; The total number of modes; Indexed by category; Total number of categories; For iteration step index; Let be the initial gating vector, given by get, Let m be the initial gating vector for mode m. Scoring results after injecting credibility into the modality; For cross-modal support vectors, Let l be the cross-modal support vector of mode m, where m and l are mode indices. Let be the final gating vector of mode l, where This is the scoring vector after adding cross-modal support. Update the coefficients for the weights; For the final gating / fusion weights, satisfy ; This is the scoring vector after adding cross-modal support; For each The vector formed This represents the scoring result after the confidence level of each modality injection, the th Each component is , Let be the final gating vector of the (t+1)th iteration. Let be the final gating vector for the t-th iteration.

[0148] Preferably, the multimodal intrusion detection model building module is also specifically used for: By sharing a transformation function, high-dimensional fusion features are used to obtain unified common features; A unified common feature is input in parallel into k structurally distinct directional subnetworks to obtain an independent intermediate representation and log risk score for each directional subnetwork. The k structurally distinct directional subnetworks are constructed using at least two families of neural network operators with different computational paradigms to capture different time scales and related patterns. By using an off-diagonal label correlation matrix, the log risk score output by each targeted sub-network is collaboratively corrected to obtain a log risk vector, thereby obtaining the co-occurrence or mutual exclusion relationship between different attack labels. Convert the logarithmic risk vector into a multi-label risk probability vector; The shared transformation function is as follows:

[0149] To share the transformation function, As a high-dimensional fusion feature, To unify public characteristics; The intermediate formula is as follows:

[0150] This represents the intermediate feature representation output by the k-th specialized sub-network. To unify common characteristics, It is a linear discriminant layer; The formula for collaborative correction is as follows:

[0151]

[0152] To obtain the basic logarithmic risk in the dedicated subnetwork, To correct for logarithmic risk, To learn in conjunction with sub-network parameters, This indicates the co-occurrence tendency of i and j. This indicates the mutual exclusion tendency of i and j. As the first fundamental logarithmic risk component, This is the fourth fundamental logarithmic risk component; The formula for calculating the multi-label risk probability vector is as follows:

[0153] As a high-dimensional fusion feature, The parameters of the first classifier for the k-th class are... The parameters of the second classifier for the k-th class are... The Sigmoid function outputs multi-label probabilities for each category. Let be the probability of the k-th type of risk occurring at time n. This represents the total number of tags.

[0154] Preferably, the system also includes an adaptive risk assessment building module for: The real-time multi-label risk probability vector is weighted according to preset risk weights to obtain the comprehensive risk index under the current situation; Based on the comprehensive risk index and adaptively set high and low thresholds, the risk level is divided into three levels: high, medium and low, and an adaptive risk assessment algorithm for adaptive risk assessment is constructed. The formula for calculating the comprehensive risk index is as follows:

[0155] Let n be the comprehensive risk index at time n. Let n be the weight coefficient of the k-th type of risk at time n. Let be the probability of the k-th type of risk occurring at time n.

[0156] Preferably, the system also includes a hierarchical uplink strategy construction module, used for: When the current risk level is high, a high-priority uplink strategy is generated. The high-priority uplink strategy includes: a high-priority alarm data packet, and uplinking to the ground terminal by preempting resources through the communication link. The high-priority alarm data packet encapsulates at least the following data: risk token, multi-label attack vector, comprehensive risk index, and key forensic summary data for post-event analysis. When the current risk level is medium, a medium-priority uplink strategy is generated; the medium-priority uplink strategy includes: a lightweight alarm data packet and sending it to the ground terminal; the lightweight alarm data packet encapsulates a risk token and a multi-tag attack vector; When the current risk level is low, a low-priority uplink strategy is generated. The low-priority uplink strategy includes: not triggering real-time uplink alarms, recording the current risk index, status slice and timestamp in the drone's local memory, and transmitting them back to the ground in batch compression when the communication link is idle or bandwidth is sufficient.

[0157] Preferably, the system further includes a buffered return module for: A buffer zone is set up on the drone. The buffer zone adopts a circular queue and is equipped with metadata index to store alarm and evidence data to be reported when the communication link is interrupted. When the communication link is restored, the priority of the backhaul is calculated based on the freshness and risk level of the alarm and evidence data to be reported, and the compensation backhaul is performed in order of priority.

[0158] Example 3: Based on the same inventive concept, such as Figure 5 As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.

[0159] The processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, and it is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in a readable storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of the deep learning-based UAV inspection intrusion detection method in the above embodiments.

[0160] Example 4: Based on the same inventive concept, this invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory). This readable storage medium is a memory device within an electronic device used to store programs and data. It is understood that the readable storage medium here can include both the built-in storage medium within the electronic device and extended storage media supported by the electronic device. The storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more executable programs (including program code). It should be noted that the storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the storage medium to implement the steps of the deep learning-based UAV inspection intrusion detection method described in the above embodiments.

[0161] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0162] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0163] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0164] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading the present invention, they can still make various changes, modifications or equivalent substitutions to the specific implementation of the application, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending approval.

Claims

1. A deep learning-based unmanned aerial vehicle (UAV) intrusion detection method, characterized in that, include: Acquire real-time status data of the UAV, including real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data. The real-time status data is fused and analyzed using a pre-built multimodal intrusion detection model to generate a real-time multi-label risk probability vector representing the type of attack suffered. Adaptive risk assessment is performed based on the real-time multi-label risk probability vector, and a hierarchical uplink strategy is executed to obtain the UAV inspection intrusion detection results. The tiered uplink strategy includes: classifying the assessed current risks into different risk levels, and determining the content and priority of data to be reported to the ground based on the risk level.

2. The method according to claim 1, characterized in that, The construction of the multimodal intrusion detection model includes: The system collects raw data from drones under four modes: network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit, under normal inspection and pre-set attack scenarios. It then performs time-synchronized preprocessing to form time-aligned unified sample slices. Based on the credibility weight, the unified sample slice is subjected to quality detection and credibility evaluation optimization to obtain the optimized unified sample slice. A coupled gated sensing mechanism is used to dynamically weight and fuse the modal features of each modality in the optimized unified sample slice to generate a unified high-dimensional fusion feature. The high-dimensional fusion features are processed using a heterogeneous specialized cooperative array to output a multi-label risk probability vector representing the type of attack suffered. The heterogeneous specialized cooperative array contains multiple parallel specialized sub-networks, each of which is used to identify a specific type of attack. By conducting multi-label-guided end-to-end joint training, the coupled gating sensing mechanism and the heterogeneous dedicated cooperative array are jointly optimized to obtain the mapping relationship from the original data to the multi-label risk probability vector, and a multimodal intrusion detection model is constructed based on the mapping relationship.

3. The method according to claim 2, characterized in that, The data collection drone performs time-synchronized preprocessing on raw data from four modes—network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit—under normal inspection and pre-defined attack scenarios, generating time-aligned unified sample slices, including: The system collects raw data from drones under four modes: network traffic, dynamic network topology, geospatial positioning, and inertial measurement unit, both during normal inspections and under pre-set attack scenarios. Based on the original timestamps corresponding to each modality data in the original data, the neighboring observations are aligned to the same time window to form a preliminary correspondence and obtain the preliminary corresponding timestamps of each modality. Using geospatial positioning time data as a reference time base, the preliminary corresponding timestamps of each mode are corrected by estimating a linear clock mapping algorithm to obtain the corrected data of each mode. The corrected modal data are resampled and aligned at equally spaced time grid points. Within an adaptive sliding time window centered on the current time, the corrected modal data are stitched together to form a time-aligned unified sample slice.

4. The method according to claim 3, characterized in that, The method uses geospatial positioning time data as a reference time base, and corrects the initial corresponding timestamps of each mode through an estimation linear clock mapping algorithm to obtain corrected modal data, including: Using geospatial positioning time data as a reference time base, the least squares method is used to estimate the clock scaling factor and offset of each mode at the selected anchor point. By estimating the linear clock mapping algorithm, the clock scaling factor and offset on the initial corresponding timestamp of each mode are corrected, and the corrected mode data are obtained based on the corrected timestamp. The original timestamps for each modality are as follows: , Various modes NetFlow is network traffic data, Topo is dynamic network topology data, GPS is geospatial positioning data, and IMU is inertial measurement unit data. The algorithm for estimating linear clock mapping is as follows: , in, This represents the initial corresponding timestamp for mode m. Indicates to Corrected timestamp mapped to GPS reference time base; The formula for the least squares method is as follows: For least squares at anchor point pairs Upper fitting, This is the time index of the anchor point in mode m. This is the time index corresponding to the GPS sequence anchor point of mode m. Let m be the clock scaling factor for mode m. Let m be the clock offset of mode m. This is the index of the anchor point in mode m. This is the index of the corresponding anchor point in the GPS sequence. Number the anchor points. This is the clock scaling factor. This is the clock offset coefficient.

5. The method according to claim 2, characterized in that, The process of optimizing the unified sample slice based on credibility weights through quality detection and credibility assessment to obtain an optimized unified sample slice includes: Based on statistical sufficiency and temporal consistency, independent quality indices for each modality are calculated. By integrating the independent quality metrics of each modality with the cross-modal alignment residuals, the modal credibility of each modality in the current sample slice is calculated, and the overall credibility of the current sample slice is calculated. Based on the comprehensive credibility and the preset quality threshold, quality gating is performed to obtain an optimized unified sample slice. The formula for calculating the modal confidence level is as follows: ; Let m be the confidence level of the m-th mode at time n. Let m be the coverage of the m-th mode. Let m be the stability of the m-th mode. To align residuals based on cross-modal temporal consistency, To adjust the power exponent of quality discrimination, The preset lower limit of credibility weight, To truncate the result to an interval ; The formula for calculating the overall credibility is as follows: The overall confidence level of all modes at time n is given by: Let m be the confidence level of the m-th mode at time n. Let be the confidence weight of the m-th mode.

6. The method according to claim 2, characterized in that, The coupled gated sensing mechanism is used to dynamically weight and fuse the modal features of each segment in the optimized unified sample slice to generate unified high-dimensional fused features, including: Encode each modal data in the optimized unified sample slice to obtain the encoding features of each modality; Based on the modality coding features, the baseline importance score of each modality is calculated, and the confidence level of each modality is injected into the baseline importance score to obtain the corrected score result. Calculate the similarity of each modality's encoded features in the common projection space to form a cross-modal support matrix; The corrected scoring results are coupled and updated based on the cross-modal support matrix to obtain the final gating vector; Based on the final gating vector, the modality coding features are weighted and summed to generate a unified high-dimensional fusion feature; The formulas for each modality coding feature are as follows: For the encoding features of the m-th modality, For the preprocessed features of the m-th mode, For the encoder of the m-th mode; The formula for scoring the baseline importance is as follows: Assign a baseline importance score to the m-th mode on sample n. Encode features for each modality Projected onto common dimensions linear mapping, For the encoding features of the m-th modality, For context vectors, Let be the linear readout parameters for the m-th mode. for transpose, This is the bias for the m-th mode; The revised scoring formula is as follows: The revised scoring results Assign a baseline importance score to the m-th mode on sample n. The injection strength coefficient, Let m be the confidence level of the m-th mode at time n; The formula for calculating the cross-modal support matrix is ​​as follows: For time window indexing, For the mode within the nth time window For modes Support For the mode within the nth time window Status coding, For the mode within the nth time window Status coding, , For modal indexing, For the mode within the nth time window With mode Self-loop; The calculation process for the final gating vector includes the following: The initial gating vector is as follows: The cross-modal support is as follows: The final gating vector is as follows: in, Index for time windows; ,and For modal indexing; The total number of modes; Indexed by category; Total number of categories; For iteration step index; Let be the initial gating vector, given by get, Let m be the initial gating vector for mode m. Scoring results after injecting credibility into the modality; For cross-modal support vectors, Let l be the cross-modal support vector of mode m, where m and l are mode indices. Let be the final gating vector of mode l, where This is the scoring vector after adding cross-modal support. Update the coefficients for the weights; For the final gating / fusion weights, satisfy ; This is the scoring vector after adding cross-modal support; For each The vector formed This represents the scoring result after the confidence level of each modality injection, the th Each component is , Let be the final gating vector of the (t+1)th iteration. Let be the final gating vector for the t-th iteration.

7. The method according to claim 2, characterized in that, The process employs a heterogeneous specialized cooperative array to process the high-dimensional fused features, outputting a multi-label risk probability vector representing the type of attack suffered, including: The high-dimensional fusion features are then transformed using a shared transformation function to obtain unified common features; The unified common features are input in parallel into k structurally different directional subnetworks to obtain independent intermediate representations and log risk scores for each directional subnetwork; the k structurally different directional subnetworks are constructed using at least two families of neural network operators with different computational paradigms to capture different time scales and related patterns. By using an off-diagonal label correlation matrix, the log risk score output by each targeted sub-network is collaboratively corrected to obtain a log risk vector, thereby obtaining the co-occurrence or mutual exclusion relationship between different attack labels. The logarithmic risk vector is converted into a multi-label risk probability vector; The shared transformation function is as follows: To share the transformation function, As a high-dimensional fusion feature, To unify public characteristics; The intermediate representation formula is as follows: This represents the intermediate feature representation output by the k-th specialized sub-network. To unify common characteristics, It is a linear discriminant layer; The collaborative correction formula is as follows: To obtain the basic logarithmic risk in the dedicated subnetwork, To correct for logarithmic risk, To learn in conjunction with sub-network parameters, This indicates the co-occurrence tendency of i and j. This indicates the mutual exclusion tendency of i and j. As the first fundamental logarithmic risk component, This is the fourth fundamental logarithmic risk component; The formula for calculating the multi-label risk probability vector is as follows: As a high-dimensional fusion feature, The parameters of the first classifier for the k-th class are... The parameters of the second classifier for the k-th class are... The Sigmoid function outputs multi-label probabilities for each category. Let be the probability of the k-th type of risk occurring at time n. This represents the total number of tags.

8. The method according to claim 1, characterized in that, Before performing adaptive risk assessment based on the real-time multi-label risk probability vector and executing a tiered uplink strategy to obtain the UAV inspection intrusion detection result, the method further includes: The real-time multi-label risk probability vector is weighted according to preset risk weights to obtain the comprehensive risk index under the current situation; Based on the comprehensive risk index and the adaptively set high and low thresholds, the risk level is divided into three levels: high, medium and low, and an adaptive risk assessment algorithm for adaptive risk assessment is constructed. The formula for calculating the comprehensive risk index is as follows: Let n be the comprehensive risk index at time n. Let n be the weight coefficient of the k-th type of risk at time n. Let be the probability of the k-th type of risk occurring at time n.

9. The method according to claim 1, characterized in that, Before performing adaptive risk assessment based on the real-time multi-label risk probability vector and executing the tiered uplink strategy, the method further includes: When the current risk level is high, a high-priority uplink strategy is generated; the high-priority uplink strategy includes: a high-priority alarm data packet, and uplinking to the ground terminal by preempting resources through the communication link; the high-priority alarm data packet encapsulates at least the following data: risk token, multi-label attack vector, comprehensive risk index, and key forensic summary data for post-event analysis; When the current risk level is medium, a medium-priority uplink strategy is generated; the medium-priority uplink strategy includes: a lightweight alarm data packet and sending it to the ground terminal; the lightweight alarm data packet encapsulates a risk token and a multi-tag attack vector; When the current risk level is low, a low-priority uplink strategy is generated. The low-priority uplink strategy includes: not triggering real-time uplink alarms, recording the risk index, status slice and timestamp of the current moment in the local memory of the UAV, and transmitting them back to the ground end in batch compression when the communication link is idle or the bandwidth is sufficient.

10. The method according to any one of claims 1-9, characterized in that, After performing adaptive risk assessment based on the real-time multi-label risk probability vector and executing a tiered uplink strategy, the method further includes: A buffer zone is set up on the drone terminal. The buffer zone adopts a circular queue and is equipped with a metadata index. It is used to store alarm and evidence data to be reported when the communication link is interrupted. When the communication link is restored, the priority of the backhaul is calculated based on the freshness and risk level of the alarm and evidence data to be reported, and the compensation backhaul is performed in order of priority.

11. A deep learning-based unmanned aerial vehicle (UAV) inspection and intrusion detection system, characterized in that, include: The status data acquisition module is used to acquire real-time status data of the UAV, including real-time network traffic data, real-time dynamic network topology data, real-time geospatial positioning data, and real-time inertial measurement unit data. The risk probability vector generation module is used to perform fusion analysis on the real-time status data using a pre-built multimodal intrusion detection model to generate a real-time multi-label risk probability vector representing the type of attack suffered. The intrusion detection result acquisition module is used to perform adaptive risk assessment based on the real-time multi-label risk probability vector and execute a hierarchical uplink strategy to obtain the intrusion detection result of the UAV inspection. The tiered uplink strategy includes: classifying the assessed current risks into different risk levels, and determining the content and priority of data to be reported to the ground based on the risk level.

12. An electronic device, characterized in that, include: At least one processor and memory; The memory and processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the deep learning-based drone inspection intrusion detection method as described in any one of claims 1 to 10 is implemented.

13. A readable storage medium, characterized in that, It contains an executable program, which, when executed, implements the deep learning-based drone inspection intrusion detection method as described in any one of claims 1 to 10.