Unmanned aerial vehicle inspection equipment identification method and system based on multi-mode equipment fingerprints
Through the multimodal device fingerprint recognition method, combined with GraphSAGE and deep separable convolutional network, the problems of anti-counterfeiting, discrimination accuracy and adaptability in drone inspection are solved, and high-precision, low-resource consumption real-time identity authentication is achieved in complex environments.
Patent Information
- Application Number
- CN202510908653.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-16
AI Technical Summary
Existing drone inspection technologies suffer from poor anti-counterfeiting capabilities, insufficient discrimination accuracy, inability to adapt to feature drift, and difficulty in real-time deployment on the edge, resulting in low recognition accuracy and high maintenance costs in complex industrial inspection scenarios.
A multimodal device fingerprint recognition method is adopted to collect hardware, software, network and behavioral features, and the neighbor node information is fused through the GraphSAGE graph convolutional network. The feature fusion is performed using a deep separable convolutional network and adaptive weights. Combined with INT8 quantization and micro-step incremental learning, adaptive identity authentication is achieved.
It improves the recognition accuracy and anti-interference capability of drone inspection equipment, reduces hardware resource consumption, and realizes real-time and efficient recognition and adaptive feature drift processing on edge devices.
Smart Images

Figure CN120658479A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of drone safety technology, and specifically relates to a drone inspection equipment identification method and system based on multimodal device fingerprints. Background Art
[0002] In the drone inspection sector, mainstream identity authentication still revolves around single-modality device fingerprinting. The first approach is based on physical-layer hardware fingerprinting, building a hash matching database by collecting inherent defects such as CPU serial numbers, crystal oscillator deviation, or RF IQ distortion. The second approach focuses on network-layer protocol features, recording fields such as MAC addresses, TTL values, and packet intervals, and using rule engines or naive Bayesian methods for identification. The third approach focuses on behavioral-layer statistical information, utilizing one-dimensional sequences such as takeoff and landing times, flight duration, and battery discharge curves, and employing dynamic time warping or threshold comparison for identification.
[0003] The aforementioned single-modality solution can meet basic identification needs in laboratory environments, but it exhibits significant shortcomings in complex industrial inspection scenarios. First, hardware identification can be forged through firmware flashing or RF cloning, and network protocol fields can also be overwritten, resulting in weak anti-counterfeiting capabilities. Second, the physical differences between drones of the same model are minimal, making a single feature dimension insufficient for fine-grained differentiation. Field-measured misclassification rates often exceed 10%. Third, feature drift caused by factors such as battery aging and firmware upgrades cannot be adapted in real time. The system must be manually recalibrated to the baseline, which is costly and prone to leaving gaps. Fourth, conditions such as strong electromagnetic interference and extreme class imbalance can distort RF fingerprints or cause model collapse, resulting in suboptimal generalization performance. Fifth, existing research generally uses simple concatenation or voting to process cross-domain information. Sixth, common convolutional or recurrent network models often have over five million parameters, making real-time inference in the 100s of milliseconds impossible on low-power processors such as the Cortex-A53, making it difficult to meet the resource constraints of edge deployments. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides a method and system for identifying drone inspection equipment based on multimodal device fingerprints, which can solve the problems of poor anti-counterfeiting ability, insufficient distinction accuracy, inability to adapt to feature drift, and difficulty in real-time deployment on the edge side in the existing technology, thereby improving the recognition accuracy of drone inspection equipment.
[0005] The present invention provides the following technical solutions:
[0006] In a first aspect, a method for identifying drone inspection equipment based on multimodal device fingerprints is provided, comprising the following steps:
[0007] S1: Collect multimodal features of the unmanned aerial vehicle inspection equipment to be identified and perform preprocessing and dimensionality reduction. The multimodal features include hardware features, software features, network features, and behavioral features;
[0008] S2: Construct a k-nearest neighbor undirected graph based on the spatial proximity or topological proximity of drone inspection devices in the same scene and the same batch, and fuse the multimodal features of all neighboring nodes of the current node to update the current node;
[0009] S3: For each UAV inspection device, the multimodal features integrated with neighbor information are mapped into query, key, and value matrices, and the multimodal features are fused using adaptive weights.
[0010] S4: Based on the fusion feature vector output after fusion, the trained deep learning classification model is used to identify the identity of the drone and output the classification probability vector. Based on the information entropy of the probability vector, it is determined whether the current drone inspection equipment is a legal device.
[0011] Optionally, the hardware features include CPU serial number, RF transmit crystal oscillator relative deviation, IQ distortion vector and power supply ripple peak; the software features include firmware version number, operating system kernel hash and application signature; the network features include MAC address, IP message identifier, survival time sequence and adjacent data packet sending interval; the behavioral features include three-axis acceleration, three-axis angular velocity, battery terminal voltage and current.
[0012] Optionally, in step S1, the preprocessing and dimensionality reduction include: integrity check, anomaly removal, normalization, and principal component analysis dimensionality reduction;
[0013] After normalizing the behavioral features, information enhancement is performed on the behavioral features and then principal component analysis dimensionality reduction is performed; the information enhancement of the behavioral features is specifically as follows: the behavioral features are converted into a spectrum by short-time Fourier transform, and one-third octave pooling is applied to the spectrum.
[0014] Optionally, in step S2, a GraphSAGE graph convolutional network is used to perform multimodal feature fusion of neighbor nodes. The specific fusion formula is:
[0015]
[0016] in, is the hidden representation of node v in the l+1 layer, σ is the activation function, W l is the weight matrix of the lth layer, Mean is the mean aggregation function, N k (v) is the set of neighbor nodes of node v, is the hidden representation of node v’s neighbor node u in layer l.
[0017] Optionally, step S3 specifically comprises: for each UAV inspection device, mapping the hardware feature vector integrated with neighbor information into a query matrix, concatenating the software features, network features, and behavior features integrated with neighbor information into a key and value matrix, and fusing the multimodal features using adaptive weights;
[0018]
[0019] Z = σ(Attention(Q,K,V))
[0020] in, is the query matrix, n is the number of hardware feature vectors, d k is the dimension of the key vector; is the key matrix, m is the number of concatenated eigenvectors; is the value matrix, d v is the dimension of the value vector; the softmax function is normalized by row, is the scaling factor used to scale the dot product result to prevent the gradient from disappearing or exploding; σ is the activation function, and Z is the fused feature vector.
[0021] Optionally, in step S4, the deep learning classification model is based on a four-layer fully connected neural network, and two hidden layers are replaced by two depthwise separable convolutional layers; each depthwise separable convolutional layer includes pointwise convolution, depthwise convolution and activation function, extracts channel features by pointwise convolution in the spatial dimension, uses depthwise convolution to learn spatial correlation, and finally activates by activation function;
[0022] In step S4, the depthwise separable convolutional layer is compressed using the INT8 quantization method;
[0023] In step S4, when the information entropy of the probability vector exceeds the threshold, the current drone inspection device is an unknown device; otherwise, the current drone inspection device is a legitimate device; the information entropy H(p) of the probability vector is:
[0024]
[0025] Among them, p i is the probability vector of category i, and C is the total number of categories.
[0026] Optionally, the method further includes step S5: when feature drift occurs between the mean of the real-time fused feature vectors of a batch of drone inspection devices and the mean of their historical fused feature vectors, updating the parameters of the second depthwise separable convolutional layer and the last fully connected neural network layer of the deep learning classification model through micro-step incremental learning;
[0027] Specifically, for the same batch of UAV inspection equipment, when the Mahalanobis distance between the mean of their real-time fusion feature vector and the mean of their historical fusion feature vector exceeds the set threshold, feature drift occurs. The formula for obtaining the Mahalanobis distance is:
[0028]
[0029] Among them, μ0 is the mean of the fusion feature vectors of the same batch of UAV inspection equipment in the historical period, μ t is the mean of the real-time fusion feature vectors of a batch of UAV inspection equipment; D M μ t and the Mahalanobis distance of μ0, ∑ is the feature covariance matrix.
[0030] In the second aspect, a drone inspection equipment identification system based on multimodal device fingerprint is provided, comprising:
[0031] Acquisition module: collects multimodal features of the drone inspection equipment to be identified, including hardware features, software features, network features, and behavioral features;
[0032] Preprocessing and dimensionality reduction module: preprocesses and reduces the dimensionality of the collected multimodal features;
[0033] Fusion module: This module constructs a k-nearest neighbor undirected graph based on the spatial or topological proximity of drone inspection devices in the same scene and batch, and fuses the multimodal features of all neighboring nodes of the current node to update the current node. For each drone inspection device, the multimodal features that incorporate neighbor information are mapped into query, key, and value matrices, and the multimodal features are fused using adaptive weights.
[0034] Identification module: Based on the fusion feature vector output after fusion, the trained deep learning classification model is used to identify the drone's identity and output the classification probability vector. Based on the information entropy of the probability vector, it is determined whether the current drone inspection equipment is a legitimate device;
[0035] The adaptive update module updates the parameters of the second depth-wise separable convolutional layer and the last fully connected neural network of the deep learning classification model through micro-step incremental learning when feature drift occurs between the mean of the real-time fusion feature vector of a batch of drone inspection equipment and the mean of its historical fusion feature vector.
[0036] In a third aspect, a computer device is provided, comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the steps of the method for identifying drone inspection equipment based on multimodal device fingerprints as described in any one of the first aspects are implemented.
[0037] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program; when the computer program is executed by a processor, the steps of the method for identifying drone inspection equipment based on multimodal device fingerprints described in any one of the first aspects are implemented.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] (1) This application collects multimodal features including hardware features, software features, network features and behavioral features, and fuses the multimodal features by fusing neighbor information and cross-modal self-attention fusion, thereby avoiding the technical problems in the existing technology such as the identity of drone inspection equipment being easily forged, the high misclassification rate of equipment of the same model, authentication failure caused by feature drift, and insufficient reliability in complex attack scenarios, thereby achieving high-precision, anti-interference and adaptive identity authentication of drone inspection equipment in complex industrial environments.
[0040] (2) This application reduces the dimensionality of multimodal data through principal component analysis and also performs information enhancement on behavioral characteristics. In addition, the deep learning classification model of this application is based on a four-layer fully connected neural network, and two hidden layers are replaced by two depthwise separable convolutional layers; at the same time, the INT8 quantization method is used to compress the depthwise separable convolutional layers, so that this application can run stably for a long time under conditions of extremely low hardware resource overhead, which is particularly suitable for drone supervision scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow chart of the method for identifying drone inspection equipment based on multimodal device fingerprints of the present invention;
[0042] Figure 2 This is a structural block diagram of the drone inspection equipment identification system based on multimodal device fingerprints of the present invention. DETAILED DESCRIPTION
[0043] The present invention will be further described below with reference to the accompanying drawings. The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. It should be noted that the term "comprising" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0044] Example 1
[0045] like Figure 1As shown, a method for identifying drone inspection equipment based on multimodal device fingerprints is provided, comprising the following steps:
[0046] S1: Collect multimodal features of the drone inspection equipment to be identified and perform preprocessing and dimensionality reduction.
[0047] Specifically, multimodal features include hardware features, software features, network features, and behavioral features. Drones in complex environments can experience instantaneous frame loss, electromagnetic glitches, and aerodynamic disturbances. To eliminate non-steady-state anomalies, the present invention uses triple standard deviation threshold detection to delete outlier frames, then performs z-score normalization to suppress cross-channel amplitude differences, and finally uses principal component analysis to compress the 128-dimensional original features to 32 dimensions, with the cumulative contribution rate remaining at 95%. The proportion of abnormal frames has dropped from the original 3.9% to 0.17%, and the end-side inference delay has been simultaneously shortened by 24 milliseconds, achieving the dual benefits of noise reduction and lightweighting.
[0048] In this embodiment, the hardware features include: CPU serial number, RF transmit crystal oscillator relative deviation, IQ distortion vector, and power supply ripple peak. The CPU serial number is usually a 32-bit integer, and the unit of RF transmit crystal oscillator relative deviation is usually ppm.
[0049] Software features (OS kernel fingerprints) include: firmware version number, OS kernel hash, and application signature;
[0050] Network characteristics include: MAC address, IP packet identifier, survival time sequence and interval between adjacent data packets;
[0051] Behavioral characteristics include: three-axis acceleration, three-axis angular velocity, battery terminal voltage and current.
[0052] When collecting multimodal features, a unified time base t0 is established through a 25-40MHz TCXO, and the hardware clock count t clk , operating system timestamp t os , Network instruction conversion time t net and the behavioral incentive triggering time t beh Adjusted to:
[0053] t sync =t clk +Δt1=t os +Δt2=t net +Δt3=t beh +Δt4,|Δt i |≤0.2μs
[0054] Hardware Domain raw ={CPU serial number id cpu, relative deviation of RF crystal oscillator Δf xtal (ppm), IQ distortion vector Power supply ripple peak-to-peak value V rpp};
[0055] Software Domains raw ={Firmware version FW ver , operating system hash OS hash_256 、Application Signature APP sig_256};
[0056] Network domain n raw ={MAC, message identification sequence ID seq , TTL vector Interval between sending adjacent data packets };
[0057] Behavioral Domain Synchronous sampling,
[0058] The behavior domain sampling rate f s =200Hz, window length L win =2s.
[0059] The collected multimodal data is checked for integrity using CRC-32. Specifically, the CRC-32 is used to generate the polynomial:
[0060] G(x)=x 32 +x 26 +x 23 +x 22 +x 16 +x 12 +x 11 +x 10 +x 8 +x 7 +x 5 +x 4 +x 2 +1
[0061] Calculate the redundant code R(x) = x·x for each frame bit stream x 32 mod G(x), check condition R(x) = 0 is considered as CRC pass =1 and encapsulated into a 128-byte payload and sent to the edge node via CAN-FD (1 Mbit / s backplane rate).
[0062] Preprocessing and dimensionality reduction include: applying the three-sigma criterion to each modality data to perform outlier removal, z-score normalization, and principal component analysis for dimensionality reduction. Specific methods can refer to existing technologies. As an optional method:
[0063] |x-μ|>3σ
[0064] After deleting abnormal frames, the proportion of abnormal frames dropped from 3.9% to 0.17%. The remaining samples were then normalized. After principal component analysis compression, the dimension of the normalized sample matrix was compressed by 75%, and the inference delay was reduced by 24ms.
[0065] In this embodiment, after normalizing the behavioral features, information enhancement is performed on the behavioral features and then principal component analysis dimensionality reduction is performed; wherein, information enhancement on the behavioral features is specifically performed as follows: short-time Fourier transform is performed on the behavioral features to convert them into a spectrum, and one-third octave pooling is applied to the spectrum.
[0066] Information enhancement is specifically as follows: Apply short-time Fourier transform (window length N = 256, shift H = 64, Hanning window w[m]) to each IMU sequence x(t):
[0067]
[0068] Then the center frequency satisfies f c,n+1 =f c,n 2 1 / 3 ,f c,0 =300Hz, Divide the power spectrum into one-third octave bands and calculate the average energy within the bands:
[0069]
[0070] Motor vibration and power-frequency magnetic fields in interference scenarios can cause artifacts. This method first performs a short-time Fourier transform on the behavioral domain waveform and then applies one-third octave energy pooling to the spectrum, evenly spreading narrowband peaks to adjacent bands and enhancing robustness to vibration and electromagnetic noise.
[0071] S2: Construct a k-nearest neighbor undirected graph based on the spatial proximity or topological proximity of drone inspection equipment in the same scene and the same batch, and fuse the multimodal features of all neighbor nodes of the current node to update the current node.
[0072] The method of constructing a k-nearest neighbor undirected graph can refer to the existing technology. To address the problem of scarcity of early samples, the GraphSAGE neighborhood aggregation algorithm is introduced to supplement the graph structure information of a small number of similar nodes into the feature space.
[0073] In step S2, the GraphSAGE graph convolutional network is used to fuse the multimodal features of neighbor nodes. The specific fusion formula is:
[0074]
[0075] in, is the hidden representation of node v in the l+1 layer, σ is the activation function, W l is the weight matrix of the lth layer, Mean is the mean aggregation function, N k (v) is the set of neighbor nodes of node v, is the hidden representation of node v’s neighbor node u in layer l.
[0076] GraphSAGE uses a mean aggregator with a convolution depth of 2 layers and a hidden dimension of 128 to integrate historical interaction relationships between devices and improve recognition accuracy in cold start scenarios.
[0077] S3: For each drone inspection device, the multimodal features integrated with neighbor information are mapped into query, key, and value matrices, and the multimodal features are fused through adaptive weights.
[0078] Specifically, the hardware feature vector that integrates neighbor information is mapped into a query matrix. The software features, network features, and behavior features that integrate neighbor information are concatenated and mapped into a key and value matrix. The multimodal features are then fused using adaptive weights.
[0079]
[0080] Z = σ(Attention(Q,K,V))
[0081] in, is the query matrix, n is the number of hardware feature vectors, d k is the dimension of the key vector; is the key matrix, m is the number of concatenated eigenvectors; is the value matrix, d v is the dimension of the value vector; the softmax function is normalized by row, is the scaling factor used to scale the dot product result to prevent the gradient from disappearing or exploding; σ is the activation function, and Z is the fused feature vector.
[0082] S4: Based on the fusion feature vector output after fusion, the trained deep learning classification model is used to identify the identity of the drone and output the classification probability vector. Based on the information entropy of the probability vector, it is determined whether the current drone inspection equipment is a legal device.
[0083] The deep learning classification model is based on a four-layer fully connected neural network, with two hidden layers replaced by two depthwise separable convolutional layers. Specifically, the model consists of an input layer, a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, and a classifier (fully connected layer). Global average pooling is introduced before the classifier, and each depthwise separable convolutional layer includes pointwise convolution, depthwise convolution, and an activation function. First, pointwise convolution is performed in the spatial dimension to extract channel features, then depthwise convolution is used to learn spatial correlations, and finally, activation is performed using an activation function, typically ReLU + BatchNorm.
[0084] Furthermore, in step S4, the depthwise separable convolutional layers are compressed using the INT8 quantization method, which applies eight-bit integer quantization to weights and activations. Peak memory usage is 3.8MB. Based on a Cortex-A53 quad-core platform, single-frame inference latency is 78 milliseconds, throughput is 12.8 frames per second, and Top-1 accuracy reaches 99.43%, equivalent to a desktop floating-point model while consuming only 0.4 watts.
[0085] The specific formula for INT8 quantization is:
[0086]
[0087] The depthwise separable convolutional layer uses only 38% of the parameters of an equivalent fully connected layer, yet significantly reduces inference latency while maintaining expressiveness. Actual deployment data shows that the total model parameter count has been reduced from 3.2M to 1.8M, and the time required for single-sample recognition on edge devices has been reduced from 120ms to 80ms, while maintaining virtually unchanged overall accuracy. The introduction of the depthwise separable convolutional layer not only enhances the model's multi-scale perception capabilities but also meets the stringent real-time requirements of industrial sites.
[0088] The training of deep learning classification models can refer to existing technologies.
[0089] In this embodiment, in step S4, when the information entropy of the probability vector exceeds the threshold, the current drone inspection device is an unknown device; otherwise, the current drone inspection device is a legitimate device; the information entropy H(p) of the probability vector is:
[0090]
[0091] Among them, p i is the probability vector of category i, and C is the total number of categories.
[0092] When the information entropy value exceeds the threshold δ, the system determines that the target is an unknown or malicious device and immediately disconnects the link, triggering an abnormal alarm; when the entropy value does not exceed the threshold δ, it is a legitimate device and is allowed to access.
[0093] This application uses CRC-32 check to ensure that the frame integrity rate = 100% and the orthogonality of multi-domain original vectors ≥ 90%, providing a high signal-to-noise ratio data base for subsequent dimensionality reduction and attention fusion; the dimension of PCA compression in this application is 128 before and 32 after compression, and the cumulative contribution rate is still 95%, and the model side inference delay is reduced from 0.102s to 0.078s; when the number of sampling neighbors k = 5 of GraphSAGE in this application is compared with k = 0 (no graph aggregation), the Top-1 accuracy in the few-sample 10-shot scenario is improved by 7.5% (from 91.8% to 98.7%); this application also uses a cross-modal self-attention module with hardware vectors as queries and other vectors as key / values, and the Top-1 accuracy of the deep learning classification model is improved from 93.7% to 99.34%; the number of parameters of the deep separable convolutional network quantized by INT8 is only 1.97×10 6 The single-frame inference latency on the client side is kept at 7.8×10 -2 s, taking into account both high precision and low resources. The number of parameters of the depthwise separable convolutional network after INT8 quantization in this application is 1.97×10 6 The entire model has 3.8MB of resident memory and a peak power consumption of 0.4W, which meets the power consumption budget of less than 5W for inspection drones in the power, petrochemical, and urban infrastructure industries.
[0094] Example 2
[0095] A method for identifying drone inspection equipment based on multimodal device fingerprints, based on Example 1, also includes step S5: when feature drift occurs between the mean of the real-time fusion feature vectors of a batch of drone inspection equipment and the mean of their historical fusion feature vectors, the parameters of the second depth-separable convolutional layer and the last layer of the fully connected neural network of the deep learning classification model are updated through micro-step incremental learning.
[0096] Specifically, the latest batch of fused feature vectors F(t) output by the cross-modal self-attention module is cached in a sliding window W = {F(t-M+1)…F(t)} of size M. The current batch mean μ is calculated for this window t , and calculate the Mahalanobis distance with the training period mean μ0:
[0097]
[0098] Among them, μ0 is the mean of the fusion feature vectors of the same batch of UAV inspection equipment in the historical period, μ t is the mean of the real-time fusion feature vectors of a batch of UAV inspection equipment; D M μ t and the Mahalanobis distance of μ0, ∑ is the feature covariance matrix.
[0099] Mahalanobis distance measures the distribution difference more accurately by considering the correlation between features. M >θ, θ is the preset drift detection threshold, triggering micro-step incremental learning, and only the last two layers of parameters θ of the deep learning classification model tail and a learning rate η to perform a slow gradient update: Among them, the input data is the feature vector set W in the sliding window t Its corresponding label
[0100] Only the last two layers of the deep learning classification model are updated at a low rate, avoiding significant changes to early classification boundaries. During six months of flight aging testing, the model's accuracy rapidly recovered from 84.6% to 98.02% when the capacity decayed by 18 percentage points, achieving a drift suppression gain of 13.4%.
[0101] Example 3
[0102] Provide a set of measured data based on different data volumes for drone inspection equipment identification:
[0103] 1. Comparative Example 1 (Traditional Solution)
[0104] No CRC check, direct DMA writing to FIFO; no exception rejection; single-domain feature input ResNet-18 floating-point model; no open set entropy judgment; no online adaptation.
[0105] 2. Example 1 (Lower limit parameter section)
[0106] Master clock frequency f clk =25MHz, PCA dimension d pca =24, the number of GraphSAGE sampling neighbors k=3, the quantization bit width q=8, and the experimental conditions: 18 drones of the same model were selected, 800 frames were collected for each drone, and recognition was performed using the method of Example 2.
[0107] 3. Example 2 (Median Parameter Segment)
[0108] Master clock frequency f clk =30MHz, PCA dimension d pca =32, number of GraphSAGE sampling neighbors k=5, quantization bit width q=8, test conditions: 60 drones, 1900 frames per drone, ambient temperature -10°C to 55°C. The data of Example 2 are the indicators given in Examples 1 and 2.
[0109] Example 3 (Upper Parameter Section)
[0110] Master clock frequency f clk f clk =40MHz, PCA dimension dpca =48, the number of GraphSAGE sampling neighbors k = 8, the quantization bit width q = 6 (to further reduce the model size), the experimental conditions: 90 drones, 2700 frames per drone, and an additional 300m long-distance remote control link interference.
[0111] The main test instruments and parameters used for the results of Comparative Example 1 and Examples 1-3 are:
[0112] Keysight DSOX3104T high-precision oscilloscope, 1 GHz bandwidth, for verifying ADC sampling jitter. Espec SU-262 constant temperature chamber, 40–85°C temperature range. NI-PXIe-4480 power analyzer, 204.8 kS / s sampling rate, for evaluating real-time power consumption. Saleae Logic-Pro 16-channel logic analyzer, for monitoring CAN-FD bus error frame rate.
[0113] Table 1 Performance comparison data of different examples
[0114] plan Frame completeness rate / % Top-1 / % <![CDATA[Miss detection rate / (10 -4 )]]> Power consumption / W Model size / MB Comparative Example 1 94.3 93.7 32.4 1.7 45 Example 1 100 98.6 3.6 0.42 3.4 Example 2 100 99.34 0.91 0.4 3.8 Example 3 100 99.42 0.88 0.38 4.5
[0115] As shown in Table 1, Examples 1-3 stabilized the integrity rate to 100% through CRC-32 check. Comparative Example 1, due to the lack of check, produced approximately 5.7% damaged frames. Open Set Security: Information entropy threshold judgment reduced the missed detection rate to a minimum of 0.88×10 -4 , which is much better than the 32.4×10 -4 Resource consumption: INT8 quantization combined with depthwise separable convolution reduces the model size by over 90%, with end-side power consumption of only 0.4W.
[0116] Table 1 fully demonstrates the feasibility and superiority of the present invention under different parameter ranges. The comparative data in Comparative Example 1 demonstrates that the present invention's CRC checksum, anomaly rejection, GraphSAGE enhancement, and self-attention fusion offer significant technological advancements in improving integrity, recognition accuracy, and open set security. This invention can operate stably for extended periods while maintaining extremely low hardware resource overhead, making it suitable for drone surveillance scenarios.
[0117] The optimal dimension range for PCA dimensionality reduction in this application is: 24≤d pca ≤48, the inference delay is between 62-96ms, the recognition accuracy of the deep learning classification model is between 98.6%-99.42%, and the size of the deep learning classification model is between 3.4-4.5MB. pca When ≥24, the 128-dimensional original features can be compressed to the above interval, and the cumulative contribution rate remains ≥95%. pca When the value is ≤48, the model size does not exceed 4.5MB, and a single-frame inference latency of ≤96ms can be achieved.
[0118] This application uses GraphSAGE to sample the optimal number of neighbors k, which is 3≤k≤8. At this time, the recognition accuracy of small samples is 96.5%-99%, a relative improvement of 5.1%-7.8%. The implementation shows that when the number of sampled neighbors k≥3, GraphSAGE neighborhood aggregation significantly improves the recognition accuracy of small samples, with a relative improvement of ≥5%. When k>8 continues to increase, the contribution to accuracy tends to saturation and the computational overhead increases.
[0119] Example 4
[0120] like Figure 2 As shown, a drone inspection equipment identification system based on multimodal device fingerprints includes:
[0121] Acquisition module: collects multimodal features of the drone inspection equipment to be identified, including hardware features, software features, network features, and behavioral features;
[0122] Preprocessing and dimensionality reduction module: preprocesses and reduces the dimensionality of the collected multimodal features;
[0123] Fusion module: This module constructs a k-nearest neighbor undirected graph based on the spatial or topological proximity of drone inspection devices in the same scene and batch, and fuses the multimodal features of all neighboring nodes of the current node to update the current node. For each drone inspection device, the multimodal features that incorporate neighbor information are mapped into query, key, and value matrices, and the multimodal features are fused using adaptive weights.
[0124] Identification module: Based on the fusion feature vector output after fusion, the trained deep learning classification model is used to identify the drone's identity and output the classification probability vector. Based on the information entropy of the probability vector, it is determined whether the current drone inspection equipment is a legitimate device;
[0125] The adaptive update module updates the parameters of the second depth-wise separable convolutional layer and the last fully connected neural network of the deep learning classification model through micro-step incremental learning when feature drift occurs between the mean of the real-time fusion feature vector of a batch of drone inspection equipment and the mean of its historical fusion feature vector.
[0126] For more specific details about the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be described again here.
[0127] Example 5
[0128] The present invention provides a computer device comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the steps of the above-mentioned method for identifying drone inspection equipment based on multimodal device fingerprints are implemented.
[0129] For more specific details about the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be described again here.
[0130] Example 6
[0131] The present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the above-mentioned method for identifying drone inspection equipment based on multimodal device fingerprints are implemented.
[0132] For more specific details about the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be described again here.
[0133] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments will be sufficient. The systems, devices, and storage media disclosed in the embodiments are described briefly because they correspond to the methods disclosed in the embodiments. For relevant details, refer to the method description.
[0134] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention or certain portions of the embodiments.
[0135] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for identifying drone inspection equipment based on multimodal device fingerprints, characterized in that: The following steps are involved: S1: Collect multimodal features of the unmanned aerial vehicle inspection equipment to be identified and perform preprocessing and dimensionality reduction. The multimodal features include hardware features, software features, network features, and behavioral features; S2: Construct a k-nearest neighbor undirected graph based on the spatial proximity or topological proximity of drone inspection devices in the same scene and the same batch, and fuse the multimodal features of all neighboring nodes of the current node to update the current node; S3: For each UAV inspection device, the multimodal features integrated with neighbor information are mapped into query, key, and value matrices, and the multimodal features are fused using adaptive weights. S4: Based on the fusion feature vector output after fusion, the trained deep learning classification model is used to identify the identity of the drone and output the classification probability vector. Based on the information entropy of the probability vector, it is determined whether the current drone inspection equipment is a legal device.
2. The method for identifying drone inspection equipment based on multimodal device fingerprint according to claim 1 is characterized in that: The hardware features include CPU serial number, RF transmit crystal oscillator relative deviation, IQ distortion vector and power supply ripple peak; the software features include firmware version number, operating system kernel hash and application signature; the network features include MAC address, IP message identifier, survival time sequence and adjacent data packet sending interval; the behavioral features include three-axis acceleration, three-axis angular velocity, battery terminal voltage and current.
3. The method for identifying drone inspection equipment based on multimodal device fingerprint according to claim 1 is characterized in that: In step S1, the preprocessing and dimensionality reduction include: integrity check, anomaly removal, normalization and principal component analysis dimensionality reduction; After normalizing the behavioral features, information enhancement is performed on the behavioral features and then principal component analysis dimensionality reduction is performed; the information enhancement of the behavioral features is specifically as follows: the behavioral features are converted into a spectrum by short-time Fourier transform, and one-third octave pooling is applied to the spectrum.
4. The method for identifying drone inspection equipment based on multimodal device fingerprint according to claim 1 is characterized in that: In step S2, the GraphSAGE graph convolutional network is used to fuse the multimodal features of neighbor nodes. The specific fusion formula is: in, is the hidden representation of node v in the l+1 layer, σ is the activation function, W l is the weight matrix of the lth layer, Mean is the mean aggregation function, N k (v) is the set of neighbor nodes of node v, is the hidden representation of node v’s neighbor node u in layer l.
5. The method for identifying drone inspection equipment based on multimodal device fingerprint according to claim 1 is characterized in that: Step S3, specifically: for each UAV inspection device, the hardware feature vector integrated with neighbor information is mapped into a query matrix, the software features, network features, and behavior features integrated with neighbor information are concatenated and mapped into a key and value matrix, and the multimodal features are fused using adaptive weights; Z = σ(Attention(Q,K,V)) in, is the query matrix, n is the number of hardware feature vectors, d k is the dimension of the key vector; is the key matrix, m is the number of concatenated eigenvectors; is the value matrix, d v is the dimension of the value vector; the softmax function is normalized by row, is the scaling factor used to scale the dot product result to prevent the gradient from disappearing or exploding; σ is the activation function, and Z is the fused feature vector.
6. The method for identifying drone inspection equipment based on multimodal device fingerprint according to claim 1 is characterized in that: In step S4, the deep learning classification model is based on a four-layer fully connected neural network, and two hidden layers are replaced by two depthwise separable convolutional layers; each depthwise separable convolutional layer includes pointwise convolution, depthwise convolution and activation function, extracts channel features by pointwise convolution in the spatial dimension, uses depthwise convolution to learn spatial correlation, and finally activates by activation function; In step S4, the depthwise separable convolutional layer is compressed using the INT8 quantization method; In step S4, when the information entropy of the probability vector exceeds the threshold, the current drone inspection device is an unknown device; otherwise, the current drone inspection device is a legitimate device; the information entropy H(p) of the probability vector is: Among them, p i is the probability vector of category i, and C is the total number of categories.
7. The method for identifying drone inspection equipment based on multimodal device fingerprint according to claim 6 is characterized in that: The method further includes step S5: when feature drift occurs between the mean of the real-time fused feature vectors of a batch of drone inspection devices and the mean of their historical fused feature vectors, updating the parameters of the second depthwise separable convolutional layer and the last fully connected neural network layer of the deep learning classification model through micro-step incremental learning; Specifically, for the same batch of UAV inspection equipment, when the Mahalanobis distance between the mean of their real-time fusion feature vector and the mean of their historical fusion feature vector exceeds the set threshold, feature drift occurs. The formula for obtaining the Mahalanobis distance is: Among them, μ0 is the mean of the fusion feature vectors of the same batch of UAV inspection equipment in the historical period, μ t is the mean of the real-time fusion feature vectors of a batch of UAV inspection equipment; D M μ t and the Mahalanobis distance of μ0, ∑ is the feature covariance matrix.
8. A drone inspection equipment identification system based on multimodal device fingerprints, characterized in that: include: Acquisition module: collects multimodal features of the drone inspection equipment to be identified, including hardware features, software features, network features, and behavioral features; Preprocessing and dimensionality reduction module: preprocesses and reduces the dimensionality of the collected multimodal features; Fusion module: This module constructs a k-nearest neighbor undirected graph based on the spatial or topological proximity of drone inspection devices in the same scene and batch, and fuses the multimodal features of all neighboring nodes of the current node to update the current node. For each drone inspection device, the multimodal features that incorporate neighbor information are mapped into query, key, and value matrices, and the multimodal features are fused using adaptive weights. Identification module: Based on the fusion feature vector output after fusion, the trained deep learning classification model is used to identify the drone's identity and output the classification probability vector. Based on the information entropy of the probability vector, it is determined whether the current drone inspection equipment is a legitimate device; The adaptive update module updates the parameters of the second depth-wise separable convolutional layer and the last fully connected neural network of the deep learning classification model through micro-step incremental learning when feature drift occurs between the mean of the real-time fusion feature vector of a batch of drone inspection equipment and the mean of its historical fusion feature vector.
9. A computer device, characterized in that: The invention comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the method for identifying drone inspection equipment based on multimodal device fingerprints described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that Used to store computer programs; when the computer program is executed by the processor, the steps of the drone inspection equipment identification method based on multimodal device fingerprint according to any one of claims 1 to 7 are implemented.