LoRa terminal identification method fusing double-frame carrier frequency offset characteristics
By integrating dual-frame carrier frequency offset features and ensemble learning methods, the robustness problem of LoRa devices in complex interference environments is solved, achieving highly reliable device identification and unknown device rejection, which is suitable for resource-constrained LoRa network security access scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-05
AI Technical Summary
Existing LoRa device identification methods lack robustness in complex interference environments, and single-frame features are easily affected by noise and channel disturbances, making it difficult to meet the high-reliability access requirements of resource-constrained devices.
A method that integrates dual-frame carrier frequency offset features is adopted. By constructing dual-frame signals continuously transmitted by the same terminal, and combining signal preprocessing, integrated learning classification and dual threshold rejection mechanism, carrier frequency offset features are extracted and device identification is performed.
It improves the stability and robustness of LoRa terminal identification, making it suitable for network security access under resource-constrained conditions, reducing the risk of misidentification of unknown devices, and enhancing identification accuracy and system security.
Smart Images

Figure CN121985334A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of Internet of Things (IoT) security technology, specifically relating to a LoRa device identification method based on radio frequency fingerprinting (RFF), which is particularly suitable for low power wide area network (LPWAN) scenarios with limited resources and time-varying channels. Background Technology
[0002] With the rapid expansion of the Internet of Things (IoT) field, LoRa technology has established its key position in low-power wide-area networks due to its significant characteristics such as long transmission distance, low power consumption, and open source code, and has been widely adopted in many fields such as environmental monitoring, smart metering, and asset management and tracking.
[0003] However, the large-scale application of LoRa networks has also triggered significant security crises, such as covert channel intrusion and beacon spoofing attacks, posing a major threat to the operational security and stability of the Internet of Things (IoT) system. Software identifiers like MAC addresses are easily tampered with, cloned, or forged, failing to effectively support access standards for highly trusted devices. Although encryption and authentication mechanisms can enhance security, they often come with additional computing and storage resource consumption, which is highly disadvantageous for LoRa terminal devices that prioritize low cost, low energy consumption, and limited resources.
[0004] Radio frequency fingerprinting technology utilizes inherent physical layer differences in wireless device hardware for authentication, offering advantages such as difficulty in forgery, no need to modify upper-layer protocols, and low implementation overhead. Therefore, it has become an important research direction for security authentication of LoRa devices. Existing radio frequency fingerprinting methods are mainly divided into two categories:
[0005] Deep learning-based methods: Automatically extract fingerprint features using convolutional neural networks or recurrent neural networks. These methods have the following shortcomings: (1) They are sensitive to channel time-varying characteristics, and the recognition accuracy drops significantly when collecting data across time periods; (2) The model has poor interpretability and is difficult to meet the requirements of patent application for clarity of technical principles; (3) The training computation overhead is large and it is not suitable for IoT scenarios with frequent updates.
[0006] Feature-engineered methods: Device identification is performed using physical layer features such as carrier frequency offset and phase noise. These methods have the advantages of simple implementation and good interpretability. However, existing feature-engineered methods still have the following shortcomings: (1) Single-frame features are easily affected by noise and instantaneous channel disturbances, resulting in insufficient stability; (2) Carrier frequency offset is prone to drift with changes in temperature and local oscillator state, and its distinguishing ability is limited when used alone; (3) Existing discrimination methods mostly use simple classifiers, which lack effective suppression of misidentification of unknown devices and complex scenes.
[0007] In feature engineering-based research, based on the temporal stability of signal features, existing studies typically categorize them into two types: steady-state features, which are hardware fingerprints that remain relatively stable over a longer timescale; and semi-steady-state features, which are hardware features that are relatively stable in a short time but may change over a longer period. Based on this classification, differential adjacent two-frame RF fingerprinting has been proposed as an existing technique. This method utilizes the differential phase noise between adjacent two frames as the primary fingerprint feature.
[0008] It is particularly important to point out that although existing dual-frame methods, such as differential adjacent dual-frame RF fingerprinting, utilize the hardware correlation between adjacent frames, their technical solutions mainly rely on a single differential logic to extract fingerprint features. In complex interference environments, they cannot adaptively adjust the differential strategy according to real-time channel quality and signal status, which limits their robustness in complex IoT scenarios.
[0009] To address the aforementioned issues, this invention proposes a LoRa terminal identification method that integrates dual-frame carrier frequency offset features and ensemble learning. This method overcomes the shortcomings of single differential logic in terms of robustness under complex interference. Furthermore, by using ensemble learning, the decision boundaries of different types of classifiers are complementary, and the combined result covers a more complex fingerprint space, further improving the accuracy and stability of identification. Summary of the Invention
[0010] This invention relates to the field of wireless communication device identification technology, specifically proposing a LoRa terminal identification method that fuses dual-frame carrier frequency offset features. This method constructs dual-frame LoRa signals continuously transmitted by the same terminal, extracts fused carrier frequency offset features, and combines signal preprocessing, ensemble learning classification, and a dual-threshold rejection mechanism to achieve highly reliable identification of registered LoRa terminals and effective rejection of unknown terminals. Compared to existing technologies, this invention can effectively suppress the impact of channel environment changes and receive gain fluctuations, improving the stability and robustness of identification, and is suitable for secure LoRa network access scenarios under resource-constrained conditions.
[0011] This invention proposes a LoRa terminal identification method that integrates dual-frame carrier frequency offset features. The method includes:
[0012] Step (1) Signal acquisition: The wireless signal transmitted by the LoRa terminal is acquired using software radio equipment to obtain a complex baseband signal sequence, and the LoRa frame start position is determined and the signal frame is parsed by correlation detection with the local reference Chirp signal.
[0013] Step (2) Continuous double frame construction, signal preprocessing and feature extraction: Two consecutive frames of signals sent by the same device are used to form a double frame sample, and they are synchronized, carrier frequency offset estimated and compensated and amplitude normalized. The carrier frequency offset of the double frame and its mean and difference are extracted to form a feature vector.
[0014] Step (3) Integrated learning and recognition: The feature vector is input into multiple basic classifiers for joint decision-making, and weighted fusion is performed based on classifier accuracy and prediction entropy. Combined with confidence threshold, device identity recognition and unknown device rejection are achieved.
[0015] Furthermore, step (1) specifically includes:
[0016] (11) Use the software-defined radio device USRP to collect the wireless signals transmitted by the LoRa terminal, and set the sampling rate. The following is a sequence of complex baseband signals:
[0017]
[0018] in For in-phase components, These are orthogonal components.
[0019] (12) LoRa frame detection is performed on the acquired continuous signal. The frame start position is determined based on the correlation results between the received signal and the local reference Chirp signal, thereby parsing out the complete LoRa signal frame.
[0020] Furthermore, step (2) specifically includes:
[0021] (21) After obtaining a single LoRa signal frame, two consecutive frames of signals sent by the same device in chronological order are combined to form a continuous double-frame sample. By constructing a continuous double-frame sample, the two frames retain the same hardware characteristics of the device under approximately consistent short-time channel conditions, providing a basis for subsequent RF fingerprint feature extraction.
[0022] (22) Synchronization, carrier frequency offset estimation and compensation, and amplitude normalization are performed on the consecutive two-frame signals respectively to obtain a standardized signal sequence after time alignment, frequency offset compensation and amplitude normalization, so as to reduce the impact of receiver gain fluctuation and environmental changes on subsequent feature extraction.
[0023] (23) After completing the signal preprocessing, the carrier frequency offset is estimated for the first and second consecutively received signals to obtain the corresponding frequency offset estimates. and Based on the frequency offset estimate, a continuous two-frame frequency offset feature vector is constructed:
[0024]
[0025] in:
[0026] This represents the estimated carrier frequency offset of the first frame of the signal;
[0027] This represents the estimated carrier frequency offset of the second frame signal;
[0028] It represents the mean characteristic of frequency offset between consecutive two frames, and is used to suppress random noise interference;
[0029] It represents the frequency deviation characteristics between consecutive two frames, used to characterize the frequency drift characteristics of the device during short-term continuous transmission;
[0030] The continuous dual-frame frequency offset feature vector is used to characterize the local oscillator frequency deviation characteristics of LoRa devices and the frequency stability differences under continuous transmission conditions, and serves as the device's RF fingerprint feature input to the subsequent classification and identification module. The CFO reflects the minute hardware differences in the device's oscillator, possessing device uniqueness and long-term stability. Traditional single-frame CFOs are easily affected by channel and noise interference, resulting in insufficient stability. Continuous dual-frames enhance feature stability, better preserve the inherent frequency offset characteristics of the device, and improve cross-environment consistency and identification accuracy.
[0031] Furthermore, step (3) specifically includes:
[0032] (31) Using the fused fingerprint feature vector F output in step (23) as input, select four heterogeneous basic classifiers: support vector machine (SVM), random forest (RF), linear discriminant analysis (LDA) and K nearest neighbors (KNN), and train them independently on the training set so that each classifier can output the posterior probability vector of the device category.
[0033] (32) Joint estimation of static weights and dynamic confidence: The static basis weights are determined by the recognition accuracy of each classifier on the validation set, and the dynamic confidence factors are calculated by the current prediction entropy values of each classifier. The normalized product is used as the joint weight, and the weights of different dimensional features in the classification decision are dynamically adjusted according to the signal quality;
[0034] (33) Weighted probability fusion and device identity output: The posterior probabilities output by the four classifiers are weighted and summed according to their joint weights to obtain the ensemble posterior probability:
[0035]
[0036] The ensemble posterior probability integrates the judgments of all base classifiers on each category and applies differentiated weighting based on reliability, effectively reducing the impact of errors from a single classifier on the final result. The device category with the highest ensemble posterior probability is selected as the recognition output.
[0037]
[0038] (34) Unknown device rejection mechanism based on confidence: Calculate the difference between the maximum and second maximum values of the integrated posterior probability, and combine it with the double threshold judgment to reject unknown devices or illegal access; otherwise, output the device identity. Perform multi-level classification logic in the discrimination space to ensure the accuracy and robustness of the identification. Compared with the existing technology, this mechanism can prevent the forced classification of new devices that do not exist in the model database, which may cause network security impact.
[0039] Furthermore, step (31), training the heterogeneous base classifier, specifically includes:
[0040] (311) Support Vector Machine (SVM) uses radial basis kernel function to establish nonlinear decision boundary and transforms decision function values into posterior probabilities through Platt scaling. It is suitable for high-dimensional small sample classification.
[0041] (312) Random Forest (RF), which is composed of multiple decision trees ensembled through Bagging, outputs the posterior probability based on the proportion of votes for each class. It has good robustness to noise and feature fluctuations.
[0042] (313) Linear Discriminant Analysis (LDA) constructs a linear projection that maximizes the ratio of between-class divergence to within-class divergence, and performs classification based on Gaussian posterior probability. It has low computational cost and strong interpretability.
[0043] (314) The K-Nearest Neighbors (KNN) classifier uses Mahalanobis distance to estimate the posterior probability based on the class distribution of the nearest neighbor samples, effectively preserving local feature distribution information. The four classifiers described above employ different learning paradigms, ensuring the diversity of the ensemble framework and effectively reducing the probability of common errors among the classifiers.
[0044] The joint estimation of static weights and dynamic confidence in step (32) specifically includes:
[0045] (321) Static basis weight calculation: Let the recognition accuracy of the k-th classifier on the validation set be 1 / 2. Its static basis weights are defined as:
[0046]
[0047] Static basis weights reflect the overall discriminative ability of each classifier during the training phase.
[0048] (322) Calculation of prediction entropy: The k-th classifier outputs a posterior probability vector for the current input sample F, and its prediction entropy is calculated:
[0049]
[0050] Where M is the total number of registered device categories, and the prediction entropy is... The smaller the value, the more concentrated the classifier's confidence in a particular category; the larger the value, the more ambiguous the judgment, and the lower the weight should be assigned.
[0051] (323) Calculation of dynamic confidence factor:
[0052]
[0053] in This is a hyperparameter that controls the sensitivity of the confidence factor to the entropy value.
[0054] (324) Joint weight normalization:
[0055]
[0056] Joint weight This method simultaneously considers the classifier's historical accuracy (static robustness) and the predictive certainty of the current sample (dynamic confidence), ensuring that weight allocation takes into account both global and local information. By comprehensively considering static performance and dynamic output information, effective weight allocation helps suppress the adverse effects of low confidence on experimental classification, making the results more reliable.
[0057] Furthermore, the unknown device rejection mechanism based on confidence level in step (34) specifically includes:
[0058] (341) To effectively distinguish between unknown devices and unauthorized access, a double-threshold rejection test is performed on the integrated posterior probability before outputting the device identity. Define the inter-class confidence difference:
[0059]
[0060] Inter-class confidence difference Measuring the confidence advantage of the best class relative to the second-best class reflects the discriminative power of the identification results.
[0061] (342) Rejection Decision Rule: If any of the following conditions are met, the current sample shall be determined as an unknown device or illegal access: (i) The absolute confidence condition is not met, i.e. Less than the preset absolute threshold (ii) The relative confidence condition is not met, i.e. Less than the preset relative threshold If both conditions are met, then output the device identification identifier. The dual-threshold mechanism makes judgments based on both absolute confidence and inter-class discrimination. Compared with the single-threshold scheme of existing technologies, it has a stronger ability to reject unknown devices and a lower false recognition rate. Moreover, the decision mechanism is simple. Compared with the application of related algorithms to large models such as deep learning, this method is more lightweight and easier to deploy in engineering while ensuring accuracy.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] (1) By constructing continuous dual-frame signals and extracting fused dual-frame carrier frequency offset (CFO) features, this invention can characterize the frequency drift characteristics in a short time while characterizing the local oscillator frequency deviation of LoRa terminals. Compared with traditional single-frame features, it has better stability and distinguishability.
[0064] (2) In the identification stage, this invention introduces an integrated learning framework consisting of support vector machine, random forest, linear discriminant analysis and K nearest neighbor classifier, and combines static weight and dynamic confidence weighting mechanism to realize multi-classifier collaborative decision-making. Compared with the single classifier method, it can further improve the accuracy and stability of LoRa terminal identity recognition.
[0065] (3) The present invention sets up a dual-threshold open set rejection mechanism, which can jointly judge the identification results from two dimensions: absolute confidence and inter-class discrimination, thereby achieving effective rejection of unknown devices and illegal access behaviors, reducing the risk of unknown samples being misjudged as registered devices, and improving system security.
[0066] (4) Experimental results show that the method of the present invention has good recognition performance and robustness under different acquisition times and different environmental conditions; verifying the effectiveness and stability of the proposed fusion of dual-frame carrier frequency offset features and integrated learning method.
[0067] (5) Based on interpretable physical layer features and their statistical modeling, this invention is more suitable for resource-constrained LoRa network security access scenarios compared to the large-scale training samples and computing power costs of deep neural networks. Attached Figure Description
[0068] Figure 1 This is the overall architecture diagram. Detailed Implementation
[0069] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The present invention proposes a LoRa device identification method based on continuous dual-frame radio frequency fingerprinting, which reliably identifies different LoRa terminal devices by constructing continuous dual-frame signals and extracting fused carrier frequency offset features. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0070] Example 1:
[0071] See Figure 1This embodiment provides a LoRa terminal identification method that integrates dual-frame carrier frequency offset features, including signal acquisition and frame parsing, continuous dual-frame construction and feature extraction, and device identification and unknown device rejection based on ensemble learning. In the experiment, a USRP device was used to acquire LoRa signals, with a center frequency of 433 MHz, a sampling rate of 1 MSps, and a bandwidth of 125 kHz. Multiple LoRa terminals were selected to construct a dataset, which was divided into training, validation, and test sets. The classification model parameters and rejection threshold were calibrated using the validation set. Experimental results show that this method outperforms methods based on single-frame features or a single classifier in both device identification accuracy and unknown device rejection rate, with an accuracy exceeding 95%, demonstrating good robustness and stability.
[0072] 1. Signal acquisition,
[0073] This section mainly describes the LoRa signal acquisition process and the method for constructing continuous dual-frame samples. Based on this, signal preprocessing and collaborative radio frequency fingerprint feature extraction are completed to provide basic data and feature representation for subsequent device identification.
[0074] In this embodiment, LoRa wireless signals are acquired through a software-defined radio platform. The receiver uses a USRP N210 software-defined radio device to receive radio frequency signals, and GNU Radio is used to build the signal acquisition and processing flow.
[0075] During signal acquisition, the center frequency of the receiving device was set to 433 MHz, the sampling rate to 1 MSps, and the LoRa communication bandwidth to 125 kHz. A signal acquisition module was constructed using the GNU Radio flowchart to continuously receive LoRa signals from the wireless channel and save the received raw I / Q signal data as a digital file. The complex baseband signal obtained by the receiver can be represented as:
[0076]
[0077] in For in-phase components, These are orthogonal components.
[0078] During the signal parsing phase, the open-source module gr-lora is modified to enable the system to perform LoRa frame parsing simultaneously with signal acquisition, and the parsed signals are saved frame by frame. Specifically, the acquired continuous signals are parsed using the LoRa frame structure, and the frame start position is detected based on the repeating chirp structure of the LoRa preamble. The correlation value between the received signal and the local reference chirp signal is calculated using a correlation detection method.
[0079]
[0080] in For the local reference chirp signal, when The starting position of the LoRa frame is determined when the maximum value is obtained, thus parsing out the complete LoRa signal frame. By integrating the signal acquisition module and the frame parsing module, integrated processing of signal acquisition and frame parsing can be achieved.
[0081] At the transmitting end, various models of LoRa terminal devices were used for signal transmission, and device control programs were written using the Arduino IDE to enable continuous data frame transmission. Communication parameters, including spreading factor, bandwidth, and carrier frequency, were uniformly configured across different devices to ensure experiments were conducted under identical communication conditions.
[0082] The above methods can collect a large amount of LoRa signal data from different devices, providing a data foundation for subsequent radio frequency fingerprint feature extraction.
[0083] 2. Construction of original features from consecutive two frames.
[0084] 2.1 Construction of consecutive two frames,
[0085] After completing signal acquisition and LoRa single-frame parsing, a single-frame LoRa signal arranged in chronological order can be obtained. Considering that the same LoRa terminal usually sends multiple data frames continuously during actual communication, and the time interval between two adjacent frames is short, it usually experiences approximately consistent channel conditions within a short period of time and retains strong homogeneous hardware characteristics.
[0086] Based on the above characteristics, this embodiment combines two adjacent frames of signals from the same device in chronological order to form a continuous two-frame sample, which is represented as follows:
[0087]
[0088] in This represents the first frame of the signal. This indicates the second frame of the signal.
[0089] During data organization, the system classifies and stores dual-frame samples according to device number and constructs an RF fingerprint database. The continuous dual-frame structure not only preserves the device RF fingerprint information within a single frame but also characterizes the correlation changes between adjacent frames at the frequency and phase levels. Since the dual frames have similar propagation environments within a short time, and the device's local oscillator frequency error and hardware non-ideal characteristics remain correlated between the two frames, continuous dual-frame samples can more stably characterize the device's hardware features, providing a foundation for subsequent collaborative RF fingerprint feature extraction.
[0090] 2.2 Signal preprocessing,
[0091] Since wireless signals may be affected by factors such as channel noise, frequency offset, and equipment hardware errors during propagation and reception, preprocessing of the signal is necessary before feature extraction.
[0092] In this embodiment, signal preprocessing mainly includes the following steps:
[0093] The first step is frame synchronization. Since the signal acquisition produces a continuous I / Q data stream, a synchronization algorithm is needed to detect the start position of each LoRa data frame. This embodiment utilizes the characteristic of the LoRa preamble having a repeating chirp structure and uses a related detection method to determine the preamble position to complete frame synchronization. The synchronized signal is as follows:
[0094]
[0095] in This is the detected start position of the frame. The raw I / Q complex data stream acquired by the receiver. The position of the LoRa frame preamble start is determined by a correlation detection algorithm (in this embodiment, Schmidl-Cox autocorrelation detection based on the preamble repeat chirp structure), where n is the local sampling index of the synchronized signal. For This refers to the nth sampling point of the synchronized single-frame signal, intercepted from the starting point. Through this offset operation, the original data stream... Aligned to the start of the LoRa frame, providing a unified reference time base for subsequent preamble extraction and feature calculation.
[0096] After synchronization is complete, the preamble portion of the LoRa frame is extracted. Since the preamble has a fixed structure in each frame and its modulation characteristics are significantly affected by the device hardware, it is often used for RF fingerprint feature extraction.
[0097] Subsequently, carrier frequency offset is estimated for the synchronized signal. Due to frequency deviation between the local oscillators at the transmitting and receiving ends, the received signal typically exhibits a certain carrier frequency offset. To reduce the impact of this factor on subsequent feature extraction, this embodiment estimates the carrier frequency offset by analyzing the phase change between adjacent sampling points, expressed as:
[0098]
[0099] Where T is the sampling period. The total number of sampling points involved in frequency offset estimation is typically the length of a single frame of signal after synchronization. ; For the synchronized signal The complex conjugate at the (n-1)th sampling point.
[0100] Based on the estimated carrier frequency offset, the signal is compensated, and the compensated signal can be expressed as:
[0101]
[0102] After frequency offset compensation, to reduce the impact of signal amplitude differences under different acquisition conditions on subsequent feature extraction, amplitude normalization of the signal is also required. The normalized signal is represented as follows:
[0103]
[0104] After the above processing, a standardized signal sequence with time alignment, frequency offset compensation, and amplitude normalization is obtained.
[0105] It should be noted that the signal acquisition, frame synchronization processing, carrier frequency offset estimation, and carrier frequency offset compensation processes can all be implemented using conventional techniques well known to those skilled in the art. The implementation methods described in this embodiment are merely illustrative and do not constitute a limitation on the scope of protection of this invention.
[0106] 2.3 Feature Extraction of Continuous Two-Frame Signals
[0107] After completing signal preprocessing, this embodiment extracts frequency offset radio frequency fingerprint features based on continuous two-frame signals to construct the device's continuous two-frame original radio frequency fingerprint feature set.
[0108] Unlike traditional single-frame frequency offset characteristics, continuous double-frame frequency offset characteristics retain the carrier frequency offset characteristics of a single frame while also depicting the frequency drift relationship between two adjacent frames during short-term continuous transmission, thus reflecting the local oscillator frequency error characteristics of the equipment more stably.
[0109] For the constructed consecutive two-frame samples, carrier frequency offset estimation is performed on the first frame signal and the second frame signal respectively. Let the corresponding frequency offset estimates be respectively. and .
[0110] To simultaneously characterize the absolute degree of carrier frequency offset and the frequency change relationship between consecutive two frames, this embodiment constructs the following consecutive two-frame frequency offset feature vector:
[0111]
[0112] in:
[0113] This represents the estimated carrier frequency offset of the first frame of the signal;
[0114] This represents the estimated carrier frequency offset of the second frame signal;
[0115] It represents the mean characteristic of frequency offset between consecutive two frames, and is used to suppress random noise interference;
[0116] It represents the frequency deviation characteristics between consecutive two frames, used to characterize the frequency drift characteristics of the device during short-term continuous transmission;
[0117] The above feature extraction methods can be used to construct an RFID fingerprint database, providing feature input for subsequent device identification.
[0118] 3 Device Identification Based on Ensemble Learning
[0119] After completing feature extraction, this embodiment will fuse the fingerprint feature vector F (i.e., The input is fed into an ensemble learning framework for classification and recognition. This framework is based on the outputs of four heterogeneous base classifiers and achieves reliable open-set recognition through a confidence-weighted voting mechanism and a double-threshold rejection strategy. The specific process is as follows.
[0120] 3.1 Training the basic classifier
[0121] This embodiment selects four representative heterogeneous base classifiers, using the fused fingerprint feature vector F output in step 2.3 as a unified input, and trains them independently on the training set, so that each classifier can generate a posterior probability vector. The classification results are output in the form of M, where M is the total number of registered device categories and k∈{1,2,3,4} is the classifier index.
[0122] (1) Support Vector Machine (SVM): Employs radial basis kernel function
[0123] A nonlinear decision boundary is established, and the decision function values are converted into class posterior probabilities using Platt scaling. SVM has good generalization ability for small-sample classification problems in high-dimensional feature spaces, making it suitable for scenarios with high feature dimensions, such as RF fingerprints.
[0124] (2) Random Forest (RF): It is composed of multiple decision trees ensembled by Bagging, and the posterior probability is estimated by the ratio of the number of samples of each class in each leaf node. Random forest is highly resistant to noise and feature fluctuations and can capture nonlinear interaction relationships between features.
[0125] (3) Linear Discriminant Analysis (LDA): A linear projection is constructed by maximizing the ratio of between-class divergence to within-class divergence, and the posterior probability of each class is calculated based on the Gaussian distribution assumption. LDA has low computational cost and strong interpretability, and performs well when the features have good linear separability.
[0126] (4) K-Nearest Neighbors (KNN): The KNN uses Mahalanobis distance to measure the distance between samples and estimates the posterior probability based on the class distribution of the nearest neighbor samples. KNN preserves the local feature distribution information completely and can effectively capture the clustering structure of samples from the same device in the feature space.
[0127] The four classifiers mentioned above have different learning paradigms (boundary optimization, ensemble tree, linear discriminant, and instance-based), which ensures the diversity of the ensemble framework and helps reduce the probability of common errors among the classifiers.
[0128] 3.2 Joint Estimation of Static Weights and Dynamic Confidence
[0129] (1) Static basis weight calculation: Let the recognition accuracy of the k-th base classifier on the hold-out validation set be 1 / 2. Its static basis weights are defined as:
[0130]
[0131] Static basis weights reflect the comprehensive discriminative ability of each classifier during the training phase; the higher the accuracy of the classifier, the greater its contribution to the fusion.
[0132] (2) Calculation of prediction entropy: For the current input sample F, the k-th classifier outputs the posterior probability vector, and its prediction entropy is calculated:
[0133]
[0134] Predicting entropy The uncertainty of the classifier's prediction for the current sample is measured. The smaller the value, the more concentrated the classifier's confidence in a certain category; the larger the value, the more ambiguous the classifier's judgment of the current sample, and it should be assigned a lower weight.
[0135] (3) Calculation of dynamic confidence factor:
[0136]
[0137] in This is a hyperparameter that controls the sensitivity of the confidence factor to the entropy value.
[0138] (4) Joint weight normalization: Combine the static basis weights with the dynamic confidence factor to calculate the final joint weights:
[0139]
[0140] Joint weight It takes into account both the classifier's historical accuracy (static robustness) and the predictive certainty of the current sample (dynamic confidence), so that the weight allocation takes into account both global and local information.
[0141] 3.3 Weighted Probability Fusion and Device Identity Output
[0142] The ensemble posterior probability is obtained by weighting and summing the posterior probability vectors output by the four base classifiers according to their joint weights.
[0143]
[0144] The ensemble posterior probability integrates the judgments of all base classifiers on each category and differentiates the classifiers according to their reliability, effectively reducing the impact of errors by a single classifier on the final result. The device category with the highest ensemble posterior probability is taken as the recognition output.
[0145]
[0146] 3.4 Confidence-Based Unknown Device Rejection Mechanism
[0147] (1) To achieve discriminative identification of unknown devices and unauthorized access, a double-threshold rejection test is first performed on the integrated posterior probability before outputting the device identity. The inter-class confidence difference is defined as follows:
[0148]
[0149] Inter-class confidence difference The confidence advantage of the best class relative to the second-best class is measured, reflecting the discriminative power of the identification results.
[0150] (2) Rejection Decision Rule: If any of the following conditions are met, the current sample is determined to be an unknown device or an unauthorized access, and the device identity is rejected: (i) The absolute confidence condition is not met: Less than the preset absolute threshold (ii) The relative confidence condition is not met: Less than the preset relative threshold .in , These are threshold parameters pre-calibrated on the verification set. If both of the above conditions are met, the device identity is output. The dual-threshold mechanism makes a joint judgment based on two dimensions: absolute confidence and inter-class discrimination. Compared with the single-threshold scheme, it has a stronger ability to reject unknown devices and a lower false recognition rate.
[0151] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.
Claims
1. A LoRa terminal identification method integrating dual-frame carrier frequency offset features, characterized in that, The method includes the following steps: Step (1) Signal Acquisition: The wireless signal transmitted by the LoRa terminal is acquired using a software-defined radio device to obtain a complex baseband signal sequence. The LoRa frame start position is determined and the signal frame is parsed by correlation detection with the local reference chirp signal. Step (2) involves constructing consecutive dual frames, preprocessing signals, and extracting features. Two consecutive frames of signals transmitted by the same device are used to form a dual-frame sample. Synchronization, carrier frequency offset estimation and compensation, and amplitude normalization are performed on the sample. The carrier frequency offset of the dual frames, along with their mean and difference, are extracted to form a feature vector. Step (3) Integrated learning and recognition: The feature vector is input into multiple basic classifiers for joint decision-making, and weighted fusion is performed based on classifier accuracy and prediction entropy. Combined with confidence threshold, device identity recognition and unknown device rejection are achieved.
2. The LoRa terminal identification method based on fused dual-frame carrier frequency offset features according to claim 1, characterized in that, Step (1) specifically includes: (11) Use the software-defined radio device USRP to collect the wireless signals transmitted by the LoRa terminal, and set the sampling rate. The following is a sequence of complex baseband signals: in For in-phase components, For orthogonal components, Represents the imaginary unit (12) LoRa frame detection is performed on the acquired continuous signal. The frame start position is determined based on the correlation results between the received signal and the local reference Chirp signal, thereby parsing out the complete LoRa signal frame.
3. The LoRa terminal identification method based on the fusion of dual-frame carrier frequency offset features according to claim 1, characterized in that, Step (2) specifically includes: (21) After obtaining a single LoRa signal frame, two consecutive frames of signals sent by the same device in chronological order are combined to form a continuous double-frame sample. By constructing continuous double-frame samples, the two frames retain the same hardware characteristics of the device under approximately consistent short-time channel conditions, providing a basis for subsequent RF fingerprint feature extraction. (22) Synchronization, carrier frequency offset estimation and compensation, and amplitude normalization are performed on the consecutive two-frame signals to obtain a standardized signal sequence after time alignment, frequency offset compensation, and amplitude normalization, so as to reduce the impact of receiver gain fluctuations and environmental changes on subsequent feature extraction. (23) After completing the signal preprocessing, the carrier frequency offset is estimated for the first and second consecutively received signals to obtain the corresponding frequency offset estimates. and Construct a continuous two-frame frequency offset feature vector based on the frequency offset estimate: in: This represents the estimated carrier frequency offset of the first frame of the signal; This represents the estimated carrier frequency offset of the second frame signal; It represents the mean characteristic of frequency offset between consecutive two frames, and is used to suppress random noise interference; It represents the frequency deviation characteristics between consecutive two frames, used to characterize the frequency drift characteristics of the device during short-term continuous transmission; The continuous two-frame frequency offset feature vector is used to characterize the local oscillator frequency deviation characteristics of LoRa devices and the frequency stability differences under continuous transmission conditions, and serves as the input of the device's radio frequency fingerprint features into the subsequent classification and identification module.
4. The LoRa terminal identification method based on the fusion of dual-frame carrier frequency offset features according to claim 1, characterized in that, Step (3) specifically includes: (31) Basic classifier training: Using the fused fingerprint feature vector F output in step (23) as input, select four heterogeneous basic classifiers: Support Vector Machine (SVM), Random Forest (RF), Linear Discriminant Analysis (LDA), and K Nearest Neighbors (KNN), and train them independently on the training set so that each classifier can output the posterior probability vector of the device category. (32) Joint estimation of static weights and dynamic confidence: The static base weights are determined by the recognition accuracy of each base classifier on the validation set; the dynamic confidence factor is calculated by the entropy value of the current prediction output of each classifier, and the lower the entropy value, the more certain the prediction; the normalized product of the static base weights and the dynamic confidence factor is used as the final joint weight. (33) Weighted probability fusion and device identity output: The posterior probabilities output by the four classifiers are weighted and summed according to their joint weights to obtain the ensemble posterior probability: Where F is the feature vector of the current sample to be identified; c is the candidate device category identifier, with a value range of {1, 2, ..., C}, and C is the total number of devices in the training set; K is the number of basic classifiers (K=4 in this embodiment). These are SVM, Random Forest, LDA, and KNN, respectively. Let the joint weights of the k-th classifier satisfy the following condition: =1; This is the posterior rate estimate for the k-th classifier that the sample F belongs to class c; This represents the posterior probability after integration. The ensemble posterior probability integrates the judgments of all base classifiers on each category and applies differentiated weighting based on reliability, effectively reducing the impact of errors by a single classifier on the final result. The device category with the highest ensemble posterior probability is selected as the recognition output. (34) Unknown device rejection mechanism based on confidence: Calculate the difference between the maximum probability value and the second largest probability value of the integrated posterior probability vector to form the inter-class confidence difference; if the maximum integrated posterior probability is lower than the preset absolute threshold, or the inter-class confidence difference is lower than the preset relative threshold, it is determined to be an unknown device or illegal access; otherwise, output the device identity identifier.
5. The LoRa terminal identification method based on fused dual-frame carrier frequency offset features according to claim 4, characterized in that, The training of the basic classifier in step (31) specifically includes: (311) Support Vector Machine (SVM) uses radial basis kernel function to establish nonlinear decision boundary and transforms decision function values into posterior probabilities through Platt scaling. It is suitable for high-dimensional small sample classification. (312) Random Forest (RF), which is composed of multiple decision trees ensembled through Bagging, outputs the posterior probability based on the proportion of votes for each class. It has good robustness to noise and feature fluctuations. (313) Linear Discriminant Analysis (LDA) constructs a linear projection that maximizes the ratio of between-class divergence to within-class divergence, and performs classification based on Gaussian posterior probability. It has low computational cost and strong interpretability. (314) The K-Nearest Neighbors (KNN) classifier uses Mahalanobis distance to estimate the posterior probability based on the class distribution of the nearest neighbor samples, effectively preserving local feature distribution information. The above four classifiers have different learning paradigms, ensuring the diversity of the ensemble framework and effectively reducing the probability of common errors among the classifiers.
6. The LoRa terminal identification method based on fused dual-frame carrier frequency offset features according to claim 4, characterized in that, The joint estimation of static weights and dynamic confidence in step (32) specifically includes: (321) Static basis weight calculation: Assume that the device recognition accuracy of the k-th classifier on the validation set is 100%. Its static basis weights are defined as: Static basis weights reflect the overall discriminative ability of each classifier during the training phase. This represents the individual recognition accuracy of the j-th classifier. The static basis weights of the k-th classifier are given by... The ratio of the accuracy of all classifiers to the sum of their accuracies is determined, satisfying the normalization condition. The larger the value, the stronger the overall discriminative ability of the classifier on the validation set. (322) Calculation of prediction entropy: The k-th classifier outputs a posterior probability vector for the current input sample F, and its prediction entropy is calculated: Where M is the total number of registered device categories, and the prediction entropy is... The smaller the value, the more concentrated the classifier's confidence in a particular category; the larger the value, the more ambiguous the judgment, and the lower the weight should be assigned. (323) Calculation of dynamic confidence factor: in This is a hyperparameter that controls the sensitivity of the confidence factor to the entropy value. (324) Joint weight normalization: Joint weight It simultaneously considers the classifier's historical accuracy (static robustness) and the predictive certainty of the current sample (dynamic confidence), ensuring that weight allocation takes into account both global and local information. The static basis weights for the k-th classifier are determined by the validation set accuracy. Its dynamic confidence factor ( The information entropy of the posterior probability vector. > 0 represents the temperature hyperparameter), K is the total number of basic classifiers, and the denominator is... is the normalization term, and j is the classifier index variable, which iterates through all K base classifiers; The overall contribution of the j-th classifier to the current sample is represented by both its historical discriminative ability on the validation set (static) and its degree of certainty in predicting the current sample (dynamic).
7. The LoRa terminal identification method based on fused dual-frame carrier frequency offset features according to claim 4, characterized in that, The unknown device rejection mechanism based on confidence level in step (34) specifically includes: (341) To achieve effective discrimination of unknown devices and unauthorized access, a double-threshold rejection test is performed on the integrated posterior probability before outputting the device identity, and the inter-class confidence difference is defined as follows: Inter-class confidence difference The confidence advantage of the best class relative to the second-best class is measured, reflecting the discriminative power of the recognition results. Here, F represents the continuous two-frame carrier frequency offset feature vector of the current sample. The ensemble posterior probability is obtained after weighted fusion using formula (7). For the optimal category, To maximize the ensemble posterior probability, The posterior probability of the suboptimal class; The larger the value, the more reliable the identification result; the smaller the value, the more likely the sample belongs to an unknown device or illegal access. (342) Rejection Decision Rule: If any of the following conditions are met, the current sample shall be determined as an unknown device or illegal access: (i) The absolute confidence condition is not met, i.e. Less than the preset absolute threshold (ii) The relative confidence condition is not met, i.e. Less than the preset relative threshold If both conditions are met, then output the device identification identifier. The dual-threshold mechanism makes a joint judgment based on two dimensions: absolute confidence and inter-class discrimination. Compared with the single-threshold scheme, it has a stronger ability to reject unknown devices and a lower false recognition rate.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the LoRa terminal identification method that integrates dual-frame carrier frequency offset features as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instruction is executed by the processor, it implements the LoRa terminal identification method that integrates dual-frame carrier frequency offset features as described in any one of claims 1-7.