A frequency hopping modulation mode identification method
By removing noise through SPWVD, adaptive Wiener filtering, and morphological filtering, and extracting the geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy features of the frequency hopping signal, and combining reinforcement learning feature selection algorithm and SVM classifier, the problems of low accuracy and high computational complexity in frequency hopping modulation mode recognition under low signal-to-noise ratio are solved, achieving high recognition rate and robustness.
Patent Information
- Application Number
- CN202211607172.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-12-14
AI Technical Summary
Existing technologies suffer from low accuracy, poor robustness, and high computational complexity in frequency hopping modulation under low signal-to-noise ratio conditions.
Noise is removed by SPWVD time-frequency transformation, adaptive Wiener filtering, and morphological filtering. Geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy features of the time-frequency energy grayscale image are extracted. Combined with reinforcement learning feature selection algorithm and support vector machine classifier, the optimal feature set is selected and multiple types of FH modulation mode are identified.
Under low signal-to-noise ratio conditions, the recognition accuracy is high and the robustness is strong, which reduces the computational complexity of the algorithm. Simulation experiments show that the recognition rate can reach 95% at a signal-to-noise ratio of 5dB.
Smart Images

Figure CN117411752B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of frequency hopping communication technology, and in particular to a method for identifying frequency hopping modulation patterns. Background Technology
[0002] Frequency hopping (FH) communication has become the dominant technology for anti-jamming and counter-reconnaissance wireless communication in major military powers today due to its advantages such as good security, superior multiple access capability, and low probability of interception. As an important aspect of reconnaissance of enemy FH communication, FH signal modulation mode can provide strong support for tasks such as enemy-friend identification, intelligence interception, and jamming suppression.
[0003] In the existing technology:
[0004] YKYadav, G. Jajoo, and SKYadav published "Modulation scheme detection of blind signal using constellation graphical representation" at the 2017 International Conference on Computer, Communications and Electronics (Comptelix), pp:231-235. This scheme first estimates parameters such as signal carrier frequency, symbol rate, and phase offset, then extracts constellation diagram features, and uses linear regression and circle fitting analysis to identify ASK, PSK, and QAM modulation. However, this method has a large parameter estimation error under low signal-to-noise ratio (SNR) conditions, and the recognition accuracy drops significantly (YKYadav, G. Jajoo and SKYadav. Modulation scheme detection of blind signal using constellation graphical representation [C]. 2017 International Conference on Computer, Communications and Electronics (Comptelix), 2017, pp:231-235).
[0005] Another approach is: Huang Sai, Lin Chunsheng, Zhou Kai, et al. Identifying physical-layer attacks for IoT security: An automatic modulation classification approach using multi-module fusion neural network[J]. Physical Communication, 2020, 43(4):101180.
[0006] Therefore, it is necessary to design a frequency hopping modulation method for recognition that has high recognition accuracy, strong robustness, and reduced algorithm computational complexity under low signal-to-noise ratio conditions. Summary of the Invention
[0007] Therefore, to address the above problems, this invention proposes a frequency hopping modulation scheme identification method, which achieves the selection of the optimal feature set and the identification of multiple types of FH modulation schemes. Under low signal-to-noise ratio conditions, it has high identification accuracy, strong robustness, and reduces the computational complexity of the algorithm.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A method for identifying frequency hopping modulation schemes includes the following steps:
[0010] Step 1: Receive the FH signal and perform data preprocessing on the FH signal, including:
[0011] The SPWVD time-frequency transformation algorithm is used to extract the time-frequency energy map of the FH signal; adaptive Wiener filtering and morphological filtering algorithms are used to remove noise from the time-frequency energy map; and then the time-frequency energy map is converted into a time-frequency energy grayscale map.
[0012] Step 2, Feature Extraction, including:
[0013] The modulation recognition feature parameters of three types of features, namely geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy, are extracted from the time-frequency energy grayscale image and used as the feature set for FH modulation recognition.
[0014] Step 3, Feature Selection and Recognition, including:
[0015] The optimal feature set and the identification of multiple FH modulation modes are achieved by using a reinforcement learning feature selection algorithm and a support vector machine classifier.
[0016] Furthermore, in step one, the SPWVD time-frequency transformation algorithm is used to perform smoothing filtering on the FH signal in both the time and frequency domains using window functions, forming a time-frequency matrix F;
[0017] F(x,y) represents the energy value in the x-th row and y-th column of the time-frequency matrix F, where the time-frequency matrix F is of size M×N and its domain is D. F Then (x,y)∈D F 0≤x<M, 0≤y<N, take the neighborhood of the time frequency point F(x,y). If its size is m×n, then the neighborhood Expected value μ and variance σ 2 for:
[0018]
[0019]
[0020] Then, after Wiener filtering, the time-frequency point F(x,y) can be obtained as F′(x,y):
[0021]
[0022] υ 2 Take the mean of the variances of all neighborhoods of the time-frequency matrix F;
[0023] Morphological filtering is used to perform opening operations on the time-frequency matrix F, including erosion followed by dilation of F′(x,y):
[0024]
[0025] Where H(x,y) represents the energy value of the time-frequency matrix after morphological filtering, b(τ,ν) represents the structuring element, and (τ,ν)∈D b D b Describe the domain of the struct element b(τ,ν), and the operators Θ and ν. These represent the erosion and expansion operations, respectively:
[0026]
[0027]
[0028] After filtering, a connected component detection algorithm is used to extract the single-hop time-frequency matrix of the FH signal. To improve the stability of the extracted features, the single-hop time-frequency matrix is grayscaled to form a time-frequency energy grayscale image.
[0029]
[0030] Where G(x,y) represents the gray value in the x-th row and y-th column of the time-frequency gray matrix G. This indicates the floor function, and I represents the gray level.
[0031] Furthermore, the characteristic formula for the geometrically invariant moment is defined as follows:
[0032]
[0033]
[0034] Where, η c,d χ represents the (c+d)th order normalized central distance. c,d Let the center distance be (c+d), where c+d = 2, 3, ..., c = 0, 1, 2, ..., d = 0, 1, 2, ..., and the center of reference (x0, y0) be:
[0035]
[0036]
[0037] Seven geometrically invariant moments were calculated using second- and third-order normalized center moments. The modulation recognition feature parameter is:
[0038]
[0039] Furthermore, the pseudo-Zernike moment feature is defined as follows:
[0040] Map the time-frequency energy grayscale image G(x,y) onto the unit circle of the new coordinate system I(δcosθ,δsinθ), with the center and origin both being the center of the square (x0,y0), and the radius of the unit circle set to... The polar coordinates (δ, θ) are then expressed as:
[0041]
[0042] Where δ represents the vector length from pixel G(x,y) in the time-frequency energy grayscale image to the center of the moment (x0,y0), and θ represents the angle of counterclockwise rotation to that vector, which needs to be converted according to the specific quadrant of the polar coordinates of the pixel.
[0043] Based on the new coordinate system I(δcosθ,δsinθ), the pseudo Zernike moments z of the time-frequency energy grayscale image g,h It can be represented as:
[0044]
[0045] W g,h (δcosθ,δsinθ)=X g,h (δ)exp(jhθ);
[0046] Where g can be zero or a positive integer, h is a positive integer and h≤g, W g,h (δcosθ, δsinθ) is z g,h Basis functions, * denotes complex conjugate, X g,h (δ) is a basis function W g,h The radial polynomial of (δcosθ,δsinθ);
[0047]
[0048] Select z 2,0 z 2,1 z 3,0 z 3,1 z 3,2 z 3,3 z 4,0 and z 4,1 Eight pseudo-Zernike moment parameters are used as modulation recognition feature parameters.
[0049] Furthermore, the discrete definition of the Ruili entropy feature is as follows:
[0050]
[0051] in, This represents the order of the entropy ξ in Ruili.
[0052] choose Five integer-order Ruili entropy parameters of 3, 5, 7, 9, and 11 are used as modulation recognition feature parameters.
[0053] Furthermore, the reinforcement learning feature selection algorithm in step three includes the following sub-steps:
[0054] 3.1 Markov Decision Process:
[0055] A Markov decision process model is used, which is described by a quintuple {S,A,P,R,γ};
[0056] S represents the state space of the feature set; S u S represents the state of a subset of the feature set of the agent at time u. u ∈S, u≥1; A represents the action state space for retaining or deleting features; a u Let a represent the action taken by the agent at time u. u ∈A;S u+1 This indicates that the agent is in the current feature set subset state S. uNext, execute action a u The state after; P represents the state transition probability matrix of the feature set; p u This indicates that the agent is in state S with the current feature set. u Next, execute action a u Later transferred to S u+1 The probability of a state, p u ∈P; R represents the reward value, i.e., the modulation recognition accuracy of the feature set, R u (S u ,a u ) represents the state S of the agent in the current feature set subset. u Next, execute action a u The recognition accuracy after recognition; γ represents the discount coefficient, 0≤γ≤1, which is used to measure the importance of future state rewards to the current agent's action selection;
[0057] The optimal subset of features is sought by maximizing the expected cumulative reward V through the action selection strategy ψ.
[0058]
[0059] The Q-value function is used to represent the cumulative expected reward of an agent performing action a in feature set state S;
[0060]
[0061] According to Bellman's criterion, we can obtain:
[0062] Q ψ (S,a)=E[R u (S u ,a u )+γQ(S u ,a u )|S u =S,a u =a];
[0063] 3.2 Q-learning algorithm:
[0064] Employing a time-difference-based Q-learning algorithm, the agent utilizes the training dataset and, by adopting appropriate action selection strategies, executes actions to delete or retain features, updating the Q(S,a) value accordingly. Through continuous exploration of the feature set, the agent's understanding of the feature set is strengthened. Finally, based on the Q(S,a) value record table, the action selection that maximizes the Q(S,a) value for each feature set subset is identified as the optimal action selection strategy, thus yielding the optimal feature set. The Q(S,a) value update expression is:
[0065]
[0066] in, Indicates the learning rate, This indicates that the agent is in state S of a subset of the feature set. u Q u+1 (S u ,a u Action selection that maximizes the value;
[0067] The Q-learning algorithm employs an ε-greedy action selection strategy, randomly selecting actions to delete or retain features with probability ε, and selecting the action with the largest Q(S,a) value with probability 1-ε. This can be expressed as:
[0068]
[0069] Where rand represents a random number, 0 < rand < 1, and ψ(s) represents the action selection of the agent in the feature set subset state S;
[0070] A dynamic threshold ε-greedy strategy is adopted, where ε can be expressed as:
[0071]
[0072] in, This represents the average recognition accuracy of the first u training iterations.
[0073] In the early stages of training, due to the limited number of training iterations (u), the agent acquires relatively little information. The relatively small value of u leads to a larger threshold ε(u), which encourages the agent to explore more information. As the number of training iterations u gradually increases, the average recognition accuracy increases. The continuous increase in threshold value leads to a smaller threshold ε(u), which encourages the agent to make better use of historical information, accelerates training convergence, and thus obtains higher values. The optimal action selection strategy with the highest value;
[0074] 3.3 A feature selection algorithm based on reinforcement learning, including the following sub-steps:
[0075] 3.3.1: Calculate and normalize the dataset, initialize the feature set subset S1 as an empty set, set the Q(S,a) value table to 0, set the relevant parameters of the SVM classifier, and set the discount factor γ and learning rate. Maximum number of training iterations u max ;
[0076] 3.3.2: The agent randomly selects a dataset with a specific feature for training;
[0077] 3.3.3: Based on the feedback recognition accuracy R u (S u ,au ), calculate the ε-greedy threshold ε(u), and update Q. u (S u ,a u Value table;
[0078] 3.3.4: Update the feature set subset S u and training dataset;
[0079] 3.3.5: Use SVM to train and recognize the new dataset, incrementing the training iterations by 1;
[0080] 3.3.6: If the number of training iterations equals u max If the learning ends, the optimal feature set is obtained according to the Q(S,a) value table; otherwise, proceed to sub-step 3.3.3 to start a new round of training.
[0081] By adopting the aforementioned technical solution, the beneficial effects of the present invention are:
[0082] This frequency-hopping modulation scheme identification method first extracts the time-frequency energy map of the frequency-hopping signal and removes noise using an adaptive Wiener filtering algorithm. Then, it extracts 20-dimensional feature vectors from three types of features—geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy—from the time-frequency energy map as a feature set. By establishing a Markov reinforcement learning feature selection model and employing an improved Q-learning algorithm, the feature set is explored and learned to achieve optimal feature set selection and identification of multiple FH modulation schemes. Under low signal-to-noise ratio (SNR) conditions, it exhibits high accuracy, strong robustness, and reduced computational complexity. Simulation experiments show that this reinforcement learning-based feature selection frequency-hopping modulation identification algorithm achieves a 95% recognition rate for 10 types of frequency-hopping signal modulations under a 5dB SNR condition. Attached Figure Description
[0083] Figure 1 These are the time-frequency energy diagrams for QPSK modulation FH, ASK4 modulation FH, QAM64 modulation FH, and MSK modulation FH.
[0084] Figure 2 These are reinforcement learning curves for SVM feature selection using three kernel functions.
[0085] Figure 3 These are the confusion matrices identified by the RBF kernel function SVM, the Linear kernel function SVM, and the Sigmoid kernel function SVM.
[0086] Figure 4 These are the reinforcement learning curves for four action selection strategies.
[0087] Figure 5The SVM recognition accuracy is 4 for four feature sets under different signal-to-noise ratio conditions.
[0088] Figure 6 These are the average recognition rate and computation time of the four algorithms. Detailed Implementation
[0089] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0090] refer to Figures 1-6 This embodiment provides a method for identifying frequency hopping modulation schemes, including:
[0091] Step 1: Receive the FH signal and perform data preprocessing on the FH signal.
[0092] The SPWVD (Smoothed Pseudo Wigner-Ville Distribution) time-frequency transform algorithm is based on the Wigner-Ville distribution. It adds window functions in both the time and frequency domains for smoothing filtering. Selecting an appropriate window function based on the specific application scenario can effectively suppress the influence of cross terms. To improve the effectiveness and representational power of feature extraction under low signal-to-noise ratio conditions, adaptive Wiener filtering and morphological filtering algorithms are used for denoising the time-frequency matrix.
[0093] Suppose F(x,y) represents the energy value in the x-th row and y-th column of the time-frequency matrix F, and the time-frequency matrix F has a size of M×N and a domain of D. F Then (x,y)∈D F 0≤x<M, 0≤y<N, take the neighborhood of the time frequency point F(x,y). If its size is m×n, then the neighborhood Expected value μ and variance σ 2 As shown in equations (1) and (2).
[0094]
[0095]
[0096] Then, the time-frequency point F(x,y) can be obtained as F′(x,y) after Wiener filtering.
[0097]
[0098] Among them, since the noise variance of the actual FH signal is difficult to estimate accurately, υ 2 The mean of the variances of all neighborhoods of the time-frequency matrix F can be taken.
[0099] In order to further remove noise from the time-frequency energy map and smooth the time-frequency energy distribution characteristics, morphological filtering is used to perform opening operation on the time-frequency matrix, that is, to first erode and then dilate F′(x,y), as shown in Equation (4).
[0100]
[0101] Where H(x,y) represents the energy value of the time-frequency matrix after morphological filtering, b(τ,ν) represents the structuring element, and (τ,ν)∈D b D b Describe the domain of the struct element b(τ,ν), and the operators Θ and ν. These represent erosion and expansion operations, respectively.
[0102] F′(x,y)Θb(τ,ν)=min{F′(x+τ,y+ν)-
[0103] b(τ,ν)|(x+τ,y+ν)∈D F ,(τ,ν)∈D b} (5)
[0104]
[0105] After filtering, the single-hop time-frequency matrix of the FH signal is extracted using a connected component detection algorithm. In order to improve the stability of the extracted features, the single-hop time-frequency matrix is grayscaled using equation (7).
[0106]
[0107] Where G(x,y) represents the gray value in the x-th row and y-th column of the time-frequency gray matrix G. This indicates the floor function, and I represents the gray level.
[0108] Figure 1 This presents the single-hop time-domain waveforms and time-frequency energy distribution diagrams of four modulated FH signals: QPSK, ASK4, QAM64, and MSK. In the experiment, the hop rate of all four modulated FH signals was 500 hops / s, the source rate was 16 kb / s, the starting frequency was 1 MHz, the FH interval was 25 kHz, the number of FH points was 256, the FH frequency transitioned pseudo-randomly, the noise was Gaussian white noise, the signal-to-noise ratio was 0 dB, a Hamming window of length 255 was selected for SPWVD, and the neighborhood of the adaptive Wiener filter was used. Set the size to 32×32, and select an ellipse for the morphological filter structuring element b(τ,ν).
[0109] Depend on Figure 1It is known that the waveform characteristics of different modulated FH signals correspond one-to-one with the time-frequency energy distribution. This paper treats the time-frequency energy matrix as an image and extracts the feature parameters that characterize the time-frequency energy distribution of various modulated FH signals to realize the identification of FH modulation mode, thereby effectively overcoming the influence of pseudo-random variation of the carrier frequency of FH signals.
[0110] Step 2, Feature Extraction, including:
[0111] 2.2.1 Characteristics of geometrically invariant moments:
[0112] Geometric invariant moments, due to their translation, rotation and scale invariance, are widely used in the field of image classification and recognition. Their definitions are shown in equations (8) and (9).
[0113]
[0114]
[0115] Where, η c,d χ represents the (c+d)th order normalized central distance. c,d Let (c+d) represent the center distance of order (c+d), c+d=2,3,…,c=0,1,2,…,d=0,1,2,…,and the center of the square (x0,y0) is shown in equations (10) and (11).
[0116]
[0117]
[0118] This paper calculates seven Hu-geometric invariant moments using second- and third-order normalized center distances. As a modulation recognition feature parameter, it is shown in Equation (12).
[0119]
[0120] 2.2.2 Pseudo-Zernike moment features:
[0121] Pseudo-Zernike moments exhibit strong robustness to noise. Due to their orthogonality, they have low redundancy. Low-order pseudo-Zernike moments can express image contour information, while high-order pseudo-Zernike moments contain more detailed image information. For ease of calculation of pseudo-Zernike moments, the time-frequency energy grayscale image G(x,y) is mapped onto the unit circle of a new coordinate system I(δcosθ,δsinθ), with the center and origin of the circle both being the moment center (x0,y0). The radius of the unit circle is set to... The polar coordinates (δ, θ) can then be expressed as:
[0122]
[0123] Where δ represents the vector length from pixel G(x,y) in the time-frequency energy grayscale image to the center of the moment (x0,y0), and θ represents the angle of counterclockwise rotation to that vector, which needs to be transformed according to the specific quadrant of the polar coordinates of the pixel. Then, according to the new coordinate system I(δcosθ,δsinθ), the pseudo Zernike moment z of the time-frequency energy grayscale image... g,h It can be represented as
[0124]
[0125] W g,h (δcosθ,δsinθ)=X g,h (δ)exp(jhθ) (15)
[0126] Where g can be zero or a positive integer, h is a positive integer and h≤g, W g,h (δcosθ, δsinθ) is z g,h Basis functions, * denotes complex conjugate, X g,h (δ) is a basis function W g,h The radial polynomial of (δcosθ,δsinθ).
[0127]
[0128] Generally, the more low-order pseudo-Zernike moment features there are, the stronger the noise resistance. 0,0 , z 0,1 and z 1,1 The three low-order pseudo-Zernike moment eigenparameters are not separable. Considering computational limitations, this paper chooses z... 2,0 z 2,1 z 3,0 z 3,1 z 3,2 z 3,3 z 4,0 and z 4,1 Eight pseudo-Zernike moment parameters are used as modulation recognition feature parameters.
[0129] 2.2.3 Characteristics of Ruili Entropy:
[0130] The Ruili entropy feature of the time-frequency energy map can effectively reflect the time-frequency energy distribution. The more complex and disordered the time-frequency energy distribution, the more information it contains, and the larger the Ruili entropy feature value. Conversely, the stronger the regularity of the time-frequency energy distribution, the smaller the Ruili entropy feature value. Its discrete definition is shown in Equation (17).
[0131]
[0132] in, This represents the order of the entropy ξ in Ruili. This article mainly selects Five integer-order Ruili entropy parameters of 3, 5, 7, 9, and 11 are used as modulation recognition feature parameters.
[0133] In summary, in order to fully reflect the characteristics of FH modulation signals and improve the feature representation capability of the dataset, this paper adopts a 20-dimensional feature vector composed of geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy features as the feature set for FH modulation identification.
[0134] 2.2.4 Feature Normalization:
[0135] To ensure that geometric invariant moment features, pseudo-Zernike moment features, and Rayleigh entropy features can all play their distinguishing role and to reduce the influence of factors such as dimensions, value range, and order of magnitude on the feature set's representational ability, this paper uses a logarithmic nonlinear normalization algorithm to process the feature set, making the feature set distribution more uniform.
[0136] Assume the feature set is Y = {Y} ρ Let Y = |ρ=1,2,…,20} min =min{Y ρ}, Y max =max{Y ρ}, then the normalized feature set It can be represented as
[0137]
[0138] Where l is a constant and lY min ≥1.
[0139] Step 3: Feature Selection
[0140] To more comprehensively reflect the characteristics of modulated signals, the feature set needs to incorporate various feature representations. However, increasing the number of feature types introduces numerous irrelevant and redundant features, severely impacting the training speed and accuracy of the recognition model. Traditional feature selection algorithms based on human experience are inefficient and have complex mathematical models. To improve selection efficiency and recognition accuracy, this paper proposes a reinforcement learning-based feature selection algorithm for selecting features from a 20-dimensional feature set composed of geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy. The algorithm first establishes a Markov feature selection model and introduces a Support Vector Machine (SVM) classifier. The agent uses the SVM recognition rate as a reward, continuously updating the feature set subset state through autonomous exploration and interactive learning, seeking the maximum cumulative reward value and the optimal action strategy, ultimately outputting the optimal feature set.
[0141] 3.1 Markov Decision Process
[0142] Reinforcement learning tasks are generally modeled using Markov Decision Processes (MDPs). The feature selection MDP problem established in this paper can be described by a quintuple {S, A, P, R, γ}, where S represents the feature set state space, S... u S represents the state of a subset of the feature set of the agent at time u. u ∈S, u≥1, A represents the action state space for retaining or deleting features, a u Let a represent the action taken by the agent at time u. u ∈A,S u+1 This indicates that the agent is in the current feature set subset state S. u Next, execute action a u The subsequent states, P represents the state transition probability matrix of the feature set, p u This indicates that the agent is in the current feature set substate S. u Next, execute action a u Later transferred to S u+1 The probability of a state, p u ∈P, R represents the reward value, i.e., the modulation recognition accuracy of the feature set, R u (S u ,a u ) represents the state S of the agent in the current feature set subset. u Next, execute action a u The recognition accuracy is denoted by γ, which represents the discount factor, 0≤γ≤1, and is used to measure the importance of future state rewards to the current agent's action selection.
[0143] The goal of reinforcement learning in this paper is for an agent to seek the optimal subset of features by maximizing the expected cumulative reward V through the action selection strategy ψ.
[0144]
[0145] The Q-value function (which is an existing technology) can be used to represent the cumulative expected reward of an agent performing action a in a feature set state S.
[0146]
[0147] According to the Bellman criterion, we can obtain from equation (20):
[0148] Q ψ (S,a)=E[R u (S u ,a u )+γQ(S u ,a u )|S u =S,a u =a] (21)
[0149] 3.2Q Learning Algorithm
[0150] This paper employs a time-difference-based Q-learning algorithm. The agent utilizes the training dataset and, by adopting appropriate action selection strategies, executes actions to delete or retain features, updating the Q(S,a) value accordingly. Through continuous exploration of the feature set, the agent's understanding of the feature set is strengthened. Finally, based on the Q(S,a) value record table, the action selection that maximizes the Q(S,a) value for each feature set subset is identified as the optimal action selection strategy, thus yielding the optimal feature set. The Q(S,a) value update expression is as follows:
[0151]
[0152] in, Indicates the learning rate, This indicates that the agent is in state S of a subset of the feature set. u Q u+1 (S u ,a u The action selection that maximizes the value.
[0153] Q-learning algorithms typically employ an ε-greedy action selection strategy, randomly selecting actions to delete or retain features with probability ε, and selecting the action with the largest Q(S,a) value with probability 1-ε. This can be represented as...
[0154]
[0155] Where rand represents a random number, 0 < rand < 1, and ψ(s) represents the action selection of the agent in the feature set subset state S.
[0156] In the ε-greedy strategy, the threshold ε is fixed and needs to be set based on prior information about the electromagnetic environment and human experience during sample collection. This makes it difficult to effectively balance the agent's exploration and utilization during learning and training. In the early stages of learning, the agent needs to explore as much unknown information as possible to improve the comprehensiveness of the learned information. As learning and training progresses, the agent needs to focus on utilizing historical known information to obtain the highest cumulative reward expectation value and thus the optimal action selection strategy. When the threshold ε is set high, the agent severely ignores historical known information, resulting in a significant increase in the number of training iterations and a slower convergence speed. When the threshold ε is set low, the agent, lacking exploration of unknown information, easily gets trapped in local optima. To effectively balance the agent's exploration and utilization of dataset information and improve the robustness of the learning algorithm in unknown electromagnetic environments, this paper adopts a dynamic threshold ε-greedy strategy, where ε can be expressed as:
[0157]
[0158] in, This represents the average recognition accuracy of the first u training iterations.
[0159] As can be seen from equation (24), in the early training process, due to the relatively small number of training iterations u, the agent acquires relatively little information. The relatively small value of u leads to a larger threshold ε(u), which encourages the agent to explore more information. As the number of training iterations u gradually increases, the average recognition accuracy increases. The continuous increase in threshold value leads to a smaller threshold ε(u), which encourages the agent to make better use of historical information, accelerates training convergence, and thus obtains higher values. The optimal action selection strategy with the highest value.
[0160] 3.3 Sub-steps of the feature selection algorithm based on reinforcement learning:
[0161] 3.3.1: Calculate and normalize the dataset, initialize the feature set subset S1 as an empty set, set the Q(S,a) value table to 0, set the relevant parameters of the SVM classifier, and set the discount factor γ and learning rate. Maximum number of training iterations u max .
[0162] 3.3.2: The agent randomly selects a dataset with a specific feature for training.
[0163] 3.3.3: Based on the feedback recognition accuracy R u (S u ,a u The ε-greedy threshold ε(u) is calculated using equation (24), and Q is updated using equation (22). u (S u ,a u Value table.
[0164] 3.3.4: Based on equation (23), the agent selects the action to be executed and updates the feature set subset S. u and training dataset.
[0165] 3.3.5: The agent uses SVM to train and recognize the new dataset, and the training iteration count is incremented by 1.
[0166] 3.3.6: If the number of training iterations equals u max If the learning ends, the optimal feature set is obtained according to the Q(S,a) value table; otherwise, proceed to step 3.3.3 to start a new round of training.
[0167] Experimental Results and Analysis
[0168] In the experiment, the FH signal hopping rate was 500 hops / s, the starting frequency was 1MHz, the FH interval was 25kHz, the number of FH frequency points was 256, the FH frequency hopping was pseudo-random, the source rate was 16kb / s, and ten modulation schemes were used: BPSK, QPSK, 8PSK, MSK, QAM16, QAM32, QAM64, ASK2, ASK4, and ASK8. The noise was Gaussian white noise with a mean of zero, and the sampling rate was 32MHz. Within the signal-to-noise ratio range of -5dB to 10dB, 500 single-hop time-frequency energy maps of each modulation scheme were collected every 1dB using a preprocessing algorithm. The training and test samples were allocated in a 3:2 ratio, resulting in a total of 48,000 training samples and 32,000 test samples. The recognition accuracy P was used. * The performance of an algorithm is measured by what it is defined as:
[0169]
[0170] Among them, L mc For the number of experiments, This represents the number of test samples in the i-th experiment. This represents the number of samples correctly identified in the i-th experiment.
[0171] Experiment 1 analyzes the impact of different kernel functions on the recognition performance of SVM classifiers. 5000 feature set samples of ten FH modulated signals under a signal-to-noise ratio of 5dB were calculated and allocated to the training and test sets in a 3:2 ratio. SVMs using three different kernel functions—Linear, Gaussian Radial Basis Function (RBF), and Sigmoid—were used as classifiers. Figure 2 It is a reinforcement learning process for SVM feature selection using three kernel functions. Figure 2 It can be seen that the SVM recognition performance of the three kernel functions tends to stabilize with the increase of training times. The SVM using the RBF kernel function has a relatively high convergence speed and recognition rate, reaching 95.6%, while the stable recognition rates of SVM using the Linear and Sigmoid kernel functions are about 93% and 84% respectively, and the convergence speed is relatively slow and the recognition stability is poor.
[0172] Figure 3 This is the confusion matrix identified by SVM using three different kernel functions. Figure 3It can be seen that the SVM using the RBF kernel function generally outperforms the SVMs using the Linear and Sigmoid kernel functions. The three kernel function SVMs show relatively high recognition accuracy for MSK and ASK8, with the RBF kernel function SVM achieving 100% accuracy for MSK, while its accuracy for 8PSK and QAM32 is relatively low. The Sigmoid kernel function SVM's accuracy for QAM32 drops to 82%. In summary, this paper adopts the RBF kernel function for the SVM classifier in the reinforcement learning feature selection algorithm.
[0173] Experiment 2 analyzes the impact of different action selection strategies on recognition performance. The feature set samples from Experiment 1 were used in the experiment. The SVM classifier employed the RBF kernel function. The Q-learning algorithm used both the proposed dynamic threshold ε-greedy strategy and the traditional fixed threshold ε-greedy strategy for action selection. In the fixed threshold ε-greedy strategy, ε was set to values of 0.3, 0.5, and 0.8, respectively. Figure 4 These are the reinforcement learning curves for different action selection strategies. Figure 4 It can be seen that the Q-learning algorithms for the four action selection strategies gradually converge with the increase of training iterations. However, the agent using the ε=0.8 greedy strategy explores too much, resulting in slower convergence and poorer recognition stability. When ε=0.5 or ε=0.3, the agent balances exploration and utilization to some extent, and the algorithm converges around 200 episodes. However, the agent using the ε=0.3 greedy strategy explores relatively less, causing the algorithm to gradually converge to a local optimum. The Q-learning algorithm using the dynamic ε-greedy strategy has relatively higher convergence speed and recognition rate. This is mainly because the dynamic ε-greedy strategy essentially uses the historical average recognition rate as a threshold setting reference, without requiring prior information about the electromagnetic environment of the samples. When the agent reports a low recognition rate, it automatically increases the threshold ε, and the agent focuses on searching for unknown information. When the agent reports a high recognition rate, it automatically decreases the threshold ε, and the agent focuses on utilizing known information. By providing real-time feedback on the recognition results, the agent can automatically adjust its action selection strategy, effectively improving training efficiency and recognition rate, and enhancing the algorithm's robustness.
[0174] Experiment 3 analyzes the recognition performance of the feature set, FCBF, FCBF_NMI, and the proposed FM modulation scheme recognition method under different signal-to-noise ratio (SNR) conditions. The FCBF feature selection algorithm removes irrelevant redundant features by calculating the mutual information parameter between feature vectors. The FCBF_NMI feature selection algorithm normalizes the mutual information based on FCBF, alleviating the multi-value bias problem. Experimental data uses a 20-dimensional feature set of 80,000 single-hop time-frequency energy map samples within an SNR range of -5dB to 10dB, with an SNR interval of 1dB. The training and test sets are allocated in a 3:2 ratio. An RBF kernel function SVM classifier is used to recognize the feature sets selected by the feature set, FCBF, and FCBF-NMI feature selection algorithms. The Q-learning algorithm employs a dynamic threshold ε-greedy strategy. 100 recognition experiments are conducted on the four feature sets under each SNR interval condition.
[0175] Figure 5 This represents the recognition accuracy of four feature sets under different signal-to-noise ratio conditions. Figure 5 It can be seen that the SVM recognition accuracy of the four feature sets gradually increases with the increase of signal-to-noise ratio. Under the same signal-to-noise ratio, the recognition rate of the original feature set is relatively the lowest. The SVM recognition accuracy of the feature set output by the reinforcement learning feature selection algorithm in this paper is better than that of the FCBF and FCBF-NMI feature selection algorithms. This is mainly because the feature selection criteria of the FCBF and FCBF-NMI feature selection algorithms are single and have poor robustness. The reinforcement learning feature selection algorithm in this paper does not rely on the selection criteria designed by human experience. Instead, the agent learns autonomously and interactively through data samples, and continuously adjusts the learning strategy according to the real-time feedback recognition results to gradually achieve the optimal feature set selection, which effectively improves the feature set representation ability and recognition accuracy.
[0176] Experiment 4 analyzes the recognition efficiency of the FCBF feature selection algorithm, the FCBF-NMI feature selection algorithm, the Q-learning algorithm based on a fixed threshold, and the reinforcement learning feature selection algorithm presented in this paper. For each experiment, 1000 samples with different signal-to-noise ratios were randomly selected; 900 samples were used for training, and 100 samples were used for testing. A total of 10 test experiments were completed, and the results were averaged. The average recognition rate and computation time of the four algorithms are shown below. Figure 5 As shown.
[0177] Depend on Figure 5As can be seen, the average recognition rate of the algorithm presented in this paper is relatively high, but the computation time is significantly increased compared to the FCBF-NMI feature selection algorithm. This is mainly because the algorithm in this paper requires multiple rounds of autonomous exploration and learning based on the feedback results of training to seek the maximum recognition rate, while the FCBF-NMI feature selection algorithm can complete feature selection by calculating the symmetric uncertainty index between each feature, and the algorithm's computation time is relatively less. The computation time of the algorithm presented in this paper is reduced compared to the fixed threshold Q-learning algorithm, mainly because the dynamic adjustment of the threshold ε improves the efficiency of the algorithm's early exploration and learning and the convergence speed in the later stages.
[0178] This reinforcement learning-based feature selection algorithm for FH modulation scheme recognition extracts geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy features from the grayscale image of the time-frequency energy distribution of each hop, forming a 20-dimensional feature set. To further improve training speed and recognition accuracy, a reinforcement learning feature selection algorithm is used to select features from the feature set and remove redundant features. Experimental results show that the proposed algorithm has good recognition performance under different signal-to-noise ratio conditions, and the optimal feature set has strong feature representation ability and robustness. Reducing the computational complexity of the algorithm is the focus of future research.
[0179] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.
Claims
1. A method for identifying frequency hopping modulation patterns, characterized in that, Includes the following steps: Step 1: Receive the FH signal and perform data preprocessing on the FH signal, including: The SPWVD time-frequency transformation algorithm is used to extract the time-frequency energy map of the FH signal; adaptive Wiener filtering and morphological filtering algorithms are used to remove noise from the time-frequency energy map; and then the time-frequency energy map is converted into a time-frequency energy grayscale map. Step 2, Feature Extraction, including: The modulation recognition feature parameters of three types of features, namely geometric invariant moments, pseudo-Zernike moments, and Rayleigh entropy, are extracted from the time-frequency energy grayscale image and used as the feature set for FH modulation recognition. Step 3, Feature Selection and Recognition, including: The optimal feature set selection and multi-type FH modulation scheme recognition are achieved by using reinforcement learning feature selection algorithm and support vector machine classifier; The reinforcement learning feature selection algorithm in step three includes the following sub-steps: 3.1 Markov Decision Process: Modeling is performed using Markov decision processes, consisting of quintuples. describe; Represents the state space of the feature set; Indicates that the agent is in the first... The feature set subset of states at time 1 , ; The action state space represents whether features are retained or deleted; Indicates that the agent is in the first... Actions performed at all times ; This indicates the state of the agent in the current subset of features. Next action The state after; The state transition probability matrix represents the feature set; This indicates that the agent is in the state of the current feature set. Next action Later transferred to The probability of a state. ; This represents the reward value, which is the modulation recognition accuracy of the feature set. This indicates the state of the agent in the current subset of features. Next action The accuracy of the recognition after the process; Indicates the discount factor. This is used to measure the importance of future state rewards to the current agent's action selection; By implementing action selection strategies Maximize the expected cumulative reward To find the optimal subset of features; ; The Q-value function is used to represent the agent's state in the feature set. Next action The expected cumulative reward value; ; According to Bellman's criterion, we can obtain: ; 3.2 Q-learning algorithm: Employing a time-difference-based Q-learning algorithm, the agent utilizes the training dataset to execute actions that delete or retain features by adopting appropriate action selection strategies, and... The values are updated and recorded. Through continuous exploration of the feature set, the agent's understanding of the feature set is strengthened, and ultimately, based on... Value record table, find the maximum value under each feature set subset state. The action selection based on the value is the optimal action selection strategy, thus yielding the optimal feature set; The value update expression is: ; in, Indicates the learning rate, , This indicates the state of the agent in a subset of the feature set. lower envoy Action selection that maximizes value; Q-learning algorithm uses Action selection strategy, in order to The probabilistic random selection of feature actions to delete or retain, in order to Probability Selection The action with the largest value can be represented as: ; in, Represents a random number. , This indicates the state of the agent in a subset of the feature set. The next action selection; A dynamic threshold is used Strategies, in which It can be represented as: ; in, Indicates the preceding Average recognition accuracy after 1 training cycle ; During the initial training phase, due to the number of training sessions... The information acquired by the agent is relatively limited. Relatively small, resulting in a threshold Increasing the value can encourage the agent to explore more information, and this effect increases with the number of training iterations. Gradually increase, average recognition accuracy The threshold is continuously increased, leading to... A smaller value can encourage the agent to make better use of historical information, thus accelerating training convergence and enabling it to acquire more advanced skills. The optimal action selection strategy with the highest value; 3.3 A feature selection algorithm based on reinforcement learning, including the following sub-steps: 3.3.1: Calculate and normalize the dataset, and initialize the feature set subset. For empty set, The value table is 0, and the relevant parameters of the SVM classifier are set with discount coefficients. learning rate Maximum number of training sessions ; 3.3.2: The agent randomly selects a dataset with a specific feature for training; 3.3.3: Based on the feedback recognition accuracy ,calculate threshold ,renew Value table; 3.3.4: Update Feature Set Subset and training dataset; 3.3.5: Use SVM to train and recognize the new dataset, incrementing the training iterations by 1; 3.3.6: If the number of training iterations equals Then the learning ends, according to If the optimal feature set is obtained from the value table, otherwise proceed to sub-step 3.3.3 to start a new round of training.
2. The frequency hopping modulation mode identification method according to claim 1, characterized in that: In step one, the SPWVD time-frequency transformation algorithm is used to perform smoothing filtering on the FH signal in both the time and frequency domains using window functions, forming a time-frequency matrix. ; Representing the time-frequency matrix No. OK Energy values of columns, time-frequency matrix Size is The domain is ,but , , Take time and frequency points neighborhood Its size is Then the expectation of the neighborhood and variance for: ; ; but After Wiener filtering, we can obtain : ; Take the time-frequency matrix The mean of the variances of all neighborhoods; Morphological filtering of the time-frequency matrix Perform opening operations, including... Corrosion followed by expansion treatment: ; in, This represents the energy value of the time-frequency matrix after morphological filtering. Represents a structural element. , Represents structural element Domain, Operator and These represent the erosion and expansion operations, respectively: ; ; After filtering, a connected component detection algorithm is used to extract the single-hop time-frequency matrix of the FH signal. To improve the stability of the extracted features, the single-hop time-frequency matrix is grayscaled to form a time-frequency energy grayscale image. ; in, Represents the time-frequency grayscale matrix No. OK List grayscale values, This indicates the floor function. This indicates the grayscale level.
3. The frequency hopping modulation mode identification method according to claim 2, characterized in that, The characteristic formula for the geometrically invariant moment is: ; ; in, express Normalized center distance, express center distance of the steps , , , center of the square for: ; ; Seven geometrically invariant moments were calculated using second- and third-order normalized center moments. The modulation recognition feature parameter is: 。 4. The frequency hopping modulation mode identification method according to claim 3, characterized in that, The pseudo-Zernike moment feature is defined as follows: Time-frequency energy grayscale image Mapping to the new coordinate system On the unit circle, the center and the origin are both the square center. The radius of the unit circle is set to Then polar coordinates Represented as: ; in, Represents the pixels of the time-frequency energy grayscale image To the center of the rectangle The length of the vector. This indicates the angle of counterclockwise rotation to the vector, which requires angle conversion based on the specific quadrant of the polar coordinates of the pixel. According to the new coordinate system Pseudo Zernike moments of time-frequency energy grayscale images It can be represented as: ; ; in, It can take the value of zero or a positive integer. Take positive integers and , for basis functions, This indicates taking the complex conjugate. basis functions The radial polynomial; ; choose , , , , , , and Eight pseudo-Zernike moment parameters are used as modulation recognition feature parameters.
5. The frequency hopping modulation mode identification method according to claim 4, characterized in that, The discrete definition of the Ruili entropy feature is as follows: ; in, Represents Ruili entropy The order of , ; choose Five integer-order Ruili entropy parameters of 3, 5, 7, 9, and 11 are used as modulation recognition feature parameters.