Optimal Frequency Hopping Sequence Generation Method Based on Deep Reinforcement Learning

Through a deep reinforcement learning method, using deep neural networks and Monte Carlo trees to optimize frequency hopping sequence generation, the problem of insufficient dynamic adjustment and optimization capabilities in the existing technology is solved, and the optimal frequency hopping sequence that meets the Lempel world is efficiently and accurately found.

CN119995630BActive Publication Date: 2025-06-13XIHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510452123.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing frequency hopping sequence generation method has insufficient ability to dynamically adjust and optimize the generation sequence, which makes it impossible to meet the needs of the user's communication system in a complex environment.

Method used

Using a deep reinforcement learning method, the Monte Carlo tree is guided to fill in sequence through the deep neural network DNN to generate a frequency hopping sequence, and through repeated training and update of the deep neural network, the sequence is gradually expanded and optimized until the optimal frequency hopping sequence is obtained.

Benefits of technology

It realizes efficiently and accurately finding the optimal q element frequency hopping sequence that meets the Lempel world in complex high-dimensional environments, enhances the decision-making ability of frequency hopping sequence generation, and improves anti-interference performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995630B_ABST
    Figure CN119995630B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimal frequency hopping sequence generation method based on deep reinforcement learning, belonging to the fields of artificial intelligence and communication technology. The steps are as follows: Obtain the frequency gap q of the target generated frequency hopping sequence, the sequence length N, and the step size filled each time for the sequence, where the step size is a prime factor of the sequence length N; Use a deep neural network DNN to guide the Monte Carlo tree for sequence filling to generate the frequency hopping sequence result; Repeat generating the frequency hopping sequence result and construct a frequency hopping sequence training set; Train and update the deep neural network DNN to obtain a trained deep neural network DNN; According to the frequency gap q of the target generated frequency hopping sequence, the sequence length N, and the step size filled each time for the sequence, use the Monte Carlo tree and the trained deep neural network for sequence extension filling until the optimal frequency hopping sequence is obtained. The present invention solves the problem of insufficient ability of the existing frequency hopping sequence generation method to dynamically adjust and improve the generated sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and communication technology, and particularly relates to an optimal frequency hopping sequence generation method based on deep reinforcement learning. Background Art

[0002] Frequency hopping technology is a wireless communication technology that transmits information by continuously changing frequencies during communication. It has remarkable characteristics such as strong concealment, high anti-interception ability, strong anti-jamming ability, and excellent system compatibility. With the rapid development of modern information technology, frequency hopping technology has been widely applied in military communication, civilian communication, and the Internet of Things field, becoming a key means to improve anti-jamming ability, gain utilization efficiency, and enhance communication security. In a frequency hopping communication system, if different users use the same frequency slot within the same time slot, collisions will occur, that is, there will be mutual interference, thus affecting the communication quality. The number of collisions between frequency hopping sequences is usually represented by Hamming correlation. Therefore, designing frequency hopping sequences with low Hamming correlation is beneficial to improving the performance of communication systems.

[0003] The construction of frequency hopping sequences can generally be divided into two categories: mathematical construction and machine learning generation. In 1974, Lempel and Greenberger proposed a theoretical bound for frequency hopping sequence design - the Lempel bound, which defines the minimum Hamming correlation that can be achieved in the ideal case of frequency hopping sequence design, providing a theoretical reference for the evaluation of frequency hopping sequence performance. Constructing optimal frequency hopping sequences that satisfy the Lempel bound generally uses methods such as algebra, combinatorial mathematics, and recursive construction. Although traditional mathematical construction methods have certain advantages in theory, when faced with complex sequence structures, they are often restricted by parameter limitations, which limits their flexibility and effectiveness in some practical applications. In contrast, the frequency hopping sequence generation technology based on machine learning shows more powerful advantages. With the flexibility of intelligent algorithms and powerful computing capabilities, it can not only deeply explore a richer and more diverse sequence parameter space but also significantly improve the anti-jamming performance of sequences, thus better meeting the needs of user communication systems in complex environments. However, the current methods for generating optimal frequency hopping sequences based on intelligent algorithms still have problems of insufficient dynamic adjustment and optimization efficiency. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, an optimal frequency hopping sequence generation method based on deep reinforcement learning provided by the present invention solves the problem of insufficient ability to dynamically adjust and improve the generated sequences in the existing frequency hopping sequence generation methods.

[0005] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0006] An optimal frequency hopping sequence generation method based on deep reinforcement learning provided by the present invention includes the following steps:

[0007] S1. Obtain the frequency gap q of the target generated frequency hopping sequence, the sequence length N, and the step size for each filling of the sequence , where the step size is a prime factor of the sequence length N;

[0008] S2. According to the frequency gap q of the target generated frequency hopping sequence, the sequence length N, and the step size for each filling of the sequence , use a deep neural network DNN to guide the Monte Carlo tree for sequence filling to generate a frequency hopping sequence result;

[0009] S3. Repeat S2 for a preset number of times, and construct a frequency hopping sequence training set based on the generated frequency hopping sequence results;

[0010] S4. Use the frequency hopping sequence training set to train and update the deep neural network DNN to obtain a trained deep neural network DNN;

[0011] S5. According to the frequency gap q of the target generated frequency hopping sequence, the sequence length N, and the step size for each filling of the sequence , use the Monte Carlo tree and the trained deep neural network to repeat S2 - S4 for sequence extension filling until an optimal frequency hopping sequence is obtained.

[0012] The beneficial effects of the present invention are as follows: An optimal frequency hopping sequence generation method based on deep reinforcement learning provided by the present invention converts the frequency hopping sequence generation process into a sequence filling process, constructs a Monte Carlo tree corresponding to the sequence filling process, realizes a search algorithm based on Monte Carlo tree simulation, searches for a better sequence filling decision path, and by observing the filling state of the sequence at each time step and using the deep neural network DNN, realizes learning the optimal sequence filling action in a complex high-dimensional environment corresponding to the sequence filling state; in the present invention, the deep neural network DNN is used to guide the evaluation of the expansion and simulation of the Monte Carlo tree, and the deep neural network DNN is trained using a frequency hopping sequence training set composed of a large number of frequency hopping sequence results generated by the Monte Carlo tree, realizing mutual promotion, enhancing the decision-making ability of the entire frequency hopping sequence generation, and being able to more efficiently and accurately find the optimal q-ary frequency hopping sequence that satisfies the Lempel bound.

[0013] Further, the S2 includes the following steps:

[0014] S201. Take the initial state of the current sequence with length N as the root node of the Monte Carlo tree, where the value of the unfilled sequence position is set to 0, and the value of the filled sequence position is an integer, and the value range is [1, q];

[0015] The calculation expression for the initial state of the current sequence is as follows:

[0016] ,

[0017] where, represents the initial state of the current sequence, represents the filling state of the i-th sequence, where, , and i is a natural number. When the sequence filling reaches the -th sequence filling state, the sequence filling is completed;

[0018] S202. Determine whether the current sequence is filled. If so, go to S211; otherwise, go to S203;

[0019] S203. Starting from the root node according to the Monte Carlo tree method, respectively, through paths, each vertex can be reached, and each path stores the corresponding number of visits , average reward and prior probability , where, represents the -th state of the current sequence, represents the j-th filling action of the current sequence, ;

[0020] S204. Based on the number of visits , average reward and prior probability corresponding to each path, select the optimal path as the current node based on the Upper Confidence Bound (UCB) model, and determine whether the current node is a leaf node. If so, go to S205; otherwise, return to S204, where the path with the highest UCB value is the optimal path;

[0021] S205. Determine whether the current node has been visited. If so, go to S206; otherwise, take the current node as the expansion node and go to S207;

[0022] S206. Expand the current node, extract features according to the Deep Neural Network (DNN) to obtain the probability distribution and value evaluation of the current node, and select the expansion node according to the probability distribution of the current node, then go to S207;

[0023] S207. Starting from the expansion node, randomly select filling actions to simulate the process of filling the sequence to completion in all cases, and obtain the complete simulation sequence result set and the corresponding complete branch simulation results;

[0024] S208. According to the complete simulation sequence result set, backpropagate layer by layer from the child nodes of the complete branch simulation result to the root node, and update the node access times and the corresponding average reward values of all paths on the complete branch simulation result;

[0025] S209. Based on the node access times of all paths on the complete branch simulation result, calculate the node movement selection probability according to the node selection model;

[0026] S210. Select the filling action in the current sequence filling state according to the node movement selection probability to obtain the next sequence filling state, and use the next sequence filling state as the initial state of the current sequence, then return to S202;

[0027] S211. Use the current sequence, the probability of the current sequence, and the reward value corresponding to the current sequence as the frequency hopping sequence result.

[0028] The beneficial effects of adopting the above further scheme are as follows: The present invention constructs a Monte Carlo tree corresponding to the sequence state and the path of sequence filling expansion, realizes the correspondence between the sequence filling state and the sequence state, and provides a basis for simulating and filling all possible complete q-ary frequency hopping sequences; The present invention, based on the storage access times, average reward, and prior probability, realizes selecting the optimal path for the current node through the upper confidence bound UCB model, and combines the probability distribution and value evaluation obtained by feature extraction through the deep neural network DNN, providing a basis for confirming the optimal frequency hopping sequence elements; Through the complete simulation filling and backpropagation of the Monte Carlo tree, the present invention calculates the node movement selection probability, realizes gradually filling to obtain a complete q-ary frequency hopping sequence that satisfies the Lempel bound optimally, and provides a basis for feedback training of the deep neural network DNN.

[0029] Further, the calculation expression of the upper confidence bound UCB model in S204 is as follows:

[0030] ,

[0031] where represents the th path corresponding to the highest UCB value, represents taking the j value corresponding to the highest UCB value, and c represents the path exploration weighting factor.

[0032] The beneficial effects of adopting the above further scheme are as follows: The present invention provides a calculation method for the upper confidence bound UCB model, which provides a basis for selecting the optimal path from the root node of the Monte Carlo tree to each node, and also provides a basis for further expanding the nodes for optimal selection.

[0033] Further, S206 includes the following steps:

[0034] S2061. Obtain the sequence image P corresponding to the current node. Here, the size of the sequence image P is 1*N. The sequence image P corresponding to the current node is an image composed of sequences corresponding to the sequence filling states of the current node. Each of the N sequence elements in the image occupies a unit length, and the total length is N;

[0035] S2062. Expand the sequence image P into q + 1 sub-sequence images of 1*N. Here, each sub-sequence image corresponds to a sub-sequence. When is a positive number and , the th sub-sequence image only keeps consistent with the sequence positions where the sequence element values in the sequence image P are , and the sequence element values at the remaining sequence positions are all taken as 0. And the (q + 1)th sub-sequence image takes the value 1 at the sequence positions where the sequence elements in the corresponding sequence image P are 0, and takes 0 at the sequence positions where the sequence elements in the corresponding sequence image P are not 0;

[0036] S2063. Input the q + 1 sub-sequence images of 1*N into the deep neural network DNN for spatial feature extraction, and obtain the probability distribution and value evaluation corresponding to the sequence image P, which are used as the probability distribution and reward value of the current node;

[0037] S2064. According to the probability distribution of the current node, select the path corresponding to the sub-sequence image with the highest probability and add it to the Monte Carlo tree to obtain the expanded node.

[0038] The beneficial effect of adopting the above further solution is as follows: By constructing sub-sequence images and using the deep neural network DNN to perform feature extraction, probability and reward prediction on each sub-sequence image, the present invention realizes the selection of the target expanded node, providing a guarantee for gradually filling to obtain a complete q-ary frequency hopping sequence that satisfies the Lempel bound and is optimal.

[0039] Further, the calculation expression of the node selection model in S209 is as follows:

[0040] ,

[0041] where represents the node movement selection probability, represents the softmax function, represents the node selection scaling factor, represents the access times of the root node.

[0042] The beneficial effects of adopting the above further scheme are as follows: The present invention provides a calculation method for a node selection model, which can select the next state of sequence filling and determine the best choice for filling sequence elements in the current round according to the access update that occurs when the simulation results of the complete branch are uploaded back to the root node, providing a basis for gradually generating a q-ary frequency hopping sequence that satisfies the Lempel bound optimally.

[0043] Further, the calculation expression of the reward value in S211 is as follows:

[0044] ,

[0045] ,

[0046] ,

[0047] ,

[0048] ,

[0049] Among them, represents the reward function value corresponding to the current sequence S, represents the metric function value corresponding to the sequence image S, S represents the current sequence S, represents the average maximum metric threshold obtained through noiseless experiments, lempel represents the minimum metric threshold, represents the maximum period Hamming autocorrelation function of the current sequence S at time delay , represents taking the maximum value in the range where the time delay is greater than or equal to 1 and less than or equal to N, r represents the smallest non-negative residue of N modulo q, represents rounding up, represents the period Hamming autocorrelation function of the current sequence S at time delay , represents a function for measuring the similarity between the th sequence element and the th sequence element. Among them, if , then , if is satisfied, then the sequence satisfies the Lempel bound optimally.

[0050] The beneficial effects of adopting the above further scheme are as follows: The present invention provides a calculation method for the reward value, and sets the metric function of the target-generated frequency hopping sequence to the maximum period Hamming autocorrelation value of the sequence, that is, the maximum number of collisions of the sequence at different time delays, providing a basis for generating a q-ary frequency hopping sequence that satisfies the Lempel bound optimally.

[0051] Other advantages of the present invention will be analyzed in more detail in the subsequent embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of the steps of a method for generating an optimal frequency hopping sequence based on deep reinforcement learning in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the drawings below is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0055] Embodiment 1:

[0056] As Figure 1 shown, in an embodiment of the present invention, the present invention provides a method for generating an optimal frequency hopping sequence based on deep reinforcement learning, including the following steps:

[0057] S1. Obtain the frequency gap q, sequence length N, and step size for each filling of the target generated frequency hopping sequence , where the step size is a prime factor of the sequence length N;

[0058] In this solution, the generation process of the target generated sequence can be regarded as a process of filling q-ary sequence elements of length N. Among them, the unfilled value is set to 0, and the filled value ranges from [1, q].

[0059] S1.5. Initialize the learning parameters of the deep neural network DNN, and based on the frequency gap q, sequence length N, and step size for each filling of the target generated sequence , conduct noiseless frequency hopping sequence generation experiments for a preset number of times, and calculate the average metric value of the experimentally generated sequences as the average maximum metric threshold.

[0060] In this embodiment, the preset number of experiments is 50 times. In this solution, when generating the frequency-hopping sequence in the noise-free frequency-hopping sequence generation experiment, the method of S2 is used to expand element by element to obtain the experimental generation sequence. Each time a noise-free frequency-hopping sequence generation experiment generates a frequency-hopping sequence, after 50 frequency-hopping sequences are completely generated, the obtained average maximum metric threshold provides a basis for calculating the reward value corresponding to the sequence.

[0061] S2. According to the frequency gap q of the target-generated frequency-hopping sequence, the sequence length N, and the step size filled each time in the sequence , use the deep neural network DNN to guide the Monte Carlo tree for sequence filling to generate the frequency-hopping sequence result;

[0062] The S2 includes the following steps:

[0063] S201. Use the initial state of the current sequence with length N as the root node of the Monte Carlo tree. Among them, the value of the unfilled sequence position is set to 0, and the value of the filled sequence position is an integer, and the value range is [1, q];

[0064] The calculation expression of the initial state of the current sequence is as follows:

[0065] ,

[0066] Among them, represents the initial state of the current sequence, represents the i-th sequence filling state, where , and i is a natural number. When the sequence filling reaches the -th sequence filling state, the sequence filling is completed;

[0067] S202. Determine whether the current sequence is filled. If so, enter S211; otherwise, enter S203;

[0068] S203. According to the Monte Carlo tree method, starting from the root node, paths can reach each vertex respectively, and each path stores the corresponding access times , average reward and prior probability . Among them, represents the -th state of the current sequence, represents the j-th filling action of the current sequence, ; In this solution, the state of the sequence corresponds to the filled part of the sequence formed from the root node to the current node, and the filling action of the sequence is the path selection corresponding to the node, that is, the integer value that the sequence can be filled, and the value range of the fillable integer value is [1, q];

[0069] S204. According to the access times corresponding to each path , average reward and prior probability , select the optimal path as the current node based on the Upper Confidence Bound (UCB) model, and determine whether the current node is a leaf node. If so, enter S205; otherwise, return to S204, where the path corresponding to the highest UCB value is the optimal path;

[0070] The calculation expression of the Upper Confidence Bound (UCB) model in the above-mentioned S204 is as follows:

[0071] ,

[0072] where represents the th path corresponding to the highest UCB value, represents taking the j value corresponding to the highest UCB value, and c represents the path exploration weighting factor. In this embodiment, c is a constant.

[0073] S205. Determine whether the current node has been visited. If so, enter S206; otherwise, take the current node as the expansion node and enter S207;

[0074] S206. Expand the current node, perform feature extraction according to the Deep Neural Network (DNN) to obtain the probability distribution and value evaluation of the current node, and select the expansion node according to the probability distribution of the current node, then enter S207;

[0075] The above-mentioned S206 includes the following steps:

[0076] S2061. Obtain the sequence image P corresponding to the current node, where the size of the sequence image P is 1*N. The sequence image P corresponding to the current node is an image composed of the sequences corresponding to the sequence filling states of the current node, and the N sequence elements in the image each occupy a unit length, with a total length of N;

[0077] S2062. Expand the sequence image P into q + 1 sub-sequence images of 1*N. Each sub-sequence image corresponds to a sub-sequence. When is a positive number and , the th sub-sequence image only keeps the same as the sequence positions where the sequence element values in the sequence image P are , and the sequence element values at the remaining sequence positions are all taken as 0, while the (q + 1)th sub-sequence image Take the value 1 at the sequence positions where the sequence elements in the corresponding sequence image P are 0, and take 0 at the sequence positions where the sequence elements in the corresponding sequence image P are not 0;

[0078] S2063. Input these q + 1 sub-sequence images of 1*N into the deep neural network DNN for spatial feature extraction to obtain the probability distribution and value evaluation corresponding to the sequence image P, which are used as the probability distribution and reward value of the current node; in this solution, the value range of the reward value corresponding to the sequence image P is [-1, 1];

[0079] S2064. According to the probability distribution of the current node, select the path corresponding to the sub-sequence image with the highest probability and add it to the Monte Carlo tree to obtain an expanded node.

[0080] S207. Starting from the expanded node, randomly select filling actions to simulate the process of sequence filling until sequence filling is completed in all cases, and obtain the complete set of simulation sequence results and the corresponding complete branch simulation results;

[0081] S208. According to the complete set of simulation sequence results, backpropagate from the sub-nodes of the complete branch simulation results layer by layer to the root node, and update the node access times and the corresponding average reward values of all paths on the complete branch simulation results;

[0082] S209. According to the node access times of all paths on the complete branch simulation results, calculate the node movement selection probability based on the node selection model;

[0083] The calculation expression of the node selection model in S209 is as follows:

[0084] ,

[0085] where, represents the node movement selection probability, represents the softmax function, represents the node selection scaling factor, represents the access times of the root node.

[0086] S210. Select the filling action in the current sequence filling state according to the node movement selection probability to obtain the next sequence filling state, and use the next sequence filling state as the initial state of the current sequence, and return to S202;

[0087] S211. Use the current sequence, the probability of the current sequence, and the reward value corresponding to the current sequence as the frequency hopping sequence result.

[0088] The calculation expression of the reward value in S211 is as follows:

[0089] ,

[0090] ,

[0091] ,

[0092] ,

[0093] ,

[0094] Among them, represents the reward function value corresponding to the current sequence S, represents the metric function value corresponding to the sequence image S, where S represents the current sequence S, represents the average maximum metric threshold obtained through noise-free experiments, and lempel represents the minimum metric threshold, represents the maximum periodic Hamming autocorrelation function of the current sequence S at the time delay , represents taking the maximum value in the range where the time delay is greater than or equal to 1 and less than or equal to N, and r represents the least non-negative residue of N modulo q, represents rounding up, represents the periodic Hamming autocorrelation function of the current sequence S at the time delay , represents a function for measuring the similarity between the th sequence element and the th sequence element. Among them, if , then , and if is satisfied, the sequence satisfies the Lempel bound optimality.

[0095] S3. Repeat S2 a preset number of times, and construct a frequency hopping sequence training set based on the generated frequency hopping sequence results;

[0096] S4. Use the frequency hopping sequence training set to train and update the deep neural network DNN to obtain the trained deep neural network DNN;

[0097] S5. According to the frequency gap q of the target-generated frequency hopping sequence, the sequence length N, and the step size filled in each sequence, use the Monte Carlo tree and the trained deep neural network to repeat S2 - S4 for sequence extension and filling until the optimal frequency hopping sequence is obtained.

[0098] Example 2:

[0099] On the basis of Example 1, in a practical example of the present invention, taking the experiment of finding the optimal frequency hopping sequence with a frequency gap of 3 and a length within 30 as an example, the parameter settings are shown in Table 1:

[0100]

[0101] In this experiment, the input parameter is the frequency gap q = 3 of the sequence, the length of the sequence to be searched is from q + 1 to 30, and the step size for each sequence filling is a prime factor of N. After 50 rounds of noise-free experiments, the average metric value E[M] of the sequence is calculated, and the metric function value and reward value corresponding to the sequence are calculated. In MCTS, each sequence filling state s i is set as the root node of the search tree, and m = 900 look-ahead simulations are performed. In each simulation, Dirichlet noise α = 0.05 is added to the prior probability of the root node to introduce exploration. After 900 simulations, by calculating the node movement selection probability , where τ = 1 is set in the first one-third of the time steps, and the probability of selecting a movement is proportional to its visit count; τ = 10 is set in other time steps -4 , and the one with the most visit counts is deterministically selected for movement. In the deep neural network DNN, the update period G = 50 is set, that is, 6 mini-batch data are extracted for training the deep neural network DNN after every 50 experiments. In this embodiment, the mini-batch data is set to 64.

[0102] The solution of the present invention uses the sum of minimizing the mean square error and the cross-entropy loss as the loss function, trains the deep neural network DNN when q = 3 and N = 17, and records the loss values corresponding to different training steps. The loss value decreases significantly in the initial stage of training, indicating that the model quickly learns the basic features. During the training process, the loss value gradually tends to be stable, showing the convergence of the model. The decreasing trend of the loss value with the increase of the training steps verifies the effectiveness of the proposed optimization algorithm and model structure, which can significantly reduce the training error and improve the model performance. In addition, the slight fluctuations of the loss value with the increase of the training steps are caused by using the stochastic gradient descent method for optimization, which is a normal phenomenon and does not affect the final convergence effect of the model.

[0103] In this embodiment, when the solution of the present invention is compared with the existing frequency hopping sequence generation solutions, the present invention can generate frequency hopping sequences in a machine learning manner and can find more sequence parameters than traditional sequence construction solutions. Taking the example of finding the optimal frequency hopping sequence with a frequency gap of 3 and a sequence length within 30, the solution of the present invention can find the optimal frequency hopping sequences with a sequence length of 27 and a sequence length of 30, while the parameters of the existing frequency hopping sequence generation solutions cannot generate them. Among them, the existing construction solutions include: Solution 1: "Construction of Optimal Uniform Wide-Bandgap Frequency Hopping Sequences", published in IEEE Transactions on Information Theory, Vol. 68, No. 1, pp. 692-700, January 2022;

[0104] Solution 2: "A Novel Optimal Frequency Hopping Sequence Based on Circle Division" was published in the 13th International Computer Conference on Wavelet Active Media Technology and Information Processing, 2016;

[0105] Solution 3: "New Families of Composite-Length Optimal Frequency Hopping Sequences", published in IEEE Transactions on Information Theory, pp. 3688 - 3697, 2014;

[0106] Solution 4: "New Constructions of Optimal Sets of Frequency Hopping Sequences", published in IEEE Transactions on Information Theory, pp. 3831 - 3840, 2011;

[0107] Solution 5: "On the Sidon Sequences as Frequency Hopping Sequences", published in IEEE Transactions on Information Theory, pp. 4279 - 4285, 2009;

[0108] Solution 6: "Further Combinatorial Constructions of Optimal Frequency Hopping Sequences" was published in Journal of Combinatorial Theory, Series A, pp. 1699 - 1718, 2006.

[0109] The comparison construction results between the solution of the present invention and Comparative Solutions 1 to 6 are shown in Tables 2 - 1, 2 - 2, 2 - 3, and 2 - 4 as follows:

[0110]

[0111]

[0112]

[0113]

[0114] It can be found from Tables 2 - 1, 2 - 2, 2 - 3, and 2 - 4 that with the change of sequence length and frequency gap, under most different sequence lengths and different conditions satisfying the optimal conditions of the Lempel bound, both the solution of the present invention and Comparative Solutions 1 - 6 can generate optimal frequency hopping sequences. However, when the sequence length is 27 and the frequency gap is 3, and when the sequence length is 30 and the frequency gap is 3, only the solution of the present invention can generate optimal frequency hopping sequences, while Comparative Solutions 1 - 6 cannot generate optimal frequency hopping sequences. This shows that compared with the sum of the existing comparative solutions, the solution of the present invention still has the ability to find more sequence parameters.

[0115] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for generating an optimal frequency hopping sequence based on deep reinforcement learning, characterized in that: The steps include: S1. Obtain the frequency gap q, sequence length N and the step length of each filling of the target generated frequency hopping sequence , where the step length is the prime factor of the sequence length N; S2, generate the frequency gap q of the frequency hopping sequence according to the target, the sequence length N and the step length of each filling of the sequence , using deep neural network DNN to guide Monte Carlo tree to fill in the sequence and generate frequency hopping sequence results; The S2 comprises the following steps: S201, taking the initial state of the current sequence of length N as the root node of the Monte Carlo tree, wherein the value of the unfilled sequence position is set to 0, and the value of the filled sequence position is an integer in the range of [1, q]; The calculation expression of the initial state of the current sequence is as follows: , in, Indicates the initial state of the current sequence, represents the filling state of the ith sequence, where , and i is a natural number, when the sequence is filled to the When the sequence is in the filling state, the sequence filling is completed; S202, determine whether the current sequence is filled, if so, proceed to S211, otherwise proceed to S203; S203, according to the Monte Carlo tree method, starting from the root node, respectively through There are paths that can reach each vertex, and each path stores the corresponding number of visits , average reward and prior probability ,in, Indicates the current sequence status, represents the jth filling action of the current sequence, ; S204: According to the number of visits corresponding to each path , average reward and prior probability , based on the upper confidence interval UCB model, the optimal path is selected as the current node, and it is determined whether the current node is a leaf node. If so, it goes to S205, otherwise it returns to S204, wherein the path corresponding to the highest UCB value is the optimal path; S205, determine whether the current node has been visited, if yes, proceed to S206, otherwise, take the current node as an extended node and proceed to S207; S206, expand the current node, and obtain the probability distribution and value evaluation of the current node after performing feature extraction based on the deep neural network DNN, and select the expansion node based on the probability distribution of the current node, and enter S207; S207, starting from the expansion node, randomly selecting a filling action to simulate the process from sequence filling to sequence filling completion in all cases, and obtaining a complete simulation sequence result set and a corresponding complete branch simulation result; S208. According to the complete simulation sequence result set, the sub-nodes of the complete branch simulation result are transmitted back to the root node layer by layer, and the node access times and corresponding average reward values ​​of all paths on the complete branch simulation result are updated; S209, according to the node access times of all paths in the complete branch simulation results, the node movement selection probability is calculated based on the node selection model; S210, selecting a filling action in the current sequence filling state according to the node movement selection probability, obtaining the next sequence filling state, and taking the next sequence filling state as the initial state of the current sequence, and returning to S202; S211, taking the current sequence, the probability of the current sequence and the reward value corresponding to the current sequence as the frequency hopping sequence result; S3, repeat S2 for a preset number of times, and construct a frequency hopping sequence training set based on the generated frequency hopping sequence results; S4, using the frequency hopping sequence training set to train and update the deep neural network DNN to obtain a trained deep neural network DNN; S5, generate the frequency gap q of the frequency hopping sequence, the sequence length N and the step length of each filling of the sequence according to the target , using the Monte Carlo tree and the trained deep neural network to repeat S2-S4 for sequence expansion and filling until the optimal frequency hopping sequence is obtained.

2. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 1, characterized in that: The calculation expression of the upper confidence interval UCB model in S204 is as follows: , in, Indicates the highest UCB value corresponding to the Path, represents the j value that gives the highest UCB value, and c represents the path exploration weighting factor.

3. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 1, characterized in that: The S206 comprises the following steps: S2061, obtaining a sequence image P corresponding to the current node, wherein the size of the sequence image P is 1*N, and the sequence image P corresponding to the current node is an image composed of a sequence corresponding to the sequence filling state corresponding to the current node, and the N sequence elements in the image each occupy a unit length, and the total length is N; S2062, expand the sequence image P into q+1 1*N subsequence images, where each subsequence image corresponds to a subsequence. is a positive number and At that time, Subsequence images Only with sequence element values ​​in sequence image P The corresponding sequence positions of remain the same, the values ​​of the sequence elements of the other sequence positions are all 0, and the q+1th subsequence image The values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are 0 are set to 1, and the values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are not 0 are set to 0; S2063, input the q+1 1*N subsequence images into the deep neural network DNN to extract spatial features, and obtain the probability distribution and value evaluation corresponding to the sequence image P as the probability distribution and reward value of the current node; S2064. According to the probability distribution of the current node, a path corresponding to the subsequence image with the largest probability is selected and added to the Monte Carlo tree to obtain an extended node.

4. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 1, characterized in that: The calculation expression of the node selection model in S209 is as follows: , in, represents the node movement selection probability, represents the softmax function, Indicates the node selection scaling factor, Indicates the number of visits to the root node.

5. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 1, characterized in that: The calculation expression of the reward value in S211 is as follows: , , , , , in, represents the reward function value corresponding to the current sequence S, represents the metric function value corresponding to the sequence image S, S represents the current sequence S, represents the average maximum metric threshold obtained through noise-free experiments, lempel represents the minimum metric threshold, Indicates that the current sequence S is delayed The maximum periodic Hamming autocorrelation function of Indicates the delay Take the maximum value in the range greater than or equal to 1 and less than or equal to N, r represents the minimum non-negative remainder value of N modulo q, Indicates rounding up. Indicates that the current sequence S is delayed The periodic Hamming autocorrelation function of Indicates the measurement The sequence element and A function of the similarity between sequence elements, where if ,but , if satisfied , then the sequence satisfies the Lempel bound optimality.

Citation Information

Patent Citations

  • Broadband spectrum sensing method based on reinforcement learning

    CN112202514A

  • Enhanced frequency hopping for data transmissions

    WO2022192444A1