Optimal frequency hopping sequence generation method based on deep reinforcement learning

Through a deep reinforcement learning method, using deep neural networks and Monte Carlo trees to optimize frequency hopping sequence generation, the problem of insufficient dynamic adjustment and optimization capabilities in the existing technology is solved, and more efficient optimal frequency hopping sequence generation is achieved.

CN119995630AActive Publication Date: 2025-05-13XIHUA UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510452123.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing frequency hopping sequence generation methods have insufficient ability to dynamically adjust and optimize the generation of sequences, resulting in poor anti-interference performance in complex environments.

Method used

Using a deep reinforcement learning method, the Monte Carlo tree is guided to fill in sequence through the deep neural network DNN to generate frequency hopping sequences, and the sequence generation is gradually optimized by repeatedly training and updating the deep neural network.

Benefits of technology

It realizes efficiently finding the optimal frequency hopping sequence that meets the Lempel world in complex high-dimensional environments, improving the anti-interference performance and decision-making ability of the sequence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995630A_ABST
    Figure CN119995630A_ABST
Patent Text Reader

Abstract

The invention discloses an optimal frequency hopping sequence generation method based on deep reinforcement learning, and belongs to the technical field of artificial intelligence and communication, and the method comprises the following steps: obtaining a frequency slot q of a target generated frequency hopping sequence, a sequence length N and a step length # imgabs0 # of each filling of the sequence, the step length # imgabs1 # being a prime factor of the sequence length N; using a deep neural network DNN to guide a Monte Carlo tree to perform sequence filling, and generating a frequency hopping sequence result; repeatedly generating a frequency hopping sequence result, and constructing a frequency hopping sequence training set; training and updating the deep neural network DNN to obtain a trained deep neural network DNN; and according to the frequency gap q of the target generated frequency hopping sequence, the sequence length N and the step length # imgabs2 # of each filling of the sequence, performing sequence extension filling by using the Monte Carlo tree and the trained deep neural network until the optimal frequency hopping sequence is obtained. According to the invention, the problem that the existing frequency hopping sequence generation method is insufficient in capability of dynamically adjusting and improving the generated sequence is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and communication technology, and in particular relates to an optimal frequency hopping sequence generation method based on deep reinforcement learning. Background Art

[0002] Frequency hopping technology is a wireless communication technology that transmits information by constantly changing the frequency during the communication process. It has the remarkable characteristics of strong concealment, high anti-interception capability, strong anti-interference and excellent system compatibility. With the rapid development of modern information technology, frequency hopping technology has been widely used in military communications, civil communications and the Internet of Things, becoming a key means to improve anti-interference capabilities, gain utilization efficiency and enhance communication security. In a frequency hopping communication system, if different users use the same frequency slot in the same time slot, collisions will occur, that is, there will be mutual interference, which will affect the communication quality. The number of collisions between frequency hopping sequences is usually expressed by Hamming correlation. Therefore, designing a frequency hopping sequence with low Hamming correlation is conducive to improving the performance of the communication system.

[0003] The construction of frequency hopping sequences can generally be divided into two categories: mathematical construction and machine learning generation. In 1974, Lempel and Greenberger proposed a theoretical boundary for frequency hopping sequence design, the Lempel boundary, which defined the minimum Hamming correlation that can be achieved in an ideal case for frequency hopping sequence design, and provided a theoretical reference for the evaluation of frequency hopping sequence performance. The construction of the optimal frequency hopping sequence that meets the Lempel boundary generally uses methods such as algebra, combinatorial mathematics, and recursive construction. Although the traditional mathematical construction method has certain advantages in theory, it is often constrained by parameter restrictions when facing complex sequence structures, which limits its flexibility and effect in some practical applications. In contrast, the frequency hopping sequence generation technology based on machine learning shows more powerful advantages. With the flexibility and powerful computing power of intelligent algorithms, it can not only deeply explore a richer and more diverse sequence parameter space, but also significantly improve the anti-interference performance of the sequence, thereby better meeting the needs of user communication systems in complex environments. However, the current method of generating the optimal frequency hopping sequence based on intelligent algorithms still has the problem of insufficient dynamic adjustment and optimization efficiency. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides an optimal frequency hopping sequence generation method based on deep reinforcement learning, which solves the problem that the existing frequency hopping sequence generation method has insufficient ability to dynamically adjust and improve the generated sequence.

[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0006] The present invention provides an optimal frequency hopping sequence generation method based on deep reinforcement learning, comprising the following steps:

[0007] S1. Obtain the frequency gap q, sequence length N and the step length of each filling of the target generated frequency hopping sequence , where the step length is the prime factor of the sequence length N;

[0008] S2, generate the frequency gap q of the frequency hopping sequence according to the target, the sequence length N and the step length of each filling of the sequence , using deep neural network DNN to guide Monte Carlo tree to fill in the sequence and generate frequency hopping sequence results;

[0009] S3, repeat S2 for a preset number of times, and construct a frequency hopping sequence training set based on the generated frequency hopping sequence results;

[0010] S4, using the frequency hopping sequence training set to train and update the deep neural network DNN to obtain a trained deep neural network DNN;

[0011] S5, generate the frequency gap q of the frequency hopping sequence, the sequence length N and the step length of each filling of the sequence according to the target , using the Monte Carlo tree and the trained deep neural network to repeat S2-S4 for sequence expansion and filling until the optimal frequency hopping sequence is obtained.

[0012] The beneficial effects of the present invention are as follows: the present invention provides an optimal frequency hopping sequence generation method based on deep reinforcement learning, which converts the frequency hopping sequence generation process into a sequence filling process, and constructs a Monte Carlo tree corresponding to the sequence filling process, implements a search algorithm based on Monte Carlo tree simulation, finds a better sequence filling decision path, and by observing the filling state of the sequence at each time step and using a deep neural network DNN, it is possible to learn the optimal sequence filling action under the corresponding sequence filling state in a complex high-dimensional environment; in the present invention, the deep neural network DNN is used to guide the evaluation of the expansion and simulation of the Monte Carlo tree, and the deep neural network DNN is trained by using a frequency hopping sequence training set composed of a large number of frequency hopping sequence results generated by the Monte Carlo tree, so as to achieve mutual promotion, enhance the decision-making ability of the entire frequency hopping sequence generation, and more efficiently and accurately find the optimal q-ary frequency hopping sequence that meets the Lempel bound.

[0013] Furthermore, the S2 comprises the following steps:

[0014] S201, taking the initial state of the current sequence of length N as the root node of the Monte Carlo tree, wherein the value of the unfilled sequence position is set to 0, and the value of the filled sequence position is an integer in the range of [1, q];

[0015] The calculation expression of the initial state of the current sequence is as follows:

[0016] ,

[0017] in, Indicates the initial state of the current sequence, represents the filling state of the ith sequence, where , and i is a natural number, when the sequence is filled to the When the sequence is in the filling state, the sequence filling is completed;

[0018] S202, determine whether the current sequence is filled, if so, proceed to S211, otherwise proceed to S203;

[0019] S203, according to the Monte Carlo tree method, starting from the root node, respectively through There are paths that can reach each vertex, and each path stores the corresponding number of visits , average reward and prior probability ,in, Indicates the current sequence status, represents the jth filling action of the current sequence, ;

[0020] S204: According to the number of visits corresponding to each path , average reward and prior probability , based on the upper confidence interval UCB model, the optimal path is selected as the current node, and it is determined whether the current node is a leaf node. If so, it goes to S205, otherwise it returns to S204, wherein the path corresponding to the highest UCB value is the optimal path;

[0021] S205, determine whether the current node has been visited, if yes, proceed to S206, otherwise, take the current node as an extended node and proceed to S207;

[0022] S206, expand the current node, and obtain the probability distribution and value evaluation of the current node after performing feature extraction based on the deep neural network DNN, and select the expansion node based on the probability distribution of the current node, and enter S207;

[0023] S207, starting from the expansion node, randomly selecting a filling action to simulate the process from sequence filling to sequence filling completion in all cases, and obtaining a complete simulation sequence result set and a corresponding complete branch simulation result;

[0024] S208. According to the complete simulation sequence result set, the sub-nodes of the complete branch simulation result are transmitted back to the root node layer by layer, and the node access times and corresponding average reward values ​​of all paths on the complete branch simulation result are updated;

[0025] S209, according to the node access times of all paths in the complete branch simulation results, the node movement selection probability is calculated based on the node selection model;

[0026] S210, selecting a filling action in the current sequence filling state according to the node movement selection probability, obtaining the next sequence filling state, and taking the next sequence filling state as the initial state of the current sequence, and returning to S202;

[0027] S211. The current sequence, the probability of the current sequence, and the reward value corresponding to the current sequence are used as a frequency hopping sequence result.

[0028] The beneficial effects of adopting the above further scheme are as follows: the present invention constructs a Monte Carlo tree corresponding to the sequence state and the path of the sequence filling extension, realizes the corresponding connection between the sequence filling state and the sequence state, and provides a basis for simulating the filling of all possible complete q-ary frequency hopping sequences; based on the number of storage accesses, the average reward and the prior probability, the present invention realizes the selection of the optimal path as the current node through the upper confidence interval UCB model, and combines the probability distribution and value evaluation obtained by feature extraction with the deep neural network DNN, which provides a basis for confirming the optimal frequency hopping sequence element; the present invention calculates the node movement selection probability through the complete simulation filling and back transmission of the Monte Carlo tree, realizes the step-by-step filling to obtain a complete q-ary frequency hopping sequence that meets the optimal Lempel bound, and provides a basis for feedback training of the deep neural network DNN.

[0029] Furthermore, the calculation expression of the upper confidence interval UCB model in S204 is as follows:

[0030] ,

[0031] in, Indicates the highest UCB value corresponding to the Path, represents the j value that gives the highest UCB value, and c represents the path exploration weighting factor.

[0032] The beneficial effect of adopting the above further scheme is: the present invention provides a calculation method for the upper confidence interval UCB model, provides a basis for selecting the optimal path from the Monte Carlo tree root node to each node, and also provides a basis for further optimal selection of node expansion.

[0033] Furthermore, the S206 includes the following steps:

[0034] S2061, obtaining a sequence image P corresponding to the current node, wherein the size of the sequence image P is 1*N, and the sequence image P corresponding to the current node is an image composed of a sequence corresponding to the sequence filling state corresponding to the current node, and the N sequence elements in the image each occupy a unit length, and the total length is N;

[0035] S2062, expand the sequence image P into q+1 1*N subsequence images, where each subsequence image corresponds to a subsequence. is a positive number and At that time, Subsequence images Only with sequence element values ​​in sequence image P The corresponding sequence positions of remain the same, the values ​​of the sequence elements of the other sequence positions are all 0, and the q+1th subsequence image The values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are 0 are set to 1, and the values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are not 0 are set to 0;

[0036] S2063, input the q+1 1*N subsequence images into the deep neural network DNN to extract spatial features, and obtain the probability distribution and value evaluation corresponding to the sequence image P as the probability distribution and reward value of the current node;

[0037] S2064. According to the probability distribution of the current node, a path corresponding to the subsequence image with the largest probability is selected and added to the Monte Carlo tree to obtain an extended node.

[0038] The beneficial effect of adopting the above further scheme is: the present invention constructs subsequence images and uses the deep neural network DNN to extract features, predict probabilities and rewards for each subsequence image, thereby realizing the selection of target expansion nodes, and providing a guarantee for gradually filling in a complete q-ary frequency hopping sequence that meets the optimal Lempel bound.

[0039] Furthermore, the calculation expression of the node selection model in S209 is as follows:

[0040] ,

[0041] in, represents the node movement selection probability, represents the softmax function, Indicates the node selection scaling factor, Indicates the number of visits to the root node.

[0042] The beneficial effect of adopting the above further scheme is: the present invention provides a calculation method for the node selection model, which can select the next state of sequence filling according to the access update that occurs when the complete branch simulation result is passed back to the root node, determine the best choice for filling the sequence elements in the current round, and provide a basis for gradually generating a q-ary frequency hopping sequence that meets the optimal Lempel bound.

[0043] Furthermore, the calculation expression of the reward value in S211 is as follows:

[0044] ,

[0045] ,

[0046] ,

[0047] ,

[0048] ,

[0049] in, represents the reward function value corresponding to the current sequence S, represents the metric function value corresponding to the sequence image S, S represents the current sequence S, represents the average maximum metric threshold obtained through noise-free experiments, lempel represents the minimum metric threshold, Indicates that the current sequence S is delayed The maximum periodic Hamming autocorrelation function of Indicates the delay Take the maximum value in the range greater than or equal to 1 and less than or equal to N, r represents the minimum non-negative remainder value of N modulo q, Indicates rounding up. Indicates that the current sequence S is delayed The periodic Hamming autocorrelation function of Indicates the measurement The sequence element and A function of the similarity between sequence elements, where if ,but , if satisfied , then the sequence satisfies the Lempel bound optimality.

[0050] The beneficial effect of adopting the above further scheme is: the present invention provides a method for calculating the reward value, and sets the metric function of the target generated frequency hopping sequence to the maximum period Hamming autocorrelation value of the sequence, that is, the maximum number of collisions of the sequence under different time delays, providing a basis for generating a q-ary frequency hopping sequence that satisfies the optimal Lempel bound.

[0051] Other advantages of the present invention will be analyzed in more detail in subsequent embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0053] Figure 1 This is a flowchart of the steps of a method for generating an optimal frequency hopping sequence based on deep reinforcement learning in Example 1 of the present invention. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.

[0055] Embodiment 1:

[0056] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides an optimal frequency hopping sequence generation method based on deep reinforcement learning, comprising the following steps:

[0057] S1. Obtain the frequency gap q, sequence length N and the step length of each filling of the target generated frequency hopping sequence , where the step length is the prime factor of the sequence length N;

[0058] In this scheme, the generation process of the target generation sequence can be regarded as the filling process of q-ary sequence elements of length N, where unfilled values ​​are set to 0 and the value range of filled values ​​is [1, q].

[0059] S1.5. Initialize the learning parameters of the deep neural network DNN based on the frequency gap q of the target generated sequence, the sequence length N, and the step length of each sequence filling , perform a noise-free frequency hopping sequence generation experiment for a preset number of experiments, and calculate the average metric value of the experimentally generated sequence as the average maximum metric threshold.

[0060] In this embodiment, the preset number of experiments is 50 times. In this scheme, when the noiseless frequency hopping sequence generation experiment generates a frequency hopping sequence, the S2 method is used to expand the sequence elements one by one to obtain the experimental generation sequence. Each noiseless frequency hopping sequence generation experiment generates a frequency hopping sequence. After 50 frequency hopping sequences are completely generated, the average maximum metric threshold obtained provides a basis for calculating the reward value corresponding to the sequence.

[0061] S2, generate the frequency gap q of the frequency hopping sequence according to the target, the sequence length N and the step length of each filling of the sequence , using deep neural network DNN to guide Monte Carlo tree to fill in the sequence and generate frequency hopping sequence results;

[0062] The S2 comprises the following steps:

[0063] S201, taking the initial state of the current sequence of length N as the root node of the Monte Carlo tree, wherein the value of the unfilled sequence position is set to 0, and the value of the filled sequence position is an integer in the range of [1, q];

[0064] The calculation expression of the initial state of the current sequence is as follows:

[0065] ,

[0066] in, Indicates the initial state of the current sequence, represents the filling state of the ith sequence, where , and i is a natural number, when the sequence is filled to the When the sequence is in the filling state, the sequence filling is completed;

[0067] S202, determine whether the current sequence is filled, if so, proceed to S211, otherwise proceed to S203;

[0068] S203, according to the Monte Carlo tree method, starting from the root node, respectively through There are paths that can reach each vertex, and each path stores the corresponding number of visits , average reward and prior probability ,in, Indicates the current sequence status, represents the jth filling action of the current sequence, ; In this scheme, the state of the sequence corresponds to the filled part of the sequence from the root node to the current node, and the filling action of the sequence is the path selection corresponding to the node, that is, the integer value that can be filled in the corresponding sequence, and the range of the integer value that can be filled is [1, q];

[0069] S204: According to the number of visits corresponding to each path , average reward and prior probability , based on the upper confidence interval UCB model, the optimal path is selected as the current node, and it is determined whether the current node is a leaf node. If so, it goes to S205, otherwise it returns to S204, wherein the path corresponding to the highest UCB value is the optimal path;

[0070] The calculation expression of the upper confidence interval UCB model in S204 is as follows:

[0071] ,

[0072] in, Indicates the highest UCB value corresponding to the Path, represents the j value corresponding to the maximum UCB value, and c represents the path exploration weighting factor. In this embodiment, c is a constant.

[0073] S205, determine whether the current node has been visited, if yes, proceed to S206, otherwise, take the current node as an extended node and proceed to S207;

[0074] S206, expand the current node, and obtain the probability distribution and value evaluation of the current node after performing feature extraction based on the deep neural network DNN, and select the expansion node based on the probability distribution of the current node, and enter S207;

[0075] The S206 comprises the following steps:

[0076] S2061, obtaining a sequence image P corresponding to the current node, wherein the size of the sequence image P is 1*N, and the sequence image P corresponding to the current node is an image composed of a sequence corresponding to the sequence filling state corresponding to the current node, and the N sequence elements in the image each occupy a unit length, and the total length is N;

[0077] S2062, expand the sequence image P into q+1 1*N subsequence images, where each subsequence image corresponds to a subsequence. is a positive number and At that time, Subsequence images Only with sequence element values ​​in sequence image P The corresponding sequence positions of remain the same, the values ​​of the sequence elements of the other sequence positions are all 0, and the q+1th subsequence image The values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are 0 are set to 1, and the values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are not 0 are set to 0;

[0078] S2063, input the q+1 1*N subsequence images into the deep neural network DNN for spatial feature extraction, and obtain the probability distribution and value evaluation corresponding to the sequence image P as the probability distribution and reward value of the current node; in this solution, the value range of the reward value corresponding to the sequence image P is [-1,1];

[0079] S2064. According to the probability distribution of the current node, a path corresponding to the subsequence image with the largest probability is selected and added to the Monte Carlo tree to obtain an extended node.

[0080] S207, starting from the expansion node, randomly selecting a filling action to simulate the process from sequence filling to sequence filling completion in all cases, and obtaining a complete simulation sequence result set and a corresponding complete branch simulation result;

[0081] S208. According to the complete simulation sequence result set, the sub-nodes of the complete branch simulation result are transmitted back to the root node layer by layer, and the node access times and corresponding average reward values ​​of all paths on the complete branch simulation result are updated;

[0082] S209, according to the node access times of all paths in the complete branch simulation results, the node movement selection probability is calculated based on the node selection model;

[0083] The calculation expression of the node selection model in S209 is as follows:

[0084] ,

[0085] in, represents the node movement selection probability, represents the softmax function, Indicates the node selection scaling factor, Indicates the number of visits to the root node.

[0086] S210, selecting a filling action in the current sequence filling state according to the node movement selection probability, obtaining the next sequence filling state, and taking the next sequence filling state as the initial state of the current sequence, and returning to S202;

[0087] S211. The current sequence, the probability of the current sequence, and the reward value corresponding to the current sequence are used as a frequency hopping sequence result.

[0088] The calculation expression of the reward value in S211 is as follows:

[0089] ,

[0090] ,

[0091] ,

[0092] ,

[0093] ,

[0094] in, represents the reward function value corresponding to the current sequence S, represents the metric function value corresponding to the sequence image S, S represents the current sequence S, represents the average maximum metric threshold obtained through noise-free experiments, lempel represents the minimum metric threshold, Indicates that the current sequence S is delayed The maximum periodic Hamming autocorrelation function of Indicates the delay Take the maximum value in the range greater than or equal to 1 and less than or equal to N, r represents the minimum non-negative remainder value of N modulo q, Indicates rounding up. Indicates that the current sequence S is delayed The periodic Hamming autocorrelation function of Indicates the measurement The sequence element and A function of the similarity between sequence elements, where if ,but , if satisfied , then the sequence satisfies the Lempel bound optimality.

[0095] S3, repeat S2 for a preset number of times, and construct a frequency hopping sequence training set based on the generated frequency hopping sequence results;

[0096] S4, using the frequency hopping sequence training set to train and update the deep neural network DNN to obtain a trained deep neural network DNN;

[0097] S5, generate the frequency gap q of the frequency hopping sequence, the sequence length N and the step length of each filling of the sequence according to the target , using the Monte Carlo tree and the trained deep neural network to repeat S2-S4 for sequence expansion and filling until the optimal frequency hopping sequence is obtained.

[0098] Embodiment 2:

[0099] On the basis of Embodiment 1, in a practical example of the present invention, taking an experiment of finding an optimal frequency hopping sequence with a frequency gap of 3 and a length of less than 30 as an example, the parameter settings are shown in Table 1:

[0100]

[0101] In this experiment, the input parameters are the frequency gap of the sequence q=3, the length of the sequence to be found is q+1 to 30, and the step length of each sequence filling is is a prime factor of N. After 50 rounds of noise-free experiments, the average metric value E[M] of the sequence is calculated, and the metric function value and reward value corresponding to the sequence are calculated. In MCTS, each sequence is filled with state s i Set the search tree root node and perform m = 900 forward simulations. In each simulation, add Dirichlet noise α = 0.05 to the prior probability of the root node to introduce exploration. After 900 simulations, the node selection probability is calculated by , where τ=1 is set in the first one-third of the time steps, and the probability of selecting a move is proportional to the number of visits; τ=10 is set in the other time steps -4 , deterministically select the one with the most visits to move. In the deep neural network DNN, the update period G is set to 50, that is, 6 small batches of data are extracted after every 50 experiments to train the deep neural network DNN. In this embodiment, the small batch of data is set to 64.

[0102] The scheme of the present invention adopts the sum of minimizing the mean square error and the cross entropy loss as the loss function, trains the deep neural network DNN when q=3, N=17, and records the loss values ​​corresponding to different training steps. The loss value drops significantly in the early stage of training, indicating that the model quickly learns the basic features. During the training process, the loss value gradually stabilizes, showing the convergence of the model. The downward trend of the loss value with the increase of the training steps verifies the effectiveness of the proposed optimization algorithm and model structure, which can significantly reduce the training error and improve the model performance. In addition, the slight fluctuation of the loss value with the increase of the training steps is caused by the optimization using the stochastic gradient descent method, which is a normal phenomenon and does not affect the final convergence effect of the model.

[0103] In this embodiment, the scheme of the present invention is compared with the existing frequency hopping sequence generation scheme. The present invention can use machine learning to generate frequency hopping sequences, and can find more sequence parameters than the traditional sequence construction scheme. Taking the search for the optimal frequency hopping sequence with a frequency gap of 3 and a sequence length of 30 as an example, the scheme of the present invention can find the optimal frequency hopping sequence with a sequence length of 27 and a sequence length of 30, while the parameters of the existing frequency hopping sequence generation scheme cannot be generated, wherein the existing construction scheme includes: Scheme 1: "Construction of Optimal Uniform Wideband Gap Frequency Hopping Sequence", published in "IEEE Journal of Information Theory", Vol. 68, No. 1, pp. 692-700, January 2022; Scheme 2: “A New Optimal Frequency Hopping Sequence Based on Circle Cutting”, 13th International Computer Conference on Wavelet Active Media Technology and Information Processing, 2016; Scheme 3: “A New Family of Composite-Length Optimal Frequency-Hopping Sequences,” IEEE Transactions on Information Theory, pp. 3688-3697, 2014. Scheme 4: “A New Construction of Optimal Sets of Frequency Hopping Sequences”, IEEE Transactions on Information Theory, pp. 3831-3840, 2011; Scheme 5: “On Sidnikov Sequences as Frequency Hopping Sequences”, IEEE Transactions on Information Theory, pp. 4279-4285, 2009; Scheme 6: "Further Combinatorial Construction of Optimal Frequency-Hopping Sequences", Journal of Combinatorial Theory, Series A, pp. 1699-1718, 2006.

[0104] The comparative construction results of the scheme of the present invention and comparative schemes 1 to 6 are shown in Table 2-1, Table 2-2, Table 2-3 and Table 2-4:

[0105]

[0106]

[0107]

[0108]

[0109] According to Table 2-1, Table 2-2, Table 2-3 and Table 2-4, it can be found that with the changes in sequence length and frequency gap, under most different sequence lengths and different conditions satisfying the optimal condition of the Lempel bound, the scheme of the present invention and comparative schemes 1-6 can generate the optimal frequency hopping sequence, but when the sequence length is 27 and the frequency gap is 3, and when the sequence length is 30 and the frequency gap is 3, only the scheme of the present invention can generate the optimal frequency hopping sequence, and comparative schemes 1-6 cannot generate the optimal frequency hopping sequence, which shows that compared with the sum of the existing comparative schemes, the scheme of the present invention still has the ability to find more sequence parameters.

[0110] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for generating an optimal frequency hopping sequence based on deep reinforcement learning, characterized in that: The steps include: S1. Obtain the frequency gap q, sequence length N and the step length of each filling of the target generated frequency hopping sequence , where the step length is the prime factor of the sequence length N; S2, generate the frequency gap q of the frequency hopping sequence according to the target, the sequence length N and the step length of each filling of the sequence , using deep neural network DNN to guide Monte Carlo tree to fill in the sequence and generate frequency hopping sequence results; S3, repeat S2 for a preset number of times, and construct a frequency hopping sequence training set based on the generated frequency hopping sequence results; S4, using the frequency hopping sequence training set to train and update the deep neural network DNN to obtain a trained deep neural network DNN; S5, generate the frequency gap q of the frequency hopping sequence, the sequence length N and the step length of each filling of the sequence according to the target , using the Monte Carlo tree and the trained deep neural network to repeat S2-S4 for sequence expansion and filling until the optimal frequency hopping sequence is obtained.

2. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 1, characterized in that: The S2 comprises the following steps: S201, taking the initial state of the current sequence of length N as the root node of the Monte Carlo tree, wherein the value of the unfilled sequence position is set to 0, and the value of the filled sequence position is an integer in the range of [1, q]; The calculation expression of the initial state of the current sequence is as follows: , in, Indicates the initial state of the current sequence, represents the filling state of the ith sequence, where , and i is a natural number, when the sequence is filled to the When the sequence is in the filling state, the sequence filling is completed; S202, determine whether the current sequence is filled, if so, proceed to S211, otherwise proceed to S203; S203, according to the Monte Carlo tree method, starting from the root node, respectively through There are paths that can reach each vertex, and each path stores the corresponding number of visits , average reward and prior probability ,in, Indicates the current sequence status, represents the jth filling action of the current sequence, ; S204: According to the number of visits corresponding to each path , average reward and prior probability , based on the upper confidence interval UCB model, the optimal path is selected as the current node, and it is determined whether the current node is a leaf node. If so, it goes to S205, otherwise it returns to S204, wherein the path corresponding to the highest UCB value is the optimal path; S205, determine whether the current node has been visited, if yes, proceed to S206, otherwise, take the current node as an extended node and proceed to S207; S206, expand the current node, and obtain the probability distribution and value evaluation of the current node after performing feature extraction based on the deep neural network DNN, and select the expansion node based on the probability distribution of the current node, and enter S207; S207, starting from the expansion node, randomly selecting a filling action to simulate the process from sequence filling to sequence filling completion in all cases, and obtaining a complete simulation sequence result set and a corresponding complete branch simulation result; S208. According to the complete simulation sequence result set, the sub-nodes of the complete branch simulation result are transmitted back to the root node layer by layer, and the node visit counts and corresponding average reward values ​​of all paths on the complete branch simulation result are updated; S209, according to the node access times of all paths in the complete branch simulation results, the node movement selection probability is calculated based on the node selection model; S210, selecting a filling action in the current sequence filling state according to the node movement selection probability, obtaining the next sequence filling state, and taking the next sequence filling state as the initial state of the current sequence, and returning to S202; S211. The current sequence, the probability of the current sequence, and the reward value corresponding to the current sequence are used as a frequency hopping sequence result.

3. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 2, characterized in that: The calculation expression of the upper confidence interval UCB model in S204 is as follows: , in, Indicates the highest UCB value corresponding to the Path, represents the j value that gives the highest UCB value, and c represents the path exploration weighting factor.

4. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 2, characterized in that: The S206 comprises the following steps: S2061, obtaining a sequence image P corresponding to the current node, wherein the size of the sequence image P is 1*N, and the sequence image P corresponding to the current node is an image composed of a sequence corresponding to the sequence filling state corresponding to the current node, and the N sequence elements in the image each occupy a unit length, and the total length is N; S2062, expand the sequence image P into q+1 1*N subsequence images, where each subsequence image corresponds to a subsequence. is a positive number and At that time, Subsequence images Only with sequence element values ​​in sequence image P The corresponding sequence positions of remain the same, the values ​​of the sequence elements of the other sequence positions are all 0, and the q+1th subsequence image The values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are 0 are set to 1, and the values ​​at the sequence positions where the sequence elements in the corresponding sequence image P are not 0 are set to 0; S2063, input the q+1 1*N subsequence images into the deep neural network DNN to extract spatial features, and obtain the probability distribution and value evaluation corresponding to the sequence image P as the probability distribution and reward value of the current node; S2064. According to the probability distribution of the current node, a path corresponding to the subsequence image with the largest probability is selected and added to the Monte Carlo tree to obtain an extended node.

5. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 2, characterized in that: The calculation expression of the node selection model in S209 is as follows: , in, represents the node movement selection probability, represents the softmax function, Indicates the node selection scaling factor, Indicates the number of visits to the root node.

6. The optimal frequency hopping sequence generation method based on deep reinforcement learning according to claim 2, characterized in that: The calculation expression of the reward value in S211 is as follows: , , , , , in, represents the reward function value corresponding to the current sequence S, represents the metric function value corresponding to the sequence image S, S represents the current sequence S, represents the average maximum metric threshold obtained through noise-free experiments, lempel represents the minimum metric threshold, Indicates that the current sequence S is delayed The maximum periodic Hamming autocorrelation function of Indicates the delay Take the maximum value in the range greater than or equal to 1 and less than or equal to N, r represents the minimum non-negative remainder value of N modulo q, Indicates rounding up. Indicates that the current sequence S is delayed The periodic Hamming autocorrelation function of Indicates the measurement The sequence element and A function of the similarity between sequence elements, where if ,but , if satisfied , then the sequence satisfies the Lempel bound optimality.

Citation Information

Patent Citations

  • Broadband spectrum sensing method based on reinforcement learning

    CN112202514A

  • Construction method of frequency hopping sequence set without collision zone

    CN115833871A

  • Block agile frequency hopping method and system

    CN116366093A

  • Enhanced frequency hopping for data transmissions

    WO2022192444A1