Multi-platform formation frequency equipment spectrum allocation method based on dual-frequency backtracking DQN algorithm

By using the spectrum allocation method of the dual-frequency backtracking DQN algorithm in multi-platform formations, the electromagnetic interference problem caused by limited spectrum resources is solved, and efficient utilization of spectrum resources and optimization of system performance is achieved.

CN120165795APending Publication Date: 2025-06-17CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312549.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In multi-platform formations, limited spectrum resources lead to serious electromagnetic interference problems between frequency-using devices. The traditional static spectrum assignment method cannot adapt to dynamically changing task requirements, resulting in low spectrum resource utilization.

Method used

The spectrum allocation method of multi-platform frequency equipment based on dual-frequency backtracking DQN algorithm is adopted. Through greedy strategies and deep learning technology, spectrum resources are dynamically allocated, and the interference margin and priority of the equipment are comprehensively considered, and the spectrum allocation results are optimized.

Benefits of technology

It significantly improves the utilization efficiency of spectrum resources, reduces the probability of frequency conflict between devices, optimizes the overall performance of the system, and ensures the electromagnetic compatibility status of high-priority equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165795A_ABST
    Figure CN120165795A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-platform formation frequency device spectrum allocation method based on a dual-frequency backtracking DQN algorithm. The method comprises the following steps: acquiring working parameters of all frequency devices in a formation; adopting a greedy strategy to carry out working frequency point distribution on frequency equipment; when the frequency points are distributed, the interference allowance and the priority of the frequency equipment are comprehensively considered, quantitative evaluation is carried out on the distribution result, and a corresponding reward value is given accordingly; after working frequency points are distributed based on a greedy strategy every time, states before distribution, actions, reward values and states after distribution are stored in an experience pool in a tetrad mode; when samples accumulated in the experience pool reach a set amount, certain samples are extracted from the experience pool to train a DQN network, and a DQN spectrum allocation model is obtained; and applying the DQN spectrum allocation model to an actual formation for guiding spectrum allocation of frequency-using equipment. According to the invention, spectrum resources can be allocated to all devices as many as possible, and reasonable allocation and maximum utilization of the spectrum resources are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the spectrum allocation of frequency-using devices, and particularly to a spectrum allocation method for multi-platform formation frequency-using devices based on a dual-frequency backtracking DQN algorithm, aiming to effectively solve the electromagnetic interference problem among frequency-using devices within the formation by effectively and reasonably allocating electromagnetic resources, and belongs to the field of electromagnetic compatibility technology. Background Art

[0002] With the rapid development of radio technology, the number of various wireless devices, including wireless communication, satellite communication, broadcasting, radar, etc., has increased sharply, and the demand for electromagnetic spectrum resources has also grown rapidly. However, there is an obvious imbalance between this demand growth and the limited available frequency-domain electromagnetic resources. Since the available electromagnetic spectrum resources are relatively limited, it has led to the problem of spectrum congestion. The spectrum congestion problem is manifested as numerous frequency-using devices having to operate together on the limited frequency-domain electromagnetic resources, which may cause mutual interference between frequency-using devices, thereby reducing signal quality, hindering information transmission, affecting the performance of frequency-using devices, and a series of other problems. Especially in multi-platform formations, such as multiple devices carried on platforms such as airplanes, vehicles, and ships need to work together to complete complex tasks, the spectrum congestion problem is particularly prominent. When these devices use the same frequency band resources simultaneously, it makes it impossible to effectively avoid interference between devices by simply allocating different operating frequencies. Mutual interference will not only disrupt the stable operation of the formation, may cause difficulties in task execution, and in severe cases, may even cause the entire formation to be paralyzed.

[0003] Traditional spectrum resource allocation methods mainly adopt spectrum static assignment, that is, a single spectrum resource is fixedly assigned to a designated frequency-using device for use. Although this method can effectively avoid interference from other current frequency-using devices, its disadvantages are also obvious. On the one hand, it results in low utilization rate of frequency-domain electromagnetic resources, and a large amount of available spectrum resources are idle or wasted; on the other hand, once the spectrum resources are allocated, it is difficult to make flexible adjustments according to actual needs. In complex task scenarios, the working states and requirements of frequency-using devices on the platform are often dynamically changing, and the traditional spectrum static assignment method obviously cannot adapt to this change and is difficult to meet the complex and changeable electromagnetic environment and task requirements. At the same time, considering that some devices in engineering practice often have anti-interference capabilities, the interference margin threshold can be increased by moderately occupying the time-frequency working resources of the devices. Therefore, how to reasonably allocate and utilize limited spectrum resources to ensure the effective operation of multi-platform formation frequency-using devices, enhance anti-interference capabilities, and improve overall comprehensive capabilities has become a key problem to be solved urgently. Summary of the Invention

[0004] In view of the above deficiencies in the prior art, the object of the present invention is to propose a spectrum allocation method for multi-platform formation frequency-using devices based on a dual-frequency backtracking DQN algorithm. While considering the anti-interference ability of the devices and ensuring the spectrum resource requirements of key devices, the present invention allocates as much spectrum resource as possible to all devices, achieving reasonable allocation and maximized utilization of spectrum resources.

[0005] The technical solution of the present invention is implemented as follows:

[0006] A spectrum allocation method for multi-platform formation frequency-using devices based on a dual-frequency backtracking DQN algorithm, comprising the following steps:

[0007] S1. Obtain the working parameters of all frequency-using devices within the multi-platform formation;

[0008] S2. In the first stage, use a greedy strategy to allocate working frequency points to the frequency-using devices within the formation; when allocating frequency points, comprehensively consider the interference margin and priority of the frequency-using devices, quantitatively evaluate the allocation results, and give corresponding reward values accordingly;

[0009] S3. Then, after each allocation of working frequency points based on the greedy strategy, store the state before allocation, the allocation action, the reward value corresponding to the allocation action, and the state after allocation in the experience pool in the form of a quadruple;

[0010] S4. When the samples accumulated in the experience pool reach the set capacity, extract a certain number of samples from the experience pool to train the initialized DQN network, update the DQN network parameters until the set conditions are met, and the training ends to obtain the DQN spectrum allocation model;

[0011] S5. In the second stage, when allocating working frequency points to the frequency-using devices whose working frequency points need to be allocated within the formation, first input the working parameters of the current state s1 of all frequency-using devices within the formation into the DQN spectrum allocation model, and then the DQN spectrum allocation model automatically outputs the new state s2 of all frequency-using devices within the formation. The frequency points of the frequency-using devices whose working frequency points do not need to be allocated in the new state s2 remain unchanged, and the frequency points of the frequency-using devices whose working frequency points need to be allocated will be updated, and the updated frequency points are the allocated frequency points.

[0012] Further, in step S1, the working parameters include the working frequency point, bandwidth, transmit power, position, interference margin, and device priority of the frequency-using device;

[0013] First, use the set ME to represent the total number of frequency-using devices carried by each member in the formation, as shown in Equation (1).

[0014]

[0015] Among them, N Mis the number of formation members, then me i is the number carried by the frequency-using device of formation member i;

[0016] Use the set EP to represent the set of all frequency-using devices carried by all formation members, as shown in Equation (2);

[0017]

[0018]

[0019] In Equation (3), EP i represents the set of all frequency-using devices carried by member i in the formation; ep ij is the device number, representing the j-th device carried by member i;

[0020] Secondly, define the positions of all members within the formation and the transmission power, bandwidth, and interference margin thresholds for all frequency-using devices; represent them using the sets MP, PT, BW, and SF respectively, as shown in Equations (4) to (7).

[0021]

[0022]

[0023]

[0024]

[0025] In the formula, mp i is the position of member i, pt i , bw i , sf i represent the transmission power, device bandwidth, and interference margin threshold of device i respectively;

[0026] Then, use the set AR to represent the working frequency point result matrix assigned to all frequency-using devices within the formation, as shown in Equation (8);

[0027] AR = {ar1,…,ar i ,…,ar M} (8)

[0028] In the formula, M is the total number of frequency-using devices; ar i represents the working frequency point of frequency-using device i;

[0029] Use the set IP to represent the priority of the device, as shown in Equation (9);

[0030]

[0031] In the formula, For device ep ij The allocation order, the smaller the number, the higher the device allocation priority;

[0032] Finally, use AR m 、MP m 、PT m 、BW m 、SF m and IP m to represent the working frequency points, platform positions, transmission powers, bandwidths, interference margin thresholds, and device priorities of all devices in the formation under state s m as shown in Equation (10);

[0033] s m ={AR m ,MP m ,PT m ,BW m ,SF m ,IP m} (10).

[0034] Furthermore, in step S2, the specific method for quantitatively evaluating the allocation result and giving corresponding reward values accordingly is,

[0035] Define the interference margin threshold of the corrected device i as the corrected interference margin threshold, denoted as sff i ; Assume the device to be allocated is the receiving device, and calculate the interference margin im ij of the remaining devices on the device to be allocated and the corresponding reward r ij in turn according to Equation (12), and then obtain the current decision action reward value r ij by taking the mean of r i through Equation (13);

[0036]

[0037]

[0038] In the formula, M is the total number of devices, ip i is the priority of device i; sf i is the interference margin threshold of device i; k is a fixed coefficient used to amplify the reward value; σ is the standard deviation of the Gaussian function used to change the sensitivity of the interference margin change to the reward.

[0039] Furthermore, in step S3, the specific method for storing the state before allocation, the allocation action, the reward value corresponding to the allocation action, and the state after allocation in the experience pool in the form of a quadruple is,

[0040] If the result matrix of the working frequency points allocated to all frequency-using devices in the formation is AR, then the subset corresponding to the device i to be allocated is AR i , expand its corresponding available frequency point resources into three dimensions, as shown in Equation (14);

[0041] AR i ={ar1, ar2, ar3} (14)

[0042] In the formula, AR i represents the frequency points carried by device i; where ar1 represents the frequency point selected through the current decision-making action; ar2 represents the frequency point selected in the previous state; ar3 represents the frequency point corresponding to the decision-making action with the largest reward value in the set;

[0043] Update AR after each decision-making action i , store the frequency point selected by the decision-making action as ar1, store the frequency point selected by the previous decision-making action as ar2, and store the frequency point corresponding to the largest reward value among the decision-making actions made as ar3; at the same time, store the frequency point selected by the current decision-making action and the corresponding reward value into the set of decision-making actions made;

[0044] After each device to be allocated selects a frequency point, the state of the environment will transfer to a new state; the state s before allocation t , the allocation action a i , the reward value r corresponding to the allocation action ti , and the state s after allocation t+1 , are combined in the form of a quadruple <s t , a i , r ti , s t+1 > and stored as a sample in the experience pool.

[0045] Furthermore, in step S4, initialize the DQN network, including: 1) Set the number of iterations, exploration rate ε, learning rate α, and discount factor γ; where the discount factor γ is used to measure the importance of future rewards; the discount factor γ approaching 1 indicates a higher importance of future rewards, and approaching 0 indicates only emphasizing current rewards; 2) Initialize the experience pool, set the capacity N R and the sample batch_size parameter; 3) Initialize the Q estimation network and the Q target network, set the number of neurons in the hidden layer N HL , network parameters θ, network parameter replication period T c .

[0046] The spectrum allocation of the present invention is divided into two stages. In the first stage, since the model has not been established, the allocation is carried out according to the greedy strategy. The focus of the first stage is to generate a large number of sample data for the experience pool. In the second stage, the allocation is carried out according to the output result of the DQN model. The greedy strategy is used in the first stage to avoid overly random decisions, which may lead to low data quality and make it quite difficult to train the model in the second stage.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] 1. The present invention proposes a method for calculating the decision action reward based on the fuzzy function. This method can evaluate the impact of the device to be allocated within the formation on the overall electromagnetic interference state of the formation when making a decision action, and preferentially ensure the electromagnetic compatibility state of high-priority devices. By comprehensively considering factors such as the interference margin and priority of the device, this impact is quantified into corresponding reward values, significantly improving the reliability of the decision action and ensuring the electromagnetic compatibility state of high-priority devices within the formation.

[0049] 2. The present invention proposes a DQN (Deep Q-Network) algorithm based on dual-frequency backtracking, that is, when making a decision action, not only the current decision action frequency point is referred to, but also the decision action frequency points of the previous round and the decision action frequency point with the historical maximum reward value are retained to optimize the current decision action, thereby providing more optional frequency points for the device, further reducing the probability of frequency usage conflicts between devices. This not only suppresses decision fluctuations and improves the stability of decision actions, but also deeply explores the value of historical data, significantly enhancing the learning and adaptation capabilities of the algorithm. It effectively improves the utilization efficiency of spectrum resources in a dynamic electromagnetic environment and optimizes the overall performance of the system.

[0050] 3. The present invention uses the DQN algorithm for spectrum allocation. Through its powerful deep learning ability and self-adaptability, it can efficiently allocate the spectrum resources of multi-platform frequency-using devices in a dynamic and complex spectrum environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 - Schematic diagram of the dual-frequency backtracking DQN algorithm structure of the present invention.

[0052] Figure 2 - Schematic diagram of the aircraft formation scenario in the embodiment of the present invention.

[0053] Figure 3 - Curve graph showing the change of the decision action reward value based on the fuzzy function of the present invention with the interference margin.

[0054] Figure 4 - Curve graph showing the change of the average reward of the present invention with the number of iterations of the dual-frequency backtracking DQN algorithm. DETAILED DESCRIPTION OF THE INVENTION

[0055] The specific implementation manners of the present invention will be further described in detail below with reference to the accompanying drawings.

[0056] A spectrum allocation method based on a dual-frequency backtracking DQN algorithm proposed by the present invention aims to improve the stability of a single decision-making and manage the overall electromagnetic interference level. First, it is required to obtain the detailed working parameters of all frequency-using devices in the formation, including key information such as the working frequency points, bandwidths, transmission powers, positions, interference margins, and device priorities of the devices; secondly, a preliminary frequency point allocation method using a greedy strategy is adopted to preliminarily allocate the working frequency points of the frequency-using devices in the formation. In this process, the interference margin and priority of the devices are comprehensively considered, and the allocation results are quantitatively evaluated, and corresponding reward values are given accordingly; then, after each allocation, the current state, action, reward value, and updated state are stored in the experience pool in the form of a quadruple. When a certain number of samples have accumulated in the experience pool, a certain number of samples can be drawn from the experience pool to initialize and train the DQN network, and the Q network parameters are updated until the learning process ends; finally, the Q network model obtained by training is applied to the actual formation to guide the spectrum allocation of the frequency-using devices.

[0057] The specific steps are as follows, and reference can also be made to Figure 1 :

[0058] (1) Obtain the working parameters such as the working frequency points of all frequency-using devices in the formation

[0059] The working parameters of the devices mainly include key information such as the working frequency points, bandwidths, transmission powers, positions, interference margins, and device priorities of the devices.

[0060] First, use the set ME to represent the total number of frequency-using devices carried by each member in the formation, as shown in Equation (1).

[0061]

[0062] Among them, N M is the number of formation members, then me i is the number of frequency-using devices carried by formation member i.

[0063] Use the set EP to represent the set of all frequency-using devices carried by all members of the formation, as shown in Equation (2).

[0064]

[0065]

[0066] In Equation (3), EP i represents the set of all frequency-using devices carried by formation member i; ep ijIt is the device number, representing the j-th device carried by member i.

[0067] Secondly, define the positions of all members within the formation and the transmit power, bandwidth, and interference margin for all frequency-using devices. They are represented by sets MP, PT, BW, and SF respectively, as shown in Equations (4) to (7).

[0068]

[0069]

[0070]

[0071]

[0072] In the formula, mp i is the position of member i (km), pt i , bw i , sf i represent the transmit power (dBm), device bandwidth (MHz), and interference margin threshold (dB) of device i respectively.

[0073] Then, use the set AR to represent the working frequency point result matrix allocated to all frequency-using devices within the formation, as shown in Equation (8).

[0074] AR = {ar1,..., ar i ,..., ar M} (8)

[0075] In the formula, M is the total number of frequency-using devices; ar i represents the working frequency point of frequency-using device i.

[0076] Use the set IP to represent the priority of the device, as shown in Equation (9).

[0077]

[0078] In the formula, is the allocation order of device ep ij , and the smaller the number, the higher the priority of the allocated device.

[0079] Finally, use AR m , MP m , PT m , BW m , SF m , and IP m to represent the working frequency points, platform positions, transmit power, bandwidth, interference margin thresholds, and device priorities of all devices within the formation in state s m respectively, as shown in Equation (10).

[0080] s m ={AR m ,MP m ,PT m ,BW m ,SF m ,IP m} (10)

[0081] (2) Initialize DQN network parameters

[0082] 1) Set parameters such as the number of iterations, exploration rate ε, learning rate α, and discount factor γ

[0083] The number of iterations is the total number of cycles during the training process. Each iteration represents a training cycle, including sampling from the experience pool, calculating losses, updating network parameters, etc. According to experience, the value can be 5000 to 100000, and the specific value is determined by the complexity of the problem and the training time.

[0084] The exploration rate ε represents the probability of exploring new actions. The initial value is high, and it will gradually decrease with training, promoting the transition from exploration to experience utilization. According to experience, the value of ε gradually decreases from 1 to 0.1 or 0.01 as the number of iterations increases.

[0085] The learning rate α is used to control the step size of each parameter update. A learning rate that is too high will lead to unstable training, while a learning rate that is too low will slow down the training speed. According to experience, the value range of α is 0.01 to 0.0001.

[0086] The discount factor γ is used to measure the importance of future rewards. A value close to 1 indicates that future rewards are more important, and a value close to 0 indicates that only current rewards are valued. Based on experience, the value range of γ is 0.99 to 0.8.

[0087] 2) Initialize the experience pool and set the capacity N R and sample batch_size parameter

[0088] The capacity of the experience pool N R Determines how many samples can be stored. A larger capacity value can store more samples, which helps the diversity and stability of the training process, but also increases memory requirements. It is generally set between 10000 and 100000, and its specific value can be adjusted according to computer resources and memory size.

[0089] The size of the sample batch_size is the number of samples drawn from the experience pool each time, which is used to update the Q network parameters. Larger sample sizes help stabilize the gradient estimate, and based on experience, the value range is generally between 32 and 256.

[0090] (3) Initialize the Q estimation network and the Q target network, and set the number of neurons N in the hidden layer HL , network parameters as θ, network parameter replication period T c and other parameters

[0091] The number of neurons N in the hidden layer HL mainly affects the expression ability and training effect of the Q network. For the number of hidden layers, usually 1 to 3 layers are selected, and the number of neurons in each layer can be selected from 50 to 500.

[0092] The network parameter replication period T c is the frequency of copying the parameters of the Q estimation network to the Q target network. A higher replication frequency helps to stabilize the training process. Generally, according to experience, it is copied once every 500 to 5000 times.

[0093] (3) The device to be allocated makes a decision action a according to the greedy policy i , and performs double-frequency backtracking according to the historical decision actions and the corresponding reward values

[0094] (1) The device to be allocated i observes the current system state s m , and selects an action frequency point a according to the greedy policy among all available working frequency points A i , and defines A i as its action space, as shown in Equation (11). i

[0095] A i = [a1,..., a j ,..., a n (11)

[0096] In the formula, n is the total number of available frequency points of the frequency-using device i, and a j represents the jth available working frequency point of the frequency-using device i.

[0097] (2) According to the priority of the device and the interference situation it receives, use the fuzzy function to quantitatively evaluate the allocation result, and give corresponding reward values accordingly.

[0098] Considering that appropriately occupying the time-frequency working resources of the device in actual engineering can improve the anti-interference ability of the device, that is, improve the interference margin threshold of the device. The present invention defines the interference margin threshold of the modified device i as the "modified interference margin threshold", denoted as sff i . At the same time, in order to reflect the comprehensive impact of the interference margins of different priority devices on the formation, the reward values are divided into three intervals according to the interference margin threshold and the modified interference margin threshold.

[0099] Assume that the device to be allocated is a receiving device, and calculate the interference margin im of the remaining devices on the device to be allocated in turn according to Equation (12)​ij and the corresponding reward r ij , and then the mean value of r is obtained through Equation (13) ij to get the current decision action reward value.

[0100]

[0101]

[0102] In the formula, M is the total number of devices, and ip i is the priority of device i; sf i is the interference margin threshold of device i; k is a fixed coefficient used to amplify the reward value; σ is the standard deviation of the Gaussian function, and its value can be used to change the sensitivity of the change in the interference margin to the reward.

[0103] i) When the interference margin im of the device ij is lower than the interference margin threshold sf i , the device can work normally. Set its reward value to be the largest among the three intervals, and the calculation formula is ((M + 1 - ip i )k / M), and its reward value interval is [k / M, k], where k is a fixed coefficient and M is the total number of devices;

[0104] ii) When the device interference margin im ij exceeds the interference margin threshold sf i , but is lower than the corrected interference margin threshold sff i , it means that the device can improve the interference margin threshold by occupying additional working resources to achieve interference-free operation. Among them, when the interference margin is closer to the interference margin threshold, it means that the device working state is better, so a larger reward should be given; while when the interference margin is close to the corrected interference margin threshold, the electromagnetic compatibility state of the device deteriorates, and the reward should also be reduced accordingly. To make the reward value in this interval transition smoothly and avoid sudden changes, the Gaussian function is used to define r in this interval ij . Similarly, by multiplying the device priority weight ((M + 1 - ip i ) / M), the adjustment of the reward value corresponding to the device priority is realized.

[0105] iii) When the interference margin im ij exceeds the corrected interference margin threshold sff of the device i , the device is severely interfered. Therefore, the device will be punished in inverse proportion to its priority (1 / ip i ), and it is defined as a linear function of the part where the interference margin exceeds the threshold to amplify the punishment effect of bad decision actions and guide the device to avoid selecting the current frequency point.

[0106] 3) If the result matrix of the working frequency points allocated to all the frequency-using devices in the formation is AR, then the subset corresponding to the device i to be allocated is AR i , and its corresponding available frequency point resources are expanded into three dimensions, as shown in Equation (14).

[0107] AR i ={ar1, ar2, ar3} (14)

[0108] In the formula, AR i represents the frequency points carried by device i. Among them, ar1 represents the frequency point selected through the current decision-making action; ar2 represents the frequency point selected in the previous state; ar3 represents the frequency point corresponding to the decision-making action with the largest reward value in the set.

[0109] 4) Update AR i after each decision-making action, store the frequency point selected by the decision-making action as ar1, store the frequency point selected by the previous decision-making action as ar2, and store the frequency point corresponding to the largest reward value among the decision-making actions made as ar3.

[0110] 5) Store the frequency point selected by the current decision-making action and the corresponding reward value into the set of decision-making actions already made. If the number in the set does not meet the requirement for dual-frequency backtracking, then do not perform dual-frequency backtracking.

[0111] (4) After each device to be allocated selects a frequency point, the state of the environment will transfer to a new state. Combine the current state s t , action a i , reward r ti and the new state s t+1 in the form of a quadruple <s t , a i , r ti , s t+1 > and store it in the experience pool for subsequent training and learning.

[0112] (5) Randomly extract a sample set of size batch_size from the experience pool, and then use the extracted sample set to train the Q target network and continuously update the network parameters θ'.

[0113] (6) After a certain number of iterations T c , copy the trained Q target network parameters θ' to the Q estimation network.

[0114] (7) Repeat steps 3) to 6) until the learning process ends to obtain the DQN spectrum allocation model.

[0115] (8)Finally, according to the trained DQN spectrum allocation model, the optimal spectrum allocation scheme for the frequency-using devices within the formation in the current scenario can be output. When specifically performing the working frequency point allocation, first input the working parameters of all the frequency-using devices in the current state s1 within the formation into the DQN spectrum allocation model, and then the DQN spectrum allocation model automatically outputs the new state s2 of all the frequency-using devices within the formation. For the frequency-using devices whose working frequency points do not need to be allocated in the new state s2, the frequency points remain unchanged, and for the frequency-using devices whose working frequency points need to be allocated, the frequency points will be updated, and the updated frequency points are the allocated frequency points.

[0116] The present invention will be further described below in conjunction with specific examples and drawings.

[0117] Taking the spectrum allocation of the frequency-using devices of a certain platform aircraft formation as an example, the specific implementation manner of this invention patent is demonstrated.

[0118] This formation consists of 9 aircraft. Aircraft 1 is a single aircraft, and aircraft 2 to aircraft 9 are in a two-aircraft formation. Their specific positions are as Figure 2 shown. There are a total of 9 types of frequency-using devices within the formation. Their specific information is shown in Table 1. Types 1-7 are in the L band, type 8 is in the S band, and type 9 is in the X band. It is assumed that they are all transceiver-in-one devices, and device type 1 has only one working frequency point. Since anti-same-frequency interference measures are adopted among such devices, there will be no mutual interference among the same type of devices.

[0119] Table 1 Information on the types of frequency-using devices in the formation

[0120]

[0121] Assume that each aircraft carries 8 frequency-using devices, so the formation has a total of 72 frequency-using devices. Equation (15) gives the numbers of all the devices within this formation. Each row represents the numbers of all the devices carried by the aircraft members. Among them, devices 1 to 7 carried by each aircraft are the corresponding device types 1 to 7; ep 18 is device type 8, and ep 28 to ep 98 are device type 9.

[0122]

[0123] Set the priorities of all the frequency-using devices within the formation as shown in Equation (16). The devices that are more forward from left to right and from top to bottom have higher priorities.

[0124]

[0125] Set the interference margin threshold as 15 dB and the corrected interference margin threshold as 25 dB. Figure 3 Taking ep 11 being affected by ep 21Taking the interference margin as an example, the corresponding reward values for the interference margin from 0 dB to 40 dB are given.

[0126] When the interference margin is less than 15 dB, it indicates that the current decision-making action enables the device to operate normally at this frequency point. A maximum reward value of 20 is given to encourage continued making of similar favorable decision-making actions. When the interference margin is between 15 dB and 25 dB, the reward for the decision-making action starts to gradually decrease. This indicates that as the interference margin increases, the device needs to occupy working resources to enhance its anti-interference ability to achieve normal operation. The reduction of the reward is to remind that the current decision-making action is feasible but not optimal. When the interference margin exceeds 25 dB, it means that the device has been severely interfered, which may not only cause the device itself to be unable to work normally but also cause significant interference to surrounding devices. At this time, the reward value becomes negative, indicating that this is an unfavorable decision-making action. As the interference margin continues to increase, the magnitude of the penalty (negative reward) also increases significantly to reflect the negative impact of the decision-making action on the system performance and ensure that the device does not adopt such unfavorable decision-making actions.

[0127] Furthermore, the spectrum allocation for all devices within the formation is further carried out using the dual-frequency backtracking DQN algorithm. First, the number of iterations is set to 100,000 times to ensure that the algorithm has sufficient time to deeply explore the parameter space and avoid premature convergence. Setting every 5,000 iterations as an "episode" helps to regularly evaluate the performance of the algorithm and adjust the parameters if necessary. The neural network is set to have two hidden layers, and the number of neurons N in each hidden layer HL is set to 128, which can balance the expression ability and computational complexity of the model, avoid overfitting, and ensure that the model can capture the key features of the data at the same time. The capacity N of the experience replay pool R is set to 3,000. This moderate capacity stores enough historical data for learning, while avoiding excessive storage burden and data redundancy. The initial value of the exploration rate ε is set to 0.99 and gradually decreases to 0.01 with the number of iterations, allowing the algorithm to conduct extensive random exploration in the initial stage and then gradually reduce randomness and focus on refined search. The network parameter replication period T c is set to 1,000 iterations, enabling the algorithm to regularly learn from historical experience and update its parameters. The learning rate α is set to 0.01, the discount factor γ is set to 0.9, and the sample size batch_size is set to 64. Secondly, the iteration is carried out. In each episode of the algorithm, the average reward harvested by the formation is as follows Figure 4 shown. Finally, the spectrum allocation scheme of the frequency-using devices of this aircraft formation is output using the trained model, and Table 2 gives the allocation results of the frequency-using device types 5 and 7 within the formation.

[0128] From Figure 4It can be seen that during the iteration process, the algorithm gradually approaches the optimal value by continuously adjusting the spectrum selection decision actions of the devices. The average reward value of the formation reaches 1.19 in the last episode, approaching the optimal reward value of 1.35. This gradual improvement indicates that the algorithm can effectively learn how to adjust the operating frequency points of the devices to reduce the interference received by the devices. As can be seen from Table 2, in the current scenario, the model algorithm provides 3 available frequency points for the devices.

[0129] Table 2 Spectrum Allocation Results of Frequency-using Device Types 5 and 7 / MHz

[0130]

[0131] The present invention quantifies the decision-action reward value composed of the device interference margin and priority, strengthens the integrity among the devices in the formation, and preferentially guarantees the electromagnetic compatibility status of high-priority devices. Through the dual-frequency backtracking mechanism, the decision-action frequency points of the previous round and the decision-action frequency points of the historical maximum reward value are retained, providing more optional frequency points for the devices and significantly improving the stability of the algorithm.

[0132] The present invention can be directly used for the spectrum resource allocation of frequency-using devices in multi-platform formations in practice, helps to better solve the electromagnetic interference problem among frequency-using devices in the formation, and provides strong technical support for improving the electromagnetic compatibility among frequency-using devices in multi-platform formations.

[0133] Finally, it should be noted that the above examples of the present invention are merely illustrations given to explain the present invention and are not limitations on the implementation manners of the present invention. Although the applicant has described the present invention in detail with reference to the preferred embodiments, for those of ordinary skill in the art, other different forms of changes and modifications can still be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A spectrum allocation method for multi-platform formation frequency-using equipment based on dual-frequency backtracking DQN algorithm, characterized in that: The following steps are involved: S1. Obtain the working parameters of all frequency-using devices in the multi-platform formation; S2. In the first stage, a greedy strategy is used to allocate working frequencies to the frequency-using devices in the formation. When allocating frequencies, the interference margin and priority of the frequency-using devices are comprehensively considered, and the allocation results are quantitatively evaluated, and corresponding reward values ​​are given accordingly. S3, then after each allocation of working frequency points based on the greedy strategy, the state before allocation, the allocation action, the reward value corresponding to the allocation action, and the state after allocation are stored as samples in the experience pool in the form of a four-tuple; S4. When the samples accumulated in the experience pool reach the set capacity, a certain number of samples are extracted from the experience pool to train the initialized DQN network, and the DQN network parameters are updated until the set conditions are met. The training ends and the DQN spectrum allocation model is obtained. S5. In the second stage, when allocating working frequencies to the frequency-using devices that need to be allocated working frequencies within the formation, first input the working parameters of the current state s1 of all frequency-using devices in the formation into the DQN spectrum allocation model, and then the DQN spectrum allocation model automatically outputs the new state s2 of all frequency-using devices in the formation. In the new state s2, the frequencies of the frequency-using devices that do not need to be allocated working frequencies remain unchanged, and the frequencies of the frequency-using devices that need to be allocated working frequencies will be updated, and the updated frequencies are the allocated frequencies.

2. According to claim 1, a method for allocating spectrum of multi-platform formation frequency-using equipment based on dual-frequency backtracking DQN algorithm is characterized in that: In step S1, the working parameters include the working frequency, bandwidth, transmission power, location, interference margin and device priority of the frequency-using device; First, the set ME is used to represent the total number of frequency-using devices carried by each member in the formation, as shown in formula (1). Among them, N M is the number of team members, then me i is the number of frequency-using devices carried by formation member i; The set EP represents the set of all frequency-using devices carried by all members of the formation, as shown in formula (2); In formula (3), EP i represents the set of all frequency-using devices carried by member i in the formation; ep ij is the device number, indicating the jth device carried by member i; Secondly, define the positions of all members in the formation and define the transmission power, bandwidth and interference margin threshold of all frequency-using devices; they are represented by the sets MP, PT, BW and SF respectively, as shown in equations (4) to (7). In the formula, mp i is the position of member i, pt i 、bw i 、sf i represent the transmit power, device bandwidth and interference margin threshold of device i respectively; Then, the set AR is used to represent the result matrix of the working frequency points allocated to all frequency-using devices in the formation, as shown in formula (8); AR={ar1,…,ar i ,…,ar M } (8) Where, M is the total number of frequency-using devices; i Indicates the operating frequency of frequency-using device i; The priority of the device is represented by the set IP, as shown in formula (9); In the formula, For equipment ep ij The smaller the number, the higher the device allocation priority; Finally, AR m 、MP m PT m , BW m , SF m and IP m Respectively represent the state s m The working frequency, platform location, transmission power, bandwidth, interference margin threshold and device priority of all devices in the formation are shown in formula (10); s m ={AR m ,MP m ,PT m ,BW m ,SF m ,IP m } (10)。 3. According to claim 1, a method for allocating spectrum of multi-platform formation frequency-using equipment based on dual-frequency backtracking DQN algorithm is characterized in that: In step S2, the specific method of quantitatively evaluating the allocation results and giving corresponding reward values ​​is as follows: The interference margin threshold of device i after correction is called the corrected interference margin threshold, denoted as sff i ; Assume that the device to be allocated is the receiving device, and calculate the interference margin im of the remaining devices to the device to be allocated in turn according to formula (12) ij And the corresponding reward r ij , and then use formula (13) to calculate r ij Calculate the average value to get the reward value r of the current decision action i ; Where M is the total number of devices, ip i is the priority of device i; sf i is the interference margin threshold of device i; k is a fixed coefficient used to amplify the reward value; σ is the standard deviation of the Gaussian function, which is used to change the sensitivity of the interference margin change to the reward.

4. According to claim 1, a method for allocating spectrum of multi-platform formation frequency-using equipment based on dual-frequency backtracking DQN algorithm is characterized in that: In step S3, the specific method of storing the state before allocation, the allocation action, the reward value corresponding to the allocation action and the state after allocation as samples in the form of a four-tuple in the experience pool is: The result matrix of the working frequency points assigned to all frequency-using devices in the formation is AR, and the subset corresponding to the device i to be assigned is AR i , and expand its corresponding optional frequency resources into three dimensions, as shown in formula (14); AR i ={ar1,ar2,ar3} (14) In the formula, AR i Represents the frequency point carried by device i; ar1 represents the frequency point selected by the current decision action; ar2 represents the frequency point selected in the previous state; ar3 represents the frequency point corresponding to the decision action with the largest reward value in the set; Update AR after each decision action i , and store the frequency point selected for the decision action as ar1, the frequency point selected for the last decision action as ar2, and the frequency point corresponding to the maximum reward value in the decision action made as ar3; at the same time, store the frequency point selected for the current decision action and the corresponding reward value in the decision action set made; Each time the device to be assigned selects a frequency, the state of the environment will be transferred to the new state; Assign the previous state s t , assign action a i , assign the reward value r corresponding to the action ti And the state after allocation s t+1 , in the form of a quad t ,a i ,r ti ,s t+1 > are combined in the form of samples and stored in the experience pool.​ 5. According to claim 1, a method for allocating spectrum of multi-platform formation frequency-using equipment based on dual-frequency backtracking DQN algorithm is characterized in that: In step S4, the DQN network is initialized, including: 1) setting the number of iterations, exploration rate ε, learning rate α and discount factor γ; the discount factor γ is used to measure the importance of future rewards; a discount factor γ close to 1 indicates that future rewards are more important, and a discount factor close to 0 indicates that only current rewards are valued; 2) initializing the experience pool and setting the capacity N R and sample batch_size parameters; 3) Initialize the Q estimation network and Q target network, and set the number of hidden layer neurons N HL , network parameter θ, network parameter replication period T c .