Self-adaptive task scheduling method for multi-channel interference system based on deep Q network

Through the deep Q network, the task allocation of multi-task interference system is optimized, and task merging, DBF and aperture division are comprehensively considered, which solves the problems of low resource utilization and insufficient real-time performance in the existing technology, and realizes efficient resource management of multi-task interference system.

CN120357994APending Publication Date: 2025-07-22UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510444327.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art fails to effectively comprehensively consider task merging, DBF and aperture division in multi-task electronic interference systems, resulting in low resource utilization and difficult to meet real-time requirements.

Method used

Adaptive task scheduling method based on deep Q network is adopted, combined with task merging, DBF and aperture division, and task allocation is optimized through deep Q network training to meet the time, channel, array, airspace and energy resource constraints, and realize the simultaneous execution of multitasks.

Benefits of technology

The resource utilization rate of multi-task interference system is improved, real-time requirements are met, and the system execution benefits are maximized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357994A_ABST
    Figure CN120357994A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of electronic system resource management, and particularly relates to a self-adaptive task scheduling method of a multi-channel interference system based on a deep Q network. Aiming at the characteristics of a multi-channel interference system, the multi-task simultaneous execution is realized by comprehensively considering task merging, DBF and aperture division modes, and in order to effectively solve the multi-interference task scheduling under the high-dimensional resource constraint, a deep Q network method is adopted, specific actions, states and reward functions are set, and network parameters are trained, so that the multi-interference task scheduling efficiency is improved. And task scheduling is implemented by using the trained network. Simulation results show that the method meets the real-time requirement while having high execution benefits, and is a real-time and effective adaptive task scheduling method for the multi-channel interference system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic system resource management, and particularly relates to an adaptive task scheduling method for a multi-channel interference system based on a deep Q-network. Background Art

[0002] Modern warfare is becoming increasingly complex, and electronic jamming technology has gradually become an indispensable part of electronic warfare. In actual electronic jamming, the resources of the jamming system are usually limited. The traditional single-beam platform or fixed resource allocation method will result in problems such as unsatisfactory jamming effects on the enemy and low resource utilization, and it is difficult to maximize the limitation of the threats brought by multiple radars and multiple communication systems. Therefore, the resource scheduling problem of the jammer system is a key to achieving effective electronic jamming.

[0003] Early electronic jamming systems emit a single beam at each moment. For a single-beam system, when faced with multiple tasks to be executed, it adopts a time-division multiplexing method. The task scheduling algorithm based on the heuristic method by Orman AJ et al. realizes beam dwell scheduling according to the task priority (Orman AJ, Potts C N, Shahani AK, et al. Scheduling for a multifunction phased array radar system[J]. European Journal of Operational Research, 1996, 90(1): 13-25). With the gradual increase in the number of tasks and the gradual development of electronic technology, the jammer system has become increasingly complex and diversified. The single-beam system has gradually introduced aperture division technology and digital beamforming technology and developed into a multi-beam system. It not only needs to allocate time resources to tasks but also needs to consider dividing resources such as the array surface, airspace, and frequency domain so that the jammer can execute multiple tasks at the same moment. X. Zhu considered the time resource constraint and array surface resource constraint for the task scheduling problem of the multi-beam system (X. Zhu, R. Yang, X. Li, K. Yuan, Multifunctional integrated system resource scheduling based on improved GA-PSO, J. Air Space Early Warn. Res. 35(3)(2021)). Qi Wenchao et al. adopted a heuristic scheduling algorithm, namely the multi-task parallel EDF (MTPEDF) algorithm, and gave a method for resource management of a multifunctional integrated radar system based on aperture segmentation, which uniformly allocates time and aperture resources. (Qi Wenchao, Yang Ruijuan, Li Xiaobo, et al. Research on task scheduling algorithms for multifunctional integrated radars[J]. Radar Science and Technology, 2012, 10(2): 150-155. DOI: 10.3969 / j.issn.1672-2337.2012.02.006).

[0004] In view of the non-convex, non-linear, and multi-constrained characteristics of the task scheduling problem, more and more researchers have adopted swarm intelligence algorithms to solve it. S. Yang et al. presented a scheduling model based on value optimization and obtained the optimal scheduling sequence corresponding to this problem using a genetic algorithm (S. Yang, K. Tian and R. Liu, “Task Scheduling Algorithm Based on Value Optimization for Anti-missile Phased Array Radar,” IET Radar Sonar Navig., vol. 13, no. 11, pp. 1883–1889, 2019). F. Meng et al. solved this problem using a beam dwell scheduling algorithm based on a particle swarm-simulated annealing algorithm (F. Meng and K. Tian, “Phased-Array Radar Task Scheduling Method for Hypersonic-Glide Vehicles,” IEEE Access, vol. 8, pp. 221288-221298, 2020). Sun Jun et al. used a particle swarm algorithm to solve the joint optimization problem of interference beam and power with interference resource constraints, realizing the reasonable allocation of interference resources (Sun Jun, Zhang Dalin, Yi Wei. Resource Scheduling Method for Multi-aircraft Cooperative Interference Networking Radar [J]. Radar Science and Technology, 2022, 20(03): 237-244+254).

[0005] The above research has achieved certain results in the multi-task scheduling of interference systems, but only considered a single way of multi-task merging or aperture division during the simultaneous execution of multiple tasks, without comprehensively combining methods such as task merging, DBF, and aperture division. At the same time, heuristic methods are difficult to effectively solve non-convex, non-linear, and multi-constrained task scheduling problems, while swarm intelligence algorithms have the problem of being difficult to meet the real-time requirements of scheduling. Therefore, the present invention proposes a simultaneous multi-beam adaptive task scheduling method for a multi-channel interference system based on a deep Q-network, which considers ways of task merging, DBF, and aperture division, as well as time, channel, array surface, airspace, and energy resource constraints. By effectively controlling the multi-channel interference frequency band, the tasks assigned to each channel for execution, and the specific execution methods of each task, the system can maximize the execution benefit of the system and meet the real-time requirements. Summary of the Invention

[0006] Assume a self-defense single jammer platform with M channels, which receives N tasks to be scheduled at time k. Denote the i-th radar interference task to be scheduled as T i,k , i = 1, 2,..., N.

[0007] For task Ti,k , the threat level of the corresponding interference target is w i,k , w i,k ∈N + , the greater the threat level, the more urgent the need for the task to be scheduled. The start and end times of the interference request are respectively and The transmit power, antenna gain, combined loss, and RCS of the detected target of the target radar are respectively P i,k , G i,k , L i,k and σ i,k . The suppression coefficient of the jammer to the detected target radar is K i,k , the distance, azimuth angle, and elevation angle information of the target relative to the jammer system are respectively expressed as R i,k , θ i,k and α i,k , the moving speed of the target where and respectively correspond to the component velocities in the polar coordinate system. The expected center frequency and bandwidth of the task are respectively and Therefore, the expected interference frequency band of the task is

[0008] The steps of an adaptive task scheduling method for a multi-channel interference system based on a deep Q-network (DQN) at time k are as follows:

[0009] Step 1: Reorder all tasks based on the threat level.

[0010]

[0011] where, ΔT i is the interference request time of task i, μ is the influence factor of the interference request time, 0 < μ ≤ Δt, where Δt is the scheduling time interval. is the newly generated threat level, ensuring that tasks with longer interference request times are scheduled more preferentially to improve the stability of scheduling between different times. Based on reorder all tasks at time k from largest to smallest. The reordered tasks are denoted as

[0012] Step 2: Initialize

[0013] Step 2.1: Channel initialization

[0014] If the number of scheduled tasks is not more than the number of channels, there is no need to use the network, and each task is assigned to occupy one channel; if the number of scheduled tasks is more than the number of channels, the M tasks with the highest threat levels respectively occupy each channel, specifically:

[0015] s M,k = [T1 T2...T M (2)

[0016] Among them, the state s M,k represents the occupancy of the M channels of the jammer by the M tasks before the k-th moment, T i , i = 1,..., M are the tasks corresponding numbers, specifically T i = 2 i-1 , i = 1, 2,..., N, so the occupancy of the M channels of the jammer by the first i tasks before the k-th moment s i,k is specifically expressed as

[0017]

[0018] Among them, i m,n ∈{1, 2,..., i}, n = 1, 2,..., N m , m = 1, 2,..., M, N m is the total number of tasks processed by the jammer channel m and ensures that N m ≤ N m,max , N m,max is the maximum number of tasks that the jammer channel m can process.

[0019] Step 2.2: Initialization of the online network θ and the target network θ - Initialization

[0020] Construct two BP neural networks for calculating the reward value with the state and action as inputs, namely the online network θ and the target network θ - , and randomly select actions in the action space A for task scheduling to initialize the online network and the target network. The action space is specifically expressed as

[0021]

[0022] Among them [·] is the floor symbol, mod is the modulo symbol, means occupying the j-th channel, means not occupying the j-th channel.

[0023] Step 3: Based on all the re-ordered tasks to be scheduled at the k-th moment, the online network θ is trained and updated in turn, and a total of T rounds of loop training are carried out.

[0024] Step 3.1: Select the action a for the task based on the ε-greedy algorithm i,k . ​

[0025] The algorithm randomly selects an action in the action space A with a probability of ε, and selects the action that can obtain the maximum Q value with a probability of 1-ε, which is specifically expressed as follows:

[0026] a i,k = argmaxQ(s i-1,k ,A;θ) (5)

[0027] Among them, the exploration probability ε = e -δt , the non-negative real number δ is the influencing factor of exploration, and t is the measurement value of the training round. As the number of training rounds increases, the action selection of the task increases the probability of using the network and reduces the probability of exploration, thus making the network more stable. Q(s i-1,k ,A;θ) is the output value obtained using the online network θ.

[0028] Step 3.2: State s i,k Updates.

[0029] s i,k =s i-1,k +T i a i,k (6)

[0030] where s i,k and i+1,k are the states before and after the update, T i For the task The number and T i =2 i-1 ,i=1,2,...,N,a i,k The action selected in step 2.1.

[0031] Step 3.3: Reward value r i,k Get

[0032]

[0033] Among them, the penalty value Ω is a very small negative real number, For task concurrency conflict judgment, when When it is 0, there is at least one channel whose task scheduling result does not meet the resource constraints. When is 1, the resource constraints are satisfied, that is, for any channel, it satisfies:

[0034]

[0035] in, is the total number of tasks processed by channel m after the first i tasks are scheduled at time k, i * is the number of all tasks on channel m. Let \(n_{m,i}^k\) be the minimum number of feasible beams for channel \(m\) after scheduling the first \(i\) tasks at time \(k\), and \(B\) be the maximum number of beams that the channel can generate. \(\zeta(s i,k )\) and \(\zeta(s i-1,k )\) are the task execution benefits corresponding to before and after the state update, specifically expressed as

[0036]

[0037] Among them, \(u j,k = \{1, 0\}\) represents a Boolean variable indicating whether the \(j\)-th task is executed at time \(k\), \(w j,k \) represents the threat level of task \(j\) at time \(k\), \(\varPsi j,k \) and respectively represent the actual execution efficiency and the expected execution efficiency of task \(j\) at time \(k\), specifically expressed as

[0038]

[0039] Among them, is the transmission power allocated to task \(j\) on jammer channel \(m\) at time \(k\), is the transmission gain of task \(j\) on jammer channel \(m\) at time \(k\), is the frequency band coverage rate of task \(j\) on jammer channel \(m\) at time \(k\), specifically expressed as

[0040]

[0041] Among them, \(f m \) and \(K\) are the center frequency and bandwidth of channel \(m\), respectively, and

[0042] Step 3.4: Beam combination and subarray division

[0043] For two tasks in the same channel, set the azimuth threshold \(\theta th \) and the elevation threshold \(\alpha th .

[0044] If the azimuth angles \(\theta i,k , \theta j,k and the elevation angles \(\alpha i,k , \alpha j,k of the two tasks satisfy

[0045]

[0046] Then the two tasks can be combined into one task, that is, generate one beam to jam two targets simultaneously.

[0047] If the azimuth angles \(\theta i,k , \theta j,k and the elevation angles \(\alpha i,k , \alphaj,k Meet

[0048]

[0049] Then, the DBF technology can be used to execute these two tasks on the same array surface, and at this time, this array surface may be the entire array surface or a sub-array surface obtained by dividing the array surface.

[0050] If the azimuth angles θ i,k 、θ j,k and the elevation angles α i,k 、α j,k Meet

[0051]

[0052] At this time, the DBF technology cannot be used to execute these two tasks simultaneously, and these two tasks can only be assigned to different sub-array surfaces of the channel for execution.

[0053] Step 3.5: Store the experience φ in the experience replay pool and extract part of the experience from the experience replay pool for subsequent training.

[0054] φ = [s i-1,k a i,k r i,k s i,k (16)

[0055] Among them, s i-1,k 、a i,k 、r i,k 、s i,k Correspond to the state before action execution, the action taken, the reward value for executing the action, and the state after action execution respectively.

[0056] Extract L experiences to form a set Φ, which contains L1 randomly extracted experiences and L2 newly added experiences to balance robustness and real-time performance.

[0057] Step 3.6: Use the Bellman equation for Q-value prediction

[0058]

[0059] Among them, Is the predicted Q-value of the l-th experience in the obtained experience set Φ. γ is the discount factor, and the future reward value is discounted by the ratio of γ, γ ∈ (0, 1). And Are the action execution reward value and the state after action execution of the l-th experience in the experience set Φ respectively. The target network θ - Specifically, it is

[0060]

[0061] Among them, mod is the modulo operation. τ is a positive integer, and τ ≤ N.

[0062] Step 3.7: Network training and update

[0063] Train the current network to minimize the gap between the output of the online network θ and the predicted Q value, that is, update the network parameters through the backpropagation algorithm to minimize the loss function. The loss function adopts the mean square error method, specifically expressed as

[0064]

[0065] Use the gradient descent method to solve the online network that minimizes the loss function, specifically expressed as

[0066]

[0067] where θ t and θ t-1 are the updated and pre-updated online networks respectively, and μ θ is the gradient descent step size.

[0068] Step 4: Use the trained network to schedule and allocate tasks at time k

[0069] Use the updated online network to schedule the N - M tasks that are not allocated at time k, and take the action that obtains the maximum Q value in each task allocation as the action taken when allocating this task, specifically expressed as

[0070] a i,k = argmaxQ(s i-1,k , A; θ), i = 1,..., N - M (21)

[0071] where a i,k , i = 1,..., N - M are the actions taken by all tasks to be scheduled, and θ is the finally trained network.

[0072] Use the task allocation action obtained by formula (21) to obtain the final channel allocation result, that is, the final state at time k, specifically expressed as

[0073]

[0074] where s k is the final state at time k, that is, the scheduling situation of all tasks. s 0,k is the initial state before the scheduling starts.

[0075] Principle of the invention

[0076] The present invention provides an adaptive task scheduling method for a multi-channel interference system based on DQN. This method is applicable to a multi-channel interference system, and can comprehensively consider the characteristics of simultaneously performing multiple interference tasks in the ways of beam combining, aperture division, and DBF, and achieve multi-task adaptive scheduling under the system resource constraints. The principle is elaborated below.

[0077] Suppose there are M independent channels in the multi-channel interference system, the execution frequency band width of each channel is K, the maximum number of beams is B, and the starting operating frequency of this frequency band can be flexibly configured. Suppose that at time k, there are N tasks participating in the scheduling execution, then the execution benefit of the system can be defined as

[0078]

[0079] where, u j,k = {0, 1}, representing a Boolean variable indicating whether the j-th task is executed at time k, w j,k represents the threat degree of task j at time k, Ψ j,k and respectively represent the actual execution efficiency and the expected execution efficiency of task j at time k, the same as in equations (10) and (11).

[0080] In the DQN algorithm, we are more concerned about the interference benefit during the entire interference process rather than a single moment. We not only need to consider the reward value at each moment, but also need to consider the impact of the current action on the subsequent reward values. Therefore, the goal of the agent is to maximize the rewards in all subsequent interference frames:

[0081]

[0082] where, γ represents the discount factor of the impact of future rewards on the current, r k+t represents the immediate reward value at time k + t, and K represents the number of the entire tracking frames.

[0083] Taking equation (24) as the objective function, the optimization problem is established as follows:

[0084]

[0085] Constraint 1 is the time resource constraint. For the interference tasks scheduled for execution, the selected interference time must be within its interference effective time window. [k, k + 1]·Δt represents the time range of the sub-stage corresponding to time k, and the set I stores all the tasks actually executed at time k.

[0086] Constraint 2 is the time and channel resource combination constraint. It is necessary to satisfy that the bandwidth amount covered by the actually executed tasks cannot exceed the upper limit of the bandwidth coverage supported by the system. F i,k represents task T at time k iThe execution frequency range, where f i,k and B i,k are the main execution frequency and bandwidth of task T i respectively, and the union of the execution frequency ranges of all tasks in the I set can be further expressed as several continuous execution frequency bands [s j , e j , j = 1, 2,..., J, where J represents the number of continuous frequency bands in the set. The sum of the J continuous frequency bands can describe the total bandwidth occupied by the multiple tasks executed at time k, which is the total frequency band resource requirement.

[0087] Constraint three is the time, channel, and array surface resource combination constraint. Simultaneous multi-beam interference can be achieved through full-array DBF in one channel, or by dividing the array surface to form a one-to-one relationship between the array surface and the beam. The total number of beams formed in the same channel cannot exceed the maximum number of beams that the channel can generate. B k (i, m) is a Boolean matrix, and B k (i, m) = {1, 0}. When B k (i, m) = 1, it means that task i occupies channel m at time k; when B k (i, m) = 0, it means that task i does not occupy channel m at time k. Summing over i with m fixed represents that the number of tasks executed simultaneously on channel m cannot exceed the maximum number of beams B of the channel.

[0088] Constraint four is the time, channel, array surface, and spatial domain resource combination constraint. Among them, η i,k (m) represents the array surface resource amount allocated by the system to task i on channel m, and the maximum array surface resource amount is expressed as 1; B k (i, m) characterizes whether task i occupies channel m. This constraint is for the task beams i and j executed simultaneously in the same channel. At the same time, the thresholds θ th and α th are used to characterize the position relationship thresholds of the azimuth and elevation angles of tasks i and j respectively. If it is less than the threshold, it is defined that the azimuth or elevation angles of the two tasks are close; otherwise, it is defined that the azimuth or elevation angles of the two tasks are far apart. In addition, this constraint also introduces the matrix sη i,k (m) to characterize whether the DBF technology is applied when the system emits a beam to irradiate target i on channel m at time k; according to the above three situations, for the beam combination and sub-array division rules in step 3.3, sη i,k (m) = 0 reflects situation (1), where the two task beams are combined into one beam, and sη i,k (m) = 1 reflects situation (2), indicating that the beam of task i is formed by the DBF technology, and sη i,k(m) = 2 reflects the situation (3), and the task i beam is emitted separately by the divided sub - arrays. However, the array resource allocation in one channel does not only include one of the above three situations, but there are also cases where multiple situations exist simultaneously among the three situations.

[0089] Constraint five is the combination constraint of time, channel, array, airspace, and energy resources. It restricts the power upper limit of the entire array and the power upper limit of the sub - arrays in each channel under the condition of introducing DBF technology or array division. The first inequality represents that when task i is executed in channel m and DBF technology is applied, the upper limit of the execution power of task i is [η i,k (m)] 2 ·P total (m) / (n + 1), where n is the total number of other tasks transmitted in the form of DBF together with task i. The second inequality represents that when task i is executed alone on the entire array or sub - arrays in channel m, the upper limit of the execution power of the task is [η i,k (m)] 2 ·P total (m).

[0090] For the solution of this optimization problem, the tasks that are given priority consideration have a greater impact on the solution of the objective function. Therefore, in step 1, tasks with a greater threat level and a long period are given priority consideration. The greedy algorithm adopted in step 3.1 has a relatively large random probability in the initial stage of network training, which can explore all situations as much as possible; while in the late stage of training, the probability of using the network is relatively large, which can ensure the stability of training. The limitation of the constraint conditions and the calculation of the execution benefits are shown in step 3.3. Step 3.5 uses the Bellman equation to consider the influence of the current state and action on subsequent moments, so as to maximize the rewards in all subsequent interference frames. At the same time, it calculates the Q - value using a periodically updated objective function, which can reduce the fluctuation of the target value and improve the stability of the algorithm. Brief Description of the Drawings

[0091] Figure 1 is the algorithm flow framework of the deep Q - network

[0092] Figure 2 is the convergence curve of network training as the number of episodes increases

[0093] Figure 3 is the comparison of the execution benefits of each algorithm system in simulation scenario 1

[0094] Figure 4 is the comparison of the execution benefits of each algorithm system in simulation scenario 2

[0095] Figure 5 is the comparison of the average running time of each algorithm for each task scheduling in two simulation scenarios Detailed Implementation Manner

[0096] Within the simulation duration of 1 - 50 s, analyzed at 1 - s intervals, there are 10 radar targets (only radar targets are considered currently, communication targets are not considered), and the 10 radar targets move over time, making the targets to be jammed random for each analysis interval. Two scenarios with different properties are set in the simulation: In Scenario 1, there are fewer task conflicts and the task frequency band coverage is wide; in Scenario 2, there are more task conflicts and the task frequency band coverage is narrow. The specific parameters of the two simulation scenarios are shown in Tables 1 - 2 respectively.

[0097] Table 1 Parameters of Simulation Scenario 1

[0098]

[0099] Table 2 Parameters of Simulation Scenario 2

[0100]

[0101]

[0102] To illustrate the effectiveness of the present invention, we compare the performance of the present invention with two heuristic methods (Threat Level First method, Combination First method) and the Particle Swarm Optimization (PSO) algorithm. The Threat Level First (TLF) method means that resources are allocated to tasks in the order of the threat levels of multiple tasks; the Combination First (CF) method regards tasks that can be executed by the same beam as tasks with higher priority and allocates resources preferentially; the Particle Swarm Optimization (PSO) algorithm represents using the particle swarm optimization method to solve the optimization problem shown in Equation (25).

[0103] The comparison of the system execution benefits of each algorithm under the two simulation scenarios is as Figure 3-4 shown. The comparison of the execution benefits of the four algorithms in Simulation Scenario 1 is as Figure 3 shown. In the case of fewer task conflicts and wide task frequency band coverage, the differences in the execution benefits of the four algorithms are not obvious. The execution benefit of the method of the present invention is slightly worse than that of the PSO algorithm, but better than the TLF algorithm and the CF algorithm. The comparison of the execution benefits of the four algorithms in Simulation Scenario 2 is as Figure 4 shown. In the case of more task conflicts and narrow task frequency band coverage, there is no obvious difference in the execution benefits between the method of the present invention and the PSO algorithm, but it shows stronger stability relative to the PSO algorithm and is better than the TLF algorithm and the CF algorithm. In summary, under the two simulation scenarios, the execution benefit of the method of the present invention has little difference from that of the PSO algorithm and is better than the TLF algorithm and the CF algorithm.

[0104] The comparison of the average running time of each algorithm for each task scheduling in two simulation scenarios using an Intel(R) Core(TM) i7-10750H CPU and 8.00GB RAM on the Matlab R2018b platform is as follows Figure 5 As shown, in each simulation scenario, the average time consumption of the method of the present invention, the TLF algorithm, and the CF algorithm for each scheduling is significantly lower than 1 s, meeting the real-time requirement. However, the average time consumption of the PSO algorithm for each scheduling exceeds 1 s in the time interval of the second simulation scenario, not meeting the real-time requirement.

[0105] In summary, although the TLF algorithm and the CF algorithm meet the real-time requirement, their execution benefits are significantly lower than those of the other two algorithms. The PSO algorithm has the best execution benefit but cannot meet the real-time requirement. The method of the present invention meets the real-time requirement while having little difference in execution benefit from the PSO algorithm. It can be seen that the algorithm of the present invention is a real-time and effective adaptive task scheduling method for multi-channel interference systems.

Claims

1. An adaptive task scheduling method for a multi-channel interference system based on a deep Q-network, and the specific technical solution is as follows: Suppose there is a self-defense single jammer platform with M channels, and at time k, N tasks to be scheduled are received. Denote the i-th radar jamming task to be scheduled as T i,k , where i = 1, 2,..., N. For task T i,k , the threat level of the corresponding interference target is w i,k , w i,k ∈N + , the greater the threat level, the more urgent the need for the task to be scheduled. The start and end times of the interference request are respectively and The transmit power, antenna gain, combined loss, and RCS of the detected target of the target radar are P i,k , G i,k , L i,k and σ i,k . The suppression coefficient of the jammer on the detected target radar is K i,k , and the distance, azimuth angle, and elevation angle information of the target relative to the jammer system are respectively expressed as R i,k , θ i,k and α i,k , and the moving speed of the target where and correspond to the component speeds in the polar coordinate system respectively. The expected center frequency and bandwidth of the task are respectively and Therefore, the expected interference frequency band of the task is The steps of an adaptive task scheduling method for a multi-channel interference system based on a deep Q-network (DQN) at time k are as follows: Step 1: Reorder all tasks based on threat level. Among them, ΔT i is the interference request time for task i μ is the influence factor of the interference request time, 0 < μ ≤ Δt, where Δt is the scheduling time interval is the regenerated threat level, ensuring that tasks with longer interference request times are scheduled with higher priority to improve the stability of scheduling between different times. Based on all tasks at time k are re - sorted from largest to smallest, and the re - sorted tasks are denoted as Step 2: Initialize Step 2.1: Channel initialization If the number of scheduling tasks is not more than the number of channels, there is no need to use the network, and each task is assigned to occupy one channel; if the number of scheduling tasks is more than the number of channels, the M tasks with the highest threat level respectively occupy each channel, which is specifically manifested as: s M,k = [T1 T2... T M (2) Among them, the state s M,k represents the occupancy of M channels of the jammer by the M tasks before the k-th moment, T i , i = 1, ..., M are the task corresponding numbers, specifically T i = 2 i-1 , i = 1, 2, ..., N, so the occupancy of the M channels of the jammer by the first i tasks before the k-th moment s i,k is specifically manifested as where \(i\) m,n \(\in\{1,2,\cdots,i\},n = 1,2,\cdots,N\) m , \(m = 1,2,\cdots,M\), \(N\) m is the total number of tasks processed by jammer channel \(m\) and it is ensured that \(N\) m \(\leq N\) m,max , \(N\) m,max is the maximum number of tasks that jammer channel \(m\) can process. Step 2.2: Initialization of the online network θ and the target network θ - Initialization Construct two BP neural networks for calculating reward values with states and actions as inputs, namely the online network θ and the target network θ - , randomly select actions in the action space A for task scheduling to initialize the online network and the target network. The action space Specifically wherein [·] is the floor symbol, and mod is the modulo symbol, means occupying the j-th channel, means not occupying the j-th channel. Step 3: Based on all the re-ordered tasks to be scheduled at time k, the online network θ is trained and updated in sequence, and a total of T rounds of loop training are carried out. Step 3.1: Select action a for the task based on the ε-greedy algorithm for action a i,k selection. The algorithm randomly selects an action in the action space A with a probability of ε, and selects the action that can obtain the maximum Q value with a probability of 1 - ε, which is specifically manifested as: a i,k = arg max Q(s i-1,k , A; θ) (5) Among them, the exploration probability ε = e -δt , the non - negative real number δ is the influence factor of exploration, t is the measurement value of the training rounds. As the number of training rounds increases, the probability of selecting actions for the task to use the network increases, and the exploration probability decreases, thus making the network more stable. Q(s i-1,k , A; θ) is the output value obtained using the online network θ. Step 3.2: Update of status s i,k ​ s i,k = s i-1,k + T i a i,k (6) where s i,k and s i+1,k are the states before and after the update, respectively, T i is the task number and T i = 2 i-1 , i = 1, 2,..., N, a i,k is the action selected in step 2.

1. Step 3.3: Reward value r i,k Obtain Among them, the penalty value Ω is an extremely small negative real number. is for task concurrency conflict judgment. When is 0, there is at least one channel whose task scheduling result does not meet the resource constraints. When is 1, then all meet the resource constraints, that is, for any channel, it satisfies: Among them, is the total number of tasks processed by channel m after scheduling the first i tasks at time k, and i * is the number of all tasks on channel m. is the minimum viable beam number of channel m after scheduling the first i tasks at time k, and B is the maximum number of beams that the channel can generate. ζ(s i,k ) and ζ(s i-1,k ) are the task execution benefits corresponding before and after the state update, specifically manifested as where, u j,k = {1, 0} represents a Boolean variable indicating whether the j-th task is executed at the k-th moment, w j,k represents the threat level of task j at the k-th moment, Ψ j,k and respectively represent the actual execution efficiency and the expected execution efficiency of task j at the k-th moment, specifically manifested as wherein, is the transmit power allocated to task j on jammer channel m at time k, is the transmit gain of task j on jammer channel m at time k, is the frequency band coverage rate of task j on jammer channel m at time k, specifically where f m and K are the center frequency and bandwidth of channel m, respectively, and Step 3.4: Beam combination and sub-array division Set the azimuth threshold θ th and the elevation threshold α th . If the azimuth angles θ i,k and θ j,k and the elevation angles α i,k and α j,k satisfy Then two tasks can be combined into one task, that is, generate a beam to interfere with two targets at the same time. If the azimuth angles θ i,k and θ j,k of two tasks, and the elevation angles α i,k and α j,k satisfy Then the DBF technology can be used to execute these two tasks on the same array surface. At this time, this array surface may be the full array surface and the sub-array surface obtained by dividing the array surface. If the azimuth angles θ i,k 、θ j,k and the elevation angles α i,k 、α j,k meet At this time, the DBF technology cannot be used to execute these two tasks simultaneously, and these two tasks can only be assigned to different sub-array surfaces of the channel for execution. Step 3.5: Store the experience φ in the experience replay pool, and extract part of the experience from the experience replay pool for subsequent training. φ=[s i-1,k a i,k r i,k s i,k (16) Among them, s i-1,k , a i,k , r i,k , s i,k correspond to the state before action execution, the action taken, the reward value for executing the action, and the state after action execution, respectively. Extract L experiences to form a set Φ, which contains L1 randomly extracted experiences and L2 newly added experiences to balance robustness and real-time performance. Step 3.6: Predict the Q value using the Bellman equation Among them, is the predicted Q value of the l-th experience in the obtained experience set Φ. γ is the discount factor, and future reward values are discounted by the ratio of γ, where γ ∈ (0, 1). and are the execution action reward value and the state after action execution of the l-th experience in the experience set Φ, respectively. The target network θ - Specifically, Among them, mod is the modulo operation. τ is a positive integer, and τ ≤ N. Step 3.7: Network training and update Train the current network to minimize the gap between the output of the online network θ and the predicted Q value, that is, update the network parameters through the backpropagation algorithm to minimize the loss function. The loss function adopts the mean square error method, which is specifically manifested as Adopt the gradient descent method to solve the online network that minimizes the loss function, which is specifically manifested as where θ t and θ t-1 are the online networks after and before the update respectively, and μ θ is the gradient descent step size. Step 4: Use the trained network to schedule and allocate tasks at time k Use the updated online network to schedule the N - M tasks that are not assigned at time k. Take the action that obtains the maximum Q value in each task allocation as the action taken when allocating this task, which is specifically manifested as a i,k = arg max Q(s i-1,k , A; θ), i = 1, ..., N - M (21) where a i,k , i = 1, ..., N - M are the actions taken by all tasks to be scheduled, and θ is the finally trained network. Use the task allocation action obtained by formula (21) to obtain the final channel allocation result, that is, the final state at time k, which is specifically manifested as Among them, s k is the final state at time k, that is, the scheduling situation of all tasks. s 0,k is the initial state before the scheduling starts.

Citation Information

Cited By

  • A multi-dimensional resource management and control algorithm for a multi-platform jamming system based on PSO

    CN122660801A