Military communication dynamic priority scheduling method based on deep reinforcement learning

A dynamic priority scheduling method combining deep reinforcement learning and Pareto optimization solves the static priority scheduling and multi-objective optimization problems of military communication systems in dynamic environments, improving the system's flexibility and anti-interference capabilities, and optimizing resource utilization and communication efficiency.

CN120916261APending Publication Date: 2025-11-07EAST CHINA INST OF COMPUTING TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510902078.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing military communication systems have shortcomings in static priority scheduling, multi-objective optimization, anti-jamming capabilities, and resource utilization. In particular, they are unable to meet real-time mission requirements and network status changes in dynamic environments, resulting in communication delays, resource waste, and insufficient anti-jamming capabilities.

Method used

A dynamic priority scheduling method based on deep reinforcement learning is adopted. By dynamically adjusting task priorities, optimizing communication paths and resource allocation, and combining Pareto optimization algorithm, the optimal balance point of multiple objectives is found in complex environments, thereby improving system flexibility and response speed.

Benefits of technology

It significantly improves the flexibility, response speed, resource utilization and anti-jamming capability of military communication systems, ensuring timely processing of critical tasks and rapid information transmission, and optimizing communication latency and bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120916261A_ABST
    Figure CN120916261A_ABST
Patent Text Reader

Abstract

The invention relates to a military communication dynamic priority scheduling method based on deep reinforcement learning, and the method comprises the following steps: dynamic priority scheduling: carrying out the real-time adjustment through a dynamic priority formula according to the related information of a military task, state space definition, action space design, reward mechanism construction, Q table initialization, initial state selection and action selection are respectively carried out; an intelligent agent selects actions by using an epsilon-greedy strategy; setting the maximum number of iterations; executing the action, and updating the Q value in the Q table; in combination with multi-objective optimization, acquiring a current state and selecting actions; selecting actions by an intelligent agent by using an epsilon-greedy strategy; executing an action, observing feedback, and updating a Q table: updating a Q value in the Q table by the intelligent agent according to the new state and the reward; and military communication dynamic priority scheduling is completed. The problems of insufficient static priority scheduling, multi-objective optimization, anti-interference capability and resource utilization rate of a military communication system are solved, and the flexibility and response speed of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a military communication technology, in particular to a dynamic priority scheduling method based on deep reinforcement learning (DRL). BACKGROUND

[0002] 1. Overview of existing related technologies

[0003] In modern military operations, the rapid processing and transmission of information and the effective scheduling of computing tasks have become one of the key factors determining the success or failure of the battle. Efficient communication and computing mechanisms not only accelerate the speed of intelligence collection, battle plan formulation, and task execution, but also significantly improve the reaction speed and operational flexibility of the troops. However, traditional static priority scheduling methods often fail to meet the needs when faced with highly uncertain and dynamic battlefield environments, as they lack the ability to respond to changes in the current situation.

[0004] Currently, domestic and foreign research in the field of military communication mainly focuses on the following aspects:

[0005] Information network architecture: Domestic research institutions and universities are committed to building efficient and reliable military communication networks. For example, the China Electronics Technology Group, the University of Defense Science and Technology, and other units have made significant achievements in the design and optimization of military communication networks. These researches not only focus on the physical layer and data link layer of the network, but also delve into the optimization methods of the network layer and transport layer.

[0006] Communication protocol: Domestic researchers focus on how to ensure the reliability and security of communication in complex electromagnetic environments. For example, the Chinese Academy of Sciences, Beijing University of Posts and Telecommunications, and other institutions have carried out a lot of research on anti-interference communication protocols, adaptive modulation and demodulation technology, etc.

[0007] Resource scheduling: Domestic researchers actively explore dynamic scheduling methods based on deep learning and reinforcement learning. For example, research teams from Tsinghua University and Peking University have made important progress in the cooperative communication and resource optimization of unmanned aerial vehicle swarms.

[0008] Although domestic research has made some achievements in deep learning and dynamic scheduling, there is still relatively limited research on multi-objective optimization. Multi-objective optimization in military communication needs to consider multiple objectives such as communication delay, bandwidth utilization, and energy consumption, and find a best compromise solution. Currently, domestic researchers mainly focus on the research of single-objective optimization algorithms such as genetic algorithm, particle swarm optimization algorithm, etc.

[0009] 2. Disadvantages of existing related technologies

[0010] In the field of modern military communication, despite numerous studies aimed at improving the efficiency and reliability of communication systems, existing technologies still face limitations and challenges. Here are some specific issues:

[0011] 1) Limitations of static priority scheduling:

[0012] Traditional static priority scheduling methods cannot dynamically adjust task priorities based on real-time task requirements and network states, leading to low-priority tasks occupying resources for long periods under high load conditions, affecting system response speed. Although Han Songyue et al. (5G Mobile Communication Technology Military Application Research) proposed an overall framework, it mainly focuses on strategic-level design and lacks specific implementation of real-time scheduling in dynamic environments.

[0013] 2) Deficiencies in multi-objective optimization:

[0014] Existing research is weak in multi-objective optimization, especially in dynamic environments, making it difficult to find the best balance point among multiple competing objectives. Yi Shan et al. (Target Detection and Distribution Method and Device Based on Multi-agent Reinforcement Learning) proposed the MADDPG algorithm to improve convergence speed, but there is insufficient discussion on multi-objective optimization and resource allocation strategies in actual battlefield environments.

[0015] 3) Limited anti-interference capability:

[0016] Existing communication systems lack sufficient anti-interference capability in complex electromagnetic environments and enemy electronic warfare interference, which can lead to communication link interruptions or information transmission failures. Yin Shengpeng et al. (Space-Time Dynamic Deployment Method for Air-Ground Ad Hoc Communication Network Based on Deep Reinforcement Learning) considered the dynamic deployment of communication networks, but there is still room for improvement in enhancing anti-interference capability in complex electromagnetic environments.

[0017] 4) Low resource utilization:

[0018] Traditional scheduling methods fail to fully utilize limited computing resources and network bandwidth, resulting in resource waste and affecting task completion rate. Xia Yi et al. (Multi-agent Knowledge Reasoning Method Based on Deep Reinforcement Learning) focused on long-distance reasoning problems in large-scale knowledge graphs and did not involve real-time scheduling and resource management in military communications. SUMMARY

[0019] Aiming at the problems of static priority scheduling, multi-objective optimization, anti-interference ability and insufficient resource utilization in military communication systems, a military communication dynamic priority scheduling method based on deep reinforcement learning is proposed. It is used to optimize the communication efficiency between ground equipment, air information transfer equipment and combat unmanned aerial vehicles, to improve the speed of information transmission and reduce communication delay. Based on the existing related technology, the invention realizes significant innovation: it can dynamically adjust the task priority according to real-time task demand and network state. Multi-objective optimization is discussed to find the best balance point of multiple competitive objectives in a dynamic environment. By optimizing the communication path and resource allocation, the flexibility and response speed of the system are improved. These innovations comprehensively improve the performance and reliability of the military communication system, providing a more flexible, efficient and reliable solution.

[0020] The technical solution of the invention is:

[0021] A military communication dynamic priority scheduling method based on deep reinforcement learning, including the following steps:

[0022] Step 1. Dynamic priority scheduling

[0023] According to the relevant information of military tasks, real-time adjustment is carried out using the dynamic priority formula, and state space definition is carried out respectively: each state contains sufficient information about the environment, so that the agent can make decisions; action space design: the agent selects the optimal action according to the current state to optimize the communication path and resource allocation; reward mechanism construction: guide the agent to learn the optimal strategy; Q table initialization: Q table is used to track state, action and its expected reward; select the initial state: select a random initial state, a fixed initial state or a diversified initial state; select action: the agent uses the ε-greedy strategy to select action, balancing exploration and utilization; set the maximum number of iterations; execute action: the agent sends the selected action to the environment, actually executes the action, and observes the feedback of the environment, including the new state and the reward, so as to update the Q value in the Q table;

[0024] Step 2. Combination with multi-objective optimization

[0025] Get the current state: the agent first gets the state of the current environment; select action: according to the current state, the agent uses the ε-greedy strategy to select an action; execute action: the agent sends the selected action to the environment, actually executes the action; observe feedback: after executing the action, the agent observes the feedback of the environment, including the new state and the reward; update Q table: according to the new state and the reward, the agent updates the Q value in the Q table; complete the military communication dynamic priority scheduling.

[0026] Further, the dynamic priority formula is as follows:

[0027]

[0028] P dynamic : is the dynamic priority, which represents the power consumption of the device or system in a certain situation, related to the static power consumption and other parameters;

[0029] P static : is the static priority, which represents the power consumption of the device in standby or inactive state;

[0030] α, β: weight coefficients, α is used to balance the contribution of static power consumption and dynamic adjustment part in total power consumption, β represents the degree of influence of dynamic adjustment part on total power consumption, their values are usually between 0 and 1;

[0031] T due , T wait : time parameters, T due represents a specific expiration time or predetermined time limit; T wait usually represents the time the system waits or has spent.

[0032] Further, the update formula of Q value in Q table is as follows:

[0033]

[0034] Q(s, a): Q value of taking action a in state s, which represents the total expected reward obtained by performing action a in state s;

[0035] η: learning rate, which controls the weight of new information and old information when updating Q value, the value is between (0, 1], higher learning rate means faster adjustment of Q value;

[0036] r is the reward obtained immediately after performing action a;

[0037] γ is the discount factor, ranging from [0, 1], which represents the relative importance of future rewards, which is used to calculate the long-term reward from the next state s';

[0038] represents the maximum Q value of all possible actions a' in the subsequent state s', which reflects the value of the best strategy in the new state.

[0039] Further, the dynamic priority scheduling is implemented as follows:

[0040] Step 1.1) State space definition: the state space includes communication delay, energy consumption, task completion rate, system cost and network bandwidth information; specifically, the state space can be represented as a vector S = [d, e, c, n, b], where:

[0041] d: communication delay; e: energy consumption; c: task completion rate; n: system cost; b: network bandwidth and delay;

[0042] Step 1.2) Action space design: The action space includes task offloading, resource allocation, and priority adjustment; specifically, the action space can be represented as a vector A = [a1, a2, a3], where: a1: task offloading; a2: resource allocation; a3: priority adjustment;

[0043] Step 1.3) Reward mechanism construction: The reward function is designed based on multiple factors such as communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction, guiding the agent to learn the optimal strategy; the specific reward function can be represented as:

[0044] R = w1·Δd + w2·Δe + w3·Δc + w4·Δn

[0045] w1, w2, w3, w4: weight coefficients, used to balance the importance of each target;

[0046] Δd, Δe, Δc, Δn: respectively, the communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction;

[0047] Step 1.4) Q table initialization: In the initial state, all Q values are set to zero, indicating that the agent knows nothing about the environment; as the training progresses, the Q table is constantly updated, and the agent gradually learns the optimal strategy;

[0048] Step 1.5) Select initial state: Select a random initial state, a fixed initial state, or a diversified initial state to ensure that the agent can adapt to various situations;

[0049] Step 1.6) Select action: The agent generates a random number ∈ between 0 and 1:

[0050] If ∈ < ∈threshold, randomly select an action;

[0051] If ∈ ≥ ∈threshold, select the action with the maximum Q value corresponding to the state in the current Q table;

[0052] where ∈threshold is the set ∈ threshold;

[0053] Step 1.7) Maximum number of iterations: Set the maximum number of iterations Nmax to prevent the algorithm from falling into an infinite loop, ensure that the training is completed within a limited time, and reach an acceptable performance level;

[0054] Step 1.8) Perform action: The agent sends the selected action to the environment, actually performs the action, and observes the feedback of the environment, including the new state and reward, thereby updating the Q value in the Q table;

[0055] Further, the combination with multi-objective optimization is specifically implemented as follows:

[0056] Step 2.1) Obtain the current state: the agent first obtains the state S of the current environment; the state S contains sufficient information about the environment, so that the agent can make decisions; in the military communication system, the state includes communication delay, energy consumption, task completion rate, system cost and network bandwidth information;

[0057] Step 2.2) Select action: according to the current state S, the agent selects an action A using the ε-greedy strategy;

[0058] Step 2.3) Perform action: the agent sends the selected action A to the environment and actually performs the action; in the military communication system, the action includes task offloading, resource allocation and priority adjustment;

[0059] Step 2.4) Observe feedback: after performing the action, the agent observes the feedback of the environment, including the new state S' and the reward R; the new state is the state of the environment after performing the action, and the reward is calculated according to the pre-defined reward function, reflecting the effect of the current action; in the military communication system, the reward is calculated based on the reduction of communication delay, the reduction of energy consumption, the increase of task completion rate and the reduction of system cost;

[0060] Step 2.5) Update Q table:

[0061] According to the new state s' and the reward r, the agent updates the Q value in the Q table; complete the dynamic priority scheduling of the military communication.

[0062] The beneficial effects of the present application are:

[0063] 1) Improve the flexibility and response speed of the communication system: through the dynamic priority scheduling technology, the system can dynamically adjust the task priority according to the real-time task demand and network state, ensure that the high-priority task is processed in time, effectively reduce the communication delay, and improve the system response speed.

[0064] 2) Optimize communication delay and bandwidth utilization: by introducing the deep reinforcement learning algorithm, the system can automatically adjust the communication path and resource allocation, reduce unnecessary data retransmission and delay, and ensure the transmission of critical information BRIEF DESCRIPTION OF DRAWINGS

[0065] Figure 1 The flowchart is specifically designed for the method of the present application;

[0066] Figure 2 A system architecture and component relationship modeling diagram for the present invention. DETAILED DESCRIPTION

[0067] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments are implemented on the basis of the technical solutions of the present invention, and detailed implementation methods and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0068] A military communication dynamic priority scheduling method based on deep reinforcement learning, as shown in Figure 1 、 2 , includes the following steps:

[0069] 1. Dynamic priority scheduling algorithm

[0070] 1.1 Dynamic priority calculation formula

[0071] The algorithm can dynamically adjust the priority of tasks according to real-time task requirements and network status in complex battlefield environments, ensuring that critical tasks have priority to obtain resources and services.

[0072] The dynamic priority scheduling algorithm needs an effective priority calculation formula to guide the adjustment of priority. In this study, we will use the dynamic priority formula to make real-time adjustments based on the relevant information of the task (such as waiting time, deadline):

[0073]

[0074] P dynamic : is the dynamic priority, which represents the power consumption of the device or system in a certain situation, related to static power consumption and other parameters.

[0075] P static : static priority, representing the power consumption of the device in standby or inactive state.

[0076] α, β: weight coefficients, α is used to balance the contribution of static power consumption and dynamic adjustment part in total power consumption, β represents the influence degree of dynamic adjustment part on total power consumption, their values are usually between 0 and 1.

[0077] T due , T wait : time parameters, T due represents a specific expiration time or predetermined time limit. T wait usually represents the time the system is waiting for or has spent.

[0078] 1.2 Specific implementation steps

[0079] 1) State space definition: The state space includes information such as communication delay, energy consumption, task completion rate, system cost, and network bandwidth. Each state contains enough information about the environment so that the agent can make decisions. Specifically, the state space can be represented as a vector S = [d, e, c, n, b], where:

[0080] d: communication delay; e: energy consumption; c: task completion rate; n: system cost; b: network bandwidth and delay.

[0081] 2) Action space design: The action space includes task offloading, resource allocation, and adjusting priorities. The agent selects the optimal action based on the current state to optimize the communication path and resource allocation. Specifically, the action space can be represented as a vector A = [a1, a2, a3], where:

[0082] a1: task offloading (selecting offloading target, offloading proportion, offloading priority)

[0083] a2: resource allocation (bandwidth allocation, computing resource allocation, storage resource allocation)

[0084] a3: adjusting priorities (task priority adjustment, resource priority adjustment, path priority adjustment)

[0085] 3) Reward mechanism construction: The reward function is designed based on multiple factors such as the amount of communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction, guiding the agent to learn the optimal strategy. The specific reward function can be represented as:

[0086] R = w1·Δd + w2·Δe + w3·Δc + w4·Δn

[0087] w1, w2, w3, w4: weight coefficients, used to balance the importance of each target.

[0088] Δd, Δe, Δc, Δn: respectively the amount of communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction.

[0089] 4) Q-table initialization: The Q-table is used to track states, actions, and their expected rewards. In the initial state, all Q-values are set to zero, indicating that the agent knows nothing about the environment. As training progresses, the Q-table is constantly updated, and the agent gradually learns the optimal strategy.

[0090] 5) Selection of initial state: The selection of the initial state has a significant impact on the learning efficiency of the agent and the performance of the final strategy. Random initial state, fixed initial state, or diversified initial state can be chosen to ensure that the agent can adapt to various situations.

[0091] 6) Select Action: The agent uses an ε-greedy policy to select an action, balancing exploration and exploitation, ensuring both the discovery of new, potentially more advantageous actions, and the full exploitation of known best actions. Specifically, the agent generates a random number ∈ between 0 and 1:

[0092] If ∈ < ∈threshold, a random action is selected.

[0093] If ∈ ≥ ∈threshold, the action with the maximum Q-value in the current Q-table for the corresponding state is selected.

[0094] Where ∈threshold is a set ∈ threshold value.

[0095] 7) Maximum Number of Iterations: Set the maximum number of iterations Nmax to prevent the algorithm from falling into an infinite loop, ensuring that training is completed within a limited time and an acceptable performance level is reached.

[0096] 8) Execute Action: The agent sends the selected action to the environment, actually executes the action, and observes the feedback from the environment, including the new state and reward, thereby updating the Q-values in the Q-table.

[0097] 2. Combination of Multi-objective Optimization

[0098] To find the best balance point among multiple competing objectives, the invention introduces a multi-objective optimization algorithm. Combining Pareto optimization and deep reinforcement learning, dynamic priority scheduling is achieved in military communication systems to solve the problem that traditional static priority scheduling methods cannot adapt to changes in battlefield environment.

[0099] Basic framework of the combination of Pareto optimization and deep reinforcement learning

[0100] 1) Get Current State: The agent first obtains the current state S of the environment. The state S contains sufficient information about the environment, enabling the agent to make decisions. In military communication systems, the state may include information such as communication delay, energy consumption, task completion rate, system cost, and network bandwidth.

[0101] 2) Select Action: Based on the current state S, the agent uses an ε-greedy policy to select an action A. (As described in 6 of 1.2)

[0102] 3) Execute Action: The agent sends the selected action A to the environment and actually executes it. In military communication systems, actions may include task offloading, resource allocation, and adjusting priorities. For example, the agent can choose to offload tasks to an edge server or adjust the priority of a certain task.

[0103] 4) Observation feedback: After performing the action, the agent observes the feedback of the environment, including the new state S' and the reward R. The new state is the environment state after performing the action, and the reward is calculated according to the predefined reward function, reflecting the effect of the current action. In military communication systems, the reward can be calculated based on multiple factors such as communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction.

[0104] 5) Update Q table:

[0105] According to the new state s' and the reward r, the agent updates the Q value in the Q table. The specific update formula is as follows:

[0106]

[0107] Q(s,a): Q value of action a in state s, which represents the total expected reward obtained by performing action a in state s.

[0108] η: learning rate, controls the weight of new information and old information when updating Q value, value in (0,1], higher learning rate means faster adjustment of Q value.

[0109] r is the reward obtained immediately after performing action a.

[0110] γ is the discount factor, ranging from [0,1], representing the relative importance of future rewards, which is used to calculate the long-term reward from the next state s'.

[0111] represents the maximum Q value of all possible actions a' in the subsequent state s', reflecting the value of the best strategy in the new state.

[0112] Complete military communication dynamic priority scheduling.

[0113] The combined algorithm significantly improves the flexibility and response speed, wide utilization rate, resource utilization rate, and anti-interference ability of the military communication system. Although previous studies have applied deep reinforcement learning to resource management and task scheduling, such as the MADDPG algorithm proposed by Yi Shan et al. and the work on air-ground self-organizing communication networks by Yin Shengpeng et al., the present invention combines Pareto optimization and deep reinforcement learning, demonstrating innovation and uniqueness in multi-objective optimization and practical battlefield applications, and its performance gain is verified through experiments.

[0114] In view of the shortcomings of the prior art, the present invention solves the following technical problems:

[0115] 1) Improve the flexibility and response speed of the communication system:

[0116] Through dynamic priority scheduling technology, the priority of tasks is adjusted in real time according to their urgency and importance, ensuring that critical tasks can obtain resources and services first. Compared with the research of Han Songyue et al. (5G mobile communication technology military application research), this invention emphasizes real-time scheduling and resource allocation in dynamic environments.

[0117] 2) Optimize communication latency and bandwidth utilization:

[0118] By introducing deep reinforcement learning algorithms, the communication path and resource allocation are automatically adjusted to reduce unnecessary data retransmission and latency, ensuring that critical information can quickly reach the destination. Compared with the research of Yin Shengpeng et al. (Space-ground self-organizing communication network space-time dynamic deployment method based on deep reinforcement learning), this invention not only focuses on the dynamic deployment of the network, but also further optimizes the communication path and resource allocation using deep reinforcement learning algorithms, improving the flexibility and response speed of the system.

[0119] 3) Improve resource utilization:

[0120] Combined with real-time monitoring data, bandwidth and computing resources are reasonably allocated to ensure that high-priority tasks obtain resources first, while minimizing unnecessary energy consumption and prolonging the system's endurance. Unlike the research of Xia Yi et al. (Multi-agent knowledge reasoning method based on deep reinforcement learning), this invention emphasizes improving the system's anti-interference ability and resource utilization in complex electromagnetic environments, which is crucial for the survivability and efficiency of military communication systems.

[0121] 4) Enhance the system's anti-interference ability and survivability:

[0122] Through intelligent scheduling strategies, potential security threats are learned and identified, and security policies are automatically adjusted to improve the system's anti-interference ability and survivability. This invention explores how to find the best balance point between multiple competing goals in a dynamic environment, especially in the optimization of communication latency, energy consumption, and task completion rate, providing new solutions to make up for the shortcomings in the research of Yi et al. (Target detection and allocation method and device based on multi-agent reinforcement learning).

[0123] The above-described embodiments only express one embodiment of the present invention, which is described in detail and specifically, but it cannot be understood as a limitation on the scope of the patent. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present invention, several modifications and improvements can be made, which are within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be subject to the appended claims.

Claims

1. A method for dynamic priority scheduling of military communications based on deep reinforcement learning, characterized in that, The steps include: Step 1. Dynamic priority scheduling According to the relevant information of military tasks, real-time adjustment is made using the dynamic priority formula, and the state space is defined respectively: each state contains sufficient information about the environment, enabling the agent to make decisions; Action space design: the agent selects the optimal action based on the current state to optimize communication path and resource allocation; reward mechanism construction: guide the agent to learn the optimal strategy; Q table initialization: Q table is used to track state, action and its expected reward; select initial state: select random initial state, fixed initial state or diversified initial state; select action: the agent uses the ε-greedy strategy to select action, balancing exploration and utilization; set the maximum number of iterations; execute action: the agent sends the selected action to the environment, actually executes the action, and observes the feedback of the environment, including the new state and reward, so as to update the Q value in the Q table; Step 2. Combined with multi-objective optimization Get the current state: the agent first gets the state of the current environment; select action: according to the current state, the agent uses the ε-greedy strategy to select an action; execute action: the agent sends the selected action to the environment and actually executes the action; observe feedback: after executing the action, the agent observes the feedback of the environment, including the new state and reward; update Q table: according to the new state and reward, the agent updates the Q value in the Q table; complete the dynamic priority scheduling of military communication.

2. The method of claim 1, wherein, The dynamic priority formula is as follows: P dynamic : is a dynamic priority, representing the power consumption of the device or system in a certain situation, in relation to static power consumption and other parameters; P static : static priority, indicating power consumption when the device is in standby or inactive state; α, β: weight coefficients, α is used to balance the contribution of static power consumption and dynamic adjustment part in total power consumption, β represents the influence degree of dynamic adjustment part on total power consumption, their values are usually between 0 and 1; T due , T wait : time parameter, T due represents a specific expiration time or a predetermined time limit; T wait usually indicates the time the system waits for or has spent.

3. The method of claim 1, wherein, The update formula of Q value in Q table is as follows: Q(s, a): the Q value of action a in state s, which represents the total expected reward obtained by executing action a in state s; η: learning rate, which controls the weight of new information and old information when updating Q value, the value is between (0, 1], higher learning rate means faster adjustment of Q value; r is the reward obtained immediately after executing action a; γ is the discount factor, ranging from [0, 1], representing the relative importance of future rewards, which is used to calculate the long-term reward from the next state s'; represents the maximum Q-value of all possible actions a' in the subsequent state s', reflecting the value of the optimal policy in the new state.

4. The method of claim 1, wherein, The specific implementation of dynamic priority scheduling is as follows: Step 1.1) State space definition: the state space includes communication delay, energy consumption, task completion rate, system cost and network bandwidth information; specifically, the state space can be represented as a vector S = [d, e, c, n, b], where: d: communication delay; e: energy consumption; c: task completion rate; n: system cost; b: network bandwidth and delay; Step 1.2) Action space design: the action space includes task offloading, resource allocation and priority adjustment; specifically, the action space can be represented as a vector A = [a1, a2, a3], where: a1: task offloading; a2: resource allocation; a3: priority adjustment; Step 1.3) Reward mechanism construction: The reward function is designed based on multiple factors such as communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction, guiding the agent to learn the optimal strategy; The specific reward function can be expressed as: R = w1·Δd + w2·Δe + w3·Δc + w4·Δn w1, w2, w3, w4: weight coefficients, used to balance the importance of each target; Δd, Δe, Δc, Δn: respectively for communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction; Step 1.4) Q table initialization: In the initial state, all Q values are set to zero, indicating that the agent knows nothing about the environment; As the training progresses, the Q table is constantly updated, and the agent gradually learns the optimal strategy; Step 1.5) Select initial state: Select a random initial state, a fixed initial state, or a diversified initial state to ensure that the agent can adapt to various situations; Step 1.6) Select action: The agent generates a random number ∈ between 0 and 1: If ∈ < ∈threshold, randomly select an action; If ∈ ≥ ∈threshold, select the action with the maximum Q value corresponding to the current state in the Q table; Where ∈threshold is the set ∈ threshold value; Step 1.7) Maximum number of iterations: Set the maximum number of iterations Nmax to prevent the algorithm from falling into an infinite loop, ensuring that the training is completed within a limited time and reaching an acceptable performance level; Step 1.8) Execute action: The agent sends the selected action to the environment, actually executes the action, and observes the feedback from the environment, including the new state and the reward, thereby updating the Q value in the Q table.

5. The method of claim 1, wherein, The combination with multi-objective optimization is implemented as follows: Step 2.1) Get current state: The agent first obtains the state S of the current environment; The state S contains sufficient information about the environment, enabling the agent to make decisions; In military communication systems, the state includes communication delay, energy consumption, task completion rate, system cost, and network bandwidth information; Step 2.2) Select action: Based on the current state S, the agent uses the ε-greedy strategy to select an action A; Step 2.3) Execute action: The agent sends the selected action A to the environment and actually executes the action; In military communication systems, actions include task offloading, resource allocation, and priority adjustment; Step 2.4) Observe feedback: After executing the action, the agent observes the feedback from the environment, including the new state S' and the reward R; The new state is the environment state after executing the action, and the reward is calculated based on the predefined reward function, reflecting the effect of the current action; In military communication systems, the reward is calculated based on multiple factors such as communication delay reduction, energy consumption reduction, task completion rate improvement, and system cost reduction; Step 2.5) Update Q table: Based on the new state s' and the reward r, the agent updates the Q value in the Q table; Complete the dynamic priority scheduling of military communication.