Ad hoc network service scheduling mechanism implementation method based on multi-queue CQF
By using the multi-queue CQF model and GSR algorithm, the problems of low latency and high reliability in concurrent access of multiple services in UAV ad hoc networks are solved, realizing low latency and reliable communication in emergency communication environments, and improving network performance and resource utilization.
Patent Information
- Application Number
- CN202511107576.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2026-01-09
AI Technical Summary
The existing queue scheduling mechanism of the MAC layer of UAV ad hoc networks is difficult to meet the requirements of critical services for low latency and high reliability when multiple services access concurrently. It is particularly ineffective in high dynamic and sudden traffic scenarios, and has a large computational overhead.
By employing a multi-queue CQF model and the deep reinforcement learning-based time resource scheduling algorithm GSR, different queue offset values and time slots are configured for different types of service flows. Combined with BE queues, the forwarding of best-effort flows is taken into account, thereby optimizing network resource utilization.
It achieves low latency and reliable communication in emergency communication environments, improves the overall network performance and resource utilization, and solves the packet loss problem caused by dead time.
Smart Images

Figure CN121309501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for implementing a self-organizing network service scheduling mechanism based on multi-queue CQF, belonging to the field of communication technology. Background Technology
[0002] In the MAC layer of UAV ad hoc networks, queue scheduling mechanisms play a crucial role in QoS assurance as a key means of managing wireless channel access and resource allocation. However, facing channel conflicts and network congestion caused by concurrent access of multiple services, the traditional best-effort (BE) mechanism is no longer sufficient to meet the stringent requirements of critical services for low latency and high reliability. Therefore, research on MAC layer QoS scheduling mechanisms for multi-service scenarios, especially scheduling methods combining priority control and time slot allocation strategies, has become a key direction for improving system performance. In recent years, Time-Sensitive Networking (TSN) has proposed several key technologies for real-time communication, among which the Time-Aware Shaper (TAS) model and the Cyclic Queuing and Forwarding (CQF) model are the most representative mechanisms. TAS configures precise time windows for different priority services, allowing high-priority traffic to exclusively occupy transmission resources during specific time periods, thereby effectively avoiding conflicts and preemption and improving latency controllability; while CQF, through periodic scheduling of data packets in the buffer queue and in conjunction with a double buffer structure, achieves an end-to-end deterministic forwarding path, significantly enhancing network transmission stability and predictability.
[0003] Therefore, some scholars have begun to shift their research focus to the deterministic forwarding queue model of TSN technology, attempting to leverage the advantages of its standard queue model in deterministic forwarding to improve the predictability and stability of data transmission in wireless networks.
[0004] Regarding scheduling algorithms for the model, a time-varying flow scheduling scheme based on statistical multiplexing has been proposed, achieving efficient utilization of link resources under TSN technology through a combination of heuristic algorithms and reinforcement learning. However, the delay determinism of this planning algorithm is insufficient, resulting in high complexity in practical deployment. A hybrid scheduling mechanism combining TAS and CQF has been proposed, improving the system's scheduling success rate and resource utilization by determining the minimum scheduling slot, adjusting the flow sampling period, and employing an even-odd mapping flow classification strategy. However, this model relies on high-precision clock synchronization, posing challenges in complex dynamic environments. A deep reinforcement learning-assisted time-frequency resource scheduling algorithm, the DTF algorithm, has been designed, jointly allocating time intervals and frequency resources on a flow-by-flow basis, effectively learning the relationship between different time intervals and capacity utilization. A hybrid scheduling algorithm has also been designed, utilizing slot awareness to insert non-time-sensitive flows into the remaining slots of the even-odd queue at the outgoing port. The algorithm uses simulated annealing to generate scheduling strategies, avoiding getting trapped in local optima and obtaining a globally optimal solution.
[0005] Regarding queue models, a CQF-based model has been proposed, which effectively improves the scheduling performance and network resource utilization of periodically triggered traffic under high load through scenario-based parameter optimization algorithms and integer linear programming traffic scheduling algorithms. However, the flexibility of this algorithm is limited by the fixed time slot length, and its adaptability to burst traffic is insufficient. Furthermore, parameter tuning and implementation complexity are high in practical applications. Currently, a priority-based CQF queue has been proposed to address different service quality requirements in the network. By adjusting the queue length horizontally and increasing the number of queues vertically, finer-grained latency control and better adaptability to burst traffic are achieved. However, while this model improves on priority, it does not consider the arrival time of data streams, resulting in poor performance, especially when facing highly dynamic and bursty traffic. Additionally, its queue management complexity is high, leading to significant computational overhead in practical applications. Summary of the Invention
[0006] The purpose of this invention is to address the shortcomings and deficiencies of existing technologies by proposing a method for implementing a self-organizing network service scheduling mechanism based on multi-queue CQF. This method addresses the hybrid flow scheduling problem in UANET service flow transmission by designing an MQSM service scheduling mechanism. Firstly, by introducing and building a multi-queue CQF model and scheduling strategy, the low-latency and deterministic transmission requirements of hybrid flow scheduling are met, thus solving the dead-time problem. Specifically, during the forwarding of hybrid service flows, the multi-queue CQF model uses different offset values for different types of service flows to achieve deterministic transmission and low-latency communication, and also considers best-effort flow forwarding through BE queues. Secondly, a time resource scheduling algorithm based on deep reinforcement learning, namely the GSR algorithm, is proposed and adapted to the built queue model. This algorithm mainly improves the overall network performance through the design of a time-resource utilization matrix, convolutional neural networks, and loss functions. Simulation experiments have verified that the GSR algorithm outperforms other time resource scheduling algorithms in terms of stability, deployment speed, service flow scheduling capability, and time slot resource utilization. It meets the performance requirements of mixed service flow scheduling in emergency communication environments, and achieves low latency and reliable communication for the overall network.
[0007] The technical solution adopted by this invention to solve its technical problem is: an implementation method of a self-organizing network service scheduling mechanism based on multi-queue CQF. This method establishes a multi-queue CQF, where the forwarding queues support deterministic shared network transmission of bursty and periodic flows. By reasonably allocating queue offset values and matching transmission time slots for service flows, the multi-queue CQF configures M queues (M≥3) on the queue scheduling model of each UAV node. Each queue is equipped with a gating mechanism, which periodically opens or closes according to a preset rhythm. Within each time slot, only one queue is allowed to be in the active state for data transmission and reception, while the gating of other queues remains closed to prevent conflicts between service flows. If a queue is used for data transmission, its state is denoted as Q. i =0, 0 < i < M-1, if used for receiving, then its queue state is set to Q. j =1, i≠j, in each time slot T σ Within the queue, the gating state switches according to the set scheduling strategy, which is defined as follows:
[0008]
[0009] Based on the above queue control strategy, the receiving logic for the service flow between two adjacent hop nodes was formulated. This logic depends on the queue offset setting, queue Q. j offset value φ j This indicates that in the current time slot, the data is transmitted from the corresponding transmission queue Q. σ%M to queue Q j Distance, sending queue Q σ%M The offset value is defined as φbase For the queue receiving the service stream, its offset is based on the φ of the sending queue. base Using this as a reference point, and incrementing by 1 sequentially to form a cyclic offset pattern;
[0010] Based on the queue offset settings described above, and according to the transmit / receive status in each time slot, the offset is characterized by a mathematical expression, namely:
[0011]
[0012] To avoid sudden business flows within the dead time T d When packet loss occurs during internal transmission, burst flows, when selecting the receiving queue of the next-hop drone node, must prioritize queues with an offset value of not less than 2, i.e., satisfying φ. r ≥2, thus ensuring that the business flow can be successfully received, i.e., the following formula holds true:
[0013]
[0014] This satisfies formula (21). The established optimization conditions, multi-queue CQF, by extending the time slot of the received service stream, enable burst service streams sent within the dead time to still be received by the queues with larger offset values in the downstream nodes, effectively improving the reliability of its transmission process.
[0015] Furthermore, this invention optimizes the forwarding efficiency of mixed service flows and improves the utilization rate of time slot resources. t ,include:
[0016] (1) For bursty flows from upstream nodes, in order to avoid packet loss, a receiving queue with a specific offset value must be assigned to it, that is, an offset value must be selected. The queue is used as the next-hop receiving queue, and the offset value needs to be reasonably configured according to the traffic characteristics, latency requirements and anti-jitter capabilities of the burst flow.
[0017] (2) After the burst stream is successfully received at the first hop, its subsequent forwarding can be completed using a circular queue mechanism. To shorten the total delay in multi-hop transmission, a fast forwarding mechanism can be introduced, that is, selecting an offset value for each subsequent hop. The queue receives data, thereby ensuring that the transmission delay between every two hop nodes does not exceed one time slot length, thus improving traffic forwarding efficiency.
[0018] (3) To rationally allocate queue resources and avoid periodic flows occupying critical resources on which burst flows depend, appropriate offset values should be configured for periodic flows based on the latency and jitter requirements of each service flow. Within each hop of the UAV node, the queue offset range that a periodic flow can select is:
[0019] (4) For service flows with low latency requirements in the network (i.e. best-effort flows), or when the current CQF receive queue is full, they can be stored in the best-efforts (BE) queue. The BE queue does not set a queue offset value, but adopts an opportunistic transmission strategy. That is, in subsequent transmissions, if the current time slot still has remaining bandwidth after the CQF send queue has completed the predetermined data transmission, the BE queue can be scheduled to send.
[0020] Furthermore, according to the present invention, within time slot T0, terminal N1 simultaneously transmits periodic and burst streams to N2 through four UAVs in the UANET. The system sets the queue number to M=3. To achieve reliable transmission of burst streams, an offset value is selected for the burst streams within time slot T0. The queue, i.e., according to formula (23). Queue Q2 is selected to receive the service flow, thus satisfying formula (24). After two time slots, the burst flow in time slot T2 needs to be selected. The queue is used to reduce the cumulative latency caused by subsequent hops, while the periodic stream can choose an offset value based on its latency tolerance. Transmitting data through queues satisfies service quality requirements and avoids interference with high-priority traffic. If periodic streams do not have high latency requirements, more resource queues required by high-priority services can be released by allocating larger offset values.
[0021] Furthermore, a multi-queue CQF and its scheduling strategy suitable for mixed service scenarios were introduced, which effectively improved the network's comprehensive scheduling capability for periodic and bursty flows. However, model design alone is still insufficient to fully guarantee the network's resource utilization efficiency. In order to further improve the overall network performance, formula (20) was implemented. To optimize the GSR algorithm, it first constructs a two-dimensional time-resource matrix and merges it with the action matrix to form a three-channel input. It then extracts features and outputs the corresponding Q-values through a fusion convolutional network combining Ghost convolution and SRResNet. Subsequently, the expert guidance module evaluates and adjusts the selected actions. After the actions are executed, the system obtains the status and reward information from the environment and continuously optimizes the network parameters through an experience replay mechanism, thereby improving the efficiency of resource scheduling decisions.
[0022] Furthermore, in multi-queue CQF, time slots and queue resources are not completely independent. A two-dimensional time-resource utilization matrix is adopted to jointly model time slots and queue resources, thereby optimizing the feature representation capability of the model from a global perspective.
[0023] Furthermore, the two-dimensional time-resource matrix uses rows to represent time slots and columns to represent the circular queues of each node, dynamically recording the resource occupancy of each queue under different time slots. This matrix serves as the state input channel for the agent, updating the resource usage status in real time as the business flow arrives, realizing adaptive migration of the state space. Each column can also be expanded to include additional information such as frequency, computing power, and link occupancy rate, thereby extending to a wider range of scheduling scenarios. In order to further improve the agent's ability to extract key features, each matrix element is a u×m vector to map the resource utilization relationship between different time slots and queues. For queues without direct association, the matrix adopts a zero-filling strategy to construct a complete p×p state matrix.
[0024] Effective effects:
[0025] 1. This invention addresses the hybrid flow scheduling problem in UANET service flow transmission by designing an MQSM service scheduling mechanism. It introduces and builds a multi-queue CQF model and scheduling strategy to meet the low-latency and deterministic transmission requirements of hybrid flow scheduling, thus solving the dead-time problem. During hybrid service flow forwarding, the multi-queue CQF model uses different offset values for different types of service flows to achieve deterministic transmission and low-latency communication, and also considers best-effort flow forwarding through BE queues.
[0026] 2. This invention proposes a time resource scheduling algorithm based on deep reinforcement learning, namely the GSR algorithm, which is adapted to the constructed queue model. This algorithm improves the overall network performance through the design of a time-resource utilization matrix, convolutional neural networks, and loss functions. Simulation experiments verify that the GSR algorithm outperforms other time resource scheduling algorithms in terms of stability, deployment speed, service flow scheduling capability, and time slot resource utilization. It meets the performance requirements of mixed service flow scheduling in environments such as emergency communication, achieving low latency and reliable communication for the overall network. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the UANET service flow transmission scenario of the present invention.
[0028] Figure 2 This is a schematic diagram of a traditional CQF (Content Quality Function).
[0029] Figure 3 This is a schematic diagram of the service flow forwarding rules of the present invention.
[0030] Figure 4 This is a schematic diagram illustrating the service flow scheduling constraint optimization problem of the present invention.
[0031] Figure 5 This is an example diagram of traditional CQF burst flow scheduling.
[0032] Figure 6This is an example diagram of the multi-queue CQF burst flow scheduling of the present invention.
[0033] Figure 7 This is a schematic diagram of the multi-queue CQF working mechanism of the present invention.
[0034] Figure 8 This is a schematic diagram of the GSR algorithm framework of the present invention.
[0035] Figure 9 This is a schematic diagram of the two-dimensional time-resource utilization matrix of the present invention.
[0036] Figure 10 This is a schematic diagram comparing the ordinary convolution and the ghost convolution of this invention.
[0037] Figure 11 This is a schematic diagram of residual learning in this invention.
[0038] Figure 12 This is a schematic diagram of the depth residual block of the present invention.
[0039] Figure 13 This is a schematic diagram of the fusion convolutional neural network structure of the present invention.
[0040] Figure 14 This is a schematic diagram of the training process of the scheduling algorithm of the present invention.
[0041] Figure 15 A diagram showing the comparison of total test times for the models.
[0042] Figure 16 A diagram showing the comparison of the total number of execution steps of the model.
[0043] Figure 17 A schematic diagram illustrating the impact of time slot length on the number of service flow scheduling requests.
[0044] Figure 18 A schematic diagram illustrating the impact of time slot length on time slot resource utilization.
[0045] Figure 19 A diagram illustrating the impact of the service flow cycle on the number of service flow scheduling tasks. Detailed Implementation
[0046] The invention will now be described in further detail with reference to the accompanying drawings.
[0047] To address the shortcomings of existing models and scheduling algorithms, this invention designs an MQSM mechanism that uses a multi-queue CQF model and the GSR algorithm to achieve efficient service flow scheduling. Results show that the GSR algorithm has good scheduling capabilities in mixed data flow scenarios, can rationally allocate time resources, and improves the overall network performance.
[0048] Figure 1This demonstrates various service flow transmission scenarios in UANET, with each scenario consisting of R drone nodes, including N drone terminal nodes at both the transmitting and receiving ends. i And the drone relay node S in the data stream forwarding area j To construct an abstract network topology model, a directed graph structure G = (V, L) is adopted, where V = N∪S consists of a set of terminal nodes N = {N0, N1...N...}. I} and relay node set S = {S0, S1...S} R-I The network is composed of two links. The set of links in the network is denoted as L, and each link connects two neighboring nodes v. x and v y And L={(v x ,v y )|v x ,v y ∈V,v x ≠v y In special scenarios such as disaster relief, traditional ground base stations may be damaged and unable to operate normally, leading to the paralysis of conventional communication systems. In such cases, UANET can autonomously network via drone nodes to temporarily construct an aerial multi-hop communication network, providing basic emergency communication support. Sending terminal nodes can access UANET to transmit various types of service data, such as voice calls, video transmissions, location requests, and emergency control commands, to the network (i.e., best-effort transmission streams, periodic streams, and burst streams in the diagram), and forward them to receiving terminal nodes multiple times. The set of service streams is denoted as F, where |F| represents the total number of service streams, and each service stream is represented by f. i ∈F is used for identification. After receiving various service flows, the drone node will transmit in its allocated TDMA time slot T. i Internally, various types of service data are scheduled and sent through a queuing forwarding mechanism at the MAC layer. Through this resource scheduling method, the system can effectively ensure low-latency transmission of service flows while rationally scheduling the transmission of other low-priority services. This enables multi-service collaboration and communication quality assurance in resource-constrained environments, improving the communication reliability and service capabilities of UANET in emergency scenarios. Therefore, the UAV relay nodes in the network adopt CQF as the queuing model for the MAC layer when forwarding data. The main symbol definitions of this invention are shown in Table 1.
[0049] Table 1 Definitions of main symbols in this invention
[0050]
[0051]
[0052] (1) Queue scheduling model
[0053] CQF model, such as Figure 2As shown, traditional CQF uses a two-queue model, where queues 1 and 2 jointly handle the reception and transmission of the service flow. Each queue is equipped with receive-side gating and transmit-side gating to control the entry and transmission of data. The on / off states of these gatings are uniformly scheduled by a gating list to ensure the transmission of the data flow. Assuming that time slots 1 and 2 form a scheduling cycle, the CQF model must meet the following requirements:
[0054] 1) Data stream transmission and reception are performed alternately. For example... Figure 2 As shown, in even-numbered time slots T0, queue Q1 opens the transmit gate but closes the receive gate, only transmitting data; queue Q2, conversely, only receives data. In the next odd-numbered time slot T0, the queue states are reversed, with Q1 only receiving data and Q2 only transmitting data. This alternating mode continues to operate, ensuring that the service flow is always transmitted in an orderly manner. When a period contains multiple consecutive time slots, the queues still alternate between transmit and receive modes according to the odd-even time slot pattern to maintain stable transmission efficiency.
[0055] 2) The transmission time of a service flow between two adjacent hops must be less than one time slot. In a real network, upstream and downstream nodes are interconnected via links to transmit data. To ensure that the data transmission delay between adjacent nodes is deterministic and controlled, CQF requires the upstream node to transmit data within the current time slot T. i The receiving node must complete the transmission of all data packets buffered in the previous time slot, while the downstream node must complete the reception of all data packets within the same time slot. Data stored by the receiving node must wait until the next time slot T. i+1 Only then can it be sent to the next hop node.
[0056] like Figure 3 As shown, in time slot T0, node S1 performs batch data packet transmission through queue Q1. Simultaneously, downstream node S2, based on the dynamic buffering mechanism of queue Q2, completes zero-packet-loss data transmission at the receiving end within the same time slot. Constrained by the CQF gating policy, data received by node S2 can only be sent to the next-hop node in the next time slot T1. Within the same time slot, the next-hop node will receive this data. Through this strict time slot synchronization, CQF ensures that the transmission delay of the service flow between any two adjacent hops is strictly controlled to be less than the length of one time slot, thereby achieving deterministic transmission.
[0057] 3) Service flows have end-to-end deterministic delay and jitter boundaries. In the CQF mechanism, the end-to-end transmission delay of a service flow has strict upper and lower bounds (i.e., maximum and minimum delay limits). Assuming a single timeslot length is T, the end-to-end delay of the service flow depends on the hop count H of the network relay nodes and the timeslot length T. When the source node is in timeslot T... iWhen a data packet is sent at the start of the transmission, it is first written into the receive buffer queue of that node as it passes through each layer of nodes. Only after the time slot duration threshold is reached can the queue scheduler execute the forwarding instruction for forwarding. Therefore, the transmission delay of service flows between adjacent nodes is less than the length of a single time slot T.
[0058] Since the last-hop node only needs to complete data reception to the receiving terminal, when the data transmission delay from the last-hop node to the receiving terminal is exactly the slot length T, the CQF single-slot forwarding requirement is met, and the maximum end-to-end delay limit of the service flow is also determined, i.e.:
[0059] T Dmax = (H+1)×T (4)
[0060] When the service flow is in time slot T i When the data packet is sent from the source terminal, the delay at which it reaches the first-hop node is approximately zero, and the delay at the last-hop node to the receiver is also approximately zero. Therefore, the current system can minimize end-to-end transmission delay. However, because the hop-by-hop routing of services at relay nodes still needs to meet the strict slot alternation rules of the CQF mechanism, the transmission delay of service flows between adjacent nodes will not be less than the length of a single slot. Ultimately, the formula for the minimum end-to-end delay is:
[0061] T Dmin = (H-1)×T (5)
[0062] End-to-end jitter measures the range of fluctuation in the transmission delay of a data stream within a network, i.e., the difference in transmission delay between different data packets. Therefore, the maximum limit of jitter is the difference between the maximum and minimum end-to-end delay of the data stream, expressed as follows:
[0063] T Jmax =T Dmax -T Dmin =2×T (6)
[0064] (2) Business flow constraints
[0065] In the CQF queuing model, a service flow scheduling cycle typically includes multiple service flow transmission time slots. Therefore, a reasonable method is needed to schedule data flows to avoid packet loss due to insufficient capacity of a single time slot, thereby ensuring network service quality. Thus, properly planning the time slot mapping and adaptation of service flows at the terminal side is crucial for improving system performance.
[0066] Figure 4 This paper describes a service flow scheduling constraint optimization problem under the CQF model using an example. The network topology is based on... Figure 1The simplified network consists of three terminal nodes N1, N2, and N3, and one UAV relay node S1. Service flows 1 and 2 are periodic flows and are latency-sensitive. Service flow 1 has a data transmission period of 1600 μs, sending two data packets numbered 1.1 and 1.2 each time. Its transmission path starts from terminal N1, passes through node S1, and reaches terminal N3. Service flow 2 has a data transmission period of 3200 μs, sending one data packet numbered 2.1 each period. The data packet is sent from the source terminal N2, forwarded by node S1, and reaches the destination terminal N3. To ensure effective scheduling of periodic service flows, the scheduling period of the entire network is set to the least common multiple of the service flow periods, 3200 μs, and divided into 8 transmission time slots, each time slot being 400 μs, with each time slot accommodating a maximum of 2 data packets. Figure 4 (a) describes an unreasonable scheduling scheme: In time slot T0, terminal N1 sends data packets 1.1 and 1.2, while terminal N2 sends data packet 2.1. According to the CQF mechanism, these data packets will be forwarded by node S1 in time slot T1. However, since time slot T1 can only hold 2 data packets, data packet 2.1 is discarded after exceeding the capacity, reducing the reliability of the system. Figure 4 (b) demonstrates an improved scheduling scheme: adjusting the timing of N2's transmission of data packet 2.1 to time slot T1, allowing UAV segment S1 to forward the packet within that time slot. This ensures successful transmission of data packets 1.1, 1.2, and 2.1, avoiding packet loss, while also achieving a more balanced load across transmission time slots, thus improving system reliability and scheduling efficiency.
[0067] As seen in the above example, in the CQF model, to achieve reasonable scheduling of business flows and avoid packet loss, corresponding constraint optimization conditions need to be established. Therefore, it is first necessary to extract and precisely define the core characteristics of the business flow. These characteristics include business flow ID, business flow period, source address src, destination address dst, end-to-end latency requirement ζ, jitter requirement J, data packet pack, number of data packets num, and data packet size size. Therefore, the business flow can be represented as:
[0068]
[0069] To ensure the scheduling mechanism is fully adapted to network resources, the network resource attributes and functions need to be clearly defined. First, define the indicator function ω. ξ (f i Used to determine business flow f i In time slot T ξ The feasibility of transmission, if the service flow f i In time slot T ξ If the transmission is completed within the time limit, then ω ξ (f iIf the value is 1, then set it to 0.
[0070] Next, define the value function V. ξ (f i The scheduling benefits of traffic flows are evaluated using a formula (K,F), where K is the total number of transmission slots within the scheduling period, and V... ξ (f i The value function V is weighted by the number of successfully scheduled data flows and the average utilization rate of time slot resources. Therefore, the A(·) function is defined to measure the number of successfully scheduled data flows, while the B(·) function is defined to calculate the average resource utilization rate of time slots. Based on this, the value function V... ξ (f i (K,F) can be further expressed as:
[0071]
[0072] Where 0 < α1 and β1 < 1 are weighting factors used to adjust the proportion between the two optimization objectives of the number of successfully scheduled service flows and the utilization rate of time slot resources. In this invention, β1 is set to 0.1, meaning the system model focuses on the number of successfully scheduled service flows. When all data packets of a service flow are successfully scheduled, the function A(f i The value of A(f,K) is 1; otherwise, it is 0. Therefore, the function A(f i The specific expression for (K) can be described as follows:
[0073]
[0074] in, Representative business flow f i The total number of data packets transmitted within a scheduling period, f i j This represents the j-th data packet of the i-th service flow. B(K,F) is used to evaluate the average time slot resource utilization. Each time slot T... t Resource utilization rate U t It can be described as:
[0075]
[0076] Among them, time slot T t The capacity is defined as T t • C, whose value is affected by the time slot duration and link bandwidth. Therefore, the average resource utilization efficiency of each time slot in a scheduling cycle can be expressed as:
[0077]
[0078] In addition, considering the periodicity of service flows and the circular queue forwarding characteristics, the scheduling period in UANET can be set to the least common multiple of the periods of each service flow. In this way, the correspondence between service flows and time slots only needs to be determined within the first scheduling period, and the scheduling of subsequent periods can be directly reused. Therefore, the constraint condition of the scheduling period can be described as:
[0079] L p =LCM(F.period) (12)
[0080] In the CQF model, the least common multiple function LCM(·) is used to calculate the scheduling period. Since the transmission slot is the basic unit of the scheduling process, the length of the scheduling period must be divisible by the length of a single slot. Furthermore, to ensure that each service flow is accurately mapped to its corresponding transmission slot, the period of each service flow should also be divisible by the slot length. Therefore, the period constraint can be specifically expressed as:
[0081] L p %T=0,period i %T = 0 (13)
[0082] From formula (10), it can be derived that the maximum transmission time slot length should be the greatest common divisor of all service flow cycles. Therefore, the following constraints should be met when designing time slots:
[0083] T = GCD(F.period) (14)
[0084] GCD(·) is used to calculate the greatest common divisor. Given the limited time slot resources of network device interfaces, service flows should be rationally scheduled to avoid resource conflicts caused by traffic from different periods on the same interface. Therefore, the sending time of a service flow should not exceed its own period length, i.e.:
[0085] 0≤f i .tx≤period i (15)
[0086] In a packet-switched network architecture, the total transmission delay between adjacent relay nodes consists of four basic delay components: interface serialization delay, physical link propagation delay, node processing delay, and queue buffering delay. The first three components typically have predictable fixed value ranges. The core constraint is to ensure the determinism of the buffering delay through a queue admission control mechanism, thereby avoiding the risk of packet delay or loss due to queue overflow. Based on the deterministic transmission requirement, the sum of the four delay types must satisfy the upper limit constraint of the time slot. Furthermore, to fit the actual scenario, the clock synchronization accuracy η also needs to be considered, as shown in the following formula:
[0087]
[0088] Where B represents the link bandwidth. To ensure the high reliability requirements of service flow transmission, the CQF model requires that the total size of the service flow transmitted in each time slot does not exceed the maximum capacity of the time slot, that is:
[0089]
[0090] Wherein, the nth data packet of the mth service flow is denoted as This service flow contains a total of num within a scheduling cycle. m A data packet. When the data packet When the transmission is successful in the t-th time slot, the indicator function is activated. Conversely, it is 0. Based on clearly defined network resource constraints, the CQF model must also adhere to the deterministic requirements of end-to-end service flow transmission. Its core indicators cover latency extreme limits, latency fluctuation range constraints, and zero-packet-loss guarantees. Specifically, the deterministic requirement of end-to-end latency means that the actual transmission delay of each service flow must not exceed the set maximum allowable delay; jitter control requires that the deviation between the actual transmission delay of each data packet and the average delay of the service flow must be within a preset range. The zero-packet-loss constraint requires zero packet loss during service flow transmission to ensure high network reliability. Therefore, the above service flow transmission constraints can be specifically expressed as:
[0091]
[0092] Among them, Node i This represents the number of nodes that the i-th service flow passes through during transmission. This represents the end-to-end latency of the service flow data packets. Since each data packet can only be allocated to a unique transmission slot, a mapping constraint between data packets and transmission slots needs to be introduced:
[0093]
[0094] Apart from Figure 4 In addition to the periodic flow scheduling problem described in the article, traditional CQF also currently faces the problem of burst flow scheduling. Figure 5 This visually demonstrates the problems of the traditional CQF model in burst flow scheduling, among which... For business flow f i The selected receive queue offset value is 1 by default in traditional CQF. Traditional CQF requires that all traffic streams sent by terminal node N1 must be received by UAV relay node S1 within a time interval. However, due to transmission delays and unpredictable transmission times, burst streams may occur within the dead time T. dData frames are lost when they are sent within the Dead Time period, which cannot be contained within a single time interval. Dead Time occurs because the arrival time of a burst of traffic is later than the closing time of the Rx gating attached to the receive queue. Furthermore, in mixed traffic transmission scenarios, traffic may accumulate in the same queue, exceeding the capacity of a single time interval, thus affecting the reliable transmission of the stream.
[0095] This demonstrates the shortcomings of traditional CQF in bursty flow scheduling. To address packet loss caused by dead time, it is necessary to constrain the offset values in the queue model. First, for time slot T... l Burst traffic flow f in internal transmission i It needs to meet the following requirements:
[0096]
[0097] Among them, f i .offset represents the queue gating start time and the business flow f. i Clock offset between actual transmission times For business flow f i In two neighboring nodes v x and v y The transmission delay between them, where ε is the clock synchronization deviation. For time slot T l For the service flow that is sent at the end time, then f l • offset = T, that is, the time offset of the business flow is T, then the following formula can be obtained:
[0098]
[0099] The contradiction between formulas (17) and (18) indicates that when the queue offset is 1, a dead time T will definitely occur within the transmission time slot. d This can lead to packet loss. Therefore, the queue offset value must satisfy the following formula:
[0100]
[0101] To solve the service flow scheduling problem at the MAC layer, it is necessary to synchronously plan the queue resources and transmission time slot resources for each service flow, thereby maximizing the number of successfully scheduled service flows and the overall resource utilization of transmission time slots in the network. Furthermore, to avoid packet loss due to dead time during service flow scheduling, the offset constraint conditions specified in formula (19) must be strictly met. Therefore, the mathematical model of the MAC layer queue scheduling problem in this invention is as follows:
[0102]
[0103] Therefore, this invention solves the problem of service flow loss during dead time by introducing a multi-queue CQF and adopting a corresponding service flow scheduling strategy to satisfy the offset constraint condition of formula (21). Furthermore, the optimization objective of formula (20) is essentially a weighted summation of the scheduling effects of all service flows on different time slots and queue resources, thereby maximizing the overall scheduling benefit, and determined by the indicator function ω. ξ (f i As can be seen from the definition and formula (16), this optimization problem is equivalent to the 0-1 knapsack problem, which is an NP-hard problem. Therefore, in view of its solution complexity, this invention designs a time resource scheduling algorithm based on deep reinforcement learning, which guides the model to learn and approximate the optimal business flow scheduling decision through various design guidance methods.
[0104] To address the issue of burst flow scheduling and achieve coordinated transmission of mixed service flows, this invention establishes a multi-queue CQF. This forwarding queue supports deterministic shared network transmission of both burst and periodic flows. By rationally allocating queue offset values to service flows and matching transmission time slots, the deterministic transmission requirements of mixed service flows are ensured.
[0105] To enable the forwarding queue to support mixed transmission of periodic and bursty streams, the multi-queue CQF configures M queues (M ≥ 3) on the queue scheduling model of each UAV node. Each queue is equipped with a gating mechanism that periodically opens or closes according to a preset rhythm. Within each time slot, only one queue is allowed to be in the active state for data transmission and reception, while the gating of other queues remains closed to prevent conflicts between service flows, thereby improving bandwidth utilization. This mechanism essentially achieves time slot-based multiplexing, ensuring the orderly transmission of data streams. If a queue is used for data transmission, its state is denoted as Q. i =0, 0 < i < M-1, if used for receiving, then its queue state is set to Q. j =1, i≠j, specifically, in each time slot T σ Within the queue, the gating state switches according to the set scheduling strategy, which is defined as follows:
[0106]
[0107] Based on the aforementioned queue control strategy, further logic for receiving service flows between adjacent hop nodes was developed. This logic relies on the queue offset setting, where queue Q... j offset value φ j This indicates that in the current time slot, the data is transmitted from the corresponding transmission queue Q. σ%M to queue Q j The distance. Sending queue Q σ%M The offset value is defined as φ base For the queue receiving the service stream, its offset is based on the φ of the sending queue.base Using this as a reference point, the value is incremented by 1 sequentially to form a cyclic offset pattern.
[0108] Based on the queue offset settings described above, and according to the transmit / receive status in each time slot, the offset is characterized by a mathematical expression, namely:
[0109]
[0110] like Figure 6 As shown, to avoid sudden business flows within the dead time T d When packet loss occurs during internal transmission, burst flows, when selecting the receiving queue of the next-hop drone node, must prioritize queues with an offset value of not less than 2, i.e., satisfying φ. r ≥2, thus ensuring that the business flow can be successfully received, i.e., the following formula holds true:
[0111]
[0112] This satisfies the optimization conditions established by formula (21). Multi-queue CQF extends the time slot of the received service stream, so that the burst service stream sent during the dead time can still be received by the queue with a larger offset value in the downstream node, effectively improving the reliability of its transmission process.
[0113] To further optimize the forwarding efficiency of mixed service flows and improve the utilization rate of time slot resources. t The following service flow scheduling strategy has been formulated:
[0114] (1) For bursty flows from upstream nodes, in order to avoid packet loss, a receiving queue with a specific offset value must be assigned to it, that is, an offset value must be selected. The queue is used as the next-hop receiving queue. The specific offset value needs to be configured reasonably based on the traffic characteristics, latency requirements, and jitter resistance of the burst flow.
[0115] (2) After the burst stream is successfully received at the first hop, its subsequent forwarding can be completed using a circular queue mechanism. To shorten the total delay in multi-hop transmission, a fast forwarding mechanism can be introduced, that is, selecting an offset value for each subsequent hop. The system receives data from a queue, ensuring that the transmission delay between any two nodes does not exceed one time slot length, thus improving traffic forwarding efficiency.
[0116] (3) To rationally allocate queue resources and avoid periodic flows consuming critical resources on which burst flows depend, appropriate offset values should be configured for periodic flows based on the latency and jitter requirements of each service flow. Within each hop of the UAV node, the queue offset range that a periodic flow can select is...
[0117] (4) For traffic flows with low latency requirements in the network (i.e., best-effort flows), or when the current CQF receive queue is full, they can be stored in the Best Efforts (BE) queue. The BE queue does not have a set queue offset value; instead, it adopts an opportunistic transmission strategy. That is, during subsequent transmissions, if there is still bandwidth remaining in the current time slot after the CQF send queue has completed its scheduled data transmission, the BE queue can be scheduled for transmission. This fully utilizes idle resources and improves latency.
[0118] Figure 7 An example illustrates the working mechanism of multi-queue CQF. Within time slot T0, terminal N1 simultaneously transmits periodic and burst streams to N2 via four UAVs in the UANET. The system is configured with queue number M = 3. To ensure reliable transmission of the burst streams, an offset value is selected for the burst streams within time slot T0. The queue, i.e., according to formula (23), is selected for receiving the service flow, thus satisfying formula (24). After two time slots, the burst flow in time slot T2 needs to be selected. The queue is used to reduce the cumulative latency caused by subsequent hops. Periodic streams can choose an offset value based on their latency tolerance. Transmitting data through queues satisfies service quality requirements while avoiding interference with high-priority traffic. If periodic streams do not have high latency requirements, more resource queues needed by high-priority services can be freed up by allocating larger offset values.
[0119] This invention introduces a multi-queue CQF and its scheduling strategy suitable for mixed service scenarios, effectively improving the network's comprehensive scheduling capability for periodic and bursty flows. However, model design alone is still insufficient to fully guarantee the network's resource utilization efficiency. To further improve the overall network performance and achieve the optimization objective of formula (20), this invention proposes a time resource scheduling algorithm based on deep reinforcement learning, namely the GSR algorithm. Figure 8 As shown, the GSR algorithm first constructs a two-dimensional time-resource matrix, which is then fused with the action matrix to form a three-channel input. Feature extraction is performed using a fused convolutional network combining Ghost convolution and SRResNet, as designed in this invention, and the corresponding Q-values are output. Subsequently, an expert guidance module evaluates and adjusts the selected actions. After the actions are executed, the system obtains the state and reward information from the environment, and continuously optimizes the network parameters through an experience replay mechanism, thereby improving the efficiency of resource scheduling decisions. Through the above methods, the GSR algorithm effectively enhances the network's support for mixed service flows and increases the number of successfully scheduled service flows.
[0120] In multi-queue CQF, time slots and queue resources are not entirely independent. However, traditional one-dimensional neural networks often input the resource usage of each time slot and queue separately, resulting in fragmented data that is difficult to effectively capture the correlation between them, thus affecting the agent's ability to learn and optimize the network state. To further enhance the agent's perception and understanding of network state characteristics, this invention employs a two-dimensional time-resource utilization matrix to jointly model time slots and queue resources, optimizing the model's feature representation capabilities from a global perspective. This approach can more fully explore the intrinsic relationship between time slots and queue resources, providing the agent with richer information to improve its decision-making efficiency and scheduling optimization capabilities.
[0121] like Figure 9 As shown, the two-dimensional time-resource matrix uses rows to represent time slots and columns to represent the circular queues of each node, dynamically recording the resource occupancy of each queue under different time slots. This matrix serves as the state input channel for the agent, updating resource usage status in real time as business flows arrive, achieving adaptive state space migration. Furthermore, each column can be expanded to include additional information such as frequency, computing power, and link occupancy rate, thereby extending to a wider range of scheduling scenarios. To further enhance the agent's ability to extract key features, each matrix element is represented as a u×m vector to map the resource utilization relationship between different time slots and queues. For queues without direct association, a zero-filling strategy is used to construct a complete p×p state matrix.
[0122] In deep learning, efficient feature extraction, low computational cost, and high convergence speed are key to improving model performance. Addressing the drawbacks of traditional convolutional neural networks (CNNs), such as high computational cost, redundant parameters, and slow convergence speed, this invention designs a CNN that overcomes these problems. By combining the efficiency of Ghost convolution with the powerful feature extraction capabilities of SRResNet, it reduces computational cost while improving the model's feature extraction ability and accelerating convergence.
[0123] (1) GhostNet
[0124] Deep convolutional neural networks typically employ multi-layer convolutional structures. While this design enhances feature extraction capabilities, it also introduces a significant computational burden. Furthermore, in traditional convolutional operations, the dual expansion of the number of filters and channel dimensions leads to an exponential increase in computational complexity, which not only increases hardware resource consumption but also reduces model efficiency. Compared to conventional convolution, Ghost convolutional modules effectively reduce computational redundancy, resulting in fewer parameters and lower complexity.
[0125] like Figure 10As shown, the Ghost convolution employs a staged convolution strategy. This module reduces the computational load while maintaining feature representation capabilities by simplifying the number of feature channels and combining it with depthwise separable convolution techniques. This design maintains model performance while improving computational efficiency.
[0126] Suppose the given input data is X∈R c×h×w Where c is the number of input channels, and h and w are the height and width of the input data, respectively, meaning the dimension of the input feature is c×h×w. Therefore, after processing with n sets of m×m convolutional kernels using a regular convolution, the output feature dimension is n×h′×w′. Its parameter count and computational complexity can be expressed as:
[0127] P = ncmm (28)
[0128] C = nh′w′cmm (29) where the number of parameters P and the computational cost C of a regular convolution are key indicators for measuring model complexity. The number of parameters mainly depends on the number of convolutional kernels n, the number of input channels c, and the kernel size m×m, and is independent of the spatial dimension h′×w′ of the output features, so it can be expressed as n×c×m×m. The total number of elements in the output features is n×h′×w′, and the computation at each output position involves c×m×m multiplication and addition operations. Therefore, the computational cost of a regular convolution can be obtained by multiplying these two values.
[0129] In the Ghost convolutional architecture, given the same input and output dimension constraints, the first step is to obtain a convolutional architecture with dimensions of... The intrinsic features are output. Then, an s-1 linear transformation is performed to generate a dimension of... The Ghost feature output is then processed. Finally, the intrinsic features are concatenated with the Ghost features to obtain the same output as a regular convolution. The parameter count and computational complexity are as follows:
[0130]
[0131] Among them, P G C represents the number of parameters in a Ghost convolution. G Let represent the computational cost of Ghost convolution, s be the number of linear transformations much smaller than the number of channels c, and t×t be the average kernel size of the linear transformations. The convolution kernel m×m mentioned above is similar in size to t×t. Based on the above formulas (25)-(28), the following formulas can be derived:
[0132]
[0133] Among them, R p R represents the ratio of the two parameter values. CThis represents the ratio of computational cost. Compared to regular convolution, the computational cost of linear transformation is negligible. Therefore, when generating feature outputs of the same size, Ghost convolution can significantly reduce the number of parameters and computational complexity.
[0134] (2)SRResnet
[0135] In improving the accuracy of neural network models, increasing the number of network layers is often used to enhance the model's feature learning ability. However, with the increase in the number of layers, the model also faces a series of challenges, such as vanishing or exploding gradients, overfitting, and excessive computational resource consumption. Although these problems can be mitigated through certain methods, increasing the number of network layers does not always lead to improved learning efficiency and accuracy. In fact, once the model's learning performance reaches a certain level, further increasing the number of layers may be counterproductive, leading to decreased learning efficiency and even increased training and testing losses. It is worth noting that the introduction of residual learning enables deep networks to learn features more effectively, thereby improving the model's learning ability and accuracy while increasing the number of layers. To improve the model's learning ability and accuracy, this invention introduces the idea of SRResNet.
[0136] like Figure 11 As shown, the core idea of residual learning is to introduce shortcut connections into the network. Unlike traditional neural networks where each layer directly learns the mapping y = f(x), residual learning uses the mapping y = f(x) + x, directly adding the input to the output. This shortcut connection plays a crucial role in the training of deep networks: when a layer experiences the vanishing gradient problem, the gradient can propagate along the shortcut, skipping the affected layer, thus effectively mitigating the vanishing gradient phenomenon, allowing deeper network structures to be trained and improving overall performance. Furthermore, if for a given input x, the mapping function f(x) is approximately zero, then the output y is also close to x. In this case, the residual block can perform an approximate identity mapping, allowing the model to flexibly learn nonlinear features without forcibly fitting overly complex mappings, thereby avoiding unnecessary computational overhead. It is worth noting that identity shortcut connections do not add extra parameters or increase computational complexity, making them easy to implement in practical applications and not imposing additional burdens on computational resources.
[0137] In SRResNet, such as Figure 12As shown, the deep residual module consists of multiple stacked residual blocks, each of which employs skip connections, i.e., the aforementioned shortcut structure. In its implementation, each residual block first performs a convolution (Conv) and batch normalization (BN) on the input data, followed by ReLU activation. After another convolution and batch normalization, it is summed element-wise with the input, thus achieving the mapping y = f(x) + x. This design not only improves the network's training stability but also further enhances the model's feature learning capability.
[0138] After the residual convolution processing described above, this invention adds an additional layer after the last convolutional layer of the convolutional neural network to apply an additional activation function and batch normalization. The hyperbolic tangent (tanh) function is used in this network, and its mathematical expression is as follows:
[0139]
[0140] Next, output Convert to This ensures that the final output value falls within the range [0,1]. Through the above processing, the network normalizes each training mini-batch, which allows for a higher learning rate and reduces the sensitivity of the results to weight initialization.
[0141] A complete fusion convolutional neural network, such as Figure 13 As shown, in this neural network, the present invention utilizes the lightweight characteristics of Ghost convolution to reduce parameter redundancy, while combining it with the deep residual structure of SRResNet to enhance the network's expressive power. This not only reduces computational overhead but also achieves better model performance in less training time, improving training stability and convergence speed.
[0142] Because traditional loss functions suffer from problems such as sensitivity to outliers, insufficient sparsity, and difficulty in optimization in practical applications, this invention introduces the RoBoss (Robust, Bounded, Sparse, and Smooth Loss Function), whose mathematical expression is shown below:
[0143]
[0144] Among them, u=1-y k (ω T x k +b) is the classification margin error, and y k It is the label of the sample, x kω is the sample vector, b is the weight vector, a is the bias term, λ is the shape parameter that controls the shape of the loss function and the intensity of the penalty, and λ is the boundary parameter that controls the upper bound of the loss function to ensure that the loss value does not increase indefinitely.
[0145] Through the design of formula (32), RoBoSS has the following characteristics:
[0146] (1) By setting an upper bound λ, RoBoSS ensures that the value of the loss function will not increase indefinitely and is robust to outliers, i.e., when u > 0, there is
[0147] (2) RoBoSS is bounded and nonconvex. Although nonconvexity increases the complexity of optimization, its smoothness allows for the use of efficient gradient optimization algorithms.
[0148] (3) RoBoSS takes the value of 0 when u≤0, which means that the loss function does not impose additional penalties on correctly classified samples. This design makes the model more sparse, and only those samples close to the decision boundary or misclassified samples will affect the loss function.
[0149] In summary, the RoBoSS loss function improves performance when dealing with outlier and high-dimensional data through its robustness, boundedness, sparsity, and smoothness.
[0150] To enhance the decision-making efficiency of the intelligent agent, a multimodal data fusion strategy is adopted: a two-dimensional matrix representing time-resource utilization is concatenated with a decision behavior tensor along the channel axis, and parameter collaborative iterative optimization is achieved through feature extraction. This action matrix consists of two sub-channels: the first sub-channel indicates the operational behavior of queues in each time slot (e.g., sending or receiving); the second sub-channel indicates the behavior of each service flow in the valid transmission time slot. Legitimate actions are marked as 1, and illegal actions are marked as 0. Through this fusion method, the constructed state matrix ultimately contains three independent information channels.
[0151] Before an agent makes a decision, the expert system evaluates the potential performance impact of its chosen action. If an action under the current policy is judged to have a potential negative impact on system performance, the expert system will intercept the execution of that action and provide a better alternative. This mechanism not only effectively avoids system performance degradation but also accelerates the agent's understanding of the environment, helping it to build efficient state-action mapping strategies more quickly.
[0152] The MQSM designed in this invention is composed of a multi-queue CQF and a GSR algorithm, aiming to achieve efficient network resource scheduling and meet the optimization objectives of this invention. The multi-queue CQF, as the underlying forwarding framework of MQSM, ensures the orderly forwarding and latency guarantee of various types of flows in the network through multiple independent circular queues and corresponding scheduling strategies, thus providing a stable and controllable execution environment for the scheduling strategy. The GSR algorithm, as the core of scheduling decisions in MQSM, runs on top of the multi-queue CQF model. It constructs a two-dimensional time-resource utilization matrix to perceive the current resource occupancy of each time slot and queue, and extracts network state features using a deep neural network fused with GhostNet and ResNet. Then, it dynamically generates time slot-queue mapping scheduling actions based on a reinforcement learning strategy. Specifically, the multi-queue CQF provides the schedulable basic unit, while the GSR algorithm is responsible for optimal time slot scheduling on this structure. Together, they constitute the complete operating system of the MQSM mechanism.
[0153] In subsequent simulation verification, considering that multi-queue CQF has been used as the underlying implementation of MQSM service forwarding queues, the experiment will focus on the optimization performance of the GSR algorithm in different environments, including stability, deployment speed, service flow scheduling capability and time slot resource utilization, so as to verify the overall performance of MQSM.
[0154] All frameworks in this experiment were tested using the RoBoSS loss function and the Adam optimizer. The input format of the network model was [3, H, W], where H and W were dynamically determined based on the environment. The initial learning rate was set to 1e-5. If there was no performance improvement within 10 episodes, the learning rate was kept constant. The gradient update batch size during training was 32. Network parameters were updated after the minimum capacity of the experience replay buffer reached 20,000. The experimental environment was based on Windows 10 operating system, equipped with NVIDIA RTX 4090 GPU, and the simulation experiment was completed using the PyTorch 1.13 deep learning framework with Python 3 interface to analyze the scheduling performance of the GSR algorithm for mixed business flows.
[0155] Table 2 details the network parameter settings, including time slot length, service flow scheduling cycle, link bandwidth, and queue management mechanism, used to simulate scheduling strategies and resource allocation under different traffic scenarios. By configuring multi-level circular queues, best-effort transmission queues, and packet sizes, the CQF mechanism is ensured to effectively manage traffic and optimize transmission performance. The number of mixed service flows scheduled is 1000, with burst flows randomly generated within each transmission time slot. Furthermore, the simulation network used in this invention consists of several UAV subnets, with UAVs at both ends completing data transmission through networking. Considering the focus on QoS assurance of mixed service flows in complex environments, to ensure the comparability of simulation results and the effectiveness of the scheduling strategy, a relatively stable network topology combining common topologies such as star, tree, and ring topologies was used in the simulation to verify scheduling performance and test the scheduling capabilities of the proposed MQSM.
[0156] Table 2 Network Parameter Settings
[0157]
[0158] Table 3 shows the specific parameter settings for the GSR algorithm. The experience pool size, exploration rate, discount factor, and learning rate control the experience replay and policy update in reinforcement learning, while the update frequency and target network update affect the model's convergence speed and resource scheduling strategy. Overall, these parameters are used to balance exploration and exploitation, training stability, and resource optimization to ensure efficient algorithm operation.
[0159] Table 3 GSR Algorithm Parameter Settings
[0160]
[0161]
[0162] Figure 14 The training process of the scheduling algorithm is illustrated, where the transmission time slot length is 400µs and the service flow period is {400µs, 1600µs, 3200µs}. As shown in the figure, the time-based scheduling algorithm, compared to TimeDRS...
[15] The algorithm achieves a higher reward value because it employs an improved SRResnet, which boasts more stable gradient propagation and stronger nonlinear enhancement. Compared to the Tabu algorithm, the GRS algorithm converges relatively faster because it can extract key features from historical experience more quickly. DeepCQF, however, uses only a single fully connected layer to input one-dimensional queue resource information, which hinders the exploration of correlations between time slots and queues, leading to a subsequent decrease in reward value. In summary, the GRS algorithm exhibits faster convergence and a higher reward value, demonstrating its ability to train and deploy rapidly, which is of great significance for scenarios requiring rapid network deployment, such as disaster relief and search and rescue.
[0163] Figure 15 The graph shows the total execution time of the training model during the test, with each training model saved every 50 iterations. As can be seen from the graph, compared to TimeDRS, the GSR algorithm maintains a lower execution time without significant fluctuations, demonstrating strong computational stability and avoiding excessive computational overhead. DeepCQF, on the other hand, uses only a single fully connected layer to input one-dimensional queue resource information; although its execution time is even lower, its computational power is weaker, making it unable to handle complex tasks. Overall, the GSR algorithm has excellent computational efficiency, avoids the high overhead of TimeDRS, and is more computationally powerful than DeepCQF.
[0164] Figure 16 The graph shows the total number of execution steps for the training model during the test. As can be seen from the graph, compared to TimeDRS, it can complete the task in fewer steps, avoiding the redundant computation problem caused by the accumulation of multiple residual blocks in TimeDRS, thus ensuring computational stability. At the same time, GSR's execution step count is close to DeepCQF, indicating that the GSR algorithm has reasonable learning efficiency and is suitable for complex environments. Overall, the GSR algorithm balances computational stability and adaptability to complex environments, demonstrating excellent performance in environments with rapid network deployment requirements.
[0165] Figure 17 The figure illustrates the number of service flows scheduled under different time slot lengths. As the transmission time slot increases, the scheduling algorithm gradually increases the number of mixed service flows it can schedule. This is not only because a larger transmission time slot makes it easier to find the optimal solution during training, but also because a larger time slot length leads to a larger capacity for the corresponding service flows, allowing it to accommodate and schedule more data flows. Compared to the other three scheduling algorithms, the GSR algorithm can schedule more mixed service flows and has superior scheduling capabilities.
[0166] Figure 18 The figure illustrates the time slot resource utilization of various algorithms under different time slot lengths. As can be observed from the figure, due to the relatively small weighting factor of 0.1 set for time slot resource utilization in this invention, the gap between the GSR algorithm and other algorithms in terms of resource utilization gradually decreases with the increase of transmission time slot length, but it still maintains a high level overall, effectively enhancing system performance. Meanwhile, Figure 18 This also indicates a strong positive correlation between the number of scheduled service flows and the utilization rate of time slot resources. This is because when time slot resources are fully utilized, the number of service flows that can be transmitted per unit time increases, thereby improving the overall service carrying capacity of the system.
[0167] Figure 19The experiment demonstrates the number of service flows scheduled under different service flow cycles. Service flows were divided into three categories, each corresponding to a different cycle combination: {400us, 1600us}, {400us, 1600us, 3200us}, and {1600us, 3200us}. Comparison revealed that in the first two groups (i.e., the cycle set includes 400us), the scheduling performance of various algorithms was not significantly different with the number of service flows. However, in the service flow group containing only longer cycles (i.e., {1600us, 3200us}), the scheduling capabilities of each algorithm were superior. This is mainly because shorter time slots occur frequently throughout the scheduling process, consuming significant resources and limiting the overall scheduling performance of the system. Furthermore, compared to other algorithms, the GSR algorithm exhibited superior performance under all service flow cycle conditions, supporting the scheduling of more mixed flows.
[0168] The simulation results show that the GSR algorithm exhibits good adaptability, stability, and mixed service flow scheduling capabilities in different scenarios. Furthermore, it can generate high-performance models with fewer training rounds, enabling rapid network deployment and efficient real-time scheduling. In conclusion, the GSR algorithm has significant practical value in real-world applications, such as rapid deployment during disaster relief.
Claims
1. A method for implementing a self-organizing network service scheduling mechanism based on multi-queue CQF, characterized in that, The method establishes a multi-queue CQF, where the forwarding queues support deterministic shared network transmission of both bursty and periodic flows. By rationally allocating queue offset values to service flows and matching transmission time slots, the multi-queue CQF configures M queues (M≥3) on the queue scheduling model of each UAV node. Each queue is equipped with a gating mechanism, which periodically opens or closes according to a preset rhythm. Within each time slot, only one queue is allowed to be in the active state for data transmission and reception, while the gating of other queues remains closed to prevent conflicts between service flows. If a queue is used for data transmission, its state is denoted as Q. i =0, 0 < i < M-1, if used for receiving, then its queue state is set to Q. j =1, i≠j, in each time slot T σ Within the queue, the gating state switches according to the set scheduling strategy, which is defined as follows: Based on the above queue control strategy, the receiving logic for the service flow between two adjacent hop nodes was formulated. This logic depends on the queue offset setting, queue Q. j offset value φ j This indicates that in the current time slot, the data is transmitted from the corresponding transmission queue Q. σ%M to queue Q j Distance, sending queue Q σ%M The offset value is defined as φ base For the queue receiving the service stream, its offset is based on the φ of the sending queue. base Using this as a reference point, and incrementing by 1 sequentially to form a cyclic offset pattern; Based on the queue offset settings described above, and according to the transmit / receive status in each time slot, the offset is characterized by a mathematical expression, namely: To avoid sudden business flows within the dead time T d When packet loss occurs during internal transmission, burst flows, when selecting the receiving queue of the next-hop drone node, must prioritize queues with an offset value of not less than 2, i.e., satisfying φ. r ≥2, thus ensuring that the business flow can be successfully received, i.e., the following formula holds true: This satisfies formula (21). The established optimization conditions, multi-queue CQF, by extending the time slot of the received service stream, enable burst service streams sent within the dead time to still be received by the queues with larger offset values in the downstream nodes, effectively improving the reliability of its transmission process.
2. The implementation method of a self-organizing network service scheduling mechanism based on multi-queue CQF according to claim 1, characterized in that, The method optimizes the forwarding efficiency of mixed service flows and improves the utilization rate of time slot resources. t ,include: (1) For bursty flows from upstream nodes, in order to avoid packet loss, a receiving queue with a specific offset value must be assigned to it, that is, an offset value must be selected. The queue is used as the next-hop receiving queue, and the offset value needs to be reasonably configured according to the traffic characteristics, latency requirements and anti-jitter capabilities of the burst flow. (2) After the burst stream is successfully received at the first hop, its subsequent forwarding can be completed using a circular queue mechanism. To shorten the total delay in multi-hop transmission, a fast forwarding mechanism can be introduced, that is, selecting an offset value for each subsequent hop. The queue receives data, thereby ensuring that the transmission delay between every two hop nodes does not exceed one time slot length, thus improving traffic forwarding efficiency. (3) To rationally allocate queue resources and avoid periodic flows occupying critical resources on which burst flows depend, appropriate offset values should be configured for periodic flows based on the latency and jitter requirements of each service flow. Within each hop of the UAV node, the queue offset range that a periodic flow can select is: (4) For service flows with low latency requirements in the network (i.e. best-effort flows), or when the current CQF receive queue is full, they can be stored in the best-efforts (BE) queue. The BE queue does not set a queue offset value, but adopts an opportunistic transmission strategy. That is, in subsequent transmissions, if the current time slot still has remaining bandwidth after the CQF send queue has completed the predetermined data transmission, the BE queue can be scheduled to send.
3. The implementation method of the self-organizing network service scheduling mechanism based on multi-queue CQF according to claim 1, characterized in that, Within time slot T0, terminal N1 simultaneously transmits periodic and burst streams to N2 via four UAVs in the UANET. The system sets the queue size to M=3. To ensure reliable transmission of the burst streams, an offset value is selected for the burst streams within time slot T0. The queue, i.e., according to formula (23). Queue Q2 is selected to receive the service flow, thus satisfying formula (24). After two time slots, the burst flow in time slot T2 needs to be selected. The queue is used to reduce the cumulative latency caused by subsequent hops, while the periodic stream can choose an offset value based on its latency tolerance. Transmitting data through queues satisfies service quality requirements and avoids interference with high-priority traffic. If periodic streams do not have high latency requirements, more resource queues required by high-priority services can be released by allocating larger offset values.
4. The implementation method of a self-organizing network service scheduling mechanism based on multi-queue CQF according to claim 1, characterized in that, A multi-queue CQF and its scheduling strategy suitable for mixed service scenarios were introduced, which effectively improved the network's comprehensive scheduling capability for periodic and bursty flows. However, model design alone is still insufficient to fully guarantee the network's resource utilization efficiency. In order to further improve the overall network performance, formula (20) was implemented. To optimize the GSR algorithm, it first constructs a two-dimensional time-resource matrix and merges it with the action matrix to form a three-channel input. It then extracts features and outputs the corresponding Q-values through a fusion convolutional network combining Ghost convolution and SRResNet. Subsequently, the expert guidance module evaluates and adjusts the selected actions. After the actions are executed, the system obtains the status and reward information from the environment and continuously optimizes the network parameters through an experience replay mechanism, thereby improving the efficiency of resource scheduling decisions.
5. The implementation method of a self-organizing network service scheduling mechanism based on multi-queue CQF as described in claim 4, characterized in that, In multi-queue CQF, time slots and queue resources are not completely independent. A two-dimensional time-resource utilization matrix is used to jointly model time slots and queue resources, thereby optimizing the feature representation capability of the model from a global perspective.
6. The implementation method of a self-organizing network service scheduling mechanism based on multi-queue CQF as described in claim 5, characterized in that, The two-dimensional time-resource matrix uses rows to represent time slots and columns to represent the circular queues of each node. It dynamically records the resource occupancy of each queue under different time slots. This matrix serves as the state input channel for the agent, updating the resource usage status in real time as the business flow arrives, thus achieving adaptive migration of the state space. Each column can also be expanded to include additional information such as frequency, computing power, and link occupancy rate, thereby extending to a wider range of scheduling scenarios. To further improve the agent's ability to extract key features, each matrix element is a u×m vector to map the resource utilization relationship between different time slots and queues. For queues without direct association, the matrix adopts a zero-filling strategy to construct a complete p×p state matrix.