A real-time resource allocation method and system for an aircraft

By building a task priority queue and resource state matrix in the real-time aircraft control software, and using a dual-deep Q network combined with dynamic planning for resource allocation, the task delay and system crash problems caused by resource limitations in the existing technology are solved, and efficient real-time resource allocation and improved security performance are achieved.

CN119806846BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293996.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Existing real-time control software for aircraft is difficult to effectively allocate limited computing resources, storage space and network bandwidth under resource limitations, resulting in mission-critical delays, system crashes and even security accidents.

Method used

By building a task priority queue and resource state matrix, the state space and action space are defined, and the dual-deep Q network combined with dynamic planning can realize real-time resource allocation of aircraft and dynamically adjust algorithm parameters to adapt to environmental changes.

Benefits of technology

It significantly improves the task completion rate, resource utilization rate and average response time, enhances the aircraft's processing capabilities in resource limitation and multi-task concurrent environments, and improves safety performance and response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806846B_ABST
    Figure CN119806846B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time resource allocation method and system for an aircraft, relating to the technical field of aircraft control, and used to solve the technical problem of resource limitation and demand conflict caused by resource allocation of the aircraft. The real-time resource allocation method for the aircraft includes: constructing a task priority queue and a resource status matrix of the aircraft, and defining a state space; defining an action space, and defining a reward function according to the dynamic resource allocation target; building a real-time control software environment of the aircraft; preparing training data for a double deep Q network, and constructing a double deep Q network structure; constructing an experience replay buffer, performing double deep Q network training, and updating the parameters of the main network; obtaining the value function between the state and the action of the aircraft, and enabling the aircraft to make an optimal decision on resource allocation through the value function; using dynamic programming to refine the allocated resources in real time, realizing the maximization of task benefits under the condition of meeting resource constraints, and completing the real-time resource allocation of the aircraft.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aircraft control, and more specifically, to a method and system for real-time resource allocation of an aircraft. Background Art

[0002] The real-time control system of an aircraft is an important part of the modern aerospace field, and its functions cover key links such as flight attitude adjustment, path planning, and navigation control. With the increasing complexity of aircraft functions, the requirements for real-time performance and task diversity faced by the control system are gradually increasing, which poses higher requirements for the dynamic resource allocation and scheduling capabilities of the real-time control software of the aircraft. The rationality of resource allocation directly affects the overall performance of the control system. Especially in a high-dynamic and complex task environment, how to efficiently allocate limited computing resources to cope with real-time challenges has become the focus and difficulty of current research.

[0003] The core requirements of the real-time control software of an aircraft are concentrated in real-time performance and efficiency. The real-time performance requires the system to be able to quickly respond to changes in the external environment and adjustments of internal commands, which is crucial for ensuring flight safety. Efficiency is reflected in that the system needs to maximize the task execution speed and resource utilization rate under the conditions of limited processing power and storage resources to ensure that various control tasks can be completed at the best time point to avoid delays or resource waste. The dynamic resource allocation of the real-time control software of an aircraft is the key technology to achieve the real-time performance and efficiency of the real-time control software of the aircraft. The core of aircraft resource allocation lies in how to reasonably schedule limited computing resources, storage space, and network bandwidth to support the parallel processing of multiple tasks and the rapid transfer of data. However, the existing real-time resource allocation of aircraft real-time control software faces many challenges. The resources of an aircraft, such as processor speed, memory size, and its energy consumption, have fixed upper limits, while the control software needs to complete complex data processing and task scheduling under these limited resources, which leads to conflicts between resource limitations and requirements, and may cause critical task delays, system crashes, and even safety accidents. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for real-time resource allocation of an aircraft to solve the technical problem of resource limitation and demand conflict caused by existing aircraft resource allocation. In view of this, the present invention is realized through the following solutions.

[0005] In a first aspect, the present invention provides a method for real-time resource allocation of an aircraft, including:

[0006] Construct a task priority queue and a resource status matrix of the aircraft, and define a state space; allocate the task requirements of the aircraft to the system overall resource definition action space, and define a reward function according to the dynamic resource allocation target;

[0007] Build the real-time control software environment of the aircraft; prepare the training data for the dual-depth Q network, and construct the dual-depth Q network structure;

[0008] Construct an experience replay buffer, perform the training of the dual-depth Q network, and update the parameters of the main network;

[0009] Obtain the value function between the state and action of the aircraft, and make the optimal decision on resource allocation for the aircraft through the value function;

[0010] Use dynamic programming to refine the resource allocation in real time, achieve the maximization of task benefits under the condition of meeting resource constraints, and complete the real-time resource allocation of the aircraft.

[0011] Compared with the prior art, in the real-time resource allocation method of the aircraft of the present invention, by constructing a task priority queue and a resource status matrix of the aircraft, defining a state space, and defining an action space by allocating the task requirements of the aircraft to the overall system resources, a reward function is defined according to the dynamic resource allocation goal; further, after constructing the dual-depth Q network structure and the experience replay buffer, perform the training of the dual-depth QNetwork training is carried out to update the parameters of the main network; furthermore, by obtaining the value function between the states and actions of the aircraft, the aircraft makes an optimal decision on resource allocation; further, dynamic programming is also used to refine the allocated resources in real time to achieve the maximization of task benefits under resource constraint conditions, and the real-time resource allocation of the aircraft is completed. Through the above technical solutions, the present invention combines the adaptive learning ability of deep reinforcement learning for complex non-linear problems and the fast optimization advantage of dynamic programming in real-time decision-making, and realizes efficient real-time resource allocation through a two-way cooperation mechanism, which has a significant improvement in terms of task completion rate, resource utilization rate, and average response time compared with the prior art, and strengthens the aircraft's processing ability for the two problems of current resource limitations and requirements, and multi-task concurrency; the above technical solutions of the present invention have an adaptive strategy that can dynamically adjust algorithm parameters to adapt to environmental changes by monitoring the operating environment in real time, improving the aircraft's adaptability in complex environments. At the same time, the real-time resource allocation strategy based on reinforcement learning can be quickly optimized and updated through training iterations in a new environment when facing the continuously developing aircraft control algorithms, so as to support the efficient operation of new algorithms and improve the performance of the aircraft; further, for emergencies encountered during the actual operation of the aircraft, the present invention can activate an emergency handling plan to ensure that the aircraft system can still continue to operate or land safely when encountering sudden tasks or single-point failures, and can respond quickly when encountering sudden tasks, significantly improving the safety performance and response ability of the aircraft. Through the above technical solutions, the present invention solves the technical problem of resource limitation and demand conflict caused by the existing aircraft resource allocation, and effectively avoids problems such as key task delay, system crash, and even safety accidents of the aircraft.

[0012] Further, in the real-time resource allocation method of the aircraft of the present invention, the construction of the task priority queue and resource status matrix of the aircraft, and the definition of the state space include:

[0013] Define the priority queue, and the priority queue is represented by the following formula:

[0014] ;

[0015] Wherein, P i represents the priority of the task, D i represents the deadline of the task, ε represents the buffer factor of task delay, which is used to avoid resource conflicts caused by excessive priority, W i represents the task weight;

[0016] According to the current aircraft task queue Q and the resource status matrix R constitute the state spaceS ; Among them, the task queue , T n represents the task queue Q in the n th aircraft task;

[0017] Resource status matrix , R c represents the system computing resources, R m represents the system memory resources, R b represents the system bandwidth resources;

[0018] Assign the aircraft task T i to the resource R j in, defined as the action space A , the action space , among them, represents the th method of assigning the aircraft task T i to the resource R j , the resource , represents the allocated system computing resources, represents the allocated system memory resources, represents the allocated system bandwidth resources.

[0019] Further, in the aircraft real-time resource allocation method of the present invention, the preparation of the double-depth Q network training data includes:

[0020] Prepare task data: Generate a task queue containing 1000 tasks, and each task is randomly assigned requirements; the requirements include task computing requirements, task memory requirements, and task bandwidth requirements; among them, the task computing requirement , unit MHz , the task memory requirement , unit is MB , the task bandwidth requirement , unit is Mbps , the task deadline , unit is s ;

[0021] Prepare system resource data: Total computing resources ; Total memory resources ; Total bandwidth resources ;

[0022] and / or, the construction of the double deep Q network structure, including:

[0023] Define the input layer: the dimension of the task queue is 10, the dimension of the system resource status is 3, and the total input layer dimension is 13;

[0024] Define the hidden layer: there are two hidden layers, each with 128 neurons, and the activation function is ;

[0025] Define the output layer: the dimension of the action set is 10.

[0026] Furthermore, in the real-time resource allocation method of the aircraft of the present invention, the construction of the experience replay buffer for the double deep Q network training includes:

[0027] Input the state of each agent into the double deep Q network, and select an action according to the greedy algorithm in the current action set space, and apply the actions of all agents to the environment. The action selection strategy can be expressed as:

[0028] ; where represents the current state, represents the optional actions in the current state, represents the action that makes maximum in the current state, represents the exploration rate, , represents finding the one that makes the function maximize, represents all possible action sets in the current state;

[0029] Introduce a target network, which has the same structure as the main network and is independent of each other;

[0030] Initialize the parameters of the main network and the target network;

[0031] At each time step, select an action according to the current state using the main network ;

[0032] Execute the selected action , observe the next state and the immediate reward obtained ;

[0033] Determine that the training experience quadruple of the agent is , and add this quadruple Stored in the experience replay buffer as training data;

[0034] And / or, updating the parameters of the main network includes:

[0035] Randomly sampling a batch of experience quadruples from the experience replay buffer ; For any sampled experience quadruple, using the double deep Q Network to estimate the current state value function And the Target value function of the next state ; Calculate the network parameters when minimizing the mean squared error loss function and update the parameters of the main network.

[0036] Furthermore, in the real-time resource allocation method of the aircraft of the present invention, randomly sampling a batch of experience quadruples from the experience replay buffer ; For any sampled experience quadruple, using the double deep Q Network to estimate the current state value function And the Target value function of the next state ; Calculate the network parameters when minimizing the mean squared error loss function and update the parameters of the main network, including:

[0037] S100, using the target network to calculate the maximum value function of the next state Corresponding to the sampled experience quadruple Action ;

[0038] S200, using the main network to estimate the value function of the current state Of the sampled experience quadruple ;

[0039] S300, using the next state Execute the action Target value function of Update the value function of the current state Value, the formula of which is: ;

[0040] ;

[0041] Wherein, Represents the current state, Represents the optional actions in the current state, Represents the Q Value estimation of the current state, Represents the updated Q Value estimation, Is the weight factor, and , represents the reward value returned by the environment after performing an action in the current state , represents the target value function represents the next state of the current state represents a state such that Q the maximum action is the discount factor, and ;

[0042] S400, obtain the mean squared error loss function, and its formula is:

[0043] ;

[0044] where represents the loss function represents the expectation over all possible data distributions represents the current state represents the available actions in the current state represents in the state performing an action the reward value returned by the environment after is the discount factor, and , represents the next state of the current state represents a state such that Q the maximum action represents the parameters of the target network represents the Q value estimate of the main network for the current state represents the parameters of the main network represents a state the Q value estimate of the target network;

[0045] S500, update the parameters of the main network according to the stochastic gradient descent method, and its formula is:

[0046] ;

[0047] where represents the new parameter value of the main network after one gradient descent update represents the parameter value of the main network at the start of the current iteration step represents the learning rate represents the loss function; is the gradient operator, representing taking the gradient with respect to the parameter ;

[0048] S600, copy the parameters of the main network to the target network;

[0049] S700, accelerate the training of the dual-depth Q network through GPU parallel training;

[0050] And / or, making the optimal decision on resource allocation by the aircraft through the value function includes:

[0051] Utilize the trained dual-depth Q network model, input the state of the system at a certain moment ; ;

[0052] Dual-depth Q The network model outputs the optimal action according to the input state , and obtains the optimal real-time resource allocation decision of the aircraft.

[0053] Furthermore, in the real-time resource allocation method of the aircraft of the present invention, after making the optimal decision on resource allocation by the aircraft through the value function, it further includes:

[0054] S100 measures the contribution of the neuron to the result according to the gradient of the loss function by the neuron; the gradient of the loss function by the neuron is expressed by the following formula:

[0055] ; where, , represents the network parameters related to a certain neuron, represents the loss function, represents the gradient operator, representing the gradient of the parameter ;

[0056] S200, set a suitable pruning threshold according to the gradients of all neurons to the loss function, and delete the neurons and links with low contribution based on this;

[0057] S300 uses linear quantization technology to convert the weights represented by 32-bit floating-point numbers in the dual-depth Q network model into weights represented by 8-bit or 16-bit integers, and its formula is: or ; where, Q represents the quantized fixed-point data, (•) represents the rounding function, used to round the calculation result, represents the original floating-point data, represents the scaling factor, represents the offset;

[0058] Set the maximum and minimum values of the floating-point parameters and fixed-point parameters to be , , , , then there is:

[0059] , where represents the scaling factor, represents the maximum value of the floating-point parameter, represents the minimum value of the floating-point parameter, represents the maximum value of the fixed-point parameter, represents the minimum value of the fixed-point parameter;

[0060] , where represents the offset, represents the maximum value of the fixed-point parameter, represents the maximum value of the floating-point parameter, represents the scaling factor.

[0061] Furthermore, in the aircraft real-time resource allocation method of the present invention, the use of dynamic programming to refine the allocation of resources in real time to achieve the maximization of task benefits under resource constraint conditions and complete the aircraft real-time resource allocation includes:

[0062] For emergencies, estimate the amount of resources required for the emergency and determine whether the reserved resources are sufficient; form a task queue with the emergency and non-critical tasks that release resources, and divide it into multiple task sub-queues according to priorities;

[0063] Initialize and define the dynamic programming algorithm; including defining state variables in dynamic programming, defining decision variables in dynamic programming, defining the state transition matrix in dynamic programming, and defining task benefits in the dynamic programming algorithm;

[0064] Adopt distributed computing, simultaneously use the dynamic programming algorithm in each task sub-queue, recursively calculate from the highest to the lowest priority order, record the optimal decision at each stage, and then perform forward optimization to finally obtain the optimal decision sequence of the original problem, and achieve the optimal strategy for the remaining resources allocation of the aircraft system for random emergency tasks of the aircraft.

[0065] Furthermore, in the aircraft real-time resource allocation method of the present invention, the step of, for emergencies, estimating the amount of resources required for the emergency and determining whether the reserved resources are sufficient; forming a task queue with the emergency and non-critical tasks that release resources, and dividing it into multiple task sub-queues according to priorities, includes:

[0066] If it is sufficient, use the reserved resources to handle the emergency; if it is not enough, release the resources of the currently executing non-critical tasks and re-allocate the resources to critical tasks;

[0067] Calculate the priority of each task according to the following formula:

[0068] ; where, P i represents the priority of the task, D i represents the deadline of the task, ε represents the buffer factor for task delay, which is used to avoid resource conflicts caused by excessive priority, W i represents the task weight;

[0069] Divide the task queue into n levels of task sub-queues according to priority;

[0070] And / or, initialize and define the dynamic programming algorithm; including defining state variables in dynamic programming, defining decision variables in dynamic programming, defining the state transition matrix in dynamic programming, and defining the task revenue in the dynamic programming algorithm, including:

[0071] Obtain the total number of tasks in the task queue N ;

[0072] Define the state variable in dynamic programming as ; represents the amount of resources allocated for the i th task to the th task;

[0073] Define the decision variable in dynamic programming as ; represents the amount of resources allocated to the i th task, ; and define the state transition matrix of dynamic programming, the formula is: ;

[0074] Define the amount of resources allocated in the dynamic programming algorithm to task as the revenue function , the formula is:

[0075] ; where, represents the minimum amount of resources required to complete the task; represents the amount of resources allocated to ; represents task weight, represents the coefficient.

[0076] Furthermore, in the real-time resource allocation method for the aircraft of the present invention, distributed computing is adopted. In each task sub-queue, the dynamic programming algorithm is used simultaneously. Starting from the highest priority in descending order, the optimal decision at each stage is recorded, and then the optimal decision sequence of the original problem is obtained by forward optimization, realizing the optimal strategy for the remaining resources allocation of the aircraft system for the random and sudden tasks of the aircraft, including:

[0077] Adopt distributed computing, and independent computing nodes process the dynamic programming algorithm to simultaneously process n priority task sub-queues, and summarize the allocation results at the same time;

[0078] Use the dynamic programming recurrence relation to calculate section by section, and finally obtain the optimal solution of the problem; the formula is as follows:

[0079] ;

[0080] Among them, represents the amount of resources allocated for the i th task to the th task, represents the amount of resources allocated to the i th task, represents the total number of tasks in the task queue, represents the maximum benefit of the i th task in state , represents the benefit function of allocating the amount of resources to task , represents the maximum benefit of the i+1 th task, represents the maximum benefit of the th task, represents that in the dynamic programming algorithm, the benefit of task is , represents the amount of resources allocated to the th task, represents the amount of resources allocated for the th task;

[0081] Adopt a strategy that combines heuristic search and dynamic programming, and give priority to high-benefit paths; the dynamic programming path is defined by the following formula: , among which, represents the total actual benefit from the task with the highest priority to the task with the current priority of i ; represents the total estimated benefit of the remaining tasks, Denote the comprehensive path benefit from the task with the highest priority to the task with the current priority i to the task with the priority i , and then from the task with the priority j to the task with the lowest priority, where

[0082] represents the number of remaining tasks.

[0083] A priority queue determination module, configured to: construct a task priority queue and a resource status matrix of the aircraft, and define a state space; define an action space by allocating the task requirements of the aircraft to the overall system resources, and define a reward function according to the dynamic resource allocation objective;

[0084] A network structure construction module, configured to: build a real-time control software environment of the aircraft; prepare training data for the double deep Q network, and construct a double deep Q network structure;

[0085] A parameter update module of the main network, configured to: construct an experience replay buffer, perform training of the double deep Q network, and update the parameters of the main network;

[0086] An optimal decision-making module, configured to: obtain a value function between the state and the action of the aircraft, and enable the aircraft to make an optimal resource allocation decision through the value function;

[0087] A dynamic programming module, configured to: use dynamic programming to refine the allocated resources in real time, achieve the maximization of task benefits under the condition of meeting resource constraints, and complete the real-time resource allocation of the aircraft.

[0088] Compared with the prior art, the beneficial effects of the real-time resource allocation system of the aircraft in the present invention are the same as those of the real-time resource allocation method of the aircraft described in the above technical solution, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:

[0090] Figure 1 is a schematic diagram of the double deep Q network structure constructed by the present invention;

[0091] Figure 2 is a schematic diagram of the training process of the double deep Q network of the present invention;

[0092] Figure 3Schematic diagram of the overall process of the real-time resource allocation method for the aircraft of the present invention;

[0093] Figure 4 Bar chart showing the performance comparison between the present invention and the prior art. Detailed implementation manners

[0094] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0095] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0096] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined. "Several" means one or more unless otherwise specifically defined.

[0097] The core requirements of the aircraft real-time control software are mainly reflected in real-time performance and high efficiency. The real-time performance requires the system to be able to quickly respond to changes in the external environment and adjustments of internal commands, which is crucial for ensuring flight safety. The high efficiency is reflected in that the system needs to maximize the task execution speed and resource utilization rate under the conditions of limited processing power and storage resources, so as to ensure that various control tasks can be completed at the best time point to avoid delays or resource waste. The dynamic resource allocation of the aircraft real-time control software is the key technology to realize the real-time performance and high efficiency of the aircraft real-time control software. The core of aircraft resource allocation lies in how to reasonably schedule limited computing resources, storage space and network bandwidth to support the parallel processing of multiple tasks and the rapid transfer of data. However, the existing real-time resource allocation of aircraft real-time control software faces many challenges. The resources of the aircraft, such as the processor speed, memory size and its energy consumption, have fixed upper limits, while the control software needs to complete complex data processing and task scheduling under these limited resources, which leads to conflicts between resource limitations and requirements, and may cause delays in critical tasks, system crashes, and even safety accidents.

[0098] To solve the above technical problems, the present invention provides a method for real-time resource allocation of an aircraft, including:

[0099] Construct a task priority queue and a resource status matrix for the aircraft, and define the state space; define the action space by allocating the task requirements of the aircraft to the overall system resources, and define the reward function according to the dynamic resource allocation goal;

[0100] Build the real-time control software environment of the aircraft; prepare the training data for the double-depth Q network, and construct the double-depth Q network structure;

[0101] Construct an experience replay buffer, perform the training of the double-depth Q network, and update the parameters of the main network;

[0102] Obtain the value function between the state and the action of the aircraft, and make the aircraft make the optimal decision on resource allocation through the value function;

[0103] Use dynamic programming to refine the allocated resources in real time, achieve the maximization of task benefits under the condition of meeting resource constraints, and complete the real-time resource allocation of the aircraft.

[0104] In the case of adopting the above technical solution, in the real-time resource allocation method of the aircraft of the present invention, by constructing a task priority queue and a resource status matrix of the aircraft, defining the state space, and defining the action space by allocating the task requirements of the aircraft to the overall system resources, the reward function is defined according to the dynamic resource allocation goal; further, after constructing the double-depth Q network structure and the experience replay buffer, perform the double-depth QNetwork training is carried out, and the parameters of the main network are updated; furthermore, by obtaining the value function between the state and action of the aircraft, the aircraft makes an optimal decision on resource allocation; further, dynamic programming is also used to refine the allocated resources in real time to achieve the maximization of task benefits under resource constraint conditions, and the real-time resource allocation of the aircraft is completed. Through the above technical solutions, the present invention combines the adaptive learning ability of deep reinforcement learning for complex non-linear problems and the fast optimization advantage of dynamic programming in real-time decision-making, and realizes efficient real-time resource allocation through a two-way cooperation mechanism. Compared with the prior art, there are significant improvements in task completion rate, resource utilization rate, and average response time, and the handling ability of the aircraft for the two problems of current resource limitations and demands, and multi-task concurrency is strengthened; the above technical solutions of the present invention have an adaptive strategy that can dynamically adjust algorithm parameters to adapt to environmental changes by real-time monitoring the operating environment, improving the adaptability of the aircraft in complex environments. At the same time, the real-time resource allocation strategy based on reinforcement learning can be quickly optimized and updated through training iterations in a new environment when facing the continuously developing aircraft control algorithms, so as to support the efficient operation of new algorithms and improve the performance of the aircraft; further, for emergencies encountered during the actual operation of the aircraft, the present invention can start an emergency handling plan to ensure that the aircraft system can still continue to operate or land safely when encountering sudden tasks or single-point failures, and can quickly respond when encountering sudden tasks, significantly improving the safety performance and response ability of the aircraft. Through the above technical solutions, the present invention solves the technical problem of resource limitation and demand conflict caused by the existing aircraft resource allocation, and can effectively avoid problems such as key task delays, system crashes, and even safety accidents of the aircraft.

[0105] To better understand the present invention, the content of the present invention will be further clarified below in conjunction with specific embodiments, but the content of the present invention is not limited to the following embodiments.

[0106] Embodiment 1

[0107] This embodiment provides a method for real-time resource allocation of an aircraft, including:

[0108] Step 1, construct a task priority queue and a resource status matrix of the aircraft, and define the state space; define the action space by allocating the task requirements of the aircraft to the overall system resources, and define the reward function according to the dynamic resource allocation target;

[0109] Step 2, build a real-time control software environment for the aircraft; prepare the training data of the double deep Q network, and construct the double deep Q network structure;

[0110] Step 3, construct an experience replay buffer and perform the double deep QPerform network training and update the parameters of the main network;

[0111] Step 4: Obtain the value function between the state and action of the aircraft, and enable the aircraft to make an optimal decision on resource allocation through the value function;

[0112] Step 5: Use dynamic programming to refine the allocated resources in real time, achieve the maximization of task benefits under resource constraint conditions, and complete the real-time resource allocation of the aircraft.

[0113] Embodiment 2

[0114] This embodiment provides a method for real-time resource allocation of an aircraft, including:

[0115] S100: Construct a task priority queue and a resource status matrix for the aircraft, and define the state space; define the action space by allocating the task requirements of the aircraft to the overall system resources, and define the reward function according to the dynamic resource allocation goal;

[0116] Further, the construction of the task priority queue and the resource status matrix for the aircraft, and the definition of the state space include:

[0117] S101: Define the priority queue, which is represented by the following formula:

[0118] ;

[0119] Where P i represents the priority of the task, D i represents the deadline of the task, ε represents the buffer factor for task delay, which is used to avoid resource conflicts caused by excessive priority, W i represents the task weight; S102: According to the current aircraft task queue Q and the resource status matrix R constitute the state space S ; where the task queue , T n represents the Q th n aircraft task in the task queue

[0120] The resource status matrix , R c represents the system computing resources, R m represents the system memory resources, R b represents the system bandwidth resources;

[0121] S103. Assign the aircraft mission T i to resources R j and define it as the action space A The action space where represents the th method of assigning the aircraft mission T i to resources R j The resources , represents the allocated system computing resources, represents the allocated system memory resources, represents the allocated system bandwidth resources;

[0122] Furthermore, for the above reward function, in this embodiment, on the basis of the traditional maximization of resource utilization rate and task completion rate, a comprehensive consideration of task completion time and task priority is introduced to define the reward function r The reward function r is represented by the following formula: ; where r represents the reward function, represents the weight coefficient, represents the task weight, which can be adjusted according to actual needs, n represents the total number of tasks, represents the weight coefficient, represents the total system resource volume, represents the system resources required to complete the task, represents the number of completed tasks; is a Boolean variable representing the task status. When it is the case, it means the task is successful. When it is the case, it means the task is failed; The weight coefficients and can be adjusted according to actual needs;

[0123] S200. Build the real-time control software environment of the aircraft; Prepare the training data of the dual-depth Q network and construct the dual-depth Q network structure;

[0124] Furthermore, the preparation of the training data of the dual-depth Q network includes:

[0125] S201. Prepare task data: Generate a task queue containing 1000 tasks, and randomly assign requirements to each task; the requirements include task computing requirements, task memory requirements, and task bandwidth requirements; among them, the task computing requirements , unit MHz , the task memory requirement , the unit is MB , the task bandwidth requirement , the unit is Mbps , the task deadline , the unit is s ;

[0126] S202. Prepare system resource data: Total computing resources ; Total memory resources ; Total bandwidth resources ;

[0127] Furthermore, please refer to Figure 1 , the constructed double-depth Q network structure, including:

[0128] S211. Define the input layer: The dimension of the task queue is 10, the dimension of the system resource status is 3, and the total dimension of the input layer is 13;

[0129] S212. Define the hidden layer: There are two hidden layers, each with 128 neurons, and the activation function is , the activation function refers to the ramp function in mathematics, which can be expressed as , where represents taking the maximum value, is the threshold, represents the input value;

[0130] S213. Define the output layer: The dimension of the action set is 10;

[0131] S300. Construct an experience replay buffer, perform the double-depth Q network training, and update the parameters of the main network;

[0132] Furthermore, please refer to Figure 2 , the construction of the experience replay buffer and the performance of the double-depth Q network training include:

[0133] S311. Input the state of each agent into the double-depth Q network, select actions according to the greedy algorithm in the current action set space, and apply the actions of all agents to the environment. The action selection strategy can be expressed as:

[0134] ; Among them, represents the current state, represents the optional actions in the current state, represents that in the current state, the action that maximizes represents the exploration rate, , represents finding the when the function is maximized, represents the set of all possible actions in the current state;

[0135] S312. Introduce a target network, which has the same structure as the main network and is independent of each other;

[0136] S313. Initialize the parameters of the main network and the target network;

[0137] S314. At each time step, according to the current state use the main network to select an action ;

[0138] S315. Execute the selected action , observe the next state and the immediate reward obtained ;

[0139] S316. Determine that the training experience quadruple of the agent is , and use this quadruple as training data and store it in the experience replay buffer;

[0140] Furthermore, the updating of the parameters of the main network includes:

[0141] S320. Randomly sample a batch of experience quadruples from the experience replay buffer ; For any sampled experience quadruple, use the double deep Q network to estimate the current state value function and the target value function of the next state ; Calculate the network parameters when minimizing the mean square error loss function, and update the parameters of the main network; ;

[0142] Furthermore, step S320 specifically includes:

[0143] S321. Use the target network to calculate the maximum value function of the next state corresponding to the sampled experience quadruple ;

[0144] S322, use the main network to estimate the value function of the current state of the sampled experience quadruple ; ;

[0145] S323, use the target value function of the action executed in the next state to update the value function of the current state ; the formula is: ; wherein, represents the current state,

[0146] ;

[0147] wherein, represents the current state, represents the optional actions in the current state, represents the Q value estimate of the current state, represents the updated Q value estimate, is the weight factor, and ; represents the reward value returned by the environment after executing the action in the current state , represents the target value function, represents the next state of the current state, represents the state at which Q is maximized, is the discount factor, and ;

[0148] S324, obtain the mean square error loss function, the formula of which is:

[0149] ;

[0150] wherein, represents the loss function, represents the expectation over all possible data distributions, represents the current state, represents the optional actions in the current state, represents the reward value returned by the environment after executing the action in the state , is the discount factor, and ; represents the next state of the current state, represents the state at which Q is maximized, Represents the parameters of the target network, Represents the Q value estimation of the current state main network, Represents the main network parameters, Represents the state of the target network under Q value estimation;

[0151] S325. Update the parameters of the main network according to the stochastic gradient descent method, and its formula is:

[0152] ;

[0153] Wherein, Represents the new parameter value of the main network after one gradient descent update, Represents the parameter value of the main network at the beginning of the current iteration step, Represents the learning rate, Represents the loss function; Is the gradient operator, indicating the gradient of the parameter ;

[0154] S326. Copy the parameters of the main network to the target network;

[0155] S327. Accelerate the training of the double deep Q network through GPU parallel training; Repeat the above steps until the stop condition is reached;

[0156] Furthermore, in the above S327, it includes: S3271. Set as the total number of time steps for each training round, as the current time step. If it satisfies , then , jump to step S300. Otherwise, let , and jump to step S3272; For step S3272: If it satisfies , is the total number of training rounds, is the current round number, then , jump to step S200. Otherwise, jump to step S3273; For step S3273: Through iterative training, the double deep Q network finally obtains the value function between the state and the action, Q which can guide the aircraft to make the optimal decision of real-time resource allocation based on the double deep

[0157] Furthermore, in this embodiment, pruning technology is used to delete the parts with less impact on the prediction results by analyzing redundant neurons and invalid connections in the network. First, the contribution degree of each neuron is calculated through an importance evaluation method, and then the neurons and connections with low contribution are deleted based on a set pruning threshold. Specifically as follows:

[0158] S400, obtain the value function between the state and action of the aircraft; utilize the trained double deep Q network model, input the state of the system at a certain moment ; the double deep Q network model outputs the optimal action according to the input state , and obtain the optimal real-time resource allocation decision of the aircraft;

[0159] S500 measures the contribution degree of the neuron to the result according to the gradient of the neuron to the loss function; the gradient of the neuron to the loss function is represented by the following formula:

[0160] ; where, , represents the network parameter related to a certain neuron, represents the loss function, represents the gradient operator, indicating to take the gradient of the parameter ;

[0161] S600, set a suitable pruning threshold according to the gradients of all neurons to the loss function, and delete the neurons and links with low contribution degree based on this;

[0162] S700 uses linear quantization technology to convert the weights represented by 32-bit floating-point numbers in the double deep Q network model into 8-bit or 16-bit integer representation, and its formula is: or ; where, Q represents the quantized fixed-point data, (•) represents the rounding function, used to round the calculation result, represents the original floating-point data, represents the scaling coefficient, represents the offset;

[0163] Set the maximum and minimum values of the floating-point parameter and the fixed-point parameter to , , , , then there is:

[0164] , where, represents the scaling coefficient, Represents the maximum value of the floating-point parameter, Represents the minimum value of the floating-point parameter, Represents the maximum value of the fixed-point parameter, Represents the minimum value of the fixed-point parameter;

[0165] , where, Represents the offset, Represents the maximum value of the fixed-point parameter, Represents the maximum value of the floating-point parameter, Represents the scaling factor;

[0166] S800, please refer to Figure 3 , use dynamic programming to refine the resource allocation in real time, achieve the maximization of task revenue under the condition of meeting resource constraints, and complete the real-time resource allocation of the aircraft;

[0167] Furthermore, the step S800 specifically includes:

[0168] S810, for emergencies, estimate the amount of resources required for the emergency, and judge whether the reserved resources are sufficient; form a task queue with the emergency and the non-critical tasks that release resources, and divide it into multiple task sub-queues according to the priority;

[0169] S820, initialize and define the dynamic programming algorithm; include defining the state variables in dynamic programming, defining the decision variables in dynamic programming, defining the state transition matrix in dynamic programming, and defining the task revenue in the dynamic programming algorithm;

[0170] S830, adopt distributed computing, use the dynamic programming algorithm in each task sub-queue at the same time, recursively calculate from the highest to the lowest priority order, record the optimal decision at each stage, and then perform forward optimization to finally obtain the optimal decision sequence of the original problem, and achieve the optimal strategy for the remaining resources allocation of the aircraft system for random emergency tasks of the aircraft;

[0171] Furthermore, the step S810 specifically includes:

[0172] S811, if it is sufficient, use the reserved resources to handle the emergency; if it is not enough, release the resources of the currently executing non-critical tasks and re-allocate the resources to the critical tasks; form a task queue with the emergency and the non-critical tasks that release resources, and divide it into multiple sub-queues according to the priority, and at the same time perform corresponding processing on the current total remaining resources;

[0173] S812, calculate the priority of each task according to the following formula:

[0174] ; where, P i Represents the priority of the task, Di Indicates the deadline of the task, ε Indicates the buffer factor for task delay, which is used to avoid resource conflicts caused by excessive prioritization, W i Indicates the task weight;

[0175] S813, set the criteria to divide the task queue into n levels of task sub-queues; among them, the criteria can be set according to actual needs. For example, in this embodiment, tasks can be classified according to the absolute value of the priority. The first 20% is the first-level task sub-queue, 20 - 50% is the second-level task sub-queue, and 50 - 100% is the third-level task sub-queue;

[0176] Furthermore, step S820 specifically includes:

[0177] S821, obtain the total number of tasks in the task queue N ;

[0178] S822, define the state variable in dynamic programming as ; Indicates the amount of resources allocated for the i th task to the th task;

[0179] S823, define the decision variable in dynamic programming as ; Indicates the amount of resources allocated to the i th task, ; and define the dynamic programming state transition matrix, the formula is: ;

[0180] S824, define the benefit function of allocating the amount of resources to task as , the formula is:

[0181] ; among them, Indicates the minimum amount of resources required to complete the task; Indicates the amount of resources allocated to ; Indicates task 's weight, Indicates the coefficient;

[0182] Furthermore, step S830 specifically includes:

[0183] S831, adopt distributed computing, and let independent computing nodes process the dynamic programming algorithm while nProcess the sub-queues of tasks with multiple priorities, and summarize the allocation results at the same time;

[0184] S832, Use the dynamic programming recurrence relation to calculate section by section, and finally obtain the optimal solution to the problem; The formula is as follows:

[0185] ;

[0186] Among them, represents the amount of resources allocated for the i th task to the th task, represents the amount of resources allocated to the i th task, represents the total number of tasks obtained in the task queue, represents the i th task in the state The maximum benefit, represents the allocation The benefit function of the amount of resources to task , represents the i+1 th task's maximum benefit, represents the th task's maximum benefit, represents the task in the dynamic programming algorithm The benefit is , represents the amount of resources allocated to the th task, represents the amount of resources allocated for the th task;

[0187] S833, Adopt a strategy that combines heuristic search and dynamic programming, and give priority to high-benefit paths; The dynamic programming path is defined by the following formula: , Among them, represents the total actual benefit from the task with the highest priority to the task with the current priority of i , represents the total estimated benefit of the remaining tasks, represents from the task with the highest priority to the task with the current priority of the i th task, and then from the task with the priority of the i th task to the comprehensive path benefit of the task with the lowest priority, j represents the number of remaining tasks.

[0188] The real-time resource allocation method for the aircraft in this embodiment adopts a distributed sensing network, divides the resource status into different regional subnets, each subnet independently monitors and uploads resource usage data, and the master node optimizes the allocation strategy according to the global resource status. The technical solution of this embodiment has an adaptive strategy, which can dynamically adjust algorithm parameters to adapt to environmental changes by real-time monitoring the operating environment. Please refer to Figure 4 , Figure 4 , which is a bar chart showing the performance comparison between this embodiment and other existing technologies. As can be seen from Figure 4 , the task completion rate of the technical solution of this embodiment reaches 97.8%, which is significantly higher than other algorithms, indicating that the dynamic allocation strategy for resources in this embodiment effectively guarantees the completion of tasks; the resource utilization rate of the allocation method in this embodiment is 92.5%, indicating that this embodiment can maximize the resource usage efficiency under limited resource conditions; the response time of the allocation method is only 18.4 milliseconds, which is much lower than other algorithms, reflecting that the technical solution of this embodiment has significant advantages in terms of real-time performance.

[0189] Furthermore, in combination with Figures 1 to 3 , the above technical solution of the present invention will be described. Figure 1 , which is a schematic diagram of the dual-depth Q network structure constructed by the present invention. As can be seen from Figure 1 , the dual-depth Q network structure includes an input layer, a hidden layer, an activation function, and an output layer. The dimension of the state vector is 13, the dimension of the action vector is 10, each hidden layer contains 128 neurons, and the activation function is the ReLU function.

[0190] Please refer to Figure 2 , Figure 2 , which is a schematic diagram of the training process of the dual-depth Q network. As can be seen from Figure 2 , training experience quadruples are sampled from the experience replay buffer. Based on the data in the experience replay buffer, the dual-depth Q network (main network) estimates the current state value function and estimates the target value function of the next state through the target network , and uses the target value function to update the value function of the current state , and updates the value function of the current state through the target value function . By comparing the target value function with the updated value function Calculate the minimized mean squared error loss function. For the obtained minimized mean squared error loss function, use the stochastic gradient descent method to update the parameters of the main network, and periodically copy the parameters of the main network to the target network. Further, during the construction of the experience replay buffer, input the state of each agent into the double deep Q network (main network), and select actions according to the greedy algorithm in the current action set space (i.e., greedy action selection), and feedback to the environment. Further, based on the trained double deep Q network model, perform model lightweight design. The lightweight design includes neuron pruning technology and quantization technology. The pruning technology is as follows: Take the gradient of the mean squared error loss function with respect to the network parameters related to neurons, obtain the contribution degree of each neuron by comparing the absolute value of the gradient, and delete redundant neurons and invalid connections with low contribution based on the set pruning threshold. The quantization technology is as follows: Convert the weights represented by 32-bit floating-point numbers in the double deep Q network model into 8-bit or 16-bit integers.

[0191] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the overall process of the real-time resource allocation method for an aircraft. Existing aircraft can include a resource perception module, an intelligent decision-making module, and a task priority scheduling module. After the aircraft is started, it is judged whether a sudden situation occurs. If the judgment is no ( N ), the optimal resource allocation strategy trained by the double deep Q network can be directly used through the intelligent decision-making module to complete resource allocation. If the judgment is yes ( Y ), on the one hand, through the resource perception module, calculate the amount of resources required for the sudden event and judge whether the reserved resources are sufficient. On the other hand, through the task priority scheduling module, perform the priority evaluation of the sudden event and the priority evaluation of the non-critical tasks of the removed resources. Then, divide the task queue into task sub-queues according to the priority evaluation, and then adopt distributed computing. The task sub-queues are processed by independent computing nodes using the dynamic programming algorithm, and at the same time, summarize the allocation results, so as to achieve real-time resource allocation. During the process of judging whether the reserved resources are sufficient, if the judgment is no ( N ), release the resources of the current non-critical tasks and re-allocate the resources to the critical tasks. If the judgment is yes ( Y ), use the reserved resources to solve the sudden tasks. After the judgment is completed, through parallel processing technology, real-time resource allocation is achieved, that is, after the judgment is completed, there is sufficient resource amount, and then distributed computing is adopted. The task sub-queues are processed by independent computing nodes using the dynamic programming algorithm, and at the same time, summarize the allocation results, so as to achieve the real-time resource allocation algorithm based on reinforcement learning and dynamic programming (i.e., the real-time resource allocation method of the present invention).

[0192] In a second aspect, this embodiment provides a real-time resource allocation system for an aircraft, including:

[0193] A priority queue determination module, configured to: construct a task priority queue and a resource status matrix for the aircraft, and define a state space; allocate the task requirements of the aircraft to the overall system resources to define an action space, and define a reward function according to the dynamic resource allocation goal;

[0194] A network structure construction module, configured to: build a real-time control software environment for the aircraft; prepare training data for a double deep Q network, and construct a double deep Q network structure;

[0195] A parameter update module for the main network, configured to: construct an experience replay buffer, perform training on the double deep Q network, and update the parameters of the main network;

[0196] An optimal decision-making module, configured to: obtain a value function between the state and actions of the aircraft, and enable the aircraft to make an optimal decision on resource allocation through the value function;

[0197] A dynamic programming module, configured to: use dynamic programming to refine the allocated resources in real time, achieve the maximization of task benefits under resource constraint conditions, and complete the real-time resource allocation for the aircraft.

[0198] In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0199] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for real-time resource allocation of an aircraft, characterized in that: include: Construct the mission priority queue and resource status matrix of the aircraft and define the state space; The action space is defined by allocating the mission requirements of the aircraft to the overall system resources, and the reward function is defined according to the dynamic resource allocation objective; Build the real-time control software environment for the aircraft; prepare the training data for the dual-depth Q network and construct the dual-depth Q network structure; Constructing an experience replay buffer, performing the dual-depth Q network training, and updating the parameters of the main network; Obtaining a value function between the state and action of the aircraft, and enabling the aircraft to make an optimal decision on resource allocation through the value function; For emergencies, estimate the amount of resources required for the emergencies and determine whether the reserved resources are sufficient; form a task queue with emergencies and non-critical tasks that release resources, and divide them into multiple task sub-queues according to priority; Perform initialization definition of dynamic programming algorithm; including definition of state variables in dynamic programming, definition of decision variables in dynamic programming, definition of dynamic programming state transfer matrix, and definition of task benefits in dynamic programming algorithm; Distributed computing is adopted, and a dynamic programming algorithm is used in each task sub-queue at the same time. The optimal decision of each stage is recorded recursively from high to low priority, and then forward optimization is performed to finally obtain the optimal decision sequence of the original problem, and realize the optimal strategy for allocating the remaining resources of the aircraft system for random burst tasks of the aircraft; The step of constructing an experience playback buffer and performing the dual-depth Q network training includes: The state of each agent is input into the dual-depth Q network, and the .... The greedy algorithm is used to select actions and the actions of all agents are applied to the environment. The action selection strategy is expressed as: ;in, Indicates the current state. Indicates the optional actions in the current state. Indicates that the current state makes The biggest move, represents the exploration rate, , Indicates that the function is found When maximized , Represents the set of all possible actions in the current state; Introducing a target network, wherein the target network has the same structure as the main network and the two are independent of each other; Initialize the parameters of the main network and the target network; At each time step, according to the current state Use the main network to select an action ; Perform the selected action , observe the next state and instant rewards ; Determine the agent's training experience quadruple as , and the four-tuple The training data is stored in the experience playback buffer.

2. The method for real-time resource allocation of aircraft according to claim 1, characterized in that: The construction of the mission priority queue and resource state matrix of the aircraft and the definition of the state space include: The priority queue is defined and is represented by the following formula: ; in, , P i Indicates the priority of the task, D i represents the deadline of the task, ε represents the buffer factor of task delay, which is used to avoid resource conflicts caused by excessive priority, and W i represents the task weight; The state space S is constructed based on the current aircraft task queue and resource state matrix R; among them, the task queue , T n Indicates the nth aircraft task in the task queue; Resource Status Matrix , R c represents the system computing resources, R m Represents system memory resources, R b Indicates system bandwidth resources; Set the aircraft mission T i Assign to resource R j In the example, the action space A is defined as the action space ,in, Indicates The aircraft mission T i Assign to resource R j Methods, resources , Indicates the allocated system computing resources, Indicates the allocated system memory resources. Indicates the allocated system bandwidth resources.

3. The method for real-time resource allocation of aircraft according to claim 2, characterized in that: The step of preparing the training data of the dual-depth Q network comprises: Prepare task data: Generate a task queue containing 1,000 tasks, and randomly assign requirements to each task; the requirements include task computing requirements, task memory requirements, and task bandwidth requirements; among them, task computing requirements , in MHz, task memory requirement , in MB, task bandwidth requirement , in Mbps, task deadline , unit is s; Preparing system resource data: total computing resources ; Total memory resources ; Total bandwidth resources ; The construction of the dual-depth Q network structure includes: Define the input layer: the task queue dimension is 10, the system resource status dimension is 3, and the total input layer dimension is 13; Define the hidden layer: There are two hidden layers, each with 128 neurons, and the activation function is ; Define the output layer: the dimension of the action set is 10.

4. The method for real-time resource allocation of aircraft according to claim 3, characterized in that: The updating of the parameters of the main network includes: Randomly sample a batch of experience quads from the experience replay buffer ; For any sampled empirical quadruple, use the dual-depth Q network to estimate the current state The value function With the next state The target value function ; Calculate the network parameters when minimizing the mean square error loss function and update the parameters of the main network.

5. The method for real-time resource allocation of aircraft according to claim 4, characterized in that: The randomly sampling of a batch of experience quadruplets from the experience playback buffer ; For any sampled empirical quadruple, use the dual-depth Q network to estimate the current state The value function and the next state The target value function ; Calculate the network parameters that minimize the mean square error loss function and update the parameters of the main network, including: S100, use the target network to calculate the next state in the sampled experience quadruple The maximum value function Corresponding actions ; S200, use the main network to estimate the current state of the sampled experience quadruple The value function ; S300, use next state Execute an action The target value function Update current status The value function The value is: ; in, Indicates the current state. Indicates the optional actions in the current state. represents the Q-value estimate of the current state, represents the updated Q-value estimate, is the weight factor, and , Indicates the current state Execute an action After that, the reward value returned by the environment is represents the target value function, Indicates the next state of the current state. Indicates status The action that maximizes Q is is the discount factor, and ; S400, obtaining a mean square error loss function, whose formula is: ; in, represents the loss function, represents the expectation of all possible data distributions, Indicates the current state. Indicates the optional actions in the current state. Indicates in status Execute an action After that, the reward value returned by the environment is is the discount factor, and , Indicates the next state of the current state. Indicates status The action that maximizes Q is represents the parameters of the target network, represents the Q value estimate of the main network in the current state, Represents the main network parameters, Indicates status Estimate the Q value of the target network; S500, update the parameters of the main network according to the stochastic gradient descent method, the formula is: ; in, represents the new parameter value of the main network after a gradient descent update, represents the parameter value of the main network at the beginning of the current iteration step, represents the learning rate, represents the loss function; is the gradient operator, which represents the parameter Find the gradient; S600, copying the parameters of the main network to the target network; S700, accelerates dual-depth Q network training through GPU parallel training; The step of enabling the aircraft to make an optimal decision on resource allocation by using the value function includes: Using the trained dual-depth Q network model, input the system at a certain moment Status ; The dual deep Q network model is based on the input state Output the best action , and obtain the optimal aircraft real-time resource allocation decision.

6. The method for real-time resource allocation of aircraft according to claim 1, characterized in that: After the aircraft makes the optimal decision on resource allocation through the value function, the method further includes: S100 measures the contribution of a neuron to the result according to the gradient of the neuron to the loss function; the gradient of the neuron to the loss function is expressed by the following formula: ;in, , represents the network parameters associated with a neuron, represents the loss function, represents the gradient operator, which represents the parameter Find the gradient; S200, setting a suitable pruning threshold according to the gradient of all neurons to the loss function, and based on this, deleting neurons and links with small contributions; S300 uses linear quantization technology to convert the weights represented by 32-bit floating point numbers in the dual-depth Q network model into 8-bit or 16-bit integer representations. The formula is: or ;in, Represents quantized fixed-point data, (•) indicates the rounding function, which is used to round the calculation result. Represents raw floating point data, represents the scaling factor, Indicates the offset; Set the maximum value of floating-point parameters and fixed-point parameters to , , , , then: ,in, represents the scaling factor, Indicates the maximum value of a floating-point parameter. Indicates the minimum value of a floating-point parameter. Indicates the maximum value of a fixed-point parameter, Indicates the minimum value of the fixed-point parameter; ,in, Indicates the offset, Indicates the maximum value of a fixed-point parameter, Indicates the maximum value of a floating-point parameter. Indicates the scaling factor.

7. The method for real-time resource allocation of aircraft according to claim 6, characterized in that: For emergencies, the amount of resources required for emergencies is estimated, and it is determined whether the reserved resources are sufficient; emergencies and non-critical tasks for releasing resources are combined into a task queue, and divided into multiple task sub-queues according to priority, including: If it is judged to be sufficient, the reserved resources are used to deal with emergencies; if it is judged to be insufficient, the resources currently executing non-critical tasks are released and reallocated to critical tasks; The priority of each task is calculated according to the following formula: ;in, , P i Indicates the priority of the task, D i represents the deadline of the task, ε represents the buffer factor of task delay, which is used to avoid resource conflicts caused by excessive priority, and W i represents the task weight; Divide the task queue into n-level task sub-queues according to priority; The initialization definition of the dynamic programming algorithm includes defining the state variables in the dynamic programming, defining the decision variables in the dynamic programming, defining the dynamic programming state transfer matrix, and defining the task benefits in the dynamic programming algorithm, including: Get the total number of tasks N in the task queue; Define the state variables in dynamic programming as ; Indicates the allocation for the i-th task to the The amount of resources for each task; Define the decision variables in dynamic programming as ; represents the amount of resources allocated to the i-th task, ; And define the dynamic programming state transfer matrix, the formula is: ; Define the amount of resources allocated in the dynamic programming algorithm Give Tasks The profit function is , the formula is: ;in, Indicates the minimum amount of resources required to complete the task; Indicates that the The amount of resources; Indicates the task The weight of Represents the coefficient.

8. The method for real-time resource allocation of aircraft according to claim 1, characterized in that: The distributed computing is used, and the dynamic programming algorithm is used in each task sub-queue at the same time. The optimal decision of each stage is recorded recursively from high to low priority, and then the optimal decision sequence of the original problem is finally obtained, so as to realize the optimal strategy for allocating the remaining resources of the aircraft system for the random burst tasks of the aircraft, including: Distributed computing is adopted, where independent computing nodes process the dynamic programming algorithm to process n priority task subqueues simultaneously and summarize the allocation results; Use the dynamic programming recursive relationship to calculate step by step, and finally find the optimal solution to the problem; the formula is as follows: ; in, Indicates the allocation for the i-th task to the The amount of resources for a task, represents the amount of resources allocated to the i-th task, Indicates the total number of tasks in the task queue. Indicates that the i-th task is in state The maximum profit under Indicates the amount of allocated resources Give Tasks The profit function, represents the maximum benefit of the i+1th task, Indicates The maximum benefit of a task, Represents the task in the dynamic programming algorithm The income is , Indicates that the The amount of resources for a task, Indicates allocation for The amount of resources for each task; A strategy combining heuristic search and dynamic programming is adopted to give priority to high-yield paths; the dynamic programming path is defined as the following formula: ,in, represents the total actual benefit from the task with the highest priority to the task with the current priority i, represents the total estimated revenue of the remaining tasks, It represents the comprehensive path benefit from the task with the highest priority to the task with the current priority i, and then from the task with the priority i to the task with the last priority, and j represents the number of remaining tasks.

9. A real-time resource allocation system for aircraft, characterized in that: The aircraft real-time resource allocation method according to any one of claims 1 to 8 is used to implement the aircraft real-time resource allocation system, comprising: The priority queue determination module is used to: construct the mission priority queue and resource state matrix of the aircraft and define the state space; define the action space by allocating the mission requirements of the aircraft to the overall system resources, and define the reward function according to the dynamic resource allocation target; The network structure building module is used to: build the real-time control software environment of the aircraft; prepare the training data of the dual-depth Q network and build the dual-depth Q network structure; The parameter update module of the main network is used to: construct an experience playback buffer, perform the dual-depth Q network training, and update the parameters of the main network; The optimal decision module is used to obtain the value function between the state and action of the aircraft, and enable the aircraft to make the optimal decision on resource allocation through the value function; The dynamic programming module is used to: use dynamic programming to refine resource allocation in real time, maximize mission benefits while satisfying resource constraints, and complete real-time resource allocation for aircraft.

Citation Information

Patent Citations

  • D3QN-based 6G super-large-scale Internet of Vehicles network resource allocation method and system

    CN117939486A

  • Intelligent scheduling method and system for aircraft jumping out of hangar

    CN119250448A