An intelligent power distribution method and system for energy guarantee based on reinforcement learning
Through the intelligent power distribution method based on DDPG reinforcement learning algorithm, the scope of application and time-consuming problems of existing energy guarantee task planning are solved, and fast and intelligent energy distribution is achieved, and task efficiency is improved.
Patent Information
- Application Number
- CN202510393573.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing energy guarantee task planning methods are limited in scope and take a long time, making it difficult to meet the needs of rapid response and real-time adjustment.
An agent is constructed using a deep determination strategy gradient (DDPG) reinforcement learning algorithm, and by training the agent to output the ratio of the electric power of heat transfer and heat power supply, intelligent power distribution is achieved.
Energy guarantee planning in multiple states has been realized, time consumption of guarantee tasks has been shortened, and the intelligence level of tasks has been improved.
Smart Images

Figure CN119904079B_ABST
Abstract
Description
Background Art
[0002] For field station equipment without preset conditions, it is necessary to solve the fuel and power usage requirements through the large aircraft platform to achieve the non-preset and non-carrying of refueling trucks and power supply vehicles. Therefore, research on cross-aircraft type hot fuel transfer technology and cross-aircraft type multi-power system hot power supply technology and equipment development are carried out.
[0003] Both fuel transfer from large aircraft to small aircraft and power supply require electric power output. However, the output power of the auxiliary power unit (APU) of the large aircraft platform is limited. Therefore, it is necessary to plan the energy guarantee task according to the fuel quantity and power demand of the small aircraft, and adaptively allocate the electric power output of hot fuel transfer and hot power supply. The current energy guarantee task planning methods generally have certain limitations in practical applications, usually only applicable to specific system states or environmental conditions, which limits their scope of application. In addition, when these methods carry out energy guarantee task planning, they often take a long time and are difficult to meet the requirements of rapid response and real-time adjustment. Therefore, exploring a more general and efficient energy guarantee task planning method has become an important issue that needs to be solved urgently.
[0004] Therefore, it is hoped that there is a technical solution to overcome or at least alleviate at least one of the above defects of the existing technology. Summary of the Invention
[0005] The purpose of this application is to provide an intelligent power distribution method and system for energy guarantee based on reinforcement learning to solve at least one problem existing in the prior art.
[0006] The technical solution of this application is as follows:
[0007] The first aspect of this application provides an intelligent power distribution method for energy guarantee based on reinforcement learning, including:
[0008] Step S10: Construct a DDPG agent, train the DDPG agent, and obtain the trained DDPG agent;
[0009] Step S20: Input the state variables of the small aircraft obtained in real time into the trained DDPG agent. The DDPG agent outputs the electric power ratio of hot fuel transfer and hot power supply, and the large aircraft conducts intelligent power distribution for the small aircraft according to the electric power ratio of hot fuel transfer and hot power supply.
[0010] In at least one embodiment of this application, in step S10, training the DDPG agent to obtain the trained DDPG agent includes:
[0011] Step S11: Obtain the state variables of the small aircraft;
[0012] Step S12: Input the state variables of the small aircraft at time t into the DDPG agent, and the DDPG agent outputs the electric power ratio of hot fuel transfer and hot power supply. ;
[0013] Step S13: Calculate the electric power ratio corresponding to the decision reward value of the DDPG agent. ;
[0014] Step S14: Update the internal parameters of the DDPG agent according to the decision reward value to obtain the trained DDPG agent;
[0015] Step S15: Repeat the above steps S11 - S14 to obtain the trained DDPG agents in different states.
[0016] In at least one embodiment of the present application, in step S12, the state variables input into the DDPG agent include: the fuel quantity of the small aircraft at time t , the power quantity of the small aircraft at time t , the fuel quantity to be transferred of the small aircraft at time t , the power quantity to be supplied of the small aircraft at time t. .
[0017] In at least one embodiment of the present application, in step S13, the decision reward value is:
[0018] ;
[0019] where is the fuel quantity of the small aircraft at time t + Δt, is the power quantity of the small aircraft at time t + Δt, is the fuel quantity of the small aircraft at time t, is the power quantity of the small aircraft at time t, is the fuel quantity to be transferred of the small aircraft at time t, is the power quantity to be supplied of the small aircraft at time t, is the total duration consumed to complete the energy guarantee of hot fuel transfer and hot power supply, is the total fuel transfer quantity of the small aircraft from time t to time t + Δt, is the total power supply quantity of the small aircraft from time t to time t + Δt.
[0020] In at least one embodiment of the present application, in step S20, the state quantity of the small aircraft obtained in real time is input into the trained DDPG agent, and the DDPG agent outputs the electric power ratio of heat oil transportation to heat power supply. The large aircraft performs intelligent power distribution for the small aircraft according to the electric power ratio of heat oil transportation to heat power supply, including:
[0021] Step S21: Obtain the state quantity of the small aircraft at the current moment in real time;
[0022] Step S22: Input the state quantity of the small aircraft at the current moment into the trained DDPG agent, and the DDPG agent outputs the electric power ratio of heat oil transportation to heat power supply of the small aircraft at the current moment;
[0023] Step S23: The large aircraft performs real-time intelligent power distribution for the small aircraft according to the electric power ratio of heat oil transportation to heat power supply of the small aircraft at the current moment.
[0024] The second aspect of the present application provides an intelligent power distribution system for energy guarantee based on reinforcement learning, including:
[0025] A small aircraft;
[0026] A DDPG agent, configured to obtain the electric power ratio of heat oil transportation to heat power supply according to the above-mentioned intelligent power distribution method for energy guarantee based on reinforcement learning;
[0027] A large aircraft, configured to perform intelligent power distribution for the small aircraft according to the electric power ratio of heat oil transportation to heat power supply.
[0028] In at least one embodiment of the present application, the large aircraft transports oil for the small aircraft through a heat oil transportation device and supplies power to the small aircraft through a heat power supply device.
[0029] The third aspect of the present application provides a computer-readable medium storing computer-executable instructions for executing the above-mentioned intelligent power distribution method for energy guarantee based on reinforcement learning.
[0030] The fourth aspect of the present application provides a computing device, including:
[0031] At least one processor, and a memory communicatively connected to the at least one processor; wherein,
[0032] The memory stores instructions executable by the at least one processor for executing the above-mentioned intelligent power distribution method for energy guarantee based on reinforcement learning.
[0033] The invention has at least the following beneficial technical effects:
[0034] The energy guarantee intelligent power distribution method based on reinforcement learning of the present application can intelligently allocate the electric power required for oil transportation and power supply based on the reinforcement learning algorithm according to the actual guarantee requirements, can realize the energy guarantee planning in multiple states, can shorten the time consumption of the guarantee task, and improve the intelligent level of the guarantee task. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of the energy guarantee intelligent power distribution method based on reinforcement learning according to an embodiment of the present application;
[0036] Figure 2 is a schematic diagram of the energy guarantee intelligent power distribution system based on reinforcement learning according to an embodiment of the present application;
[0037] Figure 3 is a schematic diagram of the hardware structure of a computing device for implementing the energy guarantee intelligent power distribution method based on reinforcement learning according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the purpose, technical solutions and advantages of the implementation of the present application clearer, the technical solutions in the embodiments of the present application will be described in more detail below with reference to the accompanying drawings in the embodiments of the present application. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, of the embodiments of the present application. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application. The embodiments of the present application will be described in detail below with reference to the drawings.
[0039] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the protection scope of the present application.
[0040] The following will further elaborate on the present application in conjunction with the attached Figures 1 to 3 to further elaborate on the present application.
[0041] The first aspect of the present application provides an energy guarantee intelligent power distribution method based on reinforcement learning, including the following steps:
[0042] Step S10: Construct a DDPG agent, train the DDPG agent, and obtain a trained DDPG agent;
[0043] Step S20: Input the state variables of the small aircraft obtained in real time into the trained DDPG agent. The DDPG agent outputs the power ratio of hot fuel transfer to hot power supply, and the large aircraft distributes power intelligently to the small aircraft according to the power ratio of hot fuel transfer to hot power supply.
[0044] For the intelligent power distribution method based on reinforcement learning in this application, the DDPG (Deep Deterministic Policy Gradient) agent is a reinforcement learning agent based on the deep deterministic policy gradient algorithm. As Figure 1 shown, in step S10, training the DDPG agent to obtain a trained DDPG agent includes:
[0045] Step S11: Obtain the state variables of the small aircraft;
[0046] Step S12: Input the state variables of the small aircraft at time t into the DDPG agent, and the DDPG agent outputs the power ratio of hot fuel transfer to hot power supply ;
[0047] Step S13: Calculate the decision reward value of the DDPG agent corresponding to the power ratio ; ;
[0048] Step S14: Update the internal parameters of the DDPG agent according to the decision reward value to obtain a trained DDPG agent;
[0049] Step S15: Repeat the above steps S11 - S14 to obtain trained DDPG agents in different states.
[0050] During the training stage of the DDPG agent, first in step S11, obtain the state variable data of the small aircraft in a continuous time period in one state, and extract multiple groups of state variable data of the small aircraft at the front and back moments with a time interval of Δt from these state variable data as the training set for training the DDPG agent. In step S12, take the fuel quantity of the small aircraft at time t , the power quantity of the small aircraft at time t , the fuel quantity to be transferred of the small aircraft at time t , and the power supply quantity to be supplied of the small aircraft at time t as the input of the DDPG agent, and use the DDPG algorithm in the DDPG agent to make a decision and output the power ratio of hot fuel transfer to hot power supply according to the input state variables of the small aircraft. , is the electric power for hot oil transportation, is the electric power for hot power supply.
[0051] Then, in step S13, the DDPG algorithm is used to calculate the decision reward value corresponding to the electric power ratio :
[0052] ;
[0053] wherein, is the fuel quantity of the small aircraft at time t + Δt, is the electric quantity of the small aircraft at time t + Δt, is the fuel quantity of the small aircraft at time t, is the electric quantity of the small aircraft at time t, is the fuel quantity to be transported of the small aircraft at time t, is the power supply quantity to be supplied of the small aircraft at time t, is the total duration consumed for completing the energy guarantee of hot oil transportation and hot power supply, is the total fuel transportation quantity of the small aircraft from time t to time t + Δt, is the total power supply quantity of the small aircraft from time t to time t + Δt.
[0054] In step S14, the internal parameters of the DDPG agent are updated according to the decision reward value to obtain a trained DDPG agent in one state. Finally, in step S15, in the same way as steps S11 - S14, the DDPG agent is trained by the state quantities of the small aircraft in different states to obtain trained DDPG agents in different states. The trained DDPG agent can take the state quantities of the small aircraft as inputs and decision - output the electric power ratio of hot oil transportation and hot power supply, so as to perform intelligent power distribution for the small aircraft.
[0055] The intelligent power distribution method for energy guarantee based on reinforcement learning of the present application, as Figure 1 shown, in step S20, the state quantities of the small aircraft obtained in real - time are input into the trained DDPG agent, and the DDPG agent outputs the electric power ratio of hot oil transportation and hot power supply. The large aircraft performs intelligent power distribution for the small aircraft according to the electric power ratio of hot oil transportation and hot power supply, including:
[0056] Step S21, obtain the state quantities of the small aircraft at the current moment in real - time;
[0057] Step S22, input the state quantities of the small aircraft at the current moment into the trained DDPG agent, and the DDPG agent outputs the electric power ratio of hot oil transportation and hot power supply of the small aircraft at the current moment;
[0058] Step S23: The large aircraft performs real-time intelligent power distribution for the small aircraft according to the electric power ratio of hot fueling and hot power supply of the small aircraft at the current moment.
[0059] In the power distribution stage of the DDPG agent, the state quantity of the small aircraft at the current moment is input into the trained DDPG agent to obtain the electric power ratio of hot fueling and hot power supply at the current moment. The large aircraft performs real-time intelligent power distribution for the small aircraft according to this electric power ratio. After a predetermined time T, the small aircraft enters the next state, and the real-time intelligent power distribution for the small aircraft is also performed according to the above process until the hot fueling and hot power supply tasks are completed.
[0060] The intelligent power distribution method for energy guarantee based on reinforcement learning in this application trains a DDPG agent to intelligently decide the electric power ratio of hot fueling and hot power supply. The large aircraft distributes electric power for the small aircraft during the hot fueling and hot power supply processes according to this electric power ratio, realizing rapid energy supply in multiple states. This application improves the energy supply speed, optimizes the guarantee duration, and simultaneously enhances the intelligent level of the guarantee task.
[0061] The second aspect of this application provides an intelligent power distribution system for energy guarantee based on reinforcement learning, as Figure 2 shown. The system includes a small aircraft, a DDPG agent, and a large aircraft. The DDPG agent obtains the electric power ratio of hot fueling and hot power supply according to the above-mentioned intelligent power distribution method for energy guarantee based on reinforcement learning. The large aircraft performs intelligent power distribution for the small aircraft according to the electric power ratio of hot fueling and hot power supply. Among them, the large aircraft supplies fuel to the small aircraft through a hot fueling device and supplies power to the small aircraft through a hot power supply device.
[0062] The third aspect of this application provides a computer-readable medium storing computer-executable instructions for executing the above-mentioned intelligent power distribution method for energy guarantee based on reinforcement learning.
[0063] For the convenience of description, the above parts are divided into various modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.
[0064] After introducing the intelligent power distribution method, system, and readable medium for energy guarantee based on reinforcement learning in the exemplary embodiments of this application, next, a computing device according to another exemplary embodiment of this application is introduced.
[0065] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.
[0066] In some possible implementation manners, the computing device of the present application may include at least one processing unit and at least one storage unit. Among them, the storage unit stores program code, and when the program code is executed by the processing unit, the processing unit executes the steps in the above-mentioned intelligent power distribution method for energy guarantee based on reinforcement learning according to various exemplary implementation manners of the present application in this specification.
[0067] Next, refer to Figure 3 to describe the computing device 30 according to this implementation manner of the present application. Figure 3 The shown computing device 30 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0068] As Figure 3 shown, the computing device 30 is presented in the form of a general computing device. The components of the computing device 30 may include but are not limited to: the above-mentioned at least one processing unit 31, the above-mentioned at least one storage unit 32, and a bus 33 connecting different system components (including the storage unit 32 and the processing unit 31).
[0069] The bus 33 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.
[0070] The storage unit 32 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322, and may further include a read-only memory (ROM) 323.
[0071] The storage unit 32 may further include a program / utility 325 having a set (at least one) of program modules 324. Such program modules 324 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0072] The computing device 30 may also communicate with one or more external devices 34 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the computing device 30, and / or communicate with any device that enables the computing device 30 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be performed through an input / output (I / O) interface 35. Also, the computing device 30 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 36. As shown in the figure, the network adapter 36 communicates with other modules for the computing device 30 through a bus 33. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computing device 30, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0073] In some possible implementation manners, each aspect of the energy guarantee intelligent power distribution method based on reinforcement learning provided in this application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the energy guarantee intelligent power distribution method based on reinforcement learning according to various exemplary implementation manners of this application described above in this specification.
[0074] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0075] The program product of the energy guarantee intelligent power distribution method based on reinforcement learning in the implementation manner of this application may adopt a portable compact disk read-only memory (CD-ROM) and include program code, and may run on a computing device. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program, and this program may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0076] A readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal can take various forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0077] The program code contained on a readable medium can be transmitted using any appropriate medium, including - but not limited to - wireless, wired, optical fiber, RF, etc., or any suitable combination of the above. The program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0078] It should be noted that although several units or subunits of the apparatus are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more of the above-mentioned units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0079] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0080] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0081] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0082] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0084] As mentioned above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An intelligent power distribution method for energy guarantee based on reinforcement learning, characterized in that, Including: Step S10: Construct a DDPG agent, train the DDPG agent, and obtain the trained DDPG agent; Step S20: Input the state variables of the small aircraft obtained in real time into the trained DDPG agent. The DDPG agent outputs the power ratio of thermal fuel delivery to thermal power supply, and the large aircraft performs intelligent power distribution for the small aircraft according to the power ratio of thermal fuel delivery to thermal power supply; In step S10, training the DDPG agent to obtain the trained DDPG agent includes: Step S11: Obtain the state variables of the small aircraft; Step S12: Input the state variables of the small airplane at time t into the DDPG agent, and the DDPG agent outputs the electric power ratio of hot oil transportation to hot power supply ; Step S13: Calculate the electric power ratio The decision reward value of the corresponding DDPG agent ; Step S14. According to the decision reward value update the internal parameters of the DDPG agent to obtain the trained DDPG agent; Step S15: Repeat the above steps S11 - S14 to obtain the trained DDPG agent in different states; In step S12, the state variables input into the DDPG agent include: the fuel quantity of the small aircraft at time t , the power quantity of the small aircraft at time t , the fuel quantity to be refueled of the small aircraft at time t , the power quantity to be supplied of the small aircraft at time t ; In step S13, the decision reward value is as follows: Among them, is the fuel quantity of the small aircraft at time t + Δt, is the power quantity of the small aircraft at time t + Δt, is the fuel quantity of the small aircraft at time t, is the power quantity of the small aircraft at time t, is the fuel quantity to be refueled of the small aircraft at time t, is the power quantity to be supplied of the small aircraft at time t, is the total duration consumed to complete the energy guarantee of hot fueling and hot power supply, is the total fuel quantity refueled by the small aircraft from time t to time t + Δt, is the total power quantity supplied by the small aircraft from time t to time t + Δt.
2. The intelligent power distribution method for energy guarantee based on reinforcement learning according to claim 1, wherein, In step S20, inputting the state variables of the small aircraft obtained in real time into the trained DDPG agent, and the DDPG agent outputs the power ratio of thermal fuel delivery to thermal power supply, and the large aircraft performs intelligent power distribution for the small aircraft according to the power ratio of thermal fuel delivery to thermal power supply, including: Step S21: Obtain the state variables of the small aircraft at the current moment in real time; Step S22: Input the state variables of the small aircraft at the current moment into the trained DDPG agent, and the DDPG agent outputs the power ratio of thermal fuel delivery to thermal power supply of the small aircraft at the current moment; Step S23: The large aircraft performs real - time intelligent power distribution for the small aircraft according to the power ratio of thermal fuel delivery to thermal power supply of the small aircraft at the current moment.
3. An intelligent power distribution system for energy guarantee based on reinforcement learning, characterized in that, Including: Small aircraft; A DDPG agent, configured to obtain the power ratio of thermal fuel delivery to thermal power supply according to the intelligent power distribution method for energy guarantee based on reinforcement learning according to any one of claims 1 to 2; Large aircraft, configured to perform intelligent power distribution for the small aircraft according to the power ratio of thermal fuel delivery to thermal power supply.
4. The intelligent power distribution system for energy guarantee based on reinforcement learning according to claim 3, wherein, The large aircraft supplies fuel to the small aircraft through a thermal fuel delivery device and supplies power to the small aircraft through a thermal power supply device.
5. A computer-readable medium storing computer-executable instructions, characterized in that, The computer - executable instructions are used to execute the intelligent power distribution method for energy guarantee based on reinforcement learning according to any one of claims 1 to 2.
6. A computing device, characterized in that, Including: At least one processor, and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are used to execute the intelligent power distribution method for energy guarantee based on reinforcement learning according to any one of claims 1 to 2.
Citation Information
Patent Citations
Multi-agent path planning method based on deep reinforcement learning
CN114815840A
Partner type thermal power supply device between airplanes
CN115513955A