Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

70 results about "Markov game" patented technology

Virtual power plant scheduling method based on large language model and deep reinforcement learning

The invention discloses a virtual power plant scheduling method based on a large language model and deep reinforcement learning, and belongs to the technical field of virtual power plant scheduling. Comprising the following steps: constructing a virtual power plant multi-agent cloud edge collaborative scheduling framework based on large language model driving; predicting wind power, photovoltaic power and load power based on a large language model; constructing a mathematical model of virtual power plant optimization scheduling; converting the virtual power plant optimization scheduling model into a Markov game process in combination with a large language model; performing initialization training on the strategy network of the edge layer intelligent agent by adopting imitation learning to obtain a pre-trained edge layer intelligent agent strategy network; and based on the pre-training strategy network of the boundary layer intelligent agent, combining with a large language model and adopting an improved multi-agent near-end strategy optimization algorithm to solve a scheduling strategy.
Owner:NANJING UNIV OF POSTS & TELECOMM

Transmission and distribution cooperative scheduling optimization method based on multi-agent deep reinforcement learning

The invention belongs to the technical field of power systems, and particularly relates to a transmission and distribution cooperative scheduling optimization method based on multi-agent deep reinforcement learning. The method comprises the following steps: firstly, establishing a transmission and distribution cooperative scheduling distributed framework, and respectively establishing distributed optimization models taking the minimization of the operation cost of a power transmission network and the minimization of the operation cost of each active power distribution network as targets; secondly, converting a transmission and distribution collaborative scheduling problem into a Markov game model with a multi-agent benefit balance characteristic, and designing a collaborative multi-agent architecture; then interacting with a transmission and distribution interconnection environment through each agent to obtain historical interaction information, storing the historical interaction information into an experience pool, and training by using an MADDPG algorithm to realize iterative updating and optimization of an Actor network and a Critic network of each agent; and finally, realizing local optimization control of the power transmission network and each active distribution network by using an Actor network.
Owner:HEFEI UNIV OF TECH

Markov game-based satellite cluster observation resource allocation method and system

The invention relates to a Markov game-based satellite cluster observation resource allocation method and system. Through a two-stage decomposition strategy, a multi-target balance problem of a task integrity rate, a resource utilization rate and a cost-efficiency ratio is effectively solved. In the first stage, a genetic algorithm (GA) is adopted to complete satellite-task-time window three-dimensional matching, and optimal distribution of limited visible windows is achieved. And in the second stage, a distributed decision-making mechanism is implemented based on an improved Improved-MADDPG framework, and after an intelligent agent adopts a random game strategy to execute exploration in a distributed manner, optimal dynamic configuration of observation resources is achieved through global information sharing and local strategy iteration, and key calculation indexes such as an average reward value and the like are remarkably improved.
Owner:NAT UNIV OF DEFENSE TECH

Active power distribution network multi-time scale rolling optimization method driven by multi-agent graph reinforcement learning

The invention provides an active power distribution network multi-time scale rolling optimization method driven by multi-agent graph reinforcement learning. The method comprises the following steps: dividing slow time scale equipment and fast time scale equipment; the slow time scale is optimized and modeled into a multi-agent Markov game, and a global state space and a combined action space corresponding to switch actions, reactive compensation equipment gears and voltage regulation equipment gears are defined; switching action candidate vectors are generated based on a basic loop of the power distribution network, and action space dimensions are compressed through a scheme of excluding conflicting actions of the same common branch; fitting an action value function of each agent by adopting a graph neural network, generating a dimensionality-reduced discrete action strategy according to the global state space, and carrying out collaborative optimization through a distributed decision and a global reward mechanism; and on the basis of the slow time scale optimization result, the fast time scale optimization problem is converted into a mixed integer second-order cone programming model, and an instruction is solved in a rolling manner.
Owner:FUZHOU UNIV

Mountain tourism route planning method based on weather and natural disasters

The invention discloses a mountain tourism route planning method based on weather and natural disasters, and the method comprises the steps: 1, building a Markov game model based on POMDP, combining Mama multi-modal feature compression and RAG semantic enhancement, designing a dynamic risk matrix integrating geological disaster probability, weather trend and ecological sensitivity, and providing partial observable environment state perception; 2, establishing a Nash equilibrium constraint optimization model, sharing an extreme weather strategy through federal learning, introducing an NSGA-II-Actor-Critic architecture to optimize unmanned aerial vehicle inspection and path adjustment, and triggering dynamic communication when rainfall exceeds a threshold value; 3, fusing a security and experience income design reward function, processing continuous action planning by using improved DDPG, and accelerating multi-target Pareto frontier convergence in combination with NSGA-II; 4, constructing a user weather sensitivity portrait, and dynamically adjusting weight balance income and AR simulation preview preference; and 5, based on a GIS closed-loop decision engine, performing agent and environment feature complementation, recording an emergency log by a block chain, and constructing a disaster case library by meta-learning.
Owner:SICHUAN TOURISM UNIV

Double-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning

The invention discloses a double-layer non-cooperative demand response method based on asynchronous multi-agent reinforcement learning, and the method comprises the steps: S1, constructing a double-layer non-cooperative game framework: constructing a double-layer dynamic game model between a public utility company and a consumer; s2, completing modeling of the partially observable Markov game, and converting the double-layer non-cooperative game into the partially observable Markov game; s3, designing an asynchronous multi-agent reinforcement learning algorithm as a game equilibrium solving engine; and S4, implementing centralized training and priority experience playback to realize asynchronous multi-agent reinforcement learning algorithm efficiency optimization. The method is suitable for a dynamic game scene between a public utility company and a consumer group in an intelligent power grid environment, and can realize efficient consumption and load fluctuation suppression of distributed energy in a power grid.
Owner:NANJING HUIQIANG NEW ENERGY TECH CO LTD

Multi-mechanical-arm accurate assembly method for double-population evolution reinforcement learning

The invention discloses a multi-mechanical-arm accurate assembly method for double-population evolution reinforcement learning, and the method comprises the steps: constructing a partially observable Markov game model, designing a mixed reward function, and building a double-population mechanism: an Exploit main population is responsible for optimizing the performance, and an Explore auxiliary population maintains the diversity; extracting unified feature representation through a shared observation embedded network, optimizing a main population by adopting conjugate gradient differential evolution, and optimizing an auxiliary population by adopting novelty driven evolution; promoting information interaction between the double populations through an elite exchange mechanism, and utilizing parameter migration fusion to reinforce learning and evolutionary optimization results; and finally, an optimal strategy set for balancing accumulated rewards and novelty is selected through Pareto optimization, and the optimal strategy set is deployed to a multi-mechanical-arm system to execute a cooperative task after convergence conditions are met. According to the method, the population diversity is effectively maintained and the cooperation robustness is improved while the high performance of the strategy is ensured; and a powerful technical support is provided for automatic assembly in a complex industrial scene.
Owner:ANHUI UNIV

Multi-agent collaborative power distribution network planning and operation integrated optimization method and system

The invention discloses a multi-agent collaborative power distribution network planning and operation integrated optimization method and system. The method comprises the steps of dividing main bodies of a power distribution network source, a power distribution network grid, a power distribution network load and a power distribution network storage into independent agents; constructing a double-layer multi-agent game model by combining a Markov game process based on the competition-cooperation relationship of each agent; the double-layer multi-agent game model comprises a planning layer game model and an operation layer game model; solving the planning layer game model to obtain the optimal configuration capacity of each agent, and calculating the constraint boundary in the action space of each agent of the operation layer; and solving the operation layer game model by using an MATD3 algorithm in which an enhanced action projection mechanism is introduced to obtain an optimal operation strategy and operation income of each agent, and feeding back the operation income to the planning layer game model to calculate a reward function of the planning layer game model. According to the invention, multi-time scale and multi-main-body cooperation characteristics can be fused, so that the system operation efficiency is improved.
Owner:WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD +1

Electric vehicle charging management method based on HA-SAC algorithm

The invention discloses an electric vehicle charging management method based on an HA-SAC algorithm, and the method comprises the steps: firstly, carrying out the modeling of an electric vehicle charging management optimization model through fully considering the life loss of a transformer and the charging preference of a user; modeling a charging management optimization problem as Markov Game (MG for short), taking each electric vehicle as an independent intelligent agent, constructing a state space according to information such as the charging state, the electricity price and the user preference, and taking the charging power as an action variable; then, on the basis of an HA-SAC algorithm, a local action network and a centralized evaluation network are constructed for each agent, and centralized training is carried out through global information; and finally, the trained strategy is deployed to each electric vehicle agent node, so that the electric vehicle agent node can autonomously decide a charging behavior only by depending on local information in actual operation, and distributed and privacy-protected cooperative charging control is realized.
Owner:ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1

Power distribution network voltage control method based on multi-agent deep reinforcement learning

The invention provides a power distribution network voltage control method based on multi-agent deep reinforcement learning, and relates to the technical field of power distribution networks, and the method comprises the steps: dividing a power distribution network into a plurality of regions, and carrying out the centralized management and optimization of the voltage of each region; a cooperative voltage control problem between regions is converted into a Markov game so as to realize effective cooperation between intelligent agents; a centralized training and decentralized execution framework is adopted, so that the intelligent agent fully utilizes global information in a training stage and independently makes a decision according to local observation in an execution stage; a self-attention mechanism is integrated, so that the agents are more focused on information related to self rewards, a multi-agent reinforcement learning technology can be combined with the self-attention mechanism, the problem of voltage fluctuation in a power distribution network is effectively solved, stable control and optimal management of voltage are realized, and the power distribution network quality is improved. The voltage regulation precision and response speed of the power distribution network are improved, and the stability and reliability of a power system are ensured.
Owner:NANJING HUIQIANG NEW ENERGY TECH CO LTD

Multi-ship cooperative autonomous collision avoidance method based on PMA-MADDPG

The invention discloses a multi-ship cooperative autonomous collision avoidance method based on PMA-MADDPG. The method comprises the following steps that ship agents and the environment where the ship agents are located are modeled into a multi-agent Markov game system; designing a reward function of the multi-agent Markov game system into a multi-target comprehensive reward function; establishing a PMA-MADDPG model based on a policy-value network structure, and introducing a long and short-term memory network and a priority experience playback mechanism into the model; the PMA-MADDPG model is trained, and after a specified training round is reached, the PMA-MADDPG model after training is completed is obtained; and inputting the current local observation information into the trained PMA-MADDPG model by each ship intelligent body, generating an optimal collision avoidance action at the next moment, and executing a collision avoidance task according to the optimal collision avoidance action. According to the method, through multi-agent collaborative decision and time sequence state perception, the ship collision avoidance safety and the navigation rule conformity in a complex maritime scene are remarkably improved.
Owner:DALIAN MARITIME UNIVERSITY

Distribution microgrid coordination control method and system based on multi-agent deep reinforcement learning

The invention discloses a distribution-microgrid coordination control method and system based on multi-agent deep reinforcement learning. The method comprises the following steps: firstly, constructing a distribution-microgrid coordination control optimization model; then, a multi-agent deep reinforcement learning framework is built, and the distribution-microgrid coordination control optimization model is converted into a Markov game; finally, an improved TD3 algorithm is used for solving the distribution-microgrid coordination control optimization model of the Markov game, an optimal distribution-microgrid coordination control strategy is generated, and the improved TD3 algorithm is based on updating rules of Taylor series optimization strategy network parameters. Coordination control optimization of the power distribution network and the micro-grid is converted into cooperation and competition of intelligent agents, the TD3 algorithm based on Taylor series optimization is used for generating the optimal power distribution-micro-grid coordination control strategy, learning efficiency and real-time performance are both considered, dynamic coordination of the power distribution network and the micro-grid is achieved, and the power distribution network and micro-grid coordination control method is suitable for large-scale popularization and application. The operation cost of the intelligent power distribution system and the multi-microgrid system is effectively reduced, and the stability of the whole system is improved.
Owner:ECONOMIC & TECH RES INST OF HUBEI ELECTRIC POWER COMPANY SGCC

Red and blue game confrontation-oriented constrained collaborative multi-agent reinforcement learning method

The invention relates to a red-blue game confrontation-oriented constrained collaborative multi-agent reinforcement learning method, which comprises the following steps of: constructing a virtual game scene and a scenario confrontation blue target, and decomposing a game confrontation task into a target collaborative allocation problem and a maneuvering strategy generation problem; respectively establishing a Markov game process of a target collaborative allocation problem and a Markov game process of a maneuvering strategy generation problem; establishing a constraint collaborative multi-agent reinforcement learning model used for solving a Markov game process of a target collaborative allocation problem, and inputting game situation information to obtain a target allocation scheme; and establishing a constraint collaborative experience-first multi-agent reinforcement learning model of a Markov game process for solving a maneuvering strategy generation problem, and inputting game situation information and maneuvering state information to obtain an optimal path of an attack target. The method can improve the utilization rate of the training sample, reduce the training time, improve the sample efficiency, and enhance the environment learning ability of the intelligent agent.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 96901

Interference frequency domain resource scheduling method and device for multi-unmanned aerial vehicle communication

The invention discloses an interference frequency domain resource scheduling method and device for multi-unmanned aerial vehicle communication. Firstly, a distributed frequency domain blocking type interference model for multi-unmanned aerial vehicle frequency hopping communication of multiple jammers based on an interference-to-signal ratio criterion is established; then taking a distributed frequency domain blocking type interference model as a starting point, taking a plurality of jammers as an agent group, constructing a process of interfering multi-unmanned aerial vehicle communication into a partially observable Markov game process, and constructing a multi-agent reliability distribution model based on a reliability distribution criterion; and finally, an improved anti-fact multi-agent strategy gradient algorithm and a finite-length elite trajectory experience playback mechanism are utilized to optimize a reliability allocation strategy in the process that the multiple jammers confront the multiple unmanned aerial vehicles, reasonable and effective interference frequency domain resource scheduling is realized, the reliability allocation efficiency is effectively improved, and the robustness of the system is improved. The problems of low reliability distribution credibility and poor distribution efficiency of a conventional multi-jammer confrontation multi-unmanned aerial vehicle frequency hopping communication system are solved.
Owner:TIANFU JIANGXI LAB

Lightweight cloud edge collaborative optimization scheduling method and device for virtual power plant

The invention discloses a virtual power plant lightweight cloud edge collaborative optimization scheduling method and device, and relates to the field of virtual power plant optimization scheduling. The method comprises the following steps: dividing a virtual power plant cluster at a cloud according to the energy coupling strength of a virtual power plant; establishing a virtual power plant optimization scheduling model; modeling the virtual power plant optimization scheduling model as a Markov game model; multi-agent neural network parameters in the Markov game process are trained; configuring an edge computing unit for each cluster according to a cluster division result, and enabling each edge computing unit to receive optimal parameters of a Markov game model and a multi-agent neural network of the cloud; and each edge computing unit acquires state information of the respective cluster by using a self-adaptive data compression and acquisition method, and performs Markov game by using the received model and the optimal parameter to obtain a power scheduling result of each node in each cluster. According to the method, the virtual power plant scheduling decision accuracy under limited time and limited observation information is improved.
Owner:SHENYANG INST OF ENG

Clustered spacecraft cooperative game control method based on CTDE architecture

This invention discloses a collaborative game control method for swarm spacecraft based on a CTDE architecture, comprising the following steps: Step S1: Constructing a swarm-based relatively nonlinear dynamic model containing multiple tracking spacecraft; Step S2: Obtaining a configuration evaluation index for guiding a collaborative encirclement configuration; Step S3: Obtaining a Markov game environment framework; Step S4: Obtaining a dense reward signal for policy optimization within the Markov game environment framework; Step S5: Obtaining a convergent collaborative game control policy based on the CTDE architecture; Step S6: Training the collaborative game control policy using a phased learning mechanism. This invention achieves stable convergence in a multi-agent environment, provides dense reward signals to alleviate reward sparsity, forms a swarm encirclement configuration, smoothly integrates hard engineering constraints, and improves the robustness of the policy under high-dynamic adversarial conditions through phased learning.
Owner:SUN YAT SEN UNIVERSITY SHENZHEN +1

Active power distribution network medium-low voltage cooperative optimal approaching control method based on double-layer reinforcement learning

The invention relates to the technical field of power systems, and particularly discloses a medium-low voltage collaborative optimal control method for an active power distribution network based on double-layer reinforcement learning. Historical data of distributed photovoltaic output, key node voltage, power flow reverse transmission rate of a transformer in a transformer area, tap state of a medium-low voltage transformer, port power of a medium-voltage intelligent soft switch and network loss are collected through a measuring device; training a two-stage neural network based on historical data; taking the two-stage neural network as a reinforcement learning off-line training environment, and constructing a double-layer Markov game model; and deploying the trained agent parameters to medium and low voltage control terminals to realize medium and low voltage multi-target cooperative control containing multiple types of adjusting devices. The method has the advantages that the double-layer reinforcement learning algorithm integrated with the attention mechanism is proposed to solve the multi-time scale equipment regulation and control problem, the sum of lower-layer agent rewards serves as a part of an upper-layer agent, and active power distribution network medium-low voltage collaborative optimal control is achieved.
Owner:SICHUAN UNIV +1

Large-model-driven multi-subject trajectory generation and city-level behavior simulation method

PendingCN122045808ALearn and reproduce interactionsovercome limitationsBiological modelsComputation complexitySimulation
The invention discloses a large-model-driven multi-subject trajectory generation and city-level behavior simulation method. According to the method, a multi-agent generation adversarial imitation learning framework is adopted, an MAGAIL-VL model is constructed, and synchronous simulation of vehicle movement and road section traffic state evolution is realized. An urban traffic system is abstracted into a Markov game, two types of agents including a vehicle agent and a road section agent are included, and expert demonstration is carried out by utilizing a real track and speed data. The model generator generates an action according to a local state, and the discriminator provides a reward signal by distinguishing generated data and expert data to guide strategy network updating. According to the framework, centralized training and distributed execution are adopted, the calculation complexity is reduced, meanwhile, the interaction influence between main bodies can be accurately described, and vivid urban-level traffic tracks and simulation scenes are generated.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Voltage reactive power control method and device, control equipment and storage medium

The invention relates to a voltage reactive power control method and device, control equipment and a storage medium, and the method comprises the steps: obtaining the real-time state data of a power distribution network, the real-time state data comprises the reactive power output, node voltage and load information of a photovoltaic inverter in each sub-region divided in advance; inputting the real-time state data into a pre-trained multi-agent cooperative control model, and outputting a reactive power control instruction through the multi-agent cooperative control model, the reactive power control instruction being used for indicating the reactive power output of the photovoltaic inverter of each sub-region; wherein the multi-agent cooperative control model is obtained by performing set training on a Markov game model constructed by taking a sub-region as an agent, and the sub-region is obtained by dividing the power distribution network through node modularity determined by an electrical distance between nodes of the power distribution network. Therefore, the real-time performance and the control precision of voltage reactive power control can be improved.
Owner:ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID

A robust adversarial training framework for multi-agent reinforcement learning energy system

ActiveCN116306903BMachine learningNeural learning methodsStrategy trainingAlgorithm
The present application relates to a kind of robust confrontation training framework for multi-agent reinforcement learning energy system, comprising: constructing an adversarial agent to generate adversarial attack, and modeling as adversarial partially observable stochastic game system;Fixed pre-trained victim multi-agent strategy, train an optimal deterministic confrontation strategy to produce bounded disturbance;Fixed optimal confrontation attack strategy, improve the robustness of victim strategy under optimal attacker through adversarial training.The beneficial effects of the present application are: the present application models adversarial attack as an attack adversary based on single-agent reinforcement learning, and learns the strongest attack strategy considering attack constraints.Mathematically, the problem is constructed as a confrontation Markov game, and the performance of the integrated energy management system based on multi-agent reinforcement learning is improved through robust confrontation training.
Owner:ZHEJIANG ZHENENG YUEQING POWER GENERATION CO LTD +1

UAV target selection method and device based on Markov game and Bayesian optimization

This application relates to a method and device for drone target selection based on Markov games and Bayesian optimization. The method includes extracting drone target feature data, sparsifying the target feature data based on the distance between the drone and the target drone, generating a global representation, constructing a Markov game model, defining a reward function as the objective function for Bayesian optimization, probabilistically modeling the objective function using a Gaussian process as a proxy model, and iteratively solving the optimal reward function output value. Based on the optimal reward function output value, the Markov game model outputs the drone target selection result. This method can optimize drone selection in complex situations.
Owner:NAT UNIV OF DEFENSE TECH

A battery pack equalization control method and system

This invention relates to a battery pack balancing control method and system. The method includes: Step S1: treating the battery pack as an intelligent agent network and converting the battery pack balancing problem into a Markov game that can be handled by the DRL algorithm; Step S2: generating online expert data using the FCS-MPC algorithm; Step S3: training the intelligent agent network and a hybrid network based on different intelligent agent networks using the FCS-MPC-QMIX algorithm, and dynamically fusing the expert data into the objective functions of the intelligent agent network and the hybrid network during training to accelerate network convergence; Step S4: using the trained intelligent agent network and the hybrid network to achieve battery pack balancing control. This invention can provide balancing control for distributed battery energy storage systems.
Owner:JIANGNAN UNIV

Virtual power plant operation strategy optimization method and equipment considering supply and demand ends

The invention discloses a virtual power plant operation strategy optimization method and equipment considering supply and demand ends, and the method comprises the steps: obtaining the supply side, demand side and market environment data of a power market, and building a double-layer mixed game model integrating the supply and demand ends based on the data; the model comprises a lower-layer market clearing model used for market clearing and an upper-layer virtual power plant decision model used for making a power supply strategy by a virtual power plant. Meanwhile, a supply and demand prediction model is established according to market data to predict power supply output and power demands of users. And in combination with the predicted output and demand data and a double-layer game model, modeling a dynamic interaction relationship between the virtual power plant and other market subjects by using a multi-agent Markov game, and solving the Markov game through a multi-agent reinforcement learning method so as to obtain an optimized operation decision strategy of the virtual power plant. The method can improve the collaboration and dynamic adaptability of the operation strategy of the virtual power plant, and can be widely applied to the technical field of electric power.
Owner:CSG POWER GENERATION (GUANGDONG) ENERGY STORAGE TECH CO LTD

A voltage control method for distribution networks based on multi-agent deep reinforcement learning

The present invention provides a distribution network voltage control method based on multi-agent deep reinforcement learning, which relates to the technical field of distribution networks. The method includes dividing the distribution network into multiple areas and centrally managing and optimizing the voltage of each area; converting the collaborative voltage control problem between areas into a Markov game to achieve effective collaboration between agents; adopting a centralized training and decentralized execution framework so that the agents can fully utilize global information in the training phase and make independent decisions based on local observations in the execution phase; integrating a self-attention mechanism so that the agents are more focused on information related to their own rewards, thereby combining multi-agent reinforcement learning technology with the self-attention mechanism to effectively deal with the voltage fluctuation problem in the distribution network, thereby achieving stable control and optimized management of the voltage, improving the voltage regulation accuracy and response speed of the distribution network, and ensuring the stability and reliability of the power system.
Owner:NANJING HUIQIANG NEW ENERGY TECH CO LTD

Macro-micro collaborative automatic driving scene generation method facing real traffic flow

The invention provides a macro-micro collaborative automatic driving scene generation method facing real traffic flow. A high-authenticity traffic flow environment and a vehicle interaction scene are constructed through SUMO-CARLA bidirectional collaborative simulation to realize high-authenticity simulation test scene generation. Comprising the following steps: deploying an automatic driving to-be-tested vehicle and a non-player control vehicle in a CARLA high-fidelity simulation platform; in an SUMO macroscopic traffic simulation engine, background traffic flow generation driven by NGSIM real traffic data is imported; modeling the interaction process of a to-be-tested vehicle and an NPC vehicle according to the Markov game, designing a deep reinforcement learning strategy network model fused with multi-modal input, training by adopting a centralized training-distributed execution multi-agent architecture, and performing strategy optimization by utilizing a reinforcement learning optimization algorithm to obtain a convergence strategy; and deploying an optimized convergence strategy, and generating an automatic driving simulation test key scene based on evaluation indexes such as scene authenticity, effectiveness and efficiency.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

An unmanned aerial vehicle path planning method, device and storage medium

The application provides a UAV path planning method, device and storage medium, and relates to the technical field of wireless communication. The method comprises the following steps: determining the fair throughput of a wireless sensor network (WSN) based on a channel model of the WSN, wherein the fair throughput indicates the fairness of data transmission of sensor nodes occupying a wireless channel; establishing a Markov game model of the WSN based on the fair throughput and a UAV energy consumption model; iteratively training a multi-agent deep deterministic policy gradient (MADDPG) algorithm neural network based on the Markov game model, so that the fair throughput and the energy consumption of the UAV meet an optimization target; and planning a flight path of the UAV by using the trained MADDPG algorithm neural network. Based on this, the flight path of the UAV for executing a data collection task can be reasonably planned, the total energy consumption of the UAV is effectively reduced while the fairness among the sensor nodes is ensured, and the execution efficiency of the data collection task is greatly improved.
Owner:MINNAN INST OF SCI & TECH

A Method for Allocating Interference Resources in UAV Swarms Based on Pre-trained Attention Encoders

This application relates to a method for allocating interference resources for a drone swarm based on a pre-trained attention encoder. The method includes: setting an objective function for the interference resource allocation problem using an interference effect evaluation model and the interference power consumed by the drone swarm; modeling the interference resource allocation problem as a constrained combinatorial optimization problem by considering the limitations on drone communication, energy, and flight speed under denial conditions; modeling the interference decision-making process in the combinatorial optimization problem as a distributed locally observable Markov game model; mapping drones as agents; and solving the model using the MAPPO algorithm based on a pre-trained attention encoder to obtain the interference resource allocation scheme for the drone swarm. This method can solve the problems of state space dimensionality explosion and environmental non-stationarity in multi-agent collaborative decision-making processes.
Owner:NAT UNIV OF DEFENSE TECH

Unmanned aerial vehicle cooperative photovoltaic material transportation method and system based on self-evolution constraint

The invention discloses an unmanned aerial vehicle cooperative photovoltaic material transportation method and system based on self-evolution constraint, and belongs to the technical field of intelligent unmanned aerial vehicle control and complex environment autonomous transportation. The method comprises the following steps: collecting multi-source sensing data of an operation scene, and constructing a unified state space including terrain, wind field and feasible navigation area constraints based on the data; inputting the unified state space into a meta-policy network, and generating a high-level decision instruction containing a task intention; modeling an unmanned aerial vehicle group into a multi-agent Markov game framework based on the high-level decision instruction, performing distributed solution by adopting a graph reinforcement learning framework, and outputting a cooperative task and path intention sequence allocated to each unmanned aerial vehicle; and taking the path intention sequence as a reference trajectory, combining unmanned aerial vehicle dynamics and an energy constraint model, constructing a model prediction control problem, performing online solving, generating a safe actual control instruction, and driving an unmanned aerial vehicle group to execute a cooperative transportation task of photovoltaic materials.
Owner:西安沣东华能热力有限公司 +1

A multi-agent power distribution network dispatching method and device based on information dynamic aggregation

The application provides a multi-agent power distribution network scheduling method and device based on information dynamic aggregation, relates to the technical field of power system operation control, and divides the power distribution network into multiple control areas, establishes a scheduling optimization model, is further converted into a Markov game model, constructs a reinforcement learning agent for distributed optimization, constructs a distributed learning architecture, each agent independently executes decision-making and model updating based on local observation, avoids global information sharing, and effectively protects the data privacy of each area. Using the optimized reinforcement learning agent, the local observation variables at the current time are dynamically encrypted to generate each encrypted local information, realize information encryption of each agent in the information interaction process, and further reduce the risk of privacy leakage. Real-time scheduling is completed by dynamically generating aggregation weights, the collaborative efficiency and decision-making performance of the multi-agent are improved, and the problems of low collaborative efficiency caused by privacy leakage risk and static aggregation strategy are solved.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD JINHUA POWER SUPPLY CO +1

A raw water system scheduling and risk early warning method based on deep reinforcement learning

The present application relates to a kind of raw water system scheduling and risk early warning method based on deep reinforcement learning, comprising the following steps: establishing the Markov game model of waterworks water intake pump station total energy consumption minimization and minimum sediment concentration problem, and the corresponding state space, action space and reward function are designed;Obtain historical operation data and hydrological data, based on the preset hydraulic model and MLP neural network, construct scheduling environment model;Through scheduling environment model and agent PPO algorithm, the deep reinforcement learning agent is trained, and based on domain randomization improves generalization performance, obtains agent strategy;Build seq2seq model, based on historical operation data, respectively on river sediment content and reservoir water supply is multi-step forecast, combined with the agent strategy obtained by training, generate risk early warning model.Compared with prior art, the present application has strong generalization performance, has the advantages such as risk early warning function.
Owner:TONGJI UNIV +1