Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

101 results about "Partially observable Markov decision process" patented technology

A partially observable Markov decision process (POMDP) is a generalization of a Markov decision process (MDP). A POMDP models an agent decision process in which it is assumed that the system dynamics are determined by an MDP, but the agent cannot directly observe the underlying state. Instead, it must maintain a probability distribution over the set of possible states, based on a set of observations and observation probabilities, and the underlying MDP.

Multi-agent cluster cooperative attack and defense game method based on near-end strategy optimization

The invention discloses a multi-agent cluster cooperative attack and defense game method based on near-end strategy optimization, and relates to the technical field of automatic control, and the method comprises the steps: constructing a partial considerable Markov decision process, and generating a multi-agent decision model; establishing a cooperative attack and defense game simulation framework including an attacker agent, a defender agent, a target area and irregular obstacles in the preset task area; parallel trajectory sampling is executed, an observation-action-reward sequence of each agent is collected, and a training data set is generated; a centralized state is constructed by cascading and splicing local observation of each agent, a time sequence difference error and a dominant function are calculated, network parameter updating is carried out by adopting cutting strategy gradient and value function regression, and a trained cooperative attack and defense game strategy network is generated when a reward curve converges. And the multi-agent cluster realizes intelligent cooperative attack and defense game and autonomous obstacle avoidance in a task environment with irregular obstacles.
Owner:BEIHANG UNIV

Unmanned aerial vehicle dynamic task allocation and path planning method and device for cluster cooperation

The invention relates to a cluster cooperation-oriented unmanned aerial vehicle dynamic task allocation and path planning method and device, and the method comprises the steps: modeling a monitoring region comprising an unmanned aerial vehicle cluster, a plurality of to-be-tracked targets and a specified moving target as a partially observable Markov decision process, and enabling each unmanned aerial vehicle to obtain local observation information based on relative parameters; then, each unmanned aerial vehicle outputs a flight orientation adjustment instruction through a strategy neural network, and the network is trained and optimized under the guidance of a value network through a multi-agent near-end strategy optimization algorithm; the unmanned aerial vehicle cluster executes a joint action and generates a comprehensive reward signal fusing global task rewards, individual behavior rewards and anti-collision punishment; and finally, driving the strategy network to synchronously complete implicit task allocation and collaborative path planning through continuous feedback of the reward signal. According to the invention, integrated cooperative solution of task allocation and path planning is realized, so that the unmanned aerial vehicle cluster has adaptive cooperative capability and high robustness under a distributed decision framework.
Owner:TAODIAN CHAIN (GUANGZHOU) INFORMATION TECH CO LTD

Vulnerability intelligent collaborative detection method and system based on multi-agent reinforcement learning

The invention discloses a vulnerability intelligent collaborative detection method and system based on multi-agent reinforcement learning. The method comprises the following steps: independently executing detection according to local observation information and a strategy generation model through an intelligent agent deployed at an edge node; the strategy generation model is constructed based on a decentralized partial observable Markov decision process, and is obtained through cloud centralized training and global shared reward function optimization. When the agent obtains a forward reward, packaging a corresponding vulnerability context and related strategy parameters into an experience package; and the experience packet is diffused in the cooperative detection network in a mode of randomly selecting neighbor nodes and forwarding step by step. And after each agent receives the data, analyzing the content of the experience package and comparing the content with the local environment so as to dynamically adjust local strategy parameters and realize collaborative optimization. According to the vulnerability intelligent collaborative detection method and system based on multi-agent reinforcement learning provided by the invention, the problems that network security vulnerability detection is poor in expandability and lacks knowledge generalization ability are solved.
Owner:SHIZUISHAN POWER SUPPLY COMPANY OF STATE GRID NINGXIA ELECTRIC POWER +1

Distributed energy storage aggregation scheduling method and system considering user fatigue effect

The invention discloses a distributed energy storage aggregation scheduling method and system considering a user fatigue effect, and the method comprises the steps: building a user response behavior model considering the long-term response preference and short-term fatigue of a user based on a partial observable Markov decision process; establishing a user response state online inference strategy based on maximum likelihood estimation, and inferring an unobservable state of a user; and based on the user response behavior model and the user response state online inference strategy, establishing an online learning aggregation scheduling method combined with the running state constraint of the distributed energy storage unit, and realizing distributed energy storage aggregation scheduling. According to the method, the user response uncertainty, the fatigue level and the actual response capacity of the energy storage equipment are comprehensively considered, a reliable model basis is provided for aggregation scheduling, effective tracking of the target power is achieved, and therefore the accuracy and reliability of aggregation scheduling are improved.
Owner:NANJING NORMAL UNIVERSITY

Communication lightweight multi-micro energy network cooperative control method for large model assisted reward recovery

The invention belongs to the technical field of multi-micro-energy-network cooperative control, and particularly relates to a communication lightweight multi-micro-energy-network cooperative control method for reward recovery assisted by a large model, and the method comprises the following steps: S1, constructing a mathematical model of distributed energy in a micro-energy network, and enabling the mathematical model to be used for describing the dynamic behaviors of energy conversion, storage and constraint; s2, modeling a micro energy network cooperative control problem as a partially observable Markov decision process POMDP; s3, solving the POMDP of the step S2 by using an optimized MAPPO method to obtain a cooperative control scheme; and issuing to each micro-energy network terminal for execution. The method can break through the limitation of the communication bottleneck on the premise of ensuring the learning quality of the control strategy.
Owner:CHONGQING UNIV

Micro-grid distributed robust optimization scheduling method considering source charge uncertainty and electricity-hydrogen-carbon coupling

The invention belongs to the technical field of power system operation and control, and particularly provides a microgrid distributed robust optimization scheduling method considering source load uncertainty and electricity-hydrogen-carbon coupling. Comprising the following steps: constructing an electricity-hydrogen-carbon-heat four-dimensional collaborative micro-grid energy flow framework comprising a multi-port alternating current and direct current bus, a hydrogen production storage and transportation subsystem, a carbon capture methanation subsystem and a cross-stage thermodynamic energy cascade network; establishing an IGDT-based source-load bilateral uncertainty dynamic pricing model, and selecting a risk avoidance or opportunity seeking pricing model as required; a multi-agent game double-layer interaction model with a micro-grid operator as a leader and a load aggregator and a charging station as followers is constructed, and a nested part can observe a Markov decision process; and a privacy protection multi-agent deep reinforcement learning algorithm in which a maximum entropy framework and a CTDE mechanism are introduced is adopted for solving. According to the method, the problem of low-carbon operation of the micro-grid in a source load uncertainty environment is solved, and collaborative improvement of economical efficiency, low-carbon performance and robustness is realized.
Owner:CHINA THREE GORGES UNIV

MARL-based flexible distributed energy network multi-agent collaborative optimization method

The invention provides a flexible distributed energy network multi-agent collaborative optimization method based on MARL, and the method comprises the steps: building a distributed energy network multi-agent interaction framework which comprises a local energy transaction model for distinguishing an intelligent soft switching path, a network cost model based on a power transmission distribution factor, and a power distribution network operator and user income model; a physical system operation constraint model is established and covers intelligent soft switch operation characteristics, network topology radial constraint, energy storage charging and discharging dynamic states, controllable distributed power supply and photovoltaic output limitation, translational load scheduling and DistFlow power flow constraint. And finally, constructing a collaborative optimization mechanism based on a multi-agent near-end strategy optimization algorithm, modeling the system into a partially observable Markov decision process, adopting a centralized training distributed execution framework, and guaranteeing transaction security and data privacy through block chain evidence storage, RSA encryption and an intelligent contract. The system economy and the operation flexibility are improved.
Owner:INST OF ELECTRICAL ENG CHINESE ACAD OF SCI

Adaptive closed-loop optimization management method for substation maintenance plan based on multiple elements

The invention relates to the technical field of intelligent power grids, and discloses a power transformation maintenance plan adaptive closed-loop optimization management method based on multiple elements, which comprises the following steps: constructing a multi-element database of power transformation equipment; based on the multi-element database, the maintenance priority of each device is calculated through a weight quantification model, and a maintenance priority sequence is generated; a mixed parthenogenesis optimization algorithm is adopted, the maintenance priority sequence is used as input, and an optimized maintenance plan is generated with the purpose of minimizing the comprehensive cost; and on the basis of an improved partially observable Markov decision process framework, modeling is performed on the execution process of the optimized maintenance plan, and a reinforcement learning mechanism is utilized to dynamically adjust and optimize a subsequent maintenance plan according to a feedback reward after the maintenance plan is executed, so that closed-loop iteration is formed. According to the method, adaptive dynamic optimization and sustainable evolution of the substation maintenance plan are realized, and the intelligent level, the safety reliability and the resource utilization efficiency of power grid operation and maintenance are remarkably improved.
Owner:HUAIAN OF JIANGSU ELECTRIC POWER CO POWER SUPPLY

Unmanned aerial vehicle cluster dynamic area coverage method based on deep reinforcement learning

The invention discloses an unmanned aerial vehicle cluster dynamic area coverage method based on deep reinforcement learning. The method comprises the following steps: constructing a multi-physical field coupling simulation environment comprising node motion, a directional antenna, a space channel and a radar detection model; based on the environment, converting a dynamic coverage problem form into a partially observable Markov decision process, a defined state, observation, action and reward mechanism, performing multi-agent training by adopting a centralized training decentralized execution framework, evaluating a joint action by using a global Critic network during training, and storing interactive experience by using an experience playback buffer area; based on a priority experience playback mechanism, according to a time sequence difference error and a reward value, an experience sampling probability is calculated, samples are extracted, Critic and Actor networks are updated in sequence, a decision model is deployed to an unmanned aerial vehicle cluster after training is completed, and each unmanned aerial vehicle executes an area coverage task in a distributed manner, so that efficient and low-risk communication coverage can be realized in a complex threat environment. The communication quality is improved; and the detected probability is reduced.
Owner:10TH RES INST OF CETC

Policy learning method with privacy protection in mobile edge computing for intelligent agent

A policy learning method with privacy protection in mobile edge computing for an intelligent agent is provided, relating to the technical field of mobile communication. The method includes: establishing an edge-collaborative computing offloading model, where the edge-collaborative computing offloading model includes a service caching model, a task offloading model, and a system cost model; establishing an optimization problem for task offloading, service caching, computing resource allocation and transmission power control based on the edge-collaborative computing offloading model for minimizing task processing costs; abstracting the optimization problem to a partially observable Markov decision process; and autonomously learning a task offloading strategy, a service caching strategy, a computing resource allocation strategy, and a transmission power control strategy by using a federated learning-based multi-agent deep reinforcement learning algorithm based on the Markov decision process.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Frequency modulation demand quantification method based on attention-guided multi-agent reinforcement learning

The invention relates to the field of power system operation control, and discloses a frequency modulation demand quantification method based on attention guidance multi-agent reinforcement learning, which comprises the following steps: constructing a main network frequency response model; based on a main network frequency response model, modeling a frequency modulation demand online quantification task as a distributed partial observable Markov decision process; embedding a historical observation-action sequence of each agent through a track encoder, obtaining role representations of a plurality of agents through comparative learning of a role encoder and InfoNCE, and constructing a multi-head attention mechanism by taking global state embedding as query and the role representations as key values; the action of each agent is output at the actor end by adopting a strategy network sharing parameters; and according to the real-time operation state, calculating the frequency modulation demand meeting the RoCoF and frequency lowest point constraint and an allocation scheme of the frequency modulation demand among various frequency modulation resources. The method has the advantage that rapid and self-adaptive online quantification of the main network frequency modulation requirement is realized.
Owner:SICHUAN UNIV

Mine multi-robot adaptive cooperative control method based on micro-kernel operating system

The application discloses a mine multi-robot adaptive cooperative control method based on a microkernel operating system and belongs to the technical field of mine automation and robot cooperative control. The method cooperatively realizes the following processes by running multiple independent service processes on the microkernel operating system: a multi-modal perception fusion process dynamically adjusts sensor weights according to environmental dust and light, and constructs a global semantic map; a dynamic task allocation process models task allocation as a partially observable Markov decision process for solution by using a multi-agent deep reinforcement learning embedded with an attention mechanism; a hybrid path planning process constructs a hierarchical cost model based on historical traffic data, and generates a path through a three-layer planning system; a dynamic conflict resolution process uses an STCN to predict a trajectory, and realizes forward-looking conflict resolution in combination with a distributed negotiation auction mechanism; and a state synchronization process uses a lightweight practical Byzantine fault tolerance consensus to ensure that the states of the whole system are consistent. The application significantly improves the real-time performance, reliability and intelligence level of the system.
Owner:CITIC HEAVY INDUSTRIES CO LTD

A self-reconfiguration method for free-form modular robots

ActiveCN116300423BArtificial lifeAdaptive controlHardware prototypingControl engineering
This invention discloses a self-reconfiguration method for free-form modular robots. First, an abstract model of the basic elements of a class of free-form modular robots is constructed, where each robot module can be connected to each other from any direction via active and passive connectors. Then, based on a hardware prototype, a reasonable action space and observation space are designed for this type of resource-constrained modular robot. The action space includes the drive parameters of two planar continuous motion motors and the motion parameters of one connector, while the observation space includes the motion information and local connectivity relationships of the robot itself and its neighboring modules. The reconfiguration problem of the free-form modular robot is modeled as a distributed partially observable Markov decision process. Based on a proximal policy optimization algorithm, an efficient and conflict-free multi-agent reinforcement learning method for reconfiguration in the continuous action space can be achieved. Finally, the timely termination of the reconfiguration process is achieved under the conditions of local observability and limited communication.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Cluster power grid dispatching method and system

The invention discloses a sub-cluster power grid dispatching method and system, and belongs to the technical field of power grid dispatching, and the method comprises the steps: obtaining a local target of a cluster based on cluster benefits, and carrying out the local training of an intelligent agent corresponding to the cluster according to the local target and the historical data of the cluster, and obtaining a process intelligent agent; acquiring a global target based on the comprehensive benefit of the power grid, and acquiring the final contribution of the local target to the global target by using a partially observable Markov decision process; obtaining a global reward function according to the global target, and carrying out global training on the process agent according to the final contribution and the global reward function to obtain an optimized agent; and constructing a global scheduling strategy updating rule through a PPO algorithm, obtaining a global scheduling strategy through an optimization agent under the constraint of the global scheduling strategy updating rule, and scheduling the power grid through the global scheduling strategy. The technical problem that local and global scheduling requirements are difficult to consider in the prior art is solved.
Owner:ZHEJIANG SIJI TECH SERVICE CO LTD

Multi-agent scheduling method, device and equipment of industrial network and storage medium

The application relates to the technical field of computer software, and discloses a multi-agent scheduling method, device and equipment of an industrial network and a storage medium, the method comprising the following steps: dividing a single synchronization domain in the industrial network into multiple sub-synchronization domains, and constructing a network model by using the multiple sub-synchronization domains, wherein each sub-synchronization domain is controlled by a corresponding single-domain controller; constructing a constraint condition set with the aim of reducing resource occupancy and scheduling failure rate; constructing a partially observable Markov decision process according to the constraint condition set and the network model, and the partially observable Markov decision process is used for formulating an optimal strategy for cross-domain collaborative scheduling of a chain service flow. The multi-agent scheduling method of the industrial network provided by the application can effectively adapt to a complex industrial scene and meet the determined transmission requirements of multi-form service flows.
Owner:BEIJING JIAOTONG UNIV

A preemptive scheduling method for heterogeneous multi-core processors based on multi-agent reinforcement learning

PendingCN122086566AStrong heterogeneous perceptiondifferentiate decision benefitsProgram initiation/switchingBiological modelsScheduling (computing)Distributed computing
This invention relates to a preemptive scheduling method for heterogeneous multi-core processors based on multi-agent reinforcement learning. It includes: constructing a system model comprising a set of sporadic real-time tasks and a heterogeneous multi-core platform; pre-setting fixed preemption points for tasks; dividing task execution into non-preemptive segments; modeling the scheduling problem as a partially observable Markov decision process; designing state representations that include task dynamic attributes and platform heterogeneity characteristics; wherein the task dynamic attributes are normalized values ​​of relative deadline pressure, execution progress ratio, and relaxation; and the platform heterogeneity characteristics include a weighted sum of core computing power level and global idle computing power; clustering processors according to computing power; assigning independent agents to each cluster; and adopting a fully cooperative mode to share global rewards; employing a centralized training and distributed execution framework; using the discrete soft actor-critic algorithm during training; with each agent's policy network collaboratively optimizing under the guidance of a centralized value function; and during execution, each agent independently deciding on preemption timing based on local observations. This invention achieves cooperative preemptive scheduling among heterogeneous cores through multi-agent reinforcement learning, effectively reducing task overdue rates and system overhead, and improving real-time performance.
Owner:SHANXI UNIV

Distributed decision-making method based on multi-agent deep reinforcement learning

The application discloses a kind of distributed decision-making methods based on multi-agent deep reinforcement learning, constructs the integrated scheduling problem model of satellite observation and data download based on decentralized partially observable Markov decision process, each satellite is regarded as intelligent agent with autonomous decision-making ability, each intelligent agent can make autonomous decision on observation task, determine the observation time of ground target and the download time of observation data, so that the total income in scheduling period is maximum, in online phase, the decision result of satellite cluster is obtained by satellite cluster scheduling network trained in real time according to observation task data, whether each satellite distributed decision-making each observation task is executed, execution time, the download time of observation data and download ground station are realized.The application can dynamically carry out satellite observation data download while carrying out satellite observation task planning, and significantly improve the observation efficiency of satellite cluster.
Owner:SHANGHAI JIAOTONG UNIV

Intelligent driving behavior decision-making method and device fusing complex network theory and partially observable markov decision process

The application discloses a complex network theory and a part observable Markov decision process intelligent driving behavior decision method and equipment, by introducing a part observable Markov decision process (POMDP), an ideal decision model for multi-agent interaction in an uncertain environment is provided, with the increase of model dimension, the calculation complexity is also greatly increased, in order to solve the calculation complexity problem, the application introduces a complex network cognitive theory based on driving environment modeling, is used to evaluate important nodes and extract risk trees, to reduce the confidence space of the POMDP model, accelerate the model solution, so as to realize the real-time behavior decision in the highly interactive dynamic uncertain driving environment, compared with the behavior decision method using the finite state machine and RL, the application can be effectively operated in the highly interactive dynamic uncertain driving environment while ensuring the algorithm interpretability, and provides a new solution for the development of high-level intelligent driving system.
Owner:JIANGSU UNIV

A Multi-Agent, Multi-System Cooperative Detection Method Based on Deep Reinforcement Learning for Radar Jamming Countermeasures

This invention discloses a multi-agent, multi-system cooperative detection method based on deep reinforcement learning for radar jamming countermeasures: Step 1, constructing a dynamic adversarial game model of aircraft swarms and jamming aircraft swarms in a jamming scenario; Step 2, treating the jamming aircraft swarm as the environment, and modeling the interaction process between multiple aircraft and the environment as a distributed partially observable Markov decision process; Step 3, using the HyPPO-MACAD multi-agent cooperative anti-jamming detection algorithm based on hybrid attention proximal policy optimization for training and decision-making to obtain a cooperative policy; the HyPPO-MACAD algorithm introduces a continuous-discrete hybrid branch in the policy network and an attention mechanism in the commentator network; Step 4, based on the trained cooperative policy, controlling the aircraft swarm to perform cooperative detection tasks. This invention can adapt to coupled decision-making tasks, perform intelligent anti-jamming and cooperative detection, reduce detection adhesion under irregular detection models, and improve anti-jamming capability and regional coverage efficiency.
Owner:TONGJI UNIV

A Location Privacy Protection Method for 3D Localization Errors Based on Recurrent Reinforcement Learning

This invention discloses a location privacy protection method based on recurrent reinforcement learning for 3D localization errors, belonging to the fields of location services and information security. The method uses the precise coordinates of landmarks around the user as anchor points; the user generates a perturbed location for the current time slot based on their own positioning error and uploads it to the server, obtaining an assessment of privacy protection level and service quality; the perturbed location of the current time slot with positioning error is modeled as a partially observable Markov decision process, employing a dynamic location privacy protection mechanism combining a long short-term memory network and a dual-delay deep deterministic policy gradient algorithm to protect user location privacy even with positioning errors. The method is simple, effective, and achieves an effective balance between service quality and privacy protection needs, maximizing user utility.
Owner:CHINA UNIV OF MINING & TECH

Unmanned aerial vehicle auxiliary resource allocation method

The invention discloses an unmanned aerial vehicle auxiliary resource allocation method, which comprises the following steps that: an unmanned aerial vehicle is used as a data receiving end, receives pilot signals of a part of nodes at the beginning of each time slot, and performs channel estimation according to the pilot signals; based on a channel estimation result, the unmanned aerial vehicle selects a plurality of nodes to perform cooperative data uploading, allocates beam forming weights for the nodes, and moves a horizontal distance in the air at the same time; afterwards, the selected nodes carry out cooperative beam forming data uploading according to the weight distributed by the unmanned aerial vehicle; in order to better describe the dynamic change of a channel and make a decision under the condition that the CSI is partially visible, the problem is modeled into a partially observable Markov decision process, the unmanned aerial vehicle position, the channel state and the remaining data volume are described as the system state, the node selection, the beam forming weight and the unmanned aerial vehicle velocity vector are described as the actions, and the unmanned aerial vehicle channel state estimation method is provided. The unmanned aerial vehicle position, the pilot signal and the instantaneous throughput are described as observations.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Distributed database parameter adaptive tuning method based on multi-agent deep reinforcement learning

The invention relates to a distributed database parameter adaptive tuning method based on multi-agent deep reinforcement learning, and belongs to the technical field of information. According to the method, the thought of a multi-agent deep reinforcement learning algorithm is introduced into the field of database parameter tuning, the cooperation and competition relation between nodes in a distributed database environment is explored, and a deep deterministic strategy gradient reinforcement learning algorithm C-MADDPG based on a centralized strategy gradient estimator is provided. A distributed database parameter tuning problem is modeled based on a partially observable Markov decision process. Problem definition takes expansion of a parameter search space as a cost, takes a complex competition effect between database nodes as a black box, carries out mathematical modeling on an observable cooperation effect, improves the universality and scalability of an algorithm, solves a relative generalization problem of the algorithm, and can be expanded to a general distributed database architecture. The C-MADDPG provided by the invention shows good performance in the parameter tuning of the distributed database.
Owner:FUJIAN NORMAL UNIV

Method and device for multi-agent power grid voltage optimization control based on large language model, electronic equipment and readable storage medium

The application belongs to the technical field of power grid optimization control and artificial intelligence, and particularly relates to a multi-agent power grid voltage optimization control method and device based on a large language model, an electronic device and a readable storage medium, which comprises the following steps: using prompt word engineering technology, inputting power grid environment information, task target optimization model and device characteristics as context information into a large language model to generate time sequence running data sets covering different working conditions; constructing a multi-agent system based on a distributed partially observable Markov decision process mechanism and training the same by using a TD3 (double-delay deep deterministic policy gradient) algorithm; and deploying the multi-agent system in a power grid control system, and each regional agent executes voltage regulation actions according to real-time collected local state information. The application can realize safe stability and network loss minimization of power grid voltage without relying on accurate system mechanism models / equations, and has good scalability and engineering application prospects.
Owner:ZHIBO ENERGY TECHNOLOGY (JIANGSU) CO LTD +2

Inter-cloud task migration decision-making method based on deep reinforcement learning

The invention relates to the technical field of cloud edge computing, and particularly provides a cloud task migration decision-making method based on deep reinforcement learning, which comprises the following steps of: constructing a target function by taking maximization of a cloud task success rate as a target; a deep reinforcement learning framework is constructed, and an LSTM network model is adopted to observe and predict the current state of the system based on the system history; describing a decision process under the condition that the system environment is uncertain based on a partially observable Markov decision process; and taking the current state of the system as the input of the improved deep reinforcement learning model, and obtaining a migration strategy of the cloud task based on the improved deep learning model. According to the method, a proper service migration mechanism is designed aiming at the cloud edge calculation problem in a mobile state, so that the low delay of the service is ensured, and modeling is carried out on the delay and reliability of the service migration process; and a DCPPOSM algorithm based on deep reinforcement learning is designed for effective migration decision making with an optimization objective of maximizing a task execution success rate.
Owner:NORTHEASTERN UNIV CHINA

CPS security evaluation scheduling method based on event-triggered reinforcement learning under time-varying channel

PendingCN122372442AEvent triggerPhysical system
A CPS security assessment and scheduling method based on event-triggered reinforcement learning under time-varying channels is proposed. The method comprises the following steps: (1) establishing a cyber-physical system model and assessment benchmark environment based on time-varying channels; (2) constructing a multi-source assessment utility index and an event-triggered mechanism with hysteresis and refractory period; (3) establishing a partially observable Markov decision process model from the perspective of the assessment end; (4) learning and training assessment and scheduling strategies based on the near-end policy optimization algorithm; and (5) executing event-triggered adaptive security assessment decisions and closed-loop state updates. This invention aims to provide a CPS security assessment and scheduling method based on event-triggered reinforcement learning under time-varying channels that can realize on-demand triggering, adaptive scheduling, and closed-loop assessment of abnormal communication intervention behaviors, and solve the problems of insufficient adaptability to time-varying channels, insufficient characterization of scheduling utility, and weak multi-mode collaborative decision-making capabilities in existing technologies.
Owner:JIANGSU OCEAN UNIV

An unmanned aerial vehicle obstacle avoidance method based on multi-agent graph reinforcement learning

The application discloses a kind of unmanned aerial vehicle obstacle avoidance methods based on multi-agent graph reinforcement learning, comprising: constructing the heterogeneous graph capable of reflecting the association between UAV and network entity in real time, so that the system can accurately capture the dynamic change of network topology and environmental state, provide more comprehensive state information for subsequent decision-making;Select the system communication rate as the main performance index, and model the communication rate and maximum problem as a partially observable Markov decision process;Construct a heterogeneous graph that can depict the interaction of multiple agents, and use a graph attention network to realize state coding;Propose a collaborative training mechanism based on the multi-agent soft actor-critic framework, which realizes the joint optimization of Actor network and Critic network.The method of the application combines multi-agent reinforcement learning and heterogeneous graph neural network technology, which can maximize the system communication rate and improve the obstacle avoidance and communication decision-making ability of UAV in high dynamic environment.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

A task scheduling method and device for sea area perception task

The application provides a task scheduling method and device for a sea area perception task. The task scheduling method provided by the application comprises the following steps: for a pre-built sea area perception system, a corresponding task scheduling model, a communication transmission model and a calculation processing model are constructed; a task scheduling problem of the sea area perception system is constructed according to the models; the task scheduling problem takes the calculation unloading, the edge node selection, the relay node selection, the container resource level and the container survival time level decision as variables, takes the maximization of the task completion number as an optimization target, and is subjected to the decision variable, the task time delay, the signal-to-noise ratio and the transmission resource constraint; the task scheduling problem is modeled as a partially observable Markov decision process, the preset deep reinforcement learning intelligent agent is trained by using offline simulation data, and a trained strategy model is obtained; and the strategy model is deployed on a perception device, so that the perception device calls the strategy model for distributed online inference according to a local observation state, and obtains a task scheduling scheme.
Owner:DONGHAI LAB

Automatic driving behavior planning method based on spatiotemporal characteristic perception and related device

The application discloses a method for automatic driving behavior planning based on space-time characteristic perception and related devices, which comprises the following steps: determining a state space, an action space and a reward function of an automatic driving model according to an automatic driving simulator, constructing a partially observable Markov decision process (POMDP), establishing a deep neural network model for space-time characteristic perception, configuring training parameters for the deep neural network model for space-time characteristic perception, executing the automatic driving simulator and interacting with a deep reinforcement learning agent, and establishing a time-sequenced vehicle running state cache and extracting training data for the automatic driving simulation process. The application can mine potential space-time correlation information between vehicles from limited perception information, limit the calculation and analysis of invalid information, and thus realize more accurate and efficient control decisions generated by the system.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Indoor substation AP deployment method, system and device based on MAPPO and medium

The invention relates to the technical field of electric power communication and intelligent deployment, in particular to an indoor substation AP deployment method, system and device based on MAPPO and a medium, and the method comprises the steps: firstly building an indoor substation scene model, and defining a data transmission process; constructing a downlink channel model fusing line-of-sight / non-line-of-sight path loss and channel overlapping interference, establishing a resource allocation model with the maximum system total rate based on the Shannon-Hartley theorem, and determining AP power, channel and bandwidth allocation rules; and then, converting the problem into a decentralized partial observable Markov decision process, and outputting an optimized deployment scheme of AP position, power, channel and bandwidth allocation through collaborative optimization of a policy network and a value network by using an MAPPO algorithm centralized training-distributed execution architecture. The method is suitable for solving the problem of collaborative optimization of AP deployment and resource allocation in a complex electromagnetic environment of an indoor substation, and the transmission performance and resource utilization efficiency of an electric power communication network are improved.
Owner:FUXIN POWER SUPPLY COMPANY STATE GRID LIAONING ELECTRIC POWER +1

Distributed priority perception collaborative unloading algorithm based on multi-agent deep reinforcement learning

The invention discloses a distributed priority perception collaborative unloading algorithm based on multi-agent deep reinforcement learning. The algorithm comprises a vehicle-mounted edge computing (VEC) architecture supporting collaboration between edge servers; a distributed partially observable Markov decision process (De-POMDP) framework; the invention discloses a distributed priority awareness cooperative unloading (DPACO) algorithm based on multi-agent deep reinforcement learning (MADRL). The VEC architecture comprises a vehicle layer, an edge layer and a cloud layer; the De-POMDP framework comprises observation and state of a vehicle, vehicle behaviors and a reward function; the DPACO algorithm adopts a distributed SAC framework, and efficient task processing is realized through a shared VEC server pool and a queue perception priority mechanism. The algorithm overcomes the limitation of unloading of a single server, maximizes the utility of resources among a plurality of servers, realizes edge-to-edge collaborative operation, ensures low-delay service by a specially designed task priority mechanism, and particularly has a remarkable guarantee effect on key tasks.
Owner:HUBEI UNIV OF AUTOMOTIVE TECH