Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

167 results about "Partially observable Markov decision process" patented technology

A partially observable Markov decision process (POMDP) is a generalization of a Markov decision process (MDP). A POMDP models an agent decision process in which it is assumed that the system dynamics are determined by an MDP, but the agent cannot directly observe the underlying state. Instead, it must maintain a probability distribution over the set of possible states, based on a set of observations and observation probabilities, and the underlying MDP.

Unmanned aerial vehicle trajectory planning method based on noise dual-depth Q network

The invention relates to the field of intelligent unmanned aerial vehicle autonomous control and decision making, in particular to an unmanned aerial vehicle trajectory planning method based on a noise dual-depth Q network, which comprises the following steps: acquiring unmanned aerial vehicle sensor data; establishing a forest fire model according to unmanned aerial vehicle sensor data; establishing an unmanned aerial vehicle constraint condition and an optimization function according to the forest fire model, and converting the unmanned aerial vehicle constraint condition and the optimization function into a partially observable Markov decision process; solving the partial observable Markov decision process by adopting a noise double-depth Q network to obtain an optimal unmanned aerial vehicle path planning strategy; according to the invention, the node areas are taken as decision objects, and value grading is carried out on different node areas, so that the decision is more in line with the reality; a noise bias item is introduced into a neural network weight to balance learning stability and exploration performance, so that the network can dynamically adapt to exploration behaviors in different environments; and a trajectory smoothness constraint is designed to ensure that the flight path is continuous, and the success rate of unmanned aerial vehicle trajectory planning is further improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Mobile charging station load situation awareness and intelligent scheduling method for electric vehicles

The invention relates to an electric vehicle-oriented mobile charging station load situation awareness and intelligent scheduling method, which comprises the following steps of: firstly, training historical load data of a fixed charging station, and acquiring a predicted value of a future load and uncertainty probability distribution of the predicted value by using a time sequence prediction model; and then, taking the prediction result as part of observation information, constructing a scheduling model based on a POMDP (Partial Observable Markov Decision Process), solving an optimal strategy of the POMDP, generating a scheduling decision, and sending a scheduling task to the mobile charging station to realize efficient distribution of charging resources.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-agent cluster cooperative attack and defense game method based on near-end strategy optimization

The invention discloses a multi-agent cluster cooperative attack and defense game method based on near-end strategy optimization, and relates to the technical field of automatic control, and the method comprises the steps: constructing a partial considerable Markov decision process, and generating a multi-agent decision model; establishing a cooperative attack and defense game simulation framework including an attacker agent, a defender agent, a target area and irregular obstacles in the preset task area; parallel trajectory sampling is executed, an observation-action-reward sequence of each agent is collected, and a training data set is generated; a centralized state is constructed by cascading and splicing local observation of each agent, a time sequence difference error and a dominant function are calculated, network parameter updating is carried out by adopting cutting strategy gradient and value function regression, and a trained cooperative attack and defense game strategy network is generated when a reward curve converges. And the multi-agent cluster realizes intelligent cooperative attack and defense game and autonomous obstacle avoidance in a task environment with irregular obstacles.
Owner:BEIHANG UNIV

Unmanned aerial vehicle dynamic task allocation and path planning method and device for cluster cooperation

The invention relates to a cluster cooperation-oriented unmanned aerial vehicle dynamic task allocation and path planning method and device, and the method comprises the steps: modeling a monitoring region comprising an unmanned aerial vehicle cluster, a plurality of to-be-tracked targets and a specified moving target as a partially observable Markov decision process, and enabling each unmanned aerial vehicle to obtain local observation information based on relative parameters; then, each unmanned aerial vehicle outputs a flight orientation adjustment instruction through a strategy neural network, and the network is trained and optimized under the guidance of a value network through a multi-agent near-end strategy optimization algorithm; the unmanned aerial vehicle cluster executes a joint action and generates a comprehensive reward signal fusing global task rewards, individual behavior rewards and anti-collision punishment; and finally, driving the strategy network to synchronously complete implicit task allocation and collaborative path planning through continuous feedback of the reward signal. According to the invention, integrated cooperative solution of task allocation and path planning is realized, so that the unmanned aerial vehicle cluster has adaptive cooperative capability and high robustness under a distributed decision framework.
Owner:TAODIAN CHAIN (GUANGZHOU) INFORMATION TECH CO LTD

Intelligent wireless control method and system for cloud robot

The invention discloses an intelligent wireless control method and system for a cloud robot. The method comprises the following steps: constructing a reconfigurable intelligent metasurface assisted cloud robot wireless control system; based on each preset constraint condition, constructing an optimization model comprehensively considering control errors and communication jitter by taking maximization of long-term stability as a target; reconstructing the optimization model by adopting a preset part observable Markov decision process model to obtain a target optimization model; based on observation state information, collected in real time, of each cloud robot in the cloud robot wireless control system, a multi-agent migration reinforcement learning model based on near-end strategy optimization is adopted to optimize the target optimization model, and target parameter values corresponding to the to-be-optimized parameters are obtained, and performing intelligent control on each cloud robot in the cloud robot wireless control system based on the target parameter value. According to the method, the strategy migration efficiency and the system stability of wireless control of the cloud robot can be improved.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Agricultural unmanned aerial vehicle heterogeneous cluster cooperative task allocation and path planning method

The invention relates to a heterogeneous cluster cooperative task allocation and path planning method for an agricultural unmanned aerial vehicle, and the method comprises the following steps: S1, carrying out the system modeling and initialization, and carrying out the modeling of a cooperative data collection task of an unmanned aerial vehicle mechanism cluster in an agricultural environment as a decentralized partially observable Markov decision process; s2, defining a multi-dimensional space, wherein the decentralized part observes a global state space, a joint action space and a local observation space of a Markov decision process; s3, designing a mixed reward function; s4, performing cluster strategy training; and S5, performing online distributed execution. The method effectively solves the problems that an existing method is poor in environment dynamic change adaptability, does not fully consider the heterogeneous characteristics of the unmanned aerial vehicles and does not fully consider decision-making and execution separation, and remarkably improves the operation efficiency, safety and collaborative intelligence level of an agricultural unmanned aerial vehicle cluster in a complex unstructured environment.
Owner:QINGDAO AGRI UNIV

Intelligent radar and jammer dynamic waveform gaming method based on curiosity driving

The invention provides an intelligent radar and jammer dynamic waveform gaming method based on curiosity driving, and belongs to the field of intelligent radar and intelligent jammer waveform gaming. The problems that the intelligent agent is difficult to effectively learn from limited environment feedback due to lack of a waveform-level dynamic game mechanism close to an actual electronic warfare and a sparse reward problem, and the ability imbalance of the intelligent agent is easy to occur in a training process are solved. The method comprises the following steps: constructing a confrontation environment model, namely modeling a dynamic game of a radar and a jammer into a partially observable Markov decision process, and defining a state space, an observation space, an action space and a reward function; designing a complex field feature extraction network; establishing a decision-evaluation network based on an MADDPG framework; a curiosity exploration strategy is introduced; implementing an opponent agent sampling strategy; and alternately training the radar and jammer agents. The device is mainly used in the electronic warfare field.
Owner:HARBIN INST OF TECH

Unmanned aerial vehicle moving edge calculation trajectory planning method based on multi-target reinforcement learning

The invention discloses an unmanned aerial vehicle moving edge calculation trajectory planning method based on multi-target reinforcement learning. Aiming at the problems of limited energy and uneven task distribution of unmanned aerial vehicles in a multi-unmanned aerial vehicle assisted mobile edge computing system, the method comprises the following steps: collecting real-time state data of the unmanned aerial vehicles, evaluating system target preferences (minimizing service interruption time or maximizing computing task load), constructing a distributed partially observable Markov decision process model, and calculating the target preferences of the unmanned aerial vehicles in the distributed partially observable Markov decision process model. And iteratively updating model parameters in combination with a crowding distance optimization experience playback mechanism so as to dynamically generate an optimal flight path. Experiments show that the method can adaptively balance and calculate the task load and the service continuity, effectively improves the system stability and the resource utilization rate, and is suitable for dynamic scenes such as disaster rescue and intelligent traffic.
Owner:WUHAN UNIV OF TECH

Caching, migration and unloading collaborative optimization method in satellite-assisted ocean network

The invention particularly relates to a cache, migration and unloading collaborative optimization method in a satellite-assisted ocean network, which comprises the following steps: constructing a satellite-assisted ocean network architecture, mapping task completion delay to system utility according to the satellite-assisted ocean network architecture, and constructing a target function with the purpose of maximizing the system utility under multi-dimensional constraints; converting an objective function optimization problem into a partially observable Markov decision process; and solving an objective function optimization problem based on attention-enhanced multi-agent reinforcement learning to obtain an optimal solution. According to the method, the problem of low-delay and efficient task calculation under the satellite-assisted ocean network architecture is solved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Vulnerability intelligent collaborative detection method and system based on multi-agent reinforcement learning

The invention discloses a vulnerability intelligent collaborative detection method and system based on multi-agent reinforcement learning. The method comprises the following steps: independently executing detection according to local observation information and a strategy generation model through an intelligent agent deployed at an edge node; the strategy generation model is constructed based on a decentralized partial observable Markov decision process, and is obtained through cloud centralized training and global shared reward function optimization. When the agent obtains a forward reward, packaging a corresponding vulnerability context and related strategy parameters into an experience package; and the experience packet is diffused in the cooperative detection network in a mode of randomly selecting neighbor nodes and forwarding step by step. And after each agent receives the data, analyzing the content of the experience package and comparing the content with the local environment so as to dynamically adjust local strategy parameters and realize collaborative optimization. According to the vulnerability intelligent collaborative detection method and system based on multi-agent reinforcement learning provided by the invention, the problems that network security vulnerability detection is poor in expandability and lacks knowledge generalization ability are solved.
Owner:SHIZUISHAN POWER SUPPLY COMPANY OF STATE GRID NINGXIA ELECTRIC POWER +1

Distributed energy storage aggregation scheduling method and system considering user fatigue effect

The invention discloses a distributed energy storage aggregation scheduling method and system considering a user fatigue effect, and the method comprises the steps: building a user response behavior model considering the long-term response preference and short-term fatigue of a user based on a partial observable Markov decision process; establishing a user response state online inference strategy based on maximum likelihood estimation, and inferring an unobservable state of a user; and based on the user response behavior model and the user response state online inference strategy, establishing an online learning aggregation scheduling method combined with the running state constraint of the distributed energy storage unit, and realizing distributed energy storage aggregation scheduling. According to the method, the user response uncertainty, the fatigue level and the actual response capacity of the energy storage equipment are comprehensively considered, a reliable model basis is provided for aggregation scheduling, effective tracking of the target power is achieved, and therefore the accuracy and reliability of aggregation scheduling are improved.
Owner:NANJING NORMAL UNIVERSITY

Graph reinforcement learning reactive voltage control method considering domain knowledge

The invention discloses a graph reinforcement learning reactive voltage control method considering domain knowledge, and the method comprises the steps: (I) enabling a reactive voltage control problem to be built into a partially observable Markov decision process, and enabling an intelligent agent corresponding to each inverter to only observe the local state quantity of a limited number of nodes around; (2) constructing a space-time diagram in which domain knowledge is embedded; (3) utilizing a graph attention neural network to carry out domain knowledge-based feature reinforcement learning of the multi-agent; and (IV) performing multi-inverter action decision making based on a GAMARL method and the like. According to the invention, by integrating domain knowledge and an advanced graph reinforcement learning technology, the voltage regulation capability of the novel power distribution system under the condition of high-proportion distributed photovoltaic access is remarkably improved, and the intelligent and sustainable development of a power distribution network is promoted.
Owner:TIANJIN UNIV

Communication lightweight multi-micro energy network cooperative control method for large model assisted reward recovery

The invention belongs to the technical field of multi-micro-energy-network cooperative control, and particularly relates to a communication lightweight multi-micro-energy-network cooperative control method for reward recovery assisted by a large model, and the method comprises the following steps: S1, constructing a mathematical model of distributed energy in a micro-energy network, and enabling the mathematical model to be used for describing the dynamic behaviors of energy conversion, storage and constraint; s2, modeling a micro energy network cooperative control problem as a partially observable Markov decision process POMDP; s3, solving the POMDP of the step S2 by using an optimized MAPPO method to obtain a cooperative control scheme; and issuing to each micro-energy network terminal for execution. The method can break through the limitation of the communication bottleneck on the premise of ensuring the learning quality of the control strategy.
Owner:CHONGQING UNIV

Micro-grid distributed robust optimization scheduling method considering source charge uncertainty and electricity-hydrogen-carbon coupling

The invention belongs to the technical field of power system operation and control, and particularly provides a microgrid distributed robust optimization scheduling method considering source load uncertainty and electricity-hydrogen-carbon coupling. Comprising the following steps: constructing an electricity-hydrogen-carbon-heat four-dimensional collaborative micro-grid energy flow framework comprising a multi-port alternating current and direct current bus, a hydrogen production storage and transportation subsystem, a carbon capture methanation subsystem and a cross-stage thermodynamic energy cascade network; establishing an IGDT-based source-load bilateral uncertainty dynamic pricing model, and selecting a risk avoidance or opportunity seeking pricing model as required; a multi-agent game double-layer interaction model with a micro-grid operator as a leader and a load aggregator and a charging station as followers is constructed, and a nested part can observe a Markov decision process; and a privacy protection multi-agent deep reinforcement learning algorithm in which a maximum entropy framework and a CTDE mechanism are introduced is adopted for solving. According to the method, the problem of low-carbon operation of the micro-grid in a source load uncertainty environment is solved, and collaborative improvement of economical efficiency, low-carbon performance and robustness is realized.
Owner:CHINA THREE GORGES UNIV

MARL-based flexible distributed energy network multi-agent collaborative optimization method

The invention provides a flexible distributed energy network multi-agent collaborative optimization method based on MARL, and the method comprises the steps: building a distributed energy network multi-agent interaction framework which comprises a local energy transaction model for distinguishing an intelligent soft switching path, a network cost model based on a power transmission distribution factor, and a power distribution network operator and user income model; a physical system operation constraint model is established and covers intelligent soft switch operation characteristics, network topology radial constraint, energy storage charging and discharging dynamic states, controllable distributed power supply and photovoltaic output limitation, translational load scheduling and DistFlow power flow constraint. And finally, constructing a collaborative optimization mechanism based on a multi-agent near-end strategy optimization algorithm, modeling the system into a partially observable Markov decision process, adopting a centralized training distributed execution framework, and guaranteeing transaction security and data privacy through block chain evidence storage, RSA encryption and an intelligent contract. The system economy and the operation flexibility are improved.
Owner:INST OF ELECTRICAL ENG CHINESE ACAD OF SCI

Distributed unmanned aerial vehicle base station low-altitude wireless signal coverage method

The invention relates to a distributed unmanned aerial vehicle base station low-altitude wireless signal coverage method, and belongs to the technical field of wireless communication. The method comprises the following steps: S1, modeling an unmanned aerial vehicle base station coverage optimization scene and proposing an optimization problem; s2, modeling the optimization problem into a Markov decision process with an observable decentralized part; s3, designing a network structure of a distributed multi-agent deep reinforcement learning algorithm; and S4, designing a training framework of a distributed multi-agent deep reinforcement learning algorithm. The method aims to solve the problem of unmanned aerial vehicle base station collaborative coverage optimization in a complex dynamic environment, and aims to overcome the defects of dependence on a global state, adaptability to a part of observable scenes and the like in a traditional method. According to the invention, an unmanned aerial vehicle base station collaborative coverage optimization problem in a partially observable environment is modeled into a decentralized partially observable Markov decision process, and a network structure and a training framework of a distributed multi-agent deep reinforcement learning algorithm are provided.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Cellular-free network-oriented end-side collaborative reasoning and dynamic resource allocation method

The invention discloses an end-side collaborative reasoning and dynamic resource allocation method for a cellular-free network, and belongs to the technical field of mobile communication. The method comprises the following steps: constructing an end-side collaborative inference system comprising a plurality of APs (Access Points), a plurality of UEs (User Equipment) and a CPU (Central Processing Unit) equipped with an MEC (Mobile Edge Computing) server under a cellular-free architecture; constructing a multi-dimensional resource allocation model with the aim of minimizing the total energy consumption of system reasoning, and jointly optimizing the UE-AP association relationship, the segmentation position of the DNN model and the computing resource allocation of the MEC server; modeling the problem as a partially observable Markov decision process; and designing a cooperative optimization algorithm based on multi-agent near-end strategy optimization, training the agents to quickly perceive dynamic task requests and wireless environment changes, and generating a resource scheduling strategy. According to the method, on the basis of meeting the time delay requirement of the UE, the overall reasoning energy consumption of the system can be remarkably reduced through reasonable multi-dimensional resource dynamic optimization.
Owner:SOUTHEAST UNIV

Synchronous adaptive link flooding attack defense method for heterogeneous networks with multi-dimensional identification

The present invention discloses a synchronous adaptive link flood attack defense method for a multi-dimensional identification heterogeneous network, comprising the following steps: S1, sampling the packet loss rate and link utilization of the network topology link to obtain the current state of the link; S2, extracting the key traffic features of the current state through the Honey Badger optimization algorithm; S3, inputting the key traffic features of the current state into a prediction model based on a spatiotemporal graph convolutional neural network to obtain a DDoS attack traffic prediction value; S4, identifying the attack action of the DDoS attack traffic prediction value, and selecting the optimal defense action through a partially observable Markov decision process model and a reward mechanism; S5, adjusting the defense strategy based on the optimal defense action selected in the current state. The present invention can accurately predict attack behavior and dynamically learn the optimal mitigation strategy in a novel multi-dimensional identification fusion network scenario to maximize normal traffic transmission and minimize the impact of attack traffic on bottleneck links.
Owner:BEIJING JIAOTONG UNIV

Unmanned aerial vehicle assisted mobile edge computing performance optimization method based on large language model and deep reinforcement learning

The invention discloses an unmanned aerial vehicle auxiliary mobile edge computing performance optimization method based on a large language model and deep reinforcement learning, and the method comprises the steps: building an unmanned aerial vehicle auxiliary mobile edge computing system model which covers an unmanned aerial vehicle motion, communication and calculation unloading model; a target optimization problem containing a task unloading strategy is formalized into a partially observable Markov decision process; a deep reinforcement learning algorithm based on a Transform module is constructed, an Actor network generates actions according to states and observation, and a Critic network evaluates the actions; and an experience playback buffer area is adopted to store interaction experience and update network parameters. According to the method, priori knowledge is generated through LLM pre-training to guide strategy updating, the model generalization ability and the multi-target optimization ability are remarkably improved, the communication time delay, energy consumption control and resource allocation are optimized, and an intelligent unloading solution is provided for unmanned aerial vehicle edge calculation.
Owner:HEBEI NORMAL UNIV

Multi-unmanned aerial vehicle cooperative data acquisition intelligent flight path planning method driven by curiosity rewards

The invention discloses a curiosity award-driven multi-unmanned aerial vehicle cooperative data acquisition intelligent flight path planning method. The invention belongs to the field of unmanned aerial vehicle control, and mainly solves the flight path planning problem that a multi-unmanned aerial vehicle system cooperatively completes Internet of Things sensor data acquisition in a complex and dynamic scene. The method is based on deep reinforcement learning, and data volume collection and task time are taken as a joint optimization target. The method comprises the following steps: firstly, modeling a multi-unmanned aerial vehicle data acquisition route planning problem in a complex and dynamic environment into a partially observable Markov decision process; then establishing a flight path planning strategy based on a deep cycle Q network, fusing historical environment information by using a long-short-term memory network, making an accurate decision by using the fused information, setting a dual-source composite reward function, combining an internal curiosity reward with an external environment guidance reward, encouraging the unmanned aerial vehicle to comprehensively explore a complex environment, and improving the flight path planning efficiency. And the training effect of the track planning strategy in a complex environment is improved. According to the method, it is ensured that the multi-unmanned-aerial-vehicle system achieves data collection of the sensor nodes, meanwhile, the higher flight path planning capacity is obtained, and the effectiveness and robustness of the method are proved through simulation experiments.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Adaptive closed-loop optimization management method for substation maintenance plan based on multiple elements

The invention relates to the technical field of intelligent power grids, and discloses a power transformation maintenance plan adaptive closed-loop optimization management method based on multiple elements, which comprises the following steps: constructing a multi-element database of power transformation equipment; based on the multi-element database, the maintenance priority of each device is calculated through a weight quantification model, and a maintenance priority sequence is generated; a mixed parthenogenesis optimization algorithm is adopted, the maintenance priority sequence is used as input, and an optimized maintenance plan is generated with the purpose of minimizing the comprehensive cost; and on the basis of an improved partially observable Markov decision process framework, modeling is performed on the execution process of the optimized maintenance plan, and a reinforcement learning mechanism is utilized to dynamically adjust and optimize a subsequent maintenance plan according to a feedback reward after the maintenance plan is executed, so that closed-loop iteration is formed. According to the method, adaptive dynamic optimization and sustainable evolution of the substation maintenance plan are realized, and the intelligent level, the safety reliability and the resource utilization efficiency of power grid operation and maintenance are remarkably improved.
Owner:HUAIAN OF JIANGSU ELECTRIC POWER CO POWER SUPPLY

Edge-cloud task unloading method based on multi-agent deep reinforcement learning

The invention belongs to the technical field of mobile communication, and particularly relates to an edge-cloud task unloading method based on multi-agent deep reinforcement learning, which is mainly embodied in the following aspects: conversion of an optimization problem, energy consumption and time delay optimization problems under an edge-cloud task unloading scene are NP-Hard and are difficult to solve by using a traditional optimization method, and the problem is solved by using a traditional optimization method. The optimization problem is converted into a partially observable Markov decision process, and an unloading decision is optimized by using an MADRL method; a reward function is designed, the reward function of the MADRL comprises energy consumption and time delay, and a weighting factor is used for balancing the proportion of the energy consumption and the time delay in the optimization process; according to the introduction of the attention mechanism, the attention mechanism is added into a loss function of MADRL, so that the model can better distinguish the importance degrees of different features of the intelligent agent. According to the method, the energy consumption and the time delay of task unloading can be effectively balanced, and meanwhile, a more accurate unloading decision is made based on an attention mechanism.
Owner:HARBIN INST OF TECH

Unmanned aerial vehicle cluster dynamic area coverage method based on deep reinforcement learning

The invention discloses an unmanned aerial vehicle cluster dynamic area coverage method based on deep reinforcement learning. The method comprises the following steps: constructing a multi-physical field coupling simulation environment comprising node motion, a directional antenna, a space channel and a radar detection model; based on the environment, converting a dynamic coverage problem form into a partially observable Markov decision process, a defined state, observation, action and reward mechanism, performing multi-agent training by adopting a centralized training decentralized execution framework, evaluating a joint action by using a global Critic network during training, and storing interactive experience by using an experience playback buffer area; based on a priority experience playback mechanism, according to a time sequence difference error and a reward value, an experience sampling probability is calculated, samples are extracted, Critic and Actor networks are updated in sequence, a decision model is deployed to an unmanned aerial vehicle cluster after training is completed, and each unmanned aerial vehicle executes an area coverage task in a distributed manner, so that efficient and low-risk communication coverage can be realized in a complex threat environment. The communication quality is improved; and the detected probability is reduced.
Owner:10TH RES INST OF CETC

Policy learning method with privacy protection in mobile edge computing for intelligent agent

A policy learning method with privacy protection in mobile edge computing for an intelligent agent is provided, relating to the technical field of mobile communication. The method includes: establishing an edge-collaborative computing offloading model, where the edge-collaborative computing offloading model includes a service caching model, a task offloading model, and a system cost model; establishing an optimization problem for task offloading, service caching, computing resource allocation and transmission power control based on the edge-collaborative computing offloading model for minimizing task processing costs; abstracting the optimization problem to a partially observable Markov decision process; and autonomously learning a task offloading strategy, a service caching strategy, a computing resource allocation strategy, and a transmission power control strategy by using a federated learning-based multi-agent deep reinforcement learning algorithm based on the Markov decision process.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Frequency modulation demand quantification method based on attention-guided multi-agent reinforcement learning

The invention relates to the field of power system operation control, and discloses a frequency modulation demand quantification method based on attention guidance multi-agent reinforcement learning, which comprises the following steps: constructing a main network frequency response model; based on a main network frequency response model, modeling a frequency modulation demand online quantification task as a distributed partial observable Markov decision process; embedding a historical observation-action sequence of each agent through a track encoder, obtaining role representations of a plurality of agents through comparative learning of a role encoder and InfoNCE, and constructing a multi-head attention mechanism by taking global state embedding as query and the role representations as key values; the action of each agent is output at the actor end by adopting a strategy network sharing parameters; and according to the real-time operation state, calculating the frequency modulation demand meeting the RoCoF and frequency lowest point constraint and an allocation scheme of the frequency modulation demand among various frequency modulation resources. The method has the advantage that rapid and self-adaptive online quantification of the main network frequency modulation requirement is realized.
Owner:SICHUAN UNIV

Large-model intelligent agent interactive question-answering task decision-making method, device, equipment and medium

The present invention relates to a large-model intelligent agent interactive question-answering task decision-making method, device, equipment and medium, the method comprising: constructing an intelligent agent strategy model based on a partially observable Markov decision process; obtaining user input instructions, and using the intelligent agent strategy model to generate an information search trajectory; using a Monte Carlo tree search algorithm to construct a decision tree of the information search trajectory to obtain a preference data set labeled with Q values; setting a large language model, and training the large language model using a step-level optimization objective function based on the preference data set labeled with Q values ​​to obtain a trained large language model; constructing a Q-value model based on the trained large language model, and then outputting a Q-value estimate based on the Q-value model; selecting an action based on the Q-value estimate and the information search trajectory to generate a question-answering process. The present invention can significantly enhance the performance and generalization ability of large-model intelligent agents in autonomous reasoning and decision-making processes.
Owner:NAT UNIV OF DEFENSE TECH

Mine multi-robot adaptive cooperative control method based on micro-kernel operating system

The application discloses a mine multi-robot adaptive cooperative control method based on a microkernel operating system and belongs to the technical field of mine automation and robot cooperative control. The method cooperatively realizes the following processes by running multiple independent service processes on the microkernel operating system: a multi-modal perception fusion process dynamically adjusts sensor weights according to environmental dust and light, and constructs a global semantic map; a dynamic task allocation process models task allocation as a partially observable Markov decision process for solution by using a multi-agent deep reinforcement learning embedded with an attention mechanism; a hybrid path planning process constructs a hierarchical cost model based on historical traffic data, and generates a path through a three-layer planning system; a dynamic conflict resolution process uses an STCN to predict a trajectory, and realizes forward-looking conflict resolution in combination with a distributed negotiation auction mechanism; and a state synchronization process uses a lightweight practical Byzantine fault tolerance consensus to ensure that the states of the whole system are consistent. The application significantly improves the real-time performance, reliability and intelligence level of the system.
Owner:CITIC HEAVY INDUSTRIES CO LTD

A self-reconfiguration method for free-form modular robots

ActiveCN116300423BArtificial lifeAdaptive controlHardware prototypingControl engineering
This invention discloses a self-reconfiguration method for free-form modular robots. First, an abstract model of the basic elements of a class of free-form modular robots is constructed, where each robot module can be connected to each other from any direction via active and passive connectors. Then, based on a hardware prototype, a reasonable action space and observation space are designed for this type of resource-constrained modular robot. The action space includes the drive parameters of two planar continuous motion motors and the motion parameters of one connector, while the observation space includes the motion information and local connectivity relationships of the robot itself and its neighboring modules. The reconfiguration problem of the free-form modular robot is modeled as a distributed partially observable Markov decision process. Based on a proximal policy optimization algorithm, an efficient and conflict-free multi-agent reinforcement learning method for reconfiguration in the continuous action space can be achieved. Finally, the timely termination of the reconfiguration process is achieved under the conditions of local observability and limited communication.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Decomposition value evaluation multi-agent cooperative navigation method based on attention mechanism

The invention discloses a decomposition type value evaluation multi-agent collaborative navigation method based on an attention mechanism. The method comprises the following steps: modeling a collaborative navigation task of multiple agents into a decentralized partially observable Markov decision process; designing local attention modules of a precoding module and an attention coding stage to extract and code environment feature information of multiple agents, and sharing the local attention modules by a value network and a strategy network; designing a multi-agent strategy-value method under decomposition type value evaluation; constructing a static target navigation environment, and guiding the mobile intelligent agent to actively search for an optimal position; constructing a cooperative hunting environment; a multi-agent cooperative navigation model is constructed, an optimal model is obtained through training and testing, and non-explicit and complete perception navigation of multi-agent cooperation with the reliability distribution problem is achieved. The collaborative navigation method provided by the invention has the advantages of faster convergence, higher success rate, shorter time and less collision, and shows good generalization ability and robustness.
Owner:ANHUI UNIV

Robot control method and system based on cyclic depth deterministic strategy gradient algorithm

The invention relates to a robot control method and system based on a cyclic depth deterministic strategy gradient (RDDPG) algorithm, and aims to improve the decision and control capability of a robot in a complex dynamic environment. The method comprises the following steps: firstly, modeling a robot control system into a partially observable Markov decision process, and defining an environment state, an action space, a state transition probability, a reward function and an observation function; secondly, constructing a cyclic encoder by adopting a cyclic neural network, taking motion time sequence data and environment observation time sequence data of the robot as input, and outputting meta-parameters for identifying environment differences; thirdly, designing an evaluation value network, an evaluation strategy network, a target value network and a target strategy network, generating a control action and evaluating the value of the control action; finally, empirical data of interaction between the robot and the environment is stored through playback memory, network parameters are updated through time sequence difference learning, and target network parameters are updated through a moving average method. According to the invention, the cyclic encoder is introduced, the time sequence characteristics of environmental information are fully utilized, the adaptability, learning ability and decision accuracy of the robot in a complex dynamic environment are enhanced, and the method can be widely applied to the fields of industrial automation, intelligent logistics, service robots and the like.
Owner:GUILIN UNIV OF ELECTRONIC TECH