Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

345 results about "Value network" patented technology

A value network is a business analysis perspective that describes social and technical resources within and between businesses. The nodes in a value network represent people (or roles). The nodes are connected by interactions that represent tangible and intangible deliverables. These deliverables take the form of knowledge or other intangibles and/or financial value. Value networks exhibit interdependence. They account for the overall worth of products and services. Companies have both internal and external value networks.

Quadruped robot robust motion control method based on deep reinforcement learning

The invention discloses a quadruped robot robust motion control method based on deep reinforcement learning, and belongs to the technical field of robot motion control, and the method comprises the steps: constructing a deep reinforcement learning model which comprises a state estimation network, a strategy network and a value network; interaction between the quadruped robot and the simulation environment is carried out, and standard observation information, historical observation information and privileged observation information of the quadruped robot at all moments are obtained; inputting the standard observation information, the historical observation information and the privilege observation information of the moment into a deep reinforcement learning model, and training based on a total loss function until convergence is carried out to obtain a trained deep reinforcement learning model; inputting standard observation information and historical observation information at corresponding moments in an actual scene into the trained deep reinforcement learning model to obtain output features of a corresponding strategy network; and the target position of each joint motor is calculated to complete the motion control of the quadruped robot. And efficient training and robust motion on various complex and unstructured terrains can be realized.
Owner:ZHEJIANG UNIV OF TECH

Information retrieval system and method based on semantic normalization

The invention discloses an information retrieval system and method based on semantic normalization, and relates to the technical field of artificial intelligence information, and the method comprises the steps: collecting a semantic query record input by a user, carrying out the preliminary semantic analysis, and generating structured data; on the basis of the structured data, entity disambiguation is carried out by utilizing a knowledge graph, abstract classes are generated through a neural network, calibration and dynamic weight adjustment are carried out, and high-confidence entity abstract classes and confidence scores are generated; entity abstract classes and confidence scores are combined with user contexts, an action-value function is calculated through a value network, and an optimal action is selected by utilizing a-greedy algorithm; executing semantic normalization mapping according to the optimal action, and obtaining an intermediate expression by using a meta-symbol dynamic generator; and performing index retrieval and multi-dimensional sorting based on the intermediate expression to generate a sorted retrieval result list. According to the method, the semantic fragmentation problem of multi-modal query is solved, and deep semantic alignment and dynamic weight calibration of heterogeneous data are realized.
Owner:上海笑聘网络科技有限公司

Crop nitrogen fertilizer management system and method based on multi-agent reinforcement learning

The invention discloses a crop nitrogen fertilizer management system and method based on multi-agent reinforcement learning, and belongs to the technical field of intelligent agriculture. Comprising the following steps: collecting and preprocessing multi-source data of a target area, calibrating a crop growth-nitrogen cycle model based on the multi-source data, and constructing a dynamic simulation environment; constructing a nitrogen fertilizer application strategy model and a multi-target award function, and performing agent reinforcement learning training by adopting a centralized training-decentralized execution architecture; introducing a large language model, and updating a strategy network and / or a value network of each agent in a training process according to a reward adjustment signal and a decision constraint; when the index fluctuation ratio in the continuous evaluation period is smaller than a preset threshold value, it is judged that the nitrogen fertilizer application strategy model is converged, and an optimal nitrogen fertilizer application strategy model is obtained; and generating a nitrogen fertilizer application scheme based on the optimal nitrogen fertilizer application strategy model in combination with the real-time state data, and generating a natural language interpretation and risk assessment report based on a large language model.
Owner:INST OF SOIL SCI CHINESE ACAD OF SCI

Relationship-driven multi-agent reinforcement learning method and system based on mixed game

The invention belongs to the field of multi-agent reinforcement learning, and discloses a relation-driven multi-agent reinforcement learning method and system based on a mixed game, and the method comprises the steps: enabling a strategy network to generate an agent action, and carrying out the interaction of an environment, so as to collect sample data; the centralized value evaluator calculates marginal influence values of the agents based on samples, and deduces a social influence weight set; assigning group external rewards as individual external social rewards based on the weights; processing the global state by using a random network distillation-driven method, updating a prediction network parameter to minimize a prediction error, and outputting an internal reward set; the external social rewards and the internal rewards are fused to form comprehensive rewards; and updating the strategy network and the value network by using the comprehensive reward, and circularly training until convergence. By adopting the method, accurate modeling of the individual interaction relationship is enhanced, invalid exploration is remarkably reduced, the learning efficiency and the strategy reliability are improved, and the overall performance of the system is enhanced through an explicit social mechanism.
Owner:XI AN JIAOTONG UNIV

Data center IT load and cooling system cooperative control method based on TD3 algorithm

The invention provides a data center IT load and cooling system cooperative control method based on a TD3 algorithm, and relates to the field of deep reinforcement learning, and the method comprises the steps: obtaining an initial data set representing the server state, environment and task characteristics of a data center, and processing the initial data set to obtain a target data set; building a deep reinforcement learning model which is based on a TD3 algorithm and comprises a strategy network Actor and a value network Critic, and determining a state space, an action space and a reward function of the model; training the model to obtain a target model; and outputting a work task allocation strategy and a cooling system regulation and control strategy through interaction of the target model and the environment. According to the target model, by outputting the work task allocation strategy and the cooling system regulation strategy, work task allocation can be guided, operation of the cooling system can be controlled, the IT load of the server and refrigeration balance of the cooling system are achieved, safe and stable operation of the server is ensured, and energy consumption of the cooling system is reduced.
Owner:HEFEI UNIV OF TECH

Robot motion control model training method, device and equipment based on deep reinforcement learning, robot and medium

The invention provides a robot motion control model training method, device and equipment based on deep reinforcement learning, a robot and a medium, and relates to the technical field of robots. The method comprises the following steps: acquiring a first potential vector obtained after a student encoder encodes robot body observation data, and a second potential vector obtained after a teacher encoder encodes privilege observation data; based on the current training step number and a preset probability function, calculating a sampling probability for controlling a fusion proportion of the first potential vector and the second potential vector; fusing the first potential vector and the second potential vector based on the sampling probability to generate a third potential vector, and inputting the third potential vector into a strategy network; and updating the parameters of the policy network based on the value estimation of the current state output by the value network and the action policy output by the policy network. According to the method, updating oscillation caused by sudden change of input distribution in the training process of the strategy network can be avoided, the training efficiency is improved, and the training cost is reduced.
Owner:SHENZHEN ZHUJI POWER TECH CO LTD

Aircraft engine maintenance policy optimization method based on nonparametric reinforcement learning

Disclosed in the present invention is an aircraft engine maintenance policy optimization method based on nonparametric reinforcement learning. Firstly, for sparse aircraft engine operation data, constructing an aircraft engine model by means of a Bayesian network and a Gaussian process; then, establishing a policy network and a value network, interacting with the aircraft engine model to form a state / action value group, storing same into a replay buffer, and performing random sampling for use in a training set; updating the policy network and the value network, and updating the training set; and finally, performing aircraft engine maintenance optimization policy decision-making. Provided in the present invention is, for the first time, a Gaussian process-based nonparametric reinforcement learning method for an aircraft engine, which improves the degree of fit between overall training data and a model by means of dynamic data updating while integrating system uncertainty, thereby improving the sampling efficiency of an algorithm. Action selection is performed on the basis of prior maintenance experience data fused with uncertainty, improving the safety and sampling efficiency of a system and thereby solving the problem of predictive maintenance of an aircraft engine system.
Owner:ZHEJIANG UNIV

Sparse reward environment optimization learning identification method and system based on demonstration data enhancement

The invention relates to the technical field of mechanical arm optimization learning, and discloses a sparse reward environment optimization learning method and system based on demonstration data enhancement, and the method comprises the steps: obtaining expert demonstration data of a grabbing task of a mechanical arm, generating an enhanced demonstration sample through track segmentation time sequence interpolation and state space neighborhood extension, and carrying out the reconstruction of the enhanced demonstration sample; constructing a demonstration experience playback buffer area; building a reinforcement learning framework containing a strategy network and a value network, and learning an optimal grabbing strategy by utilizing reinforcement demonstration; designing a lightweight collaborative knowledge fusion model to monitor an action change rate, and outputting a smooth action when the action change rate exceeds a threshold value; and a reward function containing a task target and dynamic constraints is constructed, knowledge fusion rewards are superposed when the fusion model is activated, and reward sparsity compensation and strategy optimization are realized. According to the method, the executing stability and safety of the mechanical arm are improved, and the parameter updating accuracy and the training convergence speed are improved.
Owner:CHANGCHUN UNIV OF TECH

Flow field video generation method based on policy value architecture and online physical exploration

The invention discloses a flow field video generation method based on a policy value architecture and online physical exploration, and belongs to the technical field of crossing of artificial intelligence and computational fluid dynamics (CFD), and the method comprises the following steps: step 1, constructing an unsteady flow field multi-modal training data set, step 2, constructing a generative network system based on an Actor-Critic architecture, step 3, constructing an unsteady flow field multi-modal training data set, and step 4, constructing an unsteady flow field multi-modal training data set. Step 4, supervised fine tuning training is carried out in the first stage; step 5, online physical exploration of a generator is carried out in the second stage; step 6, feedback co-evolution of a physical encoder is carried out in the third stage; and step 7, reasoning generation of an unsteady flow field video is carried out. According to the method, a reinforcement learning architecture containing an Actor and a Critic is constructed, a physical equation is packaged into a digital environment, and a training strategy of basic supervision fine tuning, online physical exploration of a generator and coevolution feedback of an encoder is adopted.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Multi-agent combat mission cooperation method of structure entropy guided graph neural network

The invention discloses a multi-agent combat task cooperation method for a structure entropy guided graph neural network, and the method comprises the steps: S10, each combat agent interacts with an environment according to an action generated by a strategy network, the environment comprises environment information, task parameters and a preset task target, and the strategy of each combat agent is completely executed in a decentralized manner; collecting complete empirical trajectory data; s20, using the collected data for centralized training; performing value evaluation on the global state of each time step by using a value network; s30, calculating strategy loss and value loss by using a multi-agent near-end strategy optimization algorithm in combination with the output of the strategy network and the value estimation of the output of the value network; updating parameters of the strategy network and the value network by using a gradient descent method; and S40, performing loop iteration. The problems that in a traditional method, the battlefield game dynamic structure sensing ability is insufficient, the hierarchical strategy learning and generalization ability is limited, the adaptability of a model in a small sample area is poor, and the migration efficiency is low are solved.
Owner:BEIHANG UNIV

Rail transit multi-scene intelligent inspection device, method and equipment based on reinforcement learning

The invention relates to a rail transit multi-scene intelligent inspection device, method and equipment based on reinforcement learning, and the device comprises a sensing and state characterization module which is responsible for collecting multi-modal information from a rail transit inspection environment, and processing and fusing the multi-modal information into a structured environment state; the reinforcement learning decision-making module is used for selecting and outputting a current optimal decision-making action by utilizing the trained strategy network and value network based on the current structured environment state and the inspection task; the action execution and control module is used for converting the decision action into a bottom layer physical instruction sequence and controlling an intelligent agent to execute; the reinforcement learning training and optimization module is used for carrying out iterative optimization on the strategy network and the value network according to the inspection task and empirical data, and updating model parameters; and the task planning and scheduling module is used for defining inspection tasks and task allocation. Compared with the prior art, the method has the advantages of improving inspection efficiency and quality, realizing risk active early warning and the like.
Owner:CASCO SIGNAL LTD

Spacecraft attitude and orbit integrated game control method and equipment based on PPO

The invention provides a PPO-based spacecraft attitude and orbit integrated game control method and device, and the method comprises the steps: building a spacecraft attitude and orbit integrated model, and carrying out the modeling of the interaction process of a tracking spacecraft and a target spacecraft as a non-cooperative zero-sum game; setting a multi-agent deep reinforcement learning framework, wherein the tracking spacecraft and the target spacecraft are respectively configured with an independent strategy network and an independent value network; training is carried out in a double-agent PPO mode, control instructions of orbit thrust and attitude torque are generated through the strategy network, and the state value is evaluated through the value network; and controlling and guiding the behavior of the spacecraft according to the trained strategy. According to the method, the autonomous decision-making capability and cooperative control performance of the spacecraft in complex tasks are effectively improved, and the method is suitable for high-precision aerospace tasks such as interception of non-cooperative targets, autonomous trajectory planning and attitude control.
Owner:HUBEI LUOJIA LAB

Cold chain transportation path optimization method based on intelligent scheduling

The invention discloses a cold chain transportation path optimization method based on intelligent scheduling, and the method comprises the steps: collecting the environmental data of a transportation path, and constructing a directed graph structure containing the temperature and length attributes of a path segment; defining a cooling capacity consumption function, calculating an estimated cooling capacity consumption value of the path section, and forming an environment diagram containing the refrigeration cost; constructing a state vector and an action space, and designing a path selection strategy network and a state value network; setting a multi-objective reward function including cold consumption, delivery completion, cold chain failure and supply behaviors, performing multi-round strategy training by adopting an improved PPO algorithm, and optimizing and updating network parameters through a trust region strategy; and finally, calling an optimal strategy output path selection action in actual transportation to realize dynamic avoidance and supply insertion regulation and control of the cold chain interruption risk. According to the invention, intelligent, self-adaptive and energy-efficient optimization of cold chain path selection can be realized.
Owner:BEIJING LONGXUNDA COLD CHAIN TRANSPORTATION CO LTD

Layout automatic generation method based on experience module and related device

The invention discloses an automatic layout generation method based on an experience module and a related device, and relates to the technical field of computers.The method comprises the steps that a template file is extracted, an independent device is generated based on the template file, and an experience module library is constructed based on the independent device; determining coordinate mapping and size constraints of the device, and performing multi-channel chip canvas hierarchical construction based on the coordinate mapping and size constraints of the device to obtain multi-channel chip canvas hierarchical representation; based on the empirical module library and the multi-channel chip canvas hierarchical representation, a reinforcement learning model is used for layout wiring, a target layout is obtained, and the reinforcement learning model comprises a state generation model, a value network and a strategy network. Layout and wiring of devices can be controlled more accurately, stray capacitance and resistance are effectively reduced, and the performance and reliability of an analog integrated circuit are effectively improved.
Owner:SUN YAT SEN UNIV +1

Intersection phase structure optimization method based on large language model

The invention belongs to the technical field of urban traffic planning and intelligent traffic systems, particularly relates to an intersection phase structure optimization method based on a large language model, and aims to improve the intelligent level and operation efficiency of traffic signal control. According to the method, semantic mapping cues of traffic flow and a phase structure are constructed, and a large language model is guided to generate a phase structure scheme adapted to an actual traffic state. Compared with a traditional scheme depending on artificial experience and a fixed structure, the method can automatically generate diversified and data-driven phase structure combinations, and has higher adaptability and generalization ability. In the aspect of technical implementation, the method fuses prompt engineering and guides a large language model to generate an initial phase scheme, and performs evaluation and feedback by using a value network, so that optimization of a phase structure is realized, and the overall operation efficiency of a traffic system is improved.
Owner:DALIAN UNIV OF TECH

Port container automatic scheduling method based on multi-agent reinforcement learning

The invention discloses a port container automatic scheduling method based on multi-agent reinforcement learning, and the method comprises the steps: S1, building a corresponding relation between equipment and agents, and constructing a task set; s2, collecting operation state data, and constructing global and local state vectors; s3, generating a scheduling constraint vector, and cutting actions according to the resource, storage yard and path state to form a feasible action set; s4, on the basis of an improved QPLEX algorithm, constructing an individual value network containing a dump structure, and calculating an individual action value; s5, constructing a joint action value hybrid network, and mixing individual values according to the global state vector to form joint action values; s6, constructing a training sample, differentiating and aggregating instant and delayed return, and updating network parameters; and S7, during online scheduling, selecting an optimal action combination according to the combined action value, and generating and issuing a scheduling instruction. According to the invention, automatic collaborative scheduling of port container operation is realized.
Owner:安徽海润信息技术有限公司

Medical image co-processing system and method based on block chain value network

The invention relates to the technical field of medical information and block chain crossing, in particular to a medical image co-processing system and method based on a block chain value network, and the system comprises a block chain network, a privacy calculation gateway, an intelligent contract module, a hybrid storage module and a communication module. The intelligent contract module is provided with a collaborative annotation sub-module, a verifiable federated learning sub-module and a dynamic value distribution sub-module; the method comprises the steps of data preprocessing and encryption uploading, collaborative labeling and quality verification, verifiable federated learning model training, dynamic value distribution, consensus evidence storage and supervision auditing. Medical data islands are broken through by means of the block chain network, data availability and invisibility are achieved in combination with the privacy computing gateway multi-encryption technology, the contradiction between data sharing and privacy security is solved, accurate contribution and return matching is achieved through the multi-dimensional contribution degree evaluation model and the intelligent contract, and the accuracy of data sharing is improved. The traceability and the compliance are ensured by full-process operation logs and data uplink evidence storage.
Owner:AFFILIATED HOSPITAL OF JIANGNAN UNIV +1

Industrial furnace temperature control optimization method and system based on model predictive control

The invention provides an industrial furnace temperature control optimization method and system based on model predictive control. The method comprises the following steps: defining a state space and an action space of a temperature control system; an MPC prediction model is established based on the state space and the action space, a future temperature trajectory is predicted, and a prediction error compensation mechanism and a self-adaptive prediction domain are introduced; constructing a deep reinforcement learning model which comprises a state value network and a strategy network and is used for learning an optimal temperature control strategy; based on the MPC prediction model and the deep reinforcement learning model, constructing a multi-objective optimization model for controlling performance, energy consumption and temperature uniformity objectives, and solving to obtain an optimal control sequence; and taking the optimal control sequence as a reference trajectory, and alternately updating parameters of the deep reinforcement learning model by adopting an Actor-Critic architecture. Through multi-level collaborative optimization, the adaptability, the control precision and the energy utilization efficiency of the industrial furnace temperature control system are improved.
Owner:CHENZHOU JIARUN CHANGFU INTELLIGENT ROBOT CO LTD

Multi-agent reinforcement learning regional energy collaborative scheduling method and system

The invention provides a multi-agent reinforcement learning regional energy collaborative scheduling method and system, and belongs to the field of regional energy system scheduling. Coupling degrees and a coupling degree matrix between agents are constructed; inputting the observation vector into a strategy network to obtain a decision action; individual basic rewards and system economic rewards are calculated, and constraint reference rewards are constructed; individual differentiation basic rewards are calculated, and rewards are distributed; inputting the decision action into a physical quantity prediction network, calculating a physical consistency reward, obtaining a final reward and a global reward, and calculating a target return; splicing observation vectors and decision actions of all agents, splicing global joint observation vectors and joint action vectors, inputting the spliced vectors into a value network, and training; inputting the local state set into the trained strategy network, outputting a scheduling instruction, inputting the scheduling instruction into the trained value network, and outputting an evaluation result; the problems of depiction rigidness of an intelligent agent coupling relation, lack of a cooperative benefit distribution mechanism and insufficient decision physical consistency are solved.
Owner:国网安徽省电力有限公司营销服务中心 +1

Geofence-based real-time coexisting user social system and method

ActiveCN121418761ALocation information based serviceCredit systemData set
The invention belongs to the technical field of geo-fences and real-time social contact, and discloses a real-time coexistence user social contact system and method based on geo-fences, and the method comprises the steps: collecting the multi-dimensional multi-modal original data of a user through a dynamic sampling strategy and an energy efficiency balance mechanism, and generating an information-desensitized four-dimensional data set; constructing four layers of fences, evaluating the value attribute of each fence, carrying out life cycle management, dynamically adjusting the fence state, and forming a four-dimensional fence value network; identifying a basic event, a composite event and a cooperative event in a grading manner, constructing a real-time value flow through event value quantification, and generating a dynamic rule set adaptive to a scene; user permissions are distributed through the three-dimensional value permission matrix, the multi-dimensional value point system is linked to complete point generation and transaction, and a value circulation record including permission states, point changes and transaction records is formed; the value health degree is analyzed through a real-time, short-term and long-term three-level optimization mechanism, optimization parameters are generated, and the optimization parameters are reversely iterated to corresponding preorder links.
Owner:SHANGHAI XIANGYUE JIANGFENG DIGITAL TECHNOLOGY CO LTD

Water-air cross-medium unmanned aerial vehicle control modeling method based on reinforcement learning and related device

The invention discloses a reinforcement learning-based control modeling method for a water-air cross-medium unmanned aerial vehicle and a related device, and the method comprises the steps: carrying out the interaction of a strategy network and a simulation environment, and generating a series of trajectory data; predicting a state value according to the trajectory data through a value network, and calculating a corresponding dominant function and a target value function based on the state value; calculating a value loss item, a strategy loss item and a strategy regular item through a strategy optimization algorithm so as to construct a total loss function, and performing joint training update on the strategy network and the value network by taking minimization of the total loss function as a target to obtain the updated strategy network and the updated value network; and judging whether the updated strategy network and the updated value network meet a preset dual convergence condition, if so, taking the updated strategy network and the updated value network as a water-air cross-medium unmanned aerial vehicle control model, and guiding the unmanned aerial vehicle to realize smooth and stable autonomous control in air, underwater and transition stages. And the intelligent autonomous control level of the cross-medium unmanned aerial vehicle is improved.
Owner:GUANGDONG UNIV OF TECH

Vertical handover method for sky-air-ground integrated cyber-physical system

PCT designated stageWO2025232049A1Wireless communicationNetwork outputPhysical system
Disclosed in the present invention is a vertical handover method for a sky-air-ground integrated cyber-physical system. The method comprises: a base station acquiring state spaces sent by mobile devices; in the base station, using a constrained proximal policy optimization algorithm to obtain gradient values of vertical handover policy networks and gradient values of value networks, and sending the gradient values to a central learner; determining in the central learner a first average gradient value and a second average gradient value, determining updated parameters of the vertical handover policy networks and updated parameters of the value networks, and sending the updated parameters to the mobile devices; and in the mobile devices, when the cumulative value of all reward values is greater than or equal to a target value and the cumulative value of the reward values is stable, enabling finally obtained updated vertical handover policy networks to output final policies. The vertical handover method for a sky-air-ground integrated cyber-physical system provided in the present invention can be directly used for the vertical handover of a sky-air-ground integrated cyber-physical system.
Owner:XIDIAN UNIV

Urban power grid coordination control method and system based on deep reinforcement learning

The invention discloses an urban power grid coordination control method and system based on deep reinforcement learning, and solves the technical problem that the current urban power grid coordination control method causes the limitation of the urban power grid power flow adjustment precision. The method comprises the following steps: after acquiring an expert experience data set, normalizing an urban power grid state vector at the current moment to obtain a normalized vector; and pre-training the initial strategy network by minimizing a loss function according to the vector and the expert experience data set to obtain a pre-trained strategy network, and calculating and outputting the active power of the unified power flow controller and the reactive power of the reactive compensation equipment. And in combination with the power parameters and the normalized vector, determining a normalized future state vector and a current reward value by using a power flow equation and a multi-target reward function, optimizing the pre-training network by using a PPO target function and a value network to obtain a target strategy network, and outputting a target action vector as a coordination control scheme.
Owner:ZHUHAI POWER SUPPLY BUREAU GUANGDONG POWER GIRD CO

High-speed communication chip macro placement method based on reinforcement learning

The invention provides a high-speed communication chip macro placement method based on reinforcement learning. The high-speed communication chip macro placement method comprises the following steps: extracting static features and dynamic features from a netlist; fusing the static features and the dynamic features, and inputting the fused static features and dynamic features into a shared graph neural network based on a reinforcement learning architecture; the strategy network outputs action probability distribution according to the current environment state, and the legal action with the maximum probability value is selected for execution; the value network outputs accumulated value estimation of the current environment state; after the action is executed, instant rewards and next state information are obtained, and the strategy network and the value network are updated according to the obtained information; and constructing a multi-objective optimization function based on macro placement constraint. According to the method, through deep fusion of reinforcement learning and constraint modeling, the automation level and design quality of macro cell placement are improved, and therefore the requirements of a high-speed communication chip for low power consumption, high performance and compact area are met.
Owner:CORE TREND (ZHUHAI) TECH CO LTD +1

Aircraft cluster multi-task scheduling system based on cooperative game and working method thereof

The invention discloses an aircraft cluster multi-task scheduling system based on a cooperative game and a working method thereof. The system comprises a simulation environment module, a hierarchical strategy network module, a centralized value network module and a training and execution module. The simulation environment module is used for constructing a multi-agent air combat confrontation environment and generating states, rewards and interaction data required by training; the hierarchical strategy network comprises a shared space-time representation encoder f theta (.), a high-layer strategy network pi H (aHz) and a low-layer strategy network pi L (aLz; aH); the centralized value network is used for receiving global state information in a training stage, estimating the overall return of our agent cluster, calculating a dominant function, and realizing the optimization of the network by minimizing the value loss; and the training and execution module optimizes a centralized value network parameter and hierarchical strategy network parameters theta H and theta L by using a global state St in a centralized training stage, and outputs an air combat decision ai according to local observation independent decisions of each agent in a distributed execution stage.
Owner:HEBEI UNIV OF TECH

Award centralization-based unmanned mine clearing equipment agent cooperation method

The invention discloses an unmanned demining equipment agent cooperation method based on reward centralization, and the method comprises the steps: taking unmanned demining equipment as an agent, and constructing a multi-agent decision framework in a combat area, and a reward centralization mechanism and an experience playback pool in a plurality of unmanned demining equipment value networks; training a plurality of deep neural networks formed by the multi-agent double-delay depth deterministic strategy gradient architecture based on the reward centralization to obtain a multi-agent double-delay depth deterministic strategy gradient model based on the reward centralization; and solving the model. The core of the method is to transform the time difference target of the global value network, and amplify the value difference between different state-action combinations by introducing an average reward mechanism. By means of the design, the value function approximator can focus on mining a real and effective cooperative strategy path more accurately and efficiently, and the individual action value estimation deviation caused by the unsteady state environment and the multi-agent coupling effect is effectively reduced.
Owner:XIAN UNIV OF TECH

Battlefield multi-agent dynamic cooperation method based on attention mechanism

The invention discloses a battlefield multi-agent dynamic cooperation method based on an attention mechanism. The method comprises the following steps: S10, acquiring local observation information; s20, establishing a strategy network and a value network, embedding a multi-head attention mechanism in each of the strategy network and the value network, allowing each agent to dynamically query the internal representation of a teammate, and converting a search problem of an exponential level joint action space into dynamic attention on key information; s30, the strategy network generates action probability distribution of the intelligent agent through the attention layer, and the intelligent agent is forced to depend on a cooperation signal; the value network outputs the value estimation of each agent and provides a reference function; the collaborative loss function calculates similarity penalty between agents based on the embedded representation output by the attention layer; and S40, the training process is updated by adopting a strategy gradient method, strategy network parameters are updated to maximize accumulated rewards, value network parameters are updated by minimizing total loss, and a highly cooperative joint strategy is output. According to the invention, the agent is ensured to always integrate team information during decision making.
Owner:BEIHANG UNIV

Service policy generation and optimization method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a service strategy generation and optimization method, device, equipment and medium, and the method comprises the steps: collecting multi-source interaction data to generate a multi-dimensional dynamic portrait, and obtaining an initial service strategy based on preset strategy set matching, and generating a to-be-executed service strategy in combination with an action combination output by the value network, constructing abnormal scene test data by using the generative adversarial network to complete verification and adjustment, obtaining a final service strategy, recording portrait change data and service indexes by executing the final service strategy, and feeding back the portrait change data and the service indexes to the value network and the portrait construction model. Through data-driven dynamic portrait modeling and strategy closed-loop optimization, the defects that in the prior art, a static rule is relied on, and different requirements of clients are difficult to adapt are overcome, dynamic evolution of the client service strategy is achieved, the real-time performance of the service and the resource configuration efficiency are improved, and the client experience is remarkably improved.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Remote computing power dynamic collaborative optimization method based on multi-agent reinforcement learning

The invention discloses a remote computing power dynamic collaborative optimization method based on multi-agent reinforcement learning. The method comprises the following steps: constructing a resource topological graph; obtaining a node-level state feature vector and a system-level state feature vector based on the resource topological graph; constructing a multi-agent environment; inputting local observation, global information and an agent action set into an improved CTDE model, outputting strategy network parameters and value network parameters, and constructing a training batch; obtaining a converged strategy network parameter and a converged value network parameter based on the training batch; obtaining an execution result; dynamically updated strategy network parameters and dynamically updated value network parameters are obtained, and dynamic collaborative optimization of task acceptance, resource allocation, task migration, copy start and stop and bandwidth ratio is achieved.
Owner:WUHU BIG DATA CONSTRUCTION INVESTMENT & OPERATION CO LTD

Multi-AUV safe path planning method based on deep reinforcement learning

The invention discloses a multi-AUV safe path planning method based on deep reinforcement learning. According to the method, a deep reinforcement learning algorithm model based on an MATD3 method is adopted, a policy network and a value network are updated by adopting security constraints, the security of policy learning is improved, the expected reward revenue is maximized under the condition that the expected security cost constraints are met, a frequent minimum and maximum optimization process is avoided by adopting first-order penalty optimization, and the security of policy learning is improved. Meanwhile, safety correction based on a safety layer is added in the training process to guarantee safety in the early stage of training, and exploratory and safety balance is brought to strategy optimization of reinforcement learning through safety constraint and safety correction. According to the multi-AUV path planning method, the time cooperation constraint and the space cooperation constraint of multi-AUV path planning are comprehensively considered, the centralized training and decentralized decision-making architecture is applied to multi-AUV path planning, the path planning method capable of ensuring cooperation safety is provided for a multi-AUV system, and the safety and the reliability of path planning are improved.
Owner:HARBIN ENG UNIV