Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

204 results about "Value network" patented technology

A value network is a business analysis perspective that describes social and technical resources within and between businesses. The nodes in a value network represent people (or roles). The nodes are connected by interactions that represent tangible and intangible deliverables. These deliverables take the form of knowledge or other intangibles and/or financial value. Value networks exhibit interdependence. They account for the overall worth of products and services. Companies have both internal and external value networks.

Robot motion control model training method, device and equipment based on deep reinforcement learning, robot and medium

The invention provides a robot motion control model training method, device and equipment based on deep reinforcement learning, a robot and a medium, and relates to the technical field of robots. The method comprises the following steps: acquiring a first potential vector obtained after a student encoder encodes robot body observation data, and a second potential vector obtained after a teacher encoder encodes privilege observation data; based on the current training step number and a preset probability function, calculating a sampling probability for controlling a fusion proportion of the first potential vector and the second potential vector; fusing the first potential vector and the second potential vector based on the sampling probability to generate a third potential vector, and inputting the third potential vector into a strategy network; and updating the parameters of the policy network based on the value estimation of the current state output by the value network and the action policy output by the policy network. According to the method, updating oscillation caused by sudden change of input distribution in the training process of the strategy network can be avoided, the training efficiency is improved, and the training cost is reduced.
Owner:SHENZHEN ZHUJI POWER TECH CO LTD

Flow field video generation method based on policy value architecture and online physical exploration

The invention discloses a flow field video generation method based on a policy value architecture and online physical exploration, and belongs to the technical field of crossing of artificial intelligence and computational fluid dynamics (CFD), and the method comprises the following steps: step 1, constructing an unsteady flow field multi-modal training data set, step 2, constructing a generative network system based on an Actor-Critic architecture, step 3, constructing an unsteady flow field multi-modal training data set, and step 4, constructing an unsteady flow field multi-modal training data set. Step 4, supervised fine tuning training is carried out in the first stage; step 5, online physical exploration of a generator is carried out in the second stage; step 6, feedback co-evolution of a physical encoder is carried out in the third stage; and step 7, reasoning generation of an unsteady flow field video is carried out. According to the method, a reinforcement learning architecture containing an Actor and a Critic is constructed, a physical equation is packaged into a digital environment, and a training strategy of basic supervision fine tuning, online physical exploration of a generator and coevolution feedback of an encoder is adopted.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Multi-agent combat mission cooperation method of structure entropy guided graph neural network

The invention discloses a multi-agent combat task cooperation method for a structure entropy guided graph neural network, and the method comprises the steps: S10, each combat agent interacts with an environment according to an action generated by a strategy network, the environment comprises environment information, task parameters and a preset task target, and the strategy of each combat agent is completely executed in a decentralized manner; collecting complete empirical trajectory data; s20, using the collected data for centralized training; performing value evaluation on the global state of each time step by using a value network; s30, calculating strategy loss and value loss by using a multi-agent near-end strategy optimization algorithm in combination with the output of the strategy network and the value estimation of the output of the value network; updating parameters of the strategy network and the value network by using a gradient descent method; and S40, performing loop iteration. The problems that in a traditional method, the battlefield game dynamic structure sensing ability is insufficient, the hierarchical strategy learning and generalization ability is limited, the adaptability of a model in a small sample area is poor, and the migration efficiency is low are solved.
Owner:BEIHANG UNIV

Port container automatic scheduling method based on multi-agent reinforcement learning

The invention discloses a port container automatic scheduling method based on multi-agent reinforcement learning, and the method comprises the steps: S1, building a corresponding relation between equipment and agents, and constructing a task set; s2, collecting operation state data, and constructing global and local state vectors; s3, generating a scheduling constraint vector, and cutting actions according to the resource, storage yard and path state to form a feasible action set; s4, on the basis of an improved QPLEX algorithm, constructing an individual value network containing a dump structure, and calculating an individual action value; s5, constructing a joint action value hybrid network, and mixing individual values according to the global state vector to form joint action values; s6, constructing a training sample, differentiating and aggregating instant and delayed return, and updating network parameters; and S7, during online scheduling, selecting an optimal action combination according to the combined action value, and generating and issuing a scheduling instruction. According to the invention, automatic collaborative scheduling of port container operation is realized.
Owner:安徽海润信息技术有限公司

Medical image co-processing system and method based on block chain value network

The invention relates to the technical field of medical information and block chain crossing, in particular to a medical image co-processing system and method based on a block chain value network, and the system comprises a block chain network, a privacy calculation gateway, an intelligent contract module, a hybrid storage module and a communication module. The intelligent contract module is provided with a collaborative annotation sub-module, a verifiable federated learning sub-module and a dynamic value distribution sub-module; the method comprises the steps of data preprocessing and encryption uploading, collaborative labeling and quality verification, verifiable federated learning model training, dynamic value distribution, consensus evidence storage and supervision auditing. Medical data islands are broken through by means of the block chain network, data availability and invisibility are achieved in combination with the privacy computing gateway multi-encryption technology, the contradiction between data sharing and privacy security is solved, accurate contribution and return matching is achieved through the multi-dimensional contribution degree evaluation model and the intelligent contract, and the accuracy of data sharing is improved. The traceability and the compliance are ensured by full-process operation logs and data uplink evidence storage.
Owner:AFFILIATED HOSPITAL OF JIANGNAN UNIV +1

Multi-agent reinforcement learning regional energy collaborative scheduling method and system

The invention provides a multi-agent reinforcement learning regional energy collaborative scheduling method and system, and belongs to the field of regional energy system scheduling. Coupling degrees and a coupling degree matrix between agents are constructed; inputting the observation vector into a strategy network to obtain a decision action; individual basic rewards and system economic rewards are calculated, and constraint reference rewards are constructed; individual differentiation basic rewards are calculated, and rewards are distributed; inputting the decision action into a physical quantity prediction network, calculating a physical consistency reward, obtaining a final reward and a global reward, and calculating a target return; splicing observation vectors and decision actions of all agents, splicing global joint observation vectors and joint action vectors, inputting the spliced vectors into a value network, and training; inputting the local state set into the trained strategy network, outputting a scheduling instruction, inputting the scheduling instruction into the trained value network, and outputting an evaluation result; the problems of depiction rigidness of an intelligent agent coupling relation, lack of a cooperative benefit distribution mechanism and insufficient decision physical consistency are solved.
Owner:国网安徽省电力有限公司营销服务中心 +1

Geofence-based real-time coexisting user social system and method

ActiveCN121418761ALocation information based serviceCredit systemData set
The invention belongs to the technical field of geo-fences and real-time social contact, and discloses a real-time coexistence user social contact system and method based on geo-fences, and the method comprises the steps: collecting the multi-dimensional multi-modal original data of a user through a dynamic sampling strategy and an energy efficiency balance mechanism, and generating an information-desensitized four-dimensional data set; constructing four layers of fences, evaluating the value attribute of each fence, carrying out life cycle management, dynamically adjusting the fence state, and forming a four-dimensional fence value network; identifying a basic event, a composite event and a cooperative event in a grading manner, constructing a real-time value flow through event value quantification, and generating a dynamic rule set adaptive to a scene; user permissions are distributed through the three-dimensional value permission matrix, the multi-dimensional value point system is linked to complete point generation and transaction, and a value circulation record including permission states, point changes and transaction records is formed; the value health degree is analyzed through a real-time, short-term and long-term three-level optimization mechanism, optimization parameters are generated, and the optimization parameters are reversely iterated to corresponding preorder links.
Owner:SHANGHAI XIANGYUE JIANGFENG DIGITAL TECHNOLOGY CO LTD

Water-air cross-medium unmanned aerial vehicle control modeling method based on reinforcement learning and related device

The invention discloses a reinforcement learning-based control modeling method for a water-air cross-medium unmanned aerial vehicle and a related device, and the method comprises the steps: carrying out the interaction of a strategy network and a simulation environment, and generating a series of trajectory data; predicting a state value according to the trajectory data through a value network, and calculating a corresponding dominant function and a target value function based on the state value; calculating a value loss item, a strategy loss item and a strategy regular item through a strategy optimization algorithm so as to construct a total loss function, and performing joint training update on the strategy network and the value network by taking minimization of the total loss function as a target to obtain the updated strategy network and the updated value network; and judging whether the updated strategy network and the updated value network meet a preset dual convergence condition, if so, taking the updated strategy network and the updated value network as a water-air cross-medium unmanned aerial vehicle control model, and guiding the unmanned aerial vehicle to realize smooth and stable autonomous control in air, underwater and transition stages. And the intelligent autonomous control level of the cross-medium unmanned aerial vehicle is improved.
Owner:GUANGDONG UNIV OF TECH

Urban power grid coordination control method and system based on deep reinforcement learning

The invention discloses an urban power grid coordination control method and system based on deep reinforcement learning, and solves the technical problem that the current urban power grid coordination control method causes the limitation of the urban power grid power flow adjustment precision. The method comprises the following steps: after acquiring an expert experience data set, normalizing an urban power grid state vector at the current moment to obtain a normalized vector; and pre-training the initial strategy network by minimizing a loss function according to the vector and the expert experience data set to obtain a pre-trained strategy network, and calculating and outputting the active power of the unified power flow controller and the reactive power of the reactive compensation equipment. And in combination with the power parameters and the normalized vector, determining a normalized future state vector and a current reward value by using a power flow equation and a multi-target reward function, optimizing the pre-training network by using a PPO target function and a value network to obtain a target strategy network, and outputting a target action vector as a coordination control scheme.
Owner:ZHUHAI POWER SUPPLY BUREAU GUANGDONG POWER GIRD CO

Battlefield multi-agent dynamic cooperation method based on attention mechanism

The invention discloses a battlefield multi-agent dynamic cooperation method based on an attention mechanism. The method comprises the following steps: S10, acquiring local observation information; s20, establishing a strategy network and a value network, embedding a multi-head attention mechanism in each of the strategy network and the value network, allowing each agent to dynamically query the internal representation of a teammate, and converting a search problem of an exponential level joint action space into dynamic attention on key information; s30, the strategy network generates action probability distribution of the intelligent agent through the attention layer, and the intelligent agent is forced to depend on a cooperation signal; the value network outputs the value estimation of each agent and provides a reference function; the collaborative loss function calculates similarity penalty between agents based on the embedded representation output by the attention layer; and S40, the training process is updated by adopting a strategy gradient method, strategy network parameters are updated to maximize accumulated rewards, value network parameters are updated by minimizing total loss, and a highly cooperative joint strategy is output. According to the invention, the agent is ensured to always integrate team information during decision making.
Owner:BEIHANG UNIV

Remote computing power dynamic collaborative optimization method based on multi-agent reinforcement learning

The invention discloses a remote computing power dynamic collaborative optimization method based on multi-agent reinforcement learning. The method comprises the following steps: constructing a resource topological graph; obtaining a node-level state feature vector and a system-level state feature vector based on the resource topological graph; constructing a multi-agent environment; inputting local observation, global information and an agent action set into an improved CTDE model, outputting strategy network parameters and value network parameters, and constructing a training batch; obtaining a converged strategy network parameter and a converged value network parameter based on the training batch; obtaining an execution result; dynamically updated strategy network parameters and dynamically updated value network parameters are obtained, and dynamic collaborative optimization of task acceptance, resource allocation, task migration, copy start and stop and bandwidth ratio is achieved.
Owner:WUHU BIG DATA CONSTRUCTION INVESTMENT & OPERATION CO LTD

Multi-AUV safe path planning method based on deep reinforcement learning

The invention discloses a multi-AUV safe path planning method based on deep reinforcement learning. According to the method, a deep reinforcement learning algorithm model based on an MATD3 method is adopted, a policy network and a value network are updated by adopting security constraints, the security of policy learning is improved, the expected reward revenue is maximized under the condition that the expected security cost constraints are met, a frequent minimum and maximum optimization process is avoided by adopting first-order penalty optimization, and the security of policy learning is improved. Meanwhile, safety correction based on a safety layer is added in the training process to guarantee safety in the early stage of training, and exploratory and safety balance is brought to strategy optimization of reinforcement learning through safety constraint and safety correction. According to the multi-AUV path planning method, the time cooperation constraint and the space cooperation constraint of multi-AUV path planning are comprehensively considered, the centralized training and decentralized decision-making architecture is applied to multi-AUV path planning, the path planning method capable of ensuring cooperation safety is provided for a multi-AUV system, and the safety and the reliability of path planning are improved.
Owner:HARBIN ENG UNIV

Three-dimensional environment multi-agent path planning method based on deep reinforcement learning

The invention relates to the technical field of path planning, in particular to a three-dimensional environment multi-agent path planning method based on deep reinforcement learning, which adopts a near-end strategy optimization algorithm and expands the algorithm to a multi-agent system, and by designing an independent strategy network and a central value network, a multi-agent path is planned. Cooperative work and barrier-free path planning of multiple agents are realized, and the problem of slow convergence caused by complex environment is solved. Moreover, a long-short-term memory network is introduced into a strategy network, and historical state information is modeled, so that an intelligent agent can better capture a time sequence dependency relationship in an environment, and the path planning capability of the intelligent agent in a complex environment is improved; an incremental collision penalty reward mechanism based on dynamic distance perception is designed in a reward function, radar ranging is simulated through a kinematic model, and an obstacle avoidance decision is optimized in real time. According to the method, the multi-agent collaborative obstacle avoidance capability of the agents in a complex three-dimensional environment is remarkably improved, and the method has relatively high efficiency and adaptability.
Owner:JIANGNAN UNIV

Robot dexterous operation method based on improved residual reinforcement learning and electronic equipment

The invention relates to the technical field of robot dexterous operation, in particular to a robot dexterous operation method and electronic equipment based on improved residual reinforcement learning, and the method comprises the steps: constructing a basic strategy and a reinforcement learning agent, the reinforcement learning agent comprises a strategy network Actor and a value network Critic, outputting a basic action through the basic strategy, and outputting the basic action through the value network Critic; and outputting a residual action through the strategy network Actor, superposing the residual action on the basic action, generating a final dexterous operation of the robot, and introducing a residual amplitude modulation function in the training process of the reinforcement learning agent to adaptively scale the amplitude of the residual action. According to the method, a training mechanism based on residual amplitude adaptive modulation, off-line residual data optimization guidance and distributed value function enhanced reward signal utilization is provided, so that rapid, stable and efficient convergence of a residual reinforcement learning algorithm is realized.
Owner:NANJING ESTON KUZHUO TECH CO LTD

Internet of vehicles cooperative unloading method based on mixed action deep reinforcement learning

The invention discloses an Internet of Vehicles cooperative unloading method based on hybrid action deep reinforcement learning, and relates to the technical field of Internet of Vehicles and mobile edge computing. In order to solve the technical problem of joint optimization of unloading target selection and resource allocation, the problem is modeled as a Markov decision process, and a mixed action soft actor-commentator (HA-SAC) algorithm is provided for solving; the core is that a multi-head strategy network outputs discrete actions and continuous actions at the same time in a decision period. A centralized training and decentralized execution architecture is adopted, and the value network receiving the global state guides the intelligent agent which only depends on the local state to make decisions to learn; compared with the prior art, the method has the advantages that the precision loss caused by action space discretization is avoided, the task completion time delay is obviously reduced, and the robustness and the adaptive decision-making capability of the system in a highly dynamic network environment are enhanced.
Owner:NANTONG UNIV

Concrete cooling water pipe dynamic flow adjusting method and device based on reinforcement learning

The invention provides a concrete cooling water pipe dynamic flow regulation method and device based on reinforcement learning, and relates to the technical field of concrete cooling water pipe dynamic flow regulation. The method comprises the following steps: reconstructing a three-dimensional temperature field based on Kriging, fusing the three-dimensional temperature field with material and environmental parameters into a state vector, inputting a near-end strategy optimized reinforcement learning model, outputting each valve opening vector by a strategy network, evaluating a state value by a value network, and performing physical constraint verification before issuing, an opening instruction realizes closed-loop flow distribution through coarse adjustment of the main valve and fine adjustment of the piezoelectric micro valve, meanwhile, flow attenuation and pressure difference are monitored in real time, and ultrasonic waves and mechanical scraper cleaning are automatically triggered to ensure long-term stable operation of the system.
Owner:HUBEI ENG CONSTR GRP THIRD CONSTR ENG CO LTD

Global task allocation and emergency scheduling method and system for multi-agent reinforcement learning

PendingCN121766708ASafe and reliable task allocationSafe and reliable dispatchBiological modelsEngineeringMulti source data
The invention discloses a global task allocation and emergency scheduling method for multi-agent reinforcement learning, and relates to the technical field of emergency command and intelligent scheduling, and the method comprises an offline pre-training stage and an online fine tuning and reasoning execution stage. An off-line pre-training stage: pre-training the policy network and the individual value network based on the expert database and the gating routing to obtain the policy network and the individual value network for deployment; and an online fine tuning and reasoning execution stage: performing coding and relation reasoning on the multi-source data to obtain multi-dimensional state features and global relation features, inputting the multi-dimensional state features and the global relation features into a multi-view feature fusion network for fusion, generating unified state representation, calling a strategy network, generating feasible masks, sampling to form a combined action, and finally outputting a global task allocation result. According to the method and the device, the defect that global optimization and real-time security are difficult to consider in the prior art is overcome, and safe and reliable task allocation and emergency scheduling are performed on heterogeneous resources in a strongly uncertain emergency scene.
Owner:HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Train handover control method based on prediction enhancement

The invention provides a train handover control method based on prediction enhancement. The method comprises the following steps: training a wireless communication quality parameter prediction model by using a time sequence data set to obtain a trained multivariable time sequence prediction model, predicting time sequence data of a time point of a future time step length, and generating state data of a future time window; the method comprises the following steps: constructing an actor-commentator framework comprising a strategy network and a value network, iteratively training a deep reinforcement learning decision model by utilizing the actor-commentator framework based on historically collected experience training data, updating parameters of the strategy network and the value network by adopting a near-end strategy optimization algorithm through minimizing a composite loss function, and obtaining the value of the strategy network and the value network. The training is converged, and a trained deep reinforcement learning decision model is obtained; and the trained deep reinforcement learning decision model is utilized to control the train to execute a handover action in actual operation. The stability of the train system is improved, and the communication service quality under high-speed movement of the train is guaranteed.
Owner:BEIJING JIAOTONG UNIV

Flexible job shop scheduling method based on generative adversarial training framework

PendingCN121882513AMathematical modelsData processing applicationsDiscriminatorNetwork decomposition
The invention discloses a flexible job shop scheduling method based on a generative adversarial training framework. The method comprises the following steps: establishing a Markov decision process model for flexible job shop scheduling, and completing the design of a state space, an action space and a reward function; collecting expert scheduling tracks through a plurality of algorithms, and carrying out data cleaning and standardization processing; constructing a state encoder based on a graph attention network, mapping a scheduling environment state into low-dimensional vector representation, and designing a hierarchical strategy network to decompose a scheduling decision task; constructing a value network to provide stable value estimation so as to accelerate a reinforcement learning process, and constructing a discriminator network to guide a strategy search direction by generating an imitation reward; and finally, hybrid training based on near-end strategy optimization and generative adversarial imitation learning is executed, and through adversarial training and strategy gradient optimization, an intelligent agent learns to obtain a high-performance scheduling strategy with expert empirical performance and environment adaptability.
Owner:GUANGDONG UNIV OF TECH +1

Optimization control method for multiple parallel devices in process industry and related devices

The invention provides an optimization control method for multiple parallel devices in the process industry and related devices, and relates to the technical field of industrial process control. Operation data of each reaction unit is acquired; encoding the operation data of the reaction unit by using a double-branch decoupling variational auto-encoder to obtain a shared feature and a private feature, and generating a unit context representation vector based on the private feature of the reaction unit; performing cross-unit feature recombination on the shared features and the private features, generating a virtual conversion sample and constructing an enhanced data set; using the unit context representation vector and the enhanced data set to train a context conditioned strategy network and a value network to obtain a trained strategy network; the operation data newly added into the reaction unit are coded and then input into the trained strategy network for decision optimization, the optimization control action of the newly added reaction unit is generated, and the control efficiency and reliability of the parallel multiple devices in the process industry are improved.
Owner:CENT SOUTH UNIV

Zero-carbon park source network load storage coordination control method based on deep reinforcement learning algorithm

The invention relates to a zero-carbon park source network load storage coordination control method based on a deep reinforcement learning algorithm, and the method comprises the following steps: S1, collecting data in a park in real time, carrying out the standardization, denoising and feature extraction of the collected data, and forming a data set; s2, modeling the environment state space to comprehensively reflect the current state of the system; s3, defining an action space for adjusting the operation state of the system, realizing real-time regulation and control and forming the action space; s4, designing a reward function to realize dynamic balance among different targets; s5, training a strategy network and a value network by adopting an entropy regularization SAC algorithm, and optimizing a control strategy; and S6, in the system operation process, continuously collecting new operation data, updating the strategy network and the value network in real time by using the online learning capability of the SAC algorithm, and dynamically adjusting the control strategy to adapt to environmental changes. According to the invention, efficient utilization of new energy, low-carbon power purchase optimization and dynamic load regulation and control are realized, and the overall operation efficiency of the park is improved.
Owner:STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +2

Radar resource intelligent scheduling method and device based on multi-agent reinforcement learning

The invention relates to a radar resource intelligent scheduling method and device based on multi-agent reinforcement learning. The method comprises the steps of obtaining a target priority evaluation matrix based on a multi-dimensional comprehensive state vector containing a self state, a target state, a time sequence evolution characteristic and agent collaborative information, a training strategy and a value network by adopting a reinforcement learning architecture of centralized training and distributed execution, and obtaining a target priority evaluation result through a greedy allocation and dynamic replacement strategy. And converting into a beam and residence time allocation scheme meeting hard constraints. And fusing observation data through a state estimation algorithm to obtain a tracking result and a performance index, performing feedback iteration optimization on the network in combination with a multi-target weighted reward function, and finally outputting a real-time scheduling result by utilizing the trained network. By adopting the method, the distribution can be dynamically adjusted under the scene that the target density fluctuates and the resources are temporarily limited, the tracking continuity is ensured, and the risk of missing tracking is greatly reduced.
Owner:NAT UNIV OF DEFENSE TECH

Personalized recommendation method and device, equipment and storage medium

The invention discloses a personalized recommendation method and device, equipment and a storage medium, and relates to the technical field of deep learning, and the method comprises the steps: carrying out the mapping of a target user feature and a target post feature through a target double-tower model, and generating a user embedding vector and a post embedding vector; constructing an initial state feature based on the user embedding vector, the post embedding vector and an interaction embedding vector corresponding to the interaction post, and determining a corresponding initial recommendation post based on the initial state feature through a target policy network; and performing value evaluation on the initial state features and the initial recommendation posts through a target value network, and determining corresponding target recommendation posts in the initial recommendation posts. The target value network and the target strategy network are combined with the interaction post to evaluate the value of the initial recommendation post of the user, so that the recommendation result is accurately matched with the instant interest of the user, the long-term preference can be considered, and the problem of content convergence of personalized recommendation is solved.
Owner:CHINA MERCHANTS BANK

A d2d user resource allocation method based on a deep reinforcement learning algorithm and a storage medium

ActiveCN116456493Bsolve decision problemsOvercome non-stationarityPower managementMathematical modelsTransmitted powerComputational model
This invention discloses a D2D user resource allocation method and storage medium based on a deep reinforcement learning algorithm, relating to the field of wireless communication technology. The method includes: constructing a wireless network model and discretizing the D2D transmit power; constructing a user signal-to-noise ratio calculation model with maximizing the communication system throughput as the optimization objective; setting a prediction policy network π, a prediction value network Q, a target policy network π′, and a target value network Q'; modeling the D2D communication environment as a Markov decision process, treating the D2D transmitter as an agent, iteratively loading the parameters of the target policy network to generate a policy that interacts with the environment, determining the state space, action space, and reward function; using the MAAC algorithm to optimize the policy for each D2D user; using a soft update method to iteratively update the parameters of the target policy network and the target value network until training is complete; and having the D2D user download the trained parameters of the target policy network and improve their policy.
Owner:WUXI UNIV

Sequentially coordinated development timing decision method for remote wind-solar energy and flexible resources

PendingCN122453215ADecision modelNew energy
The application discloses a remote scenery new energy and flexible resource sequential cooperative development timing decision method, including the following steps: inputting state information of a current to-be-decided project into a trained timing decision model to obtain a project development timing scheme formed by a project construction action sequence; the timing decision model includes multiple intelligent agents and an alpine deduction environment, the multiple intelligent agents include a distributed strategy network and a centralized value network, and dynamic timing complementation and global collaborative optimization of a heterogeneous resource cluster in a stage production mode in a harsh environment are realized.
Owner:HUNAN UNIV

DDDPG-based autonomous path planning and obstacle avoidance multi-target continuous control method

The invention provides a DDPG-based autonomous path planning and obstacle avoidance multi-target continuous control method, which can improve the model stability, improve the feature expression capability and is suitable for more application scenes. According to the method, the strategy network and the value network are not processed as black boxes, the structures of the strategy network and the value network are improved for specific problems occurring in a complex environment, the expression ability of multi-dimensional parameters is effectively improved, and the method is particularly suitable for application scenes of path planning and obstacle avoidance in the complex environment.
Owner:AUTOLINK INFORMATION TECHNOLOGY CO LTD

Intelligent agent strategy generation method based on reinforcement learning and non-instantaneous computer readable medium

The invention provides an agent strategy generation method based on reinforcement learning and a non-instantaneous computer readable medium, and the method comprises the steps: generating a first action for the execution of an agent according to an action network and a first state of the agent, and generating a first value according to a value network and the first state, obtaining a second state of the agent generated by the environment and a reward generated by a reward function, storing the first state, the first action, the first value, the second state and the reward to a buffer area, training a value network and an action network according to the buffer area, the regularization loss includes a first distance associated with the first action and a second distance associated with the first action or the first value.
Owner:INVENTEC PUDONG TECH CORPOARTION +1

A dual-resource flexible job shop scheduling method and system integrating large language model and deep reinforcement learning

This invention belongs to the field of intelligent scheduling technology in manufacturing, and particularly relates to a dual-resource flexible job shop scheduling method and system integrating a large language model and deep reinforcement learning. The method includes: constructing a dual-resource flexible job shop scheduling environment model, defining jobs, processes, machines, workers, and their constraints, and constructing a set of state features for reinforcement learning; designing an action space and feasible action masking mechanism, based on process sequence, feasible machine set, worker skill qualifications, and available time constraints, to mask infeasible actions and reduce the probability of selecting low-quality actions; designing a multi-objective reward function, including minimizing completion time and total energy consumption; training the scheduling strategy using the Asynchronous Advantage Actor-Critic (A3C) algorithm, optimizing the policy network and value network through multi-threaded parallel sampling and asynchronous update mechanisms; and triggering the large language model to generate new reinforcement learning definitions based on trend and stability analysis of training feedback, and achieving adaptive iterative optimization through a closed-loop mechanism.
Owner:LANZHOU UNIVERSITY OF TECHNOLOGY

Fair federal learning method based on multi-agent reinforcement learning

PendingCN121480761AMachine learningNative clientEngineering
The invention relates to the technical field of federated learning algorithms, in particular to a fair federated learning method based on multi-agent reinforcement learning, which comprises the following steps: respectively initializing network parameters of a server and a local client, and calculating an initial state vector of a local environment by the local client according to the initialized network parameters; the server calculates the initial state of the global environment, and a local client selects and executes an action; the server updates the classification network and broadcasts the classification network to all local clients, and the local clients update local states and rewards; the server updates the global environment state and stores the data in an experience batch; and the server updates the value network, and the local client updates the strategy network. The strategy network is trained based on centralized training and distributed execution, and the convergence speed of the strategy network can be greatly improved.
Owner:INFORMATION CENT OF YUNNAN POWER GRID CO LTD

Edge cloud computing-oriented computing power transaction incentive method

The invention relates to the field of edge cloud computing and intelligent resource scheduling, and discloses a computing power transaction incentive method. The method comprises the following steps: obtaining a transaction request, collecting computing power cost, user preference and other information, initializing a market state, and constructing a compact state vector which does not increase along with the scale of a user; inputting the state into a diffusion model to generate an initial pricing, constructing a reward signal by combining an adversarial value evaluation network with excitation constraint and the like, and updating model parameters through reinforcement learning to enable pricing convergence to be optimal; and releasing a final contract menu for a user to select so as to complete computing power distribution, and collecting, feeding back and updating a market state for iterative optimization after transaction. According to the method, on the basis of a reinforcement learning mechanism of a diffusion model and an adversarial value network, a pricing scheme meeting excitation constraint and income maximization can be generated under information asymmetry, and market pricing and resource utilization efficiency and dynamic adaptive capacity are improved.
Owner:HUNAN UNIV