Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

5 results about "Q learning algorithm" patented technology

Q-learning, is a simple incremental algorithm developed from the theory of dynamic programming [Ross,1983] for delayed reinforcement learning. In Q-learning, policies and the value function are represented by a two-dimensional lookup table indexed by state-action pairs. Formally, for each state and action let:

A method and system for frequency control of a thermal power plant based on reinforcement learning algorithm

ActiveCN115296306BFlicker reduction in ac networkAc network load balancingAlgorithmReinforcement learning algorithm
The application relates to a frequency control method for a thermal power plant based on a reinforcement learning algorithm, wherein the method comprises the following steps: obtaining primary frequency regulation evaluation data and current combined cycle unit working condition data to form state data; determining a state-action pair in a preset algorithm model, wherein the preset algorithm model is obtained by training a Q learning algorithm through reinforcement learning based on the state data of the thermal power plant; obtaining evaluation values of each state-action pair based on a preset estimation rule through the preset algorithm model, and obtaining a target state-action pair according to the evaluation values; generating a control instruction according to an environmental interaction action in the target state-action pair, and adjusting the output power and action time of an electric energy storage device through the control instruction. Through the application, the problem that the service life of an electric energy storage device of a combined cycle unit is low due to long-term high-load working state in the related art is solved, accurate frequency regulation control is realized, power waste of the electric energy storage device is reduced, and the service life of the electric energy storage device is improved.
Owner:HUADIAN ELECTRIC POWER SCI INST CO LTD

Method for evaluating node trust degree and interrupting adaptive primary and backup routing in robot swarm coordination

ActiveCN118381755BExpand coverageImprove route discovery probabilityNetwork topologiesTransmissionRouting decisionEngineering
The application discloses a kind of main and backup routing methods of node trust degree evaluation and interruption adaptation in robot group coordination, the method includes: source node initiates routing request broadcast, and available path from source node to destination node is obtained by adaptive broadcast forwarding strategy;Afterwards, destination node carries out routing reply according to reverse path transmission;If the routing establishment from source node to destination node is successful, then enter the routing decision module based on Q learning algorithm;If the routing decision result selects main routing transmission strategy, then destination node enters trust degree evaluation module, interruption adaptive credible path or no credible path transmission module;If the routing decision result selects backup DTN (Delay Tolerant Networks, time delay tolerance network) routing transmission strategy or routing establishment is not successful, then destination node enters Prophet routing module;Through the main and backup routing methods of node trust degree evaluation and interruption adaptation in robot group coordination provided by the application example, the robustness and robustness of data transmission are improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A federated transfer learning enhanced multi-agent inter-vehicle dynamic regulation method

The application discloses a kind of federal transfer learning enhanced multi-agent workshop dynamic regulation method, establishes the distributed flexible flow workshop dynamic scheduling model based on multi-agent system;Using federal learning, the feature extraction network with representation ability is obtained;According to workshop actual task demand and equipment characteristics, combined with the working data after feature processing, based on the feature extraction network, a multi-agent deep reinforcement learning model based on cross sampling is established;The multi-agent deep reinforcement learning model is trained using deep Q learning algorithm;Using federal transfer learning technology, collect the data of other workshops or factories and local similar tasks, and train Q network according to the data;Then, through adaptive weight fusion technology, the effective transfer of knowledge is realized;Complete multi-agent workshop dynamic regulation.The application combines the advantages of autonomous decision-making and group intelligence of multi-agent, can more flexibly allocate production tasks, and dynamically adjust production control according to real-time situation.
Owner:BEIHANG UNIV

Rolling mill screwdown control method based on q-learning algorithm

ActiveCN117960798BAlgorithmControl system
The application discloses a rolling mill screwdown control method based on a Q learning algorithm, and comprises the following steps: obtaining a target value and an output value of a rolling force of a rolling mill screwdown control system; determining PID parameters of a PID controller according to a PID parameter Q value table based on the obtained target value and output value of the rolling force, the PID parameter Q value table being obtained in advance by using a preset Q learning algorithm; obtaining an output of the PID controller according to the obtained target value and output value of the rolling force and the PID parameters; and superimposing the output of the PID controller to an input of the rolling mill screwdown control system to adjust the output value of the rolling force of the rolling mill screwdown control system. The rolling mill screwdown control method based on the Q learning algorithm can realize optimal control of target value tracking characteristics and disturbance suppression characteristics by using the Q learning algorithm to learn and obtain optimal PID parameters of the PID controller corresponding to different rolling force input and output states in advance, improve the control and adjustment precision of the rolling force of the rolling mill, and ensure the regulation and control precision of the loaded roll gap.
Owner:SHANXI TAIGANG STAINLESS STEEL PRECISION STRIP CO LTD