Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

96 results about "Q learning algorithm" patented technology

Q-learning, is a simple incremental algorithm developed from the theory of dynamic programming [Ross,1983] for delayed reinforcement learning. In Q-learning, policies and the value function are represented by a two-dimensional lookup table indexed by state-action pairs. Formally, for each state and action let:

Lightweight Internet of Things management system based on star flash protocol stack

The invention relates to the technical field of Internet of Things management, in particular to a lightweight Internet of Things management system based on a star flash protocol stack, which comprises a plurality of modules such as a star flash protocol communication module, a dynamic topology management module and a resource virtualization module. The satellite flash protocol communication module realizes low-power-consumption efficient connection of equipment; the dynamic topology management module optimizes the network topology by using an improved multi-agent Q learning algorithm; the resource virtualization module realizes accurate resource allocation through a weighted bipartite graph matching model; the safety protection module adopts an attention mechanism to detect abnormities and guarantee data safety; the edge co-processing module realizes intelligent task unloading by means of deep reinforcement learning; all the modules cooperatively work under the overall planning of the cross-module coordination controller, and the information barrier is broken. According to the invention, the problems of high communication energy consumption, poor resource allocation, weak security protection, low task processing efficiency, insufficient module collaboration and the like of a traditional Internet of Things system are effectively solved, the system communication efficiency, the resource utilization rate and the security are remarkably improved, and the task processing capability is enhanced.
Owner:FUJIAN MAIWEI INFORMATION ENG CO LTD

AI-based automatic production line scheduling system in industrial internet

The invention discloses an AI-based automatic production line scheduling system in an industrial internet, which relates to the technical field of production scheduling and comprises a production plan management module S1, a dynamic scheduling engine module S2, a resource scheduling module S3, a real-time monitoring system module S4, an exception handling center module S5 and a data optimization platform module S6. In the industrial internet, an AI-based automatic production line scheduling system, an X dynamic scheduling engine millisecond response and a multi-agent reinforcement learning engine based on a federated learning architecture realize millisecond response scheduling, each device is used as an autonomous decision-making unit, and dynamic coordination is performed through a distributed Q learning algorithm, so that the vacancy rate of the devices is greatly reduced, and the scheduling efficiency is improved. According to a long-short-term memory network deep analysis model of order delivery cycle compression, emergency order insertion response speed improvement, multi-modal AI quality monitoring, fusion of vibration, thermal imaging and current spectrum, the detection rate is greatly improved compared with a unified sensor, causal reasoning and root cause analysis are performed, a fault causal graph is constructed to position a deep problem, and the average repair time is shortened.
Owner:JIANGSU AOYILAN INTELLIGENT TECH CO LTD

Charging pile scheduling method and system

The invention provides a charging pile scheduling method and system, and belongs to the technical field of charging pile energy management, and the method comprises the steps: predicting the second state data of a charging pile group based on the historical state data of the charging pile group and the first state data of the charging pile group, taking the historical state data, the first state data and the second state data as state data of the charging pile group; wherein the first state data is the state data of the charging pile group in the current preset time period; the second state data is predicted state data of the charging pile group in a future target time period; processing the state data of the charging pile group, the demand data of the charging vehicle and the power grid load data based on an ant colony algorithm to obtain an initial charging scheduling strategy; and adjusting the initial charging scheduling strategy based on a Q learning algorithm to obtain a target charging scheduling strategy. According to the charging pile scheduling method and system provided by the invention, the charging efficiency can be improved.
Owner:NANYANG JINGUAN INTELLIGENT SWITCH CO LTD

Electromagnetic micromirror output feedback tracking control method and device based on adaptive Q learning

The invention discloses an electromagnetic micromirror output feedback tracking control method and device based on adaptive Q learning, and the method comprises the steps: constructing an electromagnetic micromirror model based on a physical structure of an electromagnetic micromirror; combining the electromagnetic micro-mirror model with a reference trajectory to construct an augmentation system so as to convert a trajectory tracking problem into a state adjustment problem; based on the augmented system, designing a Q function Bellman equation to establish a data-driven performance evaluation framework; based on the augmented system, the state of the augmented system is reconstructed through historical input and output data, so that angular velocity measurement is avoided; obtaining an optimal control input and a new Q function Bellman equation based on the reconstructed augmented system state and the Q function Bellman equation; and based on the optimal control input and the new Q function Bellman equation, solving a control problem by using an adaptive Q learning algorithm. The method breaks through the limitation that a traditional model depends on a control method, and achieves optimal tracking control only through historical input and output data.
Owner:SOUTH CHINA UNIV OF TECH

Method and system for preventing pressure sore pad based on deep Q learning and attention mechanism

The invention discloses a method and a system for preventing a pressure sore pad based on deep Q learning and an attention mechanism. The method comprises the steps that pressure distribution data and posture data of a patient are collected through a pressure sensor and a posture sensor, normalization processing is carried out on the collected data, and a standardized state data flow is generated; and constructing a contact surface pressure distribution model of the patient and the pressure sore prevention pad according to the standardized state data flow to obtain a contact surface state matrix and the like. According to the method, a self-attention mechanism and a deep neural network are adopted, a Q learning algorithm is combined, predefined air bag pressure adjustment actions are evaluated and optimized, and an optimal air bag adjustment strategy is generated. The local pressure of the pressure sore prevention pad is adjusted according to the strategy. The state of a patient can be intelligently sensed, pressure distribution of the pressure sore prevention pad can be adjusted in a self-adaptive mode, pressure sores are effectively prevented, and nursing quality and patient comfort are improved.
Owner:NAT REHABILITATION ASSISTIVE DEVICES RES CENT

Low-delay video stream code rate adaptation method and system based on offline element reinforcement learning

The invention provides a low-delay video stream code rate adaptation method and system based on off-line element reinforcement learning. The method comprises the following steps: constructing an off-line expert track data set; constructing a code rate adaptation strategy network, and outputting a downloading code rate selection result and a network throughput correction value by taking a video stream transmission state as input; and adopting a meta-implicit Q learning algorithm to perform double-layer alternate optimization of an inner layer and an outer layer on the policy network, the inner layer updating policy parameters of the policy network based on the offline expert trajectory data set, and the outer layer further adjusting the policy network parameters through meta reinforcement learning to improve the adaptability of the policy network to unseen tasks. According to the method, multiple expert algorithm experiences are integrated through offline reinforcement learning, high generalization and adaptive ability can be realized in multiple target time delays and complex network environments without online trial and error exploration, and finally, high-quality and smooth playing of video streams in a low-time-delay live broadcast scene is supported.
Owner:SHANGHAI JIAOTONG UNIV

Multi-agent system double-layer Q learning control method and system based on Warisstein distance

The invention discloses a multi-agent system double-layer Q learning control method and system based on Warisstein distance. The method comprises the following steps: establishing a multi-agent system state space model under fault and unknown distribution interference; constructing an inner-layer fault-tolerant control system and an outer-layer robust control system; designing an internal fault-tolerant control gain based on a Q learning algorithm; designing an outer layer robust control gain and a bias item based on the Warisstein distance and a zero-sum game framework; and designing a distributed consistency protocol in combination with neighborhood state information. According to the method, under the conditions that a system model is unknown, an executor has additive time-varying faults, external disturbance probability distribution is uncertain and only depends on limited samples, the influence of compensation faults and disturbance can be effectively reduced, the H infinity performance constraint is met, meanwhile, asymptotic state synchronization of all agents is achieved, and system stability is ensured.
Owner:NANJING UNIV OF SCI & TECH

Data center global temperature control optimization method and system based on hybrid reinforcement learning

The invention discloses a data center global temperature control optimization method based on online-offline hybrid reinforcement learning, and the method comprises the steps: constructing an interpretable global state space fusing physical information, and guaranteeing that a state variable has a clear physical meaning; a combined action space is defined, and cooperative control of the cold source side and the tail end side is achieved; designing a multi-objective reward function based on physical model driving, and comprehensively considering energy consumption, temperature stability, carbon emission and physical optimization indexes; obtaining a basic security policy through offline pre-training, and extracting reliable behaviors from historical data by using a conservative Q learning algorithm; strategy optimization is achieved through online safety fine adjustment, and gradual adjustment is conducted under multiple constraints to adapt to real-time changes; and finally, deploying an optimization strategy to realize closed-loop control of the system, and establishing a continuous learning mechanism to cope with long-term drift. According to the method, a complete offline-online mixed learning system is established, and the energy efficiency of the data center temperature control system is improved.
Owner:SOUTHEAST UNIV

Engine rotating speed compensation method for hybrid power excavator and hybrid power excavator

According to the method, a fuzzy subset is set according to the required torque, the current SOC value and the difference value between the engine rotating speed and the motor rotating speed, and the torque of an ISG motor is dynamically adjusted through a fuzzy control algorithm. The fuzzy control adopts a fixed fuzzy rule, so that a Q table is established through a Q learning algorithm, the compensation torque of the ISG motor is obtained, the final output torque of the ISG motor is determined to be the sum of the output torque of the fuzzy controller and the compensation torque output by the Q learning algorithm, and the torque output by the ISG motor is determined. According to the method, the advantage of high response speed of the torque of the motor is utilized, the difference value between the required torque and the actual torque is effectively compensated, the rotating speed fluctuation of the engine caused by overlarge change of the torque of the engine is avoided, and the fuel economy and the driving operation stability are improved.
Owner:JIANGSU XCMG CONSTRUCTION MACHINERY RESEARCH INSTITUTE LTD

Microgrid dispatching control method, device and equipment based on strategy optimization and medium

The invention provides a micro-grid dispatching control method, device and equipment based on strategy optimization and a medium, and belongs to the technical field of power grid dispatching, and the method comprises the steps: collecting the real-time operation data of new energy power generation equipment, power utilization load equipment and an energy storage system; receiving power generation and power utilization plans of a user, extracting multi-resource dynamic response characteristics through an LSTM model in combination with the real-time operation data and the historical operation data, and outputting an adjustment capability evaluation matrix; receiving a power grid dispatching plan, matching an optimal operation mode from the knowledge base based on the adjustment capability evaluation matrix, and generating a dispatching instruction set; issuing a scheduling instruction set to the target device and executing the scheduling instruction set, collecting execution data in real time and calculating operation deviation; when the operation deviation exceeds a threshold value, an online Q learning algorithm is adopted to update the instruction weight, and an optimization strategy is fed back to the knowledge base. Through real-time data acquisition and an intelligent optimization algorithm, fine scheduling of the micro-grid is realized, the resource utilization rate and the operation flexibility are improved, and the operation cost is reduced.
Owner:INSPUR ARTIFICIAL INTELLIGENCE RES INST CO LTD SHANDONG CHINA

High-order robust control method for mobile robot in combination with weighted double-Q learning algorithm

The invention provides a high-order robust control method for a mobile robot in combination with a weighted double-Q learning algorithm. The high-order robust control method comprises the following steps: step 1, constructing an intelligent trolley motion scene including time-varying friction, variable gradient, random external disturbance and sensor noise complex uncertainty; step 2, designing a weighted double-Q learning adaptive nonlinear expansion state observer; step 3, based on the step 2, constructing a reward function including a trajectory tracking error, a disturbance estimation error and a stability constraint; 4, designing an improved super-spiral sliding mode control law based on the weighted double-Q learning adaptive nonlinear extended state observer; and 5, setting a plurality of groups of working condition simulation comparison schemes, and carrying out simulation experiment verification. According to the invention, real-time adaptive control of the mobile robot on different working conditions is realized, the trajectory tracking precision of the system is improved, and sliding-mode control buffeting is inhibited at the same time.
Owner:NANTONG UNIV

Multimodal sequence recommendation method and system based on hyperbolic curvature adaptive learning

The invention discloses a multi-modal sequence recommendation method and system based on hyperbolic curvature adaptive learning, and the method comprises the steps: exploring the adaptive hyperbolic curvature of a time sequence behavior diagram in different modals based on a reinforcement learning frame through a curvature adaptive learning agent; modeling the multi-modal interest of the user in a hyperbolic space with a specific curvature through a multi-modal interest learning agent; a Nash Q learning algorithm is adopted to cooperatively update a curvature adaptive learning agent and a multi-mode interest learning agent, and when the two agents reach Nash equilibrium, adaptive curvatures in different modes are explored; and learning a user behavior interaction mode under the image mode and the text mode in the hyperbolic space, predicting the score of the user for the next multi-mode item, and carrying out multi-mode item recommendation based on the predicted score. According to the method, the hyperbolic space is used as the embedding space, and the user preference information is mined by using the text and image modal information of the project, so that the overall recommendation performance is improved.
Owner:ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY

Multi-robot collaborative operation and task distribution system and method

The invention discloses a multi-robot collaborative operation and task allocation system and method, relates to the technical field of robot operation, and is used for solving the problems that in a traditional method, task allocation efficiency is low, resources are wasted, especially in a multi-robot system, tasks cannot be accurately allocated, dynamic adaptability is lacked, task and environment changes cannot be coped, and the task allocation efficiency is low. An optimal decision strategy is obtained step by step; and a single learning algorithm has defects. An auction algorithm is used for solving the task allocation problem of the robot, the work progress and the space matching degree of the robot are considered, a Q learning algorithm is used for determining whether to switch to a Sarsa learning algorithm or not according to the task completion condition of the robot in combination with environment interaction and a continuous optimization strategy, and the task allocation efficiency is improved. The design of the flexible switching strategy makes up for the defects of single reinforcement learning strategy and lack of flexibility and adaptability in the prior art, so that the robot can dynamically adjust the optimization algorithm according to the actual situation, and the task completion quality and efficiency are improved.
Owner:SHENZHEN WANZHONG COM TECHNOLOGY CO LTD

Multi-objective optimization method and system for heating and ventilation system based on reinforcement learning

The invention discloses a multi-objective optimization method and system for a heating and ventilation system based on reinforcement learning, and relates to the technical field of building energy system optimizing.The method comprises the steps that a state space is constructed based on collected environment parameters, equipment state parameters and historical operation data of the heating and ventilation system; the state space is composed of state vectors generated after normalization and discretization; constructing an action space based on controllable equipment of the heating and ventilation system; the action space is generated by discretizing the operation mode and control parameters of the controllable equipment; constructing a multi-target reward function to evaluate a control effect based on the state vector and the action space; taking the state vector as input, and performing iterative training by using a Double Q learning algorithm until a trained control model is obtained; and inputting the currently acquired state vector into the trained control model, and outputting an optimal action to generate a control instruction for the heating and ventilation system.
Owner:GUANGZHOU SHANGHANG INFORMATION TECH CO LTD

Self-adaptive multi-path scheduling method based on deep reinforcement learning in heterogeneous network

The invention discloses a self-adaptive multi-path scheduling method based on deep reinforcement learning in a heterogeneous network, which comprises the following steps: in a link management module, introducing a multi-path transmission gain model for quantifying a trade-off relationship between aggregation bandwidth improvement and end-to-end delay increase of a plurality of links, by screening the link sets with better performance and smaller difference, the negative influence of low-performance links on the transmission performance of the system is effectively reduced; in the deep reinforcement learning module, the adaptability, learning efficiency and decision accuracy of the model to an unstable environment are improved by combining priority experience playback and an asynchronous learning mechanism, and optimal balance between network throughput and reordering delay is realized by dynamically adjusting a data packet distribution strategy and a transmission sequence. According to the method, the asynchronous priority replay double-depth Q learning algorithm is adopted, real-time perception of the link state and decision scheduling of the data packet can be realized in a dynamic network environment, the aggregation throughput of the network is remarkably improved, and the transmission delay is effectively reduced.
Owner:NANJING UNIV OF POSTS & TELECOMM

A model-free control method for human-in-the-loop multi-agent systems under denial-of-service attacks

The present invention provides a model-free control method for a human-in-the-loop multi-agent system under a denial of service attack, belonging to the field of multi-agent system control technology. This control method is based on human-in-the-loop technology and introduces external human experts to supervise the system to improve the security and reliability of the system; at the same time, the DoS attack between the intelligent agents is modeled as a time-varying switching topology, and a fully distributed designated time leader output observer is designed, which effectively reduces the adverse effects of the DoS attack jump on the observation effect and achieves the designated time convergence performance of the observation error; in addition, an optimization performance function is designed and a model-free Q learning algorithm is derived. This control method realizes model-free optimization control of a human-in-the-loop multi-agent system under the influence of DoS attacks, while taking into account the designated time convergence performance of the consistency error, effectively improving the control quality.
Owner:INST OF ELECTRONICS & INFORMATION ENG OF UESTC IN GUANGDONG

Self-learning job shop scheduling method meeting waiting time constraint

The invention belongs to the technical field of job shop scheduling, and particularly relates to a self-learning job shop scheduling method meeting waiting time constraint, which comprises the following steps: S0, constructing a job shop scheduling problem model with waiting time constraint; the method comprises the following steps: S1, acquiring job shop scheduling problem data, a configuration algorithm and operation parameters; s2, constructing chromosome individuals, generating chromosomes and initializing a population; s3, calculating the fitness value of each chromosome individual; s4, forming a new generation of population; s5, combining the fitness information of the current population, dynamically selecting a crossover rate Pc through a Q learning algorithm, and performing crossover operation on the population; s6, dynamically determining a mutation rate Pm in the same parameter combination space by using a Q learning algorithm, and performing mutation operation on the crossover progeny to generate an updated population; and S7, judging whether an iteration termination condition is met or not. According to the method, a feasible and near-optimal scheduling scheme can be efficiently generated on the premise of ensuring that the inter-process waiting time constraint is met.
Owner:CHONGQING CITY MANAGEMENT COLLEGE

A micro-grid scheduling control method and device based on policy optimization, equipment and medium

The application provides a micro-grid scheduling control method and device based on policy optimization, equipment and medium, belonging to the power grid scheduling technical field, the method: collecting real-time operation data of new energy power generation equipment, power load equipment and energy storage system; receive user power generation and power consumption plan, combine real-time operation data and historical operation data, extract multi-resource dynamic response characteristics through LSTM model, output adjustment capacity evaluation matrix; receive power grid scheduling plan, based on the adjustment capacity evaluation matrix, match the optimal operation mode from the knowledge base, generate the scheduling instruction set; issue the scheduling instruction set to the target equipment and execute, collect the execution data in real time and calculate the operation deviation; when the operation deviation is over the threshold, the instruction weight is updated by using the online Q learning algorithm, and the optimization strategy is fed back to the knowledge base. The application realizes fine scheduling of micro-grid through real-time data collection and intelligent optimization algorithm, improves resource utilization and operation flexibility, and reduces operation cost.
Owner:INSPUR ARTIFICIAL INTELLIGENCE RES INST CO LTD SHANDONG CHINA

Hybrid multi-target coordination control system based on SVG and synchronous phase modifier

The invention relates to the technical field of power systems, and discloses a hybrid multi-target coordination control system based on SVG and a synchronous phase modifier, which significantly improves the decision quality and foresight of an intelligent agent by fusing a real-time state and future prediction, and improves the decision efficiency compared with a single system state vector. The enhanced state characterization not only comprises real-time parameters such as voltage deviation and reactive power difference, but also integrates future uncertainty probability information predicted by the random forest, so that the intelligent agent can simultaneously sense the current state and the expected evolution trend of the system, and the enhanced environment sensing capability enables the Q learning algorithm to make a more predictable decision, so that the method is more suitable for popularization and application. According to the method, the coordination strategy of the SVG and the phase modifier is adjusted in advance before voltage fluctuation occurs, state characterization is enhanced, comprehensive environment information is provided, an intelligent agent can better balance the relation between quick response and steady state support, the voltage stability is guaranteed, the economical efficiency of system operation is improved, and finally multi-target collaborative optimization in the true sense is achieved.
Owner:INNER MONGOLIA ELECTRIC POWER (GRP) CO LTD WUHAI UHV POWER SUPPLY BRANCH

A workflow scheduling method considering security risk and cost optimization in cloud environment

The application provides a workflow scheduling method considering security risks and cost optimization in a cloud environment, which reduces the overall security risks and execution costs of the workflow under the given deadline constraint of a user. Firstly, the cloud computing resources are defined, the workflow structure is abstracted, and a security risk calculation method is proposed. Then, a reinforcement learning model oriented to security risks and cost optimization is constructed. Next, an action selection algorithm in the learning process is determined. Finally, iterative learning is carried out based on the Q learning algorithm idea, and finally the task scheduling scheme is determined, and the overall security risks and total costs of the workflow are obtained.
Owner:NANJING UNIV OF POSTS & TELECOMM

False alarm suppression method for mobile satellite terminal direction finding based on multi-dimensional feature joint screening

PendingCN122640007ATime domainFeature Dimension
The present application belongs to the technical field of wireless communication signal processing, and particularly relates to a mobile satellite terminal direction finding false alarm suppression method based on multi-dimensional feature joint screening, aiming to solve the technical problems in the prior art that a single power dimension is used for signal judgment, the uplink signal and the downlink signal on the mobile satellite system cannot be distinguished, and the judgment threshold is fixed and cannot adapt to complex electromagnetic environments. A multi-dimensional signal feature library is constructed; front-end signal reception and preprocessing are performed; multi-dimensional features are extracted in parallel, and the confidence levels in the time domain, the frequency domain, the modulation domain and the feature code domain are calculated based on the multi-dimensional signal feature library; adaptive weighted fusion is performed based on information entropy, the weight calculation and fusion score of each feature dimension are automatically adjusted according to the real-time electromagnetic environment complexity; and a dynamic threshold adaptive mechanism based on the Q learning algorithm is used to realize closed-loop optimization and feedback adjustment of the judgment threshold, realize multi-dimensional joint screening of signals, and effectively suppress direction finding false alarms in combination with the accurate distinguishing mechanism of uplink and downlink signals.
Owner:CHENGDU GUDE TECH CO LTD

A power transmission network outage maintenance scheduling optimization method based on Q learning

The application provides a power grid outage maintenance scheduling optimization method based on Q learning, which comprises the following steps: generating a Gantt chart by acquiring an original power grid topology file, and establishing a state judgment rule and a target function of reinforcement learning; integrating maintenance plans of subordinate units to the Gantt chart to form an initial maintenance plan; calculating an initial power flow collection by using a built-in power flow solver of Matlab, and evaluating a power grid operation state; in the agent training process, an epsilon-greedy method is used to select a maintenance action, the Gantt chart and the topology file are updated, and the Q value is updated through a Bellman equation, so that the maintenance plan is optimized; and the Gantt chart of a final outage maintenance plan is output after the training is completed. The power grid scheduling problem is abstracted as a Q learning problem, the outage scheduling is optimized by using a Q learning algorithm through reasonable constraints and a target function, the work intensity of dispatchers is reduced, the talent training cost is reduced, the scheduling efficiency and reliability are improved, and the key technical problems of power grid scheduling are solved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1

Power distribution network planning method and device based on Nash Q learning, equipment and medium

The invention is applied to the field of power distribution network planning, and discloses a power distribution network planning method and device based on Nash Q learning, equipment and a medium. The method comprises the following steps: acquiring a real-time state corresponding to a to-be-planned power distribution network, topological parameters including the number of nodes and the number of lines of the to-be-planned power distribution network, and resource parameters and policy parameters including a parameter set of new energy candidate nodes and energy storage candidate nodes; constructing a multi-target planning model according to the topological parameters and the policy parameters; constructing a multi-target planning model according to the multi-target planning model based on a Nash Q learning algorithm; initializing a multi-main-body Q matrix, and updating the multi-main-body Q matrix based on the real-time state, a preset exploration strategy and a multi-target planning model; and determining a planning strategy of the to-be-planned power distribution network based on the multi-main-body Q matrix. In this way, the global optimality and practicability of the planning strategy are improved by combining the three-dimensional coupling decision and the Nash Q learning with the multi-agent game.
Owner:STATE GRID SICHUAN ECONOMIC RES INST

Adaptive model predictive control method based on incremental double q-learning

The application belongs to the field of intelligent control, and particularly relates to an adaptive model predictive control method based on incremental double Q learning, which comprises the following steps: 1) establishing a kinematic model of a non-complete constraint differential mobile robot, and combining the kinematic model to construct a trajectory tracking error model; 2) based on the trajectory tracking error model, constructing a prediction model of a trajectory tracking model predictive controller of the mobile robot, and realizing high-performance trajectory tracking control according to the prediction model; 3) performing incremental discretization operation on the action space and the state space of the double Q learning agent; 4) through real-time interaction between the agent and the environment and online iterative learning, the state and the action space of the double Q learning agent are continuously optimized, and an online adaptive MPC controller is obtained to realize high-precision autonomous trajectory tracking control of the mobile robot. The application realizes online optimization of a weight matrix of the MPC by using an incremental double Q learning algorithm, and adjusts the MPC weight matrix in real time according to the system state to optimize the total reward.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Production line intelligent clamp control method and system

The invention relates to the field of production line intelligent clamp control, and discloses a production line intelligent clamp control method and system, and the method comprises the steps: collecting the 3D point cloud data of a workpiece, calculating the geometric features and shape errors of the workpiece through a point cloud processing algorithm, predicting and optimizing the clamping force, and optimizing the clamp task switching through combining PID control and a Q learning algorithm; compared with the prior art that the traditional fixture control method mostly depends on manual adjustment of the position and the clamping force of the fixture, the requirements of high precision and rapid adaptation are difficult to meet when the requirements of complex workpiece shapes and diversified operation are met, and particularly, the requirements of high precision and rapid adaptation are not met under the conditions that the shapes of workpieces are irregular and the surfaces have large errors. In order to solve the technical problem that the production efficiency is low due to the fact that a traditional method cannot efficiently adjust a clamp, clamping points are accurately selected through visual detection, workpiece form changes are automatically adapted, and the production efficiency of a production line and the workpiece clamping precision are remarkably improved.
Owner:JIANGSU YUXIN SENSOR TECH CO LTD

Multi-agent system double-layer q learning control method and system based on wasserstein distance

The application discloses a multi-agent system double-layer Q learning control method and system based on a Wasserstein distance, which comprises the following steps: establishing a multi-agent system state space model under faults and unknown distribution disturbances; constructing an inner fault-tolerant control system and an outer robust control system; designing internal fault-tolerant control gains based on a Q learning algorithm; designing outer robust control gains and bias terms based on a Wasserstein distance and a zero-sum game framework; and designing a distributed consistency protocol in combination with neighborhood state information. The method can effectively reduce the influence of compensation faults and disturbances under the condition that a system model is unknown, an actuator has additive time-varying faults, external disturbance probability distribution is uncertain, and only limited samples are available, and can realize asymptotic state synchronization of all agents while meeting H-infinity performance constraints, thereby ensuring system stability.
Owner:NANJING UNIV OF SCI & TECH

Photovoltaic power station generation power prediction method based on Q learning combination model

The invention relates to the technical field of power systems and intelligent power grids, and discloses a photovoltaic power station generation power prediction method based on a Q learning combination model, and the method comprises the steps: collecting meteorological data, power station operation data and time sequence data, carrying out the preprocessing of the collected original data, and screening out the stronger correlation characteristics; a photovoltaic power prediction model based on an XGBoost model is established according to the processed data, a photovoltaic power prediction model based on an LSTM model is established according to the processed data, and after preliminary prediction results of the XGBoost model and the LSTM model are obtained respectively, a Q learning algorithm is adopted to optimize the weight of the combined model to realize accurate prediction of photovoltaic power. And finally, selecting four indexes of a decision coefficient, a root-mean-square error, a mean absolute error and a mean absolute percentage error, evaluating the obtained prediction result, and integrating the steps to finally obtain the prediction method with dynamic optimization combination parameters, adaptive data distribution change and enhanced small sample generalization ability.
Owner:LANZHOU PETROCHEMICAL VOCATIONAL & TECH UNIV

Omnidirectional leveling device for hilly and mountain tractor seat and control method

The invention belongs to the technical field of tractor seat adjustment, and particularly relates to an omni-directional leveling device for a hilly and mountainous region tractor seat and a control method, and the omni-directional leveling device comprises a posture sensing unit, a central controller and a hydraulic execution unit. The posture sensing unit is used for detecting transverse and longitudinal inclination angles of the seat in real time; the hydraulic execution unit comprises an oil source system, four symmetrically-arranged leveling hydraulic cylinders and corresponding electro-hydraulic proportional direction valves. The core of the control method is that a QBP-PID composite intelligent control strategy fusing a Q learning algorithm, a BP neural network and PID control is adopted. The ECU receives an actual inclination angle signal, optimizes PID control parameters in real time by using a QBP-PID composite intelligent control strategy with a horizontal attitude as a target, and drives the electro-hydraulic proportional directional valve to accurately control the cooperative telescopic movement of each hydraulic cylinder, so that the rapid, stable and self-adaptive adjustment of the transverse and longitudinal inclination angles of the seat is realized, and the safety of the seat is improved. The problem that the sitting posture of a driver is uncomfortable due to topographic relief of a hilly and mountain tractor is solved, and the operation comfort and safety are improved.
Owner:HENAN UNIV OF SCI & TECH

A method and system for frequency control of a thermal power plant based on reinforcement learning algorithm

The application relates to a frequency control method for a thermal power plant based on a reinforcement learning algorithm, wherein the method comprises the following steps: obtaining primary frequency regulation evaluation data and current combined cycle unit working condition data to form state data; determining a state-action pair in a preset algorithm model, wherein the preset algorithm model is obtained by training a Q learning algorithm through reinforcement learning based on the state data of the thermal power plant; obtaining evaluation values of each state-action pair based on a preset estimation rule through the preset algorithm model, and obtaining a target state-action pair according to the evaluation values; generating a control instruction according to an environmental interaction action in the target state-action pair, and adjusting the output power and action time of an electric energy storage device through the control instruction. Through the application, the problem that the service life of an electric energy storage device of a combined cycle unit is low due to long-term high-load working state in the related art is solved, accurate frequency regulation control is realized, power waste of the electric energy storage device is reduced, and the service life of the electric energy storage device is improved.
Owner:HUADIAN ELECTRIC POWER SCI INST CO LTD

Security control method for networked switching system based on encryption and decryption switching q-learning replay attack detection

The application discloses a kind of networked switching system security control method based on encryption and decryption switching Q learning replay attack detection, comprising: constructing switching system model and controller model, obtain state signal and mode signal;Event-triggered mechanism based on replay attack detection is constructed, including error detection condition, mode matching condition and attack detection condition;According to Q learning algorithm, replay attack detection mechanism based on switching Q learning is constructed, event-triggered mechanism and replay attack detection mechanism are used to process the mode signal and state signal of triggered time subsystem, state estimation value, mode estimation value and state estimation residual norm are obtained, the detection result of attack is determined;State estimation value and mode estimation value are substituted into controller model, and safety controller is obtained, the equation of closed-loop system is obtained by simultaneously solving the equation of safety controller, switching system model, state estimator and mode estimator;Switching rule based on replay attack detection is constructed, according to replay attack detection result, the switching behavior of closed-loop system under replay attack is analyzed, and the judgment condition of system safety and stability is obtained, the present application can accurately analyze whether replay attack is carried out, and give the safety control discrimination condition for ensuring that switching system under replay attack, realize the safety control of switching system.
Owner:DALIAN MARITIME UNIVERSITY