Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1621 results about "State space" patented technology

In the theory of discrete dynamical systems, a state space is the set of all possible configurations of a system. For example, a system in queueing theory defining the number of customers in a line would have state space {0, 1, 2, 3, ...}. State spaces can be either infinite or finite. An example of a finite state space is that of the toy problem Vacuum World, in which there are a limited set of configurations that the vacuum and dirt can be in.

Computing power scheduling method and system based on dynamic load prediction and resource priority ranking

The invention discloses a computing power scheduling method and system based on dynamic load prediction and resource priority ranking. The computing power scheduling method comprises the following steps: collecting historical load data, task submission data and resource state data of each node in a computing power cluster; on the basis of the preprocessed multi-dimensional load feature data set, constructing an improved hybrid prediction model, optimizing model parameters through training, and predicting the load change trend of each computing power node in a future preset time period by using the trained model to obtain a node load prediction result; extracting a service level protocol parameter, a resource demand type and historical execution efficiency data of a to-be-scheduled task, and establishing a multi-dimensional resource priority evaluation index system; according to the computing power scheduling method, the problems of low resource utilization rate and high task response delay caused by low load prediction precision and mismatching of resource allocation and task priority in a traditional computing power scheduling method are solved, and the overall operation efficiency and service quality of a computing power cluster are improved.
Owner:SHAOGUAN DATA IND RESEARCH INSTITUTE

Task scheduling optimization method and device based on reinforcement learning, equipment and medium

The invention relates to a task scheduling optimization method and device based on reinforcement learning, equipment and a medium. The method comprises the steps that firstly, system resource state data are collected in real time, dynamic environment characteristics are determined through preprocessing and time sequence analysis, task characteristic data are analyzed at the same time, and a task priority sequence and a resource demand vector are generated through a priority ranking algorithm and a resource evaluation model; and then a state space and an action space are constructed by adopting a reinforcement learning algorithm, an optimal task allocation scheme is generated through strategy iteration and reward function optimization, and if the scheme meets a resource balance threshold, scheduling is executed, and performance indexes are collected. And finally, fusing real-time indexes with historical data, and updating parameters of the reinforcement learning model through experience playback and gradient descent to form a closed-loop optimized improved scheduling strategy. By adopting the method, the accurate mapping of the resource state and the task requirement can be realized, and the problem of insufficient adaptability of the traditional static scheduling to a complex scene is solved.
Owner:SHAOGUAN XINGCHENG NETWORK TECH CO LTD

Intelligent heat supply regulation and control system based on digital twinning and deep reinforcement learning

The invention discloses an intelligent heat supply regulation and control system based on digital twinning and deep reinforcement learning. The system comprises a physical layer, a control layer and a control layer, wherein the physical layer is a physical heat supply system composed of heat source equipment, a transmission and distribution pipe network and a user terminal; according to the digital twinborn layer, a virtual heat supply system mapped with the physical layer in real time is constructed, the virtual heat supply system comprises a multi-physics field coupling model based on the thermodynamics and fluid mechanics principle, operation data of the physical layer are collected through a distributed sensor network, and the state vector of the virtual system is dynamically updated; and the intelligent decision-making layer is integrated with a DRL intelligent agent, the state space of the DRL intelligent agent is defined as a virtual system state vector output by the digital twin layer, the action space of the DRL intelligent agent is a regulation and control instruction combination of heat source power and pump valve opening, and a reward function fuses an energy consumption penalty term, a room temperature comfort reward term and a pipe network stability constraint term. According to the method, global optimization, high-precision continuous regulation and control and collaborative balance are realized through deep collaboration of digital twinning and deep reinforcement learning.
Owner:TIANJIN THERMAL CO

Resource scheduling method and system based on reinforcement learning

The invention belongs to the technical field of resource scheduling, and particularly discloses a resource scheduling method and system based on reinforcement learning, and the method comprises the steps: describing the dependency and conflict relation between tasks through constructing a task causal relation graph which can be dynamically updated; constructing a state space and an action space based on a current task state and a causal relationship, training an intelligent agent by adopting a reinforcement learning model in combination with a multi-target reward function, and generating a scheduling and resource allocation strategy; linear programming and heuristic joint resource allocation are carried out under strategy guidance, and task priorities and causal relationships are dynamically adjusted in combination with task states; by collecting execution data and analyzing strategy deviation, a model structure and parameters are further adjusted, and a complete closed-loop optimization process is formed. According to the method, task priority dynamic adjustment, resource allocation strategy self-adaptive updating and scheduling process closed-loop optimization are realized, and the method is suitable for complex project management scenes under multi-task and multi-resource constraints.
Owner:INSPUR IND (CHONGQING) INTELLIGENT EQUIPMENT TECHNOLOGY CO LTD

Power distribution network bearing capacity evaluation system based on dynamic correction

The invention relates to the technical field of power distribution network evaluation, and discloses a power distribution network bearing capacity evaluation system based on dynamic correction. The system comprises a dynamic data acquisition module, a multi-dimensional state space construction module, a security domain analysis module, a partition coupling degree calculation module and a bearing capacity evaluation engine module. The dynamic data acquisition module acquires a power injection quantity sequence, a voltage deviation ratio sequence and uncontrollable parameter fluctuation data of each partition node of the power distribution network; a multi-dimensional state space construction module performs dimension raising mapping on the sequence to generate a linearized power flow state space model containing a power-voltage Jacobian matrix; the security domain analysis module corrects the boundary of the model according to uncontrollable parameter fluctuation and generates a dynamic security operation constraint set; the partition coupling degree calculation module quantifies an electrical independence index by means of a spectrum radius; and the bearing capacity evaluation engine constructs a chance constraint optimization model, outputs the photovoltaic maximum accessible capacity of each partition and a safety guarantee supply control strategy set, and improves the evaluation accuracy and practicability.
Owner:国网甘肃省电力公司金昌供电公司

Target detection method based on Mama feature fusion

The invention discloses a target detection method based on Mama feature fusion, and relates to the technical field of image target detection. According to the method, the innovative implementation of the VSSA module is utilized, a selective scanning mechanism of the state space model is applied to 2D visual data processing, the long-distance dependency relationship in the image is effectively captured through state space modeling in four directions, the limitation of a traditional state space model in the two-dimensional visual data processing process is solved through the multi-direction processing strategy, and the processing precision of the 2D visual data is improved. The model can comprehensively perceive spatial dependency relationships in different directions in an image, the VSSA adopts learnable state space parameters to dynamically model a feature sequence, the ability of the network to understand a complex space structure is enhanced, the method is particularly suitable for processing scenes needing long-distance context information, and in addition, the method is combined with MTMHSA, so that the complexity of the network is reduced. And the fusion capability of different levels of features in target detection is further enhanced. Through the innovation, the model can better understand the target in the image, and the positioning and classification precision of the target is improved.
Owner:CHONGQING UNIV OF TECH

Virtual power plant response optimization scheduling system and method based on reinforcement learning

The invention discloses a reinforcement learning-based virtual power plant response optimization scheduling system and method, and relates to the technical field of virtual power plant intelligent scheduling. The system comprises an environment modeling module, an intelligent agent module, a multi-agent coordination module and a self-adaptive optimization module which are respectively used for constructing a multi-dimensional state space and a layered action space, generating and optimizing an action strategy based on an Actor-Critic network, executing a scheduling instruction through a layered multi-agent structure and realizing conflict consensus, and dynamically adapting to state space change in combination with incremental learning and meta-learning mechanisms. The system and the method have the advantages of fine state modeling, efficient action response, adaptive strategy updating, stable agent coordination and the like, and can keep the continuity, the stability and the optimality of a scheduling strategy in an operation environment in which multi-source heterogeneous power resources participate in scheduling cooperatively, market rules change frequently and load fluctuation is violent.
Owner:NANJING HUADUN ELECTRIC POWER INFORMATION SAFETY EVALUATION CO LTD

Large sliding bearing fault detection and evaluation method, device and system

The invention relates to the field of mechanical equipment health management, in particular to a large sliding bearing fault detection and evaluation method, device and system. Comprising the following steps: collecting multi-source sensing data, and constructing a comprehensive data set; constructing a state space model based on a sliding bearing physical mechanism; the multi-source sensing data and the state space model are fused through Bayesian filtering, and hidden state parameter posterior distribution is dynamically estimated; generating a virtual fault sample by using a generative adversarial network in combination with a physical rule base; designing a Bayesian space-time sequence diagnosis model based on an attention mechanism, and generating fusion health state features; processing and fusing the health state features by using a degradation process model, and predicting the remaining service life of the bearing; and based on the health state, the fault probability and the remaining service life, setting multi-stage early warning threshold values, and triggering intelligent early warning. According to the method, the defect that a single model is insufficient in adaptability and generalization ability under complex working conditions is overcome, and the accuracy and reliability of fault detection are remarkably improved.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU +2

Power system load dynamic optimization method based on reinforcement learning

The invention discloses a power system load dynamic optimization method based on reinforcement learning. The method comprises the following steps: S1, collecting power system data to construct a state space; s2, constructing a hierarchical reinforcement learning model based on the state space, and dividing a high-level decision and a low-level execution task; s3, training a high-level decision model, and outputting a scheduling task target category instruction in a high-level state; s4, training a low-layer execution model, and outputting a control action in combination with a current node state and a high-layer instruction; s5, introducing an evolutionary mechanism to generate a strategy population and optimizing a low-layer execution model; s6, fusing an evolutionary mechanism and a strategy gradient to synchronously optimize individuals with excellent performance; s7, deploying the trained model to a power dispatching system; s8, performing model parameter fine tuning based on scheduling feedback; and S9, continuously applying the fine-tuned model to load scheduling control. According to the invention, power load accurate scheduling and strategy efficient adaptive optimization are realized, and system responsiveness and operation stability are improved.
Owner:ZHEJIANG JUHUA THERMAL POWER CO LTD

Multi-parameter MPC two-phase liquid cooling energy efficiency optimization distribution control system and method

The invention discloses a multi-parameter MPC two-phase liquid cooling energy efficiency optimization distribution control system and method, and relates to the technical field of liquid cooling energy efficiency optimization. The multi-parameter MPC two-phase liquid cooling energy efficiency optimization distribution control system and method comprises the following steps: S1, collecting and preprocessing two-phase liquid cooling control data; s2, the thermal load level of each area is evaluated periodically, a discrete time linear state space model is constructed, and the thermal load change rate is extracted; s3, the cold distribution amount of each area is quantified periodically, the pump speed and the valve opening target value are calculated, and an adjustment control instruction is generated; and S4, high heat risk and low heat stability areas are recognized, the cooling capacity compensation requirement is evaluated, a cooling capacity enhancement and recovery distribution value is output, and a corresponding control instruction is generated. The problems that an existing two-phase liquid cooling system generally only depends on temperature feedback for adjustment, so that cooling capacity distribution response lags behind, a short-time overheating phenomenon is prone to occurring in a high-heat area, and cooling capacity waste exists in a low-heat area at the same time are solved.
Owner:TIANJIN TIER TECHNOLOGY CO LTD

Elevator energy recovery application method and system based on supercapacitor

The invention provides an elevator energy recovery application method and system based on a super capacitor, and relates to the technical field of energy recovery, and the method comprises the steps: collecting elevator operation data and super capacitor state data, constructing a state space and an action space, and constructing a reward function according to peak-valley electricity price economic benefits, super capacitor life loss and energy utilization efficiency; training a deep reinforcement learning model based on the double-Q network to obtain an energy recovery optimization strategy; when the elevator is braked, regenerative braking energy is stored in the super capacitor; when the elevator is driven, the stored energy is released for traction; the optimal storage allocation scheme of the multiple elevator rooms is calculated through the virtual energy storage unit, hierarchical control is implemented based on an optimization strategy and the allocation scheme, the upper layer adopts fuzzy self-adaptive weight to determine the energy storage priority, and the lower layer adopts a distributed algorithm to calculate the energy allocation coefficient and the target power; and the charge-discharge rate is updated according to the health state of the supercapacitor.
Owner:BEIJING RUIHE DEBAO THERMAL TECH CO LTD

Ship inclination measuring method and system based on Kalman filtering fusion

The invention discloses a ship inclination measurement method and system based on Kalman filtering fusion, and relates to the technical field of ship stability measurement, and the method comprises the steps: multi-sensor synchronous data collection, state space modeling and Kalman filtering fusion, adaptive dynamic adjustment of process noise, drift compensation, and fault tolerance and output. When the method and the system provided by the invention are used for measuring inclination, inclination angle error accumulation caused by gyroscope zero offset integral error, accelerometer dynamic interference, sensor noise and temperature drift, GNSS dependency and the like can be effectively avoided, available precision can still be maintained through multi-sensor depth coupling and an intelligent filtering algorithm, and the accuracy of inclination measurement is improved. The method is simple in operation, independent of satellite data, higher in stability, capable of greatly reducing hardware cost compared with GNSS and IMU schemes, and especially has the advantages of being simple in cost and installation configuration in ship refitting; and the method is also suitable for GNSS rejection or cost sensitive scenes such as inland ships, unmanned ships, offshore buoys and the like.
Owner:CHINA SHIPPING TELECOMM

Event camera pedestrian detection method based on space-time state space model

The invention relates to the technical field of artificial intelligence computer vision, in particular to an event camera pedestrian detection method based on a space-time state space model, and the method comprises the steps: collecting a pedestrian detection data set based on an event camera, obtaining original event data, carrying out the preprocessing, determining an event tensor, and obtaining an event frame sequence; modeling is carried out through combination of a state space and an event tensor, an event-driven recursive space-time state space module is defined as a core unit, and an isomorphic deep neural network architecture is constructed; training the isomorphic deep neural network architecture by adopting the training set, and verifying through the verification set; inputting the test set into the verified isomorphic deep neural network architecture for detection, and generating a pedestrian detection result; the collaborative optimization of sparse adaptation-dynamic capture-noise suppression is realized in a unified framework, the essential characteristics of the event camera triggered based on brightness change are theoretically fit, and higher robustness and generalization ability are shown in an actual complex scene.
Owner:NANJING UNIV OF POSTS & TELECOMM

Power stealing risk threshold dynamic optimization and multistage early warning method based on reinforcement learning

The invention discloses an electricity larceny risk threshold dynamic optimization and multistage early warning method based on reinforcement learning, and the method achieves the dynamic adjustment of an electricity larceny risk threshold through constructing a multi-dimensional state space containing a risk perception percentage, a line loss rate deviation, a cost budget and a load feature vector, and employing a reinforcement learning algorithm. Safety constraints are introduced to ensure the rationality of threshold adjustment, a TD3 algorithm is adopted to train and update a strategy network and a value network, the learning efficiency is improved through priority experience playback, a multi-target weighted reward function is designed to balance the multi-target optimization requirements of detection precision, line loss control, inspection cost and the like, and a red, orange and yellow three-color early warning mechanism is constructed. Early warning actions of different levels are triggered according to the ratio of the dynamically updated reference threshold value to the risk perception percentage, and early warning visualization and strategy effect evaluation are achieved through a geographic information graph dynamic rendering module, a natural language report generation module and an anti-fact analysis module.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Hybrid load flow calculation method and device based on cross entropy algorithm

The invention discloses a hybrid power flow calculation method and device based on a cross entropy algorithm, and belongs to the technical field of optimization algorithm and machine learning crossing. The method comprises the steps that system modeling and parameter initialization are carried out, and a state space, an action space and a target function are defined according to a specific application scene; setting and initializing row parameters; entropy algorithm-guided sample generation and elite screening: iteratively optimizing power grid parameters through probability distribution, and quickly approaching a global optimal solution; the strategy optimization target of the strategy gradient algorithm is that a dynamic strategy network is constructed based on optimized power grid parameters, and the running state of the power grid is adjusted in real time to deal with uncertainty; performing sample multiplexing and distribution fusion of importance sampling: setting a double-buffer mechanism and performing sample fusion and multiplexing; and parameter updating and convergence judgment. According to the method, global search is carried out through the cross entropy algorithm, local optimization is carried out through the strategy gradient algorithm, the sample utilization rate is improved through the importance sampling method, and cooperative processing of global optimization and local optimization is achieved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1

Mechanical arm track optimization method and system based on deep learning and fuzzy algorithm

The invention relates to the technical field of intelligent mechanical arm control, and discloses a mechanical arm track optimization method and system based on deep learning and a fuzzy algorithm, and the method comprises the steps: building a kinematic model of a mechanical arm, determining the working space of the mechanical arm, carrying out the high-density random sampling, and generating a three-dimensional point cloud picture of a reachable region at the tail end of the mechanical arm; constructing a path planning model, designing a state space and an action space, and constructing a reward function; time-impact double-target optimization is carried out on the tail end path point sequence, a smooth joint trajectory is constructed, and balance between the shortest trajectory execution time and the minimum joint impact is achieved on the premise that speed and acceleration constraints are met; and tracking control is carried out on the trajectory, external disturbance and unmodeled dynamics are estimated and compensated in real time, a parameter adaptive law is designed, and the trajectory tracking precision of the system in a complex environment is improved. The autonomy, the accuracy and the anti-interference capability of the hot-line work mechanical arm in a complex environment are improved.
Owner:CHINA UNIV OF MINING & TECH

Method and system for predicting residual service life of engine based on space-time state selection

The invention discloses an engine residual service life prediction method and system based on space-time state selection, and belongs to the field of equipment residual service life prediction. According to the method, multiple technologies such as a time sequence-space double-branch coding and hierarchical fusion mechanism, a selective state space model, a cross-layer residual jump connection and parameter sharing mechanism and an efficient self-attention mechanism are integrated, and an STSSFormer network model is constructed to predict the residual service life of the aero-engine; the system mainly comprises a data selection and preprocessing module, an input sample construction module, a model training module and a residual service life prediction module. According to the method, the feature expression and fusion capability is comprehensively enhanced, and the method has remarkable advantages in the aspects of prediction accuracy of the remaining service life, model robustness and engineering practicability.
Owner:SHANDONG UNIV OF SCI & TECH

Disassembly line balance optimization method and system based on multi-objective reinforcement learning, and medium

The invention relates to the field of data processing technology prediction, and discloses a disassembly line balance optimization method and system based on multi-objective reinforcement learning, and a medium, and the method comprises the steps: S1, obtaining input data, and carrying out the preprocessing; s2, constructing a state space, an action space and a target function, and deploying a training environment; s3, training a multi-target reinforcement learning model: introducing a preference vector through an MO-DQN algorithm to dynamically adjust a target weight, calculating a target function according to the disassembly data obtained in the S1, optimizing a plurality of targets in the S2, and selecting an action strategy in combination with an epsilon-greedy strategy and an NSGA-II algorithm; and S4, applying an optimal strategy generated by the trained model to task allocation in the disassembly line. By introducing the preference vector, the target weight is dynamically adjusted according to the specific demand of the task, the conflict between the targets can be effectively processed, the optimization strategy can be automatically adjusted according to the change of different situations, and the task scheduling and resource allocation efficiency in the disassembly task is improved.
Owner:QINGDAO UNIV OF TECH

Long-range multivariable load prediction method and system based on time-frequency domain collaboration

The invention belongs to the technical field of power system load prediction, and relates to a long-range multivariable load prediction method and system based on time-frequency domain collaboration, and the system carries out the normalization and stabilization of a multivariate load time sequence through a data preprocessing module; the feature embedding module performs linear embedding on the block sequence to construct high-dimensional feature representation; the state space coding module extracts long-range dependency features and generates depth time sequence representation; the decoding prediction module maps the coding features into a preliminary prediction sequence; the time sequence alignment module identifies a leading-lagging relation among multiple variables and aligns a time sequence; the frequency domain optimization module realizes frequency domain component fusion based on adaptive filtering; and the model training optimization module is used for performing training and optimization through a signal attenuation loss function. The method can effectively improve the precision and robustness of long-range multivariable load prediction, and especially has obvious advantages in the aspects of processing complex dependency relationships and dynamic time delay.
Owner:HARBIN INST OF TECH AT WEIHAI

Actuator multi-mode failure-oriented distributed driving hovercar self-adaptive fault-tolerant control method

The invention relates to the technical field of aerocar mode switching, and discloses a distributed driving aerocar self-adaptive fault-tolerant control method for actuator multimode failure, which comprises the following steps: constructing a unified six-degree-of-freedom dual-mode state space model; residual signals are generated based on extended Kalman filtering and a sliding-mode observer, and fault types and positions are positioned in real time through a lightweight classifier; the method comprises the following steps: extracting residual time-frequency features, identifying hard faults by using a lightweight convolutional neural network, quantifying soft fault degrees through an incremental support vector machine, fusing multi-source information based on a Bayesian network to output fault types, levels and confidence coefficients, and introducing an incremental learning mechanism to realize self-evolution of a diagnosis model; a virtual control instruction is generated by adopting hierarchical sliding mode control, thrust and torque distribution of remaining actuators is optimized based on a dynamic quadratic programming algorithm, control parameters are adjusted online in combination with a Lyapunov adaptive law, aerodynamic interference and model uncertainty are inhibited, attitude stability and trajectory tracking in air-ground mode switching are guaranteed, and the method has the advantages of being high in reliability and high in reliability. And the fault-tolerant performance and the operation safety of the hovercar in the air-ground mode switching process are obviously enhanced.
Owner:HEFEI UNIV OF TECH

Course resource recommendation method and device based on improved state space model, equipment and storage medium

The invention discloses a course resource recommendation method and device based on an improved state space model, equipment and a medium, and relates to the technical field of learning resource recommendation, and the method comprises the steps: collecting user course resource data, obtaining and preprocessing an interaction sequence, mapping the user interaction sequence to a low-dimensional vector space, and inputting the low-dimensional vector space into the improved state space model; the target interest course is matched with the candidate course resource set, the recommendation score is calculated, and the recommendation list is generated, so that the calculation complexity is reduced, the user interest is accurately captured, and the efficiency and accuracy of course resource recommendation are improved.
Owner:湖南工商大学

Intelligent 5G edge cooperation power supply driving system

The invention relates to the technical field of 5G edge collaborative power supply driving, and discloses an intelligent 5G edge collaborative power supply driving system, which comprises a state sensing module used for collecting multi-source data of each power supply device in the power supply driving system in real time through a multi-source sensor; the load prediction module is used for predicting a future load by using a load prediction algorithm and outputting a load prediction value; the thermal path analysis module is used for constructing a thermal path diagram, predicting a thermal diffusion trend and adjusting power output of the power supply driving system; the power dispatching optimization module is used for adjusting the power of the power supply based on the optimization objective function; and the remote cooperative control module is used for the cloud platform to generate an optimization strategy according to the real-time states of load prediction, temperature prediction and power supply scheduling. According to the method, the integrity and the time relevance of the equipment operation state expression are improved by constructing the continuous time state evolution model in a modeling mode of combining the multi-source state perception and the state space model.
Owner:WEISHI MILITARY & CIVIL (GUAN) ELECTRONIC TECH CO LTD

Large-scale building refrigeration system intelligent energy saving method based on large language model

The invention relates to an intelligent energy-saving method for a large building refrigerating system based on a large language model, and belongs to the field of intelligence of large building refrigerating systems. A reinforcement learning framework is constructed, a multi-dimensional state space formed by building thermal loads, equipment operation states and thermal environment parameters is defined, and a dynamic reward mechanism is designed by taking improvement of energy efficiency, guarantee of thermal safety constraints and maintenance of thermal comfort as optimization objectives; developing a word embedding model special for the refrigeration field and a cross-modal decision conversion model, and realizing bidirectional analysis of a natural language instruction and a physical control parameter through feature coding and embedding mapping; and performing low-rank adaptive fine tuning on the pre-trained large language model based on the reinforcement learning experience set, and constructing a closed-loop verification system. Generalized migration and multi-target collaborative decision-making of equipment control strategies are achieved, the self-adaptive regulation and control capacity of a refrigeration system under dynamic loads is improved, and a universal intelligent solution for energy efficiency optimization, thermal safety and comfort is provided for a large building.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Water electrolysis efficiency dynamic scheduling method and system based on reinforcement learning

The invention relates to the field of information automation management, and particularly discloses a water electrolysis efficiency dynamic scheduling method and system based on reinforcement learning. According to the method, a state space is constructed through multi-parameter dynamic perception, an improved depth deterministic strategy gradient (DDPG) algorithm is used for learning, a strategy is optimized by means of a self-adaptive exploration strategy and a multi-time scale learning mechanism, model migration and rapid starting are achieved, real-time feedback and closed-loop control are constructed, the water electrolysis efficiency is improved, and energy consumption is reduced. The method comprises a reward function and an exploration rate adjustment formula, so that the strategy optimization direction is accurate, and the exploration process is reasonable. The system comprises a data acquisition module, a processing module, a reinforcement learning module, a control execution module and a feedback monitoring module, and all the modules work cooperatively to improve the system performance.
Owner:BUOREN JINGCHUANG (BEIJING) MEDICAL EQUIP CO LTD

Transient overvoltage risk assessment method, device and equipment for AC / DC system containing high-proportion new energy, and storage medium

The invention provides a transient overvoltage risk assessment method, device and equipment for an AC / DC system containing high-proportion new energy, and a storage medium. Relates to the technical field of power systems and automation thereof. The method comprises the following steps: modeling a power system into a discrete nonlinear power system based on phase change measurement unit data, constructing a state space, introducing a Koopman operator to determine an observation function, and establishing a dynamic time sequence track prediction model based on the observation function to output a predicted dynamic time sequence track; an Informer model is constructed, in response to an input predicted dynamic time sequence track, feature weights are distributed through a ProbSparse self-attention mechanism, hierarchical feature down-sampling is performed by using a distillation mechanism, and a predicted transient voltage amplitude is output; and quantifying the transient overvoltage risk level according to the predicted transient voltage amplitude. According to the method, the problems of low isolated time section feature prediction precision and poor long-sequence dynamic evolution feature evaluation timeliness in the prior art are solved.
Owner:STATE GRID JILIN ELECTRIC POWER COMPANY LIMITED +1

Ship fuel cell cooling system energy efficiency management method based on deep reinforcement learning

The invention discloses a ship fuel cell cooling system energy efficiency management method based on deep reinforcement learning, and relates to the technical field of ship energy management, and the method comprises the steps: constructing a digital simulation model of a ship fuel cell cooling system, which is used for representing the dynamic characteristics of the cooling system under different working conditions; constructing a multi-time scale prediction and thermal load prediction model, and respectively outputting cooling system state and working condition prediction parameters and thermal load prediction parameters under different time scales; a reinforcement learning framework based on the dual deep Q network is built, a state space, an action space and a reward function are defined, and the state space comprises current state parameters of the cooling system and output parameters of all prediction models; the action space comprises seawater pump rotating speed, fresh water pump rotating speed and regulating valve opening; and an optimal control strategy is generated based on a training result of deep reinforcement learning, and global collaborative optimization operation of execution equipment of the cooling system is realized, so that the system energy efficiency and environmental adaptability are improved.
Owner:CHINA SHIP SCIENTIFIC RESEARCH CENTER

Millisecond regulation and control method for hydrogen fluoride production based on reinforcement learning and model prediction control

The invention discloses a millisecond regulation and control method for hydrogen fluoride production based on reinforcement learning and model prediction control, and relates to the field of hydrogen fluoride production regulation and control. In the multi-scale neural symbol dynamics modeling step, GNN, a cellular automaton and a state space model are fused, and parameters are updated in real time; in the causal reinforcement learning decision optimization step, an effect is calculated through a causal graph, and a reward function is optimized; in the time-space fractional order sliding mode control step, a fractional order sliding mode surface and a controller are designed, and rapid and stable control is achieved; in the memory enhancement element learning adaptation step, a DNC storage strategy is utilized, new working conditions are quickly adapted through element gradient, and millisecond-level precise regulation and control are achieved. According to the method, prediction errors are greatly reduced, the response speed is increased, and overshoot is reduced; the product yield is improved, the energy consumption is reduced, and new working conditions are quickly adapted; fault detection and risk early warning are more accurate, equipment operation is more stable, production efficiency is effectively improved, cost is reduced, and safety is enhanced.
Owner:北京云桥智海科技服务有限公司 +1

Bridge group maintenance priority dynamic decision-making method and device based on reinforcement learning

The invention provides a bridge group maintenance priority dynamic decision-making method and device based on reinforcement learning, and relates to the technical field of bridge intelligent maintenance. The method comprises the following steps: constructing a topological structure of a bridge network and a road; defining a state space, a maintenance action space and a state transition matrix of the bridge; defining a reliability index corresponding to the state of the bridge, and designing a comprehensive reward function based on maintenance cost, asset risk and traffic network capacity loss risk based on the topological structure; constructing a bridge maintenance decision problem; the method comprises the following steps of: describing a bridge maintenance decision problem as a Markov decision process, establishing a pointer network strategy model by adopting a pointer network, and training the pointer network strategy model by adopting an Actor-Critic algorithm to obtain a maintenance decision model based on reinforcement learning; and training the maintenance decision model based on reinforcement learning until convergence, and outputting a bridge maintenance action sequence under limited constraints. By adopting the method, the limitation problem of traditional single bridge assessment can be solved.
Owner:UNIV OF SCI & TECH BEIJING

MBD-based reinforcement learning inverter control algorithm optimization method and system

The invention relates to the technical field of inverter control, and provides an MBD-based reinforcement learning inverter control algorithm optimization method, which comprises the steps of S1, establishing a state space model of an inverter, designing a baseline controller of double-loop control, and constructing an MBD simulation platform containing multiple load models; s2, defining a state space of the 12-dimensional state vector, designing a continuous action space, and constructing an adaptive multi-target reward function; s3, training a control parameter optimization strategy by adopting an Actor-Critic algorithm architecture in combination with an experience playback mechanism and a hybrid exploration strategy; s4, establishing a security constraint mechanism of three-layer security protection, designing a fault detection and isolation strategy, realizing a closed-loop online learning strategy of pre-training-transfer learning-online fine tuning, and meanwhile, adopting a self-adaptive updating mechanism; and S5, verifying the optimization effect of the control algorithm through quantitative index evaluation, an experimental verification scheme and a real-time performance requirement test. And the upgrade of inverter control from model driving to data and model cooperative driving is realized.
Owner:SHANGHAI SHENSILICON SEMICON CO LTD

Data center IT load and cooling system cooperative control method based on TD3 algorithm

The invention provides a data center IT load and cooling system cooperative control method based on a TD3 algorithm, and relates to the field of deep reinforcement learning, and the method comprises the steps: obtaining an initial data set representing the server state, environment and task characteristics of a data center, and processing the initial data set to obtain a target data set; building a deep reinforcement learning model which is based on a TD3 algorithm and comprises a strategy network Actor and a value network Critic, and determining a state space, an action space and a reward function of the model; training the model to obtain a target model; and outputting a work task allocation strategy and a cooling system regulation and control strategy through interaction of the target model and the environment. According to the target model, by outputting the work task allocation strategy and the cooling system regulation strategy, work task allocation can be guided, operation of the cooling system can be controlled, the IT load of the server and refrigeration balance of the cooling system are achieved, safe and stable operation of the server is ensured, and energy consumption of the cooling system is reduced.
Owner:HEFEI UNIV OF TECH