Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

393 results about "Reinforcement learning control" patented technology

Defect repairing method based on digital twinning and friction stir welding technology

The invention discloses a defect repair method based on digital twinning and friction stir welding technologies, and relates to the technical field of intelligent manufacturing and digital twinning, and the defect repair method comprises the following steps: synchronously capturing full-dimensional data of a welding area through a multi-mode sensor array integrated by an actuator; secondly, segmenting defect boundaries by adopting a deep learning algorithm, constructing a dynamic twin model in combination with thermal-force field coupling simulation, and accurately mapping defect three-dimensional features; generating a repair track according to the twinborn model, converting the repair track into a robot joint instruction through a curved surface parameterization mapping algorithm, and implanting real-time anti-collision constraint; in the repairing process, based on reinforcement learning control of the material rheological resistance and the temperature gradient, the rotating speed, the advancing speed and the down force of the tool are dynamically adjusted; and after repairing, micro-focus CT scanning is started immediately, actually measured data is compared with twinborn prediction, and when the deviation exceeds a threshold value, a re-repairing process is triggered automatically. The method solves the problems that a traditional method depends on manual intervention and the precision of a sensor is easily interfered by the environment.
Owner:SOUTHWEST TECHNICAL ENGINEERING RESEARCH INSTITUTE OF CHINA SOUTH IND GROUP

Virtual power plant intelligent control method and system based on multiple agents

The invention discloses a multi-agent-based virtual power plant intelligent control method and system, and the method comprises the steps: dividing a virtual power plant into a plurality of sub-virtual power plants, deploying an agent in each sub-virtual power plant, collecting a local resource state through each agent, and predicting a load demand and the output of a distributed power supply, a hierarchical control unit is adopted to carry out collaborative optimization among the sub-virtual power plants according to a prediction result, and an upper-layer optimization control module constructs a linear programming model according to the prediction result and solves the linear programming model to obtain an initial scheduling scheme; and the lower-layer reinforcement learning control module performs local adjustment on the preliminary scheduling scheme according to a multi-agent depth deterministic strategy gradient algorithm to obtain a decision scheme. Based on a distributed control strategy of a multi-agent architecture, the fault-tolerant capability and reliability of the system are improved, a hierarchical control architecture is adopted, global optimization and local adjustment are organically combined, and efficient coordination and real-time adjustment capability of global resources are realized.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER RESEARCH INSTITUTE CO LTD

Multi-mode adaptive industrial instrument intelligent control system based on deep learning

The embodiment of the invention discloses a multi-mode adaptive industrial instrument intelligent control system based on deep learning, and relates to the technical field of industrial instrument intelligent systems. The system comprises a multi-modal data acquisition module used for acquiring multi-modal data of an industrial instrument in real time, the multi-modal data comprising at least two of image, sound, temperature and vibration data; the heterogeneous data fusion module is used for carrying out space-time alignment and feature fusion on the multi-modal data; the deep reinforcement learning control module is used for dynamically optimizing a control strategy and generating a control instruction; the self-adaptive human-computer interaction interface is used for providing visual operation and real-time feedback; and the predictive maintenance subsystem is used for predicting faults and generating maintenance suggestions based on the equipment state. The multi-modal data fusion technology of dynamic weight distribution is adopted, the problem that a traditional single sensor loses efficacy under the complex working condition is solved, and the performance of an industrial instrument in the aspects of measurement precision, self-adaptability and operation and maintenance efficiency is improved.
Owner:北京中科润宇环保科技股份有限公司

Distributed reinforcement learning control method and system for multi-robot cooperative welding

The invention discloses a distributed reinforcement learning control method and system for multi-robot cooperative welding. A heterogeneous multi-robot task allocation algorithm, a multi-agent reinforcement learning model and welding parameters are combined. The method comprises the steps that a state space containing welding current, voltage, speed, welding gun posture, welding seam position and environment obstacle distribution is constructed, and an initial feasible operation area of the robot is determined; the overall welding task is decomposed into sub-tasks, and candidate robots are distributed; configuring a local decision module for each robot, and outputting an action decision containing a welding gun track adjustment amount and a parameter correction value; setting a global coordination module to generate coordination factors; each robot updates the decision and converts the decision into a driving instruction to execute operation, and molten pool temperature and weld size parameters are collected; the parameters are fed back to the model to adjust the state space characterization and decision weight. According to the method, decision collaborative precision is improved in combination with process parameters, self-adaptive optimization is achieved through real-time feedback, and collaborative welding efficiency and quality are improved.
Owner:ZHENGXIN (SUZHOU) TECHNOLOGY CO LTD

Backtracking analysis model construction method based on attack chain

The invention relates to the technical field of data processing, in particular to a backtracking analysis model construction method based on an attack chain, which comprises the following steps that: a kernel layer security agent acquires process, file and network behavior characteristics in a hardware isolation environment, and generates an event tuple; the tensor network pipeline performs three-dimensional decoupling mapping on the tuple into a behavior fingerprint vector, an orthogonalization noise feature and an asymmetric adjacent tensor, and compresses the behavior fingerprint vector, the orthogonalization noise feature and the asymmetric adjacent tensor into a space-time topology tensor block; the reinforcement learning controller constructs a directed acyclic graph based on the tensor blocks, calculates connectivity loss and outputs an event risk score; the dynamic routing engine constructs a decision tree model according to the risk mark, the burst frequency and the correlation entropy, and implements three-level shunting and a multiple simulation system to generate an anti-interference index; and when the deviation between the physical trajectory and the digital model exceeds the tolerance, the closed-loop feedback weight coefficient updates the loss function parameter and adjusts the channel resource weight. And the problem of threat discovery delay caused by attack chain breakage under massive events is solved.
Owner:HUANENG INFORMATION TECH CO LTD

Double-arm body operation method of humanoid robot based on reinforcement learning

The invention relates to a humanoid robot double-arm body operation method based on reinforcement learning, belongs to the field of robot cooperative control, and is characterized in that a strategy of trajectory block prediction and time integration fusion is used for double-arm cooperative control, and continuity and stability of double-arm operation are improved; in reinforcement learning control, a three-dimensional pose track generation and correction module is introduced, condition generation and denoising correction of a three-dimensional space are carried out on a track block layer, and geometric consistency and naturalness of a generated track are guaranteed; the invention further provides a reinforcement learning optimization framework and a simulation-reality migration process, and through system integration of reward, value guidance and migration processes, strategy deployability and safety are guaranteed; compared with the prior art, the method has the advantages that the naturalness, the collaboration and the success rate of double-arm operation can be remarkably improved, the generalization ability is high, and the good simulation-to-reality migration ability is achieved.
Owner:CITIC HEAVY INDUSTRIES CO LTD

Biped robot reinforcement learning control method

The invention discloses a biped robot reinforcement learning control method which comprises the following steps: after a model prediction controller receives a walking instruction, outputting expected angle data of each joint of a robot to a reinforcement learning neural network, and meanwhile, returning joint angle data of an actual strategy of the reinforcement learning neural network to the neural network by a sensing system at a bottom layer; a joint angle taking time as a sequence and planned by model prediction is compared with a joint angle actually generated by a neural network, and model prediction control is fused into a reinforcement learning training process by setting a reward function for punishment, so that the reinforcement learning training efficiency is improved, and a stable gait is more quickly achieved. Excellent gaits planned by the MPC are transplanted into reinforcement learning control, and control robustness can be improved under the condition that the excellent gaits of the MPC are reserved; and meanwhile, the joint angle data which is obtained by taking time as a sequence and is obtained by taking MPC as a planner can accelerate the reinforcement learning training process, so that the training speed and the control effect of reinforcement learning are greatly improved.
Owner:ZHEJIANG UNIV OF TECH

Online monitoring system for friction stir welding process and control method

The invention relates to the technical field of intelligent control of the welding process, and discloses an online monitoring system and control method for the friction stir welding process, and the system comprises a multi-mode sensor array module which is used for collecting temperature, pressure and displacement data in the welding process; the edge computing node module preprocesses the collected data and extracts spatial-temporal characteristics; the space-time diagram convolutional network module is used for mining a space-time association relationship between the sensors; the hierarchical reinforcement learning controller generates a control instruction according to the associated features; the execution mechanism module receives the control instruction and outputs a welding signal; the digital twinborn verification module performs simulation feedback on the control effect; and the cloud model optimization module fuses the multi-source feedback information to continuously optimize the control strategy. The welding process self-adaptive control technology based on reinforcement learning is adopted, the technical effect of adjusting the welding parameters in real time to optimize the welding quality is achieved, and the problem that in a traditional method, the welding quality is unstable due to working condition changes is solved.
Owner:BEIJING SOONCABLE TECHNOLOGY GROUP CO LTD

Mineral processing flow intelligent optimization control method and system based on knowledge graph

The invention discloses a beneficiation process intelligent optimization control method and system based on a knowledge graph, and relates to the technical field of beneficiation process control, and the method comprises the steps: collecting static attribute data and dynamic operation data, constructing a causal knowledge graph, and obtaining a causal relationship between nodes in the causal knowledge graph and an optimization target constraint; based on a causal relationship between nodes in the causal knowledge graph and optimization target constraints, screening data from a historical database to form a high-value historical data set; according to the high-value historical data set, a reward function model is constructed, a strategy function is trained, and a reinforcement learning control strategy is obtained; running the reinforcement learning control strategy, generating a control instruction, and issuing the control instruction to the target equipment to obtain dynamic optimization control; and performing real-time monitoring on the dynamic optimization control, and performing anti-fact reasoning based on a causal knowledge graph when abnormal deviation of the running state is detected to obtain an anti-fact reasoning result.
Owner:CHANGCHUN GOLD DESIGN INST

Battery management system based on adaptive digital twinning

The invention relates to the technical field of BMS and the like, and provides a battery management system based on adaptive digital twinning, a physical layer of the battery management system comprises a battery pack, a sensor network and an edge computing node, and is responsible for data acquisition and preprocessing; the digital twin layer comprises a self-adaptive multi-scale model and a real-time data engine, battery behaviors are dynamically simulated by coupling electrochemical, thermal and aging models, a future state trajectory prediction result is output, and the real-time data engine fuses sensor data, historical data and simulation data to drive model updating; the intelligent decision-making layer comprises a reinforcement learning controller and a fault prediction module which are deployed in a local server, the reinforcement learning controller dynamically optimizes a charging and discharging strategy according to a prediction result of the digital twinborn layer and issues and executes the charging and discharging strategy, and the fault prediction module analyzes multi-source time sequence data based on an LSTM network so as to early warn thermal runaway and short circuit risks in advance. According to the invention, long-term accurate mapping and adaptive adjustment between the battery physical entity and the digital model can be realized.
Owner:深圳市华芯控股有限公司

Autonomous inspection and return control method and system for blow-off pipeline robot

The invention discloses an autonomous inspection and homeward voyage control method and system for a blow-off pipeline robot, belongs to the technical field of intelligent pipeline detection and robot control, and aims to solve the technical problems of how to realize a stable, reliable and continuous clogging exploration task in a complex blow-off pipeline environment and realize safe self-rescue in an abnormal state. According to the technical scheme, the method comprises the following steps: data dimension reduction: receiving high-dimensional environment and motion state data from a multi-modal sensor, performing real-time dimension reduction on the high-dimensional environment and motion state data by adopting a dynamic principal component analysis method, extracting key features of state information of a robot and internal state information of a pipeline, and constructing a low-dimensional state vector; performing deep reinforcement learning control based on a liquid neural network: modeling dynamic environment state change by using the liquid neural network, and realizing path planning and control in combination with a deep deterministic strategy gradient algorithm; tracing the track; and a return flight strategy based on a graph attention network and reinforcement learning.
Owner:浪潮智慧城市科技有限公司

Water electrolysis efficiency dynamic scheduling method and system based on reinforcement learning

The invention relates to the field of information automation management, and particularly discloses a water electrolysis efficiency dynamic scheduling method and system based on reinforcement learning. According to the method, a state space is constructed through multi-parameter dynamic perception, an improved depth deterministic strategy gradient (DDPG) algorithm is used for learning, a strategy is optimized by means of a self-adaptive exploration strategy and a multi-time scale learning mechanism, model migration and rapid starting are achieved, real-time feedback and closed-loop control are constructed, the water electrolysis efficiency is improved, and energy consumption is reduced. The method comprises a reward function and an exploration rate adjustment formula, so that the strategy optimization direction is accurate, and the exploration process is reasonable. The system comprises a data acquisition module, a processing module, a reinforcement learning module, a control execution module and a feedback monitoring module, and all the modules work cooperatively to improve the system performance.
Owner:BUOREN JINGCHUANG (BEIJING) MEDICAL EQUIP CO LTD

Active power distribution network multi-target collaborative voltage optimization control method based on FACMAC algorithm

The invention relates to a source-containing power distribution network multi-target collaborative voltage optimization control method based on an FACMAC algorithm, and belongs to the technical field of photovoltaic inversion control. According to the technical scheme, a power distribution network physical system is composed of a plurality of feeder lines, a transformer, a line and a plurality of grid-connected photovoltaic inverters, and each inverter can measure operation information such as local voltage and current in real time; the data acquisition and communication system is used for acquiring node operation data and realizing low-delay communication; the multi-agent reinforcement learning control system is composed of a plurality of distributed agents and factorization centralized Critic modules, and whole-network voltage optimization decision can be carried out in training and execution stages. And the execution unit adjusts the reactive power output of the inverter in real time according to the control instruction. According to the method, the whole-network cooperative regulation and control capability is improved, the training efficiency bottleneck in a high-dimensional scene is relieved, the expression capability on a complex nonlinear coupling relationship is enhanced, and efficient, stable and extensible power distribution network voltage optimization control is realized.
Owner:STATE GRID JIBEI ELECTRIC POWER COMPANY LIMITED CHENGDE POWER SUPPLY

AGC collaborative optimization method based on time sequence deep reinforcement learning PID control

The invention provides an automatic generation control (AGC) collaborative optimization method based on time sequence deep reinforcement learning PID (Proportion Integration Differentiation) control. Comprising the following steps: establishing a three-area multi-source interconnected power system with participation of hydroelectric power, thermal power, wind power, photovoltaic power and energy storage, and bringing wind and light storage resources into an automatic power generation control loop. The invention provides a novel AGC adaptive PID controller based on time sequence deep reinforcement learning by utilizing the time sequence perception capability of a long short-term memory neural network and combining a double-delay depth deterministic strategy gradient algorithm, and provides an efficient time sequence deep reinforcement learning solution for intelligent power grid frequency control.
Owner:GUO JIA DIAN WANG YOU XIAN GONG SI XI NAN FEN BU +1

Water electrolysis hydrogen production intelligent control system and method based on artificial intelligence

The invention discloses a water electrolysis hydrogen production intelligent control system and method based on artificial intelligence. The system and method are suitable for a large-scale water electrolysis hydrogen production scene under the power supply condition of fluctuating renewable energy sources such as wind power and photovoltaic. The system comprises a data acquisition layer, an edge calculation layer, an intelligent control layer and an execution layer. Dynamic modeling of the running state of the hydrogen production system is achieved through multi-parameter real-time monitoring and feature extraction. The intelligent control layer fuses an LSTM prediction module and a reinforcement learning controller, the LSTM prediction module is used for predicting future renewable energy input and hydrogen demand trends, and the reinforcement learning controller calculates an optimal current density set value based on a prediction result so as to maximize hydrogen production efficiency per unit energy consumption. And meanwhile, a digital twinning technology and a safety protection mechanism are combined, thermoelectric dual regulation and control of the electrolysis process are achieved, and the control response speed and the system stability are improved. According to the invention, energy consumption can be effectively reduced, the service life of the stack is prolonged, and the operation efficiency and reliability of the hydrogen production system under complex load are improved.
Owner:BEIJING MINGYANG HYDROGEN ENERGY TECHNOLOGY CO LTD

Expressway hard shoulder dynamic opening method and system based on NFGD model and simulation platform

The invention discloses an expressway hard shoulder dynamic opening method and system based on an NFGD model and a simulation platform. Comprising the following steps: collecting and adopting a dynamic confidence mechanism to carry out weighted fusion on multi-source real-time sensing data to obtain fused traffic state characteristics; a hard road shoulder dynamic control decision is made based on an NFGD model comprising a fuzzy-neural hybrid controller, a multi-target genetic algorithm controller and a reinforcement learning controller; and real-time interaction with a simulation platform is realized, and control instruction issuing, feedback acquisition and online strategy evaluation are realized. Compared with a traditional method based on a fixed empirical threshold value, the fuzzy-neural hybrid controller, the multi-target genetic algorithm controller and the reinforcement learning controller are fused, self-adaptive modeling and open discrimination are achieved, and decision-making precision and scene adaptability are improved. And meanwhile, by combining with a prediction adjustment feedback type step length control structure, simulation time drift is effectively inhibited, the timeliness of a control strategy is enhanced, and the method has relatively high practical application value.
Owner:HEBEI TRANSPORTATION INVESTMENT GRP CO LTD +2

Cross-scene defect identification method and system driven by migration and fine tuning

The invention discloses a migration fine tuning driven cross-scene defect identification method and system, and relates to the technical field of electric power inspection, and the method comprises the steps: obtaining a basic model through employing a general electric power image data set; constructing a triple set, and forming a power defect domain knowledge graph; performing style conversion on the general power image by adopting CycleGAN, and performing semantic consistency verification on an image sample through a triple set; forming a state vector; inputting the state vector into a reinforcement learning controller, and performing fine tuning training on the basic model to obtain a fine tuning defect identification model; and evaluating the performance of the fine tuning defect identification model on the target domain image verification set, and optimizing a migration strategy selection mechanism according to a performance index. According to the method, the problem of scarcity of target domain labeling data can be solved, sufficient data support is provided for model fine adjustment, model overfitting caused by insufficient data is avoided, the model can adapt to a visual style and a defect distribution rule of a target scene, and the generalization ability of cross-scene defect recognition is improved.
Owner:INFORMATION & COMMNUNICATION BRANCH STATE GRID JIANGXI ELECTRIC POWER CO

Intelligent regulation and control method based on perception and dynamic partitioning

The invention discloses an intelligent regulation and control method based on perception and dynamic partitioning. In order to solve the problem that an existing public building air conditioning system is difficult to deal with personnel activity intensity difference, emotion fluctuation and space density dynamic change, accurate recognition of personnel behaviors and activity density is achieved by fusing YOLOv5, OpenPose and LSTM networks, and the group emotion state is obtained by combining infrared thermal imaging and expression recognition. Based on behavior-density partitioning and an emotional load correction model, space thermal load distribution is dynamically judged, an NSGA-I multi-objective optimization algorithm is introduced, the air conditioner energy efficiency ratio and the thermal comfort degree are optimized, and differential intelligent control over regular and irregular partitioning is achieved through self-adaptive regulation and control of a meta-reinforcement learning control framework. According to the method, the building energy consumption is effectively reduced, the local comfort degree is improved, good real-time performance and intelligent adjusting capacity are achieved, and the method is suitable for optimization design of an intelligent air conditioning system of a large public building.
Owner:SUZHOU UNIV OF SCI & TECH

Rapid training type unmanned aerial vehicle reinforcement learning strategy method suitable for microcontroller

The invention discloses a fast training type unmanned aerial vehicle reinforcement learning strategy method suitable for a microcontroller, and belongs to the technical field of unmanned aerial vehicle control, and the method comprises the steps: constructing a standard four-rotor dynamic model containing the first-order delay characteristic of a motor, and simulating a real physical system; an asymmetric actor-commentator architecture is adopted to ensure the robustness and mobility of the strategy; a reward function coefficient is dynamically adjusted based on a course learning mechanism, and network parameters are optimized by adopting an offline strategy algorithm; performing network training and process optimization; and carrying out lightweight processing and deployment on the trained strategy network, landing the strategy in simulation to a real hardware platform with limited resources, jointly realizing rapid training, reliable migration and lightweight deployment, and realizing high-frequency autonomous flight control of the unmanned aerial vehicle on an embedded platform. According to the method, the high-performance reinforcement learning control strategy can be stably and efficiently operated on the microcontroller with limited resources, and the hardware cost and the power consumption of the intelligent unmanned aerial vehicle are greatly reduced.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Area inspection robot autonomous navigation and path planning method based on multi-modal perception

The invention relates to the technical field of autonomous navigation of inspection robots, and provides a field inspection robot autonomous navigation and path planning method based on multi-modal sensing. The method comprises the following steps: collecting a laser radar point cloud, a camera image, inertial measurement and positioning data, obtaining an environment semantic feature set through spatio-temporal feature fusion, and generating a semantic occupation map; inputting the current position of the robot, the target point and the semantic occupation map into a trajectory generation network to obtain candidate trajectories meeting obstacle avoidance and path smoothness constraints, and completing task sorting and trajectory splicing in combination with inspection task points to form a global path; in the operation process, the reinforcement learning control model adjusts the linear speed and the angular speed in real time, and path tracking and dynamic obstacle avoidance are achieved. The navigation precision and the operation safety of the inspection robot in the complex field area are improved.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD

Distributed reinforcement learning control method and system for multi-robot cooperative welding

The invention relates to the technical field of intelligent welding, in particular to a distributed reinforcement learning control method and system for multi-robot cooperative welding, each welding robot collects vision measurement information under a global coordinate system by using a local vision sensor, and forms self observation information in combination with robot state information obtained by a self-sensor; the information is input into respective Actor-Critic networks, an Actor-Critic algorithm is executed, control actions are generated, welding operation is achieved, visual measurement information between the robots is shared through a 5G communication network, cooperation efficiency is enhanced, the Actor-Critic networks are trained with a cross entropy loss function as a target and with a two-norm of an action value as a constraint condition, and the robot vision measurement information is obtained. The target network adopts a soft update algorithm, the parameter update efficiency is improved, the central control bottleneck is eliminated by the distributed reinforcement learning architecture, and the robustness and reliability of the system are enhanced.
Owner:SHENZHEN ANAHER TECH CO LTD

Mechanical arm visual servo control method and system based on deep reinforcement learning

The invention relates to a mechanical arm visual servo control method and system based on deep reinforcement learning. Firstly, real-time state observation and collision detection are conducted on the mechanical arm, and an observation result is obtained; then, a deep reinforcement learning controller is built by means of a maximum entropy deep reinforcement learning SAC algorithm network architecture, and an action space, a state space and a reward function are built according to an observation result and a target position; and then, outputting a first action representation based on the construction content, optimizing a controller according to a calculation result of the first action representation and the reward function, smoothly optimizing the first action representation by using a real-time track interpolation strategy to obtain a second action representation, and controlling the mechanical arm joint according to the second action representation. And the process is continuously repeated until the mechanical arm successfully arrives at the target position. Compared with the prior art, the method has the advantages of high control real-time performance, large application range and the like.
Owner:SHANGHAI UNIVERSITY OF ELECTRIC POWER

Agricultural irrigation control method and system

The invention discloses an agricultural irrigation control method and system, and relates to the technical field of irrigation control, and the method comprises the following steps: constructing a crop water demand model based on the fusion of crop physiological sensing data and a data driving model, and outputting a target irrigation flow; calculating a control error between the target irrigation flow and the actually acquired flow, and generating a control signal through a model prediction control strategy; inputting a static mapping model fitted by a Sigmoid function, and superposing the output of a dynamic correction model driven by state sensing data to generate a final driving signal; feedback data is collected, a driving signal and the feedback data are subjected to closed-loop comparison, and irrigation control strategy parameters are adaptively optimized through a reinforcement learning controller; according to the method, the crop physiological sensing data and the data driving model are fused, the space and time sequence feature fused water demand model is dynamically constructed, and the problems that agricultural irrigation cannot be accurately controlled and the irrigation strategy is rigid are solved in combination with a static and dynamic composite control mapping mechanism.
Owner:ANHUI WATER TECHNOLOGY DIGITAL INFORMATION TECHNOLOGY CO LTD

Pumping method based on multi-agent reinforcement learning control

The invention provides a pumping method based on multi-agent reinforcement learning control, and belongs to the technical field of deep reinforcement learning. S2, feature extraction and fusion; s3, time sequence modeling; s4, generating an action; s5, carrying out environment interaction; s6, virtual experience generation; s7, experience pool management; s8, performing network training; and S9, performing closed-loop control. Discrete and continuous actions are controlled through the DQN and the DDPG respectively, collaborative optimization of start-stop and frequency adjustment of the water pump is achieved by combining the Critic network, and the problem that a traditional single agent cannot process a mixed action space is solved.
Owner:MECHANICS RES & DESIGN ACAD SICHUAN PROV

Intelligent scheduling system for steel plate hot dipping production process

The invention relates to the technical field of steel plate hot-dip production, and discloses an intelligent scheduling system for a steel plate hot-dip production process. The system comprises a working condition sensing module which is used for continuously collecting steel belt inlet temperature, zinc liquid component concentration fluctuation data, air knife pressure time sequence parameters and roller system operation state signals of a galvanizing pot area; the feature extraction module receives the working condition data flow, analyzes the steel strip surface temperature field distribution features and the zinc liquid viscosity change rule, and generates a process feature spectrum; the parameter optimization module dynamically adjusts membership function parameters and rule base weight coefficients of the fuzzy reinforcement learning controller according to the process characteristic spectrum, and outputs an optimized plating layer control parameter set; the self-adaptive plating layer control module generates a plating layer thickness control instruction based on the real-time zinc liquid temperature parameter and the optimized plating layer control rule set; and the dynamic roller system scheduling module realizes global cooperative control of the roller system speed based on the instruction. The system can realize intelligent scheduling of the steel plate hot-dip production process.
Owner:HANGZHOU LONGYAO POWER PARTS CO LTD

Thermal management method for floating offshore hydrogen production platform cooling system based on phase change heat transfer

The invention belongs to the technical field of ocean engineering and water electrolysis hydrogen production. In order to solve the problems that a traditional water cooling system is large in size and low in heat dissipation efficiency, a multi-scale thermal sensing network is deployed to collect data; establishing a multi-phase-change cross-scale heat transfer model; constructing a two-stage gradient composite cooling module; constructing and optimizing a cooling decision algorithm; designing a cooperative control loop; deep reinforcement learning control is implemented; creating a functional surface; and constructing a dynamic optimization system. The characteristics of part of the steps are limited, efficient heat management can be achieved, and stable operation of the floating type offshore hydrogen production platform cooling system is guaranteed.
Owner:BEIJING NORMAL UNIVERSITY

Unmanned ship and unmanned aerial vehicle system cooperative control method based on hydrodynamic parameters and deep reinforcement learning, electronic equipment and readable storage medium

The invention belongs to the technical field of unmanned system control, and particularly discloses an unmanned ship and unmanned aerial vehicle system cooperative control method based on hydrodynamic parameters and deep reinforcement learning, electronic equipment and a readable storage medium, and the method comprises the steps: S1, obtaining the real-time position, speed and external environment information of each unmanned ship and unmanned aerial vehicle; s2, modeling the hydrodynamic system of each unmanned ship and unmanned aerial vehicle to obtain a dynamic model; s3, performing linear parameterization on the kinetic equations of each unmanned ship and each unmanned aerial vehicle; s4, constructing a deep reinforcement learning control algorithm for improving artificial potential field obstacle avoidance of the unmanned ship and the unmanned aerial vehicle; s5, constructing a network cooperation algorithm of the unmanned ship and the unmanned aerial vehicle; and S6, initializing hardware modules of each unmanned ship and each unmanned aerial vehicle, and loading a network cooperation algorithm into micro-control units of each unmanned ship and each unmanned aerial vehicle.
Owner:JINING UNIV +1

Spring steel wire drawing control method based on reinforcement learning

The invention discloses a reinforcement learning-based spring steel wire drawing control method, which comprises the following steps of: acquiring real-time process parameters to form a process state data sequence; inputting the process state data sequence into the state space model, and constructing a virtual working condition sample; based on the virtual working condition sample, pre-training a reinforcement learning controller and outputting an initial control strategy; inputting the initial control strategy into a lower-layer strategy network, and outputting a wire drawing speed adjusting instruction; the current process state and the wire drawing speed adjusting instruction serve as synchronous input, and an implicit context vector is generated; extracting control strategy characteristics in the edge controller, performing compressed encoding and forming control strategy representation; uploading the control strategy representation to a cloud server, and outputting a unified global control strategy model; and carrying out anomaly detection on the current process state, and if a detected state reconstruction error exceeds an anomaly judgment threshold, triggering safety control logic. The spring steel wire drawing control device realizes spring steel wire drawing control.
Owner:SHAOXING HONGKANG NEW MATERIALS CO LTD

Adaptive control method of analgesic infusion system based on multi-modal feedback

The invention discloses a self-adaptive control method of an analgesia infusion system based on multi-modal feedback, and relates to the field of analgesia infusion automation control, and the method comprises the steps: inputting a feature vector into a multi-modal fusion decision-making layer, and calculating a comprehensive pain index PI by adopting a confidence-weighted self-adaptive fusion algorithm; taking the comprehensive pain index PI, the drug cumulant and the physiological stability index as state input, and inputting the state input into a self-adaptive controller based on reinforcement learning; the adaptive controller outputs a basic infusion rate adjustment amount, a pulse dose adjustment amount, and a locking time adjustment amount. According to the invention, through multi-source signal acquisition and confidence coefficient weighted adaptive fusion, the accuracy and robustness of pain assessment are improved; the hierarchical reinforcement learning controller realizes personalized dose optimization on the premise of safety priority, and gives consideration to both analgesic effect and physiological stability; and the safety monitoring module is combined with multi-modal cross validation and pharmacokinetic prediction, so that the safety risk is effectively reduced.
Owner:WENZHOU PEOPLES HOSPITAL

Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning

The invention discloses a virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning, and the method comprises the steps: building a multi-time scale dynamic model covering second-level frequency response and hour-level economic scheduling through collecting the operation data of a virtual power plant, constructing a mixed sensing network, and taking the multi-time scale dynamic model as the input, extracting multi-scale time sequence features by combining a multi-head hierarchical attention encoder, obtaining equipment dynamic association features through a dynamic heterogeneous graph neural network, and generating comprehensive feature representation; constructing a hierarchical deep reinforcement learning control framework based on comprehensive feature representation, optimizing energy storage charging and discharging power, a load regulation instruction and a unit output strategy by adopting a deep reinforcement learning algorithm based on an experience playback mechanism, and constructing a multi-target reward function; and finally performing training-deployment process and simulation verification. According to the method, the response time of a protection mechanism can be shortened to be within 0.1 second under the working condition that the SOC is lower than 20%, the comprehensive operation cost of the system can be reduced under the normal working condition, and the signal-to-noise ratio of input signals of a control layer is remarkably increased.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1