Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

297 results about "Reinforcement learning control" patented technology

Defect repairing method based on digital twinning and friction stir welding technology

The invention discloses a defect repair method based on digital twinning and friction stir welding technologies, and relates to the technical field of intelligent manufacturing and digital twinning, and the defect repair method comprises the following steps: synchronously capturing full-dimensional data of a welding area through a multi-mode sensor array integrated by an actuator; secondly, segmenting defect boundaries by adopting a deep learning algorithm, constructing a dynamic twin model in combination with thermal-force field coupling simulation, and accurately mapping defect three-dimensional features; generating a repair track according to the twinborn model, converting the repair track into a robot joint instruction through a curved surface parameterization mapping algorithm, and implanting real-time anti-collision constraint; in the repairing process, based on reinforcement learning control of the material rheological resistance and the temperature gradient, the rotating speed, the advancing speed and the down force of the tool are dynamically adjusted; and after repairing, micro-focus CT scanning is started immediately, actually measured data is compared with twinborn prediction, and when the deviation exceeds a threshold value, a re-repairing process is triggered automatically. The method solves the problems that a traditional method depends on manual intervention and the precision of a sensor is easily interfered by the environment.
Owner:SOUTHWEST TECHNICAL ENGINEERING RESEARCH INSTITUTE OF CHINA SOUTH IND GROUP

Backtracking analysis model construction method based on attack chain

The invention relates to the technical field of data processing, in particular to a backtracking analysis model construction method based on an attack chain, which comprises the following steps that: a kernel layer security agent acquires process, file and network behavior characteristics in a hardware isolation environment, and generates an event tuple; the tensor network pipeline performs three-dimensional decoupling mapping on the tuple into a behavior fingerprint vector, an orthogonalization noise feature and an asymmetric adjacent tensor, and compresses the behavior fingerprint vector, the orthogonalization noise feature and the asymmetric adjacent tensor into a space-time topology tensor block; the reinforcement learning controller constructs a directed acyclic graph based on the tensor blocks, calculates connectivity loss and outputs an event risk score; the dynamic routing engine constructs a decision tree model according to the risk mark, the burst frequency and the correlation entropy, and implements three-level shunting and a multiple simulation system to generate an anti-interference index; and when the deviation between the physical trajectory and the digital model exceeds the tolerance, the closed-loop feedback weight coefficient updates the loss function parameter and adjusts the channel resource weight. And the problem of threat discovery delay caused by attack chain breakage under massive events is solved.
Owner:HUANENG INFORMATION TECH CO LTD

Double-arm body operation method of humanoid robot based on reinforcement learning

The invention relates to a humanoid robot double-arm body operation method based on reinforcement learning, belongs to the field of robot cooperative control, and is characterized in that a strategy of trajectory block prediction and time integration fusion is used for double-arm cooperative control, and continuity and stability of double-arm operation are improved; in reinforcement learning control, a three-dimensional pose track generation and correction module is introduced, condition generation and denoising correction of a three-dimensional space are carried out on a track block layer, and geometric consistency and naturalness of a generated track are guaranteed; the invention further provides a reinforcement learning optimization framework and a simulation-reality migration process, and through system integration of reward, value guidance and migration processes, strategy deployability and safety are guaranteed; compared with the prior art, the method has the advantages that the naturalness, the collaboration and the success rate of double-arm operation can be remarkably improved, the generalization ability is high, and the good simulation-to-reality migration ability is achieved.
Owner:CITIC HEAVY INDUSTRIES CO LTD

Mineral processing flow intelligent optimization control method and system based on knowledge graph

The invention discloses a beneficiation process intelligent optimization control method and system based on a knowledge graph, and relates to the technical field of beneficiation process control, and the method comprises the steps: collecting static attribute data and dynamic operation data, constructing a causal knowledge graph, and obtaining a causal relationship between nodes in the causal knowledge graph and an optimization target constraint; based on a causal relationship between nodes in the causal knowledge graph and optimization target constraints, screening data from a historical database to form a high-value historical data set; according to the high-value historical data set, a reward function model is constructed, a strategy function is trained, and a reinforcement learning control strategy is obtained; running the reinforcement learning control strategy, generating a control instruction, and issuing the control instruction to the target equipment to obtain dynamic optimization control; and performing real-time monitoring on the dynamic optimization control, and performing anti-fact reasoning based on a causal knowledge graph when abnormal deviation of the running state is detected to obtain an anti-fact reasoning result.
Owner:CHANGCHUN GOLD DESIGN INST

Battery management system based on adaptive digital twinning

The invention relates to the technical field of BMS and the like, and provides a battery management system based on adaptive digital twinning, a physical layer of the battery management system comprises a battery pack, a sensor network and an edge computing node, and is responsible for data acquisition and preprocessing; the digital twin layer comprises a self-adaptive multi-scale model and a real-time data engine, battery behaviors are dynamically simulated by coupling electrochemical, thermal and aging models, a future state trajectory prediction result is output, and the real-time data engine fuses sensor data, historical data and simulation data to drive model updating; the intelligent decision-making layer comprises a reinforcement learning controller and a fault prediction module which are deployed in a local server, the reinforcement learning controller dynamically optimizes a charging and discharging strategy according to a prediction result of the digital twinborn layer and issues and executes the charging and discharging strategy, and the fault prediction module analyzes multi-source time sequence data based on an LSTM network so as to early warn thermal runaway and short circuit risks in advance. According to the invention, long-term accurate mapping and adaptive adjustment between the battery physical entity and the digital model can be realized.
Owner:深圳市华芯控股有限公司

Autonomous inspection and return control method and system for blow-off pipeline robot

The invention discloses an autonomous inspection and homeward voyage control method and system for a blow-off pipeline robot, belongs to the technical field of intelligent pipeline detection and robot control, and aims to solve the technical problems of how to realize a stable, reliable and continuous clogging exploration task in a complex blow-off pipeline environment and realize safe self-rescue in an abnormal state. According to the technical scheme, the method comprises the following steps: data dimension reduction: receiving high-dimensional environment and motion state data from a multi-modal sensor, performing real-time dimension reduction on the high-dimensional environment and motion state data by adopting a dynamic principal component analysis method, extracting key features of state information of a robot and internal state information of a pipeline, and constructing a low-dimensional state vector; performing deep reinforcement learning control based on a liquid neural network: modeling dynamic environment state change by using the liquid neural network, and realizing path planning and control in combination with a deep deterministic strategy gradient algorithm; tracing the track; and a return flight strategy based on a graph attention network and reinforcement learning.
Owner:浪潮智慧城市科技有限公司

Active power distribution network multi-target collaborative voltage optimization control method based on FACMAC algorithm

The invention relates to a source-containing power distribution network multi-target collaborative voltage optimization control method based on an FACMAC algorithm, and belongs to the technical field of photovoltaic inversion control. According to the technical scheme, a power distribution network physical system is composed of a plurality of feeder lines, a transformer, a line and a plurality of grid-connected photovoltaic inverters, and each inverter can measure operation information such as local voltage and current in real time; the data acquisition and communication system is used for acquiring node operation data and realizing low-delay communication; the multi-agent reinforcement learning control system is composed of a plurality of distributed agents and factorization centralized Critic modules, and whole-network voltage optimization decision can be carried out in training and execution stages. And the execution unit adjusts the reactive power output of the inverter in real time according to the control instruction. According to the method, the whole-network cooperative regulation and control capability is improved, the training efficiency bottleneck in a high-dimensional scene is relieved, the expression capability on a complex nonlinear coupling relationship is enhanced, and efficient, stable and extensible power distribution network voltage optimization control is realized.
Owner:STATE GRID JIBEI ELECTRIC POWER COMPANY LIMITED CHENGDE POWER SUPPLY

Water electrolysis hydrogen production intelligent control system and method based on artificial intelligence

The invention discloses a water electrolysis hydrogen production intelligent control system and method based on artificial intelligence. The system and method are suitable for a large-scale water electrolysis hydrogen production scene under the power supply condition of fluctuating renewable energy sources such as wind power and photovoltaic. The system comprises a data acquisition layer, an edge calculation layer, an intelligent control layer and an execution layer. Dynamic modeling of the running state of the hydrogen production system is achieved through multi-parameter real-time monitoring and feature extraction. The intelligent control layer fuses an LSTM prediction module and a reinforcement learning controller, the LSTM prediction module is used for predicting future renewable energy input and hydrogen demand trends, and the reinforcement learning controller calculates an optimal current density set value based on a prediction result so as to maximize hydrogen production efficiency per unit energy consumption. And meanwhile, a digital twinning technology and a safety protection mechanism are combined, thermoelectric dual regulation and control of the electrolysis process are achieved, and the control response speed and the system stability are improved. According to the invention, energy consumption can be effectively reduced, the service life of the stack is prolonged, and the operation efficiency and reliability of the hydrogen production system under complex load are improved.
Owner:BEIJING MINGYANG HYDROGEN ENERGY TECHNOLOGY CO LTD

Cross-scene defect identification method and system driven by migration and fine tuning

The invention discloses a migration fine tuning driven cross-scene defect identification method and system, and relates to the technical field of electric power inspection, and the method comprises the steps: obtaining a basic model through employing a general electric power image data set; constructing a triple set, and forming a power defect domain knowledge graph; performing style conversion on the general power image by adopting CycleGAN, and performing semantic consistency verification on an image sample through a triple set; forming a state vector; inputting the state vector into a reinforcement learning controller, and performing fine tuning training on the basic model to obtain a fine tuning defect identification model; and evaluating the performance of the fine tuning defect identification model on the target domain image verification set, and optimizing a migration strategy selection mechanism according to a performance index. According to the method, the problem of scarcity of target domain labeling data can be solved, sufficient data support is provided for model fine adjustment, model overfitting caused by insufficient data is avoided, the model can adapt to a visual style and a defect distribution rule of a target scene, and the generalization ability of cross-scene defect recognition is improved.
Owner:INFORMATION & COMMNUNICATION BRANCH STATE GRID JIANGXI ELECTRIC POWER CO

Rapid training type unmanned aerial vehicle reinforcement learning strategy method suitable for microcontroller

The invention discloses a fast training type unmanned aerial vehicle reinforcement learning strategy method suitable for a microcontroller, and belongs to the technical field of unmanned aerial vehicle control, and the method comprises the steps: constructing a standard four-rotor dynamic model containing the first-order delay characteristic of a motor, and simulating a real physical system; an asymmetric actor-commentator architecture is adopted to ensure the robustness and mobility of the strategy; a reward function coefficient is dynamically adjusted based on a course learning mechanism, and network parameters are optimized by adopting an offline strategy algorithm; performing network training and process optimization; and carrying out lightweight processing and deployment on the trained strategy network, landing the strategy in simulation to a real hardware platform with limited resources, jointly realizing rapid training, reliable migration and lightweight deployment, and realizing high-frequency autonomous flight control of the unmanned aerial vehicle on an embedded platform. According to the method, the high-performance reinforcement learning control strategy can be stably and efficiently operated on the microcontroller with limited resources, and the hardware cost and the power consumption of the intelligent unmanned aerial vehicle are greatly reduced.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Area inspection robot autonomous navigation and path planning method based on multi-modal perception

The invention relates to the technical field of autonomous navigation of inspection robots, and provides a field inspection robot autonomous navigation and path planning method based on multi-modal sensing. The method comprises the following steps: collecting a laser radar point cloud, a camera image, inertial measurement and positioning data, obtaining an environment semantic feature set through spatio-temporal feature fusion, and generating a semantic occupation map; inputting the current position of the robot, the target point and the semantic occupation map into a trajectory generation network to obtain candidate trajectories meeting obstacle avoidance and path smoothness constraints, and completing task sorting and trajectory splicing in combination with inspection task points to form a global path; in the operation process, the reinforcement learning control model adjusts the linear speed and the angular speed in real time, and path tracking and dynamic obstacle avoidance are achieved. The navigation precision and the operation safety of the inspection robot in the complex field area are improved.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD

Intelligent scheduling system for steel plate hot dipping production process

The invention relates to the technical field of steel plate hot-dip production, and discloses an intelligent scheduling system for a steel plate hot-dip production process. The system comprises a working condition sensing module which is used for continuously collecting steel belt inlet temperature, zinc liquid component concentration fluctuation data, air knife pressure time sequence parameters and roller system operation state signals of a galvanizing pot area; the feature extraction module receives the working condition data flow, analyzes the steel strip surface temperature field distribution features and the zinc liquid viscosity change rule, and generates a process feature spectrum; the parameter optimization module dynamically adjusts membership function parameters and rule base weight coefficients of the fuzzy reinforcement learning controller according to the process characteristic spectrum, and outputs an optimized plating layer control parameter set; the self-adaptive plating layer control module generates a plating layer thickness control instruction based on the real-time zinc liquid temperature parameter and the optimized plating layer control rule set; and the dynamic roller system scheduling module realizes global cooperative control of the roller system speed based on the instruction. The system can realize intelligent scheduling of the steel plate hot-dip production process.
Owner:HANGZHOU LONGYAO POWER PARTS CO LTD

Unmanned ship and unmanned aerial vehicle system cooperative control method based on hydrodynamic parameters and deep reinforcement learning, electronic equipment and readable storage medium

The invention belongs to the technical field of unmanned system control, and particularly discloses an unmanned ship and unmanned aerial vehicle system cooperative control method based on hydrodynamic parameters and deep reinforcement learning, electronic equipment and a readable storage medium, and the method comprises the steps: S1, obtaining the real-time position, speed and external environment information of each unmanned ship and unmanned aerial vehicle; s2, modeling the hydrodynamic system of each unmanned ship and unmanned aerial vehicle to obtain a dynamic model; s3, performing linear parameterization on the kinetic equations of each unmanned ship and each unmanned aerial vehicle; s4, constructing a deep reinforcement learning control algorithm for improving artificial potential field obstacle avoidance of the unmanned ship and the unmanned aerial vehicle; s5, constructing a network cooperation algorithm of the unmanned ship and the unmanned aerial vehicle; and S6, initializing hardware modules of each unmanned ship and each unmanned aerial vehicle, and loading a network cooperation algorithm into micro-control units of each unmanned ship and each unmanned aerial vehicle.
Owner:JINING UNIV +1

Spring steel wire drawing control method based on reinforcement learning

The invention discloses a reinforcement learning-based spring steel wire drawing control method, which comprises the following steps of: acquiring real-time process parameters to form a process state data sequence; inputting the process state data sequence into the state space model, and constructing a virtual working condition sample; based on the virtual working condition sample, pre-training a reinforcement learning controller and outputting an initial control strategy; inputting the initial control strategy into a lower-layer strategy network, and outputting a wire drawing speed adjusting instruction; the current process state and the wire drawing speed adjusting instruction serve as synchronous input, and an implicit context vector is generated; extracting control strategy characteristics in the edge controller, performing compressed encoding and forming control strategy representation; uploading the control strategy representation to a cloud server, and outputting a unified global control strategy model; and carrying out anomaly detection on the current process state, and if a detected state reconstruction error exceeds an anomaly judgment threshold, triggering safety control logic. The spring steel wire drawing control device realizes spring steel wire drawing control.
Owner:SHAOXING HONGKANG NEW MATERIALS CO LTD

Adaptive control method of analgesic infusion system based on multi-modal feedback

The invention discloses a self-adaptive control method of an analgesia infusion system based on multi-modal feedback, and relates to the field of analgesia infusion automation control, and the method comprises the steps: inputting a feature vector into a multi-modal fusion decision-making layer, and calculating a comprehensive pain index PI by adopting a confidence-weighted self-adaptive fusion algorithm; taking the comprehensive pain index PI, the drug cumulant and the physiological stability index as state input, and inputting the state input into a self-adaptive controller based on reinforcement learning; the adaptive controller outputs a basic infusion rate adjustment amount, a pulse dose adjustment amount, and a locking time adjustment amount. According to the invention, through multi-source signal acquisition and confidence coefficient weighted adaptive fusion, the accuracy and robustness of pain assessment are improved; the hierarchical reinforcement learning controller realizes personalized dose optimization on the premise of safety priority, and gives consideration to both analgesic effect and physiological stability; and the safety monitoring module is combined with multi-modal cross validation and pharmacokinetic prediction, so that the safety risk is effectively reduced.
Owner:WENZHOU PEOPLES HOSPITAL

Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning

The invention discloses a virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning, and the method comprises the steps: building a multi-time scale dynamic model covering second-level frequency response and hour-level economic scheduling through collecting the operation data of a virtual power plant, constructing a mixed sensing network, and taking the multi-time scale dynamic model as the input, extracting multi-scale time sequence features by combining a multi-head hierarchical attention encoder, obtaining equipment dynamic association features through a dynamic heterogeneous graph neural network, and generating comprehensive feature representation; constructing a hierarchical deep reinforcement learning control framework based on comprehensive feature representation, optimizing energy storage charging and discharging power, a load regulation instruction and a unit output strategy by adopting a deep reinforcement learning algorithm based on an experience playback mechanism, and constructing a multi-target reward function; and finally performing training-deployment process and simulation verification. According to the method, the response time of a protection mechanism can be shortened to be within 0.1 second under the working condition that the SOC is lower than 20%, the comprehensive operation cost of the system can be reduced under the normal working condition, and the signal-to-noise ratio of input signals of a control layer is remarkably increased.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1

Interpretable high-speed ramp intelligent driving vehicle confluence control method

The invention discloses an interpretable high-speed ramp intelligent driving vehicle confluence control method. The method comprises the steps that a state space is constructed, wherein the state space comprises the position and speed of a vehicle, lane information and the relative position and speed of adjacent vehicles; human-like decision guidance: training a human-like behavior discriminator based on real driving data by using a random forest model, screening action selection of reinforcement learning, and constructing a sample set containing state vectors and human-like action labels; reinforcement learning control; and carrying out interpretability analysis. Through an SHAP analysis method, an explainable DQN longitudinal acceleration control model in an automatic driving vehicle ramp merging process is provided, and an explainable automatic driving vehicle ramp merging control model is constructed in combination with an merging decision-making model constructed based on a random forest. According to the explainable high-speed ramp automatic driving vehicle confluence control method based on the confluence urgency degree and the environmental historical information, traffic conflicts during highway confluence can be relieved, and therefore the road passing efficiency is effectively improved.
Owner:CHANGAN UNIV

Multi-unmanned aerial vehicle confrontation task execution method and device based on model reinforcement learning, and medium

The invention discloses a multi-unmanned aerial vehicle confrontation task execution method and device based on model reinforcement learning, and a medium, and the method comprises the following steps: inputting the current local observation into a reinforcement learning control network for each airframe in an own unmanned aerial vehicle group, and outputting meta-actions (including movement and attack), performing confrontation interaction with the enemy unmanned aerial vehicle group to obtain finite confrontation experience data; the collected historical interaction experience is used for supervised learning of a world model, so that the world model can approach probability distribution of an interaction rule of a real environment; data expansion is carried out on real interaction data by using a world model, an action network and an evaluation network are trained on expanded empirical data, and strategy optimization and parameter updating of a multi-unmanned aerial vehicle system are realized; repeating the process, and sequentially carrying out experience collection, world model training and enhancement strategy updating; and finally, the trained action network is reserved and is used for outputting a specific action strategy in an actual confrontation task.
Owner:ZHEJIANG UNIV

Robot motion control optimization method, device, equipment and medium

The invention discloses a robot motion control optimization method, device, equipment and medium, and the method comprises the steps: inputting task description in a natural language form, robot state information, a historical execution track and environment feedback information into a preset large language model; through the large language model, selecting a target motion strategy of the robot according to the task description; when it is detected that the task of the robot fails or the action of the robot is abnormal, a reward function used for controlling a deep reinforcement learning controller of the robot is reconstructed; and when the state of the robot is abnormal, parameters of the deep reinforcement learning controller are adjusted and optimized, the motion stability, the task completion efficiency and the strategy adaptability of the robot in a high-risk, high-temperature and high-complexity scene are remarkably improved, and the self-adaptive adjusting and optimizing capability and the abnormity recovery capability of the robot in a dynamic environment are enhanced.
Owner:HUANENG SHANTOU HAIMEN POWER GENERATION CO LTD +1

Pilotless automobile queue control method based on physical information reinforcement learning

The invention discloses an unmanned vehicle queue control method based on physical information reinforcement learning, and the method comprises the steps: building a reinforcement learning environment of an unmanned vehicle queue, and enabling a queue following vehicle to plan the acceleration according to the acceleration change of a leading vehicle; state information of each vehicle is obtained, a TD3 algorithm is used for carrying out training iteration on the model, in the training iteration process, the expected acceleration obtained through calculation of the CACC model is embedded into a loss function of a strategy network Actor to serve as physical information constraint, and a strategy is guided to generate an action conforming to vehicle dynamics; and after training iteration is finished, deploying a control strategy of the strategy network Actor in an unmanned motorcade simulation environment to realize cooperative control of the motorcade. According to the method, physical information constraints are introduced, so that the model training speed is increased, the stability and interpretability of a reinforcement learning control strategy are improved, and the queue operation efficiency and stability are enhanced.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Biped robot reinforcement learning control method and system based on historical action sequence

The invention relates to the technical field of biped robot reinforcement, and discloses a historical action sequence-based biped robot reinforcement learning control method and system. The biped robot reinforcement learning control method is applied to reinforcement learning control equipment and specifically comprises the following steps that S101, a floating base frame is adopted to load a biped robot URDF model in an Isaac Gym simulation platform, ten active joint freedom degrees are configured, and a control interface is constructed to output a 10-dimensional motor position command vector; the motion robustness and adaptability of the biped robot in a dynamic environment are remarkably improved by constructing a Transform time sequence modeling architecture fusing a historical state-action sequence, and the historical observation and action sequence is combined with sine-cosine position coding to be input into a multi-layer Transform encoder, so that the motion robustness and adaptability of the biped robot in the dynamic environment are improved. The self-attention mechanism is used for capturing the long time sequence dependence relation of foot-ground contact force change, topographic relief and the like, and the problem of gait instability caused by limited local observation in a traditional method is solved.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Lithium ion battery thermal management method based on physical perception and entropy collaborative multi-agent

A lithium ion battery thermal management method based on physical perception and entropy cooperation multiple agents comprises the steps that a lithium ion battery electric-thermal coupling model containing data driving compensation is established, and a full-state space system equation describing the dynamic characteristics of a battery is established; constructing a TCN-Transformer hybrid neural network fused with a physical constraint mechanism, and carrying out online identification on parameters in the electric-thermal coupling model by adopting the network to obtain real-time parameters; based on the full-state space system equation and the real-time parameters, constructing a minimum error entropy adaptive extended Kalman filter optimized by an entropy cooperative multi-agent flexible Actor-Critic algorithm, and performing joint estimation on the state of charge of the battery and the temperature of the battery; and based on a joint estimation result, constructing a TD3 deep reinforcement learning control algorithm embedded with a microsecurity layer, and realizing the self-adaptive thermal management direct control of the lithium ion battery under rule guidance through the algorithm. According to the invention, accurate, effective and safe thermal management control of the lithium ion battery is realized.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Grinding process control method based on reinforcement learning

The invention discloses a milling process control method based on reinforcement learning, which comprises the following steps: collecting and processing multi-source process data in real time, normalizing and extracting features, and forming state feature vectors; the state feature vectors are input into a DSAC-T reinforcement learning controller, the opening degree of a feeding valve and the rotation speed adjustment of a main shaft are used as action spaces, a reinforcement learning strategy network is constructed, and a process drift online detection sub-module is integrated; training a controller to maximize the rice yield, minimize the broken rice rate and energy consumption, and combining the process safety index to obtain an optimal strategy; deploying the trained controller in an actual scene, and outputting an adjustment instruction to implement closed-loop control; when the actual rice grain size or broken rice rate exceeds the model prediction deviation, online incremental learning is automatically triggered, and control parameters are optimized; the intelligent, self-adaptive and safe control of the rice milling process is realized. According to the method, adaptive optimization control of the rice milling process is realized, and the method has efficient response capability to dynamic disturbance and abnormal states.
Owner:TANGSHAN CITY CAO THE CAOFEIDIAN AREA WO METER CO LTD

Dynamic power optimization method and system based on bidirectional multi-agent reinforcement learning

The invention discloses a dynamic power optimization method and system based on bidirectional multi-agent reinforcement learning, and the method comprises the steps: collecting space-time joint channel data, an interference prediction result and neighbor base station feedback information, and constructing a base station multi-dimensional state space through a bidirectional real-time sensing mechanism; inputting the multi-dimensional state space of the base station into a bidirectional multi-agent reinforcement learning control module so as to carry out collaborative learning through a DEI-MADDPG algorithm, and constructing an output power adjustment strategy based on a collaborative Actor-Critic; generating a physical uplink shared channel control instruction based on a reinforcement learning strategy output result, and issuing the control instruction to realize dynamic allocation of base station transmitting power; and monitoring a network state and user feedback, executing strategy weight adjustment according to monitoring data, and returning to the step of constructing the base station multi-dimensional state space to realize base station transmission dynamic power optimization iteration closed-loop control. According to the invention, the transmitting power of the base station can be adaptively adjusted, and the whole SINR of the system can be improved while the same-frequency interference is suppressed.
Owner:CHONGQING UNIV OF TECH

Self-adaptive optical system parameter adjusting method and device combining offline pre-training and online reinforcement learning

The invention discloses a self-adaptive optical system parameter adjusting method and device combining offline pre-training and online reinforcement learning, and belongs to the field of intelligent control combining self-adaptive optics and reinforcement learning. Multi-modal state modeling, off-line strategy pre-training and an online strategy enhancement mechanism guided by behavior advantages are introduced, and key control parameters in the AO system are dynamically optimized by using a reinforcement learning algorithm. According to the method, multi-modal heterogeneous data can be effectively integrated, the dynamic modeling and prediction capability of the system is improved, and stable, efficient and robust control strategy updating is realized.
Owner:SOUTHWEST JIAOTONG UNIV

Motion control method and device based on reinforcement learning

The invention provides a motion control method and device based on reinforcement learning, electronic equipment and a storage medium, and relates to the technical field of robot motion control. The motion control method based on reinforcement learning comprises the following steps: training and generating a teacher network model corresponding to a robot according to a robot model description file, road surface detection sample data and motion performance information, and determining student network training initial parameters corresponding to the teacher network model; training and generating a corresponding student network model according to the student network training initial parameters and the road surface detection sample data; deploying the student network model to the robot, analyzing a real-time image collected by the robot based on the road surface type classification detection network model, and determining a corresponding road surface type; and determining the target joint position offset of the robot based on the road surface type corresponding to the real-time image, the current road surface detection data and the student network model, and determining the joint torque of the robot according to the target joint position offset. The effect of controlling robot motion based on reinforcement learning is improved.
Owner:NR ELECTRIC CO LTD +2

An air-cooled UAV hydrogen fuel cell control system

The present application relates to the technical field of UAV power systems, and in particular to an air-cooled UAV hydrogen fuel cell control system. It includes: a spray intercooler that sprays stack wastewater as a cooling medium; a wastewater circulation system that transports stack wastewater to the spray nozzle of the spray intercooler; and a reinforcement learning control system that executes a control program to control the spray intercooler and the wastewater circulation system. The closed-loop wastewater spray cooling system significantly reduces the damage to the membrane electrode caused by the high-temperature intake air of the air compressor. At the same time, the reinforcement learning algorithm is used to achieve four-dimensional coordinated control of temperature, humidity, pressure, and flow, so that the stack always operates stably in the optimal operating range. This not only compresses the actuator response delay to the millisecond level to ensure the real-time performance of the UAV power output, but also significantly extends the service life of the stack and improves the stability of the output power while reducing the weight of the system.
Owner:SHANGHAI DAFENG TECH CO LTD

Linear motor temperature control method and system based on air cooling heat dissipation

The invention provides a linear motor temperature control method and system based on air cooling heat dissipation, and relates to the technical field of motor temperature control, and the method comprises the steps: dividing a linear motor into a plurality of temperature control regions, and configuring an independent control execution unit for each region; predicting a temperature trend based on the LSTM; a multi-agent reinforcement learning control network is constructed, and each agent calculates an optimal control strategy based on deep Q learning; generating an overall airflow regulation strategy through a collaborative optimization algorithm; dynamically adjusting the opening of the electric adjusting valve and the rotating speed of the fan. Accurate temperature control is achieved, the heat dissipation efficiency is improved, and the service life of equipment is prolonged.
Owner:SHENZHEN FERGUS ELECTROMECHANICAL EQUIP CO LTD

Metacosmic virtual-real interaction method based on causal invariance

The invention provides a meta-universe virtual-reality interaction method based on causal invariance, belongs to the field of meta-universe virtual-reality interaction technologies, and is used for solving the problems of poor cross-domain adaptation, insufficient long-tail scene coverage and low dynamic robustness in related technologies. According to the method, through sensing layer causal kernel quality evaluation and data enhancement, reasoning layer causal invariance learning and cross-domain parameter migration, decision layer causal attention intention alignment, execution layer causal reinforcement learning control and iteration layer double-threshold knowledge updating, full-link parameter collaborative circulation is realized in combination with a causal data interface; and finally, the accuracy and the stability of virtual-real interaction of the element universe are improved, and efficient adaptation of cross-domain, long-tail and dynamic scenes is realized.
Owner:MATERIAL CHAIN CORE ENGINEERING TECHNOLOGY RESEARCH INSTITUTE (BEIJING) CO LTD +2

Layered autonomous behavior decision and joint control method for expressway and ramp scenes

ActiveCN121572982AArtificial lifeControl devicesDecision makerAutonomous behavior
The invention provides a hierarchical autonomous behavior decision and joint control method for expressway and ramp scenes, and belongs to the technical field of auxiliary driving. Comprising the steps of S1, a training stage; s1.1, collecting expert data and training an IL decision maker; s1.2, an RL control module is trained; s1.3, performing combined overall verification and fine adjustment; s2, an execution stage; s2.1, environment perception and state construction are carried out; s2.2, performing upper layer imitation learning decision making; s2.3, calling a lower layer reinforcement learning control module; s2.4, a lower layer MPC control module is called; and S2.5, vehicle execution and environment updating are carried out. According to the invention, control mode switching is more adaptive and learnable, and dependence on artificial experience is reduced; strategy flexibility and control feasibility or safety are considered; complex tasks such as highway main road lane changing, ramp confluence and ramp diversion can be processed under the same framework; according to the method, smooth and comfortable control behaviors can be generated on the premise of ensuring the safety margin.
Owner:SHENZHEN AUTOMOTIVE RES INST BEIJING INST OF TECH (SHENZHEN RES INST OF NAT ENG LAB FOR ELECTRIC VEHICLES) +1