Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

73 results about "Q-learning" patented technology

Q-learning is a model-free reinforcement learning algorithm. The goal of Q-learning is to learn a policy, which tells an agent what action to take under what circumstances. It does not require a model (hence the connotation "model-free") of the environment, and it can handle problems with stochastic transitions and rewards, without requiring adaptations.

Outbound routing method and system based on multi-dimensional index coupling and large model decision

The invention provides an outbound routing method and system based on multi-dimensional index coupling and large model decision, and the method comprises the steps: carrying out the preprocessing of the historical data of each line provider, predicting the future performance index through an LSTM neural network, and optimizing the prediction result through a Kalman filter, and eliminating the noise. In the route decision-making link, the system calculates the Q value of each line provider based on the deep Q learning network, comprehensively considers the call completing rate, average cost, negative feedback rate and other factors of each line provider, and intelligently selects the optimal line provider in combination with the user portrait, scene recognition and real-time load state. The system also constructs a complete feedback closed loop, continuously optimizes the model by collecting call result data, and automatically adjusts decision parameters. The call completing rate and the service quality of the outbound service are improved, and the communication cost is reduced.
Owner:BEIJING YULORE INNOVATION TECH

Mobile robot adaptive path planning method based on enhanced Q learning and multi-strategy cooperation

The invention discloses a mobile robot adaptive path planning method based on enhanced Q learning and multi-strategy cooperation, and relates to the field of robot autonomous navigation and computing intelligence. The method comprises the following steps: firstly, establishing a two-dimensional grid map model containing obstacle information through an environment sensing module; the initial position of a gold scrubber population is generated by utilizing the low-difference characteristic of a golden section sequence to enhance the global ergodicity of a solution space, three stages of migration, gold scrubber and cooperation of a gold scrubber optimization algorithm are modeled into a discrete action space, an optimal search action is dynamically selected through a nonlinear attenuation greedy strategy, and the optimal search action is realized. According to the method, adaptive switching of exploration and development behaviors is realized, a search gradient is constructed by using a differential vector of survival of the survival of the survival of the population, and a lens imaging reverse learning strategy is introduced to dynamically disturb an elite individual so as to jump out of a local extremum trap. And finally, outputting an optimal collision-free path through a multi-objective evaluation function including path length, safety and smoothness.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Road investigation design method and system based on mobile network

The invention discloses a road investigation design method and system based on a mobile network, and relates to the technical field of road investigation, and the method comprises the steps: dividing geological anomaly probability distribution data into dynamic grid units, constructing a reinforcement learning state space, and generating an optimal sampling path instruction in combination with a dual-network deep Q learning architecture; analyzing the optimal sampling path instruction into executable parameters, executing the executable parameters, collecting multi-source spatio-temporal data, establishing data association through a spatio-temporal hash algorithm, and generating a reconnaissance data set; and on the basis of the survey data set, through parameterized spline curve modeling, generating a candidate road design scheme, and in combination with a non-dominated sorting genetic algorithm, performing optimization to generate a three-dimensional road design scheme. According to the method, the exploration path is optimized by using the dynamic grid coding and the dual-network architecture, the high-risk area coverage and the moving efficiency are balanced, meanwhile, the method adapts to real-time geological changes in combination with priority experience playback, and the exploration reliability and the resource utilization rate are improved.
Owner:JIANGSU SHIGUANG GEOGRAPHIC INFORMATION TECH CO LTD

Multi-edge task unloading method combined with task pre-sorting

The invention relates to a multi-edge task unloading method combined with task pre-sorting, and belongs to the technical field of edge calculation and task unloading. The method comprises the following steps: firstly, establishing a multi-edge, multi-user and multi-task network architecture comprising a cloud layer, an edge layer and a user layer; based on network architecture support, an application model, a transmission model and an execution model for an application task of a user are established, the application model represents any subtask as a triple, the transmission model defines a corresponding transmission rate, and the execution model comprises a local execution model, an edge execution model and an overall cost model; a task pre-sorting mechanism is introduced, and subtasks which can be parallel are sorted in a descending mode according to priority factors; and finally, based on the ordered task sequence after sorting, performing task unloading strategy decision by adopting a deep Q learning network. According to the method, joint optimization of task execution delay and energy consumption can be realized in a multi-user, multi-task and multi-edge node computing environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Slope cutting house building landslide vulnerability analysis method and system based on reinforcement learning

The invention discloses a slope-cutting house-building landslide vulnerability analysis method and system based on reinforcement learning. The method comprises the following steps: preprocessing multi-source data to obtain a disaster sample data set, a missing tag vector and a damage level tag; feature parameters of the slope and the house are extracted, and missing mark vectors are fused to generate a state vector set; constructing an action set and calculating an external return according to the difference between the predicted damage distribution and the damage level label; calculating a conservative penalty quantity for non-data actions by utilizing the action value network and the target action value network in conservative Q learning; calculating a novelty error in the random network distillation and generating an intrinsic return; and fusing the return to construct a state transition sample, updating the action value network to obtain a conservative action value function, and outputting a vulnerability conclusion of the scene to be analyzed. According to the method, stable training and iteration of the evaluation model are realized under the conditions of sample scarcity and data incompleteness.
Owner:HOHAI UNIV

A radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method

The application discloses a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method, first establishes a confrontation scene under the radar and multi-jammer environment, then sorts out the model parameters and working modes of the radar and the jammer, determines the evaluation index to complete the evaluation of the radar scheduling strategy, and then converts the problem into a Markov decision process, and uses Q learning to solve the optimal strategy of radar multi-domain resource scheduling. The optimization strategy obtained by the method has strong adaptability and good model scalability, and can be adjusted according to the actual need of the scheduled parameter resource. Through the collaborative control of the transmission parameters between the multiple domains of the radar, compared with the existing single-domain resource scheduling anti-jamming method, the radar active anti-jamming capability can be significantly enhanced while improving the task performance, and the potential task performance of the radar is excavated. Meanwhile, Q learning is used to optimize the radar node multi-domain joint active anti-jamming strategy, and the survival ability of the radar in the electronic countermeasure environment is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Unmanned ship closed-loop path planning method and system based on multi-agent potential field and deep Q learning

The invention relates to the technical field of artificial intelligence, in particular to an unmanned ship closed-loop path planning method and system based on a multi-agent potential field and deep Q learning. According to the method, a multi-agent potential field including an attraction potential field and a repulsion potential field is constructed by using state information of multiple unmanned ships, and an action evaluation function is formed based on a potential field gradient and an epsilon-greedy strategy. And a single-boat target optimal action strategy is obtained by establishing a double-Q dynamic adjustment network model and combining deep Q learning iterative updating. On the basis, a collaborative strategy sharing mechanism of potential field gradient, position relation and state interaction among multiple agents is introduced, and cluster-level collaborative path planning is achieved. And finally, driving the unmanned ship to complete forward and return tasks according to a collaborative planning result, and generating a closed-loop optimal path in combination with a layered reward mechanism in a return stage. According to the invention, the autonomous decision-making capability, the whole-process autonomous navigation capability and the global cooperation efficiency of the unmanned ship cluster in a complex task sea area are improved.
Owner:JIMEI UNIV

An aircraft fault-tolerant control method based on reinforcement learning

The application discloses an airplane fault tolerance control method based on reinforcement learning and belongs to the field of fault control. The application uses a deep Q learning network to extract features of data generated in a simulation environment, so as to identify fault information and adjust the deep Q learning network. The fault diagnosis method based on data reduces the difficulty of modeling a complex system, replaces a fault detection subsystem by extracting data features, and simplifies the deep Q learning network. The application expands the collected signals by using an enhanced algorithm SDA and then expands the training set, so as to improve the accuracy of the model and the generalization ability of the model. The application optimizes the reinforcement learning model by using the deep Q learning network, and compared with a traditional reinforcement learning method, the optimization method significantly improves the data feature extraction ability of the model.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Aviation luggage online stacking method and device based on deep reinforcement learning

The invention discloses an air luggage online stacking method and device based on deep reinforcement learning, and belongs to the technical field of air luggage stacking. The method comprises the following steps: training a dual-depth Q learning network by using a plurality of training samples; the training mode comprises the steps that a first luggage queue and a height map of a training sample are input into a network, the network constructs a state space based on the height map and the first luggage queue, to-be-stacked luggage is selected from the first luggage queue one by one, and for the to-be-stacked luggage selected each time, the following steps are executed: the action space of the to-be-stacked luggage is determined, and the action space of the to-be-stacked luggage is determined; updating the state space by using the action space; and the stacking space utilization rate is determined, and the stacking mode is output based on the maximum space utilization rate after iterative updating. According to the method, the space utilization rate of luggage packing can be greatly improved, real-time decision making can be carried out in a complex dynamic online environment, the method can adapt to heterogeneous luggage of different sizes, shapes and packing quantities, and robustness is guaranteed.
Owner:CIVIL AVIATION FLIGHT UNIV OF CHINA

Path planning method and device and electronic equipment

The invention relates to the technical field of artificial intelligence and automatic control, and discloses a path planning method and device and electronic equipment, and the method comprises the steps: carrying out the dynamic adjustment of a learning rate and an exploration rate in a Q learning model based on the Q learning model; on the basis of Monte Carlo tree search, path selection in the Q learning model is simulated, and an accumulated discount reward of a simulated path is obtained; based on the accumulated discount rewards, correcting the Q learning model to obtain a path planning correction model; and planning the path of the intelligent agent based on the path planning correction model. According to the invention, the convergence speed of intelligent agent path planning can be improved, the dynamic obstacle avoidance efficiency is enhanced, and the real-time decision-making capability of the intelligent agent in a complex environment is improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Hierarchical clustering Q learning-based ship equipment spare part demand prediction method

The invention discloses a hierarchical clustering Q learning-based ship equipment spare part demand prediction method. The method comprises the following steps of 1) obtaining historical demand data of ship equipment spare parts; 2) preprocessing the collected data set, and identifying and processing missing values and abnormal values in the data; 3) dividing the data sets into a training data set and a test data set, and respectively performing clustering operation on the two data sets; 4) constructing a spare part prediction model; 5) training a spare part prediction model of Q learning by using the training data set; and 6) based on to-be-predicted ship spare part demand data, using the trained model to carry out subsequent ship spare part demand data prediction. According to the hierarchical clustering Q learning-based spare part demand prediction method provided by the invention, the optimal prediction model can be automatically selected according to different spare part feature distribution, the prediction error of the spare part demand is reduced, and the purpose of reducing the spare part inventory cost is achieved.
Owner:CHINA SHIP DEV & DESIGN CENT

Heterogeneous resource collaborative optimization method based on parallel reinforcement learning

The invention discloses a heterogeneous resource collaborative optimization method based on parallel reinforcement learning, and the method comprises the steps: formalizing a heterogeneous resource distribution problem into a Markov decision process, defining a state space, an action space, a reward function and a state conversion mechanism, and constructing a dynamic decision model based on Q learning; classifying the task targets and the heterogeneous resources according to the types of the task targets to obtain a task target set and a heterogeneous resource set corresponding to the task target types; for each task target type, in combination with the corresponding task target set and the heterogeneous resource set, performing parallel reinforcement learning training by adopting a dynamic decision model based on Q learning to obtain an optimal dynamic decision model of each task target type; and based on the optimal dynamic decision model of each task target type, obtaining an optimal resource allocation scheme of each task target type, and finally summarizing to form a global optimal resource allocation scheme. According to the invention, the problem of large-scale multi-domain collaboration can be efficiently solved.
Owner:WUHAN UNIV OF TECH

An automatic nesting method for irregular pieces based on deep reinforcement learning

The application discloses an automatic layout method for irregular cutting pieces based on deep reinforcement learning. The method comprises the following steps: firstly, preprocessing each batch of original irregular cutting pieces to obtain corresponding preprocessed irregular cutting pieces, and arranging the preprocessed irregular cutting pieces in descending order according to areas to obtain an initial layout sequence; then, solving and obtaining a critical polygon and an inner abutting rectangle of the current batch; then, based on the initial layout sequence of the current batch of cutting pieces, the preprocessed irregular cutting pieces of the current batch are laid out to obtain an initial layout result and an initial length of a rectangular plate; finally, the initial layout result is iteratively optimized by using a deep recurrent Q learning network model, and the length of the rectangular plate is continuously reduced until the layout time exceeds a specified time, and a final layout result is output. The application is aimed at a complex decision-making task with a large-scale state and action space, such as two-dimensional irregular cutting piece layout, and realizes automatic layout of irregular cutting pieces, which is efficient and stable in an actual production environment.
Owner:ZHEJIANG UNIV

A double-constraint reinforcement fine-tuning method and system for an ultrasound robot VLA model

PendingCN122414316AData setAlgorithm
The application discloses a double-constraint reinforcement fine-tuning method and system for an ultrasonic robot VLA model, and relates to the field of embodied intelligence. The double-constraint reinforcement fine-tuning method first uses an expert demonstration data set to perform offline initialization on a VLA strategy network and a value network, and obtains an offline basic strategy through offline consistency policy Q learning optimization. When the robot is online exploring, a time-series value model is independently trained, and sparse final task rewards are converted into dense time-series rewards for each step. Then, a proximal policy optimization framework is adopted, a consistency constraint is introduced, and the online strategy is limited to deviate from the basic strategy through KL divergence, so that the safe exploration range of the probe contact process is ensured. Meanwhile, the contact force is monitored in real time, and a penalty and manual reset are triggered when the contact force exceeds a limit. The application solves the problems that the existing VLA model is difficult to utilize sparse rewards and lacks safe constraints in ultrasonic application, and improves the learning efficiency and stability under the premise of ensuring operation safety.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

Aviation heterogeneous network airborne dynamic switching method based on Q learning

The invention discloses an aviation heterogeneous network airborne dynamic switching method based on Q learning, and aims to solve the switching decision problem caused by high-speed movement and network dynamic change of a civil aircraft in a heterogeneous network composed of VHF, LDACS, GEO satellites and LEO satellites. The method comprises the following steps: firstly, predicting an aircraft trajectory by adopting a gating circulation unit neural network, and screening candidate links in combination with network node data; then, communication quality indexes of candidate links are obtained through real-time monitoring, and predictive indexes are calculated according to aircraft trajectories; constructing a comprehensive utility function by applying an analytic hierarchy process, and performing quantitative evaluation on each index; and finally, on the basis of a Q learning decision model, designing two intelligent switching schemes, namely a centralized intelligent switching scheme and a distributed intelligent switching scheme, and dynamically selecting an optimal communication link. According to the method, through active prediction and intelligent decision making, the communication interruption risk is remarkably reduced, and seamless, reliable and efficient switching of the communication links of the aircraft in the whole air route range is ensured.
Owner:BEIHANG UNIV

Deep Q learning cloud task scheduling method based on improved exploration strategy

The invention discloses a deep Q learning cloud task scheduling method based on an improved exploration strategy. The deep Q learning cloud task scheduling method comprises the steps of obtaining a to-be-scheduled cloud task, inputting the to-be-scheduled cloud task into a cloud task scheduling model, outputting a scheduling result, and performing cloud task scheduling based on the scheduling result. Wherein the cloud task scheduling model is obtained through training of a training set, and the cloud task scheduling model is constructed through a deep Q learning network. According to the invention, intelligent resource scheduling, task allocation and load balancing can be realized, so that the performance and efficiency of the system are improved.
Owner:BAICHENG NORMAL UNIV

Scheduling method for distributed assembly wait-free flow shop based on Q learning

The invention discloses a scheduling method of a distributed assembly wait-free flow shop based on Q learning. The scheduling method comprises the following steps: 1, constructing an objective function of a distributed assembly wait-free flow shop model; 2, constructing constraint conditions of a distributed assembly wait-free flow shop scheduling model; and 3, solving the distributed assembly wait-free flow shop scheduling model by using a dual-population co-evolution algorithm based on Q learning, and generating a distributed assembly wait-free flow shop scheduling scheme. According to the method, the optimal production and processing scheme with short time consumption and low carbon emission can be obtained in a distributed assembly wait-free flow workshop scene, so that the production and processing efficiency is improved, and the carbon emission is reduced.
Owner:ANHUI UNIV

Distributed energy storage scheduling method considering source load fluctuation credible deviation index

The invention relates to the technical field of power system automation, in particular to a distributed energy storage scheduling method considering a source load fluctuation credible deviation index, which comprises the steps of collecting data and preprocessing the data, calculating the source load fluctuation credible deviation index and establishing an optimal scheduling model, and establishing a multi-agent distributed energy storage optimal scheduling decision model. Solving based on an improved MADDPG algorithm; according to the method, a quantitative basis is provided for optimization decision making through the source load fluctuation credible deviation index, a multi-target optimization model fusing economy, life loss and fluctuation suppression is constructed, an improved multi-agent deep reinforcement learning algorithm is adopted, an imitation learning and Nash-Q learning mechanism is introduced to improve the convergence speed and balance, and meanwhile, the method is suitable for large-scale popularization and application. Efficient collaboration of cloud global optimization and edge-end real-time decision making is achieved through the cloud-edge collaboration architecture, a complete, efficient and high-adaptability distributed energy storage scheduling solution is finally formed, and the stability and economical efficiency of the system are effectively improved.
Owner:ANSHAN POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER COMPANY +1

A two-shift electric taxi charging guidance method and system based on Q learning

PendingCN122311768ASimulationPower grid
This invention discloses a charging guidance method and system for two-shift electric taxis based on Q-learning. The method first constructs an electric taxi operation process model to simulate the vehicle's route selection and state of charge changes under different states. Second, it establishes a charging behavior model based on the Q-learning algorithm. By defining discretized time, region, and state of charge spaces, it sets passenger-finding and charging actions and designs a sum-of-benefits penalty function that includes state of charge threshold constraints at shift change times, training to obtain the optimal charging decision strategy. Finally, the guidance effect is evaluated from three dimensions: charging cost, satisfaction of residents' travel needs, and grid power supply balance. This invention effectively reduces unnecessary charging times through intelligent guidance, smooths the charging load during shift change times, lowers operating costs, and simultaneously improves grid stability and residents' travel security.
Owner:XUZHOU COLLEGE OF INDAL TECH

Method for adjusting compensation parameters of active filter based on Q learning and related equipment

The invention provides a method for adjusting compensation parameters of an active filter based on Q learning and related equipment. The method comprises the following steps: acquiring load current in a power grid and compensation output current of the active filter; obtaining harmonic characteristics based on load current analysis, and defining the harmonic characteristics as a current environment state; defining the adjustment amount of the to-be-adjusted control parameter in the active filter as an action; in the current environment state, executing an action to adjust the to-be-adjusted control parameter, and calculating to obtain a Q value according to a difference value between the load current and the compensation output current; using a Q-learning algorithm to update and store a Q table of the corresponding relationship between the current environment state, the action and the Q value according to the Q value; according to the updated Q table, the control parameter corresponding to the action enabling the Q value to be maximum is selected in the current environment state to be used for controlling a loop, and therefore optimal compensation of harmonic waves is achieved. According to the method, the active filter has autonomous learning and optimizing capabilities, the optimal compensation strategy can be automatically found, and the compensation efficiency and effect are improved.
Owner:SHENZHEN QIDIAN NEW ENERGY TECH CO LTD

An underwater mobile node routing method, device, equipment and storage medium

The application discloses an underwater mobile node routing method, device and equipment and a storage medium, and is applied to the underwater communication field.The method comprises the following steps: a time-varying graph model of a target node is constructed, and the time-varying graph model is updated every first preset time; a Q function is calculated according to the time-varying graph model to obtain Q values of different actions performed by the target node in a current state; a maximum Q value is selected from the Q values, and an action corresponding to the maximum Q value is determined as a target action of the target node; and the target action is executed, so that the target node forwards data to a routing node of a next hop.The method realizes high reliability, low delay, high energy efficiency and long network survival time of the underwater mobile node network by constructing the time-varying graph model of the node and making a decision by Q learning according to the time-varying graph model, and the problem that a connection prediction model of an existing routing protocol is mainly aimed at the influence of water flow on the position of a node and is not suitable for a high-dynamic mobile sensor network is avoided.
Owner:JILIN UNIVERSITY

Manta ray robot fish roll control method based on q-learning and pectoral fin amplitude bias

The application relates to a manta ray robot fish roll control method based on Q learning and pectoral fin amplitude bias, compared with a traditional control method, the application does not need to establish a control object model, collects experimental data, trains a Q table offline, then transplants the table into a prototype controller, obtains a control variable through table lookup, and controls the prototype to swim at a fixed depth. Compared with other reinforcement learning control methods, the application has low requirements on hardware resources, small space demand, low power consumption and convenient realization. The application has the beneficial effects that the application has low requirements on the hardware of a controller, and an ordinary single-chip microcomputer can be used for realization, thereby saving cost, space and energy. The application does not need to establish a mathematical model of a control object, can save a large amount of time cost, and is difficult to establish a bionic model. The application can realize roll control of the robot fish without expert experience for establishing a rule base, and the Q table can learn "expert experience" during training.
Owner:NINGBO INST OF NORTHWESTERN POLYTECHNICAL UNIV +1

A chameleon service construction and optimized deployment method for power network attack robots

PendingCN122457292APathPingAttack
The application discloses a chameleon service construction and optimization deployment method for power network attack robots, and belongs to the fields of network security and artificial intelligence. The power network attack robot is an attack execution body that automatically penetrates the power grid relying on a large language model intelligent agent. The robot has a probabilistic preference and a high-frequency path tendency, and has an anti-reconnaissance avoidance capability. The application builds a "preference decoy-anti-bypass deployment" trapping mechanism for the layered penetration scene of the power grid. The layered power grid scene and a candidate service set are first built, and a chameleon service candidate form suitable for the environment of each layer is generated. Intelligent agents on both the attack and defense sides score the candidate services. A two-person zero-sum random game model is constructed, a profit function is constructed by fusing a hierarchical weight, scores and cost constraints, a minimax Q learning is used to solve a mixed strategy, and an optimized deployment scheme for each layer is output. The application can realize random strategy deployment of services, improve the trapping effect, and effectively resist anti-reconnaissance identification bypass.
Owner:GUANGZHOU UNIVERSITY HUANGPU RESEARCH INSTITUTE

Fuzzy self-adaptive Q learning control method and system for sewage treatment

The invention provides a fuzzy self-adaptive Q learning control method and system for sewage treatment, and the method comprises the steps: obtaining the dissolved oxygen concentration as a system state, and constructing a nonlinear optimization problem; performing dynamic adjustment on proportion, integral and differential coefficients by adopting a Mamdani type fuzzy inference rule in combination with fuzzy logic, performing defuzzification through a centroid method to obtain adaptive PID parameters, and forming a fuzzy adaptive control strategy; an online Q learning framework is further constructed, a Q function is approximated by using the evaluation network, a network generation strategy adjustment amount is executed, and network weight is optimized based on a Bellman equation and a particle swarm algorithm; finally, fuzzy control and Q learning output are fused, control input is generated through a coupling coefficient to adjust an oxygen transfer coefficient, and tracking control over the dissolved oxygen concentration is achieved. According to the invention, the tracking control precision of the dissolved oxygen concentration and the system operation stability can be improved.
Owner:BEIJING UNIV OF TECH

A hyper-sonic vehicle control parameter optimization method based on deep Q learning

The application provides a control parameter optimization method for a hypersonic aircraft based on a deep Q network, and the method comprises the following steps: establishing a hypersonic aircraft dynamics model; decomposing the hypersonic aircraft dynamics model into a speed subsystem and an altitude subsystem; designing a speed subsystem adaptive controller; designing an altitude subsystem adaptive controller; and optimizing the design parameters a in steps 3 and 4 by using a deep Q learning theory h 、a γ 、a α 、a Q , and designing a reward function. According to the needs of a user, the application can dynamically optimize the parameters of the controller of the hypersonic aircraft by designing the reward function, so that the control system can exert maximum efficiency.
Owner:AIR FORCE UNIV PLA

Unmanned aerial vehicle offline reinforcement learning training method combined with implicit Q learning

The invention discloses an unmanned aerial vehicle offline reinforcement learning training method in combination with implicit Q learning, and the method comprises the following steps: building an unmanned aerial vehicle simulation environment and a reinforcement learning training environment, and constructing a rule strategy agent through an unmanned aerial vehicle maneuvering decision statistical method in combination with the unmanned aerial vehicle offline reinforcement learning simulation environment; the method comprises the following steps: interacting with an unmanned aerial vehicle offline reinforcement learning simulation environment through a rule strategy agent, using the rule strategy agent to carry out maneuvering decision-making by a first unmanned aerial vehicle and a second unmanned aerial vehicle, collecting interaction training data, and storing the interaction training data into an offline reinforcement learning experience playback pool; and constructing a soft actor commentator algorithm module combined with implicit Q learning, randomly extracting batch data from the experience playback pool, and inputting the batch data into an IQL-SAC unmanned aerial vehicle offline reinforcement learning algorithm module to train a neural network. According to the invention, through combination with the soft actor commentator algorithm offline reinforcement learning training of implicit Q learning, the offline reinforcement learning training effect and maneuvering decision-making ability of the unmanned aerial vehicle are improved.
Owner:SOUTH CHINA UNIV OF TECH

Air conditioner four-way valve pipe group design method and device based on dqn network and air conditioner

PendingCN122365790ADesign completed quicklyReduce vibration effectsFour-way valveControl engineering
The application provides a DQN network-based air conditioner four-way valve pipe group design method and device and an air conditioner, which are applied to the field of air conditioner control and include the following steps: initializing a target environment and network weights of a target network model; the target network model is a model constructed based on a deep Q learning network; an action value generated by an agent of the target network model based on a current state is input into the target environment, vibration simulation is performed on the action value output by the agent by using the target environment, and a generated strategy of the agent is optimized based on an obtained observation value until air conditioner four-way valve pipe group design parameters meeting a preset requirement are obtained. The DQN network-based air conditioner four-way valve pipe group design method and device and the air conditioner provided by the application are used for quickly completing the design of an air conditioner outdoor unit four-way valve pipe group, improve the design efficiency, and greatly reduce the influence of the pipe group on the vibration of the outdoor unit.
Owner:QINGDAO HAIER AIR CONDITIONER GENERAL CORP LTD +1

A sewage pH value regulation system based on reinforcement learning

The present application relates to sewage treatment technical field, specifically to a kind of sewage pH value control system based on reinforcement learning, comprising: multi-dimensional redundant perception module gathers sewage pH value, temperature, turbidity, conductivity and flow data;Time series prediction feedforward module calculates the basic dosing flow of acid-base liquid with trend compensation;Double Q learning hierarchical control module iteratively optimizes by constructing double Q-learning model containing three-dimensional extended state space, hierarchical action space and adaptive reward function, generates hierarchical optimal control strategy;Closed-loop execution feedback module accurately dosing acid-base liquid through variable frequency stepless adjustment unit, collects actual dosing amount back and completes control strategy closed-loop calibration correction;Fault self-diagnosis fault-tolerant module and the rest of the module two-way communication, monitor core component operating state, according to preset fault threshold positioning determination fault, automatically switch standby control strategy.
Owner:SHAOXING XUESEN ENERGY TECH CO LTD

Multi-agile earth observation satellite scheduling method based on deep q-learning and ensemble heuristic

PendingCN122114423AMathematical modelsNeural learning methodsQ-learningEarth observation satellite
The application discloses an agile earth observation satellite planning method based on deep Q learning and heuristic rules, uses a DQN learning task allocation strategy, improves the efficiency and quality of task allocation, combines various heuristic rules, dynamically adjusts a scheduling strategy, and can adapt to different scene requirements; through a time relaxation task sliding insertion strategy, the execution time of the arranged task is allowed to be adjusted, task conflicts are reduced, and scheduling flexibility is improved; through an interaction mechanism between upper task allocation and lower task scheduling, the efficiency and stability of task allocation and scheduling are ensured.
Owner:NAT UNIV OF DEFENSE TECH

Q learning-based same-cluster batch scheduling method for three-stage energy-saving hybrid flow shop

The invention discloses a same-cluster batch scheduling method for a three-stage energy-saving hybrid flow shop based on Q learning, and relates to the technical field of green scheduling in the manufacturing industry. Aiming at the scheduling problem of the three-stage hybrid flow shop in scenes such as automobile engine manufacturing, the method comprises the following steps: firstly, constructing a scheduling model comprising three stages of processing, surface treatment and final testing; determining a dual-objective function and multiple constraints for minimizing the maximum completion time and the total energy consumption; and then proposing a hybrid ant colony optimization algorithm QHACO solving model based on Q learning, dynamically selecting workpiece and machine strategies through Q learning, and generating a scheduling scheme with low energy consumption and high efficiency in combination with pheromone updating and heuristic search of the ant colony algorithm. The method can effectively cope with complex constraints such as a parallel machine, a batch processor, a limited buffer area and same-cluster group batch, and provides technical support for green transformation of manufacturing enterprises.
Owner:ANHUI UNIV