Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

177 results about "Q-learning" patented technology

Q-learning is a model-free reinforcement learning algorithm. The goal of Q-learning is to learn a policy, which tells an agent what action to take under what circumstances. It does not require a model (hence the connotation "model-free") of the environment, and it can handle problems with stochastic transitions and rewards, without requiring adaptations.

Power grid monitoring method

The invention relates to the technical field of intelligent power energy monitoring, in particular to a power grid monitoring method. The method comprises the following steps: obtaining power grid monitoring nodes, and carrying out monitoring node graph structure processing to construct a physical feature fusion network graph; performing graph structure knowledge enhancement on nodes in the physical feature fusion network graph to construct a semantic enhancement physical association graph; performing composite diagram operation processing of fault propagation risk assessment and root cause analysis based on the semantic enhanced physical association diagram, and constructing four types of intelligent agents of monitoring, coordination, analysis and decision to obtain a hierarchical collaborative decision network; and executing a local strategy learning process based on Q learning based on the hierarchical collaborative decision network so as to realize agent reinforcement learning optimization and end-to-end deployment implementation. According to the method, automatic and end-to-end construction of the intelligent agent is realized through graph structure semantic fusion and multi-stage reinforcement learning, and the intelligent level of power grid monitoring is remarkably improved.
Owner:HENAN MINGERMEI ELECTRONIC TECHNOLOGY CO LTD

Robot skill learning method, system and terminal for text large model assisted reinforcement learning

The invention discloses a robot skill learning method and system for text large model assisted reinforcement learning, and a terminal. The method comprises the following steps: representing an environment by adopting a scene graph; analyzing a task instruction and an environment by utilizing a text large model; taking the large text model as a guide agent, taking a reinforcement learning agent as an exploration agent, combining jump reinforcement learning and conservative Q learning, and training the reinforcement learning agent by using a data set generated by interaction of the guide agent and the exploration agent with the environment to obtain a trained reinforcement learning agent; and selecting a target observation object for the trained reinforcement learning agent by using the text large model, and outputting a target action sequence by the trained reinforcement learning agent. According to the method, language instructions are understood through a large model technology, feasible guide strategies are generated, so that random exploration steps are reduced, parameterized skills are introduced, the training process of reinforcement learning is accelerated, the efficiency is high, and the method does not depend on the data volume of a skill library.
Owner:SHENZHEN UNIV

Cross-cabin collaborative operation method and system for dry bulk cargo cabin cleaning and ship unloading equipment cluster

The invention belongs to but is not limited to the technical field of port loading and unloading operation, and discloses a hyper-heuristic dry bulk cargo cabin cleaning and ship unloading equipment cluster cross-cabin collaborative operation method and system based on large language model-Q learning dual drive. Acquiring information such as historical data and equipment attributes of cabin clearing and ship unloading operation of the dry bulk cargo port; determining the cabin cleaning and ship unloading operation process and task information of the dry bulk cargo port; establishing a cabin-cleaning and ship-unloading equipment cluster cross-cabin collaborative operation model for determining an equipment collaborative operation plan; acquiring and utilizing a large language model expansion training sample; training and adjusting a Q learning model based on the sample; based on a hyper-heuristic algorithm, LLM and QL are fused, an LQ-HH algorithm is designed and used for optimizing and solving the equipment collaborative operation model, and the operation sequence and starting time of cabin cleaning and ship unloading equipment are obtained. According to the method, the operation efficiency, the equipment cooperation capability and the anti-interference capability can be effectively improved, and intelligent transformation and upgrading of the dry bulk port can be promoted.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Intelligent greenhouse control method and system based on Internet of Things

The invention relates to the technical field of intelligent agriculture and Internet of Things control, and discloses an intelligent greenhouse control method and system based on the Internet of Things, and the method comprises the steps: constructing a multi-factor environment state vector; strawberry image features are extracted based on YOLOv5; inputting a strawberry growth stage prediction model in combination with the time sequence to determine target environment parameters; constructing an environmental parameter prediction model; constructing a dual error term, and inputting the dual error term into a fuzzy controller to generate a preliminary control instruction; optimizing and outputting an optimization control instruction by the Q learning network; compared with a greenhouse control method which mostly depends on a fixed threshold value or a static rule in the prior art, the technical problem that dynamic target setting and self-adaptive adjustment of environmental parameters are difficult to realize especially under the condition that demands of strawberry crops in different growth stages change rapidly is solved. Intelligent greenhouse environment control in combination with the growth stage of strawberry crops is realized, and the adaptability of the greenhouse to crop growth conditions is improved.
Owner:XUZHOU DAFENG INTERNET OF THINGS RES INST CO LTD

Double-capacity water tank control method, system and equipment based on multi-target reward Q learning and medium

The invention discloses a double-capacity water tank control method based on multi-target reward Q learning, and belongs to the technical field of machine learning and automatic control. The method comprises the following steps: firstly, establishing a discretization mathematical model of a double-capacity water tank liquid level control system, and designing a multi-target reward function; secondly, empirical data are stored only when the liquid level tracking error exceeds the limit or the control quantity suddenly changes, and sampling weights are distributed according to award absolute values; and finally, updating the Q table based on the time difference error, and performing closed-loop real-time control on the double-tank water tank according to the dynamically adjusted PID parameter. According to the method, a five-dimensional reward function is designed, and multi-index dynamic balance is realized through linear weighting; the empirical tuple is stored only when the liquid level tracking error exceeds the limit, and the quick response requirement of the dynamic time-varying scene of the double-capacity water tank is met; dimensionality reduction is performed on a continuous state space by adopting a non-uniform grading strategy, so that the calculation complexity is remarkably reduced; a perception-decision-execution integrated framework is constructed, and the dependence of a traditional control method on a mathematical model of the water tank is abandoned.
Owner:LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY

Tunnel anomaly detection and response method and equipment based on rail robot

The invention provides a tunnel anomaly detection and response method and device based on an orbital robot, and the method comprises the steps: obtaining scoring and grading data from a quantitative scoring model through a preset model optimization module, carrying out the dynamic adjustment of the parameters of the scoring model through a deep Q learning network in combination with historical response records, optimizing a grading threshold value to match an actual threat level, and carrying out the detection of an actual threat level. Transmitting the adjusted hierarchical data to a response scheduling system; and according to the adjusted hierarchical data, automatically matching a preset response strategy combination, if at least one selectable strategy combination exists, determining an optimal strategy combination through a weighted optimization algorithm, a comprehensive response speed, a resource utilization rate and a disposal cost, and issuing an execution instruction to a related execution unit.
Owner:JIE XUN TECH (GUANGZHOU) CO LTD

Outbound routing method and system based on multi-dimensional index coupling and large model decision

The invention provides an outbound routing method and system based on multi-dimensional index coupling and large model decision, and the method comprises the steps: carrying out the preprocessing of the historical data of each line provider, predicting the future performance index through an LSTM neural network, and optimizing the prediction result through a Kalman filter, and eliminating the noise. In the route decision-making link, the system calculates the Q value of each line provider based on the deep Q learning network, comprehensively considers the call completing rate, average cost, negative feedback rate and other factors of each line provider, and intelligently selects the optimal line provider in combination with the user portrait, scene recognition and real-time load state. The system also constructs a complete feedback closed loop, continuously optimizes the model by collecting call result data, and automatically adjusts decision parameters. The call completing rate and the service quality of the outbound service are improved, and the communication cost is reduced.
Owner:BEIJING YULORE INNOVATION TECH

Integrated feature selection method and product based on quantum computing and Bayesian optimization

The invention provides an integrated feature selection method and product based on quantum computing and Bayesian optimization, and relates to the technical field of data processing. According to the embodiment of the invention, an original mathematical integration model is converted into a QUBO model which can be solved by quantum calculation, and the QUBO model is decomposed into a joint optimization discretization step length optimization sub-problem and a parameterization QUBO sub-problem. In a mixed quantum classical optimization algorithm framework, a self-adaptive Q learning model is designed on the upper layer, and proper sub-problems can be dynamically selected in the search process. In the lower layer, a dropout Bayesian optimization algorithm is provided for effectively optimizing the high-dimensional discretization step length in each iteration. A CIM-based quantum computing method is adopted, and a parameterized QUBO sub-problem under the given discretization step length is efficiently solved. According to the method provided by the embodiment of the invention, the feature selection problem can be successfully and efficiently solved, and the selected features of the credit classification problem and the credit classification model for classification based on the selected features can be obtained.
Owner:BEIJING INST OF TECH +1

Foot type gait and form coevolution system integrating reinforcement learning and evolutionary computation

The invention discloses a foot type gait and form co-evolution system integrating reinforcement learning and evolutionary computation, and relates to the technical field of robots and the field of artificial intelligence. Comprising the steps that a real-time motion control strategy learning module is realized based on a dynamic mask integrated double-Q learning method; the gait and form collaborative optimization module optimizes key gait and form parameters influencing the global motion performance of the foot robot based on an evolutionary computation theory; the training environment adaptive enhancement module comprises an automatic domain randomization module and a course learning module; the parameterized trajectory generator module is used for providing structured motion prior for the real-time motion control strategy learning module; the robot simulation environment is used for simulating the physical effect of the robot and the environment in the computer; according to the invention, the high-speed, stable and efficient autonomous movement capability of the foot robot in a complex and changeable unstructured environment is realized.
Owner:SHANGHAI UNIV

Heterogeneous cluster task scheduling fusion method based on Q learning and genetic algorithm

The invention discloses a heterogeneous cluster task scheduling fusion method based on Q learning and a genetic algorithm, and belongs to the technical field of intelligent scheduling. Collaborative optimization is carried out through a dynamic feedback mechanism of Q-Learning and the global search capability of the genetic algorithm: in an initialization stage, a cluster state space and an action space are defined, and a multi-Q-value table and an initial population are generated; during task scheduling, nodes are selected based on a # imgabs0 #-greedy strategy, reward values are calculated, and a Q value table is updated; meanwhile, the scheduling scheme is coded into chromosomes, a new population is generated through roulette selection, crossover and variation, and a Q value table is optimized; in the iteration process, the convergence is judged according to the Q value change or the population fitness change, and the strategy is dynamically adjusted. According to the method, the advantages of double algorithms are fused, local optimum is avoided, the task processing efficiency and the resource utilization rate are remarkably improved, adaptive strategy updating during cluster state change is supported, and the method is suitable for an efficient task scheduling scene of a large-scale heterogeneous cluster.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Data security management method and system based on cloud computing platform

The invention discloses a data security management method and system based on a cloud computing platform, and relates to the technical field of data security management, and the method comprises the steps: collecting original data of the cloud computing platform in real time, and carrying out the preprocessing; constructing a discrimination model based on a generative adversarial network (GAN) to generate a forged data sample, taking the forged data sample and a real data sample as input data, optimizing the discrimination model by using the input data based on an extreme learning machine, and discriminating attack data; and constructing a security situation model by using a hidden Markov model, optimizing parameters through a cluster particle algorithm, and outputting a predicted security situation based on real-time attack data. A discrimination model is optimized through a generative adversarial network and an extreme learning machine, the discrimination precision of attack data is effectively improved, a security situation model is established through a hidden Markov model, model parameters are optimized in combination with a cluster particle algorithm, a predicted attack path is combined with a deep Q learning network, and the security risk of attack data is improved. And attack path prediction is carried out through the dual deep Q network, and a defense strategy is adjusted in real time.
Owner:JIANGSU XILIXI TECHNOLOGY CO LTD

Limited block length optimization method and system of cellular-free communication perception integrated system

The invention discloses a limited block length optimization method and system of a cellular-free communication perception integrated system, and belongs to the technical field of wireless communication. The method comprises the following steps: constructing a cellular-free multi-station cooperative communication perception integrated architecture, and establishing a channel estimation and data transmission model; deducing a closed expression of the communication transmission delay, the decoding error probability and the sensing precision Cramer-Rao bound under the limited block length; analyzing a Pareto optimal boundary of communication and sensing performance by adopting a non-dominated sorting genetic algorithm, and determining a block length feasible region; and designing a dynamic optimization algorithm based on a double-depth Q learning network, and realizing collaborative optimization of the communication performance by adaptively adjusting the block length. The method solves the problem that high communication reliability, low time delay and high sensing precision are difficult to consider in a dynamic environment, the overall performance of the system is remarkably improved, and the method is suitable for high-requirement scenes such as automatic driving and industrial Internet of Things.
Owner:SOUTHEAST UNIV

Linear motor temperature control method and system based on air cooling heat dissipation

The invention provides a linear motor temperature control method and system based on air cooling heat dissipation, and relates to the technical field of motor temperature control, and the method comprises the steps: dividing a linear motor into a plurality of temperature control regions, and configuring an independent control execution unit for each region; predicting a temperature trend based on the LSTM; a multi-agent reinforcement learning control network is constructed, and each agent calculates an optimal control strategy based on deep Q learning; generating an overall airflow regulation strategy through a collaborative optimization algorithm; dynamically adjusting the opening of the electric adjusting valve and the rotating speed of the fan. Accurate temperature control is achieved, the heat dissipation efficiency is improved, and the service life of equipment is prolonged.
Owner:SHENZHEN FERGUS ELECTROMECHANICAL EQUIP CO LTD

Self-adaptive search self-tuning method of motor driving algorithm based on reinforcement learning

The invention relates to the technical field of reinforcement learning, motor driving and the like, provides a reinforcement learning-based adaptive search self-tuning method for a motor driving algorithm, and realizes an optimal decision of an individual search behavior by performing adaptive switching among various heuristic search operators by introducing a Q learning mechanism. And meanwhile, in combination with a neighborhood search strategy based on differential evolution, the local search capability is enhanced, and the global optimization performance is improved. The method is applied to the setting problem of key parameters of the motor driving system, the I TAE performance index serves as an optimization target, the search behavior is dynamically adjusted through a self-adaptive strategy, and therefore better control performance is obtained. Compared with an existing method, the method has the remarkable advantages in the aspects of convergence speed, self-adaptive capacity and anti-interference performance.
Owner:SHENZHEN XILIN ELECTRICAL TECH

Two-stage excitation method based on multi-agent deep reinforcement learning

The invention discloses a two-stage excitation method based on multi-agent deep reinforcement learning, relates to the technical field of agent deep reinforcement learning, and aims to solve the problems that an excitation method in the prior art is too tedious in process and poor in effect. Through a simulation experiment in a specific area, taxies are used as participants to simulate real perception task geographical distribution and trajectory planning, comparison mechanisms such as deep Q learning and historical participation rate optimization are adopted, and multi-dimensional indexes such as average participant utility, platform utility and task completion rate are combined, so that the adaptability of the scheme in a dynamic environment is verified, and the accuracy of the scheme is improved. Cost is quantified by combining different sensor characteristics, participants can accurately evaluate energy consumption of themselves, platform utility is maximized under budget constraint, meanwhile, through bidirectional counterbalance of budget and task coverage, it is ensured that'minimum feasibility coverage 'of tasks and budget are not hyperbranched, actions of other agents can be observed during training to optimize cooperation, and the training efficiency is improved. The method only depends on a local state during execution, and adapts to a multi-participant competition scene.
Owner:HENAN UNIV OF SCI & TECH

Fault traffic electrification network emergency support strategy and system based on mixed graph attention-depth double Q learning architecture

The invention discloses a fault traffic electrification network emergency support strategy and a fault traffic electrification network emergency support system based on a mixed graph attention (Graph Attention Network, GAT)-deep double Q learning (DDQN) architecture, and aims to enhance the anti-interference capability and the recovery capability of a traffic electrification network when the traffic electrification network faces faults and emergency situations. Firstly, coupling characteristics and operation requirements of a traffic network and a power distribution network are comprehensively considered, a traffic electrification network emergency support framework based on a GAT-DDQN architecture is proposed, and a double-layer multi-objective optimization model is established. And secondly, learning fault topological structure information based on the graph attention network, constructing a double-layer finite Markov Decision Process (FMMP) model, modeling a dynamic interaction process of the traffic network and the power distribution network into a state, action and reward decision sequence, and realizing fault emergency support of the power distribution network and the traffic network. And finally, proposing a DDQN solving algorithm based on an attention mechanism and experience playback optimization, and outputting an optimal traffic electrification network emergency support strategy by strengthening key node feature learning and historical experience utilization efficiency. And finally, constructing a traffic electrification network emergency support system which comprises a data acquisition and processing module, a feature extraction module, an algorithm solving module and a visualization module.
Owner:NANJING UNIV OF POSTS & TELECOMM

Mobile robot adaptive path planning method based on enhanced Q learning and multi-strategy cooperation

The invention discloses a mobile robot adaptive path planning method based on enhanced Q learning and multi-strategy cooperation, and relates to the field of robot autonomous navigation and computing intelligence. The method comprises the following steps: firstly, establishing a two-dimensional grid map model containing obstacle information through an environment sensing module; the initial position of a gold scrubber population is generated by utilizing the low-difference characteristic of a golden section sequence to enhance the global ergodicity of a solution space, three stages of migration, gold scrubber and cooperation of a gold scrubber optimization algorithm are modeled into a discrete action space, an optimal search action is dynamically selected through a nonlinear attenuation greedy strategy, and the optimal search action is realized. According to the method, adaptive switching of exploration and development behaviors is realized, a search gradient is constructed by using a differential vector of survival of the survival of the survival of the population, and a lens imaging reverse learning strategy is introduced to dynamically disturb an elite individual so as to jump out of a local extremum trap. And finally, outputting an optimal collision-free path through a multi-objective evaluation function including path length, safety and smoothness.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Industrial network working condition coupling frequency spectrum avoiding method and system

The invention relates to the field of industrial anti-interference, in particular to an industrial network working condition coupling frequency spectrum avoiding method and system. The method comprises the following steps: constructing a main / sub-network spectrum interference coupling matrix, fusing multi-node spectrum data to generate a global map, and establishing a conflict probability model to screen a high-risk frequency band; and generating an optimal frequency band matching pair and a reconfiguration instruction set based on the cost matrix and a Hungary algorithm, and finally optimizing an avoidance strategy by taking throughput improvement and interference suppression as rewards by using a Q learning model. According to the invention, the problem that the network coupling relationship is not considered in the traditional spectrum management is solved, the dynamic refined management of spectrum resources in industrial scenes such as iron and steel plants is realized, the interference influence is obviously reduced, and the network communication quality and reliability are improved.
Owner:XINJI AOSEN STEEL GRP CO LTD

Fire monitoring and predicting method based on multi-agent reinforcement learning

The invention discloses a fire monitoring and predicting method based on multi-agent reinforcement learning, and belongs to the field of artificial intelligence and disaster prevention and control. According to the method, early detection, real-time monitoring, accurate prediction and rapid response are performed on a fire based on an intelligent means; calculating and evaluating a plurality of indexes such as fire risk degree, fire spreading degree, fire hazard degree, environmental influence factor and emergency responsivity; a reinforcement learning method including deep Q learning, Monte Carlo tree search, a strategy gradient method, an actor-evaluator framework and a double-DQN network is adopted to improve the coverage range, the response speed and the prediction capability of fire monitoring. According to the method, multi-agent combination is utilized, intelligent risk assessment and spreading prediction are carried out on the fire, the fire prevention and control efficiency and accuracy are effectively improved, and the method has a wide application prospect.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Off-line reinforcement learning action exploration agent method based on expected reward regularization

The invention discloses an off-line reinforcement learning action exploration agent method based on expected reward regularization, and belongs to the field of reinforcement learning. The problem that reliable track splicing and strategy generalization are difficult to realize in a complex task by the existing method is solved. The method comprises the following steps: constructing state loss and action loss based on sequence modeling, and carrying out iterative training of states and actions; designing an RTG loss function based on a weighted square error; a double-Q learning framework is adopted to maintain two independent Q functions, conservative constraints of conservative Q learning are applied in the Q function updating process, and action exploration optimization is achieved in combination with Boltzmann distribution; combining state loss, action loss, RTG regularization loss and Q value loss to form a joint optimization objective function; performing noise disturbance sampling on the plurality of RTG candidate values to generate diversified action prediction; and value evaluation is performed on the candidate actions based on a double-conservative Q function, and the action with the highest Q value is selected for execution. The method is mainly used in the intelligent agent exploration field.
Owner:HARBIN INST OF TECH

Bellman infinite error-based robust Q learning resisting method and system

The invention relates to the technical field of data processing, and discloses a Bellman infinity error-based robust Q learning resisting method and a Bellman infinity error-based robust Q learning resisting system. The method comprises the following steps of: constructing an intrinsic state neighborhood set by utilizing a Bellman optimal Q function according to a state space, an action space and a dynamic transition probability function of a Markov decision process, depicting an optimal adversarial Q function model by utilizing a consistency adversarial robust operator, forming a training target by adopting a Bellman infinite norm, and constructing an intrinsic state neighborhood set; a projection gradient descent algorithm is combined to directly obtain an adversarial sample, or a Q value boundary is estimated through interval boundary propagation so as to obtain a substitute upper bound target of a training target, and finally network training is completed through methods such as stochastic gradient descent or adaptive momentum estimation. On the premise of ensuring the existence of the optimal robust strategy, the deep Q learning method with adversarial robustness and training stability is designed, so that the trained intelligent agent can show consistent excellent performance in a clean environment and an adversarial environment.
Owner:UNIV OF CHINESE ACAD OF SCI

Redundant executor dynamic scheduling method and system based on foreground theory and Q learning evolutionary game method

PendingCN121261962ABiological modelsSecuring communicationQ-learningIteration loop
The invention provides a dynamic scheduling method and system for a redundant executor based on a foreground theory and a Q learning evolutionary game method. The method comprises the following steps: constructing a strategy space, a profit structure and strategy evolution dynamics of the executor and a reward and punishment mechanism of an enhanced analyzer-regulator; according to the dynamic scheduling method, an executor simulates interaction under a tripartite game framework, iterative loop is carried out, and a strategy space of a mimicry defense system is output as a redundant executor after the loop is ended, the simulation interaction process comprises the following steps: the enhanced analyzer-regulator monitors and obtains objective income of each execution body in real time, calculates subjective perception utility of each execution body according to an income structure of the execution body, and then adjusts the subjective perception utility of the mimicry defense system according to a reward and punishment mechanism of the enhanced analyzer-regulator. And each execution body dynamically updates the strategy space of each execution body according to the adjusted subjective perception utility and strategy evolution.
Owner:WUHAN UNIV +4

DMP data management method based on artificial intelligence

The invention discloses a DMP data management method based on artificial intelligence, and relates to the technical field of data management.The method comprises the steps that multi-device user behavior data are collected through a cross-platform SDK, a browser plug-in and an application log, and data standardization is conducted on a central server or an edge computing node; in the user identity matching process, in combination with a device fingerprint and deep learning method, a user cross-device behavior mode is analyzed through a Transform model, static device features and dynamic behavior features are extracted, a self-attention mechanism is adopted to learn cross-device associated features, then the user matching degree is calculated through cosine similarity, and finally the user identity is judged. The stability and accuracy of identity matching are improved; and based on a user identity matching result, carrying out time sequence analysis by utilizing an LSTM (Long Short Term Memory) model, capturing evolution of user interests along with time, dynamically adjusting user portrait weights through reinforcement learning, and optimizing feature weight updating by using a Q learning mechanism, so that the portraits can be adaptively changed.
Owner:BEIJING GREY INNOVATION TECHNOLOGY CO LTD

Road investigation design method and system based on mobile network

The invention discloses a road investigation design method and system based on a mobile network, and relates to the technical field of road investigation, and the method comprises the steps: dividing geological anomaly probability distribution data into dynamic grid units, constructing a reinforcement learning state space, and generating an optimal sampling path instruction in combination with a dual-network deep Q learning architecture; analyzing the optimal sampling path instruction into executable parameters, executing the executable parameters, collecting multi-source spatio-temporal data, establishing data association through a spatio-temporal hash algorithm, and generating a reconnaissance data set; and on the basis of the survey data set, through parameterized spline curve modeling, generating a candidate road design scheme, and in combination with a non-dominated sorting genetic algorithm, performing optimization to generate a three-dimensional road design scheme. According to the method, the exploration path is optimized by using the dynamic grid coding and the dual-network architecture, the high-risk area coverage and the moving efficiency are balanced, meanwhile, the method adapts to real-time geological changes in combination with priority experience playback, and the exploration reliability and the resource utilization rate are improved.
Owner:JIANGSU SHIGUANG GEOGRAPHIC INFORMATION TECH CO LTD

Accurate control method and system for denitrification of sewage treatment plant and electronic equipment

The invention relates to the technical field of accurate control scheme design for denitrification of a sewage treatment plant, in particular to an accurate control method and system for denitrification of a sewage treatment plant and electronic equipment. Data of carbon source concentration, dissolved oxygen concentration, pH value, temperature, alkalinity and nitrate concentration in sewage are collected in real time, parameters of a prediction model are dynamically corrected in combination with a machine learning algorithm and a mechanism model, and the denitrification rate is predicted. The carbon source adding amount and the dissolved oxygen target value are calculated according to the prediction result and the emission standard, the distribution of the carbon source and the dissolved oxygen in the reaction tank is monitored in real time through a distributed sensor network, and the operation states of a carbon source adding pump and aeration equipment are accurately adjusted through feedback control. According to the method, a pH value and alkalinity are taken as constraint conditions to cooperatively control a carbon source feeding rate and a dissolved oxygen concentration, and a control strategy is dynamically optimized through deep Q learning based on historical data and operation conditions, so that cross-cycle self-adaptive adjustment is realized, and accurate control on a denitrification process of a sewage treatment plant is realized.
Owner:SHANDONG CHUANQINGQING ENVIRONMENTAL PROTECTION TECHNOLOGY CO LTD

Distributed rainbow table decryption task scheduling system based on deep learning optimization

The invention relates to the technical field of decryption task scheduling, in particular to a distributed rainbow table decryption task scheduling system based on deep learning optimization. According to the method, the deep learning algorithm and the Q learning method are introduced, and efficient task scheduling is carried out based on the decryption task scheduling feature matrix of the ciphertext and the real-time resource information of the computing node. And the deep learning optimization module automatically selects a proper computing node and optimizes task allocation by simulating a scheduling task, so that the processing efficiency of a decryption task is remarkably improved. By decrypting a task path identification algorithm, the system realizes dynamic task allocation, monitors immigration and emigration of tasks in real time, timely discovers the condition that the load of a computing node is too heavy, automatically migrates the tasks to nodes with light loads, ensures the load balance of the system, and avoids an overload phenomenon.
Owner:JINAN LANJIANJUN NEW INFORMATION TECH CO LTD

Multi-edge task unloading method combined with task pre-sorting

The invention relates to a multi-edge task unloading method combined with task pre-sorting, and belongs to the technical field of edge calculation and task unloading. The method comprises the following steps: firstly, establishing a multi-edge, multi-user and multi-task network architecture comprising a cloud layer, an edge layer and a user layer; based on network architecture support, an application model, a transmission model and an execution model for an application task of a user are established, the application model represents any subtask as a triple, the transmission model defines a corresponding transmission rate, and the execution model comprises a local execution model, an edge execution model and an overall cost model; a task pre-sorting mechanism is introduced, and subtasks which can be parallel are sorted in a descending mode according to priority factors; and finally, based on the ordered task sequence after sorting, performing task unloading strategy decision by adopting a deep Q learning network. According to the method, joint optimization of task execution delay and energy consumption can be realized in a multi-user, multi-task and multi-edge node computing environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Broadband satellite communication verification system and method

The invention discloses a broadband satellite communication verification system and method, relates to the technical field of broadband satellite communication, and remarkably improves the response efficiency and stability of the broadband satellite communication verification system through a distributed asynchronous Q learning framework and a dynamic coordination mechanism. In a high dynamic environment of a satellite-ground fusion network, a local decision agent operates independently, a resource allocation decision is quickly generated based on a real-time channel state and a user demand, and a calculation bottleneck brought by centralized optimization is avoided; according to the distributed architecture, communication overhead among nodes is reduced, decision delay is reduced, and therefore the jitter phenomenon caused by multi-user competition is effectively restrained; a penalty function mechanism integrates the conflict rate, the interference degree and other multi-dimensional costs into a learning process, so that the nodes autonomously avoid high-competition actions, and the fairness of resource allocation is enhanced; meanwhile, the coordination module adaptively adjusts the updating period through topological change measurement, and the response interval is shortened when the satellite rapidly moves.
Owner:AEROSPACE FUTURE (SHENZHEN) AEROSPACE TECHNOLOGY CO LTD

Slope cutting house building landslide vulnerability analysis method and system based on reinforcement learning

The invention discloses a slope-cutting house-building landslide vulnerability analysis method and system based on reinforcement learning. The method comprises the following steps: preprocessing multi-source data to obtain a disaster sample data set, a missing tag vector and a damage level tag; feature parameters of the slope and the house are extracted, and missing mark vectors are fused to generate a state vector set; constructing an action set and calculating an external return according to the difference between the predicted damage distribution and the damage level label; calculating a conservative penalty quantity for non-data actions by utilizing the action value network and the target action value network in conservative Q learning; calculating a novelty error in the random network distillation and generating an intrinsic return; and fusing the return to construct a state transition sample, updating the action value network to obtain a conservative action value function, and outputting a vulnerability conclusion of the scene to be analyzed. According to the method, stable training and iteration of the evaluation model are realized under the conditions of sample scarcity and data incompleteness.
Owner:HOHAI UNIV

A method, management device and storage medium for suppressing subsynchronous oscillation of a doubly-fed wind farm based on intelligent learning and predictive control

The application discloses a double-fed wind farm subsynchronous oscillation suppression method based on intelligent learning and predictive control. First, based on double-fed wind farm dominant oscillation mode analysis, double-fed wind turbine rotor side converter outer ring damping control prediction model and inner ring damping control prediction model are analyzed; a damping optimization function approximator is constructed based on a robust model predictive control, and a double-fed wind turbine outer ring damping optimization control module and an inner ring damping optimization control module are established; then, based on Q learning optimization, optimal control laws of the double-fed wind turbine outer ring damping optimization control module and the inner ring damping optimization control module are solved, and the optimal control laws are stored based on data cubes; finally, the double-fed wind turbine outer ring damping optimization control module and the inner ring damping optimization control module are incorporated into the inner and outer ring control of the rotor side converter, and double-fed wind farm subsynchronous oscillation suppression is realized. The application considers the uncertain influence of the control system parameters, can realize double-fed wind farm subsynchronous oscillation suppression, and avoids oscillation propagation.
Owner:STATE GRID SICHUAN ELECTRIC POWER CO