Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

138 results about "Q-learning" patented technology

Q-learning is a model-free reinforcement learning algorithm. The goal of Q-learning is to learn a policy, which tells an agent what action to take under what circumstances. It does not require a model (hence the connotation "model-free") of the environment, and it can handle problems with stochastic transitions and rewards, without requiring adaptations.

Cross-cabin collaborative operation method and system for dry bulk cargo cabin cleaning and ship unloading equipment cluster

The invention belongs to but is not limited to the technical field of port loading and unloading operation, and discloses a hyper-heuristic dry bulk cargo cabin cleaning and ship unloading equipment cluster cross-cabin collaborative operation method and system based on large language model-Q learning dual drive. Acquiring information such as historical data and equipment attributes of cabin clearing and ship unloading operation of the dry bulk cargo port; determining the cabin cleaning and ship unloading operation process and task information of the dry bulk cargo port; establishing a cabin-cleaning and ship-unloading equipment cluster cross-cabin collaborative operation model for determining an equipment collaborative operation plan; acquiring and utilizing a large language model expansion training sample; training and adjusting a Q learning model based on the sample; based on a hyper-heuristic algorithm, LLM and QL are fused, an LQ-HH algorithm is designed and used for optimizing and solving the equipment collaborative operation model, and the operation sequence and starting time of cabin cleaning and ship unloading equipment are obtained. According to the method, the operation efficiency, the equipment cooperation capability and the anti-interference capability can be effectively improved, and intelligent transformation and upgrading of the dry bulk port can be promoted.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Tunnel anomaly detection and response method and equipment based on rail robot

The invention provides a tunnel anomaly detection and response method and device based on an orbital robot, and the method comprises the steps: obtaining scoring and grading data from a quantitative scoring model through a preset model optimization module, carrying out the dynamic adjustment of the parameters of the scoring model through a deep Q learning network in combination with historical response records, optimizing a grading threshold value to match an actual threat level, and carrying out the detection of an actual threat level. Transmitting the adjusted hierarchical data to a response scheduling system; and according to the adjusted hierarchical data, automatically matching a preset response strategy combination, if at least one selectable strategy combination exists, determining an optimal strategy combination through a weighted optimization algorithm, a comprehensive response speed, a resource utilization rate and a disposal cost, and issuing an execution instruction to a related execution unit.
Owner:JIE XUN TECH (GUANGZHOU) CO LTD

Outbound routing method and system based on multi-dimensional index coupling and large model decision

The invention provides an outbound routing method and system based on multi-dimensional index coupling and large model decision, and the method comprises the steps: carrying out the preprocessing of the historical data of each line provider, predicting the future performance index through an LSTM neural network, and optimizing the prediction result through a Kalman filter, and eliminating the noise. In the route decision-making link, the system calculates the Q value of each line provider based on the deep Q learning network, comprehensively considers the call completing rate, average cost, negative feedback rate and other factors of each line provider, and intelligently selects the optimal line provider in combination with the user portrait, scene recognition and real-time load state. The system also constructs a complete feedback closed loop, continuously optimizes the model by collecting call result data, and automatically adjusts decision parameters. The call completing rate and the service quality of the outbound service are improved, and the communication cost is reduced.
Owner:BEIJING YULORE INNOVATION TECH

Integrated feature selection method and product based on quantum computing and Bayesian optimization

The invention provides an integrated feature selection method and product based on quantum computing and Bayesian optimization, and relates to the technical field of data processing. According to the embodiment of the invention, an original mathematical integration model is converted into a QUBO model which can be solved by quantum calculation, and the QUBO model is decomposed into a joint optimization discretization step length optimization sub-problem and a parameterization QUBO sub-problem. In a mixed quantum classical optimization algorithm framework, a self-adaptive Q learning model is designed on the upper layer, and proper sub-problems can be dynamically selected in the search process. In the lower layer, a dropout Bayesian optimization algorithm is provided for effectively optimizing the high-dimensional discretization step length in each iteration. A CIM-based quantum computing method is adopted, and a parameterized QUBO sub-problem under the given discretization step length is efficiently solved. According to the method provided by the embodiment of the invention, the feature selection problem can be successfully and efficiently solved, and the selected features of the credit classification problem and the credit classification model for classification based on the selected features can be obtained.
Owner:BEIJING INST OF TECH +1

Linear motor temperature control method and system based on air cooling heat dissipation

The invention provides a linear motor temperature control method and system based on air cooling heat dissipation, and relates to the technical field of motor temperature control, and the method comprises the steps: dividing a linear motor into a plurality of temperature control regions, and configuring an independent control execution unit for each region; predicting a temperature trend based on the LSTM; a multi-agent reinforcement learning control network is constructed, and each agent calculates an optimal control strategy based on deep Q learning; generating an overall airflow regulation strategy through a collaborative optimization algorithm; dynamically adjusting the opening of the electric adjusting valve and the rotating speed of the fan. Accurate temperature control is achieved, the heat dissipation efficiency is improved, and the service life of equipment is prolonged.
Owner:SHENZHEN FERGUS ELECTROMECHANICAL EQUIP CO LTD

Two-stage excitation method based on multi-agent deep reinforcement learning

The invention discloses a two-stage excitation method based on multi-agent deep reinforcement learning, relates to the technical field of agent deep reinforcement learning, and aims to solve the problems that an excitation method in the prior art is too tedious in process and poor in effect. Through a simulation experiment in a specific area, taxies are used as participants to simulate real perception task geographical distribution and trajectory planning, comparison mechanisms such as deep Q learning and historical participation rate optimization are adopted, and multi-dimensional indexes such as average participant utility, platform utility and task completion rate are combined, so that the adaptability of the scheme in a dynamic environment is verified, and the accuracy of the scheme is improved. Cost is quantified by combining different sensor characteristics, participants can accurately evaluate energy consumption of themselves, platform utility is maximized under budget constraint, meanwhile, through bidirectional counterbalance of budget and task coverage, it is ensured that'minimum feasibility coverage 'of tasks and budget are not hyperbranched, actions of other agents can be observed during training to optimize cooperation, and the training efficiency is improved. The method only depends on a local state during execution, and adapts to a multi-participant competition scene.
Owner:HENAN UNIV OF SCI & TECH

Fault traffic electrification network emergency support strategy and system based on mixed graph attention-depth double Q learning architecture

The invention discloses a fault traffic electrification network emergency support strategy and a fault traffic electrification network emergency support system based on a mixed graph attention (Graph Attention Network, GAT)-deep double Q learning (DDQN) architecture, and aims to enhance the anti-interference capability and the recovery capability of a traffic electrification network when the traffic electrification network faces faults and emergency situations. Firstly, coupling characteristics and operation requirements of a traffic network and a power distribution network are comprehensively considered, a traffic electrification network emergency support framework based on a GAT-DDQN architecture is proposed, and a double-layer multi-objective optimization model is established. And secondly, learning fault topological structure information based on the graph attention network, constructing a double-layer finite Markov Decision Process (FMMP) model, modeling a dynamic interaction process of the traffic network and the power distribution network into a state, action and reward decision sequence, and realizing fault emergency support of the power distribution network and the traffic network. And finally, proposing a DDQN solving algorithm based on an attention mechanism and experience playback optimization, and outputting an optimal traffic electrification network emergency support strategy by strengthening key node feature learning and historical experience utilization efficiency. And finally, constructing a traffic electrification network emergency support system which comprises a data acquisition and processing module, a feature extraction module, an algorithm solving module and a visualization module.
Owner:NANJING UNIV OF POSTS & TELECOMM

Mobile robot adaptive path planning method based on enhanced Q learning and multi-strategy cooperation

The invention discloses a mobile robot adaptive path planning method based on enhanced Q learning and multi-strategy cooperation, and relates to the field of robot autonomous navigation and computing intelligence. The method comprises the following steps: firstly, establishing a two-dimensional grid map model containing obstacle information through an environment sensing module; the initial position of a gold scrubber population is generated by utilizing the low-difference characteristic of a golden section sequence to enhance the global ergodicity of a solution space, three stages of migration, gold scrubber and cooperation of a gold scrubber optimization algorithm are modeled into a discrete action space, an optimal search action is dynamically selected through a nonlinear attenuation greedy strategy, and the optimal search action is realized. According to the method, adaptive switching of exploration and development behaviors is realized, a search gradient is constructed by using a differential vector of survival of the survival of the survival of the population, and a lens imaging reverse learning strategy is introduced to dynamically disturb an elite individual so as to jump out of a local extremum trap. And finally, outputting an optimal collision-free path through a multi-objective evaluation function including path length, safety and smoothness.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Redundant executor dynamic scheduling method and system based on foreground theory and Q learning evolutionary game method

PendingCN121261962ABiological modelsSecuring communicationQ-learningIteration loop
The invention provides a dynamic scheduling method and system for a redundant executor based on a foreground theory and a Q learning evolutionary game method. The method comprises the following steps: constructing a strategy space, a profit structure and strategy evolution dynamics of the executor and a reward and punishment mechanism of an enhanced analyzer-regulator; according to the dynamic scheduling method, an executor simulates interaction under a tripartite game framework, iterative loop is carried out, and a strategy space of a mimicry defense system is output as a redundant executor after the loop is ended, the simulation interaction process comprises the following steps: the enhanced analyzer-regulator monitors and obtains objective income of each execution body in real time, calculates subjective perception utility of each execution body according to an income structure of the execution body, and then adjusts the subjective perception utility of the mimicry defense system according to a reward and punishment mechanism of the enhanced analyzer-regulator. And each execution body dynamically updates the strategy space of each execution body according to the adjusted subjective perception utility and strategy evolution.
Owner:WUHAN UNIV +4

Road investigation design method and system based on mobile network

The invention discloses a road investigation design method and system based on a mobile network, and relates to the technical field of road investigation, and the method comprises the steps: dividing geological anomaly probability distribution data into dynamic grid units, constructing a reinforcement learning state space, and generating an optimal sampling path instruction in combination with a dual-network deep Q learning architecture; analyzing the optimal sampling path instruction into executable parameters, executing the executable parameters, collecting multi-source spatio-temporal data, establishing data association through a spatio-temporal hash algorithm, and generating a reconnaissance data set; and on the basis of the survey data set, through parameterized spline curve modeling, generating a candidate road design scheme, and in combination with a non-dominated sorting genetic algorithm, performing optimization to generate a three-dimensional road design scheme. According to the method, the exploration path is optimized by using the dynamic grid coding and the dual-network architecture, the high-risk area coverage and the moving efficiency are balanced, meanwhile, the method adapts to real-time geological changes in combination with priority experience playback, and the exploration reliability and the resource utilization rate are improved.
Owner:JIANGSU SHIGUANG GEOGRAPHIC INFORMATION TECH CO LTD

Accurate control method and system for denitrification of sewage treatment plant and electronic equipment

The invention relates to the technical field of accurate control scheme design for denitrification of a sewage treatment plant, in particular to an accurate control method and system for denitrification of a sewage treatment plant and electronic equipment. Data of carbon source concentration, dissolved oxygen concentration, pH value, temperature, alkalinity and nitrate concentration in sewage are collected in real time, parameters of a prediction model are dynamically corrected in combination with a machine learning algorithm and a mechanism model, and the denitrification rate is predicted. The carbon source adding amount and the dissolved oxygen target value are calculated according to the prediction result and the emission standard, the distribution of the carbon source and the dissolved oxygen in the reaction tank is monitored in real time through a distributed sensor network, and the operation states of a carbon source adding pump and aeration equipment are accurately adjusted through feedback control. According to the method, a pH value and alkalinity are taken as constraint conditions to cooperatively control a carbon source feeding rate and a dissolved oxygen concentration, and a control strategy is dynamically optimized through deep Q learning based on historical data and operation conditions, so that cross-cycle self-adaptive adjustment is realized, and accurate control on a denitrification process of a sewage treatment plant is realized.
Owner:SHANDONG CHUANQINGQING ENVIRONMENTAL PROTECTION TECHNOLOGY CO LTD

Multi-edge task unloading method combined with task pre-sorting

The invention relates to a multi-edge task unloading method combined with task pre-sorting, and belongs to the technical field of edge calculation and task unloading. The method comprises the following steps: firstly, establishing a multi-edge, multi-user and multi-task network architecture comprising a cloud layer, an edge layer and a user layer; based on network architecture support, an application model, a transmission model and an execution model for an application task of a user are established, the application model represents any subtask as a triple, the transmission model defines a corresponding transmission rate, and the execution model comprises a local execution model, an edge execution model and an overall cost model; a task pre-sorting mechanism is introduced, and subtasks which can be parallel are sorted in a descending mode according to priority factors; and finally, based on the ordered task sequence after sorting, performing task unloading strategy decision by adopting a deep Q learning network. According to the method, joint optimization of task execution delay and energy consumption can be realized in a multi-user, multi-task and multi-edge node computing environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Broadband satellite communication verification system and method

The invention discloses a broadband satellite communication verification system and method, relates to the technical field of broadband satellite communication, and remarkably improves the response efficiency and stability of the broadband satellite communication verification system through a distributed asynchronous Q learning framework and a dynamic coordination mechanism. In a high dynamic environment of a satellite-ground fusion network, a local decision agent operates independently, a resource allocation decision is quickly generated based on a real-time channel state and a user demand, and a calculation bottleneck brought by centralized optimization is avoided; according to the distributed architecture, communication overhead among nodes is reduced, decision delay is reduced, and therefore the jitter phenomenon caused by multi-user competition is effectively restrained; a penalty function mechanism integrates the conflict rate, the interference degree and other multi-dimensional costs into a learning process, so that the nodes autonomously avoid high-competition actions, and the fairness of resource allocation is enhanced; meanwhile, the coordination module adaptively adjusts the updating period through topological change measurement, and the response interval is shortened when the satellite rapidly moves.
Owner:AEROSPACE FUTURE (SHENZHEN) AEROSPACE TECHNOLOGY CO LTD

Slope cutting house building landslide vulnerability analysis method and system based on reinforcement learning

The invention discloses a slope-cutting house-building landslide vulnerability analysis method and system based on reinforcement learning. The method comprises the following steps: preprocessing multi-source data to obtain a disaster sample data set, a missing tag vector and a damage level tag; feature parameters of the slope and the house are extracted, and missing mark vectors are fused to generate a state vector set; constructing an action set and calculating an external return according to the difference between the predicted damage distribution and the damage level label; calculating a conservative penalty quantity for non-data actions by utilizing the action value network and the target action value network in conservative Q learning; calculating a novelty error in the random network distillation and generating an intrinsic return; and fusing the return to construct a state transition sample, updating the action value network to obtain a conservative action value function, and outputting a vulnerability conclusion of the scene to be analyzed. According to the method, stable training and iteration of the evaluation model are realized under the conditions of sample scarcity and data incompleteness.
Owner:HOHAI UNIV

A method, management device and storage medium for suppressing subsynchronous oscillation of a doubly-fed wind farm based on intelligent learning and predictive control

The application discloses a double-fed wind farm subsynchronous oscillation suppression method based on intelligent learning and predictive control. First, based on double-fed wind farm dominant oscillation mode analysis, double-fed wind turbine rotor side converter outer ring damping control prediction model and inner ring damping control prediction model are analyzed; a damping optimization function approximator is constructed based on a robust model predictive control, and a double-fed wind turbine outer ring damping optimization control module and an inner ring damping optimization control module are established; then, based on Q learning optimization, optimal control laws of the double-fed wind turbine outer ring damping optimization control module and the inner ring damping optimization control module are solved, and the optimal control laws are stored based on data cubes; finally, the double-fed wind turbine outer ring damping optimization control module and the inner ring damping optimization control module are incorporated into the inner and outer ring control of the rotor side converter, and double-fed wind farm subsynchronous oscillation suppression is realized. The application considers the uncertain influence of the control system parameters, can realize double-fed wind farm subsynchronous oscillation suppression, and avoids oscillation propagation.
Owner:STATE GRID SICHUAN ELECTRIC POWER CO

A radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method

The application discloses a radar space-time-frequency-energy multi-domain joint intelligent active anti-jamming method, first establishes a confrontation scene under the radar and multi-jammer environment, then sorts out the model parameters and working modes of the radar and the jammer, determines the evaluation index to complete the evaluation of the radar scheduling strategy, and then converts the problem into a Markov decision process, and uses Q learning to solve the optimal strategy of radar multi-domain resource scheduling. The optimization strategy obtained by the method has strong adaptability and good model scalability, and can be adjusted according to the actual need of the scheduled parameter resource. Through the collaborative control of the transmission parameters between the multiple domains of the radar, compared with the existing single-domain resource scheduling anti-jamming method, the radar active anti-jamming capability can be significantly enhanced while improving the task performance, and the potential task performance of the radar is excavated. Meanwhile, Q learning is used to optimize the radar node multi-domain joint active anti-jamming strategy, and the survival ability of the radar in the electronic countermeasure environment is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A hierarchical path planning method and system based on a three-dimensional graph network

The present application relates to the technical field of unmanned system autonomous navigation and path planning, in particular to a hierarchical path planning method and system based on three-dimensional graph network, a three-dimensional environment semantic graph containing geometric and risk semantic information is constructed; then a down-sampling strategy and an improved AStar algorithm are used in the global layer, a ray collision detection mechanism is used to quickly generate a global waypoint sequence under a coarse resolution grid, solving the search efficiency bottleneck in a large-scale three-dimensional space; in the local layer, a Q learning local decision method guided by artificial potential field is proposed, a lazy loading sparse storage is realized by initializing a Q table with expert experience, and global waypoint guidance and local obstacle avoidance logic are combined. The present application realizes the decoupling and fusion of global optimality and local real-time obstacle avoidance through hierarchical architecture, effectively solves the problems of calculation explosion, local extremum and non-connected space traversal in the three-dimensional complex environment of traditional algorithms, and improves the planning efficiency and safety of the agent.
Owner:OCEAN UNIV OF CHINA

Networking radar jamming and mutual jamming elimination method based on decentralized q learning

The application discloses a networking radar jamming and mutual jamming elimination method based on decentralized Q learning, first, according to the phased array radar signal processing flow, a target echo signal, a jamming signal and a mutual jamming signal model are established, second, the conditions of radar jamming and mutual jamming are analyzed, a multi-radar frequency domain resource scheduling optimization model based on the SIJNR criterion is established, then each radar is regarded as an intelligent agent, the multi-radar cooperative anti-jamming and anti-mutual jamming process is constructed into a generalized Markov decision process, and the frequency domain resource scheduling optimization model is converted into a value function optimization model, finally, the DQJIE algorithm is used to solve the problem, and the frequency agility strategy of each radar is obtained. The method constructs the multi-radar cooperative anti-jamming and anti-mutual jamming process into a generalized Markov decision process, effectively suppresses the mutual jamming in the system while resisting the frequency sweeping jamming, and effectively improves the networking radar detection performance.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Spectrum analyzer parameter automatic adjustment method and system based on deep learning

The application discloses a spectrum analyzer parameter automatic adjustment method and system based on deep learning, receives each parameter of a spectrum analyzer and a range thereof and signal parameters sent by a signal generator; a deep Q learning network model is constructed, training parameters and training conditions of the deep Q learning network are set according to each parameter of the spectrum analyzer and the range thereof and the signal parameters; the deep Q learning network model is trained based on the training parameters and the training conditions, each parameter of the spectrum analyzer is adaptively adjusted within the range of each parameter of the spectrum analyzer, an optimal parameter configuration action is selected as an optimal spectrum analyzer parameter combination of the spectrum analyzer through Q value prediction, key parameters of the spectrum analyzer are automatically adjusted, and measurement performance is optimized. The method can not only reduce human intervention, but also adapt to various complex signal environments, and provides a new solution for spectrum measurement automation and intelligentization.
Owner:成都华兴汇明科技有限公司

A deep q-learning-based rsu-assisted video data routing method

The application provides a kind of RSU auxiliary video data routing method based on deep Q learning, establishes system model, and mobile source vehicle collects real-time road condition video, and it is forwarded to the RSU at the intersection connected with traffic management department by SVC coding;Carry out problem description, select suitable path forwarding for different layer video, maximize the MOS metric of each GOP under the delay constraint, describe the problem as: source vehicle selects the first RSU to be reached by the data to be forwarded according to the decision information of RSU;On the basis of deep Q learning on RSU, the next RSU between adjacent road sections is selected as the forwarding direction for different layer SVC video data;The forwarding in road section is selected according to frame type based on the middle intelligence set analytic hierarchy process to select the best relay vehicle, through multiple inter-road section decision and intra-road section forwarding, until the data reaches the destination RSU.The application can effectively achieve better overall performance in different vehicle density and different network load scenarios.
Owner:HENAN UNIV OF SCI & TECH

A reinforcement learning supported collaborative learning intervention method and system

The application provides a collaborative learning intervention method and system under reinforcement learning, comprising the following steps: selecting an intervention action from an intervention action space based on a perceived state of a collaborative discussion environment; obtaining a reward and a next moment state fed back by the collaborative discussion environment after executing the intervention action, and then obtaining transition data; the state comprises collaborative sessions of all learners in the collaborative discussion environment; repeating the above steps until a preset number of transition data are obtained; training a Q network based on the preset number of transition data; in the training process, an improved Q learning method is used to update a Q value function; using the trained Q network to determine a current intervention action corresponding to a maximum Q value in a current state; and providing an intervention suggestion to the learner based on the current intervention action. The application can provide learning support for learners who do not fully master the meta-cognitive strategies of collaborative learning, and promote the collaborative activities among learners and the development of individual meta-cognitive ability.
Owner:HUAZHONG NORMAL UNIV

Unmanned ship closed-loop path planning method and system based on multi-agent potential field and deep Q learning

The invention relates to the technical field of artificial intelligence, in particular to an unmanned ship closed-loop path planning method and system based on a multi-agent potential field and deep Q learning. According to the method, a multi-agent potential field including an attraction potential field and a repulsion potential field is constructed by using state information of multiple unmanned ships, and an action evaluation function is formed based on a potential field gradient and an epsilon-greedy strategy. And a single-boat target optimal action strategy is obtained by establishing a double-Q dynamic adjustment network model and combining deep Q learning iterative updating. On the basis, a collaborative strategy sharing mechanism of potential field gradient, position relation and state interaction among multiple agents is introduced, and cluster-level collaborative path planning is achieved. And finally, driving the unmanned ship to complete forward and return tasks according to a collaborative planning result, and generating a closed-loop optimal path in combination with a layered reward mechanism in a return stage. According to the invention, the autonomous decision-making capability, the whole-process autonomous navigation capability and the global cooperation efficiency of the unmanned ship cluster in a complex task sea area are improved.
Owner:JIMEI UNIV

An aircraft fault-tolerant control method based on reinforcement learning

The application discloses an airplane fault tolerance control method based on reinforcement learning and belongs to the field of fault control. The application uses a deep Q learning network to extract features of data generated in a simulation environment, so as to identify fault information and adjust the deep Q learning network. The fault diagnosis method based on data reduces the difficulty of modeling a complex system, replaces a fault detection subsystem by extracting data features, and simplifies the deep Q learning network. The application expands the collected signals by using an enhanced algorithm SDA and then expands the training set, so as to improve the accuracy of the model and the generalization ability of the model. The application optimizes the reinforcement learning model by using the deep Q learning network, and compared with a traditional reinforcement learning method, the optimization method significantly improves the data feature extraction ability of the model.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Aviation luggage online stacking method and device based on deep reinforcement learning

The invention discloses an air luggage online stacking method and device based on deep reinforcement learning, and belongs to the technical field of air luggage stacking. The method comprises the following steps: training a dual-depth Q learning network by using a plurality of training samples; the training mode comprises the steps that a first luggage queue and a height map of a training sample are input into a network, the network constructs a state space based on the height map and the first luggage queue, to-be-stacked luggage is selected from the first luggage queue one by one, and for the to-be-stacked luggage selected each time, the following steps are executed: the action space of the to-be-stacked luggage is determined, and the action space of the to-be-stacked luggage is determined; updating the state space by using the action space; and the stacking space utilization rate is determined, and the stacking mode is output based on the maximum space utilization rate after iterative updating. According to the method, the space utilization rate of luggage packing can be greatly improved, real-time decision making can be carried out in a complex dynamic online environment, the method can adapt to heterogeneous luggage of different sizes, shapes and packing quantities, and robustness is guaranteed.
Owner:CIVIL AVIATION FLIGHT UNIV OF CHINA

Path planning method and device and electronic equipment

The invention relates to the technical field of artificial intelligence and automatic control, and discloses a path planning method and device and electronic equipment, and the method comprises the steps: carrying out the dynamic adjustment of a learning rate and an exploration rate in a Q learning model based on the Q learning model; on the basis of Monte Carlo tree search, path selection in the Q learning model is simulated, and an accumulated discount reward of a simulated path is obtained; based on the accumulated discount rewards, correcting the Q learning model to obtain a path planning correction model; and planning the path of the intelligent agent based on the path planning correction model. According to the invention, the convergence speed of intelligent agent path planning can be improved, the dynamic obstacle avoidance efficiency is enhanced, and the real-time decision-making capability of the intelligent agent in a complex environment is improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Cluster scheduling method based on nerve attention enhanced Q learning

The invention discloses a cluster scheduling method based on neural attention enhanced Q learning. The cluster scheduling method comprises the following steps: setting constraint conditions of cluster scheduling of storage robots; each storage robot comprises a neural network architecture fusing a gating circulation unit and a multi-layer sensor; by adopting a parameter sharing strategy, all robots share the same set of network weight for learning; an attention mechanism based on a variational automatic encoder is used, and the attention degree of the robot to a task related area is enhanced; calculating a joint effect value; using a resolver to generate an independent learning sample for each robot; the learning samples are used to train a neural network in a multi-agent system. According to the method, the interpretability model and the decomposer model are organically combined, so that the overall performance of the model is improved; the problem that the performance of an interpretable method is limited in a complex environment with a large number of agents is solved through the decomposer.
Owner:NAT UNIV OF DEFENSE TECH

Power monitoring system and method for carbon comprehensive management

The invention relates to the technical field of power monitoring, in particular to a power monitoring system and method for comprehensive carbon management, and the system comprises a sensing system, a network transmission system, a control system and a data monitoring system. The sensing system is used for sensing environment data in the building and power equipment operation data; the network transmission system is used for realizing communication among the sensing system, the control system and the data monitoring system; the data monitoring system is used for fusing the acquired environment data and the power equipment operation data based on a data fusion algorithm of a DS evidence theory, and processing fused information by adopting an intelligent decision algorithm based on Q learning so as to obtain an autonomous control decision; the control system is used for controlling the power equipment based on the generated autonomous control strategy; and the problems of information conflict and uncertainty between different types of sensors can be solved.
Owner:HEPU SIDER ENERGY CONSULTING (SHANDONG) CO LTD

Hierarchical clustering Q learning-based ship equipment spare part demand prediction method

The invention discloses a hierarchical clustering Q learning-based ship equipment spare part demand prediction method. The method comprises the following steps of 1) obtaining historical demand data of ship equipment spare parts; 2) preprocessing the collected data set, and identifying and processing missing values and abnormal values in the data; 3) dividing the data sets into a training data set and a test data set, and respectively performing clustering operation on the two data sets; 4) constructing a spare part prediction model; 5) training a spare part prediction model of Q learning by using the training data set; and 6) based on to-be-predicted ship spare part demand data, using the trained model to carry out subsequent ship spare part demand data prediction. According to the hierarchical clustering Q learning-based spare part demand prediction method provided by the invention, the optimal prediction model can be automatically selected according to different spare part feature distribution, the prediction error of the spare part demand is reduced, and the purpose of reducing the spare part inventory cost is achieved.
Owner:CHINA SHIP DEV & DESIGN CENT

Heterogeneous resource collaborative optimization method based on parallel reinforcement learning

The invention discloses a heterogeneous resource collaborative optimization method based on parallel reinforcement learning, and the method comprises the steps: formalizing a heterogeneous resource distribution problem into a Markov decision process, defining a state space, an action space, a reward function and a state conversion mechanism, and constructing a dynamic decision model based on Q learning; classifying the task targets and the heterogeneous resources according to the types of the task targets to obtain a task target set and a heterogeneous resource set corresponding to the task target types; for each task target type, in combination with the corresponding task target set and the heterogeneous resource set, performing parallel reinforcement learning training by adopting a dynamic decision model based on Q learning to obtain an optimal dynamic decision model of each task target type; and based on the optimal dynamic decision model of each task target type, obtaining an optimal resource allocation scheme of each task target type, and finally summarizing to form a global optimal resource allocation scheme. According to the invention, the problem of large-scale multi-domain collaboration can be efficiently solved.
Owner:WUHAN UNIV OF TECH

Maintenance troubleshooting method for complex system

PendingCN120894007AMathematical modelsForecastingObservation - actionFlight vehicle
The invention provides a complex system-oriented intelligent maintenance troubleshooting method combining a Bayesian network and double-Q learning, and aims to solve the problems of low decision-making efficiency, poor path and high maintenance cost of a traditional heuristic strategy. The method comprises the following steps: firstly, constructing a Bayesian network model according to a system structure and a causal dependency relationship, and setting a prior probability of each node in combination with an FMECA analysis result; and then, a troubleshooting sample is generated by using a Monte Carlo method, a double-Q learning model is trained offline, observation and maintenance costs are fused, and a maintenance action strategy is optimized. In order to improve strategy convergence and stability, an observation action screening mechanism and an action shielding technology are introduced, and dynamic decision optimization in the multi-step troubleshooting process is achieved. The method has a causal modeling capability and reinforcement learning adaptivity, has relatively high robustness and real-time performance, and is suitable for intelligent maintenance and fault diagnosis scenes of key equipment such as aircrafts and automobiles.
Owner:BEIHANG UNIV