Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1787 results about "Decision process" patented technology

Multi-modal fusion and reinforcement learning collaborative retrieval enhancement generation method and system

The invention relates to the technical field of information retrieval, and discloses a multi-modal fusion and reinforcement learning collaborative retrieval enhancement generation method and system. The method comprises the following steps: receiving an original query input by a user, and generating a sub-query based on a large language model in combination with a multi-modal context of a current iteration step; forming a current state in combination with the sub-query and the multi-modal context, modeling a retrieval enhancement generation task as a Markov decision process, and adaptively selecting an optimal action from a predefined action set in the current state by utilizing a large language model according to a decision strategy; executing a corresponding multi-modal retrieval operation according to the optimal action, fusing the obtained multi-modal information, generating an intermediate answer or a final answer of the sub-query, and updating a multi-modal context by using the intermediate answer; off-line training optimization is carried out on the large language model through imitation learning and a calibration chain, and decision strategies and sub-queries are inferred online through the model after fine adjustment. According to the invention, more efficient and accurate complex query processing is realized.
Owner:DATA SPACE RES INST

Intelligent flow arrangement method based on fusion expert network and deep reinforcement learning

The invention discloses an intelligent flow arrangement method based on fusion expert network and deep reinforcement learning, which comprises the following steps: collecting network node and link state data in real time, and constructing a time sequence input vector and a topological graph structure; a time sequence neural network and a graph neural network are used for extracting traffic spatial-temporal features and node topological features respectively, future traffic is predicted through a classification network after fusion, and coarse-grained arrangement of network slices of different service levels is completed; modeling resource scheduling into a multi-agent Markov decision process, and designing a state space, an action space and a reward function; a deep reinforcement learning agent is initialized, and training is carried out through interaction experience; fusing a pre-trained expert strategy network, and constructing a total loss function to optimize network parameters; and finally generating an intelligent strategy capable of dynamically optimizing the flow path and resource allocation according to the real-time state. According to the invention, efficient resource scheduling under multi-service differentiation service quality requirements can be realized.
Owner:NARI INFORMATION & COMM TECH

Power distribution network fault transfer optimization method fusing knowledge base under participation of virtual power plant

The invention relates to the technical field of power system fault recovery, in particular to a power distribution network fault transfer optimization method fusing a knowledge base under the participation of a virtual power plant, and the method comprises the steps: firstly modeling a power distribution network fault transfer process into a Markov decision process to construct a power grid environment model, and then extracting power grid topological features through a graph neural network; the method comprises the following steps: extracting and fusing time sequence features in combination with a Transform structure, then introducing expert knowledge to carry out imitation learning, providing an initial strategy for an intelligent agent, then adopting PPO and DQN cooperative training to optimize an intelligent agent strategy, finally aggregating distributed energy with the help of a virtual power plant, realizing resource coordination and fault load transfer, and dynamically correcting the strategy through closed-loop feedback. Therefore, dynamic adaptability, resource cooperation efficiency and strategy reliability of power distribution network fault recovery are improved, power supply recovery time is shortened, and safe and stable operation of a power grid is guaranteed.
Owner:HEFEI POWER SUPPLY COMPANY OF STATE GRID ANHUI ELECTRIC POWER

Heterogeneous computing resource scheduling method and apparatus based on multi-objective optimization

The present application belongs to the technical field of parallel task scheduling. Specifically, disclosed are a heterogeneous computing resource scheduling method and apparatus based on multi-objective optimization. The method comprises: selecting at least two performance indicators from both a task dimension and a resource dimension as optimization objectives, and establishing a multi-objective optimization model for heterogeneous computing resource scheduling; converting a computation task request process into a computation task waiting model on the basis of a queuing theory; on the basis of an observed resource state, constructing a multi-task adaptive scheduling model based on reinforcement learning; on the basis of the multi-objective optimization model and the computation task waiting model, constructing a Markov decision process model from a multi-task-oriented heterogeneous computing resource scheduling problem and a resource mapping process; and on the basis of the Markov decision process model and the multi-task adaptive scheduling model, realizing adaptive multi-task heterogeneous computing resource scheduling. The embodiments of the present application solve the problem of it being difficult for a homogeneous computing resource scheduling method to adapt to heterogeneous computing resource scheduling.
Owner:709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD

Cement equipment maintenance decision-making method and device based on knowledge graph and large model reasoning

The invention provides a cement equipment maintenance decision-making method and device based on a knowledge graph and large model reasoning, relates to the field of cement industry intelligent operation and maintenance, and solves the technical problem of decision-making response delay caused by knowledge fragmentation. The method comprises the following steps: extracting real-time characteristics from vibration spectrum signals, temperature curves and torque waveform data collected by an edge gateway, and extracting a work order entity triple from a natural language work order text of an EAM system; based on an equipment BOM list, a historical maintenance record and an FMEA analysis table, physical assembly constraint conditions are defined through ontology modeling to generate a cement equipment topological relation and a fault rule chain, and a knowledge graph is created to output a fault rule base with confidence coefficient weights. And inputting the real-time feature vector and the work order entity triple into a multi-modal collaborative inference engine, triggering a matched fault rule chain by combining real-time features and semantic features, outputting a fault root cause and an associated maintenance strategy ID, and labeling a logic chain. And activating the associated maintenance strategy ID and obtaining the real-time characteristic deviation degree of the maintenance strategy ID, quantifying the decision credibility through a tracing rule matching path, obtaining an executable maintenance instruction packet with a logic chain, executing the maintenance instruction packet and dynamically updating the knowledge graph based on a maintenance result. The method is used in the maintenance decision-making process of the cement equipment.
Owner:HEFEI CEMENT RESEARCH AND DESIGN INSTITUTE CO LTD

TSN scheduling optimization method and device based on flow sensing autonomous learning, equipment and medium

The invention discloses a TSN scheduling optimization method and device based on flow sensing autonomous learning, equipment and a medium. The method comprises the following steps: deploying a lightweight flow detection module in a switch or a router, and after a controller receives a data request, automatically identifying a newly arrived unknown service flow by using the lightweight flow detection module, and judging whether the newly arrived unknown service flow is a periodic TT flow or an unpredictable burst flow; the controller performs classification management on the identified service flow types, collects topological information and flow requirements of the whole network and issues the topological information and the flow requirements to the terminal nodes through a network management interface; the controller constructs an intelligent queue scheduling task based on the collected network topology information and traffic demand and converts the task into a Markov decision process MDP, network resources, queue states and priorities are used as a state space, a scheduling strategy is used as an action space, a reward function is designed in combination with throughput and delay indexes, and an intelligent queue scheduling task is obtained. Driving a dynamic environment through real-time data and training a DRL model; an enhanced queue scheduling mechanism Pro-CQF is adopted, different priority labels are configured according to classified flow types, and then mixed flow scheduling is carried out; the controller periodically collects time delay, packet loss rate and end-to-end transmission delay indexes and feeds back the indexes to the DRL model, and a scheduling strategy is updated online. According to the method, the traffic sensing and scheduling efficiency is greatly improved in a network environment in which multiple service flows coexist and end-side equipment functions are different, and the reliability and the expandability of the TSN in industrial Internet of Things, edge computing and other high-real-time application scenes are remarkably enhanced.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Large model business logic processing method and system based on workflow engine and domain knowledge fusion

The invention relates to a large model business logic processing method and system based on workflow engine and domain knowledge fusion, and is suitable for automatic processing of complex multi-node and multi-branch business processes. According to the method, natural language input of a user is analyzed through a large language model, a service intention is recognized, and the service intention is converted into an executable task process through a workflow engine. And the workflow engine dynamically adjusts an execution path according to rules and data in the domain knowledge graph to realize efficient parallel processing of tasks. The system integrates a workflow engine, a domain knowledge graph and a large language model, supports real-time data processing, rule matching, conflict resolution and decision optimization, is suitable for the fields of water conservancy, medical treatment, finance and the like, automatically generates and executes a complex business process, and optimizes a decision process. The system improves the accuracy and execution efficiency of business decision through deep fusion of domain knowledge.
Owner:JIANGHE RUITONG (BEIJING) TECH CO LTD

AI agent emergency order insertion dynamic decision production scheduling method, medium and system

The invention provides an AI agent emergency order insertion dynamic decision production scheduling method, a medium and a system, and belongs to the technical field of industrial agents. A dynamic weight adaptive optimization model is adopted to calculate a target weight coefficient and construct a multi-target function set, an improved non-dominated sorting genetic algorithm is adopted to solve and output a Pareto optimal solution set, and a delay risk assessment correlation matrix is combined to start an incremental re-planning algorithm to generate a local adjustment scheme. The Pareto optimal solution set and the local adjustment scheme are combined to generate a final production scheduling scheme, a real-time monitoring module is started to track the execution deviation condition, and when it is detected that the deviation degree exceeds a threshold value, a rapid rescheduling mechanism is triggered to conduct scheme correction; the technical problem of poor production scheduling scheme quality caused by low multi-agent cooperation efficiency in the emergency order insertion dynamic decision process is solved.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

Heterogeneous capacity constraint unmanned aerial vehicle path planning method and device based on deep reinforcement learning, and storage medium

The invention discloses a heterogeneous capacity constraint unmanned aerial vehicle path planning method and device based on deep reinforcement learning and a storage medium. According to the method, a path planning problem is modeled as a Markov decision process, and an encoder-double decoder strategy network based on an attention mechanism is adopted for solving. Wherein the encoder network is used for carrying out feature extraction on all task node information so as to generate task node embedding containing a global dependency relationship; in each decision-making step, the dual-stage decoder network firstly determines an optimal execution unmanned aerial vehicle through an unmanned aerial vehicle selection decoder, then selects a next target task node for the unmanned aerial vehicle through a node selection decoder, and generates a complete path scheme in an iteration mode. According to the scheme, a high-quality and high-robustness path planning scheme can be efficiently and intelligently generated, heterogeneous constraint and three-dimensional space problems are effectively processed, and good generalization ability is achieved.
Owner:TSINGHUA UNIVERSITY

Civil aircraft component fault diagnosis method based on condition diffusion in variable working condition scene

PendingCN120974314AData spaceEngineering
A civil aircraft component fault diagnosis method based on conditional diffusion in a variable working condition scene comprises the steps that limited high-confidence-coefficient fault samples and label-free fault data under different working conditions are collected, after hidden space feature reconstruction is carried out, a denoising diffusion probability model is adopted to carry out conditional generation augmentation on target domain high-confidence-coefficient fault samples of a target domain, and the target domain high-confidence-coefficient fault samples of the target domain are obtained; and through joint optimization of an objective function, a mapping relation between a generation feature and a diagnosis decision is restrained while fault sample condition distribution is deduced, and collaborative improvement of generation quality and model generalization ability is realized. According to the method, the fault sample generation and fault cross-domain decision-making process is jointly optimized through the conditional diffusion-based fault diagnosis algorithm (CDFD), so that the fault diagnosis classifier can fully sense noise hidden space and data space characteristics of a source domain and a target domain, and the robustness is effectively improved.
Owner:SHANGHAI JIAOTONG UNIV

Multi-source knowledge conflict detection and repair method and system based on semantic graph

The invention discloses a multi-source knowledge conflict detection and restoration method and system based on a semantic graph, and the method comprises the steps: combining a received user query, multi-source context data and a predefined prompt word template, and inputting the combination into a conflict detection model to generate a structured conflict report; analyzing the conflict report, matching a corresponding repair operator from a predefined repair operator set according to the conflict type, and executing the matched repair operator to dynamically repair the multi-source context data to obtain the repaired context data and generate a decision log; a final response to the user query is generated based on the repaired contextual data. According to the invention, a lightweight detection model based on knowledge distillation is adopted, so that the real-time requirement and the detection precision requirement of online service are met; through a dynamic operator combination framework, repair strategies are automatically selected for different types of conflicts, and the processing efficiency is improved; conflict reports and detailed decision logs are structured, and a transparent and traceable decision process is provided.
Owner:彩讯科技股份有限公司

GPU cluster scheduling strategy optimization system based on deep reinforcement learning

The invention discloses a GPU cluster scheduling strategy optimization system based on deep reinforcement learning, which comprises a simulation environment layer, a reinforcement learning layer and a strategy evaluation layer, and is characterized in that the simulation environment layer is used for providing a real and credible GPU cluster scheduling environment, reproducing a core mechanism of actual cluster scheduling and simultaneously providing controllable experiment conditions; comparative evaluation and iterative optimization of strategies are facilitated; the reinforcement learning layer converts a GPU cluster scheduling problem into a Markov decision process based on a deep neural network, and obtains an optimal scheduling strategy by applying a reinforcement learning strategy; and the strategy evaluation layer counts each key index, compares and analyzes a plurality of scheduling strategies, and visually displays an analysis result to realize optimization of the GPU cluster scheduling strategy. The system solves the time sequence short view problem of a traditional scheduler, so that the scheduling decision can consider the influence on the future, global optimization instead of local optimization is realized, and the overall resource utilization efficiency is improved.
Owner:UNIV OF SCI & TECH OF CHINA

Personified automatic driving simulation test scene construction method

The invention provides an anthropomorphic automatic driving simulation test scene construction method. The method comprises the steps that all vehicles in a traffic scene are divided into two types of intelligent agents including a test vehicle and an environment vehicle, the traffic scene is modeled into a Markov decision process, the test vehicle is a vehicle controlled by an automatic driving algorithm, and the environment vehicle is a vehicle controlled by a GAIL-GRU driving strategy model; a Markov decision process is utilized to extract a driving track from the human driving data set, and an expert track data set is generated; and training a GAIL-GRU driving strategy model by using the expert track data set, realizing interaction between an environment vehicle and a test vehicle by using the trained GAIL-GRU driving strategy model, and constructing a simulation test scene. The automatic driving simulation test scene constructed by the invention can effectively expose decision defects of an automatic driving algorithm in a complex interaction situation, has good anthropomorphism and relatively high risk scene coverage, and provides support for a decision control simulation test of a high-level automatic driving vehicle.
Owner:BEIJING JIAOTONG UNIV

Ultra-high performance concrete multi-performance prediction method based on machine learning

The invention provides an ultra-high performance concrete multi-performance prediction method based on machine learning. The ultra-high performance concrete multi-performance prediction method comprises the following steps: Step 1, establishing a data set; step 2, data preprocessing is carried out; step 3, establishing an optimal prediction model: based on the feature subset, adopting a plurality of different machine learning algorithms for training, and selecting the machine learning algorithm with the best training effect as the optimal prediction model; step 4, selecting an optimal feature subset; step 5, explaining the influence of the features on model prediction: calculating the contribution degree of each feature to a prediction result based on the optimal prediction model and the optimal feature subset, and helping to understand the decision process of the model; and Step 6, performance prediction of the ultra-high performance concrete: inputting parameters of the to-be-predicted ultra-high performance concrete into the optimal prediction model to obtain a predicted value of the performance. The technical problems that an existing UHPC performance prediction method is incomplete in data set, insufficient in consideration of data processing and feature engineering and poor in model interpretation can be solved.
Owner:XINJIANG BINGTUAN CONSTR ENG CO LTD +1

Counter measure strategy construction method and system based on attack intelligence

The invention belongs to the technical field of network security, and discloses a countering measure strategy construction method and system based on attack intelligence. According to the method, firstly, multi-source attack intelligence is fused, and a causal time sequence attack knowledge graph is constructed; then, mapping the atlas on a digital twinborn model based on the IEC 62443 standard, constructing a Bayesian attack graph, and quantifying the risk of each attack path; then, taking the risk, the cost and the operation influence as multiple targets, and adopting an NSGA-II algorithm to generate a Pareto optimal countering strategy set; further, an optimal robust strategy is selected from the strategy set by solving the Stackelberg safety game model; and finally, explaining the decision process by using an interpretable AI technology. According to the method, the problems of passive threat intelligence analysis, static risk assessment, sub-optimal strategy selection and opaque decision-making process in the prior art are solved, and active, quantitative, optimal and credible defense decision-making for the key infrastructure is realized.
Owner:GUANGXI POWER GRID CORP

Asset allocation recommendation system, method and equipment based on multi-modal data fusion

The invention discloses an asset configuration recommendation system, method and device based on multi-modal data fusion, relates to the technical field of asset configuration, and can be applied to financial science and technology business scenarios, and the method comprises the following steps: obtaining multi-source heterogeneous data including structured data and unstructured data; preprocessing the multi-source heterogeneous data and generating an environment feature vector; designing a reward function based on the environment feature vector, utilizing the reward function to drive a reinforcement learning framework to simulate an asset configuration decision process, and in the simulation process of investment decision, iteratively optimizing strategy parameters through reward feedback to realize dynamic configuration of an asset configuration strategy; and according to the personalized parameters and the transaction behavior characteristics of the target object, carrying out weight calibration on an asset configuration strategy output by the reinforcement learning strategy module, and generating a customized asset configuration recommendation result for the target object. According to the invention, dynamic updating of asset configuration can be realized, and the accuracy and adaptability of asset configuration are improved.
Owner:PING AN HEALTH INSURANCE CO LTD

Multi-dimensional regulation and control decision-making method, system and equipment for power distribution network and medium

The invention relates to the technical field of power systems, and provides a power distribution network multi-dimensional regulation and control decision method, system and device and a medium, and the method comprises the steps: inputting the preprocessed multi-source operation data into a preset state perception model, and obtaining a multi-dimensional state vector representing the operation state of a power distribution network; a multi-dimensional state vector is used as a state space, regulation and control operation is used as an action space, a composite reward function is established according to a power distribution network operation target, and modeling is carried out to obtain a Markov decision process framework; interacting with a power distribution network simulation environment by adopting a deep reinforcement learning algorithm, obtaining a current state from a state space, selecting and executing regulation and control operation in an action space according to a strategy network, updating strategy network parameters based on feedback of a composite reward function until an optimal regulation and control strategy network is obtained, and obtaining a deep reinforcement learning strategy model; and performing strategy rolling updating based on the real-time monitoring data to obtain a target regulation and control strategy. According to the invention, comprehensive optimal regulation and control of a complex operation scene can be realized.
Owner:FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID

New energy power grid look-ahead scheduling method and device

The invention provides a new energy power grid prospective scheduling method and device, and relates to the technical field of electric power system intraday economic scheduling. The method comprises the following steps: firstly, constructing an opportunity constraint optimal power flow model of prospective scheduling, analyzing the influence of new energy uncertainty on opportunity constraint, establishing a constrained Markov decision process of prospective scheduling, and then utilizing a risk evaluator network fitting risk function probability distribution and an actuator network considering extreme scene performance to determine the opportunity constraint optimal power flow model of prospective scheduling. The processing capability of the intelligent agent on a prospective scheduling scene containing a new energy extreme climbing event is enhanced; and finally, the training of the intelligent agent is accelerated by utilizing an imitation learning technology in a power grid prospective scheduling off-line simulation environment. According to the method, the solving speed and the strategy robustness and safety of the double-layer robust optimization model of the look-ahead scheduling can be considered.
Owner:WUHAN UNIV

Business travel journey automatic optimization method

The invention discloses an automatic business travel itinerary optimization method, and relates to the technical field of intelligent itinerary planning, and the method comprises the steps: integrating the multi-source heterogeneous data of enterprise policies, personal preferences and real-time traffic through a federated learning framework, and achieving the cross-domain knowledge sharing; the method comprises the following steps: constructing a staged optimization engine by adopting an attention mechanism to dynamically balance cost, time, comfort and sustainability targets: in the first stage, modularly disassembling a travel through sparse constraint linear programming, and quickly generating a Pareto frontier candidate set; in the secondary stage, on the basis of a multi-agent reinforcement learning framework, complex interaction is simulated through a Markov decision process, and strategy iteration is driven through a special reward function for quantifying a comfort index; in order to cope with real-time disturbance, event-driven edge computing nodes are deployed, flight delay and traffic jam emergencies are responded in real time, an incremental topology updating algorithm is triggered, and only affected sub-modules are reconstructed to reduce computing complexity. According to the invention, the bottleneck of dynamic adjustment efficiency and multi-target balance capability is solved.
Owner:YISHANG TRAVEL CO LTD

Quantum circuit mapping method and system based on deep reinforcement learning

The invention relates to the technical field of quantum computing, and discloses a quantum circuit mapping method and system based on deep reinforcement learning, according to the method, an original logic quantum circuit is analyzed and simplified into a ZX-graph, then the ZX-graph is converted into a quantum circuit dependency graph, the computing complexity is remarkably reduced, and the quantum circuit mapping efficiency is improved. A large-scale quantum circuit mapping task can be processed within reasonable time, a Markov decision process environment is constructed through a quantum circuit dependency graph and quantum chip parameters, a deep reinforcement learning agent is trained by using the Markov decision process environment, and an optimal mapping strategy optimization model is obtained. Through the optimal mapping strategy optimization model, the corresponding optimal quantum bit mapping scheme can be output to the current to-be-mapped logic line, so that the method can be adapted to quantum chips of various topological structures, and the generalization ability and adaptability of the mapping scheme are improved.
Owner:SUN YAT SEN UNIV

Vehicle and unmanned aerial vehicle combined dispatching method for wide-range low-cost inspection

The invention relates to a vehicle and unmanned aerial vehicle combined scheduling method for wide-range low-cost inspection. The method comprises the following steps: acquiring prior information; modeling the unmanned aerial vehicle inspection problem of each target area according to the prior information to obtain a mixed integer non-convex optimization problem with the goal of minimizing the weighted sum of the total execution time and the energy consumption of all the inspection unmanned aerial vehicles; performing linearization on a non-convex bilinear term in the mixed integer non-convex optimization problem, and performing discretization processing on a nonlinear function by adopting piecewise linear approximation; an approximate mixed integer linear programming problem is obtained and solved, and an unmanned aerial vehicle scheduling strategy is obtained; modeling according to the unmanned aerial vehicle scheduling strategy and the prior information to obtain an inspection vehicle path planning problem taking the comprehensive driving cost as a target; the routing inspection vehicle path planning problem is converted and modeled into a Markov decision process, a routing inspection vehicle is used as an intelligent agent, a state, an action and a reward function are defined, and a routing inspection vehicle scheduling strategy is obtained. Therefore, combined inspection of the inspection vehicle and the unmanned aerial vehicle is realized, and the inspection range is expanded.
Owner:GUANGDONG UNIV OF TECH

Multi-modal large language model inference engine, site restoration method and storage medium

The invention provides a multi-modal big language model inference engine which comprises a pre-training big language model adaptation layer, a pollution remediation knowledge graph enhancement module, a multi-modal information understanding and fusion unit, a technology implementation optimization module and an inference chain and decision interpretation generator. The system adopts a hierarchical modular design, all components communicate through a standardized API interface, efficient and stable data circulation is ensured, and the problems of knowledge application limitation, opaque reasoning process, difficulty in professional knowledge fusion, knowledge updating lagging, limited multi-modal data processing capability and the like in the existing contaminated site remediation decision process are solved.
Owner:HUBEI PROVINCIAL ACADEMY OF ECO-ENVIRONMENTAL SCIENCES(PROVINCIAL ECOLOGICAL ENVIRONMENT ENGINEERING ASSESSMENT CENTER)

Adaptive load balancing ground user access method for unmanned-aerial-vehicle-assisted network

An adaptive load balancing (ALB) ground user (GU) access method for an unmanned-aerial-vehicle-assisted network. By means of an unmanned aerial vehicle deployment algorithm based on a deep Q-learning network (DQN), and ALB for GU access, a GU access problem in a BS-UAV-NTN is converted into a maximization problem, and the maximization problem is converted into a Markov decision process (MDP) problem for unmanned aerial vehicle deployment in an unknown environment. The method comprises an unmanned aerial vehicle deployment algorithm based on a DQN, and an access scheme for performing priority ranking on BSs and unmanned aerial vehicles. A simulation result shows that the access scheme is superior to conventional Q-learning and random schemes in the reward aspect, and the number of access GUs.
Owner:NANJING UNIV OF POSTS & TELECOMM

Bidding and incentive combined strategy optimization method, system and device for source-load double-side peak shaving auxiliary service market and storage medium

The invention discloses a bidding and incentive combined strategy optimization method, system and device for a source-load double-side peak regulation auxiliary service market, and a storage medium, and belongs to the technical field of electricity market transaction. The method comprises the following steps: constructing a load aggregator and thermal power generating unit-oriented double-market joint optimization model according to parameter data of a peak regulation auxiliary service market and an incentive demand response market; the double-market joint optimization model is formalized into a Markov decision process; constructing a multi-task multi-agent reinforcement learning training framework; and solving the multi-task multi-agent reinforcement learning training framework through an asynchronous training multi-task multi-agent flexible action-evaluation algorithm to obtain a load aggregator bidding and incentive optimal joint strategy and an optimal bidding amount and quotation strategy of the thermal power generating unit. According to the method, game behaviors among multiple market participants and joint decision behaviors of demand-side market subjects in double markets are considered at the same time, and the method is widely applicable to multi-task agent collaborative optimization in the electricity market.
Owner:HOHAI UNIV

Power equipment full life cycle management system based on digital twinning technology

The invention relates to the technical field of electric power asset management, and discloses an electric power equipment full life cycle management system based on a digital twinning technology, and the system comprises a causal knowledge graph construction unit, a strategy rule learning unit, a decision option value quantification unit, a combined decision optimization unit, and a maintenance decision voucher management unit. The method comprises the following steps: constructing a knowledge graph containing a causal relationship among equipment, environment and management activities; carrying out anti-fact deduction based on the atlas to dynamically learn strategy rules; quantizing the value of the potential maintenance decision by adopting a physical option model; performing combination optimization under resource constraints of budget, spare parts and the like to generate an optimal decision combination; and finally, generating a standardized maintenance decision voucher for the decision in the combination. According to the method, decision-making flexibility and environment uncertainty can be quantified into specific economic values, future-oriented and globally optimal resource allocation is realized, and a decision-making process and a decision-making result are solidified into traceable and manageable digital assets.
Owner:ZHONG YI DING SHENG JIAN SHE JI TUAN YOU XIAN GONG SI

DRL-based AUV path planning and obstacle avoidance method in partially observable environment

The invention relates to a DRL-based AUV (Autonomous Underwater Vehicle) path planning and obstacle avoidance method in a partially observable environment, which comprises the following steps of: firstly, executing an AUV path planning and obstacle avoidance task under the condition that marine environment information is partially observable; secondly, constructing an underwater environment model and an AUV kinetic model, and constructing a dynamic and static separated dual-channel adaptive sensing model through local environment information acquired by detection equipment; then, path planning is modeled as a Markov decision process, and a multi-objective optimization reward function is designed for collaborative optimization; finally, an improved maximum entropy SAC algorithm is combined with a priority experience playback mechanism to train a strategy network, a gliding mode is dynamically triggered through a physical constraint layer, attitude angle and thrust output is corrected in real time, and the safety boundary condition is monitored in real time. According to the method, the real-time performance of path planning, the dynamic and static obstacle avoidance success rate and the ocean current utilization rate under the complex sea condition can be remarkably improved, and a high-robustness autonomous navigation solution is provided for ocean exploration.
Owner:HOHAI UNIV

Automatic driving bus collaborative formation dynamic scheduling method

The invention discloses an automatic driving bus collaborative formation dynamic scheduling method, which comprises the following steps of: after introducing a vehicle formation operation, constructing a vehicle energy consumption model based on vehicle specific power, and measuring the vehicle operation energy consumption of an automatic driving bus formation in any formation mode; in an intelligent network connection environment, based on a rolling time domain optimization framework, a Markov decision process model is constructed, and an automatic driving bus collaborative formation dynamic scheduling problem is explained; defining a state variable and a decision variable of the Markov decision process model, and determining constraint conditions of related variables and a target function of the Markov decision process model; the decision space of the Markov decision process model is reduced, and an approximate dynamic programming algorithm is used to solve the automatic driving bus collaborative formation scheduling problem; a multi-step look-ahead strategy based on dynamic planning is provided, the convergence efficiency of an approximate dynamic planning algorithm is improved, and an automatic driving bus collaborative formation optimization scheduling scheme is obtained. And the automatic driving bus capacity utilization rate is improved.
Owner:SOUTH CHINA UNIV OF TECH +1

Multi-objective optimized water resource allocation scheduling method

The invention discloses a multi-objective optimization water resource allocation scheduling method, which combines an LSTM-Prophet hybrid prediction model with an attention mechanism, synchronously captures climate periodicity, policy regulation sensitivity and industrial development tendency characteristics in water demand prediction, and improves water demand prediction precision through a genetic algorithm. Through an improved NSGA-III algorithm and a PSO collaborative optimization mechanism, the convergence problem of a high-dimensional target space is effectively solved in combination with a chaotic mapping technology, and the multi-target optimization efficiency is improved by 40% in cooperation with a dynamic weight adjustment strategy of deep reinforcement learning. And multi-scenario simulation verification of a scheduling scheme is realized by utilizing deep coupling of a digital twin system and a three-dimensional GIS, and a real-time feedback correction mechanism of a Markov decision process is realized. A finally constructed entropy weight-TOPSIS multi-dimensional evaluation system is combined with AR visualization and intelligent contract traceability technologies, and a Pareto optimal solution giving consideration to social fair, economic benefits and ecological integrity is provided for decision makers.
Owner:YELLOW RIVER ENG CONSULTING CO LTD

Power distribution network emergency recovery method, system and device and storage medium

The invention discloses a power distribution network emergency recovery method, system and device and a storage medium, and the method comprises the steps: obtaining the operation environment information of an active power distribution network system, constructing a feature extraction model, carrying out the graph structured representation of the active power distribution network system, and extracting the emergency recovery decision features of the power distribution network; constructing a power distribution network emergency recovery model based on the power distribution network emergency recovery decision features; and converting the emergency recovery problem of the power distribution network into a finite Markov decision process, and calculating and updating operation state parameters of the power distribution network to obtain an optimal emergency recovery strategy of the power distribution network. According to the invention, through collaborative optimization of a space-time diagram attention mechanism and a double-Q network, multi-dimensional feature fusion under topology dynamic change is realized; the capacity of capturing implicit topological association is enhanced, and the pertinence and robustness of feature fusion are improved; through parameter soft update and a dynamic punishment mechanism, optimal strategy search under security constraints is realized, voltage out-of-limit and line overload risks are reduced, and physical feasibility of a recovery strategy is ensured.
Owner:GUIZHOU POWER GRID CO LTD

Meat duck breeding environment temperature regulation and control method and system integrating quantile regression prediction and reinforcement learning

The invention belongs to the technical field of livestock and poultry breeding environment intelligent control, and particularly relates to a meat duck breeding environment temperature regulation and control method and system integrating quantile regression prediction and reinforcement learning. The method comprises the following steps: collecting multi-dimensional data of a breeding environment and operation state information of environment regulation and control equipment, and performing feature construction to obtain a feature vector; using the multi-source data set to predict in-house temperature based on quantile regression to obtain predicted temperatures under different quantiles; modeling a breeding environment temperature regulation problem into a Markov decision process, and dynamically adjusting fan power and a wet curtain equipment state; and setting an online updating model step threshold, and performing actual deployment and application on the reinforcement learning decision model obtained by training and fusing quantile regression information. According to the method, the problems that the meat duck breeding environment temperature regulation and control technology still faces insufficient predictability, regulation and control lag, control strategy static performance, multi-target optimization deficiency and the like are solved.
Owner:SHANDONG ACADEMY OF AGRICULTURAL SCIENCES