Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

88 results about "Sequential decision" patented technology

Sequential decision making. In artificial intelligence, sequential decision making refers to algorithms that take the dynamics of the world into consideration, thus delay parts of the problem until it must be solved. It can be described as a procedural approach to decision-making, or as a step by step decision theory.

Reinforcement learning method and system based on sequential decision-making, device, and medium

Provided are a reinforcement learning method and system based on sequential decision-making, a device, and a medium. The method includes: preprocessing historical trajectory data of reinforcement learning to generate preprocessed historical trajectory data to train a Transformer network model, and transforming the reinforcement learning into a language conversion model task by using a text conversion mechanism in the Transformer network model, to generate a trained Transformer network model, where the Transformer network model is used to predict action information at a next time point in a historical environment, and determine a maximum target reward value in a historical environmental state to obtain a complete trajectory in the historical environmental state; and predicting action information at the next time point in a real environmental state by using the trained Transformer network model, to obtain a complete trajectory in the real environmental state.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Diffusion-reward adversarial imitation learning

Imitation learning, or artificial intelligence-based learning from demonstration, aims to acquire an agent policy by observing and mimicking the behavior demonstrated in expert demonstrations. Imitation learning can be used to generate reliable and robust learned policies in a variety of tasks involving sequential decision-making, such as autonomous driving and robotics tasks. However, current imitation learning solutions are limited in their ability to generalize states or goals unseen from the expert's demonstrations. The present disclosure integrates a diffusion model into generative adversarial imitation learning, which, in terms of prior solutions, can provide superior performance in generalizing to states or goals unseen from the expert's demonstrations, provide data efficiency for varying the amounts of available expert data, and capture more robust and smoother rewards.
Owner:NVIDIA CORP

Matrix pre-sorting method and system based on reinforcement learning and storage medium

The invention relates to the technical field of data processing, and particularly discloses a matrix pre-sequencing method and system based on reinforcement learning and a storage medium, which are used for modeling a matrix pre-sequencing problem into a sequence decision process, and learning an optimal sequencing strategy by utilizing a reinforcement learning agent through interaction with a matrix environment. The state of the agent is defined by the structural features of the matrix, and the action corresponds to the selection of nodes that have not been eliminated. The reward function aims at minimizing the ratio of the number of filling elements to the matrix dimension in the calculation process. The performance of the method is compared with that of a traditional heuristic algorithm in an experiment, and the result shows that in specific test cases such as a large-scale circuit model, the pre-sequencing method based on reinforcement learning has the advantages in the aspects of reducing the number of filling elements and shortening the calculation time.
Owner:CHONGQING UNIV

System and method for offline data-driven discovery and distillation for sequential decision-making with large language models

PendingUS20250284965A1Biological modelsMachine learningData abstractionDecision taking
A method and system for performing a task requiring a plurality of sequential decision-making operations by an autonomous agent are disclosed. The method includes acquiring a portion of offline data from a database, and discovering a set of separate skills. The method further includes segmenting the offline data and segmenting the offline data according to skill, and abstracting the segmented offline data into reusable functions. From the abstracted reusable functions, distilling primitives for each skill and distilling reusable tips in various functions. Modifying skill based text-based and code-based policies for augmenting the autonomous agent, and processing a task requiring a set of sequential decision-making operations using the augmented autonomous agents.
Owner:JPMORGAN CHASE BANK NA

Vehicle Decision Making Using Sequential Information Probing

Vehicle decision-making is analyzed and can be used to modify a decision-making process. For subsets of features comprising a vehicle operational scenario, a first value is generated that quantifies behavior of an artificial intelligence (AI) agent as the AI agent performs a sequence of actions within a first world model based on a complete set of observations for the subset of features. A first world model is a copy of a second world model for sequential decision making. A second value is generated that quantifies behavior of an AI agent as the AI agent performs a sequence of actions in the second world model based on an incomplete set of observations for the subset of features. A difference between the first and second values determines the impact of individual features on the AI agent within the second world model. A decision-making process of the AI agent can be updated.
Owner:NISSAN MOTOR CO LTD +1

Urban rail multi-train energy storage system charging and discharging voltage threshold value adjusting method based on deep reinforcement learning

The invention discloses a dynamic multi-interval charging and discharging voltage threshold value adjusting method based on deep reinforcement learning by adopting a multi-train super-capacitor energy storage system aiming at the problem of traction network voltage safety and stability caused by high power and high energy of an urban rail transit multi-train energy storage system circuit model. According to a super-capacitor double-cascade control part of charge and discharge voltage thresholds, the charge and discharge voltage thresholds of a multi-train super-capacitor energy storage system (SC-ESSs) are dynamically adjusted according to intervals in real time by adopting a deep reinforcement learning (Soft Actor-Critic SAC) algorithm, and online strategy updating and sequential decision execution are further designed, so that the online strategy updating and sequential decision execution are realized. On-line strategy learning and optimization are carried out by combining'trial and error 'and'feedback' mechanisms of the agent, so that the training efficiency and optimization accuracy of the agent are improved, the energy-saving and voltage-stabilizing characteristics of the traction network are effectively improved, the charging and discharging current of the ground super capacitor can be stably controlled, the over-charging and over-discharging phenomena are prevented, and the service life of an energy storage element of the super capacitor is prolonged.
Owner:HUNAN UNIV OF TECH

Sequence decision strategy acquisition method and system for high-neuroplasm crowd

PendingCN121117495ASequence controlTimestamp
The invention provides a sequence decision strategy acquisition method and system for high-neuroplasm crowds. The method comprises the following steps: establishing a full-screen drawing canvas, and loading a task set comprising a decision map type, an emotion titer ratio, a reward parameter, a random seed and a sequence control parameter; collecting user gender selection and binding a corresponding emotion picture library; generating a decision map based on a Markov decision principle under the control of the parameter set, defining nodes and connecting edges, randomly scheduling and laying emotion pictures according to a titer ratio, and assigning rewards to form an emotion and reward coupling environment; based on the environment, recording a mouse position and a mouse path in an execution exploration stage, recording key pressing time, an action direction and an award in an execution stage, and establishing step sequence association according to a timestamp to form sequence decision data; a Latin square balance control task sequence is adopted, and the process is propelled under the trigger of a user; and finally, filing the data in a binary system and a text file for subsequent analysis. According to the invention, the authenticity and repeatability of data are improved.
Owner:AIR FORCE MEDICAL CENT PLA

Shield tunneling machine cutter maintenance decision-making method, equipment, medium and computer program product

The invention relates to the technical field of shield tunneling machine cutter maintenance, and discloses a shield tunneling machine cutter maintenance decision-making method and device, a medium and a computer program product, and the method comprises the steps: initializing a state; defining a sequence decision process in a simulation environment: according to a maintenance action and an updated state calculation moment award, storing an experience tuple into an experience playback buffer area; randomly sampling a batch of empirical tuples from an empirical playback buffer area, minimizing a loss function by using a gradient descent method, and updating policy network parameters of an inspection agent and a maintenance agent; and repeating the above process until a preset termination condition is reached. According to the method, the problem of passive agents in multi-agent reinforcement learning is solved by constructing a sequential simulation environment and adopting a value decomposition network architecture, and the inspection interval can be dynamically adjusted according to different degradation conditions of the shield tunneling machine cutter, so that the maintenance strategy is optimized, and the maintenance cost is reduced.
Owner:江淮前沿技术协同创新中心 +1

Laboratory intelligent material scheduling system and method based on artificial intelligence

The invention discloses a laboratory intelligent material scheduling system and method based on artificial intelligence. The system comprises an AI-driven intelligent prediction and decision-making system, an automatic scheduling and ordering execution system, a solvent recovery optimization system and a data and service support layer. According to the method, inventory prediction, replenishment decision and recovery process parameter optimization are uniformly modeled into a sequential decision problem by adopting an SAC reinforcement learning algorithm, so that the system can perform collaborative optimization and dynamic adjustment based on multi-dimensional state information; through fusion of time sequence analysis, graph database association mining and a multi-objective optimization model, material scheduling meets experimental requirements; by constructing a closed-loop autonomous system of'perception-prediction-decision-execution-optimization 'and a micro-service architecture, the system has continuous online learning and self-evolution capabilities, and finally, the intelligent level, the operation efficiency and the economic and environmental protection benefits of laboratory material management are comprehensively improved.
Owner:NINGBO XINGBOYUAN INTELLIGENT TECHNOLOGY CO LTD

A GNSS / INS tightly coupled adaptive anti-spoofing method based on generalized likelihood ratio and sequential decision

This invention belongs to the field of anti-spoofing and interference technology for satellite navigation and inertial navigation integrated systems. It provides a GNSS / INS tightly coupled adaptive anti-spoofing method based on generalized likelihood ratio and sequential decision-making, applicable to multi-satellite spoofing interference and micro-ramp spoofing scenarios in autonomous driving. The method includes: first, constructing a GNSS / INS tightly coupled model; second, constructing a GNSS spoofing and interference model; third, constructing a generalized likelihood ratio statistic and designing an iterative isolation strategy, improving detection accuracy through sequential hypothesis testing; and fourth, embedding the detection results into a filter and constructing an adaptive factor to achieve anomaly weighting and prediction adjustment. This method integrates spoofing detection and spoofing suppression, significantly improving the anti-spoofing capability of the navigation system.
Owner:LIAONING TECHNICAL UNIVERSITY

A time sequence consistent evidence organization method for aging-oriented question answering

The application provides a time sequence consistent evidence organization method for time limit question answering, comprising: obtaining time limit question answering text; obtaining a candidate evidence set; obtaining priori scores corresponding to each candidate evidence; constructing action scoring references corresponding to the candidate evidence based on the priori scores, and modeling an evidence selection process as a sequential decision process; in each round of candidate evidence selection, constructing a current state according to a current selected evidence set, an unselected candidate set, a time slot coverage state, a semantic abstract state and a remaining budget, calculating action probabilities corresponding to candidate selection actions and a stop action through a strategy network, and selecting a target candidate evidence to join the selected evidence set based on the action probabilities, or selecting the stop action based on the action probabilities; when the strategy network selects the stop action, or the current remaining budget is less than a preset budget threshold, or the current selection round number reaches a preset maximum selection round number, outputting the selected evidence set and generating a corresponding question answering result.
Owner:GUSU LAB OF MATERIALS

Single-chip microcomputer intelligent adjustment control method and system based on linkage of environmental parameters

The application relates to a single-chip intelligent adjustment control method and system based on linkage of environmental parameters. The method realizes real-time calculation of an environmental deviation vector through a lightweight atlas of a preset device mutual exclusion relationship and an environmental influence relationship of a single-chip microcomputer, and obtains a device operation urgency through weighted accumulation after influence direction matching, thereby estimating a required operation time length and expected power to form a unified request. Then, future environmental states are deduced in a prediction time domain by using the influence relationship, hard boundaries are formed by using mutual exclusion constraints and an upper limit of household power, and a sequential decision is made on all requests to generate an execution sequence and implement control, so that real-time resolution of device mutual exclusion conflicts, peak shaving and valley filling of household power loads and prospective linkage maintenance of indoor comfort are realized under extremely low computing resources.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Dynamic task scheduling method for heterogeneous multi-machine systems based on multi-agent reinforcement learning

A dynamic task scheduling method for heterogeneous multi-machine systems based on multi-agent reinforcement learning includes the following steps: 1. Inputting task information and decomposing all tasks into multiple atomic tasks with different capability requirements that can be executed by a single robot; 2. Establishing an adhesion utility evaluation model and heterogeneous priorities to create a mathematical optimization model; 3. Establishing a local perception domain model and a Markov decision model based on the sequential decision-making process of scheduling decision points; 4. Building a heterogeneous multi-machine system scheduling environment and combining multiple advanced training strategies to learn a deep reinforcement learning solver for heterogeneous multi-machine task scheduling through the interaction between the environment and the agents; 5. Using the deep reinforcement learning solver, sequentially outputting actions for each scheduling decision point for the scheduling instance to obtain a dynamic scheduling solution. This method can handle the uncertainty in the task execution process by inserting dynamic scheduling decision points for dynamic events, thereby achieving fast and robust dynamic task scheduling.
Owner:HUNAN UNIV

Multi-agent-based micro-grid energy management method and device containing flexible resources, and storage medium

The invention discloses a multi-agent-based micro-grid energy management method and device containing flexible resources, and a storage medium. The method comprises the following steps: step 1, determining each main body of a micro-grid; 2, constructing a state space, an action space and a reward function of each main body in the three main bodies, thereby constructing a sequential decision model of the micro-grid intelligent body under a GPRO algorithm framework; 3, determining an objective function and constraint conditions; step 4, generating a training sample, and performing multiple rounds of training on the sequential decision model of the micro-grid intelligent agent under a GPRO algorithm framework through the training sample to obtain a micro-grid energy management model; and 5, inputting the current state of each main body in the micro-grid to be managed into the micro-grid energy management model to realize micro-grid energy management. The device and the storage medium are used for implementing the method. According to the method, flexible resources on the load side are fully considered, memory occupation is reduced, and training resources are remarkably reduced.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Optimized control method, device and medium for drone information collection

The present application discloses an optimization control method, device and medium for unmanned aerial vehicle (UAV) information collection, which includes: defining the information collection method of the UAV, establishing a set of UAV flight actions, and constructing an optimization control problem for UAV information collection; obtaining the action information and location information taken by the UAV at a certain moment, and constructing a model for the total amount of information collected by the UAV; constructing a sequential optimization decision model based on the model for the total amount of information collected by the UAV, and modeling the UAV information collection optimization control problem as a problem of finding the optimal strategy for the sequential optimization decision problem; and reconstructing the sequential decision problem as a steady-state reinforcement learning problem and establishing a robust reinforcement learning model; based on the robust reinforcement learning model, using a robust Q-value learning algorithm to solve the problem, and obtaining the optimal control strategy for information collection, which can collect information quickly and effectively, adapt to various scenarios and environments, and have good scalability.
Owner:NAT UNIV OF DEFENSE TECH

Structural grid quality optimization method based on deep reinforcement learning

The invention discloses a structural grid quality optimization method based on deep reinforcement learning, and relates to the field of computer-aided engineering and computational fluid mechanics pretreatment. The problem of position optimization of singular points in a structural grid is modeled as a sequential decision process; grid features are extracted through a state sensing module, an intelligent agent strategy network outputs singular point position movement actions, grids are updated after the environment is executed, reward values based on quality changes are calculated, and the strategy network is updated after interaction experience is stored. A pre-trained expert strategy network is introduced to accelerate training convergence, and suggested actions of the expert strategy network are introduced into a reward function to guide an intelligent agent exploration direction. After training is completed, the fixed strategy network can be rapidly deployed to optimization of a new geometric model. According to the method, automatic global optimization of singular point layout is realized, the grid quality and optimization efficiency are remarkably improved, and the trained model has good generalization ability and can be widely applied to structured grid generation.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Automated model predictive control using a regression-optimization framework for sequential decision making

A computer-implemented method, computer program product, and computer system for automated model predictive control. The computer system trains multiple step look-ahead regression models, using historical states and historical actions for a to-be-optimized system, for each timestep of a past time horizon. Regression models may be either linear or nonlinear in order to capture process dynamics and nonlinearity. The computer system generates optimization constraints for each timestep of a future time horizon. The computer system generates optimization variables, based on the multiple step look-ahead regression models, for each timestep of the future time horizon. The computer system constructs a mixed integer linear programming based optimization model that includes an objective function, the optimization constraints, and the optimization variables. Nonlinear regression models are converted into piecewise linear approximation functions. The computer system solves the optimization model to produce actions for the to-be-optimized system, over the future time horizon, and recommend commitment-look-ahead actions.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Calculation-efficient privacy protection method and system based on data shuffling in sequential decision

The invention discloses a privacy protection method and system based on data shuffling in a sequential decision, which are efficient in calculation and are used for solving the dual technical problems of user data privacy leakage and overhigh algorithm calculation complexity when reinforcement learning is adopted to carry out the sequential decision. According to the method, on the basis of introducing a credible shuffling device to guarantee data privacy, a brand-new staged strategy iteration framework which does not need to traverse a strategy space is provided. According to the framework, through an embedded adaptive dichotomy strategy search module and an experimental design method, the exponential traversal search on an active strategy set and the dependence on an external calculation oracle machine are thoroughly avoided. According to the method, the optimal regret boundary equivalent to or even better than centralized differential privacy can be realized in a scene with extremely high privacy protection requirements, meanwhile, the performance of a traditional local differential privacy model is remarkably superior, and the actual deployment feasibility and the calculation efficiency of a privacy protection scheme are greatly improved.
Owner:ZHEJIANG UNIV

Fault risk sequential decision optimization method for high-availability system

The invention belongs to the technical field of computers, provides a fault risk sequential decision optimization method for a high-availability system, and solves the problems that the missed report / false report cost of a static model is asymmetric and a dynamic decision is missing. Extracting features from logs and performance data in a multi-scale manner, and encoding the features into a state vector z; the modeling comprises MDP, state space, action, '1' table early warning and '0' table no risk, and designing an asymmetric reward function, underreporting penalty and false reporting penalty; based on DQN training, layered playback and balanced sampling are adopted to relieve sample sparsity so as to minimize the TD error of the signal, and iteration is optimal; and inputting, selecting and outputting a judgment result in real time. The method realizes dynamic risk quantitative decision and improves system toughness.
Owner:INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER

An intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network

The present invention provides an intelligent trajectory planning method for a near-space communication network relay unmanned aerial vehicle, which includes: Step 1, designing a method for generating a communication network topology graph: constructing a directed graph called a fully connected graph, and adding directed edges E from the base station node to each relay unmanned aerial vehicle, between each relay unmanned aerial vehicle, and from each relay unmanned aerial vehicle to each ground user, where the edges between each relay unmanned aerial vehicle have two directions; Step 2, converting the trajectory planning process of the stratospheric airship into a sequential decision-making process and constructing a multi-agent Markov decision-making process model; Step 3, designing a policy network and a value network and constructing a virtual simulation environment for the trajectory planning of the relay unmanned aerial vehicle; Step 4, training the agent based on the soft actor-critic algorithm to find the optimal policy. The method of the present invention can plan the position of the relay unmanned aerial vehicle in real time, adjust the communication link, and maintain network connectivity.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Vehicle decision making using sequential information probing

Vehicle decision-making is analyzed and can be used to modify a decision-making process. For subsets of features comprising a vehicle operational scenario, a first value is generated that quantifies behavior of an artificial intelligence (AI) agent as the AI agent performs a sequence of actions within a first world model based on a complete set of observations for the subset of features. A first world model is a copy of a second world model for sequential decision making. A second value is generated that quantifies behavior of an AI agent as the AI agent performs a sequence of actions in the second world model based on an incomplete set of observations for the subset of features. A difference between the first and second values determines the impact of individual features on the AI agent within the second world model. A decision-making process of the AI agent can be updated.
Owner:NISSAN MOTOR CO LTD +1

Monte Carlo tree search-based unmanned cluster reconnaissance rescue task allocation method

The invention relates to an unmanned cluster reconnaissance rescue task allocation method based on Monte Carlo tree search, and belongs to the technical field of task planning. According to the method, the task allocation problem in reconnaissance and rescue is expressed as a sequential decision process, heuristic information of the problem is fully utilized, a state space and an action space are designed, the action completeness and the action number are balanced, and calculation resources are saved while the allocation effect is guaranteed. Besides, a self-adaptive simulation number selection mechanism is provided, simulation times can be automatically selected according to the tasks to be distributed and the number of the unmanned aerial vehicles, and a large amount of repeated simulation is avoided, so that the calculation amount of an algorithm is reduced, and the calculation efficiency of task distribution is further improved. According to the method, the task allocation effectiveness of the unmanned system can be enhanced, and the task efficiency of the unmanned system is improved.
Owner:AERONAUTICS RES INST OF CHINA

A reinforcement learning partition pre-interaction method and system

The application relates to a reinforcement learning partition pre-interaction method and system, wherein the method comprises the following steps: constructing a bias measurement function for evaluating model bias of a virtual environment relative to a real environment; dividing the real environment into multiple trust intervals according to the model bias; performing parameter coarse adjustment on an intelligent agent of the reinforcement learning based on the trust intervals; in the parameter coarse adjustment stage, sampling experience data generated by interaction of the intelligent agent with the virtual environment and the real environment respectively to update network parameters of the intelligent agent; monitoring a performance index of the intelligent agent in the parameter coarse adjustment stage, and when the performance index meets a preset interaction mode switching condition, the intelligent agent enters a parameter fine adjustment stage; in the parameter fine adjustment stage, the intelligent agent completely interacts with the real environment, and parameters of the intelligent agent are adjusted to convergence, and the parameter-converged intelligent agent is used for solving a complex sequential decision problem. The application effectively adjusts parameters of the intelligent agent of the reinforcement learning, and makes the problem-solving capability of the intelligent agent more reliable.
Owner:JIANGNAN UNIV

Intelligent optimization method for standby power consumption of electric vehicle charging system

PendingCN122092643ARealize refined managementFlexible to adaptCharging stationsBatteries circuit arrangementsSleep stateControl engineering
The invention belongs to the technical field of electric vehicle charging, and particularly relates to an intelligent optimization method for standby power consumption of an electric vehicle charging system. In order to overcome the defect that the existing electric vehicle charging system cannot give consideration to reduction of standby power consumption and guarantee of user experience, the invention adopts the following technical scheme: the intelligent optimization method for the standby power consumption of the electric vehicle charging system comprises the following steps: establishing a multi-stage dormant state mechanism; through a data driving method, the optimal state granularity in the multi-stage dormant state is automatically determined, and a state conversion threshold value is dynamically adjusted; acquiring an environment state sensing feature vector; a standby power consumption optimization problem is converted into a sequence decision problem, the sequence decision problem is solved according to a multi-stage dormant state mechanism and an environment state sensing feature vector, and the sequence decision problem is to determine which state conversion action should be executed at what time. The method has the beneficial effect that the dynamic balance between the standby power consumption reduction and the instant response of the charging system is ensured.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT +2

Method for calibrating a target metric of an ECU function

The invention relates to a method (100) for calibrating a target metric of an ECU function of a system with a plurality of operating parameters, in particular an ECU function for controlling an air-fuel ratio of an internal combustion engine, comprising at least the following steps: Using (102) at least one calibration parameter in the at least one ECU function, wherein the at least one calibration parameter was determined by means of at least one trained learning algorithm from at least one input signal provided by the ECU function; wherein the algorithm is trained beforehand by means of at least the following step: Solving (104) at least one sequential decision task describing the problem of calibrating the ECU function with an reinforcement learning agent for training the learning algorithm.Method (100) can provide robust calibration, thereby reducing time and costs.
Owner:DR ING H C F PORSCHE AG +1

System and method for implementing a sequential decision-making agent considering uncertain states

A system and method for implementing a sequential decision-making agent considering uncertain states, comprising: a prior-posterior information combination processing module, an input-distribution decision-making agent module, wherein the prior-posterior information combination processing module utilizes the prior estimated information and the posterior real feedback information to obtain the distribution of the conversion amount and the unit conversion cost parameter; the input-distribution decision-making agent module samples the distribution information of the unit conversion cost parameter to obtain the corresponding discrete distribution, and inputs the distribution into the parallel action state neural network to obtain the optimal decision under the reference uncertain state. The present application utilizes the feature distribution and the reinforcement learning method when making sequential decisions, and significantly improves the optimization effect of the agent when making sequential decisions by constructing the agent with low complexity cost.
Owner:SHANGHAI JIAOTONG UNIV

Generator set control method and apparatus, and device

A generator set control method and apparatus, and a device. Comprising: constructing a state transition model for sub-problems of a single unit, and adding as a state in the model a penalty price corresponding to a Lagrange multiplier for each time period Using a reinforcement learning algorithm to train a startup / shutdown strategy and a power increase / decrease strategy for each unit; using a surrogate sub-gradient method to relax constraints coupled to different units in a UC problem, using the surrogate sub-gradient method to perform iteration and Lagrange multiplier updating, solving sub-problems in the iteration process using a trained reinforcement learning agent to perform sequential decision-making, and iterating repeatedly until convergence, so as to obtain an optimal solution to a dual problem; and performing a feasibility operation on a resulting unit commitment state, and controlling generator set nodes.
Owner:TSINGHUA UNIVERSITY

Vehicle decision making using sequential information probing

Vehicle decision-making is analyzed and can be used to modify a decision-making process. For subsets of features comprising a vehicle operational scenario, a first value is generated that quantifies behavior of an artificial intelligence (AI) agent as the AI agent performs a sequence of actions within a first world model based on a complete set of observations for the subset of features. A first world model is a copy of a second world model for sequential decision making. A second value is generated that quantifies behavior of an AI agent as the AI agent performs a sequence of actions in the second world model based on an incomplete set of observations for the subset of features. A difference between the first and second values determines the impact of individual features on the AI agent within the second world model. A decision-making process of the AI agent can be updated.
Owner:NISSAN NORTH AMERICA INC +1

Software tool and method for analysis of cybersecurity vulnerabilities

A system and method for analysis of potential cybersecurity threats in an engineered system combines a functional simulation of the system with a cyberattack information layer. The cyberattack information layer informs the simulation with respect to the effects of a potential cyberattack on devices and links in the system. An iterative AI-based sequential decision-making optimization identifies a sequence of attacker steps for carrying out a cyberattack which has the greatest impact on a key performance indicator relating to operation of the system. The cyberattack information layer includes information relating to a topology and devices in the system from a computer network perspective and includes information on an amount of effort required to carry out possible attacks affecting each device or communication link in the system. The decision-making optimization may be run iteratively between the simulation and the cyberattack information layer to find a most impactful sequence of attacker actions.
Owner:SIEMENS CORP +1

A regional power grid source and load collaborative scheduling learning optimization method based on principal-agent game

The present application relates to the technical field of power system, especially to a regional power grid source-load collaborative scheduling learning optimization method based on master-slave game, comprising the following steps: S1, analyzing the physical architecture and logical relationship of the regional power grid source-load collaborative scheduling system, and proposing optimization strategy merit evaluation index; S2, establishing the master-slave game model between the market price setting institution and the load aggregator, the unit generation plan model and the constraint model under the condition of uncertain load response; S3, describing the scheduling decision problem as a learning optimization mechanism for random sequential decision of electricity price, and solving by using a typical reinforcement learning method.
Owner:HEFEI UNIV OF TECH +1