Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

60 results about "Sequential decision" patented technology

Sequential decision making. In artificial intelligence, sequential decision making refers to algorithms that take the dynamics of the world into consideration, thus delay parts of the problem until it must be solved. It can be described as a procedural approach to decision-making, or as a step by step decision theory.

Matrix pre-sorting method and system based on reinforcement learning and storage medium

The invention relates to the technical field of data processing, and particularly discloses a matrix pre-sequencing method and system based on reinforcement learning and a storage medium, which are used for modeling a matrix pre-sequencing problem into a sequence decision process, and learning an optimal sequencing strategy by utilizing a reinforcement learning agent through interaction with a matrix environment. The state of the agent is defined by the structural features of the matrix, and the action corresponds to the selection of nodes that have not been eliminated. The reward function aims at minimizing the ratio of the number of filling elements to the matrix dimension in the calculation process. The performance of the method is compared with that of a traditional heuristic algorithm in an experiment, and the result shows that in specific test cases such as a large-scale circuit model, the pre-sequencing method based on reinforcement learning has the advantages in the aspects of reducing the number of filling elements and shortening the calculation time.
Owner:CHONGQING UNIV

Sequence decision strategy acquisition method and system for high-neuroplasm crowd

PendingCN121117495ASequence controlTimestamp
The invention provides a sequence decision strategy acquisition method and system for high-neuroplasm crowds. The method comprises the following steps: establishing a full-screen drawing canvas, and loading a task set comprising a decision map type, an emotion titer ratio, a reward parameter, a random seed and a sequence control parameter; collecting user gender selection and binding a corresponding emotion picture library; generating a decision map based on a Markov decision principle under the control of the parameter set, defining nodes and connecting edges, randomly scheduling and laying emotion pictures according to a titer ratio, and assigning rewards to form an emotion and reward coupling environment; based on the environment, recording a mouse position and a mouse path in an execution exploration stage, recording key pressing time, an action direction and an award in an execution stage, and establishing step sequence association according to a timestamp to form sequence decision data; a Latin square balance control task sequence is adopted, and the process is propelled under the trigger of a user; and finally, filing the data in a binary system and a text file for subsequent analysis. According to the invention, the authenticity and repeatability of data are improved.
Owner:AIR FORCE MEDICAL CENT PLA

Laboratory intelligent material scheduling system and method based on artificial intelligence

The invention discloses a laboratory intelligent material scheduling system and method based on artificial intelligence. The system comprises an AI-driven intelligent prediction and decision-making system, an automatic scheduling and ordering execution system, a solvent recovery optimization system and a data and service support layer. According to the method, inventory prediction, replenishment decision and recovery process parameter optimization are uniformly modeled into a sequential decision problem by adopting an SAC reinforcement learning algorithm, so that the system can perform collaborative optimization and dynamic adjustment based on multi-dimensional state information; through fusion of time sequence analysis, graph database association mining and a multi-objective optimization model, material scheduling meets experimental requirements; by constructing a closed-loop autonomous system of'perception-prediction-decision-execution-optimization 'and a micro-service architecture, the system has continuous online learning and self-evolution capabilities, and finally, the intelligent level, the operation efficiency and the economic and environmental protection benefits of laboratory material management are comprehensively improved.
Owner:NINGBO XINGBOYUAN INTELLIGENT TECHNOLOGY CO LTD

A GNSS / INS tightly coupled adaptive anti-spoofing method based on generalized likelihood ratio and sequential decision

This invention belongs to the field of anti-spoofing and interference technology for satellite navigation and inertial navigation integrated systems. It provides a GNSS / INS tightly coupled adaptive anti-spoofing method based on generalized likelihood ratio and sequential decision-making, applicable to multi-satellite spoofing interference and micro-ramp spoofing scenarios in autonomous driving. The method includes: first, constructing a GNSS / INS tightly coupled model; second, constructing a GNSS spoofing and interference model; third, constructing a generalized likelihood ratio statistic and designing an iterative isolation strategy, improving detection accuracy through sequential hypothesis testing; and fourth, embedding the detection results into a filter and constructing an adaptive factor to achieve anomaly weighting and prediction adjustment. This method integrates spoofing detection and spoofing suppression, significantly improving the anti-spoofing capability of the navigation system.
Owner:LIAONING TECHNICAL UNIVERSITY

A time sequence consistent evidence organization method for aging-oriented question answering

The application provides a time sequence consistent evidence organization method for time limit question answering, comprising: obtaining time limit question answering text; obtaining a candidate evidence set; obtaining priori scores corresponding to each candidate evidence; constructing action scoring references corresponding to the candidate evidence based on the priori scores, and modeling an evidence selection process as a sequential decision process; in each round of candidate evidence selection, constructing a current state according to a current selected evidence set, an unselected candidate set, a time slot coverage state, a semantic abstract state and a remaining budget, calculating action probabilities corresponding to candidate selection actions and a stop action through a strategy network, and selecting a target candidate evidence to join the selected evidence set based on the action probabilities, or selecting the stop action based on the action probabilities; when the strategy network selects the stop action, or the current remaining budget is less than a preset budget threshold, or the current selection round number reaches a preset maximum selection round number, outputting the selected evidence set and generating a corresponding question answering result.
Owner:GUSU LAB OF MATERIALS

Single-chip microcomputer intelligent adjustment control method and system based on linkage of environmental parameters

The application relates to a single-chip intelligent adjustment control method and system based on linkage of environmental parameters. The method realizes real-time calculation of an environmental deviation vector through a lightweight atlas of a preset device mutual exclusion relationship and an environmental influence relationship of a single-chip microcomputer, and obtains a device operation urgency through weighted accumulation after influence direction matching, thereby estimating a required operation time length and expected power to form a unified request. Then, future environmental states are deduced in a prediction time domain by using the influence relationship, hard boundaries are formed by using mutual exclusion constraints and an upper limit of household power, and a sequential decision is made on all requests to generate an execution sequence and implement control, so that real-time resolution of device mutual exclusion conflicts, peak shaving and valley filling of household power loads and prospective linkage maintenance of indoor comfort are realized under extremely low computing resources.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Multi-agent-based micro-grid energy management method and device containing flexible resources, and storage medium

The invention discloses a multi-agent-based micro-grid energy management method and device containing flexible resources, and a storage medium. The method comprises the following steps: step 1, determining each main body of a micro-grid; 2, constructing a state space, an action space and a reward function of each main body in the three main bodies, thereby constructing a sequential decision model of the micro-grid intelligent body under a GPRO algorithm framework; 3, determining an objective function and constraint conditions; step 4, generating a training sample, and performing multiple rounds of training on the sequential decision model of the micro-grid intelligent agent under a GPRO algorithm framework through the training sample to obtain a micro-grid energy management model; and 5, inputting the current state of each main body in the micro-grid to be managed into the micro-grid energy management model to realize micro-grid energy management. The device and the storage medium are used for implementing the method. According to the method, flexible resources on the load side are fully considered, memory occupation is reduced, and training resources are remarkably reduced.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Structural grid quality optimization method based on deep reinforcement learning

The invention discloses a structural grid quality optimization method based on deep reinforcement learning, and relates to the field of computer-aided engineering and computational fluid mechanics pretreatment. The problem of position optimization of singular points in a structural grid is modeled as a sequential decision process; grid features are extracted through a state sensing module, an intelligent agent strategy network outputs singular point position movement actions, grids are updated after the environment is executed, reward values based on quality changes are calculated, and the strategy network is updated after interaction experience is stored. A pre-trained expert strategy network is introduced to accelerate training convergence, and suggested actions of the expert strategy network are introduced into a reward function to guide an intelligent agent exploration direction. After training is completed, the fixed strategy network can be rapidly deployed to optimization of a new geometric model. According to the method, automatic global optimization of singular point layout is realized, the grid quality and optimization efficiency are remarkably improved, and the trained model has good generalization ability and can be widely applied to structured grid generation.
Owner:CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT

Calculation-efficient privacy protection method and system based on data shuffling in sequential decision

The invention discloses a privacy protection method and system based on data shuffling in a sequential decision, which are efficient in calculation and are used for solving the dual technical problems of user data privacy leakage and overhigh algorithm calculation complexity when reinforcement learning is adopted to carry out the sequential decision. According to the method, on the basis of introducing a credible shuffling device to guarantee data privacy, a brand-new staged strategy iteration framework which does not need to traverse a strategy space is provided. According to the framework, through an embedded adaptive dichotomy strategy search module and an experimental design method, the exponential traversal search on an active strategy set and the dependence on an external calculation oracle machine are thoroughly avoided. According to the method, the optimal regret boundary equivalent to or even better than centralized differential privacy can be realized in a scene with extremely high privacy protection requirements, meanwhile, the performance of a traditional local differential privacy model is remarkably superior, and the actual deployment feasibility and the calculation efficiency of a privacy protection scheme are greatly improved.
Owner:ZHEJIANG UNIV

Fault risk sequential decision optimization method for high-availability system

The invention belongs to the technical field of computers, provides a fault risk sequential decision optimization method for a high-availability system, and solves the problems that the missed report / false report cost of a static model is asymmetric and a dynamic decision is missing. Extracting features from logs and performance data in a multi-scale manner, and encoding the features into a state vector z; the modeling comprises MDP, state space, action, '1' table early warning and '0' table no risk, and designing an asymmetric reward function, underreporting penalty and false reporting penalty; based on DQN training, layered playback and balanced sampling are adopted to relieve sample sparsity so as to minimize the TD error of the signal, and iteration is optimal; and inputting, selecting and outputting a judgment result in real time. The method realizes dynamic risk quantitative decision and improves system toughness.
Owner:INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER

Vehicle decision making using sequential information probing

Vehicle decision-making is analyzed and can be used to modify a decision-making process. For subsets of features comprising a vehicle operational scenario, a first value is generated that quantifies behavior of an artificial intelligence (AI) agent as the AI agent performs a sequence of actions within a first world model based on a complete set of observations for the subset of features. A first world model is a copy of a second world model for sequential decision making. A second value is generated that quantifies behavior of an AI agent as the AI agent performs a sequence of actions in the second world model based on an incomplete set of observations for the subset of features. A difference between the first and second values determines the impact of individual features on the AI agent within the second world model. A decision-making process of the AI agent can be updated.
Owner:NISSAN MOTOR CO LTD +1

A reinforcement learning partition pre-interaction method and system

The application relates to a reinforcement learning partition pre-interaction method and system, wherein the method comprises the following steps: constructing a bias measurement function for evaluating model bias of a virtual environment relative to a real environment; dividing the real environment into multiple trust intervals according to the model bias; performing parameter coarse adjustment on an intelligent agent of the reinforcement learning based on the trust intervals; in the parameter coarse adjustment stage, sampling experience data generated by interaction of the intelligent agent with the virtual environment and the real environment respectively to update network parameters of the intelligent agent; monitoring a performance index of the intelligent agent in the parameter coarse adjustment stage, and when the performance index meets a preset interaction mode switching condition, the intelligent agent enters a parameter fine adjustment stage; in the parameter fine adjustment stage, the intelligent agent completely interacts with the real environment, and parameters of the intelligent agent are adjusted to convergence, and the parameter-converged intelligent agent is used for solving a complex sequential decision problem. The application effectively adjusts parameters of the intelligent agent of the reinforcement learning, and makes the problem-solving capability of the intelligent agent more reliable.
Owner:JIANGNAN UNIV

Intelligent optimization method for standby power consumption of electric vehicle charging system

PendingCN122092643ARealize refined managementFlexible to adaptCharging stationsBatteries circuit arrangementsSleep stateControl engineering
The invention belongs to the technical field of electric vehicle charging, and particularly relates to an intelligent optimization method for standby power consumption of an electric vehicle charging system. In order to overcome the defect that the existing electric vehicle charging system cannot give consideration to reduction of standby power consumption and guarantee of user experience, the invention adopts the following technical scheme: the intelligent optimization method for the standby power consumption of the electric vehicle charging system comprises the following steps: establishing a multi-stage dormant state mechanism; through a data driving method, the optimal state granularity in the multi-stage dormant state is automatically determined, and a state conversion threshold value is dynamically adjusted; acquiring an environment state sensing feature vector; a standby power consumption optimization problem is converted into a sequence decision problem, the sequence decision problem is solved according to a multi-stage dormant state mechanism and an environment state sensing feature vector, and the sequence decision problem is to determine which state conversion action should be executed at what time. The method has the beneficial effect that the dynamic balance between the standby power consumption reduction and the instant response of the charging system is ensured.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT +2

Method for calibrating a target metric of an ECU function

The invention relates to a method (100) for calibrating a target metric of an ECU function of a system with a plurality of operating parameters, in particular an ECU function for controlling an air-fuel ratio of an internal combustion engine, comprising at least the following steps: Using (102) at least one calibration parameter in the at least one ECU function, wherein the at least one calibration parameter was determined by means of at least one trained learning algorithm from at least one input signal provided by the ECU function; wherein the algorithm is trained beforehand by means of at least the following step: Solving (104) at least one sequential decision task describing the problem of calibrating the ECU function with an reinforcement learning agent for training the learning algorithm.Method (100) can provide robust calibration, thereby reducing time and costs.
Owner:DR ING H C F PORSCHE AG +1

System and method for implementing a sequential decision-making agent considering uncertain states

A system and method for implementing a sequential decision-making agent considering uncertain states, comprising: a prior-posterior information combination processing module, an input-distribution decision-making agent module, wherein the prior-posterior information combination processing module utilizes the prior estimated information and the posterior real feedback information to obtain the distribution of the conversion amount and the unit conversion cost parameter; the input-distribution decision-making agent module samples the distribution information of the unit conversion cost parameter to obtain the corresponding discrete distribution, and inputs the distribution into the parallel action state neural network to obtain the optimal decision under the reference uncertain state. The present application utilizes the feature distribution and the reinforcement learning method when making sequential decisions, and significantly improves the optimization effect of the agent when making sequential decisions by constructing the agent with low complexity cost.
Owner:SHANGHAI JIAOTONG UNIV

Generator set control method and apparatus, and device

A generator set control method and apparatus, and a device. Comprising: constructing a state transition model for sub-problems of a single unit, and adding as a state in the model a penalty price corresponding to a Lagrange multiplier for each time period Using a reinforcement learning algorithm to train a startup / shutdown strategy and a power increase / decrease strategy for each unit; using a surrogate sub-gradient method to relax constraints coupled to different units in a UC problem, using the surrogate sub-gradient method to perform iteration and Lagrange multiplier updating, solving sub-problems in the iteration process using a trained reinforcement learning agent to perform sequential decision-making, and iterating repeatedly until convergence, so as to obtain an optimal solution to a dual problem; and performing a feasibility operation on a resulting unit commitment state, and controlling generator set nodes.
Owner:TSINGHUA UNIVERSITY

A regional power grid source and load collaborative scheduling learning optimization method based on principal-agent game

The present application relates to the technical field of power system, especially to a regional power grid source-load collaborative scheduling learning optimization method based on master-slave game, comprising the following steps: S1, analyzing the physical architecture and logical relationship of the regional power grid source-load collaborative scheduling system, and proposing optimization strategy merit evaluation index; S2, establishing the master-slave game model between the market price setting institution and the load aggregator, the unit generation plan model and the constraint model under the condition of uncertain load response; S3, describing the scheduling decision problem as a learning optimization mechanism for random sequential decision of electricity price, and solving by using a typical reinforcement learning method.
Owner:HEFEI UNIV OF TECH +1

Closed-loop supervised fine-tuning of tokenized traffic models

Imitation learning, or artificial intelligence-based learning from demonstration, aims to acquire an agent policy by observing and mimicking the behavior demonstrated in expert demonstrations. Imitation learning can be used to generate reliable and robust learned policies in a variety of tasks involving sequential decision-making, such as autonomous driving and robotics tasks. However, existing methods that use next-token-prediction (NTP) models, where the policy reduces to a classifier over a discrete set of trajectory tokens, suffer from covariate shift due to their open-loop training a closed-loop execution. The present disclosure provides closed-loop fine tuning of autonomous agent policies in a manner that can mitigate covariate shift.
Owner:NVIDIA CORP

Intelligent decision-making methods, terminal equipment and storage media for power system unit dispatching

ActiveCN117726478BOvercome the impact of schedulingImprove training efficiencyConcurrent computationTraining phase
This invention discloses an intelligent decision-making method, terminal equipment, and storage medium for power system unit scheduling. Based on historical power system operating data, it extracts typical features using dimensionality reduction methods and constructs a feature index set by configuring feature weights using an objective weighting method. A Gaussian mixture clustering model is used for multi-scenario partitioning as a front-end optimization measure for deep reinforcement learning methods, mitigating the suboptimal decision-making problem that may be caused by differences in data distribution across multiple scenarios under source-load uncertainty. The unit scheduling problem is modeled as a sequential decision Markov process, constructing a multi-scenario unit scheduling model based on deep reinforcement learning, overcoming the limitations of the original single-scenario model. Through a dynamic step-size update mechanism and parallel computing, the parameter update efficiency during the offline training phase of the decision network is improved.
Owner:HUNAN UNIV

Parameterized three-dimensional modeling method for shield tunnel based on reinforcement learning double-layer optimization

The application discloses a shield tunnel parameterized three-dimensional modeling method based on reinforcement learning double-layer optimization, which comprises the following steps: inputting planning line parameters through a global policy network, and outputting segment design parameters; constructing a layout design space according to the segment design parameters and the planning line; outputting the position of the next ring segment based on the layout space state by a sequential decision network; determining a reward and punishment value by calculating the deviation value of the segment axis and the design line, and training the sequential decision layer network; completing line layout and calculating the cumulative reward and punishment value through an iterative cycle; training the global parameter layer network according to the cumulative reward and punishment value and updating the global parameters, repeating the above process until the convergence condition is reached, and finally obtaining the optimized segment parameters and layout, and constructing a three-dimensional model. The application realizes double-layer optimization of segment design and layout, improves the design efficiency, and reduces the line deviation.
Owner:TONGJI UNIV

Intraday dispatch optimization method for power grid considering photovoltaic output uncertainty under adverse weather

The application belongs to the technical field of power systems, and more particularly to a power grid intra-day scheduling optimization method considering photovoltaic output uncertainty under severe weather, which uses an A2C algorithm framework and knowledge transfer technology to solve and optimize the power grid scheduling problem under photovoltaic output uncertainty under severe weather. First, the sequential decision-making process of power grid intra-day scheduling is determined, and a mathematical model is established, and then the optimization objectives and constraint conditions of power grid intra-day scheduling are determined; secondly, the CGAN scene generation method is used to generate training samples, and the A2C algorithm framework is used to optimize the scheduling intelligent agent that can adapt to various source and load scenes. Finally, transfer learning is used to improve the optimization of the specific intra-day scheduling intelligent agent, solving the problem that the accuracy of the prediction data of large-scale photovoltaic access to the power system greatly affects the stable operation of the power grid, and ensuring that the power grid operates in a relatively economical manner under the premise of safety.
Owner:HEFEI UNIV OF TECH

Response-driven load shedding control method for receiving-end power grid considering transient voltage recovery requirement

The application provides a kind of load shedding control method of receiving end power grid considering transient voltage recovery demand, belongs to response driven type load shedding control, solve the technical problem of insufficient system voltage regulation capability caused by large-scale access of new energy to receiving end power grid and large number of replacement of traditional fossil energy synchronous unit.Its technical scheme is as follows: containing the following steps: (1) establish the Markov sequential decision process model of transient voltage control;(2) design the state, action and reward in Markov process;(3) establish the reinforcement learning model based on Double DQN;(4) carry out off-line centralized learning and training of corresponding agent;(5) form decentralized adaptive load shedding control method.The beneficial results of the application are as follows: the load shedding controller can be deployed at each load site, and the closed-loop decentralized load shedding control strategy of "monitoring-determination-decision-control" can be realized to promote the overall voltage recovery of power grid.
Owner:NORTHEAST DIANLI UNIVERSITY +1

SUPERVISED FINE-TUNING OF TOKENIZED TRAFFIC MODELS WITH CLOSED REGULATIONS

UndeterminedDE102025142087A1Mathematical modelsArtificial lifeAlgorithmTraffic model
Imitation learning, or AI-based demonstration-based learning, aims to derive a policy for an agent by observing and imitating the behavior shown in expert demonstrations. Imitation learning can be used to generate reliable and robust learned policies for a variety of tasks requiring sequential decision-making, such as autonomous driving and robotics. However, existing methods that use next-token prediction (NTP) models, where the policy is reduced to a classifier over a discrete set of trajectory tokens, suffer from covariate shifting due to their open-loop training and closed-loop execution.The present disclosure provides a fine-tuning of the policies of autonomous closed-loop agents in a manner that can mitigate covariate shifting.
Owner:NVIDIA CORP

Distribution line intelligent planning method considering operation flexibility and based on deep reinforcement learning

The invention discloses a deep reinforcement learning-based distribution line intelligent planning method considering operation flexibility, and relates to the technical field of distribution network planning, and the method comprises the following steps: initializing a reinforcement learning environment containing operation state information, and constructing a deep reinforcement learning agent; the deep reinforcement learning agent is trained until a preset condition is met, the training process comprises the steps of obtaining operation state information of a next state based on the deep reinforcement learning agent and calculating a cost optimization reward by using a comprehensive cost mathematical model, and the comprehensive cost mathematical model comprises a flexible shortage cost; updating a strategy network parameter of the deep reinforcement learning agent based on the cost optimization reward; and inputting the initial state of the to-be-planned power grid into the trained deep reinforcement learning agent, and executing a sequential decision to generate a line planning scheme. The method is used for solving the problems that in the prior art, the operation flexibility is difficult to scientifically quantify, and the solving efficiency in multi-stage decision making is low.
Owner:ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER +1

Personalized learning path planning method and system based on online education

The invention belongs to the technical field of artificial intelligence, and particularly relates to a personalized learning path planning method and system based on online education, and the method comprises the steps: firstly, precisely perceiving the state of a learner through constructing a multi-dimensional meta-cognitive state model; secondly, based on the state and the long-term target, the reinforcement learning agent carries out sequence decision making by taking maximization of the long-term income as a target, and a macroscopic learning path is generated; then, according to the path instruction, fairly and accurately matching specific learning resources from a resource library by combining content features with a UCB algorithm; and finally, updating the state model in real time according to learning feedback, optimizing a planning strategy and calibrating resource evaluation. The method achieves the global and adaptive planning of a learning path, remarkably improves the learning continuity and long-term efficiency, effectively solves a resource cold start problem, and drives the continuous optimization of teaching resource ecology.
Owner:ZHIWANG TECHNOLOGY (GUANGZHOU) CO LTD

A multi-agent path planning method based on a distributed cooperative deep reinforcement learning model

The application discloses a multi-agent path planning method based on a distributed cooperative deep reinforcement learning model, belongs to the communication field and the technical field of reinforcement learning, and comprises the following steps: training the distributed cooperative deep reinforcement learning model designed in the application to obtain a trained distributed cooperative deep reinforcement learning model and randomly generating a map; and based on the trained distributed cooperative deep reinforcement learning model, realizing online path planning of an agent in a multi-agent environment. The distributed cooperative deep reinforcement learning method is different from a centralized deep reinforcement learning method, each agent only makes a decision according to its own observation, meanwhile, communication and sequential decision making among the agents are introduced, and coordination among the agents is promoted. The method is designed from the perspective of a single agent, without considering a group, so that the non-scalability problem of the centralized deep reinforcement learning is avoided.
Owner:NORTHEASTERN UNIV CHINA

Systems and methods for detecting and regularizing drift in sequential decisions

Embodiments of the present disclosure are directed to systems and methods to display a plurality of items in a spatial layout, monitor an eye gaze of the user while making sequential decisions about the plurality of items, determine that a drift has occurred, automatically regularize a plurality of alternative choice options, learn and identify an upper boundary and a lower boundary based on whether the user chooses or rejects at least one alternative choice, model the drift using machine learning based at least in part on a history of user decision responses, predict a drift direction of the drift and a relevance impact of different attributes into a drift state, classify the drift states into a risk-averse classification and a risk-seeking classification by clustering drift parameters over different behavioral measures to learn an attentive drift of the user, and output a desirable behavior model for machine learning.
Owner:TOYOTA JIDOSHA KK

Sparse array DOA (direction of arrival) estimation method based on adaptive decision level fusion

The invention discloses a sparse array DOA (direction of arrival) estimation method based on adaptive decision level fusion. The method comprises the following steps: establishing a steering vector model for describing receiving signals and channel transmission characteristics of a sparse array; executing exploratory rotation measurement, and constructing an initial historical candidate angle set through incoherent decision-level fusion; entering a self-adaptive verification decision stage, performing clustering analysis on the initial historical candidate angle set to determine an optimal candidate angle, and calculating the global consistency cost to judge whether the angle estimation algorithm is converged or not; and repeating the steps, and carrying out iterative updating and outputting. Direction of arrival estimation is reconstructed into a sequential decision process, and a self-adaptive feedback mechanism is introduced to intelligently plan a measurement process, so that the defect of low efficiency of an existing fixed rotation strategy is overcome. According to the invention, the measurement overhead can be automatically adjusted according to the complexity of a scene, and redundant measurement in a fixed strategy is avoided.
Owner:NANTONG UNIV

Fixed base mechanical arm point-to-point trajectory optimization method based on near-end strategy optimization algorithm

The invention discloses a fixed base mechanical arm point-to-point trajectory optimization method based on a near-end strategy optimization algorithm, solves the problem that a high-dimensional continuous action space exists when deep reinforcement learning is directly used for a fixed base mechanical arm tail end position to reach a task, and belongs to the field of robot operation control and mechanical arm autonomous motion optimization. The method comprises the steps that a mechanical arm kinematics model is established, and constraint conditions are defined; modeling a tail end point-to-point arrival task as a sequential decision process, constructing a state vector as the input of a strategy network, and taking a tail end displacement increment direction vector as the action output of the strategy network; mapping the tail end displacement increment direction vector into a joint increment through differential inverse kinematics, and performing constraint processing according to a constraint condition to generate an executable joint control instruction; and training the strategy network by adopting a near-end strategy optimization algorithm, obtaining a single-step reward through a multi-target reward function, and evaluating the performance output by the action of the strategy network according to the single-step reward until convergence.
Owner:HARBIN INST OF TECH

Photovoltaic bearing capacity assessment method considering space-time uncertainty of flexibility demand of AC / DC hybrid power distribution network

The invention relates to the technical field of power distribution system operation planning, and discloses a photovoltaic bearing capacity evaluation method considering space-time uncertainty of flexibility requirements of an AC / DC hybrid power distribution network, comprising the following steps: generating a robust confidence domain according to half-space depth distribution of data points in an uncertainty set; generating an initial extreme scene set in the robust confidence domain by using a boundary search and sequential decision technology, and building a double-layer robust optimization model taking a main problem and a sub-problem as a core based on the initial extreme scene set; and repeatedly executing iterative solution processing of the main problem and the sub-problem by using the double-layer robust optimization model until the convergence tolerance of a solution result is smaller than the target convergence tolerance, and outputting a photovoltaic bearing capacity evaluation result. According to the method, a data-driven robust uncertain domain is constructed, and a double-layer optimization model of a worst scene is generated through alternate iteration of a main problem and an antagonistic sub-problem, so that the photovoltaic bearing capacity of the alternating current and direct current hybrid power distribution network can be quickly evaluated.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +3