Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Sequential decision" patented technology

Sequential decision making. In artificial intelligence, sequential decision making refers to algorithms that take the dynamics of the world into consideration, thus delay parts of the problem until it must be solved. It can be described as a procedural approach to decision-making, or as a step by step decision theory.

A GNSS / INS tightly coupled adaptive anti-spoofing method based on generalized likelihood ratio and sequential decision

This invention belongs to the field of anti-spoofing and interference technology for satellite navigation and inertial navigation integrated systems. It provides a GNSS / INS tightly coupled adaptive anti-spoofing method based on generalized likelihood ratio and sequential decision-making, applicable to multi-satellite spoofing interference and micro-ramp spoofing scenarios in autonomous driving. The method includes: first, constructing a GNSS / INS tightly coupled model; second, constructing a GNSS spoofing and interference model; third, constructing a generalized likelihood ratio statistic and designing an iterative isolation strategy, improving detection accuracy through sequential hypothesis testing; and fourth, embedding the detection results into a filter and constructing an adaptive factor to achieve anomaly weighting and prediction adjustment. This method integrates spoofing detection and spoofing suppression, significantly improving the anti-spoofing capability of the navigation system.
Owner:LIAONING TECHNICAL UNIVERSITY

A time sequence consistent evidence organization method for aging-oriented question answering

The application provides a time sequence consistent evidence organization method for time limit question answering, comprising: obtaining time limit question answering text; obtaining a candidate evidence set; obtaining priori scores corresponding to each candidate evidence; constructing action scoring references corresponding to the candidate evidence based on the priori scores, and modeling an evidence selection process as a sequential decision process; in each round of candidate evidence selection, constructing a current state according to a current selected evidence set, an unselected candidate set, a time slot coverage state, a semantic abstract state and a remaining budget, calculating action probabilities corresponding to candidate selection actions and a stop action through a strategy network, and selecting a target candidate evidence to join the selected evidence set based on the action probabilities, or selecting the stop action based on the action probabilities; when the strategy network selects the stop action, or the current remaining budget is less than a preset budget threshold, or the current selection round number reaches a preset maximum selection round number, outputting the selected evidence set and generating a corresponding question answering result.
Owner:GUSU LAB OF MATERIALS

A reinforcement learning partition pre-interaction method and system

The application relates to a reinforcement learning partition pre-interaction method and system, wherein the method comprises the following steps: constructing a bias measurement function for evaluating model bias of a virtual environment relative to a real environment; dividing the real environment into multiple trust intervals according to the model bias; performing parameter coarse adjustment on an intelligent agent of the reinforcement learning based on the trust intervals; in the parameter coarse adjustment stage, sampling experience data generated by interaction of the intelligent agent with the virtual environment and the real environment respectively to update network parameters of the intelligent agent; monitoring a performance index of the intelligent agent in the parameter coarse adjustment stage, and when the performance index meets a preset interaction mode switching condition, the intelligent agent enters a parameter fine adjustment stage; in the parameter fine adjustment stage, the intelligent agent completely interacts with the real environment, and parameters of the intelligent agent are adjusted to convergence, and the parameter-converged intelligent agent is used for solving a complex sequential decision problem. The application effectively adjusts parameters of the intelligent agent of the reinforcement learning, and makes the problem-solving capability of the intelligent agent more reliable.
Owner:JIANGNAN UNIV

Intelligent optimization method for standby power consumption of electric vehicle charging system

PendingCN122092643ARealize refined managementFlexible to adaptCharging stationsBatteries circuit arrangementsSleep stateControl engineering
The invention belongs to the technical field of electric vehicle charging, and particularly relates to an intelligent optimization method for standby power consumption of an electric vehicle charging system. In order to overcome the defect that the existing electric vehicle charging system cannot give consideration to reduction of standby power consumption and guarantee of user experience, the invention adopts the following technical scheme: the intelligent optimization method for the standby power consumption of the electric vehicle charging system comprises the following steps: establishing a multi-stage dormant state mechanism; through a data driving method, the optimal state granularity in the multi-stage dormant state is automatically determined, and a state conversion threshold value is dynamically adjusted; acquiring an environment state sensing feature vector; a standby power consumption optimization problem is converted into a sequence decision problem, the sequence decision problem is solved according to a multi-stage dormant state mechanism and an environment state sensing feature vector, and the sequence decision problem is to determine which state conversion action should be executed at what time. The method has the beneficial effect that the dynamic balance between the standby power consumption reduction and the instant response of the charging system is ensured.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT +2

Closed-loop supervised fine-tuning of tokenized traffic models

Imitation learning, or artificial intelligence-based learning from demonstration, aims to acquire an agent policy by observing and mimicking the behavior demonstrated in expert demonstrations. Imitation learning can be used to generate reliable and robust learned policies in a variety of tasks involving sequential decision-making, such as autonomous driving and robotics tasks. However, existing methods that use next-token-prediction (NTP) models, where the policy reduces to a classifier over a discrete set of trajectory tokens, suffer from covariate shift due to their open-loop training a closed-loop execution. The present disclosure provides closed-loop fine tuning of autonomous agent policies in a manner that can mitigate covariate shift.
Owner:NVIDIA CORP

Intelligent decision-making methods, terminal equipment and storage media for power system unit dispatching

ActiveCN117726478BOvercome the impact of schedulingImprove training efficiencyConcurrent computationTraining phase
This invention discloses an intelligent decision-making method, terminal equipment, and storage medium for power system unit scheduling. Based on historical power system operating data, it extracts typical features using dimensionality reduction methods and constructs a feature index set by configuring feature weights using an objective weighting method. A Gaussian mixture clustering model is used for multi-scenario partitioning as a front-end optimization measure for deep reinforcement learning methods, mitigating the suboptimal decision-making problem that may be caused by differences in data distribution across multiple scenarios under source-load uncertainty. The unit scheduling problem is modeled as a sequential decision Markov process, constructing a multi-scenario unit scheduling model based on deep reinforcement learning, overcoming the limitations of the original single-scenario model. Through a dynamic step-size update mechanism and parallel computing, the parameter update efficiency during the offline training phase of the decision network is improved.
Owner:HUNAN UNIV

SUPERVISED FINE-TUNING OF TOKENIZED TRAFFIC MODELS WITH CLOSED REGULATIONS

UndeterminedDE102025142087A1Mathematical modelsArtificial lifeAlgorithmTraffic model
Imitation learning, or AI-based demonstration-based learning, aims to derive a policy for an agent by observing and imitating the behavior shown in expert demonstrations. Imitation learning can be used to generate reliable and robust learned policies for a variety of tasks requiring sequential decision-making, such as autonomous driving and robotics. However, existing methods that use next-token prediction (NTP) models, where the policy is reduced to a classifier over a discrete set of trajectory tokens, suffer from covariate shifting due to their open-loop training and closed-loop execution.The present disclosure provides a fine-tuning of the policies of autonomous closed-loop agents in a manner that can mitigate covariate shifting.
Owner:NVIDIA CORP

A multi-agent path planning method based on a distributed cooperative deep reinforcement learning model

The application discloses a multi-agent path planning method based on a distributed cooperative deep reinforcement learning model, belongs to the communication field and the technical field of reinforcement learning, and comprises the following steps: training the distributed cooperative deep reinforcement learning model designed in the application to obtain a trained distributed cooperative deep reinforcement learning model and randomly generating a map; and based on the trained distributed cooperative deep reinforcement learning model, realizing online path planning of an agent in a multi-agent environment. The distributed cooperative deep reinforcement learning method is different from a centralized deep reinforcement learning method, each agent only makes a decision according to its own observation, meanwhile, communication and sequential decision making among the agents are introduced, and coordination among the agents is promoted. The method is designed from the perspective of a single agent, without considering a group, so that the non-scalability problem of the centralized deep reinforcement learning is avoided.
Owner:NORTHEASTERN UNIV CHINA

AI-based landing page structure optimization method and system, and electronic device

The application provides an AI-based landing page structure optimization method and system and electronic equipment, and relates to the field of structure optimization. The method comprises the following steps: fusing a convolutional neural network and a gradient boosting tree to obtain an initial landing page structure AI deep learning model; taking historical landing page structure features and associated delivery data in a training sample set as inputs of the initial landing page structure AI deep learning model and taking conversion effect data in the training sample set as outputs of the initial landing page structure AI deep learning model to perform model training on the initial landing page structure AI deep learning model, and optimizing hyperparameters through a cross-validation method in the model training process to obtain an intermediate landing page structure AI deep learning model; regarding a landing page optimization process as a sequential decision problem and taking a conversion effect as a reward signal, and real-time receiving delivery data feedback and updating model parameters of the intermediate landing page structure AI deep learning model to obtain a final landing page structure AI deep learning model.
Owner:WEIZHONG DREAM TECH (BEIJING) CO LTD

Multi-expert module cooperative intelligent agent system for pcell layout automatic generation

The application discloses a multi-expert module cooperative intelligent agent system for automatic generation of a Pcell layout, relates to design verification or optimization of an integrated circuit layout, and comprises three expert modules, a decision module, and the like, wherein the three expert modules are trained based on a constructed sequential decision problem to obtain an expert calling sequence; the decision module is used for problem decomposition on an input problem to obtain problem analysis; a target calling sequence matched with the problem analysis is searched as a target calling sequence, and then the problem analysis is input into an expert module at a first sequence position in the target calling sequence to obtain an output result of the expert module at the first sequence position; the problem analysis and the output result of the expert module at a previous sequence position in the calling sequence are taken as inputs of an expert module at a next sequence position, and finally process information corresponding to the input problem is obtained. The application improves the development efficiency of the Pcell layout, realizes fast migration of an existing process layout to a new process, and significantly shortens the research and development cycle.
Owner:SOUTH CHINA UNIV OF TECH

Cross-platform user behavior analysis and media information matching method and system

The application provides a cross-platform user behavior analysis and media information matching method and system, relates to the technical field of user analysis, and comprises the following steps: time sequence analysis of multi-platform user behavior data, generation of a dynamic interest vector with time decay weight, and mapping to a unified semantic space to obtain user potential interest distribution. Based on this, a multi-granularity user group portrait is constructed, and a sequential decision model with long-term conversion revenue as the target is used to dynamically generate platform selection, delivery timing and content matching strategies. After the delivery is executed, the model parameters and decay rules are updated using feedback data to form an optimized closed loop. The application realizes accurate capture of dynamic evolution of user interest and adaptive optimization of cross-platform delivery strategy, and improves marketing efficiency and return on investment.
Owner:BEIJING YUNSHANG XIAOCHENG TECHNOLOGY CO LTD

Crop disease and pest recognition method based on hyperspectral and LiDAR data

The application discloses a crop disease and pest identification method based on hyperspectral and LiDAR data, and belongs to the technical field of pattern recognition. Firstly, a cascaded multi-scale convolution module is used to extract features of hyperspectral images and LiDAR data. Then, a fusion process is modeled as a sequential decision process, and the optimal classification target is promoted step by step. Secondly, four key constraint conditions, i.e., spectral fidelity, feature consistency, structure preservation and resolution matching degree, are monitored in real time at the feature level, and the fusion strategy is adjusted according to the monitored conditions, so that the disease and pest identification information of the two modalities can be fully preserved in the whole fusion process. Finally, the optimal complementary fusion strategy of the hyperspectral image and the LiDAR is obtained through a phased decision, so as to improve the disease and pest identification precision.
Owner:CHINA UNIV OF MINING & TECH

Design space exploration method, system, and computer program product for microarchitecture

PendingCN122263584ADesign optimisation/simulationNeural learning methodsAlgorithmDesign space exploration
The application discloses a design space exploration method, system and computer program product for micro-architecture. The method comprises: inputting a set of micro-architecture parameters of a target hardware, a value range of the parameters and a plurality of preset condition regions into a trained design point generation model, wherein the condition regions are preference regions divided from a predefined target space; generating a plurality of design points for each preset condition region according to the micro-architecture parameters and the value range thereof by using the trained design point generation model in a sequential decision manner, wherein a design point comprises all parameters and a set of values thereof; combining all generated design points, and performing a Pareto dominance relationship comparison to screen out a non-dominated solution set to form a Pareto front approximation. The application can quickly generate a high-quality, reasonable and interpretable Pareto front.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

An AI-driven IVF personalized treatment plan generation system

This invention discloses an AI-driven IVF personalized treatment plan generation system, relating to the fields of medical and artificial intelligence technologies. Through the deep coupling of unsupervised clustering and long-term reinforcement learning planning, this invention achieves a leap from static classification to dynamic strategy generation, constructing a globally optimal personalized treatment path for complex patient groups. It utilizes unsupervised clustering algorithms to perform deep pattern mining on patient baseline data, automatically identifying fine subtypes with similar physiological characteristics. The reinforcement learning agent module models the patient's treatment process as a sequential decision-making process. The agent's state space integrates the patient's subtype attributes and dynamically changing cyclical treatment outcomes. By iteratively building a simulation environment with historical data, a higher success probability is achieved at the global level, transforming clinical decisions based on subjective experience into an automated decision-making process based on historical evidence and global optimization.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

NL2SQL reinforcement learning method based on multi-round SQL perception reward

PendingCN122285690AEngineeringReward value
This invention specifically provides an NL2SQL reinforcement learning method based on multi-turn SQL-aware rewards. The method includes: acquiring sampled trajectory data of multi-turn dialogues in completed sessions; determining the multi-turn dense reward value and total reward value of the completed sessions based on the sampled trajectory data; and selectively outputting an optimized NL2SQL model based on the multi-turn dense reward value and total reward value of the completed sessions. This achieves the goal of modeling multi-turn dialogue NL2SQL tasks as a multi-turn sequential decision-making process using group-relative policy optimization or proximal policy optimization as the optimization framework. Furthermore, by introducing round decay weights, sessions that complete the query task in earlier rounds receive higher incentives, enabling the model to obtain more granular reward feedback in multi-turn interaction scenarios, thereby improving model training efficiency, SQL generation accuracy, and robustness in intelligent agent applications.
Owner:NINGBO TELIAN INFORMATION TECH CO LTD

Ecological environment evolution simulation and planning scheme evaluation system based on reinforcement learning

PendingCN122287286ARealize dynamic modelingPreserve ecological authenticityEcological environmentAdaptive management
This invention belongs to the field of intelligent decision-making technology for ecological environment, specifically disclosing a reinforcement learning-based system for simulating ecological environment evolution and evaluating planning schemes. This invention constructs a closed-loop technical architecture integrating mechanism-driven simulation, data-driven perception, and intelligent decision optimization. This architecture models planning schemes as intelligent agents capable of interacting with the environment and continuously learning, enabling the system to perform sequential decision optimization and dynamic performance evaluation of complex adaptive management strategies while maintaining the physical consistency of ecological processes and environmental quality evolution. The system tightly couples ecological dynamics models, environmental pollution diffusion models, and reinforcement learning intelligent agents, driving strategy optimization through a multi-objective reward function that includes both ecosystem service enhancement and environmental quality improvement, achieving dynamic intervention and adaptive evaluation within a unified spatiotemporal grid. This application supports interpretable and robust decision-making for complex ecological environment planning schemes.
Owner:NANJING ACAD OF ENVIRONMENTAL PROTECTION SCI

Reinforcement learning driven scalable 3D panorama segmentation method and system based on clustering

PendingCN122368458AVoxelAlgorithm
This invention relates to computer vision and point cloud processing technology, specifically a scalable 3D panoramic segmentation method and system based on reinforcement learning-driven clustering. The method includes the following steps: dividing the input point cloud data into semantic voxels to construct an initial super-point graph; constructing a high-level meta-controller and a low-level controller; modeling panoramic segmentation as a sequential graph clustering problem and iteratively processing the initial super-point graph; evaluating the current super-point subgraph using the high-level meta-controller, extracting global feature embeddings, and outputting the probability of whether to terminate segmentation or continue splitting the current super-point subgraph; if the decision is to terminate segmentation, using the current super-point subgraph as the final instance; if the decision is to continue splitting, traversing the edges in the super-point subgraph to be split using the low-level controller and updating the subgraph structure; and employing a near-end policy optimization algorithm with a hybrid reward function for optimization. This invention reconstructs graph clustering into a multi-step sequential decision-making process, effectively solving the problems of under-segmentation and boundary ambiguity in large-scale point clouds.
Owner:SUN YAT SEN UNIV