Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Inverse reinforcement learning" patented technology

An intent-driven communication service quality dynamic guarantee method

This invention discloses an intelligent traffic prediction and scheduling method based on user behavior data, comprising: acquiring the user's actual driving trajectory and matching it to a digital road network; extracting expert trajectory data containing state-action sequences and contextual features; constructing a generative adversarial learning model and using the expert trajectory data for adversarial training, wherein the generator learns to imitate the user's driving behavior, and the discriminator learns to distinguish between the real user trajectory and the trajectory simulated by the generator; using the output of the discriminator after training convergence, assigning a personalized cost value of the user's subjective preference to each state-action pair in the digital road network; based on the constructed personalized cost network, using the A* path search algorithm to plan an optimal personalized driving route for the user; this invention, through inverse reinforcement learning, can automatically learn the user's inherent preferences from the user's behavior, and the planned route is more in line with the user's true intention, improving user satisfaction and system adoption rate.
Owner:WUHAN XINGCHEN WENHUI TECH CO LTD

An Interactive Automated Driving Decision-Making and Planning Method and System Based on Joint Reward Networks

This invention discloses an interactive autonomous driving decision-making and planning method and system based on a joint reward network, relating to the field of autonomous driving technology. The invention includes: constructing a joint reward network, taking the scene state and evaluated candidate joint trajectories as input, and outputting the joint reward for the corresponding candidate joint trajectories; training the joint reward network using a maximum entropy inverse reinforcement learning algorithm until the network parameters converge; calculating the reward for the candidate joint trajectories of the autonomous vehicle and the interactive vehicle using the converged joint reward network; and selecting the candidate joint trajectory with the highest reward as the planning result output. This invention introduces a deep neural network to construct the joint reward function, enabling it to automatically learn complex driving behavior patterns from demonstration data. This avoids the insufficient expressive power problem caused by relying on manually designed features and linear combinations in traditional methods, significantly improving the accuracy of the reward function and its scene generalization ability.
Owner:SOUTHEAST UNIV

Reinforcement learning and dynamic game-based in-construction project risk decision management method and device

PendingCN122264523ABiological modelsOffice automationDecision managementInformation gain
This invention discloses a method and apparatus for risk decision-making and governance of projects under construction based on reinforcement learning and dynamic game theory. The method includes: constructing an engineering game state space based on multi-source heterogeneous data; inferring the implicit payoff functions of participating agents based on maximum entropy inverse reinforcement learning; constructing a dynamic game adversarial training environment including a strategic data fraud mechanism; generating robust governance strategies based on meta-game theory and multi-agent reinforcement learning; performing proactive detection and dynamic risk locking based on maximizing information gain; and deploying and implementing closed-loop evolution of the governance strategies. This invention can identify and quantify the motivation for strategic data fraud, and the generated governance strategies have the ability to predict the opponent's reactions. It overcomes the problem of static models failing during adversarial data drift and can proactively break information asymmetry, solving the technical problem of risk governance failure caused by human game theory.
Owner:TAIAN TAISHAN DEVELOPMENT INVESTMENT CO LTD

A multi-agent motor rehabilitation system and method based on tactile feedback

This invention relates to the field of rehabilitation robots and intelligent control technology, specifically disclosing a multi-agent motor rehabilitation system and method based on tactile feedback. The system includes a multimodal perception unit, a treatment plan generation unit, a multi-agent collaboration unit, and a compliant execution unit connected in sequence. The multimodal perception unit acquires patient kinematics, contact force, tissue characteristics, and pathological semantic information. The treatment plan generation unit extracts rehabilitation decision intentions from expert demonstrations based on inverse reinforcement learning, generating a baseline control target containing motion primitive parameters and impedance parameters. The multi-agent collaboration unit collaboratively corrects the baseline control target under physical coupling constraints based on a distributed partially observable Markov decision model combined with an attention mechanism. The compliant execution unit outputs drive commands and implements rehabilitation training. This invention improves the individual adaptability, coordination, and safety of rehabilitation training.
Owner:OCEAN UNIV OF CHINA

A prediction method based on inverse reinforcement learning

PendingCN122334390AAlgorithmData evolution
The application discloses a prediction method based on inverse reinforcement learning, which firstly maps a long-time sequence original space sample set to a low-dimensional hidden space; then adopts an inverse reinforcement learning algorithm to train a prediction model; takes an atmospheric or ocean data evolution track in the hidden space as an expert track, constructs a discriminator to distinguish the expert track from a prediction track of the prediction model, and optimizes the prediction model by using an implicit reward signal output by the discriminator, so that the track generated by the prediction model approaches the atmospheric or ocean evolution track in the statistical characteristics; and the original space result is output through an encoding-prediction-decoding process during prediction. The method can automatically learn an optimization target conforming to physical laws, and significantly improves long-term stability and dynamic rationality of atmospheric and ocean prediction.
Owner:Pingdu Meteorological Bureau +2

A system and method for collecting energy consumption of loads in a building

The application discloses a kind of building inner load energy consumption collection system and method, its system is by imitative plant vascular structure collection unit, distributed communication network unit, edge computing processing unit, dynamic collection strategy regulation unit, carbon footprint correlation accounting unit and multidimensional visualization unit composition, its core unit is built-in lightweight ATCN network module and carbon factor dynamic update module, realize the accurate decomposition of building load and carbon emission real-time tracking, it runs in system based on " load type - operating state " two-dimensional dynamic collection implementation method, and low-carbon decision optimization method of maximum entropy inverse reinforcement learning, respectively realizes the goal of accurate configuration of collection resource and building energy efficient low-carbon operation.This system and method are suitable for multiple types of building scene, through multidimensional data synchronous collection, edge high-precision load decomposition, dynamic adaptive collection adjustment and energy consumption-carbon emission linkage control, solve the problems, such as low data precision, load identification fuzzy, poor collection efficiency, carbon emission control disconnection, etc. of traditional technology, provide all-round technical support for building energy fine management and low-carbon transformation.
Owner:ANHUI ZHONGKE CARBON DIGITAL TECHNOLOGY CO LTD

Trajectory planning inverse reinforcement learning method, system and product

The trajectory planning inverse reinforcement learning method, system and product belong to the technical field of robots, and solve the technical problems that the existing technology has weak strategy generalization ability and insufficient environment state perception when facing more complex compliant placement tasks. The execution environment and execution parameters of the mechanical arm are initialized; the current state is loaded, and demonstration data is collected; the hierarchical coordinate perception neural network features are extracted according to the current state; the strategy function is constructed to generate the current execution action; the current total reward is calculated based on the current state, the current execution action and the demonstration data; the state value is calculated, and the value advantage of the current execution action is obtained in combination with the current total reward; after the total loss is used to update the corresponding strategy parameters and value parameters, the mechanical arm executes the current action, it is judged whether the current action reaches the target position, and the optimal strategy is output. The present application is used for realizing the human-like string fruit compliant placement of the mechanical arm.
Owner:JILIN AGRICULTURAL UNIV

An algorithm for generating personalized health recommendations based on multi-dimensional health data

The application relates to the technical field of health information, and discloses a personalized health suggestion generation algorithm based on multidimensional health data, which comprises the following steps: constructing a comprehensive context feature vector of a user; combining historical behavior data of the user to quantize personalized health resistance of the user when participating in various health activities; applying an inverse reinforcement learning algorithm to inversely deduce an implicit health preference reward function of the user based on a historical health decision sequence of the user; applying a multi-objective optimization strategy generation strategy to candidate health activities based on a health knowledge graph, personalized health resistance and the implicit health preference reward function to obtain a personalized health suggestion scheme; and iteratively updating the quantization model of the personalized health resistance and the implicit health preference reward function based on feedback data of the user after the user performs health activities according to the scheme; and the application can improve the efficiency of personalized health suggestion generation based on multidimensional health data.
Owner:XINGTUQIHANG ARTIFICIAL INTELLIGENCE TECHNOLOGY (ZIBO) CO LTD

Method for dynamic attack surface real-time discovery and management based on multi-space perception

PendingCN122457332AGraph spectraAttack
The application relates to a kind of dynamic attack surface real-time discovery and management method based on multi-space perception, method includes the following steps: obtaining multi-source heterogeneous data, time uniformity, format normalization and entity identification extraction are carried out on data;In incremental mode, the graph is continuously updated;Update causal strength parameter, evidence confidence parameter and graph structure;Based on the graph after updating, the attack surface risk score and influence range index of asset node are calculated, the risk score is obtained by fusing vulnerability vulnerability, behavior abnormality, exposure, intelligence matching degree and causal propagation probability;Using inverse reinforcement learning and causal effect evaluation generates disposal sequence, reduces the overall attack surface risk under the condition of meeting business cost constraint, and the disposal action and its evidence are written into audit record for backtracking.The application realizes real-time monitoring, active verification, adaptive defense and dynamic balance management of multi-space attack surface, and significantly improves attack discovery rate and response accuracy in complex environment.
Owner:STATE GRID HUBEI ELECTRIC POWER INFORMATION & TELECOMMUNICATION COMPANY

A Reinforcement Learning-Based Intelligent Microgrid Cooperative Control System and Method

PendingCN122315788AStored energySmart microgrid
This invention discloses a cooperative control system and method for intelligent microgrids based on reinforcement learning, comprising: collecting cooperative operation data of the microgrid and preprocessing it to output a standardized operation dataset; constructing state sequence samples to generate a training state sequence set; pairing these samples with control actions and subsequent states to form an expert trajectory library; constructing an improved Reformer model to obtain a multi-agent consistent control representation; generating an access frequency distribution and merging it into an occupancy sample set; obtaining a reward function based on maximum entropy inverse reinforcement learning; and completing policy training to form an online control strategy. This invention, by combining an improved Reformer model and maximum entropy inverse reinforcement learning, achieves real-time and stable cooperative control that satisfies voltage frequency and energy storage state of charge constraints.
Owner:ZHEJIANG ZHIYUAN ENGINEERING MANAGEMENT CO LTD

Inverse reinforcement learning based on positive and negative demonstrations for autonomous driving decision planning

The application discloses a positive and negative demonstration-based inverse reinforcement learning automatic driving decision planning method and relates to the technical field of automatic driving. The application comprises the following steps: a joint cost function is designed based on a safety cost, a traffic cost and a comfort cost, and a joint reward function is obtained by performing the inverse number processing on the joint cost function; the positive and negative demonstration-based inverse reinforcement learning method is adopted to perform training until the joint reward function converges; candidate joint trajectories are sampled according to the states of a self vehicle and an interactive vehicle, and the reward of the candidate joint trajectories of the self vehicle and the interactive vehicle is calculated by using the converged joint reward function; and the candidate joint trajectory with the maximum reward is selected as a planning result output. By introducing the positive demonstration trajectory and the negative demonstration trajectory in the process of the maximum entropy inverse reinforcement learning, the application can guarantee the human nature and effectively improve the safety of the automatic driving planning decision.
Owner:SOUTHEAST UNIV

A Safety-Enhanced Deep Inverse Reinforcement Learning-Based Decision-Making and Planning Method for Autonomous Driving

This invention discloses an autonomous driving decision-making and planning method based on safety-enhanced deep inverse reinforcement learning, relating to the field of autonomous driving technology. The invention includes: constructing a joint reward network, taking the scene state and the candidate joint trajectories to be evaluated as input, and outputting the joint reward for the corresponding candidate joint trajectories; training the network using an inverse reinforcement learning method based on positive and negative demonstrations until the joint reward network parameters converge; and calculating the reward value for each candidate joint trajectory using the converged joint reward network. This invention, through the joint reward network and the positive and negative demonstration learning mechanism, overcomes the limitations of traditional methods in reward function design and safety assurance. Furthermore, by introducing a deep neural network to construct the joint reward function, this invention can automatically learn complex driving behavior patterns from demonstration data, avoiding the insufficient expressive power of traditional methods due to their reliance on manually designed features and linear combinations.
Owner:SOUTHEAST UNIV

A method for socially navigating a mobile robot

The application discloses a kind of mobile robot social navigation method, comprising the following steps: global path planning is carried out based on static environment map, and global reference path is generated;When detecting dynamic pedestrian, start local social navigation, and local social navigation includes: learning and extracting social cost function from human demonstration behavior by inverse reinforcement learning;Using model predictive control method, according to social cost function, robot dynamics constraint and environment constraint, local control trajectory is generated by rolling optimization in limited prediction time domain;Through reinforcement learning regulator, the control cost weight in model predictive control is dynamically adjusted according to real-time environment state, to realize the balance of safety, sociality and navigation efficiency.
Owner:SUZHOU UNIV

Power plant abnormality early warning method and system fusing multi-source time sequence sensing and spatial feature coordination

PendingCN122264512AData processing applicationsBiological modelsTucker decompositionEngineering
The application relates to the technical field of power plant intelligent early warning, and discloses a power plant abnormal early warning method and system fusing multi-source time sequence sensing and spatial feature cooperation, which comprises the following steps: carrying out adaptive normalization of multi-source sensor data based on dynamic statistics; extracting multi-resolution time sequence features by using discrete wavelet transform and multi-branch Transformer, and carrying out weighted fusion through a gating network; combining process topology and data driving to construct a dynamic graph structure, and utilizing two-stage graph convolution to aggregate high-order spatial correlations among devices; carrying out deep nonlinear interaction of space-time features in a tensor space through Tucker decomposition; quantifying classification uncertainty based on an evidence network, and generating early warning actions in combination with an inverse reinforcement learning strategy. The application effectively solves the problem of multi-scale non-stationary and spatial strong coupling of power plant data, and improves the sensitivity and decision reliability of abnormal early warning by fusing expert strategies and uncertainty measurement.
Owner:HUANENG LANCANG RIVER HYDROPOWER CO LTD

Intelligent ship navigation control method and system based on maximum entropy inverse reinforcement learning

This application discloses an intelligent ship navigation control method and system based on maximum entropy inverse reinforcement learning. The method includes: collecting expert demonstration trajectory data; preprocessing the expert demonstration trajectory data to obtain an expert trajectory dataset; and then constructing a ship motion simulation model based on the expert trajectory dataset; obtaining an optimal reward function based on the maximum entropy inverse reinforcement learning algorithm, according to the expert trajectory dataset and the ship motion simulation model; obtaining an optimal policy network based on a reinforcement learning algorithm, according to the optimal reward function and the ship motion simulation model; acquiring ship state data; and generating ship motion control commands based on the optimal policy network and the ship state data. This application can generate navigation control commands that are safe, efficient, and deeply aligned with the wisdom of maritime practice without relying on manually preset rewards, and can be widely applied in the field of ship control technology.
Owner:WUHAN UNIV OF TECH

A highway new energy charging service intelligent optimization analysis method and system

The application discloses a kind of expressway new energy charging service intelligent optimization analysis method and system, method includes: constructing structured expressway travel big data set.Extract vehicle multidimensional feature to distinguish new energy vehicle and fuel vehicle, obtain basic information;Based on inverse reinforcement learning, construct Markov decision process model, combine the fitting optimal return function of kiosk data, solve trajectory strategy, reconstruct new energy vehicle space-time trajectory, and mine space-time characteristics of charging demand.Combining service area distribution, power supply and cost constraints, construct multi-objective optimization model, through reinforcement learning and traffic simulation training, iteration site selection and facility configuration, output full-cycle layout and flexible configuration scheme.Real-time acquisition vehicle and site state, construct "car-station" collaborative scheduling model, generate site matching, pile position allocation and traffic guidance strategy, realize charging demand and service ability adaptive scheduling.
Owner:HUBEI INTELLIGENT TRANSPORTATION RES INST CO LTD

Deep learning-based personalized rehabilitation nursing scheme generation method and intelligent chip

This invention relates to the field of rehabilitation data processing technology, and in particular to a method for generating personalized rehabilitation care plans based on deep learning and an intelligent chip. First, multimodal data including patient physiological signals, movement data, subjective feedback, rehabilitation therapist assessments, and environmental information is loaded. State modeling is then used to obtain a patient implicit state vector that integrates physiological, motor, and psychological data. Next, an initial personalized reward function is learned using inverse reinforcement learning. A meta-learner dynamically adjusts the reward function weights based on the patient's state and rehabilitation effect. The updated reward function and implicit state vector are input into a reinforcement learning policy network to optimize and obtain the optimal rehabilitation action, which is then sent to the rehabilitation execution device. Simultaneously, new data is collected for closed-loop iteration, achieving long-term adaptive rehabilitation plan generation.
Owner:SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL