Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

81 results about "Inverse reinforcement learning" patented technology

Track planning inverse reinforcement learning method, system and product

The invention discloses a trajectory planning inverse reinforcement learning method, system and product, belongs to the technical field of robots, and solves the technical problems of weak strategy generalization ability and insufficient environmental state perception when facing a more complex compliant placement task in the prior art. Initializing an execution environment and execution parameters of the mechanical arm; loading a current state, and collecting demonstration data; extracting hierarchical coordinate sensing neural network features according to the current state; constructing a strategy function, and generating a current execution action; calculating a current total reward based on the current state, the current execution action and the demonstration data; calculating a state value, and combining the current total reward to obtain a value advantage of the current execution action; and after the corresponding strategy parameters and value parameters are updated through the total loss, the mechanical arm is adopted to execute the current action, whether the current action reaches the target position or not is judged, and the optimal strategy is output. The mechanical arm is used for achieving humanoid fruit bunching and smooth placement.
Owner:JILIN AGRICULTURAL UNIV

Motor health management method and system based on full life cycle monitoring

The invention provides a motor health management method and system based on full life cycle monitoring, and the method comprises the steps: firstly collecting motor multi-source sensor data, and mapping the motor multi-source sensor data into a coordinate point in a low-dimensional health state space through a manifold learning model, so as to represent the comprehensive health state of the motor multi-source sensor data; and on the basis of the expert strategy database, a health cost function associated with the health coordinates is constructed by adopting inverse reinforcement learning and is used for quantifying health loss of different operations. And a decision model is constructed, and the model is trained through an optimal control algorithm, so that an operation strategy capable of minimizing the total health cost of the whole life cycle of the motor can be found out. In actual operation, the system determines current health coordinates in real time, generates an optimal operation strategy by using a decision model in combination with future task requirements, and analyzes the optimal operation strategy into specific control parameters to be issued and executed. According to the method, active health decline path management and closed-loop control of the motor can be realized, so that dynamic optimal balance between the performance and the service life of the motor is realized.
Owner:CHINA CARBON MEDIGA (WUHAN) TECHNOLOGY CO LTD

Prediction method for interaction track of right-turn vehicle and pedestrian at intersection and computer equipment

The invention discloses an intersection right-turn vehicle and pedestrian interaction trajectory prediction method and computer equipment. The method comprises the following steps: acquiring a human-vehicle historical trajectory and initial state information in intersection right-turn vehicle and pedestrian interaction; a confrontation reinforcement learning structure is improved through KL regularization, and combined learning of a reward function and a strategy is achieved; a Nash Q-learning algorithm with a KL constraint is utilized to carry out joint optimization; and inputting the optimized strategy model into a behavior simulator, driving a generated prediction trajectory by an initial state, comparing the prediction trajectory with a real trajectory, and evaluating a prediction result based on a specified index. According to the method, the performance of the model in the aspects of prediction precision, behavior consistency and environment generalization ability can be effectively improved, the stability of a reward function and a strategy network is further improved, the generalization ability of the model in different traffic scenes subsequently is enhanced, and the learned strategy has better migration potential.
Owner:NANJING UNIV OF POSTS & TELECOMM

Injection molding production optimization method based on multi-agent cooperation

InactiveCN121836000AImprove collaborative optimization capabilitiesEnsure coordination and unityForecastingArtificial lifeOptimal decisionDecision strategy
The invention discloses an injection molding production optimization method based on multi-agent cooperation, and the method comprises the following steps: decomposing a plurality of targets in an injection molding production process, and constructing a layered multi-agent structure; collecting data in the injection molding production process in real time, and constructing a global expert behavior track and a local expert behavior track; based on the global target and each local target, obtaining initial parameters of a global reward function and a local reward function; adopting an inverse reinforcement learning method to obtain an optimal global reward function and an optimal local reward function; obtaining an optimal decision strategy of each agent through a hierarchical collaboration mechanism by each hierarchical agent; when detecting that a conflict exists between the targets, dynamically adjusting the reward weight of the reward function by adopting a conflict coordination mechanism; and continuously collecting new expert behavior data, and updating the optimal decision strategy of each agent. According to the invention, a layered multi-agent inverse reinforcement learning method is adopted, and multi-target dynamic collaborative optimization of injection molding production is realized.
Owner:HEBEI QUANYUN INTELLIGENT TECH CO LTD

Task type dialogue strategy learning method and system based on reinforcement learning

The invention belongs to the technical field of intelligent task-based dialogue, and particularly relates to a task-based dialogue strategy learning method and system based on reinforcement learning, which adopts a soft actor commentator algorithm to be combined with behavior cloning and maximum entropy inverse reinforcement learning to relieve a cold start problem. According to the method, maximum entropy inverse reinforcement learning is utilized, a reward value is calculated, and a user target in a dialogue is accurately deduced according to maximum log likelihood estimation of a human dialogue and a simulated dialogue; a potential reward function is extracted from a successful track through inverse reinforcement learning, manual reward design is replaced, accumulated reward maximization is pursued during strategy optimization, diversity and exploratory performance of the strategy are encouraged, and the dialogue strategy is prevented from being converged to a single mode too early.
Owner:QUFU NORMAL UNIV

Intelligent optimization method and system for enterprise management system based on digital twinning

The invention provides an enterprise management system intelligent optimization method and system based on digital twinning, and the method comprises the following steps: constructing an organization structure diagram with a path weight by taking a post as a node, taking a superior-subordinate relationship as a directed edge, and taking a KPI (Key Performance Indicator) target vector as a node attribute; constructing an instruction conduction path model through breadth-first traversal based on the organization structure chart to generate an effective conduction path set, and calculating a target propagation vector of each path; a post is regarded as a reinforcement learning agent, an optimal post behavior strategy is recovered through inverse reinforcement learning, and a target vector matrix of each post adjusted according to the strategy is obtained; constructing an anti-fact management strategy optimization instruction conduction path, and generating an optimized target matrix and path set; and comparing the optimized target matrix with actual execution data, and calculating a target deviation as an execution effect evaluation value. According to the method, the defects of an existing system in the aspects of multi-level management collaboration, behavior deviation modeling and strategic execution consistency guarantee are overcome.
Owner:GUANGZHOU SHUANGHE SOFTWARE CO LTD +4

Supply chain fund flow intelligent scheduling method based on reinforcement learning

The invention discloses a supply chain fund flow intelligent scheduling method based on reinforcement learning, and the method comprises the steps: S1, collecting multi-dimensional heterogeneous data, constructing a dynamic digital twin environment, and carrying out the fusion to generate a standardized input tensor; s2, aggregating node features by using a heterogeneous graph attention network, and outputting a high-dimensional state feature graph; s3, analyzing market fluctuation and risk through a high-level agent, and outputting a global scheduling guidance strategy; s4, outputting a refined scheduling instruction set by using a DPPO algorithm in combination with tail risk minimization; s5, based on the improved MAPPO algorithm, noise is injected to cover privacy, and Nash equilibrium optimization is executed through inverse reinforcement learning; and S6, fund allocation is executed, and online iterative optimization is carried out. According to the invention, supply chain fund scheduling efficiency and security are improved, and individual and collective benefits are effectively balanced.
Owner:SHENGWEI (TIANJIN) TECHNOLOGY DEVELOPMENT CO LTD

Unmanned system multi-intention recognition method and system based on inverse reinforcement learning

The invention discloses an unmanned system multi-intention recognition method and system based on inverse reinforcement learning, and belongs to the technical field of artificial intelligence and autonomous decision making. According to the system, a deep multi-intention inverse reinforcement learning framework is provided for solving the problems of multi-modality, fuzziness and dynamics of unmanned system decision-making intentions in a complex dynamic environment. By designing a multi-scale intention encoder based on Transform, state-intention combined dynamic reasoning is realized; constructing an end-to-end intention perception reward network, and adaptively fusing multi-intention features through an attention mechanism; and proposing an expectation maximization optimization strategy to realize collaborative optimization of intention reasoning and strategy learning. The system can automatically separate and identify various potential intentions from expert demonstration of mixed intentions, and learn corresponding reward functions and decision strategies. According to the method, the intention recognition accuracy and the strategy reproduction capability in a multi-intention scene are remarkably improved, and an effective solution is provided for behavior understanding and autonomous decision making of an unmanned system in a complex environment.
Owner:GUILIN UNIV OF ELECTRONIC TECH +1

Decision model training method and device and driving decision method and device

The invention discloses a decision model training method and device and a driving decision method and device, and the method comprises the steps: inputting each first candidate decision behavior corresponding to a first traffic scene into a to-be-trained decision model, and obtaining the decision probability of each first candidate decision behavior; based on the decision probability of each first candidate decision behavior, performing inverse reinforcement learning on each first candidate decision behavior to obtain an inverse reinforcement learning loss value of the first traffic scene; performing imitation learning on the first optimal decision behavior with the maximum decision probability in the first candidate decision behaviors to obtain an imitation learning loss value of the first traffic scene; based on the inverse reinforcement weight and the imitation weight, performing weighted fusion on the inverse reinforcement learning loss value and the imitation learning loss value to obtain a weighted loss value of the first traffic scene; and based on the weighted loss value, optimizing model parameters of the decision model to be trained to obtain a trained decision model. The reliability of the decision model can be improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Task-oriented dialogue system reward function optimization method and system

The invention discloses a task-oriented dialogue system reward function optimization method and system, and belongs to the technical field of natural language processing. The method comprises the following steps: acquiring an expert dialogue track from a dialogue system data set, extracting a state-action-reward triple, training an initial reward function by using a maximum entropy inverse reinforcement learning framework, and initializing a strategy network; an Actor-Critic reinforcement learning algorithm is adopted to train a strategy network, and a suboptimal trajectory is collected to dynamically update a reward function; and taking the dynamic reward function as a unified evaluation signal, and optimizing the strategy network to form a dialogue strategy model. Through dynamic reward function optimization, manual rule dependence is reduced, the generalization ability, the task completion rate and the stability of a dialogue system are improved, and the method is suitable for complex dialogue scenes in multiple fields.
Owner:XIAN UNIV OF POSTS & TELECOMM

Security unmanned vehicle path planning method fusing tactical risk map and game prediction

The invention discloses a security and protection unmanned vehicle path planning method fusing a tactical risk map and game prediction, and relates to the technical field of intelligent security and protection, and the method comprises the steps: receiving a task path issued by a remote control security and protection unmanned vehicle, carrying out the semantic perception of an environment, and constructing a static tactical risk map; acquiring real-time observation track sequences of potential opponents, inferring target intentions of the opponents on line through a Bayesian inverse reinforcement learning BIRL model, and generating a high-risk region thermodynamic diagram; generating a speed and angular speed control instruction through a path tracking controller based on the balanced deterrent path; and executing a speed and angular speed control instruction, collecting a real-time observation track sequence of a potential opponent, and performing online fine adjustment of a Bayesian inverse reinforcement learning BIRL model to generate an optimized security unmanned vehicle path. According to the invention, the unmanned vehicle can continuously generate an optimized path with both safety and tactical initiative in a dynamic confrontation environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

An intent-driven communication service quality dynamic guarantee method

This invention discloses an intelligent traffic prediction and scheduling method based on user behavior data, comprising: acquiring the user's actual driving trajectory and matching it to a digital road network; extracting expert trajectory data containing state-action sequences and contextual features; constructing a generative adversarial learning model and using the expert trajectory data for adversarial training, wherein the generator learns to imitate the user's driving behavior, and the discriminator learns to distinguish between the real user trajectory and the trajectory simulated by the generator; using the output of the discriminator after training convergence, assigning a personalized cost value of the user's subjective preference to each state-action pair in the digital road network; based on the constructed personalized cost network, using the A* path search algorithm to plan an optimal personalized driving route for the user; this invention, through inverse reinforcement learning, can automatically learn the user's inherent preferences from the user's behavior, and the planned route is more in line with the user's true intention, improving user satisfaction and system adoption rate.
Owner:WUHAN XINGCHEN WENHUI TECH CO LTD

Building energy management system for source network load storage and utilization

The invention relates to the technical field of energy management, and discloses a source network load storage and utilization oriented building energy management system, which comprises at least one intelligent sensing terminal deployed in a user physical space, and a cloud data processing platform running on a cloud infrastructure, and the intelligent sensing terminal and the cloud data processing platform perform real-time interaction through a safe and encrypted bidirectional data communication link to form a complete closed loop of sensing, analysis, decision making, control and feedback. According to the invention, through multi-dimensional data acquisition of the intelligent sensing terminal and in combination with a non-intrusive load decomposition algorithm of the edge calculation unit, the total load can be decomposed to each energy consumption device in real time, and an independent operation power curve is generated; the cloud digital behavior twinborn modeling module further predicts a baseline load through a long short-term memory network, quantifies the comfort level and economic preference of the user through maximum entropy inverse reinforcement learning, and achieves deep analysis of the power consumption behavior of the user.
Owner:ZHEJIANG PROVINCE ELECTRIC POWER FUEL CORP +1

An unmanned ship dynamic collision avoidance decision method and system based on a diffusion model and adversarial inverse reinforcement learning

The application provides an unmanned ship dynamic collision avoidance decision method and system based on a diffusion model and an adversarial inverse reinforcement learning. The method comprises the following steps: 1, collecting expert collision avoidance data; 2, training a diffusion model; 3, performing data enhancement; 4, building a training environment; 5, constructing a generator and a discriminator; 6, updating the discriminator; 7, updating the generator; and 8, completing the training. The diffusion model is trained by using limited expert sample data, additional collision avoidance data is generated by using the trained diffusion model, and the additional data and the original expert data jointly participate in the training process of the adversarial inverse reinforcement learning, so that the problem of limited expert collision avoidance data can be effectively solved. Furthermore, an engineering application plan is provided for the proposed algorithm, the algorithm is verified, and a systematic solution is formed by combining theory with practice.
Owner:JIMEI UNIV

Breast cancer MRI image enhancement and focus automatic identification method based on attention mechanism

The invention discloses a breast cancer MRI image enhancement and focus automatic identification method based on an attention mechanism, and relates to the technical field of data processing and computer-aided diagnosis. The method comprises the following steps: acquiring a multi-time-point mammary gland MRI image sample and performing preprocessing; performing random window enhancement and random channel exchangeable enhancement on the preprocessed sample, inputting the enhanced image sample into an improved convolutional neural network, and performing deep metric learning loss function optimization based on proxy NCA to extract focus embedding features; according to the focus embedding features, constructing a moment matching game framework with focus state distribution marked by experts as a target; and according to a moment matching game framework, adopting an inverse reinforcement learning strategy without reinforcement learning, carrying out iterative optimization on parameters of the lesion recognition model through a regret-moment-free matching algorithm, carrying out lesion property judgment according to the optimized model, and outputting a judgment result and a corresponding confidence score. Therefore, MRI image adaptive enhancement and accurate focus automatic identification are realized.
Owner:THE AFFILIATED HOSPITAL OF XUZHOU MEDICAL UNIV +1

Multi-time scale microgrid economic dispatching method based on relative entropy inverse reinforcement learning

The invention discloses a multi-time-scale microgrid economic dispatching method based on relative entropy inverse reinforcement learning, and the method comprises the steps: obtaining a day-ahead dispatching plan and a tie line dispatching plan according to a day-ahead prediction information set of a microgrid; superposing disturbance on the day-ahead scheduling plan to carry out dimension expansion, and generating an expert strategy for inverse reinforcement learning; according to the expert strategy, a reward function is constructed by adopting a relative entropy inverse reinforcement learning algorithm, and the reward function is obtained by minimizing the relative entropy of trajectory distribution and solving a corresponding Lagrange duality function; according to the reward function, adopting a depth deterministic strategy gradient algorithm to train an intra-day scheduling agent to obtain a trained intra-day scheduling model; obtaining ultra-short-term prediction information based on the tie line scheduling plan; and inputting the ultra-short-term prediction information into an intra-day scheduling model to obtain an intra-day scheduling strategy.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Chemical text attribute extraction method and system based on inverse reinforcement learning

The invention discloses a chemical text attribute extraction method and system based on inverse reinforcement learning, and the method comprises the steps: reconstructing a chemical text attribute extraction task into a generative question and answer task, converting an input text and a target attribute into a question and answer pair, and extracting attribute information through the semantic understanding and generation capability of a question and answer model; a BioBART model is selected as a basic model, and fine tuning is carried out on a public biomedical question and answer data set, so that the model preliminarily adapts to a question and answer task; constructing a reward model comprising a plurality of sub-reward components; and performing model training by using maximum entropy inverse reinforcement learning, and alternately performing reward model updating and question and answer model updating. According to the method, the attribute extraction task is converted into the natural language question and answer task, and the BioBART model and inverse reinforcement learning are utilized to optimize the multi-target award function, so that the understanding and generation capability of the model on complex scientific texts is remarkably improved, implicit attributes in chemical texts can be effectively extracted, and powerful support is provided for customs to quickly identify suspicious substances.
Owner:JIANGNAN UNIV

Low-protein diversified daily ration design method and system based on multi-source data fusion

The invention relates to the technical field of livestock breeding and feed nutrition, and discloses a low-protein diversified daily ration design method and system based on multi-source data fusion, and the method comprises the steps: constructing a formula-environment-effect directed causal diagram through a PC algorithm, estimating the effect deviation contribution proportion of an environment factor and a formula factor through anti-factual reasoning, and calculating the effect deviation contribution proportion of the environment factor and the formula factor; and inferring the target weight preference of the farm by using inverse reinforcement learning, and carrying out conditional weight updating according to a causal attribution result. Real reasons of growth effect deviation can be distinguished, effect deviation errors caused by environmental factors are prevented from being attributed to improper formula setting, and the accuracy and stability of formula recommendation are improved.
Owner:LIAONING WELLHOPE AGRI TECH

Inverse reinforcement learning automatic parking method and system based on driving habits

The invention provides an inverse reinforcement learning automatic parking method and system based on driving habits. The method comprises the steps that an expert parking state data set is constructed according to expert tracks under typical parking working conditions; constructing a maximum entropy probability model based on an expert track according to a parameterized reward model based on linear feature combination; by utilizing a maximum entropy inverse reinforcement learning algorithm, optimizing and solving a weight vector of a reward function in the maximum entropy probability model by maximizing a log-likelihood function of an expert track, and obtaining an optimal reward function; based on the optimal reward function, adopting a depth deterministic strategy gradient algorithm to train a strategy network so as to obtain an optimal parking strategy; and according to the optimal parking strategy, a two-stage parking control architecture is adopted to execute parking control on the vehicle. According to the method, the learning efficiency and convergence stability of automatic parking are remarkably improved, the smoothness of the parking track and the accuracy of final parking are guaranteed, and the reliability and riding comfort of the automatic driving vehicle in the parking process are improved.
Owner:NANCHANG AUTOMOTIVE INST OF INTELLIGENCE & NEW ENERGY

Film and television film production actuarial and risk intelligent control system

The invention relates to the technical field of film and television investment intelligent decision, and discloses a film and television production actuarial and risk intelligent control system, which comprises seven modules, namely a data annotation module, a feature extraction and conversion module, a symbol reasoning and evaluation module, a scheme generation module, a monitoring and early warning module, a decision optimization verification module and a knowledge updating module. According to the method, semantic annotation is performed through a multi-channel adaptive method, multi-modal feature extraction is performed, a feature-to-symbol conversion mechanism is constructed, multi-level cognitive reasoning is performed by adopting an OpenNARS non-axiom reasoning engine, a personalized investment scheme is generated by adjusting cognitive parameter configuration, and a risk early warning report is obtained through real-time data monitoring and risk assessment. Decision optimization and cross validation are carried out through an imitation learning method and an inverse reinforcement learning method. According to the method, the precision, individuation and interpretability of film and television project investment decision can be realized, the return on investment is improved, and the project risk is reduced.
Owner:SHENZHEN VISION CULTURAL IND INVESTMENT MANAGEMENT CO LTD

A dispatching method and system for matching airline baggage change information

The application discloses a kind of scheduling method and system of matching aviation luggage change information, it is related to aviation transport intelligent scheduling technical field, including, real-time acquisition luggage change information in aviation transport scene, passenger itinerary data and airport equipment state data, output original data stream;In the metaverse, the airport operation digital twin is constructed, and the collaborative transfer path is imported into the digital twin environment for space-time pre-play simulation, real-time detection of luggage flow conflict points and marking of physical constraint conflict region coordinate set;By inverse reinforcement learning algorithm, the conflict events in the pre-play simulation are solved, the strategy defect features are traced, and the reward function space is reconstructed, and the gradient compensation parameters are output;Fusion physical constraint conflict region coordinate set and gradient compensation parameters, generate three-dimensional space scheduling instruction set and execute instruction issuing.The application is based on space-time pre-play simulation method, accurately identifies the safety boundary of luggage transfer path, and effectively reduces the risk of physical constraint overrun.
Owner:SHANGHAI PENGBANG IND CO LTD

A mixed traffic interaction behavior trajectory prediction method, device and storage medium

This invention discloses a method, device, and storage medium for predicting the trajectory of mixed traffic interaction behavior. The method includes: acquiring historical trajectory data and initial state information of multiple agents; inputting the historical trajectory data into a maximum entropy inverse reinforcement learning model to construct a maximum entropy behavior generation probability model based on the interaction behavior of multiple agents; employing a MA-AIRL framework with embedded dynamic reward generation for behavior backpropagation modeling and extracting implicit intentions; simultaneously utilizing Nash Q-learning based on Nash equilibrium to learn the game relationship between agents and generate a cooperative strategy; finally, inputting the strategy into a simulator to generate a predicted trajectory; comparing the predicted results with the actual trajectories; and verifying the accuracy of the prediction model based on specified indicators. This invention can significantly improve the accuracy and stability of multi-agent interaction trajectory prediction and effectively enhance the reliability of interaction behavior trajectory prediction results in mixed traffic environments.
Owner:NANJING UNIV OF POSTS & TELECOMM

Unmanned vehicle trajectory prediction method based on fuzzy inverse reinforcement learning

The application discloses an unmanned vehicle trajectory prediction method based on fuzzy inverse reinforcement learning, which comprises the following steps: calculating trajectory features, introducing fuzzy logic for inverse reinforcement learning to find an optimal weight vector, and optimal trajectory prediction simulation. The application mainly processes unmanned vehicle driving trajectory data and actual data of human drivers, combines fuzzy logic processing, trains an inverse reinforcement learning model, obtains an optimal weight vector, and thus calculates the optimal driving trajectory of the unmanned vehicle in a limited trajectory. The application solves the problems of poor adaptability to complex environments and low decision accuracy of the unmanned vehicle in the driving strategy process in the prior art, and significantly improves the driving strategy efficiency and safety of the unmanned vehicle in the actual road environment. The introduction of the fuzzy inverse reinforcement learning method enables the unmanned vehicle to better learn the driving strategy of human drivers, thereby realizing more intelligent and safe automatic driving functions.
Owner:EAST CHINA NORMAL UNIV

A multi-mode comprehensive control method and device for surgical robot body intelligence

The application provides a multi-mode comprehensive control method and device for surgical robot body intelligence, the method comprising: obtaining a first reward function based on inverse reinforcement learning processing; obtaining target position and obstacle information; constructing a second reward function based on the target position and a third reward function based on the obstacle information; training the operation strategy of the mechanical arm through reinforcement learning according to the first reward function, the second reward function and the third reward function; selecting an operation trajectory with the highest cumulative reward and controlling the mechanical arm to execute the operation trajectory. In the application, the first reward function is generated through historical operation experience, the second reward function and the third reward function are generated through visual obstacle avoidance; the overall operation strategy based on different reward functions is realized through the reinforcement learning mode, thereby realizing the comprehensive control of the mechanical arm based on historical operation experience and visual obstacle avoidance.
Owner:LONGWOOD VALLEY MEDICAL TECH CO LTD

Intelligent home energy management method based on deep adversarial inverse reinforcement learning

This invention discloses a smart home energy management method based on deep adversarial inverse reinforcement learning, comprising the following steps: (1) Modeling the energy cost minimization problem of a smart home and designing the corresponding Markov decision process environment state and actions; (2) Constructing n sets of random parameter historical data based on a set of random parameter historical data and a rolling generation method; (3) Solving the above minimization problem using the n sets of random parameter historical data and an optimization algorithm to obtain n expert trajectories; (4) The discriminator trains a reward network based on the expert trajectories and the trajectories generated by the generator agent; Under the guidance of the reward network, the generator agent is trained using experience tuples and a proximal policy optimization algorithm; (5) Repeating step (4) until a stable agent policy is obtained; (6) Deploying the trained agent policy in a real environment. Compared with existing methods, the method of this invention can effectively reduce energy costs and improve user comfort.
Owner:NANJING UNIV OF POSTS & TELECOMM

An Interactive Automated Driving Decision-Making and Planning Method and System Based on Joint Reward Networks

This invention discloses an interactive autonomous driving decision-making and planning method and system based on a joint reward network, relating to the field of autonomous driving technology. The invention includes: constructing a joint reward network, taking the scene state and evaluated candidate joint trajectories as input, and outputting the joint reward for the corresponding candidate joint trajectories; training the joint reward network using a maximum entropy inverse reinforcement learning algorithm until the network parameters converge; calculating the reward for the candidate joint trajectories of the autonomous vehicle and the interactive vehicle using the converged joint reward network; and selecting the candidate joint trajectory with the highest reward as the planning result output. This invention introduces a deep neural network to construct the joint reward function, enabling it to automatically learn complex driving behavior patterns from demonstration data. This avoids the insufficient expressive power problem caused by relying on manually designed features and linear combinations in traditional methods, significantly improving the accuracy of the reward function and its scene generalization ability.
Owner:SOUTHEAST UNIV

Guided classroom teaching content generation method and system based on inverse reinforcement learning

The invention discloses a guide type classroom teaching content generation method and system based on inverse reinforcement learning, and the method comprises the steps: firstly constructing a teaching process automaton according to expert classroom dialogues and behavior annotations, defining a teaching stage as a state set, defining interactive behaviors as an action set, and building a legal transfer relation and an acceptance state; on the basis, trainable transfer parameters are introduced, a probabilistic automaton is constructed, and track structure representation is obtained through state distribution propagation. Then constructing an expert track set and a comparison track set, establishing a reward network, and obtaining a track reward function through difference constraint training; and performing behavior recognition on the generated classroom dialogue to form a generation track, and performing reinforcement learning optimization on the generation model based on a reward signal. And finally, outputting classroom teaching content meeting the teaching process constraint, thereby realizing explicit modeling and continuous optimization of classroom structure and behavior logic, and improving the structural rationality and teaching guidance effect of the generated content.
Owner:EAST CHINA NORMAL UNIV

An unmanned driving reward learning and control method based on integrated maximum entropy deep inverse reinforcement learning

This invention discloses an autonomous driving reward learning and control method based on ensemble maximum entropy deep inverse reinforcement learning, comprising the following steps: Step 1: Learning the reward function and control operation in a highway autonomous vehicle driving environment and constructing it into a Markov decision process model; Step 2: Constructing a soft Q-learning model based on value pruning to obtain expert demonstrations, and dividing the inverse reinforcement learning task into sub-tasks according to expert preferences; Step 3: Establishing a strong learner ensemble model and recovering the reward function through maximum entropy deep inverse reinforcement learning; Step 4: Achieving the fusion of reward functions of each sub-task through linear combination, thereby improving the learning accuracy of the reward function. This invention considers the problems of gradient explosion, gradient vanishing, and data overflow in soft Q-learning. It learns expert demonstrations through an improved soft Q-learning algorithm and proposes an ensemble maximum entropy deep inverse reinforcement learning algorithm based on the learned expert demonstrations, which can better achieve decision control.
Owner:BEIJING UNIV OF CHEM TECH

Urban sewage treatment process dissolved oxygen concentration safety control method based on inverse reinforcement learning

The invention provides a municipal sewage treatment process dissolved oxygen concentration safety control method based on inverse reinforcement learning so as to realize stable regulation and control of dissolved oxygen concentration under false data invasion and guarantee stable operation of a municipal sewage treatment process. Firstly, an urban sewage treatment process dissolved oxygen concentration control system with false data injection attacks is constructed; secondly, constructing a cost function of an error dynamic system; finally, a safety controller based on inverse reinforcement learning is designed, and safe and stable control over the dissolved oxygen concentration of the urban sewage treatment system is achieved. Experimental results show that the method can realize accurate control after the urban sewage treatment system is invaded by false data, and ensures safe and stable operation of the urban sewage treatment system.
Owner:BEIJING UNIV OF TECH

Industrial wastewater discharge intelligent optimization control method based on deep reinforcement learning

The invention discloses an industrial wastewater discharge intelligent optimization control method based on deep reinforcement learning, and the method comprises the following steps: S1, collecting the water quality data of industrial wastewater, and carrying out the preprocessing of the data; s2, deducing an optimal reward function by adopting a maximum entropy inverse reinforcement learning method; s3, optimizing a control strategy through a Q-learning reinforcement learning model, and training the Q-learning reinforcement learning model to obtain an optimal control strategy; s4, optimizing control parameters by adopting a fractional order cuckoo search algorithm based on the optimal control strategy; and S5, applying the optimal control parameters to a wastewater treatment control system. According to the method, maximum entropy inverse reinforcement learning, Q-learning and fractional order cuckoo search algorithms are adopted, intelligent optimization of wastewater treatment control parameters is achieved, and the method has the advantages of being high in adaptability, high in pollutant removal rate and low in energy consumption.
Owner:HEFEI OUWO ENVIRONMENTAL PROTECTION TECH CO LTD