Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

103 results about "Inverse reinforcement learning" patented technology

Track planning inverse reinforcement learning method, system and product

The invention discloses a trajectory planning inverse reinforcement learning method, system and product, belongs to the technical field of robots, and solves the technical problems of weak strategy generalization ability and insufficient environmental state perception when facing a more complex compliant placement task in the prior art. Initializing an execution environment and execution parameters of the mechanical arm; loading a current state, and collecting demonstration data; extracting hierarchical coordinate sensing neural network features according to the current state; constructing a strategy function, and generating a current execution action; calculating a current total reward based on the current state, the current execution action and the demonstration data; calculating a state value, and combining the current total reward to obtain a value advantage of the current execution action; and after the corresponding strategy parameters and value parameters are updated through the total loss, the mechanical arm is adopted to execute the current action, whether the current action reaches the target position or not is judged, and the optimal strategy is output. The mechanical arm is used for achieving humanoid fruit bunching and smooth placement.
Owner:JILIN AGRICULTURAL UNIV

Motor health management method and system based on full life cycle monitoring

The invention provides a motor health management method and system based on full life cycle monitoring, and the method comprises the steps: firstly collecting motor multi-source sensor data, and mapping the motor multi-source sensor data into a coordinate point in a low-dimensional health state space through a manifold learning model, so as to represent the comprehensive health state of the motor multi-source sensor data; and on the basis of the expert strategy database, a health cost function associated with the health coordinates is constructed by adopting inverse reinforcement learning and is used for quantifying health loss of different operations. And a decision model is constructed, and the model is trained through an optimal control algorithm, so that an operation strategy capable of minimizing the total health cost of the whole life cycle of the motor can be found out. In actual operation, the system determines current health coordinates in real time, generates an optimal operation strategy by using a decision model in combination with future task requirements, and analyzes the optimal operation strategy into specific control parameters to be issued and executed. According to the method, active health decline path management and closed-loop control of the motor can be realized, so that dynamic optimal balance between the performance and the service life of the motor is realized.
Owner:CHINA CARBON MEDIGA (WUHAN) TECHNOLOGY CO LTD

Prediction method for interaction track of right-turn vehicle and pedestrian at intersection and computer equipment

The invention discloses an intersection right-turn vehicle and pedestrian interaction trajectory prediction method and computer equipment. The method comprises the following steps: acquiring a human-vehicle historical trajectory and initial state information in intersection right-turn vehicle and pedestrian interaction; a confrontation reinforcement learning structure is improved through KL regularization, and combined learning of a reward function and a strategy is achieved; a Nash Q-learning algorithm with a KL constraint is utilized to carry out joint optimization; and inputting the optimized strategy model into a behavior simulator, driving a generated prediction trajectory by an initial state, comparing the prediction trajectory with a real trajectory, and evaluating a prediction result based on a specified index. According to the method, the performance of the model in the aspects of prediction precision, behavior consistency and environment generalization ability can be effectively improved, the stability of a reward function and a strategy network is further improved, the generalization ability of the model in different traffic scenes subsequently is enhanced, and the learned strategy has better migration potential.
Owner:NANJING UNIV OF POSTS & TELECOMM

Injection molding production optimization method based on multi-agent cooperation

InactiveCN121836000AImprove collaborative optimization capabilitiesEnsure coordination and unityForecastingArtificial lifeOptimal decisionDecision strategy
The invention discloses an injection molding production optimization method based on multi-agent cooperation, and the method comprises the following steps: decomposing a plurality of targets in an injection molding production process, and constructing a layered multi-agent structure; collecting data in the injection molding production process in real time, and constructing a global expert behavior track and a local expert behavior track; based on the global target and each local target, obtaining initial parameters of a global reward function and a local reward function; adopting an inverse reinforcement learning method to obtain an optimal global reward function and an optimal local reward function; obtaining an optimal decision strategy of each agent through a hierarchical collaboration mechanism by each hierarchical agent; when detecting that a conflict exists between the targets, dynamically adjusting the reward weight of the reward function by adopting a conflict coordination mechanism; and continuously collecting new expert behavior data, and updating the optimal decision strategy of each agent. According to the invention, a layered multi-agent inverse reinforcement learning method is adopted, and multi-target dynamic collaborative optimization of injection molding production is realized.
Owner:HEBEI QUANYUN INTELLIGENT TECH CO LTD

Task type dialogue strategy learning method and system based on reinforcement learning

The invention belongs to the technical field of intelligent task-based dialogue, and particularly relates to a task-based dialogue strategy learning method and system based on reinforcement learning, which adopts a soft actor commentator algorithm to be combined with behavior cloning and maximum entropy inverse reinforcement learning to relieve a cold start problem. According to the method, maximum entropy inverse reinforcement learning is utilized, a reward value is calculated, and a user target in a dialogue is accurately deduced according to maximum log likelihood estimation of a human dialogue and a simulated dialogue; a potential reward function is extracted from a successful track through inverse reinforcement learning, manual reward design is replaced, accumulated reward maximization is pursued during strategy optimization, diversity and exploratory performance of the strategy are encouraged, and the dialogue strategy is prevented from being converged to a single mode too early.
Owner:QUFU NORMAL UNIV

Intelligent optimization method and system for enterprise management system based on digital twinning

The invention provides an enterprise management system intelligent optimization method and system based on digital twinning, and the method comprises the following steps: constructing an organization structure diagram with a path weight by taking a post as a node, taking a superior-subordinate relationship as a directed edge, and taking a KPI (Key Performance Indicator) target vector as a node attribute; constructing an instruction conduction path model through breadth-first traversal based on the organization structure chart to generate an effective conduction path set, and calculating a target propagation vector of each path; a post is regarded as a reinforcement learning agent, an optimal post behavior strategy is recovered through inverse reinforcement learning, and a target vector matrix of each post adjusted according to the strategy is obtained; constructing an anti-fact management strategy optimization instruction conduction path, and generating an optimized target matrix and path set; and comparing the optimized target matrix with actual execution data, and calculating a target deviation as an execution effect evaluation value. According to the method, the defects of an existing system in the aspects of multi-level management collaboration, behavior deviation modeling and strategic execution consistency guarantee are overcome.
Owner:GUANGZHOU SHUANGHE SOFTWARE CO LTD +4

Intelligent network connection vehicle social contact cooperative interaction control method and system in mixed traffic scene

The invention relates to the technical field of vehicle cooperative interaction control, in particular to an intelligent network connection vehicle social cooperative interaction control method and system in a mixed traffic scene, and the method comprises the steps: obtaining a spatial feature vector and a time sequence feature vector; obtaining a comprehensive feature vector through the spatial feature vector and the time sequence feature vector; performing linear combination according to the comprehensive feature vectors to obtain a reward function; obtaining an optimal weight vector through inverse reinforcement learning according to the reward function; the SVO angle value of the HDV is estimated through the optimal weight vector, and the SVO angle estimated value of the CAV is obtained; sVO parameters are optimized and dynamically adjusted through the Stackelberg game according to the SVO angle estimation value of the CAV, and optimal response parameters of the HDV are obtained; and according to the IDM model and the optimal response parameter of the HDV, simulating social perception and response of a human driver so as to adjust the driving behavior. According to the invention, the accuracy of vehicle cooperative interaction control is improved.
Owner:XIDIAN UNIV

A Hybridization and Optimization Method for High-Iron Strains in Auricularia auricula-judae Based on Big Data Analysis

This invention provides a method for optimizing the hybridization and selection of high-iron strains in black fungus based on big data analysis. It relates to the fields of big data analysis and black fungus variety breeding. The method includes: Step S1: Collecting basic data of strains, hybridization process data, culture environment parameters, and external reference data, and performing standardization and structuring processing; Step S2: Encoding and optimizing hybridization combinations using a biased random bond genetic algorithm, combined with a dual-elite evolutionary mechanism and dynamic parameter adjustment strategy to generate the optimal hybridization scheme; Step S3: Constructing a dynamic environmental model using a model-based hybrid inverse reinforcement learning algorithm, searching for the optimal culture strategy that maximizes iron content in a virtual environment; Step S4: Constructing a multi-index comprehensive scoring system based on the analytic hierarchy process (AHP), identifying abnormal samples and triggering model backtracking adjustments, outputting high-performance strains for the next round of breeding. This invention significantly improves breeding efficiency and the accuracy of strain selection.
Owner:JILIN AGRICULTURAL UNIV

Multimodal emotion fusion method based on inverse reinforcement learning

The embodiment of the invention relates to the field of multi-modal emotion analysis, and discloses a multi-modal emotion fusion method based on inverse reinforcement learning. Dynamic weighted fusion is carried out through modal features such as voice, facial expressions and texts, an optimal modal fusion strategy is derived from known emotion label data by adopting inverse reinforcement learning, and the contribution weight of each modal in emotion analysis is automatically adjusted, so that efficient, accurate and stable emotion recognition is realized. Through the technology of the invention, the sentiment analysis system can adaptively optimize modal fusion in various sentiment situations, the generalization ability of the sentiment analysis system in different scenes is significantly improved, the dependence on a large amount of annotated data is reduced, and the sentiment recognition precision is improved at the same time. The method has a wide application prospect and can be widely applied to the fields of intelligent customer service, emotion calculation and the like. The method and the device can be at least used for solving the problems of fixed modal weight and low emotion recognition accuracy in the prior art.
Owner:BITMAP3D TECH (SHANGHAI) CO LTD

Traffic scene closed site cluster simulation method and device, equipment and medium

The invention relates to a traffic scene closed site cluster simulation method and device, equipment and a medium. The method comprises the following steps: establishing a traffic participant game utility matrix; introducing the traffic participant game utility matrix into a maximum entropy inverse reinforcement learning model, and constructing a multi-traffic participant traffic behavior model; training the model through a traffic participant interaction data set to obtain an inverse reinforcement learning traffic participant dynamic game model; building a real object scene, a virtual simulation scene and a data interaction link; and performing traffic behavior simulation on the plurality of second traffic participants through an inverse reinforcement learning traffic participant dynamic game model in a virtual simulation scene by adopting a virtual-real combination method, and controlling action tracks of the plurality of first traffic participants through a data interaction link. According to the technical scheme of the invention, the method can achieve the vivid simulation of the interactive game traffic behaviors of a plurality of traffic participants in a real object scene through the target object, achieves the precise testing of an automatic driving system, and improves the performance of the automatic driving system.
Owner:CHONGQING UNIV OF TECH

Scheduling method and system matched with aviation luggage change information

The invention discloses a scheduling method and system matched with air luggage change information, and relates to the technical field of air transportation intelligent scheduling, and the method comprises the steps: collecting luggage change information, passenger travel data and airport equipment state data in an air transportation scene in real time, and outputting an original data stream; an airport operation digital twinborn body is constructed in the element universe, a collaborative transfer path is imported into a digital twinborn environment for space-time preview simulation, luggage flow conflict points are detected in real time, and a physical constraint conflict area coordinate set is marked; inverse solution is carried out on conflict events and traceability strategy defect features in rehearsal simulation through an inverse reinforcement learning algorithm, a reward function space is reconstructed, and gradient compensation parameters are output; and fusing the physical constraint conflict region coordinate set and the gradient compensation parameters, generating a three-dimensional space scheduling instruction set, and executing instruction issuing. Based on the space-time rehearsal simulation method, the security boundary of the luggage transfer path is accurately identified, and the physical constraint overrun risk is effectively reduced.
Owner:SHANGHAI PENGBANG IND CO LTD

Supply chain fund flow intelligent scheduling method based on reinforcement learning

The invention discloses a supply chain fund flow intelligent scheduling method based on reinforcement learning, and the method comprises the steps: S1, collecting multi-dimensional heterogeneous data, constructing a dynamic digital twin environment, and carrying out the fusion to generate a standardized input tensor; s2, aggregating node features by using a heterogeneous graph attention network, and outputting a high-dimensional state feature graph; s3, analyzing market fluctuation and risk through a high-level agent, and outputting a global scheduling guidance strategy; s4, outputting a refined scheduling instruction set by using a DPPO algorithm in combination with tail risk minimization; s5, based on the improved MAPPO algorithm, noise is injected to cover privacy, and Nash equilibrium optimization is executed through inverse reinforcement learning; and S6, fund allocation is executed, and online iterative optimization is carried out. According to the invention, supply chain fund scheduling efficiency and security are improved, and individual and collective benefits are effectively balanced.
Owner:SHENGWEI (TIANJIN) TECHNOLOGY DEVELOPMENT CO LTD

Unmanned system multi-intention recognition method and system based on inverse reinforcement learning

The invention discloses an unmanned system multi-intention recognition method and system based on inverse reinforcement learning, and belongs to the technical field of artificial intelligence and autonomous decision making. According to the system, a deep multi-intention inverse reinforcement learning framework is provided for solving the problems of multi-modality, fuzziness and dynamics of unmanned system decision-making intentions in a complex dynamic environment. By designing a multi-scale intention encoder based on Transform, state-intention combined dynamic reasoning is realized; constructing an end-to-end intention perception reward network, and adaptively fusing multi-intention features through an attention mechanism; and proposing an expectation maximization optimization strategy to realize collaborative optimization of intention reasoning and strategy learning. The system can automatically separate and identify various potential intentions from expert demonstration of mixed intentions, and learn corresponding reward functions and decision strategies. According to the method, the intention recognition accuracy and the strategy reproduction capability in a multi-intention scene are remarkably improved, and an effective solution is provided for behavior understanding and autonomous decision making of an unmanned system in a complex environment.
Owner:GUILIN UNIV OF ELECTRONIC TECH +1

Decision model training method and device and driving decision method and device

The invention discloses a decision model training method and device and a driving decision method and device, and the method comprises the steps: inputting each first candidate decision behavior corresponding to a first traffic scene into a to-be-trained decision model, and obtaining the decision probability of each first candidate decision behavior; based on the decision probability of each first candidate decision behavior, performing inverse reinforcement learning on each first candidate decision behavior to obtain an inverse reinforcement learning loss value of the first traffic scene; performing imitation learning on the first optimal decision behavior with the maximum decision probability in the first candidate decision behaviors to obtain an imitation learning loss value of the first traffic scene; based on the inverse reinforcement weight and the imitation weight, performing weighted fusion on the inverse reinforcement learning loss value and the imitation learning loss value to obtain a weighted loss value of the first traffic scene; and based on the weighted loss value, optimizing model parameters of the decision model to be trained to obtain a trained decision model. The reliability of the decision model can be improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Parameter optimization method, device and equipment for oil reservoir development and storage medium

The invention discloses a parameter optimization method and device for oil reservoir development, equipment and a storage medium, and relates to the technical field of oil reservoir development. Comprising the steps that oil field geological parameters, well location and operation parameters, injection parameters and economic evaluation parameters are obtained and input into a numerical simulator to complete initialization, and oil reservoir state change data and development effect indexes are calculated; generating a first data set representing the corresponding relationship among the parameters, the oil reservoir state and the development effect by combining a particle swarm optimization algorithm; a second data set containing multiple expert tracks is obtained through screening and data enhancement processing, and the expert tracks reflect time sequence evolution of the oil reservoir state under the action of the injection parameters; inputting the expert track into a preset inverse reinforcement learning model, and screening a target expert track according to the strategy network update amplitude and the discriminator probability value; and finally, optimizing the input parameters based on the target expert trajectory to obtain optimal parameters. Intelligent optimization of oil reservoir parameters is realized, optimization efficiency and suitability are improved, and development benefits are guaranteed.
Owner:XI'AN PETROLEUM UNIVERSITY

Task-oriented dialogue system reward function optimization method and system

The invention discloses a task-oriented dialogue system reward function optimization method and system, and belongs to the technical field of natural language processing. The method comprises the following steps: acquiring an expert dialogue track from a dialogue system data set, extracting a state-action-reward triple, training an initial reward function by using a maximum entropy inverse reinforcement learning framework, and initializing a strategy network; an Actor-Critic reinforcement learning algorithm is adopted to train a strategy network, and a suboptimal trajectory is collected to dynamically update a reward function; and taking the dynamic reward function as a unified evaluation signal, and optimizing the strategy network to form a dialogue strategy model. Through dynamic reward function optimization, manual rule dependence is reduced, the generalization ability, the task completion rate and the stability of a dialogue system are improved, and the method is suitable for complex dialogue scenes in multiple fields.
Owner:XIAN UNIV OF POSTS & TELECOMM

Security unmanned vehicle path planning method fusing tactical risk map and game prediction

The invention discloses a security and protection unmanned vehicle path planning method fusing a tactical risk map and game prediction, and relates to the technical field of intelligent security and protection, and the method comprises the steps: receiving a task path issued by a remote control security and protection unmanned vehicle, carrying out the semantic perception of an environment, and constructing a static tactical risk map; acquiring real-time observation track sequences of potential opponents, inferring target intentions of the opponents on line through a Bayesian inverse reinforcement learning BIRL model, and generating a high-risk region thermodynamic diagram; generating a speed and angular speed control instruction through a path tracking controller based on the balanced deterrent path; and executing a speed and angular speed control instruction, collecting a real-time observation track sequence of a potential opponent, and performing online fine adjustment of a Bayesian inverse reinforcement learning BIRL model to generate an optimized security unmanned vehicle path. According to the invention, the unmanned vehicle can continuously generate an optimized path with both safety and tactical initiative in a dynamic confrontation environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

An intent-driven communication service quality dynamic guarantee method

This invention discloses an intelligent traffic prediction and scheduling method based on user behavior data, comprising: acquiring the user's actual driving trajectory and matching it to a digital road network; extracting expert trajectory data containing state-action sequences and contextual features; constructing a generative adversarial learning model and using the expert trajectory data for adversarial training, wherein the generator learns to imitate the user's driving behavior, and the discriminator learns to distinguish between the real user trajectory and the trajectory simulated by the generator; using the output of the discriminator after training convergence, assigning a personalized cost value of the user's subjective preference to each state-action pair in the digital road network; based on the constructed personalized cost network, using the A* path search algorithm to plan an optimal personalized driving route for the user; this invention, through inverse reinforcement learning, can automatically learn the user's inherent preferences from the user's behavior, and the planned route is more in line with the user's true intention, improving user satisfaction and system adoption rate.
Owner:WUHAN XINGCHEN WENHUI TECH CO LTD

Building energy management system for source network load storage and utilization

The invention relates to the technical field of energy management, and discloses a source network load storage and utilization oriented building energy management system, which comprises at least one intelligent sensing terminal deployed in a user physical space, and a cloud data processing platform running on a cloud infrastructure, and the intelligent sensing terminal and the cloud data processing platform perform real-time interaction through a safe and encrypted bidirectional data communication link to form a complete closed loop of sensing, analysis, decision making, control and feedback. According to the invention, through multi-dimensional data acquisition of the intelligent sensing terminal and in combination with a non-intrusive load decomposition algorithm of the edge calculation unit, the total load can be decomposed to each energy consumption device in real time, and an independent operation power curve is generated; the cloud digital behavior twinborn modeling module further predicts a baseline load through a long short-term memory network, quantifies the comfort level and economic preference of the user through maximum entropy inverse reinforcement learning, and achieves deep analysis of the power consumption behavior of the user.
Owner:ZHEJIANG PROVINCE ELECTRIC POWER FUEL CORP +1

An unmanned ship dynamic collision avoidance decision method and system based on a diffusion model and adversarial inverse reinforcement learning

The application provides an unmanned ship dynamic collision avoidance decision method and system based on a diffusion model and an adversarial inverse reinforcement learning. The method comprises the following steps: 1, collecting expert collision avoidance data; 2, training a diffusion model; 3, performing data enhancement; 4, building a training environment; 5, constructing a generator and a discriminator; 6, updating the discriminator; 7, updating the generator; and 8, completing the training. The diffusion model is trained by using limited expert sample data, additional collision avoidance data is generated by using the trained diffusion model, and the additional data and the original expert data jointly participate in the training process of the adversarial inverse reinforcement learning, so that the problem of limited expert collision avoidance data can be effectively solved. Furthermore, an engineering application plan is provided for the proposed algorithm, the algorithm is verified, and a systematic solution is formed by combining theory with practice.
Owner:JIMEI UNIV

Breast cancer MRI image enhancement and focus automatic identification method based on attention mechanism

The invention discloses a breast cancer MRI image enhancement and focus automatic identification method based on an attention mechanism, and relates to the technical field of data processing and computer-aided diagnosis. The method comprises the following steps: acquiring a multi-time-point mammary gland MRI image sample and performing preprocessing; performing random window enhancement and random channel exchangeable enhancement on the preprocessed sample, inputting the enhanced image sample into an improved convolutional neural network, and performing deep metric learning loss function optimization based on proxy NCA to extract focus embedding features; according to the focus embedding features, constructing a moment matching game framework with focus state distribution marked by experts as a target; and according to a moment matching game framework, adopting an inverse reinforcement learning strategy without reinforcement learning, carrying out iterative optimization on parameters of the lesion recognition model through a regret-moment-free matching algorithm, carrying out lesion property judgment according to the optimized model, and outputting a judgment result and a corresponding confidence score. Therefore, MRI image adaptive enhancement and accurate focus automatic identification are realized.
Owner:THE AFFILIATED HOSPITAL OF XUZHOU MEDICAL UNIV +1

Unmanned bus trajectory optimization method for entering and leaving station based on inverse reinforcement learning

The application relates to an unmanned bus station entry and exit trajectory optimization method based on inverse reinforcement learning, which comprises the following steps: analyzing the constraints of bus station entry and exit behaviors, defining state space variables of the station entry and exit behaviors, traversing the state space variables, and generating a candidate behavior trajectory curve; a maximum entropy inverse reinforcement learning model is constructed, a reward function of the model is composed of a group of base functions, the base functions are designed from the driving danger degree, the driving efficiency, the violation punishment, the influence on surrounding vehicles and the comfort degree, so that the bus station entry and exit behaviors are safe, efficient and comfortable, and meanwhile, the coefficients of the base functions are determined through inverse reinforcement learning, so that the behaviors are humanized; the maximum entropy inverse reinforcement learning model is trained based on human driver demonstration trajectories; and the candidate behavior trajectory curve is evaluated based on the trained model, and the optimal trajectory is selected. Compared with the prior art, the application has the advantages of high safety, high efficiency and good comfort.
Owner:TONGJI UNIV

An evaluation method and system for reverse engineering of reinforcement learning rewards

The present invention belongs to the field of reinforcement learning security technology, and specifically relates to an evaluation method for reverse engineering of reinforcement learning rewards. The method comprises: calculating the standard mutual information between real reward clusters obtained in real time and reward clusters obtained by inverse reinforcement learning; generating a standard mutual information sequence based on a pre-set range of cluster numbers, a cluster sorting range, and a contamination trajectory ratio range; inputting the standard mutual information into a pre-established fourth-order tensor model to obtain an evaluation result; and evaluating the reward clusters obtained by inverse reinforcement learning based on the evaluation result.
Owner:BEIJING YUNLIANG SHUMENG TECH CO LTD

Multi-time scale microgrid economic dispatching method based on relative entropy inverse reinforcement learning

The invention discloses a multi-time-scale microgrid economic dispatching method based on relative entropy inverse reinforcement learning, and the method comprises the steps: obtaining a day-ahead dispatching plan and a tie line dispatching plan according to a day-ahead prediction information set of a microgrid; superposing disturbance on the day-ahead scheduling plan to carry out dimension expansion, and generating an expert strategy for inverse reinforcement learning; according to the expert strategy, a reward function is constructed by adopting a relative entropy inverse reinforcement learning algorithm, and the reward function is obtained by minimizing the relative entropy of trajectory distribution and solving a corresponding Lagrange duality function; according to the reward function, adopting a depth deterministic strategy gradient algorithm to train an intra-day scheduling agent to obtain a trained intra-day scheduling model; obtaining ultra-short-term prediction information based on the tie line scheduling plan; and inputting the ultra-short-term prediction information into an intra-day scheduling model to obtain an intra-day scheduling strategy.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Chemical text attribute extraction method and system based on inverse reinforcement learning

The invention discloses a chemical text attribute extraction method and system based on inverse reinforcement learning, and the method comprises the steps: reconstructing a chemical text attribute extraction task into a generative question and answer task, converting an input text and a target attribute into a question and answer pair, and extracting attribute information through the semantic understanding and generation capability of a question and answer model; a BioBART model is selected as a basic model, and fine tuning is carried out on a public biomedical question and answer data set, so that the model preliminarily adapts to a question and answer task; constructing a reward model comprising a plurality of sub-reward components; and performing model training by using maximum entropy inverse reinforcement learning, and alternately performing reward model updating and question and answer model updating. According to the method, the attribute extraction task is converted into the natural language question and answer task, and the BioBART model and inverse reinforcement learning are utilized to optimize the multi-target award function, so that the understanding and generation capability of the model on complex scientific texts is remarkably improved, implicit attributes in chemical texts can be effectively extracted, and powerful support is provided for customs to quickly identify suspicious substances.
Owner:JIANGNAN UNIV

Gait control method and system for humanoid robot based on imitation and inverse reinforcement learning

The application discloses a kind of humanoid robot gait control method and system based on imitation and inverse reinforcement learning, collect human or simulation robot walking data to construct high-quality training dataset, then through graph convolution structure, the similarity of robot and expert data is calculated.Based on task target design multidimensional linear reward function, four kinds of indexes such as stability, gait periodicity, trajectory similarity and energy efficiency are comprehensively considered, and multi-objective optimization is realized through weight distribution.Neural network is used to construct strategy model, and training is carried out in combination with proximal policy optimization algorithm, and the update amplitude is constrained using strategy ratio clipping technology to improve learning stability.Maximal entropy principle is introduced to dynamically update reward function, and the adaptability to expert strategy is enhanced.Finally, simulation verification is carried out through MuJoCo physical engine, and visual analysis and performance evaluation of high human-like gait are realized.The application solves the problems of robot gait adaptability and robustness in complex dynamic environment.
Owner:WUHAN UNIV

Low-protein diversified daily ration design method and system based on multi-source data fusion

The invention relates to the technical field of livestock breeding and feed nutrition, and discloses a low-protein diversified daily ration design method and system based on multi-source data fusion, and the method comprises the steps: constructing a formula-environment-effect directed causal diagram through a PC algorithm, estimating the effect deviation contribution proportion of an environment factor and a formula factor through anti-factual reasoning, and calculating the effect deviation contribution proportion of the environment factor and the formula factor; and inferring the target weight preference of the farm by using inverse reinforcement learning, and carrying out conditional weight updating according to a causal attribution result. Real reasons of growth effect deviation can be distinguished, effect deviation errors caused by environmental factors are prevented from being attributed to improper formula setting, and the accuracy and stability of formula recommendation are improved.
Owner:LIAONING WELLHOPE AGRI TECH

Inverse reinforcement learning automatic parking method and system based on driving habits

The invention provides an inverse reinforcement learning automatic parking method and system based on driving habits. The method comprises the steps that an expert parking state data set is constructed according to expert tracks under typical parking working conditions; constructing a maximum entropy probability model based on an expert track according to a parameterized reward model based on linear feature combination; by utilizing a maximum entropy inverse reinforcement learning algorithm, optimizing and solving a weight vector of a reward function in the maximum entropy probability model by maximizing a log-likelihood function of an expert track, and obtaining an optimal reward function; based on the optimal reward function, adopting a depth deterministic strategy gradient algorithm to train a strategy network so as to obtain an optimal parking strategy; and according to the optimal parking strategy, a two-stage parking control architecture is adopted to execute parking control on the vehicle. According to the method, the learning efficiency and convergence stability of automatic parking are remarkably improved, the smoothness of the parking track and the accuracy of final parking are guaranteed, and the reliability and riding comfort of the automatic driving vehicle in the parking process are improved.
Owner:NANCHANG AUTOMOTIVE INST OF INTELLIGENCE & NEW ENERGY

Humanoid robot gait control method and system based on imitation and inverse reinforcement learning

The invention discloses a humanoid robot gait control method and system based on simulation and inverse reinforcement learning, and the method comprises the steps: collecting the walking data of a human body or a simulation robot to construct a high-quality training data set, and calculating the similarity between the robot and expert data through a graph convolution structure; a multi-dimensional linear reward function is designed based on a task target, four types of indexes including stability, gait periodicity, trajectory similarity and energy efficiency are integrated, and multi-target optimization is achieved through weight distribution. A neural network is adopted to construct a strategy model, training is performed in combination with a near-end strategy optimization algorithm, and a strategy ratio cutting technology is utilized to constrain an update amplitude so as to improve learning stability. And meanwhile, a maximum entropy principle is introduced to dynamically update a reward function, so that the adaptability to expert strategies is enhanced. Finally, simulation verification is conducted through a MuJoCo physical engine, and visual analysis and performance evaluation of the highly-simulated human gait are achieved. According to the method, the problem that the robot gait self-adaptability and robustness are insufficient in a complex dynamic environment is solved.
Owner:WUHAN UNIV

Film and television film production actuarial and risk intelligent control system

The invention relates to the technical field of film and television investment intelligent decision, and discloses a film and television production actuarial and risk intelligent control system, which comprises seven modules, namely a data annotation module, a feature extraction and conversion module, a symbol reasoning and evaluation module, a scheme generation module, a monitoring and early warning module, a decision optimization verification module and a knowledge updating module. According to the method, semantic annotation is performed through a multi-channel adaptive method, multi-modal feature extraction is performed, a feature-to-symbol conversion mechanism is constructed, multi-level cognitive reasoning is performed by adopting an OpenNARS non-axiom reasoning engine, a personalized investment scheme is generated by adjusting cognitive parameter configuration, and a risk early warning report is obtained through real-time data monitoring and risk assessment. Decision optimization and cross validation are carried out through an imitation learning method and an inverse reinforcement learning method. According to the method, the precision, individuation and interpretability of film and television project investment decision can be realized, the return on investment is improved, and the project risk is reduced.
Owner:SHENZHEN VISION CULTURAL IND INVESTMENT MANAGEMENT CO LTD