Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

418 results about "Imitation learning" patented technology

Imitation learning is learning by imitation in which an individual observes an arbitrary behavior of a demonstrator and replicates that behavior.

Multi-modal fusion and reinforcement learning collaborative retrieval enhancement generation method and system

The invention relates to the technical field of information retrieval, and discloses a multi-modal fusion and reinforcement learning collaborative retrieval enhancement generation method and system. The method comprises the following steps: receiving an original query input by a user, and generating a sub-query based on a large language model in combination with a multi-modal context of a current iteration step; forming a current state in combination with the sub-query and the multi-modal context, modeling a retrieval enhancement generation task as a Markov decision process, and adaptively selecting an optimal action from a predefined action set in the current state by utilizing a large language model according to a decision strategy; executing a corresponding multi-modal retrieval operation according to the optimal action, fusing the obtained multi-modal information, generating an intermediate answer or a final answer of the sub-query, and updating a multi-modal context by using the intermediate answer; off-line training optimization is carried out on the large language model through imitation learning and a calibration chain, and decision strategies and sub-queries are inferred online through the model after fine adjustment. According to the invention, more efficient and accurate complex query processing is realized.
Owner:DATA SPACE RES INST

Virtual power plant scheduling method based on large language model and deep reinforcement learning

The invention discloses a virtual power plant scheduling method based on a large language model and deep reinforcement learning, and belongs to the technical field of virtual power plant scheduling. Comprising the following steps: constructing a virtual power plant multi-agent cloud edge collaborative scheduling framework based on large language model driving; predicting wind power, photovoltaic power and load power based on a large language model; constructing a mathematical model of virtual power plant optimization scheduling; converting the virtual power plant optimization scheduling model into a Markov game process in combination with a large language model; performing initialization training on the strategy network of the edge layer intelligent agent by adopting imitation learning to obtain a pre-trained edge layer intelligent agent strategy network; and based on the pre-training strategy network of the boundary layer intelligent agent, combining with a large language model and adopting an improved multi-agent near-end strategy optimization algorithm to solve a scheduling strategy.
Owner:NANJING UNIV OF POSTS & TELECOMM

Robot continuous imitation learning method and system based on double diffusion model

The invention belongs to the technical field of robot learning, and particularly discloses a robot continuous imitation learning method and system based on a double diffusion model. Comprising the following steps: generating old task data in a submerged space by utilizing a pre-trained latent diffusion model, generating a model through guidance of a task identifier in combination with conditional diffusion, generating an image and state data related to an old task, modeling to construct a diffusion model, realizing multi-modal mapping from a state to an action, and generating a diversified action sequence; high-quality images and state data are generated through a pre-trained variational auto-encoder and a conditional diffusion model, and the consistency of the generated data in time and the accuracy of global semantics are ensured through time sequence continuity constraint and semantic consistency constraint; and performing feedback optimization training on the latent diffusion model. According to the method, the loss weight is dynamically adjusted, and the model is ensured to show good generalization ability and stability in a complex task in combination with collaborative training of an imitation learning strategy and a generator.
Owner:HUAZHONG UNIV OF SCI & TECH

Mechanical arm grabbing method and system based on multi-modal information fusion

The invention provides a mechanical arm grabbing method and system based on multi-modal information fusion, and belongs to the technical field of robot intelligent control. Comprising the steps that the conversion relation between a camera coordinate system and a mechanical arm base coordinate system is established through camera calibration, a deep learning neural network is used for conducting grabbing pose estimation on an obtained RGB-D image, and multiple candidate grabbing poses are determined; analyzing a natural language instruction input by a user based on a multi-modal large model, and recognizing a target object region from the RGB-D image by combining a target detection and image segmentation technology; based on the obtained candidate grabbing poses and the target object area, an optimal grabbing pose is screened through a scoring mechanism and mapped to a mechanical arm base coordinate system; and then a dynamic grabbing path is generated by adopting an imitation learning algorithm, and the mechanical arm is controlled to execute grabbing operation. Through multi-modal semantic understanding, accurate grabbing of the mechanical arm in a complex environment can be achieved.
Owner:SHANDONG UNIV

Automatic driving method and device based on multi-dimensional reward function

The embodiment of the invention provides an automatic driving method and device based on a multi-dimensional reward function, and the method and device achieve the control of a driving motion through the combination of an imitation learning framework and a reinforcement learning framework, collection of environment information through a plurality of cameras, and construction of a strategy generation network and a discriminator network. A multi-dimensional reward function model is designed, behaviors such as red light running, line pressing, lane departure and collision are detected and evaluated in real time, and a driving behavior reward and punishment matrix is constructed. Based on an actor evaluation network architecture, environment information and a navigation instruction are input into an actor network to generate an optimal driving action, and the action value is evaluated through the evaluation network to realize dynamic parameter optimization. According to the method, the defects of the traditional technology in the aspects of driving behavior evaluation, action value judgment and the like are effectively overcome, and the safety and reliability of the automatic driving system are remarkably improved.
Owner:ZHEJIANG WUWEN ZHIXING TECHNOLOGY CO LTD

Robot time sequence imitation learning method and system based on Mama coding complete history

The invention belongs to the related technical field of artificial intelligence, and discloses a robot time sequence imitation learning method and system based on a Mama coding complete history, and the method comprises the steps: receiving a multi-modal observation sequence in a task execution process of a robot, the multi-modal observation sequence comprising observation data of at least one sensor; processing the multi-modal observation sequence by using a sequence processing module based on a state space model, and updating a time sequence output of complete historical information of one code up to the current time step at each time step; and predicting the next step or a series of future actions of the robot based on the time sequence output of the current time step so as to control the robot to simulate. The time sequence processing module based on the state space model is utilized to process and encode the complete observation history in the task execution process of the robot, so that a non-Markov decision-making imitation learning method is realized, and the learning efficiency and the execution success rate of the robot in a complex and state-dependent long time sequence operation task are improved.
Owner:HUAZHONG UNIV OF SCI & TECH

Visual language navigation method based on memory driving

The invention discloses a visual language navigation method based on memory driving, and the method comprises the following steps: S1, constructing an MDQT model which comprises a panoramic encoder, a text embedding layer, a Q-Former, an action prediction module and a memory updating module; s2, training an MDQT model by using three pre-training tasks of image-concerned mask language modeling, instruction track matching and instruction track contrast learning; and S3, performing fine tuning on the MDQT model by using imitation learning and reinforcement learning. According to the invention, a learnable memory vector with a fixed length is used to encode historical information. In each navigation step, the memory vector interacts with the extracted panoramic information and instruction information, and visual information most relevant to the current instruction information is extracted according to the current memory state of the robot for decision making. The memory of the robot to the historical navigation steps is effectively maintained under limited resources, and the navigation success rate and efficiency of the robot are improved.
Owner:NANJING UNIV

Machine learning for video game help sessions

The disclosed concepts relate to training a machine learning model to provide help sessions during a video game. For instance, prior video game data from help sessions provided by human users can be filtered to obtain training data. Then, a machine learning model can be trained using approaches such as imitation learning, reinforcement learning, and / or tuning of a generative model to perform help sessions. Then, the trained machine learning model can be employed at inference time to provide help sessions to video game players.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Robot control method based on tactile prediction pre-training

A robot control method based on tactile prediction pre-training comprises the following steps: acquiring and generating a human playing data set consisting of three-channel image tensors in an offline stage, and training a constructed conditional diffusion model comprising a tactile encoder, a tactile decoder and an action and visual encoder; in the online stage, the trained conditional diffusion model is integrated into a standard imitation learning strategy network, and an action instruction of the robot is generated according to the state of the robot, the current visual features and the tactile feature vectors extracted by the imitation learning strategy network. According to the method, a specific agent task is completed by training a deep neural network model, that is, a future tactile signal sequence is predicted according to historical information and future action intentions; the model is enabled to characterize generic haptic features contacting physical dynamic laws for further migration into downstream robot control tasks.
Owner:SHANGHAI JIAOTONG UNIV

New energy power grid look-ahead scheduling method and device

The invention provides a new energy power grid prospective scheduling method and device, and relates to the technical field of electric power system intraday economic scheduling. The method comprises the following steps: firstly, constructing an opportunity constraint optimal power flow model of prospective scheduling, analyzing the influence of new energy uncertainty on opportunity constraint, establishing a constrained Markov decision process of prospective scheduling, and then utilizing a risk evaluator network fitting risk function probability distribution and an actuator network considering extreme scene performance to determine the opportunity constraint optimal power flow model of prospective scheduling. The processing capability of the intelligent agent on a prospective scheduling scene containing a new energy extreme climbing event is enhanced; and finally, the training of the intelligent agent is accelerated by utilizing an imitation learning technology in a power grid prospective scheduling off-line simulation environment. According to the method, the solving speed and the strategy robustness and safety of the double-layer robust optimization model of the look-ahead scheduling can be considered.
Owner:WUHAN UNIV

Zero-sample cross-domain imitation learning method based on three-dimensional semantic point cloud and related equipment

The invention discloses a zero-sample cross-domain imitation learning method based on three-dimensional semantic point cloud and related equipment, and the method comprises the steps: constructing a training scene in a physical simulation engine, randomly generating an operation object pose, and enabling a robot to act to complete an operation task; an original image is observed, semantic segmentation privilege information is obtained by utilizing a simulation environment, an operation object is segmented, and semantic point cloud of the operation object is obtained by combining a depth map generated by the simulation environment; synchronously recording robot joint angles, and constructing a training data set; inputting the semantic point cloud of the operation object and the training data set into an imitation learning model, performing action blocking, extracting features, performing feature splicing, generating a predicted action sequence, and completing model training; deploying the trained model to a real environment, obtaining the semantic point cloud of an operation object, reading the current joint angle of the robot, sending the angle to the model for reasoning, outputting a robot action sequence, and completing a specified task. According to the method, the generalization ability of the strategy can be improved while the operation precision is ensured.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Adaptive pervasive edge computing task unloading method based on imitation learning

The invention relates to the technical field of edge computing, and discloses a self-adaptive pervasive edge computing task unloading method based on imitation learning, and the method comprises the steps: firstly constructing a task unloading strategy imitation model and an environment adaptability evaluation model, the former is used for learning and simulating a historical task unloading strategy, and the latter is used for evaluating the influence of environment change on the strategy; a deep reinforcement learning algorithm is introduced to train a simulation model, and real-time feedback of an environment evaluation model is combined to continuously optimize and adjust an unloading strategy so as to generate a highly adaptive task unloading decision. According to the method, the task unloading instruction can be intelligently generated and executed according to the real-time edge computing environment state and the task characteristics, the dynamic change of the environment is effectively coped with, and efficient utilization of resources and rapid execution of tasks are achieved. Compared with a traditional method depending on a static rule or a predefined strategy, the task unloading flexibility is remarkably improved, and powerful support is provided for intelligent application in the edge computing environment.
Owner:HENAN UNIV OF ANIMAL HUSBANDRY & ECONOMY

Diffusion-reward adversarial imitation learning

Imitation learning, or artificial intelligence-based learning from demonstration, aims to acquire an agent policy by observing and mimicking the behavior demonstrated in expert demonstrations. Imitation learning can be used to generate reliable and robust learned policies in a variety of tasks involving sequential decision-making, such as autonomous driving and robotics tasks. However, current imitation learning solutions are limited in their ability to generalize states or goals unseen from the expert's demonstrations. The present disclosure integrates a diffusion model into generative adversarial imitation learning, which, in terms of prior solutions, can provide superior performance in generalizing to states or goals unseen from the expert's demonstrations, provide data efficiency for varying the amounts of available expert data, and capture more robust and smoother rewards.
Owner:NVIDIA CORP

Personified virtual test scene multi-background vehicle collaborative decision-making method

The invention provides an anthropomorphic virtual test scene multi-background vehicle collaborative decision-making method, which comprises the following steps of: constructing a dynamic traffic map of multiple vehicles based on a graph structure, and constructing a parameter-shared generative adversarial imitation learning framework, inputting the node features optimized by the graph attention network to a reinforcement learning network to generate a serialized control instruction, and comparing the generated control instruction with a state-action pair of expert data by using a discriminator sharing scoring parameters to iteratively optimize a control instruction generator; potential factor variables are introduced to represent driving styles, and strategy personification is enhanced in combination with explicit rule punishment and implicit confrontation awards; and a highly anthropomorphic multi-background vehicle sequence cooperative control instruction is generated through multiple rounds of iteration. According to the method and the system, the adversarial imitation learning is iteratively optimized and generated by utilizing potential factor variables and reward enhanced imitation learning, and a decision with diversity, compliance and personification is generated, so that the reliability of an automatic driving simulation test is improved in multiple aspects.
Owner:CHANGAN UNIV

Multi-transportation equipment cooperative motion control method based on multi-stage mixed learning and storage medium

The invention relates to the technical field of automatic driving and intelligent control, in particular to a multi-transportation-equipment cooperative motion control method based on multi-stage mixed learning and a storage medium, and the method comprises the steps: generating expert demonstration data through employing a single-vehicle motion control algorithm based on an expert rule, performing initialization training on the strategy network through online iterative supervised learning to obtain a pre-training strategy model; the multi-vehicle cooperative motion control method comprises the following steps: selecting a multi-vehicle cooperative motion model, loading parameters of the model into an Actor strategy network of multi-agent reinforcement learning, performing interactive training on a plurality of agents in a simulation environment by adopting a centralized training and decentralized execution normal form, and performing online iterative optimization on the strategy based on a composite reward function and generalized advantage estimation to obtain a multi-vehicle cooperative motion control strategy. According to the method, the complementary advantages of imitation learning and reinforcement learning are exerted, the training efficiency, the strategy performance and the collaborative operation capability and robustness of the system in a complex scene are improved, and the method can be directly applied to collaborative scheduling and control of transportation equipment groups in scenes such as surface mines and ports.
Owner:SHANGHAI JIAOTONG UNIV

Robot gait training method and system based on reinforcement learning

The invention discloses a robot gait training method and system based on reinforcement learning, and relates to the technical field of electric vehicle charging, and the method comprises the steps: carrying out the simulation learning initialization of a strategy based on a reference video; selecting a plurality of evaluation indexes to design a dynamic reward function so as to guide the initial strategy network to optimize the gait performance; through an environment difficulty scheduling mechanism, the training environment difficulty is automatically adjusted according to strategy performance; performing optimization training on the strategy network by using a PPO algorithm; and judging the optimized gait strategy by constructing a cognitive load index, and outputting a final gait strategy. According to the method, imitation learning, self-adaptive reward modeling, multi-stage scheduling, reinforcement learning optimization and cognitive feedback are organically fused, and a robot gait training framework with generalization ability, stability and social adaptability is constructed. According to the method, the naturalness, the stability and the man-machine friendliness of the gait of the robot can be remarkably improved under complex terrains and interaction scenes, and the method has wide application prospects and popularization value.
Owner:NANJING KANGLONGWEI TECH IND CO LTD

Double-arm robot imitation learning method and system based on action blocking and force sensing

The invention discloses a double-arm robot imitation learning method and system based on action blocking and force perception, and the method comprises the steps: setting the pose of the tail end of a mechanical arm as an action, carrying out the action blocking, and collecting the action blocking and multi-mode observation information in a process that an expert operates robot teaching to complete a specified task; inputting the action blocks and the multi-modal observation information into a pre-established imitation learning model, constructing a mapping relationship between the action blocks and the multi-modal observation information, generating a predicted action sequence, and completing imitation learning model training; and deploying the trained imitation learning model to a real environment, obtaining real-time multi-modal observation information, dynamically generating corresponding action blocks according to the multi-modal observation information, outputting a robot action sequence according to the action blocks, and completing a specified task. According to the method, the accuracy of the position and force in the operation process can be ensured, fine and smooth operation skills can be learned only through a small amount of demonstration, and the task requirements of an unstructured environment and diversity of operation objects are met.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Mechanical arm trajectory planning method and device based on generative adversarial imitation learning

The invention discloses a mechanical arm trajectory planning method and device based on generative adversarial imitation learning, and relates to the technical field of mechanical arms. The method comprises the following steps: acquiring joint state data of the mechanical arm; inputting the mechanical arm joint state data into a trained improved generative adversarial imitation learning model to obtain a mechanical arm action, and circularly executing data acquisition-input of the improved generative adversarial imitation learning model to finally obtain a mechanical arm track; wherein in the improved generative adversarial imitation learning model, the reward function is a dynamic reward function, and parameters in the dynamic reward function are dynamically adjusted according to the task completion degree. According to the technical scheme, the learning efficiency can be improved. Compared with the prior art, the mechanical arm trajectory planning method has the advantages that the stability is obviously improved, expert trajectories can be better simulated, the success rate and accuracy of mechanical arm trajectory planning based on generative adversarial imitation learning are greatly improved, and a solid foundation is laid for practical application of mechanical arms in the fields of material synthesis and the like.
Owner:UNIV OF SCI & TECH BEIJING

Unmanned logistics vehicle-oriented dedicated end-to-end planning control method and system

The invention relates to the technical field of automatic driving, in particular to a special end-to-end planning control method and system for an unmanned logistics vehicle, and the method employs an end-to-end design, integrates the functions of sensing, decision making, planning, control and the like into a unified model, simplifies the system architecture, reduces the development complexity, and improves the development efficiency. And the real-time performance and the consistency of the system are improved. According to the invention, two methods of imitation learning and reinforcement learning are innovatively combined. Simulation learning is used for model initialization, so that the model has a safe and mild driving style quickly; and reinforcement learning is used for strategy optimization, so that the model can exceed the level of a human driver, and a better driving strategy is found. The collaborative optimization mode gives consideration to both training efficiency and performance. The method is specially designed for the unmanned logistics vehicle, and the particularity of the unmanned logistics vehicle in an urban distribution scene is fully considered, so that the unmanned logistics vehicle can better adapt to the application requirements of the unmanned logistics vehicle.
Owner:HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD

Nuclear power plant reactor core optimization power control method based on imitation learning

The invention discloses a nuclear power plant reactor core optimization power control method based on imitation learning. A reactor core optimization power control model comprising a plurality of intelligent agents and an average layer is constructed based on the imitation learning method; and averaging the output results of the trained agents as a final control result. According to the invention, nuclear power plant reactor core optimization power control is developed based on an imitation learning method, an intelligent agent is trained in a supervised manner to learn a nuclear power plant power control system and decision logic of an operator, and serious consequences caused by direct interaction between an immature intelligent agent and a power plant are avoided while real data of the power plant are effectively utilized. A plurality of agents are integrated by adopting an integrated learning method, so that the model can be suitable for multiple working conditions of peak regulation, normal start-stop and emergency shut-down of a plurality of nuclear power plants. Each agent is optimized based on the Bayesian method in the training process, the optimal hyper-parameter is searched, and the network training efficiency and the fault diagnosis precision are improved.
Owner:SHANGHAI JIAOTONG UNIV

Training generative model to generate predicted rewards and / or use thereof in reinforcement learning

Implementations relate to training a generative model. Some of those implementations include performing offline supervised fine-tuning (SFT) of a base generative model using an imitation learning dataset of successful task episodes. The SFT training utilizes both a behavioral cloning loss function, which compares offline predictions to ground truth data, and a reinforcement reward loss function, comparing predicted rewards to actual episode rewards. Subsequently, online reinforcement learning (RL) is performed on an instance of the base generative model. This online RL processes online data using the model instance to generate predictions and uses the SFT-trained model to generate online predicted rewards. These online predicted rewards are then used to train the instance of the base generative model, improving its performance for various robot and application control tasks.
Owner:GDM HOLDING LLC

Digital microfluidic biochip droplet routing method based on reinforcement learning and imitation learning

The invention discloses a liquid drop routing method under fluid constraint in a digital micro-fluidic biochip based on reinforcement learning and imitation learning, and aims to solve the technical problems of low planning efficiency, low success rate and weak generalization ability of a traditional liquid drop routing algorithm under fluid constraint. The method adopts a strategy combining imitation learning and reinforcement learning and comprises the following steps: firstly, generating expert data by utilizing an expert algorithm meeting fluid constraint after transformation, and pre-training a network model through imitation learning to enable the network model to quickly grasp a basic routing strategy; the model and environment interaction are autonomously optimized through reinforcement learning, and the adaptability to the dynamic environment is improved. And finally fusing the weights to form a final model. Through the hybrid training framework, the success rate and planning efficiency of droplet routing are remarkably improved, and the generalization ability of the model to chip environments with different sizes is enhanced.
Owner:CHANGCHUN UNIV OF TECH

A gait imitation learning method for humanoid robots combined with periodic rewards

A method for learning the gait imitation of a humanoid robot combined with periodic rewards relates to the field of robot motion control technology. In view of the problem of unstable posture of humanoid robots when walking in a humanoid posture in a plane in the prior art, this application constructs a reference action library that integrates contact information as a reference for imitation reward items and periodic contact reward items. This application creates a comprehensive reference action library for basic actions and their corresponding periodic contact information. This strategy introduces periodic reward items by imitating the style of the reference action and its contact information, which not only improves the realism and style consistency of the robot's actions, but also enhances the attention to the details of the interaction between the feet and the ground during the execution of the action, thereby ensuring the stability of the posture of the humanoid robot when walking in a humanoid posture in a plane.
Owner:HARBIN INST OF TECH

Mechanical arm intelligent control method based on demonstration video imitation learning

The invention discloses a mechanical arm intelligent control method based on demonstration video imitation learning, and the method comprises the steps: firstly obtaining a human demonstration video containing a task target, and obtaining a video marked by a key point; the method comprises the following steps: hierarchically extracting semantic information and spatial geometric information by using a multi-modal large language model, decomposing a task into a plurality of sub-task stages, and generating a sub-target constraint function and a path constraint function of each sub-task stage; a task scene similar to a human demonstration video is arranged in a mechanical arm simulation environment, and environment key points are generated through feature extraction and clustering; constructing a mapping function from the video key points to the environment key points; and solving the optimal solution of the sub-target constraint function and the path constraint function of the current working environment of the mechanical arm, and driving the mechanical arm to execute the action until the task is completed. Through fine-grained key point analysis, multi-modal information fusion and cross-scene mapping, the action understanding accuracy, the complex scene adaptability and the task generalization ability of the mechanical arm are remarkably improved.
Owner:HANGZHOU DIANZI UNIV

Cluster confrontation method and system based on expert knowledge assisted deep reinforcement learning

The invention provides a cluster confrontation method and system based on expert knowledge-assisted deep reinforcement learning, and the method and system improve the initial strategy learning speed and overall combat effectiveness of the system by introducing an expert knowledge base and an imitation learning technology and combining deep reinforcement learning to optimize the collaborative decision-making efficiency of an intelligent agent. The method aims at providing an effective initial strategy acquisition mechanism, utilizing an expert knowledge base to accelerate the early strategy learning of the agents, reducing the training time, optimizing the strategies of the agents in a complex dynamic environment through a multi-agent deep reinforcement learning algorithm, and improving the cooperative combat ability. According to the scheme, the time required for initial strategy learning can be greatly shortened, a more optimized strategy is obtained in combination with deep reinforcement learning, the high efficiency of strategy tuning is guaranteed, and then the real-time guarantee in large-scale cluster confrontation is guaranteed.
Owner:TONGJI UNIV

Mechanical arm imitation learning method and device based on DMP

The invention provides a mechanical arm imitation learning method and device based on DMP, and belongs to the technical field of robots. The method comprises the steps that a trajectory is fitted through a motion primitive method; according to the teaching or simulation track, the weight of the motion primitive is obtained through a least square method; a new trajectory regeneration method is provided, and new trajectories with the same geometric features are generated by setting a new starting point, a new ending point or a new scaling coefficient; a new track passing point strategy is provided, so that the generated new track can pass through a specific key point; applying a Gaussian process to generate a high-order smooth trajectory passing through a specific point; generating a high-order smooth curve meeting the geometric characteristics by using a weighted average method; substituting the generated new trajectory into the new DMP model, and establishing a dynamic system model; and according to the dynamic system model, the mechanical arm joints are driven through the joint speed track, and speed control is achieved. According to the method, a large amount of similar industrial programming work can be reduced, and the flexibility and efficiency of trajectory planning and control are improved.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Unified supervision fine tuning and reinforcement learning training method based on dynamic weight fusion

The invention provides a unified supervision fine tuning and reinforcement learning training method based on dynamic weight fusion, and belongs to the field of artificial intelligence. According to the unified supervision fine tuning and reinforcement learning training method based on dynamic weight fusion, knowledge imitation and strategy exploration are balanced based on unified SFT and RL training targets, and therefore the purpose of improving training stability is achieved. Comprising the following implementation steps: step (1), data preparation; step (2), performing double-path parallel processing; step (2.1), an SFT path is established; (2.2) an RL path; step (3), a dynamic weight fusion mechanism; the loss weights of the SFT path and the RL path are dynamically adjusted through the global coefficient mu, and progressive transition from imitation learning to exploration learning is achieved.
Owner:青岛蚂蚁机器人有限责任公司

Air combat strategy generation system and method based on lightweight sequence modeling imitation learning

The invention provides an air combat strategy generation method based on lightweight sequence modeling imitation learning. According to the training method, an air combat decision model based on rules can be quickly migrated into a parameter model in a neural network form. The method comprises the following steps: firstly, constructing a rule-based air combat decision model, and generating high-quality training data through adversarial simulation; and then, the training data is converted into a state-action sequence pair, and the state-action sequence pair is input into a lightweight sequence modeling network based on a Transform architecture for learning and training. According to the method, decision knowledge of a rule model is converted into a neural network model through an imitation learning method, and efficient utilization and flexible generalization of expert knowledge in air combat decision model construction are realized; the model designed by the method can accurately reproduce the decision behavior of the rule model through a small amount of training, and compared with a reinforcement learning method, the strategy construction speed is higher and the strategy performance is more stable.
Owner:BEIHANG UNIV

Simulation learning method and device based on stream matching and dynamic reward scheduling

The invention relates to a deep learning technology, and discloses a flow matching and dynamic reward scheduling-based imitation learning method, which comprises the following steps of: acquiring an expert demonstration track, and extracting an expert feature sequence in the expert demonstration track by utilizing a flow matching model; obtaining a strategy state sequence of task execution of the agent strategy network and converting the strategy state sequence into strategy features; constructing a reward function by utilizing the expert feature sequence and the strategy feature; calculating a reward result and optimizing the median function network of the agent strategy network; according to the optimized value function network, an intelligent agent strategy network is optimized, and then an action sequence of a current strategy is sampled; and generating a track fragment according to the action sequence, and iteratively optimizing the intelligent agent strategy network by taking the track fragment as a training sample, and obtaining a target intelligent agent strategy network after iteration optimization is completed. The invention also provides an imitation learning device and equipment based on stream matching and dynamic reward scheduling, and a storage medium. According to the method, the state modeling efficiency in imitation learning and the stability of the reward structure can be improved.
Owner:CHONGQING UNIV

Intelligent learning method for grabbing of three-finger dexterous hand for multiple types of workpieces

PendingCN121403390AProgramme-controlled manipulatorHand graspData set
The invention discloses an intelligent learning method for grabbing of a three-finger dexterous hand for multiple types of workpieces, and belongs to the technical field of intelligent grabbing application of manipulators. A pre-defined rule and simulation expert double-layer mode is provided, a high-quality demonstration data set is automatically generated, a simulation grabbing scene of a mechanical arm and a three-fingered dexterous hand is constructed, recognition and classification of multiple types of workpieces are achieved by analyzing workpiece contour features, grabbing hand shapes are matched according to the recognition and classification, strategy pre-training is carried out through simulation learning, and the recognition and classification efficiency is improved. And optimizing the strategy based on a PPO algorithm and a composite reward function. Aiming at the problem of lack of a standardized three-fingered dexterous hand parameter model, a modeling method based on parameterized URDF and automatic script generation is used to realize efficient configuration and optimization of the three-fingered dexterous hand in simulation. Through performance evaluation, stable and high-quality grabbing of multiple types of workpieces can be achieved.
Owner:ZHEJIANG UNIV OF TECH