Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

52 results about "Strategy training" patented technology

Strategy optimization method and device based on interaction track, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of robot strategy training, financial science and technology, medical treatment and health and the like, and discloses a strategy optimization method and device based on an interaction track, equipment and a medium. The reinforcement learning agent interacts with the environment to generate an action and state sequence and record an interaction track, an optimal action strategy is generated based on a track optimization strategy, an original empirical data set is further generated, and a pre-training model is optimized through supervised learning to obtain a target strategy model. According to the method, a task-related high-quality trajectory is generated through reinforcement learning agent and environment interaction, original empirical data is extracted on this basis, and a pre-training strategy model is optimized in combination with a supervised learning mechanism, so that the sample utilization efficiency is effectively improved, and the generalization ability and execution robustness of the model under a multi-task condition are enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Strategy generation method and device based on hierarchical reinforcement learning, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a strategy generation method and device based on hierarchical reinforcement learning, equipment and a medium. And processing environment state information to generate a sub-target and a specific action, generating a strategy reward signal in combination with state change, carrying out joint training and updating on the dynamic causal graph and the hierarchical reinforcement learning model based on the strategy reward signal, and generating an optimized action strategy. The state evolution relation is modeled by constructing the dynamic causal graph, so that the reinforcement learning can obtain causal understanding of the state change trend, decomposition and optimization of sub-targets and actions are realized in combination with a layered reinforcement learning architecture, the response precision and generalization ability of the action strategy in a complex environment are improved, and the method is suitable for application and popularization. Therefore, the task completion stability and the convergence efficiency of strategy training are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Automatic driving digital twin agent system based on reinforcement learning training

The invention relates to the technical field of artificial intelligence and automatic driving, and discloses an automatic driving digital twin agent system based on reinforcement learning training, which comprises an agent library construction module, a simulation and configuration module, a distributed simulation execution module, a scheduling and learning module, a strategy training module and a simulation migration module. The method comprises the following steps: constructing a modular digital twin agent library, combining to generate an urban traffic simulation scene, executing large-scale parallel simulation on a distributed computing cluster, implementing dynamic task scheduling and course learning according to an agent learning progress, and training a reinforcement learning strategy and a value network based on simulation data. And verifying the strategy model and implementing migration deployment from simulation to reality. According to the method, the training efficiency can be improved, the diversity of simulation environments is enriched, and the sim-to-real migration difficulty is reduced, so that the robustness of the strategy model is improved.
Owner:KUNSHAN MENGYU 3D DIGITAL TECH CO LTD

Stochastic trading system with synthetic charting and real-time price anchoring for enhanced execution control

The invention relates to a trading platform that shifts focus from market prediction to disciplined execution, capital management, and trading psychology. By disclosing the general market trend in advance, users are empowered to train and operate within a known directional framework. The system generates synthetic charts using directional random walks anchored at algorithmically or randomly selected convergence points from historical data. This recreates real trends while introducing stochastic variability between anchors. A simulated order book models supply-demand dynamics using Poisson-distributed flows, capturing sentiment, liquidity shifts, and large trades. Features include Accelerated Time Compression for fast-forwarded trade simulation and a Rewind Trade mechanism allowing entry at the start of validated candlesticks. The rewind logic applies to both futures and spot markets, including staking actions. Safeguards such as rewind limits, volume caps, and LP prioritization preserve fairness. The platform supports multi-timeframe execution and bridges strategic training with interactive trading, enhancing decision quality and behavioral discipline.
Owner:MOTEVALLI ABYAZANI HOOMAN +1

Digital twin model placement and resource allocation method based on action mask reinforcement learning

The invention discloses a digital twin model placement and resource allocation method based on action mask reinforcement learning, and solves the problems of low reinforcement learning efficiency, lack of modeling for edge resource competition and the like caused by unhooking of a deployment strategy and resource allocation and overlarge action space. According to the method, the model placement position selection and the computing power resource allocation process are uniformly modeled into the Markov decision process, and the action mask mechanism is introduced to dynamically constrain the action space, so that the strategy training efficiency and the effectiveness of the deployment decision are improved. According to the method, not only is the competition relation of a plurality of digital twin service requests under shared edge node resources modeled finely, but also a potential function-based reward function is constructed, and a strategy is guided to converge to a system time delay minimization target. Finally, collaborative optimization of digital twinborn model deployment and resource allocation in a dynamic environment is realized, the computing power resource utilization rate and the service quality are improved, and the high requirements of an industrial scene for real-time performance and intelligence are met.
Owner:BEIJING UNIV OF TECH +1

Robot agent evaluation system for lower limb exoskeleton and gait strategy training method corresponding to agent

The invention relates to a robot agent evaluation system for a lower limb exoskeleton and a gait strategy training method corresponding to an agent, and belongs to the technical field of rehabilitation medical auxiliary equipment. The device comprises a humanoid robot, an exoskeleton, an auxiliary device and a detection element, the exoskeleton is connected with the humanoid robot through the auxiliary device, a first auxiliary structure is provided with a groove, a second auxiliary structure is provided with a protrusion, the groove is connected with the humanoid robot in a clamped mode, and the protrusion is detachably connected with the humanoid robot; the detection elements are installed on the soles of the humanoid robot. The gait of the humanoid robot is adjusted by using a gait strategy training method, so that the humanoid robot can simulate the walking postures of different healthy people or patients with different hemiplegia levels. By wearing the exoskeleton on the device and deploying different walking aid strategies, related data such as angle, torque and power change of each joint, total energy consumption and gait symmetry of the humanoid robot under different walking conditions are obtained, and an exoskeleton developer is assisted to better test and optimize an exoskeleton product.
Owner:WESTLAKE UNIV

Multi-energy micro-grid real-time regulation and control method fusing topology perception and meta-learning self-adaption

The invention discloses a multi-energy micro-grid real-time regulation and control method fusing topology perception and meta-learning self-adaption, and the method comprises the steps: inputting electric-heat-gas multi-source heterogeneous data obtained in real time into a pre-trained lightweight student model, and generating a prediction multi-energy flow cooperative regulation and control strategy. When the model is trained, firstly, a corresponding multi-energy micro-grid graph structure is constructed based on the network topology of the multi-energy micro-grid, multi-source heterogeneous data used as training data is combined and input to a teacher model constructed based on a graph neural network, and a corresponding predictive multi-energy flow cooperative regulation strategy is output; calculating a loss function based on the strategy and a corresponding optimal multi-energy-flow cooperative control strategy, and optimizing network parameters of the teacher model; and finally, migrating the knowledge of the trained multi-teacher network to the lightweight student model through a method of combining knowledge distillation, meta learning and reinforcement learning to obtain a trained lightweight student model. According to the invention, the problems of multi-target optimization, real-time regulation and control and cross-scene adaptation of the multi-energy micro-grid are solved.
Owner:CHONGQING NORMAL UNIVERSITY

Vehicle network interaction regulation and control method and system based on multi-task offline element reinforcement learning

The invention discloses a vehicle network interactive regulation and control method and system based on multi-task off-line element reinforcement learning, and the method comprises the steps: constructing a regulation and control strategy training mechanism based on an off-line element reinforcement learning framework, introducing a element network for task recognition, and obtaining a regulation and control strategy through a semantic-physical fusion mechanism enhanced by a large language model. And combined modeling of historical moving trajectory data and expert scheduling text knowledge is realized. And the element network generation task specific embedding is used for guiding a downstream strategy network to quickly adapt to a dynamic regulation and control task so as to formulate a context-aware regulation and control strategy. According to the method, the bottleneck of a traditional reinforcement learning method in the aspects of task mobility and generalization ability is effectively broken through, and the adaptation efficiency of a strategy network to diversified EV distribution and regulation targets is remarkably improved.
Owner:SOUTHEAST UNIV

Bid and tendering document deep understanding and multi-rule review management system based on large model

The invention relates to the technical field of bidding and tendering, and discloses a bidding and tendering file deep understanding and multi-rule review management system based on a large model, and the system comprises an environment state representation module which is used for encoding a bidding and tendering file into a structured state vector; the multi-agent decision-making module comprises a plurality of professional agents and is used for executing a review action according to the state vector; the dynamic risk transmission module is used for generating a risk broadcast vector when a professional agent identifies the risk, and adjusting subsequent actions of other agents through a cross-agent attention mechanism; the strategy training module is used for training the strategy of the multi-agent decision-making module by utilizing an expert review track through a generative adversarial imitation learning mode; and the report generation module is used for outputting a dynamic review path report. According to the method, the dynamic optimization of the review path is realized by simulating an expert strategy, the risk collaborative review across professional fields is realized through a risk transmission mechanism, and the traceability of the review process is ensured at the same time.
Owner:BEIJING NUOLAI DIGITAL TECHNOLOGY CO LTD

Multi-spacecraft orbit game method based on MAPPO

The invention provides a multi-spacecraft orbit game method based on MAPPO. The method comprises the following steps: step 1, establishing an orbit pursuit game scene model; 2, designing reward functions of the spacecrafts of the two parties in the pursuit game; 3, remodeling a pursuit game reward function; 4, training a pursuit game strategy network; according to the invention, aiming at an orbit game scene model, a guide reward function is designed, a strategy training network is perfected, and the problem that a strategy network designed by a traditional reward function is not easy to converge is solved. According to the method, the MAPPO algorithm is combined to complete pursuit game intelligent control strategy network training, the trained network is utilized, the track pursuit game problem under the constraint conditions of considering track dynamics, maneuvering modes and capabilities, fuel consumption and the like is solved, and intelligent control over the pursuit game is achieved.
Owner:BEIHANG UNIV

Intelligent optimization system for enhancing programmed advertisement bidding based on trajectory data

The invention relates to the field of data processing, and relates to an intelligent optimization system for enhancing programmed advertisement bidding based on trajectory data. Comprising a simulation bidding module used for performing data enhancement on original advertisement bidding data so as to obtain enhanced advertisement bidding data; the track data generation module is used for arranging the enhanced advertisement bidding data into track data capable of being learned by an IQL model; the IQLTrainer module is used for generating training parameters of an IQL model, constructing an IQL model instance in combination with the trajectory data, training the IQL model, and storing optimal model parameters after training is finished; the IQL model is used for acquiring a bidding strategy according to the current state of single real-time bidding; the training parameters comprise a budget proportion, an advertiser ID, a bidding model name and an operation parameter index. By adopting the system, the convergence efficiency of a bidding algorithm can be improved, and the noise of a training track is reduced.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Video large language model security assessment and defense method, system, device and medium

The invention discloses a video big language model security assessment and defense method, system, device and medium, which are corresponding schemes, and in the scheme, a constructed video risk assessment data set is large in scale, comprehensive in coverage, hierarchical and multi-language in design, so that the data set has stronger risk generalization ability, and the security assessment and defense of the video big language model is realized. And the performance and security of the model in an actual deployment environment can be effectively improved. In the two-stage post-training strategy, the method in the first stage is a context sensing type risk detection token fine adjustment method, and the risk signal sensing capacity of the model during complex input processing can be remarkably improved; the method in the second stage is a security reward-driven dynamic strategy optimization method, and the stability and generalization of strategy training can be improved. The finally obtained optimized video big language model can efficiently realize risk detection and safety response, and is especially suitable for scenes such as real-time content auditing, online education question answering, government affair video monitoring and the like.
Owner:UNIV OF SCI & TECH OF CHINA

A multi-agent reinforcement learning training method and system

The application relates to a multi-agent reinforcement learning training method and system, which comprises the following steps: a local strategy training stage, in which an agent makes an action by using local observation information; a global strategy training stage, in which a global strategy of the agent uses the local strategy as an action module for interacting with an environment, the global strategy takes global state of the environment as input, encodes the global information in a hidden space, and uses a neural network to fit the global state by using local observation information of all agents; the local strategy makes a suitable action in the environment according to the local observation information and the output of the global strategy; and a local strategy optimization stage, in which the global strategy, the local strategy and a fitting model obtained in the previous two stages are used to optimize the existing local strategy, so that an agent with better effect is finally obtained. The application can improve the speed and accuracy of multi-agent reinforcement learning training.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Garment design system and method based on artificial intelligence

The invention provides a costume design system and method based on artificial intelligence, and relates to the field of artificial intelligence, and the method comprises the steps: collecting multi-dimensional design data through a cloud collaborative architecture of costume design resources; performing dynamic hierarchical allocation on the multi-dimensional design data to obtain a hierarchical allocation result, and performing sparse processing on the multi-dimensional design data to obtain a sparse model of the costume design resources; performing strategy training based on deep reinforcement learning on the sparse model according to a hierarchical distribution result and historical collaborative data to obtain a sharing strategy and an optimized migration parameter; constructing a knowledge migration component of the cloud collaborative architecture based on the sharing strategy and the optimized migration parameters; knowledge migration optimization is carried out on the collaborative process of the costume design resources based on the knowledge migration component through the cloud collaborative architecture, the costume shared resources for intelligent generalization are generated, dynamic hierarchical distribution and knowledge migration optimization can be carried out on the costume design resources, and intelligent generalization sharing of the costume design resources is achieved.
Owner:HUNAN ARTS & CRAFTS VOCATIONAL COLLEGE

Coding strategy digital training method and device, terminal and storage medium

The invention discloses a coping strategy digital training method and device, a terminal and a storage medium. The method comprises the following steps: receiving individual basic information and environment state information input by a user; performing demand clarification grading evaluation on the user; performing thinking feature classification analysis according to the user feature data to judge a training type; setting individual coping strategy training parameters of each parameter component containing the multi-dimensional scene factor indexes; generating a coping strategy training plan; executing multi-mode man-machine interaction preset scene coping strategy loop training of the corresponding project; calculating and evaluating an adaptation index; and outputting a training report. The invention provides a self-service, systematized, structured and continuously optimized psychological adaptive ability training scheme, and the defects of the existing psychological service in the aspects of coverage, systematicness, self-adaptability and accessibility are effectively overcome.
Owner:SHENZHEN IVY LEAGUE COUNSELING CO LTD

Dynamic experience priority management method for task type dialogue strategy optimization

The invention provides a dynamic experience priority management method and system for task type dialogue strategy optimization. The method comprises the steps that information such as states, actions and rewards is obtained in the dialogue interaction process; constructing an experience evaluation mechanism based on the main Q network and the target Q network, performing conjoint analysis on the current learning income and the historical learning trend of each experience, and quantifying the empirical value state; dividing experiences into four types of normal forms, namely an unmastered normal form, a forgotten normal form, a stable normal form and a noise normal form; a sampling priority is generated in a self-adaptive mode according to the empirical state, and the sampling proportion is dynamically adjusted through threshold value mapping, a piecewise function and a hierarchical weight strategy; a hierarchical tree structure is employed to implement weighted sampling for strategy training. According to the method, the influence of noise samples can be inhibited, forgotten high-value experience can be activated, the effective sample proportion can be improved, dialogue strategy convergence can be accelerated, and the method is suitable for a multi-domain and multi-task complex task type dialogue system and an experience playback-based reinforcement learning scene.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Personal agent strategy training method and system based on preference reinforcement learning

The invention discloses a preference reinforcement learning-based body-equipped agent strategy training method and system, and the method comprises the steps: obtaining and constructing a multi-source non-stationary environment preference data set: collecting historical interaction trajectories from different bottom-layer dynamic environments, and generating trajectory fragment pairs and corresponding human preference labels; carrying out environment dynamic predictor training, and enabling a calculation control unit to grasp dynamic evolution and dynamic change rules of the body-equipped intelligent body in different physical environments in advance; executing decoupling joint optimization based on conditional mutual information and a regret model; extracting a pure task award decoder which is absolutely immune to any dynamic physical disturbance; security underlying action strategy distillation is executed based on expectation regression, and a final action strategy for controlling physical equipment is safely generated by using pure rewards; and after the strategy network is trained and converged, the processor compiles the strategy network into an executable instruction, and directly deploys the executable instruction to a control unit of the target physical equipment to execute intelligent action control.
Owner:NANJING UNIV

An intelligent cooperative confrontation strategy generation method based on a heterogeneous unmanned aerial vehicle group

The application discloses a kind of intelligent cooperative confrontation strategy generation methods based on heterogeneous unmanned aerial vehicle group, belong to multi-agent deep reinforcement learning field.The application includes the following steps: constructing cooperative confrontation scene, constructing blue side confrontation rule strategy model, constructing red side heterogeneous multi-agent reinforcement learning model, intelligent strategy training and generation.The application is directed to intelligent cooperative confrontation strategy generation method problem, abstracts electromagnetic confrontation and fire confrontation mechanism, and innovatively constructs heterogeneous multi-agent reinforcement learning model, and realizes intelligent confrontation strategy generation by data training.
Owner:THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION

Data-efficient hierarchical reinforcement learning

Training and / or utilizing a hierarchical reinforcement learning (HRL) model for robotic control. The HRL model can include at least a higher-level policy model and a lower-level policy model. Some implementations relate to technique(s) that enable more efficient off-policy training to be utilized in training of the higher-level policy model and / or the lower-level policy model. Some of those implementations utilize off-policy correction, which re-labels higher-level actions of experience data, generated in the past utilizing a previously trained version of the HRL model, with modified higher-level actions. The modified higher-level actions are then utilized to off-policy train the higher-level policy model. This can enable effective off-policy training despite the lower-level policy model being a different version at training time (relative to the version when the experience data was collected).
Owner:GOOGLE LLC

A robust adversarial training framework for multi-agent reinforcement learning energy system

The present application relates to a kind of robust confrontation training framework for multi-agent reinforcement learning energy system, comprising: constructing an adversarial agent to generate adversarial attack, and modeling as adversarial partially observable stochastic game system;Fixed pre-trained victim multi-agent strategy, train an optimal deterministic confrontation strategy to produce bounded disturbance;Fixed optimal confrontation attack strategy, improve the robustness of victim strategy under optimal attacker through adversarial training.The beneficial effects of the present application are: the present application models adversarial attack as an attack adversary based on single-agent reinforcement learning, and learns the strongest attack strategy considering attack constraints.Mathematically, the problem is constructed as a confrontation Markov game, and the performance of the integrated energy management system based on multi-agent reinforcement learning is improved through robust confrontation training.
Owner:ZHEJIANG ZHENENG YUEQING POWER GENERATION CO LTD +1

Video large language model security evaluation and defense method, system, device and medium

The application discloses a video large language model security evaluation and defense method, system, device and medium, which are corresponding solutions, wherein the constructed video risk evaluation dataset is large in scale and comprehensive in coverage, hierarchical and multilingual design, so that the dataset has stronger risk generalization capability, and the performance and security of the model in the actual deployment environment can be effectively improved. In the two-stage post-training strategy, the method of the first stage is a context-aware risk detection token fine-tuning method, which can significantly improve the risk signal perception ability of the model when processing complex inputs; the method of the second stage is a security reward-driven dynamic strategy optimization method, which can improve the stability and generalization of strategy training. The optimized video large language model obtained finally can efficiently realize risk detection and security response, and is especially suitable for scenes such as real-time content review, online education and question answering, government video monitoring and the like.
Owner:UNIV OF SCI & TECH OF CHINA

Training system and method for teammate-aware multi-agent coordination communication strategy

The application discloses a kind of training system and method of multi-agent cooperation communication strategy based on teammate perception, with the help of teammate modeling method, generate customized, sparse message for teammates in the process of interacting with teammates. Through multi-agent reinforcement learning communication method, the agent can learn efficient cooperation communication strategy that can resist interference under the condition of limited bandwidth, and the agent cooperation communication strategy obtained by training can enable the agent to select reasonable cooperation action according to the observation information obtained by itself, and send information to teammates at appropriate time to promote cooperation, or cooperate according to the teammate information received. The application realizes close and effective cooperation through communication between a small number of nodes in a limited bandwidth actual scene, and can ensure the anti-interference ability in the process of multi-agent cooperation communication.
Owner:NANJING UNIV

Context strategy training and reasoning method for strategy action self-adaption

The invention discloses a context strategy training and reasoning method for strategy action self-adaption, and relates to the technical field of artificial intelligence and intelligent control. The invention relates to a context adaptive strategy learning method for variable actions. By introducing a Token sequence organization structure, linear action compensation matrix unified modeling and a double-pass forward reasoning mechanism, rapid strategy generalization of cross-body, cross-scale and cross-dynamics disturbance is realized.
Owner:SHANGHAI JIAOTONG UNIV

Robot-oriented artificial intelligence strategy training system and method

The invention provides a robot-oriented artificial intelligence strategy training system and method, and relates to the technical field of artificial intelligence. The system comprises a standardized interface process and a strategy training module connected with the standardized interface process, the standardized interface process is configured to trigger callback with a periodic clock; in each callback period, environment data at the target moment and action execution data of the target robot under the environment data are synchronously collected; framing the environment data at the target moment and action execution data of the target robot under the environment data to generate a first data unit; the strategy training module is configured to train the first strategy model based on the first data unit. According to the method, strict time sequence correspondence of the environment and the action data is ensured, and the Markov property of the data is effectively maintained.
Owner:AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD

Data-efficient hierarchical reinforcement learning

Training and / or utilizing a hierarchical reinforcement learning (HRL) model for robotic control. The HRL model can include at least a higher-level policy model and a lower-level policy model. Some implementations relate to technique(s) that enable more efficient off-policy training to be utilized in training of the higher-level policy model and / or the lower-level policy model. Some of those implementations utilize off-policy correction, which re-labels higher-level actions of experience data, generated in the past utilizing a previously trained version of the HRL model, with modified higher-level actions. The modified higher-level actions are then utilized to off-policy train the higher-level policy model. This can enable effective off-policy training despite the lower-level policy model being a different version at training time (relative to the version when the experience data was collected).
Owner:GOOGLE LLC

Human-machine intelligent confrontation method and system based on cloud architecture

The invention belongs to the technical field of computers, and aims to provide a man-machine intelligent confrontation method and system based on a cloud architecture. On one hand, multiple terminal devices are used for collecting user behavior data in a distributed mode and generating local confrontation samples, strategy training is conducted through the multiple terminal devices on the basis of the local confrontation samples of the terminal devices, and a global model is formed through aggregation optimization of a cloud server; the generalization ability and flexibility of the model under diversified human behavior data are remarkably improved, and the problems of single strategy and poor adaptability in the prior art are solved; and on the other hand, after the cloud server completes global model training, a strategy revenue matrix is generated through large-scale man-machine confrontation simulation, and a game optimization method is further introduced to solve an equilibrium strategy, so that the confrontation adaptability and the intelligent decision-making capability of the AI strategy in a complex human behavior mode are effectively improved.
Owner:ARMOR ACADEMY OF CHINESE PEOPLES LIBERATION ARMY

A paragraph-level neural machine translation method and system based on masking mechanism

The present invention discloses a method and system for passage-level neural machine translation based on a masking mechanism, which belongs to the technical field of natural language processing and language translation. This method uses a masking mechanism to mask part of the information in the sentence to be translated, so that the model is more dependent on contextual information when translating sentences. This method mainly makes improvements to the training strategy and adopts a two-stage training process to better enable the encoder to restore the encoding information, effectively improving the effect of passage-level neural machine translation. The system part includes a data processing module, a source language encoder module, a source language self-encoding decoder module, a target language decoder module and a generator module. Compared with the existing technology, this method adopts a masking and self-encoding strategy in the training strategy, trains a passage-level translation model that is more dependent on context when encoding, makes full use of contextual information, and further improves the effect and quality of passage-level machine translation.
Owner:BEIJING INST OF TECH

Reinforcement learning policy training method, device, and storage medium

The application provides a reinforcement learning strategy training method and device and a storage medium. The method comprises: generating a plurality of action data in response to a control instruction of a user; updating an intervention experience pool and a full experience pool according to the action data, or updating the full experience pool; updating network parameters of an intervention value network and network parameters of a value network according to the intervention experience pool and the full experience pool, and updating network parameters of a first strategy network based on the updated intervention value network and the value network; and updating network parameters of a second strategy network according to the network parameters of the first strategy network in a preset interaction period. The application can predict future risks through intervention signals themselves without directly imitating intervention actions, so that the strategy can identify hidden dangerous states that seem safe at present but will inevitably lead to future intervention, and the performance of the trained strategy is improved.
Owner:BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Strategy training method and device, and storage medium

The application provides a policy training method and device and a storage medium. The method comprises: obtaining a time-ordered action sequence; updating a full experience pool according to each action data; determining at least one target action data according to the action sequence, and updating an intervention experience pool according to the at least one target action data; updating network parameters of a value network according to the intervention experience pool and the full experience pool, and updating network parameters of a first policy network based on the updated value network. The application can determine the target action data in the updating process of the first policy network, forcibly capture the critical state before the intervention occurs, enable the first policy network to learn the safety boundary causing the intervention, identify the dangerous trend earlier, realize accurate causal attribution, and improve the performance of the first policy network obtained by training.
Owner:BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Reinforcement learning strategy training methods, equipment, and storage media

This application provides a reinforcement learning policy training method, device, and storage medium. The method includes: generating multiple action data in response to user control commands; updating an intervention experience pool and a full experience pool, or updating the full experience pool, based on each action data; updating the network parameters of an intervention value network and a value network based on the intervention experience pool and the full experience pool, and updating the network parameters of a first policy network based on the updated intervention value network and value network; and updating the network parameters of a second policy network according to a preset interaction period based on the network parameters of the first policy network. This application can predict future risks through the intervention signal itself without directly imitating intervention actions, enabling the policy to identify seemingly safe but potentially dangerous states that will inevitably lead to future intervention, thereby improving the performance of the trained policy.
Owner:BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD