Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

168 results about "Reward value" patented technology

Value of Rewards is a Total Reward, Remuneration, Compensation & Benefits activity by René Broekhuis.

Multimodal large model task processing method, device and equipment based on reinforcement learning

The invention provides a multi-modal large model task processing method, device and equipment based on reinforcement learning, and the method comprises the steps: employing a multi-modal large model, generating G groups of responses for multi-modal visual task data, carrying out the scoring of each group of responses, obtaining a reward value, carrying out the standardization of the reward value, obtaining a dominance score, building a strategy updating gradient through the dominance score, and carrying out the calculation of the strategy updating gradient. A final result is obtained after expectation reward maximization, model parameter adjustment and multi-round iteration, and an expectation reward objective function comprises an expectation calculation item and KL divergence constraint, so that a current strategy can be updated by comparing the performance of candidate strategies through group relative strategy optimization, and local optimum can be helped to be jumped out; the strategy updating amplitude is limited by the KL divergence constraint, and the model is prevented from violently changing in the optimization process, so that the training stability is improved; the dynamic strategy iteration allows the model to adjust the balance between exploration and utilization according to the learning progress on the basis of keeping the stability, thereby further ensuring the effectiveness and stability of strategy optimization.
Owner:STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2

Robot adaptive training method and device based on reinforcement learning and medium

The invention relates to the technical field of robot training. The robot self-adaptive training method based on reinforcement learning comprises the steps that task sub-target information is generated through a high-level strategy network, the task sub-target information is input into a low-level execution network, an action control instruction is generated according to the task sub-target information, interaction feedback information is collected in the execution process, and the action control instruction is sent to a robot through a robot. Calculating a reward value according to the interaction feedback information, carrying out association processing on the reward value and the scene complexity parameter, executing a dynamic reward shaping operation, generating an adjusted reward signal, generating a strategy model optimized by meta-learning based on the adjusted reward signal, loading the strategy model in a simulation environment, and carrying out dynamic reward shaping. A target strategy model optimized through simulation training is generated, the target strategy model is loaded to the robot, and the robot is controlled to execute task operation in the actual interaction scene. The method has the effect of realizing adaptive task learning of the robot in a multi-interaction scene.
Owner:SEVEN (BEIJING) EDUCATION TECH CO LTD

Multi-round automatic machine learning agent system based on reinforcement learning optimization

The invention provides a multi-round automatic machine learning agent system based on reinforcement learning optimization. Comprising a task analysis module used for generating an initial prompt for an MLE agent to call; the MLE agent module is used for generating an executable code; the code executor is used for generating an execution result; the evaluator is used for outputting a normalized value of each index and a code correctness identifier; the reward construction module is used for generating a reward value; the reinforcement learning optimizer is used for calculating group average return and candidate advantages and updating strategy parameters of the MLE intelligent agent module based on the candidate advantages; and the multi-round interaction control module is used for feeding back the execution result of the previous round and the reward value to the MLE agent module in the multi-round interaction process, and controlling code generation of the next round until a preset termination condition is met. According to the invention, strategy adaptive evolution, reinforcement learning optimization of fine-grained credit distribution and multi-round closed-loop automatic process improvement can be realized.
Owner:北京衔远有限公司 +1

Large language model optimization method and optimization device

The invention relates to the technical field of artificial intelligence, and provides a large language model optimization method and device. According to the method, a candidate answer is obtained by inputting a question in a pre-trained large language model, a first reward value, a second reward value and a third reward value of the candidate answer are calculated, the first reward value is used for representing a basic reward of the candidate answer, the second reward value is used for representing the integrating degree of the candidate answer and the question, and the third reward value is used for representing the integrating degree of the candidate answer and the question. The third reward value is used for representing the alignment degree of the candidate answers and the user preferences, then performing multi-head fusion on the first reward value, the second reward value and the third reward value to obtain a target reward value, and finally training a pre-trained large language model based on the target reward value to obtain an optimized large language model. The importance coefficients of different levels of reward mechanisms of the large language model are adaptively learned in a multi-head layering manner, so that the model training efficiency is improved, and the generation quality of the trained model is improved.
Owner:北京衔远有限公司 +1

Large language model training method and device, electronic equipment and storage medium

The invention discloses a large language model training method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence and reinforcement learning. The method comprises the steps of obtaining an initial prediction answer statement corresponding to a question statement; semantic expansion retrieval is carried out on the initial prediction answer statement, and a retrieval answer statement corresponding to the initial prediction answer statement is determined; scoring the retrieval answer statement, and determining a target reward value of the initial prediction answer statement based on a scoring result of the retrieval answer statement; and if the target reward value is smaller than a preset reward threshold value, adjusting the pre-trained large language model based on the target reward value, and inputting the question statement into the adjusted model again until the finally obtained target reward value is greater than or equal to the preset reward threshold value, thereby completing training of the pre-trained large language model. The optimization training efficiency of the large language model and the accuracy of model output can be improved, and the stability of large language model output is improved.
Owner:VOYAH AUTOMOBILE TECH CO LTD

Automatic leveling method for intelligently controlling underwater leveling machine

The invention discloses an automatic leveling method for intelligently controlling an underwater leveling machine. The method comprises the following steps: acquiring real-time attitude data and working surface flatness requirements of a plurality of execution mechanisms of the underwater leveling machine, and constructing state and operation action information characteristics; an execution mechanism is mapped into an agent with a body, information is shared through an underwater acoustic communication network, and each agent selects and executes a leveling action based on self and neighborhood interaction data; after execution, obtaining a reward value, storing the reward value in an experience pool, sampling and updating the state-action function, and performing iterative training to obtain a multi-body agent reinforcement learning optimization model; continuously monitoring the state of a mechanism when the model runs, dynamically updating the features of an affected area when a fault or hydrodynamic abnormality is identified, and recalculating a leveling strategy to obtain a dynamic re-planning scheme; the scheme is combined with model output to generate a real-time leveling instruction sequence, and the real-time leveling instruction sequence is decomposed and then issued to each execution mechanism, so that self-adaptive accurate leveling control of the underwater leveling machine is realized, and the working efficiency and reliability of a complex underwater environment are effectively improved.
Owner:CCCC FOURTH HARBOR ENG INST CO LTD

Reinforced learning training method and system for relieving hallusion of multi-modal large model

The invention discloses a reinforcement learning training method and system for relieving illusion of a multi-modal large model, and belongs to the field of reinforcement learning training of a multi-modal large language model. Firstly, a planning and visual description generation step is introduced in an early stage to guide a model to perform structured reasoning, then a grouping relative strategy optimization algorithm is used, reward values are calculated for multiple candidate responses generated by the model after cold start, and particularly, a visual perception reward mechanism is set. The reward mechanism evaluates the consistency of the generated text description and the visual information by using an external large language model. Then, based on a vision description attention score advantage distribution method, learning of the model on key vision signals is dynamically enhanced, and the perception ability of the model on the vision signals is improved; and finally, the perception and reasoning performance of the model is further improved by adopting multiple rounds of rejection sampling and supervised fine tuning. The scheme does not depend on a model architecture, the extra overhead is small, the illusion problem caused by early image-text inconsistency is effectively solved, and the accuracy and the reliability are improved.
Owner:ZHEJIANG UNIV +1

Reward model training method and device, strategy model training method and device and electronic equipment

The invention provides a reward model training method and device, a strategy model training method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a preference data pair which comprises a preferred response and a non-preferred response generated for the same prompt word, and each of the preferred response and the non-preferred response is composed of a plurality of text unit sequences; inputting each text unit sequence into a to-be-trained reward model to obtain a predicted reward value; and calculating the total training loss according to the predicted reward value, and updating the model parameters of the to-be-trained reward model. According to the method, the response text is subjected to serialized splitting, and the preference data composed of the preferred response and the non-preferred response is introduced for comparative learning, so that the target of model training is no longer to evaluate the absolute quality of a single response, but to identify a key text unit which causes one response to be superior to the other response; the fine-grained evaluation of the response content is realized, and the evaluation accuracy of the reward model and the identification capability of complex user preferences are effectively improved.
Owner:IFLYTEK SOUTH CHINA ARTIFICIAL INTELLIGENCE RES INST GUANGZHOU CO LTD

Large language model training method and device and electronic equipment

The invention provides a large language model training method and device and electronic equipment. The method comprises the steps that control information is set for a large language model through system cue words of the large language model; inputting training data to the large language model, obtaining output data of the large language model, and extracting N digits from the output data; obtaining a first reward value according to the number M of the numbers carrying the preset number label and the number source in the N numbers; in M numbers carrying preset number labels and number sources, standard values corresponding to the numbers are determined according to the number sources, and second reward values are obtained according to comparison results of the standard values and the numbers; and performing reinforcement learning training on the large language model by adopting a near-end strategy optimization mode or a group relative strategy optimization mode, and forming a training reward value in the near-end strategy optimization mode or the group relative strategy optimization mode according to the first reward value and the second reward value. According to the embodiment of the invention, the accuracy of digits generated by the large language model can be improved.
Owner:BEIJING SANKUAI CLOUD COMPUTING TECH CO LTD

Method and apparatus for preference-training language model

A method and apparatus for preference-training a language model are provided. The method according to some embodiments may include obtaining a dataset including a plurality of pieces of response data, wherein each of the plurality of pieces of response data includes multiple responses generated by the language model for a query and user preference information corresponding to each of the multiple responses, filtering out some of the plurality of pieces of response data included in the dataset using reward values for each of the plurality of pieces of response data, output from a proxy model that receives the dataset as input and training the language model using other pieces of response data that have not been filtered out.
Owner:SAMSUNG SDS CO LTD

Reinforcement learning training method and device for large language model

The embodiment of the invention provides a reinforcement learning training method and device for a large language model, and the method comprises the steps: inputting a first input text into the large language model, enabling the large language model to process the first input text based on a current strategy, and obtaining G output texts; respectively combining the G output texts with the first input text, and respectively inputting combination results into a trained reward model to obtain G reward values corresponding to the G output texts; the current strategy is updated based on a target function, the target function comprises a global penalty term, the global penalty term represents that a penalty value is applied only when the maximum value in the G reward values does not exceed a preset threshold value, and the penalty value is determined according to respective first output probabilities of the G output texts under the current strategy; therefore, the training of the large language model under the reinforcement learning framework is realized.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Model training method and device, task execution method and device, electronic equipment and storage medium

PendingCN120930803ADiscounts/incentivesMachine learningReward systemEngineering
According to the model training method and device, the task execution method and device, the electronic equipment and the storage medium, in the method, the prompt content, the reply content output by the target model for the prompt content and the process data for generating the reply content can be firstly obtained, and then the prompt content, the reply content and the process data are input into the reward system; therefore, the reward value of each reasoning step in the process data is obtained, and finally, the target model is iteratively trained based on the reward value of each reasoning step. The reward system of the method no longer generates the reward value for the token of the sample, but generates the reward value for each reasoning step in the process data, so that the target model can pay attention to the integrity and logicality of the reply content in the training process, and then the performance and stability of the target model in a complex task and the robustness of the model are improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Reinforced learning training method and system based on multi-target structured reward function

The invention provides a reinforcement learning training method and system based on a multi-target structured reward function, and the method comprises the steps: an environment simulation layer comprises a physical engine and a blood vessel environment simulator, and is used for simulating the physical movement and interaction of a catheter in a blood vessel, and providing an environment state for a reward calculation layer; the reward calculation layer performs calculation through a plurality of independent calculation modules and finally generates a comprehensive reward value through a composite reward aggregator; the reinforcement learning layer learns and optimizes a navigation strategy according to the comprehensive reward value; and the monitoring evaluation layer monitors each performance index in the training process and adaptively adjusts the weight of each reward component in the reward calculation layer according to the index change. According to the invention, a major breakthrough of reinforcement learning in medical robot training is realized. Compared with a traditional single reward function method, the intelligent agent which is really thought by experts can be trained, and the method not only pursues to quickly reach a target, but also knows how to safely and stably complete operation.
Owner:SHANGHAI OPERATION ROBOT CO LTD

Task type dialogue strategy learning method and system based on reinforcement learning

The invention belongs to the technical field of intelligent task-based dialogue, and particularly relates to a task-based dialogue strategy learning method and system based on reinforcement learning, which adopts a soft actor commentator algorithm to be combined with behavior cloning and maximum entropy inverse reinforcement learning to relieve a cold start problem. According to the method, maximum entropy inverse reinforcement learning is utilized, a reward value is calculated, and a user target in a dialogue is accurately deduced according to maximum log likelihood estimation of a human dialogue and a simulated dialogue; a potential reward function is extracted from a successful track through inverse reinforcement learning, manual reward design is replaced, accumulated reward maximization is pursued during strategy optimization, diversity and exploratory performance of the strategy are encouraged, and the dialogue strategy is prevented from being converged to a single mode too early.
Owner:QUFU NORMAL UNIV

Data enhancement method and system for Android malicious software family classification

The invention belongs to the technical field of network technology and security correlation, and discloses a data enhancement method and system for Android malicious software family classification, and the method comprises the steps: generating an adversarial sample based on a target sample, and carrying out the data enhancement of a target small family; wherein the reinforcement learning strategy network is constructed to generate a selection probability according to an original sample; establishing a reward function, wherein the reward value is higher when the classification result of the disturbed sample is closer to the target small family more easily; training the strategy network according to the reward value and selecting a target sample; and performing family tree analysis on samples in the target small family, and generating an evolution sample according to a family evolution path to perform data enhancement. According to the invention, the discrimination capability of the model near the boundary is improved by using the adversarial sample; the evolution sample improves the distribution density and the coverage range of the small categories in the space, the adversarial sample and the evolution sample are organically combined, and a new technical approach is provided for solving the family classification problem under long-tail distribution.
Owner:HUAZHONG UNIV OF SCI & TECH

Reinforced learning training method and system for small-scale large language model

The invention relates to the technical field of reinforcement learning, in particular to a reinforcement learning training method and system for a small-scale large language model, and the method comprises the steps: obtaining user query and an external tool set, and constructing a tool integration reasoning training task; driving the large language model to generate thinking content and a tool calling instruction, and forming an interaction track in combination with an external tool execution result; respectively calculating a format reward, a tool calling correctness reward, a dynamic length reward and a distillation learning reward according to the interaction track, and fusing each reward item to obtain a total reward value; performing trajectory multiplexing and trajectory data augmentation processing on the interaction trajectory to generate an extended training data set; and based on the extended training data set and the total reward value, performing parameter iterative optimization on the to-be-trained small-scale large-language model by adopting a reinforcement learning strategy. According to the method, the reasoning capability and the structured tool calling capability of a small-scale large-language model in a complex tool integrated reasoning task can be stably improved.
Owner:GUSU LAB OF MATERIALS

Large model fine tuning method and device, electronic equipment and computer storage medium

The invention provides a large model fine tuning method and device, electronic equipment and a computer storage medium, and the method comprises the steps: generating a candidate answer sequence according to question information through employing a strategy model after the question information is received; afterwards, generating an evaluation value sequence by using the reference model, and generating a reward value sequence under different reward functions by using a reward model comprising the multi-dimensional reward functions; performing group advantage evaluation according to the reward value sequence under all reward functions to obtain an advantage value sequence; performing difference analysis on the dominant value sequence by using the evaluation value sequence to obtain a difference analysis result; and finally, carrying out fine adjustment on the strategy model based on a difference analysis result to obtain an optimized strategy model. By fusing a multi-dimensional reward mechanism, the accuracy, logicality and robustness of a large model in complex logic and multi-scene tasks are effectively improved.
Owner:ASIAINFO TECH CHINA INC

Malware detection model training method and device based on itemized reward and operation index linkage

The invention discloses a malicious software detection model training method and device based on item reward and operation index linkage, and belongs to the technical field of artificial intelligence and malicious software detection. The method comprises the following steps: constructing a reinforcement learning model containing a policy network; the policy network is configured to receive software feature input and output multi-modal actions including classification results, structured interpretation text and detection confidence; based on the training sample, generating a multi-modal action by utilizing a strategy network, and calculating a composite reward value for the generated multi-modal action by utilizing a preset subitem reward; the subitem rewards at least comprise a structure consistency reward, a key element hit reward and a confidence coefficient calibration reward; and calculating advantage estimation based on the composite reward value to construct an objective function of a near-end policy optimization algorithm by using the advantage estimation, and maximizing the objective function by updating parameters of the policy network. The method can improve the explanatory and confidence of model output, and improves the recognition precision.
Owner:HARBIN ANTIY TECH

Question and answer model training method and device, electronic equipment and storage medium

The invention provides a question and answer model training method and device, electronic equipment and a storage medium. The method comprises the steps that multiple pieces of training data are acquired; in the process of training the question and answer model according to the sample question, performing repeated fragment identification on a reasoning process generated by performing multiple sampling on the sample question to determine a repeated penalty value of a repeated fragment, and performing reward value calculation on an answer generated by performing sampling on the sample question; performing fusion based on the repeated penalty value and the reward value to form a reward signal, performing normalization processing on the reward signal to obtain a dominant value of current sampling, and introducing the dominant value of the current sampling into a grouping strategy optimization algorithm to construct a loss function; and updating model parameters through gradient back propagation to inhibit repeated thinking behaviors of the question and answer model in the reasoning process, and continuing to perform iterative training on the updated question and answer model. More attention is paid to'effective thinking 'in the training process, and the ability of the question and answer model to solve complex problems is improved.
Owner:SHANGHAI ANXINCHENG NETWORK TECHNOLOGY CO LTD

Reward distribution method and device, computer equipment and storage medium

The invention relates to a reward distribution method and device, computer equipment and a storage medium, and the method comprises the steps: registering artificial intelligence bodies on a block chain, and obtaining the real-time task information of each artificial intelligence body for executing a current task; the current task is published on the block chain by a task publisher; obtaining reward rule information and a basic weight written into the smart contract by the task publisher; the smart contract is deployed in the block chain; performing summation calculation on the basic weight and the real-time task information by using an intelligent contract to obtain a dynamic weight value, and calculating a target reward value based on the reward rule information and the dynamic weight value; and assigning a target reward value to each artificial intelligence body. Through the method and the device, the problem of low flexibility of reward distribution is solved, and the accuracy and the reliability of reward distribution are improved.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY

Training a reinforcement learning agent

PCT designated stageWO2025185839A1Artificial lifeMachine learningData setEngineering
There is provided a computer-implemented method (200) for training a reinforcement learning agent to adjust or maintain a parameter of a first cell of a communication network. The method (200) comprises obtaining (202) a dataset comprising a first state of the first cell at a first time instance, a second state of the first cell at a second time instance, a first action, wherein the first action, when co-occurring with the first state at any time instance, causes the first cell to transition to the second state, and a first reward value for performing the first action, determining (204), based on the dataset, a second action, wherein the second action, when co-occurring with the second state at any time instance, causes the first cell to transition to the first state, and a second reward value for performing the second action, obtaining (206) a first information comprising the first state, the second state, the first action, and the first reward value, obtaining (208) second information comprising the second state, the first state, the second action, and the second reward value, and training (210) the reinforcement learning agent based on the first information and the second information.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Large model routing method and device combined with reward model

The invention provides a large model routing method and device combined with a reward model, and relates to the technical field of artificial intelligence model dynamic routing, and the method comprises the steps: obtaining a question inputted by a user, and determining a candidate model set according to the user authority and a preset priority rule; performing sensitive information judgment on the question through a sensitive detection model, directly returning a preset answer if sensitive information is detected, and otherwise, entering the next step; carrying out reward evaluation on each model in the candidate model set based on a routing function, training the routing function through a knowledge distillation technology, and optimizing a model selection strategy by taking KL divergence as a loss function; and according to an evaluation result of the routing function, dynamically distributing the problem to a target model with the highest reward value for reasoning, and outputting a processing result of the target model. According to the method, the calculation overhead of large model integration can be effectively reduced, the intelligent level of routing decision is improved, and the system resource utilization rate and the response efficiency are remarkably improved.
Owner:PICC INFORMATION TECH CO LTD +1

Method and apparatus for resource joint allocation in heterogeneous network with hybrid access

The application discloses a resource joint allocation method and device in a human-machine-object hybrid access heterogeneous network. The method comprises the following steps: determining constraint conditions, decision variables and optimization targets of devices in the human-machine-object hybrid access heterogeneous network; defining a state set, an action set and a reward function of an intelligent agent in the devices based on the determined constraint conditions, decision variables and optimization targets; traversing the state set of the intelligent agent based on the defined state set, action set and reward function, updating the income of the current state of the intelligent agent according to the reward value of the current state of the intelligent agent and the income estimation of the next state until the maximum reward value is found; and obtaining a resource allocation scheme that optimizes the optimization target based on the maximum reward value. The application solves the technical problems of poor network performance and high service cost caused by poor resource allocation in the network.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Advertisement putting system and method based on exposure and conversion composite charging and combined with reward and intelligent optimization mechanism

The invention discloses an advertisement putting system and method based on exposure and conversion composite charging and combined with an award and intelligent optimization mechanism. The system comprises an advertiser input module (101), a proportion parameter distribution module (102), an exposure calculation module (103), a conversion calculation module (104), an automatic adding module (105), an advertiser reward module (106), a data storage module (107), an AI intelligent optimization module (108) and a system control module (109). After the advertiser sets the total exposure amount, the system divides the exposure into two parts: the first part adopts a CPM mode for charging, and the second part adopts a CPA mode for charging; when the CPA target is not achieved, the automatic dosing module continues to add exposure until the target is achieved or a maximum threshold is reached. The system calculates a reward value based on the interaction data and the satisfaction degree and returns the reward value to the advertiser, and the AI module dynamically adjusts the proportion and the threshold value, thereby achieving the optimization of the putting effect and the platform income. The method is suitable for an internet advertisement putting platform and a related optimization system.
Owner:SHIJIAZHUANG MONKEY NEWS INFORMATION TECHNOLOGY CO LTD

Reinforcement learning method and device for large language model

The embodiment of the specification provides a reinforcement learning method and device for a large language model. The method comprises: acquiring a base large language model with a thinking chain generation capability; driving the base large language model to generate a plurality of candidate replies for a target query statement, each candidate reply comprising thinking chain content and an answer; matching the thinking chain content in each candidate reply with a preset rule set in an artificial knowledge base, so as to determine the compliance degree of the thinking chain content with the rule set as a first reward value; determining a second reward value corresponding to the answer in the candidate reply; determining a comprehensive reward value of the candidate reply according to the first reward value and the second reward value; and updating parameters of the base large language model according to the comprehensive reward values corresponding to the plurality of candidate replies respectively. The thinking process of the model can be constrained, and the convergence speed of the model is improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Large language model optimization method and optimization device

The application relates to the technical field of artificial intelligence, and provides a large language model optimization method and an optimization device. The method obtains a candidate answer by inputting a question in a pre-training large language model, calculates a first reward value, a second reward value and a third reward value of the candidate answer, wherein the first reward value is used to represent the basic reward of the candidate answer, the second reward value is used to represent the degree of fit of the candidate answer and the question, and the third reward value is used to represent the alignment degree of the candidate answer and user preferences, then the first reward value, the second reward value and the third reward value are multi-head fused to obtain a target reward value, and finally the pre-training large language model is trained based on the target reward value to obtain an optimized large language model, the importance coefficients of the reward mechanisms of different levels of the large language model are adaptively learned in a multi-head layered manner, the model training efficiency is improved, and the generation quality of the trained model is improved.
Owner:北京衔远有限公司 +1

Multi-agent electromagnetic spectrum control attack and defense method based on deep reinforcement learning

The invention particularly relates to a multi-agent electromagnetic spectrum control attack and defense method based on deep reinforcement learning, which comprises the following steps: constructing a channel model following a discrete time two-state Markov chain, and carrying out electromagnetic spectrum control multi-agent deep reinforcement learning training based on a centralized training-decentralized execution architecture; in the attack part, a quantitative trigger is designed according to the channel signal receiving strength, an induction reward is utilized to implant a back door, the situation that higher modification cost and complexity are possibly caused by modifying a global reward is avoided, a feasible action scheme is selected for a poisoning action, and multiple agents select the same channel to collide when the back door is triggered; in a defense part, multi-agent deep reinforcement learning backdoor detection is performed based on an abnormal reward value, whether an abnormal reward mode exists or not is detected by analyzing reward distribution of a strategy in different states, and whether the reward value deviates from a normal range remarkably or not under a specific triggering condition is detected, so that a trigger is reconstructed, and an input filtering method is adopted to eliminate a backdoor.
Owner:XIDIAN UNIV

Question and answer model training method, answer content generation method, device and equipment

The application provides a training method and a content generation method, device and equipment of a question and answer model. The method comprises: obtaining sample question content; based on a plurality of reasoning strategies, outputting answer content corresponding to each reasoning strategy through a question and answer model; assigning a reward value to each answer content, the reward value comprising an adaptive reward value, wherein the higher the adaptive reward value is, the more the answer content is correct and the more the corresponding reasoning strategy meets the reasoning requirement; and adjusting the model parameters of the question and answer model based on the reward value of each answer content. Through the reinforcement learning mechanism, the question and answer model can autonomously learn to select a reasoning strategy that can ensure the correctness of the answer content and better meet the reasoning requirement according to the question content, thereby improving the efficiency of the question and answer model while ensuring the accuracy of the question and answer model.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Urban rail transit train energy saving optimization method, device and equipment and storage medium

The application discloses a kind of urban rail transit train energy-saving optimization method, device, equipment and storage medium, method includes S10, obtains the state information under the train operating environment of current time and reward value Reward value is obtained using reward function, reward function includes the first reward function in the DDPG model and the second reward function that the work of traction force is done in the train operation process and the punctuality of train operation are combined;S20, based on the state information and reward value under train operating environment, select operation action and issue to train to make the train next time according to operation action train;S30, determine next time as current time, repeat the execution step S10-S20.In the application, train executes operation action every time, environment will immediately feed back a state information and reward value, guide the following manipulation sequence, to update and optimize train operation strategy, finally obtain a convergent ideal train operation strategy, reach the purpose of energy saving.
Owner:ANHUI UNIV

A path planning method based on deep reinforcement learning-fast exploration random tree

The application discloses a path planning method based on deep reinforcement learning-fast exploration random tree, comprising the following steps: S1, obtaining a starting point and an ending point; S2, calculating candidate path points and selecting an action in an action state according to a Q value; S3, calculating a reward value and a new action state after the action is executed; S4, storing the action state, the action, the reward value and the new action state to an experience pool, in response to the number of stored experience values in the experience pool being greater than a batch size, randomly selecting experience values of the batch size, and updating the Q value and a time difference error through a policy network; S5, updating policy network parameters through a mean square error loss, and calculating a target network update step according to the time difference error; S6, judging whether a searched path reaches the ending point or satisfies a set maximum path point search number, if yes, outputting a current path, and if not, returning to S2. The application improves the search efficiency of the algorithm without increasing the search time of the algorithm.
Owner:SOUTHWEAT UNIV OF SCI & TECH +2