Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

106 results about "Reward value" patented technology

Value of Rewards is a Total Reward, Remuneration, Compensation & Benefits activity by René Broekhuis.

Robot adaptive training method and device based on reinforcement learning and medium

The invention relates to the technical field of robot training. The robot self-adaptive training method based on reinforcement learning comprises the steps that task sub-target information is generated through a high-level strategy network, the task sub-target information is input into a low-level execution network, an action control instruction is generated according to the task sub-target information, interaction feedback information is collected in the execution process, and the action control instruction is sent to a robot through a robot. Calculating a reward value according to the interaction feedback information, carrying out association processing on the reward value and the scene complexity parameter, executing a dynamic reward shaping operation, generating an adjusted reward signal, generating a strategy model optimized by meta-learning based on the adjusted reward signal, loading the strategy model in a simulation environment, and carrying out dynamic reward shaping. A target strategy model optimized through simulation training is generated, the target strategy model is loaded to the robot, and the robot is controlled to execute task operation in the actual interaction scene. The method has the effect of realizing adaptive task learning of the robot in a multi-interaction scene.
Owner:SEVEN (BEIJING) EDUCATION TECH CO LTD

Reinforced learning training method and system for relieving hallusion of multi-modal large model

The invention discloses a reinforcement learning training method and system for relieving illusion of a multi-modal large model, and belongs to the field of reinforcement learning training of a multi-modal large language model. Firstly, a planning and visual description generation step is introduced in an early stage to guide a model to perform structured reasoning, then a grouping relative strategy optimization algorithm is used, reward values are calculated for multiple candidate responses generated by the model after cold start, and particularly, a visual perception reward mechanism is set. The reward mechanism evaluates the consistency of the generated text description and the visual information by using an external large language model. Then, based on a vision description attention score advantage distribution method, learning of the model on key vision signals is dynamically enhanced, and the perception ability of the model on the vision signals is improved; and finally, the perception and reasoning performance of the model is further improved by adopting multiple rounds of rejection sampling and supervised fine tuning. The scheme does not depend on a model architecture, the extra overhead is small, the illusion problem caused by early image-text inconsistency is effectively solved, and the accuracy and the reliability are improved.
Owner:ZHEJIANG UNIV +1

Reward model training method and device, strategy model training method and device and electronic equipment

The invention provides a reward model training method and device, a strategy model training method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a preference data pair which comprises a preferred response and a non-preferred response generated for the same prompt word, and each of the preferred response and the non-preferred response is composed of a plurality of text unit sequences; inputting each text unit sequence into a to-be-trained reward model to obtain a predicted reward value; and calculating the total training loss according to the predicted reward value, and updating the model parameters of the to-be-trained reward model. According to the method, the response text is subjected to serialized splitting, and the preference data composed of the preferred response and the non-preferred response is introduced for comparative learning, so that the target of model training is no longer to evaluate the absolute quality of a single response, but to identify a key text unit which causes one response to be superior to the other response; the fine-grained evaluation of the response content is realized, and the evaluation accuracy of the reward model and the identification capability of complex user preferences are effectively improved.
Owner:IFLYTEK SOUTH CHINA ARTIFICIAL INTELLIGENCE RES INST GUANGZHOU CO LTD

Reinforcement learning training method and device for large language model

The embodiment of the invention provides a reinforcement learning training method and device for a large language model, and the method comprises the steps: inputting a first input text into the large language model, enabling the large language model to process the first input text based on a current strategy, and obtaining G output texts; respectively combining the G output texts with the first input text, and respectively inputting combination results into a trained reward model to obtain G reward values corresponding to the G output texts; the current strategy is updated based on a target function, the target function comprises a global penalty term, the global penalty term represents that a penalty value is applied only when the maximum value in the G reward values does not exceed a preset threshold value, and the penalty value is determined according to respective first output probabilities of the G output texts under the current strategy; therefore, the training of the large language model under the reinforcement learning framework is realized.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Reinforced learning training method and system based on multi-target structured reward function

The invention provides a reinforcement learning training method and system based on a multi-target structured reward function, and the method comprises the steps: an environment simulation layer comprises a physical engine and a blood vessel environment simulator, and is used for simulating the physical movement and interaction of a catheter in a blood vessel, and providing an environment state for a reward calculation layer; the reward calculation layer performs calculation through a plurality of independent calculation modules and finally generates a comprehensive reward value through a composite reward aggregator; the reinforcement learning layer learns and optimizes a navigation strategy according to the comprehensive reward value; and the monitoring evaluation layer monitors each performance index in the training process and adaptively adjusts the weight of each reward component in the reward calculation layer according to the index change. According to the invention, a major breakthrough of reinforcement learning in medical robot training is realized. Compared with a traditional single reward function method, the intelligent agent which is really thought by experts can be trained, and the method not only pursues to quickly reach a target, but also knows how to safely and stably complete operation.
Owner:SHANGHAI OPERATION ROBOT CO LTD

Reinforced learning training method and system for small-scale large language model

The invention relates to the technical field of reinforcement learning, in particular to a reinforcement learning training method and system for a small-scale large language model, and the method comprises the steps: obtaining user query and an external tool set, and constructing a tool integration reasoning training task; driving the large language model to generate thinking content and a tool calling instruction, and forming an interaction track in combination with an external tool execution result; respectively calculating a format reward, a tool calling correctness reward, a dynamic length reward and a distillation learning reward according to the interaction track, and fusing each reward item to obtain a total reward value; performing trajectory multiplexing and trajectory data augmentation processing on the interaction trajectory to generate an extended training data set; and based on the extended training data set and the total reward value, performing parameter iterative optimization on the to-be-trained small-scale large-language model by adopting a reinforcement learning strategy. According to the method, the reasoning capability and the structured tool calling capability of a small-scale large-language model in a complex tool integrated reasoning task can be stably improved.
Owner:GUSU LAB OF MATERIALS

Malware detection model training method and device based on itemized reward and operation index linkage

The invention discloses a malicious software detection model training method and device based on item reward and operation index linkage, and belongs to the technical field of artificial intelligence and malicious software detection. The method comprises the following steps: constructing a reinforcement learning model containing a policy network; the policy network is configured to receive software feature input and output multi-modal actions including classification results, structured interpretation text and detection confidence; based on the training sample, generating a multi-modal action by utilizing a strategy network, and calculating a composite reward value for the generated multi-modal action by utilizing a preset subitem reward; the subitem rewards at least comprise a structure consistency reward, a key element hit reward and a confidence coefficient calibration reward; and calculating advantage estimation based on the composite reward value to construct an objective function of a near-end policy optimization algorithm by using the advantage estimation, and maximizing the objective function by updating parameters of the policy network. The method can improve the explanatory and confidence of model output, and improves the recognition precision.
Owner:HARBIN ANTIY TECH

Question and answer model training method and device, electronic equipment and storage medium

The invention provides a question and answer model training method and device, electronic equipment and a storage medium. The method comprises the steps that multiple pieces of training data are acquired; in the process of training the question and answer model according to the sample question, performing repeated fragment identification on a reasoning process generated by performing multiple sampling on the sample question to determine a repeated penalty value of a repeated fragment, and performing reward value calculation on an answer generated by performing sampling on the sample question; performing fusion based on the repeated penalty value and the reward value to form a reward signal, performing normalization processing on the reward signal to obtain a dominant value of current sampling, and introducing the dominant value of the current sampling into a grouping strategy optimization algorithm to construct a loss function; and updating model parameters through gradient back propagation to inhibit repeated thinking behaviors of the question and answer model in the reasoning process, and continuing to perform iterative training on the updated question and answer model. More attention is paid to'effective thinking 'in the training process, and the ability of the question and answer model to solve complex problems is improved.
Owner:SHANGHAI ANXINCHENG NETWORK TECHNOLOGY CO LTD

Large model routing method and device combined with reward model

The invention provides a large model routing method and device combined with a reward model, and relates to the technical field of artificial intelligence model dynamic routing, and the method comprises the steps: obtaining a question inputted by a user, and determining a candidate model set according to the user authority and a preset priority rule; performing sensitive information judgment on the question through a sensitive detection model, directly returning a preset answer if sensitive information is detected, and otherwise, entering the next step; carrying out reward evaluation on each model in the candidate model set based on a routing function, training the routing function through a knowledge distillation technology, and optimizing a model selection strategy by taking KL divergence as a loss function; and according to an evaluation result of the routing function, dynamically distributing the problem to a target model with the highest reward value for reasoning, and outputting a processing result of the target model. According to the method, the calculation overhead of large model integration can be effectively reduced, the intelligent level of routing decision is improved, and the system resource utilization rate and the response efficiency are remarkably improved.
Owner:PICC INFORMATION TECH CO LTD +1

Advertisement putting system and method based on exposure and conversion composite charging and combined with reward and intelligent optimization mechanism

The invention discloses an advertisement putting system and method based on exposure and conversion composite charging and combined with an award and intelligent optimization mechanism. The system comprises an advertiser input module (101), a proportion parameter distribution module (102), an exposure calculation module (103), a conversion calculation module (104), an automatic adding module (105), an advertiser reward module (106), a data storage module (107), an AI intelligent optimization module (108) and a system control module (109). After the advertiser sets the total exposure amount, the system divides the exposure into two parts: the first part adopts a CPM mode for charging, and the second part adopts a CPA mode for charging; when the CPA target is not achieved, the automatic dosing module continues to add exposure until the target is achieved or a maximum threshold is reached. The system calculates a reward value based on the interaction data and the satisfaction degree and returns the reward value to the advertiser, and the AI module dynamically adjusts the proportion and the threshold value, thereby achieving the optimization of the putting effect and the platform income. The method is suitable for an internet advertisement putting platform and a related optimization system.
Owner:SHIJIAZHUANG MONKEY NEWS INFORMATION TECHNOLOGY CO LTD

Reinforcement learning method and device for large language model

The embodiment of the specification provides a reinforcement learning method and device for a large language model. The method comprises: acquiring a base large language model with a thinking chain generation capability; driving the base large language model to generate a plurality of candidate replies for a target query statement, each candidate reply comprising thinking chain content and an answer; matching the thinking chain content in each candidate reply with a preset rule set in an artificial knowledge base, so as to determine the compliance degree of the thinking chain content with the rule set as a first reward value; determining a second reward value corresponding to the answer in the candidate reply; determining a comprehensive reward value of the candidate reply according to the first reward value and the second reward value; and updating parameters of the base large language model according to the comprehensive reward values corresponding to the plurality of candidate replies respectively. The thinking process of the model can be constrained, and the convergence speed of the model is improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Multi-agent electromagnetic spectrum control attack and defense method based on deep reinforcement learning

The invention particularly relates to a multi-agent electromagnetic spectrum control attack and defense method based on deep reinforcement learning, which comprises the following steps: constructing a channel model following a discrete time two-state Markov chain, and carrying out electromagnetic spectrum control multi-agent deep reinforcement learning training based on a centralized training-decentralized execution architecture; in the attack part, a quantitative trigger is designed according to the channel signal receiving strength, an induction reward is utilized to implant a back door, the situation that higher modification cost and complexity are possibly caused by modifying a global reward is avoided, a feasible action scheme is selected for a poisoning action, and multiple agents select the same channel to collide when the back door is triggered; in a defense part, multi-agent deep reinforcement learning backdoor detection is performed based on an abnormal reward value, whether an abnormal reward mode exists or not is detected by analyzing reward distribution of a strategy in different states, and whether the reward value deviates from a normal range remarkably or not under a specific triggering condition is detected, so that a trigger is reconstructed, and an input filtering method is adopted to eliminate a backdoor.
Owner:XIDIAN UNIV

Question and answer model training method, answer content generation method, device and equipment

The application provides a training method and a content generation method, device and equipment of a question and answer model. The method comprises: obtaining sample question content; based on a plurality of reasoning strategies, outputting answer content corresponding to each reasoning strategy through a question and answer model; assigning a reward value to each answer content, the reward value comprising an adaptive reward value, wherein the higher the adaptive reward value is, the more the answer content is correct and the more the corresponding reasoning strategy meets the reasoning requirement; and adjusting the model parameters of the question and answer model based on the reward value of each answer content. Through the reinforcement learning mechanism, the question and answer model can autonomously learn to select a reasoning strategy that can ensure the correctness of the answer content and better meet the reasoning requirement according to the question content, thereby improving the efficiency of the question and answer model while ensuring the accuracy of the question and answer model.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Urban rail transit train energy saving optimization method, device and equipment and storage medium

The application discloses a kind of urban rail transit train energy-saving optimization method, device, equipment and storage medium, method includes S10, obtains the state information under the train operating environment of current time and reward value Reward value is obtained using reward function, reward function includes the first reward function in the DDPG model and the second reward function that the work of traction force is done in the train operation process and the punctuality of train operation are combined;S20, based on the state information and reward value under train operating environment, select operation action and issue to train to make the train next time according to operation action train;S30, determine next time as current time, repeat the execution step S10-S20.In the application, train executes operation action every time, environment will immediately feed back a state information and reward value, guide the following manipulation sequence, to update and optimize train operation strategy, finally obtain a convergent ideal train operation strategy, reach the purpose of energy saving.
Owner:ANHUI UNIV

A path planning method based on deep reinforcement learning-fast exploration random tree

The application discloses a path planning method based on deep reinforcement learning-fast exploration random tree, comprising the following steps: S1, obtaining a starting point and an ending point; S2, calculating candidate path points and selecting an action in an action state according to a Q value; S3, calculating a reward value and a new action state after the action is executed; S4, storing the action state, the action, the reward value and the new action state to an experience pool, in response to the number of stored experience values in the experience pool being greater than a batch size, randomly selecting experience values of the batch size, and updating the Q value and a time difference error through a policy network; S5, updating policy network parameters through a mean square error loss, and calculating a target network update step according to the time difference error; S6, judging whether a searched path reaches the ending point or satisfies a set maximum path point search number, if yes, outputting a current path, and if not, returning to S2. The application improves the search efficiency of the algorithm without increasing the search time of the algorithm.
Owner:SOUTHWEAT UNIV OF SCI & TECH +2

A traffic signal control method and system based on safety reinforcement learning

This invention discloses a traffic signal control method and system based on safety reinforcement learning. The method includes: acquiring multiple historical state information and processing them using an attention network, an actor network, a reward critic network, and a cost critic network to obtain a first probability value, a first reward value, a second reward value, a first cost value, and a second cost value; then obtaining an advantage evaluation value and a cost evaluation value through a reward advantage evaluation function and a cost advantage evaluation function; optimizing the actor network, reward critic network, and cost critic network based on the advantage evaluation value, cost evaluation value, first probability value, reward value, first reward value, cost value, and first cost value; repeating the above process until a preset number of times is exceeded to obtain a trained actor network; acquiring the observation information at the current moment and inputting it into the trained actor network to obtain the current action to control the traffic signal. This invention ensures the safety of intersections while maintaining efficiency.
Owner:BEIHANG UNIV

A global interface dynamic throttling method based on reinforcement learning

The application discloses a kind of global interface dynamic flow limiting methods based on reinforcement learning, comprising the following steps: S1, construct request simulator, and send simulated request to agent by request simulator;S2, collect state information, and make decision according to current state information by agent, establish action reward function;S3, randomly select action by agent, and get reward value according to selected action by action reward function, establish Q function;S4, train Q function by deep neural network, and get trained agent according to optimal Q function;S5, make decision to simulated request by trained agent.Interface flow limiting logic based on reinforcement learning can enable agent to learn whether intelligent decision is to limit flow or release current request under the condition that there is potential statistical rule in different load, different concurrent request and access sequence of interface.
Owner:SICHUAN COMPUTING CLOUD TECHNOLOGY CO LTD

A method and apparatus for calculating the reward value of a reinforcement learning acoustic decoy.

This application discloses a method and apparatus for calculating the reward value of acoustic decoys in reinforcement learning. The method establishes variables based on the target recognition processing time and target recognition capability of dipping sonar, providing the target deception success rate and the average recognition processing time after the acoustic decoy appears. This allows for rapid calculation of the reward value, avoiding the problem of long processing times in Monte Carlo methods that affect the training speed of reinforcement learning agents. In some examples, the target success rate is weighted by considering the processing time required for subsequent actions such as attack and recognition after target recognition by an aerial platform, resulting in a more accurate calculation of the acoustic decoy reward value, thus avoiding the problem of long processing times in Monte Carlo methods that affect the training speed of reinforcement learning agents.
Owner:CHINA ACADEMY OF ELECTRONICS AND INFORMATION TECHNOLOGY OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION

Training method and device of control policy network for controlling robot double-arm end, equipment, robot and medium based on reinforcement learning

The present disclosure provides a training method and device of a control policy network for controlling the end of a robot double arm based on reinforcement learning, an equipment, a robot and a medium, relating to the technical field of robots. The method comprises: based on the target pose data and the current state data for training the control policy, calculating the position error and the orientation error of the first arm end, and the position error and the orientation error of the second arm end; calculating the first reward term according to the position error of the first arm end and the position error of the second arm end respectively, and calculating the second reward term according to the orientation error of the first arm end and the orientation error of the second arm end respectively, the first reward term and the second reward term are both constructed by an exponential function; based on the product of the first reward term and the second reward term plus a constant term, calculating the comprehensive reward value and taking the comprehensive reward value as the training signal to train the control policy network. The method provided by the present disclosure can avoid the problem of significant decline in control accuracy on one side caused by resource competition.
Owner:SHENZHEN ZHUJI POWER TECH CO LTD

Strategy model training method, device, equipment and program product

The invention relates to a strategy model training method and device, equipment and a program product. The method comprises the following steps: according to an input task instruction, obtaining a current reasoning track when a strategy model executes a target task indicated by the task instruction; determining a tool use reward value of the current reasoning track based on the current tool calling frequency in the current reasoning track; determining a total reward value of the current reasoning trajectory according to the tool use reward value; wherein the tool use reward value is in negative correlation with the tool calling difference, and the tool calling difference is related to the tool calling frequency in the historical reasoning trajectory of the strategy model for the target task; and updating the model according to the total reward value. Through the technical scheme of the invention, the strategy model can be guided to inhibit unnecessary tool calling behaviors, the reasoning efficiency and the resource utilization rate of the strategy model are improved, and the balance between the reasoning efficiency of the strategy model and the accuracy rate of an execution result is realized.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Question and answer method and device based on large language model, equipment and medium

The embodiment of the invention provides a question answering method and device based on a large language model, equipment and a medium. In the application, for each reasoning step of each answer, a reward value reflecting the quality of the reasoning step is set for the reasoning step, a weight quantifying the relative influence degree of the reasoning step on the quality of the answer is set for the reasoning step, and an information entropy representing the uncertainty degree of the reward value of the reasoning step is calculated for the reasoning step. On the basis, the reward value is cooperatively adjusted through the weight and the information entropy, and a relative advantage value capable of accurately measuring the relative value of the overall reasoning process of the answer is obtained. Relative advantage values of the M answers are used as optimization signals, and the large language model is guided to be optimized in the direction of focusing on a reasoning path for improving the accuracy of the key reasoning step and inhibiting uncertainty interference, so that the accuracy of the reasoning process of the large language model in a question and answer scene is improved, and the accuracy and reliability of generated answers are guaranteed.
Owner:NEW H3C TECH CO LTD

Precise information delivery system and method based on multi-modal perception and reinforcement learning

The invention relates to the technical field of artificial intelligence, and discloses an accurate information delivery system and method based on multi-modal perception and reinforcement learning, and the method comprises the steps: when an evaluation window ends, the system calls a content category identifier in a user behavior sequence to calculate an information entropy value, and carries out the statistics of a negative feedback rate; the system generates a long-term reward value which is in negative correlation with the information entropy value and the negative feedback rate, and the reinforcement learning model is updated in a backtracking manner in combination with the instant reward, so that the technical problem of reward index dislocation in reinforcement learning is solved, the optimization target of the model is converted into a low-entropy concentration state of cultivating a user from an induced click state, and the user experience is improved. Therefore, the degradation of the model strategy is avoided.
Owner:SHANGHAI APOKER NETWORK TECHNOLOGY CO LTD

A reinforcement learning-based adaptive training method, device, and medium for robots.

This invention relates to the technical field of robot training. Specifically, it relates to a robot adaptive training method, apparatus, and medium based on reinforcement learning. The method includes generating task sub-objective information using a high-level policy network, inputting the task sub-objective information into a low-level execution network, generating action control commands based on the task sub-objective information, collecting interaction feedback information during execution, calculating a reward value based on the interaction feedback information, associating the reward value with scene complexity parameters, performing dynamic reward shaping operations to generate an adjusted reward signal, generating a policy model optimized by meta-learning based on the adjusted reward signal, loading the policy model into a simulation environment to generate a target policy model optimized by simulation training, loading the target policy model into the robot, and controlling the robot to perform task operations in an actual interactive scenario. This invention achieves the effect of adaptive task learning for robots in multi-interaction scenarios.
Owner:SEVEN (BEIJING) EDUCATION TECH CO LTD

Robot control reinforcement learning method and device based on dynamic potential energy reward

The application relates to the technical field of computer science, in particular to a robot control reinforcement learning method and device based on dynamic potential energy reward, wherein the method comprises the following steps: obtaining robot information and initial state information of a robot; based on the robot information and the initial state information, performing at least one round of robot control interaction on the robot; in each round of control interaction, based on a control instruction and state information of the robot, calculating a dense reward value of the robot constructed based on dynamic potential energy; and combining the dense reward value constructed based on the dynamic potential energy and training data generated in the control interaction process to train an initial control strategy neural network of the robot, so as to obtain a final control strategy neural network. Thus, the problems in the prior art that the potential energy function is fixed, the reward obtained by each state in the whole training process is determined, the strategy exploration is limited, and the robot cannot be controlled in a complex manner are solved.
Owner:WUHAN UNIV

Reinforcement learning training method and system for small-scale large language model

The application relates to the technical field of reinforcement learning, in particular to a reinforcement learning training method and system for a small-scale large language model, which comprises the following steps: acquiring a user query and an external tool set and constructing a tool integrated reasoning training task; driving a large language model to generate thinking content and tool calling instructions, and combining external tool execution results to form an interaction track; respectively calculating a format reward, a tool calling correctness reward, a dynamic length reward and a distillation learning reward for the interaction track, and fusing the rewards to obtain a total reward value; performing track reuse and track data augmentation processing on the interaction track to generate an extended training data set; and based on the extended training data set and the total reward value, performing parameter iterative optimization on the small-scale large language model to be trained by using a reinforcement learning strategy. The application can enable the small-scale large language model to stably improve reasoning capability and structured tool calling capability in a complex tool integrated reasoning task.
Owner:GUSU LAB OF MATERIALS

Data interface shunting method and device, computer equipment and storage medium

The invention relates to a data interface shunting method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a target state quantity; updating the state quantity of the corresponding target data interface based on the target state quantity to obtain all data interface configuration rules; executing the following query steps: obtaining a query request; distributing the query request to different data interfaces based on a configuration rule, and receiving query data, reward values and new state quantities returned by the data interfaces; and taking the new state quantity as the interface state quantity of the next round of query, repeatedly executing the query step, and adjusting the interface state quantity of the next round according to the reward value of each round, so that the reward value of each round is located in a preset reward value interval.
Owner:CHINA TELECOM CLOUD TECH CO LTD

Order processing method and device, electronic equipment and medium

The invention provides an order processing method and device, electronic equipment and a medium, and relates to the technical field of online car-hailing, and the method comprises the steps: determining an abnormal reward value of a target user according to a reward value in an abnormal historical order of the target user; determining the user level corresponding to the numerical range to which the abnormal reward value of the target user belongs according to the corresponding relationship between the user level and the numerical range; if the order of the target user is received, processing the order of the target user according to a limiting strategy corresponding to the user level of the target user; wherein the limiting strategy is a strategy for limiting the target user to obtain a normal reward in the order. According to the embodiment of the invention, limitation can be set for the user according to the abnormal reward value obtained based on the abnormal historical order, so that the problem that the user defrauds coupons can be relieved, normal orders are easier to send normally, and the operation efficiency of online car-hailing is improved.
Owner:NANJING LINGXING TECH CO LTD

Information processing apparatus, information processing method, and information processing program

To make a robot model at the time of simulation perform a more efficient action by using reinforcement learning.SOLUTION: The information processing device includes a target information acquisition section configured to acquire, as target information, a target position of a tip portion of a robot model and a target angle of the tip portion at the target position in a simulation, a current information acquisition section configured to acquire, as current information, a current position of the tip portion and a current angle of the tip portion, an intermediate reward calculation section configured to calculate, as an immediate reward value, a reward to be given to an agent each time the robot model acts, based on a distance from the current position of the tip portion to the target position and an angle difference between the current angle of the tip portion and the target angle of the tip portion at the target position, and calculate, as an intermediate reward value, an integrated value of the immediate reward values from a start of the simulation to a current time, and a reinforcement learning section configured to learn a learning model such that the intermediate reward value at an end of the simulation is maximized.SELECTED DRAWING: Figure 1
Owner:TOYOTA PRODN ENG CORP +1

A training method and device of a task allocation system and a readable storage medium

The application provides a training method and device of a task allocation system and a readable storage medium. The method comprises: modeling a task allocation scene of an ISCC to obtain an ISCC task allocation system; repeatedly executing a target process until a first reward value of a first reward function and a second reward value of a second reward function respectively reach maximum values; the target process comprises: allocating a high-level task to a low-level agent according to a high-level agent algorithm; determining resource information contributed by the low-level agent for a target task, sending feedback information to the high-level agent, and obtaining the second reward value, updating a second network parameter in the low-level agent algorithm according to the second reward value; obtaining the first reward value, updating a first network parameter in the high-level agent algorithm according to the first reward value; and the first reward value is determined according to the feedback information of the low-level agent.
Owner:CHINA MOBILE COMM LTD RES INST +1

Hardware-in-the-loop test case generation method and device, equipment, medium and product

The embodiment of the invention provides a hardware-in-the-loop test case generation method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a test requirement; inputting the test requirements into the reinforcement learning model to obtain a test case set; the test case set is used for representing a test scene set corresponding to the test requirement; for each test case in the test case set, running the test case, and generating a test response; according to the test case and the test response, calculating to obtain a reward value; and if the reward value is smaller than a preset threshold value, performing iterative training on the reinforcement learning model according to the reward value until the reward value is not smaller than the preset threshold value, and outputting a test case set corresponding to the reward value not smaller than the preset threshold value. According to the scheme, the comprehensiveness of scene coverage of the generated test case is improved.
Owner:STARRY SKY PLAN (SHANGHAI) AUTOMOBILE TECHNOLOGY CO LTD