Method for model training, method for controlling unmanned device, and device

By obtaining the historical status data of unmanned equipment and obstacles, using long-term memory networks and attention mechanism networks to train decision models, optimize control parameters, the problem of unmanned equipment avoiding obstacle collisions during driving is solved, and safe and efficient driving is achieved.

CN115047864BActive Publication Date: 2025-07-22BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210161211.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-07-22
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

During the driving process, existing unmanned equipment has weak generalization capabilities for decision-making models that rely on human driving data training, and fails to effectively utilize the historical interaction information of obstacles, resulting in a low success rate of avoiding obstacles and insufficient safety.

Method used

By obtaining historical state data of designated devices and obstacles, using long and short-term memory networks to predict obstacle states, combining attention mechanism networks to determine data of interest, training decision models to optimize control parameters, and adjusting models to improve the rationality of control parameters using reward values.

Benefits of technology

It improves the probability of unmanned equipment avoiding obstacle collisions during driving, ensures the safe driving of the equipment, and improves the traffic efficiency and stability of the driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115047864B_ABST
    Figure CN115047864B_ABST
Patent Text Reader

Abstract

This specification discloses a method for model training, a method and device for controlling an unmanned device. First, historical state data is obtained. Secondly, the historical state data is input into a pre-trained long short-term memory network according to a time series to predict the predicted state data of each obstacle after a set historical moment. Then, the historical state data and the predicted state data are input into a decision model to be trained, and through the weights corresponding to the attention mechanism network, the data of interest is determined from the environmental data of the environment where the specified device is located at the set historical moment. Finally, the reward value corresponding to the specified device driving according to the control parameters at the set historical moment is determined, and the decision model is trained. This method can determine the data of interest from the environmental data of the environment where the specified device is located at the set historical moment, and through the reward value, measure the reasonableness of the determined control parameters, thereby effectively ensuring the safe driving of the specified device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of driverless technology, and particularly to a method for model training, a method for controlling a unmanned device, and a device therefor. Background Art

[0002] Currently, real human driving data is usually used to train a decision-making model, and a Long Short-Term Memory (LSTM) network is added to the decision-making model. According to the driving decisions determined at each historical moment by a specified device, that is, control parameters, the control parameters corresponding to the specified device at a future moment are predicted. This method relies on expert experience, requires a large amount of sample data, has weak migration ability and weak generalization ability. Moreover, although the Long Short-Term Memory network has a memory function, in the process of reinforcement learning decision-making, the historical driving decisions of the specified device are not important. Therefore, in practical applications, the success rate of avoiding obstacles by the control parameters corresponding to the specified device at a future moment predicted by the above method is not high, there is a possibility of collision with other surrounding obstacles, and the safety is relatively low.

[0003] Therefore, how a unmanned device plans a reasonable driving trajectory according to the interaction situation of surrounding traffic participants is an urgent problem to be solved. Summary of the Invention

[0004] This specification provides a method for model training, a method for controlling a unmanned device, and a device therefor, to partially solve the above problems existing in the prior art.

[0005] This specification adopts the following technical solutions:

[0006] This specification provides a method for model training, including:

[0007] Obtaining historical state data corresponding to a specified device and each obstacle at each historical moment;

[0008] Inputting the historical state data into a pre-trained Long Short-Term Memory network according to a time series to predict the state data of each obstacle after a set historical moment as predicted state data;

[0009] Inputting the historical state data and the predicted state data into a decision-making model to be trained, and determining interesting data from the environmental data of the environment where the specified device is located at the set historical moment through the weights corresponding to an attention mechanism network;

[0010] Determining the control parameter corresponding to the specified device at the set historical moment according to the interesting data;

[0011] Determine a reward value corresponding to the specified device traveling according to the control parameter at the set historical moment based on the data of interest and the control parameter, and train the decision-making model according to the reward value.

[0012] Optionally, the decision-making model includes: an evaluation sub-model;

[0013] Determining a reward value corresponding to the specified device traveling according to the control parameter at the set historical moment based on the data of interest and the control parameter includes:

[0014] Input the data of interest and the control parameter into the evaluation sub-model, predict a reward value corresponding to the specified device traveling according to the control parameter at the set historical moment as the reward value to be optimized, and determine an actual reward value corresponding to the specified device traveling according to the control parameter at the set historical moment based on the data of interest and the control parameter;

[0015] Training the decision-making model according to the reward value includes:

[0016] Train the decision-making model with the goal of approximating the actual reward value with the reward value to be optimized.

[0017] Optionally, the decision-making model includes: a target evaluation sub-model, a target policy sub-model;

[0018] Determining an actual reward value corresponding to the specified device traveling according to the control parameter at the set historical moment based on the data of interest and the control parameter includes:

[0019] Determine a reward value in the environment where the specified device is located at the set historical moment based on the data of interest and the control parameter as the reward value corresponding to the set historical moment;

[0020] Through the attention mechanism network, predict data of interest after the specified device travels according to the control parameter at the set historical moment as predicted data of interest, and input the predicted data of interest into the target policy sub-model to determine a control parameter of the specified device after the set historical moment as a predicted control parameter;

[0021] Input the predicted data of interest and the predicted control parameter into the target evaluation sub-model to determine a predicted reward value of the specified device after the set historical moment;

[0022] Determine the actual reward value based on the reward value corresponding to the set historical moment and the predicted reward value.

[0023] Optionally, the target evaluation sub-model includes: a first target evaluation sub-model, a second target evaluation sub-model;

[0024] Inputting the predicted interesting data and the predicted control parameters into the target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment, including:

[0025] Inputting the predicted interesting data and the predicted control parameters into the first target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the first candidate reward value, and inputting the predicted interesting data and the predicted control parameters into the second target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the second candidate reward value;

[0026] Taking the smaller value of the first candidate reward value and the second candidate reward value as the predicted reward value.

[0027] Optionally, determining the reward value of the specified device in the environment at the set historical moment according to the interesting data and the control parameters, including:

[0028] Determining a first influence factor corresponding to the set historical moment according to the interesting data and the control parameters;

[0029] Determining the reward value of the specified device in the environment at the set historical moment according to the first influence factor, where the first influence factor is used to characterize the time difference between the moment when the specified device reaches the specified point and the moment when each obstacle reaches the specified point, and the greater the time difference, the greater the reward value of the specified device in the environment at the set historical moment.

[0030] Optionally, determining the reward value of the specified device in the environment at the set historical moment according to the interesting data and the control parameters specifically includes:

[0031] Determining a second influence factor corresponding to the set historical moment according to the interesting data and the control parameters;

[0032] Determining the reward value of the specified device in the environment at the set historical moment according to the second influence factor, where the second influence factor is used to characterize the traffic efficiency when the specified device travels according to the control parameters, and the greater the traffic efficiency, the greater the reward value of the specified device in the environment at the set historical moment.

[0033] Optionally, according to the data of interest and the control parameters, determine the reward value of the specified device in the environment at the set historical moment, specifically including:

[0034] According to the data of interest and the control parameters, determine the third influence factor corresponding to the set historical moment;

[0035] According to the third influence factor, determine the reward value of the specified device in the environment at the set historical moment. The third influence factor is used to characterize the degree of state change of the specified device after driving according to the control parameters. The greater the degree of state change, the smaller the reward value of the specified device in the environment at the set historical moment.

[0036] Optionally, the evaluation sub-model includes: a first evaluation sub-model and a second evaluation sub-model;

[0037] Input the data of interest and the control parameters into the evaluation sub-model, and predict the corresponding reward value of the specified device after driving according to the control parameters at the set historical moment as the reward value to be optimized, specifically including:

[0038] Input the data of interest and the control parameters into the first evaluation sub-model, and predict the corresponding reward value of the specified device after driving according to the control parameters at the set historical moment as the first reward value to be optimized;

[0039] Input the data of interest and the control parameters into the second evaluation sub-model, and predict the corresponding reward value of the specified device after driving according to the control parameters at the set historical moment as the second reward value to be optimized;

[0040] Taking the approximation of the reward value to be optimized to the actual reward value as the optimization goal, train the decision model, specifically including:

[0041] Taking the approximation of the first reward value to be optimized to the actual reward value as the optimization goal, train the first evaluation sub-model in the decision model, and taking the approximation of the second reward value to be optimized to the actual reward value as the optimization goal, train the second evaluation sub-model in the decision model.

[0042] Optionally, the decision model includes: a policy sub-model and a target policy sub-model;

[0043] Taking the approximation of the reward value to be optimized to the actual reward value as the optimization goal, train the decision model, specifically including:

[0044] For each round of training, with the approximation of the first reward value to be optimized to the actual reward value as the optimization objective, based on the first parameter update step size, update the model parameters of the first evaluation sub-model in this round of training, and, with the approximation of the second reward value to be optimized to the actual reward value as the optimization objective, based on the first parameter update step size, update the model parameters of the second evaluation sub-model in this round of training;

[0045] According to the model parameters of the first evaluation sub-model in this round of training, the model parameters of the second evaluation sub-model in this round of training, and the second parameter update step size, update the model parameters of the policy sub-model in this round of training;

[0046] According to the model parameters of the policy sub-model in this round of training and the soft update coefficient, update the model parameters of the target policy sub-model in this round of training until the preset condition is met to complete the training of the target policy sub-model in the decision-making model.

[0047] This specification provides a control method for an unmanned device, including:

[0048] Obtain the state data of the unmanned device and each obstacle at the current moment as the current state data;

[0049] Input the current state data into a pre-trained long short-term memory network to predict the state data of each obstacle after the current moment as the predicted state data;

[0050] Input the current state data and the predicted state data into the trained decision-making model to determine the control parameters of the unmanned device at the current moment, and the decision-making model is obtained by the above model training method;

[0051] Control the unmanned device according to the control parameters of the unmanned device at the current moment.

[0052] This specification provides a model training device, including:

[0053] An acquisition module for acquiring the historical state data of a specified device and each obstacle at each historical moment;

[0054] A prediction module for inputting the historical state data into a pre-trained long short-term memory network according to the time series to predict the state data of each obstacle after the set historical moment as the predicted state data;

[0055] An input module, configured to input the historical state data and the predicted state data into a decision model to be trained, so as to determine, based on the weights corresponding to an attention mechanism network, the data of interest from the environmental data of the specified device in the environment at the set historical moment;

[0056] A determination module, configured to determine, based on the data of interest, the control parameter corresponding to the specified device at the set historical moment;

[0057] A training module, configured to determine, based on the data of interest and the control parameter, the reward value corresponding to the specified device after traveling according to the control parameter at the set historical moment, and train the decision model based on the reward value.

[0058] This specification provides a control device for an unmanned device, including:

[0059] An acquisition module, configured to acquire the state data of the unmanned device and each obstacle at the current moment as the current state data;

[0060] A prediction module, configured to input the current state data into a pre-trained long short-term memory network to predict the state data of each obstacle after the current moment as the predicted state data;

[0061] A determination module, configured to input the current state data and the predicted state data into a trained decision model to determine the control parameter corresponding to the unmanned device at the current moment, where the decision model is trained by the method for training the above model;

[0062] A control module, configured to control the unmanned device according to the control parameter corresponding to the unmanned device at the current moment.

[0063] This specification provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the method for training the above model or the control method for the unmanned device is implemented.

[0064] This specification provides an unmanned device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method for training the above model or the control method for the unmanned device is implemented.

[0065] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:

[0066] In the method for model training provided in this specification, historical state data corresponding to a specified device and each obstacle at each historical moment is obtained. Secondly, the historical state data is input into a pre-trained long short-term memory network according to the time series to predict the state data of each obstacle after a set historical moment as the predicted state data. Then, the historical state data and the predicted state data are input into a decision model to be trained, and through the weights corresponding to the attention mechanism network, interesting data is determined from the environmental data of the environment where the specified device is located at the set historical moment. Then, according to the interesting data, the control parameters corresponding to the specified device at the set historical moment are determined. Finally, according to the interesting data and the control parameters, the reward value corresponding to the specified device after driving according to the control parameters at the set historical moment is determined, and the decision model is trained according to the reward value.

[0067] As can be seen from the above method for model training, this method can input historical state data into a pre-trained long short-term memory network according to the time series to determine the predicted state data corresponding to each obstacle. And through the weights corresponding to the attention mechanism network in the decision model, interesting data is determined from the environmental data of the environment where the specified device is located at the set historical moment. Then, according to the interesting data and the control parameters, the reward value corresponding to the specified device after driving according to the control parameters at the set historical moment is determined to measure the reasonableness of the determined control parameters. By training the decision model in this way, the probability of the specified device colliding with surrounding obstacles can be reduced, avoiding collisions with surrounding obstacles, and thus effectively ensuring the safe driving of the specified device. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and the illustrative embodiments and descriptions thereof are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:

[0069] Figure 1 It is a schematic flowchart of the method for model training provided by an embodiment of this specification;

[0070] Figure 2 It is a schematic diagram of the model structure of the decision model provided by an embodiment of this specification;

[0071] Figure 3 It is a schematic flowchart of the control method for an unmanned device provided by an embodiment of this specification;

[0072] Figure 4 It is a schematic diagram of the structure of the model training device provided by an embodiment of this specification;

[0073] Figure 5Schematic diagram of the control device for the unmanned device provided in the embodiments of this specification;

[0074] Figure 6 Schematic diagram of the unmanned device provided in the embodiments of this specification. Detailed implementation manners

[0075] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0076] The following will detail the technical solutions provided in the embodiments of this specification with reference to the drawings.

[0077] In the embodiments of this specification, before determining the control parameters according to the current state data corresponding to the specified device, it is necessary to rely on a pre-trained decision model. The following will first introduce the process of training the decision model, as Figure 1 shown.

[0078] Figure 1 Schematic flowchart of the method for model training provided in the embodiments of this specification, which specifically includes the following steps:

[0079] S100: Obtain the historical state data corresponding to the specified device and each obstacle at each historical moment.

[0080] The execution subject of the method for training the decision model provided in this specification can be the specified device or an electronic device such as a server or a desktop computer. For ease of description, the method for training the model provided in this specification will be described below only with the server as the execution subject.

[0081] In the embodiments of this specification, the server can obtain the historical state data corresponding to the specified device and each obstacle at each historical moment. Among them, the specified device can be equipped with a variety of sensors, for example, cameras, lidars, millimeter-wave radars, satellite positioning systems, etc., to sense the environment around the specified device during driving and obtain the required state data. The obstacles mentioned here can refer to moving objects such as vehicles, bicycles, and pedestrians around the specified device during its movement, that is, obstacles that can interfere with the movement of the specified device, and the obstacles can also refer to stationary objects such as trees and buildings around the specified device during its movement.

[0082] The designated device mentioned here can refer to a device specifically used for data collection, such as a human-driven vehicle, a manned robot, etc., or it can also refer to an unmanned device.

[0083] In the embodiments of this specification, the obtained status data may include: the position data of the designated device, as well as the position data of each obstacle around the designated device, the speed data of the designated device, as well as the speed data of each obstacle around the designated device, the steering angle data of the designated device, etc.

[0084] Of course, the server can determine the relative position data between the designated device and each obstacle around the designated device, the relative speed data between the designated device and each obstacle around the designated device, etc. according to the obtained status data.

[0085] Furthermore, during the movement of the designated device, there may be multiple obstacles around it. Therefore, the designated device can collect and obtain the status data of each of these obstacles for each obstacle around it.

[0086] The unmanned device mentioned in this specification can refer to devices such as unmanned vehicles, robots, and automatic delivery devices that can achieve autonomous driving. Based on this, the unmanned device applying the model training method provided in this specification can be used to perform delivery tasks in the delivery field, such as business scenarios of using unmanned devices for express delivery, logistics, food delivery, etc.

[0087] S102: Input the historical status data into a pre-trained long short-term memory network according to the time series to predict the status data of each obstacle after a set historical moment as the predicted status data.

[0088] In practical applications, usually through the long short-term memory network (Long Short-Term Memory, LSTM), according to the driving decisions determined by the designated device at each historical moment, that is, the control parameters, the control parameters corresponding to the designated device at future moments are predicted. However, in the model training process under the reinforcement learning framework, compared with the control parameters determined by the designated device at each historical moment, the interaction between the designated device and each obstacle in the surrounding environment is more important for determining the control parameters corresponding to future moments. Based on this, the server can predict the status data of each obstacle at future moments through the long short-term memory network according to the status data corresponding to the designated device and each obstacle at each historical moment.

[0089] In the embodiments of this specification, the server can input the historical status data into a pre-trained long short-term memory network according to the time series to predict the status data of each obstacle after a set historical moment as the predicted status data.

[0090] Specifically, the server can sort the historical state data in chronological order, select the historical state data based on a preset time window, and input it into a pre-trained long short-term memory network to predict the state data of each obstacle after a set historical moment. The set historical moment mentioned here can refer to the current moment. For example, the server can input the historical state data within the last five seconds before the current moment into the pre-trained long short-term memory network to predict the state data of each obstacle within the next second.

[0091] Further, for each obstacle, the historical speed data corresponding to the obstacle sorted in chronological order can be expressed as vel i =(v t1 ,…,v tn ), vel i can represent the speed data sequence of the i-th obstacle, and v tn can represent the speed data corresponding to the i-th obstacle at time tn. The historical speed data corresponding to each obstacle sorted in chronological order can be expressed as Vel=(vel i ,…,vel n ), where n is the number of obstacles.

[0092] Then, the server can input the historical speed data corresponding to each obstacle into the pre-trained long short-term memory network in time series and output the speed data of each obstacle at a future moment. The speed data of each obstacle at a future moment can be expressed as H=(v p1 ,…,v pn ), and v pn can represent the speed data of the n-th obstacle at a future moment.

[0093] It should be noted that during the model training process of the long short-term memory network, the set historical moment mentioned here can refer to the moment with the latest time series among the historical state data selected according to a preset time window.

[0094] In the embodiments of this specification, there can be various training methods for the long short-term memory network. For example, obtain the historical state data of the obstacle sorted in chronological order, predict the state data of the obstacle after a set historical moment based on a preset time window, and train the long short-term memory network by minimizing the deviation between the predicted state data of the obstacle after the set historical moment and the true state data of the obstacle after the set historical moment. This specification does not limit the specific form of the training method of the long short-term memory network.

[0095] S104: Input the historical state data and the predicted state data into a decision model to be trained, and determine the data of interest from the environmental data of the environment where the specified device is located at the set historical moment through the weights corresponding to the attention mechanism network.

[0096] In practical applications, usually the historical state data corresponding to all obstacles around the specified device is input into the decision model. However, not all obstacles will affect the next driving decision of the specified device. For example, during the process of the specified device changing lanes, it mainly considers the state data of the obstacles on the current lane and the target lane to be changed, and determines the control parameters of the specified device in the next period of time. That is to say, the server pays more attention to the state data that can affect the lane change of the specified device, and the influence of the obstacles on other lanes on the control parameters of the specified device in the next period of time is relatively small or has no influence.

[0097] Therefore, the server needs to determine which state data in the historical state data corresponding to each obstacle has a greater influence on the control parameters of the specified device in the next period of time, and ignore the state data with a relatively small influence on the control parameters of the specified device in the next period of time.

[0098] In the embodiments of this specification, the server can input the historical state data and the predicted state data into a decision model to be trained, and determine the data of interest from the environmental data of the environment where the specified device is located at the set historical moment through the weights corresponding to the attention mechanism network. The data of interest mentioned here can be used to represent the data with a greater influence on the control parameters determined by the specified device among the state data corresponding to each obstacle. The specific formula for determining the weights corresponding to the attention mechanism network is as follows:

[0099] e t =tanh(W[CurVel,CurPos,v pt +b)

[0100] In the above formula, e t represents the data of interest determined by the t-th obstacle through the attention mechanism network. tanh() can be used to represent the activation function, making the output value between 0 and 1. W[] can be used to represent the weights corresponding to the attention mechanism network, that is, the network parameters corresponding to the attention mechanism network. CurVel can be used to represent the historical speed data of the specified device and each obstacle. CurPos can be used to represent the historical position data of the specified device and each obstacle. v pt can be used to represent the predicted state data of each obstacle at the future moment. b can be used to represent the bias.

[0101] As can be seen from the above formula, for each obstacle, the server can, through the training of the decision-making model, adjust the weights corresponding to the attention mechanism network, and determine the state data in the historical state data corresponding to the obstacle that has a greater impact on the control parameters of the specified device in a future period of time, so that the specified device can determine more accurate control parameters.

[0102] S106: Determine the control parameters corresponding to the specified device at the set historical moment according to the data of interest.

[0103] In the embodiments of this specification, the server can determine the control parameters corresponding to the specified device at the set historical moment according to the data of interest.

[0104] Specifically, the server can use its own speed change amount and its own steering angle change amount as control parameters. The specific formula is as follows:

[0105]

[0106] In the above formula, Δv can be used to represent the speed change amount of the specified device. can be used to represent the steering angle change amount of the specified device. Among them, the change amount is determined by the speed or steering angle of the specified device at the current moment and the speed or steering angle corresponding to the specified device at the previous moment. For Δv, if Δv is greater than 0, it can be used to represent the increase value of the speed of the specified device at the current moment compared with the previous moment. If Δv is greater than 0, it can be used to represent the decrease value of the current speed of the specified device compared with the previous moment. For If is greater than 0, it can be used to represent the angle of the steering angle of the specified device turning to the left at the current moment compared with the previous moment. If is less than 0, it can be used to represent the angle of the steering angle of the specified device turning to the right at the current moment compared with the previous moment.

[0107] It should be noted that the speed change amount and steering angle change amount of the specified device can be optimized through the training and update of the decision-making model to better control the specified device to drive.

[0108] S108: Determine the reward value corresponding to the specified device after driving according to the control parameter at the set historical moment according to the data of interest, and train the decision-making model according to the reward value.

[0109] In the embodiments of this specification, the server can determine the reward value corresponding to the specified device after driving according to the control parameter at the set historical moment according to the data of interest, and train the decision-making model according to the reward value.

[0110] In the embodiments of this specification, the decision-making model includes: an evaluation sub-model. The server can input the data of interest and control parameters into the evaluation sub-model to predict the reward value corresponding to the specified device after driving according to the control parameters at the set historical moment, as the reward value to be optimized, and determine the actual reward value corresponding to the specified device after driving according to the control parameters at the set historical moment based on the data of interest and control parameters.

[0111] Then, taking the approximation of the reward value to be optimized to the actual reward value as the optimization goal, the decision-making model is trained.

[0112] Furthermore, the decision-making model includes: a policy sub-model. The server can input the data of interest and control parameters into the evaluation sub-model to determine the predicted reward value of the specified device at the set historical moment.

[0113] Secondly, the server can predict the data of interest after the specified device drives according to the control parameters at the set historical moment through the attention mechanism network as the predicted data of interest, and input the predicted data of interest into the policy sub-model to determine the control parameters after the set historical moment of the specified device as the predicted control parameters.

[0114] Then, the server can input the predicted data of interest and the predicted control parameters into the evaluation sub-model to determine the predicted reward value after the set historical moment of the specified device.

[0115] Finally, the server can determine the reward value to be optimized based on the predicted reward value corresponding to the set historical moment and the predicted reward value after the set historical moment. The specific formula is as follows:

[0116]

[0117] In the above formula, s j can be used to represent the data of interest corresponding to the specified device at the set historical moment j. a j can be used to represent the control parameters corresponding to the specified device at the set historical moment j. w can be used to represent the model parameters of the evaluation sub-model. can be used to represent the reward value to be optimized of the driving trajectory obtained by the specified device driving according to the control parameters at the set historical moment j. Among them, can refer to the state-action value function, which is used to represent the expectation of the cumulative reward value of the entire driving trajectory of the specified device from the set historical moment j to the end of driving. It should be noted that although all states s j , …, s j+t and all actions a j , …, a j+t . However, since What is sought is the expectation, which can itself be used to characterize the expectation of the cumulative reward value over a future period of time, implicitly setting all states s after historical moment j j , …, s j+t and all actions a j , …, a j+t .

[0118] That is to say, the reward value to be optimized is essentially the expectation of the cumulative reward value of the specified device from the set historical moment j to the completion of the entire driving trajectory.

[0119] In the embodiments of this specification, the evaluation sub-model includes: a first evaluation sub-model and a second evaluation sub-model. The server can input the data of interest and the control parameters into the first evaluation sub-model to predict the corresponding reward value after the specified device drives according to the control parameters at the set historical moment, as the first reward value to be optimized. The specific formula is as follows:

[0120]

[0121] In the above formula, w 1,now can be used to characterize the current model parameters of the first evaluation sub-model. can be used to characterize, based on the first evaluation sub-model, the first reward value to be optimized after the specified device drives according to the control parameters at the set historical moment j. Among them, can refer to the state-action value function, which is used to characterize, based on the first evaluation sub-model, the expectation of the cumulative reward value of the specified device from the set historical moment j to the completion of the entire driving trajectory. That is to say, the first reward value to be optimized is essentially the expectation of the cumulative reward value of the specified device from the set historical moment j to the completion of the entire driving trajectory based on the first evaluation sub-model.

[0122] The server can input the data of interest and the control parameters into the second evaluation sub-model to predict the corresponding reward value after the specified device drives according to the control parameters at the set historical moment, as the second reward value to be optimized. The specific formula is as follows:

[0123]

[0124] In the above formula, w 2,now can be used to characterize the current model parameters of the second evaluation sub-model. can be used to characterize, based on the second evaluation sub-model, the corresponding second reward value to be optimized after the specified device drives according to the control parameters at the set historical moment j. Among them, It can refer to the state-action value function, which is used to characterize the expected cumulative reward value of the specified device from the set historical moment j to the completion of all driving trajectories based on the second evaluation sub-model. That is to say, the second reward value to be optimized is essentially the expected cumulative reward value of the specified device from the set historical moment j to the completion of all driving trajectories based on the second evaluation sub-model.

[0125] Furthermore, the server can train the first evaluation sub-model in the decision model with the goal of approximating the actual reward value with the first reward value to be optimized, and train the second evaluation sub-model in the decision model with the goal of approximating the actual reward value with the second reward value to be optimized.

[0126] In the embodiments of this specification, the decision model includes: a target evaluation sub-model and a target policy sub-model. The server can determine the reward value of the specified device in the environment at the set historical moment based on the interested data and control parameters as the reward value corresponding to the set historical moment.

[0127] Since the server needs to ensure that the specified device does not collide with surrounding obstacles during the driving process according to the control parameters output by the decision model, and also needs to further ensure the driving efficiency and smoothness when driving according to the control parameters output by the decision model. Based on this, the server can determine the true reward value of the specified device corresponding to the set historical moment from these three aspects.

[0128] Specifically, the server can determine the first influence factor corresponding to the set historical moment based on the interested data and control parameters. Then, the server can determine the reward value of the specified device in the environment at the set historical moment according to the first influence factor. The first influence factor mentioned here can be used to characterize the time difference between the moment when the specified device reaches the specified point and the moment when each obstacle reaches the specified point. The greater the time difference, the greater the reward value of the specified device in the environment at the set historical moment. The specified point mentioned here can be used to characterize the overlapping point of the driving trajectory corresponding to the specified device and the driving trajectories corresponding to each obstacle. The specified point mentioned here can also be used to characterize the possible position points that the specified device and each obstacle may reach in the future. Specifically, it can refer to the following formula:

[0129] r safe = ttc1

[0130] where r safeIt can be used to represent the reward value in terms of collision obtained when a specified device travels according to control parameters at a set historical moment. ttc1 can be used to represent the time difference between the moment when the specified device reaches a specified point and the moments when each obstacle reaches the specified point. There can be various specific forms. For example, during the process of the specified device traveling according to control parameters at a set historical moment, it is the average time difference between the device and the obstacles.

[0131] Of course, ttc1 can also refer to the reward value of a set numerical value corresponding to the time difference between the moment when the specified device reaches a specified point and the moments when each obstacle reaches the specified point. If the time difference is less than the set threshold, ttc1 is the reward value of a smaller set numerical value, or the reward value minus a set numerical value. If the time difference is not less than the set threshold, ttc1 is the reward value of a larger set numerical value.

[0132] It should be noted that the above ttc1 can be understood as the first influencing factor. And no matter what form ttc1 is, through the above formula, it can be characterized that if it is determined that the specified device collides with an obstacle, then r safe The reward value is subtracted by a preset maximum value, so that when making a decision using the decision model trained in this way, it can effectively avoid the specified device from colliding with an obstacle.

[0133] In the embodiments of this specification, the server can determine the second influencing factor corresponding to the set historical moment according to the interested data and control parameters. Then, the server can determine the reward value of the specified device in the environment at the set historical moment according to the second influencing factor. The second influencing factor mentioned here can be used to characterize the traffic efficiency when the specified device travels according to the control parameters. The greater the traffic efficiency, the greater the reward value of the specified device in the environment at the set historical moment. Specifically, it can refer to the following formula:

[0134]

[0135] In the above formula, r pass It can be used to represent the reward value in terms of traffic efficiency obtained when the specified device travels according to control parameters at a set historical moment. v can be used to characterize the traveling speed of the specified device at the set historical moment. (v + Δv) can be used to characterize the traveling speed when the specified device travels according to the control parameters. v max Is the maximum traveling speed of the specified device in the road scenario. It can be seen from the above formula that the greater the speed of the specified device traveling according to the control parameters at the set historical moment, the greater the reward value in terms of traffic efficiency.

[0136] In the embodiments of this specification, the server may determine a third influence factor corresponding to a set historical moment based on the data of interest and control parameters. Then, the server may determine the reward value of the specified device in the environment at the set historical moment according to the third influence factor. The third influence factor mentioned here may be used to characterize the degree of state change of the specified device after driving according to the control parameters. The greater the degree of state change, the greater the reward value of the specified device in the environment at the set historical moment. Specifically, the following formula may be referred to:

[0137]

[0138] where r soft can be used to represent the reward value of the specified device in terms of the degree of state change obtained by driving according to the control parameters at the set historical moment, and |Δv| can be used to represent the rate of change of the speed of the specified device when driving according to the control parameters at the set historical moment. Among them, it can also be used to represent the rate of change of the acceleration of the specified device when driving according to the control parameters at the set historical moment, which can be specifically determined according to service requirements. can be used to represent the rate of change of the steering angle of the specified device when driving according to the control parameters at the set historical moment. It can be seen from this formula that the greater the rate of change of the speed of the specified device, the worse the smoothness of the specified device when driving according to the control parameters at the set historical moment. Therefore, r soft is smaller. Correspondingly, if the rate of change of the steering angle of the specified device is greater, it indicates that the smoothness of the specified device when driving according to the control parameters at the set historical moment is worse, and r soft is smaller.

[0139] Of course, the above |Δv| can also be used to represent the rate of change of the acceleration of the specified device when driving according to the control parameters at the set historical moment, and the specified device can be selected according to service requirements.

[0140] Furthermore, the server may determine the actual reward value of the specified device corresponding to the set historical moment according to one or more of the above ways of determining the reward value: r safe 、r pass 、r soft .

[0141] It should be noted that there can be various specific forms of the reward function used by the server to train the decision model, as long as it can represent a positive correlation between the reward value and the above time difference, a positive correlation between the reward value and the above traffic efficiency, and a negative correlation between the reward value and the above degree of state change. This specification does not limit the specific form of the reward function.

[0142] Secondly, the server can use the attention mechanism network to predict the data of interest after the specified device travels according to the control parameters at the set historical moment, as the predicted data of interest, and input the predicted data of interest into the target policy sub-model to determine the control parameters after the set historical moment for the specified device, as the predicted control parameters. Specifically, the following formula can be referred to:

[0143]

[0144] In the above formula, can be used to represent the predicted control parameters of the specified device after the set historical moment based on the target policy sub-model. can be used to represent the current model parameters of the target policy sub-model. ξ can be used to represent noise, and ξ can be randomly selected independently from the truncated normal distribution for random selection.

[0145] Then, the server can input the predicted data of interest and the predicted control parameters into the target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment.

[0146] Finally, the server can determine the actual reward value according to the reward value corresponding to the set historical moment and the predicted reward value.

[0147] In practical applications, when the decision-making model updates the model parameters each time, it updates the model parameters of the decision-making model through the control parameter with the largest predicted reward value determined currently, so there will be an overestimation of the predicted reward value. Based on this, the server can use two sub-models to determine different predicted reward values, and update the model parameters by selecting the smaller predicted reward value to suppress the continuous overestimation.

[0148] In the embodiments of this specification, the target evaluation sub-model includes: the first target evaluation sub-model and the second target evaluation sub-model. The server can input the predicted data of interest and the predicted control parameters into the first target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment, as the first candidate reward value. Specifically, the following formula can be referred to:

[0149]

[0150] In the above formula, can be used to represent the current model parameters of the first target evaluation sub-model. can be used to represent the first candidate reward value corresponding to the specified device traveling according to the control parameters at the set historical moment j + 1 based on the first target evaluation sub-model. Among them, It can refer to the state-action value function, which is used to characterize the expected cumulative reward value of the specified device from the set historical moment j to the completion of all driving trajectories based on the first target evaluation sub-model. That is to say, the first candidate reward value is essentially the expected cumulative reward value of the specified device from the set historical moment j + 1 to the completion of all driving trajectories based on the first target evaluation sub-model.

[0151] The server can input the predicted data of interest and the predicted control parameters into the second target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the second candidate reward value. Specifically, the following formula can be referred to:

[0152]

[0153] In the above formula, It can be used to characterize the current model parameters of the second target evaluation sub-model. It can be used to characterize the second candidate reward value corresponding to the specified device driving according to the control parameters at the set historical moment j + 1 based on the second target evaluation sub-model. Among them, It can refer to the state-action value function, which is used to characterize the expected cumulative reward value of the specified device from the set historical moment j to the completion of all driving trajectories based on the second target evaluation sub-model. That is to say, the second candidate reward value is essentially the expected cumulative reward value of the specified device from the set historical moment j + 1 to the completion of all driving trajectories based on the second target evaluation sub-model.

[0154] Finally, the server can take the smaller value of the first candidate reward value and the second candidate reward value as the predicted reward value. Specifically, the following formula can be referred to:

[0155]

[0156] In the above formula, It can be used to characterize the actual reward value corresponding to the specified device driving according to the control parameters at the set historical moment. It can be used to represent selecting the smaller reward value among the first candidate reward value and the second candidate reward value. γ is the discount factor, which is used to reduce the influence of the predicted reward values predicted at other moments after the set historical moment j. r j It can be used to represent the reward value of the specified device in the environment at the set historical moment.

[0157] Further, the evaluation sub-model includes: a first evaluation sub-model and a second evaluation sub-model. The server can input the data of interest and the control parameters into the first evaluation sub-model to predict the reward value corresponding to the specified device after driving according to the control parameters at the set historical moment, as the first reward value to be optimized. Specifically, the following formula can be referred to:

[0158]

[0159] In the above formula, can be used to represent the actual reward value of the specified device after driving according to the control parameters at the set historical moment. can be used to represent the first reward value to be optimized corresponding to the specified device after driving according to the control parameters at the set historical moment j based on the first evaluation sub-model. δ 1,j can be used to represent the difference between the actual reward value and the first reward value to be optimized.

[0160] Similarly, the server can input the data of interest and the control parameters into the second evaluation sub-model to predict the reward value corresponding to the specified device after driving according to the control parameters at the set historical moment, as the second reward value to be optimized. Specifically, the following formula can be referred to:

[0161]

[0162] In the above formula, can be used to represent the actual reward value of the specified device after driving according to the control parameters at the set historical moment. can be used to represent the second reward value to be optimized corresponding to the specified device after driving according to the control parameters at the set historical moment j based on the second evaluation sub-model. δ 2,j can be used to represent the difference between the actual reward value and the second reward value to be optimized.

[0163] Then, the server can take the approximation of the first reward value to be optimized to the actual reward value as the optimization goal, and train the first evaluation sub-model in the decision model.

[0164] In the embodiments of this specification, the server can complete the training of the decision model by updating the model parameters of each sub-model in the model structure of the decision model. Specifically, as Figure 2 shown.

[0165] Figure 2 is a schematic diagram of the model structure of the decision model provided by the embodiments of this specification.

[0166] In Figure 2Among them, the server can input the data of interest into the policy model to determine the predicted control parameters output by the policy model. Secondly, the server can input the predicted control parameters output by the policy model into the first evaluation sub-model to determine the first reward value to be optimized, and input the predicted control parameters output by the policy model into the second evaluation sub-model to determine the second reward value to be optimized.

[0167] Similarly, the server can input the data of interest into the target policy model to determine the predicted control parameters output by the target policy model. The server can input the predicted control parameters output by the target policy model into the first target evaluation sub-model to determine the first candidate reward value, and input the predicted control parameters output by the target policy model into the second target evaluation sub-model to determine the second candidate reward value. Select the smaller value from the first candidate reward value and the second candidate reward value to determine the actual reward value.

[0168] Finally, the server can use the difference between the first reward value to be optimized and the actual reward value as the first difference, and use the difference between the second reward value to be optimized and the actual reward value as the second difference. The server can update the first evaluation sub-model based on the first difference and update the second evaluation sub-model based on the second difference.

[0169] Specifically, for each round of training, the server can take the approximation of the first reward value to be optimized to the actual reward value as the optimization goal, and update the model parameters of the first evaluation sub-model in this round of training based on the first parameter update step. Specifically, it can refer to the following formula:

[0170]

[0171] In the above formula, can be used to represent the model parameters of the first target evaluation sub-model in the current round. w 1,new can be used to represent the model parameters of the first target evaluation sub-model after being updated in the next round. α can be used to represent the first parameter update step. can be used to represent updating the model parameters of the first target evaluation sub-model by the method of gradient descent.

[0172] Similarly, the server can take the approximation of the second reward value to be optimized to the actual reward value as the optimization goal to train the second evaluation sub-model in the decision model.

[0173] Specifically, the server can take the approximation of the second reward value to be optimized to the actual reward value as the optimization goal, and update the model parameters of the second evaluation sub-model in this round of training based on the first parameter update step. Specifically, it can refer to the following formula:

[0174]

[0175] In the above formula, w 2,now can be used to represent the model parameters of the second target evaluation sub-model in the current round. w 2,new can be used to represent the model parameters of the second target evaluation sub-model after being updated in the next round. α can be used to represent the first parameter update step size, which is set through expert experience. can be used to represent updating the model parameters of the second target evaluation sub-model by means of gradient descent.

[0176] In the embodiments of this specification, the server can update the model parameters of the policy sub-model in the current round of training according to the model parameters of the first evaluation sub-model in the current round of training, the model parameters of the second evaluation sub-model in the current round of training, and the second parameter update step size. Specifically, the following formula can be referred to:

[0177]

[0178] In the above formula, θ now can be used to represent the model parameters of the policy sub-model in the current round. θ new can be used to represent the model parameters of the policy sub-model after being updated in the next round. β can be used to represent the second parameter update step size, which is set through expert experience. can be used to represent updating the model parameters of the policy sub-model by means of gradient descent based on the chain rule. can be used to represent the control parameters corresponding to the specified device at the set historical moment determined by the policy sub-model according to the data of interest corresponding to the specified device at the set historical moment. The specific formula is

[0179] Among them, the server can randomly select a sub-model from the model parameters of the first evaluation sub-model in the current round of training and the model parameters of the second evaluation sub-model in the current round of training to update the model parameters of the policy sub-model.

[0180] Finally, the server can update the model parameters of the target policy sub-model in the current round of training according to the model parameters of the policy sub-model in the current round of training and the soft update coefficient until the preset conditions are met to complete the training of the target policy sub-model in the decision-making model. Specifically, the following formula can be referred to:

[0181]

[0182] In the above formula, can be used to represent the model parameters of the target policy sub-model in the current round. θ new can be used to represent the model parameters of the policy sub-model after being updated in the next round. It can be used to characterize the model parameters of the target policy sub-model after update. τ is the parameter soft update coefficient, which is set based on expert experience.

[0183]

[0184] In the above formula, It can be used to characterize the model parameters of the first target evaluation sub-model in the current round. w 1,new It can be used to characterize the model parameters of the first evaluation sub-model after update in the next round. It can be used to characterize the model parameters of the first target evaluation sub-model after update. τ is the parameter soft update coefficient, which is set based on expert experience.

[0185]

[0186] In the above formula, It can be used to characterize the model parameters of the second target evaluation sub-model in the current round. w 2,new It can be used to characterize the model parameters of the second evaluation sub-model after update in the next round. It can be used to characterize the model parameters of the second target evaluation sub-model after update. τ is the parameter soft update coefficient, which is set based on expert experience.

[0187] It can be seen from the above formula that some model parameters in the current model are updated each time. Even if the target network is updated in each iteration, a certain stability can be maintained. Among them, the smaller the parameter soft update coefficient, the smaller the change in the target network parameters, and the slower the algorithm convergence speed.

[0188] It should be noted that the model structures of the policy sub-model and the target policy sub-model are the same, and the model parameters can be determined according to business requirements, and can be the same or different. The model structures of the first evaluation sub-model and the second evaluation sub-model are the same, and the model parameters can be determined according to business requirements, and can be the same or different. The model structures of the first target evaluation sub-model and the second target evaluation sub-model are the same, and the model parameters can be determined according to business requirements, and can be the same or different.

[0189] That is to say, the server can also update the model parameters of the policy sub-model in this round of training until the preset conditions are met to complete the training of the policy sub-model in the decision-making model. Specifically, it can be selected based on business requirements whether to give priority to the policy sub-model or the target policy sub-model. For example, if in the last round of training, the target policy sub-model cannot be updated through the policy sub-model and the parameter soft update coefficient, the policy sub-model can be applied in actual applications. If in the last round of training, the target policy sub-model is updated through the policy sub-model and the parameter soft update coefficient, the target policy sub-model can be applied in actual applications.

[0190] In the embodiments of this specification, the preset condition may refer to that after multiple rounds of iterative training of the decision model, the difference between the first reward value to be optimized and the actual reward value and the difference between the second reward value to be optimized and the actual reward value can be continuously reduced and converge within a numerical range, thereby completing the training process of the decision model.

[0191] Of course, in addition to training the decision model with the minimum difference as the optimization target, the preset condition can also be to use a preset difference as the optimization target and train the decision model by adjusting the model parameters included in the decision model. That is to say, during the process of multiple rounds of iterative training, it is necessary to make the difference continuously approach the preset difference. When after multiple rounds of iterative training, the target reward value fluctuates around the preset difference, it can be determined that the training of the decision model is completed.

[0192] It can be seen from the above process that this method can input historical state data into a pre-trained long short-term memory network to determine the predicted state data corresponding to each obstacle. And through the weights corresponding to the attention mechanism network in the decision model, the data of interest can be determined from the environmental data of the environment where the specified device is located at the set historical moment. Then, according to the data of interest and the control parameters, the reward value corresponding to the specified device when driving according to the control parameters at the set historical moment is determined to measure the rationality of the determined control parameters. By training the decision model in this way, the probability of the specified device colliding with surrounding obstacles can be reduced, avoiding collisions with surrounding obstacles, and thus effectively ensuring the safe driving of the specified device.

[0193] Since in the training process of the model, not only the time difference between the moment when the specified device reaches the specified point and the moment when each obstacle reaches the specified point is considered, but also the passing efficiency of the specified device and the degree of state change of the specified device are considered, the control parameters obtained by the trained decision model according to the state data can not only improve the safety during driving, but also improve the passing efficiency and smoothness.

[0194] It should be noted that the execution subject of the above model training method can also be a device such as a server or a computer. That is, the server can obtain the historical state data corresponding to the specified device and each obstacle at each historical moment, input the historical state data into the decision model to be trained, determine the control parameters corresponding to the specified device at the set historical moment, and based on the determined control parameters, train the decision model, and deploy the trained decision model to the specified device.

[0195] After the training of the decision model in the embodiments of this specification is completed, the trained decision model can be deployed to an unmanned device to achieve the control of the unmanned device, such asFigure 3 as shown

[0196] Figure 3 The figure is a schematic flowchart of the control method for the unmanned device provided in the embodiments of this specification, which specifically includes:

[0197] S300: Obtain the state data of the unmanned device and each obstacle at the current moment as the current state data.

[0198] S302: Input the current state data into a pre-trained long short-term memory network to predict the state data of each obstacle after the current moment as the predicted state data.

[0199] S304: Input the current state data and the predicted state data into the trained decision model to determine the control parameters corresponding to the unmanned device at the current moment, and the decision model is obtained by the above model training method.

[0200] S306: Control the unmanned device according to the control parameters corresponding to the unmanned device at the current moment.

[0201] In the embodiments of this specification, the unmanned device can obtain the state data of itself and each obstacle at the current moment as the current state data through various sensors (such as cameras, lidar, etc.) set on itself. Secondly, the unmanned device can input the current state data into a pre-trained long short-term memory network and predict the state data of each obstacle after the current moment as the predicted state data corresponding to the current moment based on a preset time window. Then, the unmanned device can input the predicted state data corresponding to the current moment into the decision model to determine the control parameters corresponding to itself at the current moment. Finally, the unmanned device can control itself according to the control parameters corresponding to itself at the current moment.

[0202] Specifically, the unmanned device can input the predicted state data corresponding to the current moment into the policy sub-model in the decision model to determine the control parameters corresponding to itself at the current moment. Of course, if the target policy sub-model is updated based on the policy sub-model and the parameter soft update coefficient, the unmanned device can also input the predicted state data corresponding to the current moment into the target policy sub-model in the decision model to determine the control parameters corresponding to itself at the current moment. Specifically, the policy sub-model or the target policy sub-model can be selected according to the service requirements to determine the control parameters corresponding to the current moment.

[0203] The execution subject of the control method for the unmanned device provided in this specification can be an unmanned device, or a terminal device such as a server or a desktop computer. If a terminal device such as a server or a desktop computer is the execution subject, the terminal device can obtain the status data collected and uploaded by the unmanned device, and after determining the control parameters corresponding to the current moment, can return the control parameters corresponding to the current moment to the unmanned device.

[0204] The above is the method for model training provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding model training device, as Figure 4 shown.

[0205] Figure 4 is the structural schematic diagram of the model training device provided by the embodiment of this specification, specifically including:

[0206] An acquisition module 400, configured to acquire the historical status data corresponding to a specified device and each obstacle at each historical moment;

[0207] A prediction module 402, configured to input the historical status data into a pre-trained long short-term memory network according to a time series, to predict the status data of each obstacle after a set historical moment, as predicted status data;

[0208] An input module 404, configured to input the historical status data and the predicted status data into a decision model to be trained, to determine, through the weights corresponding to the attention mechanism network, the data of interest from the environmental data of the environment where the specified device is located at the set historical moment;

[0209] A determination module 406, configured to determine the control parameters corresponding to the specified device at the set historical moment according to the data of interest;

[0210] A training module 408, configured to determine, according to the data of interest and the control parameters, the reward value corresponding to the specified device after traveling according to the control parameters at the set historical moment, and train the decision model according to the reward value.

[0211] Optionally, the decision model includes: an evaluation sub-model;

[0212] The training module 408 is specifically configured to input the data of interest and the control parameters into the evaluation sub-model, predict the reward value corresponding to the specified device after driving according to the control parameters at the set historical moment as the reward value to be optimized, and determine the actual reward value corresponding to the specified device after driving according to the control parameters at the set historical moment based on the data of interest and the control parameters. Taking the approximation of the actual reward value by the reward value to be optimized as the optimization objective, the decision model is trained.

[0213] Optionally, the decision model includes: a target evaluation sub-model and a target policy sub-model;

[0214] The training module 408 is specifically configured to determine the reward value of the specified device in the environment at the set historical moment according to the data of interest and the control parameters as the reward value corresponding to the set historical moment, predict the data of interest after the specified device drives according to the control parameters at the set historical moment through the attention mechanism network as the predicted data of interest, input the predicted data of interest into the target policy sub-model to determine the control parameters of the specified device after the set historical moment as the predicted control parameters, input the predicted data of interest and the predicted control parameters into the target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment, and determine the actual reward value according to the reward value corresponding to the set historical moment and the predicted reward value.

[0215] Optionally, the target evaluation sub-model includes: a first target evaluation sub-model and a second target evaluation sub-model;

[0216] The training module 408 is specifically configured to input the predicted data of interest and the predicted control parameters into the first target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the first candidate reward value, and input the predicted data of interest and the predicted control parameters into the second target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the second candidate reward value, and take the smaller value of the first candidate reward value and the second candidate reward value as the predicted reward value.

[0217] Optionally, the training module 408 is specifically configured to determine a first influence factor corresponding to the set historical moment according to the data of interest and the control parameter, and determine a reward value of the specified device in the environment at the set historical moment according to the first influence factor. The first influence factor is used to characterize the time difference between the moment when the specified device reaches the specified point and the moment when each obstacle reaches the specified point. The greater the time difference, the greater the reward value of the specified device in the environment at the set historical moment.

[0218] Optionally, the training module 408 is specifically configured to determine a second influence factor corresponding to the set historical moment according to the data of interest and the control parameter, and determine a reward value of the specified device in the environment at the set historical moment according to the second influence factor. The second influence factor is used to characterize the traffic efficiency when the specified device travels according to the control parameter. The greater the traffic efficiency, the greater the reward value of the specified device in the environment at the set historical moment.

[0219] Optionally, the training module 408 is specifically configured to determine a third influence factor corresponding to the set historical moment according to the data of interest and the control parameter, and determine a reward value of the specified device in the environment at the set historical moment according to the third influence factor. The third influence factor is used to characterize the degree of state change after the specified device travels according to the control parameter. The greater the degree of state change, the smaller the reward value of the specified device in the environment at the set historical moment.

[0220] Optionally, the evaluation sub-model includes: a first evaluation sub-model and a second evaluation sub-model;

[0221] The training module 408 is specifically configured to input the data of interest and the control parameter into the first evaluation sub-model to predict a corresponding reward value after the specified device travels according to the control parameter at the set historical moment as a first reward value to be optimized, and input the data of interest and the control parameter into the second evaluation sub-model to predict a corresponding reward value after the specified device travels according to the control parameter at the set historical moment as a second reward value to be optimized. The first evaluation sub-model in the decision model is trained with the optimization goal of approximating the first reward value to be optimized to the actual reward value, and the second evaluation sub-model in the decision model is trained with the optimization goal of approximating the second reward value to be optimized to the actual reward value.

[0222] Optionally, the decision model includes: a policy sub-model and a target policy sub-model;

[0223] Specifically, for each round of training, the training module 408 is configured to use the first reward value to be optimized to approximate the actual reward value as the optimization objective, and update the model parameters of the first evaluation sub-model in this round of training based on the first parameter update step size. In addition, the training module 408 is configured to use the second reward value to be optimized to approximate the actual reward value as the optimization objective, and update the model parameters of the second evaluation sub-model in this round of training based on the first parameter update step size. Then, according to the model parameters of the first evaluation sub-model in this round of training, the model parameters of the second evaluation sub-model in this round of training, and the second parameter update step size, the training module 408 updates the model parameters of the policy sub-model in this round of training. Further, according to the model parameters of the policy sub-model in this round of training and the soft update coefficient, the training module 408 updates the model parameters of the target policy sub-model in this round of training until a preset condition is met, so as to complete the training of the target policy sub-model in the decision model.

[0224] Figure 5 FIG. is a schematic structural diagram of a control device for an unmanned device provided by an embodiment of the present specification, specifically including:

[0225] An acquisition module 500, configured to acquire state data corresponding to the unmanned device and each obstacle at the current moment as current state data;

[0226] A prediction module 502, configured to input the current state data into a pre-trained long short-term memory network to predict state data of each obstacle after the current moment as predicted state data;

[0227] A determination module 504, configured to input the current state data and the predicted state data into a trained decision model to determine control parameters corresponding to the unmanned device at the current moment, where the decision model is trained by the above model training method;

[0228] A control module 506, configured to control the unmanned device according to the control parameters corresponding to the unmanned device at the current moment.

[0229] The present specification further provides a computer-readable storage medium storing a computer program, and the computer program can be used to execute the above Figure 1 provided model training method or the above Figure 3 provided control method for an unmanned device.

[0230] The present specification further provides Figure 6 a schematic structural diagram of the unmanned device shown in FIG. As Figure 6As described above, at the hardware level, the unmanned device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 method of model training or the above Figure 3 control method of the unmanned device provided. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0231] In the 1990s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system on a single PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0232] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0233] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0234] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0235] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0236] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0237] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0238] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0239] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0240] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.

[0241] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0242] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0243] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system or computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0244] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0245] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, they are described relatively simply, and for the relevant parts, reference can be made to the partial description of method embodiments.

[0246] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A method for model training, characterized in that, Including: Obtain the historical state data corresponding to the specified device and each obstacle at each historical moment; Input the historical state data into a pre-trained long short-term memory network according to the time series to predict the state data of each obstacle after a set historical moment as the predicted state data; Input the historical state data and the predicted state data into a decision model to be trained, and determine the interesting data from the environmental data of the environment where the specified device is located at the set historical moment through the weights corresponding to the attention mechanism network; Determine the control parameter corresponding to the specified device at the set historical moment according to the interesting data; Determine the reward value corresponding to the specified device after driving according to the control parameter at the set historical moment according to the interesting data and the control parameter, and train the decision model according to the reward value; The decision model includes: an evaluation sub-model; Determining the reward value corresponding to the specified device after driving according to the control parameter at the set historical moment according to the interesting data and the control parameter includes: Input the interesting data and the control parameter into the evaluation sub-model to predict the reward value corresponding to the specified device after driving according to the control parameter at the set historical moment as the reward value to be optimized, and determine the actual reward value corresponding to the specified device after driving according to the control parameter at the set historical moment according to the interesting data and the control parameter; Training the decision model according to the reward value includes: Training the decision model with the goal of approximating the actual reward value with the reward value to be optimized; the decision model is used to control the unmanned device.

2. The method according to claim 1, characterized in that, The decision model includes: a target evaluation sub-model, a target policy sub-model; Determining the actual reward value corresponding to the specified device after driving according to the control parameter at the set historical moment according to the interesting data and the control parameter includes: Determine the reward value of the environment where the specified device is located at the set historical moment according to the interesting data and the control parameter as the reward value corresponding to the set historical moment; Predict the interesting data after the specified device drives according to the control parameter at the set historical moment through the attention mechanism network as the predicted interesting data, and input the predicted interesting data into the target policy sub-model to determine the control parameter of the specified device after the set historical moment as the predicted control parameter; Input the predicted interesting data and the predicted control parameter into the target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment; Determine the actual reward value according to the reward value corresponding to the set historical moment and the predicted reward value.

3. The method according to claim 2, characterized in that The target evaluation sub-model includes: a first target evaluation sub-model, a second target evaluation sub-model; Input the predicted data of interest and the predicted control parameters into the target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment, including: Input the predicted data of interest and the predicted control parameters into the first target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the first candidate reward value, and input the predicted data of interest and the predicted control parameters into the second target evaluation sub-model to determine the predicted reward value of the specified device after the set historical moment as the second candidate reward value; Take the smaller value of the first candidate reward value and the second candidate reward value as the predicted reward value.

4. The method according to claim 2, wherein Determine the reward value of the specified device in the environment at the set historical moment according to the data of interest and the control parameters, including: Determine the first influence factor corresponding to the set historical moment according to the data of interest and the control parameters; Determine the reward value of the specified device in the environment at the set historical moment according to the first influence factor. The first influence factor is used to characterize the time difference between the moment when the specified device reaches the specified point and the moment when each obstacle reaches the specified point. The greater the time difference, the greater the reward value of the specified device in the environment at the set historical moment.

5. The method according to claim 2, characterized in that, Determine the reward value of the specified device in the environment at the set historical moment according to the data of interest and the control parameters, specifically including: Determine the second influence factor corresponding to the set historical moment according to the data of interest and the control parameters; Determine the reward value of the specified device in the environment at the set historical moment according to the second influence factor. The second influence factor is used to characterize the traffic efficiency when the specified device travels according to the control parameters. The greater the traffic efficiency, the greater the reward value of the specified device in the environment at the set historical moment.

6. The method according to claim 2, characterized in that Determine the reward value of the specified device in the environment at the set historical moment according to the data of interest and the control parameters, specifically including: Determine the third influence factor corresponding to the set historical moment according to the data of interest and the control parameters; Determine the reward value of the specified device in the environment at the set historical moment according to the third influence factor. The third influence factor is used to characterize the degree of state change after the specified device travels according to the control parameters. The greater the degree of state change, the smaller the reward value of the specified device in the environment at the set historical moment.

7. The method according to claim 1, characterized in that, The evaluation sub-model includes: a first evaluation sub-model, a second evaluation sub-model; Input the data of interest and the control parameters into the evaluation sub-model to predict the reward value corresponding to the specified device driving according to the control parameters at the set historical moment, and use it as the reward value to be optimized. Specifically, it includes: input the data of interest and the control parameters into the first evaluation sub-model to predict the reward value corresponding to the specified device driving according to the control parameters at the set historical moment, and use it as the first reward value to be optimized; input the data of interest and the control parameters into the second evaluation sub-model to predict the reward value corresponding to the specified device driving according to the control parameters at the set historical moment, and use it as the second reward value to be optimized; take the approximation of the reward value to be optimized to the actual reward value as the optimization goal, and train the decision model. Specifically, it includes: Take the approximation of the first reward value to be optimized to the actual reward value as the optimization goal, train the first evaluation sub-model in the decision model, and take the approximation of the second reward value to be optimized to the actual reward value as the optimization goal, train the second evaluation sub-model in the decision model.

8. The method according to claim 7, wherein The decision model includes: a policy sub-model and a target policy sub-model; Take the approximation of the reward value to be optimized to the actual reward value as the optimization goal, and train the decision model. Specifically, it includes: For each round of training, take the approximation of the first reward value to be optimized to the actual reward value as the optimization goal, and update the model parameters of the first evaluation sub-model in this round of training based on the first parameter update step size. Also, take the approximation of the second reward value to be optimized to the actual reward value as the optimization goal, and update the model parameters of the second evaluation sub-model in this round of training based on the first parameter update step size; According to the model parameters of the first evaluation sub-model in this round of training, the model parameters of the second evaluation sub-model in this round of training, and the second parameter update step size, update the model parameters of the policy sub-model in this round of training; according to the model parameters of the policy sub-model in this round of training and the soft update coefficient, update the model parameters of the target policy sub-model in this round of training until the preset condition is met to complete the training of the target policy sub-model in the decision model.

9. A control method for an unmanned device, characterized in that, It includes: Obtain the state data corresponding to the unmanned device and each obstacle at the current moment as the current state data; Input the current state data into the pre-trained long short-term memory network to predict the state data of each obstacle after the current moment as the predicted state data; Input the current state data and the predicted state data into the trained decision model to determine the control parameters corresponding to the unmanned device at the current moment. The decision model is trained by the method described in any one of the above claims 1 to 8; Control the unmanned device according to the control parameters corresponding to the unmanned device at the current moment.

10. An apparatus for model training, characterized in that, It includes: An acquisition module for acquiring the historical state data corresponding to the specified device and each obstacle at each historical moment; A prediction module, configured to input the historical state data into a pre-trained long short-term memory network according to a time series, so as to predict the state data of each obstacle after a set historical moment, and use it as predicted state data; An input module, configured to input the historical state data and the predicted state data into a decision model to be trained, so as to determine interesting data from the environmental data of the environment where the specified device is located at the set historical moment through the weights corresponding to the attention mechanism network; A determination module, configured to determine a control parameter corresponding to the specified device at the set historical moment according to the interesting data; A training module, configured to determine a reward value corresponding to the specified device after traveling according to the control parameter at the set historical moment according to the interesting data and the control parameter, and train the decision model according to the reward value; The decision model includes: an evaluation sub-model; Determining a reward value corresponding to the specified device after traveling according to the control parameter at the set historical moment according to the interesting data and the control parameter includes: Inputting the interesting data and the control parameter into the evaluation sub-model to predict a reward value corresponding to the specified device after traveling according to the control parameter at the set historical moment, and using it as a reward value to be optimized, and determining an actual reward value corresponding to the specified device after traveling according to the control parameter at the set historical moment according to the interesting data and the control parameter; Training the decision model according to the reward value includes: Taking the approximation of the reward value to be optimized to the actual reward value as an optimization target to train the decision model; the decision model is used to control an unmanned device.

11. A control device for an unmanned device, characterized in that, Including: An acquisition module, configured to acquire the state data corresponding to an unmanned device and each obstacle at the current moment, and use it as current state data; A prediction module, configured to input the current state data into a pre-trained long short-term memory network, so as to predict the state data of each obstacle after the current moment, and use it as predicted state data; A determination module, configured to input the current state data and the predicted state data into a trained decision model to determine a control parameter corresponding to the unmanned device at the current moment, and the decision model is trained by the method according to any one of claims 1 to 8 above; A control module, configured to control the unmanned device according to the control parameter corresponding to the unmanned device at the current moment.

12. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 or 9 above is implemented.

13. An unmanned device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 8 or 9 above is implemented.

Citation Information

Patent Citations

  • Traffic flow model training method based on attention mechanism

    CN110889546A

  • Control model training method and device and control method and device

    CN112306059A