Action prediction model training method and device, equipment and storage medium

Through the training method of the action prediction model, the vehicle is used to iteratively train the action prediction model in the ramp merging environment, and solve the problem of vehicle merging in the ramp merging area, achieving the effect of vehicle merging more safely and efficiently completing the merging.

CN120030346APending Publication Date: 2025-05-23CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510002747.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The vehicle merging problem in the ramp merging area has led to an increase in the incidence of accidents, and the existing technology is difficult to effectively solve this problem.

Method used

Through the training method of the action prediction model, the action prediction model is iteratively trained by the vehicle in the ramp merging environment to generate a target action prediction model that can predict the flow action and improve the flow effect.

Benefits of technology

The trained action prediction model can predict the combined action that enables the vehicle to successfully merge and improves the combined effect, improving the safety and efficiency of the vehicle in the ramp merging area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030346A_ABST
    Figure CN120030346A_ABST
Patent Text Reader

Abstract

The invention discloses an action prediction model training method and device, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: inputting state information of a vehicle into an action prediction model, predicting a confluence action and an award under the state information by the action prediction model, and performing model training by combining the loss of the award and the loss of the action; therefore, the confluence action predicted by the action prediction model trained by the method not only converges to the confluence action in the sample data, so that the predicted confluence action can realize successful confluence of vehicles, but also converges to a direction with a large reward, and the reward indicates a confluence effect, so that the confluence effect of the vehicles is improved. Therefore, the confluence action predicted by the action prediction model trained by the method can improve the confluence effect, so that the vehicles can complete confluence more safely and efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a training method, device, equipment and storage medium for an action prediction model. Background Art

[0002] Ramp merging is one of the typical complex traffic scenarios. The sudden merging of ramp vehicles causes the main road vehicles to stop urgently or even collide, which makes the accident rate in the ramp merging area increasingly high. Therefore, it is very important to solve the merging problem of vehicles in the ramp merging area. Summary of the invention

[0003] The embodiment of the present application provides a training method, device, equipment and storage medium for an action prediction model, wherein the merging action predicted by the target action prediction model trained by the method not only converges to the merging action in the sample data, so that the predicted merging action enables the vehicle to merge successfully, but also makes the predicted merging action converge to the direction of a large reward. Since the reward also indicates the merging effect, the merging action predicted by the model trained by the method can also improve the merging effect, so that the vehicle completes the merging more safely and efficiently. The technical solution is as follows.

[0004] On the one hand, a method for training an action prediction model is provided, the method comprising:

[0005] Acquire a plurality of first sample data at a plurality of first moments, wherein the first sample data at each first moment includes first state information, first action information, and a first reward, wherein the first state information includes ramp merging environment information of the vehicle at the first moment and the driving state of the vehicle, the first action information includes a merging action performed by the vehicle under the first state information, and the first reward includes a reward obtained by the vehicle for performing the merging action under the first state information, and the reward is used to indicate a merging effect;

[0006] Inputting the first state information in the first sample data into an action prediction model, and having the action prediction model output first predicted action information and a first predicted reward;

[0007] Determine a first loss value based on the first predicted action information and the first action information included in the first sample data;

[0008] Determine a second loss value based on the first predicted reward and the first reward included in the first sample data;

[0009] Based on the first loss value and the second loss value of each of the multiple first sample data, the action prediction model is iteratively trained to obtain a target action prediction model.

[0010] In some embodiments, the acquiring a plurality of first sample data at a plurality of first moments includes:

[0011] Input multiple first state information into the simulator corresponding to the ramp merging scenario respectively, and the simulator outputs the first merging actions corresponding to the multiple first state information respectively, obtain the first reward obtained by the vehicle when performing the first merging action under each first state information, and obtain the first sample data including the first state information, the first merging action and the first reward.

[0012] In some embodiments, the method further comprises:

[0013] Acquire a plurality of second sample data at second moments, where the second sample data is non-simulated data in a ramp merging scenario;

[0014] The iterative training of the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a target action prediction model includes:

[0015] Iteratively training the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a first action prediction model;

[0016] Inputting the second state information in the second sample data into the first action prediction model, and having the first action prediction model output second predicted action information and a second predicted reward;

[0017] Determine a third loss value based on the second predicted action information and the second action information included in the second sample data;

[0018] Determine a fourth loss value based on the second predicted reward and the second reward included in the second sample data;

[0019] Based on the third loss value and the fourth loss value of each of the plurality of second sample data, the first action prediction model is iteratively trained to obtain the target action prediction model.

[0020] In some embodiments, the first reward includes at least one of a comfort reward, an efficiency reward, and a safety reward; the comfort reward is determined based on the acceleration of the vehicle when performing the merging maneuver, the efficiency reward is determined based on the merging time of the vehicle when performing the merging maneuver, and the safety reward is determined based on at least one of the collision information after the vehicle performs the merging maneuver, the distance between the vehicle and an adjacent vehicle, and the number of lane changes.

[0021] In some embodiments, the ramp merging environment information includes at least one of lane information of a ramp merging area where the vehicle is located and a driving state of an adjacent vehicle of the vehicle;

[0022] The driving state of the vehicle includes at least one of the position, speed, acceleration and heading angle of the vehicle;

[0023] The merging action includes at least one of acceleration, deceleration and lane change.

[0024] In some embodiments, the method further comprises:

[0025] Acquire target state information of a target vehicle in a ramp merging scenario, wherein the target state information includes ramp merging environment information of the target vehicle and a driving state of the target vehicle;

[0026] The target state information is input into the target action prediction model, and the target action prediction model outputs a target merging action. The target vehicle is used to execute the target merging action to merge in the ramp merging scenario.

[0027] On the other hand, a training device for an action prediction model is provided, the device comprising:

[0028] an acquisition module, configured to acquire a plurality of first sample data at a plurality of first moments, wherein the first sample data at each first moment includes first state information, first action information, and a first reward, wherein the first state information includes ramp merging environment information of the vehicle at the first moment and the driving state of the vehicle, the first action information includes a merging action performed by the vehicle under the first state information, and the first reward includes a reward obtained by the vehicle for performing the merging action under the first state information, and the reward is used to indicate a merging effect;

[0029] An input-output module, configured to input the first state information in the first sample data into an action prediction model, and the action prediction model outputs first predicted action information and a first predicted reward;

[0030] A determination module, configured to determine a first loss value based on the first predicted action information and the first action information included in the first sample data;

[0031] The determination module is further configured to determine a second loss value based on the first predicted reward and the first reward included in the first sample data;

[0032] A training module is used to iteratively train the action prediction model based on the first loss value and the second loss value of each of the multiple first sample data to obtain a target action prediction model.

[0033] In some embodiments, the acquisition module is used to:

[0034] Input multiple first state information into the simulator corresponding to the ramp merging scenario respectively, and the simulator outputs the first merging actions corresponding to the multiple first state information respectively, obtain the first reward obtained by the vehicle when performing the first merging action under each first state information, and obtain the first sample data including the first state information, the first merging action and the first reward.

[0035] In some embodiments, the acquisition module is further used to acquire second sample data at multiple second moments, where the second sample data is non-simulated data in a ramp merging scenario;

[0036] The training module is also used for:

[0037] Iteratively training the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a first action prediction model;

[0038] Inputting the second state information in the second sample data into the first action prediction model, and having the first action prediction model output second predicted action information and a second predicted reward;

[0039] Determine a third loss value based on the second predicted action information and the second action information included in the second sample data;

[0040] Determine a fourth loss value based on the second predicted reward and the second reward included in the second sample data;

[0041] Based on the third loss value and the fourth loss value of each of the plurality of second sample data, the first action prediction model is iteratively trained to obtain the target action prediction model.

[0042] In some embodiments, the first reward includes at least one of a comfort reward, an efficiency reward, and a safety reward; the comfort reward is determined based on the acceleration of the vehicle when performing the merging maneuver, the efficiency reward is determined based on the merging time of the vehicle when performing the merging maneuver, and the safety reward is determined based on at least one of the collision information after the vehicle performs the merging maneuver, the distance between the vehicle and an adjacent vehicle, and the number of lane changes.

[0043] In some embodiments, the ramp merging environment information includes at least one of lane information of a ramp merging area where the vehicle is located and a driving state of an adjacent vehicle of the vehicle;

[0044] The driving state of the vehicle includes at least one of the position, speed, acceleration and heading angle of the vehicle;

[0045] The merging action includes at least one of acceleration, deceleration and lane change.

[0046] In some embodiments, the acquisition module is further used to acquire target state information of the target vehicle in a ramp merging scenario, wherein the target state information includes ramp merging environment information of the target vehicle and the driving state of the target vehicle;

[0047] The input-output module is further used to input the target state information into the target action prediction model, and the target action prediction model outputs a target merging action. The target vehicle is used to execute the target merging action to merge in the ramp merging scenario.

[0048] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the above-mentioned training method of the action prediction model.

[0049] On the other hand, a computer-readable storage medium is provided, in which at least one program code is stored. The at least one program code is loaded and executed by a processor to implement the above-mentioned training method of the action prediction model.

[0050] On the other hand, a computer program product is provided, wherein the product stores at least one program code, and the at least one program code is used to be executed by a processor to implement the above-mentioned training method of the action prediction model.

[0051] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present disclosure.

[0052] An embodiment of the present application provides a method for training an action prediction model, which inputs vehicle state information into the action prediction model, and the action prediction model predicts the merging action and reward under the state information, and performs model training in combination with the loss of the reward and the loss of the action. In this way, the merging action predicted by the action prediction model trained by this method not only converges to the merging action in the sample data, so that the predicted merging action can achieve the successful merging of the vehicle, but also makes the predicted merging action converge in the direction of a larger reward. Since the reward also indicates the merging effect, the merging action predicted by the action prediction model trained by this method can also improve the merging effect, so that the vehicle can complete the merging more safely and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic diagram of an implementation environment of a method for training an action prediction model shown in an exemplary embodiment of the present application;

[0054] Figure 2is a flowchart of a method for training an action prediction model shown in an exemplary embodiment of the present application;

[0055] Figure 3 is a flowchart of a method for training an action prediction model shown in an exemplary embodiment of the present application;

[0056] Figure 4 is a block diagram of a training device for an action prediction model shown in an exemplary embodiment of the present application;

[0057] Figure 5 It is a block diagram of a computer device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the technical solutions and advantages of the present application clearer, the implementation methods of the present application are described in further detail below.

[0059] The terms "first", "second", "third" and "fourth" etc. in the specification and claims of the present application and the drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0060] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions. For example, the sample data involved in this application are all obtained with full authorization.

[0061] The training method of the action prediction model provided in the embodiment of the present application can be executed by a computer device. Figure 1 , Figure 1Schematic diagram of the implementation environment of the training method of the action prediction model provided in the embodiment of the present application. The computer device in the implementation environment is provided as a terminal 10 or a server 20, or as a terminal 10 and a server 20. The terminal 10 and the server 20 can be directly or indirectly connected by wired or wireless communication, and the present application is not limited here. The training method of the action prediction model provided in the embodiment of the present application can be executed by the terminal 10 alone, or by the server 20, or by the terminal 10 and the server 20 through data interaction, and the embodiment of the present application is not limited to this. In some embodiments, the server 20 undertakes the main computing work and the terminal 10 undertakes the secondary computing work; or, the server 20 undertakes the secondary computing service and the terminal 10 undertakes the main computing work; or, the server 20 and the terminal 10 use a distributed computing architecture for collaborative computing.

[0062] In some embodiments, the action prediction model provided in the embodiments of the present application is applied to an autonomous driving vehicle or a manually driven vehicle in a ramp merging scenario, so as to predict the merging action of the vehicle in the ramp merging scenario through the action prediction model, and then guide the vehicle to complete the ramp merging based on the merging action.

[0063] The terminal 10 is at least one of a mobile phone, a tablet computer, a PC (Personal Computer), etc. The server 20 can be at least one of a server, a server cluster consisting of multiple servers, a cloud server, a cloud computing platform, and a virtualization center.

[0064] Please refer to Figure 2 , which shows a flowchart of a method for training an action prediction model shown in an exemplary embodiment of the present application. Figure 2 , the method comprises the following steps.

[0065] Step 201, the computer device obtains multiple first sample data at multiple first moments, and the first sample data at each first moment includes first state information, first action information and a first reward. The first state information includes the ramp merging environment information and the driving state of the vehicle at the first moment. The first action information includes the merging action performed by the vehicle under the first state information. The first reward includes the reward obtained by the vehicle when performing the merging action under the state information, and the reward is used to indicate the merging effect.

[0066] The ramp merging environment information may be the ramp merging environment information in a high-speed ramp scenario, the ramp merging environment information in a city expressway scenario, or the ramp merging environment information in other ramp merging scenarios. In the embodiment of the present application, the ramp merging environment information is taken as the ramp merging environment information in a high-speed ramp scenario as an example for explanation, that is, the action prediction model is used to predict the merging action in the high-speed ramp merging scenario.

[0067] The ramp merging environment information includes at least one of the lane information of the ramp merging area where the vehicle is located and the driving status of the adjacent vehicles of the vehicle. The information of each lane includes at least one of the number of vehicles on the lane, the distance between vehicles, lane lines, traffic lights, etc.

[0068] The driving state of each vehicle includes at least one of the position, speed, acceleration and heading angle of the vehicle. Further, the position of the vehicle includes a lateral position and a longitudinal position. The speed of the vehicle includes a lateral speed and a longitudinal speed. The acceleration of the vehicle includes a lateral acceleration and a longitudinal acceleration.

[0069] The merging action includes at least one of acceleration, deceleration and lane change. Further, it includes the acceleration value, the deceleration value and the number of lane changes. The merging action indicated by the first action information in each first sample data is a merging action that enables the vehicle to merge successfully under the first state information in the first sample data.

[0070] In the embodiment of the present application, the state information in the sample data includes rich content, and on this basis, the predicted merging action is more closely matched with the current state of the vehicle.

[0071] In the embodiment of the present application, when vehicles merge, not only whether the vehicles can merge successfully is considered, but also other aspects of the impact of the vehicle merging are considered, that is, the merging effect is considered, such as the efficiency, passenger comfort, safety, etc. during merging. Therefore, a first reward for indicating the merging effect is also provided, so that on the basis of the vehicles successfully merging through the merging action, a safer, more efficient and comfortable merging can be achieved.

[0072] In some embodiments, the first reward includes at least one of a comfort reward, an efficiency reward, and a safety reward.

[0073] The comfort reward is determined based on the acceleration of the vehicle when performing the merging action. Acceleration includes lateral acceleration and longitudinal acceleration. The comfort reward is negatively correlated with acceleration, that is, the greater the acceleration, the smaller the comfort reward. Optionally, the lateral acceleration and the longitudinal acceleration are weighted and summed, and the negative or reciprocal of the weighted sum value is determined as the comfort reward. Alternatively, based on the weighted sum value, the comfort reward corresponding to the weighted sum value is determined from the correspondence between acceleration and comfort reward. The correspondence between acceleration and comfort reward includes comfort rewards corresponding to different accelerations.

[0074] Among them, the efficiency reward is determined based on the merging time of the vehicle performing the merging action. The merging time refers to the time taken to merge from the ramp to the main road. The efficiency reward is negatively correlated with the merging time, that is, the shorter the merging time, the higher the efficiency reward. Optionally, the negative or reciprocal of the merging time is determined as the efficiency reward. Alternatively, based on the merging time, the efficiency reward corresponding to the merging time is determined from the corresponding relationship between the merging time and the efficiency reward. The corresponding relationship between the merging time and the efficiency reward includes efficiency rewards corresponding to different merging times.

[0075] The safety reward is determined based on at least one of the collision information after the vehicle performs the merging action, the distance between the vehicle and the adjacent vehicle, and the number of lane changes.

[0076] The collision information is used to indicate whether the vehicle will collide. If the safety reward is determined based on the collision information, if a collision occurs, the safety reward is a first value. If no collision occurs, the safety reward is a second value, and the first value is much smaller than the second value.

[0077] If the safety reward is determined based on the distance between the vehicle and the adjacent vehicle, since the greater the distance, the higher the safety, the safety reward is positively correlated with the distance, that is, the greater the distance, the greater the safety reward. Optionally, if there are multiple adjacent vehicles, the average of the multiple distances is determined to obtain the safety reward. In some embodiments, based on the distance, the safety reward corresponding to the distance is determined from the corresponding relationship between the distance and the safety reward. The corresponding relationship between the distance and the safety reward includes safety rewards corresponding to different distances.

[0078] If the safety reward is determined based on the number of lane changes, since the more lane changes, the worse the safety, the safety reward is negatively correlated with the number of lane changes, that is, the more lane changes, the smaller the safety reward. In some embodiments, based on the number of lane changes, the safety reward corresponding to the number of lane changes is determined from the correspondence between the number of lane changes and the safety reward. The correspondence between the number of lane changes and the safety reward includes safety rewards corresponding to different numbers of lane changes.

[0079] Optionally, if the safety reward is determined based on at least two of the collision information, the distance and the number of lane changes, a weighted sum is taken of the safety rewards corresponding to the at least two items to obtain a comprehensive safety reward.

[0080] In the embodiment of the present application, if the first reward includes the above three rewards, the three rewards are weighted and summed to obtain the first reward. The weight of each reward can be set as needed.

[0081] Optionally, before weighted summation, the three rewards are normalized to map the values ​​of the three rewards to the interval [0,1], thereby making the determined first reward more accurate and reasonable, and avoiding the excessive bias of the first reward due to the excessive value of a certain reward.

[0082] In the embodiment of the present application, the first reward is taken as the reward obtained after performing the current merging action under the current state information. In other embodiments, the first reward is a cumulative reward, which includes the sum of all future rewards that can be obtained by merging based on the merging action starting from the current state information and the merging action. This sum is obtained by weighted summation of each possible future reward. Among them, each reward is multiplied by the corresponding power of a discount factor to reflect its value relative to the current time point.

[0083] Among them, the ramp merging environment information is modeled to obtain a simulator. Optionally, a simulation environment of the high-speed ramp merging area is constructed using traffic simulation software, and information such as the road environment and vehicle status is obtained. Among them, the environment should include key elements such as lane lines, traffic signals, and vehicles, and set reasonable road parameters and traffic rules. And initialize the vehicle state in the simulation environment, including the vehicle's position, speed, acceleration, etc. At the same time, set the initial state of the simulation environment, such as traffic flow, vehicle distribution, etc. Then perform state perception, that is, obtain the state information of the vehicle and the surrounding environment in real time through sensors or other means. Among them, key information such as road environment and vehicle status is obtained in real time through the interface provided by the simulation software to obtain the state information in the sample data. In addition, collect data such as the driving trajectory and speed change of the vehicle in the simulation environment to obtain the merging action and reward in the sample data.

[0084] Step 202: The computer device inputs the first state information in the first sample data into the action prediction model, and the action prediction model outputs the first predicted action information and the first predicted reward.

[0085] In some embodiments, the action prediction model includes a policy network (Actor) and an evaluation network (Critic). The policy network is used to predict action information, and the evaluation network is used to predict rewards, that is, to evaluate the value of action information predicted for the current state information. Among them, the evaluation network is responsible for evaluating the quality of the current strategy, that is, evaluating the value of the current state or action in the state, while the policy network is responsible for generating confluent actions. The output of the policy network is used to update the policy network to guide the strategy to converge in the direction of greater cumulative rewards.

[0086] In this embodiment, the PPO (Proximal Policy Optimization) deep reinforcement learning method using the Actor-Critic architecture can effectively solve the problem of easily falling into state space explosion when processing complex state space.

[0087] The first predicted reward is the reward obtained by performing the merging action indicated by the first predicted action information under the first state information. The calculation method is the same as the calculation method of the first reward, which will not be repeated here.

[0088] Step 203: The computer device determines a first loss value based on the first predicted action information and the first action information included in the first sample data.

[0089] The computer device determines difference information between the first predicted action information and the first action information, and obtains a first loss value based on the difference information. The first loss value may be a mean square error loss, a mean absolute error loss, a cross entropy loss, or a logarithmic loss, etc., which are not specifically limited here.

[0090] Optionally, if the first predicted action information includes multiple information such as speed changes and lane changes, then optionally, the computer device converts these multiple information into feature vectors of the same dimension respectively, splices the feature vectors corresponding to the multiple information respectively, and obtains the feature vectors corresponding to the first predicted action information and the first action information respectively, thereby facilitating the calculation of the first loss value based on the feature vector.

[0091] Step 204: The computer device determines a second loss value based on the first predicted reward and the first reward included in the first sample data.

[0092] The computer device determines the second loss value based on the difference information between the first predicted reward and the first reward. The second loss value may be a mean square error loss, a mean absolute error loss, a cross entropy loss, or a logarithmic loss, etc., which are not specifically limited here.

[0093] Optionally, if the first reward includes a weighted sum of multiple rewards such as a comfort reward, an efficiency reward and a safety reward, and the first predicted reward also includes a weighted sum of these multiple rewards, then the difference information between the first predicted reward and the first reward is the difference between the two.

[0094] Step 205: iteratively train the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a target action prediction model.

[0095] In some embodiments, the computer device obtains a third loss value based on the first loss value and the second loss value, and iteratively trains the action prediction model based on the third loss value to obtain a target action prediction model. The third loss value is the average of the first loss value and the second loss value, or the third loss value is the weighted sum of the first loss value and the second loss value, and the weights of the first loss value and the second loss value can be set as needed, which will not be repeated here.

[0096] The computer device iteratively trains the action prediction model based on the third loss value until an iteration stop condition is reached to obtain a target action prediction model.

[0097] In some embodiments, the iteration stop condition includes at least one of the following: the number of iterations reaches a first threshold, the third loss value reaches convergence or is less than a second threshold. The iteration stop condition refers to the criterion for when to terminate the iteration process. The first threshold refers to the maximum value of the number of iterations. The second threshold represents the minimum value of the error acceptable to the action prediction model.

[0098] In an embodiment of the present application, the first sample data may be real data in a ramp merging scenario. In other embodiments, the first sample data is simulated data in a ramp merging scenario, that is, obtained by a simulator corresponding to the ramp merging scenario, wherein a plurality of first state information are respectively input into the simulator corresponding to the ramp merging scenario, and the simulator outputs a first merging action corresponding to a plurality of first state information, and obtains a first reward obtained by the vehicle for performing the first merging action under each first state information, thereby obtaining the first sample data including the first state information, the first merging action, and the first reward.

[0099] In this embodiment, since the ramp merging scenario is generally an actual application scenario, the sample data acquisition cost in this scenario is high and the acquisition process is time-consuming; in some road test scenarios, the vehicle needs to actually run in the scenario to obtain sample data. However, the sample data acquisition cost in the simulation scenario is low and the acquisition process is relatively simple. In the simulation scenario of ramp merging, the simulator can quickly obtain sample data at multiple simulation moments, improve data acquisition efficiency, and reduce data acquisition cost.

[0100] Since there is still a deviation between the simulation data and the real data, optionally, in order to further improve the performance of the action prediction model, a plurality of second sample data at the second moment is also obtained, and the second sample data is non-simulated data in the ramp merging scenario; then the above process of iteratively training the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain the target action prediction model includes the following implementation methods:

[0101] The computer device iteratively trains the action prediction model based on the first loss value and the second loss value of each of the multiple first sample data to obtain a first action prediction model; inputs the second state information in the second sample data into the first action prediction model, and the first action prediction model outputs the second predicted action information and the second predicted reward; determines the third loss value based on the second predicted action information and the second action information included in the second sample data; determines the fourth loss value based on the second predicted reward and the second reward included in the second sample data; iteratively trains the first action prediction model based on the third loss value and the fourth loss value of each of the multiple second sample data to obtain a target action prediction model.

[0102] It should be noted that obtaining sample data based on a simulator can reduce data acquisition costs and improve data acquisition efficiency, but there is still a deviation between the sample data obtained in this way and the actual sample data. Therefore, after the motion prediction model is preliminarily trained based on simulated sample data, the present application only needs to be based on a small amount of sample data in real scenes to obtain a motion prediction model suitable for real scenes, thereby reducing the amount of sample data obtained in real scenes, reducing costs, and ensuring that the motion prediction model can accurately predict real scenes.

[0103] In an embodiment of the present application, the trained target action prediction model is deployed in an actual traffic environment, so that a vehicle control system or a traffic management system can realize real-time calling and execution of the target action prediction model.

[0104] In some embodiments, the target action prediction model is also tested in an actual traffic environment to verify the performance and effect of the target action prediction model. The performance of the target action prediction model can be evaluated by comparing indicators such as traffic flow, vehicle speed, and merging time before and after the test. Accordingly, the target action prediction model can also be optimized and adjusted based on the test results. For example, the performance and generalization ability of the target action prediction model can be improved by adjusting the model parameters of the target action prediction model, improving the algorithm structure, etc.

[0105] Furthermore, since the traffic environment may change and traffic rules may be updated, the target action prediction model may optionally be updated and maintained periodically. Sample data corresponding to the changed traffic environment or traffic rules may be obtained, and the target action prediction model may be adjusted based on the new sample data to adapt the target action prediction model to the new traffic environment and traffic rules.

[0106] For example, see Figure 3 , Figure 3 A training flow chart of an action prediction model shown in an exemplary embodiment of the present application is shown. Among them, the simulation environment of ramp merging and the initialization of state information are first performed. Then, state perception and data collection are performed to obtain multiple sample data, and then an action prediction model including an Actor network and a Critic network is constructed, and the model is trained based on the sample data. The trained target action model is deployed to the actual application scenario, and the model is performance tested. Based on the test results, samples are collected to adjust the model. Then, it is determined whether the performance of the model is improved. If it is improved, the merging action in the ramp merging scenario is predicted based on the model; if it is not improved, the model is continuously optimized to update the model, and then the merging action in the ramp merging scenario is predicted based on the updated model.

[0107] The solution provided in the embodiment of the present application adopts deep reinforcement learning technology to combine the perception ability of deep learning with the decision-making ability of reinforcement learning. Through this solution, a merging action prediction system for vehicles in an intelligent ramp merging scenario can be constructed. The system can perceive traffic environment information in real time, learn and make decisions through deep reinforcement learning algorithms, generate optimal merging control strategies, and guide vehicles to complete the ramp merging process safely and efficiently, thereby improving merging efficiency and safety.

[0108] An embodiment of the present application provides a method for training an action prediction model, which inputs vehicle state information into the action prediction model, and the action prediction model predicts the merging action and reward under the state information, and performs model training in combination with the loss of the reward and the loss of the action. In this way, the merging action predicted by the action prediction model trained by this method not only converges to the merging action in the sample data, so that the predicted merging action can achieve the successful merging of the vehicle, but also makes the predicted merging action converge in the direction of a larger reward. Since the reward also indicates the merging effect, the merging action predicted by the action prediction model trained by this method can also improve the merging effect, so that the vehicle can complete the merging more safely and efficiently.

[0109] Please refer to Figure 4 , which shows a block diagram of a training device for an action prediction model shown in an exemplary embodiment of the present application. The device includes:

[0110] An acquisition module 401 is used to acquire a plurality of first sample data at a plurality of first moments, wherein the first sample data at each first moment includes first state information, first action information, and a first reward, wherein the first state information includes ramp merging environment information of the vehicle at the first moment and the driving state of the vehicle, the first action information includes a merging action performed by the vehicle under the first state information, and the first reward includes a reward obtained by the vehicle for performing the merging action under the first state information;

[0111] An input-output module 402, configured to input the first state information in the first sample data into an action prediction model, and the action prediction model outputs first predicted action information and a first predicted reward;

[0112] A determination module 403, configured to determine a first loss value based on the first predicted action information and the first action information included in the first sample data;

[0113] The determination module 403 is further configured to determine a second loss value based on the first predicted reward and the first reward included in the first sample data;

[0114] The training module 404 is used to iteratively train the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a target action prediction model.

[0115] In some embodiments, the acquisition module 401 is used to:

[0116] Input multiple first state information into the simulator corresponding to the ramp merging scenario respectively, and the simulator outputs the first merging actions corresponding to the multiple first state information respectively, obtain the first reward obtained by the vehicle when performing the first merging action under each first state information, and obtain the first sample data including the first state information, the first merging action and the first reward.

[0117] In some embodiments, the acquisition module 401 is further used to acquire second sample data at multiple second moments, where the second sample data is non-simulated data in a ramp merging scenario;

[0118] The training module 404 is further used for:

[0119] Iteratively training the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a first action prediction model;

[0120] Inputting the second state information in the second sample data into the first action prediction model, and the first action prediction model outputs the second predicted action information and the second predicted reward;

[0121] Determine a third loss value based on the second predicted action information and the second action information included in the second sample data;

[0122] Determine a fourth loss value based on the second predicted reward and the second reward included in the second sample data;

[0123] Based on the third loss value and the fourth loss value of each of the multiple second sample data, the first action prediction model is iteratively trained to obtain a target action prediction model.

[0124] In some embodiments, the first reward includes at least one of a comfort reward, an efficiency reward, and a safety reward; the comfort reward is determined based on the acceleration of the vehicle when performing the merging maneuver, the efficiency reward is determined based on the merging time of the vehicle when performing the merging maneuver, and the safety reward is determined based on at least one of the collision information after the vehicle performs the merging maneuver, the distance between the vehicle and the adjacent vehicle, and the number of lane changes.

[0125] In some embodiments, the ramp merging environment information includes at least one of information of each lane in the ramp merging area where the vehicle is located and driving status of adjacent vehicles of the vehicle;

[0126] The driving state of the vehicle includes at least one of the position, speed, acceleration and heading angle of the vehicle;

[0127] The merging maneuver includes at least one of acceleration, deceleration and lane change.

[0128] In some embodiments, the acquisition module 401 is further used to acquire target state information of the target vehicle in a ramp merging scenario, where the target state information includes ramp merging environment information where the target vehicle is located and the driving state of the target vehicle;

[0129] The input-output module 402 is further used to input the target state information into the target action prediction model, and the target action prediction model outputs the target merging action. The target vehicle is used to execute the target merging action to merge in the ramp merging scenario.

[0130] An embodiment of the present application provides a training device for an action prediction model, which inputs vehicle state information into the action prediction model, and the action prediction model predicts the merging action and reward under the state information, and performs model training in combination with the loss of the reward and the loss of the action, so that the merging action predicted based on the trained target action prediction model not only converges to the merging action in the sample data, so that the predicted merging action enables the vehicle to merge successfully, but also makes the predicted merging action converge in the direction of a larger reward. Since the reward also indicates the merging effect, the merging action predicted by the model trained by the device can also improve the merging effect, so that the vehicle can complete the merging more safely and efficiently.

[0131] It should be noted that the training device for the motion prediction model provided in the above embodiment is only illustrated by the division of the above functional modules when performing vehicle control. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the motion prediction model provided in the above embodiment and the training method embodiment of the motion prediction model belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0132] Please refer to Figure 5 , Figure 5 The block diagram of a computer device 500 provided by an exemplary embodiment of the present application is shown. The computer device 500 may be a portable mobile computer device, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The computer device 500 may also be referred to as a user device, a portable computer device, a laptop computer device, a desktop computer device, or other names.

[0133] Typically, the computer device 500 includes a processor 501 and a memory 502 .

[0134] The processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0135] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 is used to store at least one program code, which is used to be executed by the processor 501 to implement the operation performed by the computer device in the training method of the action prediction model provided in the method embodiment of the present application.

[0136] In some embodiments, the computer device 500 may also optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, the memory 502 and the peripheral device interface 503 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 503 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507 and a power supply 508.

[0137] The peripheral device interface 503 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.

[0138] The radio frequency circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 504 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 504 can communicate with other computer devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or WiFi ( WirelessFidelity In some embodiments, the radio frequency circuit 504 may also include a circuit related to NFC (Near Field Communication), which is not limited in the present application.

[0139] The display screen 505 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 505 is a touch display screen, the display screen 505 also has the ability to collect touch signals on the surface or above the surface of the display screen 505. The touch signal can be input to the processor 501 as a control signal for processing. At this time, the display screen 505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 505 can be one, arranged on the front panel of the computer device 500; in other embodiments, the display screen 505 can be at least two, respectively arranged on different surfaces of the computer device 500 or in a folding design; in other embodiments, the display screen 505 can be a flexible display screen, arranged on a curved surface or a folding surface of the computer device 500. Even, the display screen 505 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 505 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode, organic light-emitting diode).

[0140] The camera assembly 506 is used to capture images or videos. Optionally, the camera assembly 506 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the computer device, and the rear camera is set on the back of the computer device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 506 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0141] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 501 for processing, or input them into the radio frequency circuit 504 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the computer device 500. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 507 may also include a headphone jack.

[0142] The power supply 508 is used to power various components in the computer device 500. The power supply 508 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 508 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged through a wired line, and a wireless rechargeable battery is a battery that is charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0143] In some embodiments, the computer device 500 further includes one or more sensors 509 , including but not limited to: an acceleration sensor 510 , a gyroscope sensor 511 , a pressure sensor 512 , an optical sensor 513 , and a proximity sensor 514 .

[0144] The acceleration sensor 510 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the computer device 500. For example, the acceleration sensor 510 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 501 can control the display screen 505 to display the user interface in a horizontal view or a vertical view based on the gravity acceleration signal collected by the acceleration sensor 510. The acceleration sensor 510 can also be used to collect motion data of games or users.

[0145] The gyro sensor 511 can detect the body direction and rotation angle of the computer device 500, and the gyro sensor 511 can cooperate with the acceleration sensor 510 to collect the user's 3D actions on the computer device 500. Based on the data collected by the gyro sensor 511, the processor 501 can implement the following functions: motion sensing (such as changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0146] The pressure sensor 512 can be set on the side frame of the computer device 500 and / or the lower layer of the display screen 505. When the pressure sensor 512 is set on the side frame of the computer device 500, it can detect the user's grip signal of the computer device 500, and the processor 501 performs left and right hand recognition or shortcut operation based on the grip signal collected by the pressure sensor 512. When the pressure sensor 512 is set on the lower layer of the display screen 505, the processor 501 controls the operability controls on the UI interface based on the user's pressure operation on the display screen 505. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0147] The optical sensor 513 is used to collect the ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display screen 505 based on the ambient light intensity collected by the optical sensor 513. Specifically, when the ambient light intensity is high, the display brightness of the display screen 505 is increased; when the ambient light intensity is low, the display brightness of the display screen 505 is reduced. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera component 506 based on the ambient light intensity collected by the optical sensor 513.

[0148] The proximity sensor 514, also called a distance sensor, is usually disposed on the front panel of the computer device 500. The proximity sensor 514 is used to collect the distance between the user and the front of the computer device 500. In one embodiment, when the proximity sensor 514 detects that the distance between the user and the front of the computer device 500 is gradually decreasing, the processor 501 controls the display screen 505 to switch from the screen-on state to the screen-off state; when the proximity sensor 514 detects that the distance between the user and the front of the computer device 500 is gradually increasing, the processor 501 controls the display screen 505 to switch from the screen-off state to the screen-on state.

[0149] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the computer device 500, and the computer device 500 may include more or less components than those shown in the figure, or combine some components, or adopt a different arrangement of components.

[0150] The embodiment of the present application also provides a computer-readable storage medium, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to implement the training method of the action prediction model described in any of the above implementations. Optionally, the storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, and an optical data storage device.

[0151] An embodiment of the present application also provides a computer program product, which stores at least one program code, and the at least one program code is loaded and executed by a processor to implement the training method of the action prediction model shown in the above embodiments.

[0152] In some embodiments, the computer program product involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected by a communication network. Multiple computer devices distributed at multiple locations and interconnected by a communication network may constitute a blockchain system.

[0153] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0154] The above description is only for the purpose of facilitating those skilled in the art to understand the technical solution of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for training an action prediction model, characterized in that: The method comprises: Acquire a plurality of first sample data at a plurality of first moments, wherein the first sample data at each first moment includes first state information, first action information, and a first reward, wherein the first state information includes ramp merging environment information of the vehicle at the first moment and the driving state of the vehicle, the first action information includes a merging action performed by the vehicle under the first state information, and the first reward includes a reward obtained by the vehicle for performing the merging action under the first state information, and the reward is used to indicate a merging effect; Inputting the first state information in the first sample data into an action prediction model, and having the action prediction model output first predicted action information and a first predicted reward; Determine a first loss value based on the first predicted action information and the first action information included in the first sample data; Determine a second loss value based on the first predicted reward and the first reward included in the first sample data; Based on the first loss value and the second loss value of each of the multiple first sample data, the action prediction model is iteratively trained to obtain a target action prediction model.

2. The method according to claim 1, characterized in that The obtaining of a plurality of first sample data at a plurality of first moments comprises: Input multiple first state information into the simulator corresponding to the ramp merging scenario respectively, and the simulator outputs the first merging actions corresponding to the multiple first state information respectively, obtain the first reward obtained by the vehicle when performing the first merging action under each first state information, and obtain the first sample data including the first state information, the first merging action and the first reward.

3. The method according to claim 2, characterized in that The method further comprises: Acquire a plurality of second sample data at second moments, where the second sample data is non-simulated data in a ramp merging scenario; The iterative training of the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a target action prediction model includes: Iteratively training the action prediction model based on the first loss value and the second loss value of each of the plurality of first sample data to obtain a first action prediction model; Inputting the second state information in the second sample data into the first action prediction model, and having the first action prediction model output second predicted action information and a second predicted reward; Determine a third loss value based on the second predicted action information and the second action information included in the second sample data; Determine a fourth loss value based on the second predicted reward and the second reward included in the second sample data; Based on the third loss value and the fourth loss value of each of the plurality of second sample data, the first action prediction model is iteratively trained to obtain the target action prediction model.

4. The method according to claim 1, characterized in that: The first reward includes at least one of a comfort reward, an efficiency reward and a safety reward; the comfort reward is determined based on the acceleration of the vehicle when performing the merging action, the efficiency reward is determined based on the merging time of the vehicle when performing the merging action, and the safety reward is determined based on at least one of the collision information after the vehicle performs the merging action, the distance between the vehicle and the adjacent vehicle, and the number of lane changes.

5. The method according to claim 1, characterized in that The ramp merging environment information includes at least one of the lane information of the ramp merging area where the vehicle is located and the driving status of the adjacent vehicles of the vehicle; The driving state of the vehicle includes at least one of the position, speed, acceleration and heading angle of the vehicle; The merging action includes at least one of acceleration, deceleration and lane change.

6. The method according to claim 1, characterized in that The method further comprises: Acquire target state information of a target vehicle in a ramp merging scenario, wherein the target state information includes ramp merging environment information of the target vehicle and a driving state of the target vehicle; The target state information is input into the target action prediction model, and the target action prediction model outputs a target merging action. The target vehicle is used to execute the target merging action to merge in the ramp merging scenario.

7. A training device for an action prediction model, characterized in that: The device comprises: an acquisition module, configured to acquire a plurality of first sample data at a plurality of first moments, wherein the first sample data at each first moment includes first state information, first action information, and a first reward, wherein the first state information includes ramp merging environment information of the vehicle at the first moment and the driving state of the vehicle, the first action information includes a merging action performed by the vehicle under the first state information, and the first reward includes a reward obtained by the vehicle for performing the merging action under the first state information, and the reward is used to indicate a merging effect; An input-output module, configured to input the first state information in the first sample data into an action prediction model, and the action prediction model outputs first predicted action information and a first predicted reward; A determination module, configured to determine a first loss value based on the first predicted action information and the first action information included in the first sample data; The determination module is further configured to determine a second loss value based on the first predicted reward and the first reward included in the first sample data; A training module is used to iteratively train the action prediction model based on the first loss value and the second loss value of each of the multiple first sample data to obtain a target action prediction model.

8. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the training method of the action prediction model as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores at least one program code, and the at least one program code is loaded and executed by the processor to implement the training method of the action prediction model as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The product stores at least one program code, and the at least one program code is used to be executed by a processor to implement the training method of the action prediction model as described in any one of claims 1 to 6.