Information update method, route screening method, device, equipment and medium

By obtaining and filtering the waiting routes for candidate vehicles of autonomous vehicles and updating the target return function parameters, the problems of learning time and waste of computing resources in inverse reinforcement learning are solved, and efficient and accurate parameter updates are achieved.

CN114167874BActive Publication Date: 2025-06-20JINGDONG KUNPENG (JIANGSU) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111497556.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-06-20
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

When inverse reinforcement learning is used in autonomous vehicles to learn return function parameters, there are problems such as learning time, random results, and waste of computing resources.

Method used

By obtaining the initial vehicle status information of the target vehicle, a route to be driven by the candidate vehicle is generated, a route that meets the preset conditions is selected, and the parameter information of the target return function is updated according to the route.

Benefits of technology

It realizes the rapid and efficient update of the target return function parameters, reduces the waste of computing resources, and improves learning efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114167874B_ABST
    Figure CN114167874B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an information update method, a route screening method, an apparatus, an electronic device, and a computer-readable medium. A specific implementation of the method includes: obtaining at least one initial vehicle state information for a target vehicle; generating at least one candidate vehicle to-be-traveled route according to the at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled strategy set; screening out candidate vehicle to-be-traveled routes that meet preset conditions from the at least one candidate vehicle to-be-traveled route as first target candidate vehicle to-be-traveled routes, to obtain at least one first target candidate vehicle to-be-traveled route; and updating parameter information in a target reward function according to the at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route. This implementation can quickly and efficiently update the parameter information in the target reward function, reducing waste of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to an information update method, a route screening method, an apparatus, an electronic device, and a medium. Background Art

[0002] Currently, autonomous vehicles can effectively improve the safety, comfort, efficiency, and economy of vehicles through intelligent technologies, and the autonomous decision-making ability is the core embodiment of their intelligence. Among them, for the learning of the parameters of the decision-making reward function of autonomous vehicles, the learning of the parameters of the reward function is usually realized by using the traditional inverse reinforcement learning (IRL).

[0003] However, when using the above method to learn the parameters of the reward function, the following technical problems often exist:

[0004] Since inverse reinforcement learning highly depends on the reinforcement learning (RL) module for learning the optimal policy, the learning takes a long time. In addition, there may be a problem of result randomness when using reinforcement learning. Further, the amount of computation is increased, resulting in a waste of computer resources and indirectly increasing the processing load of the computer. Summary of the Invention

[0005] This section of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation section. This section of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure propose an information update method, an apparatus, an electronic device, and a computer-readable medium to solve the technical problems mentioned in the above background art section.

[0007] In a first aspect, some embodiments of the present disclosure provide an information update method, including: obtaining at least one initial vehicle state information for a target vehicle; generating at least one candidate vehicle to-be-traveled route according to the at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled policy set; screening out candidate vehicle to-be-traveled routes that meet a preset condition from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes, to obtain at least one first target candidate vehicle to-be-traveled route; updating parameter information in a target reward function according to the at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route, where the target reward function is a function for determining a vehicle route.

[0008] Optionally, generating at least one candidate vehicle to-be-traveled route according to the above at least one initial vehicle state information and the pre-selected candidate vehicle to-be-traveled strategy set includes: expanding the candidate vehicle to-be-traveled strategy set according to the pre-set action space to obtain an expanded candidate vehicle to-be-traveled strategy set; generating the at least one candidate vehicle to-be-traveled route according to the expanded candidate vehicle to-be-traveled strategy set and the above at least one initial vehicle state information.

[0009] Optionally, screening out candidate vehicle to-be-traveled routes that meet the preset conditions from the above at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes to obtain at least one first target candidate vehicle to-be-traveled route includes: determining the probability information corresponding to each candidate vehicle to-be-traveled route in the above at least one candidate vehicle to-be-traveled route according to the initial parameters of the above target reward function to obtain a probability information set; screening out candidate vehicle to-be-traveled routes whose probability information meets the target conditions from the above at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes to obtain the above at least one first target candidate vehicle to-be-traveled route.

[0010] Optionally, updating the parameter information in the target reward function according to the above at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route includes: generating a first route feature information set group corresponding to the above at least one first target candidate vehicle to-be-traveled route and a second route feature information set group corresponding to the above at least one preset vehicle to-be-traveled route; updating the parameter information in the above target reward function according to the first route feature information set group and the second route feature information set group.

[0011] Optionally, the above method further includes: screening out candidate vehicle to-be-traveled routes that meet the above preset conditions from the above at least one candidate vehicle to-be-traveled route according to the parameter information after the update of the above target reward function as the second target candidate vehicle to-be-traveled routes to obtain at least one second target candidate vehicle to-be-traveled route; determining a second difference according to the above at least one second target candidate vehicle to-be-traveled route and the above at least one preset vehicle to-be-traveled route; in response to determining that the second difference is less than or equal to the target threshold, determining the parameter information after the update of the above target reward function as the parameter information after the training of the above target reward function.

[0012] Optionally, updating the parameter information in the target return function according to the first set of route feature information and the second set of route feature information includes: inputting each first set of route feature information in the first set of route feature information into a first objective function to generate a first value, obtaining a first value set; performing weighted average processing on each first value in the first value set to obtain a first weighted value; inputting each second set of route feature information in the second set of route feature information into a second objective function to generate a second value, obtaining a second value set; performing weighted average processing on each second value in the second value set to obtain a second weighted value; determining a first difference between the first weighted value and the second weighted value; and in response to determining that the first difference is greater than a target threshold, updating the parameter information in the target return function.

[0013] Optionally, the candidate vehicle driving strategy set is selected through the following steps: sampling the target vehicle driving methods in the target vehicle driving method set to obtain at least one target vehicle driving method as the candidate vehicle driving strategy set.

[0014] Optionally, the target candidate vehicle driving strategy set is selected through the following steps: obtaining vehicle control information sets in each direction for the target vehicle, where the vehicle control information in the vehicle control information set contains vehicle feedback information; assigning different values to the vehicle feedback information corresponding to each vehicle control information in the vehicle control information set to obtain a vehicle driving method set; and screening out vehicle driving methods that match the environmental information from the vehicle driving method set according to the environmental information as the target candidate vehicle driving strategies to obtain the target candidate vehicle driving strategy set.

[0015] In a second aspect, some embodiments of the present disclosure provide a route screening method, including: obtaining at least one candidate vehicle driving route; and screening out candidate vehicle driving routes that meet preset conditions from the at least one candidate vehicle driving route according to a target return function as target candidate vehicle driving routes to obtain at least one target candidate vehicle driving route.

[0016] In a third aspect, some embodiments of the present disclosure provide an information update device, including: a first acquisition unit configured to acquire at least one initial vehicle state information for a target vehicle; a generation unit configured to generate at least one candidate vehicle to-be-traveled route according to the at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled strategy set; a first screening unit configured to screen out candidate vehicle to-be-traveled routes that meet a preset condition from the at least one candidate vehicle to-be-traveled route as a first target candidate vehicle to-be-traveled route, obtaining at least one first target candidate vehicle to-be-traveled route; and an update unit configured to update parameter information in a target reward function according to the at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route, where the target reward function is a function for determining a vehicle route.

[0017] Optionally, the generation unit is configured to: expand the candidate vehicle to-be-traveled strategy set according to a preset action space to obtain an expanded candidate vehicle to-be-traveled strategy set; and generate the at least one candidate vehicle to-be-traveled route according to the expanded candidate vehicle to-be-traveled strategy set and the at least one initial vehicle state information.

[0018] Optionally, the first screening unit is configured to: determine probability information corresponding to each candidate vehicle to-be-traveled route in the at least one candidate vehicle to-be-traveled route according to initial parameters of the target reward function, obtaining a probability information set; and screen out candidate vehicle to-be-traveled routes whose probability information meets a target condition from the at least one candidate vehicle to-be-traveled route as a first target candidate vehicle to-be-traveled route, obtaining the at least one first target candidate vehicle to-be-traveled route.

[0019] Optionally, the update unit is configured to: generate a first route feature information set group corresponding to the at least one first target candidate vehicle to-be-traveled route and a second route feature information set group corresponding to the at least one preset vehicle to-be-traveled route; and update the parameter information in the target reward function according to the first route feature information set group and the second route feature information set group.

[0020] Optionally, the device further includes: screening out candidate vehicle to-be-traveled routes that meet the preset condition from the at least one candidate vehicle to-be-traveled route according to the parameter information updated by the target reward function as a second target candidate vehicle to-be-traveled route, obtaining at least one second target candidate vehicle to-be-traveled route; determining a second difference according to the at least one second target candidate vehicle to-be-traveled route and the at least one preset vehicle to-be-traveled route; and in response to determining that the second difference is less than or equal to a target threshold, determining the parameter information updated by the target reward function as the parameter information after training of the target reward function.

[0021] Optionally, the updating unit is configured to: input each first route feature information set in the above first route feature information set group into a first objective function to generate a first numerical value, obtaining a first numerical value set; perform a weighted average process on each first numerical value in the above first numerical value set to obtain a first weighted numerical value; input each second route feature information set in the above second route feature information set group into a second objective function to generate a second numerical value, obtaining a second numerical value set; perform a weighted average process on each second numerical value in the above second numerical value set to obtain a second weighted numerical value; determine a first difference between the above first weighted numerical value and the above second weighted numerical value; in response to determining that the above first difference is greater than a target threshold, update the parameter information in the above target reward function.

[0022] Optionally, the above candidate vehicle driving strategies set is selected through the following steps: sampling the target vehicle driving methods in the target vehicle driving method set to obtain at least one target vehicle driving method as the candidate vehicle driving strategies set.

[0023] Optionally, the above target candidate vehicle driving strategies set is selected through the following steps: obtaining vehicle control information sets in each direction for the above target vehicle, where the vehicle control information in the above vehicle control information set is information containing vehicle feedback information; assigning different numerical values to the vehicle feedback information corresponding to each vehicle control information in the above vehicle control information set to obtain a vehicle driving method set; screening out vehicle driving methods matching the environmental information from the above vehicle driving method set as the target candidate vehicle driving strategies to obtain the above target candidate vehicle driving strategies set.

[0024] In a fourth aspect, some embodiments of the present disclosure provide a route screening device, including: a second obtaining unit configured to: obtain at least one candidate vehicle driving route; a second screening unit configured to: screen out candidate vehicle driving routes that meet a preset condition from the above at least one candidate vehicle driving route according to a target reward function as target candidate vehicle driving routes, obtaining at least one target candidate vehicle driving route, where the parameter information in the above target reward function is updated by using the Figure 2 method corresponding to some embodiments.

[0025] In a fifth aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect or the second aspect.

[0026] Sixth aspect, some embodiments of the present disclosure provide a computer-readable medium, on which a computer program is stored, where when the program is executed by a processor, the method described in any implementation manner of the first aspect or the second aspect is implemented.

[0027] The above various embodiments of the present disclosure have the following beneficial effects: The information update method of some embodiments of the present disclosure can update the parameter information in the target return function quickly and efficiently, reducing the waste of computing resources. Since inverse reinforcement learning highly depends on the reinforcement learning module for learning the optimal policy, the learning takes a long time. In addition, there may be a problem of result randomness when using reinforcement learning. Further, the amount of computation is increased, resulting in the waste of computer resources and indirectly increasing the processing load of the computer. Based on this, the information update method of some embodiments of the present disclosure can first obtain at least one initial vehicle state information for the target vehicle. Here, obtaining at least one initial vehicle state information is used to determine the candidate vehicle's to-be-traveled route subsequently. Then, according to the above at least one initial vehicle state information and a pre-selected candidate vehicle's to-be-traveled policy set, at least one candidate vehicle's to-be-traveled route is generated. Here, at least one initial vehicle state information can be used as the starting point, and the candidate vehicle's to-be-traveled policy set can be used as the driving policy of the target vehicle to accurately and efficiently generate at least one candidate vehicle's to-be-traveled route. Furthermore, the candidate vehicle's to-be-traveled routes that meet the preset conditions are screened out from the above at least one candidate vehicle's to-be-traveled route as the first target candidate vehicle's to-be-traveled route, obtaining at least one first target candidate vehicle's to-be-traveled route. Here, the purpose of screening the route from at least one candidate vehicle's to-be-traveled route is to screen out a more accurate and excellent route from several candidate vehicle's to-be-traveled routes corresponding to a certain initial vehicle state information. Based on this, the parameter information update of the subsequent target return function can be made more accurate and efficient. Finally, by comparing the differences between the above at least one first target candidate vehicle's to-be-traveled route and at least one preset vehicle's to-be-traveled route, the parameter information in the target return function can be updated quickly and efficiently. Among them, the above target return function is a function for determining the vehicle route. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In combination with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0029] Figure 1 is a schematic diagram of an application scenario of the information update method according to some embodiments of the present disclosure;

[0030] Figure 2It is a flowchart of some embodiments of the information update method according to the present disclosure;

[0031] Figure 3 It is a schematic diagram of the vehicle kinematic model of the information update method according to the present disclosure;

[0032] Figure 4 It is a schematic diagram of generating a route feature information set of the information update method according to the present disclosure;

[0033] Figure 5 It is a flowchart of some other embodiments of the information update method according to the present disclosure;

[0034] Figure 6 It is a flowchart of some embodiments of the route screening method according to the present disclosure;

[0035] Figure 7 It is a schematic structural diagram of some embodiments of the information update device according to the present disclosure;

[0036] Figure 8 It is a schematic structural diagram of some embodiments of the route screening device according to the present disclosure;

[0037] Figure 9 It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners

[0038] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0039] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0040] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0041] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0042] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0043] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0044] Figure 1 It is a schematic diagram of an application scenario of an information update method according to some embodiments of the present disclosure.

[0045] In Figure 1In the application scenario, the electronic device 101 can first obtain at least one initial vehicle state information 103 for the target vehicle 102. In this application scenario, the above at least one initial vehicle state information 103 may include: initial vehicle state information 1031, initial vehicle state information 1032, and initial vehicle state information 1033. Among them, the vehicle state information of the initial vehicle state information 1031, the initial vehicle state information 1032, and the initial vehicle state information 1033 are all different. Then, the above electronic device 101 can generate at least one candidate vehicle to-be-traveled route 105 according to the above at least one initial vehicle state information 103 and the pre-selected candidate vehicle to-be-traveled strategy set 104. In this application scenario, the above candidate vehicle to-be-traveled strategy set 104 may include different candidate vehicle to-be-traveled strategies. The at least one candidate vehicle to-be-traveled route 105 also includes different candidate vehicle to-be-traveled routes. Among them, each candidate vehicle to-be-traveled route is determined according to a certain initial vehicle state and a certain candidate vehicle to-be-traveled strategy. As an example, the above candidate vehicle to-be-traveled strategy set 104 includes: candidate vehicle to-be-traveled strategy 1041, candidate vehicle to-be-traveled strategy 1042, and candidate vehicle to-be-traveled strategy 1043. The above at least one candidate vehicle to-be-traveled route 105 may include: candidate vehicle to-be-traveled route 1051, candidate vehicle to-be-traveled route 1052, candidate vehicle to-be-traveled route 1053, candidate vehicle to-be-traveled route 1054, candidate vehicle to-be-traveled route 1055, candidate vehicle to-be-traveled route 1056, candidate vehicle to-be-traveled route 1057, candidate vehicle to-be-traveled route 1058, and candidate vehicle to-be-traveled route 1059. Here, according to the initial vehicle state information 1031 and the candidate vehicle to-be-traveled strategy 1041, the candidate vehicle to-be-traveled route 1051 can be generated. According to the initial vehicle state information 1031 and the candidate vehicle to-be-traveled strategy 1042, the candidate vehicle to-be-traveled route 1052 can be generated. According to the initial vehicle state information 1031 and the candidate vehicle to-be-traveled strategy 1043, the candidate vehicle to-be-traveled route 1053 can be generated. According to the initial vehicle state information 1032 and the candidate vehicle to-be-traveled strategy 1041, the candidate vehicle to-be-traveled route 1054 can be generated. According to the initial vehicle state information 1032 and the candidate vehicle to-be-traveled strategy 1042, the candidate vehicle to-be-traveled route 1055 can be generated. According to the initial vehicle state information 1032 and the candidate vehicle to-be-traveled strategy 1043, the candidate vehicle to-be-traveled route 1056 can be generated. According to the initial vehicle state information 1033 and the candidate vehicle to-be-traveled strategy 1041, the candidate vehicle to-be-traveled route 1057 can be generated. According to the initial vehicle state information 1033 and the candidate vehicle to-be-traveled strategy 1042, the candidate vehicle to-be-traveled route 1058 can be generated. According to the initial vehicle state information 1033 and the candidate vehicle to-be-traveled strategy 1043, the candidate vehicle to-be-traveled route 1059 can be generated.Furthermore, the above-mentioned electronic device 101 may screen out candidate vehicle routes to be traveled that meet the preset conditions from the above-mentioned at least one candidate vehicle route to be traveled 105 as the first target candidate vehicle routes to be traveled, and obtain at least one first target candidate vehicle route to be traveled 106. Among them, the at least one first target candidate vehicle route to be traveled 106 may include: the first target candidate vehicle route to be traveled 1061 that is the same as the candidate vehicle route to be traveled 1051, the first target candidate vehicle route to be traveled 1062 that is the same as the candidate vehicle route to be traveled 1054, and the first target candidate vehicle route to be traveled 1063 that is the same as the candidate vehicle route to be traveled 1058. Finally, the above-mentioned electronic device 101 may update the parameter information in the target reward function 108 according to the above-mentioned at least one first target candidate vehicle route to be traveled 106 and at least one preset vehicle route to be traveled 107. Among them, the above-mentioned target reward function 108 is a function for determining the vehicle route. In this application scenario, the at least one preset vehicle route to be traveled 107 includes different preset vehicle routes to be traveled. The initial vehicle states corresponding to each preset vehicle route to be traveled are all different. That is, the at least one preset vehicle route to be traveled 107 includes: the preset vehicle route to be traveled 1071 corresponding to the initial vehicle state 1031, the preset vehicle route to be traveled 1072 corresponding to the initial vehicle state 1032, and the preset vehicle route to be traveled 1073 corresponding to the initial vehicle state 1033.

[0046] It should be noted that the above-mentioned electronic device 101 may be hardware or software. When the electronic device is hardware, it may be implemented as a distributed cluster composed of multiple servers or terminal devices, or may be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it may be installed in the above-mentioned listed hardware devices. It may be implemented as, for example, multiple software or software modules for providing distributed services, or may be implemented as a single software or software module. No specific limitation is made here.

[0047] It should be understood that Figure 1 the number of electronic devices in

[0048] Continue to refer to Figure 2 , which shows the flow 200 of some embodiments of the information update method according to the present disclosure. The information update method includes the following steps:

[0049] Step 201, obtain at least one initial vehicle state information for the target vehicle.

[0050] In some embodiments, the execution subject of the above-mentioned information update method (such as Figure 1The electronic device shown can obtain at least one initial vehicle state information for the target vehicle through a wired connection method or a wireless connection method. Among them, the above-mentioned target vehicle can be an autonomous vehicle. The initial vehicle state information in the above-mentioned at least one initial vehicle state information can be the initial position information and / or the initial pose information of the target vehicle.

[0051] It should be noted that the initial position information and / or the initial pose information of the above-mentioned target vehicle can be represented in various forms. For example, for the coordinate system set for the target vehicle, the initial position information of the above-mentioned target vehicle can be the initial position coordinates of the target vehicle. The above-mentioned initial pose information can be the angle information between the straight line between the target vehicle and the origin and the horizontal axis of the above-mentioned coordinate system.

[0052] As an example, the initial vehicle state information of the above-mentioned target vehicle can include: [Initial position information: (20, 30)], [Initial pose information: 30 degrees].

[0053] Step 202, generate at least one candidate vehicle to-be-traveled route according to the above-mentioned at least one initial vehicle state information and the pre-selected candidate vehicle to-be-traveled strategy set.

[0054] In some embodiments, the above-mentioned execution subject can generate at least one candidate vehicle to-be-traveled route according to the above-mentioned at least one initial vehicle state information and the pre-selected candidate vehicle to-be-traveled strategy set. Among them, the candidate vehicle to-be-traveled strategy can be the driving method of the target vehicle, that is, the driving strategy of the target vehicle. As an example, generally speaking, the driving strategy of the above-mentioned target vehicle can refer to which direction the target vehicle will go, at what speed it will go, and at what angle it will go. Examples are not given here. The driving strategy of the above-mentioned target vehicle can include various to-be-traveled indicators of the target vehicle. Each candidate vehicle to-be-traveled strategy in the above-mentioned candidate vehicle to-be-traveled strategy set is different. The candidate vehicle to-be-traveled route can be the predicted vehicle driving route of the target vehicle.

[0055] As an example, for each initial vehicle state information in the at least one initial vehicle state information, the above-mentioned execution subject can first use the above-mentioned initial vehicle state as the initial state of the candidate vehicle to-be-traveled strategy set to generate a driving route, so as to generate a candidate vehicle to-be-traveled route set and obtain a candidate vehicle to-be-traveled route set group. Then, the above-mentioned execution subject can determine the candidate vehicle to-be-traveled route set group as at least one candidate vehicle to-be-traveled route.

[0056] In some optional implementation manners of some embodiments, the above-mentioned candidate vehicle to-be-traveled strategy set is selected through the following steps:

[0057] The above-mentioned execution entity can sample the target vehicle's to-be-traveled methods in the target vehicle's to-be-traveled method set to obtain at least one target vehicle's to-be-traveled method as a candidate vehicle's to-be-traveled strategy set. Among them, the target vehicle's to-be-traveled method set can be pre-screened.

[0058] As an example, the above-mentioned execution entity can randomly sample the target vehicle's to-be-traveled methods in the target vehicle's to-be-traveled method set to obtain at least one target vehicle's to-be-traveled method as a candidate vehicle's to-be-traveled strategy set.

[0059] Optionally, the above-mentioned target candidate vehicle's to-be-traveled strategy set is selected through the following steps:

[0060] In the first step, the above-mentioned execution entity can obtain the vehicle control information sets in various directions for the above-mentioned target vehicle. Among them, the vehicle control information in the above-mentioned vehicle control information set is information containing vehicle feedback information. As an example, the above-mentioned vehicle control information set can include: lateral vehicle control information, longitudinal vehicle control information. Among them, the longitudinal vehicle control information can be determined by the following formula:

[0061]

[0062] Among them, K Lon can be the feedback gain of the longitudinal vehicle control information, that is, a part of the vehicle feedback information. Δd Lon can be the relative longitudinal distance between the host vehicle and the leading vehicle. Δv Lon can be the relative longitudinal speed between the host vehicle and the leading vehicle. Δa Lon is the relative longitudinal acceleration between the host vehicle and the leading vehicle. a Lon is the acceleration of the autonomous vehicle.

[0063] The lateral vehicle control information can be determined by the following formula:

[0064]

[0065] Among them, K Lat is the feedback gain of the lateral vehicle control information, that is, a part of the vehicle feedback information. v Lat can be the lateral speed of the host vehicle. a Lat can be the lateral acceleration of the host vehicle.

[0066] In the second step, the above-mentioned execution entity can assign different values to the vehicle feedback information corresponding to each vehicle control information in the above-mentioned vehicle control information set to obtain a vehicle's to-be-traveled method set. Among them, the above-mentioned vehicle feedback information is information in the form of a matrix with parameters.

[0067] As an example, the above-mentioned execution entity can randomly assign values within a certain range to each parameter in the vehicle feedback information corresponding to each vehicle control information in the above-mentioned vehicle control information set, so as to obtain a set of vehicle driving methods to be executed.

[0068] In the third step, according to the environmental information, the above-mentioned execution entity filters out the vehicle driving methods to be executed that match the environmental information from the above-mentioned set of vehicle driving methods to be executed, and uses them as the target candidate vehicle driving strategies, so as to obtain the above-mentioned set of target candidate vehicle driving strategies.

[0069] Among them, the above-mentioned environmental information can be the environmental model where the above-mentioned target vehicle is located.

[0070] The above-mentioned environmental model can adopt a hybrid modeling method. For the surrounding vehicles, a data set model can be selected, that is, the behavior of the surrounding vehicles travels according to the motion trajectories recorded in the data set. For the host vehicle, a vehicle kinematic model as shown in Figure 3 can be used. The bicycle kinematic model can be represented by the following continuous nonlinear state equations:

[0071] y = v sin(β + ψ)

[0072] x = v cos(β + ψ),

[0073]

[0074] v cosβ = v f cosδ f ,

[0075] Among them, (x, y) are the coordinates of the vehicle geometric center in the inertial coordinate system. ψ is the heading angle, v is the vehicle speed, l f and l r are the distances between the centroid and the front and rear axles, β is the velocity direction, a is the centroid acceleration. The control inputs are the front wheel angle δ f and the front wheel acceleration v f .

[0076] Here, it can be assumed that the centroid of the vehicle coincides with the geometric center, that is and it is assumed that the initial angle of each trajectory is zero.

[0077] In some optional implementation manners of some embodiments, the above-mentioned generating at least one candidate vehicle driving route according to the above-mentioned at least one initial vehicle state information and the pre-selected set of candidate vehicle driving strategies may include the following steps:

[0078] First step, the above-mentioned execution entity can expand the above-mentioned candidate vehicle to-be-traveled strategy set according to a pre-set action space to obtain an expanded candidate vehicle to-be-traveled strategy set. Among them, the above-mentioned action space can be the action space involved in reinforcement learning. The above-mentioned action space includes the combined information of each to-be-traveled index of the above-mentioned target vehicle. Each of the above-mentioned to-be-traveled indexes can include, but is not limited to, at least one of the following: acceleration index, speed index, relative longitudinal distance index between the host vehicle and the preceding vehicle.

[0079] As an example, the above-mentioned execution entity can randomly select actions in the action space to expand the above-mentioned candidate vehicle to-be-traveled strategy set to obtain an expanded candidate vehicle to-be-traveled strategy set.

[0080] As another example, for the candidate vehicle to-be-traveled strategy of traveling right frontward, with a lateral traveling speed of 1 m / s and a longitudinal traveling speed of 2 m / s. According to the action space, a candidate vehicle to-be-traveled strategy of traveling left frontward, with a lateral traveling speed of 1 m / s and a longitudinal traveling speed of 2 m / s can be generated.

[0081] Second step, the above-mentioned execution entity can generate the above-mentioned at least one candidate vehicle to-be-traveled route according to the above-mentioned expanded candidate vehicle to-be-traveled strategy set and the above-mentioned at least one initial vehicle state information.

[0082] As an example, for each initial vehicle state information in the at least one initial vehicle state information, the above-mentioned execution entity can first use the above-mentioned initial vehicle state as the initial state of the expanded candidate vehicle to-be-traveled strategy set to generate a to-be-traveled route, so as to generate a candidate vehicle to-be-traveled route set and obtain a candidate vehicle to-be-traveled route set group. Then, the above-mentioned execution entity can determine the candidate vehicle to-be-traveled route set group as the above-mentioned at least one candidate vehicle to-be-traveled route.

[0083] It should be noted that the purpose of expanding the candidate vehicle to-be-traveled strategy set is that: the candidate vehicle to-be-traveled strategy set is pre-selected, and there may be no candidate vehicle to-be-traveled strategy in the candidate vehicle to-be-traveled strategy set for generating the optimal candidate vehicle to-be-traveled route subsequently. Based on this, the action space is used to further enrich the candidate vehicle to-be-traveled strategy set to ensure that there is a corresponding optimal candidate vehicle to-be-traveled route relative to the preset vehicle to-be-traveled route among the at least one first target candidate vehicle to-be-traveled routes generated by the candidate.

[0084] Step 203, screen out the candidate vehicle to-be-traveled routes that meet the preset conditions from the above-mentioned at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled route to obtain at least one first target candidate vehicle to-be-traveled route.

[0085] In some embodiments, the above-mentioned execution entity may screen out a candidate vehicle to-be-traveled route that meets a preset condition from the above-mentioned at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled route, and obtain at least one first target candidate vehicle to-be-traveled route. Among them, the first target candidate vehicle to-be-traveled route may be the route that is closest to the preset vehicle to-be-traveled route for the initial vehicle state information, or may be the route corresponding to the optimal candidate vehicle to-be-traveled strategy in the candidate vehicle to-be-traveled strategy set.

[0086] As an example, for each initial vehicle state in at least one initial vehicle state information, there is a set of candidate vehicle to-be-traveled routes in at least one candidate vehicle to-be-traveled route with the above-mentioned initial vehicle state information as the initial vehicle state. Therefore, the candidate vehicle to-be-traveled route that is closest to the preset vehicle to-be-traveled route based on the same above-mentioned initial vehicle state is screened out from the set of candidate vehicle to-be-traveled routes as the first target candidate vehicle to-be-traveled route corresponding to the above-mentioned initial vehicle state.

[0087] As another example, according to the candidate vehicle to-be-traveled strategies corresponding to the candidate vehicle to-be-traveled routes, the above-mentioned execution entity may divide at least one candidate vehicle to-be-traveled route into a set of candidate vehicle to-be-traveled route groups. Then, the above-mentioned execution entity may select a set of candidate vehicle to-be-traveled routes from the set of candidate vehicle to-be-traveled route groups through various screening methods as at least one first target candidate vehicle to-be-traveled route.

[0088] Step 204, update the parameter information in the target reward function according to the above-mentioned at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route.

[0089] In some embodiments, the above-mentioned execution entity may update the parameter information in the target reward function according to the above-mentioned at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route. Among them, the above-mentioned target reward function is a function for determining a vehicle route. For the field of autonomous driving based on reinforcement learning or inverse reinforcement learning, the above-mentioned target reward function may be the reward function involved in reinforcement learning or inverse reinforcement learning.

[0090] Among them, the above-mentioned target reward function may be composed of parameter information and multiple vehicle driving characteristics. Among them, the vehicle driving characteristics may be the driving characteristics of the target vehicle relative to the surrounding vehicles during driving. As an example, the above-mentioned multiple vehicle driving characteristics include: the time headway with the vehicle in front, the time headway with the vehicle behind, the deviation from the desired speed, the lateral displacement, the longitudinal acceleration, the lateral acceleration, and the front wheel angular velocity.

[0091] Here, the time headway with the vehicle in front can be determined by the following formula:

[0092]

[0093] where y f is the longitudinal position of the leading vehicle, v Lon is the longitudinal speed of the host vehicle, is the desired time headway to the leading vehicle. y is the longitudinal position of the host vehicle.

[0094] The time headway to the following vehicle can be determined by the following formula:

[0095]

[0096] where y r is the longitudinal position of the following vehicle, is the desired time headway to the following vehicle.

[0097] It should be noted that the time headway to the leading vehicle and the time headway to the following vehicle are characteristics related to driving safety. Longitudinal acceleration, lateral acceleration, and front wheel angular velocity are characteristics related to ride comfort.

[0098] The deviation from the desired speed can be determined by the following formula:

[0099] f des = |v - v des |

[0100] where v des can be the desired vehicle speed of the host vehicle, which is related to the current traffic flow density. This characteristic is related to driving efficiency. v can be the vehicle speed of the host vehicle.

[0101] The lateral displacement can be determined by the following formula:

[0102] f dev = |x - x des |

[0103] where x can be the lateral position of the host vehicle. x des can represent the position of the center line of the target lane.

[0104] This formula represents the deviation between the lateral position of the host vehicle and the center line of the target lane, that is, the host vehicle should try to keep driving on the center line of the lane. The center line of the target lane is related to the current driving task. For example, when the driving task is straight ahead, the target lane is the current lane; when the driving task is lane change, the target lane is the adjacent lane. This characteristic is related to the driving task.

[0105] In some optional implementation manners of some embodiments, updating the parameter information in the target reward function according to the at least one first target candidate vehicle's to-be-traveled route and the at least one preset vehicle's to-be-traveled route may include the following steps:

[0106] In the first step, the above-mentioned execution entity can generate a first set of route feature information corresponding to the at least one first target candidate vehicle's to-be-traveled route and a second set of route feature information corresponding to the at least one preset vehicle's to-be-traveled route. Among them, the first set of route feature information corresponds one-to-one with the first target candidate vehicle's to-be-traveled route. The first route feature information can be the route feature information describing the first target candidate vehicle's to-be-traveled route. As an example, the first set of feature information may include, but is not limited to, at least one of the following: the acceleration information of the first target candidate vehicle's to-be-traveled route, the total road length information of the first target candidate vehicle's to-be-traveled route, and the curvature information of the first target candidate vehicle's to-be-traveled route. Similarly, the second set of route feature information corresponds one-to-one with the preset vehicle's to-be-traveled route. The second route feature information can be the route feature information describing the preset vehicle's to-be-traveled route. As an example, the second set of route feature information may include, but is not limited to, at least one of the following: the acceleration information of the preset vehicle's to-be-traveled route, the total road length information of the preset vehicle's to-be-traveled route, and the curvature information of the preset vehicle's to-be-traveled route.

[0107] As an example, for Figure 4 the scenario shown, the state information of the three white surrounding vehicles around the black ego vehicle is used to generate a set of route feature information.

[0108] It should be noted that the purpose of choosing this scenario is that in real life, when a driver decides to change lanes, they will balance the influence of the surrounding vehicles on the ego vehicle, especially the vehicle in front in the current lane and the vehicles in front and behind in the target lane at the gap for changing lanes. Here, the extracted scenario and data do not involve the selection of the lane-changing gap, that is, each trajectory has and only has one lane-changing gap, and the lane change is successfully completed.

[0109] Among them, the processing of trajectories based on machine vision technology is sometimes relatively rough. Especially, the speed and acceleration are usually affected by measurement errors during the processing. Since the differential technology is used to calculate the speed and acceleration, the noise will be gradually amplified, thus affecting the selection of the optimal trajectory and having an impact on the subsequent inverse reinforcement learning process.

[0110] Based on this, the above-mentioned execution entity can use the Kalman filtering technology, the adopted vehicle kinematic model, and the state and actions to filter the extracted trajectories. The trajectories after filtering become smooth, conform to the vehicle kinematic model, and are also not much different from the original trajectories. Then, according to the extracted trajectories, a set of route feature information can be generated through various vehicle trajectory feature extraction methods.

[0111] In the second step, according to the above-mentioned first set of route feature information and the second set of route feature information, the parameter information in the above-mentioned target reward function is updated.

[0112] As an example, the above-mentioned execution entity can first perform weighted summation on the corresponding route feature information in each first route feature information set in the first route feature information set group to obtain a first numerical set. Then, the above-mentioned execution entity can first perform weighted summation on the corresponding route feature information in each second route feature information set in the second route feature information set group to obtain a second numerical set. Next, according to the corresponding features, the above-mentioned first numerical set and the second numerical set are subtracted from each other correspondingly to obtain a third numerical set. Finally, according to the average value corresponding to the above-mentioned third numerical set, the gradient descent method is used to update the parameter information in the above-mentioned target return function.

[0113] Optionally, updating the parameter information in the above-mentioned target return function according to the above-mentioned first route feature information set group and the above-mentioned second route feature information set group may include the following steps:

[0114] In the first step, the above-mentioned execution entity can input each first route feature information set in the above-mentioned first route feature information set group into a first target function to generate a first numerical value, thereby obtaining a first numerical set.

[0115] In the second step, the above-mentioned execution entity can perform weighted average processing on each first numerical value in the above-mentioned first numerical set to obtain a first weighted numerical value.

[0116] In the third step, the above-mentioned execution entity can input each second route feature information set in the above-mentioned second route feature information set group into a second target function to generate a second numerical value, thereby obtaining a second numerical set.

[0117] In the fourth step, the above-mentioned execution entity can perform weighted average processing on each second numerical value in the above-mentioned second numerical set to obtain a second weighted numerical value.

[0118] In the fifth step, the above-mentioned execution entity can determine a first difference between the above-mentioned first weighted numerical value and the above-mentioned second weighted numerical value.

[0119] In the sixth step, in response to determining that the above-mentioned first difference is greater than a target threshold, the parameter information in the above-mentioned target return function is updated.

[0120] As an example, in response to determining that the above-mentioned first difference is greater than a target threshold, the above-mentioned execution entity can use the gradient descent method to update the parameter information in the above-mentioned target return function.

[0121] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The information update method of some embodiments of the present disclosure can update the parameter information in the target return function quickly and efficiently, reducing the waste of computing resources. Specifically, since inverse reinforcement learning highly depends on the reinforcement learning module for learning the optimal policy, the learning time is long. In addition, there may be a problem of result randomness when using reinforcement learning. Further, the amount of computation is increased, resulting in a waste of computer resources and indirectly increasing the processing load of the computer. Based on this, the information update method of some embodiments of the present disclosure can first obtain at least one initial vehicle state information for the target vehicle. Here, obtaining at least one initial vehicle state information is used to determine the candidate vehicle's to-be-traveled route subsequently. Then, according to the above at least one initial vehicle state information and the pre-selected candidate vehicle's to-be-traveled policy set, at least one candidate vehicle's to-be-traveled route is generated. Here, at least one initial vehicle state information can be used as the starting point, and the candidate vehicle's to-be-traveled policy set can be used as the driving policy of the target vehicle to accurately and efficiently generate at least one candidate vehicle's to-be-traveled route. Furthermore, the candidate vehicle's to-be-traveled routes that meet the preset conditions are screened out from the above at least one candidate vehicle's to-be-traveled route as the first target candidate vehicle's to-be-traveled route, obtaining at least one first target candidate vehicle's to-be-traveled route. Here, the purpose of screening the route from at least one candidate vehicle's to-be-traveled route is to screen out a more accurate and excellent route from several candidate vehicle's to-be-traveled routes corresponding to a certain initial vehicle state information. Based on this, the parameter information update of the subsequent target return function can be made more accurate and efficient. Finally, by comparing the differences between the above at least one first target candidate vehicle's to-be-traveled route and at least one preset vehicle's to-be-traveled route, the parameter information in the target return function can be updated quickly and efficiently. Among them, the above target return function is a function for determining the vehicle route.

[0122] Further reference Figure 5 shows a process 500 of other embodiments of the information update method according to the present disclosure. The information update method includes the following steps:

[0123] Step 501, obtain at least one initial vehicle state information for the target vehicle.

[0124] Step 502, generate at least one candidate vehicle's to-be-traveled route according to the above at least one initial vehicle state information and the pre-selected candidate vehicle's to-be-traveled policy set.

[0125] Step 503, determine the probability information corresponding to each candidate vehicle's to-be-traveled route in the above at least one candidate vehicle's to-be-traveled route according to the initial parameters of the above target return function, obtaining a probability information set.

[0126] In some embodiments, the execution entity (e.g., Figure 1 the electronic device shown) may determine the probability information corresponding to each candidate vehicle to-be-traveled route among the at least one candidate vehicle to-be-traveled route according to the initial parameters of the above-mentioned target return function, and obtain a probability information set.

[0127] As an example, the execution entity may determine the probability information corresponding to each candidate vehicle to-be-traveled route among the at least one candidate vehicle to-be-traveled route through the following formula according to the initial parameters of the above-mentioned target return function:

[0128]

[0129] where ξ may represent a candidate vehicle to-be-traveled route. ω may be the initial parameter of the target return function. f() may be a pre-set function.

[0130] Step 504: Screen out the candidate vehicle to-be-traveled routes whose probability information meets the target conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes, and obtain the at least one first target candidate vehicle to-be-traveled route.

[0131] In some embodiments, the execution entity may screen out the candidate vehicle to-be-traveled routes whose probability information meets the target conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes, and obtain the at least one first target candidate vehicle to-be-traveled route.

[0132] As an example, the execution entity may generate the at least one first target candidate vehicle to-be-traveled route through the following formula:

[0133]

[0134] where j * may be a candidate vehicle to-be-traveled strategy for the at least one first target candidate vehicle to-be-traveled route. may be the candidate vehicle to-be-traveled route of the j-th candidate vehicle to-be-traveled strategy for the i-th initial vehicle state information. may be the at least one first target candidate vehicle to-be-traveled route.

[0135] Step 505: Update the parameter information in the target return function according to the at least one first target candidate vehicle to-be-traveled route and the at least one preset vehicle to-be-traveled route.

[0136] In some embodiments, for the specific implementation of steps 501-502 and 505 and the technical effects brought by them, reference may be made to Figure 2Steps 201-202 and 204 in the corresponding embodiments will not be elaborated herein.

[0137] Step 506: Based on the parameter information updated according to the above target reward function, select candidate vehicle to-be-traveled routes that meet the above preset conditions from the above at least one candidate vehicle to-be-traveled route as the second target candidate vehicle to-be-traveled routes, obtaining at least one second target candidate vehicle to-be-traveled route.

[0138] In some embodiments, the above execution subject may, based on the parameter information updated according to the above target reward function, select candidate vehicle to-be-traveled routes that meet the above preset conditions from the above at least one candidate vehicle to-be-traveled route as the second target candidate vehicle to-be-traveled routes, obtaining at least one second target candidate vehicle to-be-traveled route. Here, the specific implementation steps may refer to Steps 503 and 504.

[0139] Step 507: Determine a second difference based on the above at least one second target candidate vehicle to-be-traveled route and the above at least one preset vehicle to-be-traveled route.

[0140] In some embodiments, the above execution subject may determine a second difference based on the above at least one second target candidate vehicle to-be-traveled route and the above at least one preset vehicle to-be-traveled route. Here, the specific implementation steps may refer to the implementation manner involved in Step 204.

[0141] Step 508: In response to determining that the above second difference is less than or equal to the target threshold, determine the parameter information updated by the above target reward function as the parameter information after training of the above target reward function.

[0142] In some embodiments, in response to determining that the above second difference is less than or equal to the target threshold, the above execution subject may determine the parameter information updated by the above target reward function as the parameter information after training of the above target reward function. Among them, the above target threshold may be preset for the subsequent convergence of the parameter information in the target reward function.

[0143] From Figure 5 it can be seen that compared with the description of some corresponding embodiments, Figure 2 the process 500 of the information update method in some corresponding embodiments Figure 5 more prominently highlights the specific steps of generating at least one first target candidate vehicle to-be-traveled route. Thus, the solutions described in these embodiments, without using reinforcement learning to generate at least one first target candidate vehicle to-be-traveled route, reduce the waste of computer resources and alleviate the computer load, greatly improving the efficiency of determining the parameter information of the target reward function.

[0144] Continue to refer toFigure 6 , which shows a flow 600 of some embodiments of a route screening method according to the present disclosure. The route screening method includes the following steps:

[0145] Step 601, obtaining at least one candidate vehicle travel route to be traveled.

[0146] In some embodiments, the execution subject (such as Figure 1 the electronic device shown) can obtain at least one candidate vehicle travel route to be traveled by a wired or wireless manner.

[0147] Step 602, screening out candidate vehicle travel routes to be traveled that meet preset conditions from the above at least one candidate vehicle travel route to be traveled as target candidate vehicle travel routes, obtaining at least one target candidate vehicle travel route.

[0148] In some embodiments, the above execution subject can screen out candidate vehicle travel routes to be traveled that meet preset conditions from the above at least one candidate vehicle travel route to be traveled as target candidate vehicle travel routes, obtaining at least one target candidate vehicle travel route. Among them, the parameter information in the above target return function is updated by using Figure 2 the methods of corresponding some embodiments.

[0149] The above various embodiments of the present disclosure have the following beneficial effects: The route screening method of some embodiments of the present disclosure can efficiently and accurately screen out candidate vehicle travel routes to be traveled that meet preset conditions from the above at least one candidate vehicle travel route to be traveled.

[0150] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an information update device. These device embodiments correspond to Figure 2 those method embodiments shown, and the device can be specifically applied to various electronic devices.

[0151] Such as Figure 7As shown in the figure, an information update device 700 includes: a first acquisition unit 701, a generation unit 702, a first screening unit 703, and an update unit 704. Among them, the first acquisition unit 701 is configured to acquire at least one initial vehicle state information for a target vehicle; the generation unit 702 is configured to generate at least one candidate vehicle to-be-traveled route according to the at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled strategy set; the first screening unit 703 is configured to screen out candidate vehicle to-be-traveled routes that meet preset conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled route, obtaining at least one first target candidate vehicle to-be-traveled route; the update unit 704 is configured to update the parameter information in the target reward function according to the at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route, where the target reward function is a function for determining a vehicle route.

[0152] In some alternative implementation manners of some embodiments, the generation unit 702 in the device 700 may be further configured to: expand the candidate vehicle to-be-traveled strategy set according to a preset action space, obtaining an expanded candidate vehicle to-be-traveled strategy set; generate the at least one candidate vehicle to-be-traveled route according to the expanded candidate vehicle to-be-traveled strategy set and the at least one initial vehicle state information.

[0153] In some alternative implementation manners of some embodiments, the first screening unit 703 in the device 700 may be further configured to: determine the probability information corresponding to each candidate vehicle to-be-traveled route in the at least one candidate vehicle to-be-traveled route according to the initial parameters of the target reward function, obtaining a probability information set; screen out candidate vehicle to-be-traveled routes whose probability information meets the target conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled route, obtaining the at least one first target candidate vehicle to-be-traveled route.

[0154] In some alternative implementation manners of some embodiments, the update unit 704 in the device 700 may be further configured to: generate a first route feature information set group corresponding to the at least one first target candidate vehicle to-be-traveled route and a second route feature information set group corresponding to the at least one preset vehicle to-be-traveled route; update the parameter information in the target reward function according to the first route feature information set group and the second route feature information set group.

[0155] In some alternative implementations of some embodiments, the above-mentioned apparatus 700 further includes: a third screening unit, a first determination unit, and a second determination unit (not shown in the figure). Among them, the above-mentioned third screening unit may be configured to: according to the parameter information after updating the above-mentioned target reward function, screen out the candidate vehicle to-be-traveled routes that meet the above-mentioned preset conditions from the above-mentioned at least one candidate vehicle to-be-traveled route as the second target candidate vehicle to-be-traveled routes, and obtain at least one second target candidate vehicle to-be-traveled route. The first determination unit may be configured to: determine a second difference according to the above-mentioned at least one second target candidate vehicle to-be-traveled route and the above-mentioned at least one preset vehicle to-be-traveled route. The second determination unit may be configured to: in response to determining that the above-mentioned second difference is less than or equal to the target threshold, determine the parameter information after updating the above-mentioned target reward function as the parameter information after training the above-mentioned target reward function.

[0156] In some alternative implementations of some embodiments, the update unit 704 in the above-mentioned apparatus 700 may be further configured to: input each first route feature information set in the above-mentioned first route feature information set group into a first target function to generate a first numerical value, and obtain a first numerical value set; perform a weighted average process on each first numerical value in the above-mentioned first numerical value set to obtain a first weighted numerical value; input each second route feature information set in the above-mentioned second route feature information set group into a second target function to generate a second numerical value, and obtain a second numerical value set; perform a weighted average process on each second numerical value in the above-mentioned second numerical value set to obtain a second weighted numerical value; determine a first difference between the above-mentioned first weighted numerical value and the above-mentioned second weighted numerical value; in response to determining that the above-mentioned first difference is greater than the target threshold, update the parameter information in the above-mentioned target reward function.

[0157] In some alternative implementations of some embodiments, the above-mentioned candidate vehicle to-be-traveled strategy set is selected through the following steps: sampling the target vehicle to-be-traveled methods in the target vehicle to-be-traveled method set to obtain at least one target vehicle to-be-traveled method as the candidate vehicle to-be-traveled strategy set.

[0158] In some alternative implementations of some embodiments, the above-mentioned target candidate vehicle to-be-traveled strategy set is selected through the following steps: obtaining vehicle control information sets in each direction for the above-mentioned target vehicle, where the vehicle control information in the above-mentioned vehicle control information set is information containing vehicle feedback information; assigning different numerical values to the vehicle feedback information corresponding to each vehicle control information in the above-mentioned vehicle control information set to obtain a vehicle to-be-traveled method set; screening out the vehicle to-be-traveled methods that match the environmental information from the above-mentioned vehicle to-be-traveled method set as the target candidate vehicle to-be-traveled strategies, and obtaining the above-mentioned target candidate vehicle to-be-traveled strategy set.

[0159] It can be understood that the various units described in the device 700 correspond to the respective steps in the method described with reference to Figure 2 Accordingly, the operations, features, and beneficial effects described above for the method also apply to the device 700 and the units included therein, and will not be elaborated herein.

[0160] With further reference to Figure 8 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a route screening device. These device embodiments correspond to Figure 6 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0161] As shown in Figure 8 , a route screening device 800 includes: a second acquisition unit 801 and a second screening unit 802. Among them, the second acquisition unit 801 is configured to: acquire at least one candidate vehicle travel route; the second screening unit 802 is configured to: according to a target reward function, screen out candidate vehicle travel routes that meet preset conditions from the above at least one candidate vehicle travel route as target candidate vehicle travel routes, to obtain at least one target candidate vehicle travel route, where the parameter information in the above target reward function is updated by using the method of Figure 2 corresponding some embodiments.

[0162] It can be understood that the various units described in the device 800 correspond to the respective steps in the method described with reference to Figure 6 Accordingly, the operations, features, and beneficial effects described above for the method also apply to the device 800 and the units included therein, and will not be elaborated herein.

[0163] Next, with reference to Figure 9 , which shows a schematic structural diagram of an electronic device (such as the electronic device in Figure 1 ) 900 suitable for implementing some embodiments of the present disclosure. Figure 9 The electronic device shown is only an example and should not impose any limitation on the functions and usage scopes of the embodiments of the present disclosure.

[0164] As shown in Figure 9As shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0165] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device 900 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively. Figure 9 Each block shown in the figure may represent a device or, as needed, multiple devices.

[0166] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the methods of some embodiments of the present disclosure are executed.

[0167] It should be noted that in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0168] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0169] The above computer-readable medium may be included in the above electronic device; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain at least one initial vehicle state information for a target vehicle; generate at least one candidate vehicle to-be-traveled route according to the above at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled strategy set; screen out candidate vehicle to-be-traveled routes that meet preset conditions from the above at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled route, to obtain at least one first target candidate vehicle to-be-traveled route; update parameter information in a target reward function according to the above at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route, wherein the above target reward function is a function for determining a vehicle route. Obtain at least one candidate vehicle to-be-traveled route;

[0170] According to the target reward function, screen out candidate vehicle to-be-traveled routes that meet preset conditions from the above at least one candidate vehicle to-be-traveled route as the target candidate vehicle to-be-traveled route, to obtain at least one target candidate vehicle to-be-traveled route, wherein the parameter information in the above target reward function is updated by using Figure 2 corresponding methods of some embodiments.

[0171] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0173] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition unit, a generation unit, a first screening unit, and an update unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least one initial vehicle state information for a target vehicle".

[0174] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0175] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An information update method, comprising: Obtain at least one initial vehicle state information for the target vehicle; Generate at least one candidate vehicle to-be-traveled route according to the at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled strategy set, wherein the candidate vehicle to-be-traveled strategy set is selected through the following steps: sampling the target vehicle to-be-traveled methods in the target vehicle to-be-traveled method set to obtain at least one target vehicle to-be-traveled method as the candidate vehicle to-be-traveled strategy set; Screen out the candidate vehicle to-be-traveled routes that meet the preset conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes, and obtain at least one first target candidate vehicle to-be-traveled route; Update the parameter information in the target reward function according to the at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route, wherein the target reward function is a function for determining the vehicle route.

2. The method according to claim 1, wherein, The generating at least one candidate vehicle to-be-traveled route according to the at least one initial vehicle state information and a pre-selected candidate vehicle to-be-traveled strategy set includes: Expand the candidate vehicle to-be-traveled strategy set according to a preset action space to obtain an expanded candidate vehicle to-be-traveled strategy set, wherein the action space includes the combined information of each to-be-traveled index of the target vehicle; Generate the at least one candidate vehicle to-be-traveled route according to the expanded candidate vehicle to-be-traveled strategy set and the at least one initial vehicle state information.

3. The method according to claim 1, wherein, The screening out the candidate vehicle to-be-traveled routes that meet the preset conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes, and obtaining at least one first target candidate vehicle to-be-traveled route includes: Determine the probability information corresponding to each candidate vehicle to-be-traveled route in the at least one candidate vehicle to-be-traveled route according to the initial parameters of the target reward function to obtain a probability information set; Screen out the candidate vehicle to-be-traveled routes with probability information meeting the target conditions from the at least one candidate vehicle to-be-traveled route as the first target candidate vehicle to-be-traveled routes, and obtain the at least one first target candidate vehicle to-be-traveled route.

4. The method according to claim 1, wherein, The updating the parameter information in the target reward function according to the at least one first target candidate vehicle to-be-traveled route and at least one preset vehicle to-be-traveled route includes: Generate a first route feature information set group corresponding to the at least one first target candidate vehicle to-be-traveled route and a second route feature information set group corresponding to the at least one preset vehicle to-be-traveled route; Update the parameter information in the target reward function according to the first route feature information set group and the second route feature information set group.

5. The method according to claim 1, wherein, The method further includes: Screen out the candidate vehicle to-be-traveled routes that meet the preset conditions from the at least one candidate vehicle to-be-traveled route according to the parameter information after the update of the target reward function as the second target candidate vehicle to-be-traveled routes, and obtain at least one second target candidate vehicle to-be-traveled route; Determine a second difference according to the at least one second target candidate vehicle's to-be-traveled route and the at least one preset vehicle's to-be-traveled route; In response to determining that the second difference is less than or equal to a target threshold, determine the parameter information after updating the target reward function as the parameter information after training the target reward function.

6. The method according to claim 4, wherein, The updating of the parameter information in the target reward function according to the first set of route feature information and the second set of route feature information includes: Input each first route feature information set in the first set of route feature information into a first objective function to generate a first value, obtaining a first value set; Perform a weighted average process on each first value in the first value set to obtain a first weighted value; Input each second route feature information set in the second set of route feature information into a second objective function to generate a second value, obtaining a second value set; Perform a weighted average process on each second value in the second value set to obtain a second weighted value; Determine a first difference between the first weighted value and the second weighted value; In response to determining that the first difference is greater than the target threshold, update the parameter information in the target reward function.

7. The method according to claim 1, wherein, The target candidate vehicle's to-be-traveled strategy set is selected through the following steps: Obtain a set of vehicle control information for each direction of the target vehicle, where the vehicle control information in the set of vehicle control information is information containing vehicle feedback information; Assign different values to the vehicle feedback information corresponding to each vehicle control information in the set of vehicle control information to obtain a set of vehicle's to-be-traveled methods; According to the environmental information, screen out the vehicle's to-be-traveled methods that match the environmental information from the set of vehicle's to-be-traveled methods as the target candidate vehicle's to-be-traveled strategies, obtaining the target candidate vehicle's to-be-traveled strategy set.

8. A route screening method, comprising: Obtain at least one candidate vehicle's to-be-traveled route; According to the target reward function, screen out the candidate vehicle's to-be-traveled routes that meet the preset conditions from the at least one candidate vehicle's to-be-traveled route as the target candidate vehicle's to-be-traveled routes, obtaining at least one target candidate vehicle's to-be-traveled route, where the parameter information in the target reward function is updated by the method described in any one of claims 1-7.

9. An information update device, comprising: A first acquisition unit, configured to acquire at least one initial vehicle state information of the target vehicle; A generation unit, configured to generate at least one candidate vehicle's to-be-traveled route according to the at least one initial vehicle state information and a pre-selected set of candidate vehicle's to-be-traveled strategies, where the set of candidate vehicle's to-be-traveled strategies is selected through the following steps: sampling the target vehicle's to-be-traveled methods in the set of target vehicle's to-be-traveled methods to obtain at least one target vehicle's to-be-traveled method as the set of candidate vehicle's to-be-traveled strategies; A first screening unit, configured to screen out the candidate vehicle's to-be-traveled routes that meet the preset conditions from the at least one candidate vehicle's to-be-traveled route as the first target candidate vehicle's to-be-traveled routes, obtaining at least one first target candidate vehicle's to-be-traveled route; An update unit, configured to update parameter information in a target reward function according to the at least one first target candidate vehicle's to-be-traveled route and at least one preset vehicle's to-be-traveled route, wherein the target reward function is a function for determining a vehicle route.

10. A route screening device, comprising: A second acquisition unit, configured to: acquire at least one candidate vehicle's to-be-traveled route; A second screening unit, configured to: screen out, according to the target reward function, candidate vehicle's to-be-traveled routes that meet preset conditions from the at least one candidate vehicle's to-be-traveled route as target candidate vehicle's to-be-traveled routes, so as to obtain at least one target candidate vehicle's to-be-traveled route, wherein the parameter information in the target reward function is updated by the method according to any one of claims 1-7.

11. An electronic device, comprising: One or more processors; A storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-8.

12. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, the method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Prediction method and system for vehicle travelling track

    CN105718750A

  • A device, electronic equipment and a medium for controlling driving of autonomous vehicle

    CN113033925A

  • Decision-making method and device for self-driving automobile

    CN113561986A