A path planning method and device, computer equipment and storage medium
By combining variable quantum algorithm and path planning algorithm, the problems of decision accuracy and computational resource consumption of autonomous vehicles in various environments are solved, and efficient path planning is achieved.
Patent Information
- Application Number
- CN202311189716.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2043-09-14
AI Technical Summary
In the path planning of autonomous vehicles, when faced with various driving environments, existing technologies struggle to reduce computational resource consumption while ensuring decision-making accuracy.
A variable quantum algorithm is used to process vehicle state and environmental information. Combined with trajectory tracking algorithm and artificial potential field obstacle avoidance algorithm, the path planning is optimized through quantum reinforcement learning. Quantum entanglement is used to reduce the number of parameters, and a path planning algorithm suitable for different scenarios is selected.
Improve decision-making accuracy, reduce computing resource consumption, and achieve efficient path planning in various driving environments.
Smart Images

Figure CN117232546B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of artificial intelligence, and particularly relates to a path planning method and device, computer equipment and a storage medium. BACKGROUND
[0002] At present, in the path planning module design process of an autonomous vehicle, an algorithm based on reinforcement learning is widely studied. Generally, the problems encountered by a driver in the actual driving process are complex and diverse, and different rules should be based on to make the next decision. Therefore, when training a reinforcement learning algorithm for realizing an end-to-end function, a large amount of samples are required, and the solution space of the original problem is extremely large, thereby increasing the training difficulty of the reinforcement learning algorithm and consuming a large amount of computing resources. The rule-based path planning method often divides the scene, and then learns each type of scene separately. Therefore, this type of method can only be applied in a single problem or a specific few types of problem scenes, and performs poorly in a driving environment with multiple problems.
[0003] How to ensure that when facing a driving environment with multiple problems, the accuracy of the decision is improved while reducing the consumption of computer resources is a problem that needs to be solved in the prior art. SUMMARY
[0004] To solve the problems in the prior art, the embodiments of the present specification provide a path planning method, device, computer equipment and storage medium, which realize the improvement of the accuracy of the decision while reducing the consumption of computer resources when facing a driving environment with multiple problems.
[0005] To solve the above technical problems, the specific technical solutions of the present specification are as follows:
[0006] On the one hand, the embodiments of the present specification provide a path planning method, comprising,
[0007] From the received path planning request, determine the vehicle state information and the vehicle surrounding environment information corresponding to the vehicle;
[0008] Using a variational quantum algorithm, process the vehicle state information and the vehicle surrounding environment information to obtain policy information;
[0009] Determine candidate action data corresponding to the policy information; and
[0010] Based on the path planning algorithm corresponding to the candidate action data, process the vehicle surrounding environment information and the vehicle position information included in the vehicle state information to obtain planning driving information.
[0011] Further, after obtaining the planning driving information, it comprises:
[0012] In a case where a total number of driving paths of the vehicle satisfies a preset condition, safety information corresponding to each time, vehicle state information, vehicle surrounding environment information, and planned driving information corresponding to each driving path are determined, each driving path corresponding to a plurality of times;
[0013] The safety information is weighted and discounted to obtain a reward function value corresponding to each driving path; and
[0014] Based on the reward function value, the variational quantum algorithm is parameter-optimized to obtain an optimized variational quantum algorithm for path planning at a next time.
[0015] Further, the determining candidate action data corresponding to the policy information further comprises,
[0016] Determining scene information represented by the policy information;
[0017] Determining a probability of each preset action being selected in a scene corresponding to the scene information; and
[0018] Based on the probability, determining a candidate action from the preset actions and determining the candidate action data corresponding to the candidate action.
[0019] Further, the path planning algorithm comprises a trajectory tracking algorithm and an artificial potential field obstacle avoidance algorithm, and the determining of the path planning algorithm corresponding to the candidate action data further comprises,
[0020] In a case where the candidate action data is first data, the trajectory tracking algorithm is determined as the path planning algorithm; and
[0021] In a case where the candidate action data is second data, the artificial potential field obstacle avoidance algorithm is determined as the path planning algorithm.
[0022] Further, the variational quantum algorithm further comprises:
[0023]
[0024]
[0025] wherein, the <O ω > and the <O ω′ > both represent Hermitian observables obtained by the variational quantum algorithm, the Hermitian observables being determined by the vehicle state information, the vehicle surrounding environment information, and parameters of the variational quantum algorithm, the ω and the ω' both represent the candidate action data, and the P m represents a quantum state a projection onto an eigenspace M of eigenvalue m, said π Ω characterizing the policy information, said s characterizing the vehicle state information and the vehicle surrounding environment information, said Ω characterizing policy parameters, and said β characterizing a temperature coefficient of a Boltzmann exploration method.
[0026] Further, the weighting and discounting processing on the safety information to obtain a reward function value corresponding to each driving path further includes,
[0027]
[0028]
[0029] wherein said G i,t characterizing the reward function value corresponding to the i-th driving path, said H characterizing a total number of time instants corresponding to the i-th driving path, said γ characterizing a discount factor, said t characterizing a starting time instant among the time instants, said c characterizing the weighted safety information, said Γ characterizing a weight coefficient, said j characterizing a category of the safety information, and said t' characterizing the time instants.
[0030] Further, the parameter optimization on the variational quantum algorithm based on the reward function value to obtain an optimized variational quantum algorithm for path planning at a next time instant further includes,
[0031]
[0032]
[0033] wherein said Ω characterizing the policy parameters, said N characterizing a total number of the driving paths in a current optimization process, said H characterizing a total number of time instants corresponding to the i-th driving path, said ω characterizing the candidate action data, said s characterizing the vehicle surrounding environment information and the vehicle state information, said G characterizing the reward function value, said Q characterizing an evaluation function value determined according to the vehicle state information and the vehicle surrounding environment information, said a characterizing the planned driving information, said π Ω characterizing the policy information, said <O ω > and said <O ω′ > each characterizing a Hermitian observable obtained by the variational quantum algorithm, said Hermitian observable being determined by the vehicle state information, the vehicle surrounding environment information, and parameters of the variational quantum algorithm, said ω and said ω' each characterizing the candidate action data, and said β characterizing a temperature coefficient of a Boltzmann exploration method.
[0034] In another aspect, the embodiments of the present specification also provide a path planning device, comprising,
[0035] The first determining unit is configured to determine vehicle state information and vehicle surrounding environment information corresponding to the vehicle from the received path planning request.
[0036] The first processing unit is configured to process the vehicle state information and the vehicle surrounding environment information by using a variational quantum algorithm to obtain policy information.
[0037] The second determining unit is configured to determine candidate action data corresponding to the policy information.
[0038] The second processing unit is configured to process the vehicle surrounding environment information and vehicle position information included in the vehicle state information by using a path planning algorithm corresponding to the candidate action data to obtain planned driving information.
[0039] Further, the device further comprises,
[0040] The obtaining unit is configured to determine safety information corresponding to each time, the vehicle state information, the vehicle surrounding environment information and the planned driving information in a case where a total number of driving paths of the vehicle satisfies a preset condition, each of the driving paths corresponding to a plurality of times.
[0041] The third processing unit is configured to perform weighting and discounting processing on the safety information to obtain a reward function value.
[0042] The optimization unit is configured to perform parameter optimization on the variational quantum algorithm based on the reward function value to obtain an optimized variational quantum algorithm, which is used for path planning at a next time.
[0043] In another aspect, the embodiments of the present specification also provide a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the above method when executing the computer program.
[0044] In another aspect, the embodiments of the present specification also provide a computer readable storage medium, which stores computer instructions, and the computer instructions are executed by a processor to implement the above method.
[0045] In another aspect, the embodiments of the present specification also provide a computer program product, which comprises computer program / instructions, and the computer program / instructions are executed by a processor to implement the above method.
[0046] With the embodiments of the present specification, when a path planning request is received, vehicle state information and vehicle surrounding environment information corresponding to a vehicle sending the path planning request are determined; then, the vehicle state information and the vehicle surrounding environment information are processed by using a variational quantum algorithm to determine strategy information; and the vehicle surrounding environment information and vehicle position information included in the vehicle state information are processed by using a path planning algorithm corresponding to the strategy information to obtain planning driving information. Thus, based on the characteristics of quantum entanglement reducing the amount of parameters of quantum reinforcement learning algorithm, the problems of large training difficulty and large consumption of computing resources of reinforcement learning algorithm are avoided; for each training scenario, a corresponding path planning algorithm can be selected to ensure the accuracy of decision-making in the driving environment facing various problems. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0048] Figure 1 An implementation system schematic diagram of a path planning method according to an embodiment of the present specification is shown;
[0049] Figure 2 A flowchart of a path planning method according to an embodiment of the present specification is shown;
[0050] Figure 3 A flowchart of a parameter optimization method according to an embodiment of the present specification is shown;
[0051] Figure 4A A flowchart of a candidate action data determination method according to an embodiment of the present specification is shown;
[0052] Figure 4B A flowchart of a path planning algorithm determination method according to an embodiment of the present specification is shown;
[0053] Figure 4C A schematic diagram of a vehicle state model according to an embodiment of the present specification is shown;
[0054] Figure 4D A schematic diagram of a trajectory tracking algorithm according to an embodiment of the present specification is shown;
[0055] Figure 5A A principle diagram of a path planning method according to another embodiment of the present specification is shown;
[0056] Figure 5B A principle diagram of a variational quantum algorithm according to an embodiment of the present specification is shown;
[0057] Figure 6 Fig. 1 shows a structural schematic diagram of a path planning device according to an embodiment of the present specification;
[0058] Figure 7 Fig. 2 shows a structural schematic diagram of a path planning device according to another embodiment of the present specification;
[0059] Figure 8 Fig. 3 shows a structural schematic diagram of a computer device according to an embodiment of the present specification.
[0060]
Explanation of reference signs
[0061] 101, user terminal;
[0062] 102, server;
[0063] 610, first determination unit;
[0064] 620, first processing unit;
[0065] 630, second determination unit;
[0066] 640, second processing unit;
[0067] 710, acquisition unit;
[0068] 720, third processing unit;
[0069] 730, optimization unit;
[0070] 802, computer device;
[0071] 804, processing device;
[0072] 806, storage resource;
[0073] 808, driving mechanism;
[0074] 810, input / output module;
[0075] 812, input device;
[0076] 814, output device;
[0077] 816, presentation device;
[0078] 818, graphical user interface;
[0079] 820, network interface;
[0080] 822, communication link;
[0081] 824, communication bus. DETAILED DESCRIPTION
[0082] The technical solutions in the embodiments of the present specification will be described clearly and completely in combination with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present specification.
[0083] It should be noted that the terms "first", "second", and the like in the description and claims of the present specification and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present specification described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product, or apparatus that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatuses.
[0084] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.
[0085] Figure 1 The embodiment system schematic diagram of the path planning method of the embodiments of the present specification can include a user terminal 101 and a server 102.
[0086] The user terminal 101 and the server 102 communicate through a network, which can include a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof, and is connected to a website, a user device (such as a computing device), and a backend system. The vehicle can send a path planning request to the server 102 in real time through the user terminal 101. After receiving the path planning request, the server 102 determines the vehicle state information and the vehicle surrounding environment information corresponding to the vehicle from the received path planning request (which can be obtained through multiple sensors of the vehicle itself, for example). The server 102 processes the vehicle state information and the vehicle surrounding environment information using a variational quantum algorithm to obtain policy information. The server 102 determines candidate action data corresponding to the policy information. The server 102 processes the vehicle surrounding environment information and the vehicle location information included in the vehicle state information based on a path planning algorithm corresponding to the candidate action data to obtain planning driving information, and sends the planning driving information to the user terminal 101 to provide a decision reference for the driver at the next moment, or to enable the vehicle to perform corresponding actions at the next moment. It should be noted that the user terminal 101 can be a sensor for obtaining vehicle state information and vehicle surrounding environment information, and a communication device. After obtaining the vehicle state information and the vehicle surrounding environment information, the sensor communicates with the server through the communication device, and thus sends the vehicle state information and the vehicle surrounding environment information to the server 102. After determining the planning driving information, the server 102 sends the planning driving information to the communication device, which can communicate with a control device for vehicle control to send the planning driving information to the control device, so as to provide a decision reference for the driver at the next moment, or to enable the vehicle to perform corresponding actions at the next moment.
[0087] After the server 102 processes the plurality of driving paths, and the total number of driving paths meets a preset condition, the server 102 can also determine safety information, vehicle surrounding environment information, and planning driving information corresponding to each moment, each driving path corresponding to a plurality of moments. The server 102 weights and discounts the safety information to obtain a reward function value corresponding to each driving path. The server 102 optimizes the variational quantum algorithm based on the reward function value to obtain an optimized variational quantum algorithm for path planning at the next moment.
[0088] Alternatively, the server 102 can be a node of a cloud computing system (not shown in the figure), or each server 102 can be a separate cloud computing system, including multiple computers interconnected by a network and working as a distributed processing system.
[0089] In addition, it should be noted that, Figure 1The shown is only one application environment provided by the specification, and in actual application, a plurality of user terminals 101 can also be included, which is not limited by the specification.
[0090] As Figure 2 The flowchart shown is a path planning method of an embodiment of the specification. In the figure, an online path planning process is described, but more or fewer operation steps can be included based on conventional or non-inventive labor. The order of steps listed in the embodiment is only one of the many execution orders, and does not represent the only execution order. In actual system or device product execution, the method order shown in the embodiment or the figure can be executed in sequence or in parallel. Specifically Figure 2 As shown, the method can include:
[0091] S210, determining vehicle state information and vehicle surrounding environment information corresponding to the vehicle from the received path planning request;
[0092] S220, processing the vehicle state information and the vehicle surrounding environment information by using a variational quantum algorithm to obtain strategy information;
[0093] S230, determining candidate action data corresponding to the strategy information;
[0094] S240, processing the vehicle surrounding environment information and the vehicle position information included in the vehicle state information by using a path planning algorithm corresponding to the candidate action data to obtain planning driving information.
[0095] By using the embodiment of the specification, when a path planning request is received, vehicle state information and vehicle surrounding environment information corresponding to the vehicle sending the path planning request are determined; then the vehicle state information and the vehicle surrounding environment information are processed by using a variational quantum algorithm to determine strategy information; the vehicle surrounding environment information and the vehicle position information included in the vehicle state information are processed by using a path planning algorithm corresponding to the strategy information to obtain planning driving information. Thus, based on the characteristic of quantum entanglement reducing the parameter amount of quantum reinforcement learning algorithm, the problems of large training difficulty and large calculation resource consumption of reinforcement learning algorithm are avoided; for each training scene, a corresponding path planning algorithm can be selected to ensure the accuracy of decision-making in the face of various problem driving environments.
[0096] According to one embodiment of the present specification, the vehicle state information includes any information that can be used to represent the state of the vehicle, for example, can include the speed, the yaw angle of the vehicle with the lane, and the distance of the vehicle with the center line of the lane, etc. The vehicle surrounding environment information includes any information that can be used to represent the objects around the vehicle and the relationship between the objects and the vehicle, for example, can include the object information and the relationship information of the objects with the clockwise included angle of the yaw direction of the vehicle mass center in the range of 7.5° to 45°, -15° to 15°, and -45° to -7.5°, and the relationship information can include the angle and the distance between the object and the vehicle mass center, for example. It should be noted that the vehicle surrounding environment information can be obtained based on the sensors arranged on the vehicle body, and the sensors can be ultrasonic ranging sensors, for example.
[0097] The variational quantum algorithm part can include quantum superposition, quantum coding, quantum entanglement, and quantum neural network, for example. Based on the above variational quantum algorithm part, the quantum coding and prediction processing are performed on the determined vehicle state information and vehicle surrounding environment information, and the strategy information corresponding to each vehicle state information and vehicle surrounding environment information is obtained. It should be noted that for each current time, there is vehicle state information and vehicle surrounding environment information corresponding to the current time, so that after processing, there is strategy information corresponding to each current time.
[0098] The preset candidate action data corresponding to each preset strategy information can be configured in advance. The preset strategy information can be used to determine the scene where the vehicle is located at the current time, and the preset candidate action data can be the category to which the scene belongs, for example. For example, all possible scenes are classified in advance to determine the category corresponding to each scene. In addition, an algorithm (preset path planning algorithm) for determining the planned driving information is also configured for each category. The number of categories can be configured according to actual conditions, which is not limited in the present specification. The preset path planning algorithm can include any algorithm that can perform driving prediction based on vehicle state information and vehicle surrounding environment information, for example, based on artificial potential field obstacle avoidance algorithm and trajectory tracking algorithm, etc.
[0099] After the strategy information is determined, the preset candidate action data corresponding to the strategy information is determined from the preset candidate action data configured in advance, and the preset candidate action data is used as the candidate action data; and then based on the candidate action data, the preset path planning algorithm corresponding to the candidate action data is determined, and the preset path planning algorithm is used as the path planning algorithm.
[0100] After determining the path planning algorithm corresponding to the vehicle state information and the vehicle surrounding environment information, the vehicle state information and the vehicle surrounding environment information are processed by using the path planning algorithm to obtain the planning driving information corresponding to the vehicle state information and the vehicle surrounding environment information, that is, the planning driving information corresponding to the current time is obtained to guide the driving of the vehicle at the next time. When determining the planning driving information, the target position of the path of the current vehicle can also be obtained, and the path planning algorithm processes the target position, the vehicle state information and the vehicle surrounding environment information to obtain the planning driving information. It should be noted that the vehicle position information can include, for example, coordinate information determined by the vehicle state information, and specifically, for example, can include the longitudinal coordinate of the vehicle center of mass, the lateral coordinate of the vehicle center of mass, the yaw angle of the vehicle and the motion rate of the vehicle. Based on the vehicle surrounding environment information, the obstacles around the vehicle and the angle and distance between the vehicle center of mass and the obstacles are determined. The planning driving information can include, for example, the turning angle of the vehicle steering wheel, and can also include, for example, the speed and / or acceleration of the vehicle.
[0101] Figure 3 A flowchart of a parameter optimization method according to an embodiment of the present specification is shown. In this figure, a parameter optimization process is described, but more or fewer operation steps can be included based on conventional or non-creative labor. Specifically, as shown in the figure Figure 3 As shown, the method can include:
[0102] S350, in a case where the total number of driving paths of the vehicle satisfies a preset condition, determining safety information, vehicle state information, vehicle surrounding environment information and planning driving information corresponding to each time, each driving path corresponding to a plurality of times;
[0103] S360, weighting and discounting the safety information to obtain a reward function value corresponding to each driving path;
[0104] S370, based on the reward function value, performing parameter optimization on the variational quantum algorithm to obtain an optimized variational quantum algorithm for path planning at the next time.
[0105] According to another embodiment of the present specification, in the process of determining the path planning algorithm corresponding to the vehicle state information and the vehicle surrounding environment information at a plurality of times, the vehicle state information and the vehicle surrounding environment information at a plurality of times are determined based on the safety information of the vehicle at a plurality of times, the vehicle state information of the vehicle at a plurality of times and the vehicle surrounding environment information of the vehicle at a plurality of times. Figure 2After the operation, parameter optimization is performed on the variational quantum algorithm for processing for path planning in the next round. Specifically, the condition for performing parameter optimization can be that path planning of a plurality of driving paths is performed, each driving path including vehicle state information and vehicle surrounding environment information corresponding to a plurality of time points, and the total number of time points is not limited in the specification. The preset condition includes a path number threshold. In the case where the total number of driving paths is greater than or equal to the path number threshold, it is determined that the total number of driving paths of the vehicle meets the preset condition, otherwise, it is determined that the preset condition is not met. It should be noted that the vehicle can be the same vehicle, or can be different vehicles, and the path number threshold can be configured according to actual conditions, which is not limited in the specification.
[0106] The safety information corresponding to each time point is the safety index of the vehicle after the determined planning driving information is executed at the next time point, for example, including a pedestrian safety index, a vehicle turning angle index, a lane edge index, and a standard driving index. Specifically, the pedestrian safety index represents the time length of the collision between the vehicle and the pedestrian when the distance between the pedestrian and the vehicle center is less than a first fixed value; the vehicle turning angle index represents the time length when the angle between the vehicle and the lane center line is greater than a second fixed value; the lane edge index represents the time length when the vehicle crosses the lane edge; and the standard driving index represents the time length when the distance between the vehicle center and the lane center line is less than a third fixed value. It should be noted that the above first fixed value, second fixed value and third fixed value can be configured according to actual conditions. Thus, the safety information corresponding to each time point, the vehicle state information, the vehicle surrounding environment information and the planning driving information are determined.
[0107] For each driving path, the safety information corresponding to each time point included in the driving path is processed to obtain a reward function value corresponding to the driving path. A weight corresponding to each safety information is configured, and the weight can be fixed or not fixed. In the case of not fixed, a weight corresponding to each safety information corresponding to each time point is configured. Further, using the configured weight, the safety information corresponding to each time point is processed to obtain a sub-reward function value corresponding to each safety information. Further, the sub-reward function value corresponding to each safety information included in the driving path is processed by discounting and adding to obtain the reward function value corresponding to the driving path.
[0108] The reward function values corresponding to a plurality of driving paths are added to obtain a total reward function value. Using the gradient descent method, the reward function value is processed to optimize the parameters in the variational quantum algorithm, to obtain optimized parameters to form an optimized variational quantum algorithm for path planning at the next time point. The specific optimization process can be, for example, an iterative optimization process.
[0109] According to another embodiment of the present specification, the weighting and discounting processing for the safety information to obtain the reward function value corresponding to each driving path comprises determining based on the following formula (1).
[0110]
[0111]
[0112] wherein G i,t characterizes the reward function value corresponding to the i-th driving path, H characterizes the total number of time points corresponding to the i-th driving path, γ characterizes the discount factor, t characterizes the starting time point among the time points, c characterizes the weighted safety information, Γ characterizes the weight coefficient, j characterizes the category of the safety information, and t' characterizes the time point. It should be noted that the formula (1) characterizes the formula used to determine the reward function value when the corresponding weight and discount factor are configured for the safety information corresponding to each time point.
[0113] According to another embodiment of the present specification, based on the reward function value, the parameter optimization is performed on the variational quantum algorithm to obtain the optimized variational quantum algorithm for path planning in the next time point comprises optimizing based on the following formula (2) and formula (3).
[0114]
[0115]
[0116] wherein Ω characterizes the policy parameter, N characterizes the total number of driving paths in the current optimization process, H characterizes the total number of time points corresponding to the i-th driving path, ω characterizes the candidate action data, s characterizes the vehicle surrounding environment information and vehicle state information, G characterizes the reward function value, Q characterizes the evaluation function value, the evaluation function value is determined according to the vehicle state information and the vehicle surrounding environment information, a characterizes the planned driving information, π Ω characterizes the policy information, <O ω > and <O ω′The evaluation function value is a function value obtained by processing the planning driving information corresponding to each time based on the evaluation function with respect to the vehicle state information and the vehicle surrounding environment information. Specifically, the evaluation function is any function that can be used to evaluate the superiority of the planning driving information corresponding to each time with respect to the vehicle state information and the vehicle surrounding environment information, i.e., the superiority of the planning driving information corresponding to each time in the context represented by the vehicle state information and the vehicle surrounding environment information.
[0117] Figure 4A A flowchart of a candidate action data determination method according to an embodiment of the present specification is shown in FIG. 4. In the embodiment, the candidate action data determination method comprises the following steps. Figure 4A A candidate action data determination process is described in the embodiment, but more or fewer operation steps can be included based on conventional or non-inventive labor. Specifically, as shown in FIG. 4, the method can comprise the following steps. Figure 4A
[0118] S431, determining scene information represented by the policy information;
[0119] S432, determining the probability of each preset action being selected in the scene corresponding to the scene information;
[0120] S433, determining a candidate action from the preset actions based on the probability, and determining candidate action data corresponding to the candidate action.
[0121] According to another embodiment of the present specification, the policy information is information that can represent scene information, which is information representing the state of the vehicle and the state of the surrounding obstacles at the current time, which can be, for example, vehicles, double yellow lines, pedestrians, and objects. The scene information can be, for example, that a vehicle X meters ahead of the vehicle is braking.
[0122] Based on this strategy information, the current scene information is determined. Multiple preset actions are pre-configured, and corresponding candidate action data is configured for each preset action. Preset actions can include any action a driver can take while driving; this specification does not limit this, but specifically, they can include braking, acceleration, and steering. Furthermore, each preset action is categorized, and different preset candidate action data is configured for different categories. The specific categorization method and the number of categories can be configured according to actual conditions. For example, preset actions can be categorized into two types: straight-ahead and turning. Corresponding preset action data is configured for each category. Continuing the previous example, for instance, preset action data 0 is configured for the straight-ahead category, and preset action data 1 is configured for the turning category.
[0123] For each scenario, the probability of selecting a preset action is determined based on that scenario information. For example, if the scenario information is that a vehicle is braking X meters ahead, the probability of selecting the preset action "brake" is relatively high, perhaps as low as 0.8. This determines the probability corresponding to each preset action. The preset action with the highest probability is then selected as a candidate action. Finally, the preset action data corresponding to the candidate actions is used as candidate action data.
[0124] Figure 4B The diagram shown is a flowchart of a path planning algorithm determination method according to an embodiment of this specification. Figure 4B The document describes a path planning algorithm for determining a path, but based on conventional or non-creative labor, it may include more or fewer operational steps. Specifically, for example... Figure 4B As shown, the method may include:
[0125] S441, when the candidate action data is the first data, the trajectory tracking algorithm is determined to be the path planning algorithm;
[0126] S442, when the candidate action data is the second data, the artificial potential field obstacle avoidance algorithm is determined to be the path planning algorithm.
[0127] According to another embodiment of this specification, for each category, a planning algorithm is configured to determine the method to be used when planning such an action, and the planning algorithm is associated with preset action data. For example, this specification divides preset actions into a first category and a second category, and the configured planning algorithms include a trajectory tracking algorithm and an artificial potential field obstacle avoidance algorithm. For the first category of preset actions, a trajectory tracking algorithm is configured for path planning, and for the second category of preset actions, an artificial potential field obstacle avoidance algorithm is configured for path planning.
[0128] The first data represents a first type of preset action, and the second data represents a second type of preset action. For example, the first data can be 0, and the second data can be 1. When the candidate action data is 0, a path planning algorithm used is a trajectory tracking algorithm, and when the candidate action data is 1, a path planning algorithm used is an artificial potential field obstacle avoidance algorithm.
[0129] Specifically, based on the path planning algorithm corresponding to the candidate action data, the process of processing the vehicle surrounding environment information and the vehicle position information included in the vehicle state information can be that, first, a vehicle state model is constructed based on the vehicle position information included in the vehicle state information, and then the corresponding path planning algorithm is used to process the vehicle state model and the vehicle surrounding environment information to obtain corresponding planning driving information.
[0130] The vehicle state model is shown in the following formula (4) and Figure 4C .
[0131]
[0132] wherein x represents the horizontal coordinate of the vehicle center of mass, y represents the vertical coordinate of the vehicle center of mass, l represents the distance from the tire to the vehicle center of mass, l f and l r respectively represent the distance from the front wheel and the rear wheel to the vehicle center of mass, represents the current yaw angle of the vehicle, t represents the time, v represents the current speed of the vehicle, a represents the acceleration, β represents the direction of the current speed of the vehicle, δf represents the turning angle of the front wheel of the vehicle, and δr represents the turning angle of the rear wheel of the vehicle, which can be 0 by default, for example.
[0133] The trajectory tracking algorithm is specifically to calculate the lateral tracking error e y between the vehicle and the reference trajectory (i.e., the center line of the lane determined by the automatic driving system as having no obstacles), i.e., the closest distance between the center point of the rear axle of the vehicle and the reference trajectory. The trajectory tracking algorithm is shown in the following formula (5) and Figure 4D .
[0134] e y = l d sinθ e Formula (5)
[0135] wherein θ e represents the angle between the vehicle center of mass and the target position (p x , p y ), and l d represents the distance between the vehicle center of mass and the target position.
[0136] According to the PID controller, the vehicle state model and the lateral tracking error are processed to obtain a vehicle steering wheel angle δf (i.e., the angle of the front wheel of the vehicle), thereby obtaining the planned driving information. Specifically, the formula used for processing can be, for example, the following formula (6).
[0137]
[0138] wherein Kp, Ki and Kd are proportional gain, integral gain and derivative gain, respectively, and t represents the time. p , K i and K d are proportional gain, integral gain and derivative gain, respectively, and t represents the time.
[0139] The artificial potential field obstacle avoidance algorithm specifically includes: calculating a gravitational potential field of an obstacle; calculating a repulsive potential field of a target point; calculating a resultant force field; and determining a steering wheel angle of the vehicle based on the resultant force field, i.e., the planned driving information. The gravitational potential field is shown in the following formula (7) and formula (8).
[0140]
[0141]
[0142] wherein q represents the center of mass of the vehicle, q g represents a target position, U att (q) represents a gravitational potential field function, η represents a positive proportional coefficient of the gravitational potential field, and ρ represents a vector distance between the center of mass of the vehicle and the target position; F att (q) represents a gravitational potential field, represents a derivative of the gravitational potential field function.
[0143] The repulsive potential field is shown in the following formula (9) and formula (10).
[0144]
[0145]
[0146] wherein U req (q) represents a repulsive potential field function, k represents a positive proportional coefficient of the repulsive potential field, ρ0 represents an obstacle influence range, q0 represents an obstacle position coordinate, q represents the center of mass of the vehicle, q g represents a target position, F req (q) represents a repulsive potential field.
[0147] The resultant force field is a potential field obtained by adding the repulsive potential field and the gravitational potential field.
[0148] Figure 5A Fig. 1 shows a schematic diagram of a path planning method according to an embodiment of the present specification; Figure 5BA schematic diagram of a variational quantum model according to an embodiment of the present specification is shown.
[0149] According to another embodiment of the present specification, the path planning method includes a quantum computing model, a classical computing model and an acquired classical environment model. Specifically, the quantum computing model includes a variational quantum algorithm, a prediction algorithm and a path planning algorithm, Figure 5A The quantum computing unit in the quantum computing model only shows the variational quantum algorithm and the prediction algorithm, and the path planning algorithm is not shown. The quantum computing model acquires a state s from the classical environment model, the state s including vehicle state information corresponding to the vehicle and vehicle surrounding environment information, and then processes the state s by using the variational quantum algorithm, the prediction algorithm and the path planning algorithm to obtain an action a (planned driving information), and then sends the action a, the state s and an intermediate variable determined based on the state s by formula (3) to the classical computing model. The classical computing model acquires safety information from the classical environment model to determine a reward r, and then optimizes the parameters of the variational quantum algorithm in the quantum computing unit based on the reward r by using the policy gradient algorithm for planning in the next round. It should be noted that the quantum computing model is performed on a quantum computing unit, and the classical computing model is performed on a classical computing unit.
[0150] Specifically, Figure 5A The variational quantum algorithm and the prediction algorithm included in the quantum computing model are schematic diagrams, and the specific diagrams are shown in Figure 5B . Figure 5B A variational quantum model (i.e., a variational quantum circuit topology) according to an embodiment of the present specification is shown, which includes a variational quantum algorithm and a prediction algorithm. The variational quantum algorithm includes quantum superposition and a plurality of variational quantum circuits, each of which includes quantum encoding, quantum entanglement and a quantum neuron network. The prediction algorithm may, for example, be quantum measurement. The strategy information is determined based on the variational quantum algorithm.
[0151] According to another embodiment of the present specification, the variational quantum algorithm includes the following formula (11) and formula (12) to determine the strategy information.
[0152]
[0153]
[0154] wherein <O ω > and <O ω′ > each represent a Hermitian observable obtained by the variational quantum algorithm, the Hermitian observable being determined by the vehicle state information, the vehicle surrounding environment information and the parameters of the variational quantum algorithm, ω and ω' each represent candidate action data, P m represents the projection of the quantum state onto the eigenspace M with eigenvalue m, and πΩ The system represents strategy information, s represents vehicle state information and vehicle surrounding environment information, Ω represents strategy parameters, and β represents the temperature coefficient of the Boltzmann exploration method.
[0155] against Figure 5B , specifically, |0> i Let represent the state basis vector of the i-th qubit, which is converted from the ground state to a superposition state by an Adama gate H. The scaling parameter λ and the state input value s (represented in the figure) in the quantum encoding part are also shown. x (λ k,l s k,l () represents the parameters that together constitute the state input, and R represents the single-qubit Pauli rotation gate R of the d-th neural network layer. x The quantum entanglement is achieved through entangled CZ gates in adjacent circuits. The quantum neuron network part represents the single-qubit Pauli rotation gate R of the d-th state input network layer. x R y and R z Here, θ represents the rotation angle of the rotating gate, ranging from [0, 2π]. The variational quantum circuit is measured on a given quantum computing basis vector, and its expected value is estimated through repeated evaluations. <O ω >。 ω i This represents the mapping result of the quantum measurement, i.e., candidate action data.
[0156] Figure 6 The diagram shown is a structural schematic of a path planning device according to an embodiment of this specification. As shown in Figure 6, it includes:
[0157] The first determining unit 610 is used to determine the vehicle status information and the vehicle surrounding environment information corresponding to the vehicle from the received path planning request.
[0158] The first processing unit 620 is used to process vehicle state information and vehicle surrounding environment information using a variable quantum algorithm to obtain strategy information.
[0159] The second determining unit 630 is used to determine candidate action data corresponding to the strategy information; and
[0160] The second processing unit 640 is used to process the vehicle's surrounding environment information and vehicle status information, including vehicle position information, based on the path planning algorithm corresponding to the candidate action data, to obtain planned driving information.
[0161] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.
[0162] Figure 7As shown in another embodiment of the present specification is a schematic structural diagram of a path planning device. As shown in 7, comprising,
[0163] The acquisition unit 710 is configured to determine the safety information, the vehicle state information, the vehicle surrounding environment information and the planning driving information corresponding to each time point when the total number of driving paths of the vehicle meets the preset condition, each driving path corresponding to a plurality of time points;
[0164] The third processing unit 720 is configured to perform weighting and discounting processing on the safety information to obtain a reward function value; and
[0165] The optimization unit 730 is configured to perform parameter optimization on the variational quantum algorithm based on the reward function value to obtain an optimized variational quantum algorithm for path planning at the next time point.
[0166] Since the principle of solving the above-mentioned device is similar to the above-mentioned method, the implementation of the above-mentioned device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described.
[0167] As Figure 8 As shown in a schematic structural diagram of a computer device according to an embodiment of the present specification. The device in the present specification can be the computer device in the present embodiment, which executes the method of the present specification. The computer device 802 can include one or more processing devices 804, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 802 can also include any storage resource 806 for storing any kind of information such as code, settings, data, etc. Without limitation, for example, the storage resource 806 can include any one or a combination of the following: any type of RAM, any type of ROM, a flash memory device, a hard disk, an optical disk, etc. More generally, any storage resource can store information using any technology. Further, any storage resource can provide volatile or non-volatile retention of information. Further, any storage resource can represent a fixed or removable component of the computer device 702. In one case, the computer device 802 can perform any operation of the associated instructions when the processing device 804 executes the associated instructions stored in any storage resource or combination of storage resources. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any storage resource, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.
[0168] The computer device 802 can also include an input / output module 810 (I / O) for receiving input (via input device(s) 812) and for providing output (via output device(s) 814). One particular output mechanism can include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), the input device(s) 812, and the output device(s) 814 can not be included, and the computer device 802 can be used merely as a network node in a network. The computer device 802 can also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the above-described components so that each component can communicate with each other component.
[0169] The communication links 822 can be implemented in any manner, such as through local area networks, wide area networks (e.g., the Internet), point-to-point connections, etc., or any combination thereof. The communication links 822 can include any combination of hardwired links, wireless links, routers, gateway functionality, name servers, etc., governed by any protocol or combination of protocols.
[0170] The embodiments of the present specification also provide a computer-readable storage medium, the computer-readable storage medium stores a computer program, the computer program is executed by a processor to implement the above method.
[0171] The embodiments of the present specification also provide a computer program product, the computer program product includes a computer program, the computer program is executed by a processor to implement the above method.
[0172] Those skilled in the art will appreciate that the embodiments of the present specification can be provided as methods, systems, or computer program products. Therefore, the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer usable program code.
[0173] The present specification is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions, which are executed via the processor of the computer or other programmable data processing apparatus, generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one flow or multiple flows and / or blocksFigure 1 means for performing the function specified in the block or blocks.
[0174] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 flow or flows and / or blocks Figure 1 means for performing the function specified in the block or blocks.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 flow or flows and / or blocks Figure 1 Figure 1 steps for performing the function specified in the block or blocks.
[0176] The above specific embodiments, for the purpose of this description, technical solutions and beneficial effects have been further detailed, it should be understood that the above is only a specific embodiment of the present description, and is not used to limit the protection scope of the present description, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present description shall be included in the protection scope of the present description.
Claims
1. A path planning method, characterized in that, include: From the received route planning request, determine the vehicle status information and the vehicle's surrounding environment information corresponding to the vehicle; The variable quantum algorithm is used to process the vehicle state information and the vehicle surrounding environment information to obtain strategy information. Candidate action data corresponding to the strategy information is determined; based on the path planning algorithm corresponding to the candidate action data, the vehicle surrounding environment information and the vehicle position information included in the vehicle state information are processed to obtain the planned driving information; When the total number of vehicle travel paths meets the preset conditions, the safety information, vehicle status information, vehicle surrounding environment information and planned travel information corresponding to each time moment are determined, and each travel path corresponds to multiple time moments; The safety information is weighted and discounted to obtain a reward function value corresponding to each driving path; as well as Based on the reward function value, the parameters of the variable quantum algorithm are optimized to obtain an optimized variable quantum algorithm for path planning in the next time step. The variational quantum algorithm includes: Among them, the and stated All represent Hermitian observations obtained by the variable quantum algorithm, wherein the Hermitian observations are determined by the vehicle state information, the vehicle's surrounding environment information, and the parameters of the variable quantum algorithm. and stated All represent the candidate action data, the Characterizing quantum states to the eigenvalue The intrinsic space M The projection on, the Characterizing the strategy information, the The vehicle state information and the vehicle surrounding environment information characterize the vehicle state information and the vehicle surrounding environment information. The representation strategy parameters, the Temperature coefficient characterizing the Boltzmann exploration method.
2. The method according to claim 1, characterized in that, The determination of candidate action data corresponding to the strategy information includes: Determine the scenario information represented by the strategy information; Determine the probability that each preset action will be selected in the scene corresponding to the scene information; and Based on the probability, candidate actions are determined from the preset actions, and candidate action data corresponding to the candidate actions are determined.
3. The method according to claim 1, characterized in that, The path planning algorithm includes a trajectory tracking algorithm and an artificial potential field obstacle avoidance algorithm. The determination of the path planning algorithm corresponding to the candidate action data includes: When the candidate action data is the first data, the trajectory tracking algorithm is determined to be the path planning algorithm; and When the candidate action data is the second data, the artificial potential field obstacle avoidance algorithm is determined to be the path planning algorithm.
4. The method according to claim 1, characterized in that, The weighted and discounted processing of the safety information to obtain the reward function value corresponding to each driving path includes: Among them, the Characterization and the first The reward function value corresponding to each driving path, Characterization and the first The total number of time steps corresponding to each travel route, the Characterizing the discount factor, the Characterizing the start time in the stated time, the The weighted security information is represented by the following: Characterizing the weighting coefficients, the Characterizing the category of the security information, and the The time indicated.
5. The method according to claim 1, characterized in that, The step of optimizing the parameters of the variable quantum algorithm based on the reward function value to obtain an optimized variable quantum algorithm for path planning in the next time step includes: Among them, the Characterizing the policy parameters, the The total number of driving paths in the current optimization process is represented by the following. Characterization and the first The total number of time steps corresponding to each travel route, the The information representing the vehicle's surrounding environment and the vehicle's state information, the Characterizing the reward function value, the The evaluation function value is characterized by being determined based on the vehicle state information and the vehicle's surrounding environment information. Characterizing the planned driving information, the Characterizing the strategy information, the and stated All represent Hermitian observations obtained by the variable quantum algorithm, wherein the Hermitian observations are determined by the vehicle state information, the vehicle's surrounding environment information, and the parameters of the variable quantum algorithm. and stated All represent the candidate action data, the Temperature coefficient characterizing the Boltzmann exploration method.
6. A path planning device, characterized in that, include: The first determining unit is used to determine the vehicle status information and the vehicle's surrounding environment information from the received path planning request. The first processing unit is used to process the vehicle state information and the vehicle surrounding environment information using a variable quantum algorithm to obtain strategy information. The second determining unit is used to determine the candidate action data corresponding to the strategy information; The second processing unit is used to process the vehicle's surrounding environment information and vehicle position information, which are included in the vehicle's state information, based on the path planning algorithm corresponding to the candidate action data, to obtain planned driving information. The acquisition unit is used to determine the safety information, vehicle status information, vehicle surrounding environment information and planned driving information corresponding to each time point when the total number of driving paths of the vehicle meets the preset conditions, and each driving path corresponds to multiple time points; The third processing unit is used to perform weighted and discounted processing on the security information to obtain a reward function value; as well as The optimization unit is used to optimize the parameters of the variable quantum algorithm based on the reward function value, so as to obtain an optimized variable quantum algorithm for path planning in the next time step. The variational quantum algorithm includes: Among them, the and stated All represent Hermitian observations obtained by the variable quantum algorithm, wherein the Hermitian observations are determined by the vehicle state information, the vehicle's surrounding environment information, and the parameters of the variable quantum algorithm. and stated All represent the candidate action data, the Characterizing quantum states to the eigenvalue intrinsic space M The projection on, the Characterizing the strategy information, the The vehicle state information and the vehicle surrounding environment information characterize the vehicle state information and the vehicle surrounding environment information. The representation strategy parameters, the Temperature coefficient characterizing the Boltzmann exploration method.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the method of any one of claims 1-5.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Path planning method and device, computer equipment and storage medium
CN114764253A
Automatic driving control method and device, computer equipment and storage medium
CN114779776A