An Optimization Method and Device for a Path Planning Strategy
By converting the security constraints of path planning policy optimization into security level variables, and combining the optimization goals of performance gradients and security gradients, the problems of difficulty in computing multiple security constraints and low simulation efficiency in path planning policy optimization in the existing technology are solved, and more efficient and accurate strategy optimization is achieved.
Patent Information
- Application Number
- CN202410249921.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-03-05
AI Technical Summary
In the prior art, path planning strategy optimization has problems such as difficulty in accurately calculating multi-security constraints and low simulation efficiency, and it is not possible to ensure that the update of the policy is always within the scope of the security constraints, resulting in low optimization efficiency.
By establishing a value optimization model for path planning policy optimization, including performance functions and multiple security constraints, the security constraints are converted into security level variables that represent the number of security constraints that meet the number of security constraints, the initial feasible solution set is selected from large to small according to the security level, and optimization goals are established based on the performance functions, security constraints and security level variables, and the performance gradient and security gradient of the policy are used for optimization.
It improves the accuracy and efficiency of policy optimization, can better cope with uncertainty, ensures that the policy is always in security constraints, and improves the security and traffic efficiency of path planning.
Smart Images

Figure CN118070992B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of policy optimization, and particularly to an optimization method and device for a path planning policy. Background Art
[0002] For the path planning task in autonomous driving, it is crucial to ensure the safety and traffic efficiency of the policy. However, the application scenarios of autonomous driving, especially intersection scenarios, have the following characteristics: (1) The behaviors of traffic participants are highly uncertain, which makes it difficult to accurately describe the traffic participant model and there are few feasible policies at the same time. As a result, it is very difficult to measure the true value of safety based on the constraint conditions established by the traffic participant model. (2) The traffic flow situation on actual roads is complex, which leads to a large variety of safety constraint conditions, high simulation costs (such as low efficiency), and it is very difficult to ensure the safety of the optimization process.
[0003] In summary, the policy optimization for various safety constraints composed of multiple traffic participants (pedestrians, vehicles, road boundaries, etc.) faces challenges in difficult safety calculation, evaluation, and improvement.
[0004] Limited by the complexity of the traffic participant model and the uncertainty of the environment, it is difficult to accurately describe the system dynamics with differential equations. Therefore, it is necessary to evaluate the safety of the objective function through simulation. However, the performance evaluation of complex systems through simulation is usually very time-consuming. As the problem scale increases, the solution space often grows exponentially, and problems such as the curse of dimensionality will be faced when searching for the optimal solution through traversal in the solution space.
[0005] The existing path planning policy optimizations mainly include the following two types: One is the rule-based policy optimization. For example, the position line segment where the obstacle is located is inflated, and then path planning is performed according to the inflated obstacle. However, the real scenarios are ever-changing, and the rule-based method cannot effectively explore the uncertainty in the environment, resulting in a low accuracy of the path planning policy. The other is the learning-based policy planning method, such as reinforcement learning and adaptive learning. The existing learning-based methods handle a small number of safety constraints, and usually, the policy is optimized first, and then it is judged whether the policy meets the safety constraints. If not, it is corrected. Therefore, the existing learning-based methods have problems of low optimization efficiency and are not applicable to complex path planning scenarios with multiple safety constraints. In addition, rule-based and learning-based safety controls generally assume that the system model is accurate, but in actual applications, it is often difficult to obtain an accurate model. Therefore, it is impossible to ensure that the dynamic process conforms to the real safety constraints, and it is difficult to ensure that the optimization process and the final policy fully meet the safety constraints. Summary of the Invention
[0006] The present disclosure is used to solve the problems in the prior art that it is difficult to accurately calculate the safety constraints in the optimization of task strategies with multiple safety constraints, the simulation of safety constraints is inefficient, and the update of the strategy is not always guaranteed to remain within the safety constraints, resulting in low efficiency in strategy optimization.
[0007] To solve the above technical problems, on the one hand, the present disclosure provides an optimization method for a path planning strategy, including:
[0008] Establish a value optimization model for optimizing the path planning strategy, where the value optimization model includes: a performance function and multiple safety constraints;
[0009] Convert the safety constraints into safety level variables representing the number of satisfied safety constraints;
[0010] Sort the strategies in the path planning strategy space in descending order of safety level, and select a predetermined number of strategies according to the sorting result order to form an initial feasible solution set;
[0011] Establish an optimization objective based on the performance function, safety constraints, and safety level variables, where a constraint term for strategy switching when the safety level is improved is defined in the optimization objective;
[0012] Optimize the path planning strategy according to the optimization objective and the initial feasible solution set.
[0013] As a further embodiment of the present disclosure, optimizing the path planning strategy according to the optimization objective includes:
[0014] Optimize the optimization objective using the performance gradient and safety gradient of the strategy to obtain the path planning strategy.
[0015] As a further embodiment of the present disclosure, the performance function is used to measure the distance of the moving device from the target position and the speed of travel;
[0016] The safety constraints are used to ensure that no collision occurs between the obstacle and the moving device;
[0017] The safety level variable is expressed by the following formula:
[0018] S(x) = ∑S i (x);
[0019] If g i (x) ≤ 0, then S i (x) = 1;
[0020] where S(x) is the safety level variable; i is the safety constraint condition number; g i(x) ≤ 0 represents the case where the i-th safety constraint condition meets the safety constraint; S i (x) being 0 indicates non-compliance with the safety constraint, and being 1 indicates compliance with the safety constraint; x represents the decision variable for optimizing the path planning strategy.
[0021] In a further embodiment of the present disclosure, the determination process of the predetermined quantity includes:
[0022] Determine the predetermined quantity using the fine-optimization blind selection rule of the following formula:
[0023]
[0024] where N represents the quantity of strategies sampled from the strategy space; G represents the strategy set composed of the first g strategies determined from N strategies; g represents the number of strategies; O represents the initial feasible solution set, o represents the predetermined quantity; k is a positive integer less than o; P represents the probability;
[0025] where N, g, k, and P are configuration quantities.
[0026] In a further embodiment of the present disclosure, establishing an optimization objective based on the initial feasible solution set, performance function, safety constraint conditions, and safety level variable includes:
[0027] Establish a first optimization term according to the safety constraint conditions;
[0028] Convert the safety level variable into a continuous variable, and establish a second optimization term according to the continuous variable, where the second optimization term is used to constrain the switching of strategies when a nested event occurs, and the nested event is an event when the safety level is increased;
[0029] Establish a reward function according to the performance function, the first optimization term, and the second optimization term;
[0030] Establish an optimization objective according to the reward function.
[0031] In a further embodiment of the present disclosure, establishing a first optimization term according to the safety constraint conditions includes:
[0032] Represent the first constraint term using the following formula:
[0033]
[0034] where j is the safety constraint condition number; g j (x) is the j-th safety constraint condition; ω j is the weight of the j-th safety constraint condition; g j (x) > 0 represents the case where the j-th safety constraint condition does not meet the safety constraint; x represents the decision variable for optimizing the path planning strategy.
[0035] In a further embodiment of the present disclosure, establishing a second optimization term based on continuous variables includes: using a second constraint term represented by the following formula:
[0036]
[0037] where x t represents the decision variable for optimizing the path planning strategy at time t, and x t-1 represents the decision variable for optimizing the path planning strategy at time t-1, is the safety level at time t, represents the safety level at time t-1.
[0038] A second aspect of the present disclosure provides an optimization device for a path planning strategy, including:
[0039] A first modeling unit for establishing a value optimization model for optimizing the path planning strategy, where the value optimization model includes: a performance function and multiple safety constraint conditions;
[0040] A second modeling unit for converting the safety constraint conditions into safety level variables representing the number of safety constraint conditions satisfied;
[0041] A feasible solution search unit for sorting the strategies in the path planning strategy space in descending order of safety level, and screening out a predetermined number of strategies according to the sorting result order to form an initial feasible solution set;
[0042] An optimization objective establishment unit for establishing an optimization objective according to the performance function, safety constraint conditions and safety level variables, where a constraint term for switching strategies when the safety level is improved is defined in the optimization objective;
[0043] An optimization unit for optimizing the path planning strategy according to the optimization objective and the initial feasible solution set.
[0044] A third aspect of the present disclosure provides an optimization method for a target task strategy, including:
[0045] Establishing a value optimization model for optimizing the target task strategy, where the value optimization model includes: a performance function and multiple safety constraint conditions;
[0046] Converting the safety constraint conditions into safety level variables representing the number of safety constraint conditions satisfied;
[0047] Sorting the strategies in the target task strategy space in descending order of safety level, and screening out a predetermined number of strategies according to the sorting result order to form an initial feasible solution set;
[0048] An optimization objective is established based on a performance function, safety constraint conditions, and a safety level variable. Among them, a constraint term for policy switching when the safety level is improved is defined in the optimization objective.
[0049] Optimize the task policy according to the optimization objective and the initial feasible solution set.
[0050] A fourth aspect of the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the foregoing embodiments is implemented.
[0051] A fifth aspect of the present disclosure provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer device, the method described in any one of the foregoing embodiments is implemented.
[0052] A sixth aspect of the present disclosure provides a computer program product, which includes a computer program. When the computer program is executed by a processor of a computer device, the method described in any one of the foregoing embodiments is implemented.
[0053] The optimization method and device for the path planning strategy provided by the present disclosure convert the safety constraint conditions of the path planning performance function into a safety level variable representing the number of safety constraint conditions satisfied. Compared with directly calculating the safety constraint condition values in the prior art, calculating the safety level variable is easier and the calculation complexity is also lower. Therefore, the policy optimization based on the safety level variable is more capable of coping with uncertainty and improving the accuracy and efficiency of policy optimization.
[0054] By introducing the safety level variable, the optimization problem with safety constraint conditions can be transformed into an order optimization problem that satisfies the number of safety constraints. By sorting the policies in the path planning strategy space in descending order of the safety level and screening out a predetermined number of policies according to the sorting result order to form an initial feasible solution set, the policy solution space can be compressed, and sufficiently good feasible policies can be screened out from the limited policies.
[0055] An optimization objective is established based on the initial feasible solution set, the performance function, the safety constraint conditions, and the safety level variable. Among them, a constraint term for policy switching when the safety level is improved is defined in the optimization objective. The optimization objective is optimized using the performance gradient and safety gradient of the policy to obtain a path planning strategy, which not only considers the safety constraint conditions themselves but also the relationship between the safety constraint conditions, that is, the number of safety constraint conditions satisfied and the nested relationship of the constraints (reflected by the constraint term for policy switching when the safety level is improved). Therefore, the satisfaction of safety constraints and performance optimization can be combined, so that the policy always remains within the safety constraint conditions during the optimization process, and improvements are achieved in terms of efficiency, safety, and other indicators.
[0056] In order to make the above - mentioned and other objects, features and advantages of the present disclosure more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0058] Figure 1 A flowchart showing the optimization method of the path - planning strategy in the embodiment of the present disclosure;
[0059] Figure 2 A flowchart showing the determination process of the safety - level variables of each strategy in the embodiment of the present disclosure;
[0060] Figure 3 A flowchart showing the establishment process of the optimization objective in the embodiment of the present disclosure;
[0061] Figure 4 A schematic diagram showing the relationship between the safety and the strategy performance in the embodiment of the present disclosure;
[0062] Figure 5 A schematic diagram showing the effect of the optimization method in the embodiment of the present disclosure;
[0063] Figure 6 A structural diagram showing the optimization device of the path - planning strategy in the embodiment of the present disclosure;
[0064] Figure 7 A flowchart showing the optimization method of the target - task strategy in the embodiment of the present disclosure;
[0065] Figure 8 A structural diagram showing the computer device in the embodiment of the present disclosure.
[0066] Description of the Reference Signs in the Drawings:
[0067] 601, the first modeling unit;
[0068] 602, the second modeling unit;
[0069] 603, the feasible - solution search unit;
[0070] 604, the optimization - objective establishment unit;
[0071] 605, the optimization unit;
[0072] 802, the computer device;
[0073] 804. Processor;
[0074] 806. Memory;
[0075] 808. Driving mechanism;
[0076] 810. Input / output module;
[0077] 812. Input device;
[0078] 814. Output device;
[0079] 816. Rendering device;
[0080] 818. Graphical user interface;
[0081] 820. Network interface;
[0082] 822. Communication link;
[0083] 824. Communication bus. Detailed implementation manners
[0084] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0085] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0086] This specification provides the method operation steps as described in the embodiments or flowcharts, but may include more or fewer operation steps based on routine or non-creative labor. The step sequence listed in the embodiments is only one way among the execution sequences of numerous steps and does not represent the only execution sequence. When the actual system or device product is executed, it can be executed in the method sequence shown in the embodiments or the drawings, or in parallel.
[0087] It should be noted that the control optimization method and device herein can be used in the field of autonomous driving (such as autonomous vehicles), the field of intelligent mobile devices (such as automatic delivery robots, etc.). The motion devices involved include but are not limited to autonomous vehicles, automatic delivery robots, etc. In particular, any scenario involving path planning and having complex safety constraint conditions falls within the protection scope of the present disclosure.
[0088] In the prior art, existing path planning methods usually adopt rule-based strategy optimization and learning-based strategy planning methods. For rule-based methods, due to the ever-changing real scenarios, they cannot effectively explore the uncertainties in the environment. For learning-based methods, there has been insufficient research on how to analyze and solve multiple safety constraints and perform strategy optimization, and there are still the following two challenges:
[0089] On the one hand, reinforcement learning methods have achieved rich results and a large number of successful applications in dealing with unconstrained strategy optimization problems, but they have limitations in application scenarios where safety constraints are important. An effective comprehensive method for performance optimization and safety constraint handling has not been constructed. In the actual strategy optimization scenario with multiple safety constraints, performance optimization and safety constraints are generally processed separately, resulting in low data processing efficiency.
[0090] On the other hand, it cannot be guaranteed that the initial strategy and the strategy update process always remain within the (existing) safety constraint range. In the autonomous driving motion planning task, safety problems will bring greater potential hazards and costs. Therefore, the initial strategy and the strategy update process also need to be within the safety constraint range. Safety control based on models and rules generally assumes that the system model is accurate, but it is often difficult to obtain an accurate model in actual applications. Therefore, it is impossible to ensure that the dynamic process conforms to the real safety constraints, and it is difficult to ensure that the optimization process and the final strategy fully meet the safety constraints.
[0091] To solve the above technical problems in the prior art, the present disclosure provides an optimization method for path planning strategies, as Figure 1 shown, the optimization method for path planning strategies includes:
[0092] Step 101, establish a value optimization model for path planning strategy optimization, where the value optimization model includes: a performance function and multiple safety constraint conditions.
[0093] When this step is implemented, first, according to the characteristics of the operating environment of the motion device (for example, applied to roads, applied indoors, etc.), the following definitions are made:
[0094] Define the state space \(x=(x 0 ,x 1 ,…,x i ,…,x m ), where xx i and yy i represent the position coordinates; and represent the velocity coordinates; \(i\) represents the number of traffic participants, \(m\) represents the total number of traffic participants, where \(0\) represents the vehicle owner, and other values represent obstacles; \(x i represents the number of the \(i\)-th traffic participant.
[0095] Define the action space as \(a=(a 0 ,a 1 ,…,a m ), where and represent the acceleration coordinates.
[0096] Denote the relative distance and relative velocity between the host vehicle and the obstacle as \(\Delta p i =p i -p 0 and \(\Delta v i =v i -v 0 , where \(p i and \(p 0 respectively represent the positions of the vehicle owner and the obstacle, and \(v i and \(v 0 respectively represent the velocities of the vehicle owner and the obstacle. Then, \(\Delta p i \cdot\Delta v i <0 means that the host vehicle is closer to the obstacle (i.e., potentially unsafe).
[0097] According to the above-defined information, establish a performance function for measuring the distance of the motion device from the target position and the speed of its movement, and establish multiple safety constraint conditions for restricting that no collision occurs between the obstacle and the motion device. Specifically, the performance function is a function that simultaneously considers the states of the motion device such as position and speed. For example, it is desired that the motion device be closer to the target position and have a faster speed. A constraint is a key factor that represents, in an optimization problem, whether the solution of an absolute optimization problem can be accepted, expressed by an equation or an inequality. It is a series of rules imposed on decision variables, and the region that satisfies the constraints constructs the feasible region of the solution of the optimization problem.
[0098] In some embodiments, the performance function is represented by the following formula:
[0099] f(x) = Z1·p 0 + Z2·v 0 + Z3;
[0100] Wherein, Z1, Z2, and Z3 are constants that can be configured in advance; Z1·p 0 is used to characterize the current position coordinates of the vehicle owner, that is, to measure the distance from the target position; Z2·v 0 is used to measure the driving speed, that is, to measure the speed of the passing time.
[0101] The safety constraint conditions of each obstacle are represented by the following formula:
[0102] g i (x) = ‖Δv i ·Δp i ·Δt‖ - ‖Δp i ‖ - ∈ i ;
[0103] Wherein, g i (x) represents the safety constraint condition of the i-th obstacle; ‖Δv i ·Δp i ·Δt‖ represents the reduction value of the distance between the vehicle owner and the i-th obstacle within Δt time; ‖Δp i ‖ represents the distance between the i-th obstacle and the vehicle owner.
[0104] It should be noted here that the above performance function and safety constraint conditions are only examples. In specific implementation, there may be other forms of performance functions and other constraint conditions, which can be specifically set according to the actual situation. The present disclosure does not limit the specific form of the safety constraint conditions.
[0105] Step 102, convert the safety constraint conditions into safety level variables representing the number of safety constraint conditions satisfied.
[0106] Specifically, the safety level variable is represented by the following formula:
[0107] S(x) = ∑S i (x);
[0108] If g i (x) ≤ 0, then S i (x) = 1;
[0109] Wherein, S(x) is the safety level variable; i is the safety constraint condition number; g i (x) ≤ 0 is the case where the i-th safety constraint condition meets the safety constraint; S iWhen (x) is 0, it indicates non - compliance with safety constraints, and when it is 1, it indicates compliance with safety constraints; x represents the decision variable for optimizing the path - planning strategy. Taking the motion device control strategy as an example, the decision variables include, but are not limited to, the planned trajectory of the vehicle, the throttle and brake degrees of the vehicle, and the steering wheel angle, etc.
[0110] In this step, by converting the safety - constraint conditions of the path - planning performance function into safety - level variables representing the number of satisfied safety - constraint conditions, compared with directly calculating the values of safety - constraint conditions in the prior art, calculating safety - level variables is easier and the computational complexity is also lower. Therefore, the strategy optimization based on safety - level variables is more capable of coping with uncertainties and improving the accuracy and efficiency of strategy optimization.
[0111] Step 103: Sort the strategies in the path - planning strategy space in descending order of safety level, and screen out a predetermined number of strategies according to the sorting result order to form an initial feasible solution set.
[0112] When implementing this step, extract the first predetermined number of strategies in the order from front to back according to the sorting result to form an initial feasible solution set.
[0113] In some embodiments, as Figure 2 shown, the determination process of the safety level (i.e., the number of satisfied safety - constraint conditions) of each strategy includes:
[0114] Step 201: Use the pre - established path - planning simulation environment to execute each strategy in the strategy space. When implementing this step, since the number of satisfied safety - constraint conditions is determined, it is not necessary to accurately calculate the values of safety - constraint conditions. Based on this, the models of each obstacle in the path - planning simulation environment can be represented by rough models.
[0115] Specifically, the accurate model and the rough model refer to models that depict the physical and behavioral characteristics of vehicles and pedestrians to different degrees, and they are a relative concept. For example, for a vehicle, it can be depicted using a particle, or a single - track model, or even a real vehicle. The former can be a rough model, and the latter can be an accurate model. The rough model can be obtained from the accurate model. For example, by collecting the trajectory data of the accurate model and obtaining the parameters of the rough model through function fitting.
[0116] Step 202: Analyze whether each safety - constraint condition is satisfied according to the strategy execution result, where the strategy execution result includes the states of the vehicle owner and the obstacles.
[0117] Step 203: Determine the safety level of each strategy according to the number of satisfied safety - constraint conditions of each strategy. When implementing this step, use the number of satisfied safety - constraint conditions of each strategy as the safety level of each strategy.
[0118] In this step, by introducing the security level variable, the optimization problem with security constraints can be transformed into an order optimization problem that meets the number of security constraints. By sorting the strategies in the path planning strategy space in descending order of security level, a predetermined number of strategies are screened out according to the sorting result order to form an initial feasible solution set, which can compress the strategy solution space, screen out good enough feasible strategies from the limited strategies, and then improve the accuracy and efficiency of subsequent strategy optimization.
[0119] Step 104, establish an optimization objective according to the performance function, security constraints, and security level variable. Among them, a constraint term for strategy switching when the security level is increased (i.e., when the number of satisfied constraint conditions is increased) is defined in the optimization objective.
[0120] Step 105, optimize the path planning strategy according to the optimization objective and the initial feasible solution set.
[0121] The above steps 104 and 105 not only consider the security constraints themselves, but also consider the relationship between security constraints, that is, the number of satisfied security constraints and the nested relationship of constraints (reflected by the constraint term for strategy switching when the security level is increased), so that the satisfaction of security constraints can be combined with performance optimization, making the strategy always remain within the security constraints during the optimization process, and improving the indicators such as efficiency and security.
[0122] In this embodiment, the security constraint conditions are first converted into security level variables representing the number of satisfied security constraint conditions, realizing the transformation of the value evaluation of security into the order evaluation of security; then, based on the security level variables, the initial feasible solutions are quickly screened out; furthermore, an optimization objective is established by using the nested characteristics of the number of security constraints. Among them, a constraint term for strategy switching when the security level is increased is defined in the optimization objective; the nested characteristics are used as a trigger condition to guide the comprehensive performance gradient to optimize the optimization objective based on events, and further improve the security in the feasible and satisfactory strategy space.
[0123] In an embodiment of the present disclosure, as Figure 3 shown, the above step 104 establishing the optimization objective according to the initial feasible solution set, performance function, security constraint conditions, and security level variable includes:
[0124] Step 301, establish a first optimization term according to the security constraint conditions.
[0125] When implementing this step, the first constraint term can be established according to the unsatisfied security constraint conditions, so that the security constraint conditions are placed in the optimization objective and optimized synchronously with the performance function.
[0126] In some embodiments, the first constraint term is represented by the following formula:
[0127]
[0128] Among them, j is the safety constraint condition number; g j (x) is the j-th safety constraint condition; ω j is the weight of the j-th safety constraint condition; g j (x)>0 represents the situation where the j-th safety constraint condition does not meet the safety constraint; x represents the decision variable for optimizing the path planning strategy.
[0129] Step 302, convert the safety level variable into a continuous variable, and establish a second optimization term according to the continuous variable, where the second optimization term is used to constrain the switching strategy when a nested event occurs, and the nested event is an event when the safety level is increased.
[0130] Specifically, in implementation, based on the set of strategies to be simulated, construct a nested event according to the result of safety order optimization, and trigger the event e(s, s ′ ) = {<s, s ′ >|S(s ′ ) ≥ S(s)}, that is, the strategy switching is only performed when the safety level is increased.
[0131] Since S(x) is a discrete variable with a nested structure, during the policy iteration process, trigger an event when the safety of the policy crosses different safety levels within the nested safety domain: when the policy is in the completely safe area (that is, when s(x) is equal to m), the policy iteration aims at performance optimization, that is, only optimize using f(x); if starting from a policy that fails to meet all safety constraints (when s(x) is between m and 0), then the main goal is to find a feasible policy.
[0132] Introduce a surrogate model Fill the continuous area between the nested safety domains, and refine the number of safety constraint conditions with discrete values satisfied by the policy into the value of the policy in the safety constraints with continuous values.
[0133] In some embodiments, the second constraint term is represented by the following formula:
[0134]
[0135] Among them, x t represents the decision variable for optimizing the path planning strategy at time t, x t-1 represents the decision variable for optimizing the path planning strategy at time t-1, is the safety level at time t, represents the safety level at time t-1.
[0136] Step 303: Establish a reward function based on the performance function, the first optimization term, and the second optimization term.
[0137] When implementing this step, perform an addition calculation on the performance function, the first optimization term, and the second optimization term to obtain the reward function.
[0138] Continuing with the example of the foregoing implementation manner, the reward function is expressed as;
[0139]
[0140] Step 304: Establish an optimization objective based on the reward function. Specifically, the optimization objective is the expectation of the long-term cumulative discount term of the reward function.
[0141] In an embodiment of the present disclosure, when implementing step 105, utilize the performance gradient and safety gradient of the policy to optimize the optimization objective on the basis of the initial feasible solution set. Specifically, mainly determine the policy gradient direction according to the performance of the policy, and at the same time, a constraint term (which can be called a constraint term for nested events) that switches the policy when the safety level is improved provides more intermediate-level safety constraint satisfaction information to transition to the final feasible and satisfactory policy.
[0142] To illustrate the advantages of the technical solution of the present disclosure over the prior art, denote the policy optimization method used in the present disclosure as the nested event algorithm. At the same time, execute two other iterative algorithms for comparison:
[0143] The policy gradient algorithm, that is, without considering the intermediate-level safety constraint satisfaction information, but directly calculating the comprehensive performance gradient of the objective function generated by the reward function without nested events;
[0144] The safety ramp-up algorithm, that is, giving priority to considering policies with a higher safety level and then considering policies with better economic performance. Compared with the nested event algorithm, in event selection, use to replace e(s, s ′ )
[0145] The comparison results are as Figure 5 shown, Figure 5 showing the economic and safety performances of the three algorithms in the process of policy optimization. Figure 5 Among them, the dark gray line at the lower right corner is the schematic curve for finding the optimal policy of the policy gradient algorithm, the middle line is the schematic curve for finding the optimal policy of the nested event algorithm, and the light gray line at the upper left corner is the schematic curve for finding the optimal policy of the safety ramp-up algorithm to achieve the same performance as the nested event algorithm. By comparing the curves in Figure 5 , it can be seen that the nested event algorithm provided by the present disclosure can ensure the increasing relationship of safety, and can reach the optimal policy faster compared with the safety ramp-up algorithm.
[0146] Based on the same inventive concept, the present disclosure also provides an optimization device for a path planning strategy, as described in the following embodiments. Since the principle and method for the optimization device for the path planning strategy to solve problems are similar, therefore, the implementation of the optimization device for the path planning strategy may refer to the optimization method for the path planning strategy, and the repeated parts will not be elaborated.
[0147] Specifically, as Figure 6 shown, the optimization device for the path planning strategy includes:
[0148] A first modeling unit 601, configured to establish a value optimization model for optimizing the path planning strategy, where the value optimization model includes: a performance function and a plurality of safety constraint conditions;
[0149] A second modeling unit 602, configured to convert the safety constraint conditions into safety level variables representing the number of satisfied safety constraint conditions;
[0150] A feasible solution search unit 603, configured to sort the strategies in the path planning strategy space in descending order of safety level, and sequentially screen out a predetermined number of strategies according to the sorting result to form an initial feasible solution set;
[0151] An optimization objective establishment unit 604, configured to establish an optimization objective according to the performance function, the safety constraint conditions, and the safety level variables, where a constraint item for switching strategies when the safety level is increased is defined in the optimization objective;
[0152] An optimization unit 605, configured to optimize the path planning strategy according to the optimization objective and the initial feasible solution set. Specifically, the optimization objective is optimized by using the performance gradient and the safety gradient of the strategy to obtain the path planning strategy.
[0153] The optimization device for the path planning strategy provided in this embodiment can achieve the following technical effects:
[0154] (1) By using the internal structure and mutual association of the safety constraints, a multi-level safety domain (i.e., the safety level variable) that satisfies partial constraint conditions and is nested with each other is constructed, which can greatly reduce the solution complexity of the initial screening of strategies.
[0155] (2) By analyzing the multi-level safety domain nested with each other, a method for estimating the safety gradient of the strategy is provided, and a strategy iteration method driven by events of increasing the safety level (i.e., the number of satisfied safety constraint conditions) is established to ensure the safety of the strategy during the iteration process. While improving the safety signs, the performance is getting better and better.
[0156] (3) It can achieve offline training and online execution, meeting the deployment requirements in the actual scenario of autonomous driving.
[0157] In practical applications, the path planning strategy optimization algorithm can also be extended to other target task strategy optimization fields with multiple constraints besides the path planning strategy optimization scenario. The target tasks include, for example, the intelligent three-dimensional printing control strategy optimization field, and various robot control optimization fields, resource scheduling optimization in complex systems such as energy systems and logistics systems, etc. Correspondingly, as Figure 7 shown, the optimization method for the target task strategy includes:
[0158] Step 701, establish a value optimization model for the target task strategy optimization, where the value optimization model includes: a performance function and multiple safety constraint conditions.
[0159] Specifically, the target task and safety constraint conditions can be expressed as:
[0160] min -f(x);
[0161] s.t.g i (x)≤0, i = 1, …, m;
[0162] where f(x) represents the performance function; g i (x) represents different safety constraint conditions; denote g(x)=[g1(x), …, g m (x)] as an m-dimensional constraint function, then g(x)≤0 represents the case that conforms to the safety constraints; x represents the decision variable for the target task optimization; i represents the constraint condition number. The performance function and constraint conditions can be designed according to the actual situation. Specifically, when implemented, the safety constraint conditions include not only the conditions to ensure the safety after the strategy runs, but also other constraint conditions that need to be considered. The present disclosure does not limit specifically what the performance function and safety constraint conditions are.
[0163] In fact, the evaluation models corresponding to each safety constraint condition may have uncertainties, which makes it difficult to accurately characterize the constraint function. And for target tasks with high safety requirements, the optimization strategy often needs to be iterated on the basis of feasible strategies. Making the curve of g(x) for typical problems, it will be found that the constraint function is very easy to form a situation where the feasible region is not connected or even infeasible, making it difficult to perform subsequent iterative solutions.
[0164] Considering the above problems, the inventors proposed that it is necessary to address the difficulties of the uncertainty of the evaluation model of the safety constraint conditions and the sparsity of feasible strategies, and through adjusting the sampling distribution in the strategy space, achieve a balance between the strategy constraints and performance goals (such as safety and economy), so as to ensure the transition from feasible strategies to feasible and satisfactory strategies with the highest probability.
[0165] Traditional optimization searches for a single solution in the search space and hopes that the found solution coincides with the optimal solution. However, if it is possible to find a subset of sufficiently good solutions with a high probability, the probability of finding such solutions will increase as the subset grows. Based on this starting point, it is necessary to transform the constraint conditions established in step 701 so as to try to find as many and as good initial feasible solution sets as possible.
[0166] The inventors found through theoretical knowledge that order optimization has the following advantages:
[0167] First, "order" is easier to compare than "value";
[0168] Second, the goal is softened. Obtaining the theoretically optimal solution often requires many simulation evaluations and consumes a large amount of computational effort. However, in many cases, it is sufficient to obtain a truly sufficiently good solution that meets the engineering requirements. That is to say, a trade-off is made between the acceptable alignment probability and the computational burden.
[0169] Based on the above requirements and the advantages of order optimization, the following steps are proposed.
[0170] Step 702, convert the safety constraint conditions into safety level variables representing the number of satisfied constraint conditions. By introducing safety level variables in this step, the transformation from value optimization to order optimization can be achieved.
[0171] Specifically, for the constraint function g(x), S(x) is introduced to describe the number of conditions that satisfy the safety constraints. For i = 1,..., m, if g i (x) ≤ 0, then S i (x) = 1, and define S(x) = ∑S i (x), then S(x) ∈ [0, m] is a one-dimensional scalar. Where m is the number of constraint conditions and i is the constraint condition number.
[0172] Thus, the optimization problem with safety constraints is transformed into the following order optimization problem:
[0173] min -f(x);
[0174] s.t. S i (x) = 1, i = 1,..., m.
[0175] Compared with the uncertainty of value optimization, order optimization is more intuitive and naturally has the characteristic of layer-by-layer nesting. Specifically, different S(x) values correspond to different safety domains, and nested safety domains can be formed by different S(x) values. The safety domains in the nested safety domains have a progressive inclusion relationship layer by layer. For example, the safety domain with a larger value contains the safety domain with a smaller value. The optimization problem with safety constraints can be transformed into an order evaluation problem of the number of safety constraints to cope with the uncertain safety value estimation and quickly compress the strategy solution space.
[0176] Step 703: Sort the policies in the target task policy space in descending order of the constraint level, and screen out a predetermined number of policies according to the sorting result order to form an initial feasible solution set.
[0177] Specifically, as Figure 4 shown, the relationship between the security (reflected by the number of satisfied constraint conditions, i.e., reflected by S(x)) and the policy performance (reflected by f(x)) in the sample space is as follows Figure 2 shown. The points in the figure represent random samplings of the sample space, the curve represents the Pareto front of the sample points, S(x) = m represents completely feasible, and S(x) = 0 represents completely infeasible. The security rankings of each policy can be quickly evaluated by comparing the values of S(x) of each policy, and the solutions with the same S(x) are sorted in descending order of f(x), and the top predetermined number of policies are extracted from the sorting result to form the initial feasible solutions.
[0178] In specific implementation, the predetermined number is determined by the refined blind selection rule of the following formula:
[0179]
[0180] where N represents the amount of policies sampled from the policy space; G represents the policy set composed of the top g policies determined from N policies; g represents the number of policies; O represents the initial feasible solution set, o represents the predetermined number; k is a positive integer less than o; P represents the probability;
[0181] where N, g, k, and P are configuration quantities.
[0182] Specifically, the goal of the order optimization is to hope that the initial feasible solutions found are a batch that is good enough among a finite number of solutions. For example, there are 1000 solutions (corresponding to N), and it is hoped that the performance of the best several in the final batch of solutions (corresponding to O, unknown, i.e., the initial feasible solutions) is part of the top n% (e.g., the top 10%, such as the top 10, corresponding to g) of these 1000 solutions (corresponding to k, e.g., two) with a probability of (or greater than) a specified value P. Therefore, based on the above requirements, on the basis of configuring N, g, k, and P, using the formula the number of solutions in the initial feasible solutions that meet the above requirements, that is, the predetermined number, can be solved.
[0183] The formula is a formula that increases with O, that is, the larger O is, the greater this probability is. Therefore, according to N, g, k, and P, the smallest o can be solved, that is, the predetermined number is determined.
[0184] Step 704: Establish an optimization target based on the performance function, the constraint conditions and the security level variable, wherein the optimization target defines a constraint item for switching the strategy when the security level is increased.
[0185] By establishing multiple nested security areas, differential or difference information is introduced for the security assessment of the strategy. During the strategy iteration process, the performance gradient and security gradient of the strategy can be used comprehensively to improve the efficiency of the strategy iteration and ensure that the strategy meets the security constraints as much as possible in the middle of the optimization process.
[0186] Existing security policy optimization methods generally assume that security assessment is relatively easy, and do not explore the difference of the security of the policy relative to the policy behavior or the differential of the policy parameters. Based on the discovery of the above-mentioned problems of the prior art, the present disclosure extends the policy performance gradient for performance indicators to the security gradient for security indicators, and evaluates the impact of different directions of policy improvement on security and performance from the initial feasible policy. When the policy iterated according to the performance indicator violates the security constraint, not only can the policy be projected back to the safe area, but also the subsequent iteration direction can be adjusted based on the corrected policy, which can improve the efficiency of subsequent policy iterations, and promptly discover the phenomenon of slow or stagnant policy improvement when performance and security conflict with each other during the policy iteration process, and further combine it with the random sampling search method in the policy space.
[0187] Specifically, the optimization objective can be expressed using the following formula:
[0188]
[0189] Where f(x) represents the performance function; λ represents the constraint operator; ω j represents the weight of the jth constraint; g j (x) is the jth constraint; S(x) represents the number of constraints that are satisfied, which is a discrete variable with a nested structure; S(x0) is the number of constraints that are satisfied under the initial strategy.
[0190] Since S(x) is a discrete variable with a nested structure, during the policy iteration process, the security of the policy is used to trigger events across different security levels in the nested security domain: when the policy is in the completely safe area (when s(x) is equal to m), the policy iteration aims at performance optimization; if it starts from a strategy that fails to meet all security constraints (when s(x) is between m and 0), the main goal is to find a feasible strategy.
[0191] Introducing an alternative model Fill the continuous area between the nested security domains and refine the number of discrete-valued sub-security constraints satisfied by the policy to the value of the policy in the continuous-valued sub-security constraints. The optimization objective can be expressed as:
[0192]
[0193] Among them, λ 1,2 ∈ [0, ∞) represents the constraint multiplier.
[0194] Step 705, optimize the task policy according to the optimization objective and the initial feasible solution set. Comprehensively use the performance gradient and safety gradient of the policy to optimize it to transition to a feasible and satisfactory policy. Determine the policy gradient direction mainly based on the performance of the policy, and at the same time provide more information on the satisfaction of intermediate-level safety constraints.
[0195] In an embodiment of the present disclosure, a computer device is further provided, as Figure 8 shown, the computer device 802 may include one or more processors 804, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 802 may also include any memory 806 for storing any kind of information such as code, settings, data, etc. Non-limiting, for example, the memory 806 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any memory may use any technology to store information. Further, any memory may provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 802. In one case, when the processor 804 executes the associated instructions stored in any memory or combination of memories, the computer device 802 may perform any operation of the associated instructions. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.
[0196] The computer device 802 may also include an input / output module 810 (I / O) for receiving various inputs (via the input device 812) and for providing various outputs (via the output device 814). A specific output mechanism may include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), the input device 812, and the output device 814 may not be included, and it may only be a computer device in the network. The computer device 802 may also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the components described above together.
[0197] The communication link 822 can be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 822 can include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0198] Embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, and when the computer program is run by a processor, it executes the steps of the foregoing method.
[0199] Embodiments of the present disclosure also provide a computer-readable instruction, wherein when the processor executes the instruction, the program therein causes the processor to execute the method described in any one of the foregoing embodiments.
[0200] Embodiments of the present disclosure also provide a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor of a computer device, it implements the method described in any one of the foregoing embodiments.
[0201] It should be understood that in various embodiments of the present disclosure, the magnitudes of the serial numbers of the foregoing processes do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0202] It should also be understood that in the embodiments of the present disclosure, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the present disclosure, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0203] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present disclosure can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the foregoing description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present disclosure.
[0204] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0205] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.
[0206] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0207] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0208] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0209] Specific embodiments are applied in the present disclosure to elaborate on the principles and implementation manners of the present disclosure. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
Claims
1. A method for optimizing a path planning strategy, characterized in that: include: Establishing a value optimization model for path planning strategy optimization, wherein the value optimization model includes: a performance function and multiple safety constraints; Converting the safety constraint condition into a safety level variable representing the number of safety constraint conditions satisfied; Sort the strategies in the path planning strategy space in descending order of security level, and select a predetermined number of strategies according to the order of the sorting results to form an initial feasible solution set; Establishing an optimization target according to the performance function, the security constraint condition and the security level variable, wherein the optimization target defines a constraint item for switching the strategy when the security level is increased; Optimizing the path planning strategy according to the optimization objective and the initial feasible solution set; The performance function is used to measure the distance of the moving device from the target position and the speed of the moving device; The safety constraint condition is used to prevent the obstacle from colliding with the moving device; The safety level variable is expressed using the following formula: S(x)=∑S i (x); If g i (x)≤0, then S i (x) = 1; Where S(x) is the safety level variable; i is the safety constraint condition number; g i (x)≤0 means that the i-th safety constraint condition meets the safety constraint; S i When (x) is 0, it means that the safety constraint is not met, and when it is 1, it means that the safety constraint is met; x represents the decision variable for path planning strategy optimization; The process of determining the predetermined number includes: The reservation quantity is determined using the following formula for the fine-tuned blind selection rule: Where N represents the number of strategies sampled from the strategy space; G represents the strategy set consisting of the first g strategies determined from the N strategies; g represents the number of strategies; O represents the initial feasible solution set, o represents the predetermined number; k is a positive integer less than o; P represents the probability; Among them, N, g, k, and P are configuration quantities.
2. The method according to claim 1, characterized in that The path planning strategy for optimizing the optimization target includes: The performance gradient and safety gradient of the strategy are used to optimize the optimization target to obtain a path planning strategy.
3. The method according to claim 2, characterized in that The optimization objectives are established based on the initial feasible solution set, performance function, safety constraints and safety level variables, including: Establishing a first optimization item according to the safety constraint condition; Convert the security level variable into a continuous variable, and establish a second optimization item based on the continuous variable, wherein the second optimization item is used to constrain the switching strategy when a nested event occurs, and the nested event is an event when the security level is increased; Establishing a reward function according to the performance function, the first optimization item, and the second optimization item; Establish an optimization objective based on the reward function.
4. The method according to claim 3, characterized in that The first optimization item established according to the safety constraints includes: The first constraint is expressed as follows: Where j is the safety constraint number; g j (x) is the jth safety constraint; ω j is the weight of the jth safety constraint; g j (x)>0 means the jth safety constraint condition does not meet the safety constraint; x represents the decision variable for path planning strategy optimization.
5. The method according to claim 3, characterized in that Establishing the second optimization term according to the continuous variable includes: using the second constraint term expressed by the following formula: Among them, x t represents the decision variable for optimizing the path planning strategy at time t, x t-1 represents the decision variable for optimizing the path planning strategy at time t-1, is the safety level at time t, Indicates the safety level at time t-1.
6. A device for optimizing a path planning strategy, characterized in that: include: A first modeling unit is used to establish a value optimization model for path planning strategy optimization, wherein the value optimization model includes: a performance function and a plurality of safety constraints; A second modeling unit, configured to convert the safety constraint condition into a safety level variable representing the number of safety constraint conditions satisfied; A feasible solution search unit is used to sort the strategies in the path planning strategy space in descending order of security level, and select a predetermined number of strategies according to the order of sorting results to form an initial feasible solution set; An optimization target establishing unit, used to establish an optimization target according to the performance function, the security constraint condition and the security level variable, wherein the optimization target defines a constraint item for switching the strategy when the security level is increased; An optimization unit, used for optimizing the path planning strategy according to the optimization target and the initial feasible solution set; The performance function is used to measure the distance of the moving device from the target position and the speed of the moving device; The safety constraint condition is used to prevent the obstacle from colliding with the moving device; The safety level variable is expressed using the following formula: S(x)=∑S i (x); If g i (x)≤0, then S i (x) = 1; Where S(x) is the safety level variable; i is the safety constraint condition number; g i (x)≤0 means that the i-th safety constraint condition meets the safety constraint; S i When (x) is 0, it means that the safety constraint is not met, and when it is 1, it means that the safety constraint is met; x represents the decision variable for path planning strategy optimization; The process of determining the predetermined number includes: The reservation quantity is determined using the following formula for the fine-tuned blind selection rule: Where N represents the number of strategies sampled from the strategy space; G represents the strategy set consisting of the first g strategies determined from the N strategies; g represents the number of strategies; O represents the initial feasible solution set, o represents the predetermined number; k is a positive integer less than o; P represents the probability; Among them, N, g, k, and P are configuration quantities.
7. A method for optimizing a target task strategy, characterized in that: include: Establishing a value optimization model for target task strategy optimization, wherein the value optimization model includes: a performance function and multiple safety constraints; Converting the safety constraint condition into a safety level variable representing the number of safety constraint conditions satisfied; Sort the strategies in the target task strategy space in descending order of constraint levels, and select a predetermined number of strategies according to the order of the sorting results to form an initial feasible solution set; Establishing an optimization target according to the performance function, the security constraint condition and the security level variable, wherein the optimization target defines a constraint item for switching the strategy when the security level is increased; Optimizing the task strategy according to the optimization objective and the initial feasible solution set; The performance function is used to measure the distance of the moving device from the target position and the speed of the moving device; The safety constraint condition is used to prevent the obstacle from colliding with the moving device; The safety level variable is expressed using the following formula: S(x)=∑S i (x); If g i (x)≤0, then S i (x) = 1; Where S(x) is the safety level variable; i is the safety constraint condition number; g i (x)≤0 means that the i-th safety constraint condition meets the safety constraint; S i When (x) is 0, it means that the safety constraint is not met, and when it is 1, it means that the safety constraint is met; x represents the decision variable for path planning strategy optimization; The process of determining the predetermined number includes: The reservation quantity is determined using the following formula for the fine-tuned blind selection rule: Where N represents the number of strategies sampled from the strategy space; G represents the strategy set consisting of the first g strategies determined from the N strategies; g represents the number of strategies; O represents the initial feasible solution set, o represents the predetermined number; k is a positive integer less than o; P represents the probability; Among them, N, g, k, and P are configuration quantities.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
9. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor of a computer device, the method according to any one of claims 1 to 5 is implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor of a computer device, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Track planning method and device, electronic equipment and storage medium
CN113110489A
Intelligent automobile optimal decision control model building and solving method and device, and storage medium
CN113849903A