A self-driving control method and device, electronic equipment and storage medium

By using a preset rule-based scoring difference to trigger an arbitration model in an autonomous driving system, information about the vehicle and the environment is obtained, and the target trajectory is output using expert driving behavior. This solves the problem that existing technologies cannot distinguish between the merits of candidate trajectories, and achieves efficient and accurate trajectory decision-making in complex scenarios.

CN121341228BActive Publication Date: 2026-03-20ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511926049.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-20
Estimated Expiration
2045-12-19

AI Technical Summary

Technical Problem

In complex interactive scenarios of autonomous driving, existing technologies struggle to effectively distinguish the merits of candidate trajectories through pre-defined quantification rules, resulting in the inability to obtain the optimal driving strategy and limiting the system's performance in complex scenarios.

Method used

The arbitration model is triggered by determining the score difference based on preset rules, obtaining information about the vehicle and the environment, and using the pre-trained arbitration model to output target candidate trajectories based on expert driving behavior, and making decisions in combination with dynamic information.

Benefits of technology

It reduces computational overhead and response latency, improves the decision-making efficiency and quality of autonomous vehicle driving control, and ensures accurate trajectory decision-making in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121341228B_ABST
    Figure CN121341228B_ABST
Patent Text Reader

Abstract

The present disclosure provides a self-driving control method and device, electronic equipment and storage medium, the method comprising: scoring a plurality of candidate trajectories according to evaluation dimensions defined according to a preset rule, the evaluation dimensions including safety, efficiency and / or comfort, the plurality of candidate trajectories being generated based on inputting driving information at the same time during self-driving into different prediction algorithms; if the deviation of the score of at least one candidate trajectory from the highest score is less than a first preset threshold, obtaining self-vehicle information, self-vehicle interaction object information and environmental information within a preset past time at the current time; inputting the obtained information, the highest score candidate trajectory and the at least one candidate trajectory into a pre-trained arbitration model to enable the arbitration model to output a target candidate trajectory based on expert driving behavior; and controlling the self-vehicle to drive according to the target candidate trajectory. Accordingly, the arbitration model focuses on processing situations that are difficult to distinguish by rules, ensuring decision quality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One or more embodiments of the present disclosure relate to the technical field of autonomous driving, and in particular to a self-vehicle driving control method and device, an electronic device, and a storage medium. BACKGROUND

[0002] In the rule-network parallel architecture of autonomous driving, a vehicle generates candidate trajectories through rule planning algorithms and end-to-end deep network algorithms respectively. Existing technologies often use pre-set indicators (such as safety and comfort) to quantitatively score two trajectories, and select the trajectory with the highest score as the final driving trajectory.

[0003] However, in complex interactive scenarios such as turning at a signal-free intersection and gaming at a congested intersection, pre-set quantitative rules are difficult to fully depict optimal driving behavior, resulting in scores of the two candidate trajectories being very close and failing to effectively distinguish between the two. At this time, selecting a trajectory only according to the score may not obtain the optimal driving strategy, limiting the performance of the system in complex scenarios. SUMMARY

[0004] Therefore, the present disclosure provides a self-vehicle driving control method, which comprises:

[0005] scoring a plurality of candidate trajectories according to evaluation dimensions defined by pre-set rules; wherein the evaluation dimensions include safety, efficiency, and / or comfort, and the plurality of candidate trajectories are generated based on inputting driving information at the same time in the driving process of the self-vehicle into different prediction algorithms;

[0006] if the deviation of the score of at least one candidate trajectory from the highest score is less than a first pre-set threshold, obtaining self-vehicle information, self-vehicle interactive object information, and environmental information within a pre-set past time length at the current time;

[0007] inputting the obtained self-vehicle information, self-vehicle interactive object information, environmental information, highest score candidate trajectory, and at least one candidate trajectory into a pre-trained arbitration model, so that the arbitration model outputs a target candidate trajectory based on expert driving behavior;

[0008] controlling the self-vehicle to drive according to the target candidate trajectory.

[0009] Optionally, before the step of obtaining self-vehicle information, self-vehicle interactive object information, and environmental information within a pre-set past time length at the current time if the deviation of the score of at least one candidate trajectory from the highest score is less than a first pre-set threshold, the method further comprises:

[0010] if scores of the at least one candidate trajectory are all higher than the second preset threshold, determining whether deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold;

[0011] wherein the candidate trajectory with the score higher than the second preset threshold meets a bottom line requirement for driving of the ego vehicle.

[0012] Optionally, the plurality of candidate trajectories are two candidate trajectories, and if scores of the at least one candidate trajectory are all higher than the second preset threshold, determining whether deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold, comprises:

[0013] if the scores of the two candidate trajectories are both higher than the second preset threshold, determining whether a score deviation between the two candidate trajectories is less than the first preset threshold;

[0014] if the deviations between the scores of the at least one candidate trajectory and the highest score are all less than the first preset threshold, obtaining ego vehicle information, ego vehicle interactive object information and environment information within a preset past time length at a current time, comprises:

[0015] if the score deviation between the two candidate trajectories is less than the first preset threshold, obtaining the ego vehicle information, the ego vehicle interactive object information and the environment information within the preset past time length at the current time.

[0016] Optionally, the plurality of candidate trajectories are two candidate trajectories, and if scores of the at least one candidate trajectory are all higher than the second preset threshold, determining whether deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold, comprises:

[0017] if the scores of the two candidate trajectories are both higher than the second preset threshold, determining whether a score deviation between the two candidate trajectories is less than the first preset threshold;

[0018] The method further comprises:

[0019] if the score deviation between the two candidate trajectories is not less than the first preset threshold, controlling the ego vehicle to drive according to a candidate trajectory with a higher score in the two candidate trajectories.

[0020] Optionally, the plurality of candidate trajectories are two candidate trajectories, and the method further comprises:

[0021] if the scores of the two candidate trajectories are not both higher than the second preset threshold, determining whether there is a candidate trajectory in the two candidate trajectories with a score higher than the second preset threshold;

[0022] If there is a candidate trajectory whose score is higher than the second preset threshold, the ego vehicle is controlled to travel according to the candidate trajectory whose score is higher than the second preset threshold.

[0023] Optionally, the several candidate trajectories are two candidate trajectories, and the method further comprises:

[0024] If the scores of the two candidate trajectories are not both higher than the second preset threshold, it is determined whether there is a candidate trajectory whose score is higher than the second preset threshold in the two candidate trajectories;

[0025] If there is no candidate trajectory whose score is higher than the second preset threshold, a preset degradation strategy is processed;

[0026] The preset degradation strategy comprises at least one of the following: braking, prompting human intervention, and prompting human intervention to select the target candidate trajectory from the two candidate trajectories.

[0027] Optionally, the ego vehicle information, the ego vehicle interactive object information, and the environment information within the preset past time length at the current time are obtained, comprising:

[0028] The position, the acceleration, and the orientation angle of the ego vehicle within the preset past time length at the current time are obtained.

[0029] The position, the acceleration, the orientation angle, and the type of the ego vehicle interactive object within the preset past time length at the current time are obtained, wherein the type comprises at least one of the following: pedestrian, two-wheeled vehicle, three-wheeled vehicle, motor vehicle, and heavy vehicle.

[0030] The environment information within the preset past time length at the current time is obtained, wherein the environment information comprises at least one of the following: lane line, traffic sign, and traffic signal light.

[0031] The present disclosure also provides a self-driving vehicle driving control device, which comprises:

[0032] A scoring unit is configured to score several candidate trajectories according to evaluation dimensions defined by a preset rule, wherein the evaluation dimensions comprise safety, efficiency, and / or comfort, and the several candidate trajectories are generated based on driving information at the same time in the driving process of the ego vehicle input into different prediction algorithms.

[0033] An obtaining unit is configured to obtain ego vehicle information, ego vehicle interactive object information, and environment information within a preset past time length at the current time if there is at least one candidate trajectory whose score deviation from the highest score is less than a first preset threshold.

[0034] The input unit is configured to input the obtained self-vehicle information, the self-vehicle interactive object information, the environment information, the highest-scored candidate trajectory and the at least one candidate trajectory into a pre-trained arbitration model, so that the arbitration model outputs a target candidate trajectory based on expert driving behaviors;

[0035] The control unit is configured to control the self-vehicle to travel according to the target candidate trajectory.

[0036] The present disclosure also provides an electronic device, which comprises a communication interface, a processor, a memory and a bus, the communication interface, the processor and the memory are connected to each other through the bus;

[0037] The memory stores machine-readable instructions, and the processor executes the above method by invoking the machine-readable instructions.

[0038] The present disclosure also provides a machine-readable storage medium, which stores machine-readable instructions, and the machine-readable instructions realize the above method when invoked and executed by a processor.

[0039] Therefore, in the technical solution of the present disclosure, first, the candidate trajectories generated by different prediction algorithms are quantitatively scored based on the safety, efficiency, comfort and other dimensions defined by the preset rules; second, when it is detected that the score of at least one candidate trajectory deviates from the highest score by less than a first preset threshold, it is determined that the scene is a complex scene that is difficult to decide by rules, and then the self-vehicle information, interactive object information and environment information in the current and historical periods are obtained; then the above information is input into a pre-trained arbitration model together with the candidate trajectory with the highest score and other candidate trajectories with similar scores, the model learns the expert driving behavior mode, comprehensively analyzes the multi-source information and exceeds the explicit rule limit, and outputs a target candidate trajectory; finally, the self-vehicle is controlled to travel according to the target candidate trajectory.

[0040] In the above manner, on the one hand, the preset rules are used to score a plurality of candidate trajectories, and only in a complex scene where the scores cannot effectively distinguish the advantages and disadvantages, a pre-trained arbitration model is triggered to arbitrate, which can reduce the computational overhead and response delay and improve the running efficiency; on the other hand, the arbitration model focuses on processing complex situations that are difficult to distinguish by rules, and makes more accurate trajectory decisions in combination with dynamic environment information, thereby ensuring the decision quality of the self-vehicle driving control. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to make the technical solutions in the present disclosure better understood, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and all other drawings obtained by those of ordinary skill in the art without creative labor should be within the protection scope of the present disclosure.

[0042] Figure 1 FIG. 1 is a flowchart of a candidate trajectory arbitration method according to an example embodiment;

[0043] Figure 2 FIG. 2 is a schematic diagram of a candidate trajectory arbitration model training according to an example embodiment;

[0044] Figure 3 FIG. 3 is a schematic diagram of a two-candidate trajectory arbitration process according to an example embodiment;

[0045] Figure 4 FIG. 4 is a hardware structure diagram of an electronic device according to an example embodiment;

[0046] Figure 5 FIG. 5 is a block diagram of a self-vehicle driving control device according to an example embodiment. DETAILED DESCRIPTION

[0047] In order to make the technical solutions in the present disclosure better understood, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and all other drawings obtained by those of ordinary skill in the art without creative labor should be within the protection scope of the present disclosure.

[0048] It should be noted that: in other embodiments, the steps of the corresponding method do not necessarily follow the order shown and described in the present disclosure. In some other embodiments, the steps included in the method can be more or less than those described in the present disclosure. In addition, a single step described in the present disclosure may, in other embodiments, be divided into multiple steps for description; and multiple steps described in the present disclosure may, in other embodiments, be combined into a single step for description.

[0049] In the "rule-network parallel" architecture of automatic driving, the vehicle generates candidate trajectories through rule planning algorithm and end-to-end deep network algorithm respectively. The prior art often quantitatively scores two trajectories based on artificial preset indicators (such as safety, comfort), and selects the trajectory with the highest score as the final driving trajectory.

[0050] However, in complex interactive scenarios such as turning at a non-signalized intersection and gaming at a congested intersection, the preset quantization rules are difficult to comprehensively depict the optimal driving behavior, resulting in that the scores of the two candidate trajectories are very close and cannot effectively distinguish the advantages and disadvantages. At this time, selecting the trajectory only according to the score result may not obtain the optimal driving strategy, which limits the performance of the vehicle in complex scenarios.

[0051] Therefore, the present disclosure aims to provide a technical solution for determining whether to trigger an arbitration model to select a candidate trajectory based on a preset rule score difference.

[0052] The technical solution first scores a plurality of candidate trajectories according to evaluation dimensions defined by preset rules; wherein the evaluation dimensions include safety, efficiency and / or comfort, and the plurality of candidate trajectories are generated based on inputting driving information of the ego vehicle at the same time during driving into different prediction algorithms; then, if the score of at least one candidate trajectory and the highest score have a deviation less than a first preset threshold, the ego vehicle information, ego vehicle interactive object information and environment information within a preset past time are obtained; further, the ego vehicle information, the ego vehicle interactive object information, the environment information, the highest score candidate trajectory and the at least one candidate trajectory are input into a pre-trained arbitration model, so that the arbitration model outputs a target candidate trajectory based on expert driving behavior; finally, the ego vehicle is controlled to drive according to the target candidate trajectory.

[0053] For example, when an autonomous vehicle drives to a non-signalized fork in the evening rush hour, it is ready to turn left and merge into the main road. At this time, the oncoming traffic is dense, and there is a non-motor vehicle lane and a pedestrian crossing on the right. The vehicle controller (hereinafter referred to as: controller) generates three candidate trajectories based on the vehicle driving perception data at the same time through three different algorithms: trajectory A is generated by a conservative rule algorithm, which plans to stop outside the intersection and wait for a long gap in traffic flow in all directions before turning left at low speed; trajectory B is generated by an aggressive rule algorithm, which uses a short gap to quickly cut in; and trajectory C is generated by a deep network, which plans to keep crawling at low speed, predicts that the oncoming traffic will form a short gap, and plans to cut in smoothly during the short gap.

[0054] Firstly, the controller scores the above three candidate trajectories according to the evaluation dimensions (safety, efficiency and comfort) defined by the preset rules. Trajectory A scores 85 because it is completely stopped and far away from all obstacles, so it has a high safety score, but it has low comfort and efficiency because it is parked for a long time; trajectory B scores 50 because the minimum lateral distance between the vehicle and the oncoming vehicle is insufficient when cutting in, so it has low safety and comfort, although it has high efficiency; trajectory C scores 83 because all indicators meet the safety threshold, and the path is smooth and the travel time is reasonable. Since the score of trajectory C deviates from the score of the highest scoring trajectory A by 2, which is less than the first preset threshold 5, it is determined that the rule is difficult to decide the complex scene.

[0055] Then, the controller obtains the motion trajectory and state of the ego vehicle in the current and past 10 seconds, the motion trajectory and state of each vehicle in the opposite lane and the non-motor vehicle and pedestrian in the right lane, and the position of the intersection lane line and stop line and the traffic signal sign. Then, the controller inputs these information together with trajectory A and trajectory C into the pre-trained arbitration model. The arbitration model analyzes that trajectory A is too conservative and leads to low efficiency, while the vehicle gap used by trajectory C is safe and reasonable, which conforms to the travel strategy of human drivers in such a scene. Therefore, the arbitration model outputs trajectory C as the target candidate trajectory, and returns trajectory C to the vehicle controller. The vehicle controller generates precise steering, throttle and brake instructions accordingly to drive the vehicle to complete the left turn merge in a smooth and efficient manner, achieving a balance between safety and travel efficiency.

[0056] As can be seen, in the technical solution of the present disclosure, first, the candidate trajectories generated by different prediction algorithms are quantitatively scored based on the dimensions of safety, efficiency, comfort and the like defined by the preset rules; second, when the score of at least one candidate trajectory deviates from the highest score by less than the first preset threshold, it is determined that the rule is difficult to decide the complex scene, and then the ego vehicle information, interactive object information and environmental information in the current and historical period are obtained; then the above information together with the candidate trajectory with the highest score and other candidate trajectories with similar scores are input into the pre-trained arbitration model, which analyzes multiple sources of information and exceeds the limitation of explicit rules by learning expert driving behavior patterns, and outputs a target candidate trajectory; finally, the ego vehicle is controlled to travel according to the target candidate trajectory.

[0057] Through the above means, on the one hand, the preset rules are used to score a plurality of candidate trajectories, and only in the complex scene where the scores cannot effectively distinguish the advantages and disadvantages, the pre-trained arbitration model is triggered for arbitration, which can reduce the computational overhead and response delay, and improve the running efficiency; on the other hand, the arbitration model focuses on processing complex situations that are difficult to distinguish by rules, and makes more accurate trajectory decisions combined with dynamic environmental information, thereby ensuring the decision quality of the ego vehicle driving control.

[0058] The present disclosure will be described below by specific embodiments in combination with specific application scenarios.

[0059] Please refer to Figure 1 , Figure 1 is an exemplary embodiment showing a flowchart of a candidate trajectory arbitration method. The method can perform the following steps:

[0060] Step 102: score a plurality of candidate trajectories according to evaluation dimensions defined by preset rules; wherein the evaluation dimensions include safety, efficiency and / or comfort, and the plurality of candidate trajectories are generated based on inputting driving information at the same time during the driving process of the ego vehicle into different prediction algorithms.

[0061] For example, when an autonomous vehicle is driving to a fork in the road without signal control during the evening rush hour, it is ready to turn left and merge into the main road. At this time, the oncoming traffic is dense, and there is a non-motor vehicle lane and a pedestrian crossing on the right side. The vehicle controller (hereinafter referred to as: controller) generates three candidate trajectories based on the vehicle driving perception data at the same time through three different algorithms: trajectory A is generated by a conservative rule algorithm, which plans to stop outside the intersection and wait for a long gap in traffic flow in all directions before turning left at low speed; trajectory B is generated by an aggressive rule algorithm, which uses a short gap to quickly cut in; trajectory C is generated by a deep network, which plans to keep crawling at low speed, predicts that the oncoming traffic will form a short gap, and plans to cut in smoothly during the short gap. The controller scores the above three candidate trajectories according to the evaluation dimensions (safety, efficiency and comfort) defined by the preset rules.

[0062] Among them, the candidate trajectory refers to the path generated by different prediction algorithms in the autonomous driving system, which suggests the ego vehicle to drive in the future period of time, usually represented as a series of discrete points, including time-position-velocity-acceleration. The prediction algorithm for generating candidate trajectories mainly includes rule-based planning algorithm and end-to-end deep learning network. The preset rules are a set of index system for quantitatively evaluating the performance of the trajectory, and the "safety" dimension can be scored by calculating the minimum distance between the trajectory and all obstacles, the minimum collision time, whether it violates traffic regulations, etc. The "comfort" dimension can be measured by analyzing the peak and cumulative values of the longitudinal / lateral acceleration and jerk of the trajectory; the "efficiency" dimension is evaluated by estimating the travel time or total path length of the trajectory. The preset rules can be designed as a structured, artificially defined "score table" or "scoring matrix" to facilitate engineering implementation, debugging and ensure the interpretability of the decision.

[0063] Step 104: if the score of at least one candidate trajectory deviates from the highest score by less than a first preset threshold, obtain the ego vehicle information, ego vehicle interactive object information and environmental information within the preset past time at the current time.

[0064] For example, trajectory A has high safety score because it stops completely and keeps a long distance from all obstacles, but has low comfort and efficiency because it stops for a long time, so the overall score is 85; trajectory B has high efficiency but low safety and comfort because the minimum lateral distance to the oncoming vehicle is insufficient when cutting in, so the overall score is 50; trajectory C meets all safety thresholds, and the path is smooth and the travel time is reasonable, so the overall score is 83. Since the score of trajectory C deviates from the score of the highest scoring trajectory A by 2, which is less than the first preset threshold of 5, it is determined that it is a complex scene that is difficult for rules to decide. Then, the controller obtains the motion trajectory and state of the ego vehicle in the current and past 10 seconds, the motion trajectory and state of each vehicle in the opposite lane and the non-motor vehicle and pedestrian in the right lane, and the positions of the lane lines and stop lines at the intersection and the traffic signal signs.

[0065] The difference between the score of the part of the candidate trajectories and the highest score is used to quantify the "decision confidence" of the rule scoring system on the adaptive trajectory. When the difference is less than the "first preset threshold", it indicates that the overall performance of multiple trajectories under the preset rules is very close, and the rule system cannot clearly distinguish between good and bad, which is a "complex scene that is difficult for rules to decide". The ego vehicle information can include the pose (position, heading angle) of the ego vehicle, speed, acceleration, and historical trajectory point sequence, which can be obtained through the vehicle CAN bus and high-precision positioning system. The ego vehicle interaction object information refers to the information of traffic participants that have potential interaction or conflict with the ego vehicle, including its category (vehicle, pedestrian, non-motor vehicle), position, speed, acceleration, motion direction, and historical trajectory point sequence, which are output by the perception module (camera, radar) fusion. The environment information includes lane line type and topology, stop line position, traffic signal light state, roadside infrastructure, and road curvature, etc.

[0066] In this embodiment, the controller obtains the above information in the current and past period, the purpose of which is to provide sufficient context for the subsequent deep arbitration model, so that it can understand the traffic flow evolution trend and the game relationship between participants in a long time scale, and make more refined decisions that conform to human driving cognition.

[0067] Step 106: input the obtained ego vehicle information, ego vehicle interaction object information, environment information, highest scoring candidate trajectory, and at least one candidate trajectory into a pre-trained arbitration model, so that the arbitration model outputs a target candidate trajectory based on expert driving behavior.

[0068] For example, the controller inputs the current and past 10 seconds of ego motion trajectory and motion state, the motion trajectory and motion state of each vehicle in the opposite lane and non-motor vehicle and pedestrian in the right lane, the position of the intersection lane line and stop line, and the traffic signal sign and trajectory A and trajectory C into the pre-trained arbitration model. The arbitration model analyzes that trajectory A is too conservative, resulting in low efficiency, while trajectory C uses a safe and reasonable vehicle gap, which conforms to the passing strategy of human drivers in such scenarios. Therefore, the arbitration model outputs trajectory C as the target candidate trajectory and returns trajectory C to the vehicle controller.

[0069] The pre-trained arbitration model is a trained deep neural network model, and its training data is a large number of annotated complex scene samples (i.e. scenes with small rule score deviation), and the annotated label is the target trajectory considered by human experts. The input layer of the arbitration model is designed to fuse multi-modal information: ego information, interactive object information, and environment information constitute a complete dynamic description of the traffic scene; the highest scoring candidate trajectory and part of the candidate trajectory are also input as key decision options, and the arbitration model can analyze the advantages and disadvantages of each trajectory through attention mechanism and other methods. Pre-training means that the model is trained and verified on the server side using historical data before deployment on the vehicle, and the training goal is to minimize the difference between the model output and the expert annotated trajectory (such as using cross-entropy or trajectory point regression loss).

[0070] Step 108: controlling the ego vehicle to drive according to the target candidate trajectory.

[0071] For example, after receiving trajectory C output by the arbitration model, the vehicle controller generates precise steering, throttle, and brake instructions according to trajectory C to drive the vehicle to complete the left turn merge in a smooth and efficient manner, achieving a balance between safety and passing efficiency.

[0072] The target candidate trajectory refers to the spatiotemporal trajectory sequence finally decided and output by the arbitration model, which is the direct control target of the vehicle execution layer. The vehicle controller converts the target candidate trajectory into specific control instructions for the vehicle execution mechanism through the bottom layer control algorithm. For example, after receiving trajectory C, the controller analyzes its discrete trajectory points and generates a smooth control instruction sequence to ensure that the vehicle maintains a stable speed and a smooth steering wheel rotation during the crawling cut-in process, and finally safely and efficiently completes the left turn action, achieving accurate execution of the planning intent.

[0073] In the shown embodiment, before the self-vehicle information, the self-vehicle interactive object information and the environment information within the preset past time period are acquired, if the deviation of the score of the at least one candidate trajectory from the highest score is less than the first preset threshold, the method further comprises: if the score of the at least one candidate trajectory is higher than the second preset threshold, determining whether the deviation between the score of the at least one candidate trajectory and the highest score is less than the first preset threshold; wherein the candidate trajectory with the score higher than the second preset threshold meets the bottom line requirement of the self-vehicle driving.

[0074] For example, when the autonomous vehicle drives to a fork without signal control during the evening rush hour, it is ready to turn left and merge into the main road. At this time, the opposite traffic flow is dense, and there is a non-motor vehicle lane and a pedestrian crosswalk on the right side. The controller generates three candidate trajectories: trajectory A, trajectory B and trajectory C, based on the vehicle driving perception data at the same time through three different algorithms. The controller scores the above three candidate trajectories according to the evaluation dimensions (safety, efficiency and comfort) defined by the preset rules. The scoring results show that: the comprehensive score of trajectory A is 85, which is higher than the second preset threshold (comprehensive score ≥ 60); the comprehensive score of trajectory B is 50, which is lower than the second preset threshold; the comprehensive score of trajectory C is 83, which is higher than the second preset threshold. The controller calculates the score deviation of trajectory A and trajectory C with the score higher than the second preset threshold, and obtains that the score deviation of the two is 2, which is less than the first preset threshold 5, so it is determined that the rule is difficult to decide the complex scene.

[0075] The second preset threshold refers to the minimum access standard set for the comprehensive score of the candidate trajectory, which is used to judge whether the trajectory meets the basic requirement of the self-vehicle driving. Only the trajectory with the comprehensive score higher than the threshold is considered as an effective candidate scheme, which participates in the subsequent deviation judgment.

[0076] In this embodiment, by setting the front screening condition, the trajectories with poor comprehensive performance are excluded, and it is ensured that the schemes meeting the bottom line of safety and experience enter the arbitration process, so as to improve the decision efficiency and reliability.

[0077] In the shown embodiment, if the scores of the several candidate trajectories are all higher than the second preset threshold, it is determined whether the deviations between the scores of the at least one candidate trajectory and the highest score are all less than the first preset threshold, including: if the scores of the two candidate trajectories are all higher than the second preset threshold, it is determined whether the score deviation between the two candidate trajectories is less than the first preset threshold; and if the deviations between the scores of the at least one candidate trajectory and the highest score are all less than the first preset threshold, the ego vehicle information, the ego vehicle interactive object information and the environment information within the preset past time length at the current time are obtained, including: if the score deviation between the two candidate trajectories is less than the first preset threshold, the ego vehicle information, the ego vehicle interactive object information and the environment information within the preset past time length at the current time are obtained.

[0078] For example, when the autonomous vehicle drives to a fork road without signal control during the evening rush hour and is ready to turn left to merge into the main road. The controller generates two candidate trajectories based on the perception data: trajectory A (obtained by a conservative rule algorithm) has a comprehensive score of 85, and trajectory C (deep learning algorithm) has a comprehensive score of 83. The second preset threshold is set to 60. Since the scores of the two trajectories are both higher than 60, the system further calculates the score deviation of 2. If the first preset threshold is set to 5, because the score deviation of 2 is less than the first preset threshold of 5, it is determined as a complex scene, and the ego vehicle information, the interactive object information and the environment information are obtained, and the above information is input into the arbitration model for final decision.

[0079] In the shown embodiment, if the scores of the several candidate trajectories are all higher than the second preset threshold, it is determined whether the deviations between the scores of the at least one candidate trajectory and the highest score are all less than the first preset threshold, including: if the scores of the two candidate trajectories are all higher than the second preset threshold, it is determined whether the score deviation between the two candidate trajectories is less than the first preset threshold; and if the deviations between the scores of the at least one candidate trajectory and the highest score are all less than the first preset threshold, the ego vehicle information, the ego vehicle interactive object information and the environment information within the preset past time length at the current time are obtained, including: if the score deviation between the two candidate trajectories is less than the first preset threshold, the ego vehicle information, the ego vehicle interactive object information and the environment information within the preset past time length at the current time are obtained.

[0080] In the present embodiment, only when the two candidate trajectories both meet the two conditions of "being higher than the bottom line" and "having small deviation", the arbitration model with high calculation cost is called, so that the lightweight and high efficiency of the decision-making process are realized.

[0081] In order to facilitate those skilled in the art to better understand the present embodiment, the training process of the arbitration model is introduced as follows.

[0082] Please refer to Figure 2 , Figure 2 is a training schematic diagram of a candidate trajectory arbitration model according to an example embodiment. As shown in Figure 2As shown, the training system first performs a comprehensive score of two candidate trajectories (i.e., two driving schemes) in terms of safety, efficiency, and comfort, etc. based on artificial rules. Then, the training system determines whether the scores of the two candidate trajectories are both higher than a second preset threshold to ensure that they meet the basic safety and performance bottom line of the ego vehicle. If the condition is met, the training system further determines whether the score deviation between the two candidate trajectories is less than a first preset threshold for identifying complex scenarios that the rules are difficult to distinguish. When both conditions are met, the training system triggers a deep arbitration process: on the one hand, the information of the ego vehicle, the interactive object, and the environment within a preset past time are vectorized and input into the arbitration model to be trained; on the other hand, the expert arbitration result labeled by expert driving behavior is synchronously obtained as a reference label. Finally, the training system compares the network output with the expert result, and optimizes the arbitration model parameters through loss calculation back propagation, realizing a smooth transition from rule dominance to intelligent arbitration, effectively balancing the explainability of the rule system and the adaptability of the deep learning model, and improving the decision quality in complex scenarios under the premise of ensuring safety.

[0083] In an embodiment shown, the several candidate trajectories are two candidate trajectories, and if the scores of at least one candidate trajectory are both higher than the second preset threshold, the method further comprises: determining whether the score deviation between the two candidate trajectories is less than the first preset threshold if the scores of the two candidate trajectories are both higher than the second preset threshold; and the method further comprises: controlling the ego vehicle to drive according to the candidate trajectory with a higher score if the score deviation between the two candidate trajectories is not less than the first preset threshold.

[0084] For example, when an autonomous vehicle drives to a fork in the road without signal control during the evening rush hour, it is ready to turn left and merge into the main road. The controller generates two candidate trajectories based on the perception data: trajectory A (obtained by a conservative rule algorithm) has a comprehensive score of 75, and trajectory C (obtained by a deep learning algorithm) has a comprehensive score of 83. The second preset threshold is set to 60, and the first preset threshold is set to 5. Since the scores of the two candidate trajectories are both higher than 60, the bottom line requirement is met. However, the score deviation between the two trajectories is 8, which is greater than the first preset threshold 5, indicating that the rule system can clearly distinguish the advantages and disadvantages of the candidate trajectories. At this time, there is no need to call the arbitration model, and the controller directly selects the trajectory A with a higher score to control the vehicle to drive according to the trajectory A, and waits outside the intersection to ensure safe traffic.

[0085] Here, "the score deviation is not less than the first preset threshold" means that the performance difference between the candidate trajectories is significant, and the rule system has sufficient decision confidence. This determination can be realized in the vehicle controller with extremely low computational overhead through simple logical comparison.

[0086] In this embodiment, a hierarchical decision-making logic is implemented: the arbitration model is only activated when the scores are close (small deviation); otherwise, the high-scoring trajectory is used directly. This avoids ineffective calls to the arbitration model in simple scenarios, reduces computational overhead, embodies the design philosophy of "on-demand intelligence," and improves system resource utilization.

[0087] In one embodiment shown, the plurality of candidate trajectories are two candidate trajectories, and the method further includes: if the scores of the two candidate trajectories are not both higher than the second preset threshold, then determining whether there is a candidate trajectory among the two candidate trajectories whose score is higher than the second preset threshold; if there is a candidate trajectory whose score is higher than the second preset threshold, then controlling the vehicle to drive according to the candidate trajectory whose score is higher than the second preset threshold.

[0088] For example, when an autonomous vehicle approaches an uncontrolled intersection during rush hour, preparing to turn left to merge into the main road, the controller generates two candidate trajectories based on perception data: Trajectory A (obtained by a conservative algorithm) has a comprehensive score of 85, and Trajectory B (obtained by an aggressive algorithm) has a comprehensive score of 55. A second preset threshold is set at 60. Since the scores of the two candidate trajectories are not both higher than 60, the system enters a decision branch. After evaluation, Trajectory A's score is higher than the threshold, meeting the minimum safety and performance requirements, while Trajectory C does not meet the requirements. At this point, the controller directly selects Trajectory A as the target candidate trajectory, controlling the vehicle to wait safely outside the intersection, avoiding the use of a high-risk approach.

[0089] The condition "not all scores are above the second preset threshold" means that at least one of the two candidate trajectories fails to meet the system's overall performance baseline. This is used to quickly eliminate driving schemes with serious defects, ensuring that only trajectories meeting the minimum safety standards are adopted. In implementation, the controller uses logical judgment to select the only qualifying trajectory and executes it directly, without needing to initiate an arbitration model or perform complex comparisons. This design simplifies the decision-making process and improves the system's response efficiency and robustness in abnormal or high-risk scenarios while ensuring driving safety.

[0090] In one embodiment shown, the plurality of candidate trajectories are two candidate trajectories, and the method further includes: if the scores of the two candidate trajectories are not both higher than the second preset threshold, then determining whether there is a candidate trajectory whose score is higher than the second preset threshold; if there is no candidate trajectory whose score is higher than the second preset threshold, then processing is performed according to a preset degradation strategy; wherein, the preset degradation strategy includes at least one of the following: braking, prompting manual intervention, and prompting manual selection of the target candidate trajectory from the two candidate trajectories.

[0091] For example, when an autonomous vehicle is driving to a fork in the road without a traffic light control during the evening rush hour, preparing to turn left to merge onto the main road. The controller generates two candidate trajectories based on perception data: trajectory A (obtained by conservative rule algorithm) has a comprehensive score of 58, and trajectory B (obtained by aggressive algorithm) has a comprehensive score of 55. The second preset threshold is set to 60. Since the scores of both candidate trajectories are lower than 60, the rule system determines that no candidate trajectory meets the safety and performance bottom line requirements. Therefore, the controller triggers the preset degradation strategy, and the vehicle automatically slows down to a stop, and prompts the driver to take over the vehicle through the human-machine interaction interface to ensure driving safety.

[0092] Wherein, "the scores of all candidate trajectories are lower than the second preset threshold" means that all candidate trajectories do not meet the minimum operating standards set by the system in terms of safety, efficiency or comfort, which is a high-risk scenario that the decision system cannot handle autonomously. The preset degradation strategy is a key bottom mechanism to ensure the safety of autonomous driving, including the following strategies: smooth braking to safely stop the vehicle and prevent collision risks; issuing a takeover request through the human-machine interaction module to remind the driver to intervene in time; reminding the driver to choose one of the two candidate trajectories.

[0093] In this embodiment, in complex or high-risk scenarios, the system does not forcibly execute the candidate trajectory with hidden dangers, but realizes safety bottoming through "machine failure and manual intervention". While ensuring driving safety, it avoids dangerous behavior caused by decision-making deadlock of the system, and improves safety.

[0094] In order to facilitate those skilled in the art to better understand this embodiment, the following will introduce the complete arbitration process of the two candidate trajectories.

[0095] Please refer to Figure 3 , Figure 3 is a schematic diagram of an arbitration process of two candidate trajectories according to an exemplary embodiment. As shown in Figure 3 , the controller first scores the two candidate trajectories according to the artificially preset rules, and evaluates their performance in safety, efficiency and comfort. Then, the controller makes a first-level judgment: if the scores of the two schemes are both higher than the second preset threshold, it means that both of them meet the basic safety and performance requirements, and the process enters the second-level judgment - comparing whether the score difference is less than the first preset threshold. If the score difference is small, it means that the rules are difficult to distinguish between the two, which belongs to a complex scenario, and the controller triggers the deep learning arbitration model immediately, inputs the information of the ego vehicle, environment and candidate trajectories, and outputs the final optimal scheme from the arbitration model; if the score difference is large, it means that the rules can make a clear decision, and the controller directly selects the candidate trajectory with higher score as the target trajectory. If only one scheme meets the score, it is directly adopted; if neither of them meets the score, the controller executes the degradation strategy (such as emergency braking or prompting takeover).

[0096] The entire process is completed in real time by the controller on the vehicle computing platform, realizing a hierarchical decision-making mechanism of "rules first, intelligent supplementation", which balances decision-making efficiency and intelligence level while ensuring safety.

[0097] In one embodiment shown, obtaining vehicle information, vehicle interaction object information, and environmental information within the current time and a preset past time period includes: obtaining the vehicle's position, acceleration, and orientation angle within the current time and a preset past time period; obtaining the position, acceleration, orientation angle, and type of the vehicle interaction object within the current time and a preset past time period; wherein the type includes at least one of the following: pedestrian, two-wheeled vehicle, three-wheeled vehicle, motor vehicle, and heavy vehicle; obtaining environmental information within the current time and a preset past time period; wherein the environmental information includes at least one of the following: lane lines, traffic signs, and traffic lights.

[0098] For example, when an autonomous vehicle approaches an uncontrolled intersection during rush hour, preparing to turn left into the main road, the controller generates two candidate trajectories based on perception data: Trajectory A (obtained by a conservative rule algorithm) has a comprehensive score of 85, and Trajectory C (obtained by a deep learning algorithm) has a comprehensive score of 83. The score difference of 2 is less than a first preset threshold of 5, thus triggering the arbitration process. The controller then acquires dynamic information from the current and past 8 seconds: the vehicle's position moves from (116.45°E, 39.92°N) to (116.46°E, 39.92°N), the longitudinal acceleration gradually changes from -2.0 m / s² to 0.1 m / s², and the heading angle smoothly adjusts from 265° to 275°, indicating a left-turn preparation posture; regarding the interacting objects, an oncoming vehicle continues to move forward with an acceleration of -1.5 m / s² (deceleration), and its heading angle stabilizes at 90°; a two-wheeled vehicle on the right approaches the non-motorized vehicle lane stop line with an acceleration of 0.6 m / s², and its heading angle changes from 270° to 250° (preparing to turn right); a pedestrian briefly pauses before moving, with acceleration increasing to 0.7 m / s² and a heading angle of 180°. Environmental information includes a left-turn guide lane curvature of 0.005 rad / m, and the absence of traffic lights at the upcoming intersection, but the presence of a "yield" sign.

[0099] wherein, the position, acceleration, and orientation angle represent the spatial coordinates, the rate of change of velocity, and the heading direction (unit: degree) respectively, which can be calculated in real time through high-precision positioning and sensor fusion. The type is used to distinguish the category of the traffic participant, and supports the arbitration model to call the corresponding behavior prediction model. The lane line in the environment information includes the start and end positions of the lane line, the curvature, etc. The traffic signal light state in the environment information includes not only the color of the traffic light (red / yellow / green), but also the key state of "no signal light", which can be obtained through visual perception and high-precision map matching. The collection mechanism in the embodiment can align multiple source data through timestamps, construct high-fidelity scene representation, and improve the decision accuracy of the arbitration model in complex interaction scenarios.

[0100] Corresponding to the above-mentioned embodiment of the self-vehicle driving control method, the present disclosure also provides an embodiment of a self-vehicle driving control device.

[0101] Please refer to Figure 4 , Figure 4 is an exemplary embodiment showing a hardware structure diagram of an electronic device. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410, and of course can also include other required hardware. One or more embodiments of the present disclosure can be implemented in a software manner, such as reading a corresponding computer program from the non-volatile memory 410 into the memory 408 by the processor 402 and then running. Of course, in addition to the software implementation, one or more embodiments of the present disclosure do not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.

[0102] Please refer to Figure 5 , Figure 5 is an exemplary embodiment showing a block diagram of a self-vehicle driving control device. The self-vehicle driving control 500 can be applied to an electronic device as shown in Figure 4 to implement the technical solutions of the present disclosure. The device includes:

[0103] The scoring unit 502 is configured to score a plurality of candidate trajectories according to evaluation dimensions defined by a preset rule, wherein the evaluation dimensions include safety, efficiency, and / or comfort, and the plurality of candidate trajectories are generated based on inputting driving information at the same time in the driving process of the self-vehicle into different prediction algorithms;

[0104] The acquisition unit 504 is configured to acquire self-vehicle information, self-vehicle interaction object information, and environment information within a preset past time length if the deviation of the score of at least one candidate trajectory from the highest score is less than a first preset threshold.

[0105] The input unit 506 is configured to input the obtained self-vehicle information, the self-vehicle interactive object information, the environment information, the highest-score candidate trajectory, and the at least one candidate trajectory into the pre-trained arbitration model, so that the arbitration model outputs a target candidate trajectory based on expert driving behavior.

[0106] The first control unit 508 is configured to control the self-vehicle to travel according to the target candidate trajectory.

[0107] In some embodiments, before the self-vehicle information, the self-vehicle interactive object information, and the environment information within the preset past time length and the current time are obtained, the device further comprises:

[0108] The first determination unit 510 is configured to determine whether the deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold if the scores of the at least one candidate trajectory are all higher than a second preset threshold.

[0109] The candidate trajectory with a score higher than the second preset threshold meets the bottom line requirement of the self-vehicle driving.

[0110] In some embodiments, the plurality of candidate trajectories are two candidate trajectories, and the determination whether the deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold if the scores of the at least one candidate trajectory are all higher than a second preset threshold comprises:

[0111] If the scores of the two candidate trajectories are both higher than the second preset threshold, it is determined whether the score deviation between the two candidate trajectories is less than the first preset threshold.

[0112] The obtaining of the self-vehicle information, the self-vehicle interactive object information, and the environment information within the preset past time length and the current time if the deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold comprises:

[0113] If the score deviation between the two candidate trajectories is less than the first preset threshold, the self-vehicle information, the self-vehicle interactive object information, and the environment information within the preset past time length and the current time are obtained.

[0114] In some embodiments, the plurality of candidate trajectories are two candidate trajectories, and the determination whether the deviations between the scores of the at least one candidate trajectory and the highest score are all less than a first preset threshold if the scores of the at least one candidate trajectory are all higher than a second preset threshold comprises:

[0115] determining whether a score deviation between the two candidate trajectories is less than the first preset threshold value if the scores of the two candidate trajectories are both higher than the second preset threshold value;

[0116] The apparatus further includes:

[0117] The second control unit 512 is configured to control the ego vehicle to travel according to the candidate trajectory with a higher score if the score deviation between the two candidate trajectories is not less than the first preset threshold value.

[0118] In some embodiments, the plurality of candidate trajectories includes two candidate trajectories, and the apparatus further includes:

[0119] The second determination unit 514 is configured to determine whether there is a candidate trajectory with a score higher than the second preset threshold value if the scores of the two candidate trajectories are not both higher than the second preset threshold value.

[0120] The third control unit 516 is configured to control the ego vehicle to travel according to the candidate trajectory with a score higher than the second preset threshold value if there is a candidate trajectory with a score higher than the second preset threshold value.

[0121] In some embodiments, the plurality of candidate trajectories includes two candidate trajectories, and the apparatus further includes:

[0122] The second determination unit 514 is configured to determine whether there is a candidate trajectory with a score higher than the second preset threshold value if the scores of the two candidate trajectories are not both higher than the second preset threshold value.

[0123] The processing unit 518 is configured to perform processing according to a preset degradation strategy if there is no candidate trajectory with a score higher than the second preset threshold value.

[0124] The preset degradation strategy includes at least one of the following: braking, prompting human intervention, and prompting human intervention to select the target candidate trajectory from the two candidate trajectories.

[0125] In some embodiments, the obtaining of the ego vehicle information, the ego vehicle interactive object information, and the environment information within a preset past time period and at a current time includes:

[0126] obtaining a position, an acceleration, and an orientation angle of the ego vehicle within the preset past time period and at the current time;

[0127] obtaining a position, an acceleration, an orientation angle, and a type of the ego vehicle interactive object within the preset past time period and at the current time; the type includes at least one of the following: a pedestrian, a two-wheeled vehicle, a three-wheeled vehicle, a motor vehicle, and a heavy vehicle;

[0128] Obtain environmental information for the current time and a preset past time period; wherein the environmental information includes at least one of the following: lane lines, traffic signs, and traffic lights.

[0129] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0130] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0131] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0132] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0133] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0134] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0135] The user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0136] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0137] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multi-task processing and parallel processing are possible or can be advantageous.

[0138] The terminology used in this disclosure one or more embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure one or more embodiments. As used in this disclosure one or more embodiments and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0139] It is to be understood that the terms first, second, third, etc. can be employed merely for distinguishing between similar information in this disclosure one or more embodiments, and to not necessarily limit the scope of the disclosure one or more embodiments. Such terms can be interchangeable under proper context. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure one or more embodiments. For example, as used herein, the word "if" can be construed to mean "when" or "upon" or "in response to determining" or "when it is determined" or "upon determining" depending on the context.

[0140] The description above only illustrates and explains the preferred embodiments of the disclosure one or more embodiments, and is not intended to limit the disclosure one or more embodiments. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the disclosure one or more embodiments shall be included in the scope of protection of the disclosure one or more embodiments.

Claims

1. A method for controlling the movement of a vehicle, characterized in that, The method includes: According to the evaluation dimensions defined by preset rules, several candidate trajectories are scored respectively; wherein, the evaluation dimensions include safety, efficiency and / or comfort, and the several candidate trajectories are generated based on inputting driving information at the same moment during the vehicle's driving process into different prediction algorithms; If the deviation between the score of at least one candidate trajectory and the highest score is less than the first preset threshold, then obtain the vehicle information, vehicle interaction object information and environmental information for the current time and the preset past time. The acquired vehicle information, vehicle interaction object information, environmental information, highest score candidate trajectory, and at least one candidate trajectory are input into a pre-trained arbitration model so that the arbitration model outputs a target candidate trajectory based on expert driving behavior. The vehicle is controlled to move based on the target candidate trajectory.

2. The method according to claim 1, characterized in that, Before obtaining the vehicle information, vehicle interaction object information, and environment information for the current time and the preset past time period if the deviation between the score of at least one candidate trajectory and the highest score is less than a first preset threshold, the method further includes: If at least one candidate trajectory has a score higher than the second preset threshold, then determine whether the deviation between the score of the at least one candidate trajectory and the highest score is less than the first preset threshold. Among them, candidate trajectories with scores higher than the second preset threshold meet the minimum requirements for the vehicle's driving.

3. The method according to claim 2, characterized in that, The plurality of candidate trajectories consists of two candidate trajectories. If at least one candidate trajectory has a score higher than a second preset threshold, then determining whether the deviation between the score of the at least one candidate trajectory and the highest score is less than a first preset threshold includes: If the scores of both candidate trajectories are higher than the second preset threshold, then it is determined whether the score deviation between the two candidate trajectories is less than the first preset threshold. If the deviation between the score of at least one candidate trajectory and the highest score is less than a first preset threshold, then the vehicle information, vehicle interaction object information, and environmental information within the current time and a preset past time period are obtained, including: If the score deviation between the two candidate trajectories is less than the first preset threshold, then obtain the vehicle information, vehicle interaction object information, and environmental information for the current time and the preset past time period.

4. The method according to claim 2, characterized in that, The plurality of candidate trajectories consists of two candidate trajectories. If at least one candidate trajectory has a score higher than a second preset threshold, then determining whether the deviation between the score of the at least one candidate trajectory and the highest score is less than a first preset threshold includes: If the scores of both candidate trajectories are higher than the second preset threshold, then it is determined whether the score deviation between the two candidate trajectories is less than the first preset threshold. The method further includes: If the score deviation between the two candidate trajectories is not less than the first preset threshold, then the vehicle is controlled to drive according to the candidate trajectory with the higher score.

5. The method according to claim 2, characterized in that, The plurality of candidate trajectories consists of two candidate trajectories, and the method further includes: If the scores of the two candidate trajectories are not both higher than the second preset threshold, then it is determined whether there is a candidate trajectory whose score is higher than the second preset threshold. If a candidate trajectory has a score higher than the second preset threshold, the vehicle's driving is controlled based on the candidate trajectory with a score higher than the second preset threshold.

6. The method according to claim 2, characterized in that, The plurality of candidate trajectories consists of two candidate trajectories, and the method further includes: If the scores of the two candidate trajectories are not both higher than the second preset threshold, then it is determined whether there is a candidate trajectory whose score is higher than the second preset threshold. If no candidate trajectory has a score higher than the second preset threshold, then it will be processed according to the preset downgrade strategy. The preset degradation strategy includes at least one of the following: braking, prompting manual intervention, and prompting manual selection of the target candidate trajectory from the two candidate trajectories.

7. The method according to claim 1, characterized in that, The acquisition of vehicle information, vehicle interaction object information, and environmental information within the current time and a preset past time period includes: Obtain the vehicle's position, acceleration, and orientation angle at the current moment and within a preset past time period; The system obtains the position, acceleration, orientation angle, and type of the vehicle interaction object at the current time and within a preset past time period; wherein the type includes at least one of the following: pedestrian, two-wheeled vehicle, three-wheeled vehicle, motor vehicle, and heavy vehicle; Obtain environmental information for the current time and a preset past time period; wherein the environmental information includes at least one of the following: lane lines, traffic signs, and traffic lights.

8. A vehicle driving control device, characterized in that, The device includes: The scoring unit is used to score several candidate trajectories according to the evaluation dimensions defined by preset rules; wherein, the evaluation dimensions include safety, efficiency and / or comfort, and the several candidate trajectories are generated based on inputting driving information at the same moment during the vehicle's driving process into different prediction algorithms; The acquisition unit is used to acquire vehicle information, vehicle interaction object information and environmental information at the current time and within a preset past time period if the deviation between the score of at least one candidate trajectory and the highest score is less than a first preset threshold. The input unit is used to input the acquired vehicle information, vehicle interaction object information, environmental information, highest score candidate trajectory and at least one candidate trajectory into a pre-trained arbitration model, so that the arbitration model outputs a target candidate trajectory based on expert driving behavior; The control unit is used to control the vehicle's movement based on the target candidate trajectory.

9. An electronic device, characterized in that, It includes a communication interface, a processor, a memory, and a bus, wherein the communication interface, the processor, and the memory are interconnected via the bus; The memory stores machine-readable instructions, and the processor executes the method according to any one of claims 1 to 7 by invoking the machine-readable instructions.

10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-readable instructions, which, when invoked and executed by a processor, implement the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic driving comprehensive decision evaluation method and device considering multiple types of driving elements

    CN117668413A

  • Track selection method and device, electronic equipment, storage medium and program product

    CN121106347A