Method for acquiring following distance, intelligent device and vehicle
By acquiring road scene information and reinforcement learning data from the autonomous driving system, the following distance is dynamically adjusted, solving the adaptability problem of a fixed following distance threshold in complex scenarios, and achieving a balance between safety and efficiency as well as improved comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 安徽蔚来智驾科技有限公司
- Filing Date
- 2026-04-15
- Publication Date
- 2026-08-04
AI Technical Summary
Existing speed-based fixed following distance thresholds are difficult to accurately reflect human driving habits in complex road scenarios, causing autonomous vehicles to exhibit either overly conservative or overly aggressive behavior, affecting traffic efficiency and passenger comfort.
By acquiring road scene information, the following distance is dynamically adjusted based on trajectory data during the reinforcement learning process. Combining driving difficulty coefficient and historical following distance data, the baseline value is calculated using the exponential moving average method. The following distance is then updated through following distance rewards and longitudinal acceleration to achieve dynamic optimization of the following distance.
In complex road scenarios, autonomous driving systems can learn following strategies that conform to human driving habits, reduce conservative behavior, avoid following too closely, achieve a dynamic balance between safety and traffic efficiency, and improve driving comfort.
Smart Images

Figure CN122034980B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, specifically to a method for obtaining following distance, an intelligent device, and a vehicle. Background Technology
[0002] Following behavior is one of the core issues in trajectory planning and decision-making control in autonomous driving technology. Time headway (THW), as an important indicator for measuring the longitudinal safe distance of a vehicle, is usually used to constrain the minimum following distance of autonomous vehicles at different speeds. Currently, in autonomous driving systems, the time headway is usually set as a set of fixed or segmented fixed time headway thresholds based on vehicle speed, used for rule constraints and / or reward design in reinforcement learning.
[0003] However, real-world road scenarios are highly complex and dynamic. For example, in different scenarios such as congested roads, rainy or snowy weather, and low-light conditions at night, even at the same vehicle speed, the actual following distance and corresponding following time used by human drivers can vary significantly. Autonomous driving control based on a fixed following time threshold according to speed struggles to accurately reflect reasonable following behavior in different road scenarios, easily leading to overly conservative or overly aggressive behavior from autonomous vehicles, thus affecting traffic efficiency and passenger comfort. Therefore, how to dynamically adjust the following time based on real-time data from road scenarios and the autonomous driving system to improve the human-likeness and adaptability of autonomous driving systems in complex road conditions has become an urgent problem to be solved.
[0004] Accordingly, there is a need in this field for a new method to obtain following distance information in order to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned deficiencies, this application is made to solve, or at least partially solve, the technical problem of how to dynamically adjust the following distance based on road scenarios and real-time data from autonomous driving systems.
[0006] In a first aspect, a method for obtaining following distance is provided, applied to reinforcement learning in an autonomous driving system, the method comprising:
[0007] Obtain the road scene;
[0008] Based on the road scenario, determine the initial value of the following distance;
[0009] Update following distance based on trajectory data from the reinforcement learning process;
[0010] Based on the updated following distance, the autonomous driving system continues to undergo reinforcement learning.
[0011] In one technical solution of the above-mentioned method for obtaining following distance, determining the initial value of the following distance based on the road scenario includes:
[0012] Obtain historical following distance data corresponding to the road scene, wherein the historical following distance data is either human-driven historical following distance data or self-driving historical following distance data.
[0013] Based on the historical following distance data, obtain the baseline value of the following distance for this road scenario;
[0014] Based on the road scene and the following distance baseline value, an initial value of the following distance is determined, wherein the initial value of the following distance includes an initial value of the following distance threshold, an initial value of the lower limit of the following distance interval, and an initial value of the upper limit of the following distance interval. The initial value of the lower limit of the following distance interval is less than the following distance threshold, and the initial value of the upper limit of the following distance interval is greater than the following distance threshold.
[0015] In one technical solution of the above-mentioned method for obtaining following distance, the road scenario includes a driving difficulty coefficient. When the historical following distance data is the historical following distance data of the driver, the method for determining the initial value of the following distance includes:
[0016] Based on the first percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the following distance baseline value is expanded to obtain the initial value of the following distance threshold.
[0017] Based on the second percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the initial value of the following distance threshold is reduced to obtain the initial value of the lower limit of the following distance interval.
[0018] Based on the third percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the initial value of the following distance threshold is expanded to obtain the initial value of the upper limit of the following distance interval.
[0019] In one technical solution of the above-mentioned method for obtaining following distance, the method for obtaining the following distance baseline value includes:
[0020] Based on a preset quantile, the historical following distance data is filtered to obtain valid historical following distance data.
[0021] Based on the effective historical following distance data, the baseline value of the following distance is obtained by using the exponential moving average method.
[0022] In one technical solution of the above-mentioned method for obtaining following distance, the trajectory data includes the actual following distance and longitudinal acceleration, and the following distance includes a following distance threshold, an upper limit of the following distance interval, and a lower limit of the following distance interval.
[0023] The process of updating the following distance based on trajectory data during reinforcement learning includes:
[0024] Based on the actual following distance and the following distance, obtain the following distance reward;
[0025] Update the following distance based on the actual following distance and the following distance bonus; and / or,
[0026] The following distance is updated based on the actual following distance and the longitudinal acceleration.
[0027] In one technical solution of the above-mentioned method for obtaining following distance, updating the following distance based on the actual following distance and the following distance bonus includes:
[0028] Obtain the average time distance of the actual following distance within a preset first time range, and the average reward of the following distance bonus.
[0029] When the average reward statistics are less than 0, the following distance threshold is increased, wherein the larger the absolute value of the average reward statistics, the larger the increase in the following distance threshold.
[0030] When the average reward statistics are less than the first following distance reward threshold, and the average distance statistics are greater than the following distance threshold, the following distance threshold is reduced.
[0031] In one technical solution of the above-mentioned method for obtaining following distance, updating the following distance based on the actual following distance and the longitudinal acceleration includes:
[0032] Based on the longitudinal acceleration within a preset second time range, obtain the root mean square value of the longitudinal weighted acceleration;
[0033] When the longitudinal weighted root mean square value of acceleration is greater than the first weighted root mean square threshold and less than the second weighted root mean square threshold, the upper limit of the following distance interval is increased and / or the lower limit of the following distance interval is decreased.
[0034] When the longitudinal weighted root mean square value of acceleration is greater than the second weighted root mean square threshold, the following distance threshold is increased.
[0035] In one technical solution of the above-mentioned method for obtaining following distance, the method for obtaining following distance reward includes:
[0036] When the actual following distance is less than the upper limit of the following distance range and greater than or equal to the following distance threshold, the following distance reward is positive. The smaller the absolute value of the difference between the actual following distance and the following distance threshold, the larger the following distance reward.
[0037] When the actual following distance is less than the following distance threshold and is greater than or equal to the lower limit of the following distance range, the following distance reward is positive. The smaller the absolute value of the difference between the actual following distance and the following distance threshold, the larger the following distance reward.
[0038] When the actual following distance is less than the lower limit of the following distance interval, the following distance reward is negative, and the larger the absolute value of the difference between the actual following distance and the following distance threshold, the larger the absolute value of the following distance reward.
[0039] In a second aspect, a smart device is provided, comprising:
[0040] At least one processor;
[0041] And, a memory communicatively connected to the at least one processor;
[0042] The memory stores a computer program, which, when executed by the at least one processor, implements the following distance acquisition method described in any of the above technical solutions.
[0043] In a third aspect, a vehicle is provided, the vehicle including the intelligent device described in the above-described technical solution.
[0044] The above-mentioned technical solutions of this application have at least one or more of the following beneficial effects: by distinguishing the following distances corresponding to different road scenarios and continuously iterating and optimizing each following distance during the reinforcement learning process of the autonomous driving agent, and continuously feeding it back into the reinforcement learning process, the autonomous driving agent can be guided to learn following strategies that are more in line with human driving habits in different road scenarios, reduce unnecessary conservative following behavior, and avoid dangerous close following, thereby achieving a dynamic balance between safety and traffic efficiency in complex road scenarios, while also taking into account the comfort of drivers and passengers. Attached Figure Description
[0045] The disclosure of this application will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0046] Figure 1 This is a schematic diagram of a reinforcement learning architecture for an autonomous driving system according to an embodiment of this application.
[0047] Figure 2 This is a schematic flowchart of the main steps of a following distance acquisition method according to an embodiment of this application.
[0048] Figure 3 This is a detailed flowchart illustrating step S202 according to an embodiment of this application.
[0049] Figure 4 This is a detailed flowchart illustrating step S203 according to an embodiment of this application.
[0050] Figure 5 This is a schematic diagram of the main structure of a smart device according to an embodiment of this application. Detailed Implementation
[0051] Some embodiments of this application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this application and are not intended to limit the scope of protection of this application.
[0052] In the description of this application, "module" and "processor" can include hardware, software, or a combination of both. A module can include hardware circuitry, various suitable sensors, communication ports, memory, and may also include software components, such as program code, or a combination of software and hardware. A processor can be a central processing unit, microprocessor, image processor, digital signal processor, or any other suitable processor. The processor has data and / or signal processing capabilities. The processor can be implemented in software, in hardware, or a combination of both. Computer-readable storage media includes any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B. The terms "at least one A or B" or "at least one of A and B" have a similar meaning to "A and / or B" and can include only A, only B, or A and B. The singular terms "a" or "this" can also include plural forms.
[0053] First, please refer to the appendix. Figure 1 , Figure 1 This is a schematic diagram of a reinforcement learning architecture for an autonomous driving system according to an embodiment of this application. Figure 1 As shown, the reinforcement learning architecture of the autonomous driving system in this application includes a reinforcement learning module and a following distance acquisition module.
[0054] This application does not limit the architecture of the autonomous driving agent in the reinforcement learning module. As an example, an autonomous driving agent can be developed independently based on the Vision-Language-Action (VLA) model, or an autonomous driving agent can be built on the basis of an open-source autonomous driving model, such as an autonomous driving agent built based on the OpenDriveVLA model.
[0055] This application does not limit the reinforcement learning (RL) algorithm used. As an example, Q-learning algorithm, proximal policy optimization (PPO) algorithm, etc. can be used. Those skilled in the art can select the appropriate reinforcement learning algorithm according to the architecture of the autonomous driving agent.
[0056] The following distance acquisition module is used to calculate the following distance reward in real time based on the trajectory data of the autonomous driving agent during the reinforcement learning process, and to update the following distance periodically based on the following distance reward and trajectory data.
[0057] In one embodiment, the trajectory data is a data sequence τ:
[0058]
[0059] Where s represents the state, a represents the action, and r represents the reward. The state s includes information such as the position of the vehicle, the position of the vehicle in front, the speed of the vehicle, lane lines, and obstacles. The action a includes information such as longitudinal acceleration (acceleration along the direction of vehicle travel) and steering. The reward r is the immediate feedback obtained from the environment after the autonomous driving agent performs action a, which is used to guide the learning method of the autonomous driving agent.
[0060] Next reading Figure 2 and combined Figure 1 This application explains the method for obtaining following distance. Figure 2 This is a schematic flowchart illustrating the main steps of a following distance acquisition method according to an embodiment of this application. The following distance (target following distance) acquisition method in this embodiment includes:
[0061] Step S201: Obtain the road scene;
[0062] Step S202: Based on the road scenario, determine the initial value of the following distance;
[0063] Step S203: Update the following distance based on the trajectory data during the reinforcement learning process;
[0064] Step S204: Based on the updated following distance, continue to perform reinforcement learning for the autonomous driving system.
[0065] In step 201, the road scene can be directly obtained through the autonomous driving agent. That is, when the autonomous driving agent processes multimodal input data (such as image sensor data, LiDAR sensor data, voice prompt data, high-precision map data, etc.) to generate the feature codes required for autonomous driving, it also identifies the type of road scene.
[0066] In another embodiment, road scenes can also be identified using a separate scene recognition model. As an example, a multimodal scene recognition model can be built based on the Transformer model. By using input image sensor data, LiDAR sensor data, voice prompts (such as weather conditions), high-precision map data, etc., information such as road type, traffic density, weather conditions, lighting conditions, and vehicle speed can be obtained, thereby determining the type of road scene.
[0067] The road scene information includes scene type and driving difficulty coefficient corresponding to each scene type (the larger the value of the driving difficulty coefficient, the more complex the road scene and the more difficult the driving control). The number of scene types can be set by those skilled in the art according to the actual situation.
[0068] In one embodiment, the driving difficulty coefficient can be set by combining the vehicle speed to distinguish the scene type. As an example, the speed level includes high speed, medium speed and low speed. When the vehicle speed is greater than 80 km / h, the speed level is high speed; when the vehicle speed is less than or equal to 80 km / h but greater than 40 km / h, the speed level is medium speed; when the vehicle speed is less than or equal to 40 km / h, the speed level is low speed.
[0069] Regarding driving difficulty coefficients, when weather conditions and / or traffic density are the same, the higher the speed level, the greater the driving difficulty coefficient corresponding to the road scenario. When the speed level is the same, the more complex the weather conditions, lighting conditions, and traffic density, the greater the driving difficulty coefficient corresponding to the road scenario. For example, in two road scenarios, vehicle 1 and vehicle 2 are both at medium speed. Vehicle 1 is driven in a sunny daytime, and vehicle 2 is driven in a rainy nighttime. In this case, the driving difficulty coefficient of vehicle 1's road scenario (scenario type: medium speed simple scenario) will be less than the driving difficulty coefficient of vehicle 2's road scenario (scenario type: medium speed difficult scenario).
[0070] It should be noted that, in other embodiments, those skilled in the art may also set the number of speed levels to other values, or set different values for the driving difficulty coefficient, etc. Without departing from the principles of this application, those skilled in the art may make equivalent changes or substitutions to the number of scene types, the numerical range of the driving difficulty coefficient, and other related technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of this application.
[0071] Next, combine Figure 3 Explain the specific implementation method of step S202. Figure 3 This is a detailed flowchart illustrating step S202 according to an embodiment of this application.
[0072] In step S2021, historical following distance data corresponding to the road scene is obtained, including historical following distance data of human drivers (following distance of the actual trajectory of human drivers) or historical following distance data of autonomous driving (following distance of the autonomous driving trajectory of the vehicle controlled by the autonomous driving agent).
[0073] In step S2022, firstly, based on preset quantiles, the historical following data time intervals are filtered to obtain valid historical following data time intervals, thereby eliminating some data with abnormal values. For example, the historical following data time intervals can be filtered by quartiles.
[0074] Then, based on valid historical following distance data, the baseline value of following distance is obtained by exponential moving average method, which can further reduce the impact of abnormal data on the calculation results and make the calculation results more stable and reliable.
[0075] In step S2023, considering the safety of autonomous vehicles, when the following distance baseline value is obtained based on the historical following distance data of human drivers, that is, the historical following distance data is the historical following distance data of human drivers, the initial value of the following distance needs to be recalculated according to the driving difficulty coefficient of the road scenario and the following distance baseline value.
[0076] Specifically, based on the driving difficulty coefficient of the road scenario and the preset first percentage coefficient, the following distance baseline value is expanded to obtain the initial value of the following distance threshold. For example, if the first percentage coefficient is 150%, then the initial value of the following distance threshold = the following distance baseline value * 150%.
[0077] Based on the driving difficulty coefficient of the road scenario, a preset second percentage coefficient is used to reduce the initial value of the following distance threshold, thus obtaining the initial value of the lower limit of the following distance interval. For example, if the second percentage coefficient is 80%, then the initial value of the lower limit of the following distance interval = (the baseline value of the following distance * 150%) * 80%.
[0078] Based on the third percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the initial value of the following distance threshold is expanded to obtain the initial value of the upper limit of the following distance interval. For example, if the third percentage coefficient is 120%, then the initial value of the upper limit of the following distance interval = (the baseline value of the following distance * 150%) * 120%.
[0079] At this point, an allowable following distance range is determined by the initial value of the lower limit of the following distance range and the initial value of the upper limit of the following distance range [initial value of the lower limit of the following distance range, initial value of the upper limit of the following distance range].
[0080] It should be noted that after the autonomous driving system begins intensive training, the initial value of the following distance threshold is updated to the following distance threshold, the initial value of the upper limit of the following distance interval is updated to the upper limit of the following distance interval, and the initial value of the lower limit of the following distance interval is updated to the lower limit of the following distance interval.
[0081] In step S203, the following distance can be updated based solely on the actual following distance and the following distance bonus; or it can be updated solely based on the actual following distance and longitudinal acceleration; or the two methods can be combined to update the following distance using different update cycles.
[0082] Continue reading Figure 4 ,pass Figure 4 Explain the specific implementation method of step S203. Figure 4 This is a detailed flowchart illustrating step S203 according to an embodiment of this application.
[0083] Each time an autonomous driving agent completes a reasoning operation, it generates a set of trajectory data and receives a following distance reward accordingly.
[0084] Specifically, when the actual following distance (which can be calculated based on the vehicle's position, the preceding vehicle's position, and the vehicle's speed in state 's') is less than the upper limit of the following distance interval and greater than or equal to the following distance threshold, meaning the actual following distance is within the range of [following distance threshold, following distance interval upper limit), it indicates that the following distance setting is reasonable and the vehicle is within the allowable following distance interval. Therefore, the following distance reward is positive. Furthermore, the smaller the absolute value of the difference between the actual following distance and the following distance threshold, the closer the autonomous vehicle's position is to the position corresponding to the following distance threshold, indicating a smaller error in the autonomous vehicle's following control, and correspondingly, a larger following distance reward.
[0085] When the actual following distance is greater than the upper limit of the following distance range, the distance between the vehicle and the vehicle in front is too far, the following behavior is conservative, and the traffic efficiency is low. The following distance bonus in this area can be set to 0.
[0086] When the actual following distance is less than the following distance threshold and greater than or equal to the lower limit of the following distance interval, that is, the actual following distance is within the range of [lower limit of following distance interval, following distance threshold), it also indicates that the following distance setting is reasonable and the vehicle is within the allowable following distance interval. Therefore, the following distance bonus value is positive.
[0087] Furthermore, the smaller the absolute value of the difference between the actual following distance and the following distance threshold, that is, the closer the position of the autonomous vehicle is to the position corresponding to the following distance threshold, the smaller the error of the autonomous vehicle following control, and the higher the score of the following distance reward.
[0088] When the actual following distance is less than the lower limit of the following distance range, the autonomous vehicle is already quite close to the vehicle in front, exceeding the allowable following distance range, which increases the driving safety hazard. The following distance setting may be unreasonable, so the following distance bonus value is negative.
[0089] Furthermore, the larger the absolute value of the difference between the actual following distance and the following distance threshold, the closer the vehicle is to the vehicle in front, and the greater the risk of collision. In this case, the larger the absolute value of the following distance bonus, the higher the negative score (the higher the penalty).
[0090] In one embodiment, within the multiple intervals defined by the lower limit of the following distance interval, the following distance threshold, and the upper limit of the following distance interval, a linear function can be used to represent the correspondence between the following distance reward value and the actual following distance within each interval.
[0091] For example, the vertical axis can be set as the following distance reward value, and the horizontal axis as the actual following distance. The following distance reward value corresponding to the upper limit of the following distance interval is 0, and the following distance reward value corresponding to the following distance threshold is 1. These two points can determine a straight line, and the following distance reward value within the interval [following distance interval lower limit, following distance threshold] can be obtained through this line. In other embodiments, a cosine function, step function, etc., can also be selected to establish the correspondence between the following distance reward value and the actual following distance.
[0092] like Figure 4 As shown, when the timer with a duration of one hour ends, the average distance between vehicles within the preset first time range is obtained, along with the average reward for following distance bonuses. For example, the average distance between vehicles within the range of [lower limit of following distance interval, upper limit of following distance interval] is calculated to obtain the average distance, and the average reward for all following distance bonuses is calculated to obtain the average reward.
[0093] When the mean of the reward statistics is less than 0, it indicates that the following behavior is relatively aggressive, the vehicle is too close to the vehicle in front, and there is a risk of collision. In this case, the following distance threshold can be increased to keep the vehicle away from the vehicle in front, thereby improving the safety of the autonomous vehicle. Furthermore, the larger the absolute value of the mean of the reward statistics (the closer the vehicle is to the vehicle in front, the higher the risk of collision), the greater the increase in the following distance threshold.
[0094] When the average reward is greater than 0, less than the first following distance reward threshold, and the average distance reward is greater than the following distance threshold, it indicates that the vehicle was far from the vehicle in front in the previous time period, its following behavior was relatively conservative, and its traffic efficiency was low. In this case, the following distance threshold can be reduced to guide the autonomous driving agent to adopt a more aggressive following strategy and improve the vehicle's traffic efficiency.
[0095] When the timer with a duration of the second time ends, the root mean square value of the longitudinal weighted acceleration corresponding to the longitudinal acceleration within the preset second time range is calculated, and the following distance is judged from the perspective of driving smoothness to determine whether it needs to be adjusted.
[0096] According to the ISO 2631 standard, when the weighted root mean square value of acceleration is between (0.5 m / s², 1.0 m / s²), people will experience "slight discomfort"; when the weighted root mean square value of acceleration is greater than 1.0 m / s², people will feel uncomfortable. The first weighted root mean square threshold is set to 0.5 m / s², and the second weighted root mean square threshold is set to 1.0 m / s².
[0097] When the longitudinal weighted root mean square value of acceleration is greater than the first weighted root mean square threshold and less than the second weighted root mean square threshold, the occupants of the vehicle will feel slightly uncomfortable. In this case, the upper limit of the following distance interval can be increased and / or the lower limit of the following distance interval can be decreased to reduce braking and acceleration actions and improve the comfort of the driver and passengers.
[0098] When the root mean square value of longitudinal weighted acceleration is greater than the root mean square threshold of second weighted acceleration, people will feel uncomfortable. This indicates that the following distance threshold may be too small, and there are too many acceleration and deceleration actions during the following process. At this time, the following distance threshold can be increased in order to reduce the vehicle's speed change actions.
[0099] It should be noted that in the embodiments of this application, the value at the second time is greater than the value at the first time. This update strategy, which combines different evaluation conditions for short-term and long-term periods, is more in line with the actual situation of autonomous driving systems (the following distance reward has strong timeliness, while the root mean square value of longitudinal weighted acceleration usually does not need strong timeliness, but rather aims to make the evaluation more accurate); on the other hand, it takes into account both the comfort of driving and riding (root mean square value of longitudinal weighted acceleration) and the safety of following (following distance reward).
[0100] It should be noted that the detection of road scenes is also carried out in real time. When the road scene changes, step S202 will be re-executed to obtain the initial value of the following distance for the new road scene, and the following distance currently being used by the reinforcement learning of the autonomous driving system will be updated. Reinforcement learning will continue to be carried out based on the updated following distance.
[0101] After step S203 is completed, the updated following distance is obtained. In step S204, the updated following distance is used to update the following distance data in the autonomous driving agent, such as the control decision module and the environmental perception module, and the autonomous driving system continues to perform reinforcement learning based on the updated following distance.
[0102] Through the aforementioned iterative reinforcement learning process, autonomous driving agents can learn following strategies that are more in line with human driving habits in different road scenarios. This can significantly reduce unnecessary conservative following behavior and avoid dangerous close following, achieving a dynamic balance between safety and efficiency in complex road scenarios.
[0103] Another aspect of this application provides a smart device.
[0104] In one embodiment of a smart device according to this application, the smart device may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program, which, when executed by the at least one processor, implements the following distance acquisition method described in any of the above embodiments. The smart device described in this application may include driving equipment, smart vehicles, robots, and other devices. See appendix. Figure 5 , Figure 5 The example illustrates a smart device 5 including a memory 51 and a processor 52 connected via a bus.
[0105] In some embodiments of this application, the smart device 5 may further include at least one sensor for sensing information; for example, the sensor is an image sensor. The sensor is communicatively connected to any type of processor mentioned in this application. Optionally, the following distance acquisition method may also include an autonomous driving system for guiding the smart device to drive autonomously or assisting in driving. The processor communicates with the sensor and / or the autonomous driving system to perform the following distance acquisition method described in any of the above embodiments.
[0106] Another aspect of this application provides a vehicle.
[0107] The vehicle includes the intelligent device described in the above embodiments. As an example, the vehicle is a new energy vehicle.
[0108] It should be noted that although the steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of this application, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders. These adjusted solutions are equivalent to the technical solutions described in this application and therefore will also fall within the protection scope of this application.
[0109] Those skilled in the art will understand that all or part of the processes in the method of the above-described embodiment can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0110] The technical solution of this application has been described above with reference to one embodiment shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. Without departing from the principles of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of this application.
Claims
1. A method for obtaining following distance, characterized in that, The method, applied to reinforcement learning in autonomous driving systems, includes: Obtain the road scene; Based on the road scenario, determine the initial value of the target following distance; Based on the trajectory data during the reinforcement learning process, the target vehicle following distance is updated; Based on the updated target following distance, the autonomous driving system continues to undergo reinforcement learning. The trajectory data includes the actual following distance and longitudinal acceleration, and the target following distance includes a following distance threshold, an upper limit of the following distance interval, and a lower limit of the following distance interval. The process of updating the target vehicle following distance based on trajectory data during reinforcement learning includes: Based on the actual following distance and the target following distance, a following distance reward is obtained. Based on the actual following distance and the following distance bonus, the target following distance is updated, including: Obtain the average time distance of the actual following distance within a preset first time range, and the average reward of the following distance bonus. When the average reward statistics are less than 0, the following distance threshold is increased, wherein the larger the absolute value of the average reward statistics, the larger the increase in the following distance threshold. When the average reward statistics are less than the first following distance reward threshold, and the average distance statistics are greater than the following distance threshold, the following distance threshold is reduced.
2. The method for obtaining following distance according to claim 1, characterized in that, The initial value for determining the target following distance based on the road scenario includes: Obtain historical following distance data corresponding to the road scene, wherein the historical following distance data is either human-driven historical following distance data or self-driving historical following distance data. Based on the historical following distance data, obtain the baseline value of the following distance for this road scenario; Based on the road scene and the following distance baseline value, an initial value of the target following distance is determined, wherein the initial value of the target following distance includes an initial value of the following distance threshold, an initial value of the lower limit of the following distance interval, and an initial value of the upper limit of the following distance interval. The initial value of the lower limit of the following distance interval is less than the following distance threshold, and the initial value of the upper limit of the following distance interval is greater than the following distance threshold.
3. The following distance acquisition method according to claim 2, characterized in that, The road scenario includes a driving difficulty coefficient. When the historical following distance data is the same as the driver's historical following distance data, the method for determining the initial value of the target following distance includes: Based on the first percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the following distance baseline value is expanded to obtain the initial value of the following distance threshold. Based on the second percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the initial value of the following distance threshold is reduced to obtain the initial value of the lower limit of the following distance interval. Based on the third percentage coefficient corresponding to the driving difficulty coefficient of the road scenario, the initial value of the following distance threshold is expanded to obtain the initial value of the upper limit of the following distance interval.
4. The following distance acquisition method according to claim 2, characterized in that, The method for obtaining the following distance baseline value includes: Based on a preset quantile, the historical following distance data is filtered to obtain valid historical following distance data. Based on the effective historical following distance data, the baseline value of the following distance is obtained by using the exponential moving average method.
5. The method for obtaining following distance according to any one of claims 1 to 4, characterized in that, The method of updating the target vehicle following distance based on trajectory data during the reinforcement learning process further includes: Based on the actual following distance and the longitudinal acceleration, the target following distance is updated, including: Based on the longitudinal acceleration within a preset second time range, obtain the root mean square value of the longitudinal weighted acceleration; When the longitudinal weighted root mean square value of acceleration is greater than the first weighted root mean square threshold and less than the second weighted root mean square threshold, the upper limit of the following distance interval is increased and / or the lower limit of the following distance interval is decreased. When the longitudinal weighted root mean square value of acceleration is greater than the second weighted root mean square threshold, the following distance threshold is increased.
6. The method for obtaining following distance according to claim 1, characterized in that, The method for obtaining the following distance bonus includes: When the actual following distance is less than the upper limit of the following distance range and greater than or equal to the following distance threshold, the following distance reward is positive. The smaller the absolute value of the difference between the actual following distance and the following distance threshold, the larger the following distance reward. When the actual following distance is less than the following distance threshold and is greater than or equal to the lower limit of the following distance range, the following distance reward is positive. The smaller the absolute value of the difference between the actual following distance and the following distance threshold, the larger the following distance reward. When the actual following distance is less than the lower limit of the following distance interval, the following distance reward is negative, and the larger the absolute value of the difference between the actual following distance and the following distance threshold, the larger the absolute value of the following distance reward.
7. A smart device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores a computer program, which, when executed by the at least one processor, implements the following distance acquisition method as described in any one of claims 1 to 6.
8. A vehicle, characterized in that, The vehicle includes the intelligent device as described in claim 7.