Vehicle speed control method, device, equipment and medium

By calculating the optimal expected speed and reward function to optimize traffic parameters, the problem of insufficient adaptability and generalization capabilities of traditional speed control algorithms in dynamic environments is solved, multi-objective optimization is achieved, and the safety and efficiency of the traffic system are improved.

CN116572951BActive Publication Date: 2025-08-26TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310639458.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2025-08-26
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

Traditional speed control algorithms are difficult to adapt to the ever-changing dynamic environment, and only pursue a single goal such as vehicle speed. When facing increasingly complex and diverse traffic scenarios, their adaptability and generalization capabilities are weak.

Method used

By collecting the current speed of the vehicle in front and the distance difference between it and the vehicle, the optimal expected speed is calculated, the reward function is adjusted in combination with the current driving environment, and the Jerk function is introduced to optimize driving parameters other than the vehicle speed, such as the degree of bumps, to achieve multi-objective optimization.

Benefits of technology

It effectively improves the safety, efficiency and sustainability of the transportation system, and meets the needs of intelligent transportation and autonomous driving for effective and reliable speed control algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116572951B_ABST
    Figure CN116572951B_ABST
Patent Text Reader

Abstract

The present application relates to a vehicle speed control method, apparatus, device, and medium, wherein the method comprises: collecting the current speed of the preceding vehicle, and calculating the optimal expected speed of the vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; adjusting a preset reward function according to the current driving environment of the vehicle to obtain an optimal reward function; while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, introducing a preset Jerk function, and combining the preset Jerk function and the optimal reward function to obtain a reward extension function, and utilizing the reward extension function to optimize at least one driving parameter of the vehicle other than the speed, so as to control the vehicle's driving according to the at least one optimized driving parameter. This solves the problems that traditional speed control algorithms are difficult to adapt to the ever-changing dynamic environment, and only pursue a single goal such as vehicle speed, and have weak adaptability and generalization capabilities when facing increasingly complex and diverse traffic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of transportation and vehicle technology, and in particular to a vehicle speed control method, device, equipment and medium. Background Art

[0002] Intelligent Transportation Systems (ITS) aim to improve the efficiency and safety of transportation systems by using advanced technologies such as sensors, communication systems, and control algorithms. Vehicle speed control is a key aspect of ITS because it plays an important role in improving the safety, efficiency, and sustainability of transportation systems. With the increasing popularity of autonomous vehicles and the spread of connected and intelligent transportation technologies, the need for effective and reliable speed control algorithms is becoming increasingly urgent.

[0003] Traditional speed control algorithms have been widely used in various ITS applications and have shown good performance in some scenarios, but the existing technologies still have some defects.

[0004] First, traditional speed control algorithms are usually unable to adapt to various factors such as changing road conditions, traffic flow and weather in real time, which greatly limits their performance in complex dynamic environments, especially when the system dynamics are uncertain or difficult to model accurately. Second, existing speed control algorithms usually focus on a single goal, such as pursuing optimal speed, while ignoring other important parameters such as driving safety, fuel efficiency and driving comfort. This may lead to limited performance in practical applications and even a poor driving experience. In addition, traditional speed control algorithms usually do not have the ability to learn and extract useful information from data, and may face difficulties in processing large-scale actual driving data, thus limiting their adaptability and generalization capabilities in the face of increasingly complex and diverse traffic scenarios.

[0005] In summary, traditional speed control algorithms are difficult to adapt to the ever-changing dynamic environment and only pursue a single goal such as vehicle speed. When faced with increasingly complex and diverse traffic scenarios, their adaptability and generalization capabilities are weak and need to be urgently addressed. Summary of the Invention

[0006] The present application provides a vehicle speed control method, apparatus, device, and medium to address the problems that traditional speed control algorithms have difficulty adapting to ever-changing dynamic environments, only pursue a single goal such as vehicle speed, and have weak adaptability and generalization capabilities when facing increasingly complex and diverse traffic scenarios.

[0007] The first aspect of the present application provides a vehicle speed control method, comprising the following steps: collecting the current speed of a preceding vehicle, and calculating the optimal expected speed of the vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; adjusting the preset reward function according to the current driving environment of the vehicle to obtain an optimal reward function; while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, introducing a preset Jerk function, and combining the preset Jerk function and the optimal reward function to obtain a reward extension function, and utilizing the reward extension function to optimize at least one driving parameter of the vehicle other than the vehicle speed, so as to control the vehicle driving based on the optimized at least one driving parameter when the current speed of the vehicle reaches the optimal expected speed.

[0008] Optionally, in one embodiment of the present application, before adjusting the preset reward function according to the current driving environment of the vehicle, it also includes: combining different reward functions based on a preset combination strategy to obtain at least one combined reward function; performing simulation testing based on the at least one combined reward function to obtain the reward changes of each combined reward function in the at least one combined reward function, and obtaining the preset reward function based on the reward changes.

[0009] Optionally, in one embodiment of the present application, adjusting the preset reward function according to the current driving environment of the vehicle to obtain the optimal reward function includes: collecting the vehicle's surrounding environment information, and judging the current vehicle environmental conditions based on the vehicle's surrounding environment information; uploading the current vehicle environmental conditions to the automatic driving system, and the automatic driving system obtains the optimal reward function through a preset control algorithm.

[0010] Optionally, in one embodiment of the present application, the calculation formula for the optimal expected speed of the vehicle is as follows:

[0011]

[0012] Among them, PRT is the driver's reaction time, a 3. The partial braking acceleration is -3.4m / s 2 , a 7. The full braking acceleration is -7.5m / s 2 , v1(t-1) is the speed of the preceding vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

[0013] Optionally, in one embodiment of the present application, the calculation formula of the reward expansion function is as follows:

[0014] r m =r1+r2

[0015] Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

[0016] The second aspect of the present application provides a vehicle speed control device, including: an acquisition module, used to acquire the current speed of the preceding vehicle, and calculate the optimal expected speed of the vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; an adjustment module, used to adjust the preset reward function according to the current driving environment of the vehicle to obtain an optimal reward function; an optimization module, used to introduce a preset Jerk function while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, and combine the preset Jerk function and the optimal reward function to obtain a reward extension function, and use the reward extension function to optimize at least one driving parameter of the vehicle other than the speed, so as to control the vehicle driving based on the optimized at least one driving parameter when the current speed of the vehicle reaches the optimal expected speed.

[0017] Optionally, in one embodiment of the present application, it also includes: a combination module, which is used to combine different reward functions based on a preset combination strategy before adjusting the preset reward function according to the current driving environment of the vehicle to obtain at least one combined reward function; a testing module, which is used to perform simulation testing based on the at least one combined reward function to obtain the reward changes of each combined reward function in the at least one combined reward function, and obtain the preset reward function based on the reward changes.

[0018] Optionally, in one embodiment of the present application, the adjustment module includes: a judgment unit, used to collect the vehicle's surrounding environment information and judge the current vehicle environmental working condition based on the vehicle's surrounding environment information; an uploading unit, used to upload the current vehicle environmental working condition to the automatic driving system, and the automatic driving system obtains the optimal reward function through a preset control algorithm.

[0019] Optionally, in one embodiment of the present application, the calculation formula for the optimal expected speed of the vehicle is as follows:

[0020]

[0021] Among them, PRT is the driver's reaction time, a 3. The partial braking acceleration is -3.4m / s 2 , a 7. The full braking acceleration is -7.5m / s 2 , v1(t-1) is the speed of the preceding vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

[0022] Optionally, in one embodiment of the present application, the calculation formula of the reward expansion function is as follows:

[0023] r m =r1+r2

[0024] Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

[0025] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the vehicle speed control method as described in the above embodiment.

[0026] A fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program that implements the above vehicle speed control method when executed by a processor.

[0027] Therefore, the embodiments of the present application have the following beneficial effects:

[0028] The embodiments of the present application can collect the current speed of the preceding vehicle and calculate the optimal expected speed of the vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; adjust the preset reward function according to the current driving environment of the vehicle to obtain the optimal reward function; while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, introduce a preset Jerk function, and combine the preset Jerk function and the optimal reward function to obtain a reward extension function, and use the reward extension function to optimize at least one driving parameter of the vehicle other than the speed, so that when the current speed of the vehicle reaches the optimal expected speed, the vehicle driving is controlled based on the optimized at least one driving parameter, thereby meeting the needs of intelligent transportation and autonomous driving for effective and reliable speed control algorithms, and effectively improving the safety, efficiency and sustainability of the transportation system. In this way, the problems of traditional speed control algorithms being difficult to adapt to the ever-changing dynamic environment and only pursuing a single goal such as vehicle speed are solved, and their adaptability and generalization capabilities are weak when facing increasingly complex and diverse traffic scenarios.

[0029] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0031] Figure 1 This is a flow chart of a vehicle speed control method provided according to an embodiment of the present application;

[0032] Figure 2 A schematic diagram of a reward function in a speed optimization process provided by one embodiment of the present application;

[0033] Figure 3 is an exemplary diagram of a vehicle speed control device according to an embodiment of the present application;

[0034] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0035] Among them, 10 is a vehicle speed control device, 100 is a collection module, 200 is an adjustment module, 300 is an optimization module, 401 is a memory, 402 is a processor, and 403 is a communication interface. DETAILED DESCRIPTION

[0036] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0037] The following describes the vehicle speed control method, device, equipment and medium of the embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a vehicle speed control method, in which the current speed of the preceding vehicle is collected and the optimal expected speed of the vehicle is calculated based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; a preset reward function is adjusted according to the current driving environment of the vehicle to obtain an optimal reward function; while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, a preset Jerk function is introduced, and a reward extension function is obtained by combining the preset Jerk function and the optimal reward function. The reward extension function is used to optimize at least one driving parameter of the vehicle other than the speed, so that when the current speed of the vehicle reaches the optimal expected speed, the vehicle is controlled based on the optimized at least one driving parameter, thereby meeting the needs of intelligent transportation and autonomous driving for effective and reliable speed control algorithms, and effectively improving the safety, efficiency and sustainability of the transportation system. Thus, the present application solves the problems that traditional speed control algorithms are difficult to adapt to the ever-changing dynamic environment, and only pursue a single goal such as vehicle speed, and have weak adaptability and generalization capabilities when facing increasingly complex and diverse traffic scenarios.

[0038] Specifically, Figure 1 This is a flow chart of a vehicle speed control method provided in an embodiment of the present application.

[0039] like Figure 1As shown, the vehicle speed control method includes the following steps:

[0040] In step S101 , the current speed of the preceding vehicle is collected, and the optimal expected speed of the vehicle is calculated based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range.

[0041] The embodiments of the present application can detect whether there is a vehicle in front of the vehicle through radar sensors, ultrasonic sensors, and visual sensors. When there is a vehicle in front of the vehicle, the embodiments of the present application can collect the speed information of the front vehicle and the distance information between the front vehicle and the vehicle within a certain time interval through relevant sensor equipment such as vehicle speed sensors, thereby providing a reliable data basis for calculating the optimal expected speed of the vehicle.

[0042] Optionally, in one embodiment of the present application, the calculation formula for the optimal expected speed of the vehicle is as follows:

[0043]

[0044] Among them, PRT is the driver's reaction time, a 3. The partial braking acceleration is -3.4m / s 2 , a 7. The full braking acceleration is -7.5m / s 2 , v1(t-1) is the speed of the leading vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

[0045] It should be noted that, in the embodiment of the present application, the expected optimal driving speed of the vehicle is determined based on avoiding potential emergency braking of the preceding vehicle. e It can be calculated as follows:

[0046]

[0047] Among them, PRT is the driver's reaction time, a 3. The partial braking acceleration is -3.4m / s 2 , a 7. The full braking acceleration is -7.5m / s 2 , v1(-1) is the speed of the leading vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

[0048] Therefore, the embodiment of the present application calculates the optimal expected speed of the vehicle to guide the subsequent vehicles to slow down, ensuring that the vehicle maintains a safe distance from the vehicle in front (maintaining a minimum gap of at least 1.5 meters) when the vehicle in front suddenly brakes using maximum deceleration to avoid collision, thereby ensuring the driving efficiency and safety of the autonomous driving vehicle during driving.

[0049] In step S102, the preset reward function is adjusted according to the current driving environment of the vehicle to obtain the optimal reward function.

[0050] After obtaining the optimal expected speed of the vehicle, the embodiment of the present application can further adjust the reward function according to the current vehicle driving environment to optimize the current vehicle speed so that it continuously approaches the above-mentioned optimal expected speed, thereby achieving optimization of the vehicle speed.

[0051] Optionally, in one embodiment of the present application, before adjusting the preset reward function according to the current driving environment of the vehicle, it also includes: combining different reward functions based on a preset combination strategy to obtain at least one combined reward function; performing simulation tests based on at least one combined reward function to obtain the reward changes of each combined reward function in the at least one combined reward function, and obtaining the preset reward function based on the reward changes.

[0052] It should be noted that in the embodiments of the present application, the reward function used in the DRL (Deep Reinforcement Learning) algorithm can be to make the subsequent vehicle travel as close to the optimal expected speed of the vehicle as possible, and the highest reward is obtained when the subsequent vehicle reaches the optimal expected speed.

[0053] In order to select the best reward function form, the embodiment of the present application can divide the reward function form into α has three parts, and different forms of reward functions in speed optimization are as follows Figure 2 shown.

[0054] The embodiments of the present application conduct simulation tests by trying different reward function combinations, testing the reward changes of each reward function combination during the training process, and performing multiple training on all verification sample sets. Among them, the convex function performs best among the three optional functions, allowing the autonomous driving system to continue learning better speed control strategies until it approaches the optimal expected speed; the other two functions (i.e., linear function and concave function) will prematurely find a suboptimal speed control strategy. The reward function for optimizing speed can be expressed as:

[0055]

[0056] in:

[0057]

[0058]

[0059] r min1 =-α·v e

[0060] rmin2 =-α·(v m -v e )

[0061] Where, v m is the upper speed limit, r min is the minimum reward function, r max is the maximum reward function, α is the coefficient for adjusting negative rewards, which can be selected between 0, -0.5, -1, -1.5, -2, -3 and -5. The larger the α value, the higher the accuracy achieved by the model, which can achieve the effect of optimizing a single objective and then achieve the optimization effect of multiple objectives.

[0062] Optionally, in one embodiment of the present application, the preset reward function is adjusted according to the current driving environment of the vehicle to obtain the optimal reward function, including: collecting the vehicle's surrounding environment information, and judging the current vehicle environmental conditions based on the vehicle's surrounding environment information; uploading the current vehicle environmental conditions to the automatic driving system, and the automatic driving system obtains the optimal reward function through a preset control algorithm.

[0063] As a feasible method, in the process of automatic driving, the vehicle in the embodiment of the present application can obtain the vehicle's current driving environment information through visual sensors and other equipment, such as the different road conditions of the vehicle, the traffic flow of the current road, and the weather conditions in which the vehicle is located, and judge the current vehicle's environmental conditions based on the collected environmental information. The current vehicle's environmental conditions can determine the current optimal reward function. When the vehicle faces a complex and changeable traffic environment with different road conditions, different traffic flows and different weather, the embodiment of the present application can use the reward function self-optimization method to enable the vehicle to judge the current environment around the vehicle based on the vehicle's own sensors during driving, and upload it to the automatic driving system through the communication system. The automatic driving system automatically selects the optimal reward function through the control algorithm, so that the vehicle speed can adapt to the complex and changeable environmental conditions.

[0064] In step S103, while using the preset reward function and the optimal expected speed of the vehicle to optimize the current speed of the vehicle, a preset Jerk function is introduced, and the preset Jerk function and the optimal reward function are combined to obtain a reward extension function. The reward extension function is used to optimize at least one driving parameter of the vehicle other than the speed, so that when the current speed of the vehicle reaches the optimal expected speed, the vehicle driving is controlled based on the at least one optimized driving parameter.

[0065] While obtaining the optimal reward function and optimizing the current speed of the vehicle using the above reward function and the optimal expected speed of the vehicle, the embodiments of the present application can further optimize multiple parameters other than the above-mentioned vehicle speed (such as |Jerk|, etc.) through a reward function expansion method, thereby effectively ensuring the effectiveness and reliability of the vehicle speed in autonomous driving.

[0066] Optionally, in one embodiment of the present application, the calculation formula of the reward expansion function is as follows:

[0067] r m =r1+r2

[0068] Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

[0069] It should be noted that when comparing the |Jerk| of the group of models whose speed was closest to the optimal expected speed in the validation set test results for each category with the historical data of drivers in that category, it was found that because the reward function did not take into account the optimization of vehicle control indicators other than speed, the DRL model based only on optimized speed performed poorly when using |Jerk| as a measurable indicator.

[0070] Therefore, in order to optimize |Jerk|, the embodiment of the present application introduces a Jerk-based function as a reward function, that is, a part of the reward expansion function, as shown in the following formula:

[0071]

[0072] Wherein, Jerk2(t) is the degree of jerkiness of the subsequent vehicle in the time interval t.

[0073] In general, the above reward expansion function can be obtained by weighted addition, as shown in the following formula:

[0074] r m =r1+r2

[0075] Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

[0076] In order to evaluate the impact of r2 on |Jerk|, the embodiment of the present application can use the reward function r m The DRL model was trained 10 times for each category in the training set. The results showed that the introduction of the |Jerk|-based reward function and the DRL-based vehicle speed control model greatly reduced the average |Jerk| in the validation set of all categories, which was less than the driver's historical data. At the same time, the DRL model for simulated speed and multi-objective optimization had similarities with the DRL model for single-objective optimization.

[0077] Therefore, the embodiments of the present application can effectively and reliably optimize multiple objectives and single objectives at the same time through the reward function design in the form of weighted addition, such as vehicle speed, the degree of vehicle jerk during driving, and vehicle collision avoidance during driving, so as to achieve a good balance between different objectives without sacrificing the performance of any single objective, providing a valuable reference for the application of DRL in practical problems and providing guidance for the development of more advanced and effective control algorithms for wide application.

[0078] According to the vehicle speed control method proposed in the embodiment of the present application, the current speed of the preceding vehicle is collected, and the optimal expected speed of the vehicle is calculated based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; the preset reward function is adjusted according to the current driving environment of the vehicle to obtain the optimal reward function; while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, a preset Jerk function is introduced, and a reward extension function is obtained by combining the preset Jerk function and the optimal reward function, and the reward extension function is used to optimize at least one driving parameter of the vehicle other than the speed, so that when the current speed of the vehicle reaches the optimal expected speed, the vehicle driving is controlled based on the at least one optimized driving parameter, thereby meeting the requirements of intelligent transportation and autonomous driving for effective and reliable speed control algorithms, and effectively improving the safety, efficiency and sustainability of the transportation system.

[0079] Next, a vehicle speed control device according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0080] Figure 3 It is a block diagram of a vehicle speed control device according to an embodiment of the present application.

[0081] like Figure 3 As shown, the vehicle speed control device 10 includes: a collection module 100 , an adjustment module 200 and an optimization module 300 .

[0082] The acquisition module 100 is used to acquire the current speed of the preceding vehicle and calculate the optimal expected speed of the vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range.

[0083] The adjustment module 200 is used to adjust the preset reward function according to the current driving environment of the vehicle to obtain the optimal reward function.

[0084] The optimization module 300 is used to introduce a preset Jerk function while optimizing the current speed of the vehicle using a preset reward function and the optimal expected speed of the vehicle, and to obtain a reward extension function by combining the preset Jerk function and the optimal reward function. The reward extension function is used to optimize at least one driving parameter of the vehicle other than the speed, so as to control the vehicle driving based on the at least one optimized driving parameter when the current speed of the vehicle reaches the optimal expected speed.

[0085] Optionally, in one embodiment of the present application, the vehicle speed control device 10 of the embodiment of the present application further includes: a combination module and a test module.

[0086] The combination module is configured to combine different reward functions based on a preset combination strategy before adjusting the preset reward function according to the current driving environment of the vehicle to obtain at least one combined reward function.

[0087] The testing module is used to perform simulation testing based on at least one combined reward function, obtain reward changes of each combined reward function in the at least one combined reward function, and obtain a preset reward function based on the reward changes.

[0088] Optionally, in one embodiment of the present application, the adjustment module 200 includes: a judgment unit and an uploading unit.

[0089] The judgment unit is used to collect the vehicle's surrounding environment information and judge the current vehicle environment condition based on the vehicle's surrounding environment information.

[0090] The upload unit is used to upload the current vehicle environment conditions to the autonomous driving system, and the autonomous driving system obtains the optimal reward function through a preset control algorithm.

[0091] Optionally, in one embodiment of the present application, the calculation formula for the optimal expected speed of the vehicle is as follows:

[0092]

[0093] Among them, PRT is the driver's reaction time, a 3. The partial braking acceleration is -3.4m / s 2 , a 7. The full braking acceleration is -7.5m / s 2 , v1(t-1) is the speed of the leading vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

[0094] Optionally, in one embodiment of the present application, the calculation formula of the reward expansion function is as follows:

[0095] r m =r1+r2

[0096] Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

[0097] It should be noted that the above explanation of the embodiment of the vehicle speed control method is also applicable to the vehicle speed control device of this embodiment, and will not be repeated here.

[0098] According to the vehicle speed control device proposed in the embodiment of the present application, the current speed of the preceding vehicle is collected, and the optimal expected speed of the vehicle is calculated based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; the preset reward function is adjusted according to the current driving environment of the vehicle to obtain the optimal reward function; while optimizing the current speed of the vehicle using the preset reward function and the optimal expected speed of the vehicle, a preset Jerk function is introduced, and a reward extension function is obtained by combining the preset Jerk function and the optimal reward function, and the reward extension function is used to optimize at least one driving parameter of the vehicle other than the speed, so that when the current speed of the vehicle reaches the optimal expected speed, the vehicle driving is controlled based on the at least one optimized driving parameter, thereby meeting the requirements of intelligent transportation and autonomous driving for effective and reliable speed control algorithms, and effectively improving the safety, efficiency and sustainability of the transportation system.

[0099] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0100] Memory 401 , processor 402 , and computer programs stored in the memory 401 and executable on the processor 402 .

[0101] When the processor 402 executes the program, the vehicle speed control method provided in the above embodiment is implemented.

[0102] Furthermore, the electronic device further includes:

[0103] The communication interface 403 is used for communication between the memory 401 and the processor 402 .

[0104] The memory 401 is used to store computer programs that can be run on the processor 402 .

[0105] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0106] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0107] Optionally, in a specific implementation, if the memory 401 , the processor 402 and the communication interface 403 are integrated on a chip, the memory 401 , the processor 402 and the communication interface 403 can communicate with each other through an internal interface.

[0108] The processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0109] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the above vehicle speed control method when executed by a processor.

[0110] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0111] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0112] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0113] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0114] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0115] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0116] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0117] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A vehicle speed control method, characterized in that: The following steps are involved: Collecting the current speed of the preceding vehicle and calculating the optimal expected speed of the vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the vehicle within a preset time range; Adjust the preset reward function according to the current driving environment of the vehicle to obtain the optimal reward function, and While optimizing the current speed of the host vehicle using a preset reward function and the optimal expected speed of the host vehicle, introducing a preset Jerk function, and combining the preset Jerk function with the optimal reward function to obtain a reward extension function, and using the reward extension function to optimize at least one driving parameter of the host vehicle other than the speed, so as to control the vehicle based on the optimized at least one driving parameter when the current speed of the host vehicle reaches the optimal expected speed; Before adjusting the preset reward function according to the current driving environment of the vehicle, the vehicle speed control method further includes: Combining different reward functions based on a preset combination strategy to obtain at least one combined reward function; Performing simulation testing according to the at least one combined reward function to obtain a reward change of each combined reward function in the at least one combined reward function, and obtaining the preset reward function based on the reward change; The calculation formula of the optimal expected speed of the vehicle is as follows: Among them, PRT is the driver's reaction time, a 3.4 The partial braking acceleration is -3.4m / s 2 , a 7.5 The full braking acceleration is -7.5m / s 2 , v1(t-1) is the speed of the preceding vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

2. The method according to claim 1, characterized in that The step of adjusting the preset reward function according to the current driving environment of the vehicle to obtain an optimal reward function includes: Collecting vehicle surrounding environment information, and determining the current vehicle environmental operating condition based on the vehicle surrounding environment information; The current vehicle environment condition is uploaded to the autonomous driving system, and the autonomous driving system obtains the optimal reward function through a preset control algorithm.

3. The method according to claim 1, characterized in that The calculation formula of the reward expansion function is as follows: r m =r1+r2 Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

4. A vehicle speed control device, characterized in that: include: an acquisition module, configured to acquire the current speed of the preceding vehicle and calculate the optimal expected speed of the own vehicle based on the current speed of the preceding vehicle and the distance difference between the preceding vehicle and the own vehicle within a preset time range; An adjustment module, configured to adjust the preset reward function according to the current driving environment of the vehicle to obtain an optimal reward function, and an optimization module, configured to, while optimizing the current speed of the host vehicle using a preset reward function and the optimal expected speed of the host vehicle, introduce a preset Jerk function, combine the preset Jerk function and the optimal reward function to obtain a reward extension function, and utilize the reward extension function to optimize at least one driving parameter of the host vehicle other than the speed, so as to control the vehicle based on the optimized at least one driving parameter when the current speed of the host vehicle reaches the optimal expected speed; Wherein, the vehicle speed control device further includes: a combining module, configured to combine different reward functions based on a preset combination strategy before adjusting the preset reward function according to the current driving environment of the vehicle to obtain at least one combined reward function; a testing module, configured to perform a simulation test based on the at least one combined reward function, obtain a reward change of each combined reward function in the at least one combined reward function, and obtain the preset reward function based on the reward change; The calculation formula of the optimal expected speed of the vehicle is as follows: Among them, PRT is the driver's reaction time, a 3.4 The partial braking acceleration is -3.4m / s 2 , a 7.5 The full braking acceleration is -7.5m / s 2 , v1(t-1) is the speed of the preceding vehicle, and d(t-1) is the distance difference between the two vehicles in the time interval t-1.

5. The device according to claim 4, characterized in that The adjustment module includes: A judgment unit, configured to collect information about the surrounding environment of the vehicle and judge the current vehicle environmental condition based on the information about the surrounding environment of the vehicle; An uploading unit is used to upload the current vehicle environment condition to the automatic driving system, and the automatic driving system obtains the optimal reward function through a preset control algorithm.

6. The device according to claim 4, characterized in that The calculation formula of the reward expansion function is as follows: r m =r1+r2 Among them, r1 is the optimal reward function, and r2 is the reward function based on the Jerk function.

7. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the vehicle speed control method according to any one of claims 1 to 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the vehicle speed control method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Speed control multi-target optimized car following algorithm of automatic driving vehicle

    CN109709956A

  • Self-adaptive cruise control method and system, computer and storage medium

    CN114179798A