Intelligent connected vehicle cooperative control method and system based on reinforcement learning strategy

CN122821768APending Publication Date: 2026-09-25RIVOTEK TECH (JIANGSU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611083557.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,上述技术方案在复杂的行车环境中易陷入次优解,且计算复杂度高,难以有效实现多目标协同优化

Benefits of technology

[0075]1、通过构建基于头车行驶速度的数据链,并动态调整中部的车辆和尾车的速度,确保整个车辆编队能够平滑、高效地加速或减速,具体而言,系统根据头车的行驶速度变化,通过数据链实时传输给中部和尾车,中部的车辆在行驶速度连续2秒大于参考速度时进行提速,直至达到最高速度,尾车随后跟进,这种动态调整机制减少了车辆间的速度差异,提高了行驶效率;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821768A_ABST
    Figure CN122821768A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent networked vehicle cooperation control method and system based on a reinforcement learning strategy, and belongs to the technical field of intelligent transportation, and comprises the following steps: S10: a data chain is constructed, and cooperation control is performed on a target vehicle platoon according to the data chain; the step of constructing the data chain comprises the following steps: the data chain is constructed based on a head vehicle in the target vehicle platoon, and relevant data of the head vehicle is collected; the relevant data comprises a driving speed of the head vehicle; the relevant data is collected at a given reference time, and the driving speed is collected at every 3 seconds as a reference time; the data chain based on the driving speed of the head vehicle is constructed, the speeds of the vehicles in the middle and the tail vehicle are dynamically adjusted, and multi-reference period analysis and reinforcement learning optimization are combined, so that smooth acceleration and deceleration of the vehicle platoon, accurate speed prediction and control are realized, the adaptability and robustness of the platoon are enhanced, rear-end collision accidents are effectively prevented, and the driving experience and comfort are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation technology, and in particular to a cooperative control method and system for intelligent connected vehicles based on reinforcement learning strategies. Background Technology

[0002] Currently, intelligent connected vehicles face numerous challenges in platoon control technology. Traditional platoon collaborative control primarily relies on direct communication between vehicles for command transmission. However, communication networks suffer from issues such as time delays, data packet loss, and out-of-order transmission, which can easily lead to accumulated and continuous following errors. This can result in frequent acceleration and deceleration of vehicles, platoon instability, and even traffic accidents such as platoon disbandment and collisions between vehicles.

[0003] Regarding this research, application CN202311431497.8 provides a cooperative control method for intelligent connected vehicles based on reinforcement learning strategies. This technical solution includes: considering the communication anomaly problem between the vehicle and the client, designing reward functions for two DDPG algorithms based on the control objective; the upper-level algorithm uses observations as input and, based on the reward maximization principle, trains two agents; each agent has different observations. Under normal communication conditions, agent 1 plans the reference acceleration for the vehicle; when a communication anomaly occurs, according to a switching mechanism, agent 1 switches to agent 2, which plans the reference acceleration. This agent switching solves the problem of unstable vehicle operation caused by the lack of observations due to communication anomalies, preventing the agent from planning the reference acceleration.

[0004] Another application, CN202510311155.5, provides a centralized control method for intelligent connected vehicle platooning based on deep reinforcement learning. This solution constructs a vehicle-road-cloud centralized communication network architecture, offloading onboard computing tasks to edge cloud nodes and combining it with the DDPG algorithm to achieve global optimization control. Vehicle nodes acquire status information through onboard sensors and transmit it to edge cloud nodes via roadside nodes. The edge cloud integrates multi-vehicle data collected by the RSU into global platoon status information, uses a pre-trained DDPG agent to generate control commands in real time, and broadcasts them to the platooning vehicles for execution via C-V2X communication. This solution reduces the load on onboard hardware through a centralized computing architecture, improving the real-time performance and safety of platooning control; and utilizes cloud computing power to support the global optimization of complex algorithms.

[0005] Vehicle platooning needs to cope with diverse road conditions and dynamic factors such as the behavior of traffic participants during operation. This leads to frequent speed adjustments by vehicles, requiring the handling of interactions between agents (such as platoon following and intersection passage) when multiple vehicles are cooperating. However, the above-mentioned technical solutions are prone to suboptimal solutions in complex driving environments and have high computational complexity, making it difficult to effectively achieve multi-objective cooperative optimization. Summary of the Invention

[0006] In view of the problems existing in the field of intelligent transportation technology, the present invention is proposed.

[0007] Therefore, one of the objectives of this invention is to provide a cooperative control method and system for intelligent connected vehicles based on reinforcement learning strategies. By constructing a data link based on the speed of the lead vehicle, the system dynamically adjusts the speeds of the middle and rear vehicles. Combined with multi-reference time period analysis and reinforcement learning optimization, it achieves smooth acceleration and deceleration, accurate speed prediction and control of vehicle formation, enhances formation adaptability and robustness, effectively prevents rear-end collisions, and improves driving experience and comfort.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] On the one hand, the present invention provides a cooperative control method for intelligent connected vehicles based on a reinforcement learning strategy, comprising the following steps:

[0010] S10: Construct a data link and perform cooperative control in the target vehicle platoon based on the data link. The steps of constructing the data link include:

[0011] The data chain is built based on the lead vehicle in the target vehicle platoon and collects relevant data from the lead vehicle.

[0012] The relevant data includes the speed of the lead vehicle. The relevant data is collected at a given reference time, including collecting the speed at 3-second intervals.

[0013] The change in the speed of the lead vehicle is obtained by taking at least three reference times as a reference period. The speed corresponding to the first reference time is taken as the reference speed. Based on the reference speed, if the speed shows an increasing trend in the second reference time, the highest speed in the third reference time is obtained according to this trend.

[0014] S20: If the driving speed is greater than the reference speed for 2 consecutive seconds, the vehicles in the middle of the target vehicle platoon will accelerate until the driving speed reaches the maximum speed.

[0015] S30: When the vehicles in the middle accelerate, the last vehicle in the target vehicle platoon also accelerates until the speed reaches the maximum speed, and reinforcement learning is performed based on the change in the speed of the lead vehicle. The steps include:

[0016] Using each baseline time period as a learning cycle, the variation pattern of driving speed in different learning cycles is analyzed as follows:

[0017] An analysis cycle is defined as at least three reference time periods, and the intermediate speed from the reference speed to the maximum speed is obtained in each reference time period within the analysis cycle.

[0018] Calculate the time taken for the reference speed to change to the intermediate speed, and set a reference time based on the time taken, wherein the duration of the reference time includes 0.5 to 1.5 seconds;

[0019] Within the aforementioned baseline time, if the lead vehicle reaches the intermediate speed in a shorter time, then it is determined that the lead vehicle reaches the maximum speed in a shorter time.

[0020] The fusion calculation includes calculating, based on the reference time, the duration for which the lead vehicle maintains the maximum speed after reaching the maximum speed in each reference time period.

[0021] In a preferred embodiment of the present invention, in step S30, the duration for which the lead vehicle maintains the maximum speed after reaching the maximum speed in each reference time period is calculated based on the reference time, and is obtained by the following formula:

[0022] ;

[0023] In the formula, Indicates intermediate speed. Indicates reference speed;

[0024] Indicates the maximum speed;

[0025] This indicates the sequence number of the current reference time period within the analysis period. =1, 2, ..., -1;

[0026] This indicates the number of the reference time periods within the analysis period. ≥3;

[0027] ;

[0028] In the formula, Indicates the lead car is in From the reference speed within a reference time period Reaching intermediate speed The time required;

[0029] Indicates the relationship with the first The intermediate speed corresponding to each baseline time period;

[0030] Indicates the first The acceleration of the lead vehicle within a reference time period;

[0031] ;

[0032] In the formula, The reference time, ranging from 0.5 to 1.5 seconds, is used as a standard to evaluate the time it takes for the lead vehicle to reach the intermediate speed.

[0033] Indicates the first The time taken for the lead vehicle to reach the intermediate speed from the reference speed within a reference time period;

[0034] This indicates the number of the reference time periods within the analysis period. ≥3;

[0035] ;

[0036] In the formula, Indicates the first The duration during which the lead vehicle maintains its maximum speed after reaching it within a given time period;

[0037] Indicates the first The first vehicle reaches its maximum speed within a reference time period. The moment;

[0038] Indicates the first The end time of each baseline time period.

[0039] In a preferred embodiment of the present invention, the following formula is also included:

[0040] ;

[0041] In the formula, This indicates the total duration during which the lead vehicle maintains its maximum speed after reaching it within the entire analysis period;

[0042] This indicates the number of baseline time periods during which the lead vehicle reaches its maximum speed within the analysis period;

[0043] ;

[0044] In the formula, This indicates the average duration for the lead vehicle to maintain its maximum speed after reaching it within each baseline time period.

[0045] In a preferred embodiment of the present invention: the reference time period corresponding to the longest duration of maintaining the highest speed is obtained based on the calculation results, and the reference time period is marked as the reference reference time period. When the time taken for the lead vehicle to change from the reference speed to the intermediate speed in the future is the same as the reference time of the reference reference time period, it is determined that the lead vehicle will maintain the highest speed for the corresponding duration after reaching the highest speed. Based on this determination, the vehicles in the middle of the target vehicle formation will also maintain the highest speed for the same duration after reaching the highest speed.

[0046] In a preferred embodiment of the present invention: after the lead vehicle maintains its highest speed for a period of time, the change in the lead vehicle's speed is acquired. When the speed shows a decreasing trend, the decrease in the lead vehicle's speed within 2 to 4 seconds is acquired and marked as a reference value. If the decrease in the lead vehicle's speed in subsequent moments is greater than the reference decrease value, the vehicles in the middle of the target vehicle convoy brake to reduce speed; otherwise, braking to reduce speed is not performed.

[0047] In a preferred embodiment of the present invention, the decrease in speed of the lead vehicle within 2 to 4 seconds is obtained and calculated according to the following formula:

[0048] ; ;

[0049] In the formula, This represents the decrease in driving speed over the time interval;

[0050] Indicates time The speed of the lead car;

[0051] Indicates time The speed of the lead car;

[0052] Indicates the starting time of the decrease in speed ( );

[0053] Indicates the moment when the speed decreases to its end;

[0054] Indicates a time interval.

[0055] In a preferred embodiment of the present invention, the following formula is also included:

[0056] ;

[0057] In the formula, Indicates the time interval The average rate of decrease in velocity within the range;

[0058] ;

[0059] in, Indicates the time interval at subsequent moments. The decrease in driving speed within the vehicle.

[0060] In a preferred embodiment of the present invention: the driving speed corresponding to the reference value is obtained, and the driving speed is marked as driving speed I. When the vehicles in the middle of the target vehicle platoon brake and decelerate, if the driving speed of the lead vehicle is greater than driving speed I, it is determined that the lead vehicle is in a state of acceleration, and the vehicles in the middle of the target vehicle platoon also accelerate based on this determination.

[0061] On the other hand, the present invention provides a system for a cooperative control method for intelligent connected vehicles based on a reinforcement learning strategy as described above, comprising:

[0062] A data construction module is used to construct a data chain and perform cooperative control of the target vehicle platoon based on the data chain. The steps for constructing the data chain include:

[0063] The data chain is built based on the lead vehicle in the target vehicle platoon and collects relevant data from the lead vehicle.

[0064] The relevant data includes the speed of the lead vehicle. The relevant data is collected at a given reference time, including collecting the speed at 3-second intervals.

[0065] The change in the speed of the lead vehicle is obtained by taking at least three reference times as a reference period. The speed corresponding to the first reference time is taken as the reference speed. Based on the reference speed, if the speed shows an increasing trend in the second reference time, the highest speed in the third reference time is obtained according to this trend.

[0066] The acceleration determination module is used to determine whether, based on the reference speed, if the driving speed is greater than the reference speed for two consecutive seconds, the vehicles in the middle of the target vehicle platoon will accelerate until the driving speed reaches the maximum speed.

[0067] A data fusion processing module, comprising a learning unit and a computing unit;

[0068] The learning unit is used so that when the vehicles in the middle accelerate, the last vehicle in the target vehicle convoy also accelerates until the speed reaches the maximum speed, and reinforcement learning is performed based on the change in the speed of the lead vehicle. The steps include:

[0069] Using each baseline time period as a learning cycle, the variation pattern of driving speed in different learning cycles is analyzed as follows:

[0070] An analysis cycle is defined as at least three reference time periods, and the intermediate speed from the reference speed to the maximum speed is obtained in each reference time period within the analysis cycle.

[0071] Calculate the time taken for the reference speed to change to the intermediate speed, and set a reference time based on the time taken, wherein the duration of the reference time includes 0.5 to 1.5 seconds;

[0072] Within the aforementioned baseline time, if the lead vehicle reaches the intermediate speed in a shorter time, then it is determined that the lead vehicle reaches the maximum speed in a shorter time.

[0073] The computing unit is used to perform fusion calculation, which includes calculating the duration for which the lead vehicle maintains the maximum speed after reaching the maximum speed in each reference time period, based on the reference time.

[0074] Beneficial effects:

[0075] 1. By constructing a data link based on the speed of the lead vehicle and dynamically adjusting the speeds of the middle and rear vehicles, the entire vehicle formation can be accelerated or decelerated smoothly and efficiently. Specifically, the system transmits the changes in the speed of the lead vehicle to the middle and rear vehicles in real time via the data link. The middle vehicles accelerate when their speed is greater than the reference speed for two consecutive seconds until they reach the maximum speed, and the rear vehicle follows. This dynamic adjustment mechanism reduces the speed difference between vehicles and improves driving efficiency.

[0076] 2. Through reinforcement learning of the speed changes of the lead vehicle, the system can predict the future speed change trend of the lead vehicle and adjust the speed control strategies of the middle and rear vehicles accordingly. During the analysis period, the system calculates the time taken for the lead vehicle to go from the reference speed to the intermediate speed and sets the benchmark time to optimize the timing and magnitude of speed adjustment, thereby improving the overall driving safety and stability.

[0077] 3. The system sets at least three benchmark time periods as an analysis cycle. By analyzing the speed change patterns of the lead vehicle within different benchmark time periods, the system achieves more accurate speed prediction and control. For example, the system calculates the time taken for the lead vehicle to travel from the reference speed to the intermediate speed within each benchmark time period and sets the benchmark time based on these times, which serves as a standard for evaluating the speed change trend of the lead vehicle.

[0078] 4. Furthermore, the system marks the reference time period for maintaining the highest speed for the longest time based on the calculation results. When the speed change of the leading vehicle in the future is the same as the reference time of the reference time period, the system can react quickly and adjust the speed control strategies of the middle and rear vehicles, thus enhancing its adaptability to different road conditions and driving behaviors.

[0079] 5. The system also incorporates a braking deceleration mechanism. When the speed of the lead vehicle decreases and the decrease exceeds the reference decrease value, the vehicles in the middle will brake to reduce their speed. This mechanism can effectively prevent rear-end collisions and improve the robustness of vehicle platooning under complex road conditions. Attached Figure Description

[0080] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a schematic diagram of the modular structure of the intelligent connected vehicle cooperative control system based on reinforcement learning strategy according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the method flow for step S10 in an embodiment of the present invention; Figure 4 This is a schematic diagram of the method flow for step S30 in an embodiment of the present invention; The diagram is labeled as follows: 110 - Data Construction Module; 120 - Speed-up Judgment Module; 130 - Data Fusion Processing Module; 1301 - Learning Unit; 1302 - Calculation Unit. Detailed Implementation

[0081] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0082] Because existing technologies are prone to getting stuck in suboptimal solutions in complex driving environments and have high computational complexity, it is difficult to effectively achieve multi-objective collaborative optimization.

[0083] Based on this, the present invention proposes a cooperative control method and system for intelligent connected vehicles based on reinforcement learning strategy. By constructing a data link based on the speed of the lead vehicle, the system dynamically adjusts the speeds of the middle and rear vehicles. Combined with multi-reference time period analysis and reinforcement learning optimization, it achieves smooth acceleration and deceleration, accurate speed prediction and control of vehicle formation, enhances the adaptability and robustness of the formation, effectively prevents rear-end collisions, and improves the driving experience and comfort.

[0084] The present solution will be further described in detail below through embodiments and in conjunction with the accompanying drawings.

[0085] Reference Figures 1 to 4 This is one embodiment of the present invention, which provides a cooperative control method for intelligent connected vehicles based on a reinforcement learning strategy, comprising the following steps:

[0086] S10: Construct a data chain. The data chain is built based on the driving conditions of the target vehicle platoon. The driving conditions include urban roads (dense road network, many traffic lights, high traffic volume, frequent starts and stops, and relatively high speeds on main roads, but potential congestion). The data chain is used for collaborative control with the target vehicle platoon. The steps for constructing the data chain include:

[0087] S101: The data link is built based on the lead vehicle in the target vehicle platoon and collects relevant data of the lead vehicle (collecting this relevant data for the vehicles following the lead vehicle).

[0088] S102: Relevant data includes the speed of the lead vehicle. The relevant data is collected at a given reference time, including collecting the speed at 3-second intervals.

[0089] S103: Obtain the change in the speed of the lead vehicle using at least three reference times as a reference period. Take the speed corresponding to the first reference time as the reference speed. Based on the reference speed, if the speed shows an increasing trend during the second reference period, obtain the highest speed during the third reference period based on this trend.

[0090] It should be noted that road conditions are categorized according to road type, traffic conditions, and weather conditions, including:

[0091] Road types include urban roads, highways, and rural roads;

[0092] Traffic conditions, including congested and uncongested traffic;

[0093] Weather conditions, including road conditions in sunny weather, rainy weather, and snowy / icy conditions;

[0094] Admittedly, the classification of driving conditions is well known to those skilled in the art, and it also includes other road conditions, which the applicant will not elaborate on here;

[0095] In addition, collecting the speed of the lead vehicle every 3 seconds as a baseline can provide several benefits:

[0096] First, it provides a stable data source for the construction of the data chain, ensuring that the collected speed data is timely and accurate, laying the foundation for subsequent data analysis and processing;

[0097] Secondly, it facilitates the capture of speed change trends. Through continuous speed data collection, the increase or decrease in the speed of the lead vehicle can be clearly observed, providing a basis for dynamically adjusting the speeds of the middle and rear vehicles.

[0098] Finally, it supports multi-reference time period analysis and reinforcement learning optimization. With at least three reference times as a reference time period, it can analyze the speed change pattern of the lead vehicle in different time periods, and then optimize the speed control strategy through reinforcement learning to improve the efficiency and safety of vehicle platooning.

[0099] This high-frequency speed sampling (every 3 seconds) and trend prediction helps to capture the dynamics of the lead vehicle in real time, providing accurate acceleration references for subsequent vehicles and reducing speed mismatch problems caused by information lag.

[0100] S20: If the driving speed is greater than the reference speed for 2 consecutive seconds, the vehicles in the middle of the target vehicle platoon will accelerate until the driving speed reaches the maximum speed.

[0101] S30: When the vehicles in the middle of the convoy accelerate, the last vehicle in the target vehicle convoy also accelerates until it reaches its maximum speed. Reinforcement learning is then performed based on the speed change of the lead vehicle. The steps include:

[0102] S301: Taking each baseline time period as a learning cycle, the variation pattern of driving speed in different learning cycles is analyzed as follows:

[0103] S302: Using at least 3 reference time periods as an analysis cycle, obtain the intermediate speed from the reference speed to the maximum speed in each reference time period within the analysis cycle;

[0104] S303: Calculate the time taken for the reference speed to change to the intermediate speed, and set the reference time based on the time taken. The duration of the reference time includes 0.5 to 1.5 seconds.

[0105] S304: If the lead vehicle reaches the intermediate speed in a shorter time within the baseline time, the time it takes for the lead vehicle to reach the maximum speed is determined to be shorter (if the time it takes for the lead vehicle to reach the maximum speed is determined to be shorter, the vehicles in the middle of the target vehicle platoon will accelerate faster).

[0106] S305: Perform fusion calculation, which includes calculating the duration for which the lead car maintains its maximum speed after reaching it in each reference time period, based on the reference time.

[0107] It should be noted that calculating the duration for which the lead vehicle maintains its maximum speed after reaching it within each reference time period has many important implications. This data provides a clear speed maintenance reference for the middle and rear vehicles, enabling them to adjust their speeds more precisely to keep in sync with the lead vehicle, thereby improving the coordination and driving efficiency of the entire vehicle formation.

[0108] By analyzing the differences in the duration of maintaining the highest speed within different baseline time periods, the stability of the lead vehicle's driving state can be identified, providing key data support for subsequent reinforcement learning to optimize speed control strategies.

[0109] Furthermore, this data helps predict future speed trends of the lead vehicle. For example, when the duration of maintaining maximum speed changes regularly, the system can react in advance, further enhancing the adaptability and robustness of the target vehicle platoon to complex road conditions.

[0110] Meanwhile, the reinforcement learning mechanism enables the system to predict future speed changes of the lead vehicle based on historical data, adjust the control strategy of the tail vehicle in advance, and reduce reaction delay.

[0111] In S30, the duration for which the lead car maintains its maximum speed after reaching it within each reference time period is calculated based on the reference time, using the following formula:

[0112] ;

[0113] In the formula, Indicates intermediate speed. Indicates reference speed;

[0114] Indicates the maximum speed;

[0115] This indicates the sequence number of the current reference time period within the analysis period. =1, 2, ..., -1 (because we need to obtain the intermediate speed, so we don't take it) ,when When =1, the result is Relatively closer , = When -1, it is closer to );

[0116] This indicates the number of baseline time periods within the analysis period. ≥3;

[0117] ;

[0118] In the formula, Indicates the lead car is in From the reference speed within a reference time period Reaching intermediate speed The time required;

[0119] Indicates the relationship with the first The intermediate speed corresponding to each baseline time period;

[0120] Indicates the first The acceleration of the lead vehicle within a reference time period;

[0121] ;

[0122] In the formula, The reference time, ranging from 0.5 to 1.5 seconds, is used as a standard to evaluate the time it takes for the lead vehicle to reach the intermediate speed.

[0123] Indicates the first The time taken for the lead vehicle to reach the intermediate speed from the reference speed within a reference time period;

[0124] This indicates the number of baseline time periods within the analysis period. ≥3;

[0125] ;

[0126] In the formula, Indicates the first The duration during which the lead vehicle maintains its maximum speed after reaching it within a given time period;

[0127] Indicates the first The first vehicle reaches its maximum speed within a reference time period. The moment;

[0128] Indicates the first The end time of each baseline time period.

[0129] It also includes calculations based on the following formula:

[0130] ;

[0131] In the formula, This indicates the total duration during which the lead vehicle maintains its maximum speed after reaching it within the entire analysis period;

[0132] This indicates the number of baseline time periods during which the lead vehicle reaches its maximum speed within the analysis period;

[0133] ;

[0134] In the formula, This indicates the average duration for the lead vehicle to maintain its maximum speed after reaching it within each baseline time period.

[0135] Based on the calculation results, obtain the reference time period corresponding to the longest duration of maintaining the highest speed, and mark the reference time period as the reference reference time period. When the time taken for the lead vehicle to change from the reference speed to the intermediate speed in the future is the same as the reference time of the reference reference time period, it is determined that the lead vehicle will maintain the highest speed for the corresponding duration after reaching the highest speed. Based on this determination, the vehicles in the middle of the target vehicle formation will also maintain the highest speed for the same duration after reaching the highest speed.

[0136] It should be noted that the reference time period marking enables the system to reuse historical optimal control strategies, reduce redundant calculations, and improve real-time performance.

[0137] After the lead vehicle maintains its highest speed for a certain period of time, the change in the lead vehicle's speed is recorded. When the speed shows a decreasing trend, the decrease in the lead vehicle's speed over 2 to 4 seconds is recorded and marked as a reference value. If the decrease in the lead vehicle's speed in subsequent moments is greater than the reference decrease value, the vehicles in the middle of the target vehicle formation will brake to reduce their speed; otherwise, they will not brake to reduce their speed.

[0138] By converting the deceleration process into comparable numerical indicators, an objective basis is provided for braking decisions of vehicles in the middle, reducing subjective judgment errors.

[0139] The decrease in speed of the lead vehicle over 2-4 seconds is obtained and calculated using the following formula:

[0140] ; ;

[0141] In the formula, This indicates the decrease in driving speed within a time interval (2-4 seconds);

[0142] Indicates time The speed of the lead car;

[0143] Indicates time The speed of the lead car;

[0144] Indicates the starting time of the decrease in speed ( );

[0145] Indicates the moment when the speed decreases to its end;

[0146] Indicates a time interval (2 to 4 seconds).

[0147] It also includes calculations based on the following formula:

[0148] ;

[0149] In the formula, Indicates the time interval The average rate of decrease in velocity within the range;

[0150] ;

[0151] in, Indicates the time interval at subsequent moments. The decrease in driving speed within the vehicle.

[0152] Obtain the driving speed corresponding to the reference value and mark the driving speed as speed I. When the vehicles in the middle of the target vehicle platoon brake and decelerate, if the driving speed of the lead vehicle is greater than speed I, it is determined that the lead vehicle is in a state of acceleration, and the vehicles in the middle of the target vehicle platoon also accelerate based on this determination.

[0153] Based on the above, this application improves the efficiency and safety of vehicle platooning by means of dynamic speed adjustment, reinforcement learning optimization, multi-reference time period analysis, reference reference time period marking, and braking deceleration mechanism, achieving more accurate speed control and prediction, and enhancing the adaptability and robustness of vehicle platooning.

[0154] This embodiment, in conjunction with the above-mentioned cooperative control method for intelligent connected vehicles based on reinforcement learning strategies, also proposes a system applied to this method, as follows:

[0155] Data construction module 110 is used to construct a data link and perform collaborative control with the target vehicle platoon based on the data link. The steps for constructing the data link include:

[0156] The data link is built based on the lead vehicle in the target vehicle platoon and collects relevant data from the lead vehicle;

[0157] The relevant data includes the speed of the lead vehicle. The relevant data is collected at a given reference time, including collecting the speed at 3-second intervals.

[0158] The speed change of the lead vehicle is obtained by taking at least three reference times as a reference period. The speed corresponding to the first reference time is taken as the reference speed. Based on the reference speed, if the speed changes with an increasing trend in the second reference time, the highest speed in the third reference time is obtained according to this trend.

[0159] The acceleration determination module 120 is used to determine if the vehicle speed is greater than the reference speed for two consecutive seconds, and then the vehicle in the middle of the target vehicle platoon will accelerate until the vehicle speed reaches the maximum speed.

[0160] The data fusion processing module 130 includes a learning unit 1301 and a computing unit 1302.

[0161] Learning unit 1301 is used so that when the vehicles in the middle accelerate, the last vehicle in the target vehicle platoon also accelerates until the speed reaches the maximum speed, and reinforcement learning is performed based on the speed change of the lead vehicle. The steps include:

[0162] Using each baseline time period as a learning cycle, the variation pattern of driving speed in different learning cycles is analyzed as follows:

[0163] An analysis cycle is defined as at least three reference time periods, and the intermediate speed from the reference speed to the maximum speed is obtained in each reference time period within the analysis cycle.

[0164] Calculate the time taken for the reference speed to change to the intermediate speed, and set a baseline time based on the time taken. The duration of the baseline time includes 0.5 to 1.5 seconds.

[0165] Within the baseline time, if the lead vehicle reaches the intermediate speed in a shorter time, then the lead vehicle is judged to have reached the maximum speed in a shorter time.

[0166] The calculation unit 1302 is used to perform fusion calculation, which includes calculating the duration for which the lead car maintains the maximum speed after reaching the maximum speed in each reference time period based on the reference time.

[0167] In summary, this invention constructs a data chain based on the speed of the lead vehicle to dynamically adjust the speeds of the middle and rear vehicles. By combining multi-reference time period analysis and reinforcement learning optimization, it achieves smooth acceleration and deceleration, accurate speed prediction and control of vehicle formations, enhances formation adaptability and robustness, effectively prevents rear-end collisions, and improves driving experience and comfort.

[0168] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A cooperative control method for intelligent connected vehicles based on a reinforcement learning strategy, characterized in that, Includes the following steps: S10: Construct a data link and perform cooperative control in the target vehicle platoon based on the data link. The steps of constructing the data link include: The data chain is built based on the lead vehicle in the target vehicle platoon and collects relevant data from the lead vehicle. The relevant data includes the speed of the lead vehicle. The relevant data is collected at a given reference time, including collecting the speed at 3-second intervals. The change in the speed of the lead vehicle is obtained by taking at least three reference times as a reference period. The speed corresponding to the first reference time is taken as the reference speed. Based on the reference speed, if the speed shows an increasing trend in the second reference time, the highest speed in the third reference time is obtained according to this trend. S20: If the driving speed is greater than the reference speed for 2 consecutive seconds, the vehicles in the middle of the target vehicle platoon will accelerate until the driving speed reaches the maximum speed. S30: When the vehicles in the middle accelerate, the last vehicle in the target vehicle platoon also accelerates until the speed reaches the maximum speed, and reinforcement learning is performed based on the change in the speed of the lead vehicle. The steps include: Using each baseline time period as a learning cycle, the variation pattern of driving speed in different learning cycles is analyzed as follows: An analysis cycle is defined as at least three reference time periods, and the intermediate speed from the reference speed to the maximum speed is obtained in each reference time period within the analysis cycle. Calculate the time taken for the reference speed to change to the intermediate speed, and set a reference time based on the time taken, wherein the duration of the reference time includes 0.5 to 1.5 seconds; Within the aforementioned baseline time, if the lead vehicle reaches the intermediate speed in a shorter time, then it is determined that the lead vehicle reaches the maximum speed in a shorter time. The fusion calculation includes calculating, based on the reference time, the duration for which the lead vehicle maintains the maximum speed after reaching the maximum speed in each reference time period.

2. The intelligent connected vehicle cooperative control method based on reinforcement learning strategy as described in claim 1, characterized in that, In step S30, the duration for which the lead vehicle maintains the maximum speed after reaching it in each reference time period is calculated based on the reference time, using the following formula: ; In the formula, Indicates intermediate speed. Indicates reference speed; Indicates the maximum speed; This indicates the sequence number of the current reference time period within the analysis period. =1, 2, ..., -1; This indicates the number of reference time periods within the analysis period. ≥3; ; In the formula, Indicates the lead car is in From the reference speed within a reference time period Reaching intermediate speed The time required; Indicates the relationship with the first The intermediate speed corresponding to each baseline time period; Indicates the first The acceleration of the lead vehicle within a reference time period; ; In the formula, The reference time, ranging from 0.5 to 1.5 seconds, is used as a standard to evaluate the time it takes for the lead vehicle to reach the intermediate speed. Indicates the first The time taken for the lead vehicle to reach the intermediate speed from the reference speed within a reference time period; This indicates the number of reference time periods within the analysis period. ≥3; ; In the formula, Indicates the first The duration during which the lead vehicle maintains its maximum speed after reaching it within a given time period; Indicates the first The first vehicle reaches its maximum speed within a reference time period. The moment; Indicates the first The end time of each baseline time period.

3. The intelligent connected vehicle cooperative control method based on reinforcement learning strategy as described in claim 2, characterized in that, It also includes calculations based on the following formula: ; In the formula, This indicates the total duration during which the lead vehicle maintains its maximum speed after reaching it within the entire analysis period; This indicates the number of baseline time periods during which the lead vehicle reaches its maximum speed within the analysis period; ; In the formula, This indicates the average duration for the lead vehicle to maintain its maximum speed after reaching it within each baseline time period.

4. A cooperative control method for intelligent connected vehicles based on a reinforcement learning strategy as described in any one of claims 2 to 3, characterized in that, Based on the calculation results, the reference time period corresponding to the longest duration of maintaining the highest speed is obtained and marked as the reference reference time period. When the time taken for the lead vehicle to change from the reference speed to the intermediate speed in the future is the same as the reference time of the reference reference time period, it is determined that the lead vehicle will maintain the highest speed for the corresponding duration after reaching the highest speed. Based on this determination, the vehicles in the middle of the target vehicle formation will also maintain the highest speed for the same duration after reaching the highest speed.

5. The intelligent connected vehicle cooperative control method based on reinforcement learning strategy as described in claim 4, characterized in that, After the lead vehicle maintains its highest speed for a certain period of time, the change in the lead vehicle's speed is acquired. If the speed shows a decreasing trend, the decrease in the lead vehicle's speed within 2 to 4 seconds is acquired and marked as a reference value. If the decrease in the lead vehicle's speed in subsequent moments is greater than the reference decrease value, the vehicles in the middle of the target vehicle formation brake to reduce speed; otherwise, they do not brake to reduce speed.

6. The intelligent connected vehicle cooperative control method based on reinforcement learning strategy as described in claim 5, characterized in that, The decrease in speed of the lead vehicle over 2-4 seconds is obtained and calculated using the following formula: ; ; In the formula, This represents the decrease in driving speed over the time interval; Indicates time The speed of the lead car; Indicates time The speed of the lead car; Indicates the starting time of the decrease in speed ( ); Indicates the moment when the speed decreases to its end; Indicates a time interval.

7. The intelligent connected vehicle cooperative control method based on reinforcement learning strategy as described in claim 6, characterized in that, It also includes calculations based on the following formula: ; In the formula, Indicates the time interval The average rate of decrease in velocity within the range; ; in, Indicates the time interval at subsequent moments. The decrease in driving speed within the area.

8. The intelligent connected vehicle cooperative control method based on reinforcement learning strategy as described in claim 7, characterized in that, Obtain the driving speed corresponding to the reference value and mark the driving speed as I driving speed. When the vehicles in the middle of the target vehicle convoy brake and decelerate, if the driving speed of the lead vehicle is greater than the I driving speed, it is determined that the lead vehicle is in a state of acceleration, and the vehicles in the middle of the target vehicle convoy also accelerate based on this determination.

9. A system applied to the cooperative control method for intelligent connected vehicles based on a reinforcement learning strategy as described in claim 1, characterized in that, include: A data construction module is used to construct a data chain and perform cooperative control of the target vehicle platoon based on the data chain. The steps for constructing the data chain include: The data chain is built based on the lead vehicle in the target vehicle platoon and collects relevant data from the lead vehicle. The relevant data includes the speed of the lead vehicle. The relevant data is collected at a given reference time, including collecting the speed at 3-second intervals. The change in the speed of the lead vehicle is obtained by taking at least three reference times as a reference period. The speed corresponding to the first reference time is taken as the reference speed. Based on the reference speed, if the speed shows an increasing trend in the second reference time, the highest speed in the third reference time is obtained according to this trend. The acceleration determination module is used to determine whether, based on the reference speed, if the driving speed is greater than the reference speed for two consecutive seconds, the vehicles in the middle of the target vehicle platoon will accelerate until the driving speed reaches the maximum speed. A data fusion processing module, comprising a learning unit and a computing unit; The learning unit is used so that when the vehicles in the middle accelerate, the last vehicle in the target vehicle convoy also accelerates until the speed reaches the maximum speed, and reinforcement learning is performed based on the change in the speed of the lead vehicle. The steps include: Using each baseline time period as a learning cycle, the variation pattern of driving speed in different learning cycles is analyzed as follows: An analysis cycle is defined as at least three reference time periods, and the intermediate speed from the reference speed to the maximum speed is obtained in each reference time period within the analysis cycle. Calculate the time taken for the reference speed to change to the intermediate speed, and set a reference time based on the time taken, wherein the duration of the reference time includes 0.5 to 1.5 seconds; Within the aforementioned baseline time, if the lead vehicle reaches the intermediate speed in a shorter time, then it is determined that the lead vehicle reaches the maximum speed in a shorter time. The computing unit is used to perform fusion calculation, which includes calculating the duration for which the lead vehicle maintains the maximum speed after reaching the maximum speed in each reference time period, based on the reference time.

Citation Information

Patent Citations

  • Intelligent network connection vehicle cooperation control method based on reinforcement learning strategy

    CN117215203A

  • Intelligent network connection vehicle formation centralized control method based on deep reinforcement learning

    CN120161843A