Methods, devices, and vehicles for identifying road congestion conditions
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]鉴于上述问题,本发明实施例提供了一种路面拥堵工况的识别方法、装置、车辆,用于解决当前的底盘控制系统往往缺乏足够的情境感知能力,无法深度理解道路场景以输出最准确的车辆行驶拥堵工况的问题
本发明实施例通过获取车辆的图像视频帧,之后基于图像视频帧,得到车辆所在前方道路的视觉拥堵指数;然后获取车辆的行驶状态数据,基于行驶状态数据,得到前方道路的状态拥堵程度;最后基于视觉拥堵指数和状态拥堵程度,得到前方道路的拥堵工况识别结果以及融合置信度。这样本发明实施例通过图像视频帧所得到的视觉拥堵指数,使得路面拥堵工况的识别过程能够像人一样对整体交通场景的“拥堵和复杂程度”进行端到端的综合评估,避免了传统方法中手工设计特征的不完备性和短视性,同时将视觉拥堵指数和基于车辆的行驶状态数据所得到的状态拥堵程度相结合,使车辆智驾系统能智能地判断何时该“相信眼睛”,何时该“相信体感”,极大提升了系统的智能水平和可靠性,对视觉拥堵指数和状态拥堵程度的仲裁、协同的判断结果更符合人类的决策逻辑,融合效果更优。解决了相关技术中无法深度理解道路场景以输出最准确的车辆行驶拥堵工况的问题。
Smart Images

Figure CN122575134A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle driving technology, specifically to a method, device, and vehicle for identifying road congestion conditions. Background Technology
[0002] Currently, when vehicles encounter complex urban traffic congestion, their chassis control systems typically employ a recognition method based on vehicle status (such as pedals, steering wheel, and vehicle acceleration) to perceive road congestion. Essentially, this is a "post-event judgment" based on the vehicle's dynamic responses and driver actions. This method cannot perceive changes in the traffic environment ahead that lead to the current operation, exhibits significant recognition lag, and fails to provide a forward-looking time window for chassis systems requiring pre-adjustment (such as suspension and powertrain).
[0003] On the other hand, while traditional computer vision methods can identify individual targets such as vehicles and lane lines, their ability to "understand" complex traffic scenes is limited. They typically only output a series of discrete, low-level features (such as traffic density and following distance), making it difficult to provide an end-to-end, human-like comprehensive judgment on the overall "complexity and congestion level" of the scene as a whole. Furthermore, traditional vision algorithms are unstable in poor lighting, weather, and occlusion conditions, and their outputs lack a self-evaluating "confidence level," making it difficult for downstream systems to determine when to trust the results of visual perception.
[0004] Therefore, current chassis control systems often lack sufficient context awareness and are unable to deeply understand road scenarios in order to output the most accurate vehicle driving congestion conditions. Summary of the Invention
[0005] In view of the above problems, embodiments of the present invention provide a method, device, and vehicle for identifying road congestion conditions, which solves the problem that current chassis control systems often lack sufficient context awareness and are unable to deeply understand road scenarios in order to output the most accurate vehicle driving congestion conditions.
[0006] According to one aspect of the present invention, a method for identifying road congestion conditions is provided. The method includes: acquiring image video frames of a vehicle, wherein the image video frames are used to characterize the congestion conditions of the road ahead where the vehicle is located; obtaining a visual congestion index of the road ahead where the vehicle is located based on the image video frames; acquiring vehicle driving state data; obtaining the state congestion degree of the road ahead based on the driving state data; and obtaining a congestion condition identification result and a fusion confidence score of the road ahead based on the visual congestion index and the state congestion degree, wherein the fusion confidence score is used to characterize the credibility of the congestion condition identification result.
[0007] According to another aspect of the present invention, a road congestion condition identification device is provided. The device includes: a first acquisition module for acquiring image video frames of a vehicle, wherein the image video frames are used to characterize the congestion condition of the road ahead where the vehicle is located; a first obtaining module for obtaining a visual congestion index of the road ahead where the vehicle is located based on the image video frames; a second acquisition module for acquiring vehicle driving state data; a second obtaining module for obtaining the state congestion degree of the road ahead based on the driving state data; and a third obtaining module for obtaining a congestion condition identification result and a fusion confidence score of the road ahead based on the visual congestion index and the state congestion degree, wherein the fusion confidence score is used to characterize the credibility of the congestion condition identification result.
[0008] According to another aspect of the present invention, a computer device is provided, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction that causes the processor to perform the operation of the first aspect of the road congestion condition identification method.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein at least one executable instruction is stored in the storage medium, the executable instruction causing a computer device / apparatus to perform the operation of the road congestion condition identification method of the first aspect.
[0010] According to another aspect of the present invention, a vehicle is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the operation of the road congestion condition identification method of the first aspect.
[0011] The technical solution provided by this invention has the following advantages: This invention acquires vehicle image and video frames, then obtains the visual congestion index of the road ahead based on these frames; next, it acquires vehicle driving status data, and obtains the state congestion level of the road ahead based on this data; finally, it obtains the congestion condition identification result and fusion confidence score of the road ahead based on the visual congestion index and state congestion level. This invention, through the visual congestion index obtained from image and video frames, enables the road congestion condition identification process to perform an end-to-end comprehensive assessment of the "congestion and complexity" of the overall traffic scene, much like a human. This avoids the incompleteness and shortsightedness of manually designed features in traditional methods. Furthermore, by combining the visual congestion index with the state congestion level obtained from vehicle driving status data, the vehicle's intelligent driving system can intelligently determine when to "trust the eyes" and when to "trust the senses," greatly improving the system's intelligence and reliability. The arbitration and collaborative judgment results of the visual congestion index and state congestion level are more in line with human decision-making logic, resulting in a better fusion effect. This solves the problem in related technologies where a deep understanding of the road scene is insufficient to output the most accurate vehicle driving congestion condition.
[0012] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0013] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating a first embodiment of a method for identifying road congestion conditions provided by the present invention is shown. Figure 2 A schematic diagram of a second embodiment of a road congestion condition identification system provided by the present invention is shown; Figure 3 A schematic diagram of the structure of a road congestion identification device provided by the present invention is shown; Figure 4 A schematic diagram of an embodiment of a computer device provided by the present invention is shown; Figure 5 A schematic diagram of the structure of a vehicle provided by the present invention is shown. Detailed Implementation
[0014] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0015] Figure 1 The diagram shows a flowchart of a first embodiment of a method for identifying road congestion conditions provided by the present invention, which is executed by a vehicle. Figure 1 As shown, the method includes the following steps: Step S101: Acquire image and video frames of the vehicle, wherein the image and video frames are used to represent the congestion conditions on the road ahead of the vehicle.
[0016] Optionally, in this embodiment of the invention, a vehicle front-facing camera is used to capture image and video frames of the road ahead, such as real-time image frames or short image sequences. Each image and video frame is a real-world view of the road ahead during the vehicle's journey, including visual information related to traffic flow, lanes, road conditions, and other congestion-related conditions.
[0017] It should be noted that the frame rate of the image and video frames in the embodiments of the present invention can be 15-30fps, which is common for vehicle perception, and the resolution is adapted to vehicle cameras (such as 1080P). The acquisition process is real-time continuous acquisition to ensure dynamic capture of road congestion conditions.
[0018] Step S102: Based on the image video frames, obtain the visual congestion index of the road ahead where the vehicle is located.
[0019] Optionally, feature information for quantifying congestion can be extracted from image and video frames, such as traffic density (the percentage of vehicles per pixel area in a single lane) and following distance (the ratio of the pixel distance between the vehicle and the vehicle in front to the total pixel length of the lane). Analyzing these features can roughly yield a visual congestion index, which includes visual congestion level (V_Congestion_Level) and visual congestion level confidence (V_Confidence). V_Confidence represents the level of confidence in V_Congestion_Level, and is a continuous value from 0 to 1, where 0 is completely unreliable and 1 is completely reliable. This value is obtained from the model's own uncertainty assessment mechanism, such as Dropout inference and calculation of the output probability distribution entropy.
[0020] Further, step S102 includes: Step S1021: Input the image and video frames into the target visual model and output the visual congestion index. The target visual model is used to perform visual dimension congestion analysis on the image and video frames.
[0021] Optionally, embodiments of the present invention may use a “visual large model end-to-end perception” approach to achieve the output of the visual congestion index, which further refers to the output of the visual congestion level V_Congestion_Level in the visual congestion index.
[0022] Specifically: Images and videos captured by vehicle cameras are input into a pre-trained target vision model (such as VisionTransformer, large convolutional neural network, etc.). The target vision model is trained with massive amounts of complex traffic scene data. Instead of detecting a single target, it directly regresses the congestion and complexity of the entire scene and outputs the visual congestion level V_Congestion_Level, which is a continuous value from 0 to 1, where 0 represents smooth traffic and 1 represents severe congestion (vehicles are almost stationary). In subsequent applications, the congestion level is determined based on the congestion value, which directly represents the target vision model's comprehensive judgment that the road ahead is in a complex congestion condition.
[0023] Table 1 shows the relationship between visual congestion level and congestion grade:
[0024] It is important to understand that for a target visual model, it needs to analyze the input image and video frames to obtain the corresponding three indicators: traffic flow, vehicle speed, and vehicle distance. Based on these three indicators, it outputs a value indicating the degree of visual congestion, and then determines the corresponding congestion level based on the value of the visual congestion.
[0025] Step S103: Obtain vehicle driving status data.
[0026] Optionally, vehicle driving status data is collected by the vehicle status sensing unit, specifically including pedal operation signals, steering wheel angle signals, longitudinal / lateral acceleration signals of the vehicle body, vehicle speed signals, etc., which are the core sensing data reflecting the actual driving dynamics of the vehicle.
[0027] Step S104: Based on the driving status data, obtain the degree of congestion on the road ahead.
[0028] Optionally, feature extraction is performed on the collected driving status data to calculate the frequency of operation (pedal switching frequency, steering wheel fine-tuning activity rate) and the degree of dynamic change (longitudinal acceleration standard deviation, lateral angular velocity change intensity). The two types of features are then fused into a comprehensive index, namely the state congestion level S_Congestion_Index, which is also a continuous value from 0 to 1. The higher the value, the greater the possibility of congestion based on the dynamic judgment of the vehicle.
[0029] Step S105: Based on the visual congestion index and the state congestion level, obtain the congestion condition identification result and fusion confidence score of the road ahead, wherein the fusion confidence score is used to characterize the credibility of the congestion condition identification result.
[0030] Optionally, the visual congestion index of the visual flow (including the visual congestion level V_Congestion_Level and the visual congestion level confidence V_Confidence) and the state congestion level S_Congestion_Index of the state flow are fused into a dual-flow decision. The output includes rich congestion condition labels (such as smooth flow, light congestion, severe and complex congestion, etc.), i.e., congestion condition identification results, and also includes a final fusion confidence score, which is a value of 0-1, used to evaluate the credibility of the fused identification results. This is an important basis for the downstream chassis control system to execute strategies.
[0031] This invention acquires vehicle image and video frames, then obtains the visual congestion index of the road ahead based on these frames; next, it acquires vehicle driving status data, and obtains the state congestion level of the road ahead based on this data; finally, it obtains the congestion condition identification result and fusion confidence score of the road ahead based on the visual congestion index and state congestion level. This invention, through the visual congestion index obtained from image and video frames, enables the road congestion condition identification process to perform an end-to-end comprehensive assessment of the "congestion and complexity" of the overall traffic scene, much like a human. This avoids the incompleteness and shortsightedness of manually designed features in traditional methods. Furthermore, by combining the visual congestion index with the state congestion level obtained from vehicle driving status data, the vehicle's intelligent driving system can intelligently determine when to "trust the eyes" and when to "trust the senses," greatly improving the system's intelligence and reliability. The arbitration and collaborative judgment results of the visual congestion index and state congestion level are more in line with human decision-making logic, resulting in a better fusion effect. This solves the problem in related technologies where a deep understanding of the road scene is insufficient to output the most accurate vehicle driving congestion condition.
[0032] As an optional implementation, the training process of the target visual model includes: Step a1: Obtain the vehicle's image training data frame and the actual congestion label of the road ahead carried by the image training data frame; Step a2: Input the image training data into the initial visual model and output the congestion prediction label; Step a3: Compare the actual congestion label with the predicted congestion label, and when the error between the actual congestion label and the predicted congestion label is greater than the error threshold, adjust the model parameters of the initial visual model until the error is less than or equal to the error threshold to obtain the target visual model.
[0033] Optionally, model training requires a large amount of driving videos containing different levels of congestion, weather, lighting, and road conditions as training data (i.e., image training data frames), and the training data needs to be labeled with the ground truth value of the congestion index, which is the "actual congestion label" and serves as the supervision data for model training.
[0034] The labeled feature sequence / image frame is input into the untrained initial visual model. The initial visual model outputs a predicted value of the congestion level through forward propagation, corresponding to a "congestion prediction label", which has the same quantization scale (0-1 continuous value) as the actual congestion label.
[0035] Using mean squared error (MSE) as the loss function, the error between the predicted congestion index (i.e., the pre-congestion prediction label) output by the initial visual model and the true congestion index (i.e., the actual congestion label) is compared. An error threshold is set as the convergence criterion for model training, which is determined by actual engineering needs (e.g., MSE ≤ 0.01). Through the backpropagation algorithm, when the error between the actual congestion label and the predicted congestion label is greater than the error threshold, the model parameters of the initial visual model, such as the weights and biases of the initial visual model, are adjusted. The training is iterated repeatedly until the error is less than or equal to the error threshold, and the trained model is obtained, which is the final target visual model.
[0036] The embodiments of the present invention employ a supervised model training method, using actual congestion labels as supervisory data, ensuring that the training direction of the target visual model is consistent with the recognition requirements of actual road congestion conditions, thereby improving the accuracy and reliability of the model's output visual congestion index.
[0037] As an optional embodiment, step a1 above includes: Step a11: Obtain the first driving trajectory data after the vehicle has driven on the road ahead carried by the image training data frame; Step a12: Based on the first driving trajectory data, obtain the vehicle's average travel speed; Step a13: Obtain the second driving trajectory data when the road ahead is in a non-congested period, carried by the image training data frame, and the vehicle is driving at a constant speed. Step a14: Based on the second driving trajectory data, obtain the vehicle's free-flow travel speed; Step a15: Based on the ratio of average travel speed to free-flow travel speed, the actual congestion label is obtained.
[0038] Optionally, models such as YOLO can be used to automatically detect vehicles and generate the total number of other vehicle objects identified within the field of view of the vehicle's camera, while also generating the pixel position of the vehicle after each movement.
[0039] On the same road segment, such as the "current road" mentioned in the above embodiment, the first driving trajectory data of the vehicle under the current road is obtained based on the pixel position after each movement of the vehicle. Then, the vehicle-mounted high-precision GPS module collects data such as latitude and longitude, driving time, cumulative mileage, and instantaneous speed corresponding to the first driving trajectory data, and then obtains the first total driving mileage and the first total driving time of the vehicle.
[0040] The first total mileage is the actual road mileage of the current road segment (calculated by the latitude and longitude of the GPS track or determined by the actual location), and the first total travel time is the actual travel time of the vehicle from the start to the end of the road segment (excluding non-driving stationary time such as stopping at red lights). The formula for calculating the average travel speed is: Average travel speed = First total distance traveled / First total travel time.
[0041] Free-flow velocity is the average speed of a vehicle traveling freely on a road without congestion. Therefore, this embodiment of the invention acquires high-precision GPS trajectory data of the vehicle during a non-congested period (such as early morning), when there is no traffic congestion and the vehicle is traveling at a constant speed. This data is used to obtain the vehicle's second driving trajectory data. The acquisition conditions are consistent with those of the first driving trajectory data to ensure the consistency of mileage and time calculations. Based on the second driving trajectory data, the vehicle's second total mileage and second total driving time are then obtained. The calculation process is the same as that of the first total mileage and first total driving time, and will not be repeated here. It should be noted that the second total mileage and the first total mileage are fixed mileages on the same target road segment to ensure their comparability.
[0042] The formula for calculating free-flow travel speed is: Free-flow travel speed = Second total travel distance / Second total travel time.
[0043] The ratio of average travel speed to free-flow speed is then used as the true value of the congestion index, i.e., the actual congestion label. The speed ratio ranges from 0 to 1. When vehicles are stationary within a road segment, the average travel speed approaches 0, and the speed ratio approaches 0, corresponding to severe congestion; when vehicles are traveling at free-flow speed, the speed ratio is 1, corresponding to complete free flow.
[0044] In addition, in determining the actual congestion label, the embodiments of the present invention can also manually score the video by referring to the congestion level description in Table 1.
[0045] In this embodiment of the invention, the actual congestion label is calculated using the high-precision GPS trajectory data of the vehicle itself, without the need to obtain any data from other vehicles. This conforms to the actual engineering scenario of data collection at the vehicle end and reduces the difficulty and cost of ground truth labeling.
[0046] As an optional embodiment, the visual congestion index includes visual congestion level and visual congestion level credibility. Visual congestion level credibility is used to characterize the trustworthiness of visual congestion level. Visual congestion level credibility is obtained by combining the error between each actual congestion label and each predicted congestion label with the surrounding environmental information of the road ahead of the vehicle. The above step S105 includes: Step S1051: When the credibility of the visual congestion level falls within the credibility value range, the congestion condition corresponding to the visual congestion level or the state congestion level is used as the congestion condition identification result of the road ahead.
[0047] Step S1052: When the credibility of visual congestion level does not fall within the credibility value range, based on the correlation between visual congestion level, visual congestion level credibility and state congestion level, the congestion condition identification result and fusion confidence of the road ahead are obtained.
[0048] Optionally, in embodiments of the present invention, the reliability of visual congestion levels is obtained in the following manner: During model training, there will inevitably be an error between the actual congestion label and the predicted congestion label obtained in each round. Therefore, all the errors obtained during model training are aggregated to calculate the mean absolute error (MAE), using the formula: MAE = ( For actual congestion labels, (where n is the number of samples, and n is the congestion prediction label). The smaller the mean absolute error, the higher the prediction accuracy of the model and the higher the basic reliability of the visual congestion level.
[0049] Visual algorithms are unstable in poor lighting, weather, and occlusion conditions. Therefore, it is currently necessary to obtain information about the surrounding environment of the vehicle on the current road and generate an influence coefficient based on this information.
[0050] Specifically, the surrounding environment information includes light intensity (such as at night, backlight), weather conditions (such as heavy rain, heavy fog), road obstruction (such as tall buildings, trees, vehicles), etc., which are obtained by the vehicle-mounted environmental sensors (such as light sensors, rain sensors) or the scene recognition module of the visual model itself.
[0051] An influence coefficient is set based on information about the surrounding environment. The influence coefficient is a value between 0 and 1, which is used to quantify the degree of negative impact of the environment on visual perception. The worse the environment, the greater the influence coefficient (e.g., influence coefficient = 0.8 on a rainy day and influence coefficient = 0.1 on a sunny day). The influence coefficient can be generated through a preset environment-coefficient mapping table set by the vehicle perception system.
[0052] The credibility of visual congestion can then be obtained by multiplying the influence coefficient by the mean absolute error. Alternatively, the calculation formula can be used: Credibility = 1 - (Influence Coefficient × Mean Absolute Error), ensuring that the credibility increases as the mean absolute error and influence coefficient decrease, consistent with the reliability of actual visual perception. After calculation, the results are normalized to ensure that the final credibility is a continuous value between 0 and 1.
[0053] In this embodiment of the invention, the credibility of visual congestion is calculated from two dimensions: the accuracy of the model itself and the influence of the external environment. This more comprehensively considers the actual influencing factors of vehicle-mounted visual perception and improves the accuracy of credibility assessment.
[0054] This invention relates to two modes: a "high-confidence visual-dominated mode" and a "low-confidence visual-assisted mode." Specifically, a confidence range is set for the confidence level of visual congestion. When the confidence level of visual congestion falls within this range, the congestion condition corresponding to the visual congestion level or the state congestion level is used as the identification result of the congestion condition of the road ahead. If the confidence level of visual congestion does not fall within the range, the adoption level of visual flow and state flow is dynamically adjusted based on the confidence level of visual congestion. A differentiated fusion strategy is executed by combining the numerical correlation between the three (i.e., visual congestion level, visual congestion level confidence level, and state congestion level) (such as the level of visual congestion confidence level and whether visual congestion level and state congestion level conflict), and finally, the identification result and fusion confidence level are output.
[0055] As an optional embodiment, the confidence value range includes a first range and a second range, wherein the first range is used to characterize that the confidence level of visual congestion is higher than a first threshold, and the second range is used to characterize that the confidence level of visual congestion is lower than a second threshold. Step S1051 includes: Step c1: If the credibility of the visual congestion level falls into the first interval, the congestion condition corresponding to the visual congestion level is used as the congestion condition identification result of the road ahead.
[0056] Step c2: If the visual congestion level confidence falls into the second interval, the congestion condition corresponding to the state congestion level is used as the congestion condition identification result of the road ahead.
[0057] Optionally, embodiments of the present invention involve two modes: a "high-confidence visual-dominated mode" and a "low-confidence visual-assisted mode." Specifically, the first interval is a high-confidence interval (e.g., the first threshold is set to 0.9). When the visual congestion confidence (V_Confidence) is higher than the first threshold, the result given by the target visual model is used as the final judgment benchmark. Even if the state congestion level (S_Congestion_Index) is not high temporarily (e.g., just entered the congestion queue and has not been frequently operated), the congestion condition corresponding to the visual congestion level is directly adopted to achieve early warning of congestion.
[0058] The second interval is a low confidence interval (e.g., the second threshold is set to 0.3). When the confidence level of visual congestion is lower than this second threshold (e.g., at night or during heavy rain), the visual results are for reference only. The judgment mainly depends on the degree of congestion in the state to ensure the reliability of the basic function of congestion recognition. The congestion conditions corresponding to the visual congestion level are only used as auxiliary verification.
[0059] As an optional embodiment, step S1052 above includes: Step d1: If the visual congestion level confidence does not fall into the first interval or the second interval, and the correlation is a numerical conflict relationship, then the congestion condition corresponding to the numerical conflict relationship is obtained as the congestion condition identification result of the road ahead.
[0060] Step d2: If the visual congestion level confidence score does not fall into the first or second interval, and the correlation is not a numerical conflict relationship, then a comprehensive congestion value is obtained based on the visual congestion level confidence score and the state congestion level, and the congestion condition identification result and the corresponding fusion confidence score are determined. Optionally, when the visual congestion level confidence score does not fall into the first or second interval, this embodiment of the invention determines whether the correlation between the visual congestion level confidence score and the state congestion level is a numerical conflict relationship, and then calculates the congestion condition identification result and the fusion confidence score of the road ahead based on the numerical conflict relationship.
[0061] Specifically, the embodiments of the present invention involve a "conflict resolution and collaborative verification" mode, which clarifies two core numerical conflict relationship scenarios. For different conflict scenarios, it outputs refined congestion conditions corresponding to the numerical conflict relationship: ① The visual congestion level is high and the visual congestion level is highly reliable, but the state congestion level is consistently very low (such as congestion ahead but the vehicle is in a clear lane). The fusion center will analyze the reasons (such as whether it is congestion ahead but the vehicle is already in a slow-moving clear lane?), which may delay the judgment or output a refined result of "congestion ahead, this lane is clear".
[0062] ② If the level of congestion is very high, but the visual level of congestion is low and the visual level of congestion is highly reliable, the system may be triggered to check whether the vehicle status sensor is faulty, or consider aggressive driving caused by non-congestion (such as aggressive overtaking).
[0063] When the credibility of visual congestion does not fall into the first or second interval, such as when the credibility is in the middle range of 0.3-0.9, and there is no obvious numerical conflict among the three indicators, a single index is no longer used directly. Instead, the algorithm is used to fuse the credibility of visual congestion and the state congestion to obtain a comprehensive congestion value, and then the working condition label is matched and the fusion confidence is calculated.
[0064] In this embodiment of the invention, the fusion process is achieved using the Kalman filter algorithm.
[0065] Specifically, the visual congestion confidence (V_Confidence) and the state congestion index (S_Congestion_Index) are used as the two input observations of the Kalman filter. Through the "prediction-update" loop of the filter, the optimal filtered state estimate (the comprehensive confidence value of the congestion judgment) and the filter error covariance (the error quantification value of the filtering result) are obtained.
[0066] The filtered state estimate is a continuous value from 0 to 1, which is used as the adoption weight for the visual congestion level. The corresponding state congestion level adoption weight is the filtered state estimate. The weighted fusion formula is: Comprehensive congestion value = filtered state estimate × visual congestion level + (1 - filtered state estimate) × state congestion level.
[0067] The overall congestion value remains a continuous value from 0 to 1. According to the congestion level thresholds in Table 1 (0-0.3 smooth traffic, 0.3-0.5 basically smooth traffic, etc.), the corresponding congestion condition identification result is matched. This congestion condition identification result can be a simple Boolean flag (yes / no) or a richer label (such as "severe and complex congestion", "slow traffic congestion", "smooth traffic"). In addition, when generating the fusion identification result, in addition to outputting the congestion condition identification result, a final fusion confidence score will also be attached.
[0068] The fusion confidence score is calculated as follows: The filter error covariance quantifies the magnitude of the error in the Kalman filter result. The smaller the error, the more reliable the fusion result. The filter error covariance is normalized using the formula: Fusion confidence = 1 - α × Filter error covariance (α is the normalization coefficient, such as 2), ensuring that the fusion confidence is a continuous value of 0-1, which is positively correlated with the reliability of the filter result.
[0069] The final comprehensive operating condition identification result and confidence level are sent to the chassis control system. The high confidence level identification result allows the control system to switch to the corresponding comfort mode (such as "congestion comfort mode") more decisively, and can preload control parameters based on the "pre-congestion" signal to achieve a smooth and imperceptible mode transition.
[0070] In this embodiment of the invention, the filtered state estimate is used as a dynamic weighting weight to achieve adaptive fusion of visual congestion level and state congestion level. The fusion confidence is calculated based on the filter error covariance, and the error of the fusion result is quantified into an assessable confidence level, which ensures the objectivity and accuracy of the fusion confidence and provides a reliable decision basis for downstream systems.
[0071] As an optional embodiment, step S104 above includes: Step S1041: Extract operation frequency feature data and dynamic change feature data from the driving status data; Step S1042: The operation frequency feature data and dynamic change feature data are fused to obtain the state congestion level.
[0072] Optionally, in this embodiment of the invention, a vehicle state recognition module is provided: processing internal vehicle signals and calculating state indicators. Further, operation frequency feature data and dynamic change feature data are extracted from the vehicle driving state data. The operation frequency feature data includes: pedal switching frequency (number of times the accelerator / brake pedal is switched per unit time), steering wheel fine-tuning activity rate (number of times the steering wheel is adjusted at small angles per unit time), etc.; the dynamic change feature data includes: longitudinal acceleration standard deviation (reflecting the degree of fluctuation in vehicle speed), lateral angular velocity change intensity (reflecting the frequency of vehicle steering), etc. Both types of features are calculated from the driving state data collected by the vehicle state sensing unit and are core features reflecting the correlation between vehicle driving dynamics and congestion.
[0073] We employ weighted fusion methods (such as the analytic hierarchy process to determine the weights of each feature) or machine learning fusion methods (such as logistic regression or random forest) to fuse operation frequency feature data and dynamic change feature data. After fusion, we normalize the results to ensure that the state congestion level is a continuous value of 0-1, which is consistent with the quantification scale of visual congestion level.
[0074] In this embodiment of the invention, a multi-feature fusion method is used to obtain the state congestion level, which avoids the one-sidedness of a single feature (such as the inability of pedal frequency to reflect vehicle speed fluctuations). It integrates multi-dimensional dynamic information of vehicle driving, thereby improving the reliability and generalization ability of the state congestion level.
[0075] like Figure 2 As shown, Figure 2A schematic diagram of a second embodiment of a road congestion condition identification system provided by the present invention is shown. The system includes the following core modules: Visual Large Model Module: Equipped with a pre-trained visual large model, used to process images and output perception results end-to-end.
[0076] Vehicle Status Recognition Module: Receives vehicle internal signals from vehicle status sensors, processes the vehicle internal signals, and calculates status indicators.
[0077] Event-level fusion decision center: Receives high-level decision information from the visual large model module and the vehicle state recognition module, executes fusion logic, and outputs the final event judgment.
[0078] In addition, both the visual large model module and the vehicle status recognition module contain feature extraction sub-modules; the visual large model perception module contains a congestion index prediction model, which is used to output the congestion index.
[0079] Among them, fusion occurs at a high-level decision-making level. It is an arbitration and coordination of the judgment results of the two intelligent agents (visual big model perception module and vehicle status recognition module), which is more in line with human decision-making logic and has a better fusion effect.
[0080] Figure 3 A schematic diagram of an embodiment of a road congestion condition identification device according to the present invention is shown. The device includes: The first acquisition module 301 is used to acquire image and video frames of the vehicle, wherein the image and video frames are used to represent the congestion conditions of the road ahead of the vehicle. The first module 302 is used to obtain the visual congestion index of the road in front of the vehicle based on the image and video frames. The second acquisition module 303 is used to acquire vehicle driving status data; The second module 304 is used to obtain the degree of congestion on the road ahead based on driving status data; The third module 305 is used to obtain the congestion condition identification result and fusion confidence score of the road ahead based on the visual congestion index and the state congestion degree. The fusion confidence score is used to characterize the credibility of the congestion condition identification result.
[0081] This invention acquires vehicle image and video frames, then obtains the visual congestion index of the road ahead based on these frames; next, it acquires vehicle driving status data, and obtains the state congestion level of the road ahead based on this data; finally, it obtains the congestion condition identification result and fusion confidence score of the road ahead based on the visual congestion index and state congestion level. This invention, through the visual congestion index obtained from image and video frames, enables the road congestion condition identification process to perform an end-to-end comprehensive assessment of the "congestion and complexity" of the overall traffic scene, much like a human. This avoids the incompleteness and shortsightedness of manually designed features in traditional methods. Furthermore, by combining the visual congestion index with the state congestion level obtained from vehicle driving status data, the vehicle's intelligent driving system can intelligently determine when to "trust the eyes" and when to "trust the senses," greatly improving the system's intelligence and reliability. The arbitration and collaborative judgment results of the visual congestion index and state congestion level are more in line with human decision-making logic, resulting in a better fusion effect. This solves the problem in related technologies where a deep understanding of the road scene is insufficient to output the most accurate vehicle driving congestion condition.
[0082] In one alternative approach, the first obtaining module 302 is used to input image and video frames into a target visual model and output a visual congestion index, wherein the target visual model is used to perform visual dimension congestion analysis on the image and video frames.
[0083] In one alternative approach, the training process of the target visual model includes: The third acquisition module is used to acquire the vehicle's image training data frame and the actual congestion label of the road ahead carried by the image training data frame. The input module is used to input image training data into the initial visual model and output congestion prediction labels. The comparison module compares the actual congestion labels with the predicted congestion labels. When the error between the actual and predicted congestion labels exceeds an error threshold, the module adjusts the model parameters of the initial visual model until the error is less than or equal to the error threshold, thus obtaining the target visual model.
[0084] In one alternative approach, the actual congestion labels of the road ahead carried in the image training data frames are obtained, including: The fourth acquisition module is used to acquire the first driving trajectory data obtained after the vehicle has driven on the road ahead carried by the image training data frame; The fourth module is used to obtain the vehicle's average travel speed based on the first travel trajectory data; The fifth acquisition module is used to acquire the second driving trajectory data obtained when the vehicle is traveling at a constant speed during a period when the road ahead is not congested, as carried in the image training data frame. The fifth module is used to obtain the vehicle's free-flow travel speed based on the second driving trajectory data; The sixth module is used to obtain the actual congestion label based on the ratio of average travel speed to free-flow travel speed.
[0085] In one alternative approach, the visual congestion index includes visual congestion level and visual congestion level credibility. Visual congestion level credibility is used to characterize the trustworthiness of visual congestion level. Visual congestion level credibility is obtained by combining the error between each actual congestion label and each predicted congestion label with the surrounding environmental information of the road ahead of the vehicle. The third module 305 is used to identify the congestion condition of the road ahead by taking the congestion condition corresponding to the visual congestion level or the state congestion level when the visual congestion level confidence falls within the confidence value range; and to obtain the congestion condition identification result and fusion confidence of the road ahead based on the correlation between the visual congestion level, the visual congestion level confidence and the state congestion level when the visual congestion level confidence does not fall within the confidence value range.
[0086] In one alternative approach, the confidence value range includes a first range and a second range, wherein the first range is used to characterize the confidence level of visual congestion as being higher than a first threshold, and the second range is used to characterize the confidence level of visual congestion as being lower than a second threshold. The third module 305 is further configured to, if the visual congestion level confidence level falls into the first interval, use the congestion condition corresponding to the visual congestion level as the congestion condition identification result of the road ahead; if the visual congestion level confidence level falls into the second interval, use the congestion condition corresponding to the state congestion level as the congestion condition identification result of the road ahead.
[0087] In one optional approach, the third obtaining module 305 is further configured to: if the visual congestion level confidence does not fall into the first interval and the second interval, and the correlation is a numerical conflict relationship, then obtain the congestion condition corresponding to the numerical conflict relationship as the congestion condition identification result of the road ahead; if the visual congestion level confidence does not fall into the first interval and the second interval, and the correlation is not a numerical conflict relationship, then obtain a comprehensive congestion value based on the visual congestion level confidence and the state congestion level, and determine the congestion condition identification result corresponding to the comprehensive congestion value and the corresponding fusion confidence.
[0088] In an optional approach, the third module 305 is further configured to substitute the visual congestion confidence level and the state congestion level into the Kalman filter algorithm for calculation to obtain the filtered state estimate and the filter error covariance; based on the filtered state estimate, the visual congestion level and the state congestion level are weighted and fused to obtain the comprehensive congestion value, and the corresponding congestion condition identification result is determined; the filter error covariance is normalized to obtain the fusion confidence level.
[0089] In one alternative approach, the second obtaining module 304 is used to extract operation frequency feature data and dynamic change feature data from the driving status data; and to fuse the operation frequency feature data and dynamic change feature data to obtain the state congestion level.
[0090] Figure 4 The diagram illustrates a structural schematic of an embodiment of the computer device of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the computer device.
[0091] like Figure 4 As shown, the computer device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.
[0092] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408. Communication interface 404 is used to communicate with other network elements, such as clients or other servers. The processor 402 executes program 410, specifically performing the relevant steps in the above-described embodiment of the method for estimating the driving performance and power generation capacity of the motor.
[0093] Specifically, program 410 may include program code, which includes computer-executable instructions.
[0094] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The computer device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0095] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0096] This invention also provides a vehicle, such as... Figure 5 , Figure 5 This is a structural block diagram of a vehicle provided in an optional embodiment of this disclosure, such as... Figure 5 As shown, the vehicle includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.
[0097] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0098] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0099] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0100] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0101] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0102] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0103] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. Similarly, for the sake of brevity and to aid in understanding one or more aspects of the invention, in the description of exemplary embodiments of the invention above, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0104] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.
[0105] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A method for identifying road congestion conditions, characterized in that, The method includes: Acquire image and video frames of the vehicle, wherein the image and video frames are used to characterize the congestion conditions on the road ahead of the vehicle; Based on the image and video frames, the visual congestion index of the road ahead of the vehicle is obtained; Obtain the driving status data of the vehicle; Based on the driving status data, the degree of congestion on the road ahead is obtained; Based on the visual congestion index and the state congestion level, the congestion condition identification result of the road ahead and the fusion confidence score are obtained, wherein the fusion confidence score is used to characterize the credibility of the congestion condition identification result.
2. The method according to claim 1, characterized in that, The step of obtaining the visual congestion index of the road ahead of the vehicle based on the image and video frames includes: The image / video frame is input into the target visual model, and the visual congestion index is output. The target visual model is used to perform visual dimension congestion analysis on the image / video frame.
3. The method according to claim 2, characterized in that, The training process of the target visual model includes: Acquire the image training data frame of the vehicle and the actual congestion label of the road ahead carried by the image training data frame; The image training data is input into the initial visual model, which outputs congestion prediction labels. The actual congestion label is compared with the predicted congestion label. If the error between the actual congestion label and the predicted congestion label is greater than the error threshold, the model parameters of the initial visual model are adjusted until the error is less than or equal to the error threshold, thereby obtaining the target visual model.
4. The method according to claim 3, characterized in that, Obtaining the actual congestion labels of the road ahead carried in the image training data frames includes: The first driving trajectory data obtained after the vehicle has traveled on the road ahead carried by the image training data frame is acquired; Based on the first driving trajectory data, the average travel speed of the vehicle is obtained; Acquire second driving trajectory data when the vehicle is traveling at a constant speed during a period when the road ahead is not congested, as carried in the image training data frame; Based on the second driving trajectory data, the free-flow travel speed of the vehicle is obtained; The actual congestion label is obtained based on the ratio of the average travel speed to the free-flow travel speed.
5. The method according to claim 1, characterized in that, The visual congestion index includes visual congestion level and visual congestion level credibility. The visual congestion level credibility is used to characterize the trustworthiness of the visual congestion level. The visual congestion level credibility is obtained by combining the error between each actual congestion label and each predicted congestion label with the surrounding environmental information of the road ahead of the vehicle. The process of obtaining the congestion condition identification result and fusion confidence score of the road ahead based on the visual congestion index and the state congestion level includes: When the credibility of the visual congestion level falls within the credibility value range, the congestion condition corresponding to the visual congestion level or the state congestion level is used as the congestion condition identification result of the road ahead. When the credibility of the visual congestion level does not fall within the credibility value range, the congestion condition identification result of the road ahead and the fusion confidence level are obtained based on the correlation between the visual congestion level, the credibility of the visual congestion level and the state congestion level.
6. The method according to claim 5, characterized in that, The confidence level range includes a first range and a second range, wherein the first range is used to characterize that the confidence level of visual congestion is higher than a first threshold, and the second range is used to characterize that the confidence level of visual congestion is lower than a second threshold. When the credibility of the visual congestion level falls within the credibility value range, the congestion condition corresponding to the visual congestion level or the state congestion level is used as the congestion condition identification result of the road ahead, including: If the credibility of the visual congestion level falls within the first interval, the congestion condition corresponding to the visual congestion level is used as the congestion condition identification result of the road ahead. If the credibility of the visual congestion level falls within the second interval, the congestion condition corresponding to the state congestion level is used as the congestion condition identification result of the road ahead.
7. The method according to claim 5, characterized in that, When the visual congestion level confidence score does not fall within the confidence score range, based on the correlation between the visual congestion level, the visual congestion level confidence score, and the state congestion level, the congestion condition identification result of the road ahead and the fusion confidence score are obtained, including: If the credibility of the visual congestion level does not fall into the first interval or the second interval, and the correlation is a numerical conflict relationship, then the congestion condition corresponding to the numerical conflict relationship is obtained as the congestion condition identification result of the road ahead. If the visual congestion level confidence does not fall into the first interval or the second interval, and the correlation relationship is not a numerical conflict relationship, then a comprehensive congestion value is obtained based on the visual congestion level confidence and the state congestion level, and the congestion condition identification result corresponding to the comprehensive congestion value and the corresponding fusion confidence are determined.
8. The method according to claim 7, characterized in that, The process of obtaining a comprehensive congestion value based on the visual congestion level confidence level and the state congestion level, and determining the congestion condition identification result corresponding to the comprehensive congestion value and the corresponding fusion confidence level, includes: The reliability of the visual congestion level and the state congestion level are substituted into the Kalman filter algorithm for calculation to obtain the filtered state estimate and the filter error covariance. Based on the filtered state estimate, the visual congestion level and the state congestion level are weighted and fused to obtain a comprehensive congestion value, and the corresponding congestion condition identification result is determined. The covariance of the filtering error is normalized to obtain the fusion confidence score.
9. The method according to claim 1, characterized in that, The process of determining the level of road congestion ahead based on the driving status data includes: Extract the operation frequency feature data and dynamic change feature data from the driving status data; The frequency of operation feature data and the dynamic change feature data are fused to obtain the state congestion level.
10. A device for identifying road congestion conditions, characterized in that, The device includes: The first acquisition module is used to acquire image and video frames of the vehicle, wherein the image and video frames are used to characterize the congestion conditions of the road ahead of the vehicle. The first obtaining module is used to obtain the visual congestion index of the road in front of the vehicle based on the image video frame; The second acquisition module is used to acquire the driving status data of the vehicle; The second obtaining module is used to obtain the degree of congestion of the road ahead based on the driving status data; The third module is used to obtain the congestion condition identification result and fusion confidence score of the road ahead based on the visual congestion index and the state congestion degree, wherein the fusion confidence score is used to characterize the credibility of the congestion condition identification result.
11. A vehicle, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the method for identifying road congestion conditions as described in any one of claims 1 to 9.