A road environment intelligent optimization method based on diffusion model

Through the intelligent optimization method based on the diffusion model, the problem of limited applicability of rural road environment optimization is solved, efficient and accurate road environment optimization is achieved, and the safety of rural roads is improved.

CN119380304BActive Publication Date: 2025-06-06TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411445573.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-06-06
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

The existing technology lacks precise optimization methods suitable for various road environments, resulting in limited applicability of rural road environment optimization, relying on manual operations and inefficient efficiency.

Method used

The road environment intelligent optimization method based on diffusion model is adopted, and by collecting natural driving data, semantic segmentation, feature extraction and interpretability model establishment, the optimized road driving environment image is generated to assist road designers in optimization.

Benefits of technology

It realizes efficient and precise road environment optimization, and the generated image quality is excellent, which can be directly used for road design and optimization, improving the safety of rural roads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380304B_ABST
    Figure CN119380304B_ABST
Patent Text Reader

Abstract

The present invention discloses a road environment intelligent optimization method based on a diffusion model, which belongs to the technical field of road environment optimization, and includes the following steps: S1, collecting natural driving data, using a semantic segmentation network to perform semantic segmentation on the natural driving data, and obtaining different semantic elements; S2, using Python programming to extract the area and position information of the semantic elements; S3, establishing an interpretable model for quantifying the influence of semantic elements on the road driving environment; S4, establishing an intelligent optimization model for the road driving environment based on a diffusion model, completing the training and sampling of the model, and generating an adjusted and optimized road driving environment image. The present invention adopts the above-mentioned road environment intelligent optimization method based on a diffusion model, which has high efficiency and good image generation quality, and can directly generate an optimized and accurate visual road environment image, assisting road designers in designing and optimizing the visual environment of rural roads, and improving the safety of rural roads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road environment optimization, and in particular to a road environment intelligent optimization method based on a diffusion model. Background Art

[0002] With the gradual increase in the number of cars, road traffic safety has attracted great attention and concern. Speeding is an important factor leading to traffic accidents. Although the traffic volume on rural roads is small, accidents still occur and are often serious. Through the rural environment itself, drivers can be guided to accelerate or decelerate in time according to the road environment, thereby improving traffic safety. This road that induces safe driving behavior through an appropriate road environment is called a self-explanatory road.

[0003] Numerous studies have confirmed the role of visual road environment in speed guidance. For example, lane markings play a positive role in indicating linear changes in roads, guiding drivers' sight, and improving driving safety. Whether guardrails are installed on narrow shoulder sections significantly affects driving speed and lateral position. The presence of guardrails is negatively correlated with driving speed and causes drivers to drive closer to the centerline of the road. According to a public preference survey, green landscapes can reduce driving speed and improve driving quality and safety. The rural road environment mainly includes road alignment, road facilities, and surrounding landscapes. Due to the limitations of cultivated land and terrain, rural roads have many intersections and ramps, limited space for road alignment design, and limited construction funds. Therefore, considering the difficulty of changing the alignment design of rural roads, optimizing road facilities and surrounding landscapes (i.e., traffic signs, lane markings, guardrails, and vegetation) is a feasible way to ensure rural road safety. Existing studies have not yet proposed an accurate optimization method suitable for various road environment parts, resulting in limited applicability of the optimization method.

[0004] Emerging intelligent image generation algorithms, such as diffusion models, are potential solutions to this challenge. Generative adversarial networks (GANs) and diffusion models are two commonly used image generation algorithms in the field of computer vision. GAN can convert road environment images from night to day, thereby increasing background brightness and improving the accuracy of night vehicle detection. Ren et al. used CycleGAN to generate optimized images, saving labor, but there are still disadvantages such as long time consumption and unstable image quality. Although the diffusion model overcomes the above shortcomings of GAN and is superior to GAN in image generation quality and stability, there is still a lack of research specifically applying the diffusion model to road optimization, and the research on road environment optimization still relies on manual labor, and intelligent methods cannot achieve accurate optimization. Summary of the invention

[0005] The purpose of the present invention is to provide a road environment intelligent optimization method based on a diffusion model, which has high efficiency and good image generation quality. It can directly generate an optimized and accurate visual road environment image, assist road designers in the design and optimization of the rural road visual environment, and improve the safety of rural roads.

[0006] To achieve the above object, the present invention provides a road environment intelligent optimization method based on a diffusion model, comprising the following steps:

[0007] S1. Collect natural driving data and use the semantic segmentation network to perform semantic segmentation on the natural driving data to obtain different semantic elements;

[0008] S2. Use Python programming to extract the area and location information of semantic elements;

[0009] S3. Establish an interpretable model for quantifying the impact of semantic elements of road driving environment;

[0010] S4. On the basis of step S3, an intelligent optimization model of the road driving environment based on the diffusion model is established, the training and sampling of the model are completed, and an adjusted and optimized road driving environment image is generated.

[0011] Preferably, a low-cost data acquisition system is used in step S1, which includes a GARMIN GDR35 driving recorder, a GPS locator and a three-axis acceleration sensor. The main camera in the GARMIN GDR35 driving recorder is installed on the inner front windshield to collect the road environment from the driver's perspective, and its secondary camera faces the inside of the car to collect the driver's head information.

[0012] Preferably, in step S1, the semantic segmentation network includes an encoder and a decoder, and the encoder part adopts the ResNet50 network structure; the semantic segmentation network introduces a feature pyramid network, and the feature pyramid network performs lateral extraction and vertical fusion on the feature maps output by the four stages of ResNet50. The encoder of the feature pyramid network first upsamples the four branches, and finally integrates the four sampling results using the bagging integration algorithm; the semantic elements include road surface, markings, vegetation, signboards, and guardrails.

[0013] Preferably, the specific steps of step S2 are:

[0014] S201, color of the image after statistical semantic segmentation;

[0015] S202, matching various color information with road visual semantic elements;

[0016] S203: Analyze the location information of each road visual semantic element.

[0017] Preferably, in step S3, the driving environment semantic elements include 27 variables, of which 12 variables are extracted from the road linear layer, specifically: left near view curve length, left mid view curve length, left distant view curve length, right near view curve length, right mid view curve length, right distant view curve length, left near view curve curvature, left mid view curve curvature, left distant curve curvature, right near view curve curvature, right mid view curve curvature, right distant curve curvature;

[0018] The other 15 variables are extracted at the visual semantic layer: road area, shoulder area, marking area, vegetation area, corrugated beam guardrail area, concrete guardrail area, whether there is a signboard in the foreground area, whether there is a signboard in the middle ground area, whether there is a signboard in the distant ground area, the area of ​​the corrugated beam guardrail in the foreground, the area of ​​the corrugated beam guardrail in the middle ground, the area of ​​the corrugated beam guardrail in the distant ground, the area of ​​the concrete guardrail in the foreground, the area of ​​the concrete guardrail in the middle ground, and the area of ​​the concrete guardrail in the distant ground.

[0019] Preferably, step S3 specifically inputs 27 variables and driving speed information, outputs a speed regression model through an XGBoost algorithm, and then outputs an interpretability analysis of the 27 variables through a SHAP algorithm.

[0020] Preferably, the training process of the vehicle speed regression model includes the following steps:

[0021] Initialize the model: set the initial prediction value to the global average;

[0022] Iterative training: For each iteration, the gradient of the current model and the negative gradient of the loss function are calculated;

[0023] Construct a decision tree: Use a greedy algorithm to select the best split point and recursively construct a decision tree model;

[0024] Update model: Update the model's prediction value based on the negative gradient of the loss function and the output of the decision tree model;

[0025] Regularization: Adjust the complexity and parameter values ​​of the model according to the regularization term;

[0026] Termination condition: determine whether the iteration condition is met;

[0027] Output the final model: All trained decision tree models are combined into the final integrated model, and the SHAP algorithm is used for interpretability analysis.

[0028] Preferably, the specific training process of the diffusion generation model in step S4 includes two steps: training and sampling. The algorithm of the training process is as follows: extract samples from the data, select any time t from 1 to T, and set x 0and t are passed to the diffusion generation model, which samples a random noise and adds it to time x 0 And get x t , then x t The L2 loss function is used to continuously calculate the gradient and update the weights to complete the training of the neural network.

[0029] The algorithm of the sampling process is as follows: sample x from the standard normal distribution T , repeat the following process in sequence from time T, T-1, T-2, ..., 2, 1, sampling z from the standard normal distribution, combining x according to the neural network model t and z to calculate x t-1 , and finally returns x after the loop ends. 0 .

[0030] Therefore, the present invention adopts the above-mentioned road environment intelligent optimization method based on the diffusion model, performs feature extraction on semantic elements such as markings, vegetation, signboards, guardrails, etc. in the road driving environment, extracts the area and position information of various semantic elements, and interpretably analyzes the influence of the area and position information of various semantic elements on the driving speed based on the interpretable algorithm model of XGBoost combined with SHAP, and uses this as a basis and prompt to guide road designers to reasonably adjust and optimize semantic elements such as markings, vegetation, signboards, guardrails, etc. in the road driving environment at reasonable positions.

[0031] An intelligent optimization method for road driving environment based on diffusion model was established. Road designers can reasonably adjust and optimize one or more semantic elements at specific locations in the road driving environment based on the basis and prompts of the interpretable algorithm model of XGBoost combined with SHAP, and automatically generate adjusted and optimized road driving environment images through the diffusion model. The research results are helpful to assist road designers in designing and optimizing road driving environments, guiding drivers to adopt reasonable driving behaviors, and are of great significance to improving drivers' driving experience and driving safety.

[0032] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a technical roadmap of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0034] Figure 2 It is a schematic diagram of a semantic segmentation network structure of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention; Figure 2 (a) is a schematic diagram of the encoder structure; Figure 2 (b) is a schematic diagram of the decoder structure.

[0035] Figure 3 It is a schematic diagram of a road linear layer of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0036] Figure 4 It is a schematic diagram of a diffusion model generation process of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0037] Figure 5 This is a specific area vegetation migration example 1 of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0038] Figure 6 This is a second example of vegetation migration in a specific area of ​​an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0039] Figure 7 This is a specific area guardrail migration example 1 of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0040] Figure 8 This is a second example of guardrail migration in a specific area of ​​an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0041] Fig. 9 This is a specific area marking migration example 1 of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0042] Fig.10 This is a second example of migration of markings in a specific area of ​​an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0043] Fig.11 This is a specific area signboard migration example 1 of an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention;

[0044] Fig.12 This is Example 2 of specific area sign migration in an embodiment of a road environment intelligent optimization method based on a diffusion model of the present invention. DETAILED DESCRIPTION

[0045] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.

[0046] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.

[0047] Embodiment 1

[0048] like Figure 1 As shown, the present invention provides a road environment intelligent optimization method based on a diffusion model, comprising the following steps:

[0049] S1. Collect natural driving data. The driving data is road driving scene data collected on two-way two-lane rural roads in Mangkam, Nyingchi, Basu, Lhasa and other areas in Tibet.

[0050] A low-cost data acquisition system is used in the data collection process. The system includes a GARMIN GDR35 driving recorder, a GPS locator and a three-axis acceleration sensor. The main camera in the GARMIN GDR35 driving recorder is installed on the inner front windshield to collect the road environment from the driver's perspective. Its secondary camera faces the inside of the car to collect the driver's head information. Through this data acquisition system, the information collected by the GPS module, the three-axis acceleration sensor and the driving recorder is compressed into an AVI format video file and saved in the data storage module. Subsequently, the road visual image and its matching geographic coordinates, vehicle speed, three-axis acceleration and other related information are decompressed from the video file by writing software.

[0051] The driving video collected in the natural driving data collection experiment lasted more than 40 hours, which was cut and saved into multiple AVI videos of about 2 minutes and 55 seconds, and further cut and saved into multiple PNG images with a time interval of 1 second. The driving speed in the rural road scene data is retained as an integer in km / h. The resolution of the rural road visual scene image is 1920*1080 and is stored in PNG format.

[0052] The semantic segmentation network is used to perform semantic segmentation on the road environment images in the data. The schematic diagram of the semantic segmentation network structure is shown in the figure. Figure 2 As shown in the figure, it is segmented into specific semantic categories such as road surface, marking, vegetation, signboard, guardrail, etc. The semantic segmentation network includes an encoder and a decoder. The encoder part adopts the ResNet50 network structure.

[0053] In order to retain scene features of different granularities and improve the speed of feature extraction, the semantic segmentation network introduces a feature pyramid network. During the image downsampling process, the feature pyramid network extracts and fuses the feature maps output by the four stages of ResNet50 laterally, retaining the detail features in the shallow feature maps and the macro features in the deep feature maps. The encoder of the feature pyramid network first upsamples the four branches, and finally integrates the four sampling results using the bagging integration algorithm. This processing method makes full use of multi-level feature information, thereby improving the performance and accuracy of semantic segmentation.

[0054] S2. Use Python programming to extract the area and location information of semantic elements; the specific steps are:

[0055] S201, counting the colors of the image after semantic segmentation; traversing each pixel of the image and recording the RGB values ​​of all colors in the image. This can be achieved by calling the OpenCV library function integrated in Python to read the image file and traversing each pixel.

[0056] S202, matching various color information with road visual semantic elements; calculating the number of pixels of each color in the image to obtain the area of ​​the color, which is the area of ​​road visual environment elements such as roads, shoulders, markings, vegetation, and guardrails in the original image.

[0057] S203, analyzing the location information of each road visual semantic element. By calculating the area of ​​each color pixel in the image in the "near view", "mid view" and "distant view" regions, we analyze the impact of the road visual environment elements such as roads, shoulders, markings, vegetation, guardrails, etc. on the driving speed when they are in a certain area of ​​"near view", "mid view" and "distant view", so as to guide us to determine the migration and design of road visual environment elements such as roads, shoulders, markings, vegetation, guardrails, etc. in a certain area of ​​"near view", "mid view" and "distant view".

[0058] S3. Establish an interpretable model for quantifying the impact of semantic elements of road driving environment;

[0059] In step S3, the driving environment semantic elements include 27 variables, of which 12 variables are extracted from the road linear layer, specifically: left near view curve length, left mid view curve length, left distant view curve length, right near view curve length, right mid view curve length, right distant view curve length, left near view curve curvature, left mid view curve curvature, left distant curve curvature, right near view curve curvature, right mid view curve curvature, right distant curve curvature;

[0060] The other 15 variables are extracted at the visual semantic layer: road area, shoulder area, marking area, vegetation area, corrugated beam guardrail area, concrete guardrail area, whether there is a signboard in the foreground area, whether there is a signboard in the middle ground area, whether there is a signboard in the distant ground area, the area of ​​the corrugated beam guardrail in the foreground, the area of ​​the corrugated beam guardrail in the middle ground, the area of ​​the corrugated beam guardrail in the distant ground, the area of ​​the concrete guardrail in the foreground, the area of ​​the concrete guardrail in the middle ground, and the area of ​​the concrete guardrail in the distant ground.

[0061] Road Linear Layer

[0062] The road visual environment consists of the lane area and the roadside area. During driving, there is always a "lane" in the driver's field of vision. The driver can perceive this "lane" according to the specific road environment. This "lane" is called the driver's visual lane. Yu Bo et al.

[23] established a driver's visual lane model based on the Catmull-Rom spline curve. The Catmull-Rom spline curve is a cubic interpolation spline curve that can present any shape, so it can be used to fit the lane line of the road. Compared with other quadratic or cubic curves (such as Cubic Bezier spline curves, Cubic B spline curves, etc.), the Catmull-Rom spline curve has good elasticity and smoothness, can produce a smooth curve path, smoothly fit the curvature change of the lane line, and avoid the sudden change or discontinuity of the lane line. In addition, the Catmull-Rom spline curve has a high fitting accuracy, can flexibly adapt to different road shapes and lane line changes, and can accurately fit various complex road curves, including S-shaped curves. Therefore, the Catmull-Rom spline curve is used to fit the left and right boundaries of the lane of the curved section of the rural road.

[0063] like Figure 3 As shown in Figure 1, the road linear layer constructs a coordinate system with the lower left corner of the driver's field of view as the origin. The left spline curve has four control points (P 1L ,P 2L ,P 3L ,P 4L ), the right spline also has four control points (P 1R ,P 2R ,P 3R ,P 4R ) to fit the shape of the road boundary. iL is the cumulative length of the left road boundary at the control point numbered i, S iR is the cumulative length of the right road boundary at the control point numbered i. iL is the tangent slope of the left road boundary at the control point numbered i, f iR is the tangent slope of the right road boundary at the control point numbered i.

[0064] The horizontal lines connecting the corresponding control points in the left and right spline curves divide the road into three characteristic areas: "near view", "mid view" and "distant view". These characteristic areas have different visual characteristics, which help to quantify the driver's visual perception of the road line shape. At the same time, the horizontal lines connecting the control points divide the road into three characteristic areas: "near view", "mid view" and "distant view", and also play a role in locating environmental elements such as signs and guardrails.

[0065] Among them, the curve lengths and curvatures of the three characteristic areas of the left road boundary, namely the left near view curve length, the left mid view curve length, and the left far view curve length, are extracted as vS iL (i=1, 2, 3), the curvature of the left near view curve, the curvature of the left mid view curve, and the curvature of the left distant view curve, denoted as vK iL (i=1, 2, 3). At the same time, the curve lengths and curvatures of the three characteristic areas of the right road boundary, namely, the right near view curve length, the right mid view curve length, and the right far view curve length, are extracted as shape parameters, denoted as vS iR (i=1, 2, 3), the curvature of the right near view curve, the curvature of the right mid view curve, and the curvature of the right distant view curve, denoted as vK iR (i=1, 2, 3). A total of 12 variables are used as shape parameters. The specific calculation formula is as follows:

[0066] v iL =S (i+1)L -S iL ;

[0067]

[0068] v iR =S (i+1)R -S iR ;

[0069]

[0070] Among them, S iL is the cumulative length of the left road boundary at the control point numbered i, S (i+1)L -S iL (i=1, 2, 3) is the curve length of the left road boundary in the near view, mid view, and far view;

[0071] S iR is the cumulative length of the right road boundary at the control point numbered i, S (i+1)R -S iR (i=1, 2, 3) is the curve length of the right road boundary in the near view, mid view, and far view;

[0072] f iL is the tangent slope of the left road boundary at the control point numbered i, That is, the curvature of the curve of the left road boundary in the near, middle and far views;

[0073] f iR is the tangent slope of the right road boundary at the control point numbered i, That is, the curvature of the curve of the right road boundary in the near, middle and distant views.

[0074] Visual semantic layer

[0075] Road visual environment elements usually include road surface, shoulder, marking, vegetation, signboard, guardrail, etc. Road visual environment elements directly or indirectly affect the driver's driving speed through the above factors. For example, lanes and shoulders are important factors for drivers to judge the appropriate driving position and state, which directly affect the driving speed; roadside vegetation affects the driver's field of vision and attention, which indirectly affects the driving speed; guardrails, signboards, etc. are important restrictions or prompts during driving, which directly affect the driver's perception, judgment and actual operation, and thus affect the driving speed.

[0076] The present invention extracts the areas of pavement, shoulders, markings, vegetation, and two types of guardrails as parameters of the visual semantic layer, which are expressed as AP (Area of ​​pavement), AS (Area of ​​shoulders), AL (Area of ​​lane line), AV (Area of ​​vegetation), AG1. (Area of ​​the first guardrail type), and AG2. (Area of ​​the second guardrail type). In the above road linear layer, a method is proposed to divide the road into three characteristic areas of "near view", "middle view" and "far view" by horizontally connecting the four control points of the Catmull-Rom spline curve. In order to further quantitatively analyze the impact of the specific positions of semantics such as signs and guardrails on the driving speed, three pieces of information are extracted at the same time: whether there are signs in the near view, whether there are signs in the middle view, and whether there are signs in the far view. At the same time, the areas of the two types of guardrails in the three characteristic areas of "near view", "middle view" and "far view" are extracted, which are expressed as AG1.near (Area of ​​the first guardrail type in "near scene"), AG1.middle (Area of ​​the first guardrail type in "middlescene"), AG1.far (Area of ​​the first guardrail type in "far scene"), AG2.near (Area of ​​the second guardrail type in "near scene"), AG2.middle (Area of ​​the second guardrail type in "middle scene"), and AG2.far (Area of ​​the second guardrail type in "far scene"). A total of 15 variables are used as semantic parameters.

[0077] The independent variables of the model in step S3 are 27 variables extracted from the road visual environment, and the dependent variable Y is the driving speed, which is a continuous variable. In order to analyze the specific impact of the 27 variables extracted from the road visual environment on the driving speed, it is necessary to establish a regression model of the 27 variables and the driving speed based on the XGBoost algorithm, and combine the SHAP algorithm to interpretably analyze the importance, specific impact and independence of the 27 variables (focusing on the area and location information of semantic elements such as markings, vegetation, signs, and guardrails) on the driving speed, and use this as a basis to guide road designers to use the diffusion model to select locations and adjust and optimize the semantic elements of the road driving environment.

[0078] The mathematical principle of the XGBoost algorithm mainly involves gradient boosting and regularization methods. XGBoost introduces regularization technology to control the complexity of the model and improve the generalization ability of the model. There are mainly two regularization terms: L1 regularization (Lasso) and L2 regularization (Ridge).

[0079] L1 regularization penalizes the complexity of the model by adding the sum of the absolute values ​​of the parameters to the loss function:

[0080]

[0081] Among them, Ω(F) represents the regularization term, θj represents the leaf node weight of the decision tree, J represents the number of leaf nodes, and γ is the regularization hyperparameter. L1 regularization helps with feature selection and can shrink the weights of some irrelevant features to zero, thereby improving the interpretability and generalization ability of the model.

[0082] L2 regularization penalizes the complexity of the model by adding the sum of the squares of the parameters to the loss function:

[0083]

[0084] Where λ is a regularization hyperparameter. L2 regularization can shrink the value of the parameter, reduce sensitivity to noise, and improve the robustness of the model.

[0085] The objective function of XGBoost consists of a loss function, a regularization term, and a decision tree model:

[0086]

[0087] Where T is the number of decision trees. We learn the optimal model parameters by minimizing the objective function.

[0088] The XGBoost algorithm has high prediction performance, but the XGBoost model itself does not provide an explanation for the prediction results. By combining the SHAP algorithm with XGBoost, the prediction results of the XGBoost model can be explained. The SHAP algorithm helps understand the contribution of each feature to the prediction of the XGBoost model, thereby revealing the internal mechanism of vehicle speed prediction regression.

[0089] Step S3 specifically inputs 27 variables and driving speed information, outputs the speed regression model through the XGBoost algorithm, and then outputs the interpretability analysis of the 27 variables through the SHAP algorithm. The variable input distribution table is shown in Table 3.1.

[0090] Table 3.1 Variable input distribution

[0091]

[0092]

[0093] The training process of the vehicle speed regression model includes the following steps:

[0094] Initialize the model: set the initial prediction value to the global average;

[0095] Iterative training: For each iteration, the gradient of the current model and the negative gradient of the loss function are calculated;

[0096] Construct a decision tree: Use a greedy algorithm to select the best split point and recursively construct a decision tree model;

[0097] Update model: Update the model's prediction value based on the negative gradient of the loss function and the output of the decision tree model;

[0098] Regularization: Adjust the complexity and parameter values ​​of the model according to the regularization term;

[0099] Termination condition: determine whether the iteration condition is met;

[0100] Output the final model: All trained decision tree models are combined into the final integrated model, and the SHAP algorithm is used for interpretability analysis.

[0101] The present invention evaluates the accuracy of the vehicle speed regression model based on the mean square error (MSE) indicator. The mean square error is the average value of the sum of the squares of the deviations between all predicted values ​​and actual values, and its calculation formula is as follows:

[0102]

[0103] A in the formula t Indicates the actual value, F t Represents the predicted value, and n is the total number of data samples.

[0104] The mean square error eliminates the effect of the positive and negative errors offsetting each other by squaring the difference between the actual value and the predicted value, thus improving the accuracy of the regression model detection. At the same time, when the deviation between the predicted value and the actual value is large, the mean square error will increase significantly due to the existence of the square term, because the mean square error regression model has a high degree of sensitivity.

[0105] The accuracy comparison of the XGBoost algorithm and several other common algorithms in establishing a driving speed regression model is shown in the following table:

[0106] Table 3.2 Comparison of driving speed regression model accuracy

[0107] algorithm Mean Square Error (MSE) unit Linear regression algorithm 106.70 <![CDATA[(km / h) 2 ]]> Random Forest Algorithm 35.15 <![CDATA[(km / h) 2 ]]> Support Vector Regression Algorithm 43.26 <![CDATA[(km / h) 2 ]]> XGBoost Algorithm 20.13 <![CDATA[(km / h) 2 ]]>

[0108] As shown in the table above, the XGBoost algorithm outperforms the linear regression algorithm, random forest algorithm, and support vector regression algorithm in terms of the accuracy of the vehicle speed regression model.

[0109] S4. On the basis of step S3, an intelligent optimization model of the road driving environment based on the diffusion model is established, the training and sampling of the model are completed, and an adjusted and optimized road driving environment image is generated.

[0110] Diffusion models have many advantages in terms of high-quality image generation, flexible conditional input, image editing and control capabilities, and robustness of image generation. This paper uses the diffusion model to adjust and migrate road visual environment elements of rural roads, generates high-quality road visual environment images, and effectively induces drivers' driving behavior.

[0111] The diffusion model consists of two processes: the forward process and the reverse process. Both the forward and reverse processes are based on Markov chains. The forward process first adds random noise to the real data (i.e., the initial image), and the reverse process learns and infers the distribution of the noise, thereby removing the noise and obtaining the final generated data (i.e., the generated image). The image generation process of the diffusion model is as follows: Figure 4 As shown:

[0112] The purpose of random noise injection is to add random perturbations to the real data (initial image) to make the generated image more diverse and random. The diffusion generation model will use this initial image as the starting point of the generation process and gradually generate the final image through a continuous iterative diffusion process.

[0113] The forward process is the process of adding noise. Figure 4 From right to left, x 0 to x TIt is a forward process of adding noise step by step. The noise at each step is known. The noise often adopts a predefined gradually decaying noise (Gaussian distribution is generally used). The forward process gradually adds noise from the initial image to generate a set of pure noise. Because the noise at each step in the process of adding noise is known, the image x at time t in the forward process t It can be obtained from the image x at the previous moment t-1 It is concluded that the process is a Markov process, and the conditional probability formula is as follows:

[0114]

[0115] Where: β is a predefined gradually decaying noise (Gaussian distribution is generally used). Through the continuous iterative derivation of the above two formulas, we can get x t and x 0 Relationship:

[0116] in:

[0117] The inverse process is the process of removing noise. The goal of the inverse process is to restore the original data from Gaussian noise, that is, Figure 4 This process can be regarded as the reverse operation of the generation process, and the noise sequence is restored by reverse iteration. We only need to get q(x t-1 / x t ), we can use a similar Markov process from x T The target image is derived. The diffusion generation model uses a neural network p θ (x t-1 / x t ) fits the inverse process q(x t-1 / x t ). Since the noise we add each time in the forward process is very small, p(x t-1 / x t ) is also close to a Gaussian distribution, so a neural network can be used for fitting. The mathematical formula is derived as follows:

[0118] p θ (x t-1 |x t )=N(x t-1 ;μ θ (x t ,t),∑ θ (x t ,t));

[0119]

[0120] Because the variance of the probability distribution is small during the denoising process, the diffusion generation model ignores the variance when fitting with the neural network and only fits the mean μ θ , the fitting results are as follows:

[0121]

[0122] where t and x t In the forward process, we know that μ θ And through p θ (x t-1 / x t ) Reverse deduction to get x t 、x t-1 Until x 0 The inverse process can explicitly understand how the generative model generates images, and provides the ability to edit and control the images by adjusting the parameters.

[0123] The specific training process of the diffusion generation model includes two steps: training and sampling. The algorithm of the training process is as follows: extract samples from the data, select any time t from 1 to T, and set x 0 and t are passed to the diffusion generation model, which samples a random noise and adds it to time x 0 And get x t , then x t The L2 loss function is used to continuously calculate the gradient and update the weights to complete the training of the neural network.

[0124] The algorithm of the sampling process is as follows: sample x from the standard normal distribution T , repeat the following process in sequence from time T, T-1, T-2, ..., 2, 1, sampling z from the standard normal distribution, combining x according to the neural network model t and z to calculate x t-1 , and finally returns x after the loop ends. 0 .

[0125] Final training results

[0126] The following is an example of image generation based on the road driving environment optimization based on the diffusion model. Figure 5-12 , where each picture consists of four parts from left to right: P1, P2, P3, and P4, where P1 is the road driving environment image to be optimized; the gray part in P2 is the specific position planned to be optimized in the P1 image; P3 is the ideal road driving environment image, i.e., the target image; and P4 is the final result after the road environment is optimized.

[0127] (1) Vegetation migration in specific areas:

[0128] Example Figure 5 shown.

[0129] The road driving environment image to be optimized in P1 mainly consists of mountains and sand dunes, with a speed of 56km / H. The speed limit on rural roads is 70km / h, and there is a problem that the driving speed is lower than the expected value. According to the interpretable model of the influence of the semantic elements of the road driving environment on the driving speed in step S3, the migration of vegetation has a positive impact on the driving speed. The road visual environment here is relatively open, and the overall environment is selected as the optimized position (the gray part of P2). P3 is the ideal vegetation image, that is, the target image. P4 is the generated result after the optimization of the final road driving environment, which increases the driver's driving speed, enriches the road environment content from the driver's perspective, makes it less likely for the driver to have problems such as inattention due to monotonous driving, and optimizes the driving experience.

[0130] Example 2 Figure 6 shown.

[0131] In the road driving environment image to be optimized in P1, one side is a hillside, and the vehicle speed is 44km / h, which means the driving speed is lower than expected. According to the explainability model of the influence of the semantic elements of the road driving environment on the driving speed in step S3, the migration of vegetation has a positive effect on the driving speed. A specific position is selected as the optimization position (the gray part of P2). P3 is the ideal vegetation image, i.e. the target image, and P4 is the result of the final road driving environment optimization, which increases the driver's driving speed, enriches the road environment content from the driver's perspective, makes it less likely for the driver to lose concentration due to monotonous driving, and optimizes the driving experience.

[0132] (2) Relocation of guardrails in specific areas:

[0133] Example Figure 7 shown.

[0134] One side of the road driving environment image to be optimized in P1 is mainly a continuous turning section with a speed of 69 km / h, and there is a problem that the driving speed is higher than the expected value. According to the interpretable model of the influence of the semantic elements of the road driving environment on the driving speed, the migration of guardrails has a negative impact on the driving speed. According to the interpretable model, the impact of concrete guardrails on driving speed is more significant than that of corrugated beam guardrails on driving speed. The road driving environment here is relatively safe, without special environments such as steep slopes and cliffs. The corrugated beam guardrail can meet the requirements. According to the interpretable model, the migration of guardrails in the "mid-ground" position has the greatest negative impact on driving speed, so the specific area in the "mid-ground" is selected as the optimization position (the gray part of P2). P3 is the ideal guardrail image, that is, the target image (corrugated beam guardrail). P4 is the generated result after the optimization of the final road driving environment, which reduces the driver's driving speed and ensures driving safety.

[0135] Example 2 Figure 8 shown.

[0136] In the image of the road driving environment to be optimized in P1, there are cliffs on both sides and a river on one side. It is a dangerous section of road with a speed of 57 km / h. There is a problem that the driving speed is higher than the expected value. According to the interpretable model of the influence of the semantic elements of the road driving environment on the driving speed, the migration of guardrails has a negative impact on the driving speed. According to the interpretable model, the influence of concrete guardrails on driving speed is more significant than that of corrugated beam guardrails on driving speed. The road driving environment here is relatively dangerous, and there are no special environments such as cliffs and rivers. The concrete guardrail can meet the requirements. According to the interpretable model, the migration of guardrails in the "mid-view" and "near-view" positions has the greatest negative impact on driving speed, so the specific areas in the "mid-view" and "near-view" are selected as the optimization positions (the gray part of P2). P3 is the ideal guardrail image, that is, the target image (concrete guardrail). P4 is the generated result after the optimization of the final road driving environment, which reduces the driver's driving speed and ensures driving safety.

[0137] (3) Migration of road markings in specific areas:

[0138] Example Fig. 9 shown.

[0139] The road driving environment image to be optimized in P1 mainly consists of hillsides and weeds, with a speed of 81km / h. The speed limit on rural roads is 70km / h, and there is a problem that the driving speed is higher than expected. According to the interpretable model of the influence of semantic elements of the road driving environment on driving speed, the migration sign has a negative impact on driving speed. A specific area is selected as the optimization position (the gray part of P2), and P3 is the ideal speed reduction marking image, i.e., the target image. P4 is the generated result after the final road driving environment is optimized, which reduces the driver's driving speed, ensures driving safety, enriches the road environment content from the driver's perspective, and makes it less likely for the driver to lose concentration due to monotonous driving.

[0140] Example 2 Fig.10 shown.

[0141] The road driving environment image to be optimized in P1 mainly has hillsides on both sides, the speed is 79km / h, and the speed limit on rural roads is 70km / h. There is a problem that the driving speed is higher than the expected value. According to the explainable model of the influence of the semantic elements of the road driving environment on the driving speed, it may be caused by the lack of lane lines and the negative impact of the migration of lane lines on the driving speed. A specific area is selected as the optimization position (the gray part of P2), and P3 is the ideal lane line image, that is, the target image. P4 is the generated result after the optimization of the final road driving environment, which reduces the driver's driving speed and ensures driving safety. At the same time, it enriches the road environment content from the driver's perspective and reduces the driver's inattention and other problems.

[0142] (4) Relocation of signboards in specific areas:

[0143] Example Fig.11 shown.

[0144] The road driving environment image to be optimized in P1 is mainly an open plain, with a speed of 85km / H and a speed limit of 70km / h on rural roads. There is a problem that the driving speed is higher than the expected value. According to the interpretable model of the influence of semantic elements of the road driving environment on the driving speed, the migration of signs has a negative impact on the driving speed. The signs clearly indicate information such as speed limits. Here, the road driving environment image can be optimized by migrating signs. According to the interpretable model in Chapter 3, the presence or absence of signs in the "mid-view" and "near-view" areas has a greater impact on the driving speed. Here, the "near-view" area is selected as the optimization position (the gray part of P2), and P3 is the ideal speed limit sign image, that is, the target image. P4 is the generated result after the final road driving environment is optimized, which reduces the driver's driving speed, ensures driving safety, enriches the road environment content from the driver's perspective, and makes it less likely for the driver to lose concentration due to monotonous driving.

[0145] Example 2 Fig.12 shown.

[0146] The road driving environment image to be optimized in P1 is mainly an open plain, with a speed of 56km / h. There is a sharp turn on the right ahead, and the driving speed is higher than expected. According to the interpretable model of the influence of semantic elements of the road driving environment on driving speed, migrating signs has a negative impact on driving speed. Signs have very clear prompts for sharp turns and other information. Here, the road driving environment image can be optimized by migrating signs. According to the interpretable model in Chapter 3, the presence or absence of signs in the "mid-ground" area has the greatest impact on driving speed. Here, the "mid-ground" area is selected as the optimization position (the gray part of P2), and P3 is the ideal speed limit sign image, that is, the target image. P4 is the result of the final road driving environment optimization, which reduces the driver's driving speed, ensures driving safety, enriches the road environment content from the driver's perspective, and makes it less likely for the driver to lose concentration due to monotonous driving.

[0147] Different from previous studies that only focus on the areas of semantic elements such as markings, vegetation, signboards, guardrails, etc. in the road driving environment, the present invention also extracts the position information of the above semantic elements and analyzes the specific impact of their position information on driving speed, so as to guide road designers to select specific locations to adjust and optimize the above semantic elements when using the diffusion model.

[0148] The present invention proposes an intelligent optimization method for road driving environment based on a diffusion model. Different from previous methods based on traditional manual mapping (such as CAD mapping) or image generation algorithms such as generative adversarial networks (such as CycleGAN), the diffusion model has the advantages of high generation efficiency, labor saving, high degree of automation, good stability of generated images, and high resolution and picture quality of generated images.

[0149] Therefore, the present invention adopts the above-mentioned road environment intelligent optimization method based on the diffusion model, which has high efficiency and good image generation quality, and can directly generate optimized and accurate visual road environment images, assisting road designers in the design and optimization of the rural road visual environment and improving the safety of rural roads.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A road environment intelligent optimization method based on a diffusion model, characterized in that: The following steps are involved: S1. Collect natural driving data and use the semantic segmentation network to perform semantic segmentation on the natural driving data to obtain different semantic elements; S2. Use Python programming to extract the area and location information of semantic elements; S3. Establish an interpretable model for quantifying the impact of semantic elements of road driving environment; In step S3, the driving environment semantic elements include 27 variables, of which 12 variables are extracted from the road linear layer, specifically: left near view curve length, left mid view curve length, left distant view curve length, right near view curve length, right mid view curve length, right distant view curve length, left near view curve curvature, left mid view curve curvature, left distant curve curvature, right near view curve curvature, right mid view curve curvature, right distant curve curvature; The other 15 variables are extracted at the visual semantic layer: road area, shoulder area, marking area, vegetation area, corrugated beam guardrail area, concrete guardrail area, whether there is a signboard in the near view area, whether there is a signboard in the middle view area, whether there is a signboard in the far view area, the area of ​​corrugated beam guardrail in the near view, the area of ​​corrugated beam guardrail in the middle view, the area of ​​corrugated beam guardrail in the far view, the area of ​​concrete guardrail in the near view, the area of ​​concrete guardrail in the middle view, and the area of ​​concrete guardrail in the far view; Step S3 specifically inputs 27 variables and driving speed information, outputs a speed regression model through the XGBoost algorithm, and then outputs an interpretability analysis of the 27 variables through the SHAP algorithm; S4. On the basis of step S3, an intelligent optimization model of the road driving environment based on the diffusion model is established, the training and sampling of the model are completed, and an adjusted and optimized road driving environment image is generated.

2. The road environment intelligent optimization method based on diffusion model according to claim 1 is characterized in that: In step S1, a low-cost data acquisition system is used, which includes a GARMIN GDR35 driving recorder, a GPS locator and a three-axis acceleration sensor. The main camera in the GARMIN GDR35 driving recorder is installed on the inner front windshield to collect the road environment from the driver's perspective, and its secondary camera faces the inside of the car to collect the driver's head information.

3. The road environment intelligent optimization method based on diffusion model according to claim 2 is characterized in that: In step S1, the semantic segmentation network includes an encoder and a decoder. The encoder part adopts the ResNet50 network structure. The semantic segmentation network introduces a feature pyramid network. The feature pyramid network performs lateral extraction and vertical fusion on the feature maps output by the four stages of ResNet50. The encoder of the feature pyramid network first upsamples the four branches, and finally integrates the four sampling results using the bagging integration algorithm. The semantic elements include road surface, markings, vegetation, signboards, and guardrails.

4. The road environment intelligent optimization method based on diffusion model according to claim 3 is characterized in that: The specific steps of step S2 are: S201, color of the image after statistical semantic segmentation; S202, matching various color information with road visual semantic elements; S203: Analyze the location information of each road visual semantic element.

5. The road environment intelligent optimization method based on diffusion model according to claim 1 is characterized in that: The training process of the vehicle speed regression model includes the following steps: Initialize the model: set the initial prediction value to the global average; Iterative training: For each round of iteration, the gradient of the current model and the negative gradient of the loss function are calculated; Construct a decision tree: Use a greedy algorithm to select the best split point and recursively construct a decision tree model; Update model: Update the model's prediction value based on the negative gradient of the loss function and the output of the decision tree model; Regularization: Adjust the complexity and parameter values ​​of the model according to the regularization term; Termination condition: determine whether the iteration condition is met; Output the final model: All trained decision tree models are combined into the final integrated model, and the SHAP algorithm is used for interpretability analysis.

6. The road environment intelligent optimization method based on diffusion model according to claim 1 is characterized in that: The specific training process of the diffusion generation model in step S4 includes two steps: training and sampling. The algorithm of the training process is as follows: extract samples from the data, select any time t from time 1 to T, pass x0 and t to the diffusion generation model, and the diffusion generation model samples a random noise, adds it to time x0 and obtains x t , then x t The L2 loss function is used to continuously calculate the gradient and update the weights to complete the training of the neural network. The algorithm of the sampling process is as follows: sample x from the standard normal distribution T , repeat the following process in sequence from time T, T-1, T-2, ..., 2, 1, sampling z from the standard normal distribution, combining x according to the neural network model t and z to calculate x t-1 , and finally returns to x0 after the loop ends.

Citation Information

Patent Citations

  • Curve illegal road occupation prediction method and system

    CN116386044A

  • Risk road section prediction method and system, computer equipment and storage medium

    CN118015829A