Low-level road aesthetic optimization method based on machine learning and fixed diffusion model

Through methods based on machine learning and fixed diffusion model, the aesthetic quality of low-level road environment is calculated and optimized, and the problems of aesthetic feature mining and quantification in the existing technology are solved, and efficient and intelligent aesthetic optimization of road environment is achieved.

CN119314128BActive Publication Date: 2025-06-20TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411361787.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-06-20
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively explore the aesthetic characteristics of road environment elements, lacks interpretability, and the objective quantification of aesthetic characteristics has not yet been achieved, resulting in the optimization of road environment aesthetic quality depends on a large amount of experimental data and expert experience, which is costly.

Method used

The low-level road aesthetic optimization method based on machine learning and fixed diffusion model is adopted, and the aesthetic quality of the road environment is calculated and the dependence on expert experience is reduced through data preparation, the establishment of aesthetic computing models and the intelligent optimization of fixed diffusion models.

Benefits of technology

It improves the automation and intelligence level of aesthetic quality in road environments, reduces repetitive manual work, realizes objective quantitative analysis of aesthetic characteristics, has significant optimization effect, and improves the practicality of aesthetics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119314128B_ABST
    Figure CN119314128B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-grade road aesthetics optimization method based on machine learning and a fixed diffusion model, which relates to the field of low-grade road aesthetics evaluation and includes the following steps: S1: Data preparation, collecting low-grade highway driving environment data and performing data processing; S2: Determining the optimization object based on aesthetics calculation, combining XGBoost and SHAP to establish an aesthetics calculation model for the low-grade road environment, analyzing the diversity, unity, and symmetry of road environment elements, calculating the overall aesthetics score and aesthetics feature score of the low-grade road environment, and determining the optimization object; S3: Using the fixed diffusion model to perform intelligent optimization on the filtered optimization object, the fixed diffusion model includes a diffusion model, a variational autoencoder VAE, and a conditional control module; S4: Verifying the optimization effect, using the aesthetics calculation model to compare the optimization scheme effect based on the fixed diffusion model with the optimization scheme effect based on CycleGAN. The present invention improves the automation and intelligence level in optimizing the aesthetics quality of the road environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aesthetic evaluation of low - grade roads, and particularly to a low - grade road aesthetic optimization method based on machine learning and fixed diffusion models. Background Art

[0002] In the low - grade road environment, the aesthetic preferences of drivers are mainly affected by road environment elements such as semantic, color, and texture information. These information stimuli different regions of the brain's visual cortex, resulting in different aesthetic preferences. Drivers show different aesthetic preferences for different semantic information. Natural semantic information elements such as vegetation, water, and meadows are more favored. Richer color and texture information is positively correlated with drivers' aesthetic preferences for the road environment. The combination and distribution of colors not only enhance the visual attractiveness of the road environment but also provide clear visual guidance for drivers to understand the road environment. Road environment elements are an important part of the aesthetic quality of the road environment and the information source for the computational analysis of road environment aesthetics.

[0003] Current road environment aesthetic calculation models are usually realized by extracting corresponding road environment elements. However, these models do not further explore the aesthetic features of road environment elements and lack interpretability. In addition, existing research on subjective aesthetic evaluation of road environments has found that aesthetic features are closely related to the aesthetic quality of the road environment, but these features rely on empirical scoring and have not been objectively quantified.

[0004] The optimization methods for the aesthetic quality of the road environment rely on a large amount of experimental data and expert experience, and it is difficult to verify the optimization effect through intelligent generation. The low - grade road environment aesthetic quality optimization method based on human factors engineering can improve driving safety and comfort, but due to the need for a large amount of psychological and physiological data, its cost is high. The urban and rural road environment landscape layout optimization method based on design specifications relies on expert experience and has a low degree of automation in scheme modification.

[0005] Therefore, it is necessary to provide a low - grade road aesthetic optimization method based on machine learning and fixed diffusion models to solve the above problems. Summary of the Invention

[0006] The object of the present invention is to provide a low - grade road aesthetic optimization method based on machine learning and fixed diffusion models, which objectively calculates the aesthetic quality of the low - grade road environment starting from the aesthetic features of road environment elements and emphasizes the importance of improving its aesthetic quality. In addition, it improves the automation and intelligence level in optimizing the aesthetic quality of the road environment, reduces the repetitive manual work in the road design process and the dependence on expert experience.

[0007] To achieve the above object, the present invention provides a low - grade road aesthetic optimization method based on machine learning and fixed diffusion models, including the following steps:

[0008] S1: Data preparation, collecting the driving environment data of low - grade roads and performing data processing;

[0009] S2: Determining the optimization object based on aesthetic calculation, combining XGBoost and SHAP to establish an aesthetic calculation model for the low - grade road environment, analyzing the diversity, unity, and symmetry of road environment elements, calculating the overall aesthetic score and aesthetic feature score of the low - grade road environment, and determining the optimization object;

[0010] S3: Using a fixed diffusion model to perform intelligent optimization on the filtered optimization object, where the fixed diffusion model includes a diffusion model, a variational auto - encoder VAE, and a conditional control module;

[0011] S4: Verifying the optimization effect, using the aesthetic calculation model to compare the optimization effect of the optimization scheme based on the fixed diffusion model with the optimization effect of the optimization scheme based on CycleGAN.

[0012] Preferably, in step S1, a GARMIN GDR35 driving recorder is used to collect data. Through data cleaning and screening, 2000 clear and unobstructed driving environment images are obtained, covering five types of landscapes: forest, grassland, snow mountain, wasteland, and cliff. Experimental participants rate 1500 of the environmental images at six levels to evaluate the overall aesthetic score and aesthetic feature score of the road environment.

[0013] Preferably, in step S1, the low - grade road environment elements are quantified from semantic information, color information, and texture information. The acquisition process of semantic information is as follows:

[0014] S11: Performing semantic segmentation on the original driving environment image according to the constructed semantic segmentation network;

[0015] S12: The semantic segmentation network uses ResNet50 as the backbone network for sampling;

[0016] S13: Connecting to an FPN network to perform further feature fusion on the features;

[0017] S14: Integrating the sampled features through the Bagging algorithm. The semantic information obtained after semantic segmentation includes roads, vegetation, protective facilities, and the sky;

[0018] The acquisition process of color information is as follows: The driving environment image is converted to grayscale, and then the grayscale values are mapped into the original color image to form a color heat map. The larger the grayscale value, the brighter the color in the heat map. The color information is represented by three attributes: Hue, Saturation, and Value in the HSV color space;

[0019] The Gabor filtering method is used to extract the texture information of the driving environment. The specific process is as follows:

[0020] Define the orientation angle, spatial aspect ratio, standard deviation, and frequency of the Gabor filter, generate a filter that matches the size of the environmental image, and apply it to the driving environment image. After filtering, a filtered response image at a specific direction and scale is obtained.

[0021] Preferably, in step S2, the diversity of the driving environment is quantified by diversity, which includes semantic diversity, color diversity, and texture diversity. The semantic diversity value is calculated by weighted summation of the information entropy of various semantic information. The larger the value, the more complex and diverse the semantic information; the color diversity is quantified by the RGB channels. The larger the value, the richer the color; the texture diversity is calculated by statistically averaging the variances of all local region features of the texture information image. The larger the value, the more diverse the texture information. The calculation formula is as follows:

[0022]

[0023] Div col =σ rgyb +0.3μ rgyb

[0024]

[0025] Among them, Div sem represents the semantic diversity value; p(c i ) represents the number of all pixel points of a certain semantics in the image; C represents the number of semantic categories; Div col represents the color diversity value; r, g, b respectively represent the values of the three color channels; rg represents the difference between the red channel and the green channel; yb represents half of the sum of the red and green channels minus the blue channel; μ rgyb represents the average value of the combined rg and yb channels; σ rgyb represents the standard deviation of the combined rg and yb channels; Div tex represents the texture diversity value; μ i represents the mean value of the local region; Var(μ i ) represents the variance of the mean values of all local regions; represents the variance of the local region; represents the variance of the variance values of all local regions; H i represents the entropy value of the local region; Var(H i ) represents the variance of the entropy values of all local regions.

[0026] Preferably, in step S2, the unity of the driving environment is reflected in semantic unity, color unity, and texture unity;

[0027] Process the semantic segmentation image through the Canny edge detection algorithm to obtain the edge features of the semantic segmentation image, and use the ratio of the number of edge pixel points to the total number of pixel points as the semantic unity value. The larger the semantic unity value, the more cluttered the semantics and the worse the unity;

[0028] The color unity value is obtained by matching the color unity atlas that best matches the color heat map. The color unity value is obtained by calculating the hue distance and saturation difference. The smaller the color unity value, the more uniform the color;

[0029] The texture unity value is quantified by the mean standard deviation of the pixel values of the texture information image. The larger the texture unity value, the greater the difference in texture information. The calculation formula is as follows:

[0030]

[0031] Among them, Uni sem represents the semantic unity value; M represents the number of rows of pixel points in the image; N represents the number of columns of pixel points in the image; Canny(i, j) represents the pixel point (i, j) representing the edge detected by the Canny edge detection method; Uni col represents the color unity value; H(p) and S(p) respectively represent the hue value and saturation value of pixel point p; S t (α) represents the most suitable color unity atlas; t represents the atlas type; α represents the atlas rotation angle; the hue distance represents the arc length distance (radian) on the color wheel; Uni tex represents the texture unity value; Z represents the number of local windows; Tex i represents the texture information image corresponding to the i-th local window; std(Tex i ) represents the standard deviation of the pixel values of the texture information image corresponding to the i-th local window.

[0032] Preferably, in step S2, the symmetry of the driving environment is reflected in the visual balance of semantic symmetry, color symmetry, and texture information symmetry;

[0033] The semantic symmetry value is obtained by extracting the number of pixel points of four main semantic information through k-means clustering and calculating the difference in the pixel ratio of these semantic information between the left and right parts of the semantic segmentation image. The larger the semantic symmetry value, the greater the difference in the distribution of semantic information and the weaker the symmetry;

[0034] The color symmetry value is obtained by accumulating the difference in color pixel values between the left and right parts of the color heat map. The smaller the color symmetry value, the smaller the color difference of the symmetric pixels and the better the symmetry;

[0035] The texture symmetry is represented by the mean of the absolute differences in texture between the left and right parts of the texture information image. The larger the texture symmetry value, the greater the texture difference between the left and right parts, and the worse the symmetry. The calculation formula is as follows:

[0036]

[0037] Among them, Sym sem represents the semantic symmetry value; C represents the number of semantic categories; represents the pixel proportion of the i-th semantic in the left half; represents the pixel proportion of the i-th semantic in the right half; Sym col represents the color symmetry value; A represents the height of the image; B represents the width of the image; Col(P(i, j)) represents the color pixel value of the pixel point P(i, j) on the color heat map at the position (i, j); Col(P(A - i, j)) represents the color pixel value of the pixel point P(A - i, j) that is symmetric to the pixel point P(i, j) about the center line of the color heat map; Sym tex represents the texture symmetry value; Tex(P(i, j)) represents the pixel value of the pixel point P(i, j) in the texture information image at the position (i, j); Tex(P(A - i, j)) represents the pixel value of the symmetric point P(A - i, j) of the pixel point P(i, j) about the center line of the texture information image.

[0038] Preferably, in step S2, the optimization method for the optimization object is as follows:

[0039] S21: XGBoost is used to predict the overall aesthetic score and aesthetic feature score of the low-level road environment to determine the optimization object and the key points of optimization, and SHAP is used to interpret the model output and analyze the specific influence degrees of semantic diversity, color diversity, and texture diversity on the diversity;

[0040] S22: In XGBoost, the decision tree is split according to the optimal w tj and stops when the node depth reaches the maximum depth:

[0041]

[0042] In the formula, G tj and H tj are the first-order derivative and the second-order derivative of the leaf node respectively; J is the number of leaf nodes; w tj is the optimal value of the j-th leaf node of the t-th decision tree, λ is a trade-off parameter, and γ represents the complexity of the leaf;

[0043] S23: The aesthetic calculation model is evaluated by the mean square error MSE. The smaller the MSE, the smaller the difference between the calculated value and the observed value, and the better the performance of the model. The calculation formula is as follows:

[0044]

[0045] where N is the number of samples; y i is the true value of the i-th sample; is the calculated value of the i-th sample;

[0046] S24: Use SHAP to interpret the output of the aesthetic calculation model, and the calculation formula is as follows:

[0047]

[0048] where φ i represents the contribution of factor i; N represents the set of all input factors; n represents the total number of samples of the factor; ν is the given model; S represents the set containing all observed factors.

[0049] Preferably, in step S3, the process of the fixed diffusion model performing intelligent optimization on the filtered optimization object is as follows:

[0050] S31: The input image x is transformed into the latent space through the encoder ε to obtain the latent image z in the latent space;

[0051] S32: The diffusion model continuously adds noise to the latent image z and performs forward diffusion to obtain the final noise map z T ;

[0052] S33: The conditional control module further processes the input image x and the text prompt as the constraint condition τ θ (y), and interprets it through the noise predictor, and obtains the required image by controlling the reverse diffusion process;

[0053] S34: The U-Net noise predictor of the diffusion model denoises the noise map z T to obtain the latent image z without noise, realizing the reverse diffusion process;

[0054] S35: The decoder transforms the latent space output image z into the pixel space to obtain the final image generator

[0055] Preferably, in step S3, the loss function L of the fixed diffusion model SDM represents the expected value of the losses of all noise samples of the aesthetic learning model under certain image and condition information, and the calculation formula of L SDM is as follows:

[0056]

[0057] Among them, y represents the text prompt condition information; ε(x) represents the latent image compressed and mapped to the latent space by the encoder ε; ∈~N(0,1) means that the noise sample ∈ conforms to the standard normal distribution; E ε(x),y,∈~N(0,1),t represents the expected value of the loss of the model over all possible noise samples of the image that conform to the normal distribution; represents the square of the L2 norm of the difference between the predicted noise calculated by the noise predictor and the generated noise; t is the uniformly sampled time step; z t represents the latent representation obtained from the encoder ε; θ represents the parameters of the model; τ θ (y) represents the transformation of the conditional information y through the conditional control module with parameters θ.

[0058] Therefore, the present invention adopts the above-mentioned low-grade road aesthetics optimization method based on machine learning and fixed diffusion model, and has the following beneficial effects:

[0059] (1) The present invention generates corresponding text prompts according to the optimization object and inputs them into the fixed diffusion model to intelligently generate optimized low-grade road environment images, and the overall aesthetics score and the average aesthetics feature score are increased by 64.5% and 121.5% respectively.

[0060] (2) The present invention determines the optimization object by calculating the overall aesthetics score and the aesthetics feature score of the road environment through the aesthetics calculation model, provides an interpretable framework for road environment design, and enables designers to optimize the road environment more targeted.

[0061] (3) The present invention improves the automation and intelligence level in optimizing the aesthetics quality of the road environment, reduces the repetitive manual work and the dependence on expert experience in the road design process.

[0062] (4) The present invention links road environment elements such as semantics, color, and texture in the road environment with aesthetics features, realizes objective quantitative analysis of aesthetics calculation, and provides better optimization results compared with the popular image generation algorithm CycleGAN.

[0063] (5) The intelligent optimization technology of road environment based on aesthetics calculation proposed by the present invention realizes efficient scheme generation, has remarkable optimization effect, improves the practicality of aesthetics, and is of great significance to the optimization, design and transformation of road environment.

[0064] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0065] Figure 1 is the flowchart of the low-grade road aesthetics optimization method of the present invention based on machine learning and fixed diffusion model;

[0066] Figure 2 is the overall framework diagram of the low - grade road aesthetics optimization method based on machine learning and fixed diffusion model in the present invention;

[0067] Figure 3 is the structural diagram of the semantic segmentation network in the present invention;

[0068] Figure 4 is the flow chart of the fixed diffusion model in the present invention;

[0069] Figure 5 is the fitting curve diagram of the observed value and the calculated value in the present invention;

[0070] Figure 6 is the result diagram of the shape interpretation of the aesthetic feature scores of each type in the present invention. Detailed implementation manners

[0071] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0072] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.

[0073] The terms "including" or "comprising" and the like used in the present invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation to the present invention. When the absolute position of the described object changes, the relative position relationship may also change accordingly. In the present invention, unless otherwise clearly defined and limited, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0074] Embodiment

[0075] As Figure 1 shown, the present invention provides a low - grade road aesthetics optimization method based on machine learning and fixed diffusion model, including the following steps: The framework of the optimization method is as Figure 2 shown;

[0076] S1: Data preparation, collecting low-class highway driving environment data and performing data processing;

[0077] Use a GARMIN GDR35 driving recorder to collect data. Through data cleaning and screening, 2000 clear and unobstructed driving environment images are obtained, covering five types of landscapes: forest, grassland, snow mountain, wasteland, and cliff. 1500 of the environmental images are scored by experimental participants at six levels to evaluate the overall aesthetic score and aesthetic feature score of the road environment, as shown in Table 1. The gender ratio of the experimental participants is balanced, and their ages are between 23 and 50 years old, with an average age of 30.6 years and a standard deviation of 6.1 years.

[0078] Table 1

[0079]

[0080] In step S1, the low-class road environment elements are quantified from semantic information, color information, and texture information. The semantic segmentation network structure is as Figure 3 shown. The process of obtaining semantic information is as follows:

[0081] S11: Perform semantic segmentation on the original driving environment images according to the constructed semantic segmentation network;

[0082] S12: The semantic segmentation network uses ResNet50 as the backbone network for sampling;

[0083] S13: Connect to the FPN network for further feature fusion of the features;

[0084] S14: Integrate the sampled features through the Bagging algorithm. The semantic information obtained after semantic segmentation includes roads, vegetation, protective facilities, and the sky;

[0085] The process of obtaining color information is as follows: Convert the driving environment images to grayscale, and then map the grayscale values to the original color images to form a color heat map. The larger the grayscale value, the brighter the color in the heat map. The color information is represented by the three attributes of Hue, Saturation, and Value in the HSV color space;

[0086] The Gabor filtering method is used to extract the texture information of the driving environment. The specific process is as follows:

[0087] Define the orientation angle, spatial aspect ratio, standard deviation, and frequency of the Gabor filter, generate a filter that matches the size of the environment image, and apply it to the driving environment image. After filtering, a filtered response image at a specific direction and scale is obtained.

[0088] S2: Determine the optimization object based on aesthetic calculation. Combine XGBoost and SHAP to establish an aesthetic calculation model for low-level road environments. Analyze the diversity, unity, and symmetry of road environment elements, calculate the overall aesthetic score and aesthetic feature score of the low-level road environment, and determine the optimization object;

[0089] In step S2, the diversity of the driving environment is quantified by diversity, which includes semantic diversity, color diversity, and texture diversity. The semantic diversity value is calculated by weighted summation of the information entropy of various semantic information. The larger the value, the more complex and diverse the semantic information; color diversity is quantified by the RGB channels. The larger the value, the richer the color; texture diversity is calculated by statistically averaging the variances of all local region features of the texture information image. The larger the value, the more diverse the texture information. The calculation formula is as follows:

[0090]

[0091] Div col =σ rgyb +0.3μ rgyb

[0092]

[0093] Among them, Div sem represents the semantic diversity value; p(c i ) represents the number of all pixel points of a certain semantics in the image; C represents the number of semantic categories; Div col represents the color diversity value; r, g, b respectively represent the values of the three color channels; rg represents the difference between the red channel and the green channel; yb represents half of the sum of the red and green channels minus the blue channel; μ rgyb represents the average value of the combined rg and yb channels; σ rgyb represents the standard deviation of the combined rg and yb channels; Div tex represents the texture diversity value; μ i represents the mean of the local region; Var(μ i ) represents the variance of the means of all local regions; represents the variance of the local region; represents the variance of the variance values of all local regions; H i represents the entropy value of the local region; Var(H i ) represents the variance of the entropy values of all local regions.

[0094] In step S2, the unity of the driving environment is reflected in semantic unity, color unity, and texture unity;

[0095] Process the semantic segmentation image through the Canny edge detection algorithm to obtain the edge features of the semantic segmentation image, and use the ratio of the number of edge pixel points to the total number of pixel points as the semantic unity value. The larger the semantic unity value, the more cluttered the semantics and the worse the unity;

[0096] The color unity value is obtained by matching the color unity atlas that best matches the color heat map. The color unity value is obtained by calculating the hue distance and saturation difference. The smaller the color unity value, the more unified the colors;

[0097] The texture unity value is quantified by the mean standard deviation of the pixel values of the texture information image. The larger the texture unity value, the greater the difference in texture information. The calculation formula is as follows:

[0098]

[0099] Among them, Uni sem represents the semantic unity value; M represents the number of rows of pixel points in the image; N represents the number of columns of pixel points in the image; Canny(i, j) represents the pixel point (i, j) representing the edge detected by the Canny edge detection method; Uni col represents the color unity value; H(p) and S(p) respectively represent the hue value and saturation value of pixel point p; S t (α) represents the color unity atlas that best matches; t represents the atlas type; α represents the atlas rotation angle; the hue distance represents the arc length distance (in radians) on the color wheel; Uni tex represents the texture unity value; Z represents the number of local windows; Tex i represents the texture information image corresponding to the i-th local window; std(Tex i ) represents the standard deviation of the pixel values of the texture information image corresponding to the i-th local window.

[0100] In step S2, the symmetry of the driving environment is reflected in the visual balance of semantic symmetry, color symmetry, and texture information symmetry;

[0101] The semantic symmetry value is obtained by extracting the number of pixel points of four main semantic information through k-means clustering, and by calculating the difference in the pixel proportion of these semantic information in the left and right parts of the semantic segmentation image. The larger the semantic symmetry value, the greater the difference in the distribution of semantic information and the weaker the symmetry;

[0102] The color symmetry value is obtained by accumulating the difference in color pixel values between the left and right parts of the color heat map. The smaller the color symmetry value, the smaller the color difference of symmetric pixels and the better the symmetry;

[0103] The texture symmetry is represented by the mean of the absolute differences in texture between the left and right parts of the texture information image. The larger the texture symmetry value, the greater the texture difference between the left and right parts, and the worse the symmetry. The calculation formula is as follows:

[0104]

[0105] where Sym sem represents the semantic symmetry value; C represents the number of semantic categories; represents the pixel proportion of the i-th semantics in the left half; represents the pixel proportion of the i-th semantics in the right half; Sym col represents the color symmetry value; A represents the height of the image; B represents the width of the image; Col(P(i,j)) represents the color pixel value of the pixel point P(i,j) on the color heat map at the position (i,j); Col(P(A-i,j)) represents the color pixel value of the pixel point P(A-i,j) that is symmetric to the pixel point P(i,j) about the center line of the color heat map; Sym tex represents the texture symmetry value; Tex(P(i,j)) represents the pixel value of the pixel point P(i,j) in the texture information image at the position (i,j); Tex(P(A-i,j)) represents the pixel value of the symmetric point P(A-i,j) of the pixel point P(i,j) about the center line of the texture information image.

[0106] In step S2, the optimization method for the optimization object is as follows:

[0107] S21: XGBoost is used to predict the overall aesthetic score and aesthetic feature score of the low-level road environment to determine the optimization object and optimization focus. SHAP is used to interpret the model output and analyze the specific impact degrees of semantic diversity, color diversity, and texture diversity on diversity;

[0108] S22: The core of XGBoost is to minimize the objective function. Through the second-order Taylor expansion and removing the constant term, the objective function is simplified. In XGBoost, the decision tree is split according to the optimal w tj and stops when the node depth reaches the maximum depth:

[0109]

[0110] In the formula, G tj and H tj are the first-order derivative and second-order derivative of the leaf node respectively; J is the number of leaf nodes; w tj is the optimal value of the j-th leaf node of the t-th decision tree, λ is a compromise parameter, and γ represents the complexity of the leaf;

[0111] S23: Evaluate the aesthetic calculation model through the mean squared error (MSE). The smaller the MSE, the smaller the difference between the calculated value and the observed value, and the better the model performance. The calculation formula is as follows:

[0112]

[0113] where N is the number of samples; y i is the true value of the i-th sample; is the calculated value of the i-th sample;

[0114] S24: Use SHAP to interpret the output of the aesthetic calculation model. The calculation formula is as follows:

[0115]

[0116] where φ i represents the contribution of factor i; N represents the set of all input factors; n represents the total number of samples of the factors; ν is the given model; S represents the set containing all observed factors.

[0117] S3: Use a fixed diffusion model to perform intelligent optimization on the filtered optimization object. The fixed diffusion model includes a diffusion model, a variational autoencoder (VAE), and a conditional control module. As Figure 4 shown, in step S3, the process of the fixed diffusion model performing intelligent optimization on the filtered optimization object is as follows:

[0118] S31: The input image x is transformed into the latent space through the encoder ε to obtain the latent image z in the latent space;

[0119] S32: The diffusion model continuously adds noise to the latent image z and performs forward diffusion to obtain the final noise map z T ;

[0120] S33: The conditional control module further processes the input image x and the text prompt as the constraint condition τ θ (y), and interprets it through the noise predictor to obtain the required image by controlling the reverse diffusion process;

[0121] S34: The U-Net noise predictor of the diffusion model denoises the noise map z T to obtain the latent image z without noise, realizing the reverse diffusion process;

[0122] S35: The decoder transforms the latent space output image z into the pixel space to obtain the final image generator

[0123] In step S3, the loss function L of the fixed diffusion model SDMRepresents the expected value of the loss of all noise samples of the aesthetic learning model under certain image and conditional information, \(L\). SDM The calculation formula is as follows:

[0124]

[0125] Among them, \(y\) represents the text prompt conditional information; \(\epsilon(x)\) represents the latent image compressed and mapped to the latent space by the encoder \(\epsilon\); \(\epsilon\sim N(0,1)\) means that the noise sample \(\epsilon\) conforms to the standard normal distribution; \(E\) ε(x),y,∈~N(0,1),t Represents the expected value of the loss of the model on all possible noise samples of the image that conform to the normal distribution; Represents the square of the L2 norm of the difference between the predicted noise and the generated noise calculated by the noise predictor; \(t\) is the uniformly sampled time step; \(z\) t Represents the latent representation obtained from the encoder \(\epsilon\); \(\theta\) represents the parameters of the model; \(\tau\) θ \(\tau(y)\) represents the transformation of the conditional information \(y\) through the conditional control module with parameters \(\theta\).

[0126] S4: Verify the optimization effect, and use the aesthetic calculation model to compare the optimization scheme effect based on the fixed diffusion model with the optimization scheme effect based on CycleGAN.

[0127] Performance of the aesthetic calculation model

[0128] 1500 samples are randomly divided into 80% for training and 20% for testing. After standardizing the input features, the parameters of the overall aesthetic score calculation model and the three aesthetic feature score calculation models are optimized by grid search and five-fold cross-validation, as shown in Table 2. The mean square errors of these models are 0.10, 0.06, 0.12, and 0.06 respectively. In addition, through the analysis of random sampling, the calculated values of these models are basically consistent with the observed values, as Figure 5 shown. Therefore, the aesthetic calculation model proposed in this scheme performs well.

[0129] Table 2

[0130]

[0131]

[0132] Road environment and aesthetic features to be optimized

[0133] According to the data distribution characteristics of the annotation sample evaluation indicators, the average scores of the overall aesthetics of the road environment and the average scores of the aesthetic feature scores are both less than 3 points, as shown in Table 3. In addition, as shown in Table 1, this scheme defines 3 points as a situation with better aesthetic quality. Therefore, among the unannotated samples after aesthetic calculation, the samples with the overall aesthetic score of the road environment less than 3 points are identified as road environments that need to be optimized. In addition, the aesthetic feature with the lowest score in the road environment is identified as the aesthetic feature that needs to be optimized.

[0134] Table 3

[0135]

[0136] The road environment elements with the greatest impact

[0137] Using the SHAP method, the influencing factors of each aesthetic feature score are analyzed from two aspects: global feature contribution and individual feature contribution, as Figure 6 shown to describe the global feature contribution and individual feature contribution, Figure 6 (a) represents the global feature contribution, which illustrates the relative importance of the independent variables sorted by the absolute SHAP value. Figure 6 (b) represents the individual feature contribution. Each point represents a sample, the color corresponds to the feature value, and the abscissa value of this point is the SHAP value. When the SHAP value of the black solid point is positive, it indicates that the independent variable with a larger feature value has a positive impact on the aesthetic feature score.

[0138] In terms of the diversity score, color diversity is the most important variable, with the highest absolute SHAP value of 0.30. In the Unity score, color unity is the most important variable, with an absolute SHAP value of 0.25, which is much larger than texture unity and semantic unity. Finally, semantic symmetry is the feature that contributes the most to the symmetry score, with an absolute SHAP value of 0.39. While the absolute SHAP values of texture symmetry and color symmetry are both less than 0.08. From Figure 6 (b), it can be seen that most of the black solid points of the diversity, unity, and symmetry values of each road environment element (i.e., semantic, color, and texture information) are distributed on the positive axis of the SHAP value, while the white hollow points are distributed on the negative axis. This indicates that the aesthetic feature values of these road environment elements are positively correlated with various aesthetic feature scores.

[0139] Example 1

[0140] When the road environment is a barren land without vegetation cover, its overall aesthetic score is 1.9, indicating that there is still a large room for improvement in the aesthetic performance of the road environment. The diversity score, unity score, and symmetry score are 1.5, 3.1, and 2.4 respectively, indicating that the aesthetic feature that needs to be optimized is diversity. According to the SHAP results, color diversity is the road environment element that has the greatest impact on the diversity score. Therefore, in this embodiment, the aesthetic performance of the road environment is mainly optimized by increasing color diversity.

[0141] Scenario 1 and Scenario 2 are optimization cases generated by the fixed diffusion model, and Scenario 3 is a comparative optimization case generated by CycleGAN.

[0142] Scenario 1 optimizes the road environment by adding grass. The text prompt words include adding grass, rich colors, realistic, natural, and road environment. The fixed diffusion model iterates 30 times. The overall aesthetic score of the optimized road environment is increased to 3.4 points, and the diversity score is increased to 3.6 points, with increases of 78.9% and 140.0% respectively. This indicates that there has been a significant improvement in aesthetic performance.

[0143] Trees will also bring some optimization to the color of the road environment. Scenario 2 improves color diversity by increasing trees. The text prompt words modify "adding grass" in Scenario 1 to "increasing green trees" and iterate 30 times. The overall aesthetic score of the optimized road environment is 3.7 points, and the road environment diversity score is 4.1 points, with increases of 94.7% and 173.3% respectively.

[0144] For comparison, CycleGAN uses the road environment image with rich vegetation on the roadside as the target to generate Scenario 3. The overall aesthetic score of the road environment in Scenario 3 is 3.1 points, and the road environment diversity score is 3.3 points, both of which are less than those of Optimization Case 1 and Optimization Case 2.

[0145] Therefore, the optimization cases generated based on the fixed diffusion model have significantly improved the overall aesthetic score and diversity score of the road environment, and the improvement effect is better than that of the optimization cases generated based on CycleGAN.

[0146] Embodiment 2

[0147] In the original road environment, there are vegetation-free rocks on the left and shrubs on the right. The overall aesthetic score of the road environment is 2.6, the diversity score is 3.6, the unity score is 2.1, and the symmetry score is 2.5. Therefore, the aesthetic feature that needs to be optimized in the original road environment is unity. Since color unity is the most important variable affecting the unity score, this embodiment considers adjusting the color unity performance to make the road environment more harmonious and unified.

[0148] Scenario 1 is to increase the color unity by adding vegetation to cover the bare rocks. The text prompt words include adding vegetation on the left side, with colors similar to the vegetation on the right side, realistic, natural, and the road environment. After 30 iterations of the fixed diffusion model, the overall aesthetic score of the road environment is increased to 4.1, and the unity score is increased to 4.3, with increases of 57.7% and 104.8% respectively.

[0149] Scenario 2 improves the color unity by adding shrubs and flowers to cover the bare rocks. Change "adding vegetation on the left side" in Scenario 1 to "adding flowers on the left side", and the number of iterations is the same as in Scenario 1. The overall aesthetic score of the optimized road environment is 4.5, and the unity score is 4.1. The improvement effect exceeds 73.0%. Both of the two optimization scenarios generated based on the fixed diffusion model effectively optimize the aesthetic performance of the current road environment.

[0150] Scenario 3 is a comparative optimization case generated by CycleGAN, targeting the road environment image where the roadside rocks are already covered with rich vegetation. The overall aesthetic score of the road environment is 3.3, and the unity score is 3.1. The improvement effects are both less than 48.0%.

[0151] Therefore, the improvement effects of the two optimization cases generated based on the fixed diffusion model are better, and the effects of the optimization scenarios are more realistic.

[0152] Example 3

[0153] The overall aesthetic score of the road environment to be optimized is 2.8. There are basically no trees on the left side, and there are several yellow trees on the right side. The diversity score, unity score, and symmetry score are 3.5, 3.1, and 1.9 respectively. It shows that symmetry is the road aesthetic feature that needs to be optimized, and semantic symmetry is the most important factor affecting the symmetry score. Therefore, the main purpose of this example is to improve the semantic symmetry of the road environment to make it more balanced and symmetrical.

[0154] Scenario 1 enhances the symmetry of the road environment by adding semantic information of vegetation on the left side. The text prompt includes adding many trees to cover the mountain on the left side, adding more trees on the left side similar to the right side, symmetric and realistic, and the number of iterations is 30 times. The overall aesthetic score of the optimized road environment is increased by 50.0% to reach 4.2, and the symmetry score is increased by 115.8% to reach 4.1.

[0155] Scenario 2 optimizes the semantic information of the vegetation on both sides of the road. The text prompt includes adding trees of the same color and height on both sides of the road, multiple trees covering the mountains on the left and right sides, realistic, and the number of iterations is the same as in Scenario 1. The overall aesthetic score of the optimized road environment is 3.7 points, and the symmetry score is 3.8 points. The improvement effects are 32.1% and 100.0% respectively.

[0156] In the third solution, the comparison solution generated by CycleGAN uses the road environment image with rich vegetation on both sides of the road as the target. The overall aesthetic score of the road environment is 3.2, and the symmetry score of the road environment is 2.8. The improvement effects of the two scores are 14.3% and 47.4% respectively, which are significantly less than those of the first and second solutions. Therefore, both optimization solutions generated based on the fixed diffusion model can more effectively make the road environment more balanced and symmetrical by adjusting semantic information.

[0157] Considering the display effect of the optimization solution, the solution three generated based on CycleGAN is significantly less realistic and more random.

[0158] Conclusion

[0159] In the three embodiments, the intelligent road environment optimization technology based on aesthetic calculation proposed in this solution has increased the overall aesthetic score and aesthetic feature score by 64.5% and 121.5% respectively on average. The optimization solution generated based on CycleGAN has increased the overall aesthetic score and aesthetic feature score by 34.8% and 71.7% respectively on average. Therefore, the optimization cases generated based on the fixed diffusion model can more effectively improve the aesthetic performance of the road environment. In addition, the fixed diffusion model can generate multiple more targeted optimization solutions under customized text prompts. In contrast, the optimization cases generated by CycleGAN based on the target image tend to be more random and less realistic.

[0160] Therefore, the optimization cases generated based on the fixed diffusion model are more in line with the wishes of road environment designers and provide a convenient and intuitive demonstration method. Combined with the aesthetic calculation model, it can directly reflect the effect of road environment optimization.

[0161] The data source in this solution is the static image of the road environment, which can be extended to dynamic video data in the future.

[0162] Therefore, the present invention adopts the above-mentioned low-grade road aesthetic optimization method based on machine learning and fixed diffusion model. The proposed intelligent road environment optimization technology based on aesthetic calculation realizes efficient solution generation, has significant optimization effects, improves the practicality of aesthetics, and is of great significance to the optimization, design and transformation of road environment.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A low-level road aesthetic optimization method based on machine learning and fixed diffusion model, characterized by: The following steps are involved: S1: Data preparation, collecting low-grade highway driving environment data and processing the data; In step S1, low-level road environment elements are quantified from semantic information, color information and texture information. The process of obtaining semantic information is as follows: S11: Perform semantic segmentation on the original driving environment image according to the constructed semantic segmentation network; S12: The semantic segmentation network uses ResNet50 as the backbone network for sampling; S13: access the FPN network to further fuse the features; S14: The sampled features are integrated through the Bagging algorithm. The semantic information obtained after semantic segmentation includes roads, vegetation, protective facilities and sky; The process of obtaining color information is as follows: convert the driving environment image into grayscale, and then map the grayscale value to the original color image to form a color heat map. The larger the grayscale value, the brighter the color in the heat map. The color information is represented by the three attributes of Hue, Saturation and Value in the HSV color space. The Gabor filtering method is used to extract the texture information of the driving environment. The specific process is as follows: Define the direction angle, spatial aspect ratio, standard deviation and frequency of the Gabor filter, generate a filter that matches the size of the environment image and apply it to the driving environment image, and obtain a filter response image at a specific direction and scale after filtering; S2: Determine the optimization object based on aesthetic calculation, combine XGBoost and SHAP to establish an aesthetic calculation model for low-level road environment, analyze the diversity, unity and symmetry of road environment elements, calculate the overall aesthetic score and aesthetic feature score of low-level road environment, and determine the optimization object; S3: A fixed diffusion model is used to intelligently optimize the filtered optimization object. The fixed diffusion model includes a diffusion model, a variational encoder VAE, and a conditional control module. In step S3, the process of intelligently optimizing the filtered optimization object by the fixed diffusion model is as follows: S31: The input image x is converted to the latent space through the encoder ε to obtain the latent image z of the latent space; S32: The diffusion model continuously adds noise to the potential image z and diffuses forward to obtain the final noise image z T ; S33: The conditional control module further processes the input image x and text prompt as a constraint condition τ θ (y), and make the noise predictor interpret it, and obtain the desired image by controlling the reverse diffusion process; S34: U-Net noise predictor for diffusion model on noise map z T De-noising is performed to obtain a noise-free latent image z, thus realizing the reverse diffusion process; S35: The decoder converts the latent space output image z to the pixel space to obtain the final image generator In step S3, the loss function L of the diffusion model is fixed SDM It shows the expected value of all noise sample losses of the aesthetic learning model under certain image and condition information, L SDM The calculation formula is as follows: Among them, y represents the text prompt condition information; ε(x) represents the potential image compressed and mapped to the latent space by the encoder ε; ∈~N(0,1) represents the noise sample ∈ conforms to the standard normal distribution; E ε(x),y,∈~N(0,1),t Represents the expected value of the model's loss on all possible noise samples of the image that conform to the normal distribution; represents the square of the L2 norm of the difference between the predicted noise calculated by the noise predictor and the generated noise; t is the uniformly sampled time step; z t represents the potential representation obtained from the encoder ε; θ represents the parameters of the model; τ θ (y) represents the transformation of condition information y through the condition control module with parameter θ; S4: Verify the optimization effect, and use the aesthetic calculation model to compare the optimization effect based on the fixed diffusion model and the optimization effect based on CycleGAN; The intelligent road environment optimization technology based on aesthetic calculation improved the overall aesthetic score and aesthetic feature score by an average of 64.5% and 121.5%, respectively.

2. The low-level road aesthetic optimization method based on machine learning and fixed diffusion model according to claim 1, characterized in that: In step S1, a GARMIN GDR35 driving recorder was used to collect data. Through data cleaning and screening, 2000 clear and unobstructed driving environment images were obtained, covering five types of landscapes: forests, grasslands, snow-capped mountains, wastelands, and cliffs. The experimental participants scored 1500 of the environmental images at six levels to evaluate the overall aesthetic score and aesthetic feature score of the road environment.

3. The low-level road aesthetic optimization method based on machine learning and fixed diffusion model according to claim 1, characterized in that: In step S2, the diversity of the driving environment is quantified by diversity, and the diversity includes semantic diversity, color diversity and texture diversity; The semantic diversity value is calculated by weighted summation of information entropy of various semantic information. The larger the value, the more complex and diverse the semantic information. Color diversity is quantified by the RGB channels, with larger values ​​indicating richer colors; Texture diversity is calculated by counting the variance mean of all local area features of the texture information image. The larger the value, the more diverse the texture information. The calculation formula is as follows: Div. col =s rgyb +0.3m rgyb Among them, Div sem represents the semantic diversity value; p(c i ) represents the number of all pixels of a certain semantics in the image; C represents the number of semantic categories; Div col Represents the color diversity value; r, g, b represent the values ​​of the three color channels respectively; rg represents the difference between the red channel and the green channel; yb represents half of the sum of the red and green channels minus the blue channel; μ rgyb represents the average value of the combined rg and yb channels; σ rgyb Indicates the standard deviation of the combined rg and yb channels; Div tex Represents the texture diversity value; μ i Represents the mean value of the local area; Var(μ i ) represents the variance of the mean of all local regions; Represents the variance of the local area; represents the variance of all local area variance values; H i Represents the entropy value of the local area; Var(H i ) represents the variance of all local region entropy values.

4. The low-level road aesthetic optimization method based on machine learning and fixed diffusion model according to claim 1, characterized in that: In step S2, the uniformity of the driving environment is reflected in the uniformity of semantics, color and texture; The semantic segmentation image is processed by the Canny edge detection algorithm to obtain the edge features of the semantic segmentation image, and the ratio of the number of edge pixels to the total number of pixels is used as the semantic unity value. The larger the semantic unity value, the more chaotic the semantics and the worse the unity. The color uniformity value is obtained by matching the color uniformity map that best matches the color heat map. The color uniformity value is obtained by calculating the hue distance and saturation difference. The smaller the color uniformity value, the more uniform the color. The texture uniformity value is quantified by the mean standard deviation of the pixel values ​​of the texture information image. The larger the texture uniformity value, the greater the difference in texture information. The calculation formula is as follows: Among them, Uni sem Represents the semantic unity value; M represents the number of rows of pixels in the image; N represents the number of columns of pixels in the image; Canny(i,j) represents the pixel point (i,j) detected by the Canny edge detection method to represent the edge; Uni col represents the color uniformity value; H(p) and S(p) represent the hue value and saturation value of pixel p respectively; S t (α) represents the most consistent color uniformity map; t represents the map type; α represents the rotation angle of the map; hue distance Indicates the arc length on the color wheel (radians); Uni tex Represents the texture uniformity value; Z represents the number of local windows; Tex i Represents the texture information image corresponding to the i-th local window; std(Tex i ) represents the standard deviation of the pixel values ​​of the texture information image corresponding to the i-th local window.

5. The low-level road aesthetic optimization method based on machine learning and fixed diffusion model according to claim 1, characterized in that: In step S2, the symmetry of the driving environment is reflected in the visual balance of semantic symmetry, color symmetry and texture information symmetry; The semantic symmetry value is obtained by extracting the number of pixels of four main types of semantic information through k-means clustering, and calculating the difference in pixel proportions of these semantic information in the left and right parts of the semantic segmentation image. The larger the semantic symmetry value, the greater the difference in semantic information distribution and the weaker the symmetry; The color symmetry value is obtained by accumulating the color pixel value difference between the left and right parts of the color heat map. The smaller the color symmetry value, the smaller the color difference of the symmetrical pixels, and the better the symmetry; Texture symmetry is represented by the mean of the absolute difference between the left and right parts of the texture information image. The larger the texture symmetry value, the greater the difference between the left and right parts of the texture, and the worse the symmetry. The calculation formula is as follows: Among them, Sym sem represents the semantic symmetry value; C represents the number of semantic categories; Indicates the pixel ratio of the i-th semantic in the left half; Indicates the pixel ratio of the i-th semantic in the right half; Sym col Represents the color symmetry value; A represents the height of the image; B represents the width of the image; Col(P(i,j)) represents the color pixel value of the pixel point P(i,j) at position (i,j) on the color heat map; Col(P(Ai,j)) represents the color pixel value of the pixel point P(i,j) symmetrical to the center line of the color heat map; Sym tex Represents the texture symmetry value; Tex(P(i,j)) represents the pixel value of the pixel point P(i,j) at the position (i,j) of the texture information image; Tex(P(Ai,j)) represents the pixel value of the symmetric point P(Ai,j) of the pixel point P(i,j) about the center line of the texture information image.

6. The low-level road aesthetic optimization method based on machine learning and fixed diffusion model according to claim 1, characterized in that: In step S2, the optimization method of the optimization object is as follows: S21: XGBoost is used to predict the overall aesthetic score and aesthetic feature score of low-level road environments to determine the optimization object and optimization focus. SHAP is used to interpret the model output and analyze the specific impact of semantic diversity, color diversity, and texture diversity on diversity. S22: In XGBoost, the decision tree is based on the optimal w tj Make a split and stop when the node depth reaches the maximum depth: In the formula, G tj and H tj are the first-order derivative and second-order derivative of the leaf nodes respectively; J is the number of leaf nodes; w tj is the optimal value of the jth leaf node of the tth decision tree, λ is a trade-off parameter, and γ represents the complexity of the leaf; S23: The aesthetic calculation model is evaluated by the mean square error (MSE). The smaller the MSE, the smaller the difference between the calculated value and the observed value, and the better the model performance. The calculation formula is as follows: Where N is the number of samples; y i is the true value of the i-th sample; is the calculated value of the i-th sample; S24: SHAP is used to explain the output of the aesthetic computational model. The calculation formula is as follows: Among them, φ i represents the contribution of factor i; N represents the set of all input factors; n represents the total number of samples of the factor; ν is the given model; S represents the set of all observed factors.

Citation Information

Patent Citations

  • Image data set expansion method based on diffusion model, medium and equipment

    CN116883545A

  • Face beauty prediction method and device based on double diffusion model, equipment and medium

    CN117373077A