SD-ML-based urban sustainable development innovation driving mechanism analysis method
By deeply coupling the SD-ML method, secondary data is generated using system dynamics and combined with machine learning models to identify key innovation drivers, quantify nonlinear responses and threshold effects, and solve the problems of dynamic mechanism characterization and threshold identification in urban sustainable development research, providing accurate quantitative support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-02
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional system dynamics methods have limitations in terms of the intensity of nonlinear response between quantitative variables and the accurate identification of mutation thresholds. Machine learning methods lack mechanistic constraints and are highly dependent on large-scale training data, resulting in difficulties in characterizing dynamic mechanisms, inefficient threshold identification, and insufficient interpretability in urban sustainable development research.
This study employs an SD-ML-based analytical method for the innovation-driven mechanism of urban sustainable development. By generating secondary data through a system dynamics model and combining it with a machine learning model, key innovation dimension indicators are identified, the nonlinear response of urban sustainability to innovation indicators is quantified, and the innovation-driven mechanism is explored.
This study achieves deep coupling between the SD model and the ML method, making up for the shortcomings of a single model, providing accurate quantitative support, and offering theoretical and practical value for urban sustainable development planning and policy optimization.
Smart Images

Figure CN121745776A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of environmental management, and particularly relates to a sustainable development innovation driving mechanism analysis method based on SD-ML. BACKGROUND
[0002] A town system is coupled by multiple subsystems such as population and public services, industrial structure and economic growth, resource and energy consumption and ecological environment bearing, innovation input and technology diffusion, and has obvious complexity, nonlinearity, feedback and time lag. The traditional system dynamics (SD) method is good at depicting the causal feedback structure and dynamic evolution mechanism of the system, but has limitations in quantifying the non-linear response strength between variables and accurately identifying the mutation threshold. Although the machine learning (ML) method can fit high-dimensional non-linear relationships and effectively identify complex interaction patterns in data, it lacks mechanism constraints and is strongly dependent on large-scale training data, and the sample size of town annual statistical data is small, so the effect of directly training the model is not good.
[0003] At present, there is a lack of effective coupling between mechanism models and data-driven models, and a method is needed to deeply couple the mechanism advantages of SD and the non-linear mining capabilities of ML to solve the problems of difficult depiction of dynamic mechanism, inefficient threshold identification and insufficient explanation in the research of sustainable development of towns, and to provide quantitative support for planning management and policy optimization. SUMMARY
[0004] To solve the above problems, the application provides a sustainable development innovation driving mechanism analysis method based on SD-ML, which deeply couples the mechanism modeling of SD and the non-linear mining capabilities of ML, solves the defects of single model in dynamic mechanism depiction, non-linear quantification and data dependence, provides precise quantitative support for sustainable development planning and policy optimization of towns, and has good popularization and application prospect.
[0005] To achieve the above purpose, the application adopts the following technical solutions: The sustainable development innovation driving mechanism analysis method based on SD-ML comprises the following steps: S1, obtaining research area data and performing data preprocessing to form a basic database of four dimensions of society, economy, ecological environment and innovation; S2, constructing a system dynamics model covering four subsystems of society, economy, ecological environment and innovation to realize dynamic simulation of the sustainable development of the town; S3, embedding technical innovation elements under the SSP-RCP scenario framework, setting different innovation levels of the scene, and generating secondary data sets through system dynamics multi-scenario simulation; S4, standardize the secondary data, take the urban sustainability index as the response variable, construct the feature matrix, and use the machine learning model to establish the quantitative nonlinear response relationship between the driving factors including the innovation index and the urban sustainability index; S5, use the machine learning model interpretation method to identify the key innovation dimension index affecting urban sustainability, quantify the nonlinear response of urban sustainability to the innovation index, analyze the innovation-driven urban sustainable development mechanism, and explore whether there is a threshold effect of innovation impact.
[0006] Preferably, the specific process of step S1 is: S11, determine the boundary of the modeling research area, set the time range and unify the statistical caliber; S12, obtain and organize social data, economic data, ecological environment data and innovation data; the social data includes population size and structure, population growth rate and total number of employees; the economic data includes total GDP, per capita GDP and fixed asset investment; the ecological environment data includes energy consumption, water consumption, pollutant emission, air quality and green land rate; the innovation data includes scientific research fund, scientific research personnel, patent authorization quantity and high-tech output value; S13, unify the units of the original basic data, use the natural spline interpolation method to fill in the missing values in the time series data, identify the outliers by the quartile range method and replace them by the spline interpolation method, and obtain the basic database.
[0007] Preferably, the specific process of step S2 is: S21, determine the system boundary and subsystem structure, establish the social subsystem, economic subsystem, ecological environment subsystem and innovation subsystem of the town, and draw the causal loop diagram to show the positive feedback or negative feedback relationship between variables, and clarify the influence path between each subsystem; S22, build the stock-flow diagram based on the causal loop diagram, establish the system dynamics equation and the corresponding mathematical relationship between variables; the stock-flow diagram includes state variables for describing the cumulative characteristics of the system, rate variables for reflecting the transfer rate of matter or information, auxiliary variables for describing the intermediate conversion process in the system, and constant parameter system for maintaining the overall steady-state structure of the system; The state variable is expressed by an integral equation as follows: Wherein, X(t) represents the state variable of the system at time t; X(0) represents the state variable of the system at the initial time; F(s) represents the inflow rate of the system at historical time s; O(s) represents the outflow rate of the system at historical time s; t represents time t; s represents historical time; The expression of the rate variable is: wherein, represents the rate of change of the state variable at time t; represents the inflow rate of the system at time t; represents the outflow rate of the system at time t; S23, parameterizing the quantitative relationship between the key variables using a functional equation or a regression relationship to realize dynamic simulation of the sustainable development of the city and town; S24, calibrating the model parameters using historical data of the research area, verifying the reliability of the model through historical fitting test, sensitivity analysis and consistency test.
[0008] Preferably, the specific process of step S3 is: S31, setting the socio-economic and climate constraint scenarios based on the SSP-RCP scenario framework, giving the boundary conditions of population, economic growth, energy structure and carbon emission constraints; S32, embedding technical innovation elements to adjust the innovation-related parameters to form benchmark innovation, low innovation, medium innovation and high innovation scenarios; the technical innovation elements include R&D investment intensity and patent output growth rate; S33, running the system dynamics model under multiple scenarios to simulate the future, outputting secondary data sets containing population, total GDP, per capita GDP, energy consumption and scientific research personnel indicators under different innovation levels.
[0009] Preferably, the specific process of step S4 is: S41, forming an index matrix of the annual data of the four-dimensional indicators of society, economy, ecological environment and innovation in the secondary data set; S42, attribute determination is performed on each index to obtain a direction label ; wherein, represents a positive index, represents a negative index; S43, constructing a sustainable development index excluding the innovation dimension, using equal weight average for the dimension index, and the calculation formula is: wherein, represents the dimension index; represents the th dimension; represents the th sample; represents the th dimension; represents the th index; represents the set of normalized social, economic and ecological environment indicators; represents the The sample at the th Standardized values for each indicator; The comprehensive sustainability index uses an equally weighted average across dimensions, and the calculation formula is as follows: ,in, This represents the comprehensive sustainability index; This corresponds to the three dimensions of society, economy, and ecological environment; S44. Select innovation dimension indicators as core feature variables and key indicators of social, economic, and ecological environment dimensions as auxiliary feature variables to construct a full-factor feature matrix. The calculation formula is as follows: , ,in, This represents the stack of all sample feature vectors; Indicates the preceding The input feature vector of each sample; This represents the inversion of the eigenvector; Indicates the first The input feature vector of each sample; Indicates the first The sample at the th Normalized values of indicators under each feature dimension; Indicates the number of the feature dimension; This represents the total number of feature dimensions; Indicates the first One sample in front Normalized values of indicators under each feature dimension; S45. The XGBoost model is used as a machine learning model to fit the nonlinear relationship between total factor drivers and the comprehensive sustainability index. The calculation formula is as follows: ,in, This indicates that the XGBoost model is for the first... The input feature vector of each sample The predicted value; Indicates the regression tree number; This indicates the total number of regression trees; Indicates the first A regression tree; Indicates the learning rate; A model with optimal generalization ability is obtained through hyperparameter optimization, and then... , and Evaluate the model's effectiveness; among which, It is the coefficient of determination that reflects the overall effectiveness of the model; It is the root mean square error, which measures the average deviation between the model's predicted values and the actual values. It is the average absolute error between the predicted value and the actual value.
[0010] Preferably, the specific process of step S5 is as follows: S51. Calculate the contribution of each feature using the SHAP method: ,in, Indicates the benchmark term; Indicates the first The index under the feature dimension is at the first The contribution value of each sample to the model output; Indicates the number of indicators; The formula for calculating global importance is as follows: ,in, Indicates the first The global importance of indicators under each feature dimension; N represents the total number of samples; the average absolute contribution value is used to measure the overall impact of each indicator on the model output, key innovation indicators affecting sustainability are identified and screened, and then the samples are fitted with local weighted regression to obtain a continuous contribution curve, which is used to characterize the nonlinear impact trend of key innovation indicators on the sustainability index. S52. Construct the PDP model curve, and the calculation formula is as follows: ,in, This means that, assuming the sample distribution remains unchanged for all other indicators, the... Each innovation indicator takes the value of The average impact on the overall sustainability index; Indicates the values of innovation indicators; Indicates sample In addition to the first The value vectors of the other innovation indicators besides the one indicator; Indicates the first Each innovation indicator takes the value of At that time, all other innovation indicators used samples The original values are the model's predicted values; Construct the ALE model curve, divide it into Q interval boundaries, and calculate the local effects using the following formula: ,in, Indicates the first The features are in the interval Local effects within; Indicates the first The number of samples in the interval; Represents all samples that meet the conditions. The condition is the first The first sample eigenvalues Falling in the range Inside; Indicates the first In the nth sample, the nth The feature values are: The predicted value of the time model; Indicates the first In the nth sample, the nth The feature values are: The predicted value of the time model; The formula for calculating the cumulative local effect is: ,in, To represent the cumulative local effect, the interval from 1 to... The sum of local effects; right The calculation formula is as follows: ,in, This represents the cumulative local effect after centralization. Indicates the total number of intervals. This represents the interval index for calculating the mean. This is used to iterate through all interval boundary points; Indicates the first The cumulative local effects of each interval; S53. Smooth the curves of the PDP model and the centralized ALE model, identify the points of abrupt change in the slope of the curves, explore the threshold range, and realize the mechanism analysis of innovation-driven sustainable urban development.
[0011] By adopting the above technical solution, this invention has the following beneficial effects: Through deep coupling of SD and ML methods, this invention utilizes the SD model to generate a large amount of secondary data that conforms to the system mechanism, solving the problem of insufficient training data for the ML model; leveraging the powerful nonlinear fitting capability of the ML model, it compensates for the shortcomings of the SD method in quantitatively characterizing complex response relationships. Through model interpretation methods, it achieves the identification of key innovation-driving factors, the quantification of nonlinear responses, and the exploration of threshold effects, providing precise quantitative support for urban sustainable development planning and policy formulation, and possessing significant theoretical and practical value. Attached Figure Description
[0012] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0014] like Figure 1 As shown, the analytical method for the innovation-driven mechanism of urban sustainable development based on SD-ML includes the following steps: S1. Acquire data from the study area and perform data preprocessing to form a basic database covering four dimensions: society, economy, ecological environment, and innovation. The specific process of step S1 is as follows: S11. Determine the boundaries of the research area for modeling, set the time range, and standardize the statistical criteria; S12. Acquire and organize social data, economic data, ecological environment data, and innovation data; the social data includes population size and structure, population growth rate, and total number of employees; the economic data includes total GDP, GDP per capita, and fixed asset investment; the ecological environment data includes energy consumption, water consumption, pollutant emissions, air quality, and green space ratio; the innovation data includes research funding, research personnel, number of patents granted, and high-tech output value. S13. The original basic data is processed to unify the units, and the missing values in the time series data are filled by the natural spline interpolation method. Outliers are identified by the interquartile range method and replaced by spline interpolation to obtain the basic database. S2. Construct a system dynamics model covering four subsystems: society, economy, ecological environment, and innovation, to achieve dynamic simulation of sustainable urban development; The specific process of step S2 is as follows: S21. Determine the system boundary and subsystem structure, establish the town's social subsystem, economic subsystem, ecological environment subsystem and innovation subsystem, and draw causal loop diagrams to show the positive or negative feedback relationships between variables, and clarify the influence paths between each subsystem; S22. Construct a stock-flow diagram based on the causal loop diagram, and establish the system dynamic equations and corresponding mathematical relationships between variables; the stock-flow diagram includes state variables for characterizing the cumulative characteristics of the system, rate variables for reflecting the transfer rate of matter or information, auxiliary variables for describing the intermediate transformation process of the system, and a constant parameter system for maintaining the overall steady-state structure of the system. The expression for the rate variable is: ,in, This represents the rate of change of the state variable at time t; This represents the inflow rate of the system at time t; This represents the outflow rate of the system at time t; S23. Use functional equations or regression relationships to parametrically express the quantitative relationships between key variables, thereby achieving dynamic simulation of sustainable urban development; S24. Use historical data of the study area to calibrate the model parameters, and verify the reliability of the model through historical fit test, sensitivity analysis and consistency test. S3. Embed technological innovation elements within the SSP-RCP scenario framework, set up scenarios with different levels of innovation, and generate secondary datasets through multi-scenario simulation of system dynamics. The specific process of step S3 is as follows: S31. Based on the SSP-RCP scenario framework, set up socio-economic and climate constraint scenarios, and give boundary conditions for population, economic growth, energy structure and carbon emission constraints. S32. Embedding technological innovation elements and adjusting innovation-related parameters to form baseline innovation, low innovation, medium innovation, and high innovation scenarios; the technological innovation elements include R&D investment intensity and patent output growth rate; S33. Run the system dynamics model under multiple scenarios to perform future simulations and produce secondary datasets covering different levels of innovation, including indicators such as population, total GDP, GDP per capita, energy consumption, and scientific researchers. S4. Standardize the secondary data, construct a feature matrix with the urban sustainability index as the response variable, and use a machine learning model to establish a quantitative nonlinear response relationship between the driving factors, including innovation indicators, and the urban sustainability index. The specific process of step S4 is as follows: S41. Form an indicator matrix from the annual data of the four dimensions of secondary dataset: social, economic, ecological environment and innovation. S42. Perform attribute determination on each indicator to obtain the direction label. ;in, Indicates a positive indicator. Indicates a negative indicator; S43. Construct a sustainable development index that does not include the innovation dimension. The dimension index adopts an equal-weighted average, and the calculation formula is as follows: ,in, Indicates dimensional index; Indicates the first One dimension; Indicates the first One sample; Indicates the first The number of indicators under each dimension; Indicates the first One indicator; A set of standardized indicators representing the social, economic, and ecological environment; Indicates the first The sample at the th Standardized values for each indicator; The comprehensive sustainability index uses an equally weighted average across dimensions, and the calculation formula is as follows: ,in, This represents the comprehensive sustainability index; This corresponds to the three dimensions of society, economy, and ecological environment; S44. Select innovation dimension indicators as core feature variables and key indicators of social, economic, and ecological environment dimensions as auxiliary feature variables to construct a full-factor feature matrix. The calculation formula is as follows: , ,in, This represents the stack of all sample feature vectors; Indicates the preceding The input feature vector of each sample; This represents the inversion of the eigenvector; Indicates the first The input feature vector of each sample; Indicates the first The sample at the th Normalized values of indicators under each feature dimension; Indicates the number of the feature dimension; This represents the total number of feature dimensions; Indicates the first One sample in front Normalized values of indicators under each feature dimension; S45. The XGBoost model is used as a machine learning model to fit the nonlinear relationship between total factor drivers and the comprehensive sustainability index. The calculation formula is as follows: ,in, This indicates that the XGBoost model is for the first... The input feature vector of each sample The predicted value; Indicates the regression tree number; This indicates the total number of regression trees; Indicates the first A regression tree; Indicates the learning rate; A model with optimal generalization ability is obtained through hyperparameter optimization, and then... , and Evaluate the model's effectiveness; among which, It is the coefficient of determination that reflects the overall effectiveness of the model; It is the root mean square error, which measures the average deviation between the model's predicted values and the actual values. It is the mean absolute error between the predicted value and the actual value; S5. Using machine learning models to explain the methods, identify key innovation dimension indicators that affect urban sustainability, quantify the nonlinear response of urban sustainability to innovation indicators, analyze the mechanism of innovation-driven urban sustainable development, and explore whether there is a threshold effect in the impact of innovation. The specific process of step S5 is as follows: S51. Calculate the contribution of each feature using the SHAP method: ,in, Indicates the benchmark term; Indicates the first The index under the feature dimension is at the first The contribution value of each sample to the model output; Indicates the number of indicators; The formula for calculating global importance is as follows: ,in, Indicates the first The global importance of indicators under each feature dimension; N represents the total number of samples; the average absolute contribution value is used to measure the overall impact of each indicator on the model output, key innovation indicators affecting sustainability are identified and screened, and then the samples are fitted with local weighted regression to obtain a continuous contribution curve, which is used to characterize the nonlinear impact trend of key innovation indicators on the sustainability index. S52. Construct the PDP model curve, and the calculation formula is as follows: ,in, This means that, assuming the sample distribution remains unchanged for all other indicators, the... Each innovation indicator takes the value of The average impact on the overall sustainability index; Indicates the values of innovation indicators; Indicates sample In addition to the first The value vectors of the other innovation indicators besides the one indicator; Indicates the first Each innovation indicator takes the value of At that time, all other innovation indicators used samples The original values are the model's predicted values; Construct the ALE model curve, divide it into Q interval boundaries, and calculate the local effects using the following formula: ,in, Indicates the first The features are in the interval Local effects within; Indicates the first The number of samples in the interval; Represents all samples that meet the conditions. The condition is the first The first sample eigenvalues Falling in the range Inside; Indicates the first In the nth sample, the nth The feature values are: The predicted value of the time model; Indicates the first In the nth sample, the nth The feature values are: The predicted value of the time model; The formula for calculating the cumulative local effect is: ,in, To represent the cumulative local effect, the interval from 1 to... The sum of local effects; right The calculation formula is as follows: ,in, This represents the cumulative local effect after centralization. Indicates the total number of intervals. This represents the interval index for calculating the mean. This is used to iterate through all interval boundary points; Indicates the first The cumulative local effects of each interval; S53. Smooth the curves of the PDP model and the centralized ALE model, identify the points of abrupt change in the slope of the curves, explore the threshold range, and realize the mechanism analysis of innovation-driven sustainable urban development.
[0015] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for analyzing the innovation-driven mechanism of sustainable urban development based on SD-ML, characterized in that, Includes the following steps: S1. Acquire data from the study area and perform data preprocessing to form a basic database covering four dimensions: society, economy, ecological environment, and innovation. S2. Construct a system dynamics model covering four subsystems: society, economy, ecological environment, and innovation, to achieve dynamic simulation of sustainable urban development; S3. Embed technological innovation elements within the SSP-RCP scenario framework, set up scenarios with different levels of innovation, and generate secondary datasets through multi-scenario simulation of system dynamics. S4. Standardize the secondary data, construct a feature matrix with the urban sustainability index as the response variable, and use a machine learning model to establish a quantitative nonlinear response relationship between the driving factors, including innovation indicators, and the urban sustainability index. S5. Using machine learning models to explain the methods, identify key innovation dimension indicators that affect urban sustainability, quantify the nonlinear response of urban sustainability to innovation indicators, analyze the mechanism of innovation-driven urban sustainable development, and explore whether there is a threshold effect in the impact of innovation.
2. The method for analyzing the innovation-driven mechanism of sustainable urban development based on SD-ML as described in claim 1, characterized in that, The specific process of step S1 is as follows: S11. Determine the boundaries of the research area for modeling, set the time range, and standardize the statistical criteria; S12. Acquire and organize social data, economic data, ecological environment data, and innovation data; the social data includes population size and structure, population growth rate, and total number of employees; the economic data includes total GDP, GDP per capita, and fixed asset investment; the ecological environment data includes energy consumption, water consumption, pollutant emissions, air quality, and green space ratio; the innovation data includes research funding, research personnel, number of patents granted, and high-tech output value. S13. The original basic data is processed to unify the units. The missing values in the time series data are filled by natural spline interpolation. Outliers are identified by interquartile range and replaced by spline interpolation to obtain the basic database.
3. The method for analyzing the innovation-driven mechanism of urban sustainable development based on SD-ML as described in claim 1, characterized in that, The specific process of step S2 is as follows: S21. Determine the system boundary and subsystem structure, establish the town's social subsystem, economic subsystem, ecological environment subsystem and innovation subsystem, and draw causal loop diagrams to show the positive or negative feedback relationships between variables, and clarify the influence paths between each subsystem; S22. Construct a stock-flow diagram based on the causal loop diagram, and establish the system dynamic equations and corresponding mathematical relationships between variables; the stock-flow diagram includes state variables for characterizing the cumulative characteristics of the system, rate variables for reflecting the transfer rate of matter or information, auxiliary variables for describing the intermediate transformation process of the system, and a constant parameter system for maintaining the overall steady-state structure of the system. The state variables are expressed by an integral equation as follows: ,in, Represents the state variables of the system at time t; This represents the state variables of the system at the initial moment; This represents the inflow rate of the system at historical time s; This represents the outflow rate of the system at historical time s; The initial time is represented by 't'; the time 's' is represented by 's'; and the historical time is represented by 's'. The expression for the rate variable is: ,in, This represents the rate of change of the state variable at time t; This represents the inflow rate of the system at time t; This represents the outflow rate of the system at time t; S23. Use functional equations or regression relationships to parametrically express the quantitative relationships between key variables, thereby achieving dynamic simulation of sustainable urban development; S24. Use historical data of the study area to calibrate the model parameters, and verify the reliability of the model through historical fit test, sensitivity analysis and consistency test.
4. The method for analyzing the innovation-driven mechanism of urban sustainable development based on SD-ML as described in claim 1, characterized in that, The specific process of step S3 is as follows: S31. Based on the SSP-RCP scenario framework, set up socio-economic and climate constraint scenarios, and give boundary conditions for population, economic growth, energy structure and carbon emission constraints. S32. Embedding technological innovation elements and adjusting innovation-related parameters to form baseline innovation, low innovation, medium innovation, and high innovation scenarios; the technological innovation elements include R&D investment intensity and patent output growth rate; S33. Run the system dynamics model under multiple scenarios to perform future simulations and produce secondary datasets covering different levels of innovation, including indicators such as population, total GDP, GDP per capita, energy consumption, and scientific researchers.
5. The method for analyzing the innovation-driven mechanism of urban sustainable development based on SD-ML as described in claim 1, characterized in that, The specific process of step S4 is as follows: S41. Form an indicator matrix from the annual data of the four dimensions of secondary dataset: social, economic, ecological environment and innovation. S42. Perform attribute determination on each indicator to obtain the direction label. ;in, Indicates a positive indicator. Indicates a negative indicator; S43. Construct a sustainable development index that does not include the innovation dimension. The dimension index adopts an equal-weighted average, and the calculation formula is as follows: ,in, Indicates dimensional index; Indicates the first One dimension; Indicates the first One sample; Indicates the first The number of indicators under each dimension; Indicates the first One indicator; A set of standardized indicators representing the social, economic, and ecological environment; Indicates the first The sample at the th Standardized values for each indicator; The comprehensive sustainability index uses an equally weighted average across dimensions, and the calculation formula is as follows: ,in, This represents the comprehensive sustainability index; This corresponds to the three dimensions of society, economy, and ecological environment; S44. Select innovation dimension indicators as core feature variables and key indicators of social, economic, and ecological environment dimensions as auxiliary feature variables to construct a full-factor feature matrix. The calculation formula is as follows: , ,in, This represents the stack of all sample feature vectors; Indicates the preceding The input feature vector of each sample; This represents the inversion of the eigenvector; Indicates the first The input feature vector of each sample; Indicates the first The sample at the th Normalized values of indicators under each feature dimension; Indicates the number of the feature dimension; This represents the total number of feature dimensions; Indicates the first One sample in front Normalized values of indicators under each feature dimension; S45. The XGBoost model is used as a machine learning model to fit the nonlinear relationship between total factor drivers and the comprehensive sustainability index. The calculation formula is as follows: ,in, This indicates that the XGBoost model is for the first... The input feature vector of each sample The predicted value; Indicates the regression tree number; This indicates the total number of regression trees; Indicates the first A regression tree; Indicates the learning rate; A model with optimal generalization ability is obtained through hyperparameter optimization, and then... , and Evaluate the model's effectiveness; among which, It is the coefficient of determination that reflects the overall effectiveness of the model; It is the root mean square error, which measures the average deviation between the model's predicted values and the actual values. It is the average absolute error between the predicted value and the actual value.
6. The method for analyzing the innovation-driven mechanism of urban sustainable development based on SD-ML as described in claim 5, characterized in that, The specific process of step S5 is as follows: S51. Calculate the contribution of each feature using the SHAP method: ,in, Indicates the benchmark term; Indicates the first The index under the feature dimension is at the first The contribution value of each sample to the model output; Indicates the number of indicators; The formula for calculating global importance is as follows: ,in, Indicates the first The global importance of indicators under each feature dimension; N represents the total number of samples; the average absolute contribution value is used to measure the overall impact of each indicator on the model output, key innovation indicators affecting sustainability are identified and screened, and then the samples are fitted with local weighted regression to obtain a continuous contribution curve, which is used to characterize the nonlinear impact trend of key innovation indicators on the sustainability index. S52. Construct the PDP model curve, and the calculation formula is as follows: ,in, This means that, assuming the sample distribution remains unchanged for all other indicators, the... Each innovation indicator takes the value of The average impact on the overall sustainability index; Indicates the values of innovation indicators; Indicates sample In addition to the first The value vectors of the other innovation indicators besides the one indicator; Indicates the first Each innovation indicator takes the value of At that time, all other innovation indicators used samples The original values are the model's predicted values; Construct the ALE model curve, divide it into Q interval boundaries, and calculate the local effects using the following formula: ,in, Indicates the first The features are in the interval Local effects within; Indicates the first The number of samples in the interval; Represents all samples that meet the conditions. The condition is the first The first sample eigenvalues Falling in the range Inside; Indicates the first In the nth sample, the nth The feature values are: The predicted value of the time model; Indicates the first In the nth sample, the nth The feature values are: The predicted value of the time model; The formula for calculating the cumulative local effect is: ,in, To represent the cumulative local effect, the interval from 1 to... The sum of local effects; right The calculation formula is as follows: ,in, This represents the cumulative local effect after centralization. Indicates the total number of intervals. This represents the interval index for calculating the mean. This is used to iterate through all interval boundary points; Indicates the first The cumulative local effects of each interval; S53. Smooth the curves of the PDP model and the centralized ALE model, identify the points of abrupt change in the slope of the curves, explore the threshold range, and realize the mechanism analysis of innovation-driven sustainable urban development.
Citation Information
Patent Citations
Data-driven urban sustainable development planning decision optimization method and system
CN117973641A
Urban coordinated development assessment method based on water resources and coupling factors thereof
CN121118618A
Urban carbon emission intelligent prediction method based on multi-modal data and machine learning integration
CN121503790A
Method for analyzing changes in urban economic development characteristics of urban agglomeration based on nighttime light remote sensing
US20250013961A1