Urban creative innovation enterprise distribution prediction method, system, equipment and product based on regression analysis and machine learning
Through regression analysis and machine learning methods, combined with urban facility data, enterprise distribution data and population density data, and using negative binomial regression model and distributed random forest model for analysis, the problem of low enterprise distribution prediction accuracy in the existing technology is solved, and higher accuracy prediction results are achieved, and more scientific urban planning and industrial layout are supported.
Patent Information
- Application Number
- CN202510461162.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-24
AI Technical Summary
The existing urban creative innovation enterprise distribution prediction method has low prediction accuracy, failing to fully consider the complex relationship between multiple factors, and it is difficult to effectively support urban planning and resource allocation.
Using a method based on regression analysis and machine learning, the target area is divided into grid units, urban facility data, enterprise distribution data and population density data are collected, and evaluation system is constructed, and negative binomial regression model and distributed random forest model are used for analysis and prediction, and enterprise distribution prediction results are generated and visualized.
It improves the accuracy of enterprise distribution forecasts, can support urban planning and resource allocation more effectively, help identify key influencing factors in the development of cultural and creative industries, and provides site selection suggestions to reduce entrepreneurial risks and enhance industrial competitiveness.
Smart Images

Figure CN120197777A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of urban planning and big data analysis, and particularly relates to a method, system, device and product for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning. Background Art
[0002] The distribution of creative and innovative enterprises in a city refers to the aggregation or dispersion of enterprises in the field of creative and innovative industries at different geographical locations within the urban space. Studying the distribution of urban creative and innovative enterprises helps to optimize the allocation of urban resources, promote industrial development, and provide a scientific basis for urban planning.
[0003] The distribution of urban enterprises is affected by various factors, such as the distribution of urban facilities, the accessibility of transportation facilities, and the richness of public facilities. Most of the existing methods for predicting the distribution of urban enterprises are based on simple regression models or empirical analyses, and fail to fully consider the complex relationships among multiple factors, resulting in low prediction accuracy and being difficult to effectively support urban planning and resource allocation.
[0004] In the related art, with the advancement of the urbanization process, the relationship between enterprise distribution and urban facilities has become more complex. Traditional methods are difficult to provide sufficiently accurate prediction results when facing large-scale data. Therefore, there is an urgent need to propose a method and system for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning to solve the deficiencies in the prior art. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method, system, device and product for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning, which improves the prediction accuracy and provides effective support for urban planning.
[0006] On the one hand, to achieve the above object, the present invention provides a method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning, including:
[0007] Dividing a target area into grid cells with a preset granularity, and collecting urban facility data, enterprise distribution data and population density data within each grid cell;
[0008] Constructing an evaluation system, reclassifying the urban facility data and calculating the richness of each type of facility;
[0009] Using a negative binomial regression model to analyze the influence of the variables in the evaluation system on the distribution of cultural and innovative enterprises, and extracting the regression residuals as new variables;
[0010] Inputting the urban facility data, enterprise distribution data, population density data and regression residuals into a distributed random forest model for training to generate an enterprise distribution prediction result;
[0011] Generate a visualization graph based on the prediction results to display the density and spatial patterns of enterprise distribution in each region.
[0012] Optionally, the process of dividing the target area into grid cells of a preset granularity includes: performing grid division on the target area using a 1km×1km fishing net grid to form uniform grid cells; mapping enterprise POI data, urban facility data, and road length data to the corresponding grid cells through a geographic information system tool.
[0013] Optionally, the urban facility data includes consumer facilities, public facilities, transportation facilities, and scenic area data. The process of calculating the richness of each type of facility includes:
[0014]
[0015] where R i is the richness of each type of facility; p ik is the proportion of facility type k in grid cell i; n is the total number of different facility types in grid cell i.
[0016] Optionally, the process of analyzing using the negative binomial regression model includes:
[0017] Taking the number of cultural innovation enterprises as the dependent variable, and taking facility richness, traffic accessibility, and population density as independent variables for regression fitting, and extracting the regression residuals as additional variables representing unexplained variation.
[0018] Optionally, the training process of the distributed random forest model includes:
[0019] Combining the original data with the regression residual variables, dividing the training set, test set, and validation set, and capturing the non-linear relationships and interaction effects between variables through multi-decision tree integration.
[0020] Optionally, the process of generating the visualization graph includes:
[0021] Output an enterprise distribution density map, which shows the aggregation degree of enterprises in each region of the city with a color gradient map.
[0022] Optionally, the visualization graph also includes: marking resource blank areas and aggregation areas, and generating resource allocation suggestions.
[0023] On the other hand, to achieve the above object, the present invention also provides a prediction system for the distribution of urban creative and innovative enterprises based on regression analysis and machine learning, including:
[0024] A data collection and region division module, an evaluation system construction module, a regression analysis and residual extraction module, a model training and prediction module, and a prediction result visualization module;
[0025] The data acquisition and region division module is used to divide the target region into grid cells with a preset granularity, and collect urban facility data, enterprise distribution data, and population density data within each grid cell;
[0026] The evaluation system construction module is used to construct an evaluation system, reclassify the urban facility data, and calculate the richness of each type of facility;
[0027] The regression analysis and residual extraction module is used to analyze the impact of the evaluation system variables on the distribution of cultural innovation enterprises using a negative binomial regression model, and extract the regression residuals as new variables;
[0028] The model training and prediction module is used to input the urban facility data, enterprise distribution data, population density data, and regression residuals into a distributed random forest model for training to generate an enterprise distribution prediction result;
[0029] The prediction result visualization module is used to generate a visualization graph based on the prediction result to display the density and spatial pattern of enterprise distribution in each region.
[0030] An urban creative innovation enterprise distribution prediction device based on regression analysis and machine learning includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the described method is implemented.
[0031] A computer program product, when running on an urban creative innovation enterprise distribution prediction device based on regression analysis and machine learning, causes the device to execute the described method.
[0032] Technical effects of the present invention: The present invention discloses a method, system, device, and product for predicting the distribution of urban creative innovation enterprises based on regression analysis and machine learning, comprehensively considering various influencing factors of urban facilities, and capable of providing more accurate enterprise distribution prediction results. Through these data, urban planners can identify the key influencing factors for the development of the cultural and creative industries, and optimize the industrial layout accordingly to guide industrial agglomeration. In addition, these data can also provide location selection suggestions for practitioners, reduce the risk of starting a business, and enhance industrial competitiveness. Description of the Drawings
[0033] The drawings forming a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0034] Figure 1Schematic flowchart of a method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning according to an embodiment of the present invention;
[0035] Figure 2 Schematic flowchart of the working process of the regression analysis unit according to an embodiment of the present invention;
[0036] Figure 3 Schematic flowchart of the working process of the machine learning prediction unit according to an embodiment of the present invention;
[0037] Figure 4 Schematic structural diagram of a system for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning according to an embodiment of the present invention. Detailed implementation manners
[0038] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0039] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0040] As Figure 1 shown, this embodiment provides a method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning, including:
[0041] Dividing the target area into grid cells with a preset granularity, and collecting urban facility data, enterprise distribution data, and population density data within each grid cell;
[0042] Constructing an evaluation system, reclassifying the urban facility data, and calculating the richness of each type of facility;
[0043] Using a negative binomial regression model to analyze the influence of the evaluation system variables on the distribution of cultural and innovative enterprises, and extracting the regression residuals as new variables;
[0044] Inputting the urban facility data, enterprise distribution data, population density data, and regression residuals into a distributed random forest model for training to generate an enterprise distribution prediction result;
[0045] Generating a visualization graph according to the prediction result to display the density and spatial pattern of enterprise distribution in each region.
[0046] In this implementation, the specific implementation method includes:
[0047] The target area is pre-segmented using a 1 km x 1 km fishing net grid to establish initial divided areas. Each grid has a minimum unit granularity of 1 kilometer, and a regional dataset is constructed by collecting basic data such as the distribution of urban facilities, the number of enterprises, and the road length;
[0048] In the embodiment of the present application, the 1 km×1 km fishing net grid refers to a way of finely grid-dividing the target area. The regional dataset is the collection of basic data within the target area, and the key content collected includes: the distribution of creative and innovative enterprises (including POI data related to cultural and creative enterprises), urban facility data (such as consumer facilities, public facilities, transportation facilities, etc.), population density (reflecting the distribution of residents within the area), and road length. The minimum unit granularity refers to the minimum scale requirement for dividing the area, which directly affects the accuracy of data analysis. A division accuracy of 1 km is moderate, which can not only ensure the meticulousness of data collection but also effectively balance the consumption of computing resources and ensure efficient processing within a large-scale area.
[0049] First, according to the requirements of the minimum unit granularity, the target area is grid-divided using tools such as geographic information systems. Then, combined with existing statistical data and remote sensing image data, the basic data of the target area is collected, including but not limited to enterprise distribution, population density, road length, distribution of urban public facilities, etc., to complete the acquisition of various indicators within the grid, providing reliable data support for accurate prediction and analysis.
[0050] Construct an evaluation target, and determine that the evaluation target is the distribution of cultural and creative enterprises within the area, including companies in related industries such as art, design, media, and digital culture. These variables are used to analyze the spatial distribution of the target object within the target area;
[0051] In the embodiment of the present application, the evaluation target refers to the dependent variable in subsequent experiments, that is, the spatial distribution of cultural and creative enterprises within the target area. These enterprises cover multiple creative industry fields such as art, design, media, and digital culture, including various companies and institutions in related industries. This dependent variable not only represents the existence and scale of the cultural and creative industry within the area but also reflects the spatial characteristics and aggregation patterns of enterprise distribution, providing important data support for subsequent analysis.
[0052] First, select companies and institutions related to cultural innovation from the enterprise POI data. Then, with the help of technical tools such as geographic information system (GIS), match and connect these POI point data with the 1 km grid data, and then calculate and determine the number of cultural and creative enterprises contained in each grid unit. Through this process, the spatial distribution of cultural and creative enterprises in each area can be clearly obtained, providing basic data support for subsequent analysis and model prediction.
[0053] Construct an evaluation system, including the reclassification of point data such as consumer facilities, public facilities, transportation facilities, scenic spots, etc. By calculating the quantity and diversity of various facilities, richness is introduced to measure the distribution of various facilities within the target area;
[0054] In the embodiment of the present application, the evaluation system refers to the reclassification of various urban facility data within the target area, mainly including multiple facility types such as consumer facilities, public facilities, transportation facilities, scenic spots, etc., so as to provide quantitative indicators for the distribution of various facilities within the area. By calculating the quantity and diversity of various facilities, richness is introduced as a measurement standard. Richness refers to the distribution diversity and density of various facilities within each grid cell, reflecting the adequacy and diversity of facility configuration within the area, and is used to measure the distribution of various facilities within the target area.
[0055] First, load the data of various facilities, including consumer facilities, public facilities, scenic green spaces, and transportation facilities, etc., to obtain relevant data such as the spatial location and attribute information of each facility. Then, use geographic information system (GIS) tools to match and connect these facility data with a 1km×1km unit fishing net to ensure that each facility can be accurately mapped to the corresponding grid cell. Next, according to the characteristic types of the facilities, they are classified and integrated into four categories: consumer facilities, public facilities, scenic green spaces, and transportation facilities. Each type of facility is statistically analyzed separately according to its distribution within the target area.
[0056] After classification and integration, calculate the richness of each type of facility according to the formula of richness.
[0057]
[0058] Among them, p ik is the proportion of facility type k in grid cell i (such as the proportion of this facility type in all facilities), and n is the total number of different facility types in grid cell i;
[0059] The calculation of richness is based on the quantity, type diversity, and distribution density of facilities, and measures the configuration degree and diversity of various facilities within each grid cell. For example, the richness of consumer facilities depends on the quantity and type diversity of consumer places such as shopping malls and restaurants, and the richness of transportation facilities is related to the distribution density of transportation nodes. Through this process, detailed facility richness indicators can be generated for each grid cell, providing data support for subsequent analysis and prediction.
[0060] Use the evaluation target and evaluation system variables to perform regional regression calculations, and adopt negative binomial regression fitting. Analyze the influence of each variable in the evaluation system on the evaluation target, and add the residual variable of the regression as a new variable to the subsequent experiment;
[0061] In the embodiments of the present application, negative binomial regression considers the influence of independent variables such as facility richness and traffic accessibility by fitting the distribution of cultural and creative innovation enterprises in the target area. At the same time, the residuals in the regression results reflect the additional changes in the enterprise distribution except for the known factors, and these residual variables provide further input for the subsequent prediction model to optimize the prediction results.
[0062] Furthermore, as Figure 2 shown, this step specifically includes:
[0063] Perform regional regression calculation using the evaluation objective and evaluation system variables, and adopt negative binomial regression for fitting.
[0064] In the embodiments of the present application, regional regression calculation refers to the process of quantitatively analyzing the relationship between the evaluation objective and the evaluation system variables in the target area using statistical methods. The negative binomial regression model is a regression analysis method suitable for count data, especially suitable for dealing with the situation of over-dispersed data.
[0065] First, determine the evaluation objective, that is, analyze the distribution of cultural and creative enterprises in the area, and these enterprises include but are not limited to companies in related industries such as art, design, media, and digital culture. Next, construct an evaluation system, which covers data of different types of facilities such as consumer facilities, public facilities, transportation facilities, and scenic spots. These facilities are sorted through reclassification, and the concept of richness is further introduced to measure the quantity distribution and diversity of various facilities in the target area. Based on these data, perform regional regression calculation and adopt the negative binomial regression model for fitting to capture the complex relationship between the distribution of cultural and creative enterprises and the evaluation system variables.
[0066] Analyze the influence of each variable in the evaluation system on the evaluation objective, and evaluate the significance and influence of the variables.
[0067] In the embodiments of the present application, variable influence analysis refers to the process of detailed evaluation of how each variable in the evaluation system affects the evaluation objective. Through statistical analysis methods such as correlation analysis and significance test of regression coefficients, it can be identified which variables have a significant impact on the distribution of cultural and creative enterprises, and further evaluate the magnitude of the influence of these variables. This analysis helps to reveal the main driving factors of the distribution of cultural and creative enterprises and provides a scientific basis for urban planning and industrial layout.
[0068] Optimize and adjust the model to ensure the stability and reliability of the model.
[0069] In the embodiments of the present application, model optimization and adjustment refer to the process of further improving the initially established negative binomial regression model to enhance its prediction accuracy and stability. In this process, methods such as adjusting model parameters, selecting more appropriate variables, or adopting a more complex model structure may be involved. By using technical means such as cross-validation and residual analysis, the performance of the model can be evaluated, and necessary adjustments can be made based on the evaluation results. This step is crucial for ensuring that the model can adapt to different datasets and maintain stable prediction performance, thereby enhancing the practical application value of the model.
[0070] Calculate the residuals of the regression model to evaluate the fitting effect and prediction ability of the model.
[0071] In the embodiments of the present application, residual analysis refers to the process of quantitatively analyzing the prediction errors of the regression model. Residuals are the differences between the actual observed values and the model prediction values. By calculating the residuals, the fitting effect of the model on the data can be evaluated. Smaller residuals indicate a better fitting effect of the model on the data, and the distribution of residuals can also provide further clues about the model's prediction ability. For example, if there are certain specific patterns or trends in the residuals, it may indicate that the model fails to capture some key relationships in the data. Through residual analysis, not only can the deficiencies of the model be discovered, but also a basis for model optimization and improvement can be provided.
[0072] Add the residual variables of the regression as new variables to subsequent experiments to enhance the explanatory power and prediction accuracy of the model.
[0073] In the embodiments of the present application, adding the regression residual variables as new variables to subsequent experiments refers to the process of using the prediction errors of the model to enhance the explanatory power and prediction accuracy of the model. Residual variables can capture the additional variations that the model fails to explain. By introducing them as new variables into the model, it can help the model better understand the complexity of the data and improve its prediction ability for the distribution of cultural and creative enterprises. This process plays an important role in enhancing the practical applicability and long-term reliability of the model.
[0074] Input various types of collected data into the DRF model for training the urban creative and innovative enterprise distribution prediction model. Combine various independent variables (such as facility richness, traffic accessibility, etc.) and the residual data extracted from the regression analysis to capture the complex non-linear relationships of the urban creative and innovative enterprise distribution and provide accurate enterprise distribution prediction results;
[0075] In the embodiments of the present application, the Distributed Random Forest (DRF) model is an ensemble learning method that processes large-scale data by parallelly training multiple decision trees. This model can effectively capture the complex non-linear relationships between variables and has high accuracy and stability when dealing with high-dimensional data. By using various types of facility data and residual data in regression analysis as inputs, the DRF model can optimize the prediction results of enterprise distribution, further improve the prediction ability of the model, and effectively cope with the complexity and variability in the distribution of urban creative and innovative enterprises.
[0076] Further, as Figure 3 shown, this step specifically includes:
[0077] Input various types of collected data into the DRF model
[0078] In this embodiment, the establishment of the regional dataset depends on the collection of basic data within the target region, specifically including but not limited to the distribution of creative and innovative enterprises (relevant POI data), urban facility data (such as consumption facilities, public facilities, transportation facilities, etc.), and road lengths. In addition, the residual data extracted by the negative binomial regression model during regression analysis is also added as a new variable to the subsequent model training to further optimize the prediction results. These data provide solid data support for the subsequent model training and prediction and enhance the model's adaptability to complex distribution patterns.
[0079] Train the prediction model for the distribution of urban creative and innovative enterprises
[0080] In the model training stage, the original data will be divided into a training set, a test set, and a validation set according to a predetermined ratio. This division aims to ensure that the model can effectively learn and validate on different data subsets, thereby improving its generalization ability. Through training with the DRF model, the model will learn based on the enterprise distribution patterns in the training set data. This process involves in-depth analysis and pattern recognition of the data to ensure that the model can accurately capture the internal laws of the distribution of urban creative and innovative enterprises.
[0081] Capture the complex non-linear relationships in the distribution of urban creative and innovative enterprises
[0082] Using the ensemble learning mechanism of the DRF model, this step aims to identify and fit the complex non-linear relationships in the distribution of urban creative and innovative enterprises. By integrating multiple decision trees, the DRF model can effectively handle highly complex data patterns, thereby providing more accurate prediction results. This mechanism enables the model to adapt to various changes in the data, including non-linear and interaction effects, further improving the accuracy of the prediction.
[0083] Provide accurate prediction results for enterprise distribution
[0084] After completing model training and complex relationship recognition, this step will output a density map of enterprise distribution. The density map intuitively shows the density and spatial pattern of enterprise distribution in each region, providing effective visualization information for urban planners. Through this information, urban planners can identify potential enterprise agglomeration areas and resource gaps, thereby optimizing resource allocation. This process not only supports current urban planning but also provides strong data support for future urban development and industrial layout.
[0085] Conduct visual analysis on the model operation results to display the density and spatial pattern of enterprise distribution in each region. Help urban planners identify potential enterprise agglomeration areas and resource gaps, and reasonably adjust resource allocation according to actual needs and development strategies, providing strong data support for future urban planning and industrial layout;
[0086] In the embodiment of this application, the visual analysis of the model operation results refers to the process of displaying the model prediction results in a graphical way, mainly showing the density and spatial pattern of enterprise distribution in each region. Specifically, the enterprise distribution density refers to the number of creative and innovative enterprises per unit area, reflecting the concentration degree of enterprises in the region; while the spatial pattern describes the distribution trend of enterprises in the urban area, including whether there are obvious agglomeration phenomena or dispersion trends.
[0087] Through the visual results, urban planners can clearly understand the current situation of enterprise development in each region and reasonably adjust resource allocation according to actual needs and development strategies. For example, when it is found that there is excessive agglomeration of enterprises in some regions, it can be considered to optimize the allocation of public resources or promote the agglomeration of enterprises in other regions; for resource gaps, relevant industry enterprises can be attracted to settle in through means such as policy support and infrastructure construction. This process can not only optimize the current urban resource allocation but also provide a scientific basis for future urban planning and industrial layout, promoting the reasonable planning of urban functional areas and the optimization of the industrial chain.
[0088] As Figure 4 shown, this embodiment provides an urban creative and innovative enterprise distribution prediction system based on regression analysis and machine learning, including:
[0089] Data collection and regional division module, evaluation system construction module, regression analysis and residual extraction module, model training and prediction module, prediction result visualization module;
[0090] The data collection and regional division module is used to divide the target area into grid cells with a preset granularity and collect urban facility data, enterprise distribution data, and population density data in each grid cell;
[0091] The evaluation system construction module is used to construct an evaluation system, reclassify the urban facility data, and calculate the richness of each type of facility;
[0092] The regression analysis and residual extraction module is used to analyze the impact of the evaluation system variables on the distribution of cultural innovation enterprises by using the negative binomial regression model, and extract the regression residuals as new variables;
[0093] The model training and prediction module is used to input the urban facility data, enterprise distribution data, population density data, and regression residuals into a distributed random forest model for training to generate an enterprise distribution prediction result;
[0094] The prediction result visualization module is used to generate a visualization graph according to the prediction result to display the density and spatial pattern of the enterprise distribution in each region.
[0095] An urban creative innovation enterprise distribution prediction device based on regression analysis and machine learning includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the method described above is implemented.
[0096] A computer program product, when running on an urban creative innovation enterprise distribution prediction device based on regression analysis and machine learning, causes the device to execute the method described above.
[0097] The present invention discloses an urban creative innovation enterprise distribution prediction method, system, device, and product based on regression analysis and machine learning, which comprehensively considers various influencing factors of urban facilities and can provide a higher-precision enterprise distribution prediction result. Through these data, urban planners can identify the key influencing factors for the development of the cultural and creative industries and optimize the industrial layout accordingly to guide industrial agglomeration. In addition, these data can also provide location selection suggestions for practitioners, reduce the entrepreneurial risk, and enhance the industrial competitiveness.
[0098] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning, characterized in that: include: Divide the target area into grid units of preset granularity, and collect urban facility data, enterprise distribution data and population density data in each grid unit; Constructing an evaluation system to reclassify the urban facility data and calculate the richness of each type of facility; The negative binomial regression model is used to analyze the impact of the evaluation system variables on the distribution of cultural innovation enterprises, and the regression residuals are extracted as new variables; Inputting the urban facility data, enterprise distribution data, population density data and regression residuals into a distributed random forest model for training to generate enterprise distribution prediction results; Generate visualization graphics based on the prediction results to show the density and spatial patterns of enterprise distribution in each region.
2. The method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning as claimed in claim 1, characterized in that: The process of dividing the target area into grid units of preset granularity includes: using a 1km×1km fishing net grid to grid the target area to form uniform grid units; mapping enterprise POI data, urban facility data and road length data to corresponding grid units through a geographic information system tool.
3. The method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning as claimed in claim 1, characterized in that: The urban facility data includes consumer facilities, public facilities, transportation facilities and scenic area data. The process of calculating the richness of each type of facility includes: Among them, R i is the richness of each type of facility; p ik is the proportion of facility type k in grid cell i; n is the total number of different facility types in grid cell i.
4. The method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning as claimed in claim 1, characterized in that: The process of analyzing using the negative binomial regression model includes: Regression fitting was performed with the number of cultural innovation enterprises as the dependent variable and facility richness, transportation accessibility and population density as independent variables, and the regression residuals were extracted as additional variables to characterize the unexplained variation.
5. The method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning as claimed in claim 1, characterized in that: The training process of the distributed random forest model includes: The original data is combined with the regression residual variables, divided into training set, test set and validation set, and the nonlinear relationship and interaction effect between variables are captured through multiple decision tree integration.
6. The method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning as claimed in claim 1, characterized in that: The process of generating a visualization graph includes: Output the enterprise distribution density map, which uses a color gradient map to show the concentration of enterprises in various areas of the city.
7. The method for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning as claimed in claim 1, characterized in that: The visualization graph also includes: marking resource blank areas and clustered areas, and generating resource allocation suggestions.
8. A system for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning according to any one of claims 1 to 7, characterized in that: Data collection and regional division module, evaluation system construction module, regression analysis and residual extraction module, model training and prediction module, prediction result visualization module; The data collection and area division module is used to divide the target area into grid units of preset granularity, and collect urban facility data, enterprise distribution data and population density data in each grid unit; The evaluation system construction module is used to construct an evaluation system, reclassify the urban facility data and calculate the richness of each type of facility; The regression analysis and residual extraction module is used to analyze the impact of the evaluation system variables on the distribution of cultural innovation enterprises using a negative binomial regression model, and extract regression residuals as new variables; The model training and prediction module is used to input the urban facility data, enterprise distribution data, population density data and regression residuals into a distributed random forest model for training to generate enterprise distribution prediction results; The prediction result visualization module is used to generate visualization graphics based on the prediction results to show the density and spatial pattern of enterprise distribution in each region.
9. A device for predicting the distribution of urban creative and innovative enterprises based on regression analysis and machine learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method as described in any one of claims 1 to 7 when executed by the processor.
10. A computer program product, when the computer program product is run on a city creative and innovative enterprise distribution prediction device based on regression analysis and machine learning, causes the device to execute the method as described in any one of claims 1 to 7.