An intelligent matching investment strategy system and method based on enterprise needs

By constructing an enterprise demand scoring model using a BP neural network and an improved random forest algorithm, the problems of information asymmetry and blind spots in the enterprise investment promotion process in existing technologies are solved, and accurate matching of enterprise needs and efficient investment promotion are achieved.

CN114529038BActive Publication Date: 2025-11-25SHANGHAI AFAXIDI DATA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111678574.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-11-25
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Existing technologies suffer from information asymmetry and blind spots in the process of attracting investment, resulting in a mismatch between the companies being attracted and the local economy, making it difficult to achieve precise investment attraction. Furthermore, existing algorithms have a simple logic and are unable to quickly obtain matching data for companies.

Method used

By employing the LM-BP neural network algorithm optimized based on BP neural network and the improved random forest algorithm, and through enterprise demand profiling, the regional industry, policy, location and building characteristics are analyzed to construct an enterprise demand scoring model. The rich algorithm logic is used to calculate enterprise demand assessment to ensure accurate investment attraction.

Benefits of technology

It achieves precise matching of enterprise needs, improves the targeting and efficiency of investment promotion, and quickly obtains suitable industry, policy, location and building resources through deep learning and machine learning technologies, thereby reducing investment promotion risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529038B_ABST
    Figure CN114529038B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent matching strategy system for attracting investment based on enterprise demand, comprising: an enterprise demand evaluation module, which obtains enterprise demand matching data through LM-BP neural network deep learning; an investment strategy module, which establishes an enterprise demand scoring model based on random forest improvement, and classifies and predicts enterprise demand characteristics; an industry matching module, which selects an industry suitable for the enterprise through an industry matching library; a policy matching module, which selects a policy suitable for the enterprise through a policy matching library; a regional space matching module, which selects a regional space suitable for the enterprise through a matching regional space library; and a building matching module, which selects a building suitable for the enterprise through a matching planning building library. The application also discloses a method for the intelligent matching strategy system for attracting investment. The application optimizes and improves BP neural network and random forest algorithm to obtain demand data and build an enterprise demand scoring model, more effectively calculates enterprise demand evaluation, and ensures accurate investment attraction through comprehensive evaluation of investment enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent investment promotion technology, and in particular to an intelligent matching investment promotion strategy system and method based on enterprise needs. Background Technology

[0002] Targeted investment attraction is an inevitable choice for enhancing the core competitiveness of regional economies and adapting to the regularity of investment attraction activities. How to improve our investment attraction efforts, implement targeted investment attraction, and optimize the industrial structure are key areas of current research. Among these, the most important is the issue of targeted investment attraction: how to assess enterprise needs to obtain structured data on industries, policies, location, and architecture that match those needs; how to use this data to build an enterprise needs scoring model; and how to classify and predict the characteristics of enterprise needs. These are our primary concerns.

[0003] In reality, factors such as performance evaluation and information asymmetry still lead to many regions blindly attracting investment without fully considering the precision and industry fit of the projects. This results in attracted companies failing to drive local economic development and even requiring government support, thus hindering local economic growth. The core of precise investment attraction is to improve its targeting, avoid arbitrariness in investment activities, reduce investment risks, and ensure that the attracted companies better meet the requirements of local economic development.

[0004] An invention patent application with application number 201811058230.8 discloses a business location system based on an environmental evaluation index for investment cooperation carriers. The system includes a location information collection unit to guide users in entering desired address information, a GIS unit to generate maps of recommended business location areas, a carrier query unit to provide query services for information on other investment carriers besides the recommended location options, a location selection unit to analyze business location areas, an industrial cluster analysis and evaluation database to store industrial cluster analysis and evaluation indices, a location element database to store location element scores, and a business basic information database to store the industry category, business type, sales revenue, and tax payment of businesses according to specified regional divisions. Its significant effect is that it can classify six investment methods—investment in technology, talent, intelligence, platforms, and cooperation—through the carrier environment index calculation unit. Separate cluster evaluation modules and carrier resource evaluation modules are designed for each method. If the investment method is investment or technology, the cluster evaluation module calculates the industry cluster evaluation index; if the investment method is talent, intelligence, platforms, or cooperation, the carrier resource evaluation module calculates the business environment analysis index. The industry cluster evaluation index or business environment analysis index is ranked, and the three highest-scoring carrier environments are selected to recommend to enterprises, making enterprise site selection more scientific and accurate.

[0005] However, the invention patent application with application number 201811058230.8 mainly uses open source algorithms. The algorithm logic is relatively simple. When seeking standardized decisions, it is not easy to obtain positive and negative ideal solutions through linear derivation. Therefore, it is impossible to quickly and accurately obtain enterprise matching data and it is difficult to efficiently build an investment promotion strategy scoring model. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides an intelligent matching investment promotion strategy system and method based on enterprise needs. Using this system and method, in terms of acquiring accurate matching data for enterprise needs, the LM-BP neural network algorithm is optimized and improved based on a BP neural network. Furthermore, enterprise need profiles are used to match regional industrial value characteristics, regional policy characteristics, regional location advantages, and regional building suitability characteristics. An improved random forest algorithm is used to acquire various models and indicators related to local regional industrial value, policies, location advantages, building resources, policy advantages, location advantages, and building advantages to build a more efficient enterprise need scoring model. The rich algorithmic logic enables more effective calculation of enterprise need assessments, and comprehensive evaluation of investment enterprises ensures accurate investment promotion.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0008] A smart matching investment promotion strategy system based on enterprise needs is provided, including: an enterprise needs assessment module, an investment promotion strategy module, an industry support matching module, a policy matching module, a regional space matching module, and a building matching module;

[0009] The enterprise demand assessment module is used to obtain structured data on industry, policy, location, and architecture that match enterprise demand by deep learning of regional industry value characteristics analysis, policy characteristics analysis, location characteristics analysis, and building characteristics analysis through LM-BP neural network.

[0010] The investment promotion strategy module is used to score the enterprise demand profile and the matching local industries, policies, location and building resources, construct an enterprise demand investment promotion strategy report scoring model, establish an enterprise demand scoring model based on random forest, and classify and predict the characteristics of enterprise demand.

[0011] The industry matching module is used to analyze the scale of local investment promotion areas, the total number of related enterprises in the industrial chain, the number of large-scale enterprises, customers and suppliers of enterprises in their respective sub-industries through the industry matching database, and to provide reference values ​​of the enterprise's industrial agglomeration effect and supply and sales relationship effect, so as to match and select the optimal industry for the enterprise.

[0012] The policy matching module is used to analyze policies suitable for enterprises through a policy matching library. It recommends general policies suitable for enterprises in their startup, growth, and maturity stages, industry-specific policies suitable for all sub-sectors of the enterprise, and policies that cultivate the enterprise's prospects. This allows investment promotion personnel to follow up on the specific needs of enterprises and mark them in the policy matching module, thereby matching and selecting the optimal policy suitable for the enterprise.

[0013] The regional space matching module is used to match the data formatted by the regional space database with the quantitative evaluation and scoring of enterprise needs, and to conduct an adaptability evaluation by matching the enterprise needs with macro location, meso location, micro location, land use planning, transportation and logistics, and living facilities, thereby matching and selecting the optimal regional space suitable for the enterprise.

[0014] The building matching module is used to select the optimal building carrier for the enterprise by matching the building planning library. It evaluates the building carrier based on basic building information, building structure, energy conservation and environmental protection, fire prevention and explosion protection, supporting equipment, and occupancy costs, thereby matching and selecting the optimal building plan for the enterprise.

[0015] To solve its technical problem, the present invention further adopts the following technical solution:

[0016] Furthermore, in the enterprise demand assessment module, the industry structured data for matching enterprise demand includes: industry scale, value chain customer market, industrial chain supporting facilities and enterprise supporting facilities; the policy structured data for matching enterprise demand includes: industrial policies, talent policies and financial policies; the location structured data for matching enterprise demand includes: transportation and logistics, supporting resources and planning elements; and the building structured data for matching enterprise demand includes: basic building elements, energy conservation and environmental protection, load-bearing capacity, fire prevention and explosion protection and usage costs.

[0017] Furthermore, the specific calculation steps of the LM-BP neural network described in step 1 are as follows:

[0018] Step 1.1: Initialize the network structure parameters, with the error tolerance value being ε, constants u and b, and initialize the weight and threshold vectors. Let k = 0, u = u0, and calculate the precision ε and the maximum number of learning iterations M.

[0019] Step 1.2: Input the training data of the enterprise demand profile index matrix as the input vector into the LM-BP neural network;

[0020] Step 1.3: Calculate the network output and error index function e;

[0021] Step 1.4: Calculate the Jacobi matrix J[W(k)]; where W(k) represents the vector composed of the threshold and weights in the k-th neural network iteration;

[0022] Step 1.5: Calculate ΔW; where, ΔW is the threshold change amount;

[0023] Step 1.6: If e < v, then go to Step 1.8, otherwise go to Step 1.5;

[0024] Step 1.7: Calculate the error function e with the new weight and threshold vector W(k + 1),

[0025] W(k + 1) = W(k) - {J T [W(k)]J[W(k)]} -1 J[W(k)]e[W(k)]

[0026] If e[W(k + 1)] is less than e[W(k)], then let k = k + 1, u = u * b, and go to Step 1.2, otherwise u = u / b, and go to Step 1.5; where, W(k) represents the vector composed of the threshold and weights of the k-th neural network iteration, and W(k + 1) represents the vector composed of the threshold and weights of the new (k + 1)-th iteration;

[0027] Step 1.8: The LM-BP neural network calculation ends.

[0028] Furthermore, in Step 1.1, the value range of b is: 0 < b < 1. When k = 0 and u = u0, calculate the calculation accuracy ε and the maximum number of learning times M.

[0029] Furthermore, in Step 1.4 and Step 1.7, through the deformation of the Jacobian matrix J[W(k)], calculate J T (W), calculate J T (W) with the formula:

[0030]

[0031] Furthermore, the steps of the enterprise demand scoring model established based on the improved random forest in Step 2 include:

[0032] Step 2.1: Preprocess the enterprise demand data;

[0033] Step 2.2: Calculate the enterprise demand matrix;

[0034] Step 2.3: Data weighted sampling;

[0035] Step 2.4: Select the optimal demand feature subset of the enterprise demand by the feature selection method;

[0036] Step 2.5: Optimize the algorithm parameters;

[0037] Step 2.6: Generate the evaluation result.

[0038] Furthermore, in step 2.1, the enterprise demand data preprocessing includes Min-Max normalization and Z-Score normalization.

[0039] The Min-Max standardization process involves linearly transforming the offline data in the enterprise demand data so that the transformed data falls within the range of [0-1].

[0040] The Z-Score standardization process transforms the enterprise demand data into a Gaussian distribution with a mean of 0 and a standard deviation of 1.

[0041] Furthermore, in step 2.2, the specific steps for calculating the enterprise demand matrix are as follows: Let X = {b1, b2, ..., bL} represent the set of L samples with M features, and Y = {y1, y2, ..., yL} represent the category set. Then, the enterprise demand data can be used to construct a matrix as follows:

[0042]

[0043] The size of matrix L is L(M+1), where +1 represents the set of categories, bi = {Xi1, Xi2, ..., XiM} represents the M feature values ​​of sample bi, and Xij represents the j-th feature value of sample bi.

[0044] In the enterprise demand matrix L, there are minority samples L″ and majority samples L′. Taking Q samples from the minority samples L″, the matrix form is as follows:

[0045]

[0046] If there are Q samples in the minority sample L″ and L samples in total, then there are (LQ) samples in the majority sample L′. Therefore, its matrix form is:

[0047]

[0048] Furthermore, in step 2.3, the specific steps of the data weighted sampling include:

[0049] Step 2.3.1: Divide the original enterprise demand data into training set L and training set L1;

[0050] Step 2.3.2: Divide the training set L into two subsets, namely the majority class samples L′ and the minority class samples L″;

[0051] Step 2.3.3: During the sampling process, first perform weighted sampling on the majority class samples L′, select samples of similar size from the minority class samples L″ in the majority class samples L′, calculate the proportion of the selected minority L″ samples L′, and calculate the proportion of L″ in all L samples. Then, perform weighted sampling to select the final training samples.

[0052] Step 2.3.4: Repeat step 2.3.3 multiple times until a balanced sample is selected;

[0053] Step 2.3.5: Select balanced samples and divide them into training set and test set.

[0054] Furthermore, in step 2.4, the input is the original enterprise demand dataset D = {(x1,y1),(x2,y2),……,(xn,yn)},x i ∈R m , and y n ∈{-1, 1}; Let g1, g2;

[0055] The output is the optimal feature subset f;

[0056] The specific steps for selecting the optimal subset of enterprise demand features using the feature selection method include:

[0057] Step 2.4.1: Define M enterprise demand characteristics i = 1, 2, 3, 4, ..., M;

[0058] Step 2.4.2: Calculate the corresponding value for each enterprise's demand characteristic using the following formula;

[0059] Let D be a sample dataset, x and y be arbitrary attributes of the samples, and n be the number of categories in dataset D. Then the information entropy of x is:

[0060]

[0061] Where P(xi) is the probability that the characteristic attribute x takes the value xi;

[0062] Given feature attribute y, the conditional entropy of feature attribute x is:

[0063]

[0064] Where p(yi) is the probability that the characteristic attribute y takes the value yj, and p(xi|yi) is the probability that the characteristic attribute x takes the value xi when the characteristic attribute y takes the value yj.

[0065] The information entropy obtained from the above formula is:

[0066] Gain(x,y) = Info(x) - Info(x|y)

[0067] Select the feature with the largest information gain as the splitting attribute of dataset D, create a node, use this feature as a label, create branches for each value of the feature, and divide the enterprise needs of the sample accordingly.

[0068] Step 2.4.3: Calculate the entropy comparison value un of each feature and the category variable yn using the following formula;

[0069] Step 2.4.4: If un is greater than or equal to g1, then feature xn is in the selected optimal feature subset f, i.e., x n ∈f;

[0070] The features are sorted, the selected features within set f are measured, and the correlation value S between features xi and xj is determined.

[0071] Step 2.4.5: When S is less than or equal to g2, follow up with the information entropy comparison value un in step (3) to delete the characteristics in set f;

[0072] Step 2.4.6: Obtain the optimal feature subset.

[0073] Furthermore, in step 2.5, the specific optimization steps for optimizing the algorithm parameters include:

[0074] Step 2.5.1: Set the search range and step size for the parameters to be optimized;

[0075] Step 2.5.2: Based on step 2.5.1, further calculate the mean absolute error values ​​of the two parameters S and C, and use the mean absolute error values ​​to obtain the specific range of the number of the two parameters S and C;

[0076] Step 2.5.3: Based on the value range of parameters S and C obtained in Step 2.5.2, calculate the OOB value of the random forest using the S*C combination and the following process to obtain the accuracy;

[0077] During each training run, the unsampled data is labeled as set OOBi. The number of misclassified data in the unsampled dataset OOBi is labeled as ErrorNumOOB. Finally, the error of the random forest OOB value is defined as:

[0078]

[0079] That is, the generalization error is:

[0080]

[0081] Step 2.5.4: Select the optimal parameters determined by the S*C combination based on the OOB value. If the OOB value of the random forest meets the requirements, output the S*C combination; otherwise, change the search range and step size and continue the search until the final condition is met.

[0082] Furthermore, in step 2.6, the enterprise demand scoring model based on random forest generates the optimal evaluation result through steps 2.1 to 2.5, which serves as an evaluation reference and is provided to the industrial park's investment promotion personnel to determine whether the enterprise is suitable for the park's investment promotion strategy.

[0083] A method for an enterprise intelligent matching investment promotion strategy system is also provided, comprising the following steps:

[0084] S1: The enterprise demand assessment module uses an LM-BP neural network to perform deep learning on regional industry value characteristics analysis, policy characteristics analysis, location characteristics analysis, and building characteristics analysis to obtain structured data on industry, policy, location, and building that match enterprise demand.

[0085] S2: The investment promotion strategy module scores the enterprise demand profile and the matching local industries, policies, location, and building resources, constructs an enterprise demand investment promotion strategy report scoring model, establishes an enterprise demand scoring model based on random forest, and classifies and predicts the characteristics of enterprise demand.

[0086] S3: The industry matching module analyzes the scale of local investment promotion areas, the total number of related enterprises in the industrial chain, the number of large-scale enterprises, customers and suppliers of enterprises in their respective sub-industries through the industry matching database. It also provides reference values ​​for the enterprise's industrial agglomeration effect and supply and sales relationship effect, thereby matching and selecting the optimal industry for the enterprise.

[0087] S4: The policy matching module analyzes policies suitable for enterprises through the policy matching library. It recommends general policies suitable for enterprises in the startup, growth and maturity stages, industry-specific policies suitable for all sub-sectors of the enterprise, and policies that cultivate the enterprise's prospects. This allows investment promotion personnel to follow up on the specific needs of enterprises and make recommendations in the policy matching module, thereby matching and selecting the best policy suitable for the enterprise.

[0088] S5: The regional space matching module matches the formatted data of the regional space database with the quantitative evaluation and scoring of enterprise needs. It also performs an adaptability evaluation based on the enterprise needs, including macro-location, meso-location, micro-location, land use planning, transportation and logistics, and living facilities, thereby matching and selecting the optimal regional space suitable for the enterprise.

[0089] S6: The building matching module selects the optimal building carrier for the enterprise by matching the planning building library. It evaluates the building carrier based on basic building information, building structure, energy conservation and environmental protection, fire prevention and explosion protection, supporting equipment, and occupancy costs, thereby matching and selecting the optimal building plan for the enterprise.

[0090] The beneficial effects of this invention are:

[0091] I. The intelligent matching investment promotion strategy system and method based on enterprise needs of this invention, in terms of acquiring accurate matching data for enterprise needs, optimizes and improves the LM-BP neural network algorithm based on the BP neural network, and matches regional industrial value characteristics, regional policy characteristics, regional location advantages, and regional building suitability characteristics through enterprise needs profile analysis. The improved random forest algorithm further acquires various models and indicators of local regional industrial value, policies, location advantages, building resources, policy advantages, location advantages, and building advantages to build a more efficient matching enterprise needs scoring model. The rich algorithm logic can more effectively calculate enterprise needs assessment, and ensure accurate investment promotion through comprehensive evaluation of investment enterprises.

[0092] II. The present invention optimizes the LM-BP neural network algorithm based on the BP neural network. The Levenberg-Marquardt (LM) algorithm has the advantages of both the gradient method and Newton's method. In order to alleviate the singularity problem of non-optimal points, the second derivative is used to approximate quadraticity when the objective function is close to the optimal point, so as to accelerate the optimization convergence process. It is much faster than the gradient method and the BP algorithm, and optimizes the efficiency of big data analysis and processing. Through the LM-BP neural network improved by the BP neural network algorithm, after deep learning for regional industrial value characteristic analysis, policy characteristic analysis, location characteristic analysis, and building characteristic analysis, the structured data of industry, policy, location, and building that accurately match the needs of enterprises can be obtained. This can help enterprises to deeply learn and explore their needs and obtain suitable local resources according to their needs.

[0093] III. The main function of this invention is to score enterprises based on their needs profiles and matching local industries, policies, locations, and building resources. It utilizes machine learning and statistics to construct a scoring model for enterprise demand-based investment promotion strategy reports. This model tracks enterprise demand characteristics, mines the relationships between different characteristics, and establishes a scoring model based on an improved random forest. It then classifies and predicts the status of enterprise demand characteristics, identifying enterprises with high matching scores for weighted indicators as precisely suitable for investment promotion. Because the random forest algorithm is relatively simple to implement, has fast training speed, strong generalization ability, and strong robustness, the improved random forest algorithm is applied to the construction of the enterprise demand scoring model. The improved random forest enterprise demand scoring model mainly includes enterprise demand data preprocessing, enterprise demand matrix, data weighted sampling, feature selection to select the optimal subset of enterprise demand features, algorithm parameter optimization, and generating evaluation results. Based on the improved random forest enterprise demand scoring model, through the above steps, it provides an effective evaluation reference for industrial park investment promotion personnel to determine whether enterprises are suitable for the park's investment promotion strategy.

[0094] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0095] Figure 1 This is a functional architecture diagram of an intelligent matching investment promotion strategy system based on enterprise needs, as described in this invention.

[0096] Figure 2 This is an industry structured data flowchart in the enterprise demand assessment module described in this invention (taking Class I and Class II indicators as examples);

[0097] Figure 3 This is a schematic diagram illustrating the specific principle and algorithm flow of the BP neural network described in this invention;

[0098] Figure 4 This is a schematic diagram of the LM-BP neural network algorithm, which is an improvement on the BP neural network described in this invention.

[0099] Figure 5 This is a schematic diagram illustrating the specific principles and algorithm flow of the random forest algorithm described in this invention;

[0100] Figure 6 This is one of the flowcharts illustrating the process of establishing an improved enterprise demand scoring model based on random forest as described in this invention;

[0101] Figure 7 This is a schematic diagram of the data weighted sampling process described in this invention;

[0102] Figure 8This is a flowchart illustrating the process of selecting the optimal subset of enterprise demand features using the feature selection method described in this invention.

[0103] Figure 9 This is a flowchart illustrating the parameter optimization process of the random forest algorithm described in this invention.

[0104] Figure 10 This is a flowchart illustrating a method for an intelligent matching investment promotion strategy system for enterprises, as described in this invention.

[0105] Figure 11 This is the second flowchart illustrating the process of establishing an improved enterprise demand scoring model based on random forest as described in this invention.

[0106] The parts in the attached diagram are labeled as follows:

[0107] The module includes: Enterprise Needs Assessment Module 1, Investment Promotion Strategy Module 2, Industrial Support Matching Module 3, Industrial Supporting Library 31, Policy Matching Module 4, Policy Matching Library 41, Regional Spatial Matching Module 5, Regional Spatial Library 51, Building Matching Module 6, and Planning and Building Library 61. Detailed Implementation

[0108] The following specific embodiments illustrate the detailed implementation of the present invention. Those skilled in the art can easily understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented in other different ways, that is, different modifications and changes can be made without departing from the scope disclosed in the present invention.

[0109] Example 1:

[0110] A smart matching investment promotion strategy system based on enterprise needs, such as... Figure 1 As shown, it includes: Enterprise Needs Assessment Module 1, Investment Promotion Strategy Module 2, Industrial Support Matching Module 3, Policy Matching Module 4, Regional Spatial Matching Module 5, and Building Matching Module 6;

[0111] Enterprise demand assessment module 1 is used to obtain structured data on industry, policy, location and building characteristics that match enterprise demand after deep learning of regional industry value characteristics analysis, policy characteristics analysis, location characteristics analysis and building characteristics analysis through LM-BP neural network;

[0112] The Investment Promotion Strategy Module 2 is used to score the enterprise demand profile and the matching local industries, policies, location, and building resources, construct an enterprise demand investment promotion strategy report scoring model, establish an enterprise demand scoring model based on random forest, and classify and predict the characteristics of enterprise demand.

[0113] The industry matching module 3 is used to analyze the local investment promotion area industry scale, total number of related enterprises in the industrial chain, number of large-scale enterprises, customer and supplier information of enterprises in their respective sub-industries through the industry matching library 31, and to provide reference values ​​of enterprise industry agglomeration effect and supply and sales relationship effect, so as to match and select the optimal industry suitable for the enterprise.

[0114] The policy matching module 4 is used to analyze and match policies suitable for enterprises through the policy matching library 41. It recommends general policies suitable for the enterprise's development stage (start-up, growth, and maturity), industry-specific policies suitable for all sub-sectors of the enterprise, and policies that cultivate the enterprise's prospects. This allows investment promotion personnel to follow up on the specific needs of the enterprise and make recommendations in the policy matching module 4, thereby matching and selecting the optimal policy suitable for the enterprise.

[0115] The regional space matching module 5 is used to match the data formatted by the regional space library 51 with the quantitative evaluation and scoring of enterprise needs, and to conduct an adaptability evaluation by matching the enterprise needs with macro location, meso location, micro location, land use planning, transportation and logistics, and living facilities, so as to match and select the optimal regional space suitable for the enterprise.

[0116] The building matching module 6 is used to select the optimal building carrier for the enterprise by matching the planning building library 61. It evaluates the adaptability of the building carrier by using basic building information, building structure, energy conservation and environmental protection, fire prevention and explosion protection, supporting equipment, and occupancy costs, thereby matching and selecting the optimal building plan for the enterprise.

[0117] The main function of the enterprise demand assessment module is based on the precise matching of enterprise demand data. It creates a demand profile of enterprises for local industries, policies, location, and buildings. Through the LM-BP neural network, which is an improvement on the BP neural network algorithm, it performs deep learning on the analysis of regional industrial value characteristics, policy characteristics, location characteristics, and building characteristics to obtain structured data on industry, policy, location, and building characteristics that are precisely matched to enterprise demand.

[0118] like Figure 2 As shown, in the enterprise demand assessment module, the industry structured data for matching enterprise demand includes: industry scale, value chain customer market, industrial chain supporting facilities and enterprise supporting facilities; the policy structured data for matching enterprise demand includes: industrial policies, talent policies and financial policies; the location structured data for matching enterprise demand includes: transportation and logistics, supporting resources and planning elements; and the building structured data for matching enterprise demand includes: basic building elements, energy conservation and environmental protection, load, fire prevention and explosion protection and usage costs.

[0119] The main Category I and Category II indicators are as follows (Category III indicators are too numerous to list here):

[0120] A. Industry:

[0121] Industry scale: industry output value, revenue of industrial enterprises above designated size, related enterprises in the industrial chain, high-tech enterprises, and listed companies;

[0122] Value chain customers: similar enterprises, similar high-tech enterprises, similar large-scale enterprises, similar enterprise evaluation scale, downstream enterprises, and downstream enterprise scale;

[0123] Supply chain support: upstream enterprises, scale of upstream enterprises, level / scale of regional brand promotion platforms, enterprise R&D centers, and industry support service platforms;

[0124] Supporting services for businesses: level / scale of talent training institutions, scale of financial support funds, scale of professional intermediary services, one-stop agency services, online government services, direct communication between government and enterprises, and talent apartments;

[0125] B. Policy:

[0126] Industrial development policies: industrial structure policies, industrial organization policies, and industrial regional layout policies;

[0127] Tax policies: Corporate income tax policy, individual income tax policy, value-added tax policy;

[0128] Talent policies: talent settlement policies, talent subsidy policies, talent housing and development policies;

[0129] Financial policies: R&D subsidies, innovation subsidies, and tax subsidies;

[0130] C. Location:

[0131] Transportation and Logistics: Distance to airports, high-speed rail / train stations, national highways / provincial highways / expressways, ports / docks, government agencies, and logistics;

[0132] Supporting resources: residential, commercial, medical, educational, green spaces, and daily life services;

[0133] Planning elements: park construction time, land use, land area, permitted industries, plot ratio, building attributes, building classification, and rental / sale ratio;

[0134] D. Architecture:

[0135] Basic building elements: total building area, building footprint, parking spaces, single-story building area, building contents, non-motorized vehicle parking spaces, floors, floor height, etc.

[0136] Energy saving and environmental protection: safety protection distance, radiation intensity, vibration, noise level limits, corrosion prevention;

[0137] Loads: permanent loads, variable loads, accidental loads;

[0138] Fireproof and explosion-proof: fire resistance rating, fire resistance limit, fire separation distance;

[0139] Usage costs: building availability, rent, property management fees, water consumption per 10,000 yuan of output value, and energy consumption per 10,000 yuan of output value.

[0140] Introduction to the principle of the LM-BP neural network, an improvement on the BP neural network algorithm:

[0141] The LM algorithm combines the advantages of both the gradient method and Newton's method. To mitigate the singularity problem at non-optimal points, it approximates the quadraticity of the objective function near the extreme point by utilizing the properties of the second derivative, thus accelerating the convergence process. It is significantly faster than the gradient method and the backpropagation (BP) algorithm, optimizing the efficiency of big data analysis and processing.

[0142] Detailed introduction:

[0143] Define the error function as

[0144]

[0145] Where w is the vector composed of the neural network threshold and weights; ei = (w) is the error. According to the Gauss-Newton method, the calculation is as follows:

[0146] W(k+1)=W(k)-[J T (W k )J(W k )] -1 J(W k )e(W k )

[0147] Where W(k) represents the vector composed of the threshold and weights in the k-th iteration of the neural network, and W(k+1) represents the vector composed of the threshold and weights in the (k+1)-th iteration.

[0148] The LM algorithm is an improved Gauss-Newton method, as follows:

[0149] W(k+1)=W(k)-[J T (w k )J(w k )+μ k I] -1 J(w k )e(w k )

[0150] Where I is the identity matrix, J is the Jacobian matrix, and uk is a scaling factor. A key step in the LM algorithm is the calculation of the Jacobian matrix, which is performed using a variation of the BP algorithm, as follows:

[0151]

[0152] If the scaling factor u = 0, it is equivalent to the Gauss-Newton algorithm; if the scaling factor is very large, the LM algorithm is close to the gradient algorithm. With each iteration, u decreases slightly. Therefore, when approaching the error target, it is closer to the Gauss-Newton algorithm, offering faster computation and higher accuracy. u is an exploratory parameter; for a given parameter u, if the calculated threshold change Δw reduces the error function e(w), then u decreases. Conversely, u increases. The LM algorithm utilizes information from the second derivative, making it much faster than the gradient algorithm. Practical test data demonstrates that the LM-BP algorithm is tens of times faster than the traditional BP gradient descent.

[0153] Since the BP neural network is publicly available and open source, the specific principles and algorithmic logic of the BP neural network itself will not be described in detail in this patented technical solution. Figure 3 As shown; Figure 4 As shown, this section mainly introduces the optimized and improved parts. The specific calculation steps of the LM-BP neural network include the following steps:

[0154] Step 1.1: Initialize the network structure parameters, with the error tolerance value being ε, constants u and b, and initialize the weight and threshold vectors. Let k = 0, u = u0, and calculate the precision ε and the maximum number of learning iterations M.

[0155] Step 1.2: Input the training data of the enterprise demand profile index matrix as the input vector into the LM-BP neural network;

[0156] Step 1.3: Calculate the network output and error index function e;

[0157] Step 1.4: Calculate the Jacobi matrix J[W(k)]; where W(k) represents the vector composed of the threshold and weights in the k-th neural network iteration;

[0158] Step 1.5: Calculate ΔW; where ΔW is the threshold change.

[0159] Step 1.6: If e < ε, then go to step 1.8; otherwise, go to step 1.5.

[0160] Step 1.7: Calculate the error function e using the new weight and threshold vector W(k+1).

[0161] W(k+1)=W(k)-{J T [W(k)]J[W(k)]} -1 J[W(k)]e[W(k)]

[0162] If \(e[W(k + 1)]\) is less than \(e[W(k)]\), then let \(k=k + 1\), \(u = u*b\), and go to step 1.2; otherwise, \(u = u / b\), and go to step 1.5; where \(W(k)\) represents the vector composed of the threshold and weights in the \(k\)-th neural network iteration, and \(W(k + 1)\) represents the vector composed of the threshold and weights in the new \((k + 1)\)-th iteration.

[0163] Step 1.8: The calculation of the LM-BP neural network ends.

[0164] The above calculates the industrial, policy, location, and building structured data for the precise matching of enterprise needs through the LM-BP neural network, and scores the indicators and weights matched by the enterprise demand portrait through the investment promotion strategy module.

[0165] In step 1.1, the value range of \(b\) is: \(0 < b < 1\). When \(k = 0\) and \(u = u0\), calculate the calculation accuracy \(\epsilon\) and the maximum number of learning times \(M\).

[0166] In steps 1.4 and 1.7, through the deformation of the Jacobian matrix \(J[W(k)]\), calculate \(J\) T (W), calculate \(J\) T (W) The formula is:

[0167]

[0168] Function introduction of the investment promotion strategy module:

[0169] The main function is to score based on the enterprise demand portrait and the matched local industries, policies, locations, and building resources, use machine learning and statistics to construct a scoring model for the enterprise demand investment promotion strategy report, follow up the enterprise demand characteristics, mine between its various characteristics, mine the relationships between different characteristics, establish an enterprise demand scoring model improved based on random forest, classify and predict the enterprise demand characteristic status, and enterprises with a large number and high matching degree of weight indicators are precisely adapted investment promotion enterprises.

[0170] Introduction to the algorithm principle and logic:

[0171] As Figure 5 shown, the basic component unit of the random forest algorithm is the decision tree. The main idea of this algorithm is to randomly draw \(n\) samples with replacement from the original sample set \(S\) to generate a new sample set, and then generate \(n\) classification trees according to the bootstrap sample set to form a random forest. The final classification result is determined by the number of votes of the classification trees.

[0172] Since the random forest algorithm is relatively simple to implement, has a fast training speed, strong generalization ability, and strong robustness, therefore, the improved random forest algorithm in this patent technology is applied to the construction of the enterprise demand scoring model.

[0173] The improved enterprise demand scoring model based on random forest mainly consists of six parts: enterprise demand data preprocessing, enterprise demand matrix, data weighted sampling, feature selection method to select the optimal subset of enterprise demand features, algorithm parameter optimization, and generation of evaluation results. Figure 11 As shown.

[0174] Step 2, establishing an improved enterprise demand scoring model based on random forest, includes the following steps:

[0175] Step 2.1: Preprocessing enterprise demand data;

[0176] The enterprise demand data is complex, including matching analysis of various indicators related to local industries, policies, location, and construction needs. Preliminary matching and screening were performed in the previous enterprise demand matching module. However, due to frequent occurrences of missing, abnormal, and redundant data during testing, preprocessing is necessary before establishing the scoring model to reduce the impact of noisy data on the scoring logic and to meet computational requirements and result validity. The preprocessing methods used in this technical solution include... Figure 11 As shown, it mainly includes:

[0177] (1) Min-Max standardization

[0178] The main approach involves performing a linear transformation on the offline data, resulting in values ​​between [0 and 1]. The formula is as follows:

[0179]

[0180] Where min_value is the minimum value of the data sample, max_value is the maximum value of the data sample, and new_value(x) is the new data value of the sample data after Min-Max standardization.

[0181] (2) Z-Score Standardization

[0182] The main steps involve calculating the mean and standard deviation of the original data, followed by Z-score processing. The goal is to transform the data into a Gaussian distribution with a mean of 0 and a standard deviation of 1. The formula is:

[0183]

[0184] The mean of all data samples is u(data), the standard deviation of all data samples is o(data), and Z(new_data) is the data value after Z-Score processing of the data samples.

[0185] Because the enterprise demand dataset contains non-numerical features, these features are processed using One-Hot encoding, according to the following formula:

[0186]

[0187] Where A and B represent two feature attributes, ra and b represent the correlation between the two feature attributes, n is the number of tuples, ai and bi are the values ​​on A and B respectively, Amean and Bmean are the means on A and B respectively, aibi are the cross products of A and B respectively, and σ A σ B Let ra and b be the standard deviations of A and B. The larger the values ​​of ra and b, the higher the correlation between the two features. At the same time, the higher the values ​​of ra and b, the more likely feature attributes A or B can be deleted as redundant feature attributes.

[0188] Step 2.2: Calculate the enterprise demand matrix;

[0189] In step 2.2, the specific steps for calculating the enterprise demand matrix are as follows: Let X = {b1, b2, ..., bL} represent the set of L samples with M features, and Y = {y1, y2, ..., yL} represent the category set. Then, the enterprise demand data can be used to construct a matrix as follows:

[0190]

[0191] The size of matrix L is L(M+1), where +1 represents the set of categories, bi = {Xi1, Xi2, ..., XiM} represents the M feature values ​​of sample bi, and Xij represents the j-th feature value of sample bi.

[0192] In the enterprise demand matrix L, there are minority samples L″ and majority samples L′. Taking Q samples from the minority samples L″, the matrix form is as follows:

[0193]

[0194] If there are Q samples in the minority sample L″ and L samples in total, then there are (LQ) samples in the majority sample L′. Therefore, its matrix form is:

[0195]

[0196] Step 2.3: Weighted sampling of data;

[0197] The Random Forest algorithm uses the bootstrap sampling method by default. This method introduces significant errors and disrupts the structure of the original data, leading to extreme imbalance in scoring. This imbalance can cause companies whose needs are not matched to those of the investment opportunities to be mistakenly identified as suitable. To improve the accuracy of the scoring model, the original bootstrap sampling method will be modified to sample based on weights.

[0198] Therefore, in step 2.3, the specific steps of data weighted sampling include:

[0199] Step 2.3.1: Divide the original enterprise demand data into training set L and training set L1;

[0200] Step 2.3.2: Divide the training set L into two subsets, namely the majority class samples L′ and the minority class samples L″;

[0201] Step 2.3.3: During the sampling process, first perform weighted sampling on the majority class samples L′, select samples of similar size from the minority class samples L″ in the majority class samples L′, calculate the proportion of the selected minority L″ samples L′, and calculate the proportion of L″ in all L samples. Then, perform weighted sampling to select the final training samples.

[0202] Step 2.3.4: Repeat step 2.3.3 multiple times until a balanced sample is selected;

[0203] Step 2.3.5: Select balanced samples and divide them into training set and test set.

[0204] After weighted sampling based on the above data, the next step is to extract the majority class L′ of enterprise needs. Assume that each class in L′ contains H(k1), H(k2), H(k3), ..., H(kn), where k1, k2, ..., kn represent the distribution of the number of samples in L′ across classes. Then, the weight percentage of ki in the majority sample L′ is:

[0205]

[0206] Based on the above formula, the weight of the minority sample selected from the multiple samples is calculated. Then, the overall weight of ki in the minority sample is:

[0207]

[0208] Based on the above assumption that the minority sample size is Q, the weighted sampling weight of kj in the majority sample is:

[0209]

[0210] Step 2.4: Use feature selection to select the optimal subset of enterprise demand features;

[0211] In step 2.4, the input is the original enterprise demand dataset D = {(x1,y1),(x2,y2),……,(xn,yn)},x i ∈R m , and y n ∈{-1, 1}; Let g1, g2;

[0212] The output is the optimal feature subset f;

[0213] The specific steps for selecting the optimal subset of enterprise requirements features using the feature selection method are as follows: Figure 8 As shown, it includes:

[0214] Step 2.4.1: Define M enterprise demand characteristics i = 1, 2, 3, 4, ..., M;

[0215] Step 2.4.2: Calculate the corresponding value for each enterprise's demand characteristic using the following formula;

[0216] Let D be a sample dataset, x and y be arbitrary attributes of the samples, and n be the number of categories in dataset D. Then the information entropy of x is:

[0217]

[0218] Where P(xi) is the probability that the characteristic attribute x takes the value xi;

[0219] Given feature attribute y, the conditional entropy of feature attribute x is:

[0220]

[0221] Where p(yi) is the probability that the characteristic attribute y takes the value yj, and p(xi|yi) is the probability that the characteristic attribute x takes the value xi when the characteristic attribute y takes the value yj.

[0222] The information entropy obtained from the above formula is:

[0223] Gain(x,y) = Info(x) - Info(x|y)

[0224] Select the feature with the largest information gain as the splitting attribute of dataset D, create a node, use this feature as a label, create branches for each value of the feature, and divide the enterprise needs of the sample accordingly.

[0225] Step 2.4.3: Calculate the entropy comparison value un of each feature and the category variable yn using the following formula;

[0226] Step 2.4.4: If un is greater than or equal to g1, then feature xn is in the selected optimal feature subset f, i.e., xn ∈f;

[0227] The features are sorted, the selected features within set f are measured, and the correlation value S between features xi and xj is determined.

[0228] Step 2.4.5: When S is less than or equal to g2, follow up with the information entropy comparison value un in step (3) to delete the characteristics in set f;

[0229] Step 2.4.6: Obtain the optimal feature subset.

[0230] Step 2.5: Algorithm parameter optimization;

[0231] Traditional random forest algorithms have the following drawbacks:

[0232] (1) Complex algorithm parameters: Random forests have many parameters set before training, mainly including n_estimators (number of decision trees), maximum tree depth, and maximum number of features max_feature. Because these parameters are set in advance, unreasonable parameters can seriously affect the final evaluation accuracy of enterprise needs. If the parameters are set too small, underfitting is likely to occur; if the parameters are set too large, overfitting is likely to occur.

[0233] (2) Too many decision trees will result in excessively long training times; too few decision trees will result in short training times, affecting prediction accuracy. Balancing these two factors is a challenge faced in training with large amounts of data.

[0234] Improvements to the parameter selection of the random forest algorithm:

[0235] A grid search strategy is introduced to optimize the parameters n_estimators (number of decision trees) and max_feature in the random forest. Assuming the two parameters n_estimators and max_feature are S and C respectively, the random forest classifier is trained using S*C. The specific optimization steps for the algorithm parameters are described in step 2.5. Figure 9 As shown, it includes:

[0236] Step 2.5.1: Set the search range and step size for the parameters to be optimized;

[0237] Step 2.5.2: Based on step 2.5.1, further calculate the mean absolute error values ​​of the two parameters S and C, and use the mean absolute error values ​​to obtain the specific range of the number of the two parameters S and C;

[0238] Step 2.5.3: Based on the value range of parameters S and C obtained in Step 2.5.2, calculate the OOB value of the random forest using the S*C combination and the following process to obtain the accuracy;

[0239] During each training run, the unsampled data is labeled as set OOBi. The number of misclassified data in the unsampled dataset OOBi is labeled as ErrorNumOOB. Finally, the error of the random forest OOB value is defined as:

[0240]

[0241] That is, the generalization error is:

[0242]

[0243] Step 2.5.4: Select the optimal parameters determined by the S*C combination based on the OOB value. If the OOB value of the random forest meets the requirements, output the S*C combination; otherwise, change the search range and step size and continue the search until the final condition is met.

[0244] The above method introduces a grid search strategy to find the optimal parameters for the random forest, which can reduce running time, reduce algorithm complexity, improve classification accuracy, and optimize the process. Figure 6 As shown.

[0245] Step 2.6: Generate evaluation results;

[0246] In step 2.6, the enterprise demand scoring model based on random forest generates the optimal evaluation result through steps 2.1 to 2.5, which serves as an evaluation reference and is provided to the industrial park's investment promotion personnel to determine whether the enterprise is suitable for the park's investment promotion strategy.

[0247] Example 2:

[0248] A method for an enterprise intelligent matching investment promotion strategy system, such as Figure 10 As shown, it includes the following steps:

[0249] S1: The enterprise demand assessment module uses an LM-BP neural network to perform deep learning on regional industry value characteristics analysis, policy characteristics analysis, location characteristics analysis, and building characteristics analysis to obtain structured data on industry, policy, location, and building that match enterprise demand.

[0250] The principle of the LM-BP neural network, an improvement on the BP neural network algorithm, is introduced below:

[0251] The LM algorithm combines the advantages of both the gradient method and Newton's method. To mitigate the singularity problem at non-optimal points, it approximates the quadraticity of the objective function near the extreme point by utilizing the properties of the second derivative, thus accelerating the convergence process. It is significantly faster than the gradient method and the backpropagation (BP) algorithm, optimizing the efficiency of big data analysis and processing.

[0252] Detailed introduction:

[0253] Define the error function as

[0254]

[0255] Where w is the vector composed of the neural network threshold and weights; ei = (w) is the error. According to the Gauss-Newton method, the calculation is as follows:

[0256] W(k+1)=W(k)-[J T (W k )J(W k )] -1 J(W k )e(W k )

[0257] Where W(k) represents the vector composed of the threshold and weights in the k-th iteration of the neural network, and W(k+1) represents the vector composed of the threshold and weights in the (k+1)-th iteration.

[0258] The LM algorithm is an improved Gauss-Newton method, as follows:

[0259] W(k+1)=W(k)-[J T (w k )J(w k )+μ k I] -1 J(w k )e(w k )

[0260] Where I is the identity matrix, J is the Jacobian matrix, and uk is a scaling factor. A key step in the LM algorithm is the calculation of the Jacobian matrix, which is performed using a variation of the BP algorithm, as follows:

[0261]

[0262] If the scaling factor u = 0, it is equivalent to the Gauss-Newton algorithm; if the scaling factor is very large, the LM algorithm is close to the gradient algorithm. With each iteration, u decreases slightly. Therefore, when approaching the error target, it is closer to the Gauss-Newton algorithm, offering faster computation and higher accuracy. u is an exploratory parameter; for a given parameter u, if the calculated threshold change Δw reduces the error function e(w), then u decreases. Conversely, u increases. The LM algorithm utilizes information from the second derivative, making it much faster than the gradient algorithm. Practical test data demonstrates that the LM-BP algorithm is tens of times faster than the traditional BP gradient descent.

[0263] Since the BP neural network is publicly available and open source, the specific principles and algorithmic logic of the BP neural network itself will not be described in detail in this patented technical solution. Figure 3As shown, it mainly introduces the optimized and improved parts, such as Figure 4 As shown, the specific calculation steps of the LM-BP neural network include the following steps:

[0264] Step 1.1: Initialize the network structure parameters, the error tolerance value is ε, the constants u and b, initialize the weight and threshold vectors, let k = 0, u = u0, the calculation accuracy is ε and the maximum number of learning times M;

[0265] Step 1.2: Input the training data of the enterprise demand portrait index matrix as the input vector into the LM-BP neural network;

[0266] Step 1.3: Calculate the network output and the error index function e;

[0267] Step 1.4: Calculate the Jacobian matrix J[W(k)]; where, W(k) represents the vector composed of the thresholds and weights of the k-th neural network iteration;

[0268] Step 1.5: Calculate ΔW; where, ΔW is the threshold change amount;

[0269] Step 1.6: If e < ε, then go to Step 1.8, otherwise go to Step 1.5;

[0270] Step 1.7: Calculate the error function e with the new weight and threshold vector W(k + 1),

[0271] W(k + 1) = W(k) - {J T [W(k)]J[W(k)]} -1 J[W(k)]e[W(k)]

[0272] If e[W(k + 1)] is less than e[W(k)], then let k = k + 1, u = u * b, go to Step 1.2, otherwise u = u / b, go to Step 1.5; where, W(k) represents the vector composed of the thresholds and weights of the k-th neural network iteration, and W(k + 1) represents the vector composed of the thresholds and weights of the new (k + 1)-th iteration;

[0273] Step 1.8: The calculation of the LM-BP neural network ends;

[0274] The above calculates the industrial, policy, location, and building structured data that accurately matches the enterprise demand through the LM-BP neural network, and scores the indicators and weights matched by the enterprise demand portrait through the investment promotion strategy module.

[0275] In Step 1.1, the value range of b is: 0 < b < 1. When k = 0, u = u0, calculate the calculation accuracy ε and the maximum number of learning times M. <0​In steps 1.4 and 1.7, J is calculated by transforming the Jacobi matrix J[W(k)]. T (W), calculate J T The formula for (W) is:

[0277]

[0278] S2: The investment promotion strategy module scores the enterprise demand profile and the matching local industries, policies, location, and building resources, constructs an enterprise demand investment promotion strategy report scoring model, establishes an enterprise demand scoring model based on random forest, and classifies and predicts the characteristics of enterprise demand.

[0279] Introduction to the functions of the investment promotion strategy module:

[0280] The main function is to score enterprises based on their needs profiles and matching local industries, policies, locations, and building resources. It uses machine learning and statistics to build a scoring model for enterprise demand investment promotion strategy reports, follows up on enterprise demand characteristics, mines the relationships between different characteristics, establishes an enterprise demand scoring model based on random forest, classifies and predicts the status of enterprise demand characteristics, and identifies enterprises with a high degree of matching of weighted indicators as accurately matched investment promotion enterprises.

[0281] Algorithm principle and logic introduction:

[0282] like Figure 5 As shown, the basic building block of the random forest algorithm is the decision tree. The main idea of ​​this algorithm is to randomly draw n samples with replacement from the original sample set S to generate a new sample set. Then, n classification trees are generated based on the bootstrap sample set to form a random forest. The final classification result is decided by the number of votes cast by the classification trees.

[0283] Because the random forest algorithm is relatively simple to implement, has a fast training speed, strong generalization ability, and strong robustness, this patented technology applies the improved random forest algorithm to the construction of enterprise demand scoring models.

[0284] The improved enterprise demand scoring model based on random forest consists of six parts: enterprise demand data preprocessing, enterprise demand matrix, data weighted sampling, feature selection method to select the optimal subset of enterprise demand features, algorithm parameter optimization, and generation of evaluation results.

[0285] Step 2, establishing an improved enterprise demand scoring model based on random forest, includes the following steps:

[0286] Step 2.1: Preprocessing enterprise demand data;

[0287] The enterprise demand data is complex, including matching analysis of various indicators related to local industries, policies, location, and construction needs. Preliminary matching and screening were performed in the previous enterprise demand matching module. However, due to frequent occurrences of missing, abnormal, and redundant data during testing, preprocessing is necessary before establishing the scoring model to reduce the impact of noisy data on the scoring logic and to meet computational requirements and result validity. The preprocessing methods used in this patented technology mainly include:

[0288] (1) Min-Max standardization

[0289] The main approach involves performing a linear transformation on the offline data, resulting in values ​​between [0 and 1]. The formula is as follows:

[0290]

[0291] Where min_value is the minimum value of the data sample, max_value is the maximum value of the data sample, and new_value(x) is the new data value of the sample data after Min-Max standardization.

[0292] (2) Z-Score Standardization

[0293] The main steps involve calculating the mean and standard deviation of the original data, followed by Z-score processing. The goal is to transform the data into a Gaussian distribution with a mean of 0 and a standard deviation of 1. The formula is:

[0294]

[0295] The mean of all data samples is u(data), the standard deviation of all data samples is o(data), and Z(new_data) is the data value after Z-Score processing of the data samples.

[0296] Because the enterprise demand dataset contains non-numerical features, these features are processed using One-Hot encoding, according to the following formula:

[0297]

[0298] Where A and B represent two feature attributes, ra and b represent the correlation between the two feature attributes, n is the number of tuples, ai and bi are the values ​​on A and B respectively, Amean and Bmean are the means on A and B respectively, aibi are the cross products of A and B respectively, and σ A σ BLet ra and b be the standard deviations of A and B. The larger the values ​​of ra and b, the higher the correlation between the two features. At the same time, the higher the values ​​of ra and b, the more likely feature attributes A or B can be deleted as redundant feature attributes.

[0299] Step 2.2: Calculate the enterprise demand matrix;

[0300] In step 2.2, the specific steps for calculating the enterprise demand matrix are as follows: Let X = {b1, b2, ..., bL} represent the set of L samples with M features, and Y = {y1, y2, ..., yL} represent the category set. Then, the enterprise demand data can be used to construct a matrix as follows:

[0301]

[0302] The size of matrix L is L(M+1), where +1 represents the set of categories, bi = {Xi1, Xi2, ..., XiM} represents the M feature values ​​of sample bi, and Xij represents the j-th feature value of sample bi.

[0303] In the enterprise demand matrix L, there are minority samples L″ and majority samples L′. Taking Q samples from the minority samples L″, the matrix form is as follows:

[0304]

[0305] If there are Q samples in the minority sample L″ and L samples in total, then there are (LQ) samples in the majority sample L′. Therefore, its matrix form is:

[0306]

[0307] Step 2.3: Weighted sampling of data;

[0308] The Random Forest algorithm uses the bootstrap sampling method by default. This method introduces significant errors and disrupts the structure of the original data, leading to extreme imbalance in scoring. This imbalance can cause companies whose needs are not matched to those of the investment opportunities to be mistakenly identified as suitable. To improve the accuracy of the scoring model, the original bootstrap sampling method will be modified to sample based on weights.

[0309] Therefore, in step 2.3, as Figure 7 As shown, the specific steps of data weighted sampling include:

[0310] Step 2.3.1: Divide the original enterprise demand data into training set L and training set L1;

[0311] Step 2.3.2: Divide the training set L into two subsets, namely the majority class samples L′ and the minority class samples L″;

[0312] Step 2.3.3: During the sampling process, first perform weighted sampling on the majority class samples L′, select samples of similar size from the minority class samples L″ in the majority class samples L′, calculate the proportion of the selected minority L″ samples L′, and calculate the proportion of L″ in all L samples. Then, perform weighted sampling to select the final training samples.

[0313] Step 2.3.4: Repeat step 2.3.3 multiple times until a balanced sample is selected;

[0314] Step 2.3.5: Select balanced samples and divide them into training set and test set.

[0315] After weighted sampling based on the above data, the next step is to extract the majority class L′ of enterprise needs. Assume that each class in L′ contains H(k1), H(k2), H(k3), ..., H(kn), where k1, k2, ..., kn represent the distribution of the number of samples in L′ across classes. Then, the weight percentage of ki in the majority sample L′ is:

[0316]

[0317] Based on the above formula, the weight of the minority sample selected from the multiple samples is calculated. Then, the overall weight of ki in the minority sample is:

[0318]

[0319] Based on the above assumption that the minority sample size is Q, the weighted sampling weight of kj in the majority sample is:

[0320]

[0321] Step 2.4: Use feature selection to select the optimal subset of enterprise demand features;

[0322] In step 2.4, the input is the original enterprise demand dataset D = {(x1,y1),(x2,y2),……,(xn,yn)},x i ∈R m , and y n ∈{-1, 1}; Let g1, g2;

[0323] The output is the optimal feature subset f;

[0324] The specific steps for selecting the optimal subset of enterprise requirements features using the feature selection method include:

[0325] Step 2.4.1: Define M enterprise demand characteristics i = 1, 2, 3, 4, ..., M;

[0326] Step 2.4.2: Calculate the corresponding value for each enterprise's demand characteristic using the following formula;

[0327] Let D be a sample dataset, x and y be arbitrary attributes of the samples, and n be the number of categories in dataset D. Then the information entropy of x is:

[0328]

[0329] Where P(xi) is the probability that the characteristic attribute x takes the value xi;

[0330] Given feature attribute y, the conditional entropy of feature attribute x is:

[0331]

[0332] Where p(yi) is the probability that the characteristic attribute y takes the value yj, and p(xi|yi) is the probability that the characteristic attribute x takes the value xi when the characteristic attribute y takes the value yj.

[0333] The information entropy obtained from the above formula is:

[0334] Gain(x,y) = Info(x) - Info(x|y)

[0335] Select the feature with the largest information gain as the splitting attribute of dataset D, create a node, use this feature as a label, create branches for each value of the feature, and divide the enterprise needs of the sample accordingly.

[0336] Step 2.4.3: Calculate the entropy comparison value un of each feature and the category variable yn using the following formula;

[0337] Step 2.4.4: If un is greater than or equal to g1, then feature xn is in the selected optimal feature subset f, i.e., x n ∈f;

[0338] The features are sorted, the selected features within set f are measured, and the correlation value S between features xi and xj is determined.

[0339] Step 2.4.5: When S is less than or equal to g2, follow up with the information entropy comparison value un in step (3) to delete the characteristics in set f;

[0340] Step 2.4.6: Obtain the optimal feature subset.

[0341] Step 2.5: Algorithm parameter optimization;

[0342] Traditional random forest algorithms have the following drawbacks:

[0343] (1) Complex algorithm parameters: Random forests have many parameters set before training, mainly including n_estimators (number of decision trees), maximum tree depth, and maximum number of features max_feature. Because these parameters are set in advance, unreasonable parameters can seriously affect the final evaluation accuracy of enterprise needs. If the parameters are set too small, underfitting is likely to occur; if the parameters are set too large, overfitting is likely to occur.

[0344] (2) Too many decision trees will result in excessively long training times; too few decision trees will result in short training times, affecting prediction accuracy. Balancing these two factors is a challenge faced in training with large amounts of data.

[0345] Improvements to the parameter selection of the random forest algorithm:

[0346] A grid search strategy is introduced to optimize the parameters n_estimators (number of decision trees) and max_feature in the random forest. Assuming the two parameters n_estimators and max_feature are S and C respectively, the random forest classifier is trained using S*C. In step 2.5, as follows... Figure 9 As shown, the specific optimization steps for algorithm parameter optimization include:

[0347] Step 2.5.1: Set the search range and step size for the parameters to be optimized;

[0348] Step 2.5.2: Based on step 2.5.1, further calculate the mean absolute error values ​​of the two parameters S and C, and use the mean absolute error values ​​to obtain the specific range of the number of the two parameters S and C;

[0349] Step 2.5.3: Based on the value range of parameters S and C obtained in Step 2.5.2, calculate the OOB value of the random forest using the S*C combination and the following process to obtain the accuracy;

[0350] During each training run, the unsampled data is labeled as set OOBi. The number of misclassified data in the unsampled dataset OOBi is labeled as ErrorNumOOB. Finally, the error of the random forest OOB value is defined as:

[0351]

[0352] That is, the generalization error is:

[0353]

[0354] Step 2.5.4: Select the optimal parameters determined by the S*C combination based on the OOB value. If the OOB value of the random forest meets the requirements, output the S*C combination; otherwise, change the search range and step size and continue the search until the final condition is met.

[0355] The above method introduces a grid search strategy to find the optimal parameters for the random forest, which can reduce running time, reduce algorithm complexity, improve classification accuracy, and optimize the process. Figure 6 As shown.

[0356] Step 2.6: Generate evaluation results;

[0357] In step 2.6, the enterprise demand scoring model based on random forest generates the optimal evaluation result through steps 2.1 to 2.5, which serves as an evaluation reference and is provided to the industrial park's investment promotion personnel to determine whether the enterprise is suitable for the park's investment promotion strategy.

[0358] S3: The industry matching module analyzes the scale of local investment promotion areas, the total number of related enterprises in the industrial chain, the number of large-scale enterprises, customers and suppliers of enterprises in their respective sub-industries through the industry matching database. It also provides reference values ​​for the enterprise's industrial agglomeration effect and supply and sales relationship effect, thereby matching and selecting the optimal industry for the enterprise.

[0359] S4: The policy matching module analyzes policies suitable for enterprises through the policy matching library. It recommends general policies suitable for enterprises in the startup, growth and maturity stages, industry-specific policies suitable for all sub-sectors of the enterprise, and policies that cultivate the enterprise's prospects. This allows investment promotion personnel to follow up on the specific needs of enterprises and make recommendations in the policy matching module, thereby matching and selecting the best policy suitable for the enterprise.

[0360] S5: The regional space matching module matches the formatted data of the regional space database with the quantitative evaluation and scoring of enterprise needs. It also performs an adaptability evaluation based on the enterprise needs, including macro-location, meso-location, micro-location, land use planning, transportation and logistics, and living facilities, thereby matching and selecting the optimal regional space suitable for the enterprise.

[0361] S6: The building matching module selects the optimal building carrier for the enterprise by matching the planning building library. It evaluates the building carrier based on basic building information, building structure, energy conservation and environmental protection, fire prevention and explosion protection, supporting equipment, and occupancy costs, thereby matching and selecting the optimal building plan for the enterprise.

[0362] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure made using the content of the present invention specification and drawings, or directly or indirectly applied to other related technical fields, is similarly included within the patent protection scope of the present invention.

Claims

1. A smart matching investment promotion strategy system based on enterprise needs, characterized in that, include: The modules include: Enterprise Demand Assessment (1), Investment Promotion Strategy (2), Industrial Support Matching (3), Policy Matching (4), Regional Spatial Matching (5), and Building Matching (6); The enterprise demand assessment module (1) is used to obtain structured data on industry, policy, location and building characteristics that match enterprise demand after deep learning of regional industry value characteristics analysis, policy characteristics analysis, location characteristics analysis and building characteristics analysis through LM-BP neural network. The investment promotion strategy module (2) is used to score the enterprise demand profile and the matching local industries, policies, location and building resources, construct the enterprise demand investment promotion strategy report scoring model, establish an enterprise demand scoring model based on random forest, and classify and predict the enterprise demand characteristics. The industry matching module (3) is used to analyze the scale of local investment promotion areas, the total number of related enterprises in the industrial chain, the number of enterprises above the designated size, customers and suppliers of enterprises in their respective sub-industries through the industry matching library (31), and provide reference values ​​for the enterprise's industrial agglomeration effect and supply and sales relationship effect, so as to match and select the optimal industry suitable for the enterprise. The policy matching module (4) is used to analyze the policies suitable for enterprises through the policy matching library (41). Through the general policies suitable for the enterprise's development stage (start-up, growth, and maturity), the industry-specific policies suitable for all sub-sectors of the enterprise, and the policy recommendations for cultivating the enterprise's prospects, the investment promotion personnel can follow up on the specific needs of the enterprise and make recommendations in the policy matching module (4), thereby matching and selecting the optimal policy suitable for the enterprise. The regional space matching module (5) is used to match the quantitative evaluation and scoring of enterprise needs after the data formatting of the matching regional space library (51), and to conduct an adaptive evaluation by matching the enterprise needs with macro location, meso location, micro location, land use planning, transportation and logistics, and living facilities, so as to match and select the optimal regional space suitable for the enterprise. The building matching module (6) is used to select the optimal building carrier suitable for the enterprise by matching the planning building library (61). It evaluates the building carrier based on basic building information, building structure, energy conservation and environmental protection, fire prevention and explosion protection, supporting equipment, and occupancy cost, thereby matching and selecting the optimal building plan suitable for the enterprise.

2. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 1, characterized in that: In the enterprise demand assessment module (1), the industry structured data for matching enterprise demand includes: industry scale, value chain customer market, industrial chain supporting facilities and enterprise supporting facilities; the policy structured data for matching enterprise demand includes: industrial policy, talent policy and financial policy; the location structured data for matching enterprise demand includes: transportation and logistics, supporting resources and planning elements; the building structured data for matching enterprise demand includes: building basic elements, energy conservation and environmental protection, load, fire prevention and explosion protection and usage cost.

3. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 1, characterized in that: The specific calculation steps of the LM-BP neural network described in step 1 are as follows: Step 1.1: Initialize the network structure parameters, with the error tolerance value being ε, constants u and b, and initialize the weight and threshold vectors. Let k = 0, u = u0, and calculate the precision ε and the maximum number of learning iterations M. Step 1.2: Input the training data of the enterprise demand portrait index matrix as an input vector into the LM-BP neural network; Step 1.3: Calculate the network output and the error index function e; Step 1.4: Calculate the Jacobian matrix J[W(k)]; where W(k) represents the vector composed of the thresholds and weights of the k-th neural network iteration; Step 1.5: Calculate ΔW; where ΔW is the threshold change; Step 1.6: If e < ε, go to Step 1.8, otherwise go to Step 1.5; Step 1.7: Calculate the error function e with the new weight and threshold vector W(k + 1), W(k+1)=W(k)-{J T [W(k)]J[W(k)]} -1 J[W(k)]e[W(k)] If e[W(k + 1)] is less than e[W(k)], then let k = k + 1, u = u * b, go to Step 1.2, otherwise u = u / b, go to Step 1.5; where W(k) represents the vector composed of the thresholds and weights of the k-th neural network iteration, and W(k + 1) represents the vector composed of the thresholds and weights of the new (k + 1)-th iteration; Step 1.8: The calculation of the LM-BP neural network ends.

4. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 3, characterized in that: In Step 1.1, the value range of b is: 0 < b < 1. When k = 0 and u = u0, calculate the calculation accuracy ε and the maximum number of learning times M.

5. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 3, characterized in that: In steps 1.4 and 1.7, J is calculated by transforming the Jacobi matrix J[W(k)]. T (W), calculate J T The formula for (W) is:

6. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 1, characterized in that: The steps of establishing the enterprise demand scoring model improved based on the random forest described in Step 2 include: Step 2.1: Preprocess the enterprise demand data; Step 2.2: Calculate the enterprise demand matrix; Step 2.3: Data weighted sampling; Step 2.4: Select the optimal demand feature subset of the enterprise demand by the feature selection method; Step 2.5: Optimize the algorithm parameters; Step 2.6: Generate the evaluation results.

7. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 6, characterized in that: In Step 2.1, the preprocessing of the enterprise demand data includes Min-Max normalization processing and Z-Score normalization processing; The Min-Max normalization processing is to perform a linear transformation on the offline data in the enterprise demand data, so that the data after the linear transformation of the enterprise demand data is between [0 - 1]; The Z-Score normalization processing is to convert the enterprise demand data into a Gaussian distribution with a mean of 0 and a standard deviation of 1.

8. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 6, characterized in that: In Step 2.2, the specific steps of calculating the enterprise demand matrix are as follows: Let X = {b1, b2, ……, bL} represent the set composed of L samples of M features, and Y = {y1, y2, ……, yL} represent the category set. Then the enterprise demand data can be constructed into a matrix as: Where the size of matrix L is L(M + 1), +1 represents the set of categories, bi = {Xi1, Xi2, ……, XiM} represents the M feature values of sample bi, and Xij represents the j-th feature value of sample bi; In the enterprise demand matrix L, it includes a small number of samples L″ and a large number of samples L′. Take Q samples from the small number of samples L″, and the matrix form is: There are Q samples in the small number of samples L″, and the total number of samples is L. Then there are (L - Q) sample numbers in the large number of samples L′, and its matrix form is:

9. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 8, characterized in that: In Step 2.3, the specific steps of the data weighted sampling include: Step 2.3.1: Divide the original enterprise demand data into training set L and training set L1; Step 2.3.2: Divide the training set L into two subsets, namely the majority class samples L′ and the minority class samples L″; Step 2.3.3: During the sampling process, first perform weighted sampling on the majority class samples L′, select samples of similar size from the minority class samples L″ in the majority class samples L′, calculate the proportion of the selected minority L″ samples L′, and calculate the proportion of L″ in all L samples. Then, perform weighted sampling to select the final training samples. Step 2.3.4: Repeat step 2.3.3 multiple times until a balanced sample is selected; Step 2.3.5: Select balanced samples and divide them into training set and test set.

10. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 6, characterized in that: In step 2.4, the input is the original enterprise demand dataset D = {(x1,y1),(x2,y2),……,(xn,yn)},x i ∈R m , and y n ∈{-1, 1}; Let g1, g2; The output is the optimal feature subset f; The specific steps for selecting the optimal subset of enterprise demand features using the feature selection method include: Step 2.4.1: Define M enterprise demand characteristics i = 1, 2, 3, 4, ..., M; Step 2.4.2: Calculate the corresponding value for each enterprise's demand characteristic using the following formula; Let D be a sample dataset, x and y be arbitrary attributes of the samples, and n be the number of categories in dataset D. Then the information entropy of x is: Where P(xi) is the probability that the characteristic attribute x takes the value xi; Given feature attribute y, the conditional entropy of feature attribute x is: Where p(yi) is the probability that the characteristic attribute y takes the value yj, and p(xi|yi) is the probability that the characteristic attribute x takes the value xi when the characteristic attribute y takes the value yj. The information entropy obtained from the above formula is: Gain(x,y)=Info(x)-Info(x|y) Select the feature with the largest information gain as the splitting attribute of dataset D, create a node, use this feature as a label, create branches for each value of the feature, and divide the enterprise needs of the sample accordingly. Step 2.4.3: Calculate the entropy comparison value un of each feature and the category variable yn using the following formula; Step 2.4.4: If un is greater than or equal to g1, then feature xn is in the selected optimal feature subset f, i.e., x n ∈f; The features are sorted, the selected features within set f are measured, and the correlation value S between features xi and xj is determined. Step 2.4.5: When S is less than or equal to g2, follow up with the information entropy comparison value un in step (3) to delete the characteristics in set f; Step 2.4.6: Obtain the optimal feature subset.

11. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 6, characterized in that: In step 2.5, the specific optimization steps for optimizing the algorithm parameters include: Step 2.5.1: Set the search range and step size for the parameters to be optimized; Step 2.5.2: Based on step 2.5.1, further calculate the mean absolute error values ​​of the two parameters S and C, and use the mean absolute error values ​​to obtain the specific range of the number of the two parameters S and C; Step 2.5.3: Based on the value range of parameters S and C obtained in Step 2.5.2, calculate the OOB value of the random forest using the S*C combination and the following process to obtain the accuracy; During each training run, the unsampled data is labeled as set OOBi. The number of misclassified data in the unsampled dataset OOBi is labeled as ErrorNumOOB. Finally, the error of the random forest OOB value is defined as: That is, the generalization error is: Step 2.5.4: Select the optimal parameters determined by the S*C combination based on the OOB value. If the OOB value of the random forest meets the requirements, output the S*C combination; otherwise, change the search range and step size and continue the search until the final condition is met.

12. The intelligent matching investment promotion strategy system based on enterprise needs according to claim 6, characterized in that: In step 2.6, the enterprise demand scoring model based on random forest generates the optimal evaluation result through steps 2.1 to 2.5, which serves as an evaluation reference and is provided to the industrial park's investment promotion personnel to determine whether the enterprise is suitable for the park's investment promotion strategy.

13. A method for an enterprise intelligent matching investment promotion strategy system, characterized in that: Includes the following steps: S1: The enterprise demand assessment module uses an LM-BP neural network to perform deep learning on regional industry value characteristics analysis, policy characteristics analysis, location characteristics analysis, and building characteristics analysis to obtain structured data on industry, policy, location, and building that match enterprise demand. S2: The investment promotion strategy module scores the enterprise demand profile and the matching local industries, policies, location, and building resources, constructs an enterprise demand investment promotion strategy report scoring model, establishes an enterprise demand scoring model based on random forest, and classifies and predicts the characteristics of enterprise demand. S3: The industry support matching module analyzes the scale of local investment promotion areas, the total number of related enterprises in the industrial chain, the number of large-scale enterprises, customers and suppliers of enterprises in their respective sub-industries through the industry support database, and provides reference values ​​for the enterprise's industrial agglomeration effect and supply and sales relationship effect, thereby matching and selecting the optimal industry for the enterprise. S4: The policy matching module analyzes policies suitable for enterprises through the policy matching library. It recommends general policies suitable for enterprises in the startup, growth and maturity stages, industry-specific policies suitable for all sub-sectors of the enterprise, and policies that cultivate the enterprise's prospects. This allows investment promotion personnel to follow up on the specific needs of enterprises and make recommendations in the policy matching module, thereby matching and selecting the best policy suitable for the enterprise. S5: The regional space matching module matches the formatted data of the regional space database with the quantitative evaluation and scoring of enterprise needs. It also performs an adaptability evaluation based on the enterprise needs, including macro-location, meso-location, micro-location, land use planning, transportation and logistics, and living facilities, thereby matching and selecting the optimal regional space suitable for the enterprise. S6: The building matching module selects the optimal building carrier for the enterprise by matching the planning building library. It evaluates the building carrier based on basic building information, building structure, energy conservation and environmental protection, fire prevention and explosion protection, supporting equipment, and occupancy costs, thereby matching and selecting the optimal building plan for the enterprise.

Citation Information

Patent Citations

  • An enterprise location selection system based on environmental evaluation index of investment cooperation carrier

    CN109213838A

  • Overseas park investment attracting service system based on digital earth framework

    CN108470032A

  • Intelligent recommendation system and recommendation method based on industrial park investment attraction

    CN113537728A