An industry chain investment targeting matching method and system based on big data analysis
By using big data analysis to construct industry chain maps and comprehensive scores, the shortcomings of human experience-based judgment in the investment promotion process have been addressed, enabling targeted matching between enterprises and regional industries and improving the efficiency and quality of investment promotion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGLIAN HENGCHUANG (SHANXI) TECHNOLOGY CO LTD
- Filing Date
- 2025-07-02
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, the investment promotion process relies on human experience and judgment, which is inefficient and greatly affected by subjective factors. It is difficult to ensure the matching of enterprises with the regional industrial ecosystem, resulting in waste of resources and poor investment promotion results.
A targeted matching method for investment attraction based on big data analysis is adopted. By extracting the industrial characteristics of enterprise data, an industrial chain map is constructed. Combined with recommendation and evaluation models, a comprehensive business score is obtained to achieve targeted matching between enterprises and regional industries.
It improved the efficiency and quality of investment promotion, ensured the matching of enterprises with regional industries, reduced resource waste, enhanced the synergistic effect of the industrial chain, reduced investment promotion risks, and improved the reliability and stability of matching.
Smart Images

Figure CN120851920B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for targeted matching of investment promotion in the industrial chain based on big data analysis. Background Technology
[0002] With the rapid development of the industrial economy, local governments and industrial parks are increasingly emphasizing investment promotion, which has become an important means to drive regional economic growth and optimize industrial structure. However, current investment promotion efforts rely heavily on the human experience of staff to determine whether companies align with local industrial positioning. This lacks an overall framework and data-driven matching and screening. This human-targeted matching approach is not only inefficient but also highly susceptible to subjective factors, making it difficult to ensure that the screening results match the local industrial development direction. This leads to a waste of investment resources. Furthermore, human experience cannot automatically identify and recommend companies based on the specific needs of regional development, resulting in a mismatch between the introduced companies and the regional industrial ecosystem, thus failing to fully leverage their driving effect on the region.
[0003] Therefore, it is necessary to design a targeted matching method and system for investment promotion based on big data analysis to solve the problems existing in the current technology. Summary of the Invention
[0004] In view of this, the present invention proposes a targeted matching method and system for industrial chain investment promotion based on big data analysis, which aims to solve the problems that relying on human-based targeted matching is not only inefficient, but also greatly affected by subjective factors, making it difficult to ensure that the screening results are consistent with the local industrial development direction, and that the introduced enterprises are mismatched with the regional industrial ecosystem.
[0005] In one aspect, this invention proposes a targeted matching method for investment promotion in the industrial chain based on big data analysis, including:
[0006] Data from several companies is collected based on the time period, and the data from each company is preprocessed to determine the target company data. The industry characteristics of each target company data are then extracted.
[0007] Starting with the aforementioned industry characteristics, and extending the industry chain map based on all target enterprise data, the similarity between the industry characteristics to be matched and the industry chain map is determined, and matching enterprises are determined based on the similarity and selection orientation.
[0008] The historical operating information and historical business behavior of the matched enterprise are obtained. Based on the recommendation model, the historical operating information and historical business behavior, the model operating score of the matched enterprise is determined. Based on the historical operating information and historical business behavior, the evaluation operating score of the matched enterprise is determined. Based on the relationship between the model operating score and the evaluation operating score, the comprehensive operating score of the matched enterprise is determined.
[0009] The matching company is determined based on its comprehensive business performance score. If the matching company is not determined as a target matching result, the matching company is re-determined.
[0010] Furthermore, when acquiring data from several enterprises based on the collection time period and preprocessing each enterprise's data to determine the target enterprise data, this includes:
[0011] The preprocessing includes data standardization and outlier detection;
[0012] The data standardization includes unified field naming and industry classification coding;
[0013] The outlier detection involves removing data with zero registered capital and abnormal establishment time.
[0014] Abnormal data in the establishment time are defined as data that is greater than or equal to the right boundary of the collection time period.
[0015] Furthermore, when extracting industry characteristics from the data of each target enterprise, the following are included:
[0016] Based on NLP analysis of each target company's data, word segmentation, stop word removal, and stemming are performed on each target company's data. Dependency parsing is used to analyze the relationships between words in each target company's data and extract noun phrases. The noun phrases that conform to the syntactic structure are taken as the industry features.
[0017] Furthermore, when using the aforementioned industry characteristics as a starting point and extending the construction of the industry chain map based on all target enterprise data, this includes:
[0018] Starting with the industry characteristics, each data feature of each target enterprise's data is taken as a node. Whenever a data feature has a sub-data feature, the corresponding node is expanded by a sub-node. All industry characteristics, as well as the corresponding nodes and sub-nodes, are used to construct the industry chain map.
[0019] The data features include basic enterprise information, geographic information, industry attributes, operational capabilities, and cooperative relationships.
[0020] Furthermore, when determining the similarity between the characteristics of the industry to be matched and the industry chain map, and determining the matching enterprise based on the similarity and selection orientation, the process includes:
[0021] Determine the similarity between the industry feature to be matched and each industry feature;
[0022] Similar companies are identified based on preset selection criteria, which include the expected company size, whether the company possesses specific qualifications, and whether there is a preference for companies with existing cooperative relationships. The company with the highest similarity is then identified as the matching company.
[0023] Furthermore, when obtaining the historical operating information and historical business behavior of the matched enterprise, and determining the model operating score of the matched enterprise based on the recommendation model, the historical operating information, and the historical business behavior, the process includes:
[0024] Obtain the model dataset, and sample the model dataset according to the proportion to determine the training set and test set. Obtain the pre-selected random forest model, and iteratively train the random forest model according to the training set. Test the iteratively trained random forest model according to the test set.
[0025] If the mean squared error of the random forest model after the current iteration is greater than or equal to the mean squared error of the random forest model after the previous iteration, then adjust the learning rate of the random forest model after the current iteration and continue iterative training until the preset number of iterations is reached.
[0026] If the mean squared error of the random forest model after the current iteration is less than the mean squared error of the random forest model after the previous iteration, then the iteration training stops, and the random forest model after the current iteration is determined as the recommended model. Based on the historical business information and historical enterprise behavior, the model business score of the matching enterprise is output.
[0027] Furthermore, when determining the performance evaluation score of the matched enterprise based on the historical operating information and historical corporate behavior, the following steps are included:
[0028] The initial score of the matched enterprise is set to zero. Whenever there is an anomaly in the historical business information or an anomaly in the historical enterprise behavior, a score is added to the zero score, and the sum of the final added scores is determined as the evaluation business score of the matched enterprise.
[0029] Furthermore, when determining the overall business performance score of the matched enterprise based on the relationship between the model business performance score and the evaluation business performance score, the process includes:
[0030] When the model operation score and the evaluation operation score are equal, the model operation score or the evaluation operation score is determined as the overall business operation score of the matched enterprise.
[0031] When the model management score and the evaluation management score are not equal, the average of the model management score and the evaluation management score is determined as the overall management score of the matched enterprise.
[0032] Furthermore, when determining whether to identify the matched enterprise as a targeted matching result based on the enterprise's comprehensive business performance score, the process includes:
[0033] A pre-set threshold for the enterprise's overall business performance score is used to compare the enterprise's overall business performance score with the threshold.
[0034] If the enterprise's overall business score is greater than the enterprise's overall business score threshold, then the matched enterprise will not be identified as a targeted matching result.
[0035] When the enterprise's overall business score is less than or equal to the enterprise's overall business score threshold, the matched enterprise is determined as the targeted matching result.
[0036] Compared with existing technologies, the beneficial effects of this invention are as follows: By extracting the industry characteristics of each target enterprise's data, the stability and reliability of the data source are ensured, avoiding the subjectivity and uncertainty of human judgment based on impressions. Through the construction and similarity of the industrial chain map, targeted matching between enterprises and regional industrial positioning is achieved, reducing resource mismatch caused by human experience bias, thereby avoiding the waste of resources in ineffective investment promotion. It ensures that the matched enterprises and the regional industrial ecosystem form an organic connection, strengthens the synergistic effect of upstream and downstream of the industrial chain, and the comprehensive operation score achieves a comprehensive evaluation of the operating status of the matched enterprises by integrating two scoring dimensions, improving the reliability and stability of targeted matching, effectively avoiding the one-sidedness of a single evaluation dimension, reducing investment promotion risks, and thus improving the stability of targeted matching in attracting enterprises. Through a cyclical verification mechanism, the matched enterprises are continuously optimized to ensure the adaptability and compatibility of targeted matching results with regional development needs, thereby improving the efficiency and quality of investment promotion.
[0037] On the other hand, this application also provides a supply chain investment promotion targeted matching system based on big data analysis, used to apply the above-mentioned supply chain investment promotion targeted matching method based on big data analysis, including:
[0038] The data acquisition unit is configured to acquire data from several enterprises based on the data acquisition time period, preprocess each enterprise's data to determine the target enterprise's data, and extract the industry characteristics of each target enterprise's data.
[0039] The analysis unit is configured to take the industry characteristics as a starting point, extend and construct an industry chain map based on all target enterprise data, determine the similarity between the industry characteristics to be matched and the industry chain map, and determine the matching enterprise based on the similarity and selection orientation.
[0040] The processing unit is configured to acquire the historical operating information and historical business behavior of the matched enterprise, determine the model operating score of the matched enterprise based on the recommendation model, the historical operating information and historical business behavior, determine the evaluation operating score of the matched enterprise based on the historical operating information and historical business behavior, and determine the comprehensive operating score of the matched enterprise based on the relationship between the model operating score and the evaluation operating score.
[0041] The matching unit is configured to determine whether to identify the matching enterprise as a targeted matching result based on the enterprise's comprehensive business score. If it is determined that the matching enterprise should not be identified as a targeted matching result, the matching enterprise is re-identified.
[0042] It is understandable that the above-mentioned method and system for targeted matching of investment promotion based on big data analysis has the same beneficial effects, and will not be elaborated further here. Attached Figure Description
[0043] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0044] Figure 1 A flowchart illustrating a targeted matching method for investment promotion based on big data analysis, provided as an embodiment of the present invention;
[0045] Figure 2 This is a functional block diagram of a targeted matching system for investment promotion based on big data analysis, provided as an embodiment of the present invention. Detailed Implementation
[0046] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0047] In some embodiments of this application, see Figure 1 As shown, a targeted matching method for investment promotion based on big data analysis of the industrial chain includes:
[0048] S100: Acquire data from several enterprises based on the collection time period, preprocess each enterprise's data to determine the target enterprise data, and extract the industry characteristics of each target enterprise's data.
[0049] S200: Starting with industry characteristics, and extending to construct an industry chain map based on all target enterprise data, determine the similarity between the industry characteristics to be matched and the industry chain map, and determine the matching enterprises based on the similarity and selection orientation.
[0050] S300: Obtain the historical operating information and historical corporate behavior of the matching enterprise; determine the model operating score of the matching enterprise based on the recommendation model, historical operating information and historical corporate behavior; determine the evaluation operating score of the matching enterprise based on the historical operating information and historical corporate behavior; and determine the comprehensive operating score of the matching enterprise based on the relationship between the model operating score and the evaluation operating score.
[0051] S400: Determine whether to identify the matched company as a targeted matching result based on the company's comprehensive business performance score. If it is determined that the matched company should not be identified as a targeted matching result, then the matched company will be re-identified.
[0052] Specifically, enterprise data is obtained by defining a collection period (such as the past three years). The sources of enterprise data can be the National Enterprise Credit Information Publicity System (business registration information), third-party enterprise databases (such as Tianyancha, Qichacha, and Qixinbao), government industrial policy release platforms, industry association platforms, archives of established enterprises, online public opinion, recruitment websites, and supply chain platforms. During the collection process, enterprise data is inevitably affected by interference factors (such as network transmission fluctuations). After preprocessing, valid samples (target enterprise data) are selected to ensure the stability and reliability of the data source. This approach avoids the subjectivity and uncertainty of human judgment based on impressions. It extracts the industry characteristics of each target company's data. The target company data encompasses all information about the company, while the industry characteristics make the core of the company explicit. By extracting these industry characteristics, misjudgments caused by information omissions are avoided. Starting with the industry characteristics, all target company data is extended to construct an industry chain map. The industry chain map intuitively presents the actual situation of all companies within the collection period. Similarity is objectively quantified through mathematical models and other calculation methods to determine the similarity between the industry characteristics to be matched and the industry chain map. The industry characteristics to be matched indicate the type of company required, such as "high-tech enterprise," "export-oriented enterprise," "new energy vehicle enterprise," and "biopharmaceutical enterprise." Matching companies are determined based on similarity and selection preferences. Selection preferences include the expected company size, whether it possesses specific qualifications, and whether there is a preference for companies with existing cooperation. This allows for targeted screening of companies, avoiding subjective targeting based on human experience. The construction of the industry chain map achieves a structured presentation of regional industry demand, ensuring the stability and reliability of targeted matching.
[0053] Understandably, the matching companies are evaluated using two scoring mechanisms: one is based on a recommendation model to determine the model's operational score, and the other is based on historical operational information and historical corporate behavior to calculate an objective evaluation score. The recommendation model is based on big data analysis, dynamically outputting the model's operational score through intelligent learning and judgment. The comprehensive operational score of the matching companies is dynamically determined based on the results of big data analysis and objective calculation. This approach not only uncovers hidden patterns in the data through machine learning models but also verifies its stability with actual operational indicators, avoiding the one-sidedness of a single-dimensional evaluation. The comprehensive operational score is used to determine whether a matching company is included in the targeted matching. Matching companies that do not meet the criteria are discarded, and new matching companies are selected, forming a closed loop of screening-evaluation-feedback-rescreening. This continuously approaches the actual needs of the relevant regions, ensuring the accuracy of targeted matching.
[0054] In some embodiments of this application, when acquiring data from several enterprises based on a collection time period and preprocessing each enterprise's data to determine the target enterprise data, the preprocessing includes: data standardization and outlier detection. Data standardization involves unifying field naming and industry classification coding. Outlier detection involves removing data with zero registered capital and outlier establishment time. Outlier establishment time data is defined as data that is greater than or equal to the right boundary of the collection time period.
[0055] Specifically, in the process of data standardization, unified field naming can eliminate the differences in the description of enterprise data from different sources, such as the synonyms of "main business" and "core business". Unified industry classification coding maps scattered industry descriptions to a standard system, such as the national economic industry classification code, ensuring consistency in subsequent feature extraction. Outlier detection removes enterprises with zero registered capital, excluding shell or invalid entities. Removing abnormal data in establishment time means removing enterprises whose establishment time is later than the end of the collection period, avoiding the inclusion of data with insufficient operating cycle and incomplete information. For example, if the collection period is from March 5, 2020 to March 5, 2023, data after March 5, 2023 will not be considered. By standardizing the data format, analysis interference is eliminated, and invalid samples are filtered to ensure the reliability of target enterprise data. This provides a reliable foundation for the subsequent construction of the industrial chain map, reduces deviations caused by data chaos or invalid samples, and improves the accuracy of targeted matching for investment promotion.
[0056] In some embodiments of this application, the process of extracting industry features from each target enterprise data includes: analyzing each target enterprise data based on NLP, performing word segmentation, removing stop words, and stemming on each target enterprise data, using dependency parsing to analyze the relationships between words in each target enterprise data and extracting noun phrases, and using noun phrases that conform to the syntactic structure as industry features.
[0057] Specifically, natural language processing technology is used to perform word segmentation, stop word removal, and stemming on each target enterprise data. Word segmentation divides the continuous text in the target enterprise data into independent words or phrases for subsequent processing. Stop word removal removes common words that do not affect the actual semantics in the target enterprise data (such as "of", "in", etc.) to improve the accuracy of extraction. Stemming reduces semantic redundancy by reducing words to their root forms (e.g., "enterprise" and "enterprises" are unified as "enterprise"). Dependency parsing is used to analyze the grammatical and semantic relationships between words, and noun phrases are extracted from the target enterprise data. Dependency parsing can accurately identify noun phrases (such as "high-tech enterprises", "export-oriented enterprises", etc.). By combining NLP technology and dependency parsing, the accurate extraction and recognition of the industrial characteristics of each target enterprise data are achieved, and noun phrases based on grammatical structures are extracted, ensuring the semantic value of the extracted content, enabling these noun phrases to represent the actual characteristics and enterprise attributes of the enterprise, and improving the accuracy and reliability of targeted matching.
[0058] In some embodiments of the present application, when using the industrial characteristics as a starting point and constructing an industrial chain map based on all target enterprise data, it includes: taking the industrial characteristics as a starting point, using each data feature of each target enterprise data as a node, and whenever there is a sub-data feature for a data feature, expanding the corresponding node with a sub-node, and constructing all industrial characteristics and the corresponding nodes and sub-nodes into an industrial chain map. The data features include enterprise basic information, basic information, geographical information, industry attributes, operating capabilities, and cooperation relationships.
[0059] Specifically, the industry chain map constructs a hierarchical node network to present relevant information about enterprises. Taking industry characteristics as the core starting point, it transforms data features such as basic information, geographical information, and industry attributes of enterprises into independent nodes. Then, based on sub-data features (such as sub-sectors under industry attributes and upstream and downstream connections in cooperation relationships), sub-nodes are extended to form a multi-dimensional nested network structure. For example, basic information includes name, unified social credit code, legal representative, registered capital, and establishment time; geographical information includes registered address, actual business location, and affiliated industrial park; and operating capabilities include annual output value, number of employees, and percentage of professional R&D investment. Whenever an enterprise discloses a sub-data feature of a data feature (such as name, unified social credit code, legal representative, registered capital, and establishment time of basic information), its corresponding node is expanded by a sub-node. The more nodes and sub-nodes there are for an industry feature (representing an enterprise), the more comprehensive and reliable the fields covered by the enterprise are. By reconstructing the internal logic of the industry through the hierarchical relationship of nodes, the division of labor network of enterprises can be outlined. The nodes at each level map each other to form a complete enterprise industrial chain, providing a clear direction for investment promotion. The dynamic node network can reflect the changes of relevant enterprises in real time, so that the matching process can not only grasp the overall situation of enterprises, but also focus on local details, improve the pertinence of targeted matching for investment promotion, and avoid mismatch caused by blind introduction.
[0060] In some embodiments of this application, when determining the similarity between the industry feature to be matched and the industry chain map, and determining the matching enterprise based on the similarity and selection orientation, the process includes: determining the similarity between the industry feature to be matched and each industry feature, determining similar enterprises based on a preset selection orientation, the preset selection orientation including the expected enterprise size, whether it has specific qualifications, and whether it prefers enterprises with existing cooperation foundations, and determining the similar enterprise with the highest similarity as the matching enterprise.
[0061] Specifically, targeted matching of enterprises is achieved through feature comparison and demand filtering. The similarity is determined by comparing the characteristics of the industry to be matched with the characteristics of the industry in the industrial chain map. The similarity is determined by Euclidean distance or cosine similarity. The similarity reflects the degree of fit between the characteristics of the industry to be matched. Then, a preset selection orientation is introduced as a screening condition to determine similar enterprises. The enterprise size requirement ensures that the matched enterprises meet the regional carrying capacity, and specific qualifications (such as high-tech enterprise certification) guarantee the technical threshold. Enterprises with existing cooperation are preferred, and relatively stable cooperative enterprises are selected first. Finally, the enterprises with the highest similarity are selected as the matched enterprises. The quantitative similarity eliminates the bias of subjective judgment and ensures the objectivity of matching with demand. The preset selection orientation incorporates actual demand into the screening process, so that the results not only fit the industrial chain ecosystem, but also meet personalized investment promotion goals, thereby improving the accuracy and reliability of targeted matching.
[0062] In some embodiments of this application, when obtaining historical operating information and historical corporate behavior of matching enterprises, and determining the model operating score of matching enterprises based on the recommendation model, historical operating information, and historical corporate behavior, the process includes: obtaining a model dataset, sampling the model dataset proportionally to determine a training set and a test set, obtaining a pre-selected random forest model, iteratively training the random forest model based on the training set, testing the iteratively trained random forest model based on the test set, adjusting the learning rate of the currently iteratively trained random forest model if the mean squared error of the currently iteratively trained random forest model is greater than or equal to the mean squared error of the previously iteratively trained random forest model, and continuing iterative training until a preset number of iterations is reached, stopping iterative training if the mean squared error of the currently iteratively trained random forest model is less than the mean squared error of the previously iteratively trained random forest model, determining the currently iteratively trained random forest model as the recommendation model, and outputting the model operating score of matching enterprises based on historical operating information and historical corporate behavior.
[0063] Specifically, the model dataset records the historical operating information and behaviors of all enterprises, along with their corresponding sample scores. The dataset is proportionally divided into training and testing sets, typically in a 4:1 ratio, to ensure both sets contain diverse data, thus improving the model's generalization ability. A random forest model is trained using the training set. By ensembled from multiple decision trees and averaging their predictions, the random forest model reduces the risk of overfitting and improves its accuracy and stability. In each training iteration, the model attempts to learn patterns and relationships in the data to improve its prediction or classification capabilities. The test set is input into the iteratively trained random forest model to determine the mean squared error (MSE) of the current iteration. The MSE reflects the model's performance on unknown data and is an important indicator for evaluating model performance. If the MSE of the current iteration is greater than or equal to the MSE of the random forest model after the previous iteration, it indicates that the model's performance is at risk of decline or overfitting. In this case, it is necessary to reduce the magnitude of the model's gradient change (i.e., the learning rate) and continue iterative training to help the model more stably approach the global optimum. If the MSE of the current iteration is less than the MSE of the random forest model after the previous iteration, it indicates that the model's performance has improved. At this point, iterative training can be stopped, and the model is considered to have reached a satisfactory performance level. The trained random forest model is then designated as the recommendation model. Historical business information and historical enterprise behavior are then substituted into the recommendation model to output a model business score that matches the enterprise, ensuring the reliability and stability of targeted investment matching.
[0064] In some embodiments of this application, when determining the evaluation score of a matching enterprise based on historical business information and historical enterprise behavior, the following steps are taken: setting the initial score of the matching enterprise to zero, and whenever there is an anomaly in the historical business information or an anomaly in the historical enterprise behavior, adding a score to the zero score, and determining the sum of the final added scores as the evaluation score of the matching enterprise.
[0065] Specifically, the evaluation score is calculated by quantifying abnormal behaviors to measure the compliance and stability of a company's historical operations. Using zero as a baseline, it uses abnormalities in historical operational information (such as financial fluctuations and tax records) and historical corporate behavior (such as contract performance and qualification compliance) as the scoring basis. Each abnormality adds points, and the final total score reflects the degree of risk in the company's operations. A higher evaluation score indicates more problems and weaker operational compliance. By transforming scattered abnormal information into intuitive quantitative indicators and using a "score for any violation" rule, the risk in the historical performance of matched companies is objectively presented, avoiding the fuzzy judgment of human experience and ensuring the objectivity of the evaluation score. This reduces the influence of subjective preferences on targeted matching and accurately identifies matched companies with higher operational risks, avoiding the introduction of entities with potential compliance issues or operational instability, thereby ensuring the safety and sustainability of targeted matching and regional cooperation.
[0066] In some embodiments of this application, when determining the overall business performance score of a matching enterprise based on the relationship between the model business performance score and the evaluation business performance score, the method includes: when the model business performance score and the evaluation business performance score are equal, the model business performance score or the evaluation business performance score is determined as the overall business performance score of the matching enterprise; when the model business performance score and the evaluation business performance score are not equal, the average of the model business performance score and the evaluation business performance score is determined as the overall business performance score of the matching enterprise.
[0067] In some embodiments of this application, when determining whether to identify a matching enterprise as a targeted matching result based on the enterprise's comprehensive business score, the method includes: pre-setting a threshold for the enterprise's comprehensive business score, comparing the enterprise's comprehensive business score with the threshold, and determining that the matching enterprise will not be identified as a targeted matching result if the enterprise's comprehensive business score is greater than the threshold, and determining that the matching enterprise will be identified as a targeted matching result if the enterprise's comprehensive business score is less than or equal to the threshold.
[0068] Specifically, the comprehensive business performance score integrates two scoring dimensions to achieve a comprehensive assessment of the matched company's business performance. The model-based business performance score focuses on determining the risk of the matched company based on algorithms, while the evaluation business performance score focuses on the risk level reflected by historical abnormal behaviors. When the two scores are equal, it indicates that the quantitative assessment of the matched company is consistent with the model assessment, and either one can be used as the company's comprehensive business performance score. When the two scores are not equal, taking the average can balance the differences between the two dimensions, thereby improving the reliability and stability of targeted matching. Using a comprehensive business performance score threshold as a baseline, which is based on regional industrial needs and risk tolerance, the threshold reflects the minimum requirements for the operational quality of matched enterprises. When an enterprise's comprehensive business performance score exceeds the threshold, it indicates that the operational risks or insufficient potential of the matched enterprise have exceeded acceptable limits, and the enterprise is not considered a targeted match. Conversely, when the enterprise's comprehensive business performance score is less than or equal to the threshold, it indicates that the overall condition of the matched enterprise meets the investment promotion standards, and the enterprise is considered a targeted match. By comparing the enterprise's comprehensive business performance score with the threshold, the ambiguity of subjective decision-making is eliminated, ensuring that all matched enterprises are examined using the same standard. This avoids the risk of inconsistent human judgment based on experience, guarantees the reliability and stability of targeted investment promotion, and thus improves the efficiency and quality of investment promotion.
[0069] In summary, the beneficial effects of this invention are as follows: By extracting the industry characteristics of each target enterprise's data, the stability and reliability of the data source are ensured, avoiding the subjectivity and uncertainty of human judgment based on impressions. Through the construction and similarity of the industrial chain map, targeted matching between enterprises and regional industrial positioning is achieved, reducing resource mismatch caused by human experience bias, thereby avoiding the waste of resources in ineffective investment promotion. It ensures that the matched enterprises form an organic connection with the regional industrial ecosystem, strengthens the synergistic effect of upstream and downstream of the industrial chain, and the comprehensive operation score achieves a comprehensive evaluation of the operating status of the matched enterprises by integrating two scoring dimensions, improving the reliability and stability of targeted matching, effectively avoiding the one-sidedness of a single evaluation dimension, reducing investment promotion risks, and thus improving the stability of targeted matching in attracting enterprises. Through a cyclical verification mechanism, the matched enterprises are continuously optimized to ensure the adaptability and compatibility of the targeted matching results with regional development needs, thereby improving the efficiency and quality of investment promotion.
[0070] In another preferred embodiment based on the above embodiments, see [reference] Figure 2 As shown, this embodiment provides a supply chain investment promotion targeted matching system based on big data analysis, used to apply the above-mentioned supply chain investment promotion targeted matching method based on big data analysis, including:
[0071] The data acquisition unit is configured to acquire data from several enterprises based on the data acquisition time period, preprocess each enterprise's data to determine the target enterprise's data, and extract the industry characteristics of each target enterprise's data.
[0072] The analysis unit is configured to take industry characteristics as the starting point and extend them based on all target enterprise data to construct an industry chain map, determine the similarity between the industry characteristics to be matched and the industry chain map, and determine the matching enterprises based on the similarity and selection orientation.
[0073] The processing unit is configured to acquire the historical operating information and historical corporate behavior of the matching enterprise, determine the model operating score of the matching enterprise based on the recommendation model, historical operating information and historical corporate behavior, determine the evaluation operating score of the matching enterprise based on the historical operating information and historical corporate behavior, and determine the comprehensive operating score of the matching enterprise based on the relationship between the model operating score and the evaluation operating score.
[0074] The matching unit is configured to determine whether to identify the matching enterprise as a targeted matching result based on the enterprise's comprehensive business score. If it is determined that the matching enterprise should not be identified as a targeted matching result, the matching enterprise will be re-identified.
[0075] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0076] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0077] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a processFigure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0078] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A targeted matching method for investment promotion in the industrial chain based on big data analysis, characterized in that, include: Data from several companies is collected based on the time period, and the data from each company is preprocessed to determine the target company data. The industry characteristics of each target company data are then extracted. Starting with the aforementioned industry characteristics, and extending the industry chain map based on all target enterprise data, the similarity between the industry characteristics to be matched and the industry chain map is determined, and matching enterprises are determined based on the similarity and selection orientation. The process involves obtaining historical operating information and historical business behavior of the matched enterprises; determining a model operating score for the matched enterprises based on a recommendation model, the historical operating information, and the historical business behavior; determining an evaluation operating score for the matched enterprises based on the historical operating information and historical business behavior; and determining a comprehensive enterprise operating score for the matched enterprises based on the relationship between the model operating score and the evaluation operating score. Obtain the model dataset, and sample the model dataset proportionally to determine the training set and test set. Obtain the pre-selected random forest model, and iteratively train the random forest model based on the training set. Test the iteratively trained random forest model based on the test set. If the mean squared error of the random forest model after the current iteration is greater than or equal to the mean squared error of the random forest model after the previous iteration, adjust the learning rate of the random forest model after the current iteration, that is, reduce the magnitude of the model's change in the gradient direction, and continue iterative training until the preset number of iterations is reached. If the mean squared error of the random forest model after the current iteration is less than the mean squared error of the random forest model after the previous iteration, stop iterative training, determine the random forest model after the current iteration as the recommendation model, and output the model operation score of the matching enterprise based on the historical business information and historical enterprise behavior. The initial score of the matched enterprise is set to zero. Whenever there is an anomaly in the historical business information or an anomaly in the historical enterprise behavior, a score is added to the zero score, and the sum of the final added scores is determined as the evaluation business score of the matched enterprise. When the model operation score and the evaluation operation score are equal, the model operation score or the evaluation operation score is determined as the overall business operation score of the matched enterprise; when the model operation score and the evaluation operation score are not equal, the average of the model operation score and the evaluation operation score is determined as the overall business operation score of the matched enterprise. The matching company is determined based on its comprehensive business performance score. If the matching company is not determined as a target matching result, the matching company is re-determined.
2. The targeted matching method for industrial chain investment promotion based on big data analysis according to claim 1, characterized in that, When acquiring data from several enterprises based on a collection time period and preprocessing each enterprise's data to determine the target enterprise data, the following steps are included: The preprocessing includes data standardization and outlier detection; The data standardization includes unified field naming and industry classification coding; The outlier detection involves removing data with zero registered capital and abnormal establishment time. Abnormal data in the establishment time are defined as data that is greater than or equal to the right boundary of the collection time period.
3. The targeted matching method for industrial chain investment promotion based on big data analysis according to claim 2, characterized in that, When extracting industry characteristics from the data of each target enterprise, the following are included: Based on NLP analysis of each target company's data, word segmentation, stop word removal, and stemming are performed on each target company's data. Dependency parsing is used to analyze the relationships between words in each target company's data and extract noun phrases. The noun phrases that conform to the syntactic structure are taken as the industry features.
4. The targeted matching method for industrial chain investment promotion based on big data analysis according to claim 3, characterized in that, When using the aforementioned industry characteristics as a starting point and extending the industry chain map based on all target enterprise data, the following is included: Starting with the industry characteristics, each data feature of each target enterprise's data is taken as a node. Whenever a data feature has a sub-data feature, the corresponding node is expanded by a sub-node. All industry characteristics, as well as the corresponding nodes and sub-nodes, are used to construct the industry chain map. The data features include basic enterprise information, geographic information, industry attributes, operational capabilities, and cooperative relationships.
5. The targeted matching method for industrial chain investment promotion based on big data analysis according to claim 4, characterized in that, When determining the similarity between the characteristics of the industry to be matched and the industry chain map, and determining the matching enterprises based on the similarity and selection orientation, the process includes: Determine the similarity between the industry feature to be matched and each industry feature; Similar companies are identified based on preset selection criteria, which include the expected company size, whether the company possesses specific qualifications, and whether there is a preference for companies with existing cooperative relationships. The company with the highest similarity is then identified as the matching company.
6. The targeted matching method for industrial chain investment promotion based on big data analysis according to claim 1, characterized in that, When determining whether to identify the matched enterprise as a targeted matching result based on the enterprise's comprehensive business performance score, the following steps are included: A pre-set threshold for the enterprise's overall business performance score is used to compare the enterprise's overall business performance score with the threshold. If the enterprise's overall business score is greater than the enterprise's overall business score threshold, then the matched enterprise will not be identified as a targeted matching result. When the enterprise's overall business score is less than or equal to the enterprise's overall business score threshold, the matched enterprise is determined as the targeted matching result.
7. A targeted matching system for investment promotion based on big data analysis, used to apply the targeted matching method for investment promotion based on big data analysis as described in any one of claims 1-6, characterized in that, include: The data acquisition unit is configured to acquire data from several enterprises based on the data acquisition time period, preprocess each enterprise's data to determine the target enterprise's data, and extract the industry characteristics of each target enterprise's data. The analysis unit is configured to take the industry characteristics as a starting point, extend and construct an industry chain map based on all target enterprise data, determine the similarity between the industry characteristics to be matched and the industry chain map, and determine the matching enterprise based on the similarity and selection orientation. The processing unit is configured to acquire the historical operating information and historical business behavior of the matched enterprise, determine the model operating score of the matched enterprise based on the recommendation model, the historical operating information and historical business behavior, determine the evaluation operating score of the matched enterprise based on the historical operating information and historical business behavior, and determine the comprehensive operating score of the matched enterprise based on the relationship between the model operating score and the evaluation operating score. The matching unit is configured to determine whether to identify the matching enterprise as a targeted matching result based on the enterprise's comprehensive business score. If it is determined that the matching enterprise should not be identified as a targeted matching result, the matching enterprise is re-identified.