Design method for constructing high-tech enterprise group portrait based on enterprise data

By constructing a "production-creation" evaluation model and combining it with SE-DEA and GBRT models, the innovation capabilities of high-tech enterprises are assessed. This solves the problem of difficulty in identifying high-tech enterprises in existing technologies, achieves accurate and real-time traceability of enterprise group profiles, and provides multi-faceted data support.

CN121958635APending Publication Date: 2026-05-01SICHUAN ZHIHE TECHNOLOGY CONSULTING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN ZHIHE TECHNOLOGY CONSULTING CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for constructing enterprise profiles are insufficient to accurately identify high-tech enterprises and fail to fully consider their innovation capabilities and production value creation, making it difficult to achieve accurate information acquisition and real-time tracking of high-tech enterprises.

Method used

By constructing a "production-creation" evaluation model based on enterprise data, and combining the SE-DEA model and gradient boosting regression tree GBRT, the innovation capabilities of enterprises are assessed, stratified and clustered, a profile of high-tech enterprise groups is established, and the parameter range is optimized through performance evaluation.

Benefits of technology

It enables the assessment and stratification of the innovation capabilities of high-tech enterprises, ensuring the accuracy of enterprise attributes and real-time data traceability, providing intuitive and quantitative data support, and supporting industrial policy formulation, supply chain analysis and precision services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958635A_ABST
    Figure CN121958635A_ABST
Patent Text Reader

Abstract

The invention relates to the field of enterprise group portrait design, in particular to a design method for constructing a high-tech enterprise group portrait based on enterprise data, and realizes real-time information tracing and accurate information confirmation of high-tech enterprise production data information through analysis and design of the enterprise data. The design method comprises the following steps: step 1, capturing basic information and production industry data of each enterprise from an enterprise database; 2, inputting the production industry data into a preset production-creation evaluation model for analysis, and obtaining an innovation ability score; step 3, layering each enterprise according to the innovation ability score, taking a layering result as a label, performing feature clustering in combination with enterprise process parameters and productivity utilization rate, and establishing an enterprise group portrait; and 4, performing performance evaluation on the constructed enterprise group portrait, and optimizing and adjusting a parameter interval of the enterprise group portrait according to a performance evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

A design method for constructing profiles of high-tech enterprise groups based on enterprise data Technical Field

[0001] This invention relates to the field of enterprise group profiling design, specifically to a design method for constructing high-tech enterprise group profiling based on enterprise data. Background Technology

[0002] With the rapid development of the market economy, more and more technology companies are being registered and established, resulting in a massive amount of enterprise operation data. To address this massive amount of enterprise operation data, data mining techniques are needed to analyze and study enterprise information. In particular, during the process of building enterprise profiles for high-tech enterprises, it is necessary to fully decompose enterprise information into dimensions, extract labels from each dimension, and create an enterprise profile label map for each enterprise.

[0003] In order to identify high-tech innovation and sustainability and improve policy targeting during the development and expansion of high-tech enterprises, multi-dimensional labels such as intellectual property rights, R&D expenses, technology income, and personnel structure can be used to quickly screen out enterprises that engage in "certificate arbitrage" and ensure that tax incentives and R&D subsidies flow to high-growth innovative entities.

[0004] However, existing common enterprise group profiles typically require depicting the enterprise from three dimensions: basic enterprise attributes, business operations, and risk information. However, the concept of enterprise attributes for high-tech enterprises is vague and does not fully consider the role of actual production value creation in innovative technology enterprises, making it difficult to clearly distinguish between genuine and pseudo-high-tech enterprises. Furthermore, there are many differences in the data across different dimensions for different high-tech enterprises. Therefore, how to accurately integrate and stratify the enterprise data information of high-tech enterprises, and achieve a complete real-time traceability of high-tech enterprise production data information and a precise acquisition of high-tech enterprise attributes, has become an urgent technical challenge to be solved. Summary of the Invention

[0005] The purpose of this invention is to provide a design method for constructing a profile of high-tech enterprise groups based on enterprise data, and to solve the following technical problem: how to improve the real-time information traceability and accurate information confirmation of high-tech enterprise production data information through the analysis and design of enterprise data.

[0006] The objective of this invention can be achieved through the following technical solution: a design method for constructing a high-tech enterprise group profile based on enterprise data, the method comprising: Step 1, retrieving basic information and production industry data of each enterprise from an enterprise database; the production industry data includes production efficiency, capacity utilization rate, and technology transfer rate; Step 2, inputting the production industry data into a preset "production-creation" evaluation model for analysis to obtain an innovation capability score; Step 3, stratifying each enterprise according to the innovation capability score, using the stratification results as labels, and performing feature clustering based on enterprise process parameters and capacity utilization rate to establish an enterprise group profile; Step 4, conducting performance evaluation on the constructed enterprise group profile, and optimizing and adjusting the parameter range of the enterprise group profile based on the performance evaluation results.

[0007] Preferably, the construction method of the "production-creation" evaluation model in step two is as follows: S1. Using historical production industry data of enterprises as the training set, based on the SE-DEA model, historical production efficiency and historical capacity utilization rate are used as input variables of the model, and the technology transfer rate is used as the output value to obtain the relative efficiency value of each enterprise, and the optimal relative efficiency value of each enterprise is used as the production efficiency vector; S2. Using the net profit increment of enterprises in the same training set in the expected future time period as the label, a gradient boosting regression tree GBRT is constructed, with the technology transfer rate and production efficiency vector as input, and the predicted profit increment is output to obtain the value creation vector; S3. After normalizing the production efficiency vector and the value creation vector by minimum-maximum, the weights are dynamically allocated by the entropy weight method to establish a weighted linear fusion function, thus forming the "production-creation" evaluation model.

[0008] Preferably, in step S1: for the first Solving linear programming problems for businesses: Obtaining the objective function Constraints: Input conditions: and Output conditions: ; , and ;in, Indicates the total number of enterprises; Indicates the type of input variable. Inputs representing production efficiency Inputs indicating capacity utilization rate; include and ; Indicates enterprise Historical production efficiency; Indicates enterprise Historical capacity utilization rate; Indicates enterprise Actual input variables ; Indicates enterprise The rate of technology transfer; Indicates enterprise Weight variables, The desired superefficiency value; Other companies The rate of technology transfer; Indicates the required output after amplification; optimal solution For the first The relative efficiency value of a company will individual enterprises Calculate them sequentially to form a column vector. ,and .

[0009] Preferably, the value creation vector is: constructing a training set. ;in, ; Indicates enterprise The rate of technology transfer; Indicates enterprise The production efficiency vector; through the GBRT loss function:

[0010] in, This represents the m-th CART regression tree. Its splitting characteristics, threshold, and leaf value; This refers to the learning rate, also known as the shrinkage coefficient. For exporting companies The predicted profit increment; after forecasting for all n companies, the value creation vector is obtained: Preferably, the calculation method for innovation capability score is as follows: extracting enterprise... Production efficiency vector relative efficiency value and enterprises Value creation vector Predicted profit increment ;Will and Perform min-max normalization:

[0011]

[0012] in, Represents the normalized i-th The production efficiency value of a company Represents the normalized i-th The value created by a company This represents the minimum relative efficiency value. This represents the maximum relative efficiency value. To minimize the predicted profit increment, To maximize the predicted profit increment, a weighted calculation is performed:

[0013] in, For enterprises The score for innovation ability, and ∈ ; The weighting coefficients for production efficiency indicators. Weighting coefficients for value creation metrics; and , All are greater than 0.

[0014] Preferably, the determination of the predicted profit increment also includes introducing a piecewise penalty term based on the technology transfer rate into the GBRT loss function: if the enterprise's technology transfer rate is lower than the industry's preset minimum percentile, then the penalty coefficient is increased. If a company's technology transfer rate exceeds the industry's preset highest percentile, then the penalty coefficient will be increased. If a company's technology transfer rate falls between the industry's preset highest and lowest percentiles, then the penalty coefficient will be adjusted accordingly. .

[0015] Preferably, the method for stratifying enterprises in step three is as follows: SS1, use the Jenks Natural Breaks algorithm to optimize the breakpoint division of EIS values ​​to minimize the variance within groups and maximize the variance between groups, and output three-level labels: "high-level innovation, mid-level innovation, and early-stage innovation"; SS2, after concatenating the three-level labels of "high-level innovation, mid-level innovation, and early-stage innovation" with the enterprise's process parameter vector and capacity utilization rate, use density peak clustering (DPC) to calculate local density and relative distance, and automatically determine the cluster center through the decision graph to form a group profile cluster with innovation capability gradient.

[0016] Preferably, the performance evaluation method is as follows: calculate the Calinski-Harabasz index of the hierarchical-clustered profile. With contour coefficient And through the formula Calculate comprehensive index Preferably, the comprehensive index Compared with the preset comprehensive index threshold range Compare and convert the profile coefficients With preset coefficient threshold range Compare and judge: If ≥ and ≥ If the portrait is deemed "excellent," the current parameter range is maintained; if ≤ < or ≤ < If it is deemed "qualified", a fine-tuning operation is triggered; if < or < If the result is deemed "unqualified," global optimization is initiated, and iterative calculations are re-executed until the comprehensive index is reached. or profile coefficient qualified.

[0017] The beneficial effects of this invention are as follows: This invention further analyzes the basic information of high-tech enterprises and their high-tech-related production industry data. By extracting enterprise production data information and inputting the production industry data into a preset "production-creation" evaluation model for analysis, an innovation capability score is obtained. This innovation capability score enables further hierarchical clustering of high-tech enterprises. Based on the innovation capability score, the enterprises are stratified and labeled. After labeling, feature clustering is performed by combining enterprise process parameters and capacity utilization rates. This clustering achieves the goal of constructing a profile of high-tech enterprise groups, revealing the innovative development status of high-tech enterprises. Through the constructed profile of enterprises… The system performs performance evaluation on enterprise group profiles and optimizes and adjusts the parameter ranges of these profiles based on the evaluation results. It also confirms the conceptual attributes of high-tech enterprises, ensuring full consideration of actual production value creation attributes. This allows for data integration and stratification of different high-tech enterprises across industrial creativity dimensions, enabling comprehensive traceability of high-tech enterprise production data. Furthermore, it confirms the profiles of high-tech enterprises' attributes, ensuring the construction of enterprise group profiles reflecting innovation capabilities and production characteristics. This provides intuitive and quantitative data support for various aspects, including industrial policy formulation, supply chain analysis, investment attraction, and targeted services.

[0018] Of course, any product implementing this invention does not necessarily need to achieve all the advantages described above at the same time. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 is a flowchart illustrating the design steps of a high-tech enterprise group profile based on enterprise data according to the present invention; Figure 2 is a flowchart illustrating the construction steps of the "production-creation" evaluation model according to the present invention; Figure 3 is a flowchart illustrating the steps of the method of classifying enterprises according to the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please refer to Figure 1. This invention is a design method for constructing a profile of a high-tech enterprise group based on enterprise data. The method includes: Step 1: Extracting basic information and production industry data of each enterprise from the enterprise database; the production industry data includes production efficiency, capacity utilization rate, and technology transfer rate; Step 2: Inputting the production industry data into a preset "production-creation" evaluation model for analysis to obtain an innovation capability score; Step 3: Stratifying each enterprise according to the innovation capability score, and using the stratification results as labels, combining enterprise process parameters and capacity utilization rate to perform feature clustering and establish an enterprise group profile; Step 4: Performing performance evaluation on the constructed enterprise group profile, and optimizing and adjusting the parameter range of the enterprise group profile based on the performance evaluation results.

[0023] In the above technical solution, firstly, step one involves real-time data capture from a pre-set enterprise database, which can integrate data from multiple data sources such as industry and commerce, taxation, intellectual property, statistics, and self-reported data from enterprises. Typically, for massive amounts of operational data from high-tech enterprises, data mining techniques are used to analyze and study enterprise information, decompose the information into dimensions, and extract labels from each dimension. The method of creating a label profile for each enterprise relies on the accuracy of the labels. Traditionally, labels need to fully characterize the profile of such enterprises, requiring at least three categories of information: basic enterprise attributes, business operations, and risk information. This ensures sufficient confirmation of enterprise characteristics for general enterprises. However, this design focuses on acquiring two major categories of information for high-tech enterprises: basic information about high-tech enterprises and their high-tech-related production industry data. The production industry data mainly includes production efficiency, capacity utilization rate, and technology transfer rate. Step two involves further analysis of the production industry data for this type of enterprise. By extracting enterprise production data information, a data-driven analysis of the technological production and technological creation of high-tech enterprises is conducted. Specifically, the production industry data is input into a pre-set "production-creation" evaluation model for analysis to obtain an innovation capability score. This innovation capability score allows for further hierarchical clustering of high-tech enterprises. Step three involves stratification based on the innovation capability score, followed by labeling based on the stratification results. After labeling, feature clustering is performed by combining enterprise process parameters and capacity utilization rate. This clustering achieves the goal of constructing a profile of high-tech enterprise groups, revealing the innovative development status of high-tech enterprises. Finally, step four involves performance evaluation of the constructed enterprise group profile, and optimization and adjustment of the parameter range of the enterprise group profile based on the performance evaluation results.

[0024] The above-described method for constructing enterprise group profiles of high-tech enterprises provides a method that, in addition to obtaining enterprise profiles through common enterprise information confirmation methods, further confirms the concept of enterprise attributes for high-tech enterprise types. This ensures that the actual production value creation attributes are fully considered. Through the design of the above technical methods and steps, data can be integrated and layered for different high-tech enterprises based on the dimension of industrial creativity. This achieves real-time traceability of high-tech enterprise production data information and completes the profile confirmation of high-tech enterprise attributes. It ensures that enterprise group profiles reflecting innovation capabilities and production characteristics can be constructed from the attributes of high-tech enterprises, providing intuitive and quantitative data support for various aspects such as industrial policy formulation, industrial chain analysis, investment promotion, and precision services.

[0025] As one embodiment of the present invention, please refer to Figure 2. The construction method of the "production-creation" evaluation model in step two is as follows: S1. Using the historical production industry data of enterprises as the training set, based on the SE-DEA model, the historical production efficiency and historical capacity utilization rate are used as input variables of the model, and the technology transfer rate is used as the output value to obtain the relative efficiency value of each enterprise. The optimal relative efficiency value of each enterprise is used as the production efficiency vector; S2. Using the net profit increment of enterprises in the same training set in the expected future time period as the label, a gradient boosting regression tree GBRT is constructed. The technology transfer rate and production efficiency vector are used as inputs, and the predicted profit increment is output to obtain the value creation vector; S3. After normalizing the production efficiency vector and the value creation vector by minimum-maximum, the weights are dynamically allocated by the entropy weight method to establish a weighted linear fusion function, thus forming the "production-creation" evaluation model.

[0026] In the above technical solution, the construction of the "production-creation" evaluation model in the steps realizes the evaluation process of the innovation capability of high-tech enterprises from efficiency to direct value orientation. First, step S1 uses production efficiency and capacity utilization rate as training set inputs and technology achievement conversion rate as output to evaluate the relative efficiency of enterprises in converting production resources into technological achievements. This not only eliminates the size difference of high-tech enterprises of the same level, but also enables a relatively fair measurement of the cost-effectiveness and efficiency level of the innovation process. It allows the use of the SE-DEA model to distinguish the true efficiency values ​​of different enterprises, obtain a continuous production efficiency vector, and provide a basis for accurate stratification. Furthermore, the GBRT model in step S2 is used for prediction. The prediction uses the efficiency and conversion rate obtained in step S1 as basic inputs and uses the nonlinear fitting calculation of gradient boosting regression tree to predict the potential economic value of enterprises. This process ensures that technological innovation activities are linked to the final market value and economic contribution. That is, it eliminates the analysis of the subjective efficiency of enterprise content and focuses on the objective performance of market value and economic contribution. This step can realize the data-driven representation of the attributes (innovation) of high-tech enterprises, making the evaluation dimensions more scientific.

[0027] Finally, step S3 integrates the economic value from production to creation in steps S1 and S2 and makes a comprehensive evaluation. It fully merges the vector representing current efficiency and the vector representing future value, avoiding deviations in static evaluation results due to subjectively set weights. This emphasizes that the model can be optimized according to the degree of variation of the data itself, realizing the minimum-maximum normalization of the production efficiency vector and the value creation vector, and then dynamically allocating weights through the entropy weight method to establish a weighted linear fusion function. It also enables the constructed "production-creation" evaluation model to dynamically balance the evaluation of the current production efficiency of high-tech enterprises and the predicted economic effects in the future, enhancing the dynamic predictability and causal logic of the model. It also ensures that the evaluation results are confirmed through objective data-driven modeling, resulting in a more comprehensive and objective innovation score.

[0028] As one embodiment of the present invention, in step S1: for the first Solving linear programming problems for businesses: Obtaining the objective function Constraints: Input conditions: and Output conditions: ; , and ;in, Indicates the total number of enterprises; Indicates the type of input variable. Inputs representing production efficiency Inputs indicating capacity utilization rate; include and ; Indicates enterprise Historical production efficiency; Indicates enterprise Historical capacity utilization rate; Indicates enterprise Actual input variables ; Indicates enterprise The rate of technology transfer; Indicates enterprise Weight variables, The desired superefficiency value; Other companies The rate of technology transfer; Indicates the required output after amplification; optimal solution For the first The relative efficiency value of a company will individual enterprises Calculate them sequentially to form a column vector. ,and .

[0029] In the above technical solution, during the data preparation and input phase, historical data from n companies are collected to form three datasets; each company... dataset Representing historical production efficiency and dataset This indicates the historical capacity utilization rate; the company Output dataset This represents the technology transfer rate; then, by analyzing each enterprise... ( ≠ ) loop through the target function. This involves finding the superefficiency value that maximizes the reduction of the target firm's input; this is achieved by constructing constraints, specifically including input constraints. and The purpose is to evaluate the enterprise. Each input resource consumed by the constructed reference firm must not exceed the actual resource consumption of firm k itself; the output condition is... The purpose is to serve as the output of a virtual reference enterprise, and it must at least reach the level of the enterprise. Current actual output times; and , and The above approach, for each enterprise k (k ranges from 1 to n), formulates the objective function and constraints into an independent linear programming problem. This problem is then solved using an optimization solver (such as the simplex method or interior-point method, which can be implemented using tools like MATLAB, Python's SciPy, or R's lpSolve) to obtain the optimal solution, i.e., the optimal superefficiency value. and the corresponding weight vector After solving for all n companies, n optimal superefficiency values ​​are obtained: Arrange these values ​​in order to form an n-dimensional column vector, which is the Production Efficiency Vector (PTE). Each element in this vector quantifies the relative efficiency of the corresponding enterprise in the process of "converting production resources (efficiency, capacity) into technological achievements". The higher the value, the more advantageous its "production-innovation" conversion efficiency is among the same batch of enterprises.

[0030] As one embodiment of the present invention, the value creation vector is: constructing a training set. ;in, ; Indicates enterprise The rate of technology transfer; Indicates enterprise The production efficiency vector; through the GBRT loss function:

[0031] in, This represents the m-th CART regression tree. Its splitting characteristics, threshold, and leaf value; This refers to the learning rate, also known as the shrinkage coefficient. For exporting companies The predicted profit increment; after forecasting for all n companies, the value creation vector is obtained: In the above technical solution, the construction method and technical implementation process of the value creation vector within the "production-creation" evaluation model are further explained. The training dataset is constructed by collecting historical panel data of the target enterprise group to build a supervised learning training set. Each sample in the training set corresponds to an enterprise j, and its structure is as follows: ,in, The feature vector input to the model is a two-dimensional vector, and express ; This represents the technology transfer rate of enterprise j during the historical observation period. This data comes from the production industry data captured and calculated in step one of this design. This represents the production efficiency value of firm j; this data originates from the corresponding element in the production efficiency vector PTE calculated using the SE-DEA model in step S1. ; It quantifies the relative efficiency with which the company transforms production inputs into technological achievements; The label represents the model's output, indicating the expected increase in net profit for company j within a future timeframe (e.g., the next fiscal year, or the next two years) after the historical observation period ends. This data needs to be obtained from the company's subsequent financial database (such as annual reports and tax data) and is used to train the model to learn the mapping relationship from current features to future value creation. The training set consists of n samples. .

[0032] Next, based on the constructed training set, a gradient boosting regression tree model is trained, whose prediction function is... Defined as an additive combination of a series of regression trees, the formula is:

[0033] The technical implementation process of the above model training and prediction is as follows: Set initial prediction values ​​(such as all sample labels). The mean of the training set is used to construct the learner, and then the learner is iteratively constructed for M rounds (m=1, 2, ..., M). In each round of iteration m: the negative gradient of the loss function (calculated as mean squared error MSE) with respect to the predicted value of the current model is obtained to obtain the pseudo residual, which represents the direction of the difference between the current model's prediction and the true value. This process is the existing calculation method of the GBRT model and will not be described in detail here. Then, a CART regression tree is used. This tree fits the pseudo-residual calculated in the previous step; it searches for the optimal... (i.e., splitting features, splitting threshold, and the output value of each leaf node) to minimize residuals; in this embodiment, the input features That is, a two-dimensional vector ; the newly fitted tree The update rule is added to the current model as follows: ,in, The learning rate (shrinkage coefficient) is a positive number between 0 and 1 (e.g., 0.1) used to control the contribution of each tree to the final model. Adding a learning rate is an effective regularization technique that can prevent overfitting and improve the model's generalization ability. After M iterations, the final strong learner is obtained. That is, the weighted sum of M regression trees; this model can capture the features TCR and PTE and the target The complex nonlinear relationships between them; finally, the trained GBRT model is used. Make predictions for all n firms (including the training sample and new firms to be evaluated); for each firm... Input its corresponding feature vector The model outputs its predicted profit increment. This predicted value represents the company's value creation potential over a future period, based on its current "technology transfer rate" and "production-innovation conversion efficiency." The model iterates through all n companies to obtain their corresponding predicted values. Arrange these predicted values ​​in order to form an n-dimensional column vector, which is the value creation vector VAU: And so, that concludes the discussion.

[0034] As one embodiment of the present invention, the innovation capability score is calculated as follows: extracting enterprise... Production efficiency vector relative efficiency value and enterprises Value creation vector Predicted profit increment ;Will and Perform min-max normalization:

[0035]

[0036] in, Represents the normalized i-th The production efficiency value of a company Represents the normalized i-th The value created by a company This represents the minimum relative efficiency value. This represents the maximum relative efficiency value. To minimize the predicted profit increment, To maximize the predicted profit increment, a weighted calculation is performed:

[0037] in, For enterprises The score for innovation ability, and ∈ ; The weighting coefficients for production efficiency indicators. Weighting coefficients for value creation metrics; and , All are greater than 0.

[0038] The aforementioned technical solution also includes a more detailed calculation of the innovation capability score, namely, extracting the enterprise's... Production efficiency vector relative efficiency value and enterprises Value creation vector Predicted profit increment ;Will and By performing min-max normalization, the production efficiency value and value creation value can be further integrated and processed. The final innovation capability score is determined through weighted calculation. The numerical result represents the strength of the comprehensive innovation capability of high-tech enterprises, and the result of the score calculation can be used for subsequent enterprise stratification and profiling clustering processes.

[0039] As one embodiment of the present invention, the determination of the predicted profit increment also includes introducing a piecewise penalty term based on the technology transfer rate into the GBRT loss function: if the enterprise's technology transfer rate is lower than the industry's preset minimum percentile, then the penalty coefficient is... If a company's technology transfer rate exceeds the industry's preset highest percentile, then the penalty coefficient will be increased. If a company's technology transfer rate falls between the industry's preset highest and lowest percentiles, then the penalty coefficient will be adjusted accordingly. .

[0040] The aforementioned technical solution introduces industry benchmarks (highest / lowest percentiles) through segmented penalty terms, classifying enterprises into three innovation output quality levels based on their Total Credential Ratio (TCR), and applying different prediction confidence adjustments to each level. This allows for model training on the low-quality innovation level, increasing attention to prediction errors for enterprises in this level and effectively preventing the overestimation of the value of enterprises with "production capacity but no transformation." Furthermore, it encourages high-tech enterprises corresponding to the high-quality innovation level, more fully exploring and recognizing the growth potential of high-conversion-rate enterprises, and avoiding underestimating their value. In addition, it maintains a standard evaluation for the mainstream intermediate level, making the model's prediction results more consistent with economic logic and industry reality.

[0041] As one embodiment of the present invention, please refer to Figure 3. The method of stratifying enterprises in step three is as follows: SS1, use the Jenks Natural Breaks algorithm to optimize the breakpoint division of EIS values ​​to minimize the variance within groups and maximize the variance between groups, and output three-level labels: "high-level innovation, medium-level innovation, and early-stage innovation"; SS2, after concatenating the three-level labels of "high-level innovation, medium-level innovation, and early-stage innovation" with the enterprise process parameter vector and capacity utilization rate, use density peak clustering (DPC) to calculate the local density and relative distance, and automatically determine the cluster center through the decision graph to form a group profile cluster with innovation capability gradient.

[0042] In the above technical solution, the Jenks Natural Breaks algorithm is used to process the EIS (Enterprise Innovation Index) score. This algorithm finds the combination of breakpoints that minimizes the sum of variances within each layer and maximizes the variances between layers through iterative optimization. Based on the natural clustering of EIS values, rather than subjectively setting fixed thresholds (such as equal division or percentage), the three levels of "high-level innovation," "medium-level innovation," and "early-stage innovation" are divided. The category labels generated in the previous step are encoded and concatenated with continuous enterprise process parameter vectors (such as the proportion of R&D personnel, the code of the technical field, etc.) and capacity utilization rate values ​​to form a composite feature vector that integrates "innovation capability level" and "production and operation characteristics." The clustering algorithm automatically identifies the distance center by simultaneously calculating the local density and relative distance of each data point, and the output automatically forms several "profile clusters." Enterprises within each cluster have high similarity in innovation capability level, process characteristics, and capacity utilization, while there are significant differences between different clusters.

[0043] As one embodiment of the present invention, the specific method for performance evaluation is as follows: calculating the Calinski-Harabasz index of the profile after hierarchical-clustering. With contour coefficient And through the formula Calculate comprehensive index By combining indicators Compared with the preset comprehensive index threshold range Compare and convert the profile coefficients With preset coefficient threshold range Compare and judge: If ≥ and ≥ If the portrait is deemed "excellent," the current parameter range is maintained; if ≤ < or ≤ < If it is deemed "qualified", a fine-tuning operation is triggered; if < or < If the result is deemed "unqualified," global optimization is initiated, and iterative calculations are re-executed until the comprehensive index is reached. or profile coefficient qualified.

[0044] In the above technical solution, the performance of the profile can be further evaluated by confirming comprehensive indicators. The Calinski-Harabasz index is used to measure the overall comparison between inter-cluster separation and intra-cluster compactness; the larger the value, the more obvious the differences between different profile clusters, and the more similar the enterprises within the same cluster. The silhouette coefficient measures the matching degree between each sample point and its cluster and nearest neighbor clusters, and the average value reflects the clarity and consistency of the clustering structure. Then, the harmonic mean of these two values ​​is calculated by calculating the comprehensive indicators. and When both are relatively high Only by addressing these weaknesses can a high score be achieved; and any deficiency in any aspect will significantly lower the overall score; thus, the comprehensiveness and rigor of the evaluation are ensured. This comprehensive indicator effectively guarantees that each output of enterprise group profiles has high discriminative power, clarity, and practicality, improving the model's adaptability.

[0045] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, and therefore described more simply; relevant parts can be referred to the descriptions of the method embodiments.

[0046] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this application, they should all fall within the protection scope of the present invention.

Claims

1. A design method for constructing a profile of high-tech enterprise groups based on enterprise data, characterized in that, The method includes: Step 1, retrieving basic information and production industry data of each enterprise from the enterprise database; the production industry data includes production efficiency, capacity utilization rate, and technology transfer rate; Step 2, inputting the production industry data into a preset "production-creation" evaluation model for analysis to obtain an innovation capability score; Step 3, stratifying each enterprise according to the innovation capability score, using the stratification results as labels, and performing feature clustering based on enterprise process parameters and capacity utilization rate to establish an enterprise group profile; Step 4, conducting performance evaluation on the constructed enterprise group profile, and optimizing and adjusting the parameter range of the enterprise group profile based on the performance evaluation results.

2. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 1, characterized in that, The construction method of the "production-creation" evaluation model in step two is as follows: S1. Using historical production industry data of enterprises as the training set, based on the SE-DEA model, historical production efficiency and historical capacity utilization rate are used as input variables of the model, and the technology transfer rate is used as the output value to obtain the relative efficiency value of each enterprise. The optimal relative efficiency value of each enterprise is used as the production efficiency vector; S2. Using the net profit increment of enterprises in the same training set in the expected future time period as labels, a gradient boosting regression tree GBRT is constructed. The technology transfer rate and production efficiency vector are used as inputs, and the predicted profit increment is output to obtain the value creation vector; S3. After normalizing the production efficiency vector and the value creation vector by minimum-maximum, the weights are dynamically allocated by the entropy weight method to establish a weighted linear fusion function, thus forming the "production-creation" evaluation model.

3. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 2, characterized in that, In step S1: for the first Solving linear programming problems for companies: For the first... Solving linear programming problems for businesses: Obtaining the objective function Constraints: Input conditions: and Output conditions: ; , and ;in, Indicates the total number of enterprises; Indicates the type of input variable. Inputs representing production efficiency Inputs indicating capacity utilization rate; include and ; Indicates enterprise Historical production efficiency; Indicates enterprise Historical capacity utilization rate; Indicates enterprise Actual input variables ; Indicates enterprise The rate of technology transfer; Indicates enterprise Weight variables, The desired superefficiency value; Other companies The rate of technology transfer; Indicates the required output after amplification; optimal solution For the first The relative efficiency value of a company will individual enterprises Calculate them sequentially to form a column vector. ,and 。 4. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 5, characterized in that, The value creation vector is: constructing a training set. ;in, ; Indicates enterprise The rate of technology transfer; Indicates enterprise The production efficiency vector; through the GBRT loss function: ,in, This represents the m-th CART regression tree. Its splitting characteristics, threshold, and leaf value; This refers to the learning rate, also known as the shrinkage coefficient. For exporting companies The predicted profit increment; after forecasting for all n companies, the value creation vector is obtained: 。 5. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 4, characterized in that, The innovation capability score is calculated as follows: extracting enterprise... Production efficiency vector relative efficiency value and enterprises Value creation vector Predicted profit increment ; Will and Perform min-max normalization: , ,in, Represents the normalized i-th The production efficiency value of a company Represents the normalized i-th The value created by a company This represents the minimum relative efficiency value. This represents the maximum relative efficiency value. To minimize the predicted profit increment, To maximize the predicted profit increment, a weighted calculation is performed: ,in, For enterprises The score for innovation ability, and ∈ ; The weighting coefficients for production efficiency indicators. Weighting coefficients for value creation metrics; and 、 All are greater than 0.

6. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 5, characterized in that, The determination of the predicted profit increment also includes introducing a piecewise penalty term based on the technology transfer rate into the GBRT loss function: if the enterprise's technology transfer rate is lower than the industry's preset minimum percentile, then the penalty coefficient is... If a company's technology transfer rate exceeds the industry's preset highest percentile, then the penalty coefficient will be increased. If a company's technology transfer rate falls between the industry's preset highest and lowest percentiles, then the penalty coefficient will be adjusted accordingly. 。 7. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 1, characterized in that, The method for stratifying enterprises in step three is as follows: SS1, use the Jenks Natural Breaks algorithm to optimize the breakpoint division of EIS values ​​to minimize the within-group variance and maximize the between-group variance, and output three-level labels: "high-level innovation, medium-level innovation, and early-stage innovation"; SS2, after concatenating the three-level labels of "high-level innovation, medium-level innovation, and early-stage innovation" with the enterprise's process parameter vector and capacity utilization rate, use density peak clustering (DPC) to calculate local density and relative distance, and automatically determine the cluster center through the decision graph to form a group profile cluster with innovation capability gradient.

8. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 1, characterized in that, The specific performance evaluation method is as follows: calculate the Calinski-Harabasz index of the hierarchical-clustered profile. With contour coefficient And through the formula Calculate comprehensive index 。 9. The design method for constructing a high-tech enterprise group profile based on enterprise data according to claim 8, characterized in that, Comprehensive indicators Compared with the preset comprehensive index threshold range Compare and convert the profile coefficients With preset coefficient threshold range Compare and judge: If ≥ and ≥ The image is deemed "excellent," and the current parameter range is maintained. like ≤ < or ≤ < The result is "qualified", triggering a fine-tuning operation; like < or < If the result is deemed "unqualified," global optimization is initiated, and iterative calculations are re-executed until the comprehensive index is reached. or profile coefficient Passed.