Enterprise performance capability evaluation method and system based on big data
By constructing a multi-level indicator system and information weighting method, combined with logistic regression algorithm, the problem of inaccurate evaluation of corporate performance capability in existing technologies has been solved, realizing a comprehensive and accurate assessment of corporate performance capability and supporting corporate decision-making in business activities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AEROSPACE SCI & ENG NETWORK INFORMATION DEV CO LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies cannot fully and accurately reflect a company's ability to fulfill its obligations, leading to inaccurate evaluations and impacting the company's reputation and cooperative relationships in business activities.
We construct a big data-based evaluation method for enterprise performance capabilities. We calculate the weights of each level of indicators through a multi-level indicator system and the information weighting method, and construct a scoring card by combining a logistic regression algorithm to comprehensively evaluate the enterprise's performance capabilities.
It enables a comprehensive and accurate assessment of a company's ability to fulfill its obligations, improving the accuracy and reliability of the evaluation and supporting companies' decision-making in business activities.
Smart Images

Figure CN121961204A_ABST
Abstract
Description
A method and system for evaluating enterprise performance capabilities based on big data Technical Field
[0001] This specification relates to the field of enterprise performance capability evaluation technology, and more specifically, to an enterprise performance capability evaluation method and system based on big data. Background Technology
[0002] In a market economy, while pursuing maximum sales, businesses inevitably face sales risks, namely, the inability to guarantee cash or foreign exchange transactions for all sales, which carries significant risk. However, when all transactions are conducted in cash or foreign exchange, businesses tend to refuse credit sales, essentially forfeiting opportunities to expand their market. But sales on credit may result in defaults, leading to various negative consequences for the selling company, including economic losses, reputational damage, broken partnerships, and even cash flow disruptions. Therefore, evaluating a company's ability to fulfill its obligations plays a crucial role in business activities such as bidding for projects, corporate partnerships, government procurement, and access to preferential policies.
[0003] Currently, companies' ability to fulfill their obligations is generally assessed through subjective weighting or a single algorithm, or by simply evaluating a company's ability to fulfill its obligations based on a single aspect, such as financial statements. However, these methods cannot comprehensively and accurately reflect a company's true ability to fulfill its obligations. Summary of the Invention
[0004] The purpose of this specification is to provide a big data-based method for evaluating corporate performance capabilities, which can overcome the problem that existing corporate performance capability evaluation methods cannot comprehensively and truthfully reflect corporate performance capabilities.
[0005] The embodiments described in this specification are implemented as follows:
[0006] Firstly, this specification provides a big data-based method for evaluating enterprise performance capabilities, which mainly includes:
[0007] Based on the acquired enterprise data, an evaluation index system is constructed. The evaluation index system includes multi-level indicators, and the evaluation index system includes an evaluation sub-model corresponding to the top-level indicator.
[0008] The weights of lower-level indicators are calculated using the information content weighting method.
[0009] Based on the weights of the lower-level indicators, the weights of each of the top-level indicators are calculated level by level to obtain the weights of the top-level indicators.
[0010] Based on the enterprise data and the weights of each of the top-level indicators, the enterprise's performance capability score is calculated through each of the evaluation sub-models.
[0011] Secondly, this specification provides a big data-based enterprise performance capability evaluation system, which mainly includes:
[0012] The construction module is used to construct an evaluation index system based on the acquired enterprise data. The evaluation index system includes multi-level indicators, and the evaluation index system includes an evaluation sub-model corresponding to the top-level indicator.
[0013] The first calculation module is used to calculate the weight of the lower-level indicators using the information weighting method.
[0014] The second calculation module is used to calculate the weight of each of the top-level indicators level by level according to the weight of the lower-level indicators.
[0015] The scoring module is used to calculate the enterprise's performance capability score based on the enterprise data and the weights of each of the top-level indicators, through each of the evaluation sub-models.
[0016] Thirdly, this specification also provides an electronic device, which mainly includes:
[0017] Memory, used to store computer programs;
[0018] When the processor executes the program stored in memory, it implements the above-mentioned big data-based enterprise performance evaluation method.
[0019] The embodiments described in this specification have at least the following advantages or beneficial effects:
[0020] This method constructs an evaluation index system containing multiple multi-level indicators based on the types of enterprise data. It determines the score of each top-level indicator through the constructed evaluation sub-model, and then determines the weight of the top-level indicator by determining the weight level by level. The score of each first-level indicator and its corresponding weight are used to calculate the score of the enterprise's performance capability, which can comprehensively and truthfully reflect the enterprise's true performance capability. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this specification and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a flowchart illustrating the enterprise performance capability evaluation method based on big data provided in this specification;
[0023] Figure 2 is a schematic diagram of the enterprise performance capability evaluation system based on big data provided in this manual;
[0024] Figure 3 is a schematic diagram of the electronic device provided in this specification. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments in this specification clearer, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Generally, the components of the embodiments of this specification described and shown in the accompanying drawings can be arranged and designed in various different configurations.
[0026] Please refer to Figure 1. One embodiment of this specification provides a method for evaluating enterprise performance capabilities based on big data, mainly including:
[0027] Step 102: Based on the acquired enterprise data, construct an evaluation index system. The evaluation index system includes multi-level indicators, and the evaluation index system includes an evaluation sub-model corresponding to the top-level indicator.
[0028] Step 104: Calculate the weight of the lower-level indicators using the information content weighting method;
[0029] Step 106: Calculate the weight of each of the top-level indicators by calculating the weight of each lower-level indicator level by level.
[0030] Step 108: Based on the enterprise data and the weights of each of the top-level indicators, calculate the enterprise's performance capability score through each of the evaluation sub-models.
[0031] In this embodiment, the lower-level indicators are illustrated using the second-level indicators as an example. However, in other embodiments, levels can be added, such as third-level and fourth-level indicators. That is, a first-level indicator contains multiple second-level indicators, a second-level indicator contains multiple third-level indicators, and so on.
[0032] Specifically, this method constructs an evaluation index system containing multiple primary and secondary indicators based on the types of enterprise data. It determines the score of each primary indicator through the constructed evaluation sub-model, and then determines the weight of the primary indicators by determining the weights at each level. The score of each primary indicator and its corresponding weight are used to calculate the score of the enterprise's performance capability, which can comprehensively and truthfully reflect the enterprise's true performance capability.
[0033] In this embodiment, one implementation of step 102 is as follows:
[0034] Step 112: Obtain at least the company's business registration data, electricity data, annual report, tax data, invoice data, financial data, intellectual property data, industry data, and legal data;
[0035] Step 114: Construct an evaluation sub-model corresponding to the enterprise size index based at least on the enterprise's business registration data, power data, and annual report;
[0036] Step 116: Based at least on the company's financial statements, construct an evaluation sub-model corresponding to the financial status indicators;
[0037] Step 118: Construct an evaluation sub-model corresponding to the industry status index based at least on the enterprise's business registration data, power data, annual report data, and intellectual property data;
[0038] Step 120: Based at least on the industry data and invoice data of the enterprise, construct an evaluation sub-model corresponding to the supply chain quality indicators;
[0039] Step 122: Based at least on judicial data, construct evaluation sub-models corresponding to compliance operation indicators;
[0040] Step 124: The enterprise size indicators, financial status indicators, industry position indicators, supply chain quality indicators, and compliance operation indicators are all first-level indicators, that is, the highest-level indicators.
[0041] In this embodiment, the company is scored from at least five dimensions: company size, financial condition, industry position, supply chain quality, and compliance operation, in order to obtain a comprehensive score of the company's ability to fulfill its obligations.
[0042] In this embodiment, the secondary indicators belonging to the above-mentioned primary indicators are specifically shown in the table below:
[0043] Table 1. Secondary Indicators and Data Sources
[0044]
[0045]
[0046]
[0047] In this embodiment, the secondary indicators to which each of the primary indicators belongs can be further supplemented. Similarly, parts can be deleted or replaced according to the enterprise's data type, and all can be used to calculate the weight and score of the primary indicators level by level. It is evident that by setting the above-mentioned secondary and primary indicators, the enterprise's performance capability can be evaluated from multiple dimensions, so that the score can more comprehensively reflect the enterprise's performance capability.
[0048] In this embodiment, a risk database for the enterprise can also be established based on the acquired enterprise data to facilitate the retrieval of various data and to monitor changes in risk.
[0049] In this embodiment, one implementation of step 120 is as follows:
[0050] Step 132: Based on the industry data and the invoice data, construct the optimal logistic regression model using a logistic regression algorithm;
[0051] Step 134: Based on the optimal logistic regression model, create a scorecard, which is the evaluation sub-model corresponding to the supply chain quality indicators. The scorecard can reflect the score of the enterprise's supply quality indicators.
[0052] In this embodiment, one implementation of step 132 is as follows:
[0053] Step 142: Binning is performed based on the features in the industry data and the features in the invoice data. The optimal binning and binning boundary for each feature are determined. Each bin corresponding to a feature has a WOE value.
[0054] Step 144: Replace the features with their corresponding WOE values, train the model using a logistic regression algorithm, and obtain the optimal logistic regression model.
[0055] Specifically, based on the existing basic features in the enterprise data, derived features are obtained. After extracting the above derived features, the sample data (i.e., enterprise data) is modeled, that is, the above optimal logistic regression model is constructed. This model can reflect the feature variables that are highly correlated with default risk.
[0056] In this embodiment, before building the model, the aforementioned industry data and invoice data are processed, namely, data processing such as removing duplicate values and filling in missing values.
[0057] In this embodiment, the downstream and upstream suppliers of the company to be evaluated can be determined by the aforementioned industry data and invoice data, as well as the order amount and other information, thereby determining the quality of the company's supply chain. Based on the various secondary indicators, the score of the supply chain quality indicators can be determined.
[0058] In this embodiment, the aforementioned industry data and invoice data are divided into a training set and a test set. The training set is used to train the logistic regression model to obtain the optimal logistic regression model, while the test set can verify the effect of the optimal logistic regression model.
[0059] In this embodiment, the features in the training set and test set are represented in the form of a feature matrix, and the values in the feature matrix are replaced with the WOE values of each bin.
[0060] In this embodiment, one specific implementation of step 142 is as follows:
[0061] Step 152: Determine the preset number of boxes;
[0062] Step 154: Based on the preset number of boxes, perform equal-frequency boxing based on the features in the industry data and the features in the invoice data, and calculate the WOE value of each box and the IV value of each feature.
[0063] Step 156: Based on the calculated chi-square test values, merge similar bins and calculate the WOE value and IV value of each feature for each merged bin.
[0064] Step 158: Determine the optimal bin division and bin division boundary based on the number of bins after merging and the IV value of the characteristics calculated after merging the bins.
[0065] Specifically, WOE stands for Weight of Evidence, which is an encoding form of the original independent variable. It is used in logistic regression models, which transform the form of the original variable by grouping the variable (also known as columnar or binary) and then calculating the WOE value of each group, which helps the model better understand and fit the data.
[0066] In this embodiment, the calculation of the WOE value involves comparing the proportion of bad samples (e.g., defaulting users) and good samples (normal users) in each group, as well as the difference between these two proportions in the overall population. The WOE value reflects the logarithmic form of this difference; the larger it is, the greater the likelihood that the sample in that group is a bad sample, and vice versa.
[0067] In this embodiment, the method for calculating the WOE value is as follows:
[0068]
[0069] Among them, Bad i This represents the number of bad samples in the i-th bin. T This represents the total number of bad samples; similarly, Good... i This represents the number of good samples in the i-th bin. T This represents the total number of good samples.
[0070] In this embodiment, the full name of the above-mentioned IV value is Information Value, which can objectively and quantitatively evaluate the predictive ability of each variable when selecting input variables (input variables refer to features that enter the training of the logistic regression model).
[0071] The features mentioned above that are used to train the logistic regression model are specific features of the company’s multiple features. For example, based on the company’s industry data and invoice data, it can be concluded that the company has multiple features, including at least four features: number of suppliers, transaction amount of suppliers, number of customers, and transaction amount of customers. These four features are selected as specific features to enter the logistic regression model, which are the input variables.
[0072] In this embodiment, the logistic regression model is trained using data from multiple companies, therefore, the specific features mentioned above can be common features of multiple companies.
[0073] In this embodiment, by calculating the IV value of each input variable, the input variables with a greater impact on the model's prediction results can be identified, thus determining the input variables to be included in the model. That is, the magnitude of the IV value reflects the predictive power of the input variable on the target variable (the target variable refers to the dependent variable, i.e., the target of the model's learning, which is whether the company defaults). The larger the value, the greater the contribution of that variable to the prediction results of the logistic regression model. Therefore, in the model construction process, input variables with higher IV values are usually prioritized.
[0074] In this embodiment, the method for calculating the IV value is as follows:
[0075]
[0076] Among them, Bad i This represents the number of bad samples in the i-th bin. T This represents the total number of bad samples; similarly, Good... i This represents the number of good samples in the i-th bin. T This represents the total number of good samples, and n represents the total number of bins.
[0077] In this embodiment, one specific implementation of step 134 is as follows:
[0078] Step 162: Construct a prediction function based on the binary logistic regression function;
[0079] Step 164: Based on the prediction function, obtain the formula for calculating the logarithmic probability of the event, where the event is a corporate default event;
[0080] Step 166: Based on the preset default value and preset score, substitute them into the logarithmic probability calculation formula of the event to obtain the coefficient value of the scorecard;
[0081] Step 168: Create the scorecard based on the coefficient values of the scorecard.
[0082] In this embodiment, the scoring card is determined as follows:
[0083] Score = AB * log(Odds)
[0084] Where A and B are constants, and Odds represents the probability of an event occurring, which is the ratio of the probability of the event occurring to the probability of it not occurring. If the probability of a customer defaulting is p, then Odds = p / (1-p).
[0085] Since the log function is monotonically increasing from 0 to +∞, the higher the customer default probability (Odds), the lower the score.
[0086] By proceeding through steps 162 to 166, the aforementioned constants A and B are determined as follows:
[0087] The prediction function is constructed as follows:
[0088]
[0089] Among them, h θ (χ) represents the probability that the sample belongs to the positive class, where the positive class refers to the company defaulting. In a binary classification problem, this can be represented as the probability of "1". The vector θ represents the parameters (also called weights or coefficients) of the logistic regression model. T Represents the vector θ(θ1, θ2, ..., θ) n The transpose of ) is the vector χ, where χ is (x0, x1, ..., x2). n ), x0 = 1, and the vector χ represents the WOE value of each input variable.
[0090] The derived formula for calculating the logarithmic probability of the event is as follows:
[0091]
[0092] Where Odds represents the probability of the event, and p represents the probability of the customer defaulting on the event.
[0093] Based on the logarithmic odds of the above events, in the logistic regression model, the logarithmic odds of the output Y=1 (representing corporate default) is a linear function of the input conditions (i.e., the input features). Therefore, we can obtain:
[0094] ln(Odds) = θ0 + θ1x1 + ... + θ n x n
[0095] In this embodiment, based on the preset θ0 (initial default probability), P0 (initial score, a preset score value. It represents the baseline score under a specific business condition, such as the base score for a specific default rate θ0. This score value will serve as the benchmark point for the scorecard score and will be used to calculate the specific score under different default probabilities), and PDO (doubling factor, PDO reflects the scorecard's sensitivity to changes in default probability; that is, for every doubling of the default probability, the scorecard score will decrease by the number of points specified by PDO), that is, the score corresponding to Odds(θ0) is P0, and the score corresponding to the doubling Odds(2θ0) is P0 + PDO, then:
[0096] P0 = AB * log(θ0)
[0097] P0 + PDO = AB * log(2θ0)
[0098] From the above two equations, we can obtain:
[0099] A = P0 + B * log(θ0)
[0100]
[0101] Based on A and B above, the scoring card can be determined as follows:
[0102]
[0103] Where β0, β1, ..., β n The independent variable coefficients, WOE, are obtained from training a logistic regression model. 1i This indicates that the value of the first feature falls within the WOE value corresponding to the i-th bin. WOE nm This indicates that the value of the nth feature falls within the WOE value corresponding to the mth bin.
[0104] As can be seen, the above method can accurately calculate the score of the supply chain quality indicator, so that the score of this primary indicator can truly reflect the quality of the enterprise's supply chain.
[0105] In this embodiment, for other primary indicators, data processing is performed on their respective secondary indicators, namely, outlier removal, data standardization, and equal-frequency binning based on the data and characteristics of the secondary indicators.
[0106] In this embodiment, taking the minimum classification index as a secondary index as an example, the data standardization process is as follows:
[0107] The method for handling positive indicators is as follows:
[0108]
[0109] The method for handling negative indicators is as follows:
[0110]
[0111] Where, x pq Y represents the original value of the q-th evaluation indicator for the p-th evaluation object. pq This represents the value of the standardized evaluation index.
[0112] In this embodiment, the evaluation objects mentioned above are relative. That is, for the "enterprise size" sub-model, enterprise size is the evaluation object, and registered capital and paid-in capital are the evaluation indicators; if the evaluation is of an enterprise, that is, the enterprise is the evaluation object, and the evaluation indicators are primary indicators such as enterprise size and financial status.
[0113] In this embodiment, the aforementioned positive indicators refer to those whose values increase as the phenomenon under study becomes better or more severe. For example, an increase in corporate sales reflects a good business situation and increased market demand; the higher the better. The aforementioned negative indicators refer to those whose values increase as the phenomenon under study becomes more severe or unfavorable. For example, the level of electricity consumption reflects the efficiency of energy utilization; the lower the better.
[0114] The specific method for calculating the weights is as follows:
[0115] Step 172: Calculate the average value of each evaluation indicator: This represents the average value of the q-th evaluation indicator.
[0116] Step 174: Calculate the standard deviation of each evaluation indicator: S q Let q represent the standard deviation of the q-th evaluation indicator.
[0117] Step 176: Calculate the coefficient of variation for each evaluation indicator: Let q represent the q-th evaluation indicator.
[0118] Step 178: Calculate the weights of each evaluation indicator: q represents the q-th evaluation indicator.
[0119] Step 180: Through the above calculations, the weights of each level of evaluation indicators can be obtained. Using the additivity principle of the coefficient of variation, the weights of higher-level indicators are calculated step by step; that is, the weight of a first-level indicator is equal to the sum of the weights of the second-level indicators. In other embodiments, third-level and / or fourth-level indicators can also be set to comprehensively and thoroughly evaluate the enterprise's performance capabilities.
[0120] Step 182: Calculate the overall score
[0121] In this embodiment, the comprehensive score of each sub-model (such as enterprise size, financial status, industry position, and compliance operation) is calculated, and then the comprehensive score of the enterprise is calculated based on the comprehensive score of each sub-model.
[0122] Please refer to Figure 2. Another embodiment of this specification provides a big data-based enterprise performance capability evaluation system, which mainly includes:
[0123] The construction module 202 is used to construct an evaluation index system based on the acquired enterprise data. The evaluation index system includes multiple primary indicators and multiple secondary indicators, and the multiple secondary indicators belong to the primary indicators. The evaluation index system includes evaluation sub-models constructed corresponding to the primary indicators.
[0124] The first calculation module 204 is used to calculate the weight of the secondary indicator using the information weighting method;
[0125] The second calculation module 206 is used to calculate the weight of each of the primary indicators based on the weight of the secondary indicators;
[0126] The scoring module 208 is used to calculate the enterprise's performance capability score based on the enterprise data and the weights of each of the primary indicators, through each of the evaluation sub-models.
[0127] Specifically, the system constructs an evaluation index system containing multiple primary and secondary indicators based on the types of enterprise data. It determines the score of each primary indicator through the constructed evaluation sub-model, and then determines the weight of the primary indicators by determining the weights at each level. The score of each primary indicator and its corresponding weight are used to calculate the score of the enterprise's performance capability, which can comprehensively and truthfully reflect the enterprise's true performance capability.
[0128] In this embodiment, the construction module 202 is used to acquire at least the enterprise's business registration data, electricity data, annual report, tax data, invoice data, financial data, intellectual property data, industry data, and judicial data; to construct an evaluation sub-model corresponding to the enterprise size indicator based at least on the enterprise's business registration data, electricity data, and annual report; to construct an evaluation sub-model corresponding to the financial status indicator based at least on the enterprise's financial statements; to construct an evaluation sub-model corresponding to the industry position indicator based at least on the enterprise's business registration data, electricity data, annual report, and intellectual property data; to construct an evaluation sub-model corresponding to the supply chain quality indicator based at least on the enterprise's industry data and invoice data; and to construct an evaluation sub-model corresponding to the compliance operation indicator based at least on the judicial data. The enterprise size indicator, financial status indicator, industry position indicator, supply chain quality indicator, and compliance operation indicator are all primary indicators. Based on the industry data and invoice data, an optimal logistic regression model is constructed using a logistic regression algorithm; based on the optimal logistic regression model, a scorecard is created, which is the evaluation sub-model corresponding to the supply chain quality indicator. The scorecard reflects the score of the enterprise's supply quality indicator. The data is binned based on features from the industry data and features from the invoice data. The optimal binning and binning boundaries for each feature are determined, and each bin corresponding to a feature has a Word of Entity (WOE) value. The features are replaced with their corresponding WOE values, and a logistic regression algorithm is used for training to obtain the optimal logistic regression model. A preset number of bins is determined. Based on the preset number of bins, the features from the industry data and features from the invoice data are binned using equal frequency binning. The WOE value of each bin and the IV value of each feature are calculated. Similar bins are merged based on the calculated chi-square test values, and the WOE value of each merged bin and the IV value of each feature are calculated. Based on the number of merged bins and the IV values of the features calculated after merging, the optimal binning and binning boundaries are determined. A prediction function is constructed based on the binary logistic regression function. Based on the prediction function, a formula for calculating the log probability of an event, where the event is a corporate default event, is obtained. Preset default values and preset score values are substituted into the log probability calculation formula of the event to obtain the coefficient values of the scorecard. The scorecard is created based on the coefficient values of the scorecard. The above methods enable a true assessment of supply chain quality, thereby providing a more comprehensive and accurate evaluation of a company's ability to fulfill its obligations.
[0129] Referring to Figure 3, another embodiment of the present invention provides an electronic device, including:
[0130] Memory, used to store computer programs;
[0131] The processor implements the above method when executing programs stored in memory.
[0132] This electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it, forming a big data-based enterprise performance evaluation device at the logical level. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0133] Network interfaces, processors, and memory can be interconnected via a bus system. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. These buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only a single bidirectional arrow is used in the diagram, but this does not imply that there is only one bus or one type of bus.
[0134] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include read-only memory and random-access memory, and provides instructions and data to the processor. Memory may include random-access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0135] Memory is used to store computer programs and execute them.
[0136] Step 102: Based on the acquired enterprise data, construct an evaluation index system. The evaluation index system includes multi-level indicators, and the evaluation index system includes an evaluation sub-model corresponding to the top-level indicator.
[0137] Step 104: Calculate the weight of the lower-level indicators using the information content weighting method;
[0138] Step 106: Calculate the weight of each of the top-level indicators by calculating the weight of each lower-level indicator level by level.
[0139] Step 108: Based on the enterprise data and the weights of each of the top-level indicators, calculate the enterprise's performance capability score through each of the evaluation sub-models.
[0140] The methods for evaluating enterprise performance capabilities based on big data, as disclosed in the embodiments shown in the figures of this specification, can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions formed by software. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this specification can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0141] Based on the same invention, another embodiment of this specification provides a computer-readable storage medium storing one or more programs that, when executed by an electronic device including multiple applications, cause the electronic device to perform the big data-based enterprise performance evaluation method provided in the embodiment corresponding to FIG1.
[0142] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0143] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0144] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0145] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to create a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one block of the block diagram.
[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0148] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0149] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0150] The above description is merely an embodiment of this application and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.
Claims
1. A method for evaluating enterprise performance capabilities based on big data, characterized in that, include: Based on the acquired enterprise data, an evaluation index system is constructed. This system comprises multi-level indicators, each with a corresponding evaluation sub-model. Each sub-model reflects the enterprise's situation across various dimensions. The weights of lower-level indicators are calculated using an information weighting method. Based on the weights of the lower-level indicators, the weights of each upper-level indicator are calculated level by level. Finally, based on the enterprise data and the weights of each upper-level indicator, the enterprise's performance capability score is calculated using each evaluation sub-model.
2. The enterprise performance capability evaluation method based on big data according to claim 1, characterized in that, The process of constructing an evaluation indicator system based on the acquired enterprise data includes: acquiring at least the enterprise's business registration data, electricity data, annual report data, tax data, invoice data, financial data, intellectual property data, industry data, and judicial data; constructing an evaluation sub-model corresponding to the enterprise size indicator based at least on the enterprise's business registration data, electricity data, and annual report data; constructing an evaluation sub-model corresponding to the financial status indicator based at least on the enterprise's financial statements; constructing an evaluation sub-model corresponding to the industry position indicator based at least on the enterprise's business registration data, electricity data, annual report data, and intellectual property data; constructing an evaluation sub-model corresponding to the supply chain quality indicator based at least on the enterprise's industry data and invoice data; and constructing an evaluation sub-model corresponding to the compliance operation indicator based at least on the judicial data. The enterprise size indicator, financial status indicator, industry position indicator, supply chain quality indicator, and compliance operation indicator are all first-level indicators, i.e., the highest-level indicators.
3. The enterprise performance capability evaluation method based on big data according to claim 2, characterized in that, The step of constructing an evaluation sub-model corresponding to the supply chain quality indicators based at least on the company's industry data and invoice data includes: constructing an optimal logistic regression model based on the industry data and invoice data using a logistic regression algorithm; and creating a scorecard based on the optimal logistic regression model, which is the evaluation sub-model corresponding to the supply chain quality indicators, and the scorecard can reflect the score of the company's supply chain quality indicators.
4. The enterprise performance capability evaluation method based on big data according to claim 3, characterized in that, The step of constructing an optimal logistic regression model based on the industry data and the invoice data using a logistic regression algorithm includes: binning the features in the industry data and the invoice data, determining the optimal binning and binning boundaries for each feature, with each bin corresponding to a feature having a WOE value; replacing the features with their corresponding WOE values, and training the model using a logistic regression algorithm to obtain the optimal logistic regression model.
5. The enterprise performance capability evaluation method based on big data according to claim 3, characterized in that, The step of creating a scorecard based on the optimal logistic regression model includes: constructing a prediction function based on a binary logistic regression function, wherein the prediction function reflects whether the enterprise has defaulted; obtaining a formula for calculating the logarithmic probability of an event, wherein the event is an enterprise default event, based on the prediction function; substituting a preset default value and a preset score value into the formula for calculating the logarithmic probability of the event to obtain the coefficient value of the scorecard; and creating the scorecard based on the coefficient value of the scorecard.
6. The enterprise performance capability evaluation method based on big data according to claim 5, characterized in that, The prediction function is: Among them, h θ (χ) represents the probability of the result being 1, and the vector θ represents the parameters of the logistic regression model. T Represents the vector θ(θ1, θ2, ..., θ) n The | transpose of ) is the vector χ, where (x0, x1, ..., x2) n x0 = 1, and the vector χ represents the WOE value of each feature.
7. The enterprise performance capability evaluation method based on big data according to claim 6, characterized in that, The formula for calculating the logarithmic probability of the event is: Where Odds represents the probability of the event, and p represents the probability of the company defaulting.
8. The method for evaluating enterprise performance capabilities based on big data according to claim 4, characterized in that, The process of binning based on features in the industry data and features in the invoice data, and determining the optimal binning and binning boundaries for each feature, includes: determining a preset number of bins; performing equal-frequency binning on the features in the industry data and features in the invoice data based on the preset number of bins, and calculating the WOE value of each bin and the IV value of each feature; merging similar bins based on the calculated chi-square test values, and calculating the WOE value of each merged bin and the IV value of each feature; and determining the optimal binning and binning boundaries based on the number of merged bins and the IV values of the features calculated after merging the bins.
9. A big data-based enterprise performance capability evaluation system, characterized in that, include: The construction module is used to construct an evaluation index system based on the acquired enterprise data. The evaluation index system includes multi-level indicators, and the evaluation index system includes an evaluation sub-model corresponding to the top-level indicator. The first calculation module is used to calculate the weight of the lower-level indicators using the information weighting method. The second calculation module is used to calculate the weight of each of the top-level indicators level by level according to the weight of the lower-level indicators. The scoring module is used to calculate the enterprise's performance capability score based on the enterprise data and the weights of each of the top-level indicators, through each of the evaluation sub-models.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-8.