Enterprise credit automatic rating method and device, computer equipment and storage medium
By protecting the privacy and classifying the features of enterprise data, and combining industry-specific mapping relationships and multivariate mapping functions, the problems of data acquisition and quality in enterprise credit rating are solved, and a more comprehensive and accurate credit assessment is achieved.
Patent Information
- Application Number
- CN202511330758.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies for corporate credit rating face challenges such as difficulty in data acquisition, unstable data quality, poor model interpretability, and complexity and accuracy issues in the processing.
By acquiring enterprise data, performing privacy protection and anonymization processing, screening key credit features and classifying them, and combining industry-specific mapping relationships and multivariate mapping functions, a credit rating model is constructed, and ratings are performed using linear, logarithmic, and exponential functions.
It achieves more comprehensive and accurate credit assessment, improves the industry adaptability and differentiation of ratings, and enhances the interpretability of the model and the accuracy of rating results.
Smart Images

Figure CN121504587A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computers, and more specifically to methods, apparatus, computer equipment, and storage media for automatic enterprise credit rating. Background Technology
[0002] In the process of corporate credit rating prediction, current technologies mainly rely on machine learning, big data analytics, and knowledge graphs. These technologies have been widely used in financial institutions for credit approval, risk assessment, and supplier selection, and have achieved certain results in improving the efficiency and accuracy of credit rating, but they also have some limitations and shortcomings.
[0003] First, machine learning techniques predict a company's credit rating by training classification or regression models using large amounts of historical data. Common models include logistic regression, decision trees, random forests, and neural networks. These models take various feature data of the company as input and output credit rating results based on the parameters obtained from training. However, some data is difficult to obtain due to reasons such as company privacy or industry confidentiality, resulting in incomplete data. In addition, the lag in data updates can also affect the accuracy of the model. At the same time, machine learning models are very sensitive to the distribution and features of the data. If the data is biased or the feature selection is inappropriate, it will lead to overfitting or underfitting, affecting the reliability of the prediction.
[0004] Secondly, big data analytics technology collects, stores, and processes massive amounts of enterprise-related data, using statistical analysis and data mining to extract valuable information and assess a company's creditworthiness. Credit rating agencies typically integrate data from multiple data sources to analyze and identify potential risks and credit trends. However, acquiring big data faces privacy and security challenges; some data is restricted by laws and regulations and cannot be accessed. Furthermore, inconsistent data quality, noise, and erroneous data can affect the accuracy of the analysis. Moreover, the sheer volume and complexity of the data can lead to information overload, thus impacting the effectiveness of the analysis.
[0005] Finally, knowledge graph technology constructs a network of relationships between entities such as enterprises, people, and events, forming a semantic knowledge system that links various types of enterprise information. Its applications include construction, querying, and reasoning. In financial regulation, knowledge graphs can be used to identify related-party transactions and potential risks between enterprises. However, the construction of knowledge graphs requires the participation of domain experts, and manual annotation is costly. At the same time, updating and maintaining knowledge graphs also requires significant resources, and the relationships and attributes within the graph may be uncertain, affecting the accuracy of reasoning results. Furthermore, the complexity and scale of the knowledge graph also impact the efficiency of reasoning and the accuracy of results.
[0006] While these technologies have achieved some success in corporate credit rating, they still face many challenges and limitations, such as difficulties in data acquisition, data quality assurance, processing complexity, and the accuracy of analysis results. Furthermore, practical applications also face challenges in interpretability and generalization capabilities.
[0007] Therefore, it is necessary to design a new method to achieve a more comprehensive and accurate credit rating system, and to solve the problems of difficulty in data acquisition, unstable data quality, poor model interpretability, and complexity and accuracy of the processing in existing technologies. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, apparatus, computer equipment and storage medium for automatic enterprise credit rating.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: an automatic enterprise credit rating method, comprising: Obtain data on companies to be rated; The data of the companies to be rated is input into the credit rating model to calculate the credit score and obtain the calculation result; The industry to which the enterprise belongs is determined based on the data of the enterprise to be rated; The mapping relationship is determined based on the industry to which the enterprise belongs; The credit rating of the enterprise is determined based on the calculation results and the mapping relationship. Output the credit rating of the company.
[0010] The further technical solution is as follows: the training process of the credit rating model includes: Obtain relevant enterprise data; The relevant enterprise data is subjected to privacy protection and anonymization processing to obtain the processing result; Key credit features are selected from the processing results and then classified to obtain classification results. The classification results are weighted to obtain the allocation results; Quantiles are constructed based on the data distribution of the allocation results, and a mapping function is constructed based on the quantiles to obtain the credit rating model.
[0011] The further technical solution is as follows: the enterprise-related data includes basic enterprise information, operating conditions, financial status, and credit status.
[0012] The further technical solution is as follows: the step of filtering key credit features from the processing results and performing feature classification to obtain classification results includes: Key credit features are selected from the processing results. Based on the nature of the features, the key credit features are classified into continuous features and discrete features. The continuous features are assigned scores, and the discrete features are processed using binning technology to obtain the classification results.
[0013] The further technical solution is as follows: The weight allocation of the classification results to obtain the allocation result includes: Based on historical data analysis and expert experience, a percentage-based weighting system is used to divide the features in the classification results according to their importance, so as to assign a weight to each feature in the classification results and obtain the allocation result.
[0014] The further technical solution is as follows: The step of constructing quantiles based on the data distribution of the allocation results, and constructing a mapping function based on the quantiles to obtain a credit rating model, includes: The quantiles of the allocation results are calculated through statistical analysis, and initial scores and special value processing are set according to the quantiles. A mapping function suitable for the feature distribution is designed based on the quantiles to obtain the credit rating model.
[0015] The further technical solution is as follows: the mapping function includes linear functions, logarithmic functions and exponential functions.
[0016] The present invention also provides an automatic enterprise credit rating device, comprising: The data acquisition unit is used to acquire data on the companies to be rated. The calculation unit is used to input the data of the enterprise to be rated into the credit rating model to calculate the credit score and obtain the calculation result; The industry determination unit is used to determine the industry to which the enterprise belongs based on the data of the enterprise to be rated. A mapping relationship determination unit is used to determine a mapping relationship based on the industry to which the enterprise belongs; A rating determination unit is used to determine the credit rating of an enterprise based on the calculation results and the mapping relationship; The output unit is used to output the credit rating of the enterprise.
[0017] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0018] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0019] The beneficial effects of this invention compared to existing technologies are as follows: By inputting the data of the companies to be rated into the credit rating model and determining specific mapping relationships in conjunction with industry classifications, this invention ensures the comprehensiveness and accuracy of the rating process. First, it accurately acquires and cleans company data, solving the shortcomings of existing technologies in data acquisition and quality. Second, through industry-specific mapping relationships, it effectively improves the industry adaptability and differentiation of the rating, enhancing the interpretability of the model. Finally, by combining multivariate mapping functions and dynamic weight adjustment mechanisms, it ensures the accuracy and stability of the rating results, simplifies the rating process, solves the complexity and accuracy problems of traditional technologies, and achieves a more comprehensive and accurate credit assessment.
[0020] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram illustrating an application scenario of the automatic enterprise credit rating method provided in this embodiment of the invention. Figure 2 A flowchart illustrating the automatic enterprise credit rating method provided in an embodiment of the present invention; Figure 3 A schematic diagram of a sub-process of the automatic enterprise credit rating method provided in an embodiment of the present invention; Figure 4 A schematic block diagram of an automatic enterprise credit rating device provided in an embodiment of the present invention; Figure 5 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0025] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0026] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the automatic enterprise credit rating method provided in an embodiment of the present invention. Figure 2 This is a schematic flowchart illustrating the automatic enterprise credit rating method provided in this embodiment of the invention. The automatic enterprise credit rating method is applied in a server. The server interacts with the terminal, acquiring relevant data of the enterprise to be rated and performing calculations based on a credit rating model. This effectively solves problems such as difficulty in data acquisition, unstable data quality, and poor model interpretability in existing technologies. The method first performs privacy protection and anonymization processing on the enterprise data, filters and classifies key credit features, then assigns weights to the features and constructs quantiles and mapping functions to achieve accurate credit rating. By employing linear, logarithmic, and exponential mapping functions, the accuracy and interpretability of the rating model are further improved. Furthermore, this method, through industry mapping relationships combined with expert experience and historical data analysis, achieves a more comprehensive and stable rating system, effectively addressing the challenges of data processing complexity and accuracy.
[0028] Figure 2 This is a flowchart illustrating the automatic enterprise credit rating method provided in an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S150.
[0029] S110, Obtain data on companies to be rated.
[0030] In this embodiment, before conducting enterprise credit rating, it is necessary to collect relevant data on the enterprise to be rated. This data includes, but is not limited to, the enterprise's basic information, operating status, financial status, and creditworthiness. This step is mainly achieved through the following operations: Data Sources: A company's data can come from multiple sources, including internal systems, third-party data providers, and public databases. This data covers basic company information, financial statements, sales records, tax records, court cases, and administrative penalty records.
[0031] Data integration and cleaning: Through data preprocessing techniques, duplicate data is removed, formats are standardized, and abnormal data is eliminated to ensure data quality and provide clean and high-quality data for subsequent credit scoring models.
[0032] Data anonymization and privacy protection: To ensure the privacy and security of corporate data, sensitive data (such as tax ID, company name, etc.) is anonymized, and sensitive information is hashed using encryption algorithms such as SHA-256 to ensure the irreversibility of data.
[0033] S120. Input the data of the enterprise to be rated into the credit rating model to calculate the credit score and obtain the calculation result.
[0034] In this embodiment, the calculation result refers to the credit score of the company to be rated.
[0035] Specifically, in this stage, the cleaned and anonymized enterprise data will be used as input to the credit rating model for further credit score calculation. The specific implementation process is as follows: Based on the requirements of the credit rating model, the input enterprise data will first undergo feature selection and feature engineering. This process selects 28 key features that are most representative and predictive.
[0036] For key features, first classify and process them accordingly: Continuous features, such as sales revenue and debt-to-equity ratio, will be processed using the quantile method. The quantiles of each feature will be calculated, and the data will be mapped to a score interval.
[0037] Discrete characteristics, such as whether the enterprise is a technology company or whether it has court case records, are processed using binning technology.
[0038] Based on the importance of each feature, different weights are assigned to each feature to ensure that each feature reflects its importance to the company's credit in the final score.
[0039] Based on the input enterprise characteristics and corresponding quantiles, combined with the set weights, the credit scoring model is used to calculate and finally output the enterprise's credit score.
[0040] In one embodiment, such as Figure 3 As shown, the training process of the credit rating model described above includes steps S121 to S125.
[0041] S121. Obtain relevant enterprise data.
[0042] In this embodiment, the enterprise-related data includes basic enterprise information, operating conditions, financial status, and credit status.
[0043] First, obtain relevant enterprise data from internal company systems, public platforms, and other authorized channels. This data includes, but is not limited to: Basic company information: such as company name, date of establishment, registered capital, legal representative, registered address, etc.
[0044] Operating performance: This includes the company's main business, historical sales, market share, industry position, and whether it is involved in multinational business.
[0045] Financial status: such as various financial data in the income statement, balance sheet, and cash flow statement, including but not limited to revenue, net profit, debt ratio, and total assets.
[0046] Credit status: The company's tax records, court case records, whether there have been administrative penalties, and whether there has been any bad credit history.
[0047] Data deduplication is a crucial step in ensuring the uniqueness of each piece of information. During data collection, the same data may be recorded multiple times. To avoid duplicate calculations, deduplication is essential. For example: Basic company information: The same company may appear in different data sources, so it is essential to ensure that unique identifiers such as company name and unified social credit code are unique.
[0048] Financial data: Financial data at the same point in time may be entered multiple times, and duplicates need to be removed to ensure that the same data point is not counted twice in the analysis.
[0049] The purpose of data cleaning is to remove erroneous, redundant, and inconsistent data. This process includes, but is not limited to: Missing value handling: In the actual collected data, there may be some fields that are empty (e.g., some companies did not provide complete financial reports). Missing values can be handled by imputing them (e.g., using the mean, median, or values from the previous period) or by removing samples containing a large number of missing values.
[0050] Outlier detection: Data may contain extreme outliers, such as an unusually high sales volume for a company during a particular period. These outliers may be due to data entry errors or other reasons. Statistical analysis or machine learning methods are needed to identify and handle these outliers to prevent them from interfering with model training.
[0051] Data format standardization: Different data sources may use different formats (e.g., date formats may be "YYYY / MM / DD" or "DD-MM-YYYY"). In order to process data uniformly, these formats must be standardized to a standard format (e.g., "YYYY-MM-DD") to avoid errors in subsequent analysis caused by inconsistent formats.
[0052] During data cleaning, it is also necessary to remove some unnecessary or irrelevant data. For example: Irrelevant data: Business information of some companies may not meet the needs of credit scoring analysis, and this data can be removed.
[0053] Low-quality data: For example, some data may lack context or have no analytical significance and may need to be removed.
[0054] After unifying and formatting the data, it is often necessary to standardize and normalize it. This is because the data volume can vary significantly depending on the characteristics. For example, some financial data (such as revenue and total assets) might be in the millions, while other characteristics (such as the company's years of operation) might only be single digits. If these characteristics are not standardized or normalized, they may bias subsequent models. Common methods include: Standardization: This involves subtracting the mean from the data and then dividing by the standard deviation, resulting in data with a zero mean and a unit standard deviation.
[0055] Normalization: scaling data proportionally so that it falls into a uniform range (such as [0,1] or [-1,1]).
[0056] Data consistency checks ensure that data from different data sources and at different points in time can be correctly compared and analyzed. For example, whether total assets and total liabilities in financial data are consistent, or whether there are discrepancies between tax records and court records. In this process, developers typically write validation rules to automatically check for inconsistencies in the data.
[0057] Finally, data from different sources are integrated and correlated as necessary. For example, it may be necessary to merge basic information, financial data, and operating conditions of the same company using its unique identifier (such as the unified social credit code). In this way, it is ensured that each piece of company data contains multi-dimensional information, forming a complete company profile.
[0058] After data processing is complete, the final step is to store the cleaned, deduplicated, and formatted data in the database and set up a regular update mechanism. As time goes on, a company's operating conditions, financial situation, and creditworthiness will change; therefore, it is essential to ensure that the data is continuously updated and supplemented to guarantee the timeliness of the scoring model.
[0059] This series of processing steps ensures the quality and reliability of the data, providing high-quality input data for subsequent credit scoring models, reducing errors caused by data quality issues, and thus improving the accuracy and reliability of the scoring results.
[0060] S122. Perform privacy protection and desensitization processing on the enterprise-related data to obtain the processing result.
[0061] In this embodiment, the processing result refers to the enterprise-related data after privacy protection and de-identification processing.
[0062] Specifically, first, it's necessary to identify which data constitutes sensitive information. Among a company's various data, key data related to the company's identity typically includes: Enterprise tax identification number (such as unified social credit code): This information is a unique identifier for an enterprise in the government and other regulatory agencies. Disclosure may lead to the impersonation of the enterprise or fraud.
[0063] Company Name: A company name is usually a unique identifier for a company. Disclosure or alteration of a company name may lead to confusion or malicious use.
[0064] Company registered address, legal representative, etc.: This information can be directly or indirectly linked to the company's identity. If misused, it may bring legal and security risks to the company.
[0065] To protect this sensitive information, hash encoding (also called hash encryption) is used during data storage and processing to convert the original data into an irreversible encrypted form. Key characteristics of hash encoding include: Irreversibility: Hash algorithms are one-way, meaning that once data is hashed, the original data cannot be recovered from the hash result. This implies that even if data is leaked, hackers cannot reconstruct the specific sensitive information.
[0066] Data uniqueness: Different input data processed by the same hash algorithm will generate a unique output (i.e., a "hash value"). Even a tiny change (such as a single different character in a company name) will produce a completely different hash value. Therefore, a hash value can be used to uniquely identify a company, but it cannot be used to deduce specific information.
[0067] SHA-256 (Secure Hash Algorithm 256-bit) is a widely used encryption algorithm, belonging to the SHA family of algorithms. It has the following characteristics: Fixed output length: Regardless of the size of the input data, SHA-256 always generates a 256-bit (32-byte) hash value, and the length of the hash value is fixed, which is beneficial for data storage and management.
[0068] High collision resistance: A collision occurs when two different input data produce the same hash value. The SHA-256 algorithm has strong collision resistance; with current computing power, it is almost impossible for different data to generate the same hash value, thus ensuring the uniqueness of the data.
[0069] Efficiency and security: The SHA-256 algorithm is highly efficient and secure in the calculation process, making it suitable for hash calculations of large-scale data.
[0070] Therefore, using encryption algorithms such as SHA-256 to hash sensitive data can ensure data security while avoiding the risks of leaking the original information.
[0071] First, collect sensitive data related to the company's identity (such as the company's tax ID and name). Then, input the sensitive data (such as the company's tax ID and name) into the SHA-256 algorithm to generate its corresponding hash value.
[0072] For example, the company tax number "1234567890" will generate a fixed-length hash value after being processed by the SHA-256 algorithm: e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855.
[0073] This means that the original tax ID number "1234567890" has been converted into an irreversible hash value, making it impossible to recover the original tax ID information from the hash value.
[0074] The generated hash value is stored in the database and used as a unique identifier for the enterprise. In subsequent queries, comparisons, or matching processes, the hash value is used directly for comparison, avoiding the exposure of the original data.
[0075] Although hash values are irreversible, the values generated through hash encoding are still unique and distinguishable, thus ensuring that different companies can be correctly distinguished when querying the hash values of the same company.
[0076] Sensitive data such as company tax identification numbers and names, after being encrypted using SHA-256, are no longer stored in plaintext, greatly enhancing data security and preventing direct damage to the company in the event of a data breach. Since the hash value is generated from the original data, any changes to the data will also change the hash value, ensuring data integrity. The system can detect inconsistencies when any data is tampered with, allowing for necessary measures to be taken. Even if a data breach occurs during storage or transmission, because the data exists in hash form, external attackers cannot reconstruct the original sensitive information, avoiding the risk of misuse of corporate identity information.
[0077] By using encryption algorithms such as SHA-256 to hash-encode sensitive information such as company tax identification numbers and company names, not only is the irreversibility and uniqueness of the data guaranteed, but the risks of data leakage are also effectively avoided. Enterprises can still perform effective identity verification and data comparison while ensuring information security, thus enhancing data protection and ensuring the smooth operation of business.
[0078] S123. Select key credit features from the processing results and perform feature classification to obtain classification results.
[0079] In this embodiment, the classification result refers to the result obtained after classifying the key information features and performing corresponding processing.
[0080] Specifically, key credit features are selected from the processing results, and the key credit features are classified into continuous features and discrete features according to their properties. The continuous features are assigned scores, and the discrete features are processed using binning technology to obtain the classification results.
[0081] In corporate credit assessment, the number of corporate characteristics is often enormous, potentially reaching hundreds or even more. However, only a few dozen characteristics typically have a significant impact on corporate credit. Therefore, characteristic screening is a crucial step, involving the extraction of the most representative and predictive indicators from a massive dataset by combining the experience of business experts. The selected characteristics include: the company's establishment date, number of affiliated companies, changes in shareholders or legal representatives within the past two years, the most recent tax credit rating, sales amount over the past 12 months, year-on-year growth rate of sales over the past 12 months, tax-related violations and irregularities within the past 36 months, and court case records within the past 36 months. After screening and analysis by business experts, 28 key characteristics with a significant impact on corporate credit rating were ultimately identified, as shown in Table 1.
[0082] Table 1. Key Credit Features
[0083] When analyzing the characteristics in corporate credit assessment, the first step is to classify the characteristics through statistical analysis, mainly into two categories: continuous characteristics and discrete characteristics. This classification not only helps to understand the nature of each characteristic more accurately, but also enables different processing methods to be adopted for different types of characteristics, thereby ensuring the effectiveness and accuracy of the final scoring model.
[0084] Continuous features are variables that can take any value and are usually expressed numerically, such as sales amount, years of operation, and sales growth rate. Because continuous features have a wide and variable range of values, directly scoring these features may affect the stability and predictive ability of the model.
[0085] To better handle these features, quantile techniques are employed. Specifically, quantile techniques divide the values of a continuous feature into several intervals (such as one-quarter, fifty percentile, etc.) and assign a corresponding score to each interval. For example, the values of a feature can be sorted from low to high, and then, based on the data distribution, the data can be divided into several segments according to quantiles, assigning a score to each segment. The advantage of this method is that it can handle skewed data distributions, ensuring that different ranges of feature values do not affect the fairness of the scoring due to the existence of extreme values.
[0086] Discrete features are variables that can only take a limited number of values. They are usually categorical or classifiable data, such as a company's industry type, whether it has committed tax violations, or whether it is involved in litigation. Discrete features have a limited number of values and are usually represented by specific categories or labels.
[0087] For discrete features, binning is employed. The core idea of binning is to group different values of a discrete feature and assign a score to each group. Binning allows multiple similar or related categories to be merged into larger groups, simplifying the model's complexity. For example, for the discrete feature "tax-related violations," it can be divided into several categories based on the frequency or severity of the violations, with each category assigned a corresponding score. This approach reduces data noise and improves the stability and interpretability of the scoring model.
[0088] By employing different processing techniques for continuous and discrete features, feature standardization and comparability are ensured. This means that regardless of the original form of the features, all features can be compared and calculated using a uniform standard in the scoring model, thus avoiding the impact of differences in the dimensions of different feature types on the scoring results. Furthermore, this classification approach lays a solid foundation for subsequent weighted scoring. Weighted scoring assigns different weights based on the importance and predictive ability of features; this process relies on standardized features to ensure that the contribution of each feature to the scoring model is reasonably assessed.
[0089] In summary, by adopting targeted processing methods for continuous and discrete features, we can effectively improve the expressiveness of features in the scoring model, making the final score mapping more accurate and fair, and providing strong data support for subsequent credit assessment.
[0090] S124. The classification results are weighted to obtain the allocation results.
[0091] In this embodiment, the allocation result refers to the result obtained after assigning weights to the classification result.
[0092] Specifically, based on historical data analysis and expert experience, a percentage-based weighting is used to divide the features in the classification results according to their importance, so as to assign a weight to each feature in the classification results and obtain the allocation result.
[0093] Specifically, in conducting corporate credit assessments, the setting of feature weights is crucial, as it determines the proportion of each feature in the final scoring model. Feature weights are typically set using a percentage system to ensure that the influence and importance of each feature are reasonably quantified and reflected. The weight allocation process usually combines historical data analysis with the experience and judgment of business experts to ensure that the influence of each feature is closely related to the company's creditworthiness.
[0094] The setting of feature weights should first be based on an assessment of the importance of each feature. Each feature (such as a company's sales revenue, debt ratio, industry category, tax record, etc.) has a different degree of impact on a company's creditworthiness; some features may be better predictors of a company's default risk, while others may have a smaller impact. To allocate weights reasonably, each feature needs to be analyzed, which can typically be done in the following ways: By analyzing historical data, we can observe the correlation between various characteristics and a company's creditworthiness (such as whether it has defaulted and its repayment ability). Data mining and statistical analysis (such as correlation coefficient analysis and regression analysis) can help identify which characteristics have strong predictive power when predicting a company's creditworthiness.
[0095] Even though data analytics provides important insights, in some cases, the industry experience and knowledge of business experts remain crucial for judging features. Experts can assign weights to each feature based on their long-term experience, industry trends, and business behavior patterns. Expert judgment helps to fill in any potential biases or gaps in the data.
[0096] Weights are typically assigned on a percentage basis, meaning the total weight of all features equals 100%. For example, if a scoring model has 10 features, 100% can be allocated to each feature based on its importance. More important features are assigned larger weights, while less important features are assigned smaller weights. For instance, a company's financial health (such as profits and liabilities) is often more important than its industry type, and therefore financial features might be given higher weights.
[0097] For some features, there may be different dimensions of measurement (such as numerical values in financial data and industry category classifications). To avoid certain features unfairly influencing the final score, these features need to be standardized to have the same dimensions. After standardization, the weight of each feature is then assigned based on the results of data analysis.
[0098] The initial weight allocation is not the final result and usually requires repeated optimization and adjustment. After the weights are set, the impact of different feature weights on the scoring results can be examined by building a model and performing cross-validation. If the weights of certain features are set too high, the model may become overly reliant on those features, ignoring other potentially important features; while if the weights are too low, the model may become insensitive to certain key features. To improve the accuracy and stability of the model, the weights of each feature can be adjusted based on the validation results, making the model more predictive in practical applications.
[0099] Because the business environment, market conditions, and industry dynamics are constantly changing, feature weights should be dynamically adjustable. As new historical data accumulates and business practices evolve, it is essential to periodically reassess and adjust the weights. Enterprises can leverage machine learning models and artificial intelligence tools for continuous performance monitoring and optimization, adjusting feature weights based on the latest evaluation results.
[0100] By appropriately allocating feature weights, the scoring model can more accurately reflect a company's creditworthiness. Specific objectives include: Each feature contributes differently to a company's credit rating, and the weighting should accurately reflect this.
[0101] By allocating weights appropriately, the scoring model can accurately predict a company's creditworthiness in various scenarios, reducing the risk of misjudgment and omission.
[0102] When businesses or financial institutions use this scoring model, they can make reasonable credit decisions based on the scoring results, thereby reducing the risk of bad debts and improving business efficiency.
[0103] In summary, setting feature weights is a crucial step in corporate credit assessment models. It relies not only on historical data analysis and the output of statistical models but also on the professional judgment and industry knowledge of business experts. Appropriate weight settings can effectively enhance the performance and predictive ability of the scoring model, providing strong data support and decision-making basis for corporate credit assessment.
[0104] S125. Construct quantiles based on the data distribution of the allocation results, and construct a mapping function based on the quantiles to obtain a credit rating model.
[0105] In this embodiment, the quantiles of the allocation results are calculated through statistical analysis, and initial scores and special value processing are set according to the quantiles. A mapping function suitable for the feature distribution is designed based on the quantiles to obtain a credit rating model.
[0106] The mapping functions include linear functions, logarithmic functions, and exponential functions.
[0107] After completing the preliminary work on the scoring model, the next focus is on its implementation to ensure the scores have good discriminative power. The goal is to effectively utilize the characteristic information of enterprises to ensure that enterprises with different credit statuses show significant differences in scores, thereby improving the effectiveness and interpretability of the model.
[0108] According to statistical analysis, it is necessary to first calculate the quantile values of different characteristics, and calculate the 5th quantile, 25th quantile, 50th quantile, 75th quantile, 95th quantile, etc. The number of quantiles to be calculated is different for different characteristics, which needs to be set by business experts (null values are ignored in the process of calculating quantiles).
[0109] Let the sample data for a certain credit feature be... Then the p quantile The calculation is as follows: Where p is the quantile value (e.g., 0.05, 0.25, 0.50), and n is the sample size. This indicates rounding up to the nearest integer.
[0110] Each quantile endpoint is assigned an initial score based on different features and weights, and scores are also assigned to some special values or values with special markings (such as null values, 0 values, -99999, etc.). These initial scores are determined based on the experience of business experts.
[0111] Some special quantiles, such as those below the 5th percentile and those above the 95th percentile, are considered special values. These special values need to be assigned separate scores. For example: ;in It is the score of the current feature. , , , This is a special score assigned to this feature, and this process ensures that extreme values do not affect the overall score.
[0112] Based on the characteristics of different features, corresponding mapping functions are constructed. By combining weights and quantiles, the overall feature is divided into quantiles, and function mapping is applied within each interval. This design ensures that the numerical distribution of the same feature presents a non-smooth curve, so that different values can reflect both differences and stages, thereby improving the discriminative power and the sense of hierarchy.
[0113] For the intrinsic linear function, the mapping is as follows: Where x is the input data, and These are the maximum and minimum values of the input data, respectively. and These are the maximum and minimum values within the target range. It is a switching parameter, if Then the formula becomes a linearly increasing function, if Then the formula becomes a linear decreasing function. Using this method, different scores can be assigned to different features, which can increase the discriminative power.
[0114] The mapping for the logarithmic function is as follows: Where x is the input data, , , and These are the maximum and minimum values within the target range. and These are the minimum and maximum values of the input data, respectively, and d is the base of the objective function. After calculation, where... , is most suitable as the base of the feature objective function.
[0115] For the exponential function, its mapping is as follows: Where x is the input data, , , and These are the maximum and minimum values within the target range. and These are the minimum and maximum values of the input data, respectively. That is the base of the objective function, which, after calculation, is... It is most suitable as the base of the feature target mapping function.
[0116] By dividing the sample data into quantiles and calculating different quantiles (e.g., 5%, 25%, 50%, 75%, 95%), the characteristics of the data distribution can be understood more accurately. This refined statistical analysis helps the model better capture the different distributions of the sample data and allows for the design of suitable mapping functions based on the different characteristics of the data, thereby improving the model's discriminative power and accuracy.
[0117] By combining different types of mapping functions (such as linear, logarithmic, and exponential functions) with weights to handle quantiles of different features, the model can become more sensitive to differences between different values. For example, using logarithmic and exponential mapping functions can better capture changes in extreme data, thus avoiding the impact of extreme values on the score. It can also improve the hierarchy of the score, allowing the model to show clear differences when evaluating companies with different credit ratings.
[0118] By assigning separate scores to special values (such as null values, 0 values, -99999, etc.), it is ensured that these extreme or outlier values will not affect the overall score. This avoids the disturbance of scoring results by extreme data, ensuring that the model can maintain stability and accuracy when faced with outliers. This is also an important measure for data quality control in credit scoring models.
[0119] Through statistical analysis and the proper design of mapping functions, every step of the model's changes and the scoring process can be clearly explained. For example, the setting of quantile values and the design of mapping functions can be adjusted based on the experience of business experts, making the scoring process not only efficient but also clearly explaining to users and decision-makers the contribution of each feature to the final credit rating, thereby improving the model's transparency and interpretability.
[0120] The scoring model was designed with extensive input from business experts, particularly in handling special values and setting initial scores. This expertise is crucial for ensuring the model's effectiveness in practical applications, making it more aligned with real-world credit assessment needs and thus enhancing its usability and operability.
[0121] The design of mapping functions (such as switching parameters to control linear increments and decrements) enables the model to adapt to the characteristics of different features. Through different types of function mappings, the model can flexibly adjust according to the distribution of different features. This flexibility and adaptability allow the model to maintain high prediction accuracy and stability when facing different types of data.
[0122] Applying special handling to quantiles below 5% and above 95% helps to effectively manage outliers in the data and prevent these values from affecting the model. Specialized outlier handling reduces the interference of these extreme data points on the overall scoring results, improving the model's robustness and effectiveness.
[0123] By employing reasonable quantile calculations and adaptive mapping function design, the credit rating model can be made more accurate, interpretable, and discriminative. This method fully considers the differences in various data characteristics, enabling it to provide more reliable and actionable scoring results when assessing the creditworthiness of different enterprises, thereby enhancing the model's effectiveness and practical application value.
[0124] Because the weighting is based on a 100-point scale, but the output is based on a 900-point scale, the final score needs to be adjusted. Additionally, for negative characteristics that the company has identified, corresponding points will be deducted, such as for legal cases or administrative penalties. ,in This is the final score (out of 900). It is the base score (out of 100). ,in Based on the base score (out of 100). For the first The weights of each feature, For the first The score after feature mapping, where n is the total number of continuous features.
[0125] The P deduction items depend on whether the company has identified any negative characteristics, and can be defined as follows: ,in Is the company in the first The hit situation on a negative feature (if hit, then) ,otherwise ), This is the score weight corresponding to the negative feature.
[0126] If the company does not exhibit any negative characteristics, then The score will not be affected by the deduction; however, if a company has multiple negative characteristics, the deductions will accumulate, and the final score will be reduced accordingly.
[0127] Final score The data is mapped to the corresponding corporate performance rating, and the interpretation of different statistics is shown in Table 2. The higher the rating, the better the corporate credit, the stronger the performance capability, and the better the business performance.
[0128] Table 2. Enterprise Performance Rating Score grade Definition of level [750 (inclusive) — 900 points] Grade A The company has strong contract performance capabilities. It has a good credit record, sound business operations, and minimal impact from uncertainties on its operations and development. [700 (inclusive) - 750 points] B+ level The company has a strong ability to fulfill its obligations. It has a good credit record, sound business operations, and is less affected by uncertainties in its operations and development. [650 (inclusive) - 700 points] Grade B The company has a good ability to fulfill its obligations. No or very few instances of general credit default have been found, and the company is operating in a virtuous cycle. However, there may be some uncertainties that could affect its future operations and development, thereby impacting its debt repayment ability. [600 (inclusive) - 650 points] C+ level The company's ability to fulfill its obligations is generally average. No major defaults or only minor issues have been identified. Its operations and future development are relatively susceptible to uncertainties, and its debt repayment ability is generally average. [550 (inclusive) - 600 points] Class C The company's ability to fulfill its obligations is relatively weak. It may have a small number of negative credit records, an uncertain future outlook, and a weak debt repayment capacity. [500 (inclusive) - 550 points] D+ level The company has poor debt repayment ability. It may have a negative credit history, poor operating conditions, and weak debt repayment capacity. [450 (inclusive) - 500 points] Class D The company has a very poor ability to fulfill its obligations. It may have a negative credit history, poor operating conditions, and weak debt repayment ability. Below 450 points Class E The company has extremely poor ability to fulfill its obligations. It may have a significant number of negative credit records, be in very poor financial condition, and have virtually no capacity to repay its debts. By considering multiple features and allocating them according to weights, a comprehensive assessment of a company's repayment ability and creditworthiness can be achieved, including its operating status, credit history, and negative characteristics. The negative characteristic deduction mechanism promptly reflects negative information about the company, ensuring a more accurate and fair final score. This avoids misclassifying companies with poor credit histories as having high credit ratings. By mapping scores to different repayment levels, a company's repayment ability can be clearly defined, helping external institutions, investors, or partners better understand the company's creditworthiness and potential risks. A company's credit rating provides debtors or partners with a basis for decision-making, helping them assess the company's repayment ability and future operational risks. The negative characteristic deduction mechanism ensures that the company's score is dynamically adjusted as circumstances change, reflecting changes in the company's credit history and helping external stakeholders obtain more real-time and accurate information.
[0129] Overall, this scoring and rating system provides a quantitative tool for corporate credit assessment and risk management, helping stakeholders make more informed decisions.
[0130] S130. Determine the industry to which the enterprise belongs based on the data of the enterprise to be rated.
[0131] In this embodiment, during the credit score calculation process, the industry's characteristics differ across industries, affecting the quantile division of the scoring model. Therefore, it is first necessary to determine the industry to which a company belongs based on its industry information. The implementation method is as follows: Companies to be rated are categorized by industry. Industry information can typically be obtained through the company's industry code, description of its main business, or through third-party databases.
[0132] Once the industry to which a company belongs is determined, the company is categorized into the corresponding industry category, and the quantile range of that industry is used for calculations in the subsequent scoring model.
[0133] S140. Determine the mapping relationship based on the industry to which the enterprise belongs.
[0134] In this embodiment, this step involves differentiated scoring for companies in different industries to improve the accuracy and comparability of the scores. The implementation method is as follows: Different industries may require different quantile rules. For each industry, the model needs to adjust the feature distribution based on the industry's characteristic data and establish industry-specific mapping relationships.
[0135] Each industry may have specific risk and credit characteristics, and industry-specific mapping relationships can ensure that the scoring model is better adapted to the characteristics of each industry. By analyzing the data distribution of different industries, a data mapping function suitable for that industry can be constructed.
[0136] S150. Determine the credit rating of the enterprise based on the calculation results and the mapping relationship.
[0137] In this embodiment, based on the credit score calculation results and industry-specific mapping relationships, the model further determines the enterprise's credit rating. The specific steps are as follows: Based on the calculated corporate credit score, companies are classified into different credit ratings. For example: A, B+, B, C+, etc. Each rating represents a different aspect of the company's ability to fulfill obligations, credit history, and operational status.
[0138] Based on the previously established mapping rules, define the correspondence between score ranges and credit ratings. For example, a score of 750 or above is rated A, and a score between 700 and 750 is rated B+, etc.
[0139] Based on industry mapping relationships and scoring calculation models, also known as credit rating models, the final credit rating of enterprises is further adjusted to ensure the accuracy and reliability of the rating results.
[0140] S160, Output the credit rating of the enterprise.
[0141] In this embodiment, based on the calculated credit rating, the system will output the enterprise's credit rating. The output is typically a rating identifier, including: Credit Rating Report: Automatically generates a credit rating report for the company, detailing the company's credit rating and related explanations.
[0142] Data visualization: Provides stakeholders with charts or visualizations of credit rating results to facilitate a quick understanding of a company's creditworthiness.
[0143] Through these steps, the entire corporate credit rating process is automated, enabling real-time dynamic calculation and updating of credit ratings based on a company's characteristic data. This automation significantly improves the efficiency and accuracy of credit rating, allowing companies to obtain accurate credit assessments in a shorter time, thereby assisting decision-makers in risk assessment and credit management.
[0144] The above steps provide a clear implementation framework for automating the corporate credit rating process. Through data collection, processing, feature selection, determination of mapping relationships, and final rating allocation, the entire system achieves automated calculation and updating of credit ratings, exhibiting a high degree of adaptability and optimization capabilities.
[0145] In this embodiment, before starting the enterprise credit rating, it is necessary to collect relevant feature data for each enterprise. This data will serve as input to the model for predicting credit ratings. An enterprise dataset is defined, containing multiple enterprise features, each representing a specific indicator for a given enterprise. Enterprise dataset definition: in, It is a collection of enterprise characteristics, each This represents a characteristic variable of a company. Since companies in different industries have different operating characteristics and credit risk profiles, quantiles need to be assigned based on industry. By determining a company's industry code, its industry affiliation can be identified. Each company has an industry code, which can be processed using a classification function to determine the specific industry category to which the company belongs. Industry Affiliation , ,in It is a classification function. It refers to the industry category to which the company belongs. This is the company's industry code; once the company's industry affiliation is confirmed, the system merges the new company's data with that of other companies in the same industry and automatically generates quantiles for each feature. These quantiles are used for the final score calculation, and the resulting score is converted into a corresponding credit rating according to a preset mapping relationship. In this way, the company's credit rating is calculated and output in real time. This method automates corporate credit rating, not only dynamically classifying corporate credit ratings in real time but also possessing the ability to continuously optimize as the amount of data increases, thereby improving the accuracy and efficiency of the rating.
[0146] By deduplicated, cleaned, formatted, and anomaly removed data, combined with hashing of sensitive information using encryption algorithms such as SHA-256, the irreversibility and anonymity of the data are ensured. This processing method not only guarantees data quality but also overcomes the limitations of existing technologies in data acquisition and privacy protection, allowing companies to participate in credit rating while protecting trade secrets, significantly improving data availability and rating participation.
[0147] Twenty-eight key features were selected from a large pool of characteristics, covering four major areas: basic business conditions, operational status, financial condition, and creditworthiness. By combining expert experience with data analysis, the high correlation and predictive power of the selected features with credit ratings were ensured. This precise selection overcomes the overfitting or underfitting problems caused by inappropriate feature selection in existing technologies, improving the accuracy and reliability of the ratings.
[0148] Based on the different properties of the features, quantile and binning techniques are used for processing, with different processing methods applied to continuous and discrete features, and special values are assigned separate scores. This mechanism addresses the lack of flexibility in data processing in existing technologies, enabling various types of data to be better integrated into the rating system and improving the model's adaptability and inclusiveness.
[0149] By designing various mapping functions, such as linear, logarithmic, and exponential functions, features can be transformed into scores in the most appropriate way, and the increasing or decreasing functions can be flexibly switched by adjusting parameters. This multivariate mapping mechanism solves the problem of uneven data distribution in existing technologies, ensures fair and reasonable scoring, and significantly improves the stability and adaptability of the model.
[0150] Based on historical data and expert experience, importance weights are assigned to each feature, and a deduction mechanism is designed to handle negative characteristics of enterprises. This mechanism solves the problems of information overload and data bias in existing technologies, making the rating results more comprehensive, in line with actual business needs, and significantly improving the interpretability and applicability of the model.
[0151] Based on the assessment and classification of the industry to which enterprises belong, data from enterprises within the same industry are integrated to generate industry-specific quantiles, enabling differentiated ratings. This mechanism addresses the insufficient generalization ability of existing technologies, making the rating results more reflective of industry characteristics and enhancing the comparability and reference value of rating results across different industries.
[0152] The system achieves full automation of the rating process, including data collection, processing, and final rating result output, and possesses the ability to automatically optimize as the data volume increases. This automated update mechanism solves the problems of lagging knowledge graph updates and untimely data updates in existing technologies, enabling the rating system to continuously learn and optimize, ensuring the timeliness and accuracy of rating results.
[0153] By combining business expert experience with data analysis, 28 key features were selected from hundreds of characteristics, covering four dimensions: basic enterprise status, operational status, financial status, and credit status. This multi-dimensional feature selection method breaks through the dependence of traditional technologies on a single data source, constructing a comprehensive and insightful credit assessment framework.
[0154] By calculating multiple quantiles (such as 5%, 25%, 50%, 75%, 95%, etc.) and handling special and extreme values separately, a feature mapping mechanism based on statistical analysis was established. This innovative application enables the rating model to effectively handle outliers and anomalies, enhances the model's adaptability to different feature distributions, and ensures the reasonable setting of dynamic rating thresholds.
[0155] This system designs various mapping functions, including linear, logarithmic, and exponential functions, and allows for flexible switching between increasing and decreasing functions by adjusting parameters. This mapping system can select the optimal mapping method for the distribution characteristics of different features, thereby improving the discrimination and accuracy of the scoring and overcoming the limitations of a single mapping method.
[0156] A two-tiered scoring mechanism is employed, consisting of a base score of 100 points and a final score of 900 points. A deduction system is designed for negative characteristics, ensuring both internal consistency and external comparability of the scoring results. This mechanism reflects the company's basic creditworthiness and, by comprehensively considering special circumstances and risk factors, provides more detailed and differentiated rating results.
[0157] By categorizing companies according to their industry, integrating data from companies within the same industry, generating industry-specific quantiles, and implementing differentiated ratings, this technology addresses the issue of significant differences in characteristics between industries, making rating results more comparable within an industry, while also reflecting the overall credit environment and development trends of the industry.
[0158] An adaptive deduction weighting system was designed to precisely adjust the rating results based on the company's hit rate on negative characteristics, making them closer to the company's actual risk level. This system can sensitively capture changes in corporate risk signals and performance capabilities, enhancing the rating's foresight and early warning capabilities, and effectively reducing the lag in credit assessment.
[0159] A complete automated rating process has been built, forming a closed loop from data collection and processing to final result output, and possessing self-optimization capabilities. This closed-loop architecture improves rating efficiency, reduces labor costs, and continuously evolves the rating system through the accumulation of data feedback, ensuring the timeliness and market synchronization of rating results.
[0160] The aforementioned automated corporate credit rating method ensures the comprehensiveness and accuracy of the rating process by inputting the data of the companies to be rated into the credit rating model and determining specific mapping relationships based on industry classification. First, it accurately acquires and cleans corporate data, addressing the shortcomings of existing technologies in data acquisition and quality. Second, through industry-specific mapping relationships, it effectively improves the industry adaptability and differentiation of the rating, enhancing the interpretability of the model. Finally, by combining a multivariate mapping function with a dynamic weight adjustment mechanism, it ensures the accuracy and stability of the rating results, simplifies the rating process, solves the complexity and accuracy problems of traditional technologies, and achieves a more comprehensive and precise credit assessment.
[0161] Figure 4 This is a schematic block diagram of an automatic enterprise credit rating device 300 provided in an embodiment of the present invention. Figure 4As shown, corresponding to the above-described automatic enterprise credit rating method, the present invention also provides an automatic enterprise credit rating device 300. This automatic enterprise credit rating device 300 includes a unit for executing the above-described automatic enterprise credit rating method, and the device can be configured in a server. Specifically, please refer to... Figure 4 The enterprise credit automatic rating device 300 includes a data acquisition unit 301, a calculation unit 302, an industry determination unit 303, a mapping relationship determination unit 304, a rating determination unit 305, and an output unit 306.
[0162] The data acquisition unit 301 is used to acquire data of the enterprise to be rated; the calculation unit 302 is used to input the data of the enterprise to be rated into the credit rating model to calculate the credit score and obtain the calculation result; the industry determination unit 303 is used to determine the industry to which the enterprise belongs based on the data of the enterprise to be rated; the mapping relationship determination unit 304 is used to determine the mapping relationship based on the industry to which the enterprise belongs; the rating determination unit 305 is used to determine the credit rating of the enterprise based on the calculation result and the mapping relationship; and the output unit 306 is used to output the credit rating of the enterprise.
[0163] The training process of the credit rating model includes: acquiring relevant enterprise data; performing privacy protection and anonymization processing on the relevant enterprise data to obtain processing results; selecting key credit features from the processing results and classifying the features to obtain classification results; assigning weights to the classification results to obtain allocation results; constructing quantiles based on the data distribution of the allocation results, and constructing a mapping function based on the quantiles to obtain the credit rating model.
[0164] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned automatic enterprise credit rating device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0165] The aforementioned automatic enterprise credit rating device 300 can be implemented as a computer program, which can, for example... Figure 5 It runs on the computer device shown.
[0166] Please see Figure 5 , Figure 5 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0167] See Figure 5The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0168] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an automatic enterprise credit rating method.
[0169] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0170] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an automatic enterprise credit rating method.
[0171] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0172] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps: Obtain data on the companies to be rated; input the data into the credit rating model to calculate a credit score and obtain the result; determine the industry to which the company belongs based on the data; determine the mapping relationship based on the industry to which the company belongs; determine the credit rating of the company based on the calculation result and the mapping relationship; and output the credit rating of the company.
[0173] The training process of the credit rating model includes: Obtain relevant enterprise data; perform privacy protection and anonymization processing on the relevant enterprise data to obtain processing results; filter key credit features from the processing results and perform feature classification to obtain classification results; assign weights to the classification results to obtain allocation results; construct quantiles based on the data distribution of the allocation results, and construct a mapping function based on the quantiles to obtain a credit rating model.
[0174] The enterprise-related data includes basic enterprise information, operating conditions, financial status, and credit status.
[0175] In one embodiment, when the processor 502 implements the step of filtering key credit features from the processing result and performing feature classification to obtain a classification result, the following steps are specifically implemented: Key credit features are selected from the processing results. Based on the nature of the features, the key credit features are classified into continuous features and discrete features. The continuous features are assigned scores, and the discrete features are processed using binning technology to obtain the classification results.
[0176] In one embodiment, when the processor 502 performs the step of weighting the classification result to obtain the allocation result, it specifically implements the following steps: Based on historical data analysis and expert experience, a percentage-based weighting system is used to divide the features in the classification results according to their importance, so as to assign a weight to each feature in the classification results and obtain the allocation result.
[0177] In one embodiment, when implementing the step of constructing quantiles based on the data distribution of the allocation results and constructing a mapping function based on the quantiles to obtain the credit rating model, the processor 502 specifically implements the following steps: The quantiles of the allocation results are calculated through statistical analysis, and initial scores and special value processing are set according to the quantiles. A mapping function suitable for the feature distribution is designed based on the quantiles to obtain the credit rating model.
[0178] The mapping function includes linear functions, logarithmic functions, and exponential functions.
[0179] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0180] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0181] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps: Obtain data on the companies to be rated; input the data into the credit rating model to calculate a credit score and obtain the result; determine the industry to which the company belongs based on the data; determine the mapping relationship based on the industry to which the company belongs; determine the credit rating of the company based on the calculation result and the mapping relationship; and output the credit rating of the company.
[0182] The training process of the credit rating model includes: Obtain relevant enterprise data; perform privacy protection and anonymization processing on the relevant enterprise data to obtain processing results; filter key credit features from the processing results and perform feature classification to obtain classification results; assign weights to the classification results to obtain allocation results; construct quantiles based on the data distribution of the allocation results, and construct a mapping function based on the quantiles to obtain a credit rating model.
[0183] The enterprise-related data includes basic enterprise information, operating conditions, financial status, and credit status.
[0184] In one embodiment, when the processor executes the computer program to implement the step of filtering key credit features from the processing results and performing feature classification to obtain classification results, it specifically implements the following steps: Key credit features are selected from the processing results. Based on the nature of the features, the key credit features are classified into continuous features and discrete features. The continuous features are assigned scores, and the discrete features are processed using binning technology to obtain the classification results.
[0185] In one embodiment, when the processor executes the computer program to implement the step of weighting the classification results to obtain the allocation results, it specifically implements the following steps: Based on historical data analysis and expert experience, a percentage-based weighting system is used to divide the features in the classification results according to their importance, so as to assign a weight to each feature in the classification results and obtain the allocation result.
[0186] In one embodiment, when the processor executes the computer program to implement the steps of constructing quantiles based on the data distribution of the allocation results and constructing a mapping function based on the quantiles to obtain the credit rating model, the processor specifically implements the following steps: The quantiles of the allocation results are calculated through statistical analysis, and initial scores and special value processing are set according to the quantiles. A mapping function suitable for the feature distribution is designed based on the quantiles to obtain the credit rating model.
[0187] The mapping function includes linear functions, logarithmic functions, and exponential functions.
[0188] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0189] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0190] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0191] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0193] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An automatic enterprise credit rating method, characterized in that, include: Obtain data on companies to be rated; The data of the companies to be rated is input into the credit rating model to calculate the credit score and obtain the calculation result; The industry to which the enterprise belongs is determined based on the data of the enterprise to be rated; The mapping relationship is determined based on the industry to which the enterprise belongs; The credit rating of the enterprise is determined based on the calculation results and the mapping relationship. Output the credit rating of the company.
2. The automatic enterprise credit rating method according to claim 1, characterized in that, The training process of the credit rating model includes: Obtain relevant enterprise data; The relevant enterprise data is subjected to privacy protection and anonymization processing to obtain the processing result; Key credit features are selected from the processing results and then classified to obtain classification results. The classification results are weighted to obtain the allocation results; Quantiles are constructed based on the data distribution of the allocation results, and a mapping function is constructed based on the quantiles to obtain the credit rating model.
3. The automatic enterprise credit rating method according to claim 2, characterized in that, The enterprise-related data includes basic enterprise information, operating conditions, financial status, and credit status.
4. The automatic enterprise credit rating method according to claim 2, characterized in that, The step of filtering key credit features from the processing results and classifying these features to obtain classification results includes: Key credit features are selected from the processing results. Based on the nature of the features, the key credit features are classified into continuous features and discrete features. The continuous features are assigned scores, and the discrete features are processed using binning technology to obtain the classification results.
5. The automatic enterprise credit rating method according to claim 2, characterized in that, The weight allocation of the classification results to obtain the allocation result includes: Based on historical data analysis and expert experience, a percentage-based weighting system is used to divide the features in the classification results according to their importance, so as to assign a weight to each feature in the classification results and obtain the allocation result.
6. The automatic enterprise credit rating method according to claim 2, characterized in that, The step of constructing quantiles based on the data distribution of the allocation results, and constructing a mapping function based on the quantiles to obtain the credit rating model, includes: The quantiles of the allocation results are calculated through statistical analysis, and initial scores and special value processing are set according to the quantiles. A mapping function suitable for the feature distribution is designed based on the quantiles to obtain the credit rating model.
7. The automatic enterprise credit rating method according to claim 6, characterized in that, The mapping functions include linear functions, logarithmic functions, and exponential functions.
8. An automatic enterprise credit rating device, characterized in that, include: The data acquisition unit is used to acquire data on the companies to be rated. The calculation unit is used to input the data of the enterprise to be rated into the credit rating model to calculate the credit score and obtain the calculation result; The industry determination unit is used to determine the industry to which the enterprise belongs based on the data of the enterprise to be rated. A mapping relationship determination unit is used to determine a mapping relationship based on the industry to which the enterprise belongs; A rating determination unit is used to determine the credit rating of an enterprise based on the calculation results and the mapping relationship; The output unit is used to output the credit rating of the enterprise.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.