Small and micro enterprise credit risk prediction method and device based on graph neural network
By constructing a small and micro enterprise association relationship graph based on graph neural networks and using graph convolutional neural network models to predict credit risk, the problem of inaccurate credit risk assessment of small and micro enterprises in traditional credit assessment methods is solved, more accurate and timely risk warnings are achieved, and the loan approval efficiency and accuracy of financial institutions are improved.
Patent Information
- Application Number
- CN202510598125.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional credit assessment methods are difficult to accurately assess the credit risk of small and micro enterprises, especially because the financial information of related enterprises in existing technologies is opaque. Traditional methods are unable to identify the financial information of small enterprises and the relationships between small enterprises are often incomplete, resulting in financial information asymmetry. Traditional methods are unable to accurately assess the relationships between small and micro enterprises, resulting in one-sided and inaccurate credit risk assessment results.
Based on graph neural network, a small and micro enterprise association relationship graph is constructed, and a small and micro enterprise credit risk prediction model is constructed through the graph convolutional neural network framework. The graph neural network model is used to learn the association relationship between small and micro enterprises, and risk prediction is performed through the node-level and semantic-level attention mechanism to construct a small and micro enterprise credit risk prediction model.
It improves the accuracy and timeliness of credit risk assessment for small and micro enterprises, reduces the time and cost of manual review, improves the accuracy and efficiency of financial institutions' loan approval decisions for small and micro enterprises, and reduces the non-performing loan rate.
Smart Images

Figure CN120672453A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a method and device for predicting credit risk of small and micro enterprises based on a graph neural network. Background Art
[0002] Small and micro enterprises (SMEs), including small and micro businesses, family-run businesses, and individual businesses, are an integral part of the national economy and play an irreplaceable role in driving economic growth, employment, tax revenue, and technological innovation. However, with the evolving market environment and intensifying competition, SMEs face numerous challenges in financing, particularly in credit risk management.
[0003] Currently, traditional banks rely primarily on small and micro-enterprise credit risk assessments based on their financial statements, collateral, and the subjective judgment of relationship managers. While these methods can help banks identify credit risk to a certain extent, the often opaque financial information and incomplete financial statements of small and micro-enterprises, coupled with limited collateral assets, make it difficult for banks to accurately assess their true probability of default. Furthermore, traditional credit assessment methods overlook the interdependencies among small and micro-enterprises, such as capital flows, investments, guarantees, and interpersonal relationships—information that is crucial for a comprehensive assessment of their credit risk.
[0004] In recent years, with the development of big data and artificial intelligence technologies, some financial institutions have begun experimenting with using machine learning models to predict credit risk. These methods extract multi-dimensional characteristics of small and micro enterprises (SMEs) to train credit risk assessment models. While these methods have improved prediction accuracy to some extent, they still have many shortcomings. For example, traditional machine learning models perform poorly when dealing with complex relationships and struggle to capture the dynamic interactions and risk transfer processes among SMEs. Traditional methods rely on limited financial data and collateral information, ignoring the complex relationships among SMEs, resulting in one-sided and inaccurate assessment results. Summary of the Invention
[0005] The present invention provides a method and device for predicting the credit risk of small and micro enterprises based on graph neural networks, which is used to address the shortcomings of the existing technology in terms of insufficient accuracy in predicting the credit risk of small and micro enterprises and difficulty in comprehensively assessing enterprise-related risks and credit risks, thereby improving the accuracy and timeliness of credit risk assessment.
[0006] The present invention provides a method for predicting credit risk of small and micro enterprises based on a graph neural network, comprising the following steps: Obtaining original data of the small and micro enterprises to be tested within the loan duration period, performing data preprocessing on the original data of the small and micro enterprises to be tested, and obtaining data of the small and micro enterprises to be tested; Constructing a correlation diagram of the small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested; Simplifying the association relationship diagram of the small and micro enterprises to be tested to obtain a simplified association relationship diagram; Input the simplified correlation relationship graph into the trained small and micro enterprise credit risk prediction model to perform prediction, and obtain the default risk probability of the small and micro enterprise to be tested; Determining the classification of the small and micro enterprises to be tested based on the default risk probability of the small and micro enterprises to be tested; The trained small and micro enterprise credit risk prediction model is trained based on a small and micro enterprise association relationship diagram, and the small and micro enterprise association relationship diagram is constructed based on data samples of small and micro enterprises that have purchased loan products.
[0007] According to a graph neural network-based credit risk prediction method for small and micro enterprises provided by the present invention, a small and micro enterprise association relationship graph is constructed based on small and micro enterprise data samples of purchased loan products. The method specifically includes: obtaining small and micro enterprise data samples of purchased loan products that meet preset conditions; establishing a time window for the small and micro enterprise data samples; determining the prediction variables and target variables of the small and micro enterprises as nodes; based on the time window, determining the association relationships between small and micro enterprises according to the prediction variables; and constructing a small and micro enterprise association relationship graph based on the association relationships.
[0008] According to a small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention, the association relationship graph of the small and micro enterprises to be tested is simplified to obtain a simplified association relationship graph, specifically including: the small and micro enterprise association relationship graph is a directed multi-heterogeneous graph; the directed multi-heterogeneous graph is composed of a directed graph, a multi-graph and a heterogeneous graph; the directed graph is converted into an undirected graph; and the multi-heterogeneous graph is split into several semantically homogeneous networks.
[0009] According to a small and micro enterprise credit risk prediction method based on a graph neural network provided by the present invention, the small and micro enterprise credit risk prediction model is trained according to the small and micro enterprise association relationship graph to obtain a trained small and micro enterprise credit risk prediction model, specifically comprising: simplifying the small and micro enterprise association relationship graph to obtain a semantic homogeneous network, and projecting all nodes in the semantic homogeneous network into a unified feature space; in each semantic homogeneous network, the small and micro enterprise credit risk prediction model learns the first importance weights of different neighbor nodes through a node-level attention mechanism; the neighbor nodes are nodes directly connected to the current node; the small and micro enterprise credit risk prediction model jointly learns the second importance weight of each semantic homogeneous network through a semantic-level attention mechanism; based on the first importance weight and the second importance weight, the small and micro enterprise credit risk prediction model performs classification prediction on all nodes in the small and micro enterprise association relationship graph through a multi-layer perceptron; calculating the classification loss according to the classification prediction result, and optimizing the small and micro enterprise credit risk prediction model according to the classification loss, and obtaining a trained small and micro enterprise credit risk prediction model after the optimization is completed.
[0010] According to a small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention, the correlation relationship between small and micro enterprises is determined according to the prediction variables, and a small and micro enterprise correlation relationship graph is constructed according to the correlation relationship, which specifically includes: determining attribute data on each node in the small and micro enterprise correlation relationship graph according to small and micro enterprise data; the node represents a small and micro enterprise; the attribute data corresponds to the characteristic value of the prediction variable; based on the time window, determining the correlation relationship between nodes according to the attribute data; and generating a small and micro enterprise correlation relationship graph according to the correlation relationship between the nodes.
[0011] According to a graph neural network-based credit risk prediction method for small and micro enterprises provided by the present invention, splitting multiple heterogeneous graphs into several semantically homogeneous networks specifically includes: converting the heterogeneous graphs into a single-node multi-relationship heterogeneous network; dividing the single-node multi-relationship heterogeneous network into several relationship graphs according to different relationship categories; and treating each relationship graph as a semantically homogeneous network.
[0012] According to a small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention, the association relationship between the nodes includes a capital flow relationship; the capital flow relationship between the nodes is determined based on the attribute data, specifically including: querying and incorporating small and micro enterprises that have indirect capital transaction relationships with the small and micro enterprise center sample into the small and micro enterprise association relationship graph; aggregating the capital flow transaction data between two nodes in the small and micro enterprise association relationship graph by month, and generating an edge based on this as the capital flow relationship between the nodes.
[0013] The present invention also provides a small and micro enterprise credit risk prediction device based on graph neural network, which includes the following modules: The module for acquiring data of small and micro enterprises to be tested is used to acquire original data of small and micro enterprises to be tested during the loan duration period, and to perform data preprocessing on the original data of small and micro enterprises to be tested to obtain data of small and micro enterprises to be tested; A module for constructing a correlation diagram of small and micro enterprises to be tested, configured to construct a correlation diagram of small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested; The small and micro-enterprise credit risk prediction module is used to simplify the association relationship diagram of the small and micro-enterprises to be tested to obtain a simplified association relationship diagram; input the simplified association relationship diagram into a trained small and micro-enterprise credit risk prediction model for prediction to obtain the default risk probability of the small and micro-enterprises to be tested; determine the classification of the small and micro-enterprises to be tested based on the default risk probability of the small and micro-enterprises to be tested; the trained small and micro-enterprise credit risk prediction model is trained based on the small and micro-enterprise association relationship diagram, and the small and micro-enterprise association relationship diagram is constructed based on the small and micro-enterprise data sample of the purchased loan products.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements any of the above-mentioned methods for predicting credit risks of small and micro enterprises based on graph neural networks.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements any of the above-mentioned methods for predicting credit risks of small and micro enterprises based on graph neural networks.
[0016] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for predicting credit risks of small and micro enterprises based on graph neural networks.
[0017] The present invention provides a method and device for predicting the credit risk of small and micro enterprises based on a graph neural network. By acquiring data on small and micro enterprises in real time or periodically and constructing an association relationship graph, it is possible to promptly discover and track the operating conditions and credit risk changes of small and micro enterprises, thereby identifying possible credit risks in advance and providing financial institutions with earlier risk warnings. By learning and predicting on the association relationship graph of small and micro enterprises, the credit risk of small and micro enterprises can be more accurately reflected, thereby improving the accuracy of the prediction. The credit risk prediction model based on a graph neural network can provide financial institutions with more accurate and comprehensive credit risk assessments, helping financial institutions make more informed decisions during the loan approval process and reduce the non-performing loan rate. Through automated credit risk prediction, the time and cost of manual review can be greatly reduced, and the efficiency of financial services can be improved. At the same time, it can also provide small and micro enterprises with faster and more convenient loan services. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a schematic diagram of a capital flow association network constructed based on small and micro enterprise data provided by the present invention.
[0020] Figure 2 It is a data structure diagram of the capital flow relationship provided by the present invention.
[0021] Figure 3 It is a data structure diagram of the investment relationship provided by the present invention.
[0022] Figure 4 It is a data structure diagram of the guarantee relationship provided by the present invention.
[0023] Figure 5 It is a data structure diagram of the person-enterprise relationship provided by the present invention.
[0024] Figure 6 This is a schematic diagram of the deduplication logic of capital flow data provided by the present invention that matches the data acquisition logic.
[0025] Figure 7 This is the association relationship network splitting logic diagram provided by the present invention.
[0026] Figure 8 It is a schematic diagram of the node-level and semantic-level aggregation process provided by the present invention.
[0027] Figure 9This is a structural diagram of credit risk prediction for small and micro enterprises based on graph neural networks provided by the present invention.
[0028] Figure 10 It is a schematic diagram showing the AUC value and KS value of each model provided by the present invention.
[0029] Figure 11 It is a flow chart of the small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention.
[0030] Figure 12 It is a structural diagram of the small and micro enterprise credit risk prediction device based on graph neural network provided by the present invention.
[0031] Figure 13 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0032] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0033] This paper is about the application of deep learning technology in the context of credit risk prediction for small and micro enterprises. Specifically, it is about the technical solution, implementation method, and practical application of how to use graph convolutional neural network models to dynamically monitor the credit default risks of small and micro enterprises.
[0034] Small and micro enterprises (SMEs), a collective term for small businesses, micro-enterprises, family-run businesses, and individual businesses, play a vital role in national economic and social development. However, they face significant challenges. Due to issues such as opaque financial information, insufficient collateral, and information asymmetry, SMEs generally face high credit risk, leading to severe financing constraints. Consequently, the difficulty and high cost of financing for SMEs have become a focus of widespread public concern and heated debate. Effectively addressing these challenges and helping SMEs thrive is a key issue requiring immediate attention.
[0035] Small and micro enterprises (SMEs) generally face deficiencies in operational stability, financial transparency, and internal oversight mechanisms. To address the financing challenges faced by SMEs (and, for banks, their clients and merchants) and effectively control loan default rates, credit risk management is crucial. In the absence of sufficient collateral, how can banks leverage their extensive databases accumulated over many years to dynamically monitor customer behavioral risks, accurately predict default probabilities, and provide timely warnings to potential defaulters? These are pressing challenges in the development of SME credit lending by banks.
[0036] In traditional lending, banks determine risk premiums based on their assessment of loan default risk and profit from the interest rate spread between deposits and loans. However, due to their small scale and incomplete financial statements, small and micro enterprises face significant information asymmetry with banks, making it difficult for banks to accurately identify their true probability of default, mitigate losses caused by default risk, and achieve reasonable pricing. Therefore, establishing default risk management methods and technologies suitable for small and micro enterprises, effectively identifying and measuring default risk for small and micro loans, and thus supporting the implementation of reasonable risk pricing, is key to solving their financing challenges. Accurately estimating the probability of default based on credit risk forecasts is crucial for improving default risk management capabilities.
[0037] Currently, there are several issues with managing the loan lifecycle for small and micro enterprises. These include a significant human factor, the need for significant staff input, high demands on relationship managers, numerous uncontrollable factors, and the potential for misjudgment, leading to significant operational risks. These issues can reduce relationship managers' willingness to promote relevant products, thereby impacting banks' ability to provide inclusive finance services to small and micro enterprises.
[0038] To effectively address the financing difficulties faced by small and micro enterprises, this paper utilizes graph theory and a graph convolutional neural network framework to construct a small and micro enterprise relationship graph (network) and, based on this graph, establishes a small and micro enterprise credit risk prediction model. This model can predict the lifetime default risk probability of small and micro enterprise customers to whom loans have been granted. This method can bring multiple benefits to businesses, including: 1. Through the graph convolutional neural network-based model, the default risk of small and micro enterprises can be predicted more accurately and early warnings can be issued in a timely manner, thereby improving the accuracy of post-loan risk warnings.
[0039] 2. Through automated and intelligent processing, the need for large-scale manpower input can be reduced, the workload of account managers can be reduced, and thus human resource costs can be reduced.
[0040] 3. By reducing human errors and subjective biases, the probability of operational risks can be reduced and the robustness of risk management can be improved.
[0041] 4. Accurate risk prediction and management can improve the quality of banks' lending decisions, enhance their ability and confidence in providing inclusive financial services to small and micro enterprises, and thus better support the development of small and micro enterprises.
[0042] In summary, the small and micro enterprise credit risk prediction model based on graph convolutional neural network can provide effective support for solving the financing difficulties of small and micro enterprises and enhance the inclusive financial service capabilities of banks.
[0043] The overall idea of the system construction of the present invention is to construct a dynamic small and micro enterprise association relationship diagram, and establish a small and micro enterprise credit risk prediction model based on the network, predict the probability of default of loan customers during the loan period, and provide timely warnings to potential defaulting customers, thereby improving the customers' post-loan risk warning rate, precision rate and completeness rate.
[0044] Taking into account the information asymmetry of small and micro enterprises and other problems, we rationally and fully utilize all available resources within the framework of legality and compliance. By introducing graph theory technology and graph convolutional neural network models, the core concept is to improve the early warning ability of post-loan risks through end-to-end learning of information in the relationship graph of small and micro enterprises.
[0045] For the business scenario of credit risk prediction for small and micro enterprises, a small and micro enterprise association relationship graph is established based on natural months, with small and micro enterprise data samples (including small and micro enterprises and related enterprises and individuals screened according to certain rules) as nodes, and the capital flow relationship, investment relationship, guarantee relationship and person-enterprise relationship in the "current month" as edges. The association relationship graphs of different months share parameters during the model training process, laying a good data foundation for the predictive ability of training graph neural network models.
[0046] Based on a message-passing framework and semantically homogeneous networks, a graph convolutional network node classification method is designed, incorporating heterogeneous semantic information and network structure information. First, the complete small and micro enterprise relationship graph (network) is simplified and converted into multiple semantically homogeneous graphs. A node-level attention mechanism is then applied to the semantically homogeneous graphs to learn neighbor weights and generate heterogeneous semantic representations for the nodes. Semantic-level attention is then used to aggregate node representations under different semantics. Finally, a fully connected layer is used to project the node representations into a category space, calculate the category probability distribution of the nodes, and determine the node's category label based on the category probability distribution.
[0047] The following combination Figures 1-13 The embodiments of the present invention are described in detail.
[0048] Before applying the small and micro enterprise credit risk prediction model provided by the present invention, it is necessary to train it with data to enable it to have prediction capabilities.
[0049] According to a graph neural network-based credit risk prediction method for small and micro enterprises provided by the present invention, a small and micro enterprise association relationship graph is constructed based on small and micro enterprise data samples of purchased loan products. Specifically, the method includes: obtaining small and micro enterprise data samples of purchased loan products that meet preset conditions; establishing a time window for small and micro enterprise data samples; determining the prediction variables and target variables of small and micro enterprises as nodes; based on the time window, determining the association relationships between small and micro enterprises according to the prediction variables; and constructing a small and micro enterprise association relationship graph based on the association relationships.
[0050] Specifically, a modeling dataset (data samples) is required before training the model.
[0051] In one embodiment provided by the present invention, the modeling data set is selected as a basis from small and micro-enterprise merchant customers who have purchased merchant loan products in the bank's business quick loan product system. The preset conditions that meet the specific requirements are as follows: a. The merchant transaction volume of small and micro business merchants in the past three months is greater than 0, which means that these merchants have had business activities in the past three months.
[0052] b. The merchant's transaction volume in the past 12 months is greater than 200,000. This condition ensures that the selected merchants have achieved a certain level of business scale in the past year.
[0053] c. The merchant has had transaction volume for two consecutive years, which reflects the stability and sustainability of the merchant's operations.
[0054] d. The merchant's transaction volume in the past 12 months has not declined by more than 80% compared to the previous 12 months, which means that the merchant's operating conditions in the past year have not declined significantly compared to the previous year.
[0055] e. The merchant's status is normal. Specifically, the enterprise's abnormal merchant transaction volume does not exceed 25% of the total merchant transaction volume under its name. This condition ensures that the selected merchants are in good overall operating condition and do not have excessive abnormal transactions.
[0056] In one embodiment of the present invention, after obtaining a small and micro enterprise data set, the small and micro enterprise data set is divided into training samples and verification samples; based on the small and micro enterprise data sample modeling time window, the training samples and verification samples within the observation period window, as well as the training samples and verification samples within the observation period window are determined.
[0057] Specifically, a time window is modeled.
[0058] Based on the above modeled dataset, the observation period and performance period of the modeled training samples and validation samples are specified as follows: For training samples: Observation period: January 1, 2018 to December 31, 2019; Performance period: January 1, 2020 to February 29, 2020.
[0059] This means that during the observation period, the present invention collected all relevant merchant data and used it to train the model. The performance period of the model was from January 1, 2020 to February 29, 2020, during which the model's predictive ability was verified.
[0060] For the validation sample: Observation period: July 1, 2018 to June 30, 2020; Performance period: July 1, 2020 to August 31, 2020.
[0061] Similar to the training sample, the observation period of the validation sample covers a longer time range, from July 1, 2018 to June 30, 2020. The performance period is from July 1, 2020 to August 31, 2020, which is used to verify the prediction accuracy and stability of the model.
[0062] Such time division ensures that the performance of the model in different time periods can be fully evaluated, and the model can be adjusted and optimized in a timely manner to improve its prediction accuracy and stability.
[0063] Based on the above embodiment, the “good” and “bad” performances of small and micro enterprises within the performance period window are used as target variables; and the data samples of small and micro enterprises within the observation period window are labeled with the target variables.
[0064] Specifically, define the target variable (Y variable).
[0065] Based on the above modeling, the "good" and "bad" performance of customers during the performance period are used as label Y values, and "bad" customers are clearly defined and labeled based on the customer's situation during the observation period. The details are as follows: (1) If a small and micro business merchant has overdue payments on a merchant loan product in the Quick Loan product system, the merchant will be marked as a “bad” customer.
[0066] (2) If a merchant’s corporate loan from a bank is classified as substandard, doubtful, or loss-making, the merchant is labeled as a “bad” customer.
[0067] (3) If at the end of the observation period, the merchant’s credit report of the People’s Bank of China contains substandard, doubtful or loss-making loans, the merchant will also be regarded as a “bad” customer.
[0068] (4) If at the end of the observation period, a merchant’s credit report from the People’s Bank of China shows a negative score (non-performing loan amount + bad debt amount + compensation amount is greater than 0), the merchant will also be considered a “bad” customer.
[0069] (5) If a business owner’s bank credit card is overdue for more than 30 days, the business owner will be labeled as a “bad” customer.
[0070] (6) If a business owner’s personal loan at a bank is classified as substandard, doubtful, or loss-making, the business owner is labeled as a “bad” customer.
[0071] (7) If at the end of the observation period, the business owner has a loan rated as substandard, doubtful or loss-making in the People's Bank of China's credit report, the business owner will also be regarded as a "bad" customer.
[0072] (8) If at the end of the observation period, the business owner’s credit report on the People’s Bank of China shows a negative credit score (non-performing loan amount + bad debt amount + compensation amount is greater than 0), the business owner will also be regarded as a “bad” customer.
[0073] According to a small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention, the correlation between small and micro enterprises is determined according to the prediction variables, and a small and micro enterprise correlation diagram is constructed according to the correlation. Specifically, the method includes: determining the attribute data on each node in the small and micro enterprise correlation diagram according to the small and micro enterprise data; the node represents the small and micro enterprise; the attribute data corresponds to the characteristic value of the prediction variable; based on the time window, the correlation between the nodes is determined according to the attribute data; and the small and micro enterprise correlation diagram is generated according to the correlation between the nodes.
[0074] Specifically, define the predictor variables (X variables).
[0075] Based on the above modeling, and centered around small and micro-enterprise credit risk control, we aim to design and extract data indicators and relationships from multiple dimensions. These dimensions cover not only the relevant attributes of individuals (business owners) and legal entities, but also those of legal entities. In this way, the present invention can accurately label small and micro-enterprises and their individuals and legal entities, and construct a relationship map for these enterprises accordingly.
[0076] To ensure the model accurately reflects the actual situation, the present invention has deeply mined a large amount of data to find key indicators that can truly reflect the business conditions, financing needs and repayment ability of merchants. After thoroughly studying the characteristics of small and micro enterprises and the current status of the bank's small and micro merchant Yidai corporate card business, the following indicators were selected: (1) In order to solve the problem that the governance structure of small and micro enterprises is relatively simple, and there are few information disclosure channels and the reliability is poor, the present invention adopts the following measures: First, this model collects merchant acquisition details from the 12 months prior to card issuance. By analyzing the number and amount of transactions, it assesses the business's operating scale and stability. This provides a more accurate understanding of the business's operating status and thus determines its credit risk. Furthermore, the model summarizes and analyzes the number of card-swipe customers at the merchant, breaking down the statistics by card number. Furthermore, it analyzes the average number and amount of transactions per card to prevent merchants from falsifying or inflating transaction volumes through POS (Point of Sale) transactions.
[0077] To further enhance the accuracy of the assessment, the present invention also collects merchant bank account statements. By analyzing information such as the number of transactions, transaction amounts, and counterparties in these statements, it can assist in determining the size, operational stability, and authenticity of the trade background of the enterprise.
[0078] By comprehensively applying these methods, the present invention can more comprehensively understand the operating conditions and credit risks of small and micro enterprises, thereby providing strong support for credit decision-making.
[0079] (2) In view of the characteristics of small and micro enterprises, which have a wide range of industries and flexible and diverse business operations, the present invention adopts the following measures in the model: First, we gather sufficient basic information about the company as a foundation for assessing its credit risk. Furthermore, we incorporate the company's industry category into the model to account for differences in business continuity and seasonality. This allows for a more accurate assessment of a company's risk profile within a specific industry context.
[0080] At the same time, the present invention analyzes the daily transaction records of merchants, especially their upstream and downstream enterprises, and strives to screen out high-quality target customers with potential by marking merchants of core banking enterprises or large-scale high-quality enterprises in the same industry.
[0081] Through such a method, the present invention can better grasp the risk characteristics of small and micro enterprises in different industries and operating environments, and provide them with financial services that better meet their needs.
[0082] (3) In order to solve the problem of complex relationships between natural persons and the existence of "multiple credit grants", the present invention adopts the following measures: First, this model uses the legal representative as the basis for overall credit assessment for all businesses under the same legal representative. This approach prevents the risk of excessive credit limits arising from separate card issuance to multiple affiliated businesses under the same legal representative.
[0083] This model also integrates with the GCMS (Comprehensive Credit Management System) to collect relevant data, such as a company's bank credit rating and annual credit approvals. This data can be used to conduct comprehensive credit assessments for small and micro businesses with existing bank loans. By integrating data and information from different systems, this invention provides a more comprehensive understanding of a company's credit status and risk level, enabling more accurate credit decisions.
[0084] Through the above measures, the present invention can more effectively solve the problems of complex natural person relationships and multiple credit grants, and improve the accuracy and reliability of credit assessment.
[0085] (4) In response to the relatively weak comprehensive debt repayment capacity of small and micro enterprises, the present invention adopts the following measures: First, the present invention collects credit data from the People's Bank of China for small and micro enterprises and their legal representatives, including historical defaults over the past five years. By analyzing this data, the present invention can assess the credit status of the enterprise and its legal representative, thereby determining their debt repayment capacity.
[0086] The present invention also collects bank asset data for the enterprise and its legal representative. Combining the enterprise's bill collection details and settlement account flow information, the present invention can assess the stability of the enterprise's primary repayment source. Considering that many small and micro enterprises are privately owned and some settlements are conducted through individual accounts, the present invention also specifically collects bank asset data for the enterprise's legal representative as an auxiliary indicator for assessing the enterprise's repayment capacity.
[0087] The above-collected indicators are used as attribute data of nodes in the small and micro enterprise association diagram.
[0088] According to a small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention, the association relationship between nodes includes a capital flow relationship; the capital flow relationship between nodes is determined based on attribute data, specifically including: querying and incorporating small and micro enterprises that have indirect capital transaction relationships with the small and micro enterprise center sample into the small and micro enterprise association relationship graph; aggregating the capital flow transaction data between two nodes in the small and micro enterprise association relationship graph by month, and generating a connecting edge based on this as the capital flow relationship between the nodes.
[0089] Specifically, with the small and micro enterprise data sample as the center, the logic for extracting association relationships is as follows: (1) Determine the capital flow relationship. In order to build the capital flow relationship network of small and micro enterprises, the following steps were taken: 1. Determine the central sample: Take the small and micro enterprise data sample as the center, and this center point is the starting point of the analysis.
[0090] 2. Identify Second-Degree Financial Transaction Counterparties: During the observation window, identify merchants with indirect financial transaction relationships with the central sample. For example, if merchant A has a financial transaction with merchant B (denoted as AB), and merchant B has a financial transaction with merchant C (denoted as BC), then merchants A and C are considered to have a second-degree financial transaction counterparty (AC), and these second-degree counterparties will be included in the connected network.
[0091] 3. Data Aggregation and Edge Generation: The transaction data for fund flows between two merchants is aggregated (summed) by month, and an edge is generated based on this. This means that within a month's network, there is only one transaction edge between the two merchants, avoiding data redundancy.
[0092] 4. Consider data scale and efficiency: Due to the sheer volume of financial transaction data and the relatively small sample size of small and micro-enterprise customers, traditional operations such as joins across all data tables not only place a heavy burden on computing resources but also generate significant data redundancy. To improve processing speed and efficiency, the database operation logic needs to be redesigned.
[0093] 5. Optimize database operations: By optimizing database query and storage mechanisms, secondary fund transaction data can be completed more quickly and efficiently, improving the speed and accuracy of data processing, thereby better constructing and analyzing the fund flow association network.
[0094] The specific operation logic is as follows Figure 1 As shown, in Figure 1 In the example, Customer A represents the small and micro-enterprise merchant customers who have purchased the merchant loan product from the bank's operating quick loan system. These customers constitute the small and micro-enterprise data sample used in the modeling process of this invention. Customer A's first-degree counterparties can be divided into three categories: (i) Merchants who are also small and micro loan customers of the bank and also belong to the customer A group. These transaction data are Figure 1 In the figure, it is represented as “edge 1”.
[0095] (ii) Customers who are not micro-enterprises and have loans within the bank are defined as customer B. These transaction data are Figure 1 In the figure, it is represented as "edge 2" and "edge 4".
[0096] (iii) Non-bank customers, defined as the customer set C. These transaction data are Figure 1 In the example, it is represented as “edge 3” and “edge 5”.
[0097] After identifying all first-degree counterparties, a similar approach can be used to identify all second-degree counterparties. For the customer set B, the logic is similar to that for the customer set A. This can be further expanded to include a new set of in-bank customers D (non-micro-enterprise merchants within the bank with loans and no direct transactions with customer set A) and a new set of out-of-bank customers E.
[0098] In addition, customer C can also conduct direct transactions with customers in multiple other sets. Specifically, it can directly trade with customers in customer B and customer D, as well as with other in-bank customer F that are not in sets A, B, or D.
[0099] Through the above logic, we can completely find all the second-degree trading counterparties of the target sample, as well as some of the second-degree and above trading counterparties.
[0100] It's worth noting that, in actual database operations, since no index is set for external customers, searching for internal counterparties using the C and E sets as indexes can significantly slow down queries. To avoid this, internal customers can be used as indexes to perform a reverse search for eligible external counterparties and related transaction data. This process can also include querying transaction data for DE and FE, further enhancing the density of the transaction network.
[0101] (2) Investment Relationship: With the data sample of small and micro-enterprise customers as the center, all investment data within the observation window are extracted. The counterparties involved in this investment data are considered as affiliated businesses and expanded into the association network. For each investment behavior, an edge is generated in the association network, connecting the two businesses. In this way, each investment behavior will leave an edge in the association network, indicating the investment relationship between the two businesses.
[0102] (3) Guarantee relationship: With the data sample of small and micro-enterprise customers as the center, the data of all guarantee circles in which they are located during the corresponding observation period are extracted. Merchants in these guarantee circles are regarded as related enterprises and expanded into the related network.
[0103] To better analyze relationships within a guarantee circle, it's necessary to break it down so that each target customer has an edge connected to any other merchant within the guarantee circle. This ensures that the relationships between each merchant are clearly represented, helping to reveal potential risks and opportunities.
[0104] (4) Person-Enterprise Relationship: With the data sample of small and micro-enterprise customers as the center, all the person-enterprise related data within the corresponding observation window are extracted. These related categories include the target customer's second person in charge, legal representative, insurance legal beneficiary, financial director, shareholder, general manager, unit contact person, chairman, and other persons in charge. These person-enterprise related opponents are expanded into the network, and an edge is generated between each related node and the central sample node.
[0105] The constructed small and micro-enterprise association graph primarily contains two main data components: First, the attribute data for each node in the association graph. This data corresponds to the eigenvalues of the X variables themselves and describes the characteristics and properties of each node. Second, the edge data in the graph, that is, the associations between the X variables. This data primarily describes the relationships between enterprises. The node attribute data structure is shown in Table 1.
[0106] Table 1
[0107] Figure 1 The data structures of different types of edges are as follows Figure 2-Figure 5 As shown: 1. Capital Flow Relationship Figure 2 The data structure of the capital flow relationship. When processing the counterparty category, two main methods are used to determine whether the counterparty belongs to a compliant industry, as follows: Regular expression matching: Use regular expressions to match whether the competitor's name contains industry keywords.
[0108] Screening for specific industry codes: In addition to name matching, we further determine counterparty compliance by screening for specific industry codes. This involves screening and analyzing industry codes to ensure that counterparties are aligned with compliant industries.
[0109] 2. The data structure of investment relationship can be found in Figure 3 .
[0110] 3. The data structure of the guarantee relationship can be found in Figure 4 .
[0111] 4. The data structure of the person-enterprise relationship can be found in Figure 5 .
[0112] For linked individual customers, all their fund flow transaction data within the corresponding observation window is also extracted. This data is processed in the same manner as the fund flow relationship data described above. This means that individual customer fund flow transaction data is aggregated and analyzed to identify their counterparties. Furthermore, this data is filtered for regular expression matching and specific industry codes.
[0113] After obtaining the attribute data and edge data of the above-mentioned small and micro enterprises, they need to be processed to avoid redundancy in the constructed small and micro enterprise association graph.
[0114] In the data samples of small and micro enterprises, there are common duplications such as one customer corresponding to multiple merchants, multiple households with a single card, and multiple households with multiple cards, as well as missing or abnormal feature values due to information asymmetry among small and micro enterprises. It is necessary to make corresponding adjustments mainly for the different characteristics of different data. The main methods include record deduplication, feature filling and exception handling.
[0115] (1) Record deduplication: In the original record of fund flow data, if both parties of the transaction are in-bank customers, the same transaction will be recorded twice. The difference between the two records is the different lending and borrowing directions. Therefore, it is necessary to unify the lending and borrowing directions of all records and remove duplicate records. This can avoid data redundancy and errors and ensure that each transaction is recorded only once. The fund flow data deduplication logic that matches the data acquisition logic is as follows: Figure 6 During the data retrieval process, you need to specify the primary key customer for the query and the debit / credit direction. The primary key customer can only be selected from customers registered within the bank, not from outside the bank. Figure 6 Account 1 represents the primary key customer category when querying, and account 2 represents the counterparty to which the transaction belongs. Figure 1 In the category, "1, 2" indicates the lending direction selected during the query, and "internal, external" indicates whether the customer is an internal or external customer. Figure 6 The first row in the query represents all transaction records with Class A customers as the primary key, whose counterparties are also Class A customers, and the transaction direction is "debit", that is, Figure 1 All edges marked as category 1 in the table, and ensure that there is no duplication of data. Figure 6 Each line of instructions in the table can be used to take numbers in turn. Figure 1 All 25 types of edges are extracted without duplication.
[0116] (2) Feature filling: To address the problem of missing features, targeted filling methods need to be adopted according to the reasons for the missing features. Commonly used processing methods include deleting samples, deleting feature variables, and filling missing values. Specifically, for non-bank customers, there is no ID primary key in the data. Since the number of samples involved is large, it is not appropriate to delete samples for processing. To solve this problem, the missing values can be filled by encrypting the customer account with MD5 instead of the ID primary key during modeling. By encrypting the customer account with MD5, a unique identifier can be generated to fill the missing ID primary key value. This method not only avoids data loss caused by deleting a large number of samples, but also ensures the integrity and accuracy of the data.
[0117] In the case of missing variables of the People's Bank of China's credit investigation, based on past experience and business recommendations, the samples are divided into two categories for processing: a) Credit Information: This sample type refers to individuals or legal entities with no records in the PBOC's credit reporting system. To distinguish this sample from other samples, we recommend using '-1' to fill in the missing credit information variables for this sample type.
[0118] b) Credit Novice: This type of sample refers to individuals who have credit records in the PBOC credit reporting system, but all extracted fields contain missing values. In this modeling, no special treatment was given to this type of sample; instead, the missing values were filled in as usual.
[0119] (3) Abnormal processing: The processing of data outliers is also very important. First, it is necessary to combine the results of business and exploratory data analysis to determine the outliers. Secondly, some commonly used statistical methods can be used to further identify the outliers. Then, the outliers are processed accordingly to eliminate the impact of outliers on model training.
[0120] For example, transaction amounts in fund flow transaction data should be positive. If a negative transaction amount occurs, the fund flow transaction data should be deleted. Some counterparty IDs contain characters such as "business occurrence"; such records should be deleted. If characters such as "3.10E+14" appear in the company ID field, this may be due to incorrect field formatting when entering the original data, resulting in scientific notation. Such non-compliant records should be deleted.
[0121] According to a small and micro enterprise credit risk prediction method based on graph neural network provided by the present invention, the association relationship graph of the small and micro enterprises to be tested is simplified to obtain a simplified association relationship graph, specifically including: the small and micro enterprise association relationship graph is a directed multi-heterogeneous graph; the directed multi-heterogeneous graph is composed of a directed graph, a multi-graph and a heterogeneous graph; the directed graph is converted into an undirected graph; and the multi-heterogeneous graph is split into several semantically homogeneous networks.
[0122] Specifically, after constructing the association graph based on the aforementioned predictor variables and edge data, a series of preprocessing operations are required to better facilitate subsequent analysis and modeling. The association graph for small and micro enterprises can be viewed as a directed multi-heterogeneous graph, meaning that the graph is not only directional but also contains nodes and edges that encompass multiple categories.
[0123] Specifically, the capital flow relationship and investment relationship in the diagram are both directed edges, with clear directionality. The direction of the capital flow relationship represents the inflow and outflow of funds, while the direction of the investment relationship represents the relationship between the investor and the investee. Furthermore, since two customers may have multiple relationships, this may result in multiple edges connecting two nodes.
[0124] Furthermore, the nodes in this graph include not only small and micro businesses but also potentially other businesses related to these businesses. Therefore, the edges in the graph also encompass multiple categories. This makes the entire graph a directed, multi-heterogeneous graph. To account for the multi-dimensional information contained in the graph, some preprocessing is required.
[0125] Given the properties of directed graphs, risk transmission may not completely align with capital flows or investment relationships; risks can be transmitted in different directions. This means that risks can be transmitted not only from upstream to downstream companies, but also from downstream to upstream companies. To ensure the correct transmission of risk information, the directed graph must first be converted to an undirected graph. This conversion is achieved by reversing all directed edges and adding them to the graph. This way, regardless of the direction of risk transmission, the undirected graph can accurately represent it.
[0126] According to a graph neural network-based credit risk prediction method for small and micro enterprises provided by the present invention, multiple heterogeneous graphs are split into several semantically homogeneous networks, specifically including: converting the heterogeneous graphs into single-node multi-relationship heterogeneous networks; dividing the single-node multi-relationship heterogeneous networks into several relationship graphs according to different relationship categories; and treating each relationship graph as a semantically homogeneous network.
[0127] Specifically, for multi-heterogeneous graphs, they can be viewed as networks containing multiple semantic information. In this context, different edge categories of the network can be viewed as different semantic information. In order to better fit the message passing framework, it is necessary to simplify these multi-heterogeneous networks into several semantically homogeneous networks while retaining multiple semantic information, such as Figure 7 The following is a logical diagram of the relationship network splitting.
[0128] First, since the external affiliated merchants of the bank only serve as connecting pathways in the graph and have no node attribute data, the heterogeneous network can be converted into a single-node multi-relationship heterogeneous network, ignoring the distinction between merchants within the bank and those outside the bank.
[0129] Next, let's consider the relationship network G, assuming that the set of relationships it contains is {Φ_0, Φ_1, […, Φ]_R}. Based on the different relationships, G is divided into several relationship graphs G_(Φ_0), G_(Φ_1), …, G_(Φ_R). Each relationship graph G_(Φ_r) contains only the edges with the relationship Φ_r in the original network. In G_(Φ_r), the edge type between any two connected nodes is the same, and there is at most one edge. Therefore, G_(Φ_r) is a semantically homogeneous network. This approach allows us to preserve various semantic information in a heterogeneous network through the different relationship types between nodes. This not only simplifies the network structure but also enables better processing and utilization of this semantic information.
[0130] According to a small and micro enterprise credit risk prediction method based on a graph neural network provided by the present invention, a small and micro enterprise credit risk prediction model is trained according to a small and micro enterprise association relationship graph to obtain a trained small and micro enterprise credit risk prediction model, specifically comprising: simplifying the small and micro enterprise association relationship graph to obtain a semantic homogeneous network, and projecting all nodes in the semantic homogeneous network into a unified feature space; in each semantic homogeneous network, the small and micro enterprise credit risk prediction model learns the first importance weights of different neighbor nodes through a node-level attention mechanism; a neighbor node is a node directly connected to the current node; the small and micro enterprise credit risk prediction model jointly learns the second importance weight of each semantic homogeneous network through a semantic-level attention mechanism; based on the first importance weight and the second importance weight, the small and micro enterprise credit risk prediction model is enabled to perform classification prediction on all nodes in the small and micro enterprise association relationship graph through a multi-layer perceptron; the classification loss is calculated according to the classification prediction result, and the small and micro enterprise credit risk prediction model is optimized according to the classification loss, and a trained small and micro enterprise credit risk prediction model is obtained after the optimization is completed.
[0131] Specifically, the overall structure of the small and micro enterprise credit risk prediction model based on graph convolutional neural network consists of three parts: (a) Node Feature Projection and Neighbor Node Importance Learning: First, all nodes in the enterprise association network are projected into a unified feature space. Within each semantically homogeneous network, a node-level attention mechanism is applied to learn the importance weights of different neighboring nodes (nodes directly connected to the current node). This method determines the contribution of each node to the current node, and the node's hidden layer representation is obtained by weighted summation of the neighboring nodes (the set of neighboring nodes).
[0132] (b) Semantic-level attention mechanism and node representation fusion: Next, the semantic-level attention mechanism is used to jointly learn the weights of each semantically homogeneous network. This step aims to understand the importance of different semantics (i.e., different types of edges or relationships) in the network. The node latent representations under different semantics are then fused to obtain the final node representation vector. This vector comprehensively reflects the characteristics and role of the node in credit risk prediction.
[0133] (c) Classification Prediction and Optimization: Finally, a multi-layer perceptron (MLP) is used to classify the nodes. Based on these predictions, the classification loss is calculated and end-to-end optimization is performed. The goal of this step is to minimize the difference between the predicted results and the actual credit risk, thereby enabling the model to more accurately predict the credit risk of small and micro enterprises.
[0134] The application of node-level attention mechanism and semantic-level attention mechanism in the present invention is introduced below.
[0135] 1) Node-level attention mechanism Before aggregating neighbor node information from a semantically homogeneous network for each node, it is important to note that the neighbors of each node in different semantic networks play different roles and exhibit different importance when learning a specific task. Based on this, the present invention introduces a node-level attention mechanism. The purpose of this mechanism is to learn the importance of neighbors under the same semantics to each node in the associated network. In this way, the neighbor nodes that are most critical to the current node can be identified and given greater weights. Then, these important neighbor representations are aggregated to form the representation of the node's hidden layer.
[0136] First, feature transformation is performed using a shared multilayer perceptron network. Theoretically, multilayer perceptrons have powerful representation capabilities and can approximate any measurable function. To achieve a unified representation of node features, a specific multilayer perceptron model is designed to project different node features into the same feature space. This process can be expressed as follows: in, and Node The original feature representation and the projected feature representation. Obviously, and They only contain information about a single node itself, but not the structural information of the associated network.
[0137] In order to better understand the importance of nodes in the association network, this paper introduces the self-attention mechanism, which can learn the importance relationship between nodes in the same semantic network. Connected node pairs , node-level attention Can evaluate the node For Node Importance Based on the association relationship Node pairs The node-level attention of the node pair is calculated as follows: in, Represents a deep neural network that implements a node-level attention mechanism. Under this condition, all pairs of nodes based on this relationship share the same identity network. , that is, they share in the same semantic network .
[0138] The above formula shows that the node pair The importance of each other depends only on their associated categories and their respective feature hidden layer representations. The importance of is asymmetric, that is, the nodes For Node Importance and nodes For Node There can be large differences in the importance of . This suggests that node-level attention can preserve asymmetry, which is a key property of heterogeneous graphs.
[0139] Then, the structural information is injected into the model through masked attention. The masked attention mechanism means that only the nodes Importance weight ,in Representation node In relationship The neighboring nodes under (including itself).
[0140] After obtaining the importance between all neighbor node pairs, these importances are normalized by the softmax function to obtain the aggregation weight coefficient : in represents a nonlinear activation function, Represents vector concatenation operation, It is an association relationship The parameter vector of the node-level attention mechanism. Similar to the previous analysis, the node pair The weight coefficient depends only on their relationship categories and their respective feature representations; and the weight coefficient This asymmetry is not only due to the different splicing order in the molecule, but also because they have different neighboring nodes, which further leads to significant differences in the contributions of the nodes to each other.
[0141] Final Node Based on association The hidden layer representation of can be aggregated through the feature representation of the neighbors and the corresponding aggregation weight coefficients as follows: in is the node learned by the model For the association relationship The hidden layer representation of . Figure 8 (a) briefly illustrates the node-level aggregation process. In this process, the hidden layer representation of each node is aggregated from its neighbors. It is generated for a single association relationship, so the process is able to capture a specific kind of semantic information.
[0142] Since heterogeneous graphs are scale-free, i.e. the connectivity (degree) between nodes is severely unevenly distributed, to address this issue, node-level attention can be extended to multi-head attention, making the training process more stable.
[0143] Specifically, repeat the above node-level attention operation Each operation will generate a hidden layer and learn The hidden layer representations are concatenated into a semantically specific hidden layer representation: The set of association relationships in a given network , after inputting the node feature representation into the node-level attention, we can obtain The node hidden layer representation of group semantics is expressed as , each group contains the hidden layer representations of all nodes in the network under the corresponding semantics.
[0144] 2) Semantic-level attention mechanism In heterogeneous graphs, nodes may contain multiple semantic information. However, semantically specific node-level attention mechanisms can only reflect one aspect of node information. To learn a more comprehensive node representation, it is necessary to integrate the multiple semantics contained in heterogeneous graphs. Therefore, semantic-level attention is used to automatically learn the importance of different semantics.
[0145] Learned from node-level attention The group semantic specific node hidden layer representation is taken as input, and each semantic weight learned by the model It can be expressed as: in Represents a deep neural network that performs semantic-level attention, which can capture various types of semantic information behind heterogeneous graphs.
[0146] In order to calculate the importance of each semantic category, the semantic-specific hidden layer representation of each node is first transformed by a nonlinear transformation function. Then, the transformed representation and the semantic-level attention parameter vector are calculated. The similarity is used as the importance of the specific semantic representation of each node. Finally, all nodes under the same semantics are averaged to obtain the importance of each semantics.
[0147] The relationship between the network (i.e. semantics), and its semantic importance is expressed as The calculation process is as follows: in is the weight matrix, is the bias vector, is the semantic-level attention parameter vector. To make meaningful comparisons, all semantic categories and semantic-specific node representations share all the above parameters. After obtaining the importance of each semantic, it is also normalized by the softmax function.
[0148] Semantics The weight of , can be obtained by normalizing the above importance of all semantics using the softmax function: The weights of different semantics can be understood as different semantics The degree of contribution to downstream tasks (such as credit risk prediction). Obviously, The higher the weight, the more semantic The learned semantic weights are used as coefficients to fuse the semantically specific node hidden layer representations to obtain the final node representation matrix , as shown below: Figure 8 (b) briefly describes the semantic-level aggregation process. The final node representation is the aggregation of all the hidden layer representations of a specific semantic. This example is a node classification problem, and the cross-entropy loss function is used to complete the model training process. Specifically, the cross-entropy between the true value and the predicted value of all labeled nodes is minimized: in are the parameters of the classifier, is the set of node indices with labels, and is the label and node representation of the label node. Under the guidance of labeled data, the model is optimized through back propagation and the category of the node is predicted.
[0149] like Figure 9 Shown is a structural diagram of small and micro enterprise credit risk prediction based on graph neural network constructed based on node-level attention mechanism and semantic-level attention mechanism.
[0150] The business scenario for this modeling is to predict post-loan defaults. Taking this business application into consideration, we decided to use the ROC value as the primary evaluation metric for model comparison. In principle, the model with the highest ROC value is considered the champion model. In addition to the ROC value, this paper also uses the KS, recall, and precision metrics as auxiliary evaluation indicators.
[0151] (1) Model’s ability to distinguish Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC) are often used to evaluate the performance of a binary classifier, primarily reflecting the classifier's ability to rank samples. The AUC essentially indicates the probability that, for a randomly selected positive and negative sample, the classifier's probability of assigning a positive probability to the positive sample is greater than the probability of assigning a positive probability to the negative sample.
[0152] When evaluating different models, plotting their ROC curves on the same coordinate system allows for a more intuitive comparison of their performance. The ROC curve near the upper left corner represents the classifier with the highest accuracy. Furthermore, ROC curves can reflect not only a model's discriminatory power but also its ranking ability.
[0153] When the volume of business after an early warning is too high and the account manager cannot investigate all warnings one by one, they can prioritize investigating customers with a high probability of default. In this case, a model with strong sorting capabilities is needed.
[0154] (2) KS (Kolmogorov-Smirnov) is used to evaluate the risk differentiation ability of the model. This indicator measures the difference between the cumulative distribution of positive and negative samples, and measures the difference between the positive and negative sample distributions from a probability perspective. The larger the cumulative difference between positive and negative samples, the larger the KS indicator, and the stronger the risk differentiation ability of the model.
[0155] On this modeling dataset, the AUC and KS values of each model are as follows: Figure 10ICBC-LR is a tuned credit risk prediction model, and its results are derived from the validation set in the model design report. The results of other models are reproduced based on the same data used in this modeling.
[0156] from Figure 10 The results show that compared to traditional machine learning models, the small and micro enterprise credit risk prediction model proposed in this solution achieved an average improvement of 2.13% in AUC. Compared to the optimized ICBC-LR model, the model presented in this paper achieved a 1.22% improvement in AUC. This modeling effort was primarily data-driven, employing only simple feature engineering techniques. For this reason, the logistic regression (LR) model's performance did not meet expectations, falling short of ICBC-LR. To further improve model performance, subsequent research could leverage existing industry-leading feature engineering techniques, using them as input features for deep learning to further leverage the advantages of deep learning and achieve even better prediction results.
[0157] Figure 11 This is a flow chart of the method for predicting credit risk of small and micro enterprises based on graph neural network provided by the present invention. Figure 11 As shown, the method includes the following steps: S1110. Obtain original data of the small and micro enterprises to be tested within the loan term, perform data preprocessing on the original data of the small and micro enterprises to be tested, and obtain the data of the small and micro enterprises to be tested.
[0158] First, during the loan's lifespan, comprehensive raw data on the small and micro enterprises under assessment is obtained through multiple channels. This data includes, but is not limited to, the enterprise's financial statements, tax records, transaction flows, market conditions, industry trends, and the personal credit history of the legal representative. This data can provide a comprehensive, multi-dimensional profile of the enterprise, facilitating a deeper understanding of the enterprise's operating conditions and loan risks.
[0159] After collecting raw data, data preprocessing is performed to ensure its accuracy and validity, thereby improving the precision of subsequent credit assessments. This preprocessing process includes data cleaning, data transformation, and data reduction. Specifically, it involves removing duplicate, invalid, or abnormal data, normalizing or standardizing data, and reducing data dimensionality and complexity through feature selection and extraction.
[0160] After data preprocessing, the optimized and organized data on small and micro enterprises (SMEs) under test is not only of higher quality but also easier to conduct subsequent credit assessments and risk analysis. Through in-depth mining and analysis of this data, we can more accurately determine the credit status and risk level of SMEs, providing stronger data support for loan decisions.
[0161] S1120. Construct a correlation diagram of the small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested.
[0162] After obtaining and preprocessing the data of the small and micro enterprises to be tested, these data are used to construct a complex network relationship between small and micro enterprises and other entities (such as suppliers, customers, partners, financial institutions, etc.).
[0163] First, obtain the capital flow relationship, investment relationship, guarantee relationship and human-enterprise relationship of small and micro enterprises.
[0164] Fund flow relationships: Analyze the fund transaction records of small and micro enterprises during the loan period and identify fund transactions between them and other enterprises and individuals. This includes direct fund transaction counterparties and indirect fund transaction relationships formed through intermediaries.
[0165] Investment relations: Extract investment behaviors between small and micro enterprises and other entities from investment data, clarify investors and investees, and form an investment relationship network.
[0166] Guarantee relationship: Analyze the guarantee circle that small and micro enterprises participate in during the loan or other financing process, clarify the roles of each party in the guarantee relationship, and establish a guarantee network.
[0167] Human-enterprise relationship: Integrate the personal relationship information of key personnel such as legal persons, shareholders, and senior executives of small and micro enterprises, including family relationships, professional connections, etc., as well as the relationship between these personnel and small and micro enterprises.
[0168] Association diagram construction: Node definition: Small and micro enterprises, legal persons, shareholders, executives, partners, customers, etc. are regarded as nodes in the graph.
[0169] Edge definition: Define the edges between nodes based on capital flow, investment, guarantee, and enterprise-person relationship, and specify the type and direction (undirected or directed) of the edge.
[0170] Consider the time dimension: Build a monthly or quarterly SME relationship graph to reflect the dynamic changes in relationships. Graphs at different time points share some parameters to support predictive analysis over time series. Store the constructed SME relationship graph in a graph database or dedicated graph processing framework to support fast graph traversal and querying.
[0171] S1130. Simplify the association relationship diagram of the small and micro enterprises to be tested to obtain a simplified association relationship diagram.
[0172] Specifically, converting the original directed multi-heterogeneous graph (containing multiple types of nodes and edges) into an undirected graph can simplify the complexity of risk transmission and ensure that no matter which direction the risk propagates, it is correctly represented in the graph. This is done by reversing all directed edges and appending them to the graph, thus forming an undirected graph.
[0173] The nodes in the graph are classified according to different relationship types (such as capital flow, investment, guarantee, and enterprise-person relationships). Based on these relationships, the original heterogeneous network is divided into multiple semantically homogeneous networks. Each homogeneous network contains only edges of the corresponding relationship type, and any two connected nodes can only have one edge.
[0174] The heterogeneous network is converted into a single-node multi-relationship network, that is, the attributes of the nodes are not considered and only the association relationships between the nodes are retained.
[0175] Remove duplicate and unnecessary edges, such as multiple edges between the same pair of nodes due to multiple relationship types, and retain only one representative edge to simplify the network structure.
[0176] By simplifying the complex original small and micro enterprise relationship graph into multiple homogeneous subgraphs with clear semantics and concise structure, it facilitates the subsequent graph convolutional neural network model processing and also improves the efficiency and accuracy of the model.
[0177] S1140. Input the simplified correlation relationship graph into the trained small and micro enterprise credit risk prediction model to perform prediction, and obtain the default risk probability of the small and micro enterprise to be tested.
[0178] Specifically, the association graph data of the SMEs under test is converted into the required input format for the SME credit risk prediction model. This formatted association graph data is then passed as input to the trained SME credit risk prediction model. The model processes the input graph data through steps such as node feature projection, neighbor node importance learning (node-level attention mechanism), and semantic-level attention fusion with node representations. This generates a hidden layer representation for each node. A multi-layer perceptron is then used to classify and predict the node hidden layer representations to determine the default risk probability of the SME node. The SME credit risk prediction model outputs a default risk probability value for the SME under test, which reflects the likelihood of the enterprise defaulting during the loan period.
[0179] S1250. Determine the classification of the small and micro enterprises to be tested based on the default risk probability of the small and micro enterprises to be tested; the trained small and micro enterprise credit risk prediction model is trained based on the small and micro enterprise association relationship graph, and the small and micro enterprise association relationship graph is constructed based on the data samples of small and micro enterprises that have purchased loan products.
[0180] The primary application of this model is model early warning, with the application scenario being to categorize and label existing small and micro-enterprise credit loan customers. The specific application plan is as follows: Based on the model's predicted probability, recall rate, and business experience, customers are categorized into 14 tiers. The recall rate increases by 10 percentage points in the first nine tiers, while the recall rate increases by 2 percentage points in the last five tiers. This categorization approach is designed to ensure high early warning accuracy while facilitating targeted risk management and customer service.
[0181] If the model's predicted score (default probability) is in the top four tiers, immediate warnings can be issued for these customers. When the model's predicted score is in the top eight tiers, in addition to warnings, a more detailed assessment can be conducted by referencing certain business rules. For the remaining customers, if business personnel have sufficient time, exploratory warnings can be issued based on strong business rules.
[0182] Throughout the model's lifecycle, regular testing and monitoring are required to ensure that its results align with policy and risk appetite. If model results no longer meet these requirements, the model needs to be re-diagnosed and optimized to ensure it continues to meet business needs. This maintenance and testing process is key to ensuring the model's continued effectiveness and accuracy.
[0183] Diagnostic conditions: (1) Business threshold: When the business practice score of the list exceeds the originally set score threshold, it is necessary to re-evaluate and optimize the model, or adjust the threshold to ensure that it matches the current business needs.
[0184] (2) Model performance: If the model results (such as default probability prediction value, KS value, AUC value) drop significantly, then consider whether to reconstruct or re-optimize the model. This may be due to changes in data distribution or other external factors that cause performance degradation.
[0185] (3) Variable stability: If the distribution of the input variable values changes significantly, this may also affect the stability and accuracy of the model. In this case, it is also necessary to consider reconstructing or optimizing the model to ensure that it can adapt to the new data distribution and business environment.
[0186] This modeling effort employed a graph neural network to construct a predictive model for small and micro enterprise credit risk. This model aims to predict the probability of default for small and micro enterprise loan customers (both individual and legal entity credit loans) over their lifetimes, providing timely early warnings for potential defaulters. Application of this model significantly improved the precision and recall rates of post-loan risk checks, significantly enhancing the accuracy of predictions. This significantly facilitates relationship managers, helping them more accurately assess small and micro enterprise credit risk, while also saving banks human resource costs and effectively mitigating operational risks. This innovative model provides strong support for inclusive financial services, helping banks better serve small and micro enterprises and promote their stable development.
[0187] The following describes the small and micro enterprise credit risk prediction device based on graph neural network provided by the present invention. The small and micro enterprise credit risk prediction device based on graph neural network described below and the small and micro enterprise credit risk prediction method based on graph neural network described above can be referenced to each other.
[0188] like Figure 12 The present invention provides a device for predicting credit risk of small and micro enterprises based on a graph neural network, comprising: The module 1210 for acquiring data of the small and micro enterprises to be tested is used to acquire the original data of the small and micro enterprises to be tested during the loan term, perform data preprocessing on the original data of the small and micro enterprises to be tested, and obtain the data of the small and micro enterprises to be tested; The module 1220 for constructing a correlation diagram of small and micro enterprises to be tested is used to construct a correlation diagram of small and micro enterprises to be tested based on the data of small and micro enterprises to be tested; The small and micro enterprise credit risk prediction module 1230 is used to simplify the correlation diagram of the small and micro enterprises to be tested to obtain a simplified correlation diagram; input the simplified correlation diagram into the trained small and micro enterprise credit risk prediction model for prediction to obtain the default risk probability of the small and micro enterprises to be tested; determine the classification of the small and micro enterprises to be tested based on the default risk probability of the small and micro enterprises to be tested; the trained small and micro enterprise credit risk prediction model is trained based on the small and micro enterprise correlation diagram, and the small and micro enterprise correlation diagram is constructed based on the small and micro enterprise data samples of the purchased loan products.
[0189] Figure 13 An example of a physical structure diagram of an electronic device is shown below. Figure 13As shown, the electronic device may include: a processor 1310 , a communication interface 1320 , a memory 1330 and a communication bus 1340 , wherein the processor 1310 , the communication interface 1320 and the memory 1330 communicate with each other via the communication bus 1340 . The processor 1310 can call the logic instructions in the memory 1330 to execute a small and micro enterprise credit risk prediction method based on a graph neural network, which includes: obtaining the original data of the small and micro enterprise to be tested during the loan period, performing data preprocessing on the original data of the small and micro enterprise to be tested, and obtaining the data of the small and micro enterprise to be tested; constructing a correlation relationship graph of the small and micro enterprise to be tested based on the data of the small and micro enterprise to be tested; simplifying the correlation relationship graph of the small and micro enterprise to be tested, and obtaining a simplified correlation relationship graph; inputting the simplified correlation relationship graph into a trained small and micro enterprise credit risk prediction model for prediction, and obtaining the default risk probability of the small and micro enterprise to be tested; determining the classification of the small and micro enterprise to be tested based on the default risk probability of the small and micro enterprise to be tested; the trained small and micro enterprise credit risk prediction model is trained based on the small and micro enterprise correlation relationship graph, and the small and micro enterprise correlation relationship graph is constructed based on the small and micro enterprise data samples of purchased loan products.
[0190] Furthermore, the logic instructions in the aforementioned memory 1330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0191] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the small and micro enterprise credit risk prediction method based on graph neural network provided by the above methods. The method includes: obtaining the original data of the small and micro enterprises to be tested during the loan period, performing data preprocessing on the original data of the small and micro enterprises to be tested, and obtaining the data of the small and micro enterprises to be tested; constructing a correlation relationship graph of the small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested; simplifying the correlation relationship graph of the small and micro enterprises to be tested, and obtaining a simplified correlation relationship graph; inputting the simplified correlation relationship graph into a trained small and micro enterprise credit risk prediction model for prediction to obtain the default risk probability of the small and micro enterprises to be tested; determining the classification of the small and micro enterprises to be tested based on the default risk probability of the small and micro enterprises to be tested; the trained small and micro enterprise credit risk prediction model is trained based on the small and micro enterprise correlation relationship graph, and the small and micro enterprise correlation relationship graph is constructed based on the small and micro enterprise data samples of purchased loan products.
[0192] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the small and micro enterprise credit risk prediction method based on graph neural network provided by the above methods, the method comprising: obtaining the original data of the small and micro enterprises to be tested during the loan period, performing data preprocessing on the original data of the small and micro enterprises to be tested, and obtaining the data of the small and micro enterprises to be tested; constructing a correlation relationship graph of the small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested; simplifying the correlation relationship graph of the small and micro enterprises to be tested, and obtaining a simplified correlation relationship graph; inputting the simplified correlation relationship graph into a trained small and micro enterprise credit risk prediction model for prediction, and obtaining the default risk probability of the small and micro enterprises to be tested; determining the classification of the small and micro enterprises to be tested based on the default risk probability of the small and micro enterprises to be tested; the trained small and micro enterprise credit risk prediction model is trained based on the small and micro enterprise correlation relationship graph, and the small and micro enterprise correlation relationship graph is constructed based on the small and micro enterprise data samples of purchased loan products.
[0193] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0194] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for predicting credit risk of small and micro enterprises based on graph neural network, characterized by: include: Obtaining original data of the small and micro enterprises to be tested within the loan duration period, performing data preprocessing on the original data of the small and micro enterprises to be tested, and obtaining the data of the small and micro enterprises to be tested; Constructing a correlation diagram of the small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested; Simplifying the association relationship diagram of the small and micro enterprises to be tested to obtain a simplified association relationship diagram; Input the simplified correlation relationship graph into the trained small and micro enterprise credit risk prediction model to perform prediction, and obtain the default risk probability of the small and micro enterprise to be tested; Determining the classification of the small and micro enterprises to be tested based on the default risk probability of the small and micro enterprises to be tested; The trained small and micro enterprise credit risk prediction model is trained based on a small and micro enterprise association relationship diagram, and the small and micro enterprise association relationship diagram is constructed based on data samples of small and micro enterprises that have purchased loan products.
2. The method for predicting credit risk of small and micro enterprises based on graph neural network according to claim 1 is characterized in that: Based on the data samples of small and micro enterprises that have purchased loan products, a small and micro enterprise association relationship diagram is constructed, including: Obtain data samples of small and micro enterprises that have purchased loan products that meet pre-set conditions; Establishing a time window for the small and micro enterprise data sample; Identify small and micro enterprises as predictor variables and target variables for nodes; Based on the time window, determining the association relationship between small and micro enterprises according to the prediction variables; A small and micro enterprise association relationship diagram is constructed based on the association relationship.
3. The method for predicting credit risk of small and micro enterprises based on graph neural network according to claim 1 is characterized in that: The simplified processing of the association relationship diagram of the small and micro enterprises to be tested to obtain a simplified association relationship diagram specifically includes: The small and micro enterprise association relationship graph is a directed multi-heterogeneous graph; the directed multi-heterogeneous graph comprises a directed graph, a multi-graph and a heterogeneous graph; Convert a directed graph into an undirected graph; Split multiple heterogeneous graphs into several semantically homogeneous networks.
4. The method for predicting credit risk of small and micro enterprises based on graph neural network according to claim 1 is characterized in that: The small and micro enterprise credit risk prediction model is trained according to the small and micro enterprise association relationship diagram to obtain a trained small and micro enterprise credit risk prediction model, specifically including: Simplifying the small and micro enterprise association relationship graph to obtain a semantic homogeneous network, and projecting all nodes in the semantic homogeneous network into a unified feature space; In each semantically homogeneous network, the small and micro enterprise credit risk prediction model learns the first importance weights of different neighboring nodes through the node-level attention mechanism; the neighboring nodes are nodes directly connected to the current node; Through the semantic-level attention mechanism, the small and micro enterprise credit risk prediction model jointly learns the second importance weight of each semantically homogeneous network; Based on the first importance weight and the second importance weight, a small and micro enterprise credit risk prediction model is used to perform classification prediction on all nodes in the small and micro enterprise association relationship graph through a multi-layer perceptron; The classification loss is calculated based on the classification prediction results, and the small and micro enterprise credit risk prediction model is optimized based on the classification loss. After the optimization is completed, a trained small and micro enterprise credit risk prediction model is obtained.
5. The method for predicting credit risk of small and micro enterprises based on graph neural network according to claim 2 is characterized in that: Determining the association relationships between small and micro enterprises based on the prediction variables, and constructing a small and micro enterprise association relationship graph based on the association relationships specifically includes: Determining attribute data on each node in the micro-enterprise association graph based on the micro-enterprise data; the node represents a micro-enterprise; the attribute data corresponds to the characteristic value of the prediction variable; Based on the time window, determining the association relationship between nodes according to the attribute data; A small and micro enterprise association relationship graph is generated based on the association relationships between the nodes.
6. The method for predicting credit risk of small and micro enterprises based on graph neural network according to claim 3 is characterized in that: The splitting of multiple heterogeneous graphs into several semantically homogeneous networks specifically includes: Convert heterogeneous graphs into single-node multi-relation heterogeneous networks; Dividing the single-node multi-relationship heterogeneous network into a plurality of relationship graphs according to different relationship categories; Treat each relationship graph as a semantically homogeneous network.
7. The method for predicting credit risk of small and micro enterprises based on graph neural network according to claim 5 is characterized in that: The association relationship between the nodes includes a capital flow relationship; Determining the capital flow relationship between nodes based on the attribute data specifically includes: Query and include small and micro enterprises that have indirect capital transaction relationships with the sample of the small and micro enterprise center into the small and micro enterprise association relationship diagram; The capital flow transaction data between two nodes in the small and micro enterprise association relationship diagram are aggregated by month, and a connecting edge is generated based on this data as the capital flow relationship between the nodes.
8. A small and micro enterprise credit risk prediction device based on graph neural network, characterized in that: include: The module for acquiring data of small and micro enterprises to be tested is used to acquire original data of small and micro enterprises to be tested during the loan duration period, and to perform data preprocessing on the original data of small and micro enterprises to be tested to obtain data of small and micro enterprises to be tested; A module for constructing a correlation diagram of small and micro enterprises to be tested, configured to construct a correlation diagram of small and micro enterprises to be tested based on the data of the small and micro enterprises to be tested; The small and micro enterprise credit risk prediction module is used to simplify the association relationship diagram of the small and micro enterprises to be tested to obtain a simplified association relationship diagram; the simplified association relationship diagram is input into the trained small and micro enterprise credit risk prediction model for prediction to obtain the default risk probability of the small and micro enterprises to be tested; The classification of the small and micro enterprises to be tested is determined according to the default risk probability of the small and micro enterprises to be tested; the trained small and micro enterprise credit risk prediction model is obtained by training based on the small and micro enterprise association relationship graph, and the small and micro enterprise association relationship graph is constructed based on the small and micro enterprise data samples of the purchased loan products.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for predicting credit risks of small and micro enterprises based on graph neural networks as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting credit risks of small and micro enterprises based on graph neural networks as described in any one of claims 1 to 7 is implemented.