Real-time credit risk early warning system based on multi-source heterogeneous data fusion

Through multi-source data fusion and relationship analysis, multi-source heterogeneous data are collected and processed in real time, which solves the problems of data singularity and lag in the existing technology, and improves the accuracy and real-timeness of credit risk warnings.

CN120471706APending Publication Date: 2025-08-12BEIJING ZHONGWANG ZHICE TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510598103.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the existing real-time credit risk warning system, data singularity and lag are unable to fully reflect customer credit status and potential risks, and the implicit relationship between multi-source heterogeneous data cannot be handled, affecting the accuracy of early warning.

Method used

Multi-source data acquisition module is used to collect multi-source heterogeneous data in real time, and fuse it to the same dimension through dynamic alignment and fusion module. The relationship building module is used to capture implicit relationships, generate relationship maps, and conduct real-time evaluation and early warning through the risk warning module.

Benefits of technology

It realizes the reflection of customer credit status and potential risks from multiple angles, and improves the accuracy and real-timeness of risk warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471706A_ABST
    Figure CN120471706A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time credit risk early warning system based on multi-source heterogeneous data fusion. The real-time credit risk early warning system comprises a multi-source data acquisition module, a data processing module, a dynamic alignment fusion module, a relationship construction module, an output module and a risk early warning module, the multi-source data acquisition module is used for acquiring credit-related multi-source heterogeneous data in real time. The invention belongs to the technical field of credit risk early warning, and aims to solve the problems that data for early warning in the prior art has singleness and hysteresis quality, the credit condition and potential risk of a customer cannot be reflected from multiple angles, and meanwhile, the implicit relationship among multi-source heterogeneous data cannot be processed and analyzed. The method has the technical effects that the multi-source heterogeneous data related to credit can be collected and fused in real time, the credit condition and potential risk of the customer can be reflected from multiple angles, and the implicit relationship among the multi-source heterogeneous data can be processed and analyzed conveniently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of credit risk early warning technology, and specifically relates to a real-time credit risk early warning system based on multi-source heterogeneous data fusion. Background Art

[0002] With the rapid development of financial markets and the continuous expansion of credit operations, credit risk has also increased significantly. To attract a wider customer base and increase market share, many financial institutions have relaxed lending standards or streamlined approval processes to rapidly expand their business. However, this can easily attract borrowers with poor credit standing, thereby increasing overall credit risk. With increasingly complex and volatile market conditions and evolving customer demands, financial institutions need to carefully manage their asset quality and strengthen internal control systems to ensure effective control of credit risk while maximizing profits. Therefore, establishing a robust, real-time credit risk early warning system is a key measure to prevent potential crises.

[0003] However, most current real-time credit risk warning systems tend to focus only on partial information such as customer credit data, while ignoring the impact of multi-source heterogeneous data such as market transaction data and social media data on credit risk. This results in the data used for warning being single and lagging, and unable to reflect the customer's credit status and potential risks from multiple perspectives. At the same time, it is unable to process and analyze the implicit relationship between multi-source heterogeneous data, which can easily affect the accuracy of the warning. Summary of the Invention

[0004] This application provides a real-time credit risk early warning system based on the fusion of multi-source heterogeneous data, aiming to solve the problems that the data used for early warning in the existing technology is single and lagging, and cannot reflect the customer's credit status and potential risks from multiple perspectives. At the same time, it cannot process and analyze the implicit relationship between multi-source heterogeneous data.

[0005] A real-time credit risk early warning system based on multi-source heterogeneous data fusion, including a multi-source data acquisition module, a data processing module, a dynamic alignment and fusion module, a relationship building module, an output module, and a risk early warning module;

[0006] The multi-source data acquisition module is used to collect multi-source heterogeneous data related to credit in real time; the multi-source heterogeneous data includes standardized data, non-standardized data and time series data;

[0007] The data processing module is used to clean and convert the collected multi-source heterogeneous data;

[0008] The dynamic alignment and fusion module can use cross-modal contrast learning to align standardized data, non-standardized data, and time series data, and fuse them into the same dimension to form a sample set;

[0009] The relationship building module is used to perform in-depth analysis on the fused sample set, capture the implicit relationship between sample data, and generate a relationship map;

[0010] The output module can output the risk value p of the risk factor according to the real-time dynamic risk assessment model and the relationship map;

[0011] The risk warning module is used to preset the corresponding relationship between the risk value p and the risk level, and compare the output risk value p with the preset risk level threshold to determine the risk level.

[0012] Furthermore, the data cleaning of the standardized data and the non-standardized data both involves processing missing values by filling in the mean.

[0013] Furthermore, the dynamic alignment and fusion module includes a text processing unit, a table processing unit, a time series data processing unit and an alignment unit;

[0014] The text processing unit is used to segment the processed original text into words or subwords, and convert the word segmentation results into corresponding vocabulary indexes;

[0015] The table processing unit is constructed based on a convolutional neural network model and is used to extract data features in the table;

[0016] The time series data processing unit is used to process sequence data, capture dynamic features therein, and perform feature extraction on the processed time series data;

[0017] The alignment unit is used to construct a sample ID index using the first dimension of the standardized data, the second dimension of the non-standardized data, and the third dimension of the time series data, and to fuse the first dimension, the second dimension, and the third dimension into the same dimension through a fully connected layer to form a sample set.

[0018] Furthermore, the specific content of the relationship model is as follows:

[0019] a) Node definition;

[0020] b) Model construction;

[0021] c) Model optimization.

[0022] Furthermore, the node definition is used to treat each data sample in the sample set as a node in the graph, define the relationship and interaction mode between the data according to the business logic, define the edges in the graph and the connection method between the nodes, and construct a sample graph based on the definitions of nodes and edges.

[0023] Furthermore, the relational model is constructed based on graph neural network and attention mechanism to capture the implicit relationship between sample data.

[0024] Furthermore, the risk warning module includes a preset unit and a trigger unit;

[0025] The preset unit is used to preset the corresponding relationship between the risk value p and the risk level, and preset specific response measures examples under different risk levels.

[0026] Furthermore, the trigger unit can compare the output risk value p with a preset risk level threshold. When the risk value p falls within the preset threshold range, the risk level warning mechanism corresponding to the threshold range is triggered, and a warning email is automatically sent to the mailbox of the risk management personnel.

[0027] Compared with the prior art, this application has at least the following beneficial effects:

[0028] Based on further analysis and research of existing technical issues, this application utilizes a multi-source data acquisition module and a dynamic alignment and fusion module to collect and fuse multi-source heterogeneous credit-related data in real time, providing a richer source of information and facilitating the reflection of a customer's credit status and potential risks from multiple perspectives. Furthermore, a relationship building module enables in-depth analysis of the fused sample set, exploring implicit relationships between data and identifying risk factors existing between multi-source heterogeneous data, thereby significantly improving the accuracy of risk warnings. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A module architecture diagram of a real-time credit risk early warning system based on multi-source heterogeneous data fusion, provided in accordance with one embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this application more clear, this application is further described in detail below with reference to the accompanying drawings and embodiments.

[0031] like Figure 1 As shown, the present application provides a real-time credit risk warning system based on multi-source heterogeneous data fusion, including a multi-source data acquisition module, a data processing module, a dynamic alignment and fusion module, a relationship construction module, an output module and a risk warning module;

[0032] The multi-source data acquisition module is used to collect multi-source heterogeneous data related to credit in real time. Multi-source heterogeneous data includes standardized data, non-standardized data and time series data. Among them, standardized data refers to transaction records and financial statements recorded in a unified format; non-standardized data, as opposed to standardized data, has no fixed format and is difficult to store and manage using traditional database tables, including social media texts and complaint tickets; time series data includes stock market fluctuations, has time series characteristics, and can identify upward and downward trends in stock prices.

[0033] Standardized data comes from various business systems of financial institutions, such as the core transaction system of banks, the transaction database of e-commerce platforms, etc. Through the internal API interface of financial institutions, customers' transaction records and financial statements can be obtained in real time.

[0034] Non-standardized data is widely distributed across social media platforms and corporate customer service centers. Leveraging APIs from social media platforms and corporate customer service centers, we collect non-standardized data on customer comments, behavior patterns, and complaints on social media. This data comes in a variety of forms, including mixed text, images, and videos, and is both massive and rapidly updated.

[0035] Time series features are obtained through the API interface of the online financial data platform. Calling the API interface can obtain a large amount of stock market time series data, which is convenient for subsequent further processing and analysis.

[0036] The data processing module is used to clean and convert the collected multi-source heterogeneous data, providing high-quality data support for the subsequent analysis and decision-making of the real-time credit risk early warning system. The specific contents are as follows:

[0037] a) Data cleaning

[0038] For missing values in standardized data, we use mean filling. This method ensures the integrity of the dataset by filling missing values with the average value of the data. When converting non-standardized data to a unified format, we also use mean filling to handle missing values in the non-standardized data. When cleaning time series data, we verify and correct abnormal timestamps to ensure that the timestamps in the time series data are accurate and continuous.

[0039] b) Data conversion

[0040] Standardized data from different sources are converted into the same dimensions and units according to predetermined standards for subsequent processing.

[0041] For example, if one of the numeric fields in two datasets is in "yuan" and the other is in "hundred yuan", they need to be unified into the same unit.

[0042] Natural language processing technology is used to identify and extract text data from non-standardized data, remove HTML tags, special characters, stop words, perform lexical analysis, and extract text features.

[0043] According to the time sequence and data characteristics, the time series data is organized into time series objects to ensure that the timestamps in the time series data are accurate and continuous to ensure the consistency of the time series data.

[0044] The dynamic alignment and fusion module uses cross-modal contrastive learning to align standardized data, non-standardized data, and time series data, fusing them into the same dimension to form a sample set. The dynamic alignment and fusion module includes a text processing unit, a table processing unit, a time series data processing unit, and an alignment unit. The specific contents are as follows:

[0045] a) Text processing unit

[0046] The processed raw text is segmented into words or subwords, and the segmentation results are converted into corresponding vocabulary indices. The BERT model is then used to extract features from the vocabulary indices of the processed text data. The BERT model maps the text data to the first dimension, with each vector representing a latent semantic feature of the text. The extracted feature vectors are used as the vector representation of the text data for subsequent data analysis or model training.

[0047] b) Form processing unit

[0048] The table processing unit is built based on the convolutional neural network (CNN) model. The convolutional neural network can effectively extract the data features in the table while maintaining the spatial structure of the image. The convolutional neural network uses global average pooling (GAP): the feature map output by the convolution layer is averaged in the spatial dimension to obtain a vector of fixed length, the feature map is flattened into a one-dimensional vector, and then the vector is mapped to the second dimension through a fully connected layer or PCA.

[0049] c) Time series data processing unit

[0050] The time series data processing unit is built based on the recurrent neural network (RNN) model. The recurrent neural network (RNN) can process sequence data and capture the dynamic features therein. The time series data processing unit is used to extract features from the processed time series data. The recurrent neural network model uses the recurrent unit to capture the dynamic patterns and trends in the time series. The output of the recurrent neural network model is used as a vector representation of the time series data and mapped to the third dimension for subsequent data analysis or model training.

[0051] d) Alignment unit

[0052] Construct a sample ID index using the first dimension of the normalized data, the second dimension of the unnormalized data, and the third dimension of the time series data. Each sample is assigned a unique sample ID, which can be a number, string, or other suitable format. Ensure that the sample ID is unique and corresponding across all three dimensions of data to accurately concatenate data from different dimensions.

[0053] According to the sample ID, the standardized data of the first dimension, the non-standardized data of the second dimension, and the time series data of the third dimension are spliced together, and then the first, second, and third dimensions are fused into the same dimension through the fully connected layer to form a sample set.

[0054] The relationship building module is used to conduct in-depth analysis on the fused sample set, capture the implicit relationships between sample data, and generate a relationship map. The specific contents are as follows:

[0055] a) Node definition

[0056] Each data sample in the sample set is treated as a node in the graph. Based on the business logic, the relationships and interaction patterns between the data are defined. The edges in the graph and the connections between nodes are defined. Edges represent the correlation between data. Based on the definitions of nodes and edges, the sample graph is constructed.

[0057] b) Model construction

[0058] A relational model is built based on graph neural networks and attention mechanisms. The relational model can calculate the attention scores between nodes. The attention mechanism can help the model focus on important nodes and edges, highlighting key relationships and features. Each node receives information from neighboring nodes and updates its identity based on its own features.

[0059] Graph neural networks can learn embedded representations of nodes. By aggregating information from neighboring nodes to update the feature identifier of each node, the node identifier can contain both its own features and the features of neighboring nodes, thereby better capturing structural information and local relationships between nodes.

[0060] The attention mechanism is introduced based on the embedding representation of graph neural networks. The attention mechanism can dynamically focus on the nodes and edges in the graph that are more important to the current task, enhancing the expression of important information while suppressing the interference of irrelevant information.

[0061] The relational model can represent node features based on the sample graph, add, delete, modify edges in the sample graph, and adjust the weights of node features. The degree of match between the mined implicit relationships and the real labels improves the accuracy of credit risk prediction.

[0062] c) Model optimization

[0063] The results from the graph embedding layer, attention mechanism layer, and relational model processing are fused to obtain the final node representation and sample graph structure representation. The backpropagation algorithm optimizes the parameters of the entire model and minimizes cross-entropy loss to improve the relational model's ability to mine implicit relationships and generalize performance. The implicit relationships mined in real time are output in an intuitive manner, generating a relational graph that displays the potential strength and type of associations between nodes.

[0064] The structure and node characteristics of the sample graph are updated in real time based on changes in relationships in new data. For example, if a new customer establishes a transaction relationship with the enterprise, a corresponding edge is added to the sample graph; if the financial status of the enterprise changes, the financial characteristics of the enterprise node are updated.

[0065] The output module can output the risk value p of the risk factor based on the real-time dynamic risk assessment model and relationship map, facilitating subsequent real-time assessment of credit risk. The specific contents are as follows:

[0066] The collected historical dataset, containing implicit relationship data and credit risk labels, was divided into training, validation, and test sets in a 7:2:1 ratio. The training set was used to train the dynamic risk assessment model, the validation set was used to adjust the model's hyperparameters and verify its performance, and the test set was used to ultimately evaluate the model's generalization capabilities. The three datasets were ensured to have similar data distributions to avoid model bias caused by improper data partitioning. For example, the partitioning process ensured that customer samples of different risk levels were reasonably proportioned across the three datasets. The dynamic risk assessment model was built based on a linear regression machine learning model.

[0067] Train the selected dynamic risk assessment model using the training set. During training, adjust the model's hyperparameters to optimize performance. Use the trained model to make predictions on the validation and test sets. Analyze the output of the linear regression dynamic risk assessment model to assess the impact of each implicit relationship variable on credit risk. Calculate the SHAP (SHapley Additive exPlanations) value for each variable to measure its contribution to the prediction results.

[0068] The SHAP value is an explanatory tool used to output whether each feature of each sample contributes positively or negatively to the model prediction result and the size of the contribution, so as to facilitate the determination of the risk value p of each latent relationship variable as a risk factor based on the SHAP value.

[0069] The risk warning module can preset the corresponding relationship between the risk value p and the risk level, and compare the output risk value p with the preset risk level threshold to determine the risk level. The risk warning module includes a preset unit and a trigger unit. The specific contents are as follows:

[0070] a) Preset unit

[0071] Based on historical data, find the critical value for distinguishing different risk levels as a reference for the risk value p, preset the corresponding relationship between the risk value p and the risk level, and preset specific response measures examples for different risk levels.

[0072] For example: Low risk level: 0 < risk value p ≤ 0.3, indicating good credit status and low default risk;

[0073] Medium risk level: 0.3 < risk value p ≤ 0.75, indicating a certain default risk and a certain degree of uncertainty in repayment ability;

[0074] High risk level: 0.75<risk value p≤1, indicating high default risk, poor credit status, and weak repayment ability.

[0075] b) Trigger unit

[0076] The output risk value p is compared with the preset risk level threshold. When the risk value p falls within the preset threshold range, the risk level warning mechanism corresponding to the threshold range is triggered, and a warning email is automatically sent to the risk management personnel's mailbox. A warning message pops up in the financial institution's internal business system, clearly indicating the risk level of the credit, listing the relevant data that led to the risk warning, and listing examples of specific response measures under the risk level.

[0077] The aforementioned real-time credit risk warning system based on multi-source heterogeneous data fusion utilizes a multi-source data acquisition module and a dynamic alignment and fusion module to collect and fuse multi-source heterogeneous credit-related data in real time, providing a richer source of information and facilitating the reflection of customers' credit status and potential risks from multiple perspectives. Furthermore, the relationship building module enables in-depth analysis of the fused sample set, exploring implicit relationships between data and identifying risk factors within the multi-source heterogeneous data, thereby significantly improving the accuracy of risk warnings.

[0078] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A real-time credit risk early warning system based on multi-source heterogeneous data fusion, characterized by: It includes multi-source data acquisition module, data processing module, dynamic alignment and fusion module, relationship building module, output module and risk warning module; The multi-source data acquisition module is used to collect multi-source heterogeneous data related to credit in real time; the multi-source heterogeneous data includes standardized data, non-standardized data and time series data; The data processing module is used to clean and convert the collected multi-source heterogeneous data; The dynamic alignment and fusion module can use cross-modal contrast learning to align standardized data, non-standardized data, and time series data, and fuse them into the same dimension to form a sample set; The relationship building module is used to perform in-depth analysis on the fused sample set, capture the implicit relationship between sample data, and generate a relationship map; The output module can output the risk value p of the risk factor according to the real-time dynamic risk assessment model and the relationship map; The risk warning module is used to preset the corresponding relationship between the risk value p and the risk level, and compare the output risk value p with the preset risk level threshold to determine the risk level.

2. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The data cleaning of the standardized data and the non-standardized data both processes missing values by filling in the mean.

3. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The dynamic alignment and fusion module includes a text processing unit, a table processing unit, a time series data processing unit and an alignment unit; The text processing unit is used to segment the processed original text into words or subwords, and convert the word segmentation results into corresponding vocabulary indexes; The table processing unit is constructed based on a convolutional neural network model and is used to extract data features in the table; The time series data processing unit is used to process sequence data, capture dynamic features therein, and perform feature extraction on the processed time series data; The alignment unit is used to construct a sample ID index using the first dimension of the standardized data, the second dimension of the non-standardized data, and the third dimension of the time series data, and to fuse the first dimension, the second dimension, and the third dimension into the same dimension through a fully connected layer to form a sample set.

4. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The specific content of the relationship model is as follows: a) Node definition; b) Model construction; c) Model optimization.

5. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 4, characterized in that: The node definition is used to treat each data sample in the sample set as a node in the graph, define the relationship and interaction mode between the data according to the business logic, define the edges in the graph and the connection method between the nodes, and construct a sample graph based on the definitions of nodes and edges.

6. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 4, characterized in that: The relational model is built based on graph neural network and attention mechanism to capture the implicit relationship between sample data.

7. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The risk warning module includes a preset unit and a trigger unit; The preset unit is used to preset the corresponding relationship between the risk value p and the risk level, and preset specific response measures examples under different risk levels.

8. A real-time credit risk early warning system based on multi-source heterogeneous data fusion according to claim 7, characterized in that: The trigger unit can compare the output risk value p with the preset risk level threshold. When the risk value p falls within the preset threshold range, the risk level warning mechanism corresponding to the threshold range is triggered, and a warning email is automatically sent to the mailbox of the risk management personnel.

Citation Information

Cited By

  • Loan approval algorithm optimization method based on multi-source data fusion

    CN121616391A

  • Multi-source data entity identification and context grading method

    CN121981815A

  • Multi-source data entity recognition and context classification method

    CN121981815B